匹配相同的行

Question 1

我每天编写的快速脚本类型：

#!/usr/bin/perl
#

use strict;
use warnings;

#data structures we're gonna need
my %positions; #how many times have we seen a given position
my %registered_lines; #the concatenated lines for the given position 
my $dn; # the current dn section we're in

while (<>)
{
    if (/^dn:/) #beginning of a new dn section (and end of the previous one)
    {
        my $printed = 0; #we want to print the dn line only once
        foreach my $key (keys %positions) #we look at all positions seen in last section
        {
            if ($positions{$key} gt 1) # has the current position been seen more than once
            {
                print $dn unless $printed;
                $printed = 1;
                #print "position $key is repeated $positions{$key} times\n";
                print $registered_lines{$key}; #print all the lines with the position
            }
        }

        #reset variables for the next section
        $dn = $_;
        %positions = ();
        %registered_lines = ();
    }

    if (/^rdcPosition/) #new line
    {
        /(\d+)$/; #have a look at the digits at the end of the line
        my $pos = $1;
        if (exists $positions{$pos}) #have we already seen this position
        {
            $positions{$pos} += 1; #increment the counter
            $registered_lines{$pos} .= $_; #record the line
        }
        else
        {
            $positions{$pos} = 1;
            $registered_lines{$pos} = $_;
        }
    }
}

运行它作为：

perl script.pl < input_data_file

Answer

我每天编写的快速脚本类型：

#!/usr/bin/perl
#

use strict;
use warnings;

#data structures we're gonna need
my %positions; #how many times have we seen a given position
my %registered_lines; #the concatenated lines for the given position 
my $dn; # the current dn section we're in

while (<>)
{
    if (/^dn:/) #beginning of a new dn section (and end of the previous one)
    {
        my $printed = 0; #we want to print the dn line only once
        foreach my $key (keys %positions) #we look at all positions seen in last section
        {
            if ($positions{$key} gt 1) # has the current position been seen more than once
            {
                print $dn unless $printed;
                $printed = 1;
                #print "position $key is repeated $positions{$key} times\n";
                print $registered_lines{$key}; #print all the lines with the position
            }
        }

        #reset variables for the next section
        $dn = $_;
        %positions = ();
        %registered_lines = ();
    }

    if (/^rdcPosition/) #new line
    {
        /(\d+)$/; #have a look at the digits at the end of the line
        my $pos = $1;
        if (exists $positions{$pos}) #have we already seen this position
        {
            $positions{$pos} += 1; #increment the counter
            $registered_lines{$pos} .= $_; #record the line
        }
        else
        {
            $positions{$pos} = 1;
            $registered_lines{$pos} = $_;
        }
    }
}

运行它作为：

perl script.pl < input_data_file

Question 2

Answer

Question 3

awk '/^dn:/ {d=1} {if (d) {print buf | "sort|uniq -d"; d=0; buf=""} else {buf=buf$0"\n"}} END {print buf | "sort|uniq -d"}'|grep -v '^$'

比 perl 版本少得多的打字 =)。可能更简单，但我似乎无法“在任何模式上或在末尾”执行 awk 规则，因此它包含一点 shell 代码重复。

Answer

awk '/^dn:/ {d=1} {if (d) {print buf | "sort|uniq -d"; d=0; buf=""} else {buf=buf$0"\n"}} END {print buf | "sort|uniq -d"}'|grep -v '^$'

比 perl 版本少得多的打字 =)。可能更简单，但我似乎无法“在任何模式上或在末尾”执行 awk 规则，因此它包含一点 shell 代码重复。

匹配相同的行

答案1

答案2

答案3

相关内容