Labels

Showing posts with label Data Structures And Algorithms. Show all posts
Showing posts with label Data Structures And Algorithms. Show all posts

Sunday, October 21, 2012

Rabin Carp String Matching Algorithm

     Rabin Carp Algorithm is also one of the string matching algorithm. This algorithm is also an improved version of Naive or Brute Force String Matching Algorithm. Because this algorithm is the a very basic sub-string matching algorithm, but it’s good for some reasons. For example it doesn’t require preprocessing of the text or the pattern. The problem is that it’s very slow. That is why in many cases brute force matching can’t be very useful.
     Michael O. Rabin and Richard M. Karp came up with the idea of hashing the pattern and to check it against a hashed sub-string from the text in 1987.Rabin Carp algorithm is one of the better string matching algorithm than Brute Force algorithm. This algorithm avoid comparison of every character of pattern with characters of text in each position. To do that this algorithm uses hashing. Hashing is the process of converting our data in to numerical value. To convert data in to numerical value we should have assign a numerical value for each text in the alphabet and using any hashing functions we can calculate hash value of any pattern. 
As an example if our pattern P=AABCB and we get ascii values of characters.
Then ascii value of A=65,B=66,C=67 and we can get hash value of above pattern using any hashing functions.
Hash function 1
     hash value h(P) = 65+65+66+67+66 = 329
Hash function 2
     hash value h(P) = 65*1+65*2+66*3+67*4+66*5 =991
Like above you can use any hash functions to calculate hash values.
We can say that if two strings are equal, then hash values of these two string must be same. But if hash values of two strings are equal then we can't say these two strings are equal. I may be or not.  This is the basic idea of the Rabin Carp algorithm.

Algorithm
 Compute hash value of pattern[h(P)]  
    Compute hash value of sub string of text[h(t)]  
    If h(P)=h(t)  
      Compare pattern and sub string character by character  
      If mismatch found  
         move by one position and go to second step  
      Else  
         matching sub string found in the text  
    Else  
      move by one position and go to second step  
Lets learn the algorithm using an example.
     Alphabet S = {A,B,C,D,E,F,G,H}
     Text T = AABDCFFACGABBCABHGADEAADG
     Pattern P = BBCABHG

Here i am going to assign value of character in the alphabet instead of assigning ascii values.
Value(A) = 1
Value(B) = 2
Value(C) = 3
Value(D) = 4
Value(E) = 5
Value(F) = 6
Value(G) = 7
Value(H) = 8
     Then hash value of pattern  h(P) = 2+2+3+1+2+8+7 = 25




Hash value of [AABDCFF] is 23. Hash values are not equal. Sub string is not matched. Move by one position.
Hash value of [ABDCFFA] is 23. Hash values are not equal. Sub string is not matched. Move by one position. 

Hash value of [BDCFFAC] is 25. Hash values are equal. Sub string may be matched. Compare pattern with sub string in the text.
First letter B matched with sub string. But second one mismatched. Move by one position.
Hash value of [DCFFACG] is 29. Hash values are not equal. Sub string is not matched. Move by one position. 
Hash value of [CFFACGA] is 26. Hash values are not equal. Sub string is not matched. Move by one position. 
Hash value of [FFACGAB] is 25. Hash values are equal. Sub string may be matched. Compare pattern with sub string in the text.
First letter B  mismatched. Move by one position. Like these you can find matching sub string.
C implementation of Rabin Carp Algorithm

Complexity

The Rabin-Karp algorithm has the complexity of O(nm) where n, of course, is the length of the text, while m is the length of the pattern. So where it is compared to brute-force matching? Well, brute force matching complexity is O(nm), so as it seems there’s no much gain in performance. However it’s considered that Rabin-Karp’s complexity is O(n+m) in practice, and that makes it a bit faster, as shown on the chart below.
Rabin-Karp's complexity is O(nm), but in practice it's O(n+m)!
Note that the Rabin-Karp algorithm also needs O(m) preprocessing time.

Advantages

  1. Not faster than brute force matching in theory, but in practice its complexityis O(n+m)
  2. Good hashing function it can be quite effective and it’s easy to implement!
  3. Multiple pattern matching support
  4. Good for plagiarism, because it can deal with multiple pattern matching!

Disadvantages

  1. There are lots of string matching algorithms that are faster than O(n+m)
  2. It’s practically as slow as brute force matching and it requires additional space
Rabin-Karp is a great algorithm for one simple reason – it can be used to match against multiple pattern. This makes it perfect to detect plagiarism even for larger phrases.

Saturday, September 15, 2012

How to find Z values of a string

             Finding Z value of a string is requires for Z algorithm. Z algorithm is one of the linear string matching algorithm. As pre-processing of Z algorithm we find Z value of given string. We will denote string by S an arbitrary string. Let S be a string and i >1 one of its positions. Starting at i, we consider all the substrings S[i..j], where i ≤ j ≤ |S|, so that S[i..j] matches a prefix of S itself. Among all the possible values j may take, select the greatest one, max(j) and denote the number [max(j) − i + 1] by Zi(S). If S[i] is different from S[1], then such a j does not exist and Zi(S) is set to 0.
              I think it is very hard to understand. Lets consider finding Z value using an example. Finding Z value starts from 2 to length of the string. That means if our string is "ABCDEF", we only find Z2 to Z6 except Z1. Because we can not find value for Z1.
Lets get our string as S = AABCDAABCXYAABCDAABCDX
Zi is equal to the length of the longest common prefix of S string and S[i to |S|] string. 

Finding Z2
        Devide string in to two sub-strings (first letter as a sub-string and rest as another sub-string)
             S1 =  AABCDAABCXYAABCDAABCDX
             S2 =  ABCDAABCXYAABCDAABCDX
        Then find length of the longest common prefixes of these two substrings. Look How to find prefixes of  a string if you are not well with prefixes.

Prefix of S1 sub-string    -  A, AA , AAB , AABC , AABCD , AABCDA , AABCDAA , AABCDAAB ,        AABCDAABC , AABCDAABCX , AABCDAABCXY , AABCDAABCXYA , AABCDAABCXYAA , AABCDAABCXYAAB , AABCDAABCXYAABC , AABCDAABCXYAABCD , AABCDAABCXYAABCDA , AABCDAABCXYAABCDAA , AABCDAABCXYAABCDAAB , AABCDAABCXYAABCDAABC , AABCDAABCXYAABCDAABCD , AABCDAABCXYAABCDAABCDX

Prefix of S2 substring - A , AB , ABC , ABCD , ABCDA , ABCDAA , ABCDAAB , ABCDAABC , ABCDAABCX , ABCDAABCXY , ABCDAABCXYA , ABCDAABCXYAA , ABCDAABCXYAAB , ABCDAABCXYAABC , ABCDAABCXYAABCD , ABCDAABCXYAABCDA , ABCDAABCXYAABCDAA , ABCDAABCXYAABCDAAB , ABCDAABCXYAABCDAABC , ABCDAABCXYAABCDAABCD , ABCDAABCXYAABCDAABCDX

Highlighted by green are the common prefixes of these two sub-strings. Length of the longest common prefix become as Z value of that position. Therefore, Z2 = 1
Finding Z3
        Devide string in to two sub-strings (first two letters as a sub-string and rest as another sub-string)
             S1 =  AABCDAABCXYAABCDAABCDX
             S2 =  BCDAABCXYAABCDAABCDX


        Then find length of the longest common prefixes of these two substrings.

Prefix of S1 sub-string    -  A, AA , AAB , AABC , AABCD , AABCDA , AABCDAA , AABCDAAB ,        AABCDAABC , AABCDAABCX , AABCDAABCXY , AABCDAABCXYA , AABCDAABCXYAA , AABCDAABCXYAAB , AABCDAABCXYAABC , AABCDAABCXYAABCD , AABCDAABCXYAABCDA , AABCDAABCXYAABCDAA , AABCDAABCXYAABCDAAB , AABCDAABCXYAABCDAABC , AABCDAABCXYAABCDAABCD , AABCDAABCXYAABCDAABCDX


Prefix of second substring - B , BC ,BCD , BCDA , BCDAA , BCDAAB , BCDAABC , BCDAABCX , BCDAABCXY , BCDAABCXYA , BCDAABCXYAA , BCDAABCXYAAB , BCDAABCXYAABC , BCDAABCXYAABCD , BCDAABCXYAABCDA , BCDAABCXYAABCDAA , BCDAABCXYAABCDAAB , BCDAABCXYAABCDAABC , BCDAABCXYAABCDAABCD , BCDAABCXYAABCDAABCDX

There are no any common prefixes in these two sub-strings. Therefore, Z3 = 0
Likewise we can find other Z values.
Finding Z4
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  CDAABCXYAABCDAABCDX                            Z4=0

Finding Z5
          S1 =  AABCDAABCXYAABCDAABCDX
          S2 =  DAABCXYAABCDAABCDX                                 Z5=0

Finding Z6
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 = AABCXYAABCDAABCDX                                    Z6=4

Finding Z7
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  ABCXYAABCDAABCDX                                      Z7=1

Finding Z8
           S1 =  AABCDAABCXYAABCDAABCDX
           S2 =  BCXYAABCDAABCDX                                          Z8=0

Finding Z9
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  CXYAABCDAABCDX                                            Z9=0

Finding Z10
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  XYAABCDAABCDX                                               Z10=0

Finding Z11
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  YAABCDAABCDX                                                 Z11=0

Finding Z12
           S1 =  AABCDAABCXYAABCDAABCDX
           S2 =  AABCDAABCDX                                                     Z12=9

Finding Z13
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  ABCDAABCDX                                                       Z13=1

Finding Z14
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  BCDAABCDX                                                         Z14=0

Finding Z15
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  CDAABCDX                                                            Z15=0

Finding Z16
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  DAABCDX                                                               Z16=0

Finding Z17
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  AABCDX                                                                  Z17=5

Finding Z18
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  ABCDX                                                                     Z18=1

Finding Z19
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  BCDX                                                                        Z19=0

Finding Z20
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 = CDX                                                                            Z20=0

Finding Z21
           S1 =  AABCDAABCXYAABCDAABCDX
           S2 =  DX                                                                               Z21=0

Finding Z22
            S1 =  AABCDAABCXYAABCDAABCDX
            S2 =  X                                                                                Z22=0

Wednesday, September 12, 2012

Boyer Moore string Matching algorithm

             Boyer Moore string matching algorithm is one of the improved version of Naive or Brute Force String Matching Algorithm. In this method we try to get more than one position movement like in Brute Force method to minimize running time of string matching process.This method use two running time heuristics to minimize running time.
  1. Looking-Glass Heuristics
  2. Character-Jump Heuristic
Now lets get idea of these two Heuristic

Looking-Glass Heuristics

          There we start comparison of our pattern with the text from the end of pattern and
move backward to the front of pattern. This comparison technique may not good all the time than comparing from first letter to last letter of the pattern. Sometimes it may be better than front to the backward comparing and sometimes may not.
          As an example get pattern as FOOD and comparing letters of the text as GOOD.
then front to the backward comparison get only one comparison to move to the next position. Because F and G gets mismatch. But backward to the front comparison get 4 comparisons to move to the next position. Therefore, we can't say that Looking-Glass Heuristics is not always time effective. It may be or may not be.

 Character-Jump Heuristic

          In this technique, we find more number of movement positions than moving one by one position when we get mismatches. Let text array as T[] and pattern array as P[]. If we find that T[i] letter mismatch with P[j] letter. Then we find that T[i] letter is in anywhere of P that does not take comparison in that position. If it is, then we move P as get align with the T[i] letter. If it is not,then move pattern completely past T[i] letter.

According to the above two heuristics, we can say that,
  1. Looking-Glass heuristic enables to get to the destination faster by going backward if there is a mismatch during the consideration of P at a certain location in T.
  2. Character-Jump heuristic enables to avoid lots of needless comparisons by significantly shifting P relative to T.
There are three types of Boyer Moore string Matching algorithms.
  1. Bad Character Rule
  2. Extended Bad Character Rule
  3. Good Suffix Rule

Bad Character Rule

      Mismatching character of the text aligned with the pattern is called Bad Character. In this rule we should have to find right most occurrences of each letters of the alphabet. As an example lets get
     Alphabet S={A,B,C,D,R,T}
     Pattern P= ARDCARA
     Text T=ABATARADABARDAARADABADATATABAT

Now lets find right most occurrence positions(R) of alphabet
       R(A) = 7               R(B) = 0             R(C) = 4           R(D) = 3           R(R) = 6            R(T) = 0

ALGORITHM:
        if T[k] letter mismatched with P[i] letter
               shift P along by max[1,i-R(T[k])]

Red Boxes - Mismatches               Green Boxes - Matches
Explanation of each raw.
  1. First raw is the text and second raw is the pattern. 
  2. Last three letters matched and C mismatched. Bad Character is T and i is 4. Find right most occurrence of  T in Pattern. That is R(T) = 0. Then find (i- R(T) ) = 4-0 = 4. Then find max(1,4). That is 4. Then move by 4 positions.   
  3. Last letter matched and R mismatched. Bad Character is B and i is 6. Find right most occurrence of B in Pattern. That is R(B) = 0. Then find (i- R(B) ) = 6-0 = 6. Then find max(1,6). That is 6. Then     move by 6 positions. 
  4. Last three letters matched and C mismatched. Bad Character is A and i is 4. Find right most occurrence of  A in Pattern. That is R(A) = 7. Then find (i- R(A)) = 4-7 = -3. Then find max(1,-3). That is 1. Then move by 1 positions.  
  5. Last letter A mismatched. Bad Character is D and i is 7. Find right most occurrence of  D in Pattern. That is R(D) = 3. Then find (i- R(D) ) = 7-3 = 4. Then find max(1,4). That is 4. Then move by 4 positions. 
  6. Last letter A mismatched. Bad Character is D and i is 7. Find right most occurrence of  D in Pattern. That is R(D) = 3. Then find (i- R(D) ) = 7-3 = 4. Then find max(1,4). That is 4. Then move by 4 positions.
  7. Last letter A mismatched. Bad Character is T and i is 7. Find right most occurrence of  T in Pattern. That is R(T) = 0. Then find (i- R(T) ) = 7-0 = 7. Then find max(1,7). That is 7. Then move by 7 positions.
Above example we can see,       
      We get more than one shift only when  R(T[k]) + 1 < i. Otherwise we shift the pattern by only one position. Therefore this rule is not useful when R(T[k]) >= i. Therefore we move on to extended bad character rule.

Extended Bad Character Rule 

         In this rule we does not consider right most occurrences of letters of alphabet and we should have to find all the occurrences of each letters of the alphabet. As an example lets get
     Alphabet S={A,B,C,D,R,T}
     Pattern P= ARDCARA
     Text T=ABATARADABARDAARADABADATATABAT

Now lets find  occurrence positions of alphabet
     A_list = {7,5,1}            B_list = {Ø}           C_list = {4}        D_list = {3}
     R_list = {2,6}             T_list = {Ø}

Red Boxes - Mismatches             Green Boxes - Matches

Explanation of each raw.
  1. First raw is the text and second raw is the pattern. 
  2. Last three letters matched and C mismatched. Bad Character is T and i is 4. View occurrence list of T. That is T_list = {Ø}.  Then move by 4 positions.   
  3. Last letter matched and R mismatched. Bad Character is B and i is 6. View occurrence list of B. That is B_list = {Ø}.  Then move by 6 positions. 
  4. Last three letters matched and C mismatched. Bad Character is A and i is 4. View occurrence list of A. That is A_list = {7,5,1}. Then find nearest less value than i of A_list. That is 1 and i is 4. Then move by (4-1)positions. That means move by 3 positions. 
  5. Last letter A mismatched. Bad Character is B and i is 7. View occurrence list of B. That is B_list = {Ø}.  Then move by 7 positions.
  6. Last letter matched and R mismatched. Bad Character is T and i is 6. View occurrence list of T. That is T_list = {Ø}.  Then move 6 positions.
Above example we can see,       
      We get more than one position shift (three position) in fifth raw using extended bad character rule. But using bad character rule we get only one position shift in fifth raw. Therefore we can say that Extended Bad Character Rule is efficient that Bad Character Rule.
       
You can find C Source Code for Boyer Moore String Matching Algorithm using extended bad character rule  from following link.
                   Boyer Moore String Matching Algorithm In C 

Pre-processing

             In above string matching algorithm using Bad character rule, there is a process done before starting string matching process. That is finding right most occurrence or finding all the occurrences of the alphabet in the pattern. This process is called as pre-processing. This process also get  running time. Pre-processing process is used to avoid (n-m+1) shifts in Naive algorithm.
             Also we can see that pre-processing time is depend on the alphabet. Because we find occurrences of the alphabet in the pattern. So we can say that woks the fastest when the alphabet is moderately sized and the pattern is relatively long.

 Running Time Analysis

             Boyer Moore String matching algorithm  is not like Naive algorithm. Becouse above mentioned that this algorithm take time for pre-processing. Therefore, addition of pre-processing time and string matching time is the running time of this algorithm. 
Lets get : n-length of text                  m-length of pattern                    |Σ| -length of alphabet
  1. Pre-processing time                                                                                                                                     In worst case using bad character rule takes ( |Σ| + m) comparisons. Because in each letter of alphabet(|Σ|), we find right most occurrence of pattern. Pattern length is m. So that takes  ( |Σ| + m) comparisons.Therefore, running time becomes O( |Σ| + m).                                                                
  2. String matching time                                                                                                                                      In worst case using bad character rule takes (mn) comparisons. Because in each letter of pattern(m) compare with text and move by one position each time, So we should have to make (n-m+1) positions. So all the comparisons becomes  ( m*(n-m+1)). Simply we say (mn) comparisons. Therefore, running time becomes O(nm).
Therefore the worst case running time of Boyer-Moore algorithm is O(nm + |∑|)   

Advantages

  1. Boyer-Moore algorithm is extremely fast on large alphabet (relative to the length of the pattern).