Code clone detection method based on frequent sequence mining
By a method based on frequent sequence mining, the Token sequences in the source code are parsed and processed, and the closed pattern mining algorithm is used to find frequent subsequences, and candidate cloning pairs are constructed and matched. The problem of difficulty in taking into account both accuracy and detection efficiency in the existing technology is solved, and duplicate codes of insertion and deletion types are efficiently identified.
Patent Information
- Application Number
- CN202210911699.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-07-28
AI Technical Summary
Existing code cloning detection methods are difficult to balance between accuracy and detection efficiency, and they cannot effectively identify duplicate codes of insertion and deletion types, and the method based on abstract syntax trees has a high time complexity.
Using a method based on frequent sequence mining, a lexical tool and syntax parsing tool is generated through ANTLR4 tool, the source code is parsed to generate abstract syntax trees, filter useless code, obtain the token sequence in the function, hash and index establishment, use a closed pattern mining algorithm to find frequent subsequences, build candidate cloning pairs, and obtain the final cloning pairs by setting thresholds and grouping matching.
It improves the accuracy and detection efficiency of code cloning detection, can effectively identify duplicate codes of insertion and deletion types, and reduces the high time complexity of abstract syntax tree matching.
Smart Images

Figure CN115203053B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of computer detection, and particularly relates to a code clone detection method for two files and multiple files, and can be applied to the detection of plagiarism of student codes in teaching and the detection of copy-paste codes in conventional software. Background Art
[0002] Code clone detection, also known as plagiarism detection or copy detection, is widely used in the fields of education, intellectual property, and information retrieval. In recent years, the emergence of new programming languages and a wide variety of plagiarism methods have led to more and more frequent code plagiarism and cloning. As researchers continue to deepen their research on code clone detection, new ideas for code detection are constantly being proposed. Existing code clone detection methods mainly include text-based clone detection methods, token-based code clone detection methods, and abstract syntax tree-based code clone detection methods, among which:
[0003] The text-based clone detection method treats the entire code line as a long string and uses a string matching algorithm to detect whether the code is cloned. Johnson was the first to do relevant work on text-based clone detection methods. In his paper Identifying redundancy in source code using fingerprints, he proposed using a fingerprint algorithm on source code line substrings and comparing the fingerprints calculated for each line to identify matching substrings. This method is highly efficient because it is independent of the specific language, and when there are slight changes in the code, duplicate code is difficult to detect.
[0004] The token-based code clone detection method is to parse the code into token sequences and then compare them through tools. Dup proposed by Baker in the paper On finding duplication and near-duplication in large software systems can find all matching pairs of parameterized code snippets. If two code snippets are a continuous sequence of source code lines with consistent identifier mapping, then the two code snippets match. CCFinder proposed by Kamiya et al. in the paper CCFinder: A multilinguistic token-based code clone detection system for largescale source code uniformly formats the token sequences in the source code, thereby ignoring tokens with human factors such as variables and method names when matching. Although these two methods have achieved good results in large-scale software and can identify duplicate code after identifier renaming, they cannot identify duplicate code after statement insertion and deletion.
[0005] The code clone detection method based on abstract syntax tree is to convert the code into an abstract syntax tree and search for similar subtrees through the tree matching algorithm. Baxter et al. proposed a simple and practical method in the paper Clone detection using abstract syntax trees, which uses abstract syntax trees to detect accurate clones and close clones on any program fragment in the program source code. Deckard proposed in the paper Deckard: Scalable and accurate tree-based detection of code clones to divide the nodes of the syntax tree into important nodes and unimportant nodes, represent the subtrees with vectors, and then merge different subtrees to form the whole tree as the final vector. Then use the local sensitive hashing algorithm to complete vector clustering and identify code clones. Although these two methods can detect copies of reordered statements, they have high time complexity.
[0006] In summary, due to the characteristics of the code itself, the defects of these existing methods result in a poor balance between the accuracy and efficiency of code clone detection. Summary of the invention
[0007] The purpose of the present invention is to address the deficiencies of the above-mentioned existing basis and propose a code clone detection method based on frequent sequence mining, aiming to improve the accuracy of code clone detection while improving its detection efficiency.
[0008] To achieve the above object, the technical solution of the present invention includes the following:
[0009] A code clone detection method based on frequent sequence mining, characterized by comprising the following steps:
[0010] (1) Use the metalanguage of the ANTLR4 tool to define the grammar files Grammer with the suffix g4 of different programming languages, and generate lexical tools and grammar parsing tools respectively according to Grammer;
[0011] (2) Use the above tools to parse the source code file and generate an abstract syntax tree;
[0012] (3) Parse the abstract syntax tree to filter out header files, blank lines, and comments in the source code, obtain the start and end position sets of the token set contained in each function in the source code, and slice functions that are too long;
[0013] (4) Traverse the start and end position sets to obtain all tokens contained in each function, standardize the tokens in the function, and merge the tokens in the same line of the source code into a sequence;
[0014] (5) Hash each sequence in each function, create an index for each sequence and save it to a function sequence set, assign a unique ID to each function sequence set, and use multiple function sequence sets in the source code to form a sequence database;
[0015] (6) Using the closed pattern mining algorithm ClaSp to search for frequent subsequences with a support of at least 2 in the sequence database, and replacing the sequence ID generated in ClaSp with the function sequence set ID in the above sequence database to obtain a frequent result set consisting of frequent subsequences and function sequence set IDs;
[0016] (7) For each frequent subsequence of the frequent result set, candidate clone pairs are constructed two by two using the function sequence set ID, and they are saved in the clone pair set;
[0017] (8) Expand and remove duplicate candidate clone pairs by constructing unique keywords for the two function sequence set IDs of each candidate clone pair in the clone pair set, and then remove duplicate clone rows from the remaining unique candidate clone pairs;
[0018] (9) Setting the clone size, clone size difference, gap threshold, and clone size minimum.
[0019] (10) Filter the candidate clone pairs in which any candidate clone set size is less than size and the result pairs in which the difference in size between two candidate clone sets in the candidate clone pairs is greater than diff;
[0020] (11) Sort the clone rows in the candidate clone pairs in the filtered clone pair set from small to large, and group the clone rows by the set gap and min, that is, record the rows whose difference between two consecutive clone rows is less than the gap until it is greater than the gap, generate row groups from the recorded clone rows, and filter out the row groups whose difference is less than the min;
[0021] (12) Traverse each candidate clone pair in the clone pair set, match the row groups in the candidate clone pairs, and form the matched row groups into final clone pairs to obtain the final clone pair set, thus completing the clone detection of the code.
[0022] Compared with the prior art, the present invention has the following advantages:
[0023] First, the present invention performs detection based on the function or method in the source code as the basic granularity, removes useless codes such as header files other than the function or method, reduces the size of the candidate sequence set, and improves the detection efficiency.
[0024] Second, the present invention uses the closed pattern mining algorithm ClaSp to complete code clone detection, avoiding the high time complexity of finding similar subtrees between abstract syntax trees and improving the detection performance.
[0025] Third, the present invention realizes detection of cloned codes of insertion and deletion types by setting a gap threshold, thereby improving the detection capability of different code clone types. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a flow chart for realizing the present invention;
[0027] Figure 2 is a sample diagram of the abstract syntax tree in the present invention;
[0028] Figure 3 The simulation results of code clone detection using the existing JPlag method are shown;
[0029] Figure 4 This is a simulation result diagram of code clone detection performed by the present invention;
[0030] Figure 5 It is a simulation comparison diagram of the present invention and the existing JPlag method for detection on the code data set;
[0031] Figure 6 It is a simulation comparison diagram of the detection of the present invention and the existing CloneDR method on the code data set;
[0032] Figure 7 It is a simulation diagram for comparing the performance of the present invention and the existing JPlag and CloneDR methods in detecting on code data sets. DETAILED DESCRIPTION
[0033] The embodiments and effects of the present invention are further described in detail below with reference to the accompanying drawings:
[0034] In order to solve the problems of the inability to identify insertion, deletion, and duplication of code and the high algorithm complexity in the above-mentioned existing Token-based and abstract syntax tree-based code clone detection methods, this example mainly consists of three parts: acquisition of token sequence, closed pattern sequence mining, and post-processing of mining results. First, the lexical parser is used to parse the source code to obtain the abstract syntax tree, and the abstract syntax tree is traversed and pruned to obtain the token sequence; then a sequence database is constructed and closed pattern mining is performed to obtain duplicate items; finally, the mining results are deduplicated, merged, and other post-processing to complete the code clone detection.
[0035] Reference Figure 1 , the implementation steps of this example are as follows:
[0036] Step 1: Generate lexical tools and grammar parsing tools.
[0037] Use the metalanguage of ANTLR4 (ANother Tool for Language Recognition4) to define grammar files of different programming languages with the suffix g4, which can be obtained from public websites;
[0038] According to the grammar file Grammer generates lexical tools and grammar parsing tools respectively.
[0039] The grammar file Grammer in this example is from https: / / github.com / antlr / grammars-v4 Get.
[0040] Step 2: Generate an abstract syntax tree.
[0041] Use the above tools to parse the source code file and generate an abstract syntax tree, such as Figure 2 As shown in the figure, the abstract syntax tree of the code statement "int a,b,c;" is displayed, where compilationUnit represents the source code parsing entry rule, translationUnit represents the parsing granularity, declaration is a general matching rule that can match various code statements, the one with the declarator suffix is the sub-matching rule of declaration, typeSpecifier is the declared variable type, Identifier represents the variable wildcard, Comma represents the comma, and Semi represents the semicolon.
[0042] Step 3: Parse the abstract syntax tree.
[0043] Parsing the abstract syntax tree is to filter out header files, blank lines and comments in the source code, rewrite the function interface in the parsing tool, obtain the start and end position set of the token set contained in each function in the source code, and slice the overly long functions.
[0044] Step 4: Sequence normalization.
[0045] Traverse the start and end position sets to obtain all tokens contained in each function, and standardize the tokens in the function as follows:
[0046] Number type processing: convert integer and floating point numbers to be represented by the symbol "$", such as 3.1415926 is converted to "$";
[0047] Literal constant processing: convert characters in single quotes or strings in double quotes to the symbol "#", such as converting "Hello World" to "#";
[0048] Identifier variable processing: convert the identifier variable to be represented by the symbol "@", such as converting str in String str to "@";
[0049] Merge tokens on the same line of source code into a sequence.
[0050] Step 5: Establish a sequence database.
[0051] Hash each sequence in each function, create an index for each sequence and save it to a function sequence set, assign a unique ID to each function sequence set, and use multiple function sequence sets in the source code to form a sequence database. The specific steps are as follows:
[0052] 5.1) Use the HashPJW algorithm to hash each sequence in each function:
[0053] 5.1.1) Input string;
[0054] 5.1.2) Read an 8-bit character from the string;
[0055] 5.1.3) Shift the original hash value left by 4 bits and add the character in 5.1.2);
[0056] 5.1.4) Determine whether the highest 4 bits are 0000. If not all 0, mix the highest 8 bits with bits 1-8, and then rewrite the highest 4 bits to 0000;
[0057] 5.1.5) Repeat steps 5.1.2)-5.1.4) until the end of the string to get the hash value;
[0058] 5.2) Save the token of each function to the hash table to create an index, and assign it a unique ID to form a function sequence set;
[0059] 5.3) Multiple function sequence sets are used to form a sequence database according to the input method of the closed pattern mining algorithm ClaSp.
[0060] Step 6: Closed pattern mining to obtain frequent result sets.
[0061] The existing closed pattern mining algorithm ClaSp is used to search for frequent subsequences with a support of at least 2 on the sequence database and output the function sequence set ID. A frequent result set consisting of frequent subsequences and function sequence set IDs will be obtained. The specific implementation steps are as follows:
[0062] 6.1) Read the data in the sequence database and build a sequence dictionary tree;
[0063] 6.2) Replace the sequence ID generated in ClaSp with the function sequence set ID in the above sequence database;
[0064] 6.3) Set the minimum support to 2, mine the frequent 1 sequence set L1 from the sequence database, and generate the frequent closed sequence candidate set FCC;
[0065] 6.4) Take all frequent 1 sequences in L1 as the starting pattern strings, perform sequence expansion and item set expansion on the starting pattern strings to obtain the pattern string p;
[0066] 6.5) Execute the recursive pruning search algorithm on the pattern string p and prune it according to the following rules:
[0067] For the recursively generated pattern string p′, if the frequent sequence p is found before p′, and p contains p′, if and only if the support of p and p′ is the same, then stop searching all descendant subtrees of p′ in the sequence dictionary tree;
[0068] For the recursively generated pattern string p′, if the frequent sequence p is found before p′, and p′ contains p, if and only if the support of p and p′ is the same, then all subtrees in the sequence dictionary tree are transplanted to p′;
[0069] 6.6) Update the frequent pattern strings obtained after pruning to the frequent closed sequence candidate set FCC;
[0070] 6.7) Remove the non-closed frequent pattern set in FCC and obtain the closed frequent result set FCS;
[0071] 6.8) Rewrite the interface for obtaining results in the ClaSp algorithm to save the frequent result set, where the frequent result set contains the frequent subsequence and function sequence set IDs.
[0072] Step 7: Construct candidate clone pairs.
[0073] 7.1) Traverse each result in the frequent result set and combine each result in pairs according to the function sequence set ID;
[0074] 7.2) Obtain the cloned row set by obtaining the original row number of the sequence in the function sequence set in the above combination through the index in step 5;
[0075] 7.3) The respective clone row sets and function sequence set IDs in the combination are used together to form candidate clone pairs, and each candidate clone pair can be saved in the clone pair set.
[0076] Step 8: De-duplicate candidate clones.
[0077] The duplicate candidate clone pairs are expanded and removed by constructing a unique keyword for the two function sequence set IDs of each candidate clone pair in the clone pair set. The specific steps are as follows:
[0078] 8.1) Construct a unique key for the two function sequence set IDs of each candidate clone pair in the clone pair set:
[0079] 8.1.1) Get two function sequence sets, denoted by their IDs x and y, where x <y;
[0080] 8.1.2) Use x and y to generate a unique keyword represented as x|y;
[0081] 8.2) Traverse the clone pair set, and for candidate clone pairs that generate the same keyword, use the repeated candidate clone pairs to expand the candidate clone pairs that generated the keyword for the first time, so as to ensure that the candidate clone pairs are unique;
[0082] 8.3) Remove duplicate clone rows from the remaining unique candidate clone pairs.
[0083] Step 9: Set the threshold to filter candidate clone pairs.
[0084] 9.1) Set the clone size, clone size difference, gap threshold, and clone size minimum.
[0085] 9.2) Remove any candidate clone pairs in the clone pair set whose size is less than size and any candidate clone pairs whose size difference between two candidate clone sets in the candidate clone pair is greater than diff to obtain a filtered clone pair set.
[0086] Step 10: Group the clone rows in the candidate clone pairs.
[0087] The clone rows in the candidate clone pairs in the filtered clone pair set are sorted from small to large, and the clone rows are grouped by the set gap and min, that is, the rows with a difference of two consecutive clone rows less than the gap are recorded until the difference is greater than the gap, and row groups are generated with the recorded clone rows, and the row groups with a difference less than the min are filtered out.
[0088] Step 11: Match the row groups to obtain the final clone pairs.
[0089] 11.1) Traverse each candidate clone pair in the clone pair set and match the row groups in the candidate clone pair:
[0090] 11.1.1) Obtain the hash value of the row sequence through the index in (5);
[0091] 11.1.2) Match the hash values in the row group according to the following matching principles:
[0092] If the first and last items in two row groups are the same, the row groups match successfully;
[0093] If the first item in a row group is the same as the first or second item in another row group, and the size difference between the two row groups is less than the clone size difference diff, the row group matches successfully;
[0094] If the first item or the second item in a row group is the same as the first item in another row group, and the size difference between the two row groups is less than the clone size difference diff, the row group matches successfully.
[0095] 11.2) Group the matched rows into final clone pairs, obtain the final clone pair set, and complete the clone detection of the code.
[0096] The beneficial effects of the present invention can be further illustrated by the following simulation:
[0097] 1. Simulation conditions
[0098] Software and hardware environment: Winsows10, 16 cores, 32G memory, 2.90GHz i7CPU, IntelliJ-IDEA2019.3 and java1.8.
[0099] Dataset: The dataset is the programming answer codes submitted by students downloaded from the OJ platform, namely Oj-1(27), Oj-2(32), Oj-3(45), and Oj-4(123), where the number of code files are 27, 32, 45, and 123 respectively.
[0100] 2. Simulation Content
[0101] Simulation 1, using the existing JPlag method to simulate code clone detection on two student codes, the results are shown in 3, where A.java and B.java are two java code files. From the detection results, we can see that the A.java file detects clones on lines 1-10, and the B.java file detects clones on lines 1-11. It can be seen that after encountering the code "else" statement, although there are still cloned codes in the code after line 10 of A.java and the code after line 11 of B.java file, JPlag did not detect them.
[0102] Simulation 2, using the present invention to simulate code clone detection on two student codes, the result is as shown in 4, where A.java and B.java are two java code files, and the method of the present invention can detect the clone code that cannot be detected in simulation 1, that is, lines 11-18 in the A.java file and lines 12-19 in the B.java file are detected as clone codes (codes marked with underscores).
[0103] from Figure 3 and Figure 4 It can be seen from the comparison that the present invention can detect duplicate items that the existing JPlag method cannot detect, indicating that the detection accuracy of the present invention is higher than that of JPlag.
[0104] Simulation 3, comparative simulation of code clone detection between the present invention and the existing JPlag method on a code data set, the result is as shown in 5, where the ordinate is the number of code file clone pairs detected in the data set, the abscissa is the number of code files in a single data set, and the similarity is the ratio of cloned lines in a single file to the total number of lines in the current file, which is currently set to 80%.
[0105] from Figure 5 It can be seen that under the condition of 80% similarity threshold, in the student code data set, the code file clone pairs detected by the present invention are more than those by JPlag, indicating that the present invention can detect some clones that JPlag cannot detect, indicating that the detection accuracy of the present invention is higher than that of JPlag.
[0106] Simulation 4, comparative simulation of code clone detection on code data set between the present invention and the existing CloneDR method, the result is as shown in 6, where the ordinate is the ratio of all clone lines detected in a data set to the total number of code lines in the current data set, the abscissa is the number of code files in a single data set, and the similarity is the threshold set for running code clone detection by the existing CloneDR method, which is currently set to 80%.
[0107] from Figure 6It can be seen that under the condition of a threshold of 80% similarity, in the student code dataset, the proportion of cloned lines detected by the present invention to the total number of code lines is comparable to that of the CloneDR code clone detection method based on the abstract syntax tree, indicating that the detection accuracy of the present invention is comparable to that of CloneDR.
[0108] Simulation 5, using the present invention and the existing JPlag and CloneDR methods to simulate the code clone detection time on the student code data set, the result is shown in 7, where the ordinate is the execution time (ms) and the abscissa is the number of code files in a single data set.
[0109] from Figure 7 It can be seen that the detection time of the present invention is less than that of Clone-DR, indicating that the detection performance of the present invention is higher than that of CloneDR. Although the detection time of the present invention is longer than that of JPlag, the detection accuracy of the present invention is higher than that of the JPlag method, and the detection time is also within an acceptable range.
Claims
1. A code clone detection method based on frequent sequence mining, It is characterized in that The steps include: (1) Use the metalanguage of the ANTLR4 tool to define the grammar files Grammer with the suffix g4 of different programming languages, and generate lexical tools and grammar parsing tools respectively according to Grammer; (2) Use the above tools to parse the source code file and generate an abstract syntax tree; (3) Parse the abstract syntax tree to filter out header files, blank lines, and comments in the source code, obtain the start and end position sets of the token set contained in each function in the source code, and slice functions that are too long; (4) Traverse the start and end position sets to obtain all tokens contained in each function, standardize the tokens in the function, and merge the tokens in the same line of the source code into a sequence; (5) Hash each sequence in each function, create an index for each sequence and save it to a function sequence set, assign a unique ID to each function sequence set, and use multiple function sequence sets in the source code to form a sequence database; (6) Using the closed pattern mining algorithm ClaSp to search for frequent subsequences with a support of at least 2 in the sequence database, and replacing the sequence ID generated in ClaSp with the function sequence set ID in the above sequence database to obtain a frequent result set consisting of frequent subsequences and function sequence set IDs; (7) For each frequent subsequence of the frequent result set, candidate clone pairs are constructed two by two using the function sequence set ID, and they are saved in the clone pair set; (8) Expand and remove duplicate candidate clone pairs by constructing unique keywords for the two function sequence set IDs of each candidate clone pair in the clone pair set, and then remove duplicate clone rows from the remaining unique candidate clone pairs; (9) Setting the clone size, clone size difference, gap threshold, and clone size minimum. (10) Filter the candidate clone pairs in which any candidate clone set size is less than size and the result pairs in which the difference in size between two candidate clone sets in the candidate clone pairs is greater than diff; (11) Sort the clone rows in the candidate clone pairs in the filtered clone pair set from small to large, and group the clone rows by the set gap and min, that is, record the rows whose difference between two consecutive clone rows is less than the gap until it is greater than the gap, generate row groups from the recorded clone rows, and filter out the row groups whose difference is less than the min; (12) Traverse each candidate clone pair in the clone pair set, match the row groups in the candidate clone pairs, and form the matched row groups into final clone pairs to obtain the final clone pair set, thus completing the clone detection of the code.
2. The method according to claim 1, It is characterized in that The step (3) of parsing the abstract syntax tree is to rewrite the interface in the parsing tool to record the start and end positions of the tokens contained in each function in the source code, generate the abstract syntax tree and traverse it to obtain the token of each function.
3. The method according to claim 1, It is characterized in that The step (4) standardizes the tokens in the function by converting numeric types, including integers and floating-point types, into "$", converting characters in single quotes or strings in double quotes into "#", and converting identifier variables into "@".
4. The method according to claim 1, It is characterized in that The step (5) hashes each sequence in each function using the HashPJW algorithm, which is specifically implemented as follows: (5a) Input string; (5b) Read an 8-bit character from the string; (5c) Shift the original hash value left by 4 bits and add the character in step (5b); (5d) Determine whether the highest 4 bits are 0000. If not all 0, mix the highest 8 bits with bits 1-8, and then rewrite the highest 4 bits to 0000; (5e) Repeat (5b)-(5d) until the end of the string to obtain the hash value.
5. The method according to claim 1, It is characterized in that The step (6) uses the closed pattern mining algorithm ClaSp to search for frequent subsequences with a support of at least 2 in the sequence database, and is implemented as follows: (6a) Read the data in the sequence database and build a sequence dictionary tree; (6b) Set the minimum support to 2, mine the frequent 1 sequence set L1 from the sequence database, and generate the frequent closed sequence candidate set FCC; (6c) Take all frequent 1 sequences in L1 as the starting pattern strings, perform sequence expansion and item set expansion on the starting pattern strings to obtain the pattern string p; (6d) The recursive pruning search algorithm is executed on the pattern string p to perform pruning. The pruning rules are as follows: For a recursively generated pattern string p′, if the frequent sequence p precedes p ′ is found, and p contains p ′ , if and only if p and p ′ If the support of ′ All descendant subtrees in the sequence dictionary tree; For the recursively generated pattern string p ′ , if the frequent sequence p precedes p ′ is found, and p ′ contains p if and only if p and p ′ If the support of the sequence dictionary tree is the same, all subtrees in the sequence dictionary tree will be transplanted to p ′ ; (6e) Update the frequent pattern string obtained after pruning to FCC; (6f) Remove the non-closed frequent pattern set in FCC and obtain the closed frequent result set FCS.
6. The method according to claim 1, It is characterized in that In step (7), for each frequent subsequence of the frequent result set, candidate clone pairs are constructed two by two through the function sequence set ID, which is implemented as follows: (7a) Traverse each result in the frequent result set and combine each result in pairs according to the function sequence set ID; (7b) Obtain the cloned row set by obtaining the original row number of the sequence in the function sequence set in the above combination through the index in step (5); (7c) The respective clone row sets and function sequence set IDs in the combination are used together to form a candidate clone pair.
7. The method according to claim 1, It is characterized in that The step (8) expands and removes repeated candidate clone pairs by constructing a unique keyword for the two function sequence set IDs of each candidate clone pair in the clone pair set, which is implemented as follows: (8a) Obtain two function sequence sets, denoted by their IDs x and y, where x <y; (8b) Use x and y to generate a unique keyword x|y; (8c) Traverse the clone pair set, and for candidate clone pairs that generate the same keyword, use the repeated candidate clone pairs to expand the candidate clone pairs that generated the keyword for the first time, so as to ensure that the candidate clone pairs are unique.
8. The method according to claim 1, It is characterized in that The step (12) matches the row groups in the candidate clone pairs, and is implemented as follows: (12a) Obtain the hash value of the row sequence through the index in (5); (12b) Match the hash values in the row group according to the following matching principles: If the first and last items in two row groups are the same, the row groups are matched successfully; If the first item in a row group is the same as the first or second item in another row group, and the size difference between the two row groups is less than the clone size difference diff, the row group matches successfully; If the first item or the second item in a row group is the same as the first item in another row group, and the size difference between the two row groups is less than the clone size difference diff, the row group matches successfully.
Citation Information
Patent Citations
Peptide or peptide complex binding to 2 integrin and methods and uses involving the same
CN103298488A
KR1017802330000B1