Software source code block fingerprint generation method based on text

By code preprocessing and hash value selection, the software source code block fingerprint is generated, which solves the problem of insufficient efficiency and accuracy of code reuse detection in the existing technology and realizes efficient and accurate code reuse detection.

CN120743343APending Publication Date: 2025-10-03BEIJING INST OF COMP TECH & APPL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510814524.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing technology in code reuse detection has the following problems: the coarse-grained comparison method has low accuracy and efficiency, and the fine-grained comparison method is inefficient, which cannot meet the detection needs in large-scale data scenarios, and code modifications have a great impact on detection.

Method used

Through code preprocessing, invalid characters such as comments, spaces, tabs, carriage returns, and line feeds are removed, the code lines are divided into blocks according to their relative positions, the hash values ​​of the code blocks are calculated, and representative hash values ​​are selected as fingerprints, which reduces the fingerprint size and improves detection efficiency and accuracy.

Benefits of technology

It improves the detection rate and accuracy of code reuse detection, reduces the impact of code modification on detection, and improves detection efficiency in large-scale data scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743343A_ABST
    Figure CN120743343A_ABST
Patent Text Reader

Abstract

The invention relates to a text-based software source code block fingerprint generation method, and belongs to the technical field of software code reuse detection. According to the method, through code preprocessing, the influence of code modification modes such as adding or deleting annotations, spaces, table making characters, car returning and line feed codes on code reuse detection in the code reuse process is eliminated, and the detection rate of code reuse detection is increased; small code blocks are filtered by setting a code block character number threshold value, so that the influence of the small code blocks on code reuse detection is reduced, and the accuracy of code block reuse detection is improved; according to the method, the code block hash window is set, the hash value of the code block with the most characters is selected from the code block hash window to serve as the final code block fingerprint, the code block fingerprint set scale is reduced, the retrieval and comparison workload in the code block reuse detection process is reduced, and the code block reuse detection efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of software code reuse detection, and in particular relates to a text-based software source code block fingerprint generation method. Background Art

[0002] Code reuse refers to the practice of reusing existing code to build new software during software development. Code reuse can reduce repetitive tasks during development, thereby improving development efficiency, reducing costs, and shortening development cycles. Furthermore, reused code has generally undergone rigorous testing. Reusing this code to build new software can reduce testing time and improve software quality and reliability.

[0003] Code reuse detection technology is primarily used to assess software code reuse rates. It has broad applications in a variety of fields, including software code plagiarism, software intellectual property protection, and security vulnerability detection. Both academia and industry have conducted extensive research on code reuse detection, proposing numerous methods and techniques for detecting code reuse. However, existing methods and techniques employ either coarse-grained comparison methods, such as file-based comparison, which offers high accuracy and efficiency but also a high rate of missed detections; or fine-grained comparison methods, such as line-based comparison, which offers a low missed detection rate but suffers from low efficiency. These methods are only suitable for small-scale data scenarios and cannot meet the needs of code reuse detection in large-scale data scenarios. Summary of the Invention

[0004] (1) Technical issues to be solved

[0005] The technical problem to be solved by the present invention is: how to generate a set of software source code block fingerprints in a text-based manner, reduce the size of code block fingerprints while ensuring the code reuse detection rate, and reduce the workload of code block reuse detection, thereby improving the efficiency and accuracy of software source code reuse detection.

[0006] (2) Technical solution

[0007] In order to solve the above technical problems, the present invention provides a text-based software source code block fingerprint generation method, comprising the following steps:

[0008] Step 1: Code preprocessing

[0009] Perform code preprocessing on the source code file to obtain valid code;

[0010] Step 2: Code Blocking

[0011] Code block division is to divide the valid code obtained in the first step into code blocks according to the order of code lines and the set number of code lines k to obtain a code block set;

[0012] Step 3: Code Block Hash

[0013] The code block hash is based on the second step. The number of characters in each code block is calculated in sequence, and the code block hash value is calculated using the hash algorithm to form a tuple set of the code block character number and hash value;

[0014] Step 4: Fingerprint selection

[0015] Fingerprint selection is based on the third step, and a set of representative hash values ​​is selected from the obtained code block character count and hash value tuple set as the final code block fingerprint set.

[0016] The present invention also provides a method for realizing software code reuse detection based on the method described above. The method uses the software source code feature of the code block fingerprint set as a comparison object for software code reuse detection, and realizes software code reuse detection by comparing the software source code features.

[0017] The present invention also provides a method for evaluating software code reuse rate based on the software code reuse detection method, wherein the method evaluates the software code reuse rate through software code reuse detection.

[0018] (3) Beneficial effects

[0019] (1) The present invention eliminates the impact of code modification methods such as adding or deleting comments, spaces, tabs, carriage returns, and line feeds during code reuse on code reuse detection through code preprocessing, thereby improving the detection rate of code reuse detection;

[0020] (2) The present invention divides the code into blocks according to the relative positions of the code lines, thereby improving the rationality and effectiveness of the code block division, reducing the impact of code modification methods such as adding, modifying, deleting part of the code, or adjusting the order of part of the code during the code reuse process on the code reuse detection, and improving the detection rate of code reuse detection;

[0021] (3) The present invention filters small code blocks by setting a code block character count threshold, thereby reducing the impact of small code blocks on code reuse detection and improving the accuracy of code block reuse detection;

[0022] (4) The present invention sets a code block hash window and selects the hash value of the code block with the largest number of characters as the final code block fingerprint, thereby reducing the size of the code block fingerprint set, reducing the retrieval and comparison workload in the code block reuse detection process, and improving the efficiency of code block reuse detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 Schematic diagram of the software source code block fingerprint generation process of the present invention;

[0024] Figure 2 Schematic diagram of the software source code preprocessing process of the present invention;

[0025] Figure 3 Schematic diagram of the software code segmentation and hashing process of the present invention;

[0026] Figure 4 Schematic diagram of the software code block fingerprint selection process of the present invention. DETAILED DESCRIPTION

[0027] In order to make the purpose, content and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below with reference to the accompanying drawings and examples.

[0028] The present invention provides a text-based software source code block fingerprint generation method. The method obtains valid code through code preprocessing, eliminating the impact of code modification methods such as adding or deleting comments, spaces, tabs, carriage returns, and line feeds during code reuse on code reuse detection; obtains a code block set through code segmentation, reducing the impact of code modification methods such as adding, modifying, deleting part of the code, or adjusting the order of part of the code during code reuse on code reuse detection; obtains a code block character count and hash value tuple set by calculating the number of characters and hash value of each code block, providing a basis for the next step of code block fingerprint selection; obtains a final code block fingerprint set through code block fingerprint selection, thereby improving the efficiency of code block reuse detection. A schematic diagram of the text-based software source code block fingerprint generation process is shown in FIG. Figure 1 As shown, the specific implementation steps of the text-based software source code block fingerprint generation method are described in detail below.

[0029] Step 1: Code preprocessing

[0030] Using hash algorithms such as MD5 or SHA, code blocks can be mapped into a set of fixed-length hash values, thereby obtaining a code block fingerprint. Hash algorithms are sensitive to input data; even slight modifications to the original code can result in significant changes in the hash value. To eliminate the influence of irrelevant content on the calculation of the code block fingerprint set, the source code file must be preprocessed before calculating the code block fingerprint set to obtain valid code. The main contents of code preprocessing include:

[0031] (1) Removing comments: There are many types of comments in the code text, such as inline comments, out-of-line comments, single-line comments, and multi-line comments. In order to eliminate the impact of no comments on code reuse detection, these comments need to be removed.

[0032] (2) Remove spaces, tabs, carriage returns, and line feeds: During the development process, many spaces, tabs, carriage returns, and line feeds are added to make the code more organized. In order to eliminate the impact of spaces, tabs, carriage returns, and line feeds on code reuse detection, these spaces, tabs, carriage returns, and line feeds need to be removed.

[0033] (3) Recording code line positions: During the development process, organizing the code into blocks using comments or blank lines can improve code readability and comprehensibility. For example, function blocks can be separated by comments and blank lines. Code line positions can be used as the basis for block division in the second step, code block division. For example, code lines separated by blank lines before and after can be divided into a code block, representing a complete grammatical meaning, such as the entire code is the definition of a complete function, thereby improving the rationality and effectiveness of code block division.

[0034] Based on the above analysis of code preprocessing content, the code preprocessing process is as follows Figure 2 As shown in the figure, the code preprocessing process is as follows:

[0035] (1) Let the total number of lines in the source code file F be n, then the source code file F can be expressed as F = {S i |1≤i≤n}, that is, the text string of the i-th row is S i ; The code line set obtained after preprocessing the source code file F is F'; initialize the variable i = 1;

[0036] (2) If 1≤i≤n, jump to step (3) to continue processing; otherwise, the code preprocessing ends normally;

[0037] (3) For the text string S in row i i , if S i If it is a comment line or a blank line, set S' i Add code line set F', execute i=i+1, jump to step (2) to continue processing; otherwise, delete the text string S i Invalid characters such as comments, spaces, tabs, carriage returns, and line feeds are included in the result, and S' i , S' i Add code line set F', execute i=i+1, and jump to step (2) to continue processing.

[0038] Step 2: Code Blocking

[0039] During software development or maintenance, reused code files often need to be edited for reasons such as adding new features, fixing software defects, or refactoring. This includes adding, modifying, or deleting functions, adding, modifying, or deleting statements, and adjusting the order of functions or statements. Such editing of reused files often renders code reuse detection methods based on file fingerprints or code fingerprints ineffective.

[0040] Code segmentation is to divide the valid code obtained in "the first step, code preprocessing" into code blocks according to the order of code lines and the set number of code lines k to obtain a set of code blocks. The code segmentation process diagram is as follows Figure 3 As shown in the figure, the process of code segmentation is as follows: according to the number of code lines k, the code line set F' obtained after preprocessing the source code file F is segmented into blocks, and the i-th code block block i Expressed as:

[0041] block i ={S' i ,S' i+1 ,…,S' i-k+1}, 1≤i≤n-k+1

[0042] Among them, S' i ,S' i+1 ,…,S' i-k+1 They represent the preprocessed text strings of line i, line i+1, and line i-k+1 in the code line set F' respectively; finally, the source code file F with a total of n lines is divided into code blocks according to the number of code lines k, and the code block set B = {block i |1≤i≤n-k+1}.

[0043] Step 3: Code Block Hash

[0044] Code block hashing is based on the "second step, code segmentation". The number of characters in each code block is calculated in sequence, and the code block hash value is calculated using a hash algorithm such as MD5 or SHA to form a set of code block character count and hash value tuples.

[0045] Code block hashing maps code blocks of arbitrary length into fixed-length hash values, reducing the transmission and storage costs of code block fingerprints and improving the efficiency of code block reuse detection. However, due to hash collision issues with hash algorithms such as MD5 and SHA, the code block hashing process requires the use of at least two hash algorithms to calculate the code block hash value. This ensures code block consistency through multiple hash values, thereby reducing the false positive rate of code block reuse detection.

[0046] The code block hashing process diagram is as follows Figure 3As shown, the code block hashing process is as follows: for each code block i ∈B, 1≤i≤n-k+1, calculation code block block i The number of characters (or bytes) contained in nChar i , and the code block hash value (hash all characters in the code block) vHash i , get the code block block i The number of characters and hash value tuple (nChar i ,vHash i ); Finally, the source code file F with a total number of lines of n is divided into code blocks according to k lines and the code block hash values ​​are calculated, and the code block character count and hash value tuple set H = {(nChar i ,vHash i )|1≤i≤n-k+1}.

[0047] Step 4: Fingerprint selection

[0048] Fingerprint selection is based on the third step, code block hashing, and the obtained code block character count and hash value tuple set H = {(nChar i ,vHash i )|1≤i≤n-k+1}, a set of representative hash values ​​is selected as the final code block fingerprint set.

[0049] The fingerprint selection method is to filter the hash values ​​of small code blocks according to the number of code block characters. By setting the code block character number and hash value tuple window (code block hash window), the code block hash value with the largest number of code block characters is selected from the code block character number and hash value tuple window as the code block fingerprint representing the code block character number and hash value tuple window, thereby removing duplicate code block fingerprints, reducing the size of code block fingerprints, and improving the efficiency of software code block reuse detection. The code block fingerprint selection process is shown in the figure below. Figure 4 The specific algorithm for selecting code block fingerprints is described as follows:

[0050] (1) Set the code block character count and hash value tuple window w, the code block character count threshold T; initialize the final code block fingerprint set as Variables i=1, j=1;

[0051] (2) According to the code block character number and hash value tuple window w, the code block character number and hash value tuple set H is divided into m-w+1 code block character number and hash value tuple subsets, where m=n-k+1, then the i-th code block character number and hash value tuple subset W i It can be expressed as:

[0052] W i={(nChar i ,vHash i ),(nChar i+1 ,vHash i+1 ),…,(nChar i+w-1 ,vHash i+w-1 )}

[0053] (3) If 1≤i≤m-w+1, jump to step (4) to continue processing; otherwise, P is the selected code block fingerprint set, and the code block fingerprint selection ends normally;

[0054] (4) From the subset W of the number of characters and hash value tuples of the i-th code block i Select the code block with the largest number of characters and the hash value tuple (nChar j ,vHash j ), that is, nChar j =max{nChar i ,nChar i+1 ,…,nChar i+w-1}, if there are multiple identical nChar j , then take the one with the largest j value (that is, take the last code block character number and hash value tuple); if nChar j ≤T, which means that the minimum requirement for the content (i.e., the number of characters) of the code block fingerprint is not met, and the code block hash value vHash j If the code block fingerprint cannot be selected as the final one, execute i=i+1 and jump to step (3) to continue processing; otherwise, jump to step (5) to continue processing;

[0055] (5) If nChar j The corresponding hash value vHash j ∈P, which means that the code block fingerprint is already included in the set P, so there is no need to add the hash value vHash to P j , execute i=i+1, and jump to step (3) to continue processing; otherwise, it indicates the hash value vHash j It is the fingerprint of the newly selected code block, and the code block hash value vHash j Add to the code block fingerprint set P, execute i=i+1, and jump to step (3) to continue processing.

[0056] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A text-based software source code block fingerprint generation method, characterized in that: The following steps are involved: Step 1: Code preprocessing Perform code preprocessing on the source code file to obtain valid code; Step 2: Code Blocking Code block division is to divide the valid code obtained in the first step into code blocks according to the order of code lines and the set number of code lines k to obtain a code block set; Step 3: Code Block Hash The code block hash is based on the second step. The number of characters in each code block is calculated in sequence, and the code block hash value is calculated using the hash algorithm to form a tuple set of the code block character number and hash value; Step 4: Fingerprint selection Fingerprint selection is based on the third step, and a set of representative hash values ​​is selected from the obtained code block character count and hash value tuple set as the final code block fingerprint set.

2. The method according to claim 1, wherein The code preprocessing process is as follows: (1) Assume that the total number of lines in the source code file F is n, then the source code file F is represented by F = {S i |1≤i≤n}, that is, the text string of the i-th row is S i ; The code line set obtained after preprocessing the source code file F is F'; Initialize variable i = 1; (2) If 1≤i≤n, jump to step (3) to continue processing; Otherwise, code preprocessing ends normally; (3) For the text string S in row i i , if S i If it is a comment line or a blank line, set S' i Add code line set F', execute i=i+1, and jump to step (2) to continue processing; Otherwise, delete the text string s i Invalid characters such as comments, spaces, tabs, carriage returns, and line feeds are included in the result, and S' i , S' i Add code line set F', execute i=i+1, and jump to step (2) to continue processing.

3. The method according to claim 2, wherein The specific process of code block is as follows: according to the number of code lines k, the code line set F' is divided into blocks, and the i-th code block block i Expressed as: block i ={S' i ,S' i+1 ,…,S' i-k+1 },1≤i≤n-k+1 Among them, S' i ,S' i+1 ,…,S' i-k+1 They represent the text strings of line i, line i+1, and line i-k+1 in the code line set F' respectively; the final code block set B = {block i |1≤i≤n-k+1}.

4. The method according to claim 1, wherein Code block hashing specifically maps code blocks of arbitrary length into hash values ​​of fixed length, and uses at least two hash algorithms to calculate the code block hash values, and confirms the consistency of the code blocks through multiple hash values.

5. The method according to claim 3, wherein The specific process of code block hashing is as follows: for each code block i ∈B, 1≤i≤n-k+1, calculation code block block i The number of characters nChar contained i , and the code block hash value vHash i , get the code block block i The number of characters and hash value tuple (nChar i ,vHash i ); Finally, we get the code block character count and hash value tuple set H = {(nChar i ,vHash i )|1≤i≤n-k+1}.

6. The method according to claim 1, wherein The fingerprint selection method is to filter the hash values ​​of small code blocks according to the number of code block characters, set the code block character number and hash value tuple window, and select the code block hash value with the largest number of code block characters from the code block character number and hash value tuple window as the code block fingerprint representing the code block character number and hash value tuple window.

7. The method according to claim 5, wherein The fingerprint selection process is as follows: (1) Set the code block character count and hash value tuple window w, the code block character count threshold T; initialize the final code block fingerprint set as Variables i=1, j=1; (2) According to the code block character number and hash value tuple window w, the code block character number and hash value tuple set H is divided into m-w+1 code block character number and hash value tuple subsets, where m=n-k+1, then the i-th code block character number and hash value tuple subset W i Expressed as: W i ={(nChar i ,vHash i ),(nChaar i+1 ,VHash i+1 ),…,(nChar i+w-1 ,vHash i+w-1 )} (3) If 1≤i≤m-w+1, jump to step (4) of fingerprint selection and continue processing; Otherwise, P is the selected code block fingerprint set, and the code block fingerprint selection ends normally; (4) From the subset W of the number of characters and hash value tuples of the i-th code block i Select the code block with the largest number of characters and the hash value tuple (nChar j ,vHash j ), that is, nChar j =max{nChar i ,nChar i+1 ,…,nChar i+w-1 }, if there are multiple identical nChar j , then take the one with the largest j value, that is, take the last code block character number and hash value tuple; if nChar j ≤T, which means that the minimum requirement for the number of characters of the code block fingerprint is not met, and the code block hash value vHash j If the code block fingerprint cannot be selected as the final one, execute i=i+1 and jump to step (3) of fingerprint selection to continue processing; otherwise, jump to step (5) of fingerprint selection to continue processing; (5) If nChar j The corresponding hash value vHash j ∈P, which means that the code block fingerprint is already included in the set P, so there is no need to add the hash value vHash to P i , execute i=i+1, and jump to step (3) of fingerprint selection to continue processing; otherwise, it indicates that the hash value vHash j It is the fingerprint of the newly selected code block, and the code block hash value vHash j Add the code block fingerprint set P, execute i=i+1, and jump to step (3) of fingerprint selection to continue processing.

8. The method according to claim 1, wherein The hash algorithm includes MD5 and SHA algorithms.

9. A method for implementing software code reuse detection based on the method according to any one of claims 1 to 8, characterized in that: The method uses the software source code feature of the code block fingerprint set as a comparison object for software code reuse detection, and realizes software code reuse detection through comparison of software source code features.

10. A method for evaluating software code reuse rate based on the method according to claim 9, characterized in that: This method evaluates the software code reuse rate through software code reuse detection.