Software similar project screening method based on similarity hash technology
By using Simhash and Minhash hash algorithms to filter feature vectors of software projects, the problem of low efficiency in similarity screening among massive projects is solved, enabling efficient software source tracing analysis and code reuse detection, thereby improving development efficiency and compliance management.
Patent Information
- Application Number
- CN202511019887.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies struggle to efficiently filter out highly similar projects when faced with a massive number of software projects, resulting in low efficiency in software source tracing analysis and code reuse detection.
Simhash and Minhash hash algorithms are used to extract feature vectors of software projects. Candidate projects are initially screened by calculating the Simhash Hamming distance and Minhash signature similarity. Then, by combining weighted fusion and hierarchical screening strategies, fine-grained source tracing analysis is carried out at the file level or code fragment level.
It significantly improves the efficiency of software project similarity screening, reduces the risk of redundant development, enhances code traceability, improves software supply chain management and security review, supports version management and compliance review, and reduces the risk of intellectual property disputes.
Smart Images

Figure CN120995116A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a software similar item screening method based on a similarity hash technology, and belongs to the technical field of intelligent software analysis and processing. BACKGROUND
[0002] With the wide use of open source software and the popularity of code reuse, different sources of open source frameworks, components or reused software codes of similar projects are often contained in software projects. For software developers, researchers and legal practitioners, it is of great significance to track the source and evolution process of the code, which is usually referred to as code provenance analysis. Traditional software component analysis systems are mainly based on file-level provenance analysis or code snippet-level provenance analysis. When facing a large number of projects and contained files and codes, how to efficiently screen out similar projects is a challenge.
[0003] At present, the source code provenance of software projects is usually realized by directly comparing the feature vectors of each file or the feature vectors of code snippets in the project. Through the retrieval of existing patents, three patents with main classification numbers G06F21 / 56, G06F21 / 57 and G06F21 / 16 can be found. The method of the three patents is to compare files or code snippets. Although this method is accurate, it is low in efficiency when facing hundreds of millions or more project files. Therefore, how to improve the retrieval efficiency while maintaining accuracy is an important problem in current research.
[0004] Principle and advantage of Simhash: Simhash is an algorithm based on the locality sensitive hashing (LSH, Locality Sensitive Hashing) technology, which is used to convert large texts or codes into fixed-length "hash fingerprints" or "feature vectors" for similarity judgment. Its core advantage is high calculation efficiency, especially suitable for rapid approximate similarity detection of massive data. Simhash generates low-dimensional hash values to retain the global features of the text, so it performs particularly well when dealing with small projects or short text files. Since the hash value can be directly used to calculate the Hamming distance, Simhash can quickly identify the similarity between two files or projects, and occupies less computing resources, which is very suitable for comparing a large number of projects.
[0005] Principle and advantage of Minhash: Minhash also belongs to the local sensitive hashing technology, but it is mainly used for Jaccard similarity evaluation. Minhash generates the overall hash signature by dividing the file or text into multiple small phrase fragments or code blocks and calculating the hash value of each small fragment. This method is particularly suitable for large-scale data sets or files with complex content structure. Minhash can accurately capture the local similarity between different documents, even if two files or projects have large differences on the surface, Minhash can still detect potential similarities through its fine-grained splitting method. Therefore, in the case of large projects, complex structures and the need to accurately capture content details, Minhash is particularly effective.
[0006] But in practical application, the size, structure and file quantity of different projects are different, a single algorithm may not be able to handle all situations. For a detected project, trace back from the project code base to one or a group of high similarity projects, due to the particularity of each project, we cannot predict in advance which attributes of the project can better reduce the data dimension and retain the characteristics of the project. When extracting the project feature vector, you can choose a variety of project attributes, such as extracting a feature vector based on the content of each file, or extracting a feature vector based on the physical structure attributes of each file, such as file name, file path, function information in the file, AST (abstract syntax tree) information, etc. A method is needed to efficiently find the similarity between projects while saving storage space, and considering that different detection projects cannot determine whether the feature information of different projects in the project code base is better extracted by Simhash feature vector or Minhash feature vector. In the case of comparison and verification, multiple methods can be used. SUMMARY
[0007] The technical problem of the present application is to overcome the shortcomings of the prior art and provide a software similar project preselection method based on similarity hashing technology, which is used for software source code traceability analysis, source code reuse analysis and other similar project screening methods. Simhash and Minhash are used to compare the similarity of projects, and the technology is used to improve the efficiency of project similarity screening.
[0008] The technical solution of the present application is: a software similar project screening method based on similarity hashing technology, comprising:
[0009] Selecting a plurality of sample projects to construct a knowledge base; preparing a detected project;
[0010] The file names of each sample project are spliced to generate a sample project text string; Simhash feature vectors and Minhash feature vectors of each sample project are extracted using Simhash hash algorithm and Minhash hash algorithm respectively on the sample project text string;
[0011] The file name of the detected project is spliced to generate a detected project text string, and Simhash feature vectors and Minhash feature vectors of the detected project are extracted using Simhash hash algorithm and Minhash hash algorithm respectively on the detected project text string;
[0012] The Simhash feature vectors of the detected project are compared with the Simhash feature vectors of each sample project to obtain corresponding Simhash Hamming distance values; the Minhash feature vectors of the detected project are compared with the Minhash feature vectors of each sample project to obtain corresponding Minhash signature similarity;
[0013] According to the Simhash Hamming distance values and the Minhash signature similarity, a plurality of candidate projects are preliminarily screened out;
[0014] Each candidate project is subjected to file-level or code snippet-level traceability analysis, and finally a target project is screened out.
[0015] Preferably, when a plurality of sample projects are selected to construct the knowledge base:
[0016] The number of sample projects in the knowledge base is greater than or equal to 30, the sizes of the sample projects are different, and the sample projects have different versions similar to the detected project.
[0017] Preferably, before the file names are spliced into the detected project text string, the files are sorted in alphabetical order, and the number of effective files participating in the analysis of each project and the size information of the effective files are recorded.
[0018] Preferably, when a plurality of candidate projects are preliminarily screened out according to the Simhash Hamming distance values and the Minhash signature similarity:
[0019] A Simhash Hamming distance threshold is determined to screen out sample projects with a Simhash Hamming distance less than or equal to the threshold, and a Minhash signature similarity threshold is determined to screen out sample projects with a Minhash signature similarity greater than or equal to the Minhash signature similarity threshold; the intersection of the sample projects screened out by the two methods is taken as a candidate project;
[0020] Preferably, when multiple groups of candidate items are preliminarily screened out according to the Simhash Hamming distance value and the Minhash signature similarity:
[0021] The comprehensive similarity score is obtained by weighted fusion of the Simhash Hamming distance and the Minhash signature similarity of each sample item, and the sample items are sorted and the candidate items are screened out according to the comprehensive similarity score.
[0022] Preferably, when the Simhash Hamming distance threshold is determined:
[0023] According to the local sensitive hashing theory, the Simhash Hamming distance threshold is determined by the quartile range method or by introducing a Gaussian mixture model;
[0024] The Simhash Hamming distance threshold determined by introducing a Gaussian mixture model is specifically:
[0025] The visualization analysis of the distribution of the Simhash feature vector is increased to determine the distribution characteristics of the Simhash feature vector; if the Simhash feature vector presents a mixed distribution characteristic, the Simhash feature vector is separated into sub-distributions by introducing a Gaussian mixture model, and different threshold calculation methods are set for different sub-distributions.
[0026] Preferably, when the Minhash signature similarity threshold is determined:
[0027] The Jaccard similarity of the file name overlap degree is calculated, and the Minhash signature similarity threshold is determined by normal distribution.
[0028] Preferably, when the Minhash signature similarity is calculated, the signature length and the arrangement number can be dynamically adjusted, and specifically:
[0029] The optimal arrangement number k is: k = ln(1-t) / ln(2), where t is the expected Jaccard similarity threshold;
[0030] The signature length is 128 bits by default, and is extended to 256 bits or 512 bits when the similarity precision needs to be improved.
[0031] Preferably, the candidate items are compared at a finer granularity of file level and code snippet level to finally screen out the most similar target item, and specifically:
[0032] The fingerprint feature vector is extracted from all valid file contents in each candidate item, and the same fingerprint feature vector is extracted from all file contents in the detected item;
[0033] The fingerprint feature vector of each file in the detected item is compared with the fingerprint feature vector of each file in the candidate item.
[0034] If equal, the number of valid lines of the file is accumulated, and the two equal files no longer participate in subsequent comparison; if not equal, the following steps are performed:
[0035] The window size and step length are selected, the code in each window of the detected project file is compared with the code in the window of the candidate project file, the feature vector is matched, and the number of code lines corresponding to the feature vector that matches the file is accumulated; and then the total number of matched lines of the entire candidate project is counted as the project similarity;
[0036] The project corresponding to the maximum project similarity is the most similar project.
[0037] Compared with the prior art, the present application has the following advantages:
[0038] (1) Accelerate the preliminary screening process: through the similarity hashing technology, similar projects or code libraries can be quickly identified in a large number of projects. This screening method significantly improves the efficiency of enterprises and developers in screening projects in the early stages of the project, avoiding tedious comparison work by manual work. Through the hashing technology, the evaluation time can be shortened, and irrelevant or repetitive projects can be quickly filtered out.
[0039] (2) Reduce the risk of code reuse and repeated development: the similarity hashing technology helps the development team quickly detect repeated modules or similar codes in existing code libraries, thereby avoiding the situation of "reinventing the wheel". This automated code similarity detection can significantly reduce unnecessary repeated development of the development team, improve code reuse rate, and ultimately reduce development cost and time consumption.
[0040] (3) Improve code tracking and tracing ability: another significant role of similarity hashing is to enhance the code tracing ability. It can identify the source of the code in the project and quickly identify third-party code or open source code involved in the code library. This code tracking function can help enterprises or developers understand the composition of the project, ensure code compliance, and reduce the risk of intellectual property disputes.
[0041] (4) Improve software supply chain management efficiency: for projects that use a large number of third-party libraries and external dependencies, the similarity hashing technology can more easily identify which external code or library is used in the project, thereby helping enterprises better manage their software supply chain. This ability helps improve supply chain management efficiency and ensures that the third-party software code used by the enterprise is verified and meets safety and compliance requirements.
[0042] (5) Strengthening project security review: In the preliminary screening stage, similarity hashing technology can also help find potential security vulnerabilities or unauthorized code usage. By quickly detecting whether the code in the project is similar to known third-party libraries or open source projects, enterprises can identify potential security risks in advance and avoid introducing risky code into the formal development process.
[0043] (6) Support project version management and comparison: Similarity hashing technology can also be applied to version management, quickly identifying changes in projects by comparing the hash values of different versions. This capability allows developers to easily track differences between versions and understand the specific changes made, facilitating better version control and change management.
[0044] (7) Improve R&D efficiency and quality: Project similarity hashing technology helps enterprises identify and reuse high-quality code modules, ensuring that developers can innovate and improve based on existing code snippets. This not only improves R&D efficiency but also ensures steady improvement in code quality, reducing the likelihood of code errors and security vulnerabilities.
[0045] (8) Assist in compliance and legal risk management: The application of similarity hashing technology plays an important role in compliance review. Enterprises can use similarity detection to confirm whether the code in the project complies with relevant open source license requirements, avoiding intellectual property infringement issues. This function is of great significance for legal compliance and effectively reduces the risk of intellectual property and legal disputes for enterprises. BRIEF DESCRIPTION OF DRAWINGS
[0046] Figure 1 Processing flow diagram for comparing the detected project with the sample project. DETAILED DESCRIPTION
[0047] The present application discloses a software similar project screening method based on similarity hashing technology, aiming to solve the problem of low efficiency in searching target projects in software source code traceability analysis and code reuse detection in existing large-scale project code libraries. The system extracts project file name information, uses Simhash and Minhash hashing algorithms, and extracts feature vectors for calculating the similarity between projects. This method first sorts the target project files, extracts each valid file name, and then splices and calculates the feature vectors to compare the project feature vectors and quickly determine similar projects. At the same time, the system introduces multi-threading technology for parallel processing, effectively shortening the time for project file name splicing and hashing calculation, and is suitable for large-scale project traceability analysis and similarity detection. The present application has the advantages of high efficiency, accuracy, and easy expansion, and is suitable for software traceability analysis, independent research and development rate evaluation and identification, code reuse detection, and intellectual property protection fields.
[0048] The method of the application comprises the following steps:
[0049] The candidate items are screened according to the Simhash Hamming distance value and the Minhash signature similarity value, specifically as follows:
[0050] Data selection and preprocessing
[0051] Thirty sample items are selected, including C, C++, Python, Verilog, VHDL and Java six development languages, which can be mixed language projects of these development languages. These sample items include multiple versions of a detected item, one of which is the detected item. The compressed package of the thirty sample items is placed in the same directory. The common file suffix types corresponding to the six development languages are the main program file types of the development languages.
[0052] Reading and splicing the file names of the sample items
[0053] When processing the script to process the thirty sample items and the detected item, first, the file paths and file names are sorted in the same order according to the letters, and each file name in the sample item is read and spliced one by one.
[0054] Extracting feature vectors using Simhash and Minhash
[0055] The feature vectors are extracted from the spliced text strings using Simhash and Minhash. The feature vector information is extracted from the spliced text strings of all effective file names of each item using Simhash and Minhash.
[0056] Comparing the detected item with the sample items
[0057] The detected item extracts the feature vectors in the same way as the sample items, and the Simhash feature vectors of the detected item are compared with the Simhash feature vectors of the sample items stored persistently, and the Hamming distance values are recorded. The Minhash feature vectors of the detected item are compared with the Minhash feature vectors of the thirty sample items stored persistently, and the signature similarity values are recorded.
[0058] Obtaining a group of preliminary selected items through similarity analysis
[0059] The detected item is compared with the two groups of feature vectors of the sample items, and a group of preliminary selected items is obtained through a dynamic threshold adaptive decision mechanism, normality test enhancement and weighted scoring mechanism.
[0060] The feature vectors of the effective files in each project in the above set of preliminary selection projects are compared with the feature vectors of the effective files of the detected project, and if they are the same, the total number of matching lines in the corresponding project is counted. For files that are not equal, the feature vectors are compared in 10 lines of code windows with a step of 1 line of code, and the number of code lines corresponding to the equal feature vector windows is accumulated. Finally, the number of matching lines of these files is added to the total number of matching lines,
[0061] Code snippet similarity = total number of matching lines / number of effective code lines of the detected project. The code snippet with the maximum similarity is the final most similar project.
[0062] The Simhash is used to calculate the Hamming distance of the file names of the detected project and the sample projects. According to the local sensitive hashing (LSH) theory, the interquartile range (IQR) method is used to dynamically determine the Hamming distance threshold, and projects with a Hamming distance less than or equal to the threshold are selected. Minhash is used to calculate the signature similarity of the file names of the detected project and the sample projects. Based on the probability estimation of Jaccard similarity, the file name overlap is reflected, and the signature similarity threshold is determined by assuming a normal distribution. Projects with a signature similarity greater than or equal to the threshold are selected. When selecting candidate sample projects, the intersection of the projects selected by the two methods can be used as the candidate set of similar projects.
[0063] When determining the Hamming distance threshold, in addition to using the interquartile range method, visual analysis of data distribution can be added to more accurately determine whether the data conforms to a normal distribution. If the data presents a mixed distribution characteristic, a Gaussian mixture model (GMM) can be introduced to separate the data into sub-distributions. Different thresholds can be calculated for different sub-distributions to improve the adaptability of the threshold.
[0064] For signature similarity calculation, the parameters of Minhash can be optimized. The signature length and the number of permutations can be dynamically adjusted according to the size of the data set and the expected similarity accuracy. For example, the optimal number of permutations (k) can be calculated by the formula k = ln(1-t) / ln(2), where t is the expected Jaccard similarity threshold. The signature length can be extended from the default 128 bits to 256 bits or 512 bits according to actual needs, reducing the probability of hash collisions and improving the accuracy of signature similarity calculation.
[0065] When selecting candidate sample projects, a weighted fusion model can be constructed to assign dynamic weights to the Hamming distance of Simhash and the signature similarity of Minhash. For example, by training a large number of known similarity project samples, the importance of different features in similarity judgment can be calculated using random forest algorithm, and the weight distribution of Hamming distance and signature similarity can be determined, and a comprehensive similarity score can be obtained by combining the scores of the two, and similar projects can be screened according to the score, which can more comprehensively utilize the advantages of the two algorithms and improve the accuracy of similar project screening.
[0066] When selecting candidate sample projects, a hierarchical screening strategy can also be implemented. First, use Simhash for rapid rough screening and set a relatively loose Hamming distance threshold (such as based on mean-2 standard deviation), quickly filter out obviously dissimilar projects, and greatly reduce the subsequent calculation amount; then, use Minhash for fine screening of the candidate set after rough screening, and further filter out highly similar projects according to the optimized signature similarity threshold and the weighted fusion model, to improve the screening efficiency and accuracy.
[0067] For candidate projects, more detailed file-level and code snippet-level comparison is performed to finally screen out the most similar projects, which are as follows:
[0068] Extract the fingerprint feature vector of all valid file contents in the selected group of projects, and extract the same fingerprint feature vector from all file contents in the detected project;
[0069] Take each project in the selected group of projects as a unit, compare each file feature vector in the detected project with each file feature vector in the preliminary selected project, if they are equal, count the number of valid lines of the file, and these two files will not participate in subsequent comparison. If they are not equal, proceed to the next step;
[0070] Take the code feature vector as a window of 10 lines, and take each 10 lines in the file in the detected project and the 10 lines in the file in the preliminary selected project as feature vectors for comparison, and count the number of code lines corresponding to the matching feature vectors. The number of matching lines of the current file is added to the total number of matching lines of the entire project;
[0071] Calculate the code snippet similarity of each project in the group of preliminary selected projects and the detected project, which is the total number of matching lines / the number of valid code lines in the detected project;
[0072] The project with the highest similarity in the above group of preliminary selected projects is the final most similar project.
[0073] In summary:
[0074] Threshold decision of multi-modal feature fusion, using Simhash and MinHash two kinds of feature vectors at the same time, through complementary verification, compared with single feature method, the accuracy is improved.
[0075] Adaptive threshold domain migration, a general framework based on data distribution to automatically select threshold calculation method (normal -> mean ± standard deviation, non-normal -> IQR extension) is proposed, which can be applied to code tracing in different fields.
[0076] Efficiency optimization of hierarchical screening, through the two-stage process of coarse screening (fast filtering) + fine screening (multi-dimensional verification), the calculation time is shortened while maintaining the accuracy.
[0077] Embodiments:
[0078] As shown in Figure 1 The method of the application comprises the following steps:
[0079] Step 1: Data selection and preprocessing
[0080] In engineering practice, software source code traceability analysis generally adopts the method of building a "knowledge base", that is, before traceability branching, a knowledge base for comparative analysis is established, which may have a size of tens of millions or even larger, including various development languages, various file types, and different sizes of projects and files. The process of building a knowledge base can be separated from the process of traceability analysis and completed before traceability analysis. In order to verify the method of the application, the sample projects are analyzed and stored before analysis. When selecting software source code projects, projects of different development languages and different project sizes are selected to reflect the above-mentioned characteristics of software projects in engineering practice. In this embodiment, 30 projects of different development languages and different project sizes are selected as sample projects to simulate the characteristics of the knowledge base. Six development language projects including C, C++, Python, Verilog, VHDL and Java are selected. These sample projects include multiple versions of a project to verify the similarity of projects. At the same time, a project is selected as the detected project. The 30 sample project compressed packages are placed in the same directory.
[0081] The sample projects are one of six development languages, including C, C++, Java, Python, Verilog and VHDL, and the selection range is small and focused on the algorithm itself, without paying attention to the processing of annotations and spaces due to different development languages.
[0082] The 30 sample projects are selected to be of different sizes, which aims to verify the exclusion of the influence of different project sizes on project similarity.
[0083] There are multiple different versions of the project to be analyzed in the sample project, so a version of the project can be selected as the detected project, and the most similar project to the detected project can be found.
[0084] The common file suffix types corresponding to the six development languages are defined, see Table 1.
[0085] Table 1 Six development languages
[0086] Serial number Development language Corresponding file suffix 1. Java .java;.xml;.html 2. C .c;.h 3. C++ .cpp;.h 4. Python .py 5. Verilog .v 6. VHDL .vhd;.vhdl
[0087] Meanwhile, the detected project is prepared, and an project is selected as the detected project in this experiment. The detected project is apollo_v0.4.0, and the single-source similarity recognition is verified. It is a lower version of the apollo project among the 30 sample projects, so the target project of the traceability analysis result in this experiment is three versions of the apollo project. Because they are different versions of the same project, the project similarity is the highest.
[0088] Step 2: Read the file name of the sample project to splice
[0089] The 30 sample projects are decompressed one by one into a certain directory. In order to prevent the difference of the experimental results caused by the different reading order of the files, the file names in the sample project are read one by one after sorting the file names in alphabetical order according to the file path and file name. The file name is spliced into a text string without including the path level information of the file. In order to provide more reference information for open source software project traceability analysis, the number of effective files and the size of the effective files of each project participating in the analysis can be recorded at the same time in the above processing.
[0090] Step 3: Extract feature vectors using Simhash and Minhash
[0091] The text string spliced by the Simhash and Minhash is extracted from the feature vector. The text string spliced by the Simhash and Minhash is extracted from the feature vector information of each project. In order to facilitate the reuse of sample projects for similarity comparison, all project feature vectors and file quantity and size information are stored persistently, and are respectively corresponding to the project compressed package name. In this way, when multiple projects are detected, the sample project only needs to extract the feature vector once.
[0092] Step 4: Compare the detected project with the sample project
[0093] The detected project extracts two kinds of feature vectors in the same way as the sample project. The following steps are taken to compare.
[0094] The following is the sample project experimental steps:
[0095] 1) In turn, the Simhash feature vector of the detected project is compared with the Simhash feature vector of the sample project stored in the persistent storage, and the Hamming distance value is recorded.
[0096] 2) In turn, the Minhash feature vector of the detected project is compared with the Minhash feature vector of the 30 sample projects stored in the persistent storage, and the signature similarity value is recorded. The signature difference value used by Minhash, the greater the signature difference value of two projects, the more similar the two projects. In the calculation, the Minhash signature difference value is less than 1, in order to facilitate reading, we expand 100 times and take integer. The comparison results are recorded in the table for easy comparison of project similarity.
[0097] The following Table 2 is the result statistics table of the project in the experimental process.
[0098] Table 2 is the detection result of the detected project apollo_v0.4.0, in which the three projects 2-4 in the detected project serial number are the target projects of the traceability.
[0099] Table 2: Similarity comparison statistics table of the detected project apollo_v0.4.0 and 30 sample projects
[0100]
[0101]
[0102] Step five: Similarity analysis to get the preliminary selected project
[0103] The four-part range method is used for the two groups of feature vectors of the detected project and the sample project respectively, and the threshold value is calculated through the analysis of whether the data is normally distributed. The analysis method is as follows:
[0104] Table 2 analysis: The Hamming distance of Simhash reflects the change of the file name in the project, which is based on the local sensitive hash (LSH) theory, and the threshold value is dynamically determined by the quartile range (IQR) method.
[0105] In Table 2, 30 Hamming distance values are analyzed using the quartile range method, mean = 17.35, standard deviation = 5.50, p-value > 0.05, the data conforms to the normal distribution, and the threshold is calculated using "mean-1.5*standard deviation". The calculated threshold is 9.10. This formula covers about 93.3% of similar data. According to the Hamming distance less than or equal to 9.10, project numbers 2, 3, and 4 can be found, and these three projects are exactly the target projects. The detected project and the target project file name fingerprint difference is very small, belonging to the "highly similar" group, which needs to be prioritized for subsequent refined traceability analysis.
[0106] The signature similarity of Minhash is a probability estimate based on Jaccard similarity, and the signature similarity reflects the overlap of the file name. The threshold is determined by the normal distribution. In Table 1, Q3 = 29.75, IQR = 10.50, p-value < 0.05, the data is not normally distributed, and the threshold is calculated using "Q3+1.5*IQR", which is 45.50. This method can identify outliers that are significantly higher than the core data, that is, high similarity projects, by using the upper quartile + 1.5 standard deviation. According to the signature threshold greater than or equal to 45.50, projects 2, 3, and 4 can be found, which are also the target projects.
[0107] The projects found by the two methods in Table 2 are highly consistent, and the effectiveness of finding similar projects is confirmed by the two methods.
[0108] The quartile range method for Hamming distance and signature similarity is a dynamic threshold obtained from actual calculation results, which can adapt to different project sizes and complexities. This method dynamically adjusts the threshold by standard deviation, avoiding the subjectivity of manual setting.
[0109] After comparing the two similarity algorithms, the more intersection projects in the two groups of project results, the better the effect of the two algorithms. Due to the difference between the detected project and the sample project, some projects in these two groups of project results may not be the target projects we want to find. A threshold can be set according to the detection time performance requirements to determine the range of projects selected for subsequent file-level or code snippet-level traceability analysis with finer granularity. For example, take the union of the top 5 or top 10 in the intersection, and the proportion of these selected projects will affect the time consumption of subsequent refinement analysis. In addition, after comparing the two feature vectors, take the projects with Hamming distance less than 10 from the Simhash similarity comparison result, and take the projects with signature similarity greater than 50 from the Minhash similarity comparison result. The value can be dynamically taken according to the size of the detected project. This can find a best balance point according to the actual situation.
[0110] The detection results of the two sample projects are analyzed using the above method as follows.
[0111] In Table 2, the Simhash feature vector of the apollo_v0.4.0 project compared with 30 sample projects can be seen. The Hamming distance with the other three versions of the project (target project) is 1, 5, and 9, respectively, which is obviously smaller than the Hamming distance with the other 27 sample projects. The comparison results using Minhash. The signature similarity of the three target projects is 94, 77, and 55, respectively, which is obviously larger than the other 27 sample projects. See Table 3 below, which extracts the results in Table 1.
[0112] Table 3: Comparison results of target projects of project 1
[0113]
[0114] Step six: project fine similarity analysis to find the most similar project
[0115] The selected group of preliminary projects is compared in more detail at the file level and code snippet level to finally find a most similar project. The specific steps are as follows:
[0116] (6-1) Extract the fingerprint feature vector of all valid file contents in the selected group of projects. Extract the same fingerprint feature vector for all file contents in the detected project;
[0117] (6-2) Take each project in the selected group of projects as a unit, compare each file feature vector in the detected project with the feature vector of each file in the preliminary project. If they are equal, count the number of valid lines of the file. These two files will not participate in subsequent comparison. If they are not equal, proceed to the following step;
[0118] (6-3) Take the code feature vector as a window of 10 lines, and use the 10 lines in the detected project file to compare with the 10 lines in the preliminary project file. Count the number of code lines corresponding to the matching feature vectors. The number of matching lines in the current file is added to the total number of lines matched in the entire project;
[0119] (6-4) Count the code snippet similarity of each project in the group of preliminary projects and the detected project. The method is the total number of matching lines / the number of valid code lines in the detected project;
[0120] (6-5) The project with the largest similarity in the above group of preliminary projects is the most similar project finally found.
[0121] The method of the present application uses project-level similarity comparison as a solution to improve the analysis performance in software source code traceability analysis, significantly reducing the project range and quantity of direct file-level or code fragment-level comparison in traditional analysis methods. Through the method of the present application, a group of projects with high similarity to the detected project can be quickly and efficiently screened, reducing the scope of subsequent refined traceability analysis and improving the overall efficiency of software source code traceability analysis.
[0122] The application of project similarity hashing in preliminary screening not only improves the efficiency of project screening and code management, but also enhances the enterprise's ability in intellectual property protection, compliance management, supply chain security, version control and project management. These effects bring significant results, helping enterprises save time and cost while improving project development quality and market competitiveness, creating greater value for enterprises in technological innovation and market expansion.
[0123] The contents not described in detail in the specification of the present application belong to the prior art known to those skilled in the art.
Claims
1. A software similar item screening method based on a similarity hashing technique, characterized by The application relates to a method for detecting a target project from a detected project, comprising the following steps: selecting a plurality of sample projects to construct a knowledge base; preparing the detected project; reading the sample projects in the knowledge base, splicing the file names of each sample project to generate a sample project text string; using a Simhash hash algorithm and a Minhash hash algorithm to extract a Simhash feature vector and a Minhash feature vector of each sample project from the sample project text string; splicing the file names of the detected project to generate a detected project text string, using a Simhash hash algorithm and a Minhash hash algorithm to extract a Simhash feature vector and a Minhash feature vector of the detected project from the detected project text string; comparing the Simhash feature vector of the detected project with the Simhash feature vectors of the sample projects to obtain corresponding Simhash Hamming distance values; comparing the Minhash feature vector of the detected project with the Minhash feature vectors of the sample projects to obtain corresponding Minhash signature similarities; preliminarily screening a plurality of candidate projects according to the Simhash Hamming distance values and the Minhash signature similarities; performing file-level or code snippet-level traceability analysis on each candidate project to finally screen a target project. 2.The software similar item screening method based on the similarity hashing technology according to claim 1, characterized in that: When the plurality of sample projects are selected to construct the knowledge base: the number of the sample projects in the knowledge base is greater than or equal to 30, the sizes of the sample projects are different, and the sample projects have different versions similar to the detected project. 3.The software similar item screening method based on the similarity hashing technology according to claim 1, characterized in that: Before the file names are spliced into the detected project text string, the files are sorted in alphabetical order, and the number of effective files participating in the analysis of each project and the size information of the effective files are recorded. 4.The software similar item screening method based on the similarity hashing technology according to claim 1, characterized in that: When the plurality of candidate projects are preliminarily screened according to the Simhash Hamming distance values and the Minhash signature similarities: a Simhash Hamming distance threshold value is determined to screen the sample projects with a Simhash Hamming distance less than or equal to the threshold value, and a Minhash signature similarity threshold value is determined to screen the sample projects with a Minhash signature similarity greater than or equal to the Minhash signature similarity threshold value; the intersection of the sample projects screened by the two methods is taken as the candidate projects.
5. The software similar item screening method based on the similarity hashing technology according to claim 1, characterized in that: When the plurality of candidate projects are preliminarily screened according to the Simhash Hamming distance values and the Minhash signature similarities: the Simhash Hamming distance and the Minhash signature similarity of each sample project are weighted and fused to obtain a comprehensive similarity score, and the sample projects are sorted according to the comprehensive similarity score to screen the candidate projects. 6.The software similar item screening method based on the similarity hashing technology according to claim 4, characterized in that: When the Simhash Hamming distance threshold value is determined: the Simhash Hamming distance threshold value is dynamically determined by a quartile range method or introduced by a Gaussian mixture model according to the local sensitive hash theory; the Gaussian mixture model is introduced to determine the Hamming distance threshold value, and the method comprises the following steps: Increase the visual analysis of the distribution of Simhash feature vectors to determine the distribution characteristics of Simhash feature vectors; if the Simhash feature vectors present mixed distribution characteristics, introduce a Gaussian mixture model to separate the sub-distribution of Simhash feature vectors, and set different threshold calculation methods for different sub-distributions.
7. The software similar item screening method based on the similarity hashing technology according to claim 4, characterized in that: When determining the Minhash signature similarity threshold: Calculate the Jaccard similarity of the reaction file name overlap, and determine the Minhash signature similarity threshold through the normal distribution of the Jaccard similarity. 8.The software similar item screening method based on the similarity hashing technology according to claim 1, characterized in that: When calculating the Minhash signature similarity, the signature length and the number of permutations can be dynamically adjusted, specifically: The optimal number of permutations k is: k = ln(1-t) / ln(2), where t is the expected Jaccard similarity threshold; The signature length is 128 bits by default, and can be extended to 256 bits or 512 bits when the similarity precision needs to be improved. 9.The software similar item screening method based on the similarity hashing technology according to claim 1, characterized in that: Perform more detailed file-level and code snippet-level comparisons on candidate projects to finally select the most similar target project, specifically: Extract the fingerprint feature vector from all valid file contents in each candidate project, and extract the same fingerprint feature vector from all file contents in the detected project. Compare the fingerprint feature vector of each file in the detected project with the fingerprint feature vector of each file in the candidate project on a per-project basis: If they are equal, count the number of valid lines in the file, and the two equal files will not participate in subsequent comparison; if they are not equal, proceed to the following step: Select the window size and step length, compare the feature vectors of the code in each window of the file in the detected project with the code in the window of the file in the candidate project, and accumulate the number of code lines corresponding to the matching feature vectors of the file; then count the total number of matching lines of the entire candidate project as the project similarity; The project corresponding to the maximum project similarity is the most similar project.