Method and system for analyzing similarity of software codes

By combining multi-dimensional analysis and adaptive backtracking analysis, the multi-dimensional features of software code are extracted and analyzed, and the problems of low efficiency and accuracy of code similarity analysis in the prior art are solved, and more accurate and efficient code similarity analysis is achieved.

CN119988185APending Publication Date: 2025-05-13BEIJING SANHE ANXIN TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510131243.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has low efficiency and precision when analyzing software code similarity, and cannot effectively identify similarities between code snippets, especially when dealing with complex code structures and logic.

Method used

By combining multi-dimensional analysis and adaptive backtracking analysis, the multi-dimensional features of code fragments are extracted, including syntax structure, code style, functional characteristics, quality characteristics, code performance characteristics, code language and platform characteristics, and the context backtracking analysis is used to perform context backtracking analysis, integrating static and dynamic analysis to establish similarity analysis results.

Benefits of technology

It improves the accuracy and efficiency of code similarity analysis, can more accurately identify the similarity between code fragments, and solves problems such as plagiarism, copyright and security vulnerabilities caused by code reuse.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988185A_ABST
    Figure CN119988185A_ABST
Patent Text Reader

Abstract

The invention discloses a method for analyzing software code similarity and an analysis system, and relates to the field of software engineering and machine learning. The method comprises the steps of establishing a code snippet database, extracting multi-dimensional features of code snippets, executing similarity analysis of the code snippets, and establishing a similarity analysis result; performing context backtracking analysis on the code snippets in the database by using an adaptive analysis network, and establishing a context similarity result; performing static and dynamic fusion analysis on the database, establishing a joint function similarity, and establishing a function similarity result; code version change analysis is executed, and a version similarity result is established; and performing code similarity retrieval management according to the four analysis results. The technical problem that the code similarity analysis efficiency and precision are low due to the fact that the code snippets are high in complexity in multi-dimensional feature analysis is solved, and the code similarity analysis accuracy and efficiency are improved in the mode that multi-dimensional analysis and self-adaptive backtracking analysis are combined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of software engineering and machine learning, and is specifically a method and an analysis system for analyzing the similarity of software codes. In the current information society, software development has become an indispensable part of all walks of life. However, with the increasing size and complexity of software, code reuse is widely used as a means to improve development efficiency. However, code reuse may also bring about a series of problems such as plagiarism, copyright disputes, and security vulnerabilities. Therefore, it is particularly important to effectively and accurately analyze the similarity of software codes. Background Art

[0002] Code reuse is a common phenomenon in the software development process. In order to improve development efficiency, developers often use similar or identical code snippets in different projects. However, code reuse may also bring some problems, such as plagiarism, copyright issues, security vulnerabilities, etc. Therefore, it is particularly important to analyze the similarity of software codes. Existing code similarity analysis methods mainly rely on simple text matching or syntax analysis, which are ineffective when dealing with complex code structures and logics, and cannot accurately identify the similarities between code snippets. Summary of the invention

[0003] The present application provides a method and an analysis system for analyzing software code similarity, thereby solving the technical problem of low efficiency and precision of code similarity analysis due to the high complexity of code snippets in multi-dimensional feature analysis, and improving the accuracy and efficiency of code similarity analysis by combining multi-dimensional analysis with adaptive backtracking analysis.

[0004] The present application provides a method for analyzing software code similarity, which is applied to a system for analyzing software code similarity, including: establishing a code snippet database, extracting multidimensional features of the code snippets, performing similarity analysis of the code snippets, and establishing similarity analysis results, wherein the multidimensional features include grammatical structure features, code style features, functional features, quality features, code performance features, code language and platform features; using an adaptive analysis network to perform context backtracking analysis on the code snippets in the code snippet database, and establishing context similarity results based on the context backtracking analysis results; performing static and dynamic fusion analysis on the code snippet database, establishing joint functional similarity, and establishing functional similarity results based on the joint functional similarity; performing code version change analysis in the code snippet database, and establishing version similarity results; and performing code similarity retrieval management based on the similarity analysis results, the context similarity results, the functional similarity results and the version similarity results.

[0005] The present application also provides a system for analyzing software code similarity, including: a code snippet analysis unit: establishing a code snippet database, extracting multidimensional features of code snippets, performing similarity analysis of code snippets, and establishing similarity analysis results, wherein the multidimensional features include grammatical structure features, code style features, functional features, quality features, code performance features, code language and platform features; a context backtracking analysis unit: using an adaptive analysis network to perform context backtracking analysis on code snippets in the code snippet database, and establishing a context similarity result fusion analysis based on the context backtracking analysis results.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages: The present application provides a method and system for analyzing software code similarity, which relate to the field of machine learning technology, solve the technical problem of low code similarity analysis efficiency and analysis accuracy due to the high complexity of analyzing software code similarity, and achieve the technical effect of improving the accuracy of code similarity analysis by combining multi-dimensional feature extraction with adaptive backtracking analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In order to more clearly illustrate the technical solution of the embodiment of the present application, the following drawings required for describing the embodiment are described. Figure 1 The following instructions are given: This application uses Figure 1 The flowchart is used to illustrate the operations performed by the system according to the embodiment of the present application. It should be understood that the previous or following operations are not necessarily performed precisely in order. On the contrary, various steps can be processed in reverse order or simultaneously as needed. At the same time, other operations can also be added to these processes, or one or more operations can be removed from these processes. Figure 1 A flowchart of a method for analyzing software code similarity provided in an embodiment of the present application. DETAILED DESCRIPTION

[0008] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.

[0009] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.

[0010] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict, and the terms "first\second" involved are merely to distinguish similar objects and do not represent a specific ordering of objects. The terms "including" and "having" and any variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by technicians in the technical field of this application. The terms used herein are for the purpose of describing the embodiments of the present application only.

[0011] The embodiment of the present application provides a method for analyzing software code similarity, and the method is applied to a system for analyzing software code similarity. As shown in the figure, the method includes: Multidimensional characteristics of code: The multi-dimensional features of the code cover many important aspects, including grammatical structure features, code style features, functional features, quality features, code performance features, code language and platform features. These features comprehensively describe the characteristics of the code from different angles and provide rich information for subsequent accurate analysis of code similarity.

[0012] The grammatical structure features mainly involve the grammatical rules and structural composition of the code. For example, in the Python language, the grammatical structure of function definition is "def function name (parameter list): function body". The structure of function definition in different code snippets, the inheritance relationship of classes, the nesting level of statements, etc. all belong to the category of grammatical structure features. By parsing the syntax tree of the code, these features can be accurately extracted. The syntax tree is a tree-like data structure that represents the grammatical structure of the code. The nodes represent the syntax units, and the edges represent the relationship between the syntax units. Taking a simple Python function "def add (a, b): return a + b" as an example, the root node of its syntax tree may be the function definition node, and the child nodes include the function name "add", the parameter list node (including the parameters "a" and "b"), and the function body node (including the return statement node and the addition operation node).

[0013] Code style features reflect the programming habits and preferences of developers. This includes the code indentation method (whether to use spaces or tabs, and the length of the indentation), naming conventions (naming styles of variable names, function names, and class names, such as camel case naming, underscore naming, etc.), and the style and number of comments. For example, some developers are accustomed to adding detailed comments before the function definition to explain the function, parameter meaning, and return value, while some developers have fewer comments. Although code style features do not directly affect the function of the code, they reflect the readability and maintainability of the code to a certain extent, and are also one of the important bases for judging code similarity.

[0014] Functional features focus on the specific functions implemented by the code. For a piece of code, it may implement various functions such as data sorting, file reading and writing, image recognition, etc. Extracting functional features requires in-depth analysis of the logic of the code. Its functional features can be determined by analyzing the algorithms used in the code, the data processing flow, and the interaction with external systems. For example, a piece of code that uses the quick sort algorithm to sort an array, its functional features include the application of the quick sort algorithm and the sorting operation on the array data.

[0015] Quality characteristics are used to evaluate the quality level of the code. This includes aspects such as code complexity (such as cyclomatic complexity, which measures the number of independent paths in the code. The higher the cyclomatic complexity, the more complex the code and the more difficult it is to maintain), code maintainability (such as the modularity of the code, whether it follows the design pattern, etc.), and code testability (whether the code is easy to write test cases for testing). High-quality code usually has low complexity, good modular design, and high testability. Through some code quality analysis tools and indicators, these quality characteristics can be quantified to provide a reference for code similarity analysis.

[0016] Code performance characteristics reflect the performance of the code at runtime. This includes aspects such as code execution time, memory usage, and resource utilization. For example, if a piece of code takes too long to execute or takes too much memory when processing a large amount of data, it means that there may be performance issues. In code similarity analysis, considering code performance characteristics can help determine the performance differences of similar codes in actual operation, which is of great significance for selecting a better code reuse solution.

[0017] The code language and platform characteristics specify the programming language used by the code and the platform environment on which it runs. Different programming languages ​​have different syntax rules, data types, and programming paradigms. For example, Python is a dynamically typed, interpreted language, while Java is a statically typed, compiled language. The platform environment on which the code runs will also affect its behavior. For example, the code running on Windows and Linux systems may have differences in file path representation, system calls, etc. These language and platform-related characteristics are factors that cannot be ignored in code similarity analysis, because even if two pieces of code have similar functions, their similarity will be greatly reduced if the languages ​​used and the platforms they run on are different.

[0018] Adaptive Analysis Network: The adaptive analysis network is an intelligent analysis model based on machine learning. It can automatically adjust the analysis strategy according to the characteristics of the input data to better adapt to complex and changing situations. In code similarity analysis, the adaptive analysis network is mainly used to perform contextual backtracking analysis on code snippets.

[0019] Contextual backtracking analysis aims to gain a deeper understanding of the meaning and function of code by analyzing the usage context and historical information of code snippets, so as to more accurately judge the similarity between codes. In the actual software development process, the usage scenario and context information of the code are crucial to understanding the function and intention of the code. For example, a function may be used for different purposes in different modules, and its true function can only be accurately grasped by combining its usage context.

[0020] Obtaining the call records of code snippets in the code snippet database is the first step in context backtracking analysis. The call records include two important pieces of information: usage context and usage feedback. The usage context covers the environment information when the code snippet is called, such as the function that calls it, the module it is in, and the parameters passed when calling it. The usage feedback records the performance and results of the code snippet during use, such as whether it is executed successfully, whether an exception occurs, and the return value.

[0021] Clustering similar usages based on the usage context in the call record is a key step for further analysis. Based on the usage context, a usage scenario is established. The usage scenario is an abstract description of the environment and conditions of a code snippet in a specific usage situation. For example, for a file reading function, its usage scenario may include the type of file read (text file, binary file, etc.), the storage location of the file, the purpose of reading (whether it is for data processing, display, or other purposes), etc. By performing similarity calculations on the usage scenarios of different code snippets, a first similarity result can be established. Similarity calculations can be performed in a variety of ways, such as cosine similarity calculations based on a vector space model, which represents each feature in the usage scenario as a dimension of a vector and measures the similarity of the usage scenarios by calculating the cosine similarity between vectors.

[0022] At the same time, user backtracking is performed based on the usage context to establish user usage identification. User usage identification records the usage habits and preferences of different users for code snippets. For example, some users may prefer to use a certain code snippet in a specific module, or always pass a specific type of parameter when calling a code snippet. A second similar result is established based on the user usage identification. The first similar result and the second similar result are combined to perform similar usage clustering and establish a similar usage clustering result. In this way, code snippets with similar usage patterns can be clustered together, which is convenient for subsequent more in-depth analysis.

[0023] It is an important part of contextual backtracking analysis to mark the usage context with time and use the time mark to construct the backtracking trust coefficient. The time mark records the time sequence of each use of the code snippet. By analyzing the time series, we can understand the usage frequency and usage time interval of the code snippet. The backtracking trust coefficient is determined according to the usage of the code snippet at different time points and the usage feedback. If a code snippet performs well in multiple uses and has a high usage frequency, then its backtracking trust coefficient will be high; conversely, if there are frequent exceptions or the usage frequency is very low, the backtracking trust coefficient will be low. The backtracking trust coefficient and usage feedback are used to perform contextual backtracking trust marking of similar usage clustering results to construct contextual similarity results. The contextual similarity results can more accurately reflect the similarity of code snippets in actual usage scenarios, providing a more practical reference for code similarity analysis.

[0024] Static analysis and dynamic analysis: Static analysis and dynamic analysis are two important methods of code analysis, each with its own unique advantages and limitations. Static analysis is to analyze the text of the code without executing the code to obtain information such as the structure, syntax, and semantics of the code. Static analysis can find problems such as syntax errors, potential logical errors, and unused variables in the code. For example, static analysis tools can be used to check whether there are undefined variable references, function call parameter mismatches, and other problems in the code.

[0025] Dynamic analysis is to obtain information by monitoring the execution process and behavior of the code during its execution. Dynamic analysis can observe the running status of the code in real time, including the value of variables, the order of function calls, memory usage, etc. For example, dynamic analysis tools can monitor the memory usage of the code when processing large data sets, as well as the execution time of the function under different input conditions.

[0026] In code similarity analysis, the fusion of static analysis and dynamic analysis can give full play to the advantages of both and analyze the functional similarity of the code more comprehensively and accurately. When performing static and dynamic fusion analysis on the code snippet database, the code snippet database is first input into the joint analysis channel. The joint analysis channel includes a static analysis sub-channel and a dynamic analysis sub-channel.

[0027] The static analysis subchannel of the joint analysis channel is called to parse the code, extract function, class, and module information, and generate the first functional features. In this process, the static analysis subchannel will parse the grammatical structure of the code, identify function definitions, class declarations, and module organization in the code, etc. For example, for a piece of Python code, the static analysis subchannel can extract all functions defined therein and their parameter lists, class attributes and methods, etc. This information constitutes the static functional features of the code, reflecting the function of the code at the structural level.

[0028] The dynamic analysis subchannel of the joint analysis channel is called, and the dynamic analysis subchannel is used to perform dynamic execution path analysis of the code snippet database to establish the second functional feature. Dynamic execution path analysis is one of the key contents of dynamic analysis. It tracks the execution flow of the code when the code is running, and records the various statements and function call paths passed during the code execution process. For example, for a program with multiple branches and loops, dynamic execution path analysis can record in detail the execution path of the program under different input conditions. These dynamic execution path information constitute the dynamic functional characteristics of the code, reflecting the behavior and function of the code when it is actually running.

[0029] After the first functional feature and the second functional feature are fused, a joint functional similarity is established. Feature fusion can be performed in a variety of ways, such as weighted fusion, neural network-based fusion, etc. Weighted fusion assigns different weights to the first functional feature and the second functional feature according to the importance of different features, and then combines the weighted features. Neural network-based fusion uses the powerful learning ability of neural networks to automatically learn how to fuse two features to obtain more accurate results. By establishing a joint functional similarity, the degree of functional similarity between codes can be more comprehensively measured, thereby establishing functional similarity results and providing a more reliable basis for code similarity analysis. Code version change analysis During the software development process, the code will be constantly updated and iterated, and version changes will occur frequently. Code version change analysis conducts in-depth research on the version change history of code snippets in the code snippet database, and mines the change information between different versions of the code, thus providing an important reference for code similarity analysis.

[0030] Obtaining the submission history data of code snippets in the code snippet database is the basis for code version change analysis. Submission history data records detailed information such as the time, submitter, and content of each code submission. A submission timing chain is established based on the submission history data. The submission timing chain connects the various submission versions of the code in chronological order to show the evolution of the code. For example, in a project using the Git version control system, the submission timing chain can be easily obtained through Git's submission log. Each submission corresponds to a node on the submission timing chain, and the order between the nodes reflects the order of submission.

[0031] It is also important to obtain the merge history of code snippets in the code snippet database. The merge history records the merge of code between different branches, including the time of the merge, the source of the merged branch, and other information. Based on the merge history, the merge identifier of the commit timing chain is used to establish an update timing chain. The update timing chain not only includes the submission order of the code, but also takes into account the changes in the code during the merge process, which more accurately reflects the actual evolution path of the code. For example, when a functional branch is merged with the main branch, the update timing chain will record the relevant information of the merge and the changes in the code after the merge.

[0032] It is also necessary to obtain the developer tags of the code snippets in the code snippet database. Developer tags are some identification information added by developers in the code to indicate the special purpose of the code, the module it belongs to, or the development stage, etc. Additional tags are built based on developer tags. Additional tags can provide more semantic information for code version change analysis. For example, developers may add the tag "# experimental" to the code to indicate that the code snippet is experimental, or add the tag "# for performance optimization" to indicate that the code snippet is used for performance optimization.

[0033] The version similarity results are established using the commit timing chain, update timing chain, and additional tags. By analyzing this information, the similarity between different versions of the code can be determined. If the two versions of the code are in similar positions on the commit timing chain and update timing chain and have similar additional tags, then the version similarity between them is high. The version similarity results can help developers understand the inheritance relationship and changes between different versions of the code, which is of great significance for code similarity analysis and code maintenance and management.

[0034] Finally, a comprehensive similarity analysis is conducted based on the analysis of each module.

[0035] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules may be used and run on the user terminal and / or server, and the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present application.

[0036] The above specific implementations do not constitute a limitation on the protection scope of this application. It should be understood by those skilled in the art that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this application should be included in the protection scope of this application.

Claims

1. A method for analyzing software code similarity, characterized in that: The method comprises: Establishing a code snippet database, extracting multidimensional features of the code snippets, performing similarity analysis on the code snippets, and establishing similarity analysis results, wherein the multidimensional features include grammatical structure features, code style features, functional features, quality features, code performance features, code language and platform features; Using an adaptive analysis network to perform context backtracking analysis on the code snippets in the code snippet database, and establishing context similarity results based on the context backtracking analysis results; Performing static and dynamic fusion analysis on the code snippet database to establish a joint functional similarity, and establishing a functional similarity result according to the joint functional similarity; Perform code version change analysis in the code snippet database and establish version similarity results; Code similarity retrieval management is performed based on the similarity analysis result, the context similarity result, the function similarity result and the version similarity result.

2. A method for analyzing software code similarity according to claim 1, characterized in that: The method of using the adaptive analysis network to perform context backtracking analysis on the code snippets in the code snippet database includes: Acquire a call record of a code snippet in the code snippet database, wherein the call record includes a usage context and a usage feedback; Performing similar usage clustering according to the usage context in the call record, and establishing similar usage clustering results; The usage context is time-marked, and a retrospective trust coefficient is constructed using the time mark. The context retrospective trust identification of the similar usage clustering result is performed using the retrospective trust coefficient and the usage feedback, and the context similarity result is constructed using the context retrospective trust identification.

3. A method for analyzing software code similarity as claimed in claim 2, characterized in that: The performing similar usage clustering according to the usage context in the call record and establishing similar usage clustering results includes: Establishing a usage scenario based on the usage context, performing scenario similarity calculation using the usage scenario, and establishing a first similarity result; Perform user backtracking based on the usage context, establish a user usage identifier, and establish a second similar result according to the user usage identifier; Similar usage clustering is performed according to the first similarity result and the second similarity result to establish a similar usage clustering result.

4. A method for analyzing software code similarity according to claim 1, characterized in that: The performing static and dynamic fusion analysis on the code snippet database to establish a joint functional similarity, and establishing a functional similarity result according to the joint functional similarity, further includes: Input the code snippet database into the joint analysis channel, call the static analysis sub-channel of the joint analysis channel to perform code parsing, extract function, class, and module information, and generate a first functional feature; Calling the dynamic analysis sub-channel of the joint analysis channel, using the dynamic analysis sub-channel to perform dynamic execution path analysis of the code snippet database, and establishing a second functional feature; After the first functional feature and the second functional feature are fused, a joint functional similarity is established.

5. A method for analyzing software code similarity according to claim 1, characterized in that: The code version change analysis in the execution code snippet database and establishing version similarity results include: Acquire submission history data of code snippets in the code snippet database, and establish a submission timing chain according to the submission history data; Acquire the merge history of the code snippets in the code snippet database, perform a merge identification of the submission timing chain based on the merge history, and establish an update timing chain; Obtaining developer tags of code snippets in the code snippet database, and constructing additional tags based on the developer tags; A version similarity result is established using the submission timing chain, the update timing chain, and the additional mark.

6. A method for analyzing software code similarity according to claim 1, characterized in that: The method further comprises: When executing a user's similarity search, obtaining the user's search data, and setting the independent search terms in the search data as independent search data; Sending a trust confirmation of independent retrieval data to the user, and receiving the trust confirmation result from the user; The trust confirmation result is used to set the fuzzy range of the independent search data, and the similarity search is completed based on the fuzzy range setting result and the mapped independent search data.

7. A method for analyzing software code similarity according to claim 6, characterized in that: The method further comprises: Acquire similarity search results, and establish a search enhancement factor and a search weakening factor based on the similarity search results, wherein the search weakening factor is a coefficient for weakening the user's search feature; The retrieval enhancement factor and the retrieval weakening factor are integrated into a retrieval option, and user similarity retrieval management is performed based on the retrieval option.

8. A system for analyzing software code similarity, characterized in that: The system is used to implement a method for analyzing software code similarity according to any one of claims 1 to 7, and the system comprises: Code snippet analysis unit: establish a code snippet database, extract multi-dimensional features of code snippets, perform similarity analysis of code snippets, and establish similarity analysis results, wherein the multi-dimensional features include grammatical structure features, code style features, functional features, quality features, code performance features, code language and platform features; Context backtracking analysis unit: uses an adaptive analysis network to perform context backtracking analysis on code snippets in the code snippet database, and establishes context similarity results based on the context backtracking analysis results; Fusion analysis unit: performing static and dynamic fusion analysis on the code snippet database, establishing joint functional similarity, and establishing functional similarity results according to the joint functional similarity; Code version change analysis unit: performs code version change analysis in the code snippet database and establishes version similarity results; Similarity retrieval management unit: performs code similarity retrieval management according to the similarity analysis result, the context similarity result, the function similarity result and the version similarity result.

Citation Information

Cited By

  • Automatic evidence obtaining method and system for software copyright infringement

    CN120372581A

  • An automated evidence collection method and system for software copyright infringement

    CN120372581B

  • Intelligent comparison and analysis method and system for similarity of examination answer codes

    CN120631735A