A security diagnosis and evaluation method and system for open-source software

By generating code attribute graphs and building code relationship trees, combining context slicing and feature fusion technology, the problem of insufficient dependency coverage in open source software security evaluation is solved, and more accurate vulnerability identification and risk reduction is achieved.

CN119918067BActive Publication Date: 2025-07-08CHINA ELECTRONICS RELIABILITY AND ENVIRONMENTAL TESTING INSTITUTE ((THE FIFTH INSTITUTE OF ELECTRONICS MINISTRY OF INDUSTRY AND INFORMATION TECHNOLOGY) (CHINA SAIBAO LABORATORY)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510422316.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-08
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing open source software security assessment methods are difficult to fully cover deep dependencies and cross-dependencies, resulting in the security vulnerabilities of certain dependencies being ignored, increasing the risk of supply chain attacks.

Method used

The Joern tool generates a code attribute graph (CAG), extracts the program dependency graph and builds a code relationship tree. Combining context slices and feature extraction techniques, the anonymous code sequence is processed using Word2Vec and a bidirectional gating recurrent unit model, and fuses the multi-layer perceptron model for security evaluation.

Benefits of technology

In-depth analysis of the structure of open source software, comprehensively capture code dependencies, improve the accuracy and speed of vulnerability identification, and reduce the risk of supply chain attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918067B_ABST
    Figure CN119918067B_ABST
Patent Text Reader

Abstract

The present invention discloses a security diagnosis and evaluation method and system for open-source software, which relates to the technical field of security diagnosis; after extracting dependencies from the code property graph, a code relationship tree is constructed; by performing context slicing to extract the context code related to vulnerabilities in the source code to obtain a target code segment, and feature extraction is performed on the target code segment to obtain a first feature; after anonymizing the variable names and function names in the source code, word embedding is performed on the anonymized code sequence through Word2Vec, and a bidirectional gated recurrent unit model is used to extract the syntactic information of the target code segment to obtain a second feature; substituting the first feature and the second feature into a multi-layer perceptron model to obtain the probability of the existence of vulnerabilities in the source code. By comprehensively capturing the dependencies between codes and then extracting the target code segment through context slicing, potential vulnerability locations and types can be more accurately identified, improving security and reducing the risk of being attacked by the supply chain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of security diagnosis, and particularly relates to a security diagnosis and evaluation method and system for open-source software. Background Art

[0002] With the rapid development of modern software development, open-source software has become an important part for enterprises and developers to build applications. Open-source software has the advantages of being free, open, and customizable, and is widely used in various projects. However, since the source code of open-source software is public, the security risks also increase. Attackers can analyze the public source code, discover vulnerabilities and exploit them. In addition, the development and maintenance status of open-source software projects vary, and some projects may lack continuous maintenance or fail to repair known vulnerabilities in a timely manner, resulting in security hazards during use.

[0003] In the software supply chain, the complexity of open-source software and its dependencies further amplifies security risks. Modern software often consists of multiple open-source components, and vulnerabilities in any one component may affect the security of the entire system. For example, if there is a serious vulnerability in an open-source library and the library is widely used, its potential threats will spread rapidly, possibly triggering large-scale security incidents. In addition, many industries require that the software systems used must comply with strict security and privacy standards. This means that when enterprises adopt open-source software, they must conduct a detailed security assessment on it to ensure compliance requirements and avoid security incidents such as data leakage.

[0004] Existing methods often have difficulty in comprehensively covering deep dependencies and cross-dependencies, which may lead to the omission of security vulnerabilities in some dependencies and increase the risk of supply chain attacks. Summary of the Invention

[0005] The purpose of the present invention is to solve the problem that security vulnerabilities in some dependencies are overlooked, increasing the risk of supply chain attacks, and to propose a security diagnosis and evaluation method and system for open-source software.

[0006] In the first aspect of the implementation of the present invention, a security diagnosis and evaluation method for open-source software is first proposed. The method includes:

[0007] Obtain the source code of the open-source software, parse the source code through the Joern tool to obtain a code property graph, extract dependencies from the code property graph to obtain a program dependency graph, and construct a code relationship tree according to the program dependency graph;

[0008] According to the code relationship tree, extract the context code related to vulnerabilities in the source code through context slicing to obtain a target code segment, and extract the first feature from the target code segment;

[0009] After anonymizing the variable names and function names in the source code, perform word embedding on the anonymized code sequence using Word2Vec, and use a bidirectional gated recurrent unit model to extract the syntactic information of the target code snippet to obtain a second feature;

[0010] Fuse the first feature and the second feature to obtain a target feature, substitute the target feature into a multi-layer perceptron model to obtain the vulnerability existence probability of the source code, and perform a security assessment on the open-source software according to the vulnerability existence probability.

[0011] Optionally, perform dependency extraction on the code property graph to obtain a program dependency graph, and construct a code relationship tree according to the program dependency graph, including:

[0012] Find the usage relationship between each variable by tracing the data flow graph in the code property graph to obtain data dependencies, obtain control dependencies by tracing the control flow graph in the code property graph, and obtain structural dependencies by analyzing the scope and hierarchical structure in the source code;

[0013] Integrate the data dependencies, the control dependencies, and the structural dependencies to obtain a program dependency graph;

[0014] Obtain the main function in the source code as the root node, and according to the program dependency graph, identify each code block in the source code as a child node and add it under the root node, and recursively add corresponding child nodes to each child node according to the program dependency graph until all code blocks are identified to obtain a code relationship tree.

[0015] Optionally, according to the code relationship tree, extract the context code related to the vulnerability in the source code through context slicing to obtain a target code snippet, including:

[0016] Determine the vulnerability in the source code through vulnerability detection rules, and mark the vulnerability in the code relationship tree to obtain the starting location;

[0017] According to the code relationship tree, starting from the starting location, extract the nodes related to the vulnerability forward and backward along the code relationship tree, and obtain the code snippets corresponding to all nodes related to the vulnerability and perform forward and backward slicing along the control flow and data flow to obtain the target code snippet.

[0018] Optionally, perform feature extraction on the target code snippet to obtain a first feature, including:

[0019] Obtain the local abstract syntax tree and the local control flow graph corresponding to the target code snippet, and extract the node information in the local abstract syntax tree to obtain the first information;

[0020] Extract the control flow paths related to faults in the local control flow graph to obtain the second information; extract the usage frequency, assignment chain, and usage chain of critical variables in the code to obtain the third information;

[0021] Substitute the first information, the second information, and the third information into a machine learning model for feature extraction to obtain the first feature.

[0022] Optionally, feature fusion of the first feature and the second feature to obtain the target feature includes:

[0023] Perform linear transformation on the first feature and the second feature respectively to obtain the first linear feature and the second linear feature;

[0024] Divide the first linear feature into multiple subspaces, perform parallel attention operations to obtain multiple attention heads, and splice all attention heads to obtain the first target feature;

[0025] Perform weighted fusion on the first target feature and the second linear feature to obtain the target feature.

[0026] In the second aspect of the implementation of the present invention, a security diagnosis and evaluation system for open source software is proposed, including:

[0027] A code relationship tree determination module, configured to obtain the source code of open source software, parse the source code through the Joern tool to obtain a code attribute graph, perform dependency extraction on the code attribute graph to obtain a program dependency graph, and construct a code relationship tree according to the program dependency graph;

[0028] A first feature extraction module, configured to, according to the code relationship tree, extract the context code related to vulnerabilities in the source code through context slicing to obtain a target code snippet, and perform feature extraction on the target code snippet to obtain the first feature;

[0029] A second feature extraction module, configured to anonymize the variable names and function names in the source code, perform word embedding on the anonymized code sequence through Word2Vec, and use a bidirectional gated recurrent unit model to extract the syntactic information of the target code snippet to obtain the second feature;

[0030] A feature fusion module, configured to perform feature fusion on the first feature and the second feature to obtain the target feature, substitute the target feature into a multi-layer perceptron model to obtain the vulnerability existence probability of the source code, and perform security evaluation on the open source software according to the vulnerability existence probability.

[0031] Optionally, the code relationship tree determination module includes:

[0032] A dependency extraction module, which is used to find the usage relationships between each variable by tracking the data flow graph in the code property graph to obtain data dependencies, obtain control dependencies by tracking the control flow graph in the code property graph, and obtain structural dependencies by analyzing the scopes and hierarchical structures in the source code;

[0033] A dependency integration module, which is used to integrate the data dependencies, the control dependencies, and the structural dependencies to obtain a program dependency graph;

[0034] A code relationship tree building module, which is used to obtain the main function in the source code as the root node, and according to the program dependency graph, identify each code block in the source code as a child node and add it under the root node, and recursively add corresponding child nodes to each child node according to the program dependency graph until all code blocks are identified, obtaining a code relationship tree.

[0035] Optionally, the first feature extraction module includes:

[0036] A vulnerability marking module, which is used to determine the vulnerabilities in the source code through vulnerability detection rules, and mark the vulnerabilities in the code relationship tree to obtain the starting locations;

[0037] A target code snippet determination module, which is used to, according to the code relationship tree, extract the nodes related to the vulnerability in the forward and backward directions along the code relationship tree starting from the starting location, and obtain the code snippets corresponding to all the nodes related to the vulnerability and perform forward and backward slicing along the control flow and data flow to obtain the target code snippets.

[0038] Optionally, the first feature extraction module further includes:

[0039] A node information extraction module, which is used to obtain the local abstract syntax tree and the local control flow graph corresponding to the target code snippet, and extract the node information in the local abstract syntax tree to obtain the first information;

[0040] An information extraction module, which is used to extract the control flow paths related to the fault in the local control flow graph to obtain the second information; extract the usage frequency, assignment chain, and usage chain of key variables in the code to obtain the third information;

[0041] A first feature fusion module, which is used to substitute the first information, the second information, and the third information into a machine learning model for feature extraction to obtain the first feature.

[0042] Optionally, the feature fusion module includes:

[0043] A linear transformation module, which is used to perform linear transformations on the first feature and the second feature respectively to obtain the first linear feature and the second linear feature;

[0044] The first target feature construction module is used to divide the first linear feature into multiple subspaces, perform parallel attention operations to obtain multiple attention heads, and splice all the attention heads to obtain the first target feature;

[0045] The feature weighted fusion module is used to perform weighted fusion on the first target feature and the second linear feature to obtain the target feature.

[0046] Advantages of the present invention:

[0047] The present invention proposes a security diagnosis and evaluation method for open-source software. The source code of the open-source software is obtained, and the code property graph is obtained by parsing the source code through the Joern tool. The program dependency graph is obtained by dependency extraction from the code property graph, and the code relationship tree is constructed according to the program dependency graph; according to the code relationship tree, the context code related to the vulnerability in the source code is extracted by context slicing to obtain the target code segment, and the first feature is obtained by feature extraction of the target code segment; after anonymizing the variable names and function names in the source code, word embedding is performed on the anonymized code sequence through Word2Vec, and the syntactic information of the target code segment is extracted using a bidirectional gated recurrent unit model to obtain the second feature; the first feature and the second feature are feature-fused to obtain the target feature, the target feature is substituted into the multi-layer perceptron model to obtain the vulnerability existence probability of the source code, and the open-source software is security-evaluated according to the vulnerability existence probability. By using the Joern tool to generate the code property graph (CAG) and extract the program dependency graph, the solution can deeply analyze the structure of the source code, comprehensively capture the dependencies between codes, and then extract the target code segment by context slicing, reducing the amount of code to be analyzed and improving the speed of the overall solution. By fusing the obtained first feature and second feature, the solution can more comprehensively and accurately identify potential vulnerability locations and types, improving security and reducing the risk of being attacked by the supply chain. Description of the Drawings

[0048] The present invention will be further described below with reference to the drawings.

[0049] Figure 1 It is a flowchart of a security diagnosis and evaluation method for open-source software provided by an embodiment of the present invention;

[0050] Figure 2 It is a framework diagram of a security diagnosis and evaluation system for open-source software provided by an embodiment of the present invention. Detailed Embodiments

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the descriptions such as "first" and "second" in the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0052] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0053] The embodiments of the present invention provide a security diagnosis and evaluation method for open-source software. See Figure 1 , Figure 1 which is a flowchart of a security diagnosis and evaluation method for open-source software provided by the embodiments of the present invention. The method includes the following steps:

[0054] S101, obtain the source code of the open-source software, parse the source code through the Joern tool to obtain a code property graph, extract dependencies from the code property graph to obtain a program dependency graph, and construct a code relationship tree according to the program dependency graph.

[0055] S102, according to the code relationship tree, extract the context code related to the vulnerability in the source code through context slicing to obtain a target code segment, and extract the first feature from the target code segment.

[0056] S103, after anonymizing the variable names and function names in the source code, perform word embedding on the anonymized code sequence through Word2Vec, and use a bidirectional gated recurrent unit model to extract the syntactic information of the target code segment to obtain a second feature.

[0057] S104, fuse the first feature and the second feature to obtain a target feature, substitute the target feature into a multi-layer perceptron model to obtain the probability of the existence of a vulnerability in the source code, and perform a security evaluation on the open-source software according to the probability of the existence of the vulnerability.

[0058] Based on a security diagnosis and evaluation method for open-source software provided by an embodiment of the present invention, a code attribute graph (CAG) is generated through the Joern tool and a program dependency graph is extracted. The solution can deeply analyze the structure of the source code, comprehensively capture the dependency relationships between codes, and then extract target code fragments through context slicing, reducing the amount of code to be analyzed and improving the speed of the overall solution. By fusing the obtained first feature and second feature, the solution can more comprehensively and accurately identify potential vulnerability locations and types, improving security and reducing the risk of being attacked by the supply chain.

[0059] In one implementation, through the code attribute graph and program dependency graph generated by the Joern tool, the solution can deeply understand the structure of the source code and the dependency relationships between each component. The comprehensive analysis helps to identify potential security problems and ensures comprehensive coverage of complex dependency relationships; context slicing is used to extract context code related to vulnerabilities, reducing the amount of code that needs to be analyzed, thereby improving processing efficiency, helping to accurately locate code fragments that may have vulnerabilities, and avoiding a full analysis of the entire code library.

[0060] In one implementation, by extracting structural features (i.e., the first feature) and syntactic features (i.e., the second feature) from the target code fragment, the solution combines the context information and semantic information of the code. The structural features help to identify the organization and dependency relationships of the code, while the syntactic features provide specific implementation details of the code, improving the accuracy of vulnerability detection.

[0061] In one implementation, by substituting the fused target features into a multi-layer perceptron model, the solution can automatically predict the probability of the existence of vulnerabilities. The application of the deep learning model improves the accuracy of vulnerability detection and can adapt to different types of codes and vulnerability scenarios.

[0062] In one implementation, Joern is a tool for static code analysis that can extract an abstract representation of the code from the source code. Joern helps to analyze the structure and logic in the code by constructing data structures such as an abstract syntax tree, a control flow graph, and a data flow graph. Among them, the abstract syntax tree is a tree-like structure representation of the code, showing the syntax structure of the code; the control flow graph describes the paths of the control flow in the code, such as conditional branches and loops; the data flow graph represents the data flow relationships between variables; the code attribute graph combines these graphs to form a comprehensive representation that captures information such as functions, variables, control flow, and data flow in the code.

[0063] In one implementation, in the dependency graph, nodes represent entities in the code such as functions, variables, and data, and edges represent the dependency relationships between these entities; according to the program dependency graph, each part of the code is organized into a tree structure. The code relationship tree can help understand the hierarchical structure of the code and its dependencies. Through the tree structure, the dependency levels between different code parts can be clearly viewed, which helps identify complex dependency relationships and potential security risks.

[0064] In one implementation, the purpose of anonymization is to remove specific identifiers in the source code such as variable names and function names, thereby eliminating the interference of naming differences on the analysis model and making the model more focused on the structure and semantics of the code; the step of replacing all variable names and function names in the source code with common identifiers usually keeps the logical structure of the code unchanged, but makes the names in the code no longer specific, and the syntax structure and logical relationships of the code remain consistent so that subsequent analysis can correctly understand the function of the code.

[0065] In one implementation, Word2Vec is a technology that converts words into vector representations and can capture the semantic relationships between words. In code analysis, Word2Vec maps the anonymized names of identifiers in the code to vectors, thereby representing the semantics of the code.

[0066] In one implementation, the word vectors obtained by Word2Vec are used as inputs and passed to a bidirectional gated recurrent unit. The bidirectional gated recurrent unit considers both the forward and backward information of the input sequence, thereby being able to capture more context information. The bidirectional gated recurrent unit extracts the syntactic features of the target code snippet through its gating mechanisms such as the update gate and the reset gate. These features include the structural information, logical flow, and syntactic relationships in the code.

[0067] In one implementation, the multi-layer perceptron model includes an input layer, a hidden layer, and an output layer. The number of nodes in the input layer is equal to the dimension of the target features, and the target features are extracted from the source code; there are 5 hidden layers, and the number of neurons in each hidden layer decreases sequentially from input to output, with the number of neurons being 100, 80, 60, 40, and 20 respectively. The activation function used is ReLU, and the output layer uses the Sigmoid function. The output value is between 0 and 1. If the output value is greater than or equal to 0.472, it indicates that there is a vulnerability in the code snippet, otherwise there is no vulnerability.

[0068] In one embodiment, dependency extraction is performed on the code property graph to obtain a program dependency graph. Constructing a code relationship tree according to the program dependency graph includes:

[0069] The usage relationships between each variable are found by tracking the data flow graph in the code property graph to obtain data dependencies. The control dependencies are obtained by tracking the control flow graph in the code property graph. The structural dependencies are obtained by analyzing the scopes and hierarchical structures in the source code.

[0070] The data dependencies, control dependencies, and structural dependencies are integrated to obtain a program dependency graph.

[0071] The main function in the source code is obtained as the root node. According to the program dependency graph, each code block in the source code is identified as a child node and added under the root node. And corresponding child nodes are recursively added to each child node according to the program dependency graph until all code blocks are identified, obtaining a code relationship tree.

[0072] In one implementation, by integrating data dependencies, control dependencies, and structural dependencies, various dependency relationships of the code can be comprehensively understood, and complex interactions and potential problems in the code can be revealed. The combination of data dependencies, control dependencies, and structural dependencies can provide detailed context information to help understand the execution flow and data flow of the code, thereby better locating and analyzing problems.

[0073] In one implementation, the data dependency graph can help track data flow and find potential data leakage or tampering points. The control dependency graph can reveal the impact of conditional judgments and control flow on vulnerabilities and help analyze the trigger conditions and execution paths of vulnerabilities. The structural dependency graph provides information about the hierarchy and scope of the code to help understand the relationships between different modules and functions.

[0074] In one implementation, the data dependency graph is the read-write dependency relationship between different statements or operations in the program, reflecting how data flows in the program. The source code is parsed and an abstract syntax tree is generated. The definition and usage locations of each variable are identified according to the abstract syntax tree. By tracking the definition and usage relationships of variables, the definition location and usage location of variables are connected. If the value of a variable flows between statements, a dependency relationship needs to be added, and finally a data dependency graph is obtained.

[0075] In one implementation, the control dependency graph describes the conditional dependencies on the program execution path, that is, the execution of some statements depends on the result of conditional judgments. The control flow graph of the program is generated, where nodes represent basic blocks or statements, and edges represent control flow transfers. Control structures are identified from the control flow graph, and the postdominance tree is used to determine the control dependency relationship. The postdominance tree represents that in the program control flow, on the path from a certain node to the end of the program, a certain node always executes before another node. If there are control dependencies of the subsequent child nodes on node B behind node B.

[0076] In one implementation, the structure dependency graph describes the function call relationships in a program, the block structures such as the hierarchical relationships between functions and classes, identifies the function call relationships in the program, generates nodes representing functions, with edges representing function calls, recognizes nested structures in the code such as nested statement blocks in classes, methods, and functions, records the dependency relationships between parent and child nodes, and connects nested nodes such as the structures between function calls and class-member functions to form a hierarchical dependency structure.

[0077] In one implementation, the control flow graph, data flow graph, call graph, and structure dependency are integrated into a unified program dependency graph. The nodes of the program dependency graph represent operations in the program such as variable definitions, statements, functions, etc., while the edges represent three types of dependency relationships: data dependency, control dependency, and structure dependency.

[0078] In one implementation, the code relationship tree organizes each code block of the source code in a tree structure, which can clearly display the relationships and dependencies between code blocks. Through the dependency graph and the code relationship tree, auditors can quickly identify and analyze potential problems in the code, improving the efficiency and accuracy of code auditing.

[0079] In one embodiment, according to the code relationship tree, the context code related to vulnerabilities in the source code is extracted through context slicing to obtain the target code fragment, including:

[0080] Determine the vulnerabilities in the source code through vulnerability detection rules, and mark the vulnerabilities in the code relationship tree to obtain the starting locations;

[0081] According to the code relationship tree, starting from the starting location, forward and backward extract the nodes related to the vulnerabilities along the code relationship tree, and obtain the code fragments corresponding to all the nodes related to the vulnerabilities. Then, perform forward and backward slicing along the control flow and data flow to obtain the target code fragment.

[0082] In one implementation, determining the vulnerabilities in the source code through vulnerability detection rules and marking the starting locations in the code relationship tree helps to accurately locate the vulnerability positions; by extracting the nodes related to the vulnerabilities forward and backward from the vulnerability starting locations, all the context code affecting the vulnerabilities can be comprehensively obtained, which helps to understand the root causes and scopes of influence of the vulnerabilities.

[0083] In one implementation, the control flow is based on the control dependency graph, and the data flow is based on the data dependency graph.

[0084] In one implementation, vulnerability detection tools can use static code analysis tools such as SonarQube, CodeQL, FindBugs, etc.; the detection rules are common vulnerability rules such as the vulnerability classification of OWASP Top 10, CWE, etc. Once a vulnerability is detected, record the specific location where the vulnerability is located, such as the line number of the code, the file name, and the type of vulnerability, such as unvalidated user input. This is the starting point.

[0085] In one implementation, forward slicing starts from the vulnerability point, traces the call chain of the function, and captures all affected functions and code blocks along the call path. For example, if the vulnerability point is in a function, forward slicing will find all subsequent parts called by this function, as well as other code parts that call this function; forward-backward slicing traces how the variable data at the vulnerability point is passed, identifies which variables are modified at the vulnerability point, and traces their usage scenarios. For example, if the vulnerability point affects a certain global variable, forward slicing will capture all subsequent code that uses this variable.

[0086] In one implementation, by analyzing the control flow related to the vulnerability, the triggering conditions and code paths of the vulnerability can be understood, which helps to identify which conditions and branches may lead to the vulnerability. By analyzing the data flow, it can be understood how data propagates in the code and how it affects the manifestation of the vulnerability, so as to identify potential problems in data processing and transmission. Combining control flow and data flow slicing can comprehensively analyze the vulnerability from multiple perspectives.

[0087] In one implementation, the target code fragment obtained through context slicing is directly related to the vulnerability, which helps developers quickly locate and fix the vulnerability, improve the repair efficiency, focus on the code fragments related to the vulnerability, reduce the analysis and inspection of irrelevant code, and optimize the repair process.

[0088] In one embodiment, the feature extraction of the target code fragment to obtain the first feature includes:

[0089] Obtain the local abstract syntax tree and local control flow graph corresponding to the target code fragment, and extract the node information in the local abstract syntax tree to obtain the first information;

[0090] Extract the control flow paths related to the fault in the local control flow graph to obtain the second information; extract the usage frequency, assignment chain, and usage chain of key variables in the code to obtain the third information;

[0091] Substitute the first information, the second information, and the third information into the machine learning model for feature extraction to obtain the first feature.

[0092] In one implementation, the local abstract syntax tree provides detailed syntax structure information of the target code snippet, which helps to understand the syntax and structural features of the code; by analyzing the abstract syntax tree nodes, potential code patterns, structures, and possible errors can be identified, assisting in a deeper understanding of the code's functionality and logic.

[0093] In one implementation, the local control flow graph provides control flow information of the target code snippet. Extracting relevant control flow paths can help analyze the execution logic of the code and identify the control flow paths that may be affected by vulnerabilities; extracting the usage frequency of key variables in the code can identify the importance and influence of variables in the code, helping to understand how variables affect the behavior of the code, and analyzing the assignment chain and usage chain of variables can track the life cycle and data flow of variables, assisting in identifying the role of variables in vulnerabilities and their scope of influence.

[0094] In one implementation, combining the analysis results of the local abstract syntax tree, local control flow graph, and key variables can extract the features of the target code snippet from multiple dimensions, improving the accuracy of feature extraction.

[0095] In one implementation, the extracted features are standardized and normalized to ensure that the numerical ranges of different features are consistent. The attribute values of nodes, the conditions of control flow paths, and the usage frequencies of variables are standardized to a unified scale, and discrete features are encoded. For example, node types, conditions of control flow paths, etc. are converted into numerical representations, and one-hot encoding is used to process categorical features. The features extracted from the local abstract syntax tree, local control flow graph, and key variable analysis are integrated into a feature vector. A support vector machine is selected as the training model, and the machine learning model is trained using training data including the features of the target code snippet and labeled vulnerability information, learning the relationship between the overlearned features and labels, and adjusting the internal parameters to minimize the prediction error.

[0096] In one embodiment, feature fusion of the first feature and the second feature to obtain the target feature includes:

[0097] Performing linear transformations on the first feature and the second feature respectively to obtain the first linear feature and the second linear feature;

[0098] Dividing the first linear feature into multiple subspaces, performing parallel attention operations to obtain multiple attention heads, and concatenating all the attention heads to obtain the first target feature;

[0099] Performing weighted fusion on the first target feature and the second linear feature to obtain the target feature.

[0100] In one implementation, the linear transformation can better adapt to the learning requirements of the model by adjusting the weights and biases of the features, making the features more suitable for subsequent processing and fusion. Performing linear transformations on the first feature and the second feature respectively can optimize the representation of the features, making their distribution in the feature space more in line with the learning requirements of the model.

[0101] In one implementation, the weight of the first target feature is 0.6, and the weight of the second target feature is 0.4.

[0102] In one implementation, the attention mechanism can process features from multiple attention heads, capturing different information and patterns in the features. The concatenation of multiple attention heads can fuse different feature information, improving the expression ability and representativeness of the features. Weighted fusion of the first target feature generated by the attention mechanism and the second linear feature can combine the advantages of both and enhance the overall expression ability of the features.

[0103] Based on the same inventive concept, the embodiments of the present invention also provide a security diagnosis and evaluation system for open-source software. Refer to Figure 2 , Figure 2 which is a framework diagram of a security diagnosis and evaluation system for open-source software provided by the embodiments of the present invention, including:

[0104] A code relationship tree determination module, configured to obtain the source code of the open-source software, parse the source code through the Joern tool to obtain a code attribute graph, extract dependencies from the code attribute graph to obtain a program dependency graph, and construct a code relationship tree according to the program dependency graph;

[0105] A first feature extraction module, configured to, according to the code relationship tree, extract context code related to vulnerabilities in the source code through context slicing to obtain a target code segment, and extract a first feature from the target code segment;

[0106] A second feature extraction module, configured to anonymize variable names and function names in the source code, perform word embedding on the anonymized code sequence through Word2Vec, and use a bidirectional gated recurrent unit model to extract syntactic information of the target code segment to obtain a second feature;

[0107] A feature fusion module, configured to perform feature fusion on the first feature and the second feature to obtain a target feature, substitute the target feature into a multi-layer perceptron model to obtain the probability of the existence of vulnerabilities in the source code, and perform a security evaluation on the open-source software according to the probability of the existence of vulnerabilities.

[0108] A security diagnosis and evaluation system for open-source software provided by an embodiment of the present invention generates a code attribute graph (CAG) through the Joern tool and extracts a program dependency graph. The solution can deeply analyze the structure of the source code, comprehensively capture the dependencies between codes, and then extract target code fragments through context slicing, reducing the amount of code to be analyzed and improving the speed of the overall solution. By fusing the obtained first feature and second feature, the solution can more comprehensively and accurately identify potential vulnerability locations and types, improving security and reducing the risk of being attacked by the supply chain.

[0109] In one embodiment, the code relationship tree determination module includes:

[0110] A dependency extraction module, which is used to find the usage relationship between each variable through tracking the data flow graph in the code attribute graph to obtain data dependencies, obtain control dependencies by tracking the control flow graph in the code attribute graph, and obtain structure dependencies by analyzing the scopes and hierarchical structures in the source code;

[0111] A dependency integration module, which is used to integrate data dependencies, control dependencies, and structure dependencies to obtain a program dependency graph;

[0112] A code relationship tree establishment module, which is used to obtain the main function in the source code as the root node, identify each code block in the source code as a child node according to the program dependency graph, add it under the root node, and recursively add corresponding child nodes to each child node according to the program dependency graph until all code blocks are identified to obtain a code relationship tree.

[0113] In one embodiment, the first feature extraction module includes:

[0114] A vulnerability marking module, which is used to determine vulnerabilities in the source code through vulnerability detection rules and mark the vulnerabilities in the code relationship tree to obtain the starting locations;

[0115] A target code fragment determination module, which is used to forward and backward extract nodes related to vulnerabilities along the code relationship tree with the starting location as the starting point according to the code relationship tree, obtain code fragments corresponding to all nodes related to vulnerabilities, and perform forward and backward slicing along the control flow and data flow to obtain target code fragments.

[0116] In one embodiment, the first feature extraction module further includes:

[0117] A node information extraction module, which is used to obtain the local abstract syntax tree and local control flow graph corresponding to the target code fragment, and extract node information in the local abstract syntax tree to obtain the first information;

[0118] An information extraction module for extracting the control flow paths related to faults in the local control flow graph to obtain the second information; and extracting the usage frequency, assignment chain, and usage chain of key variables in the code to obtain the third information.

[0119] A first feature fusion module for substituting the first information, the second information, and the third information into a machine learning model for feature extraction to obtain the first feature.

[0120] In one embodiment, the feature fusion module includes:

[0121] A linear transformation module for respectively performing linear transformations on the first feature and the second feature to obtain the first linear feature and the second linear feature;

[0122] A first target feature construction module for dividing the first linear feature into multiple subspaces, performing parallel attention operations to obtain multiple attention heads, and concatenating all the attention heads to obtain the first target feature;

[0123] A feature weighted fusion module for performing weighted fusion on the first target feature and the second linear feature to obtain the target feature.

[0124] The above has described in detail one embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention shall still fall within the scope covered by the patent of the present invention.

Claims

1. A security diagnosis and evaluation method for open-source software, characterized in that, The method includes: Obtain the source code of the open-source software, parse the source code through the Joern tool to obtain a code property graph, extract dependencies from the code property graph to obtain a program dependency graph, and construct a code relationship tree according to the program dependency graph; According to the code relationship tree, extract the context code related to the vulnerability in the source code through context slicing to obtain a target code snippet, and extract features from the target code snippet to obtain a first feature; After anonymizing the variable names and function names in the source code, perform word embedding on the anonymized code sequence through Word2Vec, and use a bidirectional gated recurrent unit model to extract the syntactic information of the target code snippet to obtain a second feature; Fuse the first feature and the second feature to obtain a target feature, substitute the target feature into a multi-layer perceptron model to obtain the probability of the existence of a vulnerability in the source code, and perform a security assessment on the open-source software according to the probability of the existence of the vulnerability; Extracting features from the target code snippet to obtain a first feature includes: Obtain the local abstract syntax tree and the local control flow graph corresponding to the target code snippet, and extract the node information in the local abstract syntax tree to obtain a first piece of information; Extract the control flow paths related to faults in the local control flow graph to obtain a second piece of information; extract the usage frequency, assignment chain, and usage chain of key variables in the code to obtain a third piece of information; Substitute the first piece of information, the second piece of information, and the third piece of information into a machine learning model for feature extraction to obtain a first feature.

2. The security diagnosis and evaluation method for open-source software according to claim 1, wherein Extracting dependencies from the code property graph to obtain a program dependency graph and constructing a code relationship tree according to the program dependency graph includes: Find the usage relationship between each variable by tracking the data flow graph in the code property graph to obtain data dependencies, obtain control dependencies by tracking the control flow graph in the code property graph, and obtain structure dependencies by analyzing the scope and hierarchical structure in the source code; Integrate the data dependencies, the control dependencies, and the structure dependencies to obtain a program dependency graph; Obtain the main function in the source code as the root node, and according to the program dependency graph, identify each code block in the source code as a sub-node and add it under the root node, and recursively add corresponding sub-nodes to each sub-node according to the program dependency graph until all code blocks are identified to obtain a code relationship tree.

3. The security diagnosis and evaluation method for open-source software according to claim 1, wherein According to the code relationship tree, extracting the context code related to the vulnerability in the source code through context slicing to obtain a target code snippet includes: Determine the vulnerability in the source code through vulnerability detection rules, and mark the vulnerability in the code relationship tree to obtain the starting location; According to the code relationship tree, forward and backward extract the nodes related to the vulnerability along the code relationship tree starting from the starting location, and obtain the code snippets corresponding to all nodes related to the vulnerability and perform forward and backward slicing along the control flow and data flow to obtain a target code snippet.

4. A security diagnosis and evaluation method for open-source software according to claim 1, characterized in that, Fusing the first feature and the second feature to obtain a target feature includes: Perform linear transformations on the first feature and the second feature respectively to obtain a first linear feature and a second linear feature; Divide the first linear feature into multiple subspaces, perform parallel attention operations to obtain multiple attention heads, and splice all the attention heads to obtain a first target feature; Perform weighted fusion on the first target feature and the second linear feature to obtain a target feature.

5. A security diagnosis and evaluation system for open-source software, characterized in that, The system includes: A code relationship tree determination module, configured to obtain the source code of the open-source software, parse the source code through the Joern tool to obtain a code property graph, extract dependencies from the code property graph to obtain a program dependency graph, and construct a code relationship tree according to the program dependency graph; A first feature extraction module, configured to, according to the code relationship tree, extract context code related to vulnerabilities in the source code through context slicing to obtain a target code segment, and extract a first feature from the target code segment; A second feature extraction module, configured to anonymize variable names and function names in the source code, perform word embedding on the anonymized code sequence through Word2Vec, and use a bidirectional gated recurrent unit model to extract syntactic information of the target code segment to obtain a second feature; A feature fusion module, configured to perform feature fusion on the first feature and the second feature to obtain a target feature, substitute the target feature into a multi-layer perceptron model to obtain the vulnerability existence probability of the source code, and perform security evaluation on the open-source software according to the vulnerability existence probability.

6. A security diagnosis and evaluation system for open-source software according to claim 5, characterized in that The code relationship tree determination module includes: A dependency extraction module, configured to find the usage relationship between each variable by tracking the data flow graph in the code property graph to obtain data dependencies, obtain control dependencies by tracking the control flow graph in the code property graph, and obtain structure dependencies by analyzing the scope and hierarchical structure in the source code; A dependency integration module, configured to integrate the data dependencies, the control dependencies, and the structure dependencies to obtain a program dependency graph; A code relationship tree establishment module, configured to obtain the main function in the source code as the root node, and according to the program dependency graph, identify each code block in the source code as a sub-node and add it under the root node, and recursively add corresponding sub-nodes to each sub-node according to the program dependency graph until all code blocks are identified, to obtain a code relationship tree.

7. The security diagnosis and evaluation system for open-source software according to claim 5, characterized in that The first feature extraction module includes: A vulnerability marking module, configured to determine vulnerabilities in the source code through vulnerability detection rules, and mark the vulnerabilities in the code relationship tree to obtain a starting location; A target code segment determination module, configured to, according to the code relationship tree, forward and backward extract nodes related to vulnerabilities along the code relationship tree starting from the starting location, and obtain code segments corresponding to all nodes related to vulnerabilities, and perform forward and backward slicing along the control flow and data flow to obtain a target code segment.

8. A security diagnosis and evaluation system for open-source software according to claim 5, characterized in that, The feature fusion module includes: A linear transformation module, configured to perform linear transformations on the first feature and the second feature respectively to obtain a first linear feature and a second linear feature; The first target feature construction module is used to divide the first linear feature into multiple subspaces, perform parallel attention operations to obtain multiple attention heads, and splice all the attention heads to obtain the first target feature; The feature weighted fusion module is used to perform weighted fusion on the first target feature and the second linear feature to obtain the target feature.

Citation Information

Patent Citations

  • Source code vulnerability detection method based on data dependence enhancement program slicing

    CN116702160A

  • Code vulnerability detection method based on efficient parameter fine tuning of code pre-training model

    CN119441006A