Code vulnerability automatic mining and identification method, system and equipment
By building source code control flow diagrams and performing vulnerability analysis, dividing control flow nodes and calling preset vulnerability codes and symbol execution modules, the problems of low code vulnerability mining and identification and high false alarm rate in the existing technology are solved, and more efficient vulnerability identification and lower false alarm rate are achieved.
Patent Information
- Application Number
- CN202510533598.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
In the existing technology, code vulnerability mining and identification efficiency is low and false alarm rate is high. It cannot fully cover all potential vulnerabilities, and is easily affected by program size and code complexity.
By collecting the source code files of the target software, building a source code control flow chart, conducting a confidence analysis of the risk and vulnerability code ratio of vulnerability, generating risk indicators and code ratio confidence indicators, dividing control flow nodes, building a collection of control flow jump paths of different categories, and calling preset vulnerability codes and symbol execution modules for vulnerability identification.
It improves the accuracy of vulnerability identification, reduces the false alarm rate, and improves the efficiency of code vulnerability mining and identification.
Smart Images

Figure CN120068093A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method, system and device for automatically mining and identifying code vulnerabilities. Background Art
[0002] With the rapid development of information technology, the complexity of software systems has been increasing day by day, and security vulnerabilities have become one of the key factors affecting system stability and data security. At present, the existing vulnerability detection methods mainly rely on manual review or static analysis-based tools. However, these methods have obvious limitations, cannot comprehensively cover all potential vulnerabilities, and are easily affected by the program scale and code complexity, resulting in low vulnerability detection efficiency. In addition, when dealing with large-scale code, the existing automated vulnerability detection tools often face high false positive rates and false negative rates, and thus developers need to invest a lot of time in manual screening and repair, which greatly affects the development efficiency. Summary of the Invention
[0003] This application provides a method, system and device for automatically mining and identifying code vulnerabilities, which solves the technical problems of low efficiency in mining and identifying code vulnerabilities and high false positive rate in the prior art.
[0004] In the first aspect of this application, a method for automatically mining and identifying code vulnerabilities is provided. The method includes: Collecting source code files of a target software and constructing a source code control flow graph; performing vulnerability introduction risk and vulnerability code comparison confidence analysis on multiple control flow nodes in the source code control flow graph to generate multiple introduction risk indicators and multiple code comparison confidence indicators; dividing the multiple control flow nodes into a first type of nodes and a second type of nodes based on the multiple introduction risk indicators and the multiple code comparison confidence indicators; constructing a first type of control flow jump path set and a second type of control flow jump path set for the first type of nodes and the second type of nodes; calling a first type of preset vulnerability code to perform jump path code comparison on the first type of control flow jump path set to generate a first vulnerability identification result, and calling a symbolic execution module to perform path vulnerability identification on the second type of control flow jump path set to generate a second vulnerability identification result.
[0005] In the second aspect of this application, a system for automatically mining and identifying code vulnerabilities is provided. The system includes: A control flow graph construction module is used to collect the source code files of the target software and construct a source code control flow graph; a confidence analysis module is used to perform vulnerability introduction risk and vulnerability code comparison confidence analysis on multiple control flow nodes in the source code control flow graph, and generate multiple introduction risk indicators and multiple code comparison confidence indicators; a division module is used to divide the multiple control flow nodes into a first type of node and a second type of node based on the multiple introduction risk indicators and the multiple code comparison confidence indicators; a jump path construction module is used to construct a first type of control flow jump path set and a second type of control flow jump path set for the first type of node and the second type of node; a comparison module is used to call a first type of preset vulnerability code to perform jump path code comparison on the first type of control flow jump path set, generate a first vulnerability identification result, and call a symbolic execution module to perform path vulnerability identification on the second type of control flow jump path set, generate a second vulnerability identification result.
[0006] In a third aspect of the present application, there is provided an electronic device, including: a memory for storing executable instructions; a processor for implementing the code vulnerability automatic mining and identification method provided by the present application when executing the executable instructions stored in the memory.
[0007] One or more technical solutions provided in the present application have at least the following technical effects or advantages: First, collect the source code files of the target software and construct a source code control flow graph. Then, perform vulnerability introduction risk and vulnerability code comparison confidence analysis on multiple control flow nodes in the source code control flow graph, and generate multiple introduction risk indicators and multiple code comparison confidence indicators. Further, divide the multiple control flow nodes into a first type of node and a second type of node based on the multiple introduction risk indicators and the multiple code comparison confidence indicators. Then, construct a first type of control flow jump path set and a second type of control flow jump path set for the first type of node and the second type of node. Finally, call a first type of preset vulnerability code to perform jump path code comparison on the first type of control flow jump path set, generate a first vulnerability identification result, and call a symbolic execution module to perform path vulnerability identification on the second type of control flow jump path set, generate a second vulnerability identification result. This solves the technical problems of low efficiency and high false alarm rate in code vulnerability mining and identification in the prior art, and achieves the technical effects of improving the accuracy of vulnerability identification and reducing the false alarm rate. Description of the Drawings
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0009] Figure 1 Schematic flowchart of the method for automatically mining and identifying code vulnerabilities provided by the embodiments of the present application; Figure 2 Schematic structural diagram of the system for automatically mining and identifying code vulnerabilities provided by the embodiments of the present application; Figure 3 Schematic structural diagram of an exemplary electronic device of the present application.
[0010] Explanation of reference numerals: module 11 for constructing a control flow graph, confidence analysis module 12, partitioning module 13, module 14 for constructing jump paths, comparison module 15, processor 21, memory 22, input device 23, output device 24. Detailed implementation manners
[0011] By providing a method, system, and device for automatically mining and identifying code vulnerabilities, the present application solves the technical problems of low efficiency and high false alarm rate in code vulnerability mining and identification in the prior art.
[0012] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0013] It should be noted that the terms "including" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0014] Embodiment 1, as Figure 1 shown, the embodiments of the present application provide a method for automatically mining and identifying code vulnerabilities, where the method includes: Collect the source code files of the target software and construct a source code control flow graph.
[0015] Retrieve the source code file from the code library of the target software and convert it into an Abstract Syntax Tree (AST). An Abstract Syntax Tree is a tree-like data structure used to represent the syntax structure in the source code. Each node represents a syntax unit in the source code, such as statements, expressions, variables, etc. Identify the control flow structure in the source code based on the Abstract Syntax Tree, divide the program into multiple basic blocks, and each basic block represents a continuous piece of executable code. Construct the control flow graph of the source code according to the jump relationships between the basic blocks. The nodes in the graph represent the basic blocks, and the edges represent the jump relationships of the program execution paths, thereby providing structured data support for subsequent vulnerability mining and identification.
[0016] Furthermore, collect the source code file of the target software and construct the source code control flow graph, including: Convert the source code of the source code file into an Abstract Syntax Tree; identify the control flow structure based on the Abstract Syntax Tree, divide the basic blocks, use the basic blocks as control flow nodes, and generate the source code control flow graph according to the jump relationships of the basic blocks.
[0017] Retrieve the source code file from the code library of the target software, call the compiler front-end tools for the corresponding language (such as Clang's LibTooling for C / C++, Java's Eclipse JDT for Java, Python's AST module for Python) to perform syntax parsing on the source code and generate an Abstract Syntax Tree. Each node of the Abstract Syntax Tree represents a syntax structure in the source code (such as expressions, statements, function definitions, etc.), and the syntax hierarchy is represented by the parent-child relationship between the nodes.
[0018] Based on the constructed Abstract Syntax Tree, identify the control flow structure in the source code. The control flow structure refers to the elements that control the execution path in the program, such as conditional statements (e.g., if, else statements), loop statements (e.g., for, while statements), and jump statements (e.g., break, continue statements). Based on the identified control flow structure, divide the source code into multiple basic blocks. A basic block is a continuous code segment in the source code without jump instructions or branches, representing an execution unit in the program. Construct the source code control flow graph according to the jump relationships of the basic blocks. The source code control flow graph is a graphical representation of the program structure, where each node in the graph represents a basic block, and the edges between the nodes represent the jump relationships from one basic block to another.
[0019] Perform vulnerability introduction risk and vulnerability code comparison confidence analysis on multiple control flow nodes in the source code control flow graph to generate multiple introduction risk indicators and multiple code comparison confidence indicators.
[0020] In the source code control flow graph, vulnerability introduction risk and vulnerability code comparison confidence analysis are performed on multiple control flow nodes, that is, by evaluating the code characteristics of each control flow node, potential security hazards are identified to generate introduction risk indicators; at the same time, historical vulnerability codes related to the nodes are mined, similarity comparison is carried out, the similarity with known vulnerability codes is calculated, and a code comparison confidence indicator is generated. The code comparison confidence indicator is used to indicate whether vulnerabilities can be effectively identified only through code comparison methods.
[0021] Furthermore, in the source code control flow graph, vulnerability introduction risk and vulnerability code comparison confidence analysis are performed on multiple control flow nodes to generate multiple introduction risk indicators and multiple code comparison confidence indicators, including: For each of the multiple control flow nodes, a one-hop path is constructed to generate multiple hop paths; path complexity, code complexity, and data flow risk analysis and fusion are performed on the multiple hop paths to generate the multiple introduction risk indicators; the code structure and path type of the multiple hop paths are collected, and historical vulnerability codes are mined based on the path type, and a discrimination comparison is performed to generate the multiple code comparison confidence indicators.
[0022] During the process of performing vulnerability introduction risk and vulnerability code comparison confidence analysis on multiple control flow nodes in the source code control flow graph, a one-hop path is constructed for each control flow node respectively, and multiple hop paths are generated. Each hop path represents the path passed along the execution flow of the program starting from a control flow node. Comprehensive analysis is performed on these hop paths, including path complexity, code complexity, and data flow risk.
[0023] Path complexity analysis involves counting the length of the path and the number of branches in the path. Paths with a large path length and the number of branches usually increase the risk of errors or vulnerability introduction during execution. Code complexity analysis evaluates the complexity of the code by identifying long and complex function call chains on the path. A complex code structure is likely to conceal potential vulnerabilities. Data flow risk analysis evaluates the security of the data flow by analyzing the data flow dependency relationships in the path, especially the risks of issues such as lack of input data verification and exposure of sensitive data. The results of these three analyses are fused to generate multiple introduction risk indicators for evaluating the potential vulnerability introduction risk of each hop path.
[0024] The code structure and path type of multiple hop paths are collected, historical vulnerability codes are mined based on the path type, and a discrimination comparison is performed to generate multiple code comparison confidence indicators. Specifically, a code comparison confidence indicator is generated by comparing the similarity between the historical secure code and historical vulnerability codes of the hop path.
[0025] Furthermore, perform a fusion of path complexity, code complexity, and data flow risk analysis on the multiple jump paths to generate the multiple introduced risk indicators, including: Construct a path analysis mechanism, a code analysis mechanism, and a data flow analysis mechanism; among them, the path analysis mechanism performs complexity conversion mapping by counting the path length and the number of path branches. Among them, the path complexity is proportional to the path length and the number of branches; the code analysis mechanism performs code complexity mapping by identifying code verbosity and security risks in the function call chain; the data flow analysis mechanism performs data flow risk conversion mapping by performing input validation missing, data flow dependency analysis, and sensitive data exposure analysis; call the path analysis mechanism, the code analysis mechanism, and the data flow analysis mechanism to process the multiple jump paths respectively, and generate the path complexity, code complexity, and data flow risk corresponding to the multiple jump paths respectively; perform weighted fusion on the path complexity, code complexity, and data flow risk corresponding to the multiple jump paths respectively to generate the multiple introduced risk indicators.
[0026] The path analysis mechanism evaluates the complexity of the path by counting the length and the number of path branches of each jump path. Among them, the path length represents the number of code segments included in the path, and the number of path branches indicates the possible number of jump points in the path. Path complexity calculation formula: C = W1×L + W2×B, where C is the complexity value of the path, W1 and W2 are the weight coefficients of the path length and the number of branches respectively, L is the path length, and B is the number of branches in the path. The setting of the weight coefficient can be adjusted according to the actual situation. Usually, long paths and paths with many branches will be given higher weights.
[0027] The code analysis mechanism evaluates the complexity of the code by identifying the possible verbose code and complex function call chains in the code. The more verbose the code in the path and the more complex the logic, the greater the possibility of omitting or introducing vulnerabilities; whether there are security risks in the function calls involved in the path, such as calling third-party libraries, external systems, etc. The code analysis mechanism maps these verbose parts and potential security risks and generates corresponding code complexity values to quantify the complexity of the code.
[0028] The data flow analysis mechanism evaluates the risk of the data flow by analyzing risk factors such as input validation missing, data flow dependency, and sensitive data exposure in the jump path. Missing input validation may lead to unsafe data processing, incorrect data flow dependencies may lead to information leakage or execution errors, and sensitive data exposure may lead to data leakage or being attacked. The data flow analysis mechanism generates a data flow risk map through these risk factors to quantify the risk related to the data flow in the path.
[0029] The path analysis mechanism, code analysis mechanism, and data flow analysis mechanism process each jump path respectively, generating the path complexity, code complexity, and data flow risk value for each jump path. The path analysis mechanism counts the length and number of branches of the path to calculate the path complexity; the code analysis mechanism calculates the code complexity by identifying verbose code and potential security hazards; the data flow analysis mechanism generates the data flow risk value by checking input validation, data flow dependencies, and sensitive data exposure issues. The path complexity, code complexity, and data flow risk value are weighted and fused. Weighted fusion assigns different weights according to the importance of each risk factor in the overall vulnerability introduction, and comprehensively calculates each complexity value and risk value to obtain the introduction risk index for each jump path. The introduction risk index reflects the overall risk of the jump path potentially triggering vulnerabilities during execution. The path complexity, code complexity, and data flow risk are all taken into account to comprehensively evaluate the vulnerability risk of each path.
[0030] Furthermore, collect the code structures and path types of the multiple jump paths, mine historical vulnerable codes, conduct discrimination comparison, and generate the multiple code comparison confidence indicators, including: Using the code structures and path types of the multiple jump paths as index elements respectively, retrieve the historical vulnerable codes and historical secure codes corresponding to various historical vulnerability types; based on the historical vulnerable codes and the historical secure codes, conduct similarity analysis of the vulnerable codes and secure codes to generate multiple similarity indicators corresponding to the multiple jump paths; subtract each of the multiple similarity indicators from 1 to generate the multiple code comparison confidence indicators.
[0031] Preferably, extract the code structure and path type of each jump path and use them as index elements. The code structure reflects the specific code implementation method in the jump path, while the path type indicates the execution mode of the path, such as conditional jump, loop, or exception handling, etc.; based on this information, the system will retrieve the historical vulnerable codes and historical secure codes, which respectively correspond to different vulnerability types and security standards; based on the retrieved historical vulnerable codes and historical secure codes, conduct similarity analysis, and use a suitable similarity measurement method (such as cosine similarity, Jaccard similarity, or edit distance) to quantify the similarity between the historical vulnerable codes and historical secure codes. The higher the similarity, the more difficult it is to identify the vulnerable code through code comparison. For each jump path, subtract the similarity indicator from 1 to generate the code comparison confidence indicator.
[0032] Based on the multiple introduction risk indicators and the multiple code comparison confidence indicators, divide the multiple control flow nodes into first-class nodes and second-class nodes.
[0033] For each control flow node, its potential vulnerability risk and the difficulty of vulnerability identification are judged by evaluating its introduced risk index and code comparison confidence index. The introduced risk index is used to quantify the risk of potential vulnerabilities introduced during the execution of the node, and is usually comprehensively evaluated based on factors such as path complexity, code verbosity, and data flow dependencies. The code comparison confidence index is used to indicate whether vulnerabilities can be effectively identified only through code comparison methods. The larger the code comparison confidence index, the easier it is to identify through code comparison.
[0034] By comprehensively considering the introduced risk index and the code comparison confidence index, multiple control flow nodes are divided into the first type of nodes and the second type of nodes. Among them, the first type of nodes identify vulnerabilities only by comparing codes, and the second type of nodes use symbolic execution with higher precision.
[0035] Furthermore, based on the multiple introduced risk indexes and the multiple code comparison confidence indexes, dividing the multiple control flow nodes into the first type of nodes and the second type of nodes includes: Based on the confidence index threshold and the risk index threshold, the multiple control flow nodes are classified to generate the first type of nodes and the second type of nodes. Among them, the first type of nodes are control flow nodes whose introduced risk index is less than the risk index threshold and whose code comparison confidence index is greater than the confidence index threshold; among them, when there is a classification contradiction between the confidence index threshold and the risk index threshold during classification, the classification priority of the risk index threshold is greater than the classification priority of the risk index threshold.
[0036] Classify the control flow nodes based on the risk index threshold and the confidence index threshold. Specifically, if the introduced risk index of a certain control flow node is less than the preset risk index threshold and the code comparison confidence index of this node is greater than the preset confidence index threshold, then this node is classified as the first type of nodes. The first type of nodes usually indicates that the vulnerability risk of these nodes is relatively low, and at the same time, vulnerabilities can be effectively identified through code comparison methods. Therefore, vulnerability detection can be performed only by comparing codes without using other complex vulnerability detection methods.
[0037] If the introduced risk index of a certain node is greater than the risk index threshold and its code comparison confidence index is less than the confidence index threshold, then this node is classified as the second type of nodes. The second type of nodes usually indicates that the difficulty of vulnerability identification of this node is relatively high, and its potential vulnerability risk is relatively large, and it cannot be effectively identified through simple code comparison methods. For the second type of nodes, more precise vulnerability detection methods such as symbolic execution are used to improve the accuracy and precision of vulnerability detection.
[0038] During the classification process, if there is a classification conflict between the confidence index threshold and the risk index threshold (i.e., a certain node meets the criteria of the first type of node under certain conditions, and meets the criteria of the second type of node under other conditions), the classification priority of the risk index threshold is set to be higher than that of the confidence index threshold. That is to say, if the introduced risk index of a node is high (i.e., the risk is large), the node will be preferentially classified as the second type of node, although the confidence level of code comparison is high, which can ensure the detection of higher-risk vulnerabilities.
[0039] Construct a first type of control flow jump path set and a second type of control flow jump path set for the first type of node and the second type of node.
[0040] Based on the first type of node and the second type of node, analyze each jump path and assign it to the corresponding path set. By analyzing the type of each jump path, it is divided into a first type of control flow jump path set and a second type of control flow jump path set. For the first type of node, since its vulnerability risk is low and vulnerabilities can be effectively identified through code comparison, the jump paths related to these nodes are classified into the first type of path set for subsequent code comparison vulnerability detection. For the second type of node, the vulnerability risk of these nodes is high and it is difficult to identify through code comparison, so the jump paths related to them are classified into the second type of path set for subsequent precise detection methods such as symbolic execution.
[0041] Furthermore, constructing a first type of control flow jump path set and a second type of control flow jump path set for the first type of node and the second type of node includes: Distinguish and mark the first type of node and the second type of node in the source code control flow graph to generate a marked control flow graph; in the marked control flow graph, generate the first type of control flow jump path set according to the covered paths of the first type of marked nodes, and generate the second type of control flow jump path set according to the covered paths of the second type of marked nodes.
[0042] In the control flow graph of the source code, distinguish and mark the control flow nodes that have been divided into the first type of node and the second type of node to form a marked control flow graph; in the marked control flow graph, identify and extract the covered paths related to these nodes according to the marks of the first type of nodes. Each covered path represents all possible paths during the program execution starting from the first type of node, and all these paths will be combined to form the first type of control flow jump path set. Similarly, according to the marks of the second type of nodes, extract the covered paths related to the second type of nodes from the marked control flow graph, and all paths involving the second type of nodes will be combined to form the second type of control flow jump path set.
[0043] Invoke the first type of preset vulnerability code to compare the jump path codes in the first type of control flow jump path set, generate the first vulnerability identification result, and invoke the symbolic execution module to identify path vulnerabilities in the second type of control flow jump path set, generating the second vulnerability identification result.
[0044] For the first type of control flow jump path set and the second type of control flow jump path set, different vulnerability identification methods are adopted for processing. For the first type of control flow jump path set, which contains paths that can effectively identify vulnerabilities through code comparison, the system invokes the preset first type of vulnerability code, which is constructed based on common vulnerability patterns in the historical vulnerability database. By comparing with the code structure, syntax features, and function implementation of the first type of paths, the system identifies potential vulnerability paths and generates the first vulnerability identification result, indicating the possible types of vulnerabilities. For the second type of control flow jump path set, which contains complex paths with high vulnerability risks and cannot be accurately identified through code comparison, the system invokes the symbolic execution module for path vulnerability identification. Symbolic execution analyzes potential vulnerabilities, such as memory leaks and buffer overflows, by simulating program execution and tracking each branch path. Through the symbolic execution module, the system generates the second vulnerability identification result, including the possible vulnerabilities in the path and their specific locations.
[0045] Furthermore, the first type of preset vulnerability code is historical vulnerability code retrieved based on the path types and path code structures in the first type of control flow jump path set; the symbolic execution module is a symbolic execution engine constructed based on a preset symbolic execution tool.
[0046] The first type of preset vulnerability code is historical vulnerability code retrieved based on the path types and path code structures in the first type of control flow jump path set. Specifically, the system analyzes the type of each path in the first type of path set (such as conditional judgment, loop, exception handling, etc.) and its code structure (such as variable operations and function calls in the code), and retrieves known vulnerability codes from the historical vulnerability database that match these path types and code structures. This historical vulnerability code represents common vulnerability patterns that have occurred in similar code structures and path types and is used as the preset vulnerability code for subsequent vulnerability comparison and analysis.
[0047] The symbolic execution module is a symbolic execution engine built based on a preset symbolic execution tool. Existing symbolic execution tools include KLEE, Angr, Manticore, SymbEX, etc. Symbolic execution is a program analysis technique that explores all possible execution paths of a program under different input conditions by simulating the program's execution process and replacing concrete values with symbolic variables. The symbolic execution module uses a symbolic execution engine built with a preset symbolic execution tool, which can deeply analyze complex paths in the second type of control flow jump path set and identify potential security vulnerabilities.
[0048] In summary, the embodiments of the present application have at least the following technical effects: First, collect the source code files of the target software and construct a source code control flow graph. Then, perform vulnerability introduction risk and vulnerability code comparison confidence analysis on multiple control flow nodes in the source code control flow graph to generate multiple introduction risk indicators and multiple code comparison confidence indicators. Further, based on the multiple introduction risk indicators and multiple code comparison confidence indicators, divide the multiple control flow nodes into the first type of nodes and the second type of nodes. Then, construct a first type of control flow jump path set and a second type of control flow jump path set for the first type of nodes and the second type of nodes. Finally, call the first type of preset vulnerability code to perform jump path code comparison on the first type of control flow jump path set to generate a first vulnerability identification result, and call the symbolic execution module to perform path vulnerability identification on the second type of control flow jump path set to generate a second vulnerability identification result. This solves the technical problems of low efficiency and high false positive rate in code vulnerability mining and identification in the prior art, and achieves the technical effects of improving the accuracy of vulnerability identification and reducing the false positive rate.
[0049] Embodiment 2, based on the same inventive concept as the code vulnerability automatic mining and identification method in the foregoing embodiment, as Figure 2 shown, the present application provides a code vulnerability automatic mining and identification system, wherein the system includes: A control flow graph construction module 11 is used to collect the source code files of the target software and construct a source code control flow graph; a confidence analysis module 12 is used to perform vulnerability introduction risk and vulnerability code comparison confidence analysis on multiple control flow nodes in the source code control flow graph, and generate multiple introduction risk indicators and multiple code comparison confidence indicators; a division module 13 is used to divide the multiple control flow nodes into a first type of nodes and a second type of nodes based on the multiple introduction risk indicators and the multiple code comparison confidence indicators; a jump path construction module 14 is used to construct a first type of control flow jump path set and a second type of control flow jump path set for the first type of nodes and the second type of nodes; a comparison module 15 is used to call a first type of preset vulnerability code to perform jump path code comparison on the first type of control flow jump path set, generate a first vulnerability identification result, and call a symbolic execution module to perform path vulnerability identification on the second type of control flow jump path set, and generate a second vulnerability identification result.
[0050] Further, the confidence analysis module 12 is used to execute the following method: Construct a one-time jump path for each of the multiple control flow nodes to generate multiple jump paths; perform path complexity, code complexity, and data flow risk analysis and fusion on the multiple jump paths to generate the multiple introduction risk indicators; collect the code structures and path types of the multiple jump paths, mine historical vulnerability codes based on the path types, and perform discrimination comparison to generate the multiple code comparison confidence indicators.
[0051] Further, the confidence analysis module 12 is used to execute the following method: Construct a path analysis mechanism, a code analysis mechanism, and a data flow analysis mechanism; wherein, the path analysis mechanism performs complexity conversion mapping by counting the path length and the number of path branches, and the path complexity is proportional to the path length and the number of branches; the code analysis mechanism performs code complexity mapping by identifying code verbosity and security hazards in the function call chain; the data flow analysis mechanism performs data flow risk conversion mapping by performing input validation missing, data flow dependency analysis, and sensitive data exposure analysis; call the path analysis mechanism, the code analysis mechanism, and the data flow analysis mechanism to process the multiple jump paths respectively, and generate the path complexity, code complexity, and data flow risk corresponding to each of the multiple jump paths; perform weighted fusion on the path complexity, code complexity, and data flow risk corresponding to each of the multiple jump paths to generate the multiple introduction risk indicators.
[0052] Further, the confidence analysis module 12 is used to execute the following method: Using the code structures and path types of the multiple jump paths as index elements respectively, retrieve the historical vulnerability codes and historical security codes corresponding to various historical vulnerability types; perform similarity analysis on the vulnerability codes and security codes based on the historical vulnerability codes and the historical security codes, and generate multiple similarity metrics corresponding to the multiple jump paths; subtract 1 from each of the multiple similarity metrics to generate the multiple code comparison confidence metrics.
[0053] Further, the partitioning module 13 is configured to perform the following method: Based on a confidence metric threshold and a risk metric threshold, classify the multiple control flow nodes to generate the first type of nodes and the second type of nodes, where the first type of nodes are control flow nodes with an introduced risk metric less than the risk metric threshold and a code comparison confidence metric greater than the confidence metric threshold; when there is a classification contradiction between the confidence metric threshold and the risk metric threshold during classification, the classification priority of the risk metric threshold is higher than that of the confidence metric threshold.
[0054] Further, the jump path construction module 14 is configured to perform the following method: Distinguish and mark the first type of nodes and the second type of nodes in the source code control flow graph to generate a marked control flow graph; in the marked control flow graph, generate the first type of control flow jump path set according to the covered paths of the first type of marked nodes, and generate the second type of control flow jump path set according to the covered paths of the second type of marked nodes.
[0055] Further, the comparison module 15 is configured to perform the following method: The first type of preset vulnerability code is a historical vulnerability code retrieved based on the path types and path code structures in the first type of control flow jump path set; the symbolic execution module is a symbolic execution engine constructed based on a preset symbolic execution tool.
[0056] Further, the control flow graph construction module 11 is configured to perform the following method: Convert the source code of the source code file into an abstract syntax tree; perform control flow structure recognition based on the abstract syntax tree, partition basic blocks, use the basic blocks as control flow nodes, and generate the source code control flow graph according to the jump relationships of the basic blocks.
[0057] Embodiment III Figure 3 It is a schematic structural diagram of an electronic device provided in Embodiment III of the present invention, showing a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present invention. Figure 3 The displayed electronic device is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present invention. As Figure 3As shown, the electronic device includes a processor 21, a memory 22, an input device 23, and an output device 24; the number of processors 21 in the electronic device can be one or more. Figure 3 Taking one processor 21 as an example, the processor 21, the memory 22, the input device 23, and the output device 24 in the electronic device can be connected through a bus or other means. Figure 3 Taking connection through a bus as an example.
[0058] The memory 22, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the code vulnerability automatic mining and identification method in the embodiments of the present invention. The processor 21 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 22, that is, implements the above-mentioned code vulnerability automatic mining and identification method.
[0059] It should be noted that the above-mentioned sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above describes specific embodiments of this specification. The processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0060] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0061] This specification and the drawings are only exemplary descriptions of the present application and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications.
Claims
1. A method for automatically mining and identifying code vulnerabilities, characterized in that: The method comprises: Collect the source code files of the target software and build a source code control flow graph; Performing vulnerability introduction risk and vulnerability code comparison confidence analysis on multiple control flow nodes in the source code control flow graph to generate multiple introduction risk indicators and multiple code comparison confidence indicators; Based on the multiple introduction risk indicators and the multiple code comparison confidence indicators, the multiple control flow nodes are divided into a first type of nodes and a second type of nodes; Constructing a first-type control flow jump path set and a second-type control flow jump path set for the first-type nodes and the second-type nodes; The first type of preset vulnerability code is called to perform jump path code comparison on the first type of control flow jump path set to generate a first vulnerability identification result, and the symbolic execution module is called to perform path vulnerability identification on the second type of control flow jump path set to generate a second vulnerability identification result.
2. The method for automatically mining and identifying code vulnerabilities according to claim 1, characterized in that: In the source code control flow graph, vulnerability introduction risk and vulnerability code comparison confidence analysis are performed on multiple control flow nodes to generate multiple introduction risk indicators and multiple code comparison confidence indicators, including: Constructing a jump path for each of the multiple control flow nodes to generate multiple jump paths; Performing path complexity, code complexity and data flow risk analysis integration on the multiple jump paths to generate the multiple introduced risk indicators; The code structures and path types of the multiple jump paths are collected, historical vulnerability codes are mined based on the path types, discrimination comparisons are performed, and the multiple code comparison confidence indicators are generated.
3. The method for automatically mining and identifying code vulnerabilities according to claim 2, characterized in that: Performing path complexity, code complexity and data flow risk analysis integration on the multiple jump paths to generate the multiple introduced risk indicators, including: Build path analysis mechanism, code analysis mechanism and data flow analysis mechanism; Among them, the path analysis mechanism performs complexity conversion mapping by counting the path length and the number of path branches, wherein the path complexity is proportional to the path length and the number of branches; the code analysis mechanism performs code complexity mapping by identifying code redundancy and function call chain security risks; the data flow analysis mechanism performs data flow risk conversion mapping by performing input verification missing, data flow dependency analysis and sensitive data exposure analysis; Calling the path analysis mechanism, the code analysis mechanism and the data flow analysis mechanism to process the multiple jump paths respectively, and generating path complexity, code complexity and data flow risk corresponding to the multiple jump paths respectively; The path complexity, code complexity and data flow risk respectively corresponding to the multiple jump paths are weightedly integrated to generate the multiple introduced risk indicators.
4. The method for automatically mining and identifying code vulnerabilities according to claim 2, characterized in that: Collecting the code structure and path type of the multiple jump paths, mining historical vulnerability codes, performing discrimination comparison, and generating the multiple code comparison confidence indicators, including: Using the code structures and path types of the multiple jump paths as index elements, respectively, to retrieve historical vulnerability codes and historical security codes corresponding to various historical vulnerability types; Performing similarity analysis of vulnerability codes and security codes based on the historical vulnerability codes and the historical security codes, and generating multiple similarity indicators corresponding to the multiple jump paths; The multiple similarity indicators are respectively subtracted from 1 to generate the multiple code comparison confidence indicators.
5. The method for automatically mining and identifying code vulnerabilities according to claim 1, characterized in that: Based on the multiple introduction risk indicators and the multiple code comparison confidence indicators, the multiple control flow nodes are divided into first-category nodes and second-category nodes, including: Based on the confidence index threshold and the risk index threshold, the multiple control flow nodes are classified to generate the first type of nodes and the second type of nodes, wherein the first type of nodes are control flow nodes whose introduced risk index is less than the risk index threshold and whose code comparison confidence index is greater than the confidence index threshold; Wherein, when classification is performed, a classification conflict between the confidence indicator threshold and the risk indicator threshold occurs, and the classification priority of the risk indicator threshold is greater than the classification priority of the risk indicator threshold.
6. The method for automatically mining and identifying code vulnerabilities according to claim 1, characterized in that: Constructing a first-type control flow jump path set and a second-type control flow jump path set for the first-type nodes and the second-type nodes, including: Marking the first type of nodes and the second type of nodes in the source code control flow graph to generate a marked control flow graph; In the marked control flow graph, the first type of control flow jump path set is generated according to the coverage path of the first type of marked nodes, and the second type of control flow jump path set is generated according to the coverage path of the second type of marked nodes.
7. The method for automatically mining and identifying code vulnerabilities according to claim 1, characterized in that: The first type of preset vulnerability code is a historical vulnerability code retrieved based on the path type and path code structure in the first type of control flow jump path set; The symbolic execution module is a symbolic execution engine built based on a preset symbolic execution tool.
8. The method for automatically mining and identifying code vulnerabilities according to claim 1, characterized in that: Collect the source code files of the target software and build the source code control flow graph, including: Convert the source code file into an abstract syntax tree; Based on the abstract syntax tree, control flow structure recognition is performed, basic blocks are divided, the basic blocks are used as control flow nodes, and the source code control flow graph is generated according to the jump relationship of the basic blocks.
9. Code vulnerability automatic mining and identification system, characterized by: For implementing the method for automatically mining and identifying code vulnerabilities according to any one of claims 1 to 8, the system comprises: Build a control flow graph module, which is used to collect the source code files of the target software and build a source code control flow graph; A confidence analysis module, used to perform vulnerability introduction risk and vulnerability code comparison confidence analysis on multiple control flow nodes in the source code control flow graph, and generate multiple introduction risk indicators and multiple code comparison confidence indicators; A division module, configured to divide the plurality of control flow nodes into first-category nodes and second-category nodes based on the plurality of introduction risk indicators and the plurality of code comparison confidence indicators; A jump path construction module is used to construct a first-type control flow jump path set and a second-type control flow jump path set for the first-type nodes and the second-type nodes; The comparison module is used to call the first type of preset vulnerability code to perform jump path code comparison on the first type of control flow jump path set to generate a first vulnerability identification result, and call the symbolic execution module to perform path vulnerability identification on the second type of control flow jump path set to generate a second vulnerability identification result.
10. An electronic device, characterized in that: The electronic device comprises: A memory for storing executable instructions; A processor, for implementing the method for automatically mining and identifying code vulnerabilities according to any one of claims 1 to 8 when executing executable instructions stored in the memory.
Citation Information
Patent Citations
Static-analysis-assisted symbolic execution vulnerability detection method
CN104794401A
Web vulnerability detection method based on fine-grained static stain analysis and symbolic execution
CN111695119A
Smart contract security detection method and related device
CN116861443A
Symbol execution intelligent contract vulnerability detection method, system and device
CN116933267A
Supply chain security risk identification method and system
CN118427829A
Cited By
Electronic warranty full-process management and control method and system
CN120746489A