Security vulnerability detection method and device, computer program product and storage medium

By constructing a structured graph of the code and generating a comprehensive vector input detection model, combined with contextual features and prompt words, the problems of missed detection and false alarm rates in code security vulnerability detection are solved, achieving more efficient vulnerability detection.

CN120805148AInactive Publication Date: 2025-10-17INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511278519.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-10-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology for detecting code security vulnerabilities has high missed detection and false alarm rates, and lacks effective detection methods.

Method used

By constructing a structured graph of the code to be tested, generating text encoding vectors and structured encoding vectors, and integrating them into a comprehensive vector input pre-trained detection model for security vulnerability detection, the reference information is set by combining project-level context features and prompt words to improve detection accuracy.

Benefits of technology

It reduces the missed detection rate and false alarm rate of code security vulnerability detection, improves the accuracy and efficiency of detection, and supports the identification of cross-function and cross-module vulnerabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805148A_ABST
    Figure CN120805148A_ABST
Patent Text Reader

Abstract

The invention discloses a security vulnerability detection method and device, a computer program product and a storage medium, and belongs to the field of code detection.A structured graph of a to-be-detected code can be constructed according to structured information of the to-be-detected code, and then the to-be-detected code is coded to obtain a text coding vector; and coding the structured graph of the to-be-detected code to obtain a structured coding vector, and finally inputting the comprehensive vector into a pre-trained detection model so as to obtain a security vulnerability detection result. The problem of how to reduce the omission ratio and the false alarm rate of security vulnerability detection of codes is solved, and the beneficial effects that security vulnerability detection is performed from two dimensions of semantics of code texts and semantics of structured information, and the omission ratio and the false alarm rate of security vulnerability detection of codes are reduced are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of code detection, and in particular, to a security vulnerability detection method, device, computer program product and storage medium. BACKGROUND

[0002] Security vulnerability detection of code is a core work in the field of software security, especially as the scale of software system continues to expand and the complexity of application program increases, security vulnerabilities in the code are more concealed and diversified, and there is a lack of a mature security vulnerability detection method in the related art, and the omission rate and false positive rate are high when detecting security vulnerabilities of code.

[0003] Therefore, how to provide a solution to the above technical problems is a problem that those skilled in the art need to solve at present. SUMMARY

[0004] The purpose of the present application is to provide a security vulnerability detection method, device, computer program product and storage medium to at least solve the problem of high omission rate and false positive rate when detecting security vulnerabilities of code in the related art.

[0005] To solve the above technical problems, the present application provides a security vulnerability detection method, comprising: According to the structured information of the to-be-detected code, a structured graph of the to-be-detected code is constructed; the structured graph is used to represent the structured information of the to-be-detected code; The to-be-detected code is encoded to obtain a text encoding vector, and the text encoding vector is used to represent the semantic features of the text of the to-be-detected code; The structured graph of the to-be-detected code is encoded to obtain a structured encoding vector, and the structured encoding vector is used to represent the semantic features of the structured information of the to-be-detected code; An integrated vector formed by fusing the text encoding vector and the structured encoding vector is input into a pre-trained detection model, so as to obtain a detection result of security vulnerability.

[0006] The present application also provides a security vulnerability detection device, comprising: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the security vulnerability detection method as described above.

[0007] The present application also provides a computer program product comprising computer programs / instructions, which, when executed by a processor, implement the steps of the security vulnerability detection method as described above.

[0008] The application further provides a computer readable storage medium, wherein a computer program is stored on the computer readable storage medium, and the computer program is executed by a processor to implement the steps of the security vulnerability detection method.

[0009] According to the application, the semantic features of the code text can reflect relatively surface security vulnerabilities, and the structured information of the code can reflect deep security vulnerabilities. The application encodes the text encoding vector (representing the semantic features of the text of the code to be tested) and the structured encoding vector (representing the semantic features of the structured information of the code to be tested) of the code to be tested, and inputs the comprehensive vector obtained by fusing the two vectors into a detection model to perform security vulnerability detection. Therefore, the problem of how to reduce the missed detection rate and the false positive rate of security vulnerability detection of the code can be solved, and the technical effect of detecting security vulnerabilities from two dimensions of the semantic features of the code text and the semantic features of the structured information can be achieved, and the missed detection rate and the false positive rate of security vulnerability detection of the code can be reduced.

[0010] The application further provides a security vulnerability detection device, a computer program product and a storage medium, which have the same beneficial effects as the security vulnerability detection method. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the application, the related technologies and the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor. Figure 1 A flowchart of a security vulnerability detection method provided by the application; Figure 2 A flowchart of another security vulnerability detection method provided by the application; Figure 3 A structural diagram of a security vulnerability detection device provided by the application; Figure 4 A structural diagram of a computer readable storage medium provided by the application. DETAILED DESCRIPTION

[0012] The technical solutions in the embodiments of the application will be described clearly and completely in combination with the drawings in the embodiments of the application. Obviously, the described embodiments are only some embodiments of the application, but not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.

[0013] It should be noted that in the description of the present application, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0014] Reference is made to Figure 1 , Figure 1 A flowchart of a security vulnerability detection method provided by the present application is shown. The security vulnerability detection method comprises: S101: constructing a structured graph of the to-be-tested code according to structured information of the to-be-tested code; the structured graph is used to represent the structured information of the to-be-tested code; Specifically, considering the technical problems in the background art, and in combination with the consideration that the semantic features of the code text can reflect relatively surface security vulnerabilities, while the structured information of the code can reflect deep security vulnerabilities, in the embodiments of the present application, the semantic features of the code text and the structured information of the to-be-tested code are comprehensively considered for security vulnerability detection, and the structured information capable of reflecting deeper security vulnerabilities is a more important part, and the structured information can be more accurately represented by a graph structure. Therefore, in this step, a structured graph of the to-be-tested code is first constructed according to the structured information of the to-be-tested code, and is used as a data basis for subsequent steps.

[0015] S102: obtaining a text encoding vector by encoding the to-be-tested code; the text encoding vector is used to represent the semantic features of the text of the to-be-tested code; Specifically, in this step, a text encoding vector is obtained by encoding the to-be-tested code, and the text encoding vector is used to represent the semantic features of the text of the to-be-tested code. The text encoding vector can be used as a data basis for subsequent steps.

[0016] The semantic features of the code text can reflect some relatively surface security vulnerabilities exhibited by the text itself, and do not involve the syntax and control flow mechanisms of the code, and cannot help the model to identify security vulnerabilities that need to be analyzed across functions or modules.

[0017] S103: obtaining a structured encoding vector by encoding the structured graph of the to-be-tested code; the structured encoding vector is used to represent the semantic features of the structured information of the to-be-tested code; Specifically, in this step, the structured graph of the to-be-tested code obtained in the foregoing step is encoded to obtain a structured encoding vector, so as to represent the semantic features of the structured information of the to-be-tested code by the structured encoding vector, and the structured encoding vector is used as a data basis for subsequent steps for security vulnerability detection.

[0018] S104: input the integrated vector obtained by fusing the text coding vector and the structured coding vector into the pre-trained detection model to obtain a detection result of the security vulnerability.

[0019] Specifically, after obtaining the text coding vector and the structured coding vector, the text coding vector and the structured coding vector can be fused into an integrated vector, and then the integrated vector is input into the pre-trained detection model to obtain a detection result of the security vulnerability. The detection result is obtained by analyzing the security vulnerability from two dimensions of the text semantic features of the code text and the semantic features of the structured information, and the accuracy is high.

[0020] Among them, the application field of the security vulnerability detection method in the embodiment of the present application can be various, for example, it can include: Supply chain security detection: the present technology can be extended to analyze the vulnerability propagation path in third-party dependencies and components, such as detecting potential malicious code risks in npm (Node Package Manager, Node package manager), pip (Pip Installs Packages, python package installation tool) and other dependent packages.

[0021] Smart contract and blockchain security: applied to multi-modal analysis of smart contract code, detecting related vulnerabilities such as fund flow and permission management; using the context module to obtain the transaction history of the blockchain as supplementary information.

[0022] In addition, the security vulnerability detection method in the embodiment of the present application can be integrated into the pipeline of code submission, thereby improving the detection efficiency, for example, when the code is submitted, the security vulnerability detection of the submitted code can be automatically performed to realize the security interception in the development process.

[0023] The present application provides a security vulnerability detection method, which considers that the semantic features of the code text can reflect the relatively surface security vulnerability, and the structured information of the code can reflect the deep security vulnerability. Therefore, in the present application, the structured graph of the to-be-tested code can be constructed according to the structured information of the to-be-tested code, and then the to-be-tested code is coded to obtain a text coding vector, and the structured graph of the to-be-tested code is coded to obtain a structured coding vector. Finally, the integrated vector (obtained by fusing the text coding vector and the structured coding vector) can be input into the pre-trained detection model to obtain a detection result of the security vulnerability. Since the integrated vector input into the detection model fuses the text coding vector and the structured coding vector, the security vulnerability can be detected from two dimensions of the semantic of the code text and the semantic of the structured information, thereby reducing the missed detection rate and the false positive rate of the security vulnerability detection of the code.

[0024] On the basis of the above embodiment: As an optional embodiment, constructing the structured graph of the to-be-tested code according to the structured information of the to-be-tested code comprises: performing syntax analysis on the to-be-tested code to obtain an abstract syntax tree of the to-be-tested code, the abstract syntax tree comprising a plurality of tree nodes, parent-child edge relationships and node features of the tree nodes, the tree nodes representing elements in the to-be-tested code; constructing the structured graph of the to-be-tested code based on the abstract syntax tree.

[0025] Specifically, considering that the structured graph needs to accurately reflect the syntax structure basis of the code, and the AST (Abstract Syntax Tree) is a direct representation of the syntax structure of the code, containing tree nodes, parent-child edge relationships and node features, which can provide core support for subsequent supplement of dependency relationships and construction of a complete structured graph, in the embodiment of the present application, the to-be-tested code can be first subjected to syntax analysis to obtain an abstract syntax tree of the to-be-tested code, and then based on the abstract syntax tree, the structured graph of the to-be-tested code is constructed, the abstract syntax tree can provide an accurate syntax structure basis for the structured graph, ensuring the accuracy of subsequent structured information extraction, laying a reliable structural foundation for deep vulnerability detection, thereby further improving the accuracy of security vulnerability detection.

[0026] Specifically, the syntax analysis on the to-be-tested code to obtain the abstract syntax tree of the to-be-tested code can comprise: first, performing a preprocessing operation on the to-be-tested code to remove comments, white spaces and complete macro expansion, simplifying the subsequent analysis process; then, calling a corresponding parser for syntax analysis according to different programming languages to obtain the AST. For example, performing a preprocessing operation on the to-be-tested code to remove comments, white spaces and complete macro expansion, simplifying the subsequent analysis process; calling a corresponding parser for syntax analysis according to different programming languages, for example, for C / C++, Python code, loading a language syntax package corresponding to the Tree-sitter (tree parser) parser for analysis, for Java code, calling the ASTParser (Abstract Syntax Tree Parser) class of the Eclipse JDT (Eclipse Java Development Tools), for C code, using the pycparser library, and extracting the AST of the to-be-tested code through the above analysis tools, the AST containing a plurality of tree nodes (representing elements such as variables, statements and functions in the code), parent-child edge relationships between the tree nodes, and auxiliary features (such as function names, variable names and API (Application Programming Interface) call information) of the tree nodes.

[0027] As an optional embodiment, based on the abstract syntax tree, constructing the structured graph of the to-be-tested code comprises: generating a program dependency graph of the to-be-tested code, the program dependency graph being used to describe dependency relationships between elements in the to-be-tested code; constructing edges between tree nodes in the abstract syntax tree according to the dependency relationships between the elements in the program dependency graph, and connecting the node list of the abstract syntax tree and the edges as the structured graph of the to-be-tested code.

[0028] Specifically, considering that the AST can reflect the syntax structure (parent-child edge relationship) of the code, but cannot reflect the control, data and call dependency relationships between code elements, and these dependency relationships are the key to identifying deep security vulnerabilities across statements and functions, therefore, in the embodiment of the present application, a program dependency graph of the to-be-tested code can be generated based on the AST, the program dependency graph being used to describe dependency relationships between elements in the to-be-tested code, then edges between tree nodes in the abstract syntax tree are constructed according to the dependency relationships between the elements in the program dependency graph, and the node list of the abstract syntax tree and the edges are connected as the structured graph of the to-be-tested code, that is, the dependency relationships can be supplemented by generating the program dependency graph, and then the AST is fused to construct the structured graph, so that the structured graph contains syntax structure and related dependency relationships at the same time, can more comprehensively depict code semantics, and improves the recognition ability of complex deep vulnerabilities.

[0029] Of course, in addition to the above examples, "constructing the structured graph of the to-be-tested code based on the abstract syntax tree" can also be in other forms, which are not limited in the embodiment of the present application.

[0030] As an optional embodiment, generating the program dependency graph of the to-be-tested code comprises: generating a control flow graph of the to-be-tested code, the control flow graph being used to represent control relationships between elements in the to-be-tested code; generating a data flow graph of the to-be-tested code, the data flow graph being used to represent data flow conversion relationships between elements in the to-be-tested code; generating a function call graph of the to-be-tested code, the function call graph being used to represent function call relationships between elements in the to-be-tested code.

[0031] Specifically, in order to comprehensively depict the execution logic and element association of the code, considering that the control flow relationship determines the code execution path, the data flow relationship reflects the data flow track, and the function call relationship embodies the cross-function interaction, the three can jointly constitute the core dimension of the code semantic dependency, therefore, in the embodiment of the present application, the control flow graph, the data flow graph and the function call graph of the code to be tested can be generated, and through the program relationship graph of the three dimensions, the execution logic and element association of the code are comprehensively depicted, the program relationship graph generated in the embodiment of the present application can completely cover the control, data and call dependencies of the code, provide rich semantic information for the structured graph, support the detection of complex vulnerabilities across functions and modules, further improve the detection accuracy of security vulnerabilities, and reduce the false negative rate and the false positive rate.

[0032] The specific implementation of generating the three program relationship graphs is as follows: Generate a control flow graph (CFG): a program analysis tool (such as Clang Static Analyzer or Joern) is used to analyze the control flow of the code to be tested, the basic blocks (such as statement blocks, conditional branch blocks and loop blocks) in the code are taken as the nodes of the CFG, and directed edges are added between the corresponding basic block nodes according to the code execution order (such as sequential execution, if-else branch jump and for / while loop jump), so as to form the CFG, which can clearly represent the control execution relationship between elements in the code, for example, if the condition is met, jump to the node of the corresponding branch.

[0033] Generate a data flow graph (DFG): a data flow analysis tool (such as PyDriller or Joern) is used to track the life cycle of variables in the code to be tested, including the definition, assignment, reference and transmission process of variables, the definition point and use point of the variable are taken as the nodes of the DFG, and directed edges are added between the corresponding nodes according to the data flow path of the variable (such as variable A is assigned to variable B, and variable B is referenced in function C), so as to form the DFG, which represents the data transmission and use relationship between elements.

[0034] Generate a function call graph: a function call analysis tool (such as Call Graph Generator (CGG) or the call chain analysis function of the integrated development environment (IDE)) is used to sort out the call relationship of all functions in the code to be tested, each function is taken as a node of the function call graph, and if function F1 calls function F2, a directed edge is added between the F1 node and the F2 node, so as to form the function call graph, which can clearly represent the function call association between elements in the code, including direct call and indirect call relationship.

[0035] As an optional embodiment, encoding the to-be-tested code to obtain a text encoding vector comprises: Splicing the to-be-tested code and preset auxiliary texts thereof; Performing word segmentation on the spliced to-be-tested code and auxiliary texts thereof by a word segmenter to obtain a word segmentation sequence; Inputting the word segmentation sequence into a pre-trained language model to obtain a text encoding vector of the to-be-tested code.

[0036] Specifically, considering that the text semantics of the to-be-tested code not only contains code body text, but also closely relates to function signatures, annotations, documents and other auxiliary texts, single code body text encoding is easy to lose key semantic information, and the text semantic features can be more accurately extracted through word segmentation and a pre-trained language model, therefore, in the embodiment of the present application, the to-be-tested code and preset auxiliary texts thereof can be first spliced, then the spliced to-be-tested code and auxiliary texts thereof are subjected to word segmentation by a word segmenter to obtain a word segmentation sequence, and finally the word segmentation sequence is input into a pre-trained language model to obtain a text encoding vector of the to-be-tested code, which can fully integrate the semantics of code text and auxiliary information, and the generated text encoding vector has more comprehensive semantic representation capability, thereby improving the identification accuracy of surface vulnerabilities.

[0037] The auxiliary texts can be of various specific types, for example, can include function signatures, annotations, documents and the like, which are not limited in the embodiment of the present application.

[0038] Specifically, for example, after the to-be-tested code and preset auxiliary texts thereof are spliced, a word segmenter can be loaded to perform word segmentation on the spliced text to disassemble the text into Token sequences recognizable by the model, for example, the code statement "int add(int a, int b) {return a+b;}" is segmented into ["int", "add", "(", "int", "a", ",", "int", "b", ")", "{", "return", "a", "+", "b", ";", "}"]. Then, the word segmentation sequence is input into a pre-trained language model, for example, a CodeBERT (Code Bidirectional Encoder Representations from Transformers) language model, the model captures semantic associations between Tokens (the smallest meaning unit of word segmentation) through a self-attention mechanism to obtain a text encoding vector representing the text semantic features of the to-be-tested code; if the length of the word segmentation sequence exceeds the maximum input length of the model (for example, the maximum input length of CodeBERT is 512), a sliding window (for example, the window size is 512 and the step size is 256) or a truncation strategy is adopted to ensure that key semantic information is not lost.

[0039] Of course, in addition to the specific form, "encoding the to-be-tested code to obtain a text encoding vector" can also be in other forms, which are not limited herein.

[0040] As an optional embodiment, encoding the structured graph of the to-be-tested code to obtain a structured encoding vector comprises: Converting the structured graph of the to-be-tested code into a node feature matrix and an edge list through a graph computing library; Inputting the node feature matrix and the edge list into a multi-layer graph convolution network so as to obtain a structured encoding sub-vector of each tree node; Summarizing the structured encoding sub-vectors of each tree node into a structured encoding vector through a global pooling operation.

[0041] Specifically, considering that the structured graph exists in the form of nodes and edges and cannot be directly input into a model for encoding, it can be processed by a graph convolution network after being converted into a node feature matrix and an edge list format; and considering that a single-layer graph network is difficult to capture deep structural features of the code, information needs to be aggregated through a multi-layer network, and node-level features need to be summarized into graph-level features to adapt to subsequent fusion, therefore, in the embodiment of the present application, the structured graph of the to-be-tested code can be converted into a node feature matrix and an edge list through a graph computing library, then the node feature matrix and the edge list are input into a multi-layer graph convolution network so as to obtain a structured encoding sub-vector of each tree node, and finally the structured encoding sub-vectors of each tree node can be summarized into a structured encoding vector through a global pooling operation, the embodiment of the present application can convert the structured graph into a vector with deep semantic representation, accurately reflect the structural dependency relationship of the code, and provide high-quality structured features for subsequent multi-modal fusion, which is beneficial to further improving the precision of security vulnerability detection.

[0042] For example, the graph format is converted first: the node list and edge list of the structured graph are read, and the structured graph is converted into a format that can be processed by the model through the Data class of PyG (PyTorch Geometric, PyTorch graph neural network extension library) or the DGLGraph class of DGL (Deep Graph Library, deep graph learning library): the features of each node (such as node type, variable type, function name embedding) are integrated into a node feature matrix (dimension is [number of nodes × feature dimension]), and the connection relationship of the edge is converted into an edge list (such as the edge_index matrix in PyG, dimension is [2 × number of edges]). Next, a multi-layer graph convolutional network (GCN) encoding is performed: a graph network model consisting of three GCN (Graph Convolutional Network) layers is constructed. The first GCNConv layer takes as input a node feature matrix and an edge list, and calculates the first-layer node features by aggregating the features of neighboring nodes. The second GCNConv layer uses the output of the first layer as input to further capture broader neighborhood relationships. The third GCNConv layer outputs the final node-level structured encoder vector. Each layer is activated with a ReLU activation function and a dropout layer (with a dropout probability of 0.3) to prevent overfitting. Finally, a global pooling operation is performed. For example, a global average pooling strategy can be used to average the structured encoder vectors of all nodes according to their feature dimensions to obtain a vector of fixed dimension (e.g., 768 dimensions). This vector is the structured encoder vector that represents the semantic features of the structured information of the code under test. Alternatively, global max pooling or attention pooling (e.g., assigning higher weights to key nodes) can be used. The appropriate pooling method should be selected based on the actual detection scenario.

[0043] As an optional embodiment, before inputting the comprehensive vector formed by fusing the text encoding vector and the structured encoding vector into the pre-trained detection model, the security vulnerability detection method further includes: Obtain project-level context features of the software project to which the code under test belongs; Encode the item-level context features to obtain the item-level context vector; The pre-trained detection model that inputs the comprehensive vector formed by fusing the text encoding vector and the structured encoding vector includes: The comprehensive vector formed by fusing the text encoding vector, structured encoding vector, and item-level context vector is input into the pre-trained detection model.

[0044] Specifically, considering that the vulnerability risks of the same code in different project contexts may be different, relying only on the characteristics of the code itself may lead to insufficient detection pertinence. If the scenario information (such as dependent library version, running environment, and historical submission) of the software project to which the code belongs can be combined for security vulnerability detection of the code to be tested, the scenario information of the software project to which the code to be tested belongs can be referred to, and the security vulnerability detection accuracy can be further improved. Therefore, in the embodiment of the present application, the project-level context vector can be introduced to make the comprehensive vector contain scenario information, and the recognition accuracy of the detection model for vulnerabilities in a specific project scenario can be improved, especially for accurate detection of cross-function and cross-module vulnerabilities.

[0045] Specifically, the context information includes the dependent library version, running environment configuration, historical submission information, type system constraint, API usage mode, and the like of the project in which the current code is located. When obtaining the project-level context vector, information extraction can be performed from at least one of the version control system, project metadata, and project configuration file of the software project, and integrated into the project-level context feature. For example, relevant context such as function call chain, variable life cycle, known patch description, and security bulletin text can be extracted from the version control system and project metadata. The dependent library and version information, running environment parameters (such as operating system and framework version), code submission history, and known security bulletin, and the like can be obtained from the project configuration file and version control system. The above information is integrated into the project-level context feature. Then, the project-level context feature can be encoded by a preset context encoding model to obtain the project-level context vector. For example, the Transformer (Transformer model, Transformer architecture model) can be used as the context encoding model, the project-level context feature is converted into a text sequence according to a preset format (such as “dependent library: [library name and version]; running environment: [environment parameters]; historical submission: [key change record]”), and the text sequence is input into the Transformer model after being segmented by a word segmenter. The [CLS] position vector output by the model is the project-level context vector.

[0046] The fusion of the text encoding vector, the structured encoding vector, and the project-level context vector into the comprehensive vector can include: The text encoding vector, the structured encoding vector, and the project-level context vector are normalized, and then the vectors are fused into the comprehensive vector by vector splicing, weighted summation, or “cross-modal attention mechanism fusion”.

[0047] As an optional embodiment, the comprehensive vector fused from the text encoding vector and the structured encoding vector is input into the pre-trained detection model, which includes: According to the prompt word setting reference information, the prompt word for inputting into the pre-trained detection model is generated by means of manual setting and / or model generation; the prompt word setting reference information includes the function description of the to-be-tested code, the project-level context features of the software project to which the code belongs, the detection instruction, and a preset prompt word template; The text coding vector and the structured coding vector are combined into a comprehensive vector; The comprehensive vector, the prompt word, and the to-be-tested code are input into the pre-trained detection model.

[0048] Specifically, considering that the inference of the pre-trained detection model (especially the LLM (Large Language Model)) needs to be explicitly guided, only inputting the comprehensive vector may lead to insufficient understanding of the detection target, code function, and project scenario by the model, and the prompt word can integrate the function description, context, and detection instruction to provide a clear inference direction for the model. Therefore, in the embodiments of the present application, the prompt word for inputting into the pre-trained detection model can be generated according to the prompt word setting reference information (including the function description of the to-be-tested code, the project-level context features of the software project to which the code belongs, the detection instruction, and the preset prompt word template) by means of manual setting and / or model generation, so as to input the prompt word together with the comprehensive vector and the to-be-tested code into the pre-trained detection model. The embodiments of the present application can enhance the understanding of the detection task by the model through the prompt word, make the detection process more targeted, and further improve the accuracy and reliability of the detection result.

[0049] As an optional embodiment, according to the prompt word setting reference information, the prompt word for inputting into the pre-trained detection model is generated by means of manual setting and / or model generation, including: The prompt word setting reference information is pushed to the user, so that the user sets the prompt word based on the preset prompt word template according to the function description of the to-be-tested code, the project-level context features of the software project to which the code belongs, and the detection instruction; The prompt word set by the user and the prompt word setting reference information are input into the pre-trained prompt word generation model, so that the prompt word generation model optimizes the prompt word set by the user.

[0050] Specifically, considering that manual setting of the prompt word can be limited by user experience, resulting in lack of pertinence of the prompt word or missing of key information, while model optimization can supplement details and correct bias based on reference information, and combining the advantages of manual and model can generate a prompt word of higher quality, therefore in the embodiment of the application, the prompt word setting reference information can be first pushed to the user, so that the user sets the prompt word based on the preset prompt word template, according to the function description of the to-be-tested code, the project-level context features of the software project to which the code belongs, and the detection instruction; then the user-set prompt word and the prompt word setting reference information are jointly input into the pre-trained prompt word generation model, so that the prompt word generation model optimizes the user-set prompt word. The prompt word generated in the embodiment of the application has both the scene adaptability of artificial experience and the information integrity of the model, and can more accurately guide the detection model reasoning, thereby improving the accuracy of vulnerability detection.

[0051] Specifically, for example, in the stage of manually setting the prompt word, the prompt word setting reference information (the function description of the to-be-tested code "the amount calculation function in the order payment module", the project-level context features of the software project to which the code belongs "the project adopts Spring Cloud microservice architecture and uses Redis to cache order data", the detection instruction "detect whether there is an integer overflow, cache penetration vulnerability" and the preset prompt word template "analyze code: function-[function description], context-[context feature], need to detect-[detection instruction], output vulnerability situation") can be pushed to the user; the user sets the prompt word based on the template, combined with his own understanding of the code, for example, "analyze code: function-the amount calculation function in the order payment module, context-the project adopts Spring Cloud microservice architecture and uses Redis to cache order data, need to detect-whether there is an integer overflow, cache penetration vulnerability, output vulnerability situation". In the stage of model optimizing the prompt word, the user-set prompt word and the prompt word setting reference information (function description, context feature, detection instruction, and preset prompt word template) can be jointly input into the pre-trained prompt word generation model (such as a model based on CodeLlama-7B fine-tuning). The model analyzes the prompt word setting reference information, identifies the optimization points (such as supplementing "the amount calculation needs to involve multi-precision operation" and "the Redis cache key value needs to be checked for legality") in the user prompt word, optimizes the user prompt word, and outputs the optimized prompt word: "analyze code: function-the amount calculation function in the order payment module (involving order amount accumulation and discount calculation), context-the project adopts Spring Cloud microservice architecture and uses Redis to cache order data (the cache key is order ID), need to detect-whether there is an integer overflow (focus on the amount accumulation process) and cache penetration vulnerability (focus on cache key verification), output vulnerability situation and key code location".

[0052] Of course, in addition to the specific form, "setting reference information according to the prompt word, generating the prompt word for input into the pre-trained detection model through manual setting and / or model generation" can also be other specific forms, such as generating the prompt word for input into the pre-trained detection model through manual setting or model generation alone, etc. The present embodiment does not limit here.

[0053] As an optional embodiment, after the integrated vector fused by the text encoding vector and the structured encoding vector is input into the pre-trained detection model, the security vulnerability detection method further includes: Based on the risk score of each function in the code to be tested output by the detection model, the threshold method is used to determine each function with a security vulnerability as a target function; Through the Common Vulnerability Scoring System, the risk level of the security vulnerability of each target function is evaluated and determined; The code location of each target function and the risk level of the security vulnerability are pushed to the user.

[0054] Specifically, considering that the detection model can give a risk score for each function in the code to be tested, and the threshold method can efficiently and accurately screen out functions with security vulnerabilities through the risk score, and by dividing the risk level of each target function with a security vulnerability, the user can prioritize handling high-risk security vulnerabilities, therefore, in the present embodiment, based on the risk score of each function in the code to be tested output by the detection model, the threshold method is used to determine each function with a security vulnerability as a target function, and then through the Common Vulnerability Scoring System, the risk level of the security vulnerability of each target function is evaluated and determined, and finally the code location of each target function and the risk level of the security vulnerability are pushed to the user. This scheme can efficiently and accurately screen out real vulnerabilities, clearly define the risk level and location of the vulnerability, help the user efficiently locate and prioritize handling high-risk vulnerabilities, and improve the vulnerability handling efficiency.

[0055] The risk level of the security vulnerability of the target function can be determined through CVSS (Common Vulnerability Scoring System) evaluation. For example, the risk scores of each function in the code under test output by the detection model can be obtained (for example, the risk score of function A is 0.85, and the risk score of function B is 0.42). A threshold method is used to set a screening threshold (for example, a threshold of 0.7 calibrated based on historical detection data). The function (for example, function A) with a risk score higher than the threshold is determined as the target function with a security vulnerability. Then, the risk level of the security vulnerability of each target function is evaluated through CVSS. The scores are given in dimensions such as attack vector (for example, network, local), attack complexity (for example, low, high), permission requirement (for example, no permission, need for administrator permission), and influence range (for example, only the current function, cross-module). According to the total score, the risk level is determined (for example, 9.0-10.0 is high-risk, 7.0-8.9 is medium-risk, 4.0-6.9 is low-risk, and 0.1-3.9 is information-level). For example, the vulnerability score of function A is 9.2, which is determined as high-risk.

[0056] Specifically, the code positions of the target functions and the risk levels of the security vulnerabilities can be pushed to the user in the form of a visual report or a system alarm, and the code positions of the target functions and the risk levels of the security vulnerabilities are recorded in an audit log.

[0057] As an optional embodiment, based on the risk scores of each function in the code under test output by the detection model, after determining each function with a security vulnerability as a target function through a threshold method, the security vulnerability detection method further comprises: determining the vulnerability type of the security vulnerability of each target function; generating a repair suggestion for the security vulnerability of each target function through a vulnerability knowledge base or a pre-trained repair suggestion generation model; pushing the vulnerability type and the repair suggestion of the security vulnerability of each target function of the code under test to the user.

[0058] Specifically, considering that on the basis of determining the code positions of the target functions and the risk levels of the security vulnerabilities, the user may also have a demand to understand the vulnerability type of the security vulnerability, and that the efficiency of vulnerability repair can be improved and the repair cost can be reduced by pushing the corresponding repair suggestion, in the embodiment of the present application, the vulnerability type of the security vulnerability of each target function can be determined. A repair suggestion for the security vulnerability of each target function is generated through a vulnerability knowledge base or a pre-trained repair suggestion generation model. The vulnerability type and the repair suggestion of the security vulnerability of each target function of the code under test are pushed to the user, which can help the user quickly understand the nature of the vulnerability and efficiently complete the repair, thereby reducing the technical threshold and time cost of vulnerability repair.

[0059] The vulnerability type of the security vulnerability of the target function can be determined by comparing the vulnerability features of the security vulnerability of the target function with the vulnerability type features in the related database, for example, the vulnerability features of the security vulnerability of the target function can be compared with the vulnerability type features (for example, the vulnerability feature of integer overflow is "the result of integer operation exceeds the range of data type representation", and the vulnerability feature of cache penetration is "frequently querying the database using a non-existent cache key") in the CVE (Common Vulnerabilities and Exposures) database and the CWE (Common Weakness Enumeration) library to determine the vulnerability type, for example, the vulnerability type of function A is "integer overflow (CWE-190)".

[0060] Specifically, the repair suggestion can be generated for each target function through a vulnerability knowledge base or a pre-trained repair suggestion generation model (such as a CodeT5 fine-tuned model), for example, for an integer overflow vulnerability, the suggestion is "use the BigDecimal class instead of int and long for amount calculation to avoid integer overflow; check the input value range before operation to ensure that it does not exceed the upper limit of the data type"; for the cache penetration vulnerability, the suggestion is "implement the cache null mechanism, cache a null value for a non-existent cache key and set a short expiration time; use a Bloom filter to filter non-existent cache keys to reduce database queries".

[0061] As an optional embodiment, the security vulnerability detection method further comprises: obtaining historical code labeled with actual security vulnerabilities and missed or misdetected by the detection model and a comprehensive vector thereof; incrementally updating the detection model through the historical code and the comprehensive vector thereof.

[0062] Specifically, in order to better illustrate the embodiments of the present application, please refer to Figure 2 , Figure 2 The flowchart of another security vulnerability detection method provided by the present application is as follows: first, input the code to be tested, generate an abstract syntax tree after processing, further construct a structured graph based on the abstract syntax tree, perform graph coding on the structured graph to obtain a structured coding vector; at the same time, perform text semantic coding on the code to obtain a text coding vector, and perform project-level context coding in combination with the project scenario to obtain a project-level context vector; then, the graph coding, text semantic coding and project-level context coding results are fused to obtain a comprehensive vector; then, the reference information is generated based on the fusion vector and the prompt word setting, the prompt word is used to drive the detection model reasoning, and the detection model can be optimized by incremental updating; finally, after the detection model reasoning, the risk level assessment and result output are performed, thereby completing the vulnerability detection process.

[0063] Specifically, considering that the detection model will encounter new vulnerability types or miss some vulnerabilities or misjudge some vulnerabilities after deployment, the fixed model cannot adapt to the changes in vulnerabilities, and the detection performance of the detection model can be enhanced through the "historical code missed or misjudged by the detection model", therefore, in the embodiment of the present application, the historical code labeled with actual security vulnerabilities and missed or misjudged by the detection model and its comprehensive vector can be obtained, and then the detection model is incrementally updated through the historical code and its comprehensive vector. This scheme can update the model using the missed or misjudged historical code data, so that the model continuously learns new vulnerability features and has self-adaptive evolution ability, and long-term maintains low false negative and low false positive detection performance.

[0064] Among them, the "historical code labeled with actual security vulnerabilities and missed or misjudged by the detection model" can be obtained by expert review or comparison of static analysis results. Experts can review "user-reported missed or misjudged code" or "detection results obtained by the detection model", so as to obtain "historical code labeled with actual security vulnerabilities and missed or misjudged by the detection model".

[0065] Specifically, first, historical data can be obtained: collect historical code labeled with actual security vulnerabilities and missed or misjudged by the detection model (such as 100 integer overflow vulnerability codes and 50 SQL (Structured Query Language, structured query language) injection vulnerability codes missed in January-March 2024), and extract the comprehensive vector (vector fused after text encoding vector and structured encoding vector) corresponding to these historical codes, and label the actual type and risk level of the vulnerability, forming an incremental training data set. Then perform incremental updating: use the incremental training data set to supervise the fine-tuning of the detection model, and use optimization algorithms (such as AdamW (Adaptive Moment Estimation with Weight Decay, adaptive moment estimation with weight decay)) and learning rate decay strategies to prevent overfitting during training.

[0066] In addition, as an optional embodiment, the security vulnerability detection method further comprises: In response to the performance evaluation instruction, evaluate the detection performance of the detection model; When the detection performance is not up to standard, optimize the performance of the detection model.

[0067] Specifically, considering that the detection model is prevented from reducing the detection performance due to various factors, in the embodiment of the present application, the detection performance of the detection model can be evaluated in response to the performance evaluation instruction; when the detection performance is not up to standard, the performance of the detection model is optimized, that is, the detection performance of the detection model can be actively evaluated and the performance of the detection model can be optimized when needed, so as to ensure that the detection model maintains optimal detection performance for a long time.

[0068] Specifically, in response to the performance evaluation instruction, the evaluation of the detection performance of the detection model includes: In response to the performance evaluation instruction, the historical detection data (including the code to be tested, the comprehensive vector, and the detection result) of the detection model within a preset time window is extracted, and a multi-dimensional performance evaluation index is determined, including the missed detection rate (the number of missed vulnerabilities and the total number of actual vulnerabilities), the false positive rate (the number of false positive vulnerabilities and the total number of detection results), the single code segment reasoning time consumption, and the detection accuracy of the preset high-risk vulnerability type (such as buffer overflow and SQL injection), and based on the industry security detection benchmark (such as the CVE (Common Vulnerabilities and Exposures, Common Vulnerabilities and Exposures) detection coverage standard) or the historical optimal performance of the detection model, the threshold of each index is set (for example, the missed detection rate is less than or equal to 5%, the false positive rate is less than or equal to 8%, the reasoning time consumption is less than or equal to 100 ms / segment, and the detection accuracy of the high-risk vulnerability is greater than or equal to 95%); by comparing each dimension index with the corresponding threshold, it is determined whether the detection performance of the detection model meets the standard.

[0069] Among them, the performance evaluation instruction can be of various types, for example, generated by a user through a man-machine interaction device, or periodically generated by a program, and the embodiments of the present application do not limit it.

[0070] For the structure of the security vulnerability detection device provided by the embodiments of the present application, please refer to the security vulnerability detection method described above. Figure 3 , Figure 3 The security vulnerability detection device provided by the embodiments of the present application includes: The memory 31 is used to store the computer program. The processor 32 is used to execute the computer program to realize the steps of the security vulnerability detection method in the foregoing embodiments.

[0071] For the security vulnerability detection device provided by the embodiments of the present application, please refer to the foregoing embodiments of the security vulnerability detection method, and the embodiments of the present application will not be repeated here.

[0072] The present application also provides a computer program product, including computer programs / instructions, which are executed by a processor to realize the steps of the security vulnerability detection method in the foregoing embodiments.

[0073] For the computer program product provided by the embodiments of the present application, please refer to the foregoing embodiments of the security vulnerability detection method, and the embodiments of the present application will not be repeated here.

[0074] For the structure of the security vulnerability detection device provided by the embodiments of the present application, please refer to the security vulnerability detection method described above. Figure 4 , Figure 4A structural diagram of a computer readable storage medium provided by the present application is shown in FIG. 41. The computer readable storage medium 41 stores a computer program 42. The computer program 42 is executed by a processor to implement the steps of the security vulnerability detection method according to the foregoing embodiments.

[0075] The computer readable storage medium provided by the embodiments of the present application is described above with reference to the foregoing embodiments of the security vulnerability detection method. The embodiments of the present application will not be described herein again.

[0076] Those skilled in the art will further appreciate that the steps of the units and algorithms described in connection with the examples disclosed herein can be embodied in electronic hardware, computer software, or combinations of both. To clearly illustrate the interchangeability of hardware and software, and to avoid obscuring the disclosure, the various illustrative components, units, and steps have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled persons can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0077] The security vulnerability detection method, device, computer program product, and storage medium provided by the present application are described in detail above. The principles and implementation manners of the present application are described herein by applying specific examples. The above descriptions of the embodiments are only used to help understand the method of the present application and its core idea. It should be noted that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application. These improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A security vulnerability detection method, characterized in that: include: Constructing a structured graph of the code to be tested based on the structured information of the code to be tested; the structured graph is used to represent the structured information of the code to be tested; Encode the code to be tested to obtain a text encoding vector, which is used to represent the semantic features of the text of the code to be tested; Encoding the structured graph of the code to be tested to obtain a structured coding vector, which is used to represent the semantic features of the structured information of the code to be tested; The comprehensive vector formed by fusing the text encoding vector and the structured encoding vector is input into the pre-trained detection model to obtain the detection results of the security vulnerability.

2. The security vulnerability detection method according to claim 1, characterized in that: The step of constructing a structured graph of the code to be tested according to the structured information of the code to be tested includes: Performing syntax parsing on the code to be tested to obtain an abstract syntax tree of the code to be tested, wherein the abstract syntax tree includes a plurality of tree nodes, parent-child edge relationships, and node features of the tree nodes, and the tree nodes represent elements in the code to be tested; Based on the abstract syntax tree, a structured graph of the code to be tested is constructed.

3. The security vulnerability detection method according to claim 2, characterized in that: The step of constructing a structured graph of the code to be tested based on the abstract syntax tree includes: Generate a program relationship diagram of the code to be tested, wherein the program relationship diagram is used to describe the dependency relationship between elements in the code to be tested; According to the dependency relationship between elements in the program relationship graph, edges between tree nodes are constructed in the abstract syntax tree, and the node list of the abstract syntax tree is connected with the edges as a structured graph of the code to be tested.

4. The security vulnerability detection method according to claim 3, characterized in that: The program relationship diagram for generating the code to be tested includes: Generate a control flow graph of the code to be tested, which is used to represent the control relationship between elements in the code to be tested; Generate a data flow diagram for the code under test. The data flow diagram is used to represent the data flow relationship between elements in the code under test. Generate a function call graph of the code to be tested, where the function call graph is used to represent the function call relationship between elements in the code to be tested.

5. The security vulnerability detection method according to claim 1, wherein: The encoding of the code to be tested to obtain a text encoding vector includes: Splice the code to be tested and its preset auxiliary text; The concatenated code to be tested and its auxiliary text are segmented by a word segmenter to obtain a word segmentation sequence; The word segmentation sequence is input into the pre-trained language model to obtain the text encoding vector of the code to be tested.

6. The security vulnerability detection method according to claim 2, characterized in that: The encoding of the structured graph of the code to be tested to obtain a structured coding vector includes: Use the graph computing library to convert the structured graph of the code to be tested into a node feature matrix and edge list; Inputting the node feature matrix and edge list into a multi-layer graph convolutional network to obtain a structured encoding sub-vector for each tree node; The structured encoding sub-vectors of each tree node are aggregated into a structured encoding vector through a global pooling operation.

7. The security vulnerability detection method according to claim 1, characterized in that: Before inputting the comprehensive vector formed by fusing the text encoding vector and the structured encoding vector into the pre-trained detection model, the security vulnerability detection method further includes: Obtain project-level context features of the software project to which the code under test belongs; Encode the item-level context features to obtain the item-level context vector; The method of inputting the comprehensive vector formed by fusing the text encoding vector and the structured encoding vector into the pre-trained detection model includes: The comprehensive vector formed by fusing the text encoding vector, structured encoding vector, and item-level context vector is input into the pre-trained detection model.

8. The security vulnerability detection method according to claim 1, wherein: The method of inputting the comprehensive vector formed by fusing the text encoding vector and the structured encoding vector into the pre-trained detection model includes: Generate prompt words for input into a pre-trained detection model based on prompt word setting reference information, through manual setting and / or model generation; the prompt word setting reference information includes a functional description of the code to be tested, project-level contextual features of the software project to which the code belongs, detection instructions, and a preset prompt word template; Combine the text encoding vector and the structured encoding vector into a comprehensive vector; Input the comprehensive vector, prompt word, and code to be tested into the pre-trained detection model.

9. The security vulnerability detection method according to claim 8, characterized in that: The step of setting reference information according to the prompt word and generating the prompt word for input into the pre-trained detection model by manual setting and / or model generation includes: Pushing prompt word setting reference information to users so that users can set prompt words based on preset prompt word templates according to the functional description of the code to be tested, the project-level context characteristics of the software project to which the code belongs, and the detection instructions; The prompt word set by the user and the prompt word setting reference information are inputted into the pre-trained prompt word generation model so that the prompt word generation model can optimize the prompt word set by the user.

10. The security vulnerability detection method according to claim 1, wherein: After inputting the comprehensive vector formed by fusing the text encoding vector and the structured encoding vector into the pre-trained detection model, the security vulnerability detection method further includes: Based on the risk scores of each function in the code to be tested output by the detection model, the functions with security vulnerabilities are determined as target functions through the threshold method; Use the common vulnerability scoring system to evaluate and determine the risk level of security vulnerabilities in each target function; The code location of each target function and the risk level of the security vulnerability are pushed to the user.

11. The security vulnerability detection method according to claim 10, characterized in that: After determining the functions with security vulnerabilities as target functions using a threshold method based on the risk scores of the functions in the code to be tested output by the detection model, the security vulnerability detection method further includes: Determine the vulnerability type of the security vulnerability of each target function; Generate security vulnerability remediation suggestions for each target function using a vulnerability knowledge base or pre-trained remediation suggestion generation model; The vulnerability types and repair suggestions of the security vulnerabilities of each target function of the code to be tested are pushed to the user.

12. The security vulnerability detection method according to any one of claims 1 to 11, characterized in that: The security vulnerability detection method further includes: Obtain historical codes and their comprehensive vectors that are marked with actual security vulnerabilities and missed or misdetected by the detection model; The detection model is incrementally updated using historical codes and their synthesis vectors.

13. A security vulnerability detection device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the security vulnerability detection method according to any one of claims 1 to 12 when executing the computer program.

14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the security vulnerability detection method according to any one of claims 1 to 12 are implemented.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the security vulnerability detection method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Intelligent contract vulnerability detection method fusing code structure and semantic information

    CN119106435A

  • Code vulnerability detection method based on efficient parameter fine tuning of code pre-training model

    CN119441006A

  • Vulnerability detection method based on pre-training code language model and convolutional neural network

    CN119442246A

  • Software supply chain detection system based on large language model

    CN120337231A

  • Mapping a vulnerability to a stage of an attack chain taxonomy

    US20210367961A1

Cited By

  • Radar signal detection method and related device

    CN121234176A

  • Code vulnerability detection method and electronic equipment

    CN121615145A