An AI-based vulnerability repair rule generation method and related device

By building a cross-modal alignment technology of abstract syntax trees and semantic graphs based on a large AI language model, the problem of inaccurate vulnerability repair in existing technologies is solved, and efficient and reliable vulnerability repair and program performance assurance are achieved.

CN120611390BActive Publication Date: 2025-10-10HANGZHOU XIAODAO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511114793.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-10-10
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing rule-matching-based vulnerability repair methods lack an in-depth understanding of the overall program structure and contextual dependencies when dealing with complex vulnerabilities, resulting in inaccurate or incomplete repair solutions.

Method used

An AI-based large language model is used to parse vulnerability code snippets to construct an abstract syntax tree and semantic graph. Repair rules are generated through cross-modal alignment. Combined with vulnerability type identification and description text, the structural characteristics and semantic information of the vulnerability are deeply understood, the best repair mode is screened out, and repair rules are generated.

Benefits of technology

It improves the accuracy and reliability of vulnerability repair rules, ensures the compatibility and coordination between the repair plan and the original code, reduces manual debugging costs, predicts potential side effects, and ensures program performance stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611390B_ABST
    Figure CN120611390B_ABST
Patent Text Reader

Abstract

The application discloses an AI-based vulnerability repair rule generation method and related equipment, and relates to the field of digital data processing. A computer device performs double analysis on vulnerability code fragments and vulnerability description texts through a large language model, constructs double representations of abstract syntax trees and semantic graphs, aligns the semantic graphs and the abstract syntax trees in a cross-modal manner, and obtains a vulnerability context relationship matrix. The computer device traverses the abstract syntax trees, extracts vulnerability code feature vectors, and determines multiple repair modes according to vulnerability type identifiers. The computer device calculates the similarity between the vulnerability code feature vectors and multiple context constraint conditions, filters target repair modes, verifies the feasibility of the target repair modes based on the vulnerability context relationship matrix, determines an optimal repair mode, and generates a repair rule. The double analysis and cross-modal alignment greatly improve the accuracy and reliability of vulnerability repair rule generation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of electric digital data processing, and in particular to an AI-based vulnerability repair rule generation method and related equipment. BACKGROUND

[0002] With the rapid development of information technology, the scale and complexity of software systems are increasing, and the use of open source components has become an important means to improve software development efficiency. However, vulnerabilities that may exist in open source components can pose a serious security risk to the entire software system. These vulnerabilities can be exploited by malicious attackers, leading to security incidents such as data breaches and system crashes. Therefore, building an efficient open source component vulnerability repair mechanism is of great significance to ensure the safe operation of software systems.

[0003] Currently, automated vulnerability repair systems mainly use rule-based matching to identify and repair vulnerabilities: First, the characteristics of known vulnerabilities, such as code patterns and variable usage, are summarized to form a vulnerability feature rule set. When repairing vulnerabilities, the feature information of the code to be repaired is extracted through static code analysis techniques, and pattern matching is performed with the vulnerability feature rule set to identify potential vulnerability locations. For vulnerability code that matches successfully, a corresponding repair plan is generated according to the pre-set repair template, including code replacement, variable renaming, and other operations.

[0004] This repair method based on fixed rules and templates has limitations in practical applications. Since the related technology mainly focuses on local code features when performing code analysis, it lacks a deep understanding of the overall structure and context-dependent relationships of the program, resulting in incomplete or inaccurate repair plans when dealing with complex vulnerabilities. SUMMARY

[0005] The present application provides an AI-based vulnerability repair rule generation method and related equipment to improve the accuracy and reliability of vulnerability repair rule generation.

[0006] In the first aspect, the present application provides an AI-based vulnerability repair rule generation method, which is applied to computer equipment. The method includes: obtaining vulnerability information of open source components to be repaired, the open source component vulnerability information includes vulnerability code snippets, vulnerability type identifiers and vulnerability description texts; parsing the vulnerability code snippets through a large language model to construct an abstract syntax tree, traversing the abstract syntax tree, and extracting vulnerability code feature vectors, the abstract syntax tree including multiple abstract syntax tree nodes; according to the vulnerability type identifier, retrieving a repair pattern set from a preset vulnerability repair knowledge base, the repair pattern set including multiple repair patterns, the repair pattern including a repair operation sequence and corresponding context constraints; calculating the similarity between the vulnerability code feature vector and multiple context constraints, and determining the corresponding The target repair pattern whose similarity exceeds the preset similarity threshold; the vulnerability description text is processed by natural language through a large language model to obtain structured text features to construct a semantic graph, which includes multiple semantic graph nodes; the semantic graph is cross-modally aligned with the abstract syntax tree to obtain a vulnerability context relationship matrix; based on the vulnerability context relationship matrix, the feasibility of the target repair pattern is verified and the best repair pattern is screened out. The vulnerability context relationship matrix is ​​used to represent the correspondence and association strength between semantic graph nodes and abstract syntax tree nodes; the best repair pattern is pattern matched with the vulnerability code snippet, and repair rules are generated according to the preset code mapping relationship. The repair rules include code replacement instructions, variable renaming mapping and dependency update configuration.

[0007] By employing the above technical solution, the computer device uses a large language model to perform dual analysis of vulnerability code snippets and vulnerability description text, constructing a dual representation of an abstract syntax tree and a semantic graph. The semantic graph is then cross-modally aligned with the abstract syntax tree to generate a vulnerability context matrix, which provides a deeper understanding of the relationship between the structural features and semantic information of the vulnerability code. The computer device traverses the abstract syntax tree, extracts the vulnerability code feature vector, and identifies multiple remediation patterns based on the vulnerability type identifier. The computer device calculates the similarity between the vulnerability code feature vector and multiple contextual constraints, selects the target remediation pattern, and verifies the feasibility of the target remediation pattern based on the vulnerability context matrix, determines the optimal remediation pattern, and generates remediation rules. This dual analysis and cross-modal alignment approach significantly improves the accuracy and reliability of vulnerability remediation rule generation.

[0008] In combination with some embodiments of the first aspect, in some embodiments, the vulnerability code snippet is parsed through a large language model to construct an abstract syntax tree, traverse the abstract syntax tree, and extract the vulnerability code feature vector, specifically including: inputting the vulnerability code snippet into the large language model, identifying different types of abstract syntax tree nodes to generate an initial feature matrix, the initial feature matrix includes the position information and type identification of the abstract syntax tree node; based on the initial feature matrix, analyzing the subordinate relationship of each abstract syntax tree node to obtain a node feature vector sequence, the node feature vector sequence includes the hierarchical information and call dependency relationship of each abstract syntax tree node; according to the node feature vector sequence, constructing an abstract syntax tree, traversing each abstract syntax tree node in the abstract syntax tree, and generating a vulnerability code feature vector.

[0009] By employing the aforementioned technical solution, the computer device performs multi-level parsing of vulnerable code snippets using a large language model: first, it identifies different types of abstract syntax tree nodes and generates an initial feature matrix, capturing the code's basic structural information; then, it analyzes the subordinate relationships between the abstract syntax tree nodes to obtain a sequence of node feature vectors, accurately expressing the code's hierarchical information and call dependencies. This outward-facing, layer-by-layer parsing approach enables the computer device to fully understand the code's static structure and dynamic characteristics. By traversing the entire abstract syntax tree to generate a vulnerability code feature vector, it not only preserves the code's local features but also incorporates global structural information, providing a more accurate and comprehensive feature representation foundation for subsequent repair pattern matching.

[0010] In combination with some embodiments of the first aspect, in some embodiments, natural language processing is performed on the vulnerability description text through a large language model to obtain structured text features to construct a semantic graph, specifically including: performing word segmentation processing on the vulnerability description text to obtain candidate word units, filtering the candidate word units based on a preset vulnerability feature dictionary to obtain word units; labeling the word units with entity types to obtain labeled entity word units; analyzing the relationship between the entity word units through a large language model to establish entity relationship pairs, the entity relationship pairs including two entity word units and the relationship type between the two entity word units; calculating the credibility score of each entity relationship pair, and constructing a semantic graph based on the entity relationship pairs whose credibility scores exceed the preset score.

[0011] By adopting the above technical solution, computer equipment achieves the conversion from unstructured vulnerability description text to structured semantic graphs through natural language processing of a large language model: first, through word segmentation processing and filtering based on a preset vulnerability feature dictionary, vulnerability-related word units are extracted; then, the semantic role of the word units is clarified through entity type annotation; then, the large language model is used to analyze the relationship between entity word units and establish entity relationship pairs; then, the credibility score of the entity relationship pairs is calculated and screened to ensure the high quality and reliability of the constructed semantic graph. The computer equipment can thus accurately understand the semantic information in the vulnerability description text, provide a structured semantic representation for subsequent cross-modal alignment, and improve the accuracy of vulnerability remediation rule generation.

[0012] In combination with some embodiments of the first aspect, in some embodiments, the semantic graph and the abstract syntax tree are cross-modally aligned to obtain a vulnerability context relationship matrix. Based on the vulnerability context relationship matrix, the feasibility of the target repair mode is verified to screen out the best repair mode. The vulnerability context relationship matrix is ​​used to represent the correspondence and association strength between the semantic graph nodes and the abstract syntax tree nodes, specifically including: calculating the semantic similarity between the semantic graph nodes and the abstract syntax tree nodes; based on the semantic similarity, constructing a vulnerability context relationship matrix, each matrix element in the vulnerability context relationship matrix is ​​used to represent the association strength between the corresponding semantic graph node and the abstract syntax tree node; according to the vulnerability context relationship matrix, check whether the variable name in the target repair mode conflicts with the variable in the vulnerability code snippet, whether the function call order in the target repair mode is consistent with the control flow of the vulnerability code snippet, and whether the data transfer in the target repair mode completely matches the data flow of the vulnerability code snippet, so as to screen out the best repair mode.

[0013] By adopting the above technical solution, first, the computer equipment calculates the semantic similarity between the semantic graph nodes and the abstract syntax tree nodes, and constructs a vulnerability context relationship matrix, so that it can deeply explore the intrinsic connection between the semantic graph and the abstract syntax tree to accurately capture the logical details of the code and provide a solid data foundation for the verification of the repair pattern. Secondly, multi-dimensional inspection of the target repair pattern can effectively avoid new code errors after the repair, and ensure the compatibility and coordination of the repair plan with the original code. This verification and screening mechanism quickly eliminates unfeasible ones from multiple target repair patterns and selects the best one, greatly reducing the cost of manual debugging and providing efficient and reliable technical support for vulnerability repair in network security protection.

[0014] In combination with some embodiments of the first aspect, in some embodiments, after the step of pattern matching the best repair pattern with the vulnerability code snippet and generating a repair rule according to the preset code mapping relationship, the method also includes: inputting the repair rule and the preset rule verification prompt template into the large language model to obtain a verification result report, and the preset rule verification prompt template is used to instruct the large language model to perform rule verification; based on the historical vulnerability repair case library and the verification result report, judging whether the repair rule has potential side effects, and potential side effects refer to negative conditions caused to the performance of the target program after the repair rule is introduced; if so, removing the unstable code structure in the repair rule and supplementing the defensive check code to obtain the final repair rule.

[0015] By adopting the above technical solution, the computer equipment inputs the repair rules and preset rule verification prompt templates into the large language model, and with the help of the model's powerful natural language processing and logical reasoning capabilities, it quickly and comprehensively verifies the correctness and rationality of the repair rules to generate a detailed verification result report to make up for the limitations of manual verification. Based on the historical vulnerability repair case library and verification result report, it is judged whether the repair rules have potential side effects, which can effectively predict the potential impact on the performance of the target program after the repair, and avoid the risk of performance loss introduced by the repair in advance. For repair rules with potential side effects, the computer equipment removes unstable code structures and supplements defensive check codes, which can enhance code stability, so that the final repair rules can eliminate vulnerabilities while minimizing the negative impact on the performance of the target program, achieving the dual goals of vulnerability repair and program performance assurance, and significantly improving the robustness of the network security protection system.

[0016] In combination with some embodiments of the first aspect, in some embodiments, based on the historical vulnerability repair case library and the verification result report, it is judged whether the repair rule has potential side effects. The potential side effects refer to the negative conditions caused to the performance of the target program after the repair rule is introduced. Specifically, it includes: extracting the historical repair rule feature vector from the historical vulnerability repair case library, the historical repair rule feature vector includes the code structure characteristics, application scenario characteristics and side effect identification of the historical repair rule; calculating the matching degree between the verification result report and the historical repair rule feature vector, and determining the target historical repair rule whose matching degree exceeds the preset matching degree threshold; counting the side effects of the target historical repair rule to determine whether the repair rule has potential side effects.

[0017] By adopting the above technical solution, the computer equipment extracts the historical repair rule feature vectors from the historical vulnerability repair case library, fully tapping the value of historical vulnerability repair experience and converting the complex information in past cases into quantifiable and comparable data. By calculating the matching degree between the verification result report and the historical repair rule feature vector, the computer equipment selects the target historical repair rule, thereby quickly locating historical scenarios similar to the current repair rule. Based on the side effect statistics of the target historical repair rule, it can accurately predict the potential impact of the current repair rule on the target program performance after its introduction, avoiding misjudgments due to subjective judgment or lack of experience, effectively improving the accuracy and efficiency of side effect judgments, and ensuring that the repair rule will not have a negative impact on the target program performance in actual application, providing strong support for building a stable and reliable vulnerability repair system.

[0018] In combination with some embodiments of the first aspect, in some embodiments, after the step of pattern matching the optimal repair pattern with the vulnerable code snippet and generating the repair rule according to the preset code mapping relationship, the method also includes: constructing a runtime environment sandbox for testing the repair rule, and configuring the running parameters of the runtime environment sandbox, the running parameters including program dependency configuration, environment variable settings and resource constraints; applying the repair rule to multiple test copies of the vulnerable code snippet to generate the code to be tested; executing the code to be tested in the runtime environment sandbox to obtain performance loss data, memory usage data and exception information data during the execution process; analyzing the performance loss data, memory usage data and exception information data according to preset evaluation rules to obtain test evaluation results.

[0019] By adopting the above technical solutions, the computer equipment builds an operating environment sandbox and accurately configures operating parameters such as program dependencies, environment variables, and resource restrictions. This can highly simulate the real operating environment of the target program and ensure the reliability and validity of the test results. The computer equipment applies the repair rules to multiple test copies of the vulnerable code snippet to generate the code to be tested, and executes the code to be tested in the operating environment sandbox to obtain multi-dimensional data such as performance loss, memory usage, and exception information. It can fully capture the actual performance after the application of the repair rules from multiple angles such as operating efficiency, resource usage, and error conditions. After an in-depth analysis of these data based on preset evaluation rules, the test evaluation results are obtained, which provides an objective and scientific basis for the feasibility and quality evaluation of the repair rules. It can not only promptly discover performance problems or stability risks that may exist in the actual operation of the repair rules, but also provide a clear direction for subsequent optimization and adjustment, greatly reducing the potential risks after the deployment of the repair rules, and effectively ensuring the normal and stable operation of the target program after the vulnerability is repaired.

[0020] In a second aspect, an embodiment of the present application provides a computer device, comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the computer device to execute the method described in the first aspect and any possible implementation of the first aspect.

[0021] In a third aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed on a computer device, enables the computer device to execute the method described in the first aspect and any possible implementation of the first aspect.

[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium comprising instructions, which, when executed on a computer device, enables the computer device to execute the method described in the first aspect and any possible implementation of the first aspect.

[0023] It is understandable that the computer device provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the methods provided in the embodiments of the present application. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods and will not be repeated here.

[0024] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0025] 1. By adopting the above technical solution, the computer device uses a large language model to perform dual analysis of vulnerability code snippets and vulnerability description text, constructing a dual representation of an abstract syntax tree and a semantic graph. The semantic graph is then cross-modally aligned with the abstract syntax tree to obtain a vulnerability context relationship matrix, which provides a deeper understanding of the relationship between the structural features of the vulnerability code and its semantic information. The computer device traverses the abstract syntax tree, extracts the vulnerability code feature vector, and identifies multiple repair patterns based on the vulnerability type identifier. The computer device calculates the similarity between the vulnerability code feature vector and multiple context constraints, screens the target repair pattern, and verifies the feasibility of the target repair pattern based on the vulnerability context relationship matrix, determines the optimal repair pattern, and generates repair rules. This dual analysis and cross-modal alignment approach greatly improves the accuracy and reliability of vulnerability repair rule generation.

[0026] 2. By adopting the above technical solution, first, the computer equipment calculates the semantic similarity between the semantic graph nodes and the abstract syntax tree nodes, and constructs a vulnerability context relationship matrix, so that it can deeply explore the intrinsic connection between the semantic graph and the abstract syntax tree to accurately capture the logical details of the code and provide a solid data foundation for repair pattern verification. Secondly, multi-dimensional inspection of the target repair pattern can effectively avoid new code errors after the repair, and ensure the compatibility and coordination of the repair solution with the original code. This verification and screening mechanism quickly eliminates unfeasible ones from multiple target repair patterns and selects the best one, greatly reducing manual debugging costs and providing efficient and reliable technical support for vulnerability repair in network security protection.

[0027] 3. By adopting the above technical solution, the computer equipment inputs the repair rules and preset rule verification prompt templates into the large language model, and with the help of the model's powerful natural language processing and logical reasoning capabilities, quickly and comprehensively verifies the correctness and rationality of the repair rules to generate a detailed verification result report to make up for the limitations of manual verification. Based on the historical vulnerability repair case library and verification result report, it is judged whether the repair rules have potential side effects, which can effectively predict the potential impact on the performance of the target program after the repair, and avoid the risk of performance loss introduced by the repair in advance. For repair rules with potential side effects, the computer equipment removes unstable code structures and supplements defensive check codes, which can enhance code stability, so that the final repair rules can eliminate vulnerabilities while minimizing the negative impact on the performance of the target program, achieving the dual goals of vulnerability repair and program performance assurance, and significantly improving the robustness of the network security protection system. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a flowchart of the AI-based vulnerability repair rule generation method in an embodiment of the present application;

[0029] Figure 2 This is another flowchart of the AI-based vulnerability repair rule generation method in an embodiment of the present application;

[0030] Figure 3 It is a schematic diagram of the physical device structure of the computer equipment in the embodiment of the present application. DETAILED DESCRIPTION

[0031] The terms used in the following examples of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application, the singular expressions "a", "an", "above", "the", and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations of one or more of the listed items.

[0032] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0033] The following is a description of the process of the method provided by this implementation. Figure 1 , which is a flow chart of the AI-based vulnerability repair rule generation method in an embodiment of the present application.

[0034] S101. Obtain vulnerability information of an open source component to be repaired, where the open source component vulnerability information includes a vulnerability code snippet, a vulnerability type identifier, and a vulnerability description text;

[0035] Among them, open source component vulnerability information refers to a complete set of information describing security defects in open source components, including vulnerability code snippets, vulnerability type identifiers and vulnerability description text; open source components refer to software components released under open source licenses, such as common open source frameworks, libraries or tools; vulnerability code snippets refer to the part of the source code that contains the vulnerability, such as the function implementation with buffer overflow risk; vulnerability type identifiers are used to represent specific categories of vulnerabilities, such as standardized vulnerability classification codes such as SQL injection and cross-site scripting attacks; vulnerability description text refers to the natural language description of the cause of the vulnerability, the scope of impact and the possible harm it may cause.

[0036] Specifically, computer devices can obtain open source component vulnerability information through authoritative vulnerability databases, open source community platforms, and public technical channels:

[0037] (1) Authoritative vulnerability database: Connect to standardized vulnerability databases such as CVE, NVD, and CNVD to obtain standardized vulnerability descriptions, severity scores, and technical reference information;

[0038] (2) Open source community platform: Establish connections with code hosting platforms such as GitHub, GitLab, and Bitbucket through API integration, configure the Webhook mechanism, and receive real-time event notifications, including information such as code submissions, problem reports, and merge requests;

[0039] (3) Public technical channels: Using a combination of RSS subscription and distributed crawlers, we monitor technical communities such as the Linux kernel mailing list and the Apache mailing list, as well as information sources such as security announcements, technical blogs, and product update logs from major manufacturers, and regularly trigger crawler tasks to capture the latest content from these channels.

[0040] This multi-source data collection system can comprehensively cover all sources of open source component vulnerability information, ensuring that computer equipment can obtain complete open source component vulnerability information in a timely manner, and providing a reliable data foundation for subsequent vulnerability analysis and repair.

[0041] S102. Parsing the vulnerable code snippet using a large language model to construct an abstract syntax tree, traversing the abstract syntax tree, and extracting a vulnerability code feature vector, wherein the abstract syntax tree includes a plurality of abstract syntax tree nodes;

[0042] Among them, the large language model refers to a deep learning model that has been trained with a large-scale vulnerability dataset and is specifically used to handle tasks related to software security vulnerabilities. The large language model has mastered professional knowledge such as vulnerability feature identification, code semantic understanding, and security repair patterns by learning a large number of vulnerability code samples, vulnerability description texts, and vulnerability repair solutions; the abstract syntax tree refers to the tree structure representation obtained after parsing the vulnerability code fragment, which is used to display the grammatical structure and semantic relationship of the code; the abstract syntax tree node refers to each node in the abstract syntax tree, including program elements such as function declarations, variable definitions, and control flow statements; the vulnerability code feature vector is used to represent the multi-dimensional feature information extracted from the vulnerability code fragment.

[0043] Specifically, the computer first calls the code parsing interface of the large language model to convert the vulnerable code snippet into a sequence of lexical units. Then, based on the grammatical rules of the programming language, the computer constructs an abstract syntax tree (AST). By traversing each AST node in a depth-first or breadth-first manner, the computer analyzes the node type, attributes, and hierarchical relationships, extracting a feature vector containing dimensions such as variable usage patterns, control flow paths, and data flow characteristics, thereby obtaining the vulnerability code feature vector.

[0044] The following is a specific buffer overflow vulnerability code example:

[0045] Vulnerable code snippet:

[0046] void copyData(char* input) {

[0047] char buffer

[10] ;

[0048] strcpy(buffer, input);

[0049] printf("Data: %s\n", buffer);

[0050] }

[0051] Input into the large language model to convert the vulnerable code snippet into a sequence of lexical units:

[0052] [VOID] [IDENTIFIER:copyData][LPAREN] [CHAR][MULTIPLY] [IDENTIFIER:input][RPAREN]

[0053] [LBRACE] [CHAR][IDENTIFIER:buffer] [LBRACKET][NUMBER:10] [RBRACKET][SEMICOLON]

[0054] [IDENTIFIER:strcpy] [LPAREN][IDENTIFIER:buffer] [COMMA][IDENTIFIER:input] [RPAREN]...

[0055] Constructing an abstract syntax tree:

[0056] 1. FunctionDefinition(copyData)

[0057] 1.1. ReturnType: void

[0058] 1.2. FunctionName: copyData

[0059] 1.3. ParameterList

[0060] 1.3.1. Parameter

[0061] 1.3.1.1. Type: char*

[0062] 1.3.1.2. Name: input

[0063] CompoundStatement (Function Body)

[0064] VariableDeclaration

[0065] 1.4.1.1. Type: char[]

[0066] 1.4.1.2. Name: buffer

[0067] 1.4.1.3. ArraySize: 10

[0068] FunctionCall(strcpy)

[0069] 1.4.2.1. FunctionName: strcpy

[0070] 1.4.2.2. ArgumentList

[0071] 1.4.2.2.1. Argument1: buffer

[0072] 1.4.2.2.2. Argument2: input

[0073] 1.4.3. FunctionCall(printf)

[0074] 1.4.3.1. FunctionName: printf

[0075] 1.4.3.2. ArgumentList

[0076] 1.4.3.2.1. Argument1: "Data: %s\n"

[0077] 1.4.3.2.2. Argument2: buffer

[0078] Traverse the abstract syntax tree and extract the vulnerability code feature vector:

[0079] {

[0080] "string_concatenation": 1, / / string concatenation operation exists

[0081] "user_input_vars": ["username"], / / User input variables

[0082] "sql_keywords": ["SELECT", "FROM", "WHERE"], / / SQL keywords

[0083] "db_operations": ["createStatement", "executeQuery"], / / Database operations

[0084] "input_sanitization": 0, / / Missing input validation

[0085] "prepared_statement": 0, / / Not using prepared statements

[0086] "error_handling": 0, / / Missing error handling

[0087] "return_type": "User", / / Return value type

[0088] "parameter_count": 1, / / Number of parameters

[0089] "query_complexity": 1 / / SQL query complexity

[0090] }

[0091] At level 1.4.1, the local variable declaration is identified, and the buffer size information is obtained;

[0092] At level 1.4.2, the unsafe function strcpy is identified;

[0093] At levels 1.3 and 1.4.2.2, it is found that the input parameter input is directly used in the strcpy operation;

[0094] By analyzing the relationship between 1.4.1 and 1.4.2, it is found that there is a lack of input size verification.

[0095] Optionally, in general, the vulnerability code snippet is parsed through a large language model to construct an abstract syntax tree. Traversing the abstract syntax tree and extracting the vulnerability code feature vector can be achieved in the following way, which is not limited here: inputting the vulnerability code snippet into the large language model, identifying different types of abstract syntax tree nodes to generate an initial feature matrix, which includes the position information and type identification of the abstract syntax tree nodes; based on the initial feature matrix, analyzing the subordinate relationship of each abstract syntax tree node to obtain a node feature vector sequence, which includes the hierarchical information and call dependency relationship of each abstract syntax tree node; according to the node feature vector sequence, constructing an abstract syntax tree, traversing each abstract syntax tree node in the abstract syntax tree, and generating a vulnerability code feature vector.

[0096] S103: Retrieve a repair pattern set from a preset vulnerability repair knowledge base according to the vulnerability type identifier, wherein the repair pattern set includes multiple repair patterns, and the repair pattern includes a repair operation sequence and corresponding context constraints;

[0097] Among them, the preset vulnerability repair knowledge base includes repair patterns corresponding to vulnerability type identifiers; repair patterns refer to standardized repair solutions for specific vulnerability types; repair operation sequences are used to represent specific code modification steps, such as deletion, insertion, or replacement operations; context constraints refer to code scenarios and environmental requirements to which repair patterns are applicable.

[0098] Specifically, the computer first locates a set of relevant repair patterns in a pre-set vulnerability repair knowledge base based on the vulnerability type identifier. This set of repair patterns includes multiple repair patterns, each designed and verified by experts. The computer then analyzes the repair operation sequence within the repair pattern, including the code location, modification method, and modification content. It also checks contextual constraints, such as code structure requirements, variable scope restrictions, and external dependency versions, to ensure that the repair pattern can be correctly applied to the target code environment.

[0099] S104, calculating the similarity between the vulnerability code feature vector and multiple context constraints, and determining a target repair mode whose similarity exceeds a preset similarity threshold;

[0100] Among them, similarity refers to the distance measurement between two vectors in the feature space, which is used to indicate the degree of matching between the vulnerability code snippet and the context constraints in the repair pattern; the preset similarity threshold refers to the minimum similarity requirement for determining whether the repair pattern is applicable; the target repair pattern refers to the repair pattern that is most suitable for the current vulnerability scenario after similarity screening.

[0101] Specifically, the computer device determines the feature vector representation of the contextual constraints corresponding to each repair pattern. It then uses an algorithm such as cosine similarity to calculate the similarity between the vulnerability code feature vector and the feature vector representation of each contextual constraint. The computer device compares the calculated similarity with a preset similarity threshold and selects target repair patterns whose similarity exceeds the preset similarity threshold.

[0102] Based on the buffer overflow vulnerability example in step S102, the process of similarity calculation and repair pattern matching is explained:

[0103] Vulnerability code feature vector (V1):

[0104] {

[0105] "buffer_size": 10, # Fixed buffer size

[0106] "unsafe_functions": 1, # Unsafe functions exist

[0107] "input_params": 1, # external input parameters

[0108] "bounds_check": 0, # no bounds check

[0109] "buffer_type": 1, # char array type

[0110] "param_type": 1, # char* type parameter

[0111] "memory_operations": 1, # memory operations

[0112] "size_validation": 0 # No size validation

[0113] }

[0114] The three repair modes in the repair mode set and their context constraints:

[0115] Safe function replacement mode (V2):

[0116] {

[0117] "buffer_size": -1, # arbitrary buffer size

[0118] "unsafe_functions": 1, # for unsafe functions

[0119] "input_params": 1, #external input

[0120] "bounds_check": 0, #No bounds check required

[0121] "buffer_type": 1, #char array type

[0122] "param_type": 1, #char* type parameter

[0123] "memory_operations": 1, # memory operation scenario

[0124] "size_validation": 0 #No size validation required

[0125] }

[0126] #Fix: Replace strcpy with strncpy and add a buffer size parameter;

[0127] Input Verification Mode (V3):

[0128] {

[0129] "buffer_size": 10, #fixed size buffer

[0130] "unsafe_functions": 0, #Don't focus on specific functions

[0131] "input_params": 1, # there is external input

[0132] "bounds_check": 0, # No bounds check required

[0133] "buffer_type": 1, # char array type

[0134] "param_type": 1, # char* type parameter

[0135] "memory_operations": 0, # Don't pay attention to memory operations

[0136] "size_validation": 0 # No size validation required

[0137] }

[0138] #Fix: Add input length check logic;

[0139] Dynamic allocation mode (V4):

[0140] {

[0141] "buffer_size": -1, #unlimited buffer size

[0142] "unsafe_functions": 0, #Don't focus on specific functions

[0143] "input_params": 1, # there is external input

[0144] "bounds_check": 0, # No bounds check required

[0145] "buffer_type": 0, # char array is not required

[0146] "param_type": 1, # char* type parameter

[0147] "memory_operations": 1, # memory operation scenario

[0148] "size_validation": 0 # Do not require size validation

[0149] }

[0150] #Fix solution: Use dynamic memory allocation;

[0151] Similarity calculation process (using the cosine similarity formula cos(θ) = (V1·V2) / (||V1||·||V2||), the preset similarity threshold is 0.8):

[0152] Sim(V1, V2)=0.92 # Exceeds the preset similarity threshold;

[0153] Sim(V1, V3)=0.85 # Exceeds the preset similarity threshold;

[0154] Sim(V1, V4)=0.65 # Does not exceed the preset similarity threshold;

[0155] Based on the calculation results, the computer equipment identified two target repair modes:

[0156] Security function replacement pattern (similarity 0.92);

[0157] Enter the verification pattern (similarity 0.85).

[0158] S105, performing natural language processing on the vulnerability description text by the large language model to obtain structured text features, to construct a semantic graph, the semantic graph including a plurality of semantic graph nodes;

[0159] The structured text features refer to normalized semantic information extracted from the vulnerability description text; the semantic graph refers to a network structure describing entities and their relationships in the vulnerability description text; and the semantic graph node is used to represent an entity in the semantic graph, including program elements, vulnerability types, and impact ranges.

[0160] Specifically, the computer device uses the natural language processing capability of the large language model to perform word segmentation, part-of-speech tagging, named entity recognition, and other processing on the vulnerability description text, and extracts key concepts and semantic relationships. Then, the computer device constructs a semantic graph based on this information, taking the identified program elements, vulnerability features, and security impact entities as semantic graph nodes, and representing various relationships between entities through semantic graph edges, such as "trigger-cause", "dependence-impact", etc. This graph structure can clearly show the semantic network of vulnerability-related concepts.

[0161] Taking the buffer overflow vulnerability in step S102 as an example, the processing process of the vulnerability description text is shown:

[0162] Vulnerability description text: A buffer overflow vulnerability is found in the function copyData, which uses the strcpy function to copy user input into a buffer array with a fixed size of 10 bytes without input length validation. When the length of the input string exceeds the buffer size, it will cause stack overflow, which may cause program crash or be exploited by attackers to execute arbitrary code. It is recommended to use the strncpy function and add boundary checks.

[0163] 1. Text processing steps:

[0164] Word segmentation and part-of-speech tagging:

[0165] In the function (n) copyData (n), a buffer overflow (n) vulnerability is found (v), the function (n) uses the strcpy (n) function to copy user input (n) to a buffer (n) array with a fixed size (adj) of 10 bytes (num) in (p), without (v) input length validation (n).

[0166] Named entity recognition:

[0167] [Program elements]: copyData, strcpy, buffer;

[0168] [Vulnerability Type]: Buffer overflow;

[0169] [Technical parameters]: 10 bytes;

[0170] [Security Impact]: Program crash, arbitrary code execution;

[0171] [Repair suggestion]: strncpy, bounds check;

[0172] 2. Build a semantic graph:

[0173] [Graph Node]

[0174] 1. Function node: copyData

[0175] 2. Vulnerability Node: Buffer Overflow

[0176] 3. Code element node:

[0177] - buffer (array)

[0178] - strcpy (function)

[0179] - strncpy (function)

[0180] 4. Feature nodes:

[0181] - Buffer size: 10 bytes

[0182] - Input Validation: Missing

[0183] 5. Impact nodes:

[0184] - Program crash

[0185] - Code Execution

[0186] 6. Repair the node:

[0187] - Bounds checking

[0188] - Safe function replacement

[0189] [Relationship Edge]

[0190] 1. Position relationship:

[0191] copyData --contains-->buffer;

[0192] copyData --use-->strcpy;

[0193] 2. Vulnerability characteristics:

[0194] Buffer overflow -- exists in -->copyData;

[0195] buffer -- feature --> fixed size (10 bytes);

[0196] strcpy -- missing --> input validation;

[0197] 3. Cause and effect:

[0198] buffer overflow -- can lead to --> program crash;

[0199] buffer overflow -- can lead to --> code execution;

[0200] 4. Repair relationship:

[0201] strcpy -- replace with --> strncpy;

[0202] copyData -- needs to add --> boundary check.

[0203] Optionally, in general, the structured text features obtained by processing the vulnerability description text through a large language model can be implemented in the following manner, without limitation: performing word segmentation processing on the vulnerability description text to obtain candidate word units, filtering the candidate word units based on a pre-set vulnerability feature dictionary to obtain word units; performing entity type labeling on the word units to obtain labeled entity word units; analyzing the relationships between the entity word units through a large language model to establish entity relationship pairs, the entity relationship pairs including two entity word units and the relationship type between the two entity word units; calculating the credibility score of each entity relationship pair, and constructing a semantic graph based on the entity relationship pairs whose credibility scores exceed a pre-set score.

[0204] S106, align the semantic graph with the abstract syntax tree across modalities to obtain a vulnerability context relationship matrix, based on the vulnerability context relationship matrix, verify the feasibility of the target repair mode, and select the best repair mode, the vulnerability context relationship matrix is used to represent the corresponding relationship and association strength between the semantic graph nodes and the abstract syntax tree nodes;

[0205] Wherein, the cross-modal alignment refers to establishing the mapping relationship between the text description and the code structure; the vulnerability context relationship matrix represents the corresponding relationship matrix between the semantic graph nodes and the abstract syntax tree nodes; the association strength is used to represent the confidence of the mapping relationship between the nodes; the best repair mode refers to the final repair mode determined after feasibility verification.

[0206] Specifically, the computer device uses methods such as semantic similarity calculation and naming pattern matching to identify the corresponding positions of semantic graph nodes in the abstract syntax tree, forming a mapping relationship between the two representations. These mapping relationships are organized into a vulnerability context relationship matrix, where each matrix element represents the strength of the association between corresponding nodes. Based on this vulnerability context relationship matrix, the computer device verifies the feasibility of the target repair pattern, eliminates infeasible target repair patterns, and ultimately determines the optimal repair pattern.

[0207] Based on the previous buffer overflow vulnerability example, the cross-modal alignment and feasibility verification process is demonstrated:

[0208] 1. Abstract syntax tree node (A):

[0209] A1: FunctionDefinition(copyData)

[0210] A2: ParameterDeclaration(char* input)

[0211] A3: ArrayDeclaration(char buffer

[10] )

[0212] A4: FunctionCall(strcpy)

[0213] A5: Argument (buffer)

[0214] A6: Argument (input)

[0215] A7: FunctionCall(printf)

[0216] 2. Semantic graph node (S):

[0217] S1: function (copyData)

[0218] S2: Vulnerability type (buffer overflow)

[0219] S3: array (buffer, 10 bytes)

[0220] S4: Function call (strcpy)

[0221] S5: Security impact (program crash)

[0222] S6: Repair suggestion (strncpy)

[0223] S7: Repair suggestions (bounds check)

[0224] Build vulnerability context relationship matrix (correlation strength range 0-1, 0 represents no correlation, 1 represents complete correspondence):

[0225] Table 1 Vulnerability context relationship matrix

[0226]

[0227] 4. Interpretation of mapping relationship:

[0228] Strong correlation pair:

[0229] (S1, A1): 0.9 - copyData function definition completely corresponds;

[0230] (S3, A3): 1.0 - buffer array declaration completely corresponds;

[0231] (S4, A4): 1.0 - strcpy function call completely corresponds;

[0232] (S6, A4): 0.9 - strncpy replacement is highly related to the position of strcpy;

[0233] (S7, A2): 0.7 - boundary check needs to verify input parameter;

[0234] (S7, A3): 0.8 - boundary check needs to verify buffer size;

[0235] 5. Verify the feasibility of two target repair modes:

[0236] Feasibility analysis of target repair mode 1: safe function replacement (strncpy):

[0237] Accurate positioning: 0.9 correlation strength with A4 (strcpy call);

[0238] Parameter complete: the current code contains the required buffer (A5) and input (A6) parameters;

[0239] Implementation is simple: only need to replace the function name and add the length parameter;

[0240] Feasibility score: 0.85;

[0241] Feasibility analysis of target repair mode 2: adding boundary check:

[0242] Accurate positioning: high correlation with A2 (input parameter) and A3 (buffer declaration);

[0243] Condition support: buffer size (A3) and input length information can be obtained;

[0244] Implementation complexity: New control flow statements need to be added;

[0245] Feasibility score: 0.75;

[0246] Determine the best repair mode: Based on the vulnerability context matrix analysis and feasibility score, determine "safe function replacement" as the best repair mode.

[0247] Optionally, in general, the semantic graph and the abstract syntax tree are cross-modally aligned to obtain a vulnerability context relationship matrix. Based on the vulnerability context relationship matrix, the feasibility of the target repair mode is verified to screen out the best repair mode. The vulnerability context relationship matrix is ​​used to represent the correspondence and association strength between the semantic graph nodes and the abstract syntax tree nodes. It can be achieved in the following ways, which are not limited here: calculating the semantic similarity between the semantic graph nodes and the abstract syntax tree nodes; based on the semantic similarity, constructing a vulnerability context relationship matrix, each matrix element in the vulnerability context relationship matrix is ​​used to represent the association strength between the corresponding semantic graph node and the abstract syntax tree node; according to the vulnerability context relationship matrix, check whether the variable name in the target repair mode conflicts with the variable in the vulnerability code snippet, whether the function call order in the target repair mode is consistent with the control flow of the vulnerability code snippet, and whether the data transfer in the target repair mode completely matches the data flow of the vulnerability code snippet, so as to screen out the best repair mode.

[0248] S107: Pattern-match the best repair pattern with the vulnerable code snippet, and generate repair rules based on a preset code mapping relationship. The repair rules include code replacement instructions, variable renaming mapping, and dependency update configuration.

[0249] Among them, pattern matching refers to the process of applying the repair operation sequence in the optimal repair pattern to specific code; the preset code mapping relationship represents the correspondence rule between the repair operation sequence in the optimal repair pattern and the specific code modification; the code replacement instruction is used to indicate the code modification location and code modification content; the variable renaming mapping is used to indicate the unified adjustment rules of the variable name; the dependency update configuration is used to indicate the update requirements of the related component versions and configurations.

[0250] Specifically, the computer device matches the repair operation sequence in the optimal repair mode with the vulnerable code snippet to determine the specific code modification location. Then, based on the preset code mapping relationship, the computer device converts the repair operation sequence into specific code modifications, including deleting unsafe code, inserting security checks, and replacing API calls. Simultaneously, the computer device generates mapping rules for variable renaming to ensure a consistent naming style for the modified code. Furthermore, the computer device generates necessary dependency update instructions, such as specifying the component version number of the secure version and updating configuration parameters, to form a complete repair rule.

[0251] By employing the above technical solution, the computer device uses a large language model to perform dual analysis of vulnerability code snippets and vulnerability description text, constructing a dual representation of an abstract syntax tree and a semantic graph. The semantic graph is then cross-modally aligned with the abstract syntax tree to generate a vulnerability context matrix, which provides a deeper understanding of the relationship between the structural features and semantic information of the vulnerability code. The computer device traverses the abstract syntax tree, extracts the vulnerability code feature vector, and identifies multiple remediation patterns based on the vulnerability type identifier. The computer device calculates the similarity between the vulnerability code feature vector and multiple contextual constraints, selects the target remediation pattern, and verifies the feasibility of the target remediation pattern based on the vulnerability context matrix, determines the optimal remediation pattern, and generates remediation rules. This dual analysis and cross-modal alignment approach significantly improves the accuracy and reliability of vulnerability remediation rule generation.

[0252] The following is a more detailed description of the process of the method provided by this implementation. Figure 2 , is another flowchart of the AI-based vulnerability repair rule generation method in an embodiment of the present application.

[0253] After step S107, the following steps may be performed or not performed, which is not limited here:

[0254] S201: Input the repair rule and the preset rule verification prompt template into the large language model to obtain a verification result report. The preset rule verification prompt template is used to instruct the large language model to perform rule verification.

[0255] Among them, the preset rule verification prompt template refers to the standardized input format used to guide the large language model to perform rule verification, which includes verification focus, verification rules and expected results; the verification result report refers to the detailed evaluation document generated after the large language model analyzes and repairs the rules, which includes analysis results in multiple dimensions such as rule correctness, completeness and potential problems.

[0256] Specifically, the computer device inputs the generated repair rule and a preset rule verification prompt template into the large language model. The preset rule verification prompt template includes the following verification points: the repair rule's syntactic correctness, compatibility with the original code, the completeness of the security check, and the possible performance impact. The large language model then analyzes the results and generates a verification report covering multiple dimensions.

[0257] S202. Based on the historical vulnerability repair case library and verification result report, determine whether the repair rule has potential side effects. Potential side effects refer to negative conditions caused by the introduction of the repair rule on the performance of the target program.

[0258] Among them, the historical vulnerability repair case library refers to a collection of data on past vulnerability repair experiences and results; potential side effects refer to the negative effects that repair rules may bring while solving the target vulnerability, such as increased performance overhead and increased memory usage.

[0259] Specifically, the computer device compares and analyzes the verification result report with historical fix cases in the historical vulnerability fix case library. By matching similar historical fix cases, the computer device evaluates whether the current fix rule may introduce similar potential side effects. For example, for buffer overflow vulnerabilities, the computer device will focus on the performance overhead caused by adding bounds checks in historical fix cases, or the increased memory usage caused by replacing them with safe functions.

[0260] Optionally, in general, based on the historical vulnerability repair case library and the verification result report, it is judged whether the repair rule has potential side effects. Potential side effects refer to the negative conditions on the target program performance after the repair rule is introduced. This can be achieved in the following ways, which are not limited here: extracting the historical repair rule feature vector from the historical vulnerability repair case library. The historical repair rule feature vector includes the code structure characteristics, application scenario characteristics and side effect identification of the historical repair rule; calculating the matching degree between the verification result report and the historical repair rule feature vector, and determining the target historical repair rule whose matching degree exceeds the preset matching degree threshold; counting the side effects of the target historical repair rule to determine whether the repair rule has potential side effects.

[0261] S203. If it exists, remove the unstable code structure in the repair rule and add defensive check code to obtain the final repair rule.

[0262] Among them, unstable code structure refers to code fragments that may cause program exceptions or performance problems, such as redundant memory operations or repeated security checks; defensive check code refers to security check code used to prevent and handle abnormal situations, including input validation, boundary checking, etc.; final repair rules refer to the final version of the repair plan after optimization and adjustment.

[0263] Specifically, when the computer device discovers that the repair rule has potential side effects, two aspects of optimization are performed: first, unstable code structures that can cause performance loss are identified and removed, such as merging duplicate check logic, optimizing memory operation sequences, etc.; second, necessary defensive check code is supplemented to ensure the safety and robustness of the code, such as adding input parameter validity verification, perfecting exception handling logic, etc. Through this optimization adjustment, the final repair rule is obtained which can effectively repair the vulnerability without significantly affecting the program performance.

[0264] S204, a running environment sandbox for testing the repair rule is constructed, and running parameters of the running environment sandbox are configured, the running parameters including program dependency configuration, environment variable setting and resource limitation condition.

[0265] Among them, the running environment sandbox refers to an isolated test environment for safely executing and verifying the repaired code; the program dependency configuration refers to the configuration information of external libraries, component versions, etc. required for code running; the environment variable setting includes system path, compilation option, etc. environment parameters required at runtime; the resource limitation condition is used to limit the CPU usage, memory upper limit, etc. resource constraints in the testing process.

[0266] Specifically, the computer device creates an independent virtual environment as a running environment sandbox and performs complete environment configuration: sets necessary program dependencies, such as specifying specific operating system version, compiler version and third-party library version; configures key environment variables, including compilation parameters, runtime parameters, etc.; at the same time, set resource usage limits, such as maximum memory usage, CPU time limit, etc., to ensure that the test is carried out under controlled conditions.

[0267] S205, apply the repair rule to multiple test copies of the vulnerability code segment to generate test code.

[0268] Among them, the test copy refers to multiple different variants of the original vulnerability code segment, used to test the effect of the repair rule in different scenarios; the test code refers to the complete test case generated after applying the repair rule to the test copy.

[0269] Specifically, the computer device generates multiple test copies with different characteristics by transforming the vulnerability code segment, such as changing variable names, adjusting code structure, adding boundary conditions, etc. Then, the computer device applies the repair rule to these test copies to generate a series of test code to verify the universality and stability of the repair rule.

[0270] S206, execute the test code in the running environment sandbox to obtain performance loss data, memory occupation data and exception information data in the execution process.

[0271] Among them, performance loss data refers to performance indicators such as CPU usage and response time during code execution; memory usage data includes indicators such as memory usage peak and memory allocation frequency; exception information data records abnormal conditions such as errors and warnings during code execution.

[0272] Specifically, the computer device runs the code to be tested in a configured operating environment sandbox, and uses monitoring tools to collect various data in real time during the operation process: recording the performance overhead of code execution, including function call time, CPU occupancy, etc.; monitoring memory usage, including memory allocation size, garbage collection frequency, etc.; capturing and recording abnormal information during the operation process, such as segmentation errors, memory leaks, etc.

[0273] S207: Analyze the performance loss data, memory usage data, and abnormal information data according to preset evaluation rules to obtain test evaluation results.

[0274] Among them, the preset evaluation rules refer to the standardized evaluation criteria used to judge the repair effect, including specific indicators such as performance benchmarks and stability requirements; the test evaluation results are a comprehensive evaluation of the actual effect of the repair rules, including analysis conclusions from multiple dimensions.

[0275] Specifically, the computer device conducts in-depth analysis of the collected data based on pre-set evaluation rules: assessing whether performance loss is within an acceptable range and comparing performance differences before and after the fix; analyzing whether memory usage patterns are reasonable and checking for memory leak risks; and compiling statistics on anomalies to assess code stability. Ultimately, the computer device generates a test evaluation report covering multiple dimensions, including performance impact, resource consumption, and stability ratings, providing a basis for final decision-making on the adoption of the fix.

[0276] The following describes the computer device in the embodiment of the present invention from the perspective of hardware processing. Figure 3 , is a schematic diagram of a physical device structure of a computer device in an embodiment of the present application.

[0277] It should be noted that Figure 3 The structure of the computer device shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0278] like Figure 3As shown, the computer device includes a CPU 301, which can perform various appropriate actions and processes according to programs stored in a read-only memory ROM 302 or programs loaded from a storage unit 308 into a random access memory RAM 303, such as executing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An I / O interface 305 is also connected to the bus 304.

[0279] The following components are connected to the I / O interface 305: an input section 306 including an audio input device, push button switches, and the like; an output section 307 including a liquid crystal display (LCD), an audio output device, indicator lights, and the like; a storage section 308 including a hard disk and the like; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. Removable media 311, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 310 as needed, so that computer programs read from the removable media can be installed in the storage section 308 as needed.

[0280] In particular, according to an embodiment of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present invention includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 309 and / or installed from the removable medium 311. When the computer program is executed by the CPU 301, the various functions defined in the present invention are performed.

[0281] It should be noted that specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0282] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. Each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings.

[0283] Specifically, the computer device of this embodiment includes a processor and a memory, and a computer program is stored in the memory. When the computer program is executed by the processor, the AI-based vulnerability repair rule generation method provided in the above embodiment is implemented.

[0284] As another aspect, the present invention further provides a computer-readable storage medium, which may be included in the computer device described in the above embodiments, or may exist independently and not incorporated into the computer device. The storage medium carries one or more computer programs, which, when executed by a processor of the computer device, enable the computer device to implement the AI-based vulnerability remediation rule generation method provided in the above embodiments.

[0285] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0286] In the above embodiments, the term "when" can be interpreted to mean "if" or "after" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "on determining" or "if detecting (a stated condition or event)" can be interpreted to mean "if determining" or "in response to determining" or "on detecting (a stated condition or event)" or "in response to detecting (a stated condition or event)" depending on the context.

[0287] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by a computer program instructing the relevant hardware to complete, and the program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. The aforementioned storage medium includes ROM or random storage memory RAM, magnetic disk or optical disk and various storage program codes.

Claims

1. A vulnerability repair rule generation method based on AI, characterized in that: Applied to a computer device, the method comprises: Obtaining vulnerability information of an open source component to be repaired, wherein the open source component vulnerability information includes a vulnerability code snippet, a vulnerability type identifier, and a vulnerability description text; Parsing the vulnerability code snippet using a large language model to construct an abstract syntax tree, traversing the abstract syntax tree to extract a vulnerability code feature vector, wherein the abstract syntax tree includes a plurality of abstract syntax tree nodes; Retrieving a repair pattern set from a preset vulnerability repair knowledge base according to the vulnerability type identifier, wherein the repair pattern set includes multiple repair patterns, and the repair pattern includes a repair operation sequence and corresponding context constraints; Calculating the similarity between the vulnerability code feature vector and the plurality of context constraints, and determining a target repair mode whose similarity exceeds a preset similarity threshold; Performing natural language processing on the vulnerability description text using the large language model to obtain structured text features to construct a semantic graph, wherein the semantic graph includes multiple semantic graph nodes; Calculating the semantic similarity between the semantic graph nodes and the abstract syntax tree nodes; Based on the semantic similarity, a vulnerability context relationship matrix is ​​constructed, where each matrix element in the vulnerability context relationship matrix is ​​used to represent the association strength between the corresponding semantic graph node and the abstract syntax tree node; According to the vulnerability context matrix, check whether the variable names in the target repair pattern conflict with the variables in the vulnerable code snippet, whether the function call order in the target repair pattern is consistent with the control flow of the vulnerable code snippet, and whether the data transfer in the target repair pattern completely matches the data flow of the vulnerable code snippet, so as to screen out the best repair pattern; The optimal repair pattern is pattern matched with the vulnerable code snippet, and a repair rule is generated according to a preset code mapping relationship. The repair rule includes a code replacement instruction, a variable renaming mapping, and a dependency update configuration.

2. The method according to claim 1, characterized in that The method of parsing the vulnerability code snippet using a large language model to construct an abstract syntax tree, traversing the abstract syntax tree, and extracting a vulnerability code feature vector specifically includes: Inputting the vulnerable code snippet into the large language model to identify different types of abstract syntax tree nodes to generate an initial feature matrix, wherein the initial feature matrix includes position information and type identifiers of the abstract syntax tree nodes; Analyzing the subordinate relationship of each of the abstract syntax tree nodes based on the initial feature matrix to obtain a node feature vector sequence, wherein the node feature vector sequence includes the hierarchical information and call dependency relationship of each of the abstract syntax tree nodes; The abstract syntax tree is constructed according to the node feature vector sequence, and each abstract syntax tree node in the abstract syntax tree is traversed to generate the vulnerability code feature vector.

3. The method according to claim 1, characterized in that The vulnerability description text is subjected to natural language processing by the large language model to obtain structured text features to construct a semantic graph, specifically including: Performing word segmentation processing on the vulnerability description text to obtain candidate word units, and filtering the candidate word units based on a preset vulnerability feature dictionary to obtain word units; Performing entity type annotation on the word unit to obtain annotated entity word unit; Analyzing the relationship between the entity word units through the large language model to establish an entity relationship pair, wherein the entity relationship pair includes two entity word units and a relationship type between the two entity word units; The credibility score of each entity relationship pair is calculated, and the semantic graph is constructed based on the entity relationship pairs whose credibility scores exceed a preset score.

4. The method according to claim 1, wherein After the step of pattern matching the optimal repair pattern with the vulnerable code snippet and generating a repair rule according to a preset code mapping relationship, the method further includes: Inputting the repair rule and a preset rule verification prompt template into the large language model to obtain a verification result report, wherein the preset rule verification prompt template is used to instruct the large language model to perform rule verification; Based on the historical vulnerability repair case library and the verification result report, determine whether the repair rule has potential side effects, wherein the potential side effects refer to negative conditions on the target program performance caused by the introduction of the repair rule; If so, remove the unstable code structure in the repair rule and add defensive check code to obtain the final repair rule.

5. The method according to claim 4, characterized in that Based on the historical vulnerability repair case library and the verification result report, it is determined whether the repair rule has potential side effects. The potential side effects refer to negative effects on the target program performance caused by the introduction of the repair rule, specifically including: Extracting a historical repair rule feature vector from the historical vulnerability repair case library, wherein the historical repair rule feature vector includes code structure features, application scenario features, and side effect identifiers of the historical repair rule; Calculating a matching degree between the verification result report and the historical repair rule feature vector, and determining a target historical repair rule whose matching degree exceeds a preset matching degree threshold; The side effects of the target historical repair rule are counted to determine whether the repair rule has potential side effects.

6. The method according to claim 1, characterized in that After the step of pattern matching the optimal repair pattern with the vulnerable code snippet and generating a repair rule according to a preset code mapping relationship, the method further includes: Constructing an operating environment sandbox for testing the repair rules, and configuring operating parameters of the operating environment sandbox, wherein the operating parameters include program dependency configuration, environment variable settings, and resource constraints; Applying the repair rule to multiple test copies of the vulnerable code snippet to generate code to be tested; Executing the code to be tested in the operating environment sandbox to obtain performance loss data, memory usage data and exception information data during the execution process; The performance loss data, the memory usage data and the abnormal information data are analyzed according to preset evaluation rules to obtain a test evaluation result.

7. A computer device, characterized in that: The computer device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the computer device to execute the method according to any one of claims 1 to 6.

8. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on a computer device, the computer device is caused to perform the method according to any one of claims 1 to 6.

9. A computer program product, characterized in that When the computer program product is run on a computer device, the computer device is caused to perform the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Application program bug repair scheme generation method and device, equipment and medium

    CN118114250A

  • Vulnerability repairing method and system based on knowledge graph and large language model

    CN120012095A