Code vulnerability detection and repair method and system based on general large model

Through multi-channel RAG technology combining vector embedding, abstract syntax tree analysis and knowledge extraction methods, the efficiency and accuracy problems of code vulnerability detection and repair in the existing technology are solved, efficient code vulnerability detection and repair are achieved, and code quality and security are improved.

CN120372631APending Publication Date: 2025-07-25HARBIN INST OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510546012.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing technology lacks efficient and low-cost methods for code vulnerability detection and repair, resulting in frequent security problems in the software and lack of a domestically produced independent and controllable code audit system.

Method used

Multi-channel RAG technology is used to search the vulnerability library based on vector embedding, and the code structure features are analyzed using abstract syntax trees, combined with knowledge extraction and analysis of user intentions, generate preliminary repair solutions, and generate final detection and repair solutions through weighted integration.

Benefits of technology

It significantly improves the accuracy and efficiency of code vulnerability detection, reduces the rate of missed and false alarms, provides an effective fix solution, and improves code quality and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372631A_ABST
    Figure CN120372631A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of code vulnerability detection and restoration, and discloses a code vulnerability detection and restoration method and system based on a general large model, and the method comprises the steps: employing a multi-channel RAG technology, searching a vulnerability library based on a vector embedding method, and retrieving similar vulnerability entries; analyzing the user code structure features based on the abstract syntax tree to obtain predicted vulnerability information; analyzing user intention and code logic based on knowledge extraction to generate a preliminary repair scheme; and carrying out weighted integration on the similar vulnerability entries, the predicted vulnerability information and the preliminary repair scheme, and inputting the integrated results into the general large model to obtain a final detection result and a repair scheme. According to the method, the multi-channel RAG technology is used, related professional content is preferentially retrieved as auxiliary input, the capacity of a general large model in the aspects of code vulnerability detection and repair is improved, and therefore the code quality is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of code vulnerability detection and repair, and specifically relates to a method and system for code vulnerability detection and repair based on a general large model. Background Art

[0002] At present, on the one hand, during the software development process, developers often pay more attention to the development efficiency and results, ignoring the code quality and security issues. This tendency has led to frequent vulnerabilities in software code, and various software frequently experiences security problems such as program crashes and data leaks, causing serious losses. There is a lack of domestic independent and controllable code auditing systems.

[0003] On the other hand, network attack methods are constantly iterating and updating, the attack threshold and cost have been significantly reduced, and the enterprise security cost has increased significantly. Detecting and repairing code vulnerabilities during the software development stage has the best effect and can significantly reduce the later security maintenance cost. However, at the present stage, there is a lack of efficient and low-cost technologies and means to address the security protection issues during the development process.

[0004] At the level of technological development, with the rapid development of artificial intelligence technology, it has become feasible to use large models for code vulnerability detection and repair. At the same time, some studies have shown that the method of combining a general large model with professional data is, to a certain extent, superior to professional large models in terms of effect. This provides new technical ideas and methods for code vulnerability detection and repair.

[0005] Therefore, in order to ensure code quality and software security, while improving the degree of automation and reducing the security cost, a scientific method for code vulnerability detection and repair based on a general large model is needed. Summary of the Invention

[0006] To solve the problems existing in the prior art, the present invention provides a method and system for code vulnerability detection and repair based on a general large model, which uses multi-channel RAG technology to preferentially retrieve relevant professional content as auxiliary input to improve the ability of the general large model in code vulnerability detection and repair, thereby ensuring code quality.

[0007] To achieve the above object, the present invention provides the following solutions:

[0008] A method for code vulnerability detection and repair based on a general large model, the method comprising:

[0009] Adopting multi-channel RAG technology to search for similar vulnerability entries in the vulnerability library based on the method of vector embedding;

[0010] Based on the abstract syntax tree, parsing the structural features of the user code to obtain predicted vulnerability information;

[0011] Analyze the user's intention and code logic based on knowledge extraction to generate a preliminary repair plan;

[0012] Weightedly integrate similar vulnerability entries, predicted vulnerability information, and the preliminary repair plan, and input them into a general large model to obtain the final detection result and repair plan.

[0013] Preferably, search for similar vulnerability entries in the vulnerability database based on the vector embedding method, including:

[0014] Clean the data of the publicly authoritative vulnerability database and establish a structured knowledge base; among them, the publicly authoritative vulnerability database includes: CVE, CNNVD, CWE, OWASP TOP10;

[0015] Based on the structured knowledge base, on the basis of the text embedding model BGE, use the code search task dataset code_search_net to incrementally train the code embedding model CodeSecure for Python and C-based languages independently, and adopt the method of contrastive learning to optimize the model by narrowing the distance from the positive sample and expanding the distance from the negative sample;

[0016] Use the trained model CodeSecure to vectorize the user code and the knowledge base to construct a code embedding space;

[0017] Perform vector similarity matching in the code embedding space, and retrieve the top_k similar vectors and corresponding vulnerability entries in the knowledge base according to the matching degree with the user code.

[0018] Preferably, analyze the user's intention and code logic based on knowledge extraction to generate a preliminary repair plan, including:

[0019] Preset rule templates for typical vulnerabilities;

[0020] Use parso to convert the user code into a highly fault-tolerant AST tree and extract the structural features of the user code;

[0021] Based on the structural features of the user code, traverse the nodes of the AST tree and compare them with the typical vulnerability rule templates for rule checking to predict existing code vulnerabilities.

[0022] Preferably, analyze the user's intention and code logic based on knowledge extraction to generate a preliminary repair plan, including:

[0023] Establish a historical database to store various code defects detected in the user code;

[0024] Automatically extract key information from the user code and user intention through a large model;

[0025] Retrieve the historical database and extract similar historical repair cases;

[0026] Based on various code defects, key information, and similar historical repair cases, use the CoT technology to gradually reason and form a preliminary repair plan.

[0027] Preferably, weight and integrate similar vulnerability entries, predicted vulnerability information, and the preliminary repair plan, and input them into a general large model to obtain the final detection result and repair plan, including:

[0028] Determine the weights according to the confidence, historical performance, and data quality of similar vulnerability entries, predicted vulnerability information, and the preliminary repair plan, and evaluate the repair suggestions to confirm whether the overall reliability of the proposed plan reaches the preset threshold;

[0029] Input the weighted and integrated information into the general large model to further optimize and generate the final repair plan.

[0030] The present invention also provides a code vulnerability detection and repair system based on a general large model. The system is used to implement the foregoing method, and the system includes: a vector matching module, an AST parsing module, a knowledge extraction module, and a weighted integration module;

[0031] The vector matching module is used to search the vulnerability library for similar vulnerability entries based on the method of vector embedding by adopting the multi-channel RAG technology;

[0032] The AST parsing module is used to parse the structural features of the user code based on the abstract syntax tree to obtain predicted vulnerability information;

[0033] The knowledge extraction module is used to generate a preliminary repair plan based on knowledge extraction and analyze the user's intention and code logic;

[0034] The weighted integration module is used to weight and integrate similar vulnerability entries, predicted vulnerability information, and the preliminary repair plan, and input them into a general large model to obtain the final detection result and repair plan.

[0035] Preferably, the vector matching module includes: a knowledge base construction unit, a model training unit, a vectorization processing unit, and a retrieval unit;

[0036] The knowledge base construction unit is used to clean the data of the publicly available authoritative vulnerability library and establish a structured knowledge base; wherein, the publicly available authoritative vulnerability library includes: CVE, CNNVD, CWE, OWASP TOP10;

[0037] The model training unit is used to incrementally train the code embedding model CodeSecure for Python and C-based languages on the basis of the text embedding model BGE using the code search task dataset code_search_net based on a structured knowledge base, and adopt the method of contrastive learning to optimize the model by reducing the distance from positive samples and increasing the distance from negative samples;

[0038] The vectorization processing unit is used to vectorize the user code and the knowledge base using the trained model CodeSecure to construct a code embedding space;

[0039] The retrieval unit is used to match the vector similarity in the code embedding space and retrieve the top_k similar vectors and corresponding vulnerability entries in the knowledge base according to the matching degree with the user code.

[0040] Preferably, the AST parsing module includes: a rule preset unit, a feature extraction unit, and a comparison and prediction unit;

[0041] The rule preset unit is used to preset rule templates for typical vulnerabilities;

[0042] The feature extraction unit is used to convert the user code into a highly fault-tolerant AST tree using parso and extract the structural features of the user code;

[0043] The comparison and prediction unit is used to traverse the nodes of the AST tree based on the structural features of the user code, compare with the typical vulnerability rule templates for rule checking, and predict the existing code vulnerabilities.

[0044] Preferably, the knowledge extraction module includes: a database construction unit, a key information extraction unit, a case extraction unit, and an inference unit;

[0045] The database construction unit is used to establish a historical database to store various code defects detected in the user code;

[0046] The key information extraction unit is used to automatically extract key information from the user code and user intent through a large model;

[0047] The case extraction unit is used to retrieve the historical database and extract similar historical repair cases;

[0048] The inference unit is used to gradually reason based on various code defects, key information, and similar historical repair cases using the CoT technology to form a preliminary repair plan.

[0049] Preferably, the weighted integration module includes: an evaluation unit and an optimization unit;

[0050] The evaluation unit is used to determine weights based on similar vulnerability entries, predicted vulnerability information, and the confidence, historical performance, and data quality of the preliminary repair plan, and evaluate the repair suggestions to confirm whether the overall reliability of the proposed plan reaches a preset threshold.

[0051] The optimization unit is used to input the weighted and integrated information into a general large model to further optimize and generate the final repair plan.

[0052] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0053] The present invention provides a method and system for code vulnerability detection and repair based on a general large model, which significantly improves the performance of the general large model in code vulnerability security detection tasks through multi-channel RAG technology, thereby further improving the efficiency and accuracy of code auditing.

[0054] The present invention relies on an authoritative vulnerability database and a user history database to provide data support; relying on the processing of the first and second channels can effectively and comprehensively detect vulnerabilities in the code, reducing the false negative and false positive rates; relying on the processing of the third channel can combine user intentions to explain and repair complex code vulnerabilities; relying on the powerful analysis and reasoning ability of the existing general large model can propose effective repair plans for vulnerabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0056] Figure 1 It is a schematic flow chart of the method for code vulnerability detection and repair based on a general large model according to an embodiment of the present invention;

[0057] Figure 2 It is an architecture diagram of the first channel according to an embodiment of the present invention;

[0058] Figure 3 It is an architecture diagram of the second channel according to an embodiment of the present invention;

[0059] Figure 4 It is an architecture diagram of the third channel according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0061] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0062] Embodiment 1

[0063] The code vulnerability detection and repair method based on a general large model proposed in this embodiment uses the multi-channel RAG technology to improve the ability of the general large model in code vulnerability detection and security, thereby providing more accurate defect detection and repair suggestions.

[0064] The RAG (Retrieval-Augmented Generation) technology refers to preferentially retrieving relevant professional content and inputting it as auxiliary knowledge to the large model. The multi-channel means extracting knowledge from multiple reliable information sources and jointly making a weighted decision to generate a repair plan. The information sources are divided into three. The first channel is the vulnerability matching channel based on code embedding. The second channel is the structural feature channel based on strongly fault-tolerant AST parsing. The third channel is the preliminary solution channel based on knowledge extraction. The specific implementation process is as follows:

[0065] As Figure 1 shown, the code vulnerability detection and repair method based on a general large model, the method includes:

[0066] Using the multi-channel RAG technology, searching for similar vulnerability entries in the vulnerability library based on the method of vector embedding;

[0067] Based on the abstract syntax tree, parsing the structural features of the user code to obtain predicted vulnerability information;

[0068] Based on knowledge extraction, analyzing the user's intention and code logic to generate a preliminary repair plan;

[0069] Weightedly integrating the similar vulnerability entries, predicted vulnerability information, and preliminary repair plan, and inputting them to the general large model to obtain the final detection result and repair plan.

[0070] In this embodiment, the first channel searches for similar vulnerability entries in the vulnerability library based on the method of vector embedding, including:

[0071] Performing data cleaning on public and authoritative vulnerability libraries (such as CVE, CNNVD, CWE, OWASP TOP10, etc.) to establish a structured knowledge base.

[0072] Based on the text embedding model BGE, the code_search_net dataset for code search tasks is used to incrementally train the CodeSecure code embedding model for Python and C-based languages. And a contrastive learning method is adopted to optimize the model by reducing the distance to positive samples and increasing the distance to negative samples.

[0073] The trained CodeSecure model is used to vectorize the user code and the knowledge base to construct a code embedding space.

[0074] In the embedding space, vector similarity matching is performed, and the top_k similar vectors and corresponding vulnerability entries in the knowledge base are retrieved according to the matching degree with the user code as the output information of the first channel.

[0075] As Figure 2 shown below, the following are the detailed steps: First channel: The vulnerability matching channel based on code embedding, including:

[0076] Step 1: Data cleaning and format unification

[0077] Retrieve the original data from public and authoritative vulnerability databases (such as CVE, CNNVD, CWE, OWASP TOP10), clean the data, remove redundant information, correct the format, and unify the vulnerability description format to form structured data records.

[0078] Step 2: Model training

[0079] The bge-small-en-v1.5 model is incrementally trained using the code_search_net dataset, and a contrastive learning method is adopted to train the Code Embedding model CodeSecure for multiple languages such as Python and C. This cross-programming language embedding model helps in project-level code processing.

[0080] The contrastive learning fine-tuning is implemented using a temperature-adjusted InfoNCE loss function to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs. Its mathematical form is defined as:

[0081]

[0082] In the formula, c i is the embedding vector of the anchor sample, t i is the embedding vector of the positive sample, t j is the embedding vector of the negative sample, s(·,·) is the cosine similarity function, N is the number of negative samples, and τ is the temperature parameter.

[0083] Adjust the temperature parameter by calculating the gradient norm in real time to ensure training stability:

[0084]

[0085] In the formula, g is the gradient vector of the model parameters.

[0086]

[0087] In the formula, α is the attenuation coefficient.

[0088] Step 3: Code embedding

[0089] 1. Vectorize the knowledge base

[0090] Use the structured data records obtained in the previous steps and the embedding model to embed the vulnerability entries in the vulnerability library, and construct the vector space of the vulnerability knowledge base, so that each vulnerability entry has a corresponding vector, which is convenient for subsequent similarity matching.

[0091] Let the set of vulnerability entries be where d i is the structured vulnerability description. Construct a mapping through the embedding encoder ε(·):

[0092]

[0093] In the formula, is the set of real numbers, and d emb is the embedding dimension.

[0094] 2. Vectorize the user code

[0095] The user uploads the code, and converts the code snippet into a structured snippet through preprocessing (removing spaces, code slicing, comment standardization, etc.), and uses the trained CodeSecure model to vectorize the user's code snippet to be detected. Obtain the vector representation of the user's code.

[0096] Similar to the vectorization process of vulnerability entries, let the set of user code snippets be where d i is the structured snippet of the user's code. Construct a mapping through the embedding encoder ε(·):

[0097]

[0098] In the formula, is the set of real numbers, and d emb is the embedding dimension.

[0099] Step 4: Vector matching

[0100] The cosine similarity between the user code vector and each vulnerability entry vector is calculated in the vector space of the knowledge base to match similar vulnerability entries.

[0101] The formula for calculating cosine similarity is:

[0102]

[0103] Where V user is the embedding vector of the user code, V cve is the embedding vector of the vulnerability entry.

[0104] According to the calculated cosine similarity value, large-scale retrieval with V based on the approximate nearest neighbor algorithm (ANN) user The most similar Top-K vulnerability records.

[0105]

[0106] In the formula, arg represents the parameter mapping relationship, which is used to find the objective function The parameter combination that reaches the maximum value, Indicates from the complete set The candidate subset selected from j Indicates the cosine similarity of the vulnerability entries. K indicates selecting the most similar K vulnerability entries (the default value is 3).

[0107] The confidence is calculated based on the degree of dispersion of the similarity of the Top-K similar vectors.

[0108]

[0109] In the formula, is the average value of the cosine similarity of the Top-K similar vectors, σ s is the standard deviation of the cosine similarity of the Top-K similar vectors.

[0110] like Figure 3 As shown, in this embodiment, the second channel parses the user code structure features based on the abstract syntax tree to obtain predicted vulnerability information, including:

[0111] Define rule templates for typical vulnerabilities.

[0112] Parso is used to convert user code into a highly fault-tolerant AST tree and extract the structural features of the user code.

[0113] The nodes of the AST tree are traversed and compared with typical vulnerability rule templates to perform rule checks and predict possible code vulnerabilities as the output information of the second channel.

[0114] The following are the detailed steps: Two-channel: The structural feature channel based on strongly fault-tolerant AST parsing, including:

[0115] Step 1: Definition of typical rule templates

[0116] Preprocess typical vulnerabilities, and through the predicate logic structured rule template, realize the formal expression of vulnerability patterns. Define the vulnerability rule as:

[0117]

[0118] In the formula, R k represents the vulnerability rule, φ i (v) represents the logical condition check of the type or attribute of the AST node v, Type(v) returns the type of node v, and Attr(v,α j ) returns the attribute of node v.

[0119] For example, the CWE-676 dangerous function call template can be simply defined as:

[0120] R CWE-676 =(Type(v)=CallExpr)∧(Attr(v,function name )∈(system,exec))

[0121] It means that the vulnerability node type of the dangerous function call class needs to be CallExpr, and the function name is system or exec.

[0122] Step 2: Convert code to AST tree

[0123] Use a highly fault-tolerant AST parser (such as parso) to convert the user code snippet into an abstract syntax tree (AST), ensuring that even when there are syntax errors or non-standard formats, a correct tree structure can be generated as much as possible, and calculate the parsing confidence.

[0124] The formula for calculating the parsing confidence is:

[0125] P(T|S)=∏ (A→β)∈Productions P(A→β) c(A→β) (10)

[0126] In the formula, S is the input code (allowing syntax errors), T is the generated AST, c(A→β) is the number of times the production rule (A→β) is used in parsing, and Productions represents the set of production rules.

[0127] For symbols that cannot be matched, insert ErrorNode and allocate confidence p error .

[0128] Step 3: Traverse and match

[0129] Traverse the nodes in the AST tree for rule comparison, output the detected vulnerability types and corresponding rule matching information, calculate the confidence level, and supplement it to the candidate vulnerability information.

[0130] First, for the AST node v, calculate the rule matching degree:

[0131]

[0132] In the formula, is the indicator function (1 for satisfying the constraint, 0 otherwise), and p (v) is the parsing confidence level of node v.

[0133] Then, for the set of rules {R k} that match the AST tree, calculate the final confidence level:

[0134]

[0135] In the formula, K represents the number of successfully matched vulnerability rules, and M is the rule matching degree calculated previously.

[0136] As Figure 4 shown, in this embodiment, the three channels generate a preliminary repair plan based on knowledge extraction and analysis of the user's intention and code logic, including:

[0137] Establish a historical database to store various code defects detected in the user's code.

[0138] Automatically extract key information from the user's code and user intention through a large model.

[0139] Retrieve the historical database and extract a small number of similar historical repair cases.

[0140] Based on the key information and a small number of examples, gradually reason with the help of the CoT (Chain of Thought) technology to form a preliminary repair plan as the output information of the three channels.

[0141] The following is a detailed description of the steps: Three channels: The preliminary plan channel based on knowledge extraction, including:

[0142] Step 1: Key information extraction

[0143] Use a large model for natural language processing on the code and comments uploaded by the user, and extract key information and risk points based on the attention mechanism.

[0144] For example: For dangerous function calls, the system extracts key information such as "calling external commands" and "user input not verified".

[0145] Step 2: Historical Case Retrieval

[0146] Use the code embedding model pre-trained on Channel 1 to encode historical cases and construct a mapping:

[0147]

[0148] where d i represents the historical case entry.

[0149] Then, retrieve the repair cases with similar vulnerabilities that have occurred in the historical database according to the cosine similarity, extract the similar code snippets and repair suggestions, and calculate the confidence of the retrieval results based on the similarity and the historical repair success rate.

[0150]

[0151] where s i represents the similarity, and r i represents the historical repair success rate of this repair case.

[0152] For example: For the above case, it is retrieved that in the historical cases, it was once recommended to replace os.system() with subprocess.run() for the command injection risk and add input validation measures.

[0153] Step 3: With the help of the Chain of Thought technique (CoT) of the large model, combine the extracted information and historical cases, and gradually reason to form a preliminary repair plan. At the same time, add a reflection mechanism to generate a confidence score for the repair plan according to the historical repair success rate. If the score is lower than the threshold, trigger the regeneration of the repair steps.

[0154] The large model generates a repair plan step by step, modeled as a chain conditional probability:

[0155]

[0156] where T represents the total number of steps to generate the repair plan, Context includes key information and historical repair cases, and Steps 1:i-1 represents the sequence of generated repair steps.

[0157] The confidence calculation formula for the repair plan is as follows:

[0158]

[0159] where r(s t ) represents the historical repair success rate corresponding to this repair step.

[0160] For example: Based on the above information, the large model recommends: Modify the code to call subprocess.run() and add parameter validation.

[0161] In this embodiment, similar vulnerability entries, predicted vulnerability information, and preliminary repair solutions are weighted and integrated, and then input into a general large model to obtain the final detection result and repair solution, including:

[0162] Determine the weights according to the confidence levels, historical performances, and data qualities of the similar vulnerability entries, predicted vulnerability information, and preliminary repair solutions, and evaluate the repair suggestions to confirm whether the overall reliability of the proposed solution reaches a preset threshold;

[0163] Input the weighted and integrated information into the general large model to further optimize and generate the final repair solution.

[0164] The following are the detailed steps: Joint module: Weighted integration of multi-channel information, including:

[0165] Step 1: Weighted integration

[0166] Weight and integrate the information output from the first channel (vector matching), the second channel (AST parsing), and the third channel (knowledge extraction). Determine the weights according to the confidence levels, historical performances, and data qualities of each channel. And evaluate the repair suggestions to confirm that the overall reliability of the proposed solution reaches a preset threshold (such as 0.85 or above).

[0167] C total = ω1C emb + ω2C AST + ω3C hist + ω4C sol (17)

[0168] Step 2: Determination of the final solution

[0169] Input the weighted and integrated information into the general large model to further optimize and generate the final repair solution.

[0170] For example: After comprehensive reasoning, the large model confirms that the preliminary solution is reasonable, adds comment explanations and boundary condition detections to the code, and finally outputs the corresponding repair code.

[0171] In this embodiment, subsequent verification and database update include:

[0172] Step 1: Conduct static analysis and necessary tests on the repaired code to ensure that the vulnerability has been effectively repaired and no new problems are introduced.

[0173] Step 2: Record the whole process of this vulnerability detection and repair, the rules used, the weighted information, and the final repair code in the historical database for faster and more accurate handling of similar vulnerabilities in the future.

[0174] For example, store the detection and repair cases for command injection in this time as structured records, including information such as detection time, code samples, weighted scores, repair solutions, etc., for subsequent query and continuous optimization of the model.

[0175] Embodiment 2

[0176] In the field of software security, the detection and repair of source code vulnerabilities are crucial. To ensure code quality and reduce audit costs, the present invention proposes a solution based on a large model combined with multi-channel RAG technology.

[0177] In terms of vulnerability detection, other related patents have provided diverse ideas for solving similar problems from different technical paths. These solutions have both differences and complementarities with the present invention in terms of technical means, application scenarios, etc. The following are alternative solutions for vulnerability detection summarized from other documents:

[0178] 1. Vulnerability detection solution based on graph structure analysis

[0179] Construct a multi-level graph for analysis. Generate an abstract syntax tree by parsing the source code, and then build a multi-level graph including a function call graph, a control flow graph, and a data flow graph. Use graph theory algorithms, such as maximum flow minimum cut calculation, connected component analysis, random walk calculation, etc., to detect potential vulnerabilities and generate reports.

[0180] A detection method integrating a code property graph. After obtaining the source code data, extract code representations and convert them into a fused code property graph, which combines various semantic information. Input the graph encoding features into a vulnerability detection model to obtain detection results, and then use an interpretable model to interpret the results.

[0181] 2. Vulnerability detection solution based on feature learning

[0182] Multi-dimensional feature fusion detection. First, represent the source code as a program dependence graph and slice it, and then perform standardization and embedding operations. Perform embedding learning on the sliced graph in the hyperbolic space and the Euclidean space respectively to capture vulnerability features from different perspectives, and then fuse the features of the two spaces and input them into a classification prediction module for vulnerability judgment.

[0183] Detection combining sequence and graph vectors. After obtaining the source code of the function to be detected, determine its first token sequence vector and first code property graph vector, and input these vectors into a source code vulnerability detection model to extract sequence features and graph features to determine the vulnerability status.

[0184] 3. Vulnerability detection solution based on model training optimization

[0185] Training of heterogeneous graph neural network model: Convert the training program code into a program dependence graph and an abstract syntax tree, generate a vulnerability detection graph based on preset key node types, and perform fine-grained classification on these graphs to obtain a training set for training the heterogeneous graph neural network.

[0186] Construction and training of large model: Train a large language model based on preset code submission information. First, use vulnerability repair information to generate the first prompt to train the initial large language model, and then optimize it based on the reward function to obtain the second large language model. Analyze the value dependence graph through a preset symbolic calculation algorithm to determine the initial vulnerability detection result, and then determine the noise transfer matrix to generate the second prompt for training the target large language model.

[0187] 4. Vulnerability detection solution combining static and dynamic analysis

[0188] Extract static data flow according to the influence degree of the generated code vulnerability on the code text of the code to be detected, and at the same time obtain the dynamic data flow generated during the code running process. Input these two data flows into the code detection model, which includes a vector extraction model, a text convolution model, and a classifier, and obtain the vulnerability detection result through a series of processes.

[0189] In terms of vulnerability repair, the method of manual repair can be adopted to achieve the repair purpose.

[0190] Embodiment III

[0191] The present invention also provides a code vulnerability detection and repair system based on a general large model. The system is used to implement the method of Embodiment I, and the system includes: a vector matching module, an AST parsing module, a knowledge extraction module, and a weighted integration module;

[0192] The vector matching module is used to search for similar vulnerability entries in the vulnerability library by using the multi-channel RAG technology based on the method of vector embedding;

[0193] The AST parsing module is used to parse the structural features of the user code based on the abstract syntax tree to obtain predicted vulnerability information;

[0194] The knowledge extraction module is used to generate a preliminary repair plan based on knowledge extraction to analyze the user's intention and code logic;

[0195] The weighted integration module is used to weight and integrate the similar vulnerability entries, predicted vulnerability information, and preliminary repair plan, and input them into the general large model to obtain the final detection result and repair plan.

[0196] In this embodiment, the vector matching module includes: a knowledge base construction unit, a model training unit, a vectorization processing unit, and a retrieval unit;

[0197] A knowledge base construction unit for cleaning data from publicly available authoritative vulnerability databases and establishing a structured knowledge base. The publicly available authoritative vulnerability databases include CVE, CNNVD, CWE, and OWASP TOP10.

[0198] A model training unit for incrementally training a code embedding model CodeSecure for Python and C-based languages autonomously on the basis of a structured knowledge base using the code search task dataset code_search_net in the text embedding model BGE, and optimizing the model by adopting a contrastive learning method to reduce the distance from positive samples and increase the distance from negative samples.

[0199] A vectorization processing unit for vectorizing user code and the knowledge base using the trained model CodeSecure to construct a code embedding space.

[0200] A retrieval unit for matching vector similarities in the code embedding space and retrieving the top_k similar vectors and corresponding vulnerability entries in the knowledge base according to the matching degree with the user code.

[0201] In this embodiment, the AST parsing module includes a rule preset unit, a feature extraction unit, and a comparison and prediction unit.

[0202] The rule preset unit for presetting rule templates for typical vulnerabilities.

[0203] The feature extraction unit for converting user code into a highly fault-tolerant AST tree using parso and extracting the structural features of the user code.

[0204] The comparison and prediction unit for traversing the nodes of the AST tree based on the structural features of the user code, comparing with the typical vulnerability rule templates for rule checking, and predicting existing code vulnerabilities.

[0205] In this embodiment, the knowledge extraction module includes a database construction unit, a key information extraction unit, a case extraction unit, and an inference unit.

[0206] The database construction unit for establishing a historical database to store various code defects detected in user code.

[0207] The key information extraction unit for automatically extracting key information from user code and user intentions through a large model.

[0208] The case extraction unit for retrieving the historical database and extracting similar historical repair cases.

[0209] A reasoning unit for gradually reasoning using the CoT technology based on various code defects, key information, and similar historical repair cases to form a preliminary repair plan.

[0210] In this embodiment, the weighted integration module includes: an evaluation unit and an optimization unit;

[0211] The evaluation unit is used to determine weights according to similar vulnerability entries, predicted vulnerability information, and the confidence, historical performance, and data quality of the preliminary repair plan, evaluate the repair suggestions, and confirm whether the overall reliability of the proposed plan reaches a preset threshold;

[0212] The optimization unit is used to input the weighted integrated information into a general large model to further optimize and generate a final repair plan.

[0213] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for code vulnerability detection and repair based on a general large model, characterized in that, The method includes: Adopting the multi-channel RAG technology, searching the vulnerability database based on the vector embedding method to retrieve similar vulnerability entries; Parsing the structural features of the user code based on the abstract syntax tree to obtain predicted vulnerability information; Generating a preliminary repair plan based on knowledge extraction to analyze the user's intention and code logic; Weightedly integrating the similar vulnerability entries, predicted vulnerability information, and preliminary repair plan, and inputting them into a general large model to obtain the final detection result and repair plan.

2. The method according to claim 1, characterized in that Searching the vulnerability database based on the vector embedding method to retrieve similar vulnerability entries, including: Cleaning the data of the publicly authoritative vulnerability database to establish a structured knowledge base; among them, the publicly authoritative vulnerability database includes: CVE, CNNVD, CWE, OWASP TOP10; Based on the structured knowledge base, on the basis of the text embedding model BGE, using the code search task dataset code_search_net to independently and incrementally train the code embedding model CodeSecure for Python and C-based languages, and adopting the contrastive learning method to optimize the model by reducing the distance to the positive sample and increasing the distance to the negative sample; Using the trained model CodeSecure to vectorize the user code and the knowledge base to construct a code embedding space; Performing vector similarity matching in the code embedding space, and retrieving the top_k similar vectors and corresponding vulnerability entries in the knowledge base according to the matching degree with the user code.

3. The method according to claim 1, characterized in that, Parsing the structural features of the user code based on the abstract syntax tree to obtain predicted vulnerability information, including: Presetting rule templates for typical vulnerabilities; Using parso to convert the user code into a highly fault-tolerant AST tree and extract the structural features of the user code; Based on the structural features of the user code, traversing the nodes of the AST tree, comparing with the typical vulnerability rule templates for rule checking, and predicting the existing code vulnerabilities.

4. The method according to claim 1, wherein Generating a preliminary repair plan based on knowledge extraction to analyze the user's intention and code logic, including: Establishing a historical database to store various code defects detected in the user code; Automatically extracting key information from the user code and user intention through a large model; Searching the historical database to extract similar historical repair cases; Based on various code defects, key information, and similar historical repair cases, using the CoT technology to gradually reason and form a preliminary repair plan.

5. The method according to claim 1, wherein Weightedly integrating the similar vulnerability entries, predicted vulnerability information, and preliminary repair plan, and inputting them into a general large model to obtain the final detection result and repair plan, including: Determining the weights according to the confidence, historical performance, and data quality of the similar vulnerability entries, predicted vulnerability information, and preliminary repair plan, and evaluating the repair suggestions to confirm whether the overall reliability of the recommended plan reaches the preset threshold; Inputting the weightedly integrated information into a general large model to further optimize and generate the final repair plan.

6. A code vulnerability detection and repair system based on a general large model, the system being used to implement the method described in any one of claims 1-5, characterized in that, The system includes: a vector matching module, an AST parsing module, a knowledge extraction module, and a weighted integration module; The vector matching module is used to adopt the multi-channel RAG technology, search the vulnerability database based on the vector embedding method to retrieve similar vulnerability entries; The AST parsing module is used to parse the structural features of user code based on the abstract syntax tree to obtain predicted vulnerability information; The knowledge extraction module is used to generate a preliminary repair plan based on knowledge extraction to analyze user intentions and code logic; The weighted integration module is used to weight and integrate similar vulnerability entries, predicted vulnerability information, and the preliminary repair plan, and input them to the general large model to obtain the final detection result and repair plan.

7. The system according to claim 6, characterized in that, The vector matching module includes: a knowledge base construction unit, a model training unit, a vectorization processing unit, and a retrieval unit; The knowledge base construction unit is used to clean the data of the publicly available authoritative vulnerability database and establish a structured knowledge base; among them, the publicly available authoritative vulnerability database includes: CVE, CNNVD, CWE, OWASP TOP10; The model training unit is used to, based on the structured knowledge base, on the basis of the text embedding model BGE, use the code search task dataset code_search_net to independently and incrementally train the code embedding model CodeSecure for Python and C series languages, and adopt the method of contrastive learning to optimize the model by reducing the distance to positive samples and increasing the distance to negative samples; The vectorization processing unit is used to vectorize the user code and the knowledge base using the trained model CodeSecure to construct a code embedding space; The retrieval unit is used to perform vector similarity matching in the code embedding space and retrieve the top_k similar vectors and corresponding vulnerability entries in the knowledge base according to the matching degree with the user code.

8. The system according to claim 6, wherein The AST parsing module includes: a rule preset unit, a feature extraction unit, and a comparison and prediction unit; The rule preset unit is used to preset rule templates for typical vulnerabilities; The feature extraction unit is used to convert the user code into a highly fault-tolerant AST tree using parso and extract the structural features of the user code; The comparison and prediction unit is used to traverse the nodes of the AST tree based on the structural features of the user code, compare with the typical vulnerability rule template for rule checking, and predict existing code vulnerabilities.

9. The system according to claim 6, wherein The knowledge extraction module includes: a database construction unit, a key information extraction unit, a case extraction unit, and an inference unit; The database construction unit is used to establish a historical database to store various code defects detected in the user code; The key information extraction unit is used to automatically extract key information from the user code and user intentions through a large model; The case extraction unit is used to retrieve the historical database and extract similar historical repair cases; The inference unit is used to gradually reason based on various code defects, key information, and similar historical repair cases using the CoT technology to form a preliminary repair plan.

10. The system according to claim 6, wherein The weighted integration module includes: an evaluation unit and an optimization unit; The evaluation unit is used to determine the weights according to the confidence, historical performance, and data quality of the similar vulnerability entries, predicted vulnerability information, and the preliminary repair plan, evaluate the repair suggestions, and confirm whether the overall reliability of the recommended plan reaches the preset threshold; The optimization unit is used to input the weighted and integrated information into the general large model to further optimize and generate the final repair solution.

Citation Information

Cited By

  • Method and system for generating open source component repair opinions and storage medium

    CN120611388A

  • Source code bug repairing method, electronic equipment and storage medium

    CN121479795A

  • Intelligent contract vulnerability detection method and system based on heterogeneous graph and large model

    CN122471465A