Similarity comparison-based large model supply chain automatic repair method and device
Through dual-channel encoding and LoRA fine-tuning of the large model, combined with similarity comparison technology, the automatic repair of large-model supply chain vulnerabilities is achieved, solving the problems of low efficiency and insufficient security in vulnerability identification and repair in existing technologies, and improving the automation and explainability of repairs.
Patent Information
- Application Number
- CN202510577169.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies make it difficult to efficiently and accurately identify and automatically repair vulnerabilities in underlying software systems in large-model supply chains, especially cross-library and cross-module vulnerabilities. Existing methods also have problems such as incomplete repairs, strong manual dependence, and insufficient security and explainability.
A large-model supply chain automatic repair method based on similarity comparison is adopted. The semantic and structural features of the vulnerability source code are extracted using a dual-channel encoding mechanism. The large model fine-tuned by LoRA is combined to classify vulnerability types and identify unknown categories. Repair patches are generated through similarity search and their effectiveness is verified.
It achieves efficient, accurate and automatic repair of large-model supply chain vulnerabilities in a local environment, improves the automation and security of repairs, and ensures the functionality and explainability of repair patches.
Smart Images

Figure CN120654241A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to fields such as artificial intelligence security technology, large-model supply chain security, and software security analysis, and in particular to a large-model supply chain automatic repair method and device based on similarity comparison. Background Art
[0002] With the advancement of artificial intelligence (AI) technology, large-scale pre-trained models (including large language models) have been widely used across various industries. The development and deployment of these models rely on a complex supply chain ecosystem, including open source model libraries, third-party components, plug-ins, and custom integration code. This large-model supply chain significantly expands the capabilities and applicability of models, but also introduces new security risks: vulnerabilities in underlying software components can be exploited by attackers to compromise the entire large-model system, resulting in serious security consequences.
[0003] Currently, the industry's focus on large-scale model security is primarily on the content security of the models themselves (such as adversarial attack patterns, prompt word escapes, and backdoor implants), while insufficient attention is paid to vulnerabilities in the underlying software systems that support large-scale model operations. However, in reality, numerous security vulnerabilities exist across all links of the large-scale model supply chain. Research has shown that these vulnerabilities are primarily located in application and model-layer code, with improper input validation and resource control being common root causes. It is worth noting that even after vulnerability patches have been released, some patches are incomplete, leading to recurring vulnerabilities. This shows that in the large-scale model supply chain environment, effectively detecting and thoroughly remediating underlying code vulnerabilities presents considerable challenges.
[0004] In this context, existing software vulnerability detection and remediation methods have exposed numerous shortcomings. Traditional static analysis tools are typically based on fixed rules or a single project scope, making it difficult to detect cross-library and cross-module supply chain vulnerabilities. They can detect some issues but are incapable of addressing new vulnerabilities. Furthermore, they often only indicate vulnerability locations, leaving developers to manually perform remediation, which is time-consuming and labor-intensive. Methods such as Software Composition Analysis (SCA) can identify known vulnerabilities and recommend component upgrades, but are ineffective for addressing unknown vulnerabilities within a project's own code. Recently emerging large-model assisted programming (such as GitHub Copilot) and chat-based models (such as ChatGPT and Deepseek) have also been explored for code vulnerability remediation. However, sending source code to the cloud raises privacy and security concerns, and general-purpose large models lack specialized security training, making their patches unreliable or difficult to interpret. Therefore, there is an urgent need for a new approach that combines program analysis with large-model intelligence to automatically and accurately classify and remediate large-model supply chain code vulnerabilities in a local environment. Summary of the Invention
[0005] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.
[0006] This paper proposes a large-scale supply chain automatic repair method based on similarity comparison. This method uses a dual-channel encoding mechanism (natural language semantic encoding and code structure summary encoding) to accurately model the semantic and structural features of the vulnerable source code and its context. Based on the Retrieval-Augmented Generation (RAG) framework, combined with a locally deployed large model fine-tuned by LoRA, it achieves efficient and accurate classification and automatic repair of large-scale supply chain code vulnerabilities. This method can automate large-scale supply chain vulnerability repair while ensuring the security, functionality, and interpretability of the repair patches.
[0007] Another object of the present invention is to provide a large-model supply chain automatic repair device based on similarity comparison.
[0008] To achieve the above objectives, the present invention proposes, on one hand, a large-scale model supply chain automatic repair method based on similarity comparison, comprising:
[0009] Use the pre-trained code model to extract natural language-level semantic representation vectors from the original vulnerability source code to be fixed;
[0010] Extract the structural summary information of the vulnerability source code and transform it into a patch-aware structural vector through a structural embedding method, and combine it with the semantic representation vector to obtain a dual-channel code representation;
[0011] The dual-channel code representation is input into a large model fine-tuned based on LoRA to automatically classify CWE vulnerability types, and zero-shot recognition of unknown categories is performed based on a semantic matching strategy to obtain vulnerability classification results.
[0012] Using the vulnerability classification result as a retrieval filter condition to filter out a case set related to the current vulnerability type from a local vector database, and using the dual-channel code representation as a query vector to perform a similarity search on the case set to retrieve patches for similar vulnerability cases;
[0013] Extract patch summaries of similar vulnerability case patches, and combine the current vulnerability code with the vulnerability classification results to construct prompt information to drive the large model to generate repair patches;
[0014] The prompt information is used as input to call the large model based on LoRA fine-tuning to generate a repair patch code and repair instructions for the current vulnerability;
[0015] The validity of the repair patch code is verified. If the verification passes, the relevant data of this repair process is structured and archived in the local knowledge base.
[0016] The large-model supply chain automatic repair method based on similarity comparison in an embodiment of the present invention may also have the following additional technical features:
[0017] In one embodiment of the present invention, a pre-trained code model is used to extract a natural language-level semantic representation vector from the original vulnerability source code to be repaired, including:
[0018] The input Python code snippet to be repaired is recorded as:
[0019] C=RawCode(l1,l2,…,l n )
[0020] Input the source code sequence C into the tokenizer of the pre-trained code model to obtain the token sequence:
[0021] T=[t1,t2,…,t m ],t i ∈V
[0022] Perform position encoding and embedding mapping, let E:V→R d is the embedding function, P:N→R d Generate the initial input vector sequence for the position encoding function:
[0023] X=[x1,x2,…,x m ],x i =E(t i )+P(i)
[0024] Input m into the encoder model and perform a multi-layer Transformer encoding process to obtain the hidden state representation corresponding to each token:
[0025] H=[h1,h2,…,h m ],h i ∈R d
[0026] The output vector at the global summary position is selected as the semantic representation of the entire code snippet, which is recorded as:
[0027] V sem =h CLS ∈R d
[0028] Among them, l i Represents the i-th line of source code, with a total length of n lines; V is the model vocabulary, and m represents the number of tokens.
[0029] In one embodiment of the present invention, the structure summary information of the vulnerability source code is extracted and converted into a patch-aware structure vector through a structure embedding method, and then combined with the semantic representation vector to obtain a dual-channel code representation, including:
[0030] Convert the code snippet into an abstract syntax tree T through the program parsing library AST :
[0031] T AST =(V,E)
[0032] Generate additional control flow graph of code as needed, namely G CFG , or data flow graph G DFG , written as:
[0033] G CFG =(N cfg ,F cfg ),G DFG =(N dfg ,F dfg )
[0034] Through graph representation learning or structural embedding methods, the structural information is converted into structural vector representation:
[0035] v struct =f struct (T AST ,G CFG ,G DFG )
[0036] The semantic vector V sem and structure vector v struct Concatenated for patch-aware representation:
[0037] v code =[V sem ;v struct ]∈R 2d
[0038] Where V is the set of syntax tree nodes, E is the set of edges between nodes, N is the set of nodes in the control flow graph or data flow graph, and F is the set of edges in the control flow graph or data flow graph; f struct Representation graph representation learning function.
[0039] In one embodiment of the present invention, the dual-channel code representation is input into a large model fine-tuned based on LoRA to automatically classify CWE vulnerability types, and zero-sample recognition of unknown categories is performed based on a semantic matching strategy to obtain vulnerability classification results, including:
[0040] Assume that the complete set of CWE labels covered by the local knowledge base is:
[0041] YCWE =[y1,y2,…,y k ]
[0042] Among them, y i Represents a CWE tag type;
[0043] For the input vulnerability code, the patch-aware representation vector v code ,Prediction is performed through a large model deployed locally and fine-tuned by LoRA, outputting independent confidence scores for each CWE label;
[0044] For each CWE label y i Extract official description text d i , and use the encoder f embed Generate semantic embedding u i :
[0045] u i =f embed (d i )
[0046] where u i ∈R d CWE label y i Semantic embedding of
[0047] Calculate the semantic embedding and semantic vector representation V of the vulnerability code to be classified sem The semantic similarity between them, the similarity score of i is s i Finally, the top N CWE labels with the largest similarity are selected as the predicted output Y pre d :
[0048] Y pred =TopN i (s i )
[0049] Combine the output results of the multi-label classification model and zero-sample matching, and use weighted fusion to determine the final predicted label set Y final .
[0050] In one embodiment of the present invention, the vulnerability classification result is used as a retrieval filter to filter out a case set related to the current vulnerability type from a local vector database, and a similarity search is performed on the case set using the dual-channel code representation as a query vector to retrieve patches for similar vulnerability cases, including:
[0051] The code represents v code Retrieve in the database as a query vector:
[0052] D topK =argTopK i{sim(v code ,code i )}
[0053] Among them, code i is the historical vulnerability code representation stored in the database, sim() is the cosine similarity function, and the retrieval result D topK Contains the K historical cases that are closest to the input.
[0054] In one embodiment of the present invention, patch summaries of similar vulnerability case patches are extracted, and combined with the current vulnerability code and vulnerability classification results to construct prompt information for driving the large model to generate repair patches, including:
[0055] Summarize the core repair ideas of historical patch cases into a patch summary description set P ref :
[0056] P ref ={p1,p2,…,p k}
[0057] where p i represents the patch summary in the i-th similar case;
[0058] Construct a hint template P based on the patch summary description set prompt To form the required prompt information.
[0059] In one embodiment of the present invention, the prompt information is used as input to call the large model based on LoRA fine-tuning to generate a repair patch code and repair instruction information for the current vulnerability, including:
[0060] Input the prompt information into the local large language model fine-tuned by LoRA, and automatically generate the patch code for fixing the vulnerability and the explanation of the repair principle:
[0061] Patch gen =LoRA-Generate(P prompt )
[0062] Final model output Patch gen Includes vulnerability patch code and repair principle explanation.
[0063] In one embodiment of the present invention, the validity of the repair patch code is verified. If the verification passes, the relevant data of the repair process is structured and archived in a local knowledge base, including:
[0064] After verifying the effectiveness of the automatically generated patch, the patch code is applied to the actual project to form a closed loop for vulnerability repair. At the same time, the repair process is automatically structured and archived and stored in the local vulnerability-patch knowledge base KnowledgeDB. The storage process is represented as follows:
[0065] KnowledgeDB←(v code ,Y final ,Patch gen ).
[0066] Includes new vulnerability code vector v code , corresponding to the CWE category label Y final , generate patch summary information Patch gen .
[0067] To achieve the above-mentioned purpose, the present invention further proposes a large-scale model supply chain automatic repair device based on similarity comparison, comprising:
[0068] The code semantic representation extraction module is used to extract natural language-level semantic representation vectors from the original source code of the vulnerability to be repaired using a pre-trained code model;
[0069] The information extraction and patch-aware building module is used to extract the structural summary information of the vulnerability source code and convert the structural summary information into a patch-aware structural vector through a structural embedding method, and then combine it with the semantic representation vector to obtain a dual-channel code representation;
[0070] The vulnerability type classification module is used to input the dual-channel code representation into the large model fine-tuned based on LoRA to automatically classify the CWE vulnerability types and perform zero-shot recognition of unknown categories based on the semantic matching strategy to obtain the vulnerability classification results;
[0071] A similar vulnerability case retrieval module is configured to use the vulnerability classification results as a retrieval filter to filter out a case set related to the current vulnerability type from a local vector database, and use the dual-channel code representation as a query vector to perform a similarity search on the case set to retrieve patches for similar vulnerability cases;
[0072] The prompt information generation module is used to extract patch summaries of similar vulnerability case patches and combine the current vulnerability code with the vulnerability classification results to construct prompt information to drive the large model to generate repair patches;
[0073] A patch code generation module is used to take the prompt information as input to call the large model based on LoRA fine-tuning to generate a patch code and repair description information for the current vulnerability;
[0074] The verification and update module is used to verify the validity of the repair patch code. If the verification passes, the relevant data of this repair process will be structured and archived to the local knowledge base.
[0075] The large-scale supply chain automatic repair method and system based on similarity comparison in the embodiments of the present invention aims to address the current problem of code vulnerabilities in large-scale supply chains being difficult to efficiently identify and automatically repair, particularly the lack of context-aware and security-specific automatic repair methods in local environments. To this end, the present invention can automatically classify, semantically retrieve, and generate patches for vulnerable code without relying on external services, thereby improving the automation and reliability of supply chain security repairs.
[0076] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0078] Figure 1 is a flow chart of a large-model supply chain automatic repair method based on similarity comparison according to an embodiment of the present invention;
[0079] Figure 2 4 is a structural diagram of a large-model supply chain automatic repair device based on similarity comparison according to an embodiment of the present invention. DETAILED DESCRIPTION
[0080] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0081] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0082] The following describes a large-model supply chain automatic repair method and apparatus based on similarity comparison according to an embodiment of the present invention with reference to the accompanying drawings.
[0083] Figure 1 Flowchart of a large-scale supply chain automatic repair method based on similarity comparison according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0084] S1 uses the pre-trained code model to extract natural language-level semantic representation vectors from the original vulnerability source code to be repaired.
[0085] Specifically, this step aims to extract a natural language-level semantic representation vector from the original source code of the vulnerability to be fixed, which can be used for downstream tasks (such as classification, retrieval, and generation). This representation should have the following properties: it can capture the semantic patterns, variable dependencies, and context of the vulnerability code, and establish a measurable vector space similarity with the patch sample.
[0086] Suppose the input is a Python code snippet to be repaired, recorded as:
[0087] C=RawCode(l1,l2,…,l n )
[0088] where l i Represents the i-th line of source code, with a total length of n lines.
[0089] The source code sequence C is then input into the tokenizer of a pre-trained code model (such as CodeBERT) to obtain a token sequence:
[0090] T=[t1,t2,…,t m ],t i ∈V
[0091] Where V is the model vocabulary and m represents the number of tokens.
[0092] Then perform position encoding and embedding mapping, let E:V→R d is the embedding function, P:N→R d Generate the initial input vector sequence for the position encoding function:
[0093] X=[x1,x2,…,x m ],x i =E(t i )+P(i)
[0094] Then m is input into the encoder model (such as CodeBERT) and a multi-layer Transformer encoding process is performed to obtain the hidden state representation corresponding to each token:
[0095] H=[h1,h2,…,h m ],h i ∈R d
[0096] The output vector at the global summary position (such as the [CLS] tag) is selected as the semantic representation of the entire code snippet, which is recorded as:
[0097] V sem =h CLS ∈R d
[0098] The vector V selected in this way sem It contains the distribution information of the input code in the semantic space, which can be used for subsequent vulnerability type identification (classification), similarity retrieval and context modeling.
[0099] S2 extracts the structural summary information of the vulnerability source code and converts the structural summary information into a patch-aware structural vector through a structural embedding method, so as to obtain a dual-channel code representation by combining it with the semantic representation vector.
[0100] Specifically, this step also focuses on the structured information of the vulnerability code to assist in vulnerability understanding and patch generation. First, the code snippet is converted into an Abstract Syntax Tree (AST) through a program parsing library (such as Python's built-in ast module). AST :
[0101] T AST =(V,E)
[0102] Among them, V is the set of syntax tree nodes (such as variables, function calls, expression nodes), and E is the set of edges between nodes, representing grammatical relationships.
[0103] Furthermore, the control flow graph (CFG) of the code can be generated as needed, that is, G CFG , or Data Flow Graph (DFG) or G DFG , written as:
[0104] G CFG =(N cfg ,F cfg ),G DFG =(N dfg ,F dfg )
[0105] Where N is the set of nodes in the control flow graph (or data flow graph), and F is the set of edges in the control flow graph (or data flow graph). Subsequently, the above structural information is converted into a structural vector representation through graph representation learning (such as graph neural network GNN) or structural embedding method:
[0106] v struct =f struct (T AST ,GCFG ,G DFG )
[0107] Among them, f struct The representation graph represents the learning function, whose goal is to embed multiple structural graphs into a vector so that they can participate in subsequent classification, retrieval and generation tasks together with the semantic vector.
[0108] Finally, the semantic vector V sem and structure vector v struct Concatenated for patch-aware representation:
[0109] v code =[V sem ;v struct ]∈R 2d
[0110] This representation combines semantic and structural information, more accurately capturing the vulnerability code and its context, laying the data foundation for subsequent steps. Steps (S1) and (S2) together construct a dual-channel encoding mechanism for the vulnerability source code (natural language semantic encoding and code structure summary encoding), accurately modeling the semantic and structural features of the vulnerability source code and its context.
[0111] S3, inputs the dual-channel code representation into the large model fine-tuned based on LoRA for automatic classification of CWE vulnerability types, and performs zero-sample recognition of unknown categories based on the semantic matching strategy to obtain the vulnerability classification results.
[0112] The goal of this step is to automatically identify the types of security vulnerabilities that the input code snippet may contain without manual labeling, and output its corresponding CWE (Common Weakness Enumeration) category label to provide contextual information constraints for subsequent similarity retrieval and patch generation.
[0113] To this end, the present invention introduces two complementary CWE intelligent classification mechanisms.
[0114] Multi-label classification mechanism: Considering that real vulnerabilities often involve multiple security weaknesses (such as missing input validation + command execution, etc.), this paper models CWE identification as a multi-label classification problem.
[0115] Assume that the complete set of CWE labels covered by the local knowledge base is:
[0116] Y CWE =[y1,y2,…,y k ]
[0117] Among them, y i Represents a CWE tag type.
[0118] For the input vulnerability code, the patch-aware representation vector v code , using a large model deployed locally and fine-tuned with LoRA to perform predictions, outputting the independent confidence level for each CWE label. To improve the model's robustness against long-tail and low-frequency CWEs, this paper optimizes the multi-label classification model by introducing sample reweighting, class-balanced sampling, and label correlation graph modeling, enabling the simultaneous identification of multiple potential vulnerability labels.
[0119] Semantic embedding matching strategy: To enhance the system's generalization ability on unknown CWE labels or those not covered by the training set, the present invention further introduces a zero-shot matching mechanism based on description semantics, that is, an approximate match is performed by comparing the code semantic vector with the embedding vector of the CWE official description statement.
[0120] Specifically, first for each CWE label y i Extract its official description text d i , and use an encoder (pre-trained models that support text embedding, such as GraphCodeBERT, CodeT5)f embed Generate its semantic vector u i :
[0121] u i =f embed (d i )
[0122] where u i ∈R d CWE label y i Semantic embedding of .
[0123] Afterwards, the semantic vector representation V of the vulnerability code to be classified is calculated sem The semantic similarity between them (such as cosine similarity), the i-th similarity score is s i Finally, the top N CWE labels with the largest similarity are selected as the predicted output Y pred :
[0124] Y pred =TopN i (s i )
[0125] This mechanism essentially transforms the classification problem into a semantic similarity matching problem. Without additional training, it can identify vulnerability types with similar semantics but unknown labels. It has good scalability and interpretability.
[0126] Finally, the present invention combines the output results of the multi-label classification model and the zero-sample matching, and adopts a weighted fusion method to determine the final predicted label set Y final .
[0127] The CWE type set output by this step is not only used to inform users of the nature of the current vulnerability, but also serves as an important context for downstream vector retrieval and repair generation modules.
[0128] S4, using the vulnerability classification result as a retrieval filter condition to filter out a case set related to the current vulnerability type from the local vector database, and using the dual-channel code representation as a query vector to perform a similarity search on the case set to retrieve similar vulnerability case patches.
[0129] Specifically, this step performs a similarity search on the current vulnerability vector representation based on a local vector database (such as Faiss or Milvus) to find the most similar historical vulnerability-patch case. code As a query vector, search the database:
[0130] D topK =argTopK i {sim(v code ,code i )}
[0131] Among them, code i is the historical vulnerability code representation stored in the database, sim() is the cosine similarity function, and the retrieval result D topK Contains K historical cases closest to the input, and their corresponding patch information is used as subsequent reference knowledge.
[0132] S5, extracts patch summaries of similar vulnerability case patches, and combines the current vulnerability code with the vulnerability classification results to construct prompt information for driving the large model to generate repair patches.
[0133] This step abstracts and summarizes the retrieved patches of similar vulnerability cases, and combines them with the current vulnerability information to form an input prompt (Prompt) for the large model to generate patch codes. The core repair ideas of historical patch cases are summarized as a patch summary description set P ref :
[0134] P ref ={p1,p2,…,p k}
[0135] where p i represents the patch summary of the i-th similarity case.
[0136] Afterwards, construct the prompt template P prompt for:
[0137] Vulnerable code snippet: <code snippet C to be fixed>
[0138] Vulnerability Type: <Identified vulnerability type>
[0139] Reference fix summary:
[0140] -Case 1: <p1>
[0141] -Case 2: <p2>
[0142] Please combine the above content to generate a repair code patch and repair instructions.
[0143] The resulting prompt information clearly expresses the vulnerability repair requirements and provides contextual reference information for the large model.
[0144] S6 takes the prompt information as input to call the large model based on LoRA fine-tuning to generate the repair patch code and repair instructions for the current vulnerability.
[0145] Specifically, the above prompt information is input into the local large language model fine-tuned by LoRA, and the patch code for fixing the vulnerability and the explanation of the repair principle are automatically generated:
[0146] Patch gen =LoRA-Generate(P prompt )
[0147] Final model output Patch gen Includes vulnerability patch code and repair principle explanation.
[0148] The output patch code can specifically and effectively fix the vulnerabilities of the input code. At the same time, the repair instructions generated by the model improve the interpretability of the results and avoid blind generation.
[0149] S7, verify the validity of the repair patch code. If the verification passes, the relevant data of this repair process is structured and archived in the local knowledge base.
[0150] Specifically, after verifying the effectiveness of the automatically generated patch, the patch code is applied to the actual project, forming a closed loop for vulnerability repair. At the same time, the system supports automatic structured archiving of the repair process and stores it in the local vulnerability-patch knowledge base KnowledgeDB, including:
[0151] New vulnerability code vector v code ;
[0152] Corresponding CWE category label;
[0153] Generate patch summary information Patch gen ;
[0154] The stored procedure is represented as:
[0155] KnowledgeDB←(v code ,Y final ,Patch gen )
[0156] Based on this, the knowledge base can be self-iterated and updated, continuously improving the accuracy and generalization of subsequent vulnerability identification and automatic repair.
[0157] In summary, the present invention first uses a dual-channel vulnerability encoding mechanism to parallel model the code to be repaired: on the one hand, a pre-trained model (such as CodeBERT) is used to extract natural language semantic vectors from the vulnerable source code, and on the other hand, its Abstract Syntax Tree (AST) or control flow information is extracted to construct a structural summary embedding, forming a comprehensive representation with patch-aware capabilities. Subsequently, the system uses a locally deployed large language model fine-tuned by LoRA to automatically identify Common Weakness Enumeration (CWE) vulnerability types on the input code, supporting multi-label prediction and zero-shot matching to improve adaptability to complex or unseen vulnerability types.
[0158] After vulnerability identification and type identification, the system uses similarity vector retrieval to search for historical cases with semantically similar vulnerability patterns within a locally maintained vulnerability-patch knowledge base, extracting their patch summaries for reference. Based on this information, the system constructs a complete prompt consisting of "code to be fixed + vulnerability type + reference case summary," which is then fed into the fine-tuned large model to generate a patch and repair instructions. Finally, the system outputs the repair results and supports user verification. If verification passes, the repair process is archived in the knowledge base, enabling continuous accumulation of model knowledge and enhanced repair performance, building an automated closed-loop repair capability for large-scale model supply chain security.
[0159] In summary, the present invention first semantically encodes the original vulnerability code through a pre-trained model, and combines program structures such as the abstract syntax tree (AST) to construct a dual-channel representation with parallel semantics and structure. Subsequently, a large language model deployed locally and fine-tuned by LoRA is used to intelligently classify the vulnerabilities into CWE types, and at the same time introduces a zero-sample semantic matching mechanism to improve the ability to identify unknown vulnerability types. On this basis, the system retrieves the closest historical cases from the local vulnerability-patch knowledge base through similarity vector comparison, extracts repair summaries, and constructs a model to generate prompts. Finally, patch code and repair instructions are generated through the large model, and automatic archiving and knowledge base updates are supported to achieve continuously enhanced automatic repair capabilities. It is mainly used in code security repair scenarios in the Python language environment, aiming to solve the problems of difficulty in identifying security vulnerabilities, low repair efficiency, and strong manual dependence in the current large model supply chain code. The invention can realize automatic understanding of vulnerability codes, type classification, similar patch retrieval and high-quality repair suggestion generation, significantly improving the localized security closed-loop capability. The method of the present invention can be widely used in automatic vulnerability repair scenarios of supply chain components such as large model plug-ins, code interfaces, and dependency packages, greatly improving the security, functionality, and explainability of code repair patches.
[0160] According to the large-model supply chain automatic repair method based on similarity comparison in an embodiment of the present invention, a dual-channel encoding mechanism (natural language semantic encoding and code structure summary encoding) is used to accurately model the semantic and structural features of the vulnerability source code and its context. Based on the Retrieval-Augmented Generation (RAG) framework, combined with a locally deployed large model fine-tuned by LoRA, efficient and accurate large-model supply chain code vulnerability classification and automatic repair are achieved. Through the method of the present invention, it is possible to automate the repair of large-model supply chain vulnerabilities and ensure the security, functionality and interpretability of the repair patch.
[0161] In order to implement the above embodiment, Figure 2 As shown, this embodiment also provides a large-scale model supply chain automatic repair device 10 based on similarity comparison, including:
[0162] A code semantic representation extraction module 100 is used to extract a natural language-level semantic representation vector from the original vulnerability source code to be repaired using a pre-trained code model;
[0163] The information extraction and patch-aware building module 200 is used to extract the structural summary information of the vulnerability source code and convert the structural summary information into a patch-aware structural vector through a structural embedding method, so as to obtain a dual-channel code representation by combining it with the semantic representation vector;
[0164] The vulnerability type classification module 300 is used to input the dual-channel code representation into the large model fine-tuned based on LoRA to automatically classify the CWE vulnerability type, and perform zero-shot recognition of unknown categories based on the semantic matching strategy to obtain the vulnerability classification result;
[0165] A similar vulnerability case retrieval module 400 is configured to use the vulnerability classification result as a retrieval filter to filter out a case set related to the current vulnerability type from a local vector database, and use the dual-channel code representation as a query vector to perform a similarity search on the case set to retrieve patches for similar vulnerability cases;
[0166] The prompt information generation module 500 is used to extract patch summaries of similar vulnerability case patches and, in combination with the current vulnerability code and vulnerability classification results, construct prompt information for driving the large model to generate repair patches;
[0167] A patch code generation module 600 is configured to use the prompt information as input to call a large model based on LoRA fine-tuning to generate a patch code and repair description information for the current vulnerability;
[0168] The verification and update module 700 is used to verify the validity of the repair patch code. If the verification is passed, the relevant data of this repair process is structured and archived to the local knowledge base.
[0169] According to an embodiment of the present invention, the large-model supply chain automatic repair device based on similarity comparison uses a dual-channel encoding mechanism (natural language semantic encoding and code structure summary encoding) to accurately model the semantic and structural features of the vulnerability source code and its context. Based on the Retrieval-Augmented Generation (RAG) framework, combined with a locally deployed and LoRA-fine-tuned large model, efficient and accurate large-model supply chain code vulnerability classification and automatic repair are achieved. Through the method of the present invention, the automation of large-model supply chain vulnerability repair can be achieved, and the security, functionality and interpretability of the repair patch can be guaranteed.
[0170] In the description of this specification, reference to the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.
[0171] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature identified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
Claims
1. A large-scale supply chain automatic repair method based on similarity comparison, characterized by: include: Use the pre-trained code model to extract natural language-level semantic representation vectors from the original vulnerability source code to be fixed; Extract the structural summary information of the vulnerability source code and transform it into a patch-aware structural vector through a structural embedding method, and combine it with the semantic representation vector to obtain a dual-channel code representation; The dual-channel code representation is input into a large model fine-tuned based on LoRA to automatically classify CWE vulnerability types, and zero-shot recognition of unknown categories is performed based on a semantic matching strategy to obtain vulnerability classification results. Using the vulnerability classification result as a retrieval filter condition to filter out a case set related to the current vulnerability type from a local vector database, and using the dual-channel code representation as a query vector to perform a similarity search on the case set to retrieve patches for similar vulnerability cases; Extract patch summaries of similar vulnerability case patches, and combine the current vulnerability code with the vulnerability classification results to construct prompt information to drive the large model to generate repair patches; The prompt information is used as input to call the large model based on LoRA fine-tuning to generate a repair patch code and repair instructions for the current vulnerability; The validity of the repair patch code is verified. If the verification passes, the relevant data of this repair process is structured and archived in the local knowledge base.
2. The method according to claim 1, characterized in that Utilize the pre-trained code model to extract natural language-level semantic representation vectors from the original vulnerability source code to be fixed, including: The input Python code snippet to be repaired is recorded as: <h2 style=";text-align:left;direction:ltr">C = RawCode(l1,l2,…,l<h2 style=";text-align:left;direction:ltr"> n <h2 style=";text-align:left;direction:ltr"> ) Input the source code sequence C into the tokenizer of the pre-trained code model to obtain the token sequence: T=[t1,t2,…,t m ],t i ∈V Perform position encoding and embedding mapping, let E:V→R d is the embedding function, P:N→R d Generate the initial input vector sequence for the position encoding function: X=[x1,x2,…,x m ],x i =E(t i )+P(i) Input m into the encoder model and perform a multi-layer Transformer encoding process to obtain the hidden state representation corresponding to each token: H=[h1,h2,…,h m ],h i ∈R d The output vector at the global summary position is selected as the semantic representation of the entire code snippet, which is recorded as: V sem =h CLS ∈R d Among them, l i Represents the i-th line of source code, with a total length of n lines; V is the model vocabulary, and m represents the number of tokens.
3. The method according to claim 1, characterized in that Extract the structural summary information of the vulnerability source code and transform it into a patch-aware structural vector through a structural embedding method. Combined with the semantic representation vector, a dual-channel code representation is obtained, including: Convert the code snippet into an abstract syntax tree T through the program parsing library AST : T AST =(V,E) Generate additional control flow graph of code as needed, namely G CFG , or data flow graph G DFG , written as: G CFG =(N cfg ,F cfg ),G DFG =(N dfg ,F dfg ) Through graph representation learning or structural embedding methods, the structural information is converted into structural vector representation: v struct =f struct (T AST ,G CFG ,G DFG ) The semantic vector V sem and structure vector v struct Concatenated for patch-aware representation: v code =[V sem ;v struct ]∈R 2d Where V is the set of syntax tree nodes, E is the set of edges between nodes, N is the set of nodes in the control flow graph or data flow graph, and F is the set of edges in the control flow graph or data flow graph; f struct Representation graph represents learning function.
4. The method according to claim 1, wherein The dual-channel code representation is input into a large model fine-tuned based on LoRA for automatic classification of CWE vulnerability types. Zero-shot recognition of unknown categories is performed based on a semantic matching strategy to obtain vulnerability classification results, including: Assume that the complete set of CWE labels covered by the local knowledge base is: Y CWE =[y1,y2,…,y k ] Among them, y i Represents a CWE tag type; For the input vulnerability code, the patch-aware representation vector v code ,Prediction is performed through a large model deployed locally and fine-tuned by LoRA, outputting independent confidence scores for each CWE label; For each CWE label y i Extract official description text d i , and use the encoder f embed Generate semantic embedding u i : u i =f embed (d i ) where u i ∈R d CWE label y i Semantic embedding of Calculate the semantic embedding and semantic vector representation V of the vulnerability code to be classified sem The semantic similarity between them, the similarity score of i is s i Finally, the top N CWE labels with the largest similarity are selected as the predicted output Y pre d : Y pred =TopN i (s i ) Combine the output results of the multi-label classification model and zero-sample matching, and use weighted fusion to determine the final predicted label set Y final .
5. The method according to claim 1, wherein The vulnerability classification result is used as a retrieval filter to filter out a case set related to the current vulnerability type from the local vector database, and the dual-channel code representation is used as a query vector to perform a similarity search on the case set to retrieve similar vulnerability case patches, including: The code represents v code Retrieve in the database as a query vector: D topK =argTopK i {yes(v co d e ,code i )} Among them, code i is the historical vulnerability code representation stored in the database, sim() is the cosine similarity function, and the retrieval result D topK Contains the K historical cases that are closest to the input.
6. The method according to claim 1, characterized in that Extract patch summaries of similar vulnerability case patches, and combine the current vulnerability code and vulnerability classification results to construct prompt information to drive the large model to generate repair patches, including: Summarize the core repair ideas of historical patch cases into a patch summary description set P ref : P ref ={p1,p2,…,p k } where p i represents the patch summary in the i-th similar case; Construct a hint template P based on the patch summary description set prompt To form the required prompt information.
7. The method according to claim 1, characterized in that The prompt information is used as input to call the large model based on LoRA fine-tuning to generate the patch code and repair instructions for the current vulnerability, including: Input the prompt information into the local large language model fine-tuned by LoRA, and automatically generate the patch code for fixing the vulnerability and the explanation of the repair principle: Pacth gen =LoRA-Generate(P prompt ) Final model output Patch gen Includes vulnerability patch code and repair principle explanation.
8. The method according to claim 1, characterized in that Verify the validity of the patch code. If the verification passes, the relevant data of the repair process will be structured and archived in the local knowledge base, including: After verifying the effectiveness of the automatically generated patch, the patch code is applied to the actual project to form a closed loop for vulnerability repair. At the same time, the repair process is automatically structured and archived and stored in the local vulnerability-patch knowledge base KnowledgeDB. The storage process is represented as follows: KnowledgeDB←(v code ,Y final ,Patch gen )。 Includes new vulnerability code vector v code , corresponding to the CWE category label Y final , generate patch summary information Patch gen .
9. A large-scale supply chain automatic repair device based on similarity comparison, characterized in that: include: The code semantic representation extraction module is used to extract natural language-level semantic representation vectors from the original source code of the vulnerability to be repaired using a pre-trained code model; The information extraction and patch-aware building module is used to extract the structural summary information of the vulnerability source code and convert the structural summary information into a patch-aware structural vector through a structural embedding method, and then combine it with the semantic representation vector to obtain a dual-channel code representation; The vulnerability type classification module is used to input the dual-channel code representation into the large model fine-tuned based on LoRA to automatically classify the CWE vulnerability types and perform zero-shot recognition of unknown categories based on the semantic matching strategy to obtain the vulnerability classification results; A similar vulnerability case retrieval module is configured to use the vulnerability classification results as a retrieval filter to filter out a case set related to the current vulnerability type from a local vector database, and use the dual-channel code representation as a query vector to perform a similarity search on the case set to retrieve patches for similar vulnerability cases; The prompt information generation module is used to extract patch summaries of similar vulnerability case patches and combine the current vulnerability code with the vulnerability classification results to construct prompt information to drive the large model to generate repair patches; A patch code generation module is used to take the prompt information as input to call the large model based on LoRA fine-tuning to generate a patch code and repair description information for the current vulnerability; The verification and update module is used to verify the validity of the repair patch code. If the verification passes, the relevant data of this repair process will be structured and archived to the local knowledge base.
Citation Information
Cited By
Source code bug repairing method, electronic equipment and storage medium
CN121479795A
Method and device for carrying out vulnerability function call analysis by utilizing large model in supply chain security
CN122221274A