Vulnerability repairing method and system based on knowledge graph and large language model

By preprocessing vulnerability data and building a vulnerability repair knowledge graph, and combining with large language models to generate repair code, the problems of insufficient semantic information extraction, low matching efficiency and low degree of automation in vulnerability repair technology are solved, and an efficient and accurate vulnerability repair process is achieved.

CN120012095APending Publication Date: 2025-05-16GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411854392.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing vulnerability repair technology has insufficient extraction of semantic information for preprocessing of vulnerable data, lack of automation and intelligent matching mechanisms for knowledge graph construction, insufficient application of large language models in automated generation and repair code, and how to achieve automation, intelligence and high efficiency of vulnerability repair process.

Method used

By preprocessing vulnerability data, provide semantic information support; build a vulnerability repair knowledge graph to match vulnerability code and repair code; generate vulnerability repair patch code based on a large language model to achieve automated repair of vulnerabilities.

Benefits of technology

It improves the accuracy and efficiency of vulnerability repair, realizes efficient matching of vulnerability code and repair code, reduces manual intervention, shortens the vulnerability repair cycle, and reduces security risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012095A_ABST
    Figure CN120012095A_ABST
Patent Text Reader

Abstract

The invention discloses a vulnerability repairing method and system based on a knowledge graph and a large language model, and relates to the technical field of network security, and the method comprises the steps: preprocessing vulnerability data, and providing semantic information support in a vulnerability repairing process; constructing a vulnerability repair knowledge graph, matching the vulnerability code with the repair code, and providing a matching result; and generating a vulnerability repair patch code based on the large language model, and automatically repairing the vulnerability. According to the method, the accuracy and the efficiency of vulnerability repair are improved, a data basis and semantic understanding are provided for subsequent vulnerability repair, the matching speed and the matching accuracy are improved, powerful support is provided for automatic repair, key features of vulnerability problems are accurately identified, and the vulnerability repair efficiency is improved. And the repair patch is automatically generated through the large language model, so that manual intervention is reduced, the repair speed and reliability are improved, efficient and reliable automatic patch repair is realized, and the safety risk is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and specifically to a vulnerability repair method and system based on a knowledge graph and a large language model. Background Art

[0002] Against the backdrop of continuous advancement of information technology, network security has become an important part of corporate and national security. As a key link in network security, vulnerability repair technology has evolved from traditional manual repair to semi-automated tools. In recent years, knowledge graphs as a structured semantic network and breakthroughs in natural language processing using large language models have provided new impetus for the development of vulnerability repair technology. The construction and application of knowledge graphs have revealed the semantic associations of vulnerability data, while the powerful language generation capabilities of large language models have made it possible to automatically generate repair code. Research in this field is gradually deepening.

[0003] However, although the existing technology has made certain progress in vulnerability repair, it still has significant shortcomings. In the vulnerability data preprocessing stage, the existing technology often fails to fully extract and utilize the semantic information of the vulnerability, resulting in insufficient information support in the subsequent repair process, affecting the accuracy and efficiency of the repair. The construction of the vulnerability repair knowledge graph mostly relies on expert experience and lacks an automated and intelligent matching mechanism, resulting in a low match between the vulnerability code and the repair code, and the effect of the repair solution is not good. Finally, although the large language model has the potential to generate code, the existing technology has not effectively integrated this capability to realize the automation of vulnerability repair, resulting in the repair process in actual applications still relying on a large amount of manual intervention and unable to meet the needs of rapid response. These problems have jointly restricted the development of vulnerability repair technology and urgently need new technical solutions to solve them. Summary of the invention

[0004] In view of the above-mentioned problems, the present invention is proposed.

[0005] Therefore, the technical problems solved by the present invention are: the existing vulnerability repair technology has insufficient semantic information extraction for vulnerability data preprocessing, lack of automation and intelligent matching mechanism in knowledge graph construction, insufficient application of large language models in automatic generation of repair code, and how to realize the automation, intelligence and high efficiency of the vulnerability repair process.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: a vulnerability repair method based on a knowledge graph and a large language model, comprising preprocessing vulnerability data to provide semantic information support in the vulnerability repair process; constructing a vulnerability repair knowledge graph, matching vulnerability code with repair code, and providing matching results; generating vulnerability repair patch code based on a large language model to automatically repair the vulnerability.

[0007] As a preferred solution of the vulnerability repair method based on knowledge graph and large language model described in the present invention, wherein: the preprocessing of vulnerability data includes obtaining vulnerability knowledge and extracting vulnerability entities, extracting required information from vulnerability reports and repair reports by using natural language processing tools, extracting vulnerability ID, components, products, summary information, and description information from the vulnerability report, searching for the corresponding report on Github according to the vulnerability ID, and extracting vulnerability code and correct code from the report, StackOverflow is an open source community communication platform that records vulnerability content, and the data storage format is .xml, and the submission ID and tag information are extracted by parsing the file.

[0008] As a preferred solution of the vulnerability repair method based on knowledge graph and large language model described in the present invention, the preprocessing of vulnerability data also includes extracting the relationship between vulnerability entities, using the TAKG model to preprocess the vulnerability report and open source platform description information, remove stop words, punctuation marks, unify capitalization, and convert words into bag-of-words vectors, and train the processed data to generate topic distribution using a neural topic model. NTM is based on the architecture of the variational autoencoder to process X bow For the reconstruction task, the encoder is responsible for estimating the prior variables μ and σ.

[0009] The latent representation z induces the intermediate topic to participate in the generation process of the generative model Decoder, and uses NTM for X bow In the reconstruction process, the model learns the topic information corresponding to different words as the guiding information of the generation process, which is expressed as:

[0010] μ=f μ (f e (X bow )),logσ=f σ (f e (X bow ))

[0011] Among them, μ is the mean value of the encoder output, f e (.) is the encoder function used to convert the input vector into a potential representation, X bow is the input bag-of-words model vector, f μ (.) represents the mapping function, f σ (.) is the neural network mapping function, and σ represents the variance of the latent variable.

[0012] The processed data is used to train the neural topic model to generate topic distribution. The optimal number of topics is determined based on topic perplexity and topic consistency. The consistency score is introduced as an evaluation criterion to measure the quality and relevance of topics generated by the topic model. The higher the score, the more semantically coherent and meaningful the topic is.

[0013] As a preferred solution of the vulnerability repair method based on knowledge graph and large language model described in the present invention, wherein: the construction of the vulnerability repair knowledge graph includes determining the data structure of the knowledge graph, including entities, relationships and attributes, entities are nodes in the knowledge graph, relationships are edges connecting nodes, and attributes are used to describe the characteristics of nodes and edges.

[0014] As a preferred solution of the vulnerability repair method based on knowledge graph and large language model described in the present invention, wherein: the providing of matching results includes converting the collected data into a triple format, each triple consisting of a subject, a predicate and an object, forming a subject, predicate, object structure, so as to be stored in Neo4j, deploying the knowledge graph in the Neo4j database, using the graph database characteristics of Neo4j, importing the converted triple data into the database, and establishing corresponding nodes and edges, showing the association between entities, and defining entity types and relationship types in the knowledge graph.

[0015] Use py2neo to create node indexes and nodes, add attributes, search for nodes based on indexes and create relationships between nodes to create a bug fixing knowledge graph BKG.

[0016] As a preferred solution of the vulnerability repair method based on knowledge graph and large language model described in the present invention, the automatic repair of the vulnerability includes using BKG to search for vulnerability issues, using natural language technology to extract keywords of the input vulnerability issues, and using a pre-trained BERT model to convert the extracted keywords into high-dimensional vector representations. The BERT model captures the semantic relationship between keywords through deep learning and natural language processing capabilities, and converts them into points in the vector space.

[0017] By calculating the cosine similarity between the keyword vector and the topic vector, we can find the topic that best matches the input vulnerability description and measure the closeness of the two vectors in direction. The closer the value is to 1, the more similar they are. is the dot product of two vectors, the modulus of vectors A and B, expressed as:

[0018]

[0019] Among them, cos_sim is the cosine similarity, A feature vector representing the input vulnerability description, The representation vector represents the feature vector of the vulnerability repair solution, For vector The model, For vector Model.

[0020] As a preferred solution of the vulnerability repair method based on knowledge graph and large language model described in the present invention, the automatic repair of vulnerabilities also includes adopting a vulnerability repair method based on Incoder pre-trained code model, using a large language model to learn and directly generate patches from a large number of complete code snippets.

[0021] Select the project with the bug, find the error line according to the vulnerability location information, obtain the entire method code and method comments including the error line, use the Javaparse library to identify the parameters and method content in the error line, and use the mask template to generate multiple mask lines. Each mask line will replace the error line and be used as the input of the code generation model together with the entire method code and method comments. Iteratively call the Incoder model to automatically generate the code for the mask part. Each patch replaces the error line in the mask line. <mask>Part of the code lines are replaced with the generated part, the generated patches are rechecked, the patches that are identical to the original error code and the duplicates are eliminated, patch verification is performed, each candidate patch is compiled and verified using the test suite, and the patches that pass the test are output. The developers check the generated patches one by one.

[0022] Another object of the present invention is to provide a vulnerability repair system based on a knowledge graph and a large language model, which can match vulnerability codes with repair codes by constructing a vulnerability repair knowledge graph and provide matching results, thereby solving the problems of incomplete knowledge graph construction and low efficiency in matching vulnerability codes with repair codes in current vulnerability repair technologies.

[0023] As a preferred solution of the vulnerability repair system based on knowledge graph and large language model described in the present invention, it includes: a semantic information support module, a vulnerability matching and repair module, and an automatic patch generation module.

[0024] The semantic information support module is used to preprocess vulnerability data and provide semantic information support in the vulnerability repair process; the vulnerability matching and repair module is used to build a vulnerability repair knowledge graph, match vulnerability code with repair code, and provide matching results; the patch automatic generation module is used to generate vulnerability repair patch code based on a large language model and automatically repair vulnerabilities.

[0025] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a step of a vulnerability repair method based on a knowledge graph and a large language model.

[0026] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a vulnerability repair method based on a knowledge graph and a large language model.

[0027] Beneficial effects of the present invention: The vulnerability repair method based on knowledge graph and large language model provided by the present invention preprocesses vulnerability data, provides semantic information support in the vulnerability repair process, provides data basis and semantic understanding for subsequent vulnerability repair, improves the accuracy and efficiency of vulnerability repair, constructs a vulnerability repair knowledge graph, matches vulnerability code with repair code, provides matching results, and establishes a structured knowledge base, so that the relationship between the vulnerability and the repair solution is clearer, thereby improving the matching speed and matching accuracy, providing strong support for automated repair, generating vulnerability repair patch code based on the large language model, automatically repairing the vulnerability, accurately identifying the key features of the vulnerability problem, and automatically generating repair patches through the large language model, which not only reduces manual intervention, but also improves the speed and reliability of repair, realizes efficient and reliable automated repair patches, shortens the cycle of vulnerability repair, and reduces security risks. The present invention achieves better results in vulnerability data preprocessing, vulnerability matching and repair, and automated repair patch generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0029] Figure 1 An overall flow chart of a vulnerability repair method based on a knowledge graph and a large language model provided for the first embodiment of the present invention.

[0030] Figure 2 A perplexity and consistency score graph of a vulnerability repair method based on a knowledge graph and a large language model provided in the second embodiment of the present invention.

[0031] Figure 3 FIG1 is an architecture diagram of a vulnerability repair method based on a knowledge graph and a large language model provided for a second embodiment of the present invention.

[0032] Figure 4 FIG2 is an architecture diagram of a vulnerability repair method based on a knowledge graph and a large language model provided for a second embodiment of the present invention.

[0033] Figure 5 A patch generation process diagram of a vulnerability repair method based on a knowledge graph and a large language model provided in the second embodiment of the present invention.

[0034] Figure 6 An overall flow chart of a vulnerability repair system based on a knowledge graph and a large language model provided for the third embodiment of the present invention. DETAILED DESCRIPTION

[0035] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0036] Example 1, reference Figure 1 , is an embodiment of the present invention, and provides a vulnerability repair method based on knowledge graph and large language model, including:

[0037] S1: Preprocess vulnerability data to provide semantic information support during vulnerability repair.

[0038] Furthermore, the vulnerability data is preprocessed, including obtaining vulnerability knowledge and extracting vulnerability entities. The required information is extracted from vulnerability reports and repair reports by using natural language processing tools. The vulnerability ID, component, product, summary information, and description information are extracted from the vulnerability report. The corresponding report is found on Github according to the vulnerability ID, and the vulnerability code and correct code are extracted from the report. StackOverflow is an open source community communication platform that records vulnerability content. The data storage format is .xml. The submission ID and tag information are extracted by parsing the file.

[0039] It should be noted that the preprocessing of vulnerability data also includes extracting the relationship between vulnerability entities. The TAKG model is used to preprocess the vulnerability reports and open source platform description information, remove stop words, punctuation marks, unify capitalization, and convert words into bag-of-words vectors. The processed data is then trained with a neural topic model to generate topic distribution. NTM is based on the architecture of variational autoencoders to process X bow For the reconstruction task, the encoder is responsible for estimating the prior variables μ and σ.

[0040] The latent representation z induces the intermediate topic to participate in the generation process of the generative model Decoder, and uses NTM for X bow In the reconstruction process, the model learns the topic information corresponding to different words as the guiding information of the generation process, which is expressed as:

[0041] μ=f μ (f e (X bow )),logσ=f σ (f e (X bow ))

[0042] Among them, μ is the mean value of the encoder output, f e (.) is the encoder function used to convert the input vector into a potential representation, X bow is the input bag-of-words model vector, f μ (.) represents the mapping function, f σ (.) is the neural network mapping function, and σ represents the variance of the latent variable.

[0043] The processed data is used to train the neural topic model to generate topic distribution. The optimal number of topics is determined based on topic perplexity and topic consistency. The consistency score is introduced as an evaluation criterion to measure the quality and relevance of topics generated by the topic model. The higher the score, the more semantically coherent and meaningful the topic is.

[0044] It should also be noted that through natural language processing tools, key data such as vulnerability ID, components, products, summary information, and description information are effectively extracted from unstructured vulnerability reports, ensuring the accuracy and completeness of the information. The TAKG model is used to extract the relationships between vulnerability entities, and the neural topic model is used to generate topic distributions. This not only improves the level of intelligent data processing, but also ensures the semantic coherence and relevance of the extracted topics through topic perplexity and consistency scores, so that the vulnerability data forms a network with rich semantics, providing a deeper understanding and more accurate information support for vulnerability repair, improving the accuracy and efficiency of vulnerability repair, reducing the time and error rate of manual data processing, and laying a solid foundation for the entire vulnerability repair process.

[0045] S2: Build a vulnerability repair knowledge graph, match vulnerability codes with repair codes, and provide matching results.

[0046] Furthermore, building a vulnerability repair knowledge graph includes determining the data structure of the knowledge graph, including entities, relationships, and attributes. Entities serve as nodes in the knowledge graph, relationships serve as edges connecting nodes, and attributes are used to describe the characteristics of nodes and edges.

[0047] It should be noted that providing matching results includes converting the collected data into a triple format, where each triple consists of a subject, a predicate and an object, forming a subject, predicate and object structure for storage in Neo4j, deploying the knowledge graph in the Neo4j database, and using the graph database feature of Neo4j to import the converted triple data into the database, establish corresponding nodes and edges, display the associations between entities, and define entity types and relationship types in the knowledge graph.

[0048] Use py2neo to create node indexes and nodes, add attributes, search for nodes based on indexes and create relationships between nodes to create a bug fixing knowledge graph BKG.

[0049] It should also be noted that defining entity types and relationship types in the knowledge graph includes constructing a vulnerability knowledge graph entity type description table and a vulnerability knowledge graph relationship type description table. There are five entity types and three relationship types in total. Refer to Table 1 for a description of the vulnerability knowledge graph entity types.

[0050] Table 1 Vulnerability knowledge graph entity type description table

[0051]

[0052] Refer to Table 2 to explain the relationship types of the vulnerability knowledge graph.

[0053] Table 2 Vulnerability knowledge graph relationship type description table

[0054]

[0055]

[0056] There is a "has_question" relationship between topics, posts and bug reports, recorded as (Topic, has_question, Post / BugReport); there is a "has_buggycode" relationship between posts, bug reports and error codes, recorded as (StackOverflow / Bugzilla, has_buggycode, Buggycode); there is a "has_fixedcode" relationship between vulnerability codes and correct codes, recorded as (Buggycode, has_fixedcode, Fixedcode).

[0057] It should also be noted that by constructing a vulnerability repair knowledge graph, efficient matching of vulnerability codes and repair codes is achieved. The construction of the knowledge graph includes the determination of entities, relationships, and attributes, so that vulnerability data is presented in a structured form for easy analysis and processing. By converting the data into a triple format and deploying it in a Neo4j database, the characteristics of the graph database are used to intuitively display the associations between entities, providing an intuitive and efficient data structure for vulnerability matching. The indexes and relationships created using py2neo further optimize the search and matching process, and improve the accuracy and speed of matching. Through the construction and application of the knowledge graph, not only the accuracy of matching vulnerabilities with repair solutions is improved, but also the matching efficiency is improved, the time and labor intensity of manual matching are reduced, and strong support is provided for automated vulnerability repair.

[0058] S3: Generate vulnerability repair patch code based on the large language model to automatically repair the vulnerability.

[0059] Furthermore, automated repair of vulnerabilities includes using BKG to search for vulnerability issues, using natural language technology to extract keywords of the input vulnerability issues, and using the pre-trained BERT model to convert the extracted keywords into high-dimensional vector representations. The BERT model captures the semantic relationship between keywords through deep learning and natural language processing capabilities and converts them into points in the vector space.

[0060] By calculating the cosine similarity between the keyword vector and the topic vector, we can find the topic that best matches the input vulnerability description and measure the closeness of the two vectors in direction. The closer the value is to 1, the more similar they are. is the dot product of two vectors, the modulus of vectors A and B, expressed as:

[0061]

[0062] Among them, cos_sim is the cosine similarity, A feature vector representing the input vulnerability description, The representation vector represents the feature vector of the vulnerability repair solution, For vector The model, For vector Model.

[0063] It should be noted that automated vulnerability repair also includes the use of a vulnerability repair method based on the Incoder pre-trained code model, which uses a large language model to learn and directly generate patches from a large number of complete code snippets.

[0064] Select the project with the bug, find the error line according to the vulnerability location information, obtain the entire method code and method comments including the error line, use the Javaparse library to identify the parameters and method content in the error line, and use the mask template to generate multiple mask lines. Each mask line will replace the error line and be used as the input of the code generation model together with the entire method code and method comments. Iteratively call the Incoder model to automatically generate the code for the mask part. Each patch replaces the error line in the mask line. <mask>Part of the code lines are replaced with the generated part, the generated patches are rechecked, the patches that are identical to the original error code and the duplicates are eliminated, patch verification is performed, each candidate patch is compiled and verified using the test suite, and the patches that pass the test are output. The developers check the generated patches one by one.

[0065] It should also be noted that by using the pre-trained BERT model, the vulnerability problem description can be converted into a high-dimensional vector representation, and by calculating the cosine similarity, the topic that best matches the input vulnerability description can be found, thereby locating the corresponding repair solution. The method based on the Incoder pre-trained code model can directly learn and generate patches from a large number of code snippets, effectively improving the automation level of the repair process. It not only shortens the time cycle for vulnerability repair, but also reduces human errors through automated code generation, improves the quality and reliability of patches, and ensures that the generated patches can pass the verification of the test suite through the patch verification process, providing developers with repair solutions that have been preliminarily screened and verified, improving the overall efficiency and security of vulnerability repair.

[0066] Example 2, reference Figure 2-Figure 5 , which is an embodiment of the present invention, provides a vulnerability repair method based on knowledge graph and large language model. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0067] First, the experiment selected 200 public software security vulnerability cases, covering the fields of Web applications, operating systems, and database management systems. The experimental environment included a workstation equipped with a high-performance GPU, as well as the necessary natural language processing and machine learning libraries. First, the reports of the 200 vulnerability cases were electronically processed, and the text in the reports was parsed using a customized natural language processing script to extract relevant information about the vulnerability, including but not limited to the vulnerability type, affected software version, vulnerability description, and possible attack vectors. Figure 2 The perplexity and consistency score graph is shown in Figure 1. Combining the results of the perplexity score and the consistency score, 180 topics are selected as the optimal number of topics. Part-of-speech tagging and dependency syntactic analysis techniques are used in the extraction process to ensure the accuracy and completeness of the information. For example, for a vulnerability description, the script can identify the key entities of "software", "version", and "vulnerability" and establish semantic relationships. In the construction phase of the vulnerability matching and repair module, a knowledge graph containing vulnerability entities and repair strategies is constructed using the extracted semantic information. Each node in the graph represents a vulnerability or repair strategy, and the edge represents the relationship between them. In order to improve the accuracy of the matching, the graph embedding technology is used to convert the nodes in the graph into vector representations, and the best vulnerability repair strategy is determined by calculating the similarity between vectors. Figure 3 and Figure 4 The overall architecture of the vulnerability repair knowledge graph is represented. For example, for a vulnerability node V, the similarity between it and all repair strategy nodes in the graph is calculated, and the repair strategy R with the highest similarity is selected. In the implementation of the automatic patch generation module, the experiment uses a pre-trained large language model to generate repair patch codes. The model receives the vulnerability description and the matched repair strategy as input, and outputs the corresponding code patch. Figure 5 This is a patch generation process diagram. In order to verify the effectiveness of the generated patch, the experiment sets up an automated testing process, including compilation test and functional test. Each generated patch needs to be verified by this process. For example, for the generated patch P, the experiment first compiles it using a compiler to ensure that there are no syntax errors, and then performs functional verification through a predefined test case set T. It can be seen from the experimental results that the vulnerability repair efficiency, accuracy and degree of automation of the present invention are advantageous, and a new solution is provided for the field of network security.

[0068] Example 3, reference Figure 6 , is an embodiment of the present invention, which provides a vulnerability repair system based on knowledge graph and large language model, including a semantic information support module, a vulnerability matching and repair module, and a patch automatic generation module.

[0069] The semantic information support module is used to preprocess vulnerability data and provide semantic information support in the vulnerability repair process; the vulnerability matching and repair module is used to build a vulnerability repair knowledge graph, match vulnerability code with repair code, and provide matching results; the patch automatic generation module is used to generate vulnerability repair patch code based on a large language model and automatically repair vulnerabilities.

[0070] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0071] The logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, apparatus or device and execute the instructions), or in conjunction with such instruction execution system, apparatus or device. For purposes of this specification, "computer-readable medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, apparatus or device, or in conjunction with such instruction execution system, apparatus or device.

[0072] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0073] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc. It should be noted that the above embodiments are only used to illustrate the technical solution of the present invention and are not limited. Although the present invention is described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention, which should be included in the scope of the claims of the present invention.

[0074] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.< / mask> < / mask>

Claims

1. A vulnerability repair method based on knowledge graph and large language model, characterized in that: include: Preprocess vulnerability data to provide semantic information support during vulnerability repair; Build a vulnerability repair knowledge graph, match vulnerability codes with repair codes, and provide matching results; Generate vulnerability repair patch code based on large language models to automatically repair vulnerabilities.

2. The vulnerability repair method based on knowledge graph and large language model according to claim 1, characterized in that: The preprocessing of vulnerability data includes obtaining vulnerability knowledge and extracting vulnerability entities, extracting required information from vulnerability reports and repair reports by using natural language processing tools, extracting vulnerability ID, components, products, summary information, and description information from the vulnerability report, searching for the corresponding report on Github according to the vulnerability ID, and extracting vulnerability code and correct code from the report. StackOverflow is an open source community communication platform that records vulnerability content. The data storage format is .xml, and the submission ID and tag information are extracted by parsing the file.

3. The vulnerability repair method based on knowledge graph and large language model as claimed in claim 2, characterized in that: The preprocessing of vulnerability data also includes extracting the relationship between vulnerability entities, using the TAKG model to preprocess vulnerability reports and open source platform description information, removing stop words, punctuation marks, unifying capitalization, and converting words into bag-of-words vectors, and training the processed data to generate topic distribution using a neural topic model. NTM is based on the architecture of a variational autoencoder to process X bow For the reconstruction task, the encoder is responsible for estimating the prior variables μ and σ; The latent representation z induces the intermediate topic to participate in the generation process of the generative model Decoder, and uses NTM for X bow In the reconstruction process, the model learns the topic information corresponding to different words as the guiding information of the generation process, which is expressed as: μ=f μ (f e (X bow )),logσ=f σ (f e (X bow )) Among them, μ is the mean value of the encoder output, f e (.) is the encoder function used to convert the input vector into a potential representation, X bow is the input bag-of-words model vector, f μ (.) represents the mapping function, f σ (.) is the neural network mapping function, σ represents the variance of the latent variable; The processed data is used to train the neural topic model to generate topic distribution. The optimal number of topics is determined based on topic perplexity and topic consistency. The consistency score is introduced as an evaluation criterion to measure the quality and relevance of topics generated by the topic model. The higher the score, the more semantically coherent and meaningful the topic is.

4. The vulnerability repair method based on knowledge graph and large language model as claimed in claim 3, characterized in that: The construction of the vulnerability repair knowledge graph includes determining the data structure of the knowledge graph, including entities, relationships, and attributes. Entities serve as nodes in the knowledge graph, relationships serve as edges connecting nodes, and attributes are used to describe the characteristics of nodes and edges.

5. The vulnerability repair method based on knowledge graph and large language model as claimed in claim 4, characterized in that: Providing matching results includes converting the collected data into a triple format, each triple consisting of a subject, a predicate and an object, forming a subject, predicate, and object structure, so as to be stored in Neo4j, deploying a knowledge graph in a Neo4j database, using the graph database feature of Neo4j, importing the converted triple data into the database, and establishing corresponding nodes and edges to display the association between entities, and defining entity types and relationship types in the knowledge graph; Use py2neo to create node indexes and nodes, add attributes, search for nodes based on indexes and create relationships between nodes to create a bug fixing knowledge graph BKG.

6. The vulnerability repair method based on knowledge graph and large language model according to claim 5, characterized in that: The automatic repair of vulnerabilities includes searching for vulnerability issues using BKG, extracting keywords of the input vulnerability issues using natural language technology, and converting the extracted keywords into high-dimensional vector representations using a pre-trained BERT model. The BERT model captures the semantic relationship between keywords through deep learning and natural language processing capabilities and converts them into points in the vector space. By calculating the cosine similarity between the keyword vector and the topic vector, we can find the topic that best matches the input vulnerability description and measure the closeness of the two vectors in direction. The closer the value is to 1, the more similar they are. is the dot product of two vectors, the modulus of vectors A and B, expressed as: Among them, cos_sim is the cosine similarity, A feature vector representing the input vulnerability description, The representation vector represents the feature vector of the vulnerability repair solution, For vector The model, For vector Model.

7. The vulnerability repair method based on knowledge graph and large language model according to claim 6, characterized in that: The automated repair of vulnerabilities also includes adopting a vulnerability repair method based on an Incoder pre-trained code model, using a large language model to learn and directly generate patches from a large number of complete code snippets; Select the project with the bug, find the error line according to the vulnerability location information, obtain the entire method code and method comments including the error line, use the Javaparse library to identify the parameters and method content in the error line, and use the mask template to generate multiple mask lines. Each mask line will replace the error line and be used as the input of the code generation model together with the entire method code and method comments. Iteratively call the Incoder model to automatically generate the code for the mask part. Each patch replaces the error line in the mask line. <mask> Part of the code lines are replaced with the generated part, the generated patches are rechecked, the patches that are identical to the original error code and the duplicates are eliminated, patch verification is performed, each candidate patch is compiled and verified using the test suite, and the patches that pass the test are output. The developers check the generated patches one by one.< / mask> 8. A system using the vulnerability repair method based on knowledge graph and large language model as described in any one of claims 1 to 7, characterized in that: Including semantic information support module, vulnerability matching and repair module, and patch automatic generation module; The semantic information support module is used to pre-process vulnerability data and provide semantic information support during vulnerability repair; The vulnerability matching and repairing module is used to construct a vulnerability repairing knowledge graph, match the vulnerability code with the repairing code, and provide a matching result; The patch automatic generation module is used to generate vulnerability repair patch codes based on the large language model to automatically repair the vulnerabilities.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the vulnerability repair method based on the knowledge graph and the large language model described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the vulnerability repair method based on knowledge graph and large language model described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Vulnerability repair rule generation method based on AI and related equipment

    CN120611390A

  • An AI-based vulnerability repair rule generation method and related device

    CN120611390B

  • Retrieval enhancement large model vulnerability detection method based on variable type knowledge graph

    CN121598393A