Method and device for intelligently repairing code defects, computer equipment and medium
By building a knowledge base and using in-depth analysis platforms and large language models, efficient and automated code defect repair is achieved, solving the problem of inability to efficiently repair code defects in the existing technology, and improving the efficiency and quality of software development.
Patent Information
- Application Number
- CN202510133770.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-17
AI Technical Summary
The existing technology cannot repair code defects efficiently and automatically, especially in modern software projects. Due to the large scale and high complexity of the project, and the lack of efficient automated repair solutions, it seriously affects the development efficiency.
By using open source code bases and defect data sets to build a knowledge base, using the deep analysis platform to obtain the defect type of code to be detected, and query the code repair pairs that are most similar to the code to be repaired in the knowledge base through the code similarity algorithm, and generate the fixed code using dynamically constructed Prompt and large language models, replace and/or integrate it into the code to be detected.
It realizes efficient and automated code defect repair, improves software development efficiency and quality, provides a solution that goes beyond traditional LLM direct repair methods, bringing innovative technical and practical value to the field of code defect repair.
Smart Images

Figure CN120162073A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic information technology, and in particular to a method, device, computer equipment and medium for intelligently repairing code defects. Background Art
[0002] With the rapid development of informatization, digitization and intelligence, computer software plays an increasingly important role in modern life and industrial production. Especially driven by the open source movement, software products are showing a trend of scale and complexity, resulting in a sharp increase in the amount of code. The quality of software directly affects the operating efficiency and stability of related products and electronic equipment. Therefore, how to improve software quality has become a key issue that needs to be urgently addressed in my country's computer software industry. Automated software code detection, testing and maintenance, especially code repair, has become one of the core ways to ensure software quality and reliability.
[0003] Coding rules (code rule lists) represented by GJB8114 play an important role in high-demand fields such as aviation. They provide unified and standardized programming guidance for software development, ensuring the readability, maintainability and security of the code, thus laying the foundation for the successful execution of the project. However, although these coding rules are crucial to software quality, the current automatic detection and repair methods for these rule defects are still relatively backward, especially in modern software projects. Due to the large scale and high complexity of the projects, there is a lack of efficient automated repair solutions, which seriously affects development efficiency. Therefore, how to automatically generate effective repair solutions based on existing detection results has become a key requirement for improving software development efficiency and quality. Summary of the invention
[0004] In view of this, an embodiment of the present invention provides a method for intelligently repairing code defects to solve the technical problem that code defects cannot be repaired efficiently and automatically in the prior art. The method includes:
[0005] A knowledge base is constructed using an open source code base and a defect dataset, wherein the knowledge base includes code repair pairs and defect types of code repair pairs, the code repair pairs include the code to be repaired and code repair samples of the code to be repaired, and the defect types include defect information and code rules violated by the code repair pairs;
[0006] The defect type of the code to be detected is obtained through the deep analysis platform, and the code to be detected is converted into the code to be repaired. The code similarity algorithm is used to query the knowledge base for the code repair pair that is most similar to the code to be repaired based on the defect type of the code to be detected, and the most similar code repair pair, defect type and code to be repaired are used as the most similar code information;
[0007] Use a dynamically constructed Prompt to generate targeted code descriptions based on the most similar code information, convert the targeted code descriptions into repaired code through a large language model, and replace and / or integrate the repaired code into the code to be detected.
[0008] An embodiment of the present invention also provides a device for intelligent code defect repair to solve the technical problem in the prior art that code defects cannot be repaired efficiently and automatically. The device includes:
[0009] A knowledge base construction module for constructing a knowledge base using an open source code library and a defect data set. The knowledge base includes code repair pairs and the defect types of the code repair pairs. The code repair pairs include the code to be repaired and the code repair examples of the code to be repaired. The defect types include defect information and the code rules violated by the code repair pairs.
[0010] A most similar code acquisition module for obtaining the defect type of the code to be detected through a deep analysis platform, converting the code to be detected into the code to be repaired, using a code similarity algorithm to query the most similar code repair pair in the knowledge base through the defect type of the code to be detected, and taking the most similar code repair pair, defect type, and code to be repaired as the most similar code information.
[0011] A code defect intelligent repair module for using a dynamically constructed Prompt to generate targeted code descriptions based on the most similar code information, converting the targeted code descriptions into repaired code through a large language model, and replacing and / or integrating the repaired code into the code to be detected.
[0012] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned method for intelligent code defect repair to solve the technical problem in the prior art that code defects cannot be repaired efficiently and automatically.
[0013] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program for executing the above-mentioned method for intelligent code defect repair to solve the technical problem in the prior art that code defects cannot be repaired efficiently and automatically.
[0014] Compared with the prior art, the at least one technical solution adopted in the embodiments of this specification can achieve at least the following beneficial effects:
[0015] By adopting the dynamic Prompt construction technology, it is possible to deeply mine and utilize the valuable data in the knowledge base, provide a solution for the LLM direct repair method, and bring innovative technologies and practical values to the field of code defect repair. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 is a flowchart of a method for intelligent repair of code defects provided by an embodiment of the present invention;
[0018] Figure 2 is a flowchart of a method for implementing the intelligent repair of code defects provided by an embodiment of the present invention;
[0019] Figure 3 is a flowchart of constructing a knowledge base provided by an embodiment of the present invention;
[0020] Figure 4 is a schematic diagram of the code similarity calculation process provided by an embodiment of the present invention;
[0021] Figure 5 is a structural block diagram of a computer device provided by an embodiment of the present invention;
[0022] Figure 6 is a structural block diagram of a device for intelligent repair of code defects provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The following will describe the embodiments of the present application in detail with reference to the drawings.
[0024] The following illustrates the embodiments of the present application through specific examples. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of them. The present application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.
[0025] In the embodiments of the present invention, a method for intelligent repair of code defects is provided, as shown in Figure 1 and Figure 2 The method includes:
[0026] Step S101: Construct a knowledge base using an open-source code library and a defect dataset. The knowledge base includes code repair pairs and the defect types of the code repair pairs. The code repair pairs include the code to be repaired and the code repair examples of the code to be repaired. The defect types include defect information and the code rules violated by the code repair pairs.
[0027] Step S102: Obtain the defect type of the code to be detected through a deep analysis platform, and convert the code to be detected into the code to be repaired. Using a code similarity algorithm, query the most similar code repair pair in the knowledge base based on the defect type of the code to be detected, and take the most similar code repair pair, defect type, and the code to be repaired as the most similar code information.
[0028] Step S103: Use a dynamically constructed Prompt to generate a targeted code description based on the most similar code information, convert the targeted code description into the repaired code through a large language model, and replace and / or integrate the repaired code into the code to be detected.
[0029] Specifically, the deep analysis platform (the integrated deep analysis detection and repair platform) is the intelligent platform relied on in the embodiments of the present invention. This platform can detect potential defects in the code segment and mark the starting and ending positions of the defects according to the detection criteria and configurations selected by the user.
[0030] In specific implementation, the following steps are used to construct a knowledge base using an open-source code library and repair code cases:
[0031] Obtain a list of code rules, where the code rules in the code rule list are related to the industry in which the code is used; according to the code rule list, obtain code repair pairs from the open-source code library and repair code cases, convert the code to be repaired in the code repair pairs into code vectors, and save the code repair pairs, the defect types of the code repair pairs, the code vectors, and the code rule list to the knowledge base; after the intelligent repair of the code to be detected is completed, update the repair success flag to the corresponding code repair pair in the knowledge base and iterate the knowledge base.
[0032] In specific implementation, in order to correspond the code rules with the code repair pairs one by one, the following steps are used to obtain code repair pairs from the open-source code library and repair code cases according to the code rule list:
[0033] Obtain repair code cases and retrieve code defect repair records from the open-source code repository; set filtering rules to screen the code defect repair records that meet the filtering rules and generate filtered code defect repair records; perform standardization processing on the filtered code defect repair records and repair code cases to generate standardized defect repair records; calculate the similarity between the current standardized defect repair record and other standardized defect repair records except the current one through a code similarity algorithm; set a benchmark duplication rate and delete other standardized defect repair records with a similarity greater than the benchmark duplication rate to generate a defect repair dataset; match the defect repair dataset with each code rule in the code rule list to generate code repair pairs and the defect types of the code repair pairs.
[0034] Specifically, as Figure 3 shown, the source code obtained from the open-source community (as the open-source code repository), including Github open-source code, code in actual projects, manually supplemented code, etc., is stored in the form of.cpp,.c,.xml files, etc., and presented and displayed in the system as a tree menu item.
[0035] Based on the detection results of multiple commits of the code in the open-source code repository, use open-source tools such as ctags to compare the code before and after in the commit records of two consecutive commits, and detect whether there is a situation where a rule violation has been repaired based on the in-depth analysis platform. If so, record the repaired code pairs for subsequent processing of saving to the knowledge base.
[0036] Classification of code rules, classify the code of different projects into categories to generate a code rule list. The project code has a strong industry color. For unit projects applying GJB8114, more similar industry codes have a greater possibility of constructing more appropriate Prompts. At the same time, save the vectorized data of some representative case codes under each code rule for querying similarity.
[0037] When specifically implementing, in order to improve the repair success rate of the code repair pairs in the knowledge base, the knowledge base is iterated through the following steps:
[0038] Extract the code repair pairs with successful repairs from the knowledge base through the repair success flag; use the code repair pairs with successful repairs as repair code cases and iterate the knowledge base through the repair code cases.
[0039] Specifically, during the iteration of the knowledge base (code knowledge base), first, extract successful code repair pair examples from the knowledge base. Then, according to the preset filtering rules, screen out the code defect repair records that meet the standards. Subsequently, standardize these records and repair code cases to facilitate the calculation of similarity. By applying the code similarity algorithm to evaluate the duplication degree between records, perform deduplication operations on the records with a high duplication rate to ensure the refinement of the knowledge base. For the deduplicated code repair cases, the system matches them with the code rule list in the knowledge base to identify appropriate rules. Finally, according to the matching results, update the code knowledge base to achieve the iteration and optimization of knowledge, thereby improving the efficiency of the knowledge base in code repair and defect management.
[0040] Specifically, in implementation, the following steps are used to utilize the code similarity algorithm to query the code repair pair most similar to the code to be repaired in the knowledge base according to the defect type of the code to be detected:
[0041] Obtain all code repair pairs with the same defect type as the code to be detected from the knowledge base as the knowledge base code repair pairs; loop through all knowledge base code repair pairs until each knowledge base code repair pair is processed: calculate the similarity between the code to be repaired and the code to be repaired in the knowledge base code repair pair through the pre-trained model to obtain the first vector cosine similarity; calculate the similarity between the code to be repaired and the code to be repaired in the knowledge base code repair pair through the embedding model to obtain the second vector cosine similarity; set the first weight and the second weight, and use the first weight and the second weight to perform weighted processing on the first vector cosine similarity and the second vector cosine similarity to generate the final vector cosine similarity; compare the final vector cosine similarity with the similarity threshold. If the final vector cosine similarity is greater than or equal to the similarity threshold, use the current knowledge base code repair pair as the code repair pair most similar to the code to be repaired.
[0042] Specifically, to improve the performance of the matching between the knowledge base and the code repair pairs, scores can be set for all code repair pairs corresponding to the defect type. When obtaining code repair pairs with the same defect type from the knowledge base, only obtain several code repair pairs with high scores as the knowledge base code repair pairs.
[0043] Specifically, as Figure 4 shown, the pre-trained models include CodeBERT, GraphCodeBERT, etc. The embedding models include Doc2Vec, etc. Through these models, word embeddings are performed on the uploaded code to be repaired to generate the word vector representation of the code to be repaired. Through the word vector representation, the cosine similarity is used to calculate the similarity between the code to be repaired and the code to be repaired in the knowledge base code repair pair.
[0044] In some embodiments, the similarity of code may be calculated only using a pre-trained model. In some embodiments, the similarity of code may be calculated only using a pre-trained model. In other embodiments, the similarities of two pieces of code are calculated separately using a pre-trained model and the pre-trained model, and then the two similarities are weighted to obtain the final code similarity.
[0045] During specific implementation, to improve the accuracy of the LLM large model, the following steps are taken to fine-tune the LLM large model (large language model) through a knowledge base:
[0046] After the knowledge base is constructed, the knowledge base is used as a training set to train the large language model, and the parameters of the large language model are adjusted according to the training results.
[0047] Specifically, using a dynamically constructed Prompt, a targeted code description is generated based on the most similar code information, and the targeted code description is converted into repaired code through the large language model, including:
[0048] First step: As a prompt for the LLM (large language model), use professional syntax for interacting with the model, such as "You are a code automatic specification system. In the following task, I will give a piece of code to be repaired. You will refer to the given rule definitions, violation examples, and compliance examples to repair the code to be repaired. The code violates at least one rule and needs to complete all repairs for all the given rules at once and directly output the repaired result. The error positions have been marked for you. Try to modify the code within the smallest range and only output the repaired result of the code to be repaired without any additional description." A brief and clear prompt can facilitate the large model to locate its own functions and output limitations. This part corresponds to the prompt at the "instruction" level in the LLM system.
[0049] Second step: The second part of the prompt consists of input data, such as "Code to be repaired: {code_to_repair}, all the rules violated by this code: {rule}. The following is a repair case, violation example: {break_sample}, compliance example: {follow_sample}", where code_to_repair and rule come from the upstream system, and break_sample and follow_sample are extracted data from the previous cosine similarity calculation results and concatenated as the final Prompt.
[0050] For the large model base in the LLM (large language model), to meet the requirements of localization, open source, and offline deployment security, after experimental verification, the open source large model Deepseeker of DeepSeek Company is selected as the final base.
[0051] In specific implementation, since replacing code has certain risks, the ability to improve the risk resistance by evaluating the confidence of the code before replacing the code is realized. The availability of the most similar code information is judged through confidence, and application recommendations are made through risk assessment by the following steps:
[0052] Before generating the targeted code expression according to the most similar code information, calculate the confidence of the code repair pair in the most similar code information; if the confidence does not meet the confidence threshold, obtain the static case from the configuration file, and use the information of the static case as the most similar code information; after generating the repaired code, conduct a risk assessment on the repaired code, and perform different degrees of replacement and / or integration on the code to be detected according to the results of the risk assessment.
[0053] Specifically, if the similar code fails to meet the requirements of the confidence threshold, use the static case in the configuration file as a substitute to construct a Prompt applicable to the large model.
[0054] Specifically, evaluate the generated repaired code, and make application recommendations according to the evaluation results, including: if the confidence meets the requirements, directly replace the original code; if there are certain risks in the repair result, provide this solution as a reference for the user to further verify and adjust.
[0055] In this embodiment, a computer device is provided, as Figure 5 shown, including a memory 501, a processor 502, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for intelligent repair of any of the above code defects is implemented.
[0056] Specifically, the computer device can be a computer terminal, a server, or a similar computing device.
[0057] In this embodiment, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program for executing the method for intelligent repair of any of the above code defects.
[0058] Specifically, a computer-readable storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, computer-readable storage media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0059] Based on the same inventive concept, an embodiment of the present invention also provides a device for intelligent code defect repair, as described in the following embodiments. Since the principle of solving problems by the device for intelligent code defect repair is similar to that of the method for intelligent code defect repair, the implementation of the device for intelligent code defect repair can refer to the implementation of the method for intelligent code defect repair, and the repeated parts will not be elaborated. As used hereinafter, the term "unit" or "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0060] Figure 6 is a structural block diagram of the device for intelligent code defect repair according to an embodiment of the present invention, as Figure 6 shown, including: a knowledge base construction module 601, a most similar code acquisition module 602, and a code defect intelligent repair module 603. The following describes this structure.
[0061] The knowledge base construction module 601 is used to construct a knowledge base by using an open-source code library and a defect data set. The knowledge base includes code repair pairs and the defect types of the code repair pairs. The code repair pairs include the code to be repaired and the code repair examples of the code to be repaired. The defect types include defect information and the code rules violated by the code repair pairs;
[0062] The most similar code acquisition module 602 is used to obtain the defect type of the code to be detected through the deep analysis platform, convert the code to be detected into the code to be repaired, use the code similarity algorithm, query the most similar code repair pair to the code to be repaired in the knowledge base according to the defect type of the code to be detected, and use the most similar code repair pair, defect type and code to be repaired as the most similar code information;
[0063] The code defect intelligent repair module 603 is used to use the dynamically constructed Prompt to generate a targeted code description according to the most similar code information, convert the targeted code description into the repaired code through the large language model, and replace and / or integrate the repaired code into the code to be detected.
[0064] In one embodiment, the knowledge base construction module includes:
[0065] The code rule list acquisition unit is used for the code rules in the code rule list to be related to the industry in which the code is used;
[0066] The repair pair acquisition unit is used to obtain code repair pairs from the open source code library and repair code cases according to the code rule list, convert the code to be repaired in the code repair pairs into code vectors, and save the code repair pairs, the defect types of the code repair pairs, the code vectors and the code rule list to the knowledge base;
[0067] The iterative knowledge base unit is used to update the repair success flag to the corresponding code repair pair in the knowledge base after the intelligent repair of the code to be detected is completed, and iterate the knowledge base.
[0068] In one embodiment, the repair pair acquisition unit is used to obtain repair code cases and obtain code defect repair records from the open source code library; set filtering rules to screen the code defect repair records that meet the filtering rules to generate filtered code defect repair records; perform standardization processing on the filtered code defect repair records and repair code cases to generate standardized defect repair records; calculate the similarity between the current standardized defect repair record and other standardized defect repair records except the current standardized defect repair record through the code similarity algorithm; set a benchmark repetition rate, delete other standardized defect repair records with a similarity greater than the benchmark repetition rate to generate a defect repair data set; match the defect repair data set with each code rule in the code rule list to generate code repair pairs and the defect types of the code repair pairs.
[0069] In one embodiment, the iterative knowledge base unit is used to extract the code repair pairs with successful repairs from the knowledge base through the repair success flag; use the code repair pairs with successful repairs as repair code cases, and iterate the knowledge base through the repair code cases.
[0070] In one embodiment, the most similar code acquisition module includes:
[0071] A filtered code repair pair unit for obtaining all code repair pairs with the same defect type as the code to be detected from the knowledge base as knowledge base code repair pairs;
[0072] A loop unit for looping through all knowledge base code repair pairs until each knowledge base code repair pair has been processed:
[0073] A first vector cosine similarity calculation unit for calculating the similarity between the code to be repaired and the code to be repaired in the knowledge base code repair pair through a pre-trained model to obtain a first vector cosine similarity;
[0074] A second vector cosine similarity calculation unit for calculating the similarity between the code to be repaired and the code to be repaired in the knowledge base code repair pair through an embedding model to obtain a second vector cosine similarity;
[0075] A final vector cosine similarity calculation unit for setting a first weight and a second weight, and generating a final vector cosine similarity after performing weighted processing on the first vector cosine similarity and the second vector cosine similarity using the first weight and the second weight;
[0076] A code repair pair acquisition unit for comparing the final vector cosine similarity with a similarity threshold, and if the final vector cosine similarity is greater than or equal to the similarity threshold, taking the current knowledge base code repair pair as the code repair pair most similar to the code to be repaired.
[0077] In one embodiment, the above device further includes: a model training module.
[0078] In one embodiment, the model training module includes:
[0079] A model adjustment unit for, after the knowledge base is constructed, using the knowledge base as a training set to train a large language model, and adjusting the parameters of the large language model according to the training results.
[0080] In one embodiment, the above device further includes: an evaluation module.
[0081] In one embodiment, the evaluation module includes:
[0082] A confidence calculation unit for calculating the confidence of the code repair pair in the most similar code information before generating a targeted code expression according to the most similar code information;
[0083] A static case acquisition unit for, if the confidence does not meet the confidence threshold, obtaining a static case from a configuration file and using the information of the static case as the most similar code information;
[0084] A risk assessment unit, which is used to generate the repaired code and then conduct a risk assessment on the repaired code, and perform different degrees of replacement and / or integration on the code to be detected according to the results of the risk assessment.
[0085] The embodiments of the present invention achieve the following technical effects:
[0086] The method for intelligent repair of code defects in the embodiments of the present invention improves the existing code repair technology based on large language models (LLMs), and effectively overcomes the challenges of automatic code defect repair under multiple standards; the method for intelligent repair of code defects in the present invention can give full play to the excellent ability of the code large model in code understanding without relying on a large-scale data set, and achieve efficient and accurate defect repair; by integrating the rich resources of the knowledge base, not only the repair quality is improved, but also a good interaction and iterative optimization with the knowledge base are realized; the dynamic Prompt construction technology is adopted, which can deeply excavate and utilize the valuable data in the knowledge base, providing a solution that surpasses the traditional direct repair method of LLMs, bringing innovative technologies and practical values to the field of code defect repair.
[0087] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present invention can be implemented by a general computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in the storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. Thus, the embodiments of the present invention are not limited to any specific combination of hardware and software.
[0088] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the embodiments of the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for intelligently repairing code defects, characterized in that: include: A knowledge base is constructed using an open source code base and code repair cases, wherein the knowledge base includes code repair pairs and defect types of the code repair pairs, the code repair pairs include codes to be repaired and code repair samples of the codes to be repaired, and the defect types include defect information and code rules violated by the code repair pairs; Obtaining the defect type of the code to be detected through the deep analysis platform, and converting the code to be detected into the code to be repaired, using the code similarity algorithm, searching the knowledge base for the code repair pair that is most similar to the code to be repaired through the defect type of the code to be detected, and taking the most similar code repair pair, the defect type and the code to be repaired as the most similar code information; Using a dynamically constructed prompt, a targeted code description is generated according to the most similar code information, the targeted code description is converted into a repaired code through a large language model, and the repaired code is replaced and / or integrated into the code to be detected.
2. The method for intelligently repairing code defects according to claim 1, characterized in that: Build a knowledge base using open source code base and fixed code cases, including: Obtaining a code rule list, wherein the code rules in the code rule list are related to the industry in which the code is used; According to the code rule list, the code repair pair is obtained from the open source code base and the repair code case, the code to be repaired of the code repair pair is converted into a code vector, and the code repair pair, the defect type of the code repair pair, the code vector and the code rule list are saved in a knowledge base; After the intelligent repair of the code to be detected is completed, a repair success flag is updated to the code repair pair corresponding to the knowledge base, and the knowledge base is iterated.
3. The method for intelligently repairing code defects according to claim 2, characterized in that: According to the code rule list, obtaining the code repair pair from the open source code base and the repair code case includes: Obtain the repair code case, and obtain the code defect repair record from the open source code library; Setting filtering rules, screening the code defect repair records that meet the filtering rules, and generating filtered code defect repair records; Standardizing the filtered code defect repair records and the repair code cases to generate standardized defect repair records; Calculate the similarity between the current standardized defect repair record and other standardized defect repair records other than the current standardized defect repair record by using a code similarity algorithm; Setting a baseline repetition rate, deleting the other standardized defect repair records whose similarity is greater than the baseline repetition rate, and generating a defect repair data set; The defect repair data set is matched one by one with each of the code rules in the code rule list to generate the code repair pair and the defect type of the code repair pair.
4. The method for intelligently repairing code defects according to claim 2, characterized in that: Iterating the knowledge base, including: Extracting the successfully repaired code repair pair from the knowledge base through the successful repair mark; The successfully repaired code repair pair is used as the repair code case, and the knowledge base is iterated through the repair code case.
5. The method for intelligently repairing code defects according to claim 1, characterized in that: Using a code similarity algorithm, searching the knowledge base for the code repair pair that is most similar to the code to be repaired according to the defect type of the code to be detected, including: Acquire from the knowledge base all code repair pairs with the same defect type as the defect type of the code to be detected as knowledge base code repair pairs; Loop through all knowledge base code repair pairs until each of the knowledge base code repair pairs is processed: Calculating the similarity between the code to be repaired and the code to be repaired of the knowledge base code repair pair through the pre-trained model to obtain a first vector cosine similarity; Calculating the similarity between the code to be repaired and the code to be repaired of the knowledge base code repair pair through the embedding model to obtain a second vector cosine similarity; Setting a first weight and a second weight, and performing weighted processing on the first vector cosine similarity and the second vector cosine similarity using the first weight and the second weight to generate a final vector cosine similarity; The final vector cosine similarity is compared with a similarity threshold, and if the final vector cosine similarity is greater than or equal to the similarity threshold, the current knowledge base code repair pair is taken as the code repair pair that is most similar to the code to be repaired.
6. The method for intelligently repairing code defects according to any one of claims 1 to 5, characterized in that: Also includes: After the knowledge base is constructed, the knowledge base is used as a training set, the large language model is trained through the training set, and the parameters of the large language model are adjusted according to the training results.
7. The method for intelligently repairing code defects according to any one of claims 1 to 5, characterized in that: Also includes: Before generating the targeted code representation according to the most similar code information, calculating the confidence of the code repair pair in the most similar code information; If the confidence does not meet the confidence threshold, a static case is obtained from a configuration file, and information of the static case is used as the most similar code information; After the repaired code is generated, a risk assessment is performed on the repaired code, and the code to be detected is replaced and / or integrated to varying degrees according to the result of the risk assessment.
8. A device for intelligently repairing code defects, characterized in that: include: A knowledge base construction module, used to construct a knowledge base using an open source code base and a repair code case, wherein the knowledge base includes code repair pairs and defect types of the code repair pairs, the code repair pairs include code to be repaired and code repair samples of the code to be repaired, and the defect types include defect information and code rules violated by the code repair pairs; The most similar code acquisition module is used to obtain the defect type of the code to be detected through the deep analysis platform, and convert the code to be detected into the code to be repaired, and use the code similarity algorithm to query the code repair pair that is most similar to the code to be repaired in the knowledge base through the defect type of the code to be detected, and use the most similar code repair pair, the defect type and the code to be repaired as the most similar code information; The code defect intelligent repair module is used to use a dynamically constructed prompt to generate a targeted code description according to the most similar code information, convert the targeted code description into a repaired code through a large language model, and replace and / or integrate the repaired code into the code to be detected.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for intelligently repairing code defects according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for executing the method for intelligently repairing code defects according to any one of claims 1 to 7.