AI code defect repair data set construction method, defect repair method and equipment

By building a high-quality AI code defect repair dataset, the AI ​​code defect repair model performs poorly when dealing with complex programming languages ​​and multi-scene code defects, achieving more efficient code defect detection and repair effects.

CN119938492APending Publication Date: 2025-05-06BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411861567.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing AI code defect repair models perform poorly when dealing with complex and diverse programming languages ​​and multi-scene code defects, making it difficult to cope with the complexity of programming languages ​​and multi-scene code defects.

Method used

A method for constructing AI code defect repair data sets is proposed. By obtaining multiple sets of initial AI code data, mutation processing, deduplication and context extraction are performed, and structured AI code fragments are generated according to the CWE standard, and thinking chain data sets are constructed to systematically collect and annotate data.

Benefits of technology

Generate high-quality data sets, enhance the generalization ability of the model, improve the accuracy and repair effect of code defect detection, and better deal with complex and diverse programming languages ​​and multi-scenario code defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938492A_ABST
    Figure CN119938492A_ABST
Patent Text Reader

Abstract

The invention provides an AI code defect repair data set construction method, a defect repair method and equipment. The data set construction method comprises the steps of obtaining multiple groups of initial AI code data; each group of initial AI code data comprises defect data and corresponding repair data; performing variation processing on the multiple groups of initial AI code data; each group of variation AI code data comprises variation defect data and corresponding variation repair data; performing duplicate removal and context extraction processing on the multiple groups of initial AI code data and the multiple groups of variant AI code data to obtain multiple groups of AI code snippets; classifying and labeling the multiple groups of AI code snippets according to a CWE standard to obtain multiple groups of structured AI code snippets; the multiple sets of structured AI code snippets are preprocessed, so that the data formats corresponding to the multiple sets of structured AI code data are consistent; and respectively generating thinking chain data for each group of preprocessed structured AI code snippets to obtain an AI code data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method for constructing an AI code defect repair dataset, a defect repair method and a device. Background Art

[0002] Traditional methods of detecting and repairing code defects mainly rely on technical means such as static analysis, dynamic analysis, and manual testing. However, with the increase in code size and the complexity of programming languages, the efficiency of defect identification and repair in complex scenarios by traditional methods has gradually decreased, showing problems such as incomplete detection and high false positive rate.

[0003] In order to improve the automation level of code defect detection and repair, large models based on deep learning technology have been applied in recent years. The performance of these large models in AI code defect repair tasks is still unsatisfactory, and it is difficult to deal with complex and diverse programming languages ​​and code defects in multiple scenarios. Summary of the invention

[0004] In view of this, the purpose of this application is to propose an AI code defect repair data set construction method, defect repair method and equipment.

[0005] Based on the above objectives, this application provides a method for constructing an AI code defect repair dataset, including:

[0006] Acquire multiple groups of initial AI code data; each group of initial AI code data includes defect data and corresponding repair data; the multiple groups of initial AI code data have multiple types of programming languages, and each group of initial AI code data corresponds to one type of programming language;

[0007] Performing mutation processing on the multiple groups of initial AI code data to obtain multiple groups of mutated AI code data to expand the initial AI code data; each group of mutated AI code data includes mutated defect data and corresponding mutated repair data;

[0008] Deduplication and context extraction are performed on the multiple groups of initial AI code data and the multiple groups of variant AI code data to obtain multiple groups of AI code snippets; each group of AI code snippets includes a vulnerability AI code snippet and a corresponding repair AI code snippet;

[0009] Classifying and annotating the multiple groups of AI code snippets according to the CWE standard to obtain multiple groups of structured AI code snippets;

[0010] Preprocessing the multiple groups of structured AI code fragments to make the data formats corresponding to the multiple groups of structured AI code data consistent;

[0011] For each group of structured AI code snippets after preprocessing, corresponding thought chain data is generated to obtain a data set of the AI ​​code; the thought chain data includes logical reasoning from the vulnerability AI code snippet to the corresponding repair AI code snippet.

[0012] In some embodiments, performing mutation processing on the multiple groups of initial AI code data to obtain multiple groups of mutated AI code data includes:

[0013] For each group of initial AI code data, at least one of grammatical symbol replacement, semantic equivalence transformation and control flow structure adjustment is performed to obtain corresponding variant AI code data.

[0014] In some embodiments, the multiple groups of initial AI code data and the multiple groups of variant AI code data are subjected to deduplication and context extraction processing respectively to obtain multiple groups of AI code snippets, including:

[0015] Performing syntax analysis on each set of defect data and corresponding repair data in the multiple sets of initial AI code data and the multiple sets of variant AI code data, performing text comparison on functions with the same name, and obtaining defect code blocks related to the defects and corresponding repair code blocks;

[0016] Performing semantic analysis on the functions in the code block to obtain defective code segments and corresponding repair code segments that are directly related to the defect, and marking defective code segments and corresponding repair code segments that are not related to the defect;

[0017] Compare the defective code segments directly related to the defect and the corresponding repair code segments, identify mutations, extract the vulnerability code segments and the corresponding repair code segments, delete irrelevant code segments, and delete the marked defective code segments and the corresponding repair code segments that do not involve the defect.

[0018] In some embodiments, the multiple groups of AI code snippets are classified and annotated according to the CWE standard to obtain multiple groups of structured AI code snippets, including:

[0019] Identify and annotate the defect types of each set of AI code snippets;

[0020] Add a label that corresponds to the defect type.

[0021] In some embodiments, the defect type includes a memory error, a missing validation, or an incorrect incoming data type; or

[0022] The preprocessing includes replacing symbols, filtering comments or standardizing variable names.

[0023] In some of the embodiments, the thought chain data includes vulnerability AI code snippets, corresponding problem descriptions, reasoning process descriptions, and example repair codes and repair suggestions.

[0024] The present application also provides a method for repairing defects in AI code, including:

[0025] Get the AI ​​code to be fixed.

[0026] The AI ​​code to be defect-repaired is input into a defect-repair model of the AI ​​code; wherein the defect-repair model of the AI ​​code is fine-tuned based on a data set of the AI ​​code constructed by the method for constructing a data set of defect-repairing AI code as described in any of the preceding items.

[0027] An embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the methods described above when executing the program.

[0028] An embodiment of the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute any of the methods described above.

[0029] An embodiment of the present application further provides a computer program product, comprising computer program instructions, which, when executed on a computer, enable the computer to execute any of the methods described above.

[0030] From the above, it can be seen that the AI ​​code defect repair data set construction method provided by the present application obtains multiple groups of initial AI code data; each group of initial AI code data includes defect data and corresponding repair data; the multiple groups of initial AI code data have multiple types of programming languages, and each group of initial AI code data corresponds to one type of programming language; the multiple groups of initial AI code data are mutated to obtain multiple groups of mutated AI code data to expand the initial AI code data; each group of mutated AI code data includes mutated defect data and corresponding mutated repair data; the multiple groups of initial AI code data and the multiple groups of mutated AI code data are deduplicated and context extracted to obtain multiple groups of AI code fragments; each group of AI code fragments includes missing or missing data. vulnerabilities AI code snippets and corresponding repair AI code snippets; classify and annotate the multiple groups of AI code snippets according to the CWE standard to obtain multiple groups of structured AI code snippets; preprocess the multiple groups of structured AI code snippets to make the data formats corresponding to the multiple groups of structured AI code data consistent; generate corresponding thinking chain data for each group of structured AI code snippets after preprocessing to obtain the AI ​​code data set; the thinking chain data includes logical reasoning from the vulnerability AI code snippet to the corresponding repair AI code snippet; can systematically carry out data collection, deduplication, annotation and prompt design process, generate high-quality data sets, enhance the generalization ability of the model obtained based on the data set, and improve the accuracy of code defect detection and repair effect of the obtained model. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the present application or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0032] Figure 1 A flowchart of a method for constructing an AI code defect repair dataset according to an embodiment of the present application;

[0033] Figure 2 A schematic diagram of a data deduplication processing flow chart of an embodiment of the present application;

[0034] Figure 3 A schematic diagram of a thought chain data format according to an embodiment of the present application;

[0035] Figure 4 A flowchart of a defect repair method for AI code according to an embodiment of the present application;

[0036] Figure 5A schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0037] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0038] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be understood by people with ordinary skills in the field to which the present application belongs. The "first", "second" and similar words used in the embodiments of the present application do not represent any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing in front of the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.

[0039] In the related art, in order to improve the level of automation of code defect detection and repair, large model applications based on deep learning technology have emerged in recent years. These large models use a large amount of code data for pre-training and have good potential in identifying code defects. However, the proprietary nature of AI code is relatively strong, resulting in a scarcity of relevant public data sets, and the performance of existing large models in AI code defect repair tasks is still unsatisfactory. The uneven quality of data sets for code defect repair tasks and the lack of contextual understanding and reasoning capabilities often make it difficult for models to cope with complex and diverse programming languages ​​and code defects in multiple scenarios. Therefore, in the related art, there are problems such as the difficulty of AI code defect repair models in coping with the complexity of programming languages ​​and the difficulty in dealing with code defects in multiple scenarios.

[0040] Based on this, the embodiments of the present application provide a data set construction method for AI code defect repair tasks, an AI code defect repair method and related equipment, which generates a high-quality fine-tuning training data set through a systematic data collection, deduplication and annotation process to enhance the performance of large models in AI code defect detection and repair tasks. It can take into account the preprocessing requirements of multi-language codes, and ultimately generate a data set format that meets the requirements of large model fine-tuning, thereby improving the generalization ability and detection accuracy of AI models in actual code repair tasks. It can solve the problems that AI code defect repair models are difficult to cope with the complexity of programming languages, and are difficult to deal with code defects in multiple scenarios.

[0041] like Figure 1 As shown, the embodiment of the present application provides a method for constructing an AI code defect repair dataset, which may include:

[0042] Step S100, obtaining multiple groups of initial AI code data; each group of initial AI code data includes defect data and corresponding repair data; the multiple groups of initial AI code data have multiple types of programming languages, and each group of initial AI code data corresponds to one type of programming language;

[0043] Step S200, performing mutation processing on the multiple groups of initial AI code data to obtain multiple groups of mutated AI code data to expand the initial AI code data; each group of mutated AI code data includes mutated defect data and corresponding mutated repair data;

[0044] Step S300, performing deduplication processing and context extraction processing on the multiple groups of initial AI code data and the multiple groups of variant AI code data, respectively, to obtain multiple groups of AI code snippets; each group of AI code snippets includes a vulnerability AI code snippet and a corresponding repair AI code snippet;

[0045] Step S400, classifying and annotating the multiple groups of AI code snippets according to the CWE standard to obtain multiple groups of structured AI code snippets;

[0046] Step S500, preprocessing the multiple groups of structured AI code snippets to make the data formats corresponding to the multiple groups of structured AI code data consistent;

[0047] Step S600, for each group of structured AI code snippets after preprocessing, generate corresponding thought chain data respectively to obtain the fine-tuning data set for the AI ​​code defect repair task; the thought chain data includes logical reasoning from the vulnerability AI code snippet to the corresponding repair AI code snippet.

[0048] The method for constructing an AI code defect repair data set provided in an embodiment of the present application obtains multiple groups of initial AI code data; each group of initial AI code data includes defect data and corresponding repair data; the multiple groups of initial AI code data have multiple types of programming languages, and each group of initial AI code data corresponds to one type of programming language; the multiple groups of initial AI code data are mutated to obtain multiple groups of mutated AI code data to expand the initial AI code data; each group of mutated AI code data includes mutated defect data and corresponding mutated repair data; the multiple groups of initial AI code data and the multiple groups of mutated AI code data are deduplicated and context extracted to obtain multiple groups of AI code snippets; each group of AI code snippets includes vulnerability AI code snippets code snippets and corresponding repair AI code snippets; classify and annotate the multiple groups of AI code snippets according to the CWE standard to obtain multiple groups of structured AI code snippets; preprocess the multiple groups of structured AI code snippets to make the data formats corresponding to the multiple groups of structured AI code data consistent; generate corresponding thinking chain data for each group of structured AI code snippets after preprocessing to obtain a data set of the AI ​​code; the thinking chain data includes logical reasoning from the vulnerability AI code snippet to the corresponding repair AI code snippet; can systematically carry out data collection, deduplication, annotation and prompt design process, generate high-quality data sets, enhance the generalization ability of the model obtained based on the data set, and improve the accuracy of code defect detection and repair effect of the obtained model.

[0049] In some of the embodiments, in step S100, multiple sets of initial AI code data can be automatically acquired through crawlers and data interfaces. Among them, the defect data and corresponding repair data of the AI ​​code can be acquired from resources that can publicly acquire data, such as open source code libraries, vulnerability databases, and online programming platforms. The multiple sets of initial AI code data acquired can include multiple types of programming languages. For any set of initial AI code data, there is usually only one type of programming language, for example, it can be any one of multiple languages ​​such as C, Java, and Python. Exemplarily, multiple sets of initial AI code data can be acquired from the CrossVul dataset.

[0050] In some embodiments, in step S200, performing mutation processing on the multiple groups of initial AI code data to obtain multiple groups of mutated AI code data may include:

[0051] For each group of initial AI code data, at least one of grammatical symbol replacement, semantic equivalence transformation and control flow structure adjustment is performed to obtain corresponding variant AI code data.

[0052] In this way, by expanding the acquired AI code-related defect data and corresponding repair data through mutation technology, the diversity of AI code data can be increased to cover more scenarios.

[0053] In some embodiments, the symbol replacement may be to perform symbol replacement on at least one set of initial AI code data at the syntax level to obtain at least one set of first variant AI code data. By performing symbol replacement at the syntax level, the code can still maintain grammatical correctness after the variant.

[0054] In some embodiments, the semantic equivalent transformation may be a renaming operation. The renaming operation may be to rename at least one of the at least one set of initial AI code data and the at least one set of first variant AI code data by renaming variables or functions, generating a new version of code, and obtaining at least one set of second variant AI code data. The logic of the second variant AI code data is the same as the logic of the corresponding initial AI code data. By renaming the variables or functions, the code logic can be kept unchanged and a new version of the code can be obtained.

[0055] In some embodiments, the control flow structure adjustment may be to adjust the control flow structure of at least one of at least one set of initial AI code data, at least one set of first variant AI code data, and at least one set of second variant AI code data to obtain at least one set of third variant AI code data. The format of the third variant AI code data is logically equivalent to the corresponding initial AI code data / first variant AI code data / second variant AI code data. In this way, variant AI code fragments that are logically equivalent but have different structures can be obtained.

[0056] In some of the embodiments, a UniversalMutator tool may be used to perform mutation processing on multiple groups of initial AI code data to obtain corresponding mutated AI code data.

[0057] In some embodiments, step S300 can be understood as performing deduplication processing based on the text difference algorithm of git diff and AST parsing technology. Use the diff tool to compare functions with the same name to determine the code change area, then use the AST parsing tool to analyze the grammatical structure of the function to identify the changed parts in the grammatical level, use the Context2Vec algorithm to filter out irrelevant contexts, and only retain the core vulnerabilities and repair codes. Figure 2 As shown, the context extraction processing is performed on the multiple groups of initial AI code data and the multiple groups of variant AI code data respectively to obtain multiple groups of AI code fragments, which may include:

[0058] For each set of defect data and corresponding repair data in the multiple sets of initial AI code data and the multiple sets of variant AI code data, syntax parsing is performed, and text comparison is performed on functions with the same name to obtain defect code blocks related to the defects and corresponding repair code blocks.

[0059] Perform semantic analysis on the functions in the code block to obtain defective code segments and corresponding repair code segments that are directly related to the defect, and mark defective code segments and corresponding repair code segments that are not related to the defect. Usually, defective code segments and corresponding repair code segments that are not related to the defect can be auxiliary functions or initialization codes. When marking, they can be marked as "irrelevant context".

[0060] Compare the defective code segments and corresponding repair code segments that are directly related to the defect, identify mutations, extract the vulnerable code segments and corresponding repair code segments, delete irrelevant code segments, and delete the marked defective code segments and corresponding repair code segments that are not related to the defect. Generally, this step can be achieved by analyzing differences, function dependencies, and variable usage, so that the remaining code is concentrated on the part directly related to the defect.

[0061] In this way, the core parts of the vulnerability code and the repair code in the AI ​​code data can be obtained. And by extracting the modified parts from the repair and defect files, the obtained vulnerability code fragments can have good correspondence with the corresponding repair code fragments.

[0062] In some of the embodiments, it may also include reorganizing and formatting the code segments after deleting the defective code segments that do not involve the defect, and reorganizing and formatting the corresponding repair code segments after deleting the repair code segments that do not involve repairing the defect, so that each group of initial AI code data and each group of mutated AI code data can still maintain grammatical integrity after removing irrelevant information, thereby facilitating normal input and learning of the model.

[0063] In some embodiments, the deduplication process may be to remove duplicate data. The deduplication process may be implemented by a text difference algorithm based on git diff and an AST parsing technique.

[0064] In some of the embodiments, the grammatical structure of the code can be parsed by using an AST parsing tool, the syntax tree of the function can be deeply analyzed, the specific syntax level of the mutation and modification can be identified, and the functions involved in the design defects or repairs can be found, usually the functions with the same name. By matching the function name and parameters, the code blocks directly related to the defects can usually be screened out. In practical applications, each set of defect data and corresponding repair data, as well as each set of variant defect data and corresponding variant repair data can be composed of multiple functions or code segments. In the steps of deduplication processing and context extraction processing, the first step can be to extract each function or logical code block in each set of initial AI code data or each set of variant AI code data through an AST parser.

[0065] In some of the embodiments, the git diff tool can be used to perform text comparison on different versions of the same-name function - vulnerability code and repair code (that is, multiple sets of initial AI code data and each set of defect data and corresponding repair data in the multiple sets of variant AI code data), identify the modified areas in the code (that is, the difference between each defect code and the repair code), and extract the specific code lines involved in the repair part.

[0066] In some of these instances, the Context2Vec algorithm can be used to perform semantic analysis on the code context, filter out code context that is not related to vulnerability repair, and retain only the core vulnerabilities and repair codes, ensuring that the most representative code snippets are used for subsequent model training through the dataset.

[0067] In some of the examples, in step S400, the classification and annotation of the multiple groups of AI code snippets according to the CWE standard to obtain multiple groups of structured AI code snippets may include:

[0068] Identify and annotate the defect types of each set of AI code snippets;

[0069] Add a label that corresponds to the defect type.

[0070] In this way, we can get structured AI code snippets.

[0071] In some embodiments, the defect type may include incorrect input data type, memory error, missing verification, etc. By classifying specific defect types, different types of defect data can be correctly labeled, which facilitates the formation of structured training data.

[0072] In some of the embodiments, appropriate labels can be assigned to different vulnerability types based on the annotation results to facilitate the supervised learning of subsequent training of the AI ​​code defect repair model.

[0073] In some embodiments, in step S500, the preprocessing may include operations such as symbol replacement, comment filtering, or variable name standardization, and the preprocessing may make the data formats of different programming languages ​​consistent. In specific processing, corresponding symbol replacement, comment filtering, or variable name standardization operations may be performed according to the characteristics of different programming languages.

[0074] In some embodiments, in step S600, Figure 3 As shown, the thought chain data may include the corresponding reasoning data in the process of obtaining the corresponding repair AI code snippet from the vulnerability AI code snippet. The reasoning data may include the vulnerability AI code snippet, the problem description corresponding to the vulnerability AI code snippet, the reasoning process description corresponding to the vulnerability AI code snippet, and the sample repair code and repair suggestion corresponding to the vulnerability AI code snippet. Generally, the thought chain data may be data in JSON format.

[0075] In some embodiments, the problem description is used to provide a brief and clear description to help the model understand the nature of the error. For example, for a vulnerability in which the incoming data type is incorrect, the problem description can be "This code has a data processing API compatibility error in which the incoming data type is incorrect." In this way, this description can help the subsequent AI code defect repair model focus on this specific defect type during analysis.

[0076] In some embodiments, the reasoning process description may include guiding the model to identify the type and location of the vulnerability and analyze possible solutions. For example, in code analysis, the model may first be prompted to find the type of vulnerability, then find the specific location that triggers the vulnerability, and finally propose feasible repair measures.

[0077] In some embodiments, the sample repair code and repair suggestions may include relevant code snippet examples to enhance the model's understanding of the problem after the problem is prompted. The sample repair code may include a revised version corresponding to the original error code, enabling the model to reason in a clear context and generate a correct repair solution.

[0078] In this way, through this kind of "thinking chain" prompt design, the large model obtained by subsequent fine-tuning can understand and repair complex code defects in stages. By combining the real code context, the model obtained by subsequent fine-tuning can more accurately reason about the code repair process, and the model can reason in a clear context and generate the correct repair plan.

[0079] The AI ​​code defect repair data set construction method provided in the embodiment of the present application is classified and labeled based on the CWE standard, uses an automated script to remove duplicates, removes redundant information by comparing the differences between the vulnerability code and the repair code, introduces a chain of thinking prompt design, and gradually guides the model to reason, so that it can more accurately understand and solve code defect problems. Among them, the data is labeled based on the CWE standard, and the high coverage and accuracy of the data set are ensured by classifying and marking the vulnerabilities in the code. Among them, automatic deduplication analyzes the differences between the functions of the same name in the vulnerability code and the repair code, removes redundant or irrelevant context code, and improves the purity and effectiveness of the data set. This method can generate high-quality code defect and repair data, and designs a chain of thinking prompt structure to enhance the model's understanding and repair capabilities in code defect detection. Among them, the chain of thinking prompt design can gradually guide the large model to reason and improve the model's analysis depth of code defects and the accuracy of repair decisions.

[0080] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.

[0081] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can independently choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0082] As an optional but non-limiting implementation, in response to receiving the user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0083] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0084] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the described method.

[0085] It should be noted that the above describes some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0086] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a method for repairing defects in AI code.

[0087] like Figure 4 As shown, the defect repair method of the AI ​​code may include:

[0088] Step S810, obtaining the AI ​​code to be repaired;

[0089] Step S820: Input the AI ​​code to be defect-repaired into the defect-repair model of the AI ​​code to perform defect repair. The defect-repair model of the AI ​​code is fine-tuned based on the AI ​​code dataset constructed by the AI ​​code defect-repair dataset construction method as described in any of the preceding items.

[0090] In some of the embodiments, the defect repair model of AI code obtained by fine-tuning the AI ​​code dataset constructed based on the AI ​​code defect repair dataset construction method as described in any of the preceding items may be a Codeshell-7B model.

[0091] In some of the embodiments, the defect repair model of the AI ​​code can be obtained by fine-tuning the pre-trained defect repair model of the AI ​​code based on the data set of the AI ​​code. The specific fine-tuning can be Lora fine-tuning, and the parameters of the low-rank matrix part are updated by low-rank matrix decomposition technology, and the remaining weights are kept fixed. The fine-tuning process can be implemented using the deep learning framework PyTorch, combined with the Adam optimizer for gradient descent optimization.

[0092] In some of the embodiments, in order to avoid the overfitting problem, verification can be performed during fine-tuning training of the AI ​​code defect repair model. A 5-fold cross-validation technique is used. When fine-tuning the model, the constructed code defect and repair dataset (that is, the constructed AI code dataset) is randomly divided into five subsets. Four of the subsets are used for model training each time, and the remaining subset is used as a verification set.

[0093] By selecting Lora fine-tuning technology and gradient descent optimization algorithm as the fine-tuning method, the amount of parameter adjustment for large models can be effectively reduced while maintaining the performance of the model. During the fine-tuning process, the Lora fine-tuning technology is used through supervised learning to efficiently adjust model parameters for specific tasks, overcoming the overfitting problem that may be caused by traditional full fine-tuning. Lora fine-tuning adds low-rank matrices to specific layers of the model, allowing the model to perform fine-grained optimization while maintaining its original capabilities. In this way, not only the training efficiency of the model is improved, but also the generalization performance of the model in code defect detection is effectively enhanced.

[0094] For the convenience of description, the above device is described in terms of functions divided into various modules. Of course, when implementing the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0095] The device of the above embodiment is used to implement the corresponding AI code defect repair data set construction method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0096] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for constructing an AI code defect repair dataset described in any of the above embodiments is implemented.

[0097] Figure 5 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.

[0098] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0099] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0100] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0101] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0102] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0103] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0104] The electronic device of the above-mentioned embodiment is used to implement the corresponding AI code defect repair data set construction method in any of the above-mentioned embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0105] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the AI ​​code defect repair dataset construction method described in any of the above embodiments.

[0106] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0107] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the AI ​​code defect repair data set construction method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0108] Based on the same inventive concept, corresponding to the AI ​​code defect repair dataset construction method described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer so that the computer and / or the processor execute the AI ​​code defect repair dataset construction method. Corresponding to the execution subject corresponding to each step in each embodiment of the AI ​​code defect repair dataset construction method, the processor that executes the corresponding step may belong to the corresponding execution subject.

[0109] The computer program product of the above embodiment is used to enable the computer and / or the processor to execute the AI ​​code defect repair dataset construction method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0110] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0111] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present application difficult to understand, the known power supply / ground connection with the integrated circuit (IC) chip and other components may or may not be shown in the provided drawings. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform to be implemented in the embodiments of the present application (that is, these details should be fully within the scope of understanding of those skilled in the art). In the case of elaborating specific details (e.g., circuits) to describe exemplary embodiments of the present application, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.

[0112] Although the present application has been described in conjunction with specific embodiments of the present application, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0113] The embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the present application.

Claims

1. A method for constructing an AI code defect repair dataset, characterized in that: include: Obtain multiple sets of initial AI code data; Each set of initial AI code data includes defect data and corresponding repair data; the multiple sets of initial AI code data have multiple types of programming languages, and each set of initial AI code data corresponds to one type of programming language; Performing mutation processing on the multiple groups of initial AI code data to obtain multiple groups of mutated AI code data to expand the initial AI code data; each group of mutated AI code data includes mutated defect data and corresponding mutated repair data; Deduplication and context extraction are performed on the multiple groups of initial AI code data and the multiple groups of variant AI code data to obtain multiple groups of AI code snippets; each group of AI code snippets includes a vulnerability AI code snippet and a corresponding repair AI code snippet; Classifying and annotating the multiple groups of AI code snippets according to the CWE standard to obtain multiple groups of structured AI code snippets; Preprocessing the multiple groups of structured AI code fragments to make the data formats corresponding to the multiple groups of structured AI code data consistent; For each group of structured AI code snippets after preprocessing, corresponding thought chain data is generated to obtain a data set of the AI ​​code; the thought chain data includes logical reasoning from the vulnerability AI code snippet to the corresponding repair AI code snippet.

2. The method for constructing an AI code defect repair dataset according to claim 1, characterized in that: The performing mutation processing on the multiple groups of initial AI code data to obtain multiple groups of mutated AI code data includes: For each group of initial AI code data, at least one of grammatical symbol replacement, semantic equivalence transformation and control flow structure adjustment is performed to obtain corresponding variant AI code data.

3. The method for constructing an AI code defect repair dataset according to claim 1, characterized in that: The deduplication and context extraction processing are respectively performed on the multiple groups of initial AI code data and the multiple groups of variant AI code data to obtain multiple groups of AI code fragments, including: Performing syntax analysis on each set of defect data and corresponding repair data in the multiple sets of initial AI code data and the multiple sets of variant AI code data, performing text comparison on functions with the same name, and obtaining defect code blocks related to the defects and corresponding repair code blocks; Performing semantic analysis on the functions in the code block to obtain defective code segments and corresponding repair code segments that are directly related to the defect, and marking defective code segments and corresponding repair code segments that are not related to the defect; Compare the defective code segments directly related to the defect and the corresponding repair code segments, identify mutations, extract the vulnerability code segments and the corresponding repair code segments, delete irrelevant code segments, and delete the marked defective code segments and the corresponding repair code segments that do not involve the defect.

4. The method for constructing an AI code defect repair dataset according to claim 3, characterized in that: The multiple groups of AI code snippets are classified and annotated according to the CWE standard to obtain multiple groups of structured AI code snippets including: Identify and annotate the defect types of each set of AI code snippets; Add a label that corresponds to the defect type.

5. The method for constructing an AI code defect repair dataset according to claim 3, characterized in that: The defect types include memory errors, missing validation, or incorrect incoming data types; or The preprocessing includes replacing symbols, filtering comments or standardizing variable names.

6. The method for constructing an AI code defect repair dataset according to claim 1, characterized in that: The thought chain data includes vulnerability AI code snippets, corresponding problem descriptions, reasoning process descriptions, and sample repair codes and repair suggestions.

7. A defect repair method for AI code, characterized in that: include: Get the AI ​​code to be fixed. The AI ​​code to be defect-repaired is input into a defect-repair model of the AI ​​code; wherein the defect-repair model of the AI ​​code is obtained by fine-tuning a data set of the AI ​​code constructed based on the method for constructing a data set of AI code defect-repair according to any one of claims 1 to 6.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 7.

10. A computer program product, comprising computer program instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 7.