A code defect detection method and system based on large models

Through the large-model-based code defect detection method, combined with attention mechanism, sequence pattern recognition, graph neural network and reinforcement learning algorithm, the problem of poor code defect detection effect in the existing technology is solved, high-accurate defect recognition and high-quality repair suggestions are achieved, and code quality and maintainability are improved.

CN119690512BActive Publication Date: 2025-05-30LUSTER LIGHTWAVE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510206786.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-30
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The prior art is not effective in code defect detection, it is difficult to accurately identify and repair logical errors in complex code, and lacks effective repair suggestions and code quality evaluation.

Method used

The code defect detection method based on large models is adopted, and a pre-trained large-scale language model is used to combine attention mechanism for semantic understanding, sequence pattern recognition algorithm and graph neural network to identify logical error patterns, and repair code snippets are generated through reinforcement learning algorithms. At the same time, the code quality evaluation model is used for quality scores and change annotations generation.

Benefits of technology

It improves the accuracy of code defect detection and the effectiveness of the repair plan, reduces the need for manual intervention, improves the reliability and code quality of the repair plan, and ensures that the repaired code maintains good maintainability and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119690512B_ABST
    Figure CN119690512B_ABST
Patent Text Reader

Abstract

The present application provides a code defect detection method and system based on a large model. Among them, a pre-trained large-scale language model is used in combination with an attention mechanism to generate an enhanced semantic understanding vector corresponding to the initial source code; logical error patterns are identified to determine the defect type and location of the initial source code; the reinforcement learning algorithm selects the optimal repair path by simulating multiple repair paths and evaluating the effectiveness of each repair path, and generates a repair code snippet through the optimal repair path; the repair code snippet is used to replace the corresponding part in the initial source code, and the target source code is generated; the quality of the target source code is scored based on a code quality evaluation model to obtain a quality feedback result; the target source code with change annotations and the quality feedback result is output. The present application improves the speed and accuracy of discovering and solving coding errors in the software development process, and also ensures the quality of the repaired code.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of code detection, and in particular, to a code defect detection method and system based on a large model. Background Art

[0002] During the software development process, the detection and repair of code defects are key links to ensure software quality. With the continuous expansion of software scale and the increase in complexity, traditional manual reviews or rule-based static analysis tools have difficulty meeting the requirements of efficiently and accurately identifying and solving code problems. Especially in large projects, developers often need to handle a large amount of code, which poses higher requirements for automated defect detection. Therefore, there is a strong demand in the market for technical solutions that can automatically detect and intelligently repair code defects, especially in terms of improving detection accuracy, reducing false positive rates, and providing high-quality repair suggestions.

[0003] Currently, there are various methods and technologies for code defect detection on the market, including but not limited to rule-based static analysis tools, which check whether the code has potential problems through a series of predefined rules. And by actually running the program to observe its behavior, so as to discover runtime errors. Including machine learning-assisted analysis systems, which use historical data to train models to predict possible defect types. Although these methods have improved the ability of code defect detection to a certain extent, they each have limitations.

[0004] Rule-based methods usually rely on manually written rule sets, which makes it difficult for them to capture all types of logical errors and are prone to a high false positive rate; while dynamic analysis can locate problems more accurately, but it requires specific test environment settings and is powerless against errors on non-execution paths. Most existing tools mainly focus on the defect detection stage and rarely provide effective repair suggestions. Even if a few tools try to propose modification opinions, they are usually generated based on simple templates and lack flexibility and adaptability. Traditional methods often focus on defect detection in a single dimension (such as security or functionality), ignoring the overall quality assessment of the code, such as the impact on readability, performance, etc. In addition, existing solutions rarely consider long-term influencing factors such as the maintenance cost of the code after repair. Summary of the Invention

[0005] The embodiments of the present application provide a code defect detection method and system based on a large model to solve the problem of poor code defect detection effect in the prior art.

[0006] In a first aspect, the embodiments of the present application provide a code defect detection method based on a large model, including:

[0007] Using a pre-trained large-scale language model combined with an attention mechanism, perform semantic understanding processing on the input initial source code to generate an enhanced semantic understanding vector corresponding to the initial source code;

[0008] According to the enhanced semantic understanding vector, use a sequence pattern recognition algorithm combined with a graph neural network for analysis, identify logical error patterns, and compare the logical error patterns based on a known defect library through a fuzzy matching algorithm to determine the defect type and location of the initial source code;

[0009] Based on the defect type and location of the initial source code, use a reinforcement learning algorithm to generate a repair code snippet. The reinforcement learning algorithm selects an optimal repair path by simulating multiple repair paths and evaluating the effectiveness of each repair path, and generates a repair code snippet through the optimal repair path;

[0010] Replace the corresponding part of the initial source code with the repair code snippet and generate a target source code;

[0011] Based on a code quality evaluation model, perform a quality score on the target source code to obtain a quality feedback result, and in the target source code, generate a corresponding change annotation for the code replaced by the repair code snippet;

[0012] Output the target source code with the change annotation and the quality feedback result.

[0013] Optionally, based on the defect type and location of the initial source code, use a reinforcement learning algorithm to generate a repair code snippet. The reinforcement learning algorithm selects an optimal repair path by simulating multiple repair paths and evaluating the effectiveness of each repair path, and generates a repair code snippet through the optimal repair path, including:

[0014] Use the defect type and location of the initial source code to determine code defects;

[0015] Perform initialization processing on the reinforcement learning environment to obtain a code state including the initial source code and an environment state of the code defect;

[0016] According to the environment state, define corresponding actions and form an action set, where the actions are used to repair the code defect;

[0017] Based on the action set, use a reinforcement learning algorithm to simulate multiple repair paths, each repair path corresponding to an action sequence, and the action sequence is composed of multiple actions;

[0018] For each repair path, execute the corresponding action sequence and perform evaluation processing on the code state of the initial source code after executing the action sequence to obtain a reward value reflecting the repair effect;

[0019] Evaluate the effectiveness of each repair path according to the reward value, and select the repair path with the highest reward value as the optimal repair path;

[0020] Generate a repair code snippet based on the action sequence included in the optimal repair path.

[0021] Optionally, the generating a repair code snippet based on the action sequence included in the optimal repair path includes:

[0022] Parse each action in the action sequence of the optimal repair path to obtain the operation type and scope of each action, where the operation type includes replacement, insertion or deletion;

[0023] Generate corresponding code modification instructions according to the operation type and scope of each action;

[0024] Modify the initial source code based on the code modification instructions to obtain a repair code snippet.

[0025] Optionally, the analyzing according to the enhanced semantic understanding vector, identifying the logical error pattern by using a sequence pattern recognition algorithm combined with a graph neural network, and comparing the logical error pattern with a known defect library through a fuzzy matching algorithm to determine the defect type and location of the initial source code includes:

[0026] Use the enhanced semantic understanding vector to perform sequence pattern recognition processing on the initial source code to obtain a logical error pattern;

[0027] Construct a graph structure based on the abstract syntax tree of the initial source code, where the nodes in the graph structure represent code elements, and the edges represent the relationships between the code elements. Use a graph neural network to process the graph structure to extract the relationship information between the code elements and generate a structural feature representation;

[0028] Merge the logical error pattern with the structural feature representation to form a comprehensive feature representation;

[0029] Based on the comprehensive feature representation, use a fuzzy matching algorithm to compare with the defect information in the known defect library, and traverse the similarity scores between the comprehensive feature representation and multiple pieces of the defect information to identify the defect information that matches the logical error pattern;

[0030] Generate the defect type and location of the initial source code according to the defect information.

[0031] Optionally, based on the comprehensive feature representation, use a fuzzy matching algorithm to compare with the defect information in the known defect library, and identify the defect information that matches the logical error pattern by traversing the similarity scores between the comprehensive feature representation and multiple pieces of the defect information, including:

[0032] Use the comprehensive feature representation as an input vector and perform a fuzzy matching algorithm on each piece of defect information in the known defect library;

[0033] Calculate the similarity score between the comprehensive feature representation and each piece of defect information, where the similarity score reflects the degree of matching between the two;

[0034] Traverse the similarity scores between all pieces of defect information and the comprehensive feature representation, and select the piece of defect information with the highest score as the defect information that matches the logical error pattern;

[0035] The generating the defect type and location of the initial source code according to the defect information includes:

[0036] Take the defect type and location corresponding to the defect information as the defect type and location of the initial source code.

[0037] Optionally, use a pre-trained large-scale language model combined with an attention mechanism to perform semantic understanding processing on the input initial source code and generate an enhanced semantic understanding vector corresponding to the initial source code, including:

[0038] Use a pre-trained large-scale language model to encode the input initial source code to obtain a preliminary semantic representation;

[0039] According to the preliminary semantic representation, apply the attention mechanism to weight the key code segments to generate an enhanced semantic understanding vector.

[0040] Optionally, based on a code quality assessment model, perform a quality scoring on the target source code to obtain a quality feedback result, including:

[0041] Use a pre-trained code quality assessment model to analyze the target source code and extract various indicators reflecting the code quality;

[0042] According to the various indicators reflecting the code quality, perform a quality scoring process on the target source code to obtain a quality score;

[0043] Based on the quality score, generate a quality feedback result, and the quality feedback result includes an overall quality evaluation of the code, a specific problem description, and improvement suggestions.

[0044] Second aspect, an embodiment of the present application provides a code defect detection system based on a large model, including:

[0045] A generation module, configured to use a pre-trained large-scale language model combined with an attention mechanism to perform semantic understanding processing on the input initial source code, and generate an enhanced semantic understanding vector corresponding to the initial source code;

[0046] A determination module, configured to analyze according to the enhanced semantic understanding vector by using a sequence pattern recognition algorithm combined with a graph neural network, recognize a logical error pattern, and compare the logical error pattern based on a known defect library through a fuzzy matching algorithm to determine the defect type and location of the initial source code;

[0047] The generation module is further configured to generate a repair code snippet based on the defect type and location of the initial source code by using a reinforcement learning algorithm. The reinforcement learning algorithm simulates multiple repair paths and evaluates the effectiveness of each repair path to select an optimal repair path, and generates a repair code snippet through the optimal repair path; replace the corresponding part of the initial source code with the repair code snippet, and generate a target source code; in the target source code, generate a corresponding change annotation for the code replaced by the repair code snippet;

[0048] A calculation module, configured to perform a quality score on the target source code based on a code quality evaluation model to obtain a quality feedback result;

[0049] An output module, configured to output the target source code with the change annotation and the quality feedback result.

[0050] Third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a code defect detection method based on a large model as described in the first aspect above.

[0051] Fourth aspect, an embodiment of the present application provides a computer storage medium, storing a computer program, which when executed by a computer, implements a code defect detection method based on a large model as described in the first aspect.

[0052] In the embodiments of the present application, a pre-trained large-scale language model is used in combination with an attention mechanism to perform semantic understanding processing on the input initial source code, generating an enhanced semantic understanding vector corresponding to the initial source code; according to the enhanced semantic understanding vector, a sequence pattern recognition algorithm is used in combination with a graph neural network for analysis to identify logical error patterns, and the logical error patterns are compared with a known defect library through a fuzzy matching algorithm to determine the defect type and location of the initial source code; based on the defect type and location of the initial source code, a reinforcement learning algorithm is used to generate a repair code snippet. The reinforcement learning algorithm simulates multiple repair paths and evaluates the effectiveness of each repair path to select the optimal repair path, and generates a repair code snippet through the optimal repair path; the repair code snippet is used to replace the corresponding part in the initial source code, and a target source code is generated; based on a code quality evaluation model, a quality score is given to the target source code to obtain a quality feedback result, and in the target source code, a corresponding change annotation is generated for the code replaced by the repair code snippet; the target source code with the change annotation and the quality feedback result is output.

[0053] The technical solution of the present application has the following beneficial effects:

[0054] Through the combination of a pre-trained large-scale language model and an attention mechanism, the present application can more deeply analyze the semantic information of the source code and generate an enhanced semantic understanding vector. This helps to improve the understanding of complex programming logic, thereby more accurately identifying potential defects. And a reinforcement learning algorithm is used to generate and evaluate different repair paths and select the optimal solution, providing high-quality repair suggestions for developers. This method not only reduces the need for manual intervention but also improves the effectiveness and reliability of the repair solution. After the repair is completed, by giving a quality score to the target source code and adding a change annotation, it is ensured that the repaired code not only solves the original problem but also maintains good maintainability and performance.

[0055] Furthermore, by combining sequence pattern recognition with a graph neural network, logical error patterns in the code can be captured from different perspectives. This method not only considers the linear structure of the code but also analyzes the relationships between code elements, enhancing the detection ability for complex logical errors. And by comparing with the known defect library through a fuzzy matching algorithm, the specific defect type and its location existing in the code can be more accurately determined. This pattern matching-based method can effectively reduce the false alarm rate and improve the accuracy of the detection result. And by using an abstract syntax tree to construct a graph structure and processing it through a graph neural network, the system can better understand the internal structure of the code, thereby more accurately locating the specific location where the defect is located.

[0056] Furthermore, by traversing all possible defect information and calculating similarity scores with the comprehensive feature representation, the most matching defect record can be quickly found. This method greatly improves the matching efficiency, especially when facing a large-scale defect library. And using a fuzzy matching algorithm allows for a certain degree of imperfect matching, which is particularly useful for dealing with defects that are slightly different in expression but essentially the same or similar. This can more comprehensively cover various forms of defects and improve the coverage and accuracy of identification. Also, as newly discovered defects are continuously added to the known defect library, this method can easily adapt to these changes and continuously improve its detection ability. At the same time, by continuously optimizing the matching algorithm, the overall performance of the system can be further enhanced.

[0057] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0059] Figure 1 It is a flowchart of a code defect detection method based on a large model provided by an embodiment of the present application;

[0060] Figure 2 It is a schematic structural diagram of a code defect detection system based on a large model provided by an embodiment of the present application;

[0061] Figure 3 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] In order to enable those skilled in the art to better understand the solution of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application.

[0063] In some of the processes described in the specification, claims, and the above-mentioned drawings of this application, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The operation numbers such as 101, 102, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.

[0064] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of the present application.

[0065] Figure 1 The following is a flowchart of a code defect detection method based on a large model provided for an embodiment of the present application, as Figure 1 shown. The method includes:

[0066] 101. Use a pre-trained large language model combined with an attention mechanism to perform semantic understanding processing on the input initial source code, and generate an enhanced semantic understanding vector corresponding to the initial source code;

[0067] A large language model refers to a language processing model trained through a large amount of text data, which can understand complex semantic relationships.

[0068] The attention mechanism is a technology that enables the model to focus on specific parts of the input sequence, thereby improving processing efficiency and accuracy.

[0069] An enhanced semantic understanding vector refers to a representation generated based on the source code, which captures the deep semantic information of the code for subsequent processing.

[0070] In the field of software development, pre-trained models such as CodeBERT can be used, which are specifically optimized for programming languages. When inputting a piece of Java or Python code, the model can output a vector that contains important information such as the function of this piece of code and the relationship between variables.

[0071] Optionally, the step of "using the pre-trained large-scale language model combined with the attention mechanism to perform semantic understanding processing on the input initial source code and generate an enhanced semantic understanding vector corresponding to the initial source code" in step 101 includes: using the pre-trained large-scale language model to perform encoding processing on the input initial source code to obtain a preliminary semantic representation; according to the preliminary semantic representation, applying the attention mechanism to perform weighted processing on the key code segments to generate an enhanced semantic understanding vector.

[0072] Using a pre-trained large-scale language model combined with an attention mechanism to process source code aims to understand the semantic structure of the code through deep learning techniques. The pre-trained large-scale language model is trained based on a large amount of programming language data and can capture the syntax and semantic information in the code; the attention mechanism is a technique that allows the model to focus on specific parts when processing sequences, thereby better understanding the key parts of the code. The enhanced semantic understanding vector generated by this method is an advanced representation form of the input code, which not only reflects the basic structure of the code but also highlights the parts that have an important impact on the logic or function implementation.

[0073] First, the pre-trained large-scale language model is used to convert the original source code into a preliminary semantic representation. In this step, the model will automatically encode the vocabulary, syntax, and context information in the code. Next, the attention mechanism is used to further process this preliminary representation. In particular, it will give different weights according to the importance of each code segment, so that the finally generated enhanced semantic understanding vector can more accurately reflect the core logic and key features of the code. Throughout the process, from encoding to weighted processing, it is automated to improve the efficiency and accuracy of subsequent steps (such as error detection and repair).

[0074] In the embodiment of this application, taking a Java development project as an example, assume there is a class file containing multiple method definitions. First, use the CodeBERT model trained with a large amount of Java code as the pre-trained large-scale language model, take the entire class file as the input, and the model outputs a preliminary semantic representation of each method and its surrounding context. Then, apply the attention mechanism to these representations. For example, give a higher weight to the methods involving complex conditional judgments or loop control flows because such code is more likely to have logical errors. The finally obtained enhanced semantic understanding vector not only contains information about the overall structure of the class file but also particularly emphasizes the areas that may have potential problems. Such an output can be used to support more accurate error localization and suggestion generation.

[0075] 102. According to the enhanced semantic understanding vector, use a sequence pattern recognition algorithm combined with a graph neural network for analysis, identify logical error patterns, and compare the logical error patterns with a known defect library through a fuzzy matching algorithm to determine the defect type and location of the initial source code;

[0076] Sequence pattern recognition refers to discovering frequently occurring patterns from a series of events.

[0077] A graph neural network is a deep learning method suitable for processing graph-structured data and can be used to analyze complex structures such as program control flow graphs.

[0078] In the embodiment of the present application, for a given piece of C++ code, it is first converted into an abstract syntax tree (AST) form, and then a graph neural network is applied to analyze this AST, and at the same time, a sequence pattern recognition algorithm is combined to detect whether there are common error patterns such as infinite loops and deadlocks.

[0079] Optionally, the "According to the enhanced semantic understanding vector, use a sequence pattern recognition algorithm combined with a graph neural network for analysis, identify logical error patterns, and compare the logical error patterns with a known defect library through a fuzzy matching algorithm to determine the defect type and location of the initial source code" in step 102 includes: using the enhanced semantic understanding vector to perform sequence pattern recognition processing on the initial source code to obtain logical error patterns; constructing a graph structure based on the abstract syntax tree of the initial source code, where the nodes in the graph structure represent code elements and the edges represent the relationships between the code elements, using a graph neural network to process the graph structure to extract the relationship information between the code elements and generate a structural feature representation; merging the logical error patterns with the structural feature representation to form a comprehensive feature representation; based on the comprehensive feature representation, using a fuzzy matching algorithm to compare with the defect information in the known defect library, and by traversing the similarity scores between the comprehensive feature representation and multiple pieces of defect information, to identify the defect information that matches the logical error pattern; generating the defect type and location of the initial source code according to the defect information.

[0080] The sequence pattern recognition algorithm is a technique used to discover recurring subsequences from time series data, and here it is used to analyze the logical flow in source code to identify potential error patterns. A graph neural network (GNN) is a neural network model specifically designed to process graph-structured data, capable of effectively capturing the relationships and dependencies between nodes. An abstract syntax tree (AST) is a tree-like representation of a program's source code, where each node represents a construction unit such as a variable declaration, function call, etc. The fuzzy matching algorithm allows for the comparison of two or more elements in the case of not being completely precise, and is commonly used in text search and similarity calculation. And each defect information includes at least a defect description, defect example code, defect type, location, and associated ID. Among them, the defect description refers to the detailed textual description of each defect, including problem manifestations, scope of influence, etc. The defect example code refers to specific code snippets or patterns that show how the defect appears in the code. The defect type refers to the classification according to the nature of the defect, such as logical errors, memory leaks, security vulnerabilities, etc. The location refers to the specific location of the defect in the code (e.g., file name, line number). The associated ID refers to the unique identifier related to the defect, facilitating tracking and management. This solution aims to accurately locate and classify the logical errors existing in the source code by combining these technical means.

[0081] First, perform sequence pattern recognition processing on the initial source code using the enhanced semantic understanding vector generated in step 101 to extract possible logical error patterns. Then, construct its corresponding abstract syntax tree based on the source code and convert it into a graph structure suitable for graph neural network processing, where nodes represent code elements and edges reflect the relationships between these elements. Next, use the graph neural network to learn this graph structure and extract deep structural feature representations. Then, merge the previously identified logical error patterns with the newly generated structural feature representations to form a more comprehensive integrated feature representation. Finally, use the fuzzy matching algorithm to compare this integrated feature representation with the information in the known defect library, and determine the closest defect type and its location by calculating the similarity score, thus completing the specific location of the logical errors in the source code.

[0082] In the embodiments of the present application, it is assumed that in a Python project, there is a piece of code containing complex conditional judgments. First, use the sequence pattern recognition algorithm to analyze this piece of code and identify the patterns that may misuse conditional operators. Subsequently, build the corresponding abstract syntax tree based on this piece of code and convert it into a graph structure, where each node corresponds to an operation or variable in the code, and the edges represent the control flow or data dependency relationships. Apply the pre-trained graph neural network model to process this graph structure to obtain the high-level features reflecting the internal structural characteristics of the code. Then, combine the previously identified logical error patterns with these structural features to form a comprehensive feature vector. In the last step, use the fuzzy matching algorithm to compare this comprehensive feature vector with various known defect features stored in the database, such as common problem templates like "using the wrong operator in the conditional expression". Finally, the system will report a logical error in this piece of code caused by misusing == instead of!=, and point out the specific line number to help developers quickly locate and fix the problem.

[0083] Optionally, in step 102, "using the fuzzy matching algorithm to compare with the defect information in the known defect library based on the comprehensive feature representation, and traversing the similarity scores between the comprehensive feature representation and multiple pieces of the defect information to identify the defect information that matches the logical error pattern" includes: using the comprehensive feature representation as the input vector to perform the fuzzy matching algorithm on each piece of defect information in the known defect library; calculating the similarity score between the comprehensive feature representation and each piece of defect information, where the similarity score reflects the matching degree between the two; traversing the similarity scores between all the defect information and the comprehensive feature representation, and selecting the defect information with the highest score as the defect information that matches the logical error pattern; "generating the defect type and location of the initial source code according to the defect information" includes: using the defect type and location corresponding to the defect information as the defect type and location of the initial source code.

[0084] The fuzzy matching algorithm is a technique for comparing two or more elements in a situation where exact precision is not required. It allows for matching within a certain error range. This method is widely used in fields such as text search and pattern recognition to improve the matching ability in the presence of noise or variants. In this solution, the fuzzy matching algorithm is used to compare the comprehensive feature representation (a vector that combines the code logical error pattern and structural features) with the defect information in the known defect library, and determine the most likely defect type by calculating the similarity score. The similarity score is a quantitative indicator measuring the similarity degree between two sets of data and is used here to evaluate the matching degree between a given code snippet and known defects.

[0085] First, take the comprehensive feature representation as the input and execute the fuzzy matching algorithm for each entry in the known defect library. Then, calculate a similarity score for each pair of the comprehensive feature representation and the defect information, which directly reflects the degree of matching between the two. After that, by traversing all the obtained similarity scores, select the record with the highest score, indicating that this defect information is most likely to correspond to the logical error pattern existing in the current source code. The last step is to clearly identify the specific defect type and its location in the initial source code according to the selected defect information, so as to provide accurate information support for subsequent repair work.

[0086] In the embodiment of the present application, assume that a C++ code is being processed and a comprehensive feature representation has been obtained through the previous steps. Now, use this representation to match a database containing various common programming error types. For example, the database has entries such as "dereferencing a null pointer" and "accessing an array out of bounds". For each entry, use a fuzzy matching algorithm (such as the Levenshtein distance or the Jaccard similarity coefficient) to calculate the similarity score between it and the comprehensive feature representation. If it is found that the score of "dereferencing a null pointer" is the highest, then the system will consider that there is likely such a problem in this C++ code. Next, according to this defect information, not only will it report that this is an error of "dereferencing a null pointer", but also specifically point out which line of code this problem occurs in, such as "There may be a risk of dereferencing a null pointer on line 45". In this way, developers can quickly locate the problem and take corresponding measures to solve it.

[0087] 103. Generate a repair code snippet based on the defect type and location of the initial source code by using a reinforcement learning algorithm. The reinforcement learning algorithm selects the optimal repair path by simulating multiple repair paths and evaluating the effectiveness of each repair path, and generates a repair code snippet through the optimal repair path;

[0088] Reinforcement learning is a machine learning method that learns how to make decisions through a reward / punishment mechanism.

[0089] A repair path is a series of operations from the current state to the desired state.

[0090] In the embodiment of the present application, assume that the system detects an error of an uninitialized variable. During the attempt to repair, the reinforcement learning algorithm will explore various possible modification methods (such as direct assignment, modifying function parameters, etc.) and select the most effective solution according to the simulation running results.

[0091] Optionally, the step 103 of "generating a repair code snippet by using a reinforcement learning algorithm based on the defect type and location of the initial source code, where the reinforcement learning algorithm simulates multiple repair paths, evaluates the effectiveness of each repair path to select an optimal repair path, and generates a repair code snippet through the optimal repair path" includes:

[0092] Using the defect type and location of the initial source code to determine a code defect; initializing a reinforcement learning environment to obtain a code state including the initial source code and an environment state of the code defect; defining corresponding actions according to the environment state and forming an action set, where the actions are used to repair the code defect; based on the action set, using a reinforcement learning algorithm to simulate multiple repair paths, each repair path corresponding to an action sequence, and the action sequence being composed of multiple actions; for each repair path, executing the corresponding action sequence, and evaluating the code state of the initial source code after executing the action sequence to obtain a reward value reflecting the repair effect; according to the reward value, evaluating the effectiveness of each repair path, and selecting the repair path with the highest reward value as the optimal repair path; generating a repair code snippet based on the action sequence included in the optimal repair path.

[0093] In this solution, the environment state is a vectorized representation composed of the source code text structure, defect location, and type characteristics. The action set refers to predefined atomic repair operations (such as variable initialization, conditional branch insertion, etc.). The reward function refers to a quantitative evaluation system constructed based on the test pass rate, code quality metrics, and repair integrity to obtain a reward value.

[0094] In the embodiment of the present application, first input a defective code snippet (Java language):

[0095] public class PaymentProcessor {

[0096] public void processTransaction() {

[0097] UserAccount account;

[0098] if (account.getBalance()>0) { / / Defect: The account object is not initialized

[0099] executePayment();

[0100] }

[0101] }

[0102] }

[0103] The environmental status code is:

[0104] {

[0105] "code_state": {

[0106] "ast_hash": "a1b2c3d4",

[0107] "variables": ["account"],

[0108] "method_calls": ["getBalance()"],

[0109] "defect_type": "NP_NULL_ON_SOME_PATH",

[0110] "defect_location": [3, 15]

[0111] }

[0112] }

[0113] Secondly, generate candidate actions according to the defect type, as shown in Table 1 below:

[0114] Table 1

[0115]

[0116] Furthermore, conduct repair path simulation. Assuming three typical repair paths are generated as examples:

[0117] Path α: A1 → End (assuming there is only one action A1 in path α for simplicity, and it can be a set of multiple actions in actual applications)

[0118] UserAccount account = new UserAccount();

[0119] Path β: A2 → End (assuming there is only one action A2 in path β for simplicity, and it can be a set of multiple actions in actual applications)

[0120] if (account != null&&account.getBalance()>0);

[0121] Path γ: A3 → A2 → End

[0122] UserAccount account = getCurrentAccount();

[0123] if (account != null && account.getBalance() > 0);

[0124] Next, calculate the reward value by constructing a multi-dimensional evaluation system:

[0125] def calculate_reward(modified_code):

[0126] # Test coverage (weight 40%)

[0127] test_score = run_unit_tests(modified_code) 40

[0128] # Code quality (weight 30%)

[0129] quality_score = (1 - cyclomatic_complexity_growth) 30

[0130] # Fix completeness (weight 20%)

[0131] defect_remain = static_analysis(modified_code)

[0132] completeness = (1 - defect_remain) 20

[0133] # Side effect penalty (weight 10%)

[0134] penalty = -10 if find_new_warnings() else 0

[0135] return test_score + quality_score + completeness + penalty

[0136] The evaluation results of each path are shown in Table 2 below:

[0137] Table 2

[0138] Next, perform the optimal path selection by Q-value iteration calculation (γ = 0.9):

[0139] Q(s,γ) = 0.9 (87) + 0.1 (0.95 87) ≈ 83.2;

[0140] Q(s,α) = 75 (1 - 0.9) + 0.9 0 = 7.5;

[0141] Q(s,β) = 61 (1 - 0.9) + 0.9 (-10) = -2.9;

[0142] Step 6: Generate a repair code snippet based on the action sequence included in the optimal repair path γ, which specifically includes the following process:

[0143] Parse the action sequence of path γ:

[0144] Execution of action A3:

[0145] - UserAccount account;

[0146] + UserAccount account = getCurrentAccount();

[0147] Execution of action A2:

[0148] - if (account.getBalance()>0)

[0149] + if (account != null&&account.getBalance()>0)

[0150] Final repair code:

[0151] public class PaymentProcessor {

[0152] public void processTransaction() {

[0153] UserAccount account = getCurrentAccount();

[0154] if (account != null&&account.getBalance()>0) {

[0155] executePayment();

[0156] }

[0157] }

[0158] }

[0159] In this embodiment, by constructing a code repair framework based on reinforcement learning, the whole process from defect location to repair solution generation is fully demonstrated: First, the code defect is transformed into a computable environmental state, and the set of atomic repair actions is defined; then, multiple repair paths are simulated, and the reward value is calculated through a multi-dimensional evaluation system; then, the Q-learning algorithm is used to select the optimal repair path and generate specific code modification instructions; finally, the verification shows that this method can effectively repair the null pointer exception defect, improve the code quality, and introduce no new problems, fully proving the feasibility and effectiveness of the solution and meeting the requirements of full disclosure.

[0160] This application takes into account that in the software development process, code repair not only needs to solve known defect problems but also ensure the code quality after repair. For this reason, a reward function for comprehensively evaluating the repair effect is designed , which takes into account multiple key factors: the degree of defect repair, code readability, code performance, code complexity, and maintenance cost. In this way, it can encourage the generation of effective and high-quality repair solutions.

[0161] Optionally, "for each repair path, execute the corresponding action sequence, and evaluate the code state of the initial source code after executing the action sequence to obtain a reward value reflecting the repair effect" in step 103 includes:

[0162] For each repair path, execute the corresponding action sequence;

[0163] After executing the action sequence, based on the new code state , define a complex reward function , which is used to comprehensively evaluate the repair effect, and the reward function is determined according to the degree of defect repair, code readability, code performance, code complexity, and code maintenance cost;

[0164] Specifically, the reward function can be calculated by the following formula:

[0165] ;

[0166] where, represents the change in the degree of defect repair, and the specific expression is:

[0167] ;

[0168] where, is the number of defects before repair, is the number of defects after repair. These two values can be obtained through static analysis tools or defect detection tools;

[0169] represents the change in code readability, and the specific expression is:

[0170] ;

[0171] Among them, and are the code readability scores before and after repair respectively. These scores can be obtained through static analysis tools (such as cyclomatic complexity, number of lines of code, etc.);

[0172] and are the possible maximum and minimum readability scores respectively. These values can be set through historical data or empirical values. For example, assume (best readability), (worst readability).

[0173] represents the change in code performance, and the specific expression is:

[0174] ;

[0175] Among them, and are the code performance scores before and after repair respectively. These scores can be obtained through performance testing tools (such as execution time, memory usage, etc.);

[0176] and are the possible maximum and minimum performance scores respectively. These values can be set through historical data or empirical values. For example, assume (best performance), (worst performance).

[0177] represents the change in code complexity, and the specific expression is:

[0178] ;

[0179] Among them, and are the code complexity scores before and after repair respectively. These scores can be obtained through static analysis tools (such as control flow complexity, number of lines of code, etc.);

[0180] and They are the maximum and minimum possible complexity scores respectively, and these values can be set through historical data or empirical values. For example, assume (highest complexity), (lowest complexity).

[0181] represents the change in code maintenance cost, and the specific expression is:

[0182] ;

[0183] where, and are the code maintenance cost scores before and after repair respectively; these scores can be obtained through static analysis tools (such as comment coverage, modularity, etc.);

[0184] and are the maximum and minimum possible maintenance cost scores respectively. These values can be set through historical data or empirical values. For example, assume (highest maintenance cost), (lowest maintenance cost).

[0185] and are the weight coefficients of each factor respectively, satisfying ; and are adjustment factors used to control the sensitivity of readability and performance changes; and are penalty factors used to control the penalty intensity for the increase in code complexity and maintenance cost.

[0186] This reward function aims to balance the optimization goals in different aspects. Defect repair is the top priority, so its weight is usually high. At the same time, a good code structure (such as readability and low complexity) and efficient execution efficiency are also important considerations. In addition, considering the needs of long-term maintenance, a low maintenance cost is equally important. The overall formula adopts the form of weighted summation, allowing the importance of each factor to be adjusted according to the specific project requirements.

[0187] The following briefly introduces the design reasons for each item of this formula:

[0188] ;

[0189] where, represents the change in the degree of defect repair, and the specific expression is:

[0190] ;

[0191] Indicates the change in code readability, and the specific expression is:

[0192] ;

[0193] Indicates the change in code performance, and the specific expression is:

[0194] ;

[0195] Indicates the change in code complexity, and the specific expression is:

[0196] ;

[0197] Indicates the change in code maintenance cost, and the specific expression is:

[0198] ;

[0199] Change in the degree of defect repair: As the core goal of the repair process, the degree of defect repair directly reflects the effect of the repair operation. By calculating the change ratio of the number of defects before and after repair, the effectiveness of the repair can be quantified. A higher value means that more defects are successfully repaired, thus improving the overall quality of the code. Change in code readability: Code readability is an important indicator to measure how easy the code is to understand and maintain. By introducing the Sigmoid function , even a small improvement can get positive incentives. This encourages not only solving defects during the repair process, but also paying attention to the code structure and clarity, making the code more readable and easier to maintain. Change in code performance: Performance is a key factor in the running efficiency of software. By using the Sigmoid function , it is ensured that even a small improvement in performance can obtain rewards. This can encourage the repair solution to not only solve problems, but also optimize the execution efficiency of the code, thereby improving the user experience and system response speed. Change in code complexity: Excessive code complexity will lead to difficulties in maintenance and an increase in potential error risks. Treat it as a penalty term, and control the negative impact brought by the increase in complexity through a negative weight and an adjustment factor . This encourages the repair solution to avoid introducing complex code structures as much as possible while reducing defects, and keep the code concise and clear.

[0200] The following briefly introduces the acquisition methods of the parameters in this formula:

[0201] Among them, and : It can be automatically calculated by static analysis tools or defect detection tools (such as SonarQube, FindBugs, etc.). and Evaluate code characteristics through static analysis tools, such as metrics like cyclomatic complexity and number of lines. and Measure the performance of the program during runtime using performance testing tools (such as JMeter, LoadRunner, etc.). and Evaluate metrics related to code complexity through static analysis tools. and It is evaluated based on static analysis results, such as factors like comment coverage and modularity. The weight coefficients ( , , , ) and adjustment factors ( , , , ) are set according to project requirements and determined based on expert opinions or historical data analysis.

[0202] In the embodiments of this application, assume there is a method in a Java project. After a series of repair actions, it is desired to evaluate the effects of these repair actions. The following is a specific numerical substitution example:

[0203] Number of defects ;

[0204] Readability score ;

[0205] Performance score ;

[0206] Complexity score ;

[0207] Maintenance cost score ;

[0208] Set the weight coefficient as ;

[0209] The adjustment factor is , ;

[0210] The penalty factor is , ;

[0211] Calculate each score: , , , , ;

[0212] Substitute into the formula:

[0213] ;

[0214] Calculation result: The first item: , the second item: , the third item: , the fourth item: ;

[0215] Final score: .

[0216] According to the above calculation results, the current repair action obtained a reward value of approximately 0.5206. This indicates that while the repair action reduced defects, it also improved the readability and performance of the code. Although the code complexity and maintenance cost increased slightly, overall, a good repair effect was still achieved. In particular, the significant improvement in the degree of defect repair (from 3 to 1) contributed the most to the total score, while although other factors also had a positive effect, due to their relatively small weights, their impact on the final score was relatively limited. In this case, developers can further optimize the parts that lead to the increase in complexity and maintenance cost to achieve a better balance.

[0217] Optionally, "generating a repair code snippet based on the action sequence included in the optimal repair path" in step 103 includes: parsing each action in the action sequence in the optimal repair path to obtain the operation type and scope of action of each action, where the operation type includes replacement, insertion, or deletion; generating corresponding code modification instructions according to the operation type and scope of action of each action; and modifying the initial source code based on the code modification instructions to obtain a repair code snippet.

[0218] In the process of generating a repair code snippet based on the optimal repair path, the key concepts include the operation type and scope of action. The operation type defines the specific modification method for the source code, such as replacing existing code, inserting new code, or deleting part of the code; while the scope of action specifies the specific positions in the source code to which these modifications should be applied. By combining the operation type of each action with the scope of action, detailed code modification instructions can be formed, which directly guide how to change the original code to achieve defect repair.

[0219] First, for the optimal repair path selected from the reinforcement learning algorithm, each action in the path is parsed one by one to clarify its operation type (replacement, insertion, or deletion) and scope of action (i.e., the specific location in the source code). Then, corresponding code modification instructions are generated for each action based on the parsed information, such as "insert a new if statement after line 10" or "replace the code between lines 5 and 7 with...". Finally, the initial source code is gradually adjusted according to these specific modification instructions to finally form the repaired code snippet.

[0220] In the embodiment of this application, it is assumed that a JavaScript program is being processed, and the optimal repair path has been determined in the previous step. This path indicates that an error handling logic needs to be added within the function processData to prevent exceptions caused by null values. The first step is to parse the actions in this path and find that it requires an "insertion" operation, and the scope of action is a specific location within the function processData. Next, specific code modification instructions are generated based on this information, such as: "Within the processData function, add a line of code before data processing: if(data === null){throw new Error('Invalid data');}". The last step is to implement the above modification instructions on the original code to ensure that the correct error checking logic is added at the appropriate location, thus completing the repair process and generating the final repaired code snippet. This not only solves potential problems but also improves the overall robustness of the code.

[0221] 104. Replace the corresponding part of the initial source code with the repaired code snippet and generate the target source code;

[0222] This step involves actually applying the best repair solution generated in the above steps to the original code.

[0223] If the best repair solution is to add a line of initialization code, this line of code will be precisely inserted into the original position to correct the problem.

[0224] 105. Based on the code quality evaluation model, perform a quality score on the target source code to obtain a quality feedback result, and in the target source code, generate corresponding change comments for the code replaced by the repaired code snippet;

[0225] The code quality evaluation model is a standard tool or algorithm used to measure aspects such as code readability and complexity.

[0226] In the embodiments of the present application, tools such as SonarQube are used to evaluate the quality of the repaired code, ensure that no new problems are introduced, and the code maintains good maintainability and readability.

[0227] Optionally, "performing a quality score on the target source code based on the code quality evaluation model to obtain a quality feedback result" in step 105 includes: using a pre-trained code quality evaluation model to analyze and process the target source code, and extracting various indicators reflecting the code quality; performing a quality score processing on the target source code according to the various indicators reflecting the code quality to obtain a quality score; generating a quality feedback result based on the quality score, where the quality feedback result includes an overall quality evaluation of the code, a specific problem description, and improvement suggestions.

[0228] The code quality evaluation model is a tool built based on machine learning or deep learning technology, which can automatically analyze the source code and give an evaluation according to a series of predefined quality indicators. These indicators may include but are not limited to code complexity, readability, maintainability, and whether coding specifications are followed, etc. Through this model, the overall quality of the code can be objectively quantified, and specific improvement suggestions can be provided to help developers improve the code standards.

[0229] First, a pre-trained code quality evaluation model is used to comprehensively analyze the target source code, and key indicator data reflecting the code quality are extracted from multiple dimensions. Then, according to the specific values of these indicators, a comprehensive quality score is calculated according to a certain scoring rule or algorithm, and this score intuitively reflects the current quality level of the code. The last step is to generate a detailed feedback report, which will not only give an overall quality evaluation, but also list the specific problems and their locations found, and put forward targeted improvement suggestions, so that developers can quickly locate and solve the problems.

[0230] In an embodiment of the present application, in a Python Web application development project, it is assumed that the repair work on a piece of functional module code has been completed and quality assessment is desired. Using a code quality management platform such as SonarQube as a pre-trained code quality assessment model, the repaired code is uploaded to the system for analysis. After processing, SonarQube provides detailed data on multiple aspects such as code complexity, proportion of duplicate code, and unit test coverage. Based on this information, the system automatically calculates a comprehensive quality score, such as 85 points (out of 100), indicating that the overall code quality is good but there is still room for improvement. Subsequently, SonarQube generates a comprehensive feedback report, pointing out some potential problems such as "the function calculate_discount is too long and it is recommended to split it into smaller functions", "the handling logic for exception situations is missing", etc., and gives specific improvement suggestions, such as introducing more auxiliary functions to simplify the main logic, adding try-catch blocks to improve program stability, etc. In this way, developers can further optimize the code according to these suggestions, thereby continuously improving the overall quality of the software project.

[0231] 106. Output the target source code with the change annotation and the quality feedback result.

[0232] The change annotation refers to the information recording where changes are made and the reasons.

[0233] The quality feedback result refers to the specific evaluation of the code quality provided.

[0234] In an embodiment of the present application, the final output includes not only the updated code file but also detailed documentation, such as information like "fixed the null pointer exception on line 35", "the code complexity decreased by 10%", etc., to help developers quickly understand the changed content and its impact.

[0235] Figure 2 The following is a schematic structural diagram of a code defect detection system based on a large model provided by an embodiment of the present application, as Figure 2 shown. The device includes:

[0236] A generation module 21, configured to use a pre-trained large-scale language model combined with an attention mechanism to perform semantic understanding processing on the input initial source code, and generate an enhanced semantic understanding vector corresponding to the initial source code;

[0237] A determination module 22, configured to analyze according to the enhanced semantic understanding vector, using a sequence pattern recognition algorithm combined with a graph neural network, identify logical error patterns, and compare the logical error patterns with a known defect library through a fuzzy matching algorithm to determine the defect type and location of the initial source code;

[0238] The generation module is further configured to generate a repair code snippet based on the defect type and location of the initial source code by using a reinforcement learning algorithm. The reinforcement learning algorithm simulates multiple repair paths, evaluates the effectiveness of each repair path to select the optimal repair path, and generates a repair code snippet through the optimal repair path; replace the corresponding part of the initial source code with the repair code snippet, and generate a target source code; for the code replaced by the repair code snippet in the target source code, generate a corresponding change annotation.

[0239] The calculation module 23 is configured to perform a quality score on the target source code based on a code quality evaluation model to obtain a quality feedback result.

[0240] The output module 24 is configured to output the target source code with the change annotation and the quality feedback result.

[0241] Figure 2 The described code defect detection system based on a large model can execute Figure 1 The described code defect detection method based on a large model in the illustrated embodiment, its implementation principle and technical effects will not be elaborated. For the code defect detection system based on a large model in the above embodiment, the specific ways for each module and unit to perform operations have been described in detail in the embodiment related to the method, and will not be elaborated here.

[0242] In a possible design, Figure 2 The code defect detection system based on a large model in the illustrated embodiment can be implemented as a computing device, such as Figure 3 shown, the computing device may include a storage component 31 and a processing component 32;

[0243] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32.

[0244] The processing component 32 is configured to: perform semantic understanding processing on the input initial source code by using a pre-trained large-scale language model combined with an attention mechanism to generate an enhanced semantic understanding vector corresponding to the initial source code; analyze the enhanced semantic understanding vector by using a sequence pattern recognition algorithm combined with a graph neural network to identify logical error patterns, and compare the logical error patterns with a known defect library through a fuzzy matching algorithm to determine the defect type and location of the initial source code; generate a repair code snippet based on the defect type and location of the initial source code by using a reinforcement learning algorithm, where the reinforcement learning algorithm simulates multiple repair paths and evaluates the effectiveness of each repair path to select an optimal repair path, and generates a repair code snippet through the optimal repair path; replace the corresponding part of the initial source code with the repair code snippet and generate a target source code; perform a quality score on the target source code based on a code quality evaluation model to obtain a quality feedback result, and generate a corresponding change annotation for the code replaced by the repair code snippet in the target source code; and output the target source code with the change annotation and the quality feedback result.

[0245] Among them, the processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.

[0246] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc.

[0247] Of course, the computing device may also necessarily include other components, such as an input / output interface, a display component, a communication component, etc.

[0248] The input / output interface provides an interface between the processing component and a peripheral interface module, and the above peripheral interface module may be an output device, an input device, etc.

[0249] The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.

[0250] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server. The above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from a cloud computing platform.

[0251] An embodiment of the present application also provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above-mentioned Figure 1 A method for detecting code defects based on a large model shown in the embodiment.

[0252] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0253] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0254] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0255] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A code defect detection method based on a large model, characterized in that: include: Using a pre-trained large-scale language model combined with an attention mechanism, a semantic understanding process is performed on the input initial source code to generate an enhanced semantic understanding vector corresponding to the initial source code; According to the enhanced semantic understanding vector, a sequential pattern recognition algorithm combined with a graph neural network is used for analysis to identify a logic error pattern, and the logic error pattern is compared with a fuzzy matching algorithm based on a known defect library to determine the defect type and location of the initial source code; Based on the defect type and location of the initial source code, a reinforcement learning algorithm is used to generate a repair code snippet, wherein the reinforcement learning algorithm selects an optimal repair path by simulating multiple repair paths and evaluating the effectiveness of each repair path, and generates a repair code snippet through the optimal repair path; Replacing the corresponding part of the initial source code with the repair code fragment and generating a target source code; Performing a quality score on the target source code based on a code quality assessment model to obtain a quality feedback result, and generating a corresponding change annotation in the target source code for the code replaced by the repair code snippet; Outputting a target source code with the change annotation and the quality feedback result; The method of analyzing the enhanced semantic understanding vector using a sequential pattern recognition algorithm combined with a graph neural network to identify a logic error pattern, and comparing the logic error pattern using a fuzzy matching algorithm based on a known defect library to determine the defect type and location of the initial source code includes: Using the enhanced semantic understanding vector, performing sequence pattern recognition processing on the initial source code to obtain a logic error pattern; Building a graph structure based on the abstract syntax tree of the initial source code, wherein nodes in the graph structure represent code elements, and edges represent relationships between the code elements, and processing the graph structure using a graph neural network to extract relationship information between the code elements and generate a structural feature representation; combining the logical error pattern with the structural feature representation to form a comprehensive feature representation; Based on the comprehensive feature representation, a fuzzy matching algorithm is used to compare the defect information in the known defect library, and the defect information matching the logic error pattern is identified by traversing the similarity scores between the comprehensive feature representation and the plurality of defect information; The defect type and location of the initial source code are generated according to the defect information.

2. The method according to claim 1, characterized in that The method of generating a repair code snippet based on the defect type and location of the initial source code by using a reinforcement learning algorithm, wherein the reinforcement learning algorithm selects an optimal repair path by simulating multiple repair paths and evaluating the effectiveness of each repair path, and generates a repair code snippet through the optimal repair path, including: Determine code defects using defect types and locations of the initial source code; Initializing the reinforcement learning environment to obtain a code state including the initial source code and an environment state of the code defect; According to the environment state, corresponding actions are defined and constituted into an action set, wherein the actions are used to repair the code defect; Based on the action set, a reinforcement learning algorithm is used to simulate multiple repair paths, each repair path corresponds to an action sequence, and the action sequence consists of multiple actions; For each repair path, a corresponding action sequence is executed, and a code state of the initial source code after the action sequence is executed is evaluated to obtain a reward value reflecting the repair effect; According to the reward value, evaluating the effectiveness of each repair path, and selecting the repair path with the highest reward value as the optimal repair path; A repair code snippet is generated based on the action sequence included in the optimal repair path.

3. The method according to claim 2, characterized in that The generating of the repair code snippet based on the action sequence included in the optimal repair path includes: Parsing each action in the action sequence in the optimal repair path to obtain an operation type and an action scope of each action, wherein the operation type includes replacement, insertion or deletion; Generate corresponding code modification instructions based on the operation type and scope of each action; Based on the code modification instruction, the initial source code is modified to obtain a repair code fragment.

4. The method according to claim 1, characterized in that: The method of comparing the comprehensive feature representation with the defect information in the known defect library by using a fuzzy matching algorithm, and identifying the defect information matching the logic error pattern by traversing the similarity scores between the comprehensive feature representation and a plurality of the defect information, includes: Using the comprehensive feature representation as an input vector, a fuzzy matching algorithm is performed on each defect information in the known defect library; Calculating a similarity score between the comprehensive feature representation and each defect information, wherein the similarity score reflects the degree of matching between the two; Traversing the similarity scores between all defect information and the comprehensive feature representation, and selecting the defect information with the highest score as the defect information matching the logical error pattern; The step of generating the defect type and location of the initial source code according to the defect information includes: The defect type and position corresponding to the defect information are used as the defect type and position of the initial source code.

5. The method according to claim 1, characterized in that The method uses a pre-trained large-scale language model in combination with an attention mechanism to perform semantic understanding processing on the input initial source code to generate an enhanced semantic understanding vector corresponding to the initial source code, including: Using a pre-trained large-scale language model, the initial source code is encoded to obtain a preliminary semantic representation. Based on the preliminary semantic representation, an attention mechanism is applied to perform weighted processing on key code snippets to generate an enhanced semantic understanding vector.

6. The method according to claim 1, characterized in that The quality scoring of the target source code based on the code quality assessment model to obtain a quality feedback result includes: Analyze and process the target source code using a pre-trained code quality assessment model to extract various indicators reflecting code quality; According to various indicators reflecting code quality, the target source code is scored to obtain a quality score; Based on the quality score, a quality feedback result is generated, which includes an overall quality evaluation of the code, a specific problem description, and improvement suggestions.

7. A code defect detection system based on a large model, characterized in that: include: A generation module, used to use a pre-trained large-scale language model combined with an attention mechanism to perform semantic understanding processing on the input initial source code, and generate an enhanced semantic understanding vector corresponding to the initial source code; A determination module is used to identify the logic error pattern by using a sequence pattern recognition algorithm combined with a graph neural network to analyze the enhanced semantic understanding vector, and compare the logic error pattern with a fuzzy matching algorithm based on a known defect library to determine the defect type and location of the initial source code; The generation module is further used to generate a repair code snippet based on the defect type and location of the initial source code by using a reinforcement learning algorithm, wherein the reinforcement learning algorithm selects an optimal repair path by simulating multiple repair paths and evaluating the effectiveness of each repair path, and generates a repair code snippet through the optimal repair path; replace the corresponding part of the initial source code with the repair code snippet, and generate a target source code; In the target source code, generating corresponding change annotations for the code replaced by the repair code snippet; A calculation module, used to perform a quality score on the target source code based on a code quality assessment model to obtain a quality feedback result; An output module, used for outputting a target source code with the change annotation and the quality feedback result; The method of analyzing the enhanced semantic understanding vector using a sequential pattern recognition algorithm combined with a graph neural network to identify a logic error pattern, and comparing the logic error pattern using a fuzzy matching algorithm based on a known defect library to determine the defect type and location of the initial source code includes: Using the enhanced semantic understanding vector, performing sequence pattern recognition processing on the initial source code to obtain a logic error pattern; Building a graph structure based on the abstract syntax tree of the initial source code, wherein nodes in the graph structure represent code elements, and edges represent relationships between the code elements, and processing the graph structure using a graph neural network to extract relationship information between the code elements and generate a structural feature representation; combining the logical error pattern with the structural feature representation to form a comprehensive feature representation; Based on the comprehensive feature representation, a fuzzy matching algorithm is used to compare the defect information in the known defect library, and the defect information matching the logic error pattern is identified by traversing the similarity scores between the comprehensive feature representation and the plurality of defect information; The defect type and location of the initial source code are generated according to the defect information.

8. A computing device, characterized in that It comprises a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a large model-based code defect detection method as described in any one of claims 1 to 6.

9. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, a code defect detection method based on a large model as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Source code vulnerability detection method based on sequence and graph two-channel model

    CN118296612A

  • Code vulnerability detection method based on efficient parameter fine tuning of code pre-training model

    CN119441006A