Intelligent code review method and system based on large language model
Through the intelligent code review method based on large language model, the problems of low efficiency and poor flexibility of traditional code review are solved, and more efficient and accurate code review is achieved to adapt to the software development needs of rapid iteration.
Patent Information
- Application Number
- CN202510651656.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional code review method relies on manual labor, is inefficient, is prone to miss problems, is subjective, and is difficult to meet the software development needs of rapid iteration. The existing automation tools are poor in flexibility and cannot handle complex semantic and logical problems.
An intelligent code review method based on a large language model is designed. By obtaining the code to be reviewed for lexical, grammatical and semantic analysis, code features are extracted, and features are input into the large language model to generate code review results.
It improves the accuracy of code review, reduces omissions and errors of manual review, significantly improves the efficiency of code review, shortens the software development cycle, reduces development costs, and provides more comprehensive and targeted code review suggestions.
Smart Images

Figure CN120162239A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to an intelligent code review method and system based on a large language model. Background Art
[0002] With the continuous increase in the scale and complexity of software development, code review, as an important link to ensure code quality and improve software reliability, faces huge challenges. The traditional code review method mainly relies on manual review, which has defects such as low efficiency, easy omission of problems, and strong subjectivity, and is difficult to meet the requirements of rapidly iterative software development. Although there are currently some automated code review tools based on rules and simple algorithms, these tools have poor flexibility, cannot handle complex semantic and logical problems, and are difficult to adapt to diverse code scenarios. Summary of the Invention
[0003] The purpose of the present invention is to solve the above problems, and an intelligent code review method and system based on a large language model are designed.
[0004] In the first aspect of the present invention, an intelligent code review method based on a large language model is provided. The method includes the following steps: Obtain the code to be reviewed, perform lexical analysis, syntactic analysis, and semantic analysis on the code to be reviewed, extract the key information of the code, and obtain the preprocessed code; Extract various features of the code from the preprocessed code to obtain the target code features; Input the target code features into the large language model, and generate a code review result through the large language model.
[0005] Optionally, in the first implementation manner of the first aspect of the present invention, the obtaining the code to be reviewed, performing lexical analysis, syntactic analysis, and semantic analysis on the code to be reviewed, extracting the key information of the code, and obtaining the preprocessed code includes: Obtain the code to be reviewed, perform lexical analysis on the code to be reviewed, and split the code to be reviewed into individual words; According to the word sequence obtained by lexical analysis, construct an abstract syntax tree of the code, use a dynamic syntax tree encoder to encode the preprocessed abstract syntax tree, perform semantic checking on the constructed abstract syntax tree, extract the key information of the code, and obtain the preprocessed code.
[0006] Optionally, in the second implementation of the first aspect of the present invention, the dynamic syntax tree encoder adopts a two-stream architecture. The GGNN layer captures the topological structure features of the abstract syntax tree, and the Transformer layer identifies the semantic associations across nodes in the abstract syntax tree through the attention mechanism. A gating mechanism is used to dynamically fuse the two representations obtained from the GGNN layer and the Transformer layer, and a syntax tree pruning attention mechanism is introduced to filter redundant branches in the abstract syntax tree.
[0007] Optionally, in the third implementation of the first aspect of the present invention, extracting multiple features of the code from the preprocessed code to obtain target code features includes: Constructing a control flow graph based on the preprocessed code to present the execution logic of the code in a graphical form, where each node represents a basic code block and the edge represents the jump relationship of code execution; Processing the control flow graph using a fractal convolutional network to capture multi-scale structural patterns, and inputting the result of the fractal convolutional processing into a spectral clustering algorithm to extract the modular features of the code; Based on the extracted modular features, calculating the structural entropy to obtain a measurement index of the code structure features, decoupling and analyzing the code style, and combining the measurement index of the code structure features to obtain the target code features.
[0008] Optionally, in the fourth implementation of the first aspect of the present invention, the decoupling and analysis of the code style includes: Analyzing the layout elements in the code including at least indentation and spaces to extract the layout style features; Performing word vector clustering on the identifiers in the code including at least variable names and function names to extract the naming style features; Identifying the design patterns adopted in the code, generating design pattern fingerprints, and extracting the pattern style features; Integrating the layout style features, naming style features, and pattern style features of the code to obtain style elements; Through a contrastive disentanglement learning method, separating the style elements from the functional semantics of the code to obtain the code style features.
[0009] Optionally, in the fifth implementation of the first aspect of the present invention, inputting the target code features into a large language model to generate a code review result includes: Inputting the target code features into a large language model, and performing structural review processing, logical review processing, and style review processing on the target code features through the federated architecture in the large language model, and outputting a mixed review result; Inputting the target code features into the CodeBERT layer for screening, analyzing the code features, and identifying errors and potential problems to obtain a feature review result; Integrate the mixed review results and the feature review results to comprehensively obtain the code review results.
[0010] Optionally, in the sixth implementation manner of the first aspect of the present invention, when inputting the target code features into the large language model, the structure review process, logic review process, and style review process are performed on the target code features through the federated architecture in the large language model, and the mixed review results are output, including: Process the code structure features through the anomaly detector of the Graph Transformer network in the large language model, analyze the structural pattern of the code, detect whether there are structural anomalies, and output a preliminary judgment on possible errors and potential problems in terms of code structure. Use the symbolic execution path explorer in the large language model to perform symbolic execution exploration on the execution path of the code, simulate the execution of the code under different conditions, and identify possible problems in the code logic. Compare the style features of the code with the preset good style standards through the contrastive learning style transfer discriminator in the large language model to determine whether the layout style, naming style, and pattern style of the code meet the specifications.
[0011] The second aspect of the present invention provides an intelligent code review system based on a large language model, and the system includes: An acquisition module, configured to acquire the code to be reviewed, perform lexical analysis, syntax analysis, and semantic analysis on the code to be reviewed, extract the key information of the code, and obtain the preprocessed code. An extraction module, configured to extract various features of the code from the preprocessed code to obtain the target code features. A generation module, configured to input the target code features into the large language model and generate code review results through the large language model.
[0012] Optionally, in the first implementation manner of the second aspect of the present invention, the extraction module includes: A construction sub-module, configured to construct a control flow graph based on the preprocessed code, present the execution logic of the code in a graphical form, where each node represents a basic code block and the edge represents the jump relationship of code execution. A capture sub-module, configured to process the control flow graph using a fractal convolutional network, capture multi-scale structural patterns, input the result after fractal convolution processing into a spectral clustering algorithm, and extract the modular features of the code. A calculation sub-module, configured to calculate the structural entropy based on the extracted modular features, obtain a measurement index of the code structure features, decouple and analyze the code style, and combine the measurement index of the code structure features to obtain the target code features.
[0013] Optionally, in the second implementation manner of the second aspect of the present invention, the generation module includes: An input sub-module, configured to input the target code features into a large language model, and perform structural review processing, logical review processing, and style review processing on the target code features through the federated architecture in the large language model, and output a mixed review result; A recognition sub-module, configured to input the target code features into the CodeBERT layer for screening, analyze the code features, identify errors and potential problems, and obtain a feature review result; An integration sub-module, configured to integrate the mixed review result and the feature review result to comprehensively obtain a code review result.
[0014] In the technical solution provided by the present invention, the code to be reviewed is obtained, lexical analysis, syntax analysis, and semantic analysis are performed on the code to be reviewed, key information of the code is extracted to obtain the preprocessed code; various features of the code are extracted from the preprocessed code to obtain the target code features; the target code features are input into a large language model, and a code review result is generated through the large language model; the present invention combines the large language model to be able to handle complex code semantics and logical problems, improve the accuracy of code review, reduce omissions and errors in manual review, and significantly improve the efficiency of code review through an automated code review process, shorten the software development cycle, reduce development costs, and be able to provide more comprehensive and targeted code review suggestions, which helps developers improve code quality and programming level. Description of the Drawings
[0015] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention.
[0016] Figure 1 It is a flowchart of an intelligent code review method based on a large language model provided by an embodiment of the present invention; Figure 2 It is a schematic structural diagram of an intelligent code review system based on a large language model provided by an embodiment of the present invention. Detailed Embodiments
[0017] In the description, claims, and above-mentioned drawings of the present invention, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the term "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or equipment.
[0018] For ease of understanding, the specific process of the embodiments of the present invention will be described below. Please refer to Figure 1 The flowchart of the intelligent code review method based on the large language model provided by the embodiments of the present invention specifically includes the following steps: Step 101: Obtain the code to be reviewed, perform lexical analysis, syntax analysis, and semantic analysis on the code to be reviewed, extract the key information of the code, and obtain the preprocessed code; In this embodiment, the code to be reviewed is obtained, lexical analysis is performed on the code to be reviewed, and the code to be reviewed is segmented into individual words; according to the word sequence obtained by lexical analysis, an abstract syntax tree of the code is constructed, the abstract syntax tree obtained by preprocessing is encoded using a dynamic syntax tree encoder, semantic checking is performed on the constructed abstract syntax tree, the key information of the code is extracted, and the preprocessed code is obtained.
[0019] In this embodiment, the dynamic syntax tree encoder adopts a two-stream architecture. The GGNN layer captures the topological structure features of the abstract syntax tree, the Transformer layer identifies the semantic associations across nodes in the abstract syntax tree through the attention mechanism, a gating mechanism is used to dynamically fuse the two representations obtained by the GGNN layer and the Transformer layer, and a syntax tree pruning attention mechanism is introduced to filter out the redundant branches in the abstract syntax tree.
[0020] In this embodiment, the code to be reviewed is obtained: The code to be reviewed is obtained from a specified code storage location, such as a code repository, a local file system, etc.; A lexical analyzer is used to scan the obtained code, splitting the code into individual words or symbols, such as identifiers, keywords, operators, etc., to prepare for subsequent syntax analysis; According to the word sequence obtained from lexical analysis, an abstract syntax tree (AST) of the code is constructed based on the syntax rules of the programming language; This process will identify the statement structure, expression hierarchy, etc. in the code, converting the text form of the code into a tree structure for subsequent semantic analysis and feature extraction; Semantic checks are performed on the constructed abstract syntax tree, including type checking, scope analysis, etc., to ensure that the code is semantically legal; At the same time, key information in the code, such as variable declarations, function calls, class definitions, etc., is extracted to provide a basis for subsequent feature extraction and review; The abstract syntax tree obtained from preprocessing is encoded using a dynamic syntax tree encoder; Here, a two-stream architecture is adopted, and Gated Graph Neural Networks (GGNN) and Transformer are used for processing respectively.GGNN Encoding: GGNN is responsible for capturing the topological structure features of the abstract syntax tree; it iteratively updates the nodes and edges in the abstract syntax tree, and through the message passing mechanism, spreads the feature information of the nodes in the graph, so as to learn the topological relationships and structural features between nodes; Transformer Encoding: The Transformer layer identifies the semantic associations across nodes in the abstract syntax tree through the attention mechanism; it processes the features of each node, calculates the attention weights between nodes, so as to capture the semantic dependencies between different nodes; Use a gating mechanism to dynamically fuse the two representations obtained from the GGNN layer and the Transformer layer; The gating mechanism will adaptively adjust the weights of the two representations according to different task requirements and data characteristics, and fuse them into a unified code representation for subsequent review and analysis; Introduce the Pruned AST Attention mechanism, which automatically filters redundant branches in the abstract syntax tree based on the importance of node types; When calculating the attention weights, the importance of nodes will be evaluated, and those nodes and branches that have little impact on the overall semantics and review results will be ignored, so as to improve the model's ability to capture long-distance dependencies; Construct an Extended Symbolic Execution Graph (ESEG): Based on LLVM IR (Low Level Virtual Machine Intermediate Representation), construct an Extended Symbolic Execution Graph; This graph will integrate multiple key components, including a data flow-sensitive state transition matrix, a Z3 expression tree for path constraints, and a Petri net model for memory alias analysis; Analyze the data flow in the code, record the state changes of variables on different execution paths, and construct a state transition matrix to track the execution process and data dependencies of the code; Analyze the conditional statements and branches in the code, and convert the path constraints into a Z3 expression tree for subsequent symbolic execution and constraint solving; Petri Net Model for Memory Alias Analysis: Use the Petri net model to analyze the memory aliases in the code, identify the memory sharing relationships between different variables, and provide support for symbolic execution and memory management; Use the Dafny verifier to analyze the code and generate formal properties as an additional annotation layer; These formal properties can describe the properties of the code in terms of correctness, security, etc., and provide additional information for subsequent review and verification.
[0021] Step 102: Extract various features of the code from the preprocessed code to obtain the target code features; In this embodiment, a control flow graph is constructed based on the preprocessed code, presenting the execution logic of the code in a graphical form, where each node represents a basic code block and the edges represent the jump relationships of code execution. The control flow graph is processed using a fractal convolutional network to capture multi-scale structural patterns, and the result of the fractal convolution processing is input into a spectral clustering algorithm to extract the modular features of the code. Based on the extracted modular features, the structural entropy is calculated to obtain a measurement index of the code structure features, decoupling and analyzing the code style, and combining the measurement index of the code structure features to obtain the target code features.
[0022] In this embodiment, the typesetting elements including at least indentation and spaces in the code are analyzed to extract the layout style features; the identifiers including at least variable names and function names in the code are subjected to word vector clustering to extract the naming style features; the design patterns adopted in the code are identified to generate design pattern fingerprints and extract the pattern style features; the layout style features, naming style features and pattern style features of the code are integrated to obtain style elements; and through a contrastive disentanglement learning method, the style elements are separated from the functional semantics of the code to obtain the code style features.
[0023] In this embodiment, based on the preprocessed code, a control flow graph (CFG) is constructed through a compiler or a dedicated analysis tool to present the execution logic of the code in a graphical form. Each node represents a basic code block, and the edges represent the jump relationships of code execution. The control flow graph is processed using a fractal convolutional network (Fractal CNN). The fractal convolutional network adopts self-similar convolutional kernels to scan the control flow graph at different scales. The small-scale convolutional kernels focus on the local code structure, and the large-scale convolutional kernels capture the overall structural patterns, thereby achieving the capture of multi-scale structural patterns. The result of the fractal convolution processing is input into the spectral clustering algorithm. The spectral clustering algorithm divides the graph into different modules according to the similarity between the nodes of the control flow graph, and extracts the modular features of the code. These features reflect the division and organization methods of the code functions. Based on the extracted modular features, a structural entropy index is defined. By calculating the complexity of each module and the degree of association between modules, the overall structural complexity of the code is quantified, and an important measurement index of the code structure features is obtained. The preprocessed code is dynamically analyzed to simulate the execution process of the code and generate a series of execution path sequences. These sequences record the execution trajectories of the code under different inputs and conditions. A temporal convolutional network (TCN) is used to analyze the execution path sequences. The TCN mines the time-dependent relationships in the sequences through convolutional operations to capture the changing patterns and rules of the execution paths over time. On the basis of the TCN analysis, a causal attention mechanism is introduced. This mechanism focuses on the causally related parts of the execution path sequences, identifies potential logical problems such as deadlocks and race conditions, and highlights the areas that may have logical defects by paying attention to the critical paths and events. For complex logical paths, a quantum circuit simulator is introduced. The different states and branches in the logical paths are encoded as the probability amplitudes of quantum states, and by simulating the behavior under the quantum superposition state, a more in-depth analysis of the complex logic is carried out to detect abnormal behaviors that are difficult to discover by traditional methods, thereby extracting the complete code logical features.
[0024] In this embodiment, traverse the preprocessed code text to identify typesetting elements such as indentation and spaces therein; indentation usually manifests as the number of spaces or tabs at the beginning of each line of code, while spaces are distributed in various parts of the code, such as on both sides of operators, at comma separations, etc.; regular expressions or simple character matching methods can be used to locate and count the occurrence positions and quantities of these typesetting elements; for indentation, count the indentation levels of each line of code, for example, represented by the number of spaces or tabs; statistical quantities such as the average indentation level and the distribution range of indentation levels can be calculated; for spaces, analyze their usage frequencies and position rules in different code structures (such as function definitions, conditional statements, loop statements, etc.); for example, count the usage of spaces on both sides of operators to see if consistent rules are followed; use the above quantified statistical information as layout style features; for example, use the average indentation level, the standard deviation of the indentation level distribution, the usage frequencies of spaces in different code structures, etc. as elements of the feature vector to form a layout style feature representation; extract identifiers such as variable names and function names from the preprocessed code; through syntax analysis, locate these identifier nodes in the abstract syntax tree (AST) and obtain their text contents; use a pre-trained word vector model (such as Word2Vec, GloVe, etc.) to convert each identifier into a corresponding word vector; these word vectors map identifiers into a high-dimensional vector space, making identifiers with similar semantics closer in the vector space; apply a clustering algorithm (such as K-Means, DBSCAN, etc.) to cluster the extracted word vectors; the purpose of clustering is to group identifiers with similar naming styles into one category; for example, identifiers named in camel case may be clustered into one category, and identifiers named in underscore case may be clustered into another category; analyze the features of each cluster, such as the center vector of the cluster, the size of the cluster, the common prefix or suffix of the identifiers within the cluster, etc.; use these features as naming style features to describe the naming style of the code; adopt a pattern matching algorithm or a machine learning method to analyze the preprocessed code to identify the design patterns adopted therein; common design patterns include the singleton pattern, the factory pattern, the observer pattern, etc.; predefined pattern templates or rule sets can be used for matching, or a classifier can be trained to identify different design patterns; generate a unique fingerprint for each identified design pattern; the fingerprint can be the name of the pattern, the hash value of the key structural features of the pattern, etc.; for example, for the singleton pattern, the way of creating an object and the uniqueness check of the object can be used as key features to generate a fingerprint; count the occurrence frequencies and combination methods of different design pattern fingerprints; use these statistical information as pattern style features to reflect the characteristics of the code in architectural design and programming paradigms; combine the layout style features, naming style features, and pattern style features extracted in the previous steps into a unified feature vector;These features can be arranged in a certain order. For example, place the layout style feature first, then the naming style feature, and finally the mode style feature; perform normalization on the merged feature vector. For example, use the Z-score normalization method to convert the value of each feature into a standard normal distribution with a mean of 0 and a standard deviation of 1; this helps to eliminate the dimensionality differences between different features and improve the accuracy of subsequent analysis; construct a comparison dataset that includes code pairs with different styles but the same function, as well as code pairs with the same style but different functions; these code pairs can be obtained by modifying or generating existing codes; use the comparison dataset to train a disentangled learning model, such as a model based on variational autoencoder (VAE) or generative adversarial network (GAN); the goal of the model is to learn how to separate the style elements and functional semantics of the code; during the training process, the model will learn the different representations of style elements and functional semantics in the feature space, and gradually achieve their separation by comparing the feature differences of different code pairs; input the integrated style elements into the trained disentangled learning model, and the model will output the separated code style features; these features only reflect the style information of the code and do not contain the interference of functional semantics.
[0025] Step 103: Input the target code features into the large language model to generate a code review result through the large language model.
[0026] In this embodiment, input the target code features into the large language model, and perform structural review processing, logical review processing, and style review processing on the target code features through the federated architecture in the large language model to output a mixed review result; input the target code features into the CodeBERT layer for screening, analyze the code features, identify errors and potential problems, and obtain a feature review result; integrate the mixed review result and the feature review result to comprehensively obtain a code review result.
[0027] In this embodiment, process the code structure features through the anomaly detector of the Graph Transformer network in the large language model, analyze the structural pattern of the code, detect whether there are structural anomalies, and output a preliminary judgment on possible errors and potential problems in terms of code structure; perform symbolic execution exploration on the execution path of the code through the symbolic execution path explorer in the large language model, simulate the execution of the code under different conditions, and identify possible problems in the code logic; compare the style features of the code with the preset good style standards through the contrastive learning style transfer discriminator in the large language model to determine whether the layout style, naming style, and mode style of the code meet the specifications.
[0028] In this embodiment, the code structure features in the target code features are input into an anomaly detector (structural review expert) based on the Graph Transformer network; the Graph Transformer network processes the code structure features, analyzes the structural patterns of the code, and detects whether there are structural anomalies, such as overly deep loop nesting, chaotic function call relationships, etc.; outputs a preliminary judgment on possible errors and potential problems in terms of the code structure; provides the code logic features in the target code features to a reinforcement learning-driven symbolic execution path explorer (logical review expert); the explorer uses the method of reinforcement learning to perform symbolic execution exploration on the execution path of the code, simulating the execution of the code under different conditions; identifies problems such as deadlocks, race conditions, and logical contradictions that may exist in the code logic, and outputs corresponding review information; allows the code style features in the target code features to pass through a contrastive learning style transfer discriminator (style review expert); the discriminator, based on the contrastive learning method, compares the code style features with a preset good style standard to determine whether the layout style, naming style, and pattern style of the code meet the specifications; generates suggestions regarding the code style, such as whether the indentation is standard, whether the naming is clear, and whether the design pattern is used appropriately; inputs the target code features into the distilled CodeBERT model for a quick preliminary screening; the distilled CodeBERT uses its pre-trained language understanding ability to quickly analyze the code features, initially identifying obvious errors and potential problems; outputs a preliminary review result, marking the areas that may have problems; for the parts marked as requiring further review in the intuitive channel, input their code features into a deep verification module based on the Lean theorem prover; the Lean theorem prover deeply verifies the logic and properties of the code, and proves whether the code meets specific specifications and requirements through a formal method; outputs the review result after deep verification, clearly pointing out the errors and potential problems in the code; encodes the property specifications of the code into a differentiable logic layer, and uses a fuzzy type system and probabilistic Hoare logic to represent the correctness and security requirements of the code; during the forward propagation process of the large language model, fuse the differentiable logic layer with the calculations of the neural network; enabling the model to perform formal verification while performing statistical learning to further confirm whether the code meets the property specifications; integrate the results of the differentiable formal verification with the review results of the multi-expert hybrid model and the dual-process decision-making mechanism; comprehensively obtain the code review result, including the errors, potential problems, and code style suggestions in the code.
[0029] In this embodiment, the code structure features extracted previously are input into the anomaly detector of the Graph Transformer network. These code structure features may be obtained by analyzing the abstract syntax tree (AST), control flow graph (CFG), etc. of the code, and contain the structure information of the code, such as node relationships, hierarchical structures, module divisions, etc.; The Graph Transformer network will first perform an embedding operation on each node in the code structure, converting the feature information of the node into a low-dimensional vector representation. In this way, the information of the code structure can be mapped into a continuous vector space for subsequent processing. Using the multi-head attention mechanism, the model calculates the attention weights between nodes to capture the dependency relationships and interaction information between nodes. Through different attention heads, the model can focus on different parts of the code structure from multiple perspectives, thus understanding the structural patterns of the code more comprehensively. After being processed by the multi-head attention mechanism, the features of the nodes will be propagated and updated between the layers of the Graph Transformer network. Each layer will further transform and fuse the features of the nodes, enabling the model to learn more advanced structural features. During the training phase, the anomaly detector of the Graph Transformer network will learn the structural patterns of a large number of normal codes and establish a model of the normal structure. During actual detection, the model will compare the input code structure features with the learned normal patterns. If it is found that some structural features deviate significantly from the normal patterns, they will be marked as parts that may have anomalies. According to the results of the anomaly detection, a preliminary judgment on the possible errors and potential problems in the code structure is output. These judgments can be presented in text form, indicating the code locations and problem types where problems may exist, such as overly deep loop nesting, abnormal function call relationships, etc.
[0030] In this embodiment, the code is input into the symbolic execution path explorer, and symbolic values are assigned to the variables and expressions in the code; the symbolic values can represent any possible values, so that the execution of the code under different input conditions can be simulated; the symbolic execution path explorer will start from the starting point of the code and select different execution paths to explore according to the control flow and conditional statements of the code; when encountering a conditional judgment statement, it will consider the cases where the condition is true and false respectively, and generate different execution paths; during each step of execution, the symbolic executor will update the state of the code, including the values of variables, memory state, etc.; for an assignment statement, it will update the symbolic value of the corresponding variable; for a function call, it will handle the parameter passing and return value of the function; during the exploration process, a series of path constraint conditions will be generated, which describe the conditions that need to be satisfied for the code to execute to the current path; use a constraint solver (such as Z3) to solve these constraint conditions to determine whether the path is reachable; during the execution path exploration process, it will check whether there are logical errors, such as deadlocks, race conditions, null pointer references, etc.; for example, by analyzing the synchronization operations and resource access of threads, detect the possibility of deadlocks; pay attention to the execution of the code under boundary conditions, and check whether there are problems such as out-of-bounds access and division-by-zero errors; output the possible problems in the identified code logic, including the description of the problem, the code location where the problem may occur, and the relevant execution path information.
[0031] In this embodiment, the previously extracted code style features (including layout style, naming style, and pattern style) are input into the contrastive learning style transfer discriminator; load the preset good style standards, which can be summarized by analyzing a large number of excellent codes, or can be formulated according to programming specifications and best practices; the good style standards are also represented in the form of feature vectors; the contrastive learning style transfer discriminator will match the input code style features with the preset good style standards and calculate the similarity between them; cosine similarity, Euclidean distance, etc. can be used as measurement methods to measure the similarity between features; analyze the differences between the code style features and the good style standards to find out the parts that do not meet the specifications; for example, for the layout style, check whether the indentation is consistent and whether the use of spaces is standardized; for the naming style, check whether the variable names and function names conform to the naming convention; for the pattern style, check whether the design pattern is used correctly; according to the results of the contrastive learning, judge whether the layout style, naming style, and pattern style of the code meet the specifications; output the judgment results, and for the parts that do not meet the specifications, give specific suggestions and improvement directions, such as suggesting to modify the indentation method and adjust the naming rules.
[0032] Please refer to Figure 2 , the structural schematic diagram of the intelligent code review system based on the large language model provided by the embodiment of the present invention, the system includes: An acquisition module, configured to acquire the code to be reviewed, perform lexical analysis, syntax analysis, and semantic analysis on the code to be reviewed, extract the key information of the code, and obtain the preprocessed code; An extraction module, configured to extract various features of the code from the preprocessed code to obtain target code features; A generation module, configured to input the target code features into a large language model, and generate a code review result through the large language model.
[0033] In this embodiment, the extraction module includes: A construction sub-module, configured to construct a control flow graph based on the preprocessed code, present the execution logic of the code in a graphical form, where each node represents a basic code block and the edge represents the jump relationship of code execution; A capture sub-module, configured to process the control flow graph using a fractal convolutional network, capture multi-scale structural patterns, input the result after fractal convolution processing into a spectral clustering algorithm, and extract the modular features of the code; A calculation sub-module, configured to calculate the structural entropy based on the extracted modular features, obtain a measurement index of the code structure features, decouple and analyze the code style, and combine the measurement index of the code structure features to obtain the target code features.
[0034] In this embodiment, the generation module includes: An input sub-module, configured to input the target code features into a large language model, perform structural review processing, logical review processing, and style review processing on the target code features through the federated architecture in the large language model, and output a mixed review result; An identification sub-module, configured to input the target code features into the CodeBERT layer for screening, analyze the code features, identify errors and potential problems, and obtain a feature review result; An integration sub-module, configured to integrate the mixed review result and the feature review result to comprehensively obtain the code review result.
[0035] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent code review method based on a large language model, characterized in that: The method comprises the following steps: Obtain the code to be reviewed, perform lexical analysis, syntax analysis, and semantic analysis on the code to be reviewed, extract key information of the code, and obtain the preprocessed code; Extracting multiple features of the code from the preprocessed code to obtain target code features; The target code features are input into the large language model, and the code review results are generated through the large language model.
2. The intelligent code review method based on a large language model as described in claim 1, characterized in that: The step of obtaining the code to be reviewed, performing lexical analysis, grammatical analysis and semantic analysis on the code to be reviewed, extracting key information of the code and obtaining the preprocessed code includes: Obtain the code to be reviewed, perform lexical analysis on the code to be reviewed, and split the code to be reviewed into individual words; According to the word sequence obtained by lexical analysis, the abstract syntax tree of the code is constructed, and the abstract syntax tree obtained by preprocessing is encoded using a dynamic syntax tree encoder. The constructed abstract syntax tree is semantically checked to extract the key information of the code and obtain the preprocessed code.
3. The intelligent code review method based on a large language model as described in claim 2, characterized in that: The dynamic syntax tree encoder adopts a dual-stream architecture. The GGNN layer captures the topological structure features of the abstract syntax tree. The Transformer layer identifies the semantic associations across nodes in the abstract syntax tree through an attention mechanism. The two representations obtained by the GGNN layer and the Transformer layer are dynamically fused using a gating mechanism, and a syntax tree pruning attention mechanism is introduced to filter redundant branches in the abstract syntax tree.
4. The intelligent code review method based on a large language model as claimed in claim 1, characterized in that: The method extracts multiple features of the code from the preprocessed code to obtain target code features, including: Construct a control flow graph based on the preprocessed code to present the execution logic of the code in a graphical form. Each node represents a basic code block, and the edge represents the jump relationship of the code execution; Use fractal convolutional networks to process control flow graphs to capture multi-scale structural patterns, and input the results of fractal convolution processing into the spectral clustering algorithm to extract the modular features of the code; Based on the extracted modular features, the structural entropy is calculated to obtain the measurement indicators of the code structure characteristics, the code style is decoupled and analyzed, and the target code characteristics are obtained by combining the measurement indicators of the code structure characteristics.
5. The intelligent code review method based on a large language model as described in claim 4, characterized in that: The decoupling and analysis of the code style includes: Analyze the layout elements in the code, including at least indentation and spaces, and extract layout style features; Perform word vector clustering on identifiers in the code, including at least variable names and function names, to extract naming style features; Identify the design patterns used in the code, generate design pattern fingerprints, and extract pattern style features; Integrate the layout style features, naming style features and pattern style features of the code to obtain style elements; By comparing the disentanglement learning method, the style elements are separated from the functional semantics of the code to obtain the code style features.
6. The intelligent code review method based on a large language model as claimed in claim 1, characterized in that: The step of inputting the target code features into the large language model and generating the code review results through the large language model includes: Input the target code features into the large language model, perform structural review, logical review, and style review on the target code features through the federated architecture in the large language model, and output a mixed review result; Input the target code features into the CodeBERT layer for screening, analyze the code features, identify errors and potential problems, and obtain feature review results; Integrate the mixed review results and feature review results to obtain the comprehensive code review results.
7. The intelligent code review method based on a large language model as claimed in claim 6, characterized in that: The target code features are input into the large language model, and the target code features are subjected to structural review processing, logical review processing, and style review processing through the federated architecture in the large language model, and a mixed review result is output, including: The anomaly detector of the Graph Transformer network in the large language model processes the code structure features, analyzes the code structure pattern, detects whether there are structural anomalies, and outputs a preliminary judgment on possible errors and potential problems in the code structure; The symbolic execution path explorer in the large language model is used to perform symbolic execution exploration on the execution path of the code, simulate the execution of the code under different conditions, and identify possible problems in the code logic; The contrastive learning style transfer discriminator in the large language model compares the style features of the code with the preset good style standards to determine whether the layout style, naming style and pattern style of the code meet the specifications.
8. An intelligent code review system based on a large language model, characterized in that: The system includes: The acquisition module is used to obtain the code to be reviewed, perform lexical analysis, syntax analysis and semantic analysis on the code to be reviewed, extract key information of the code, and obtain the preprocessed code; An extraction module is used to extract various features of the code from the preprocessed code to obtain target code features; The generation module is used to input the target code features into the large language model and generate code review results through the large language model.
9. The intelligent code review system based on a large language model as claimed in claim 8, characterized in that: The extraction module comprises: The construction submodule is used to construct a control flow graph based on the preprocessed code, presenting the execution logic of the code in a graphical form. Each node represents a basic code block, and the edge represents the jump relationship of the code execution; The capture submodule is used to process the control flow graph using a fractal convolutional network to capture multi-scale structural patterns, and input the results of the fractal convolution processing into the spectral clustering algorithm to extract the modular features of the code; The calculation submodule is used to calculate the structural entropy based on the extracted modular features, obtain the measurement indicators of the code structure characteristics, decouple and analyze the code style, and obtain the target code characteristics in combination with the measurement indicators of the code structure characteristics.
10. The intelligent code review system based on a large language model as claimed in claim 8, characterized in that: The generation module comprises: An input submodule is used to input the target code features into the large language model, perform structural review, logical review and style review on the target code features through the federated architecture in the large language model, and output a mixed review result; The recognition submodule is used to input the target code features into the CodeBERT layer for screening, analyze the code features, identify errors and potential problems, and obtain feature review results; The integration submodule is used to integrate the mixed review results and the feature review results to obtain the comprehensive code review results.
Citation Information
Patent Citations
Large-scale code data feature extraction method and system
CN117113347A
Code review method and device, electronic equipment and medium
CN117648931A
Code review method and system
CN118349454A
Cited By
Programming review method and system based on multi-modal data
CN120804726A
Center knowledge base-based general manuscript checking method and device
CN121210520A
Java code intelligent review method based on call chain analysis
CN122527020A
Java code intelligent review method based on call chain analysis
CN122527020B
A hybrid intelligent code review system and method based on model agent layer packaging
CN122816599A