Code review method and device, electronic equipment and storage medium

By using a pre-stored knowledge base and a large language model (LLM) for in-depth analysis in code review, the problems of low efficiency and insufficient accuracy in existing code review technologies are solved, achieving more efficient and accurate code review.

CN120872780APending Publication Date: 2025-10-31YOUKA (CHANGZHOU) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510923408.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing code review methods rely on manual inspection, which is inefficient and inaccurate, makes it difficult to understand the code context, and traditional tools have a high false alarm rate, making them unsuitable for multiple programming languages ​​and dynamic business needs.

Method used

By acquiring the code to be reviewed, the target context information is retrieved using a pre-stored knowledge base, and then input into a pre-defined large language model (LLM) along with the review instructions for in-depth analysis to generate the target review results.

Benefits of technology

It improves the accuracy and efficiency of code review, reduces the reliance on experience and subjectivity in manual review, and provides explanatory and actionable review comments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872780A_ABST
    Figure CN120872780A_ABST
Patent Text Reader

Abstract

Embodiments of the invention disclose a code review method and apparatus, an electronic device and a storage medium, and can solve the problem of how to effectively improve the accuracy and efficiency of code review. The method comprises the steps of obtaining a to-be-reviewed code, wherein the to-be-reviewed code comprises code content and a review instruction; in a pre-stored knowledge base, retrieving the code content to obtain target context information; obtaining target input content according to the code content and the target context information; and inputting the target input content and the review instruction into a preset large language model to obtain a target review result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of large language model technology, and in particular to a code review method, apparatus, electronic device and storage medium. Background Technology

[0002] The current code review process relies heavily on manual processes. After developers submit PRs / MRs, team members must meticulously review the code line by line for logic, style, and potential defects. Since manual review depends heavily on the reviewer's experience, its efficiency is low. To address this, rule-based automated static analysis tools and traditional machine learning models are used to check code, which can uncover potential coding style issues, error patterns, security vulnerabilities, and performance risks. However, these automated solutions still have a certain false positive rate, and their inability to adequately understand the context of complex code leads to a significant cognitive burden. Therefore, effectively improving the accuracy and efficiency of code reviews has become a pressing issue. Summary of the Invention

[0003] To address, or at least partially address, the aforementioned technical problems, embodiments of this application provide a code review method, apparatus, electronic device, and storage medium to solve the problem of how to effectively improve the accuracy and efficiency of code review.

[0004] To achieve the above objectives, the technical solutions provided in this application are as follows:

[0005] In a first aspect, embodiments of this application provide a code review method, the code review method comprising: obtaining code to be reviewed, the code to be reviewed comprising: code content and review instructions;

[0006] The code content is retrieved from the pre-stored knowledge base to obtain the target context information;

[0007] Based on the code content and the target context information, the target input content is obtained;

[0008] The target input content and the review instructions are input into a preset large language model to obtain the target review result.

[0009] As an optional implementation, in a first aspect of the embodiments of this application, obtaining the code to be reviewed includes:

[0010] Get the initial code input by the user;

[0011] The initial code is subjected to data extraction and preprocessing to obtain the code to be reviewed.

[0012] As an optional implementation, in a first aspect of the embodiments of this application, the step of retrieving the code content from a pre-stored knowledge base to obtain target context information includes:

[0013] In the pre-stored knowledge base, a vector search is performed on the code content to obtain multiple initial context information;

[0014] Calculate the semantic similarity between each initial context information and the code content;

[0015] The target context information is determined based on the semantic similarity.

[0016] As an optional implementation, in a first aspect of the embodiments of this application, obtaining the target input content based on the code content and the target context information includes:

[0017] The code content and the target context information are integrated to obtain the initial input content;

[0018] If the size of the initial input content exceeds the model processing threshold, the initial input content is compressed to obtain the target input content, wherein the target input content is less than or equal to the model processing threshold.

[0019] As an optional implementation, in a first aspect of this application, the step of compressing the initial input content to obtain the target input content includes:

[0020] After preprocessing the data patches in the initial input content, the preprocessed initial input content is divided according to language type to obtain multiple data queues;

[0021] Each data queue is sorted by priority, resulting in multiple sorting results;

[0022] Based on the multiple sorting results, the data in each data queue is polled sequentially to obtain the target input content.

[0023] As an optional implementation, in a first aspect of the embodiments of this application, the step of inputting the target input content and the review instruction into a preset large language model to obtain the target review result includes:

[0024] The target input content and the review instructions are input into the preset large language model to obtain the initial review result;

[0025] The initial review results are parsed and structured to obtain the target review results.

[0026] As an optional implementation, in a first aspect of the embodiments of this application, the method further includes:

[0027] Output the target review results to the user;

[0028] Obtain feedback data from the user regarding the target review results;

[0029] Based on the feedback data and the target review results, the pre-stored knowledge base and the preset large language model are improved.

[0030] Secondly, embodiments of this application provide a code review device, the code review device comprising: an acquisition module, configured to acquire code to be reviewed, the code to be reviewed comprising: code content and review instructions;

[0031] The processing module is used to retrieve the code content from a pre-stored knowledge base to obtain target context information;

[0032] The processing module is further configured to obtain target input content based on the code content and the target context information;

[0033] The processing module is also used to input the target input content and the review instructions into a preset large language model to obtain the target review result.

[0034] As an optional implementation, in a second aspect of the embodiments of this application, the acquisition module is specifically used to acquire the initial code input by the user;

[0035] The processing module is specifically used to extract and preprocess data from the initial code to obtain the code to be reviewed.

[0036] As an optional implementation, in a second aspect of the embodiments of this application, the processing module is specifically used to perform vector search on the code content in the pre-stored knowledge base to obtain multiple initial context information;

[0037] The processing module is specifically used to calculate the semantic similarity between each initial context information and the code content;

[0038] The processing module is specifically used to determine the target context information based on the semantic similarity.

[0039] As an optional implementation, in a second aspect of the embodiments of this application, the processing module is specifically used to integrate the code content and the target context information to obtain initial input content;

[0040] The processing module is specifically used to compress the initial input content to obtain the target input content if the size of the initial input content exceeds the model processing threshold. The target input content is less than or equal to the model processing threshold.

[0041] As an optional implementation, in the second aspect of the embodiments of this application, the processing module is specifically used to preprocess the data patches in the initial input content, and then divide the preprocessed initial input content according to the language type to obtain multiple data queues;

[0042] The processing module is specifically used to perform priority sorting on each group of data queues to obtain multiple sorting results.

[0043] The processing module is specifically used to sequentially poll the data in each data queue according to the multiple sorting results to obtain the target input content.

[0044] As an optional implementation, in a second aspect of the embodiments of this application, the processing module is specifically used to input the target input content and the review instruction into the preset large language model to obtain an initial review result;

[0045] The processing module is specifically used to parse and structure the initial review results to obtain the target review results.

[0046] As an optional implementation, in a second aspect of the embodiments of this application, the processing module is further configured to output the target review result to the user;

[0047] The acquisition module is also used to acquire feedback data from the user regarding the target review result;

[0048] The processing module is further configured to improve the pre-stored knowledge base and the preset large language model based on the feedback data and the target review results.

[0049] Thirdly, embodiments of this application provide an electronic device, the electronic device comprising:

[0050] Memory containing executable program code;

[0051] A processor coupled to the memory;

[0052] The processor calls the executable program code stored in the memory to execute the code review method in the first aspect of the embodiments of this application.

[0053] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that causes a computer to execute the code review method of the first aspect of embodiments of this application. The computer-readable storage medium includes ROM / RAM, a magnetic disk, or an optical disk, etc.

[0054] Fifthly, embodiments of this application provide a computer program product that, when run on a computer, causes the computer to perform some or all of the steps of any of the methods of the first aspect.

[0055] Sixthly, embodiments of this application provide an application publishing platform for publishing computer program products, wherein when the computer program product is run on a computer, the computer performs some or all of the steps of any of the methods of the first aspect.

[0056] Compared with the prior art, the embodiments of this application have the following beneficial effects:

[0057] This application provides a code review method, apparatus, electronic device, and storage medium. The method involves acquiring code to be reviewed, which includes code content and review instructions. The code content is retrieved from a pre-stored knowledge base to obtain target context information. Based on the code content and target context information, target input content is obtained. The target input content and review instructions are then input into a pre-defined large language model to obtain the target review result. This solution dynamically retrieves and integrates content from the knowledge base, enabling the model to go beyond analyzing individual code snippets and understand the contextual meaning and impact of code content within the entire project environment. This overcomes the problems of insufficient contextual understanding in traditional review methods and insufficient contextual depth in early review tools. Furthermore, the review instructions guide the LLM (Large Language Model) to conduct comprehensive and in-depth analysis from multiple dimensions, reducing inconsistencies and subjectivity in review results caused by differences in human review experience, personal preferences, or current status. This ensures that the output review opinions are interpretable, operable, and constructive. Attached Figure Description

[0058] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0059] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Figure 1 This is a flowchart illustrating a code review method provided in an embodiment of this application. Figure 1 ;

[0061] Figure 2 This is a flowchart illustrating a code review method provided in an embodiment of this application. Figure 2 ;

[0062] Figure 3 This is a schematic diagram of an interface for a code review method provided in an embodiment of this application. Figure 1 ;

[0063] Figure 4 This is a schematic diagram of an interface for a code review method provided in an embodiment of this application. Figure 2 ;

[0064] Figure 5 This is a schematic diagram of an interface for a code review method provided in an embodiment of this application. Figure 3 ;

[0065] Figure 6 This is a flowchart illustrating a code review method provided in an embodiment of this application. Figure 3 ;

[0066] Figure 7 This is a schematic diagram of the structure of a code review device provided in an embodiment of this application;

[0067] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0068] To better understand the above-mentioned objectives, features, and advantages of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features of this application can be combined with each other. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0069] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, rather than to describe a specific order of objects.

[0070] The terms “comprising” and “having”, and any variations thereof, in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or device.

[0071] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0072] The current code review process relies primarily on manual processes. After developers submit PRs / MRs, team members must check the code line by line for logic, style, and potential defects. Although static code analysis tools (such as SonarQube) and automated testing tools (such as Jenkins) exist, they can only detect syntax errors or simple rule violations and cannot deeply analyze the rationality of code logic, business consistency, and complex security vulnerabilities.

[0073] Pull Request (PR) / Merge Request (MR): In code collaboration platforms (such as GitHub and GitLab), this is the process by which developers submit code change requests to the main branch, which then need to be reviewed and merged. It is the core platform for code review, discussion, and automated checks.

[0074] Existing code reviews can be achieved through manual code review: this is the most traditional and basic method. After developers submit code changes in a PR / MR, designated reviewers (usually other developers on the team) manually read and analyze the code, checking its logical correctness, style consistency, maintainability, potential defects, and whether it meets project requirements and coding standards. Reviewers leave feedback via comment functions on specific lines of code or in the PR / MR's discussion area. The author modifies the code based on the feedback until it passes the review before merging.

[0075] Existing code reviews can also be achieved through rule-based automated static analysis tools: without actually executing the code, these tools examine the code using techniques such as lexical analysis, syntax analysis, control flow analysis, and data flow analysis to uncover potential coding style issues, error patterns, security vulnerabilities, and performance risks. Common static analysis tools include Linters (such as ESLint, Pylint, and Checkstyle) and more sophisticated static application security testing tools.

[0076] Existing code review can also be achieved with simple AI-assisted tools: traditional machine learning models (such as random forests) are used to classify code defects, but this requires manual annotation of training data and has poor generalization ability.

[0077] While existing code review technologies and tools have improved software development quality and efficiency to some extent, they still have the following major drawbacks from a technical perspective. These drawbacks collectively constitute the technical problem that this application aims to solve:

[0078] Inefficient: Manual review is time-consuming and depends on the reviewer's experience level, especially in large-scale codebases, where response delays are significant (technical perspective: excessive time complexity).

[0079] Excessive cognitive load: Reviewers need to invest a lot of energy to understand the context of code changes, identify potential subtle errors, and recall project-specific conventions and specifications, which brings a huge cognitive burden.

[0080] Insufficient accuracy: Rule-based tools have a high false positive rate, requiring secondary manual screening. When the code logic is complex, it is difficult for both the human eye and the rules to detect potential defects or security vulnerabilities. (From a technical perspective: there is a contradiction between rule coverage and code diversity).

[0081] Lack of contextual understanding: Reviewers need to manually compare the relationship between code changes and other dependent modules.

[0082] Poor scalability: Existing solutions are difficult to adapt to multiple programming languages ​​and dynamic business needs (from a technical perspective: static rules do not match dynamic scenarios).

[0083] Traditional static analysis tools primarily rely on predefined rule sets and pattern matching. They typically lack an understanding of deep code semantics, the developer's true intent, and the project-specific context, thus failing to generate actionable improvement suggestions. Traditional static analysis tools are prone to generating numerous false positives or unimportant warnings, leading to "warning fatigue" and causing developers to overlook truly important issues. Rule set updates and customizations in traditional static analysis tools usually require manual intervention, lacking the ability to automatically learn and adapt to project evolution.

[0084] To address some or all of the aforementioned technical problems, embodiments of this application provide a code review method, apparatus, electronic device, and storage medium. The method involves acquiring code to be reviewed, which includes code content and review instructions; retrieving the code content from a pre-stored knowledge base to obtain target context information; obtaining target input content based on the code content and target context information; and inputting the target input content and review instructions into a preset large language model to obtain the target review result. In this solution, by dynamically retrieving and integrating content from the knowledge base, the model can transcend the analysis of individual code fragments and understand the contextual meaning and impact of code content within the entire project environment. This overcomes the problems of insufficient contextual understanding in traditional review methods and insufficient contextual depth in early review tools. Furthermore, the review instructions can guide the LLM to conduct comprehensive and in-depth analysis from multiple dimensions, reducing inconsistencies and subjectivity in review results caused by differences in human review experience, personal preferences, or current status, ensuring that the output review opinions are interpretable, operable, and constructive.

[0085] like Figure 1 As shown, Figure 1 A flowchart of a code review method is provided for embodiments of this application. The method may include the following steps:

[0086] 101. Obtain the code to be reviewed.

[0087] In this embodiment, the code to be reviewed may include: code content and review instructions. The code content may be code fields that the user inputs and that require review. The review instructions may also be instructions that the user inputs to review the code content. Through these review instructions, the user's desired review perspective can be determined, and these instructions can be understood as review prompts.

[0088] 102. Retrieve the code content from the pre-stored knowledge base to obtain the target context information.

[0089] In this embodiment, the pre-stored knowledge base can be dynamically updated. The knowledge base may include: project-specific documents, such as project coding style documents, architecture design documents, API usage manuals, domain knowledge bases, etc.; historical data, including previously closed PR / MR data in the project, including successful code change patterns, common defect fixing cases, valuable historical review comments and discussions; and existing codebase, including relevant code snippets outside the modules involved in the current PR / MR, used to understand the location and dependencies of code changes in the entire project.

[0090] It should be noted that this knowledge base can also store code-related data and review results after each user review. This allows the data to be retrieved for reference during subsequent reviews, meaning that the knowledge base is dynamically updated.

[0091] In this embodiment of the application, after obtaining the code content, the code content can be retrieved in the knowledge base. This allows the acquisition of information fragments related to the code content. The retrieved information can be provided to the model as additional target context information to improve its performance on code-related tasks.

[0092] 103. Based on the code content and target context information, obtain the target input content.

[0093] In this embodiment, since the code content can be understood as the code fields modified by the user, the user may only upload a small part of the code without relevant information such as the code application scenario, the purpose of the code modification, the code execution language, and the code execution effect. Therefore, if only this part of the code is reviewed, the review results may be one-sided and cannot obtain comprehensive and accurate review results. Therefore, the code content and target context information can be integrated to obtain complete target input content. This can be understood as building a complete code architecture for the code content, so that the model can review the code content from multiple perspectives.

[0094] 104. Input the target input content and review instructions into the preset large language model to obtain the target review results.

[0095] In this embodiment of the application, the preset large language model is the LLM model, which is a natural language processing model based on deep learning. It can generate or understand complex text through training on large-scale corpora, such as GPT-4, DeepSeek, Claude, Gemini, etc. By inputting the target input content into the LLM, the target input content can be automatically reviewed by the LLM model.

[0096] It should be noted that since the target input includes code-related content from various angles, the LLM model may not be clear about which specific aspects need to be reviewed or which aspects need to be emphasized. Therefore, the review instructions included in the code to be reviewed can be input into the LLM model. In this way, the LLM model can review the target input according to the content indicated by the review instructions, thereby obtaining the target review result.

[0097] In some embodiments, the review of the target input content may specifically include the following aspects:

[0098] (1) Generate a description of the PR (title, type, summary, changes, tags, etc.).

[0099] (2) Identify potential logical errors, boundary condition problems, concurrency problems, null pointer exceptions and other defects.

[0100] (3) Detect security vulnerabilities, such as injection risks, insecure API calls, and sensitive data leaks.

[0101] (4) Check whether the code follows general coding best practices and project-specific coding standards (obtained through RAG).

[0102] (5) Evaluate the readability, maintainability (e.g., complexity, duplicate code) and design rationality of the code.

[0103] (6) Propose specific improvement suggestions, alternative implementation schemes, or summarize the potential impact of code changes.

[0104] Of course, it can also include reviews from other perspectives. The review instruction will instruct the LLM to perform a specific review task. The review task can be at least one of the above six items, or it can be other review perspectives. This application embodiment does not make specific limitations.

[0105] This application provides a code review method that dynamically retrieves and integrates content from a knowledge base, enabling the model to go beyond analyzing individual code snippets and understand the contextual meaning and impact of code content within the entire project environment. This overcomes the problems of insufficient contextual understanding in traditional review methods and the lack of contextual depth in early review tools. Furthermore, review instructions can guide the LLM to conduct comprehensive and in-depth analysis from multiple dimensions, reducing inconsistencies and subjectivity in review results caused by differences in human review experience, personal preferences, or current status. This ensures that the output review opinions are interpretable, operable, and constructive.

[0106] like Figure 2 As shown, Figure 2 A flowchart of a code review method is provided for embodiments of this application. The method may include the following steps:

[0107] 201. Obtain the initial code input by the user.

[0108] In this embodiment of the application, the initial code entered by the user can be a new PR / MR created on GitLab or GitHub, or a new code commit pushed on top of an existing PR / MR.

[0109] 202. Extract and preprocess the initial code to obtain the code to be reviewed.

[0110] In this embodiment, the initial code input by the user can include a lot of information. Data extraction can be performed first, and the extracted data can specifically include: the specific content of code changes (Diffs), the title and description of the PR / MR, related commit messages, and existing comments and discussions in the PR / MR (used to understand the reviews and context). Of course, other information can also be extracted.

[0111] In this embodiment of the application, after data extraction, the data can be preprocessed. This preprocessing can specifically include: data cleaning, data parsing, data formatting, data association, etc. For example: obtaining the proportion of the main programming languages ​​in the repository and sorting files according to language importance; ignoring files according to filter configuration; obtaining additional lines before and after each code change as supplementary context; and combining the commit history to see which files are frequently committed together, constructing a relationship graph between files, integrating deeply related files together, and establishing potential cognition to understand the global structure of the entire code repository.

[0112] 203. In the pre-stored knowledge base, perform a vector search on the code content to obtain multiple initial context information.

[0113] In this embodiment of the application, when retrieving content in the knowledge base, vector search and semantic similarity matching based on code and text embedding can be used. Therefore, during retrieval, code content (such as modified functions, classes, and APIs used) can be converted into vector patterns. Through keyword matching and vector search technology, multiple initial contextual information matching the code content can be found in the knowledge base (including project documents, historical code, coding standards, etc.).

[0114] 204. Calculate the semantic similarity between each initial context information and code content.

[0115] In this embodiment of the application, in order to ensure the accuracy of data retrieval and model calculation, it is necessary to determine and integrate the context information most related to the code content. Therefore, it is necessary to find the context information most similar to the code content from multiple initial context information.

[0116] In some embodiments, there can be multiple ways to calculate semantic similarity. For example, cosine similarity can be used to calculate the cosine value of the angle between vectors, or Euclidean distance can be used to calculate the straight-line distance between vectors in the vector space. These are all existing technologies, and the embodiments of this application do not specifically limit or describe them.

[0117] 205. Determine the target context information based on semantic similarity.

[0118] In this embodiment of the application, multiple initial context information can be sorted according to the magnitude of semantic similarity, and the target context information can be determined based on the sorting result.

[0119] In some embodiments, the initial context information with the highest semantic similarity can be selected as the target context information.

[0120] In some embodiments, at least one initial context information with a semantic similarity greater than a preset threshold can be selected as the target context information.

[0121] 206. Integrate the code content and target context information to obtain the initial input content.

[0122] 207. If the size of the initial input content exceeds the model processing threshold, the initial input content is compressed to obtain the target input content.

[0123] In this embodiment of the application, since the LLM model has a certain upper limit on the number of tokens it can process, that is, if there is too much input content, the LLM model may not be able to process it or may not be able to process it simultaneously, and may need to process it in several batches, which greatly affects the review efficiency. Therefore, if the size of the initial input content obtained after integrating the code content and the target context information exceeds the model processing threshold, the initial input content needs to be compressed until the target input content obtained after compression is less than or equal to the model processing threshold.

[0124] It should be noted that the processing threshold of this model can be determined based on the size and architecture of the LLM model.

[0125] In some embodiments, data compression is performed on the initial input content to obtain the target input content. Specifically, this may include: preprocessing the data patches in the initial input content, dividing the preprocessed initial input content according to the language type to obtain multiple data queues; prioritizing each data queue to obtain multiple sorting results; and sequentially polling the data in each data queue according to the multiple sorting results to obtain the target input content.

[0126] It should be noted that user modifications to the code may involve adding or deleting code. However, the code uploaded by the user can be understood as a revision. That is, even if the user deletes part of the code, that part still occupies space. Therefore, if the size of the initial input content has exceeded the model's processing threshold, the deleted code can be discarded, and only the added code can be considered. This completes the preprocessing of the data patch, which is the part of the code modified by the user. All deleted files are merged into a separate list, which can record only the name without including code details.

[0127] In addition, since the code may include multiple language types, the preprocessed initial input content can be divided according to the language type to obtain multiple data queues. Specifically, the tiktoken model can be used to segment and tokenize the changed code, and the files can be divided into independent processing queues according to different language types to maintain the integrity of the language context.

[0128] For example, the divided data queues may include;

[0129] [[file2.py,file.py], #Python filegroup;

[0130] [file4.jsx,file3.js],#JS / JSX file group;

[0131] [readme.md]#Document file group.

[0132] Furthermore, each data queue for different language types can include at least one file with modified content. Some files have significant changes, such as a large amount of modified code or code affecting the core structure of the system; others have minor changes, such as a small amount of modified code or a change to only one parameter. Therefore, within each language type's data queue, a project language priority sort is implemented, and files are arranged in descending order of the token size of individual file patches (files with significant changes are given priority). This results in multiple sorting results, ensuring that large-scale changes to core business logic are prioritized for the review process.

[0133] For example, the sorting result can be represented as:

[0134] sorted_groups = [

[0135] sorted(py_files,key=token_count,reverse=True),

[0136] sorted(js_files,key=token_count,reverse=True),

[0137] #...Other language groups processing ]

[0139] Furthermore, after sorting each data queue, the data in each queue can be polled sequentially based on the sorting results. This involves three steps:

[0140] Step 1: Full code patching. The system polls across language groups to select the largest unprocessed file and continuously loads the complete patch content until it is occupied (MAX_TOKEN-BUFFER). This can be understood as selecting the highest-priority data in each data queue as the target input, then selecting the second-highest-priority data, and so on, until the selected data reaches a preset percentage of the model's processing threshold, such as 80%.

[0141] For example, assuming the model's processing threshold is 10,000 characters and the preset ratio is 80%, there are three data queues: data queue A, data queue B, and data queue C. Each queue contains five data items. Based on the sorting result, A1, B1, and C1 are selected as the target input content. If the target input content contains 1,000 characters, which is less than 80% of 10,000 characters, then A2, B2, and C2 are selected as the target input content, and so on. Assuming that after selecting A4, B4, and C4 as the target input content, it is detected that the target input content contains 8,000 characters, which has reached 80% of 10,000 characters, then step two continues. If after selecting A5, B5, and C5 as the target input content, it still has not reached 80% of 10,000 characters, then the data compression ends, and there is no need to continue to the next step.

[0142] For example, it can be represented as:

[0143] while current_tokens<(max_tokens-buffer):

[0144] next_file=find_largest_remaining()

[0145] prompt+=render_full_patch(next_file)

[0146] Step 2: Simplify the list (BUFFER area). Compress the remaining files into a list of ##Other Modified Files, keeping only the filename and modification type marker in a very simple format. This can be understood as compressing the remaining files in some data queues and keeping only the filename and modification marker as the target input content.

[0147] For example, it can be represented as:

[0148] ##Other Modified Files

[0149] -[CHG]utils / helper.js#Modify

[0150] -[DEL]legacy / api.py#Delete

[0151] Step 3: Residual capacity recycling. If there are still spare tokens, append the path list of deleted files to maintain a highly compact display. This can be understood as follows: if the model processing threshold is still not full after Step 2, then the path list of user-deleted code that was discarded when preprocessing the data patches in the initial input content can be added to the target input content.

[0152] For example, it can be represented as:

[0153] ##Deleted Files

[0154] -deprecated / old_module.py

[0155] -unused.config

[0156] The final target input content can be represented as:

[0157] #Full Patches

[0158] ##file2.py(352 tokens)

[0159] +++

[0160] def new_feature(): ...

[0162] ##file4.jsx(287 tokens)

[0163] +++

[0164] export const Component=()=>...

[0165] #Other Modified Files

[0166] -[CHG]file.py

[0167] -[FIX]file3.js

[0168] #Deleted Files

[0169] -obsolete.rb

[0170] It should be noted that if the initial input is particularly large, such as being a multiple of the model's processing threshold, it can be split into multiple parts for analysis and processing, and then the final LLM results can be merged.

[0171] 208. Input the target input content and review instructions into the preset large language model to obtain the initial review results.

[0172] In this embodiment, the review instruction may include not only the code change itself, but also rich background knowledge and explicit review instructions. For example, the review instruction might ask the LLM: "Please review the following Python code changes. The changes involve the user authentication module. Please check for potential authentication bypass vulnerabilities or data leakage risks, taking into account the project's secure coding standards (attached) and recent relevant security patches (attached). For each issue found, please explain its principle and provide remediation suggestions." The LLM processes this enhanced prompt, performs in-depth analysis, and the initial review results may specifically include: a description of the PR, identified issues, explanations of the issues, and code modification suggestions.

[0173] 209. Analyze and structure the initial review results to obtain the target review results.

[0174] In this embodiment of the application, since the review results need to be output to the user and displayed on the code review platform, the initial review results need to be parsed and structured to convert them into target review results that conform to the comment format of the code review platform and are easy to read and understand. The target review results will clearly point out the line of code where the problem is located, explain the reason for the problem, and provide specific modification suggestions. The suggestions will be classified according to the confidence level of the LLM output or the problem type (such as severity, type: defect / suggestion / question).

[0175] 210. Output the target review results to the user.

[0176] In this embodiment of the application, the target review result is output to the user. The target review result can be added as a comment to the corresponding line of code in the PR / MR or the overall discussion area, so that the user can easily view, discuss and respond to the suggestions generated by LLM as if they were human review comments.

[0177] For example, the output to the user can be as follows Figure 3 , Figure 4 and Figure 5 As shown, in Figure 3 In the process, scan for PR code changes and generate descriptions for the PR—title, type, summary, changes, and tags; Figure 4In the process, the system scans for PR code changes and generates a feedback list for reviewers to assist with the review process; Figure 5 In the process, it scans for changes in PR code and automatically generates meaningful suggestions to help code authors improve their PR code.

[0178] 211. Obtain user feedback data on the target review results.

[0179] In this embodiment of the application, in order to enable the system to continuously learn and evolve, it is necessary to continuously collect user feedback data on the review results.

[0180] In some embodiments, the feedback data may include explicit feedback and implicit feedback. Explicit feedback refers to developers directly evaluating the quality and effectiveness of the LLM suggestion through interface elements ("like / dislike" emoji buttons) or providing correction suggestions in text form. Implicit feedback refers to the system analyzing the developer's subsequent behavior regarding the suggestion, such as whether the suggestion was adopted or whether a discussion was held regarding the suggestion.

[0181] 212. Based on feedback data and target review results, improve the pre-stored knowledge base and preset large language model.

[0182] In this embodiment, feedback data can be used to generate RAG reference content, automatically generate accepted suggestion documents, maintain them on the code repository's wiki page, enabling users to track historical changes and evaluate the effectiveness of tools. Accepted suggestion documents will also serve as RAG reference knowledge base content, generating more targeted and context-sensitive review results. Furthermore, RAG strategies can be optimized; for example, if contextual information from a certain type of document frequently leads to high-quality suggestions, the weight of that type of document during retrieval can be increased. The LLM can also be iterated, using collected preference data to fine-tune the LLM itself, creating a continuously improving "data flywheel" effect. Additionally, suggestion templates can be improved by analyzing feedback to determine which suggestion structures or instructions produce better results, and optimizing suggestion engineering strategies accordingly.

[0183] It's important to note that the collected feedback data is used to drive the system's learning and evolution. For example, if developers frequently mark a certain type of suggestion as "irrelevant," the system may adjust the RAG retrieval strategy or LLM prompts to reduce the generation of such suggestions. Long-term accumulated preference data can be used for periodic fine-tuning and optimization of the LLM, making its review style and focus more closely aligned with the needs of specific teams, thus forming a positive, continuously optimizing closed loop.

[0184] In some embodiments, Figure 6 The flowchart illustrates the process of implementing this method through the interaction between multiple modules, specifically, as follows: Figure 6As shown, it can specifically include:

[0185] (1) Triggering review: When developers create a new PR / MR on GitLab or GitHub, or push a new code commit to an existing PR / MR, the code hosting platform sends an event notification to this system through a pre-set webhook.

[0186] (2) Data Acquisition and Preprocessing: After receiving the notification, the input / trigger module activates the data extraction and preprocessing module. This module uses the platform API to pull detailed information of PR / MR, including code differences (diffs), metadata (title, description, author), submission history and any existing comments. The data is cleaned and formatted.

[0187] (3) Contextual Retrieval (RAG): The data extraction module passes the processed PR / MR information to the RAG module. The RAG module performs a retrieval operation in the configured knowledge base (including project documents, historical code, coding standards, etc.) based on the content of the code changes (such as modified functions, classes, and APIs used) and keywords in the PR / MR description. It finds the most relevant contextual information fragments to the current changes through techniques such as vector search.

[0188] (4) LLM Prompt Construction and Analysis: The RAG module integrates the retrieved context information with the original PR / MR data, compresses the context through an adaptive token compression strategy, and the LLM core analysis module constructs an enhanced, structured prompt. This prompt not only includes the code changes themselves, but also contains rich background knowledge and clear review instructions. The LLM performs in-depth analysis based on the enhanced prompt.

[0189] (5) Review suggestion generation: LLM outputs its analysis results, including PR description, identified issues, explanations of the issues, code modification suggestions, etc.

[0190] (6) Recommendation Release and Presentation: The recommendation generation and formatting module converts the raw output of the LLM into standardized review comments. The integration and communication module then uses the platform API to release these comments to the PR / MR, usually associated with specific lines of code.

[0191] (7) Developer interaction and feedback collection: Developers can see the review suggestions generated by LLM in the PR / MR interface and respond, discuss, or directly adopt and modify the suggestions as if they were comments from human reviewers. At the same time, developers can provide feedback on the quality of suggestions through specific controls on the interface (such as like / dislike buttons or category tags for suggestions). The feedback collection and learning module records these interactions and feedback data.

[0192] (8) Learning and Adaptation (Continuous Improvement): The collected feedback data is used to drive the system’s learning and evolution. The long-term accumulated preference data can be used to periodically fine-tune and optimize the LLM, making its review style and focus closer to the needs of specific teams, thus forming a positive and continuously optimized closed loop.

[0193] Through the above process, deep context awareness based on RAG can be achieved. By dynamically retrieving and integrating project-specific documents, historical code, and specifications, LLM can go beyond analyzing isolated code snippets and understand the meaning and impact of changes in the entire project environment, thereby providing more accurate and valuable review opinions. This solves the problem of traditional static analysis and general LLM lacking domain knowledge.

[0194] Adaptive token compression strategy: Maximize the density of key information required for code review while strictly meeting the token restrictions of LLM.

[0195] Feedback-based adaptive learning capability: By establishing a closed loop from developer feedback to model optimization, the system can continuously learn and adapt to the coding style, quality standards and preferences of specific projects and teams. This enables the quality and relevance of review suggestions to continuously improve over time, overcoming the shortcomings of insufficient adaptability of existing AI tools.

[0196] A refined, multi-dimensional feedback engineering approach goes beyond simply dumping code onto an LLM. Instead, it guides the LLM to conduct a comprehensive and in-depth analysis from multiple dimensions (such as functional correctness, security, maintainability, performance, and specification compliance) through carefully designed and dynamically constructed feedback, ensuring that the review comments it outputs are interpretable, actionable, and constructive.

[0197] Seamless integration with mainstream development platforms: Fully leverage the APIs and Webhook mechanisms provided by platforms such as GitLab and GitHub to seamlessly embed intelligent review capabilities into the PR / MR workflows familiar to developers, minimizing interference with existing development processes and improving user acceptance and ease of use.

[0198] This technical solution, through the organic integration of RAG (Rich Annotation Group) and feedback learning mechanisms, aims to address the two core challenges faced by LLM (Large-Scale Management) systems when applied to complex code review tasks: lack of context and preference alignment. RAG provides LLM with "memory" and "knowledge," while feedback learning empowers the process with the ability to upgrade and iterate, enabling it to better understand and meet the needs of specific users. Furthermore, the dynamically updated RAG knowledge base ensures the system can keep pace with the evolution of project code and documentation, while adaptive cue engineering provides flexibility for handling different review scenarios. The system even has the potential to discover implicit norms or anti-patterns that are not explicitly documented but are commonly followed by team members through long-term observation of code patterns and review feedback, thus providing a higher level of insight.

[0199] This solution, through the collaborative work of a large language model, enhanced retrieval generation, and a continuous learning feedback mechanism, has achieved significant technical results in improving code review efficiency compared to existing technologies.

[0200] This solution automates code change analysis using LLM (Local Level Management), quickly identifying various issues including potential defects, security vulnerabilities, style violations, and deviations from project standards. This significantly reduces the burden on human reviewers who need to inspect every line of code from scratch. LLM acts as the first line of defense, completing a large amount of preliminary and time-consuming checks. LLM's high-speed processing capabilities, combined with seamless integration with code hosting platform APIs, allow for rapid analysis of PR / MRs. The system's review suggestions are typically specific and targeted, reducing the number and time spent on repeated communication between developers and reviewers regarding simple issues. This directly helps shorten the PR / MR review and merge cycle time, a key indicator of development efficiency. In contrast, purely manual reviews are time-consuming and prone to becoming bottlenecks.

[0201] In this solution, LLM possesses the ability to understand the deep semantics of code. When this ability is enhanced by project-specific context (such as architecture documents, historical defect data, and secure coding standards) through the RAG module, the system can discover more complex and hidden problems than traditional static analysis tools or AI tools lacking deep context awareness. These include subtle logical errors, potential performance bottlenecks, design choices that do not align with the long-term evolution of the project, and security risks arising from the use of project-specific components or configurations. RAG provides LLM with "domain knowledge" about project coding standards, common internal pitfalls, and security requirements. This enables the system to more accurately identify problems, rather than just general pattern matching. Discovering and fixing these deep-seated problems early on, before code merging, effectively prevents the accumulation of technical debt, reduces the risk of failures in the production environment, and thus improves the reliability and robustness of the final product. This overcomes the shortcomings of insufficient context understanding in traditional static analysis tools and the lack of sufficient context depth in early AI tools.

[0202] The LLM in this solution applies a unified set of review logic and standards to all PRs / MRs based on its training data, configured review criteria (obtained through RAG), and preferences learned through feedback. This reduces inconsistencies and subjectivity in review results caused by differences in human reviewers' experience levels, personal preferences, or current states. Unlike human reviews, which may have knowledge blind spots, fluctuating attention, or insensitivity to certain issues, LLM (especially with clear engineering guidance and the context provided by RAG) provides a stable and repeatable baseline level of review. This ensures that all submitted code changes are reviewed to the same degree in terms of defined key aspects, helping to implement unified quality standards throughout the team and project lifecycle.

[0203] Furthermore, the system aims to provide clear, easy-to-understand, and actionable review feedback, typically explaining the cause of the problem and offering specific remedial suggestions. By automating numerous routine and repetitive inspection tasks, it frees developers from tedious review work. By handling many routine checks and preliminary analyses, the system effectively reduces the cognitive burden on human reviewers, allowing them to focus on more complex issues requiring human wisdom and experience (addressing the pain points of manual review). Simultaneously, the continuously improving quality of suggestions, driven by a feedback learning mechanism, ensures the relevance and usability of the recommendations, avoiding the "warning fatigue" caused by numerous false alarms seen in traditional Linter tools or the frustration caused by inaccurate suggestions in general AI tools.

[0204] In this solution, the core feedback learning loop (borrowing ideas from RLHF / RLAIF) enables the system to learn from the interaction between developers and reviewers. It gradually understands the design philosophy of a specific project, the team's coding preferences, and which types of feedback are considered most valuable. This represents a significant leap forward from static rule sets or "one-size-fits-all" AI models. This continuous learning and adaptation mechanism ensures that the system's review performance continuously optimizes over time. The system can learn to prioritize reporting certain types of issues, adjust the tone and level of detail in its comments to suit the team culture, and even understand implicit coding conventions that are not explicitly documented but are followed by team members. This makes the system more like an "intelligent partner" that grows alongside the team, rather than a fixed tool, thus overcoming the limitations of existing AI tools in terms of adaptability and personalization.

[0205] like Figure 7 As shown in the figure, this application embodiment provides a code review device, which may include: an acquisition module 701, used to acquire code to be reviewed, the code to be reviewed including: code content and review instructions;

[0206] The processing module 702 is used to retrieve code content from a pre-stored knowledge base to obtain target context information;

[0207] The processing module 702 is also used to obtain the target input content based on the code content and the target context information;

[0208] The processing module 702 is also used to input the target input content and review instructions into a preset large language model to obtain the target review result.

[0209] In some embodiments, the acquisition module 701 is specifically used to acquire the initial code input by the user;

[0210] The processing module 702 is specifically used to extract and preprocess data from the initial code to obtain the code to be reviewed.

[0211] In some embodiments, the processing module 702 is specifically used to perform vector search on the code content in a pre-stored knowledge base to obtain multiple initial context information;

[0212] Processing module 702 is specifically used to calculate the semantic similarity between each initial context information and code content;

[0213] The processing module 702 is specifically used to determine the target context information based on semantic similarity.

[0214] In some embodiments, the processing module 702 is specifically used to integrate the code content and target context information to obtain the initial input content;

[0215] The processing module 702 is specifically used to compress the initial input content to obtain the target input content if the size of the initial input content exceeds the model processing threshold. The target input content is less than or equal to the model processing threshold.

[0216] In some embodiments, the processing module 702 is specifically used to preprocess the data patches in the initial input content, and then divide the preprocessed initial input content according to the language type to obtain multiple data queues.

[0217] The processing module 702 is specifically used to perform priority sorting on each group of data queues to obtain multiple sorting results;

[0218] The processing module 702 is specifically used to sequentially poll the data in each data queue based on multiple sorting results to obtain the target input content.

[0219] In some embodiments, the processing module 702 is specifically used to input the target input content and review instructions into a preset large language model to obtain the initial review result;

[0220] The processing module 702 is specifically used to parse and structure the initial review results to obtain the target review results.

[0221] In some embodiments, the processing module 702 is further configured to output the target review result to the user;

[0222] The acquisition module 701 is also used to acquire user feedback data on the target review results;

[0223] The processing module 702 is also used to improve the pre-stored knowledge base and preset large language model based on feedback data and target review results.

[0224] In this embodiment, each module can implement the code review method provided in the above method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0225] like Figure 8 As shown in the embodiments of this application, an electronic device is also provided, which may include:

[0226] Memory 801 storing executable program code;

[0227] Processor 802 coupled to memory 801;

[0228] Specifically, the processor 802 calls the executable program code stored in the memory 801 to execute the code review method executed by the electronic device in the above method embodiments.

[0229] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the code review method described in the above method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0230] This application also provides a computer program product that stores a computer program. When the computer program is executed by a processor, it implements each process of the code review method in the above method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0231] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0232] It should be understood, in the several embodiments provided in this application, that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0233] In this application, the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0234] In this application, memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0235] In this application, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium, including permanent and non-permanent, removable and non-removable storage media. The storage medium can implement information storage by any method or technology, and the information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), other types of random access memory (RAM), read-only memory (ROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information that can be accessed by a computing device. As defined in this document, computer-readable media do not include transient media, such as modulated data signals and carrier waves.

[0236] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0237] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application. The above-mentioned multiple embodiments are not necessarily multiple independent embodiments; they are divided into multiple embodiments only to highlight different technical features in different embodiments. Those skilled in the art should understand that the above-mentioned multiple embodiments can also be combined arbitrarily.

[0238] In the various embodiments of this application, it should be understood that the sequence number of each process does not necessarily imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0239] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they can be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0240] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0241] If the aforementioned integrated units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-accessible memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several requests to cause a computer device (which can be a personal computer, server, or network device, specifically a processor in the computer device) to execute some or all of the steps of the methods described in the various embodiments of this application.

[0242] The above are merely specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A code review method, characterized in that, The method includes: Obtain the code to be reviewed, which includes: code content and review instructions; The code content is retrieved from the pre-stored knowledge base to obtain the target context information; Based on the code content and the target context information, the target input content is obtained; The target input content and the review instructions are input into a preset large language model to obtain the target review result.

2. The method according to claim 1, characterized in that, The process of obtaining the code to be reviewed includes: Get the initial code input by the user; The initial code is subjected to data extraction and preprocessing to obtain the code to be reviewed.

3. The method according to claim 1, characterized in that, The step of retrieving the code content from a pre-stored knowledge base to obtain target context information includes: In the pre-stored knowledge base, a vector search is performed on the code content to obtain multiple initial context information; Calculate the semantic similarity between each initial context information and the code content; The target context information is determined based on the semantic similarity.

4. The method according to claim 1, characterized in that, The step of obtaining the target input content based on the code content and the target context information includes: The code content and the target context information are integrated to obtain the initial input content; If the size of the initial input content exceeds the model processing threshold, the initial input content is compressed to obtain the target input content, wherein the target input content is less than or equal to the model processing threshold.

5. The method according to claim 4, characterized in that, The step of compressing the initial input content to obtain the target input content includes: After preprocessing the data patches in the initial input content, the preprocessed initial input content is divided according to language type to obtain multiple data queues; Each data queue is sorted by priority, resulting in multiple sorting results; Based on the multiple sorting results, the data in each data queue is polled sequentially to obtain the target input content.

6. The method according to claim 1, characterized in that, The step of inputting the target input content and the review instructions into a preset large language model to obtain the target review result includes: The target input content and the review instructions are input into the preset large language model to obtain the initial review result; The initial review results are parsed and structured to obtain the target review results.

7. The method according to claim 1, characterized in that, The method further includes: Output the target review results to the user; Obtain feedback data from the user regarding the target review results; Based on the feedback data and the target review results, the pre-stored knowledge base and the preset large language model are improved.

8. A code review device, characterized in that, include: The acquisition module is used to acquire the code to be reviewed, which includes: code content and review instructions; The processing module is used to retrieve the code content from a pre-stored knowledge base to obtain target context information; The processing module is further configured to obtain target input content based on the code content and the target context information; The processing module is also used to input the target input content and the review instructions into a preset large language model to obtain the target review result.

9. An electronic device, characterized in that, include: Memory containing executable program code; and the processor coupled to the memory; The processor invokes the executable program code stored in the memory to execute the code review method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, include: The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the code review method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Alarm log analysis method, electronic equipment, storage medium and program product

    CN121814393A

  • Alarm log analysis method, electronic device, storage medium and program product

    CN121814393B