A method and system for intelligent code review

By introducing an intelligent code review system into the code management tool, the code review process is automated, and the problems of low code review efficiency and quality are solved, achieving more efficient and accurate code review.

CN119848881BActive Publication Date: 2025-05-30CHANGJIANG SECURITIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510343191.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-05-30
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

In the software development process, code review requires a comprehensive assessment of logic, structure, performance and security, which leads to a huge pressure on the technical leader to invest in time and energy, especially in scenarios where high frequency submissions or high module complexity, and there may be a risk of subjectivity or omission.

Method used

Provides a method and system for intelligent code review. Through the binding trigger mechanism, it automatically calls the backend interface of the intelligent code review system in the code management tool, obtains and processes the submitted code, and splits it into a single submitted code for review. Combining intelligent scoring and automatic generation of review opinions, it reduces the subjectivity and omissions of manual review.

Benefits of technology

It significantly improves the efficiency and quality of code reviews, reduces the work burden of technical leaders, improves the accuracy and comprehensiveness of reviews, and reduces the risk of omissions of potential problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119848881B_ABST
    Figure CN119848881B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for intelligent code review, which relates to the technical field of program testing. The method includes: binding a trigger mechanism based on a project that needs code review in a code management tool; obtaining the submitted local code, and automatically invoking a background interface according to the trigger mechanism, where the local code includes incremental code or full code submitted by the developer; if it is determined that the type of the local code is incremental code, filtering the local code, and writing the submission information corresponding to the target code that meets the filtering conditions into a cache to wait for a scheduled task to execute a review task; when executing the scheduled task, taking out the submission information from the cache, and splitting the target code into individual submission codes according to the submission information; scoring each submission code to obtain a code score, and giving a review opinion for the submission code; statistically calculating the average score of multiple code scores to obtain a comprehensive score corresponding to the target code. The present application can improve the efficiency of code review.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of program testing, and particularly to a method and system for intelligent code review. Background Art

[0002] In the current software development process, after developers write code according to the requirements document and complete the function implementation, the technical leader usually needs to conduct a quality review of the submitted code. This review process aims to ensure that the code complies with the established coding specifications and design principles of the team and eliminates potential security hazards. However, since code review requires a comprehensive evaluation of logic, structure, performance, and security, especially in scenarios with high-frequency submissions or high module complexity, the technical leader often faces huge pressure in terms of time and energy investment, which may affect the overall development efficiency.

[0003] In addition, when the technical leader is not familiar enough with certain modules, it may be difficult to deeply evaluate the business logic or key technical implementation of the code, resulting in a certain degree of subjectivity or omission risk in the judgment results. In this case, code review may not only reduce the reliability of overall quality control but also affect the development progress. Therefore, how to assist the code review process through intelligent tools and automation technologies, reduce the workload of the technical leader, and improve the review efficiency at the same time has become an important topic in the optimization of modern software development processes. Summary of the Invention

[0004] This application provides a method and system for intelligent code review, which can improve the efficiency of code review.

[0005] In the first aspect of this application, a method for intelligent code review is provided. The method includes:

[0006] Based on the project binding trigger mechanism for code review required by developers in the code management tool, so that after the code management tool receives the code pushed to the specified branch by the user, it can automatically call the background interface of the intelligent code review system;

[0007] Obtain the submitted local code, and automatically call the background interface according to the trigger mechanism, where the local code includes the incremental code or full-scale code submitted by the developer;

[0008] If it is determined that the type of the local code is the incremental code, filter the local code, and write the submission information corresponding to the target code that meets the filtering conditions into the cache to wait for the scheduled task to execute the review task;

[0009] When the scheduled task is executed, take out the submission information from the cache, and split the target code into individual submission codes according to the submission information;

[0010] Score each of the submitted codes to obtain a code score, and give a review opinion on the submitted code;

[0011] Statistically calculate the average score of multiple code scores to obtain the comprehensive score corresponding to the target code.

[0012] Based on the above technical solution, preferably, after obtaining the locally submitted code and automatically calling the background interface according to the trigger mechanism, the method further includes:

[0013] If it is determined that the type of the local code is a full-volume code, obtain the source file of the local code;

[0014] Filter out the target files that need to be scanned for security from the source file;

[0015] Parse the syntax tree of the target file and recursively obtain the target attributes in the target file;

[0016] Check whether the fully qualified name of the target attribute contains a preset suffix. If it is determined that the fully qualified name contains the preset suffix, perform a full-volume scan on the target file;

[0017] Based on the results of the full-volume scan, perform a method-level risk review on the target code;

[0018] Output the review result according to the review result of the method-level risk review on the target code. The review result includes having risks or no risks.

[0019] Based on the above technical solution, preferably, after taking out the submission information from the cache when executing the scheduled task and splitting the target code into individual submitted codes, the method further includes:

[0020] Use the developer's submission ID to call relevant interfaces from the code management tool to obtain the modified content for the target code;

[0021] Traverse the modified content to find multiple modified files for each of the submitted codes;

[0022] Filter out the modified files that do not require code review by determining whether the file type of the modified file is a file type that requires code review, to obtain differential files;

[0023] Filter out invalid characters in the differential files to obtain filtered files;

[0024] Score the filtered files and give a review opinion, and output the score and review opinion for the filtered files in Markdown format.

[0025] Based on the above technical solutions, preferably, after scoring the filtering file and giving a review opinion, and outputting the score and review opinion for the filtering file in Markdown format, the method further includes:

[0026] Sort the file scores obtained by scoring each of the filtering files in descending order of score;

[0027] Determine a preset number of the lowest-scoring file scores according to the sorting result;

[0028] Re-review the filtering files corresponding to the preset number of the lowest-scoring file scores to obtain a secondary review opinion;

[0029] Output the secondary review opinion.

[0030] Based on the above technical solutions, preferably, before scoring each of the submitted codes to obtain a code score and giving a review opinion for the submitted code, the method further includes:

[0031] Collect the open-source codes in the developer's enterprise internal code library;

[0032] Preprocess the open-source codes to obtain processed codes;

[0033] Perform length segmentation, string normalization, and index construction on the processed codes in sequence to obtain multiple code data blocks;

[0034] Generate corresponding vector embeddings for each of the code data blocks to obtain a first vector embedding;

[0035] Store the first vector embedding and the code data block together in a vector data block.

[0036] Based on the above technical solutions, preferably, the scoring of each of the submitted codes to obtain a code score and giving a review opinion for the submitted code specifically further includes:

[0037] Generate a corresponding vector embedding for the submitted code to obtain a second vector embedding;

[0038] Perform a similarity search in the vector database through the vector embedding, and calculate the cosine similarity between the second vector embedding and each of the first vector embeddings;

[0039] Query multiple cosine similarities that exceed a preset similarity, and determine the similar vector embeddings corresponding to the multiple cosine similarities that exceed the preset similarity;

[0040] Determine the similar data blocks corresponding to the similar vector embeddings;

[0041] Use the similar data blocks as context information and the submitted code as input to generate review comments;

[0042] Provide background information related to the submitted code in the context information, and combine the private domain knowledge of the developer's enterprise to perform semantic understanding and problem identification on the content of the submitted code, including identifying code structure, readability, and naming conventions, and giving optimization suggestions to obtain the review comments.

[0043] Based on the above technical solutions, preferably, scoring each of the submitted codes to obtain a code score and giving review comments for the submitted code specifically includes:

[0044] Determine multiple review levels for the preset code review, including function level, file level, and system level;

[0045] Determine the weight values of each of the similar data blocks for each of the review levels;

[0046] According to the multiple weight values of each review level, calculate the average value to obtain the dynamic weight value of the review level;

[0047] For the submitted code, generate corresponding level scores according to the review criteria corresponding to the review levels;

[0048] Calculate the code score based on the dynamic weight values and level scores corresponding to each review level.

[0049] In the second aspect of the present application, a system for intelligent code review is provided. The system includes an acquisition module, a processing module, and an output module, where:

[0050] The acquisition module is used to trigger a mechanism based on the project binding that needs to be reviewed in the code management tool by the developer, so that after the code management tool receives the code pushed to the specified branch by the user, it can automatically call the background interface of the intelligent code review system;

[0051] The acquisition module is used to obtain the submitted local code and automatically call the background interface according to the trigger mechanism, where the local code includes the incremental code or full-scale code submitted by the developer;

[0052] The processing module is used to filter the local code if it is determined that the type of the local code is the incremental code, and write the submission information corresponding to the target code that meets the filtering conditions into the cache to wait for the review task to be executed by the scheduled task;

[0053] The processing module is used to retrieve the submission information from the cache when executing the timing task, and split the target code into individual submission codes according to the submission information;

[0054] The processing module is used to score each of the submission codes to obtain a code score, and give a review opinion on the submission code;

[0055] The output module is used to calculate the average score of multiple code scores to obtain the comprehensive score corresponding to the target code.

[0056] Based on the above technical solutions, preferably, the acquisition module is used to obtain the source file of the local code if it is determined that the type of the local code is a full-scale code;

[0057] The processing module is used to filter out the target files that need to be scanned for security from the source file;

[0058] The acquisition module is used to parse the syntax tree of the target file and recursively obtain the target attributes in the target file;

[0059] The acquisition module is used to check whether the fully qualified name of the target attribute contains a preset suffix. If it is determined that the fully qualified name contains the preset suffix, the target file is scanned in full;

[0060] The processing module is used to perform method-level risk review on the target code based on the results of the full-scale scan;

[0061] The output module is used to output the review result according to the review result of the method-level risk review of the target code, and the review result includes having risks or no risks.

[0062] Based on the above technical solutions, preferably, the processing module is used to call relevant interfaces from the code management tool using the developer's submission ID to obtain the modified content for the target code;

[0063] The processing module is used to traverse the modified content to find multiple modified files for each of the submission codes;

[0064] The processing module is used to filter out the modified files that do not require code review by determining whether the file type of the modified file is a file type that requires code review, and obtain the differential files;

[0065] The processing module is used to filter out invalid characters in the differential files to obtain filtered files;

[0066] The processing module is used to score the filtering file and give a review opinion, and output the score and review opinion for the filtering file in Markdown format.

[0067] Based on the above technical solution, preferably, the processing module is used to sort the file scores obtained by scoring each of the filtering files in descending order of scores;

[0068] The processing module is used to determine a preset number of the lowest-scoring file scores according to the sorting result;

[0069] The processing module is used to re-review the filtering files corresponding to the preset number of the lowest-scoring file scores to obtain a secondary review opinion;

[0070] The output module is used to output the secondary review opinion.

[0071] Based on the above technical solution, preferably, the acquisition module is used to collect the open-source codes in the enterprise internal code library of the developer;

[0072] The processing module is used to preprocess the open-source codes to obtain processed codes;

[0073] The processing module is used to perform length segmentation, string normalization, and index construction on the processed codes in sequence to obtain a plurality of code data blocks;

[0074] The processing module is used to generate corresponding vector embeddings for each of the code data blocks to obtain a first vector embedding;

[0075] The processing module is used to store the first vector embedding and the code data block together into a vector data block.

[0076] Based on the above technical solution, preferably, the processing module is used to generate a corresponding vector embedding for the submitted code to obtain a second vector embedding;

[0077] The processing module is used to perform a similarity search in the vector database through the vector embedding, and calculate the cosine similarity between the second vector embedding and each of the first vector embeddings;

[0078] The processing module is used to query a plurality of cosine similarities that exceed a preset similarity, and determine the similar vector embeddings corresponding to the plurality of cosine similarities that exceed the preset similarity;

[0079] The processing module is used to determine the similar data blocks corresponding to the similar vector embeddings;

[0080] The processing module is configured to use the similar data blocks as context information and the submitted code as input to generate review opinions.

[0081] The processing module is configured to provide background information related to the submitted code in the context information, and combine the private domain knowledge of the developer's enterprise to perform semantic understanding and problem identification on the content of the submitted code, including identifying code structure, readability, and naming conventions, and giving optimization suggestions to obtain the review opinions.

[0082] Based on the above technical solutions, preferably, the processing module is configured to determine multiple review levels for the preset code review, including function level, file level, and system level.

[0083] The processing module is configured to determine the weight values of each of the similar data blocks for each of the review levels.

[0084] The processing module is configured to calculate the average value of the multiple weight values of each review level to obtain the dynamic weight value of the review level.

[0085] The processing module is configured to generate corresponding level scores for the submitted code according to the review criteria corresponding to the review levels.

[0086] The processing module is configured to calculate the code score based on the dynamic weight values and level scores corresponding to each review level.

[0087] In a third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method described in any one of the above.

[0088] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method described in any one of the above is executed.

[0089] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0090] 1. Through the binding trigger mechanism, when developers submit code in the code management tool, the background interface can be automatically called. At the same time, the classification processing of incremental code and full - scale code, as well as the centralized processing of submission information using scheduled tasks, break down complex code review tasks into fine - grained individual code submissions for step - by - step analysis, which helps to improve the accuracy of review. In addition, through intelligent scoring and automatic generation of review opinions, combined with the comprehensive statistics of multiple scores, fast and standardized review results are provided, avoiding the subjectivity and omission problems of manual review, thus significantly improving the efficiency and quality of code review.

[0091] 2. By automatically processing the modified content of the submitted code, screening out the files that need to be reviewed and filtering invalid characters, the review content is simplified and the efficiency is improved. Through intelligent scoring and standardized Markdown - format output of review opinions, clear and structured feedback can be quickly generated, helping developers more intuitively understand the review results.

[0092] 3. By conducting a secondary review of low - scoring files, further focusing on high - risk or complex problems that may exist in the code, ensuring that key issues are analyzed more deeply and optimization suggestions are provided. It effectively improves the comprehensiveness and accuracy of code review, avoids missing potential problems, and at the same time optimizes resource utilization, focusing on the files that most need attention, thereby improving code quality and team collaboration efficiency.

[0093] 4. By constructing a vectorization model and similarity matching mechanism for the enterprise - internal code library, associating the submitted code with the context information of historical code and open - source code, providing accurate semantic background support. Combining private domain knowledge, it can more accurately identify problems such as structural issues, lack of readability, and non - standard naming in the code, and give targeted optimization suggestions. It effectively improves the intelligence and accuracy of code review, enhances the adaptability to complex code scenarios, and at the same time reduces the subjectivity and omission risk of review, thus significantly improving code quality and development efficiency.

[0094] 5. By introducing multi - level review dimensions at the function level, file level, and system level, and adjusting the importance of each level according to dynamic weight values, ensuring that the review results can comprehensively reflect the quality of the local details and overall architecture of the code. Through the hierarchical scoring mechanism, a more comprehensive and accurate code score can be generated, and at the same time, the review process becomes more targeted and adaptable. It effectively improves the scientificity and objectivity of code review, provides more guiding optimization suggestions for developers, and thus promotes the overall improvement of code quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] Figure 1 is a schematic flowchart of a method for intelligent code review disclosed in an embodiment of the present application;

[0096] Figure 2 It is a schematic diagram of the modules of a system for intelligent code review disclosed in an embodiment of the present application;

[0097] Figure 3 It is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present application.

[0098] Explanation of reference numerals: 201, acquisition module; 202, processing module; 203, output module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed implementation manners

[0099] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0100] In the description of the embodiments of the present application, words such as "for example" or "for illustration" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for illustration" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "for example" or "for illustration" is intended to present relevant concepts in a specific manner.

[0101] In the description of the embodiments of the present application, the meaning of the term "a plurality" refers to two or more. For example, a plurality of systems refers to two or more systems, and a plurality of screen terminals refers to two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0102] In the current software development process, the technical leader usually needs to conduct a quality review of the code submitted by the developers to ensure that the code complies with the coding specifications, design principles, and eliminate potential security hazards. However, due to the high-frequency submissions or the existence of complex modules, the technical leader faces huge time and energy pressures. Especially when not familiar with certain modules, the review may be subjective or there may be risks of omission, thus affecting the code quality and development progress. Therefore, how to use intelligent tools and automation technologies to assist code review, reduce the burden on the technical leader and improve the review efficiency has become a key topic in the optimization of modern software development processes.

[0103] This embodiment discloses a method for intelligent code review. Referring to Figure 1 , it includes the following steps S110 - S160:

[0104] S110, based on the triggering mechanism bound to the project that needs code review by the developer in the code management tool, so that after the code management tool receives the code pushed to the specified branch by the user, it can automatically call the background interface of the intelligent code review system.

[0105] The method for intelligent code review disclosed in the embodiments of the present application is applied to a server. The server includes, but is not limited to, electronic devices such as mobile phones, tablet computers, wearable devices, PCs (Personal Computers), etc., and can also be a background server running a method for intelligent code review. The server can be implemented by an independent server or a server cluster composed of multiple servers.

[0106] With the increasing attention of IT Internet and fintech enterprises to the DevOps integrated development and operation and maintenance development model, continuously improving software quality has become an important factor in the software delivery process. Based on the requirements of system security and code quality, we hope that only after passing the code review can the modified content be merged into the main branch for online deployment after the R & D personnel submit the code. In terms of the code review method, the current review methods mainly include automatic code review and manual review. Manual code review and automated code review each have their own advantages and limitations. Manual code review can deeply understand the code logic and put forward constructive opinions. Especially when it comes to complex business logics or designs that require high creativity, manual review can better play its advantages. However, manual review takes a long time and may result in inconsistent review criteria due to individual experience differences. In contrast, automated code review such as SonarQube can complete basic checks of a large amount of code in a short time, such as format consistency, potential bug detection, etc., greatly improving the efficiency. Tools like SonarQube can also run in a continuous integration environment and provide feedback immediately before the code is submitted, helping to quickly correct problems and ensure that the code meets the established standards. However, automated tools may be difficult to identify deeper design problems or complex logical errors. They usually rely on preset rules and are not as flexible as manual review in dealing with some special scenarios.

[0107] To this end, a method for intelligent code review disclosed in the embodiments of the present application aims to assist the code review link in the R & D process through intelligent means, help the technical person in charge complete the code review work more efficiently, and improve the quality of code delivery. Through this method, the review results can be automatically fed back to the DevOps platform as the basis for quality control of task completion. Specifically, when a developer submits code, the server will automatically conduct a comprehensive review of it, promptly discover and point out potential problems so that the developer can quickly make corrections. When the technical person in charge conducts the final code merge review, they can refer to the opinions provided by the intelligent review to further enhance the accuracy and efficiency of the review. In addition, testers can also carry out subsequent testing work more targeted based on these review results, thus jointly promoting the improvement of the overall project quality.

[0108] According to code management tools used by enterprises such as GitLab, configure WebHook in the code management tool to bind it to the background interface of the intelligent code review system. The developer selects the project that needs code review in the code management tool and designates the target branch as the trigger condition. Whenever the user pushes code to the designated branch, WebHook will capture this event and trigger the pre-bound review system interface. After receiving the trigger request, the review system parses the event type to ensure that only code submission requests of the Commit type are processed. Subsequently, the review system writes the submission information into the cache queue and waits for the scheduling task to execute the review process. Through this automated binding mechanism, seamless docking between code submission and the review system can be efficiently achieved, ensuring the timeliness of code review and improving development efficiency and code quality.

[0109] S120, Obtain the local code developed and submitted by the developer, and automatically call the background interface according to the trigger mechanism.

[0110] After the developer completes the local code development, use a version control tool such as Git to submit the code to the designated branch, and include necessary annotations such as the task ID in the submission information. The WebHook configuration in the code management platform will detect the push event and trigger the bound background interface of the intelligent code review system. After receiving the trigger request, the background interface extracts information such as the submitted branch, submission ID, and modified content from the event data. Subsequently, call the API interface provided by the code management platform to obtain the specific code change files and content. Through this automated process, a full-link trigger mechanism from the developer's code submission to the background interface call can be realized, without manual intervention, greatly improving the response efficiency and accuracy of code review.

[0111] Further, in addition to performing code quality scanning on the added code, security scanning is also required for the existing code. The scanning results are used as a quality gate for online upgrades to improve the overall code quality. The scanning system mainly includes two parts of logic: parsing whether there is a method for manually operating the database connection in the class; using the large language model ability to scan for database leakage problems at the method level for the methods with manual database connection operations.

[0112] In a possible implementation, after obtaining the submitted local code and automatically calling the background interface according to the trigger mechanism, the method further includes: if it is determined that the type of the local code is full-scale code, obtaining the source file of the local code; filtering out the target files that need to be scanned for security from the source file; parsing the syntax tree of the target file and recursively obtaining the target attributes in the target file; checking whether the fully qualified name of the target attribute contains a preset suffix, and if it is determined that the fully qualified name contains the preset suffix, performing a full-scale scan on the target file; based on the results of the full-scale scan, performing a method-level risk review on the target code; according to the review results of the method-level risk review on the target code, outputting the review results, and the review results include having risks or no risks.

[0113] Specifically, after obtaining the source file of the code and filtering out the files that need to be scanned for security, pulling the repository code to the server temporary directory according to the system repository bound by the user in the DevOps development and operation integration platform, and then defining the filtering rules according to the internal R & D framework of the company. Specifically, if the path of the class file contains the prefix "IF-" or the suffix "client", such files will not be scanned.

[0114] Parse the syntax tree (abstract syntax code, AST) of the source file to recursively obtain all internal classes, member variables, imported dependency packages, methods and other attributes of the file. Obtaining the information of each class is convenient for subsequent checking whether there is an action of manually operating the database connection.

[0115] Check whether there is a manual database connection operation in the class. If so, extract the method information of the specific method for operating the database for resource leakage handling. The specific checking logic is as follows: check whether the package suffix introduced by the class contains fields such as "Connection", "PreparedStatement", "ResultSet". If it contains, check whether there are attributes ending with the fields "Connection", "PreparedStatement", "ResultSet" in the fully qualified name of the attribute type of the class. If both of these conditions are met, it can be considered that there is a manual database connection operation in this class and a full-scale scan is required.

[0116] Use the method body of the method as the backend input to check for resource leakage risks. Usually, there are size limits for the context length of the model. To avoid the impact of excessive context information on code review, the review is conducted at the method level. After detecting that the fully qualified name of the attribute type has a suffix of "Connection", "PreparedStatement", or "ResultSet", the method body using this attribute is parsed as the review input for the large language model.

[0117] After the backend service receives the method body that may have a database connection leak input by the full-code review service, it executes the code review task. The format of the returned result includes two specific formats: At risk: {"fields": "variable name", "checkResult": "check failed: specific problem description of `variable name`"}. No risk: {"fields": "variable name", "checkResult": "check passed: variable not found or no problem"}.

[0118] Write the result to the integrated R & D and operation and maintenance platform. After receiving the result returned by the backend, the full-code review service parses the returned result content, associates the code review situation with the specific system in the DevOps integrated R & D and operation and maintenance platform, and at the same time serves as a gating metric during the R & D process. Send the scan result to the task submitter. After the full-code scan is completed, the platform will send an email to notify the specific situation of this scan to facilitate R & D personnel to process the scan result in a timely manner.

[0119] S130, if it is determined that the type of the local code is the incremental code, filter the local code, and write the commit information corresponding to the target code that meets the filtering conditions into the cache to wait for the scheduled task to execute the review task.

[0120] After receiving the commit information, the server filters out non-push commit methods according to whether the event type is the Commit type, and then writes the commit information that meets the filtering conditions into the cache to wait for the scheduled task to execute the review work.

[0121] In a possible implementation, filtering the local code and writing the commit information corresponding to the target code that meets the filtering conditions into the cache to wait for the scheduled task to execute the review task specifically includes: according to the event type of the local code, determining whether the event type is a code commit type; if it is determined that the event type is not a code commit type and the event type is a tag creation type or a branch merge type, then filter the local code; if it is determined that the event type is a code commit type, then execute the review of the local code.

[0122] Specifically, after receiving a trigger request, parse the event type field such as event_type or object_kind. If the event type is not a code submission type, such as tag creation or branch merge, further determine the specific event type. If it is a code submission type, i.e., a push event, enter the code review process. Otherwise, the server will enter the filtering and processing stage to perform specific processing on non-code submission events.

[0123] For the detected non-code submission events, the server filters them according to the event type. If the event type is tag creation, the server extracts tag information such as the tag name, associated commits, and records them, but does not trigger a code review. If the event type is branch merge, the server parses the commit history of the source branch and the target branch being merged to update the code repository status, but also skips the review process. This step ensures that non-code submission events do not trigger unnecessary operations, thereby improving the code review efficiency of the server.

[0124] S140, when executing a scheduled task, retrieve the submission information from the cache and split the target code into individual submission codes according to the submission information.

[0125] When executing a scheduled task, the server retrieves the stored submission information from the cache, where each piece of information contains metadata of the code submission, such as the submission ID, branch name, and author, etc. Next, the server calls the relevant API interfaces of the code management tool according to the submission ID to obtain the detailed modification content of the corresponding submission. By parsing the submission content, the server splits the target code into individual submissions (i.e., individual Commits), and each submission may contain modification records of multiple files. For these modification records, the server analyzes the file types one by one to determine whether they need to be code-reviewed, filtering out file types that do not require review (such as non-code files or configuration files). Finally, the server sorts out the filtered files and their modification content in the dimension of individual Commits, providing data input for subsequent code scoring and review. This process ensures that the review task focuses on valid code, improving the review efficiency and accuracy.

[0126] In a possible implementation, after retrieving the submission information from the cache and splitting the target code into individual submission codes when executing a scheduled task, the method further includes: using the developer's submission ID to call relevant interfaces in the code management tool to obtain the modification content for the target code; traversing the modification content to find multiple modified files for each submission code; filtering out the modified files that do not require code review by determining whether the file type of the modified file is a file type that needs to be code-reviewed, obtaining the differential files; filtering out invalid characters in the differential files to obtain the filtered files; scoring the filtered files and giving review comments, and outputting the score and review comments for the filtered files in Markdown format.

[0127] Specifically, the server uses the commit ID in the submission information to call the API interface of a code management tool (such as GitLab) to obtain the change record corresponding to the commit ID. The information returned by the API usually includes the detailed changes in the commit, such as the list of modified files, the specific code snippets of the changes (i.e., the diff files), author information, and the timestamp of the commit. This data provides a complete context for subsequent code reviews.

[0128] From the change records returned by the API interface, the server traverses all the modified files included in the submitted code one by one. The information for each file includes the file path, file name, file type (such as file extension), and its changes. By analyzing each one, the server can determine which files are valid code files, thus preparing for the review process.

[0129] The server uses a predefined whitelist of file types, such as.java and.py, which are the code file types supported for review, to filter out files that do not meet the criteria. For example, configuration files such as.yaml, documentation files such as.md, and binary files such as.png will be excluded. The files remaining after filtering form a list of diff files, ensuring that the review is only targeted at the code files that actually need attention.

[0130] The content in the diff files may contain a large number of invalid characters used to identify changes, such as the plus and minus signs generated by Git tools. These characters are not helpful for the analysis of large language models and may even increase unnecessary context complexity. The server cleans these invalid characters through regular expressions or string processing methods, retaining only the actual changes in the code to form a more concise filtered file.

[0131] The server sends each filtered file to the large language model for analysis together with predefined prompts. The prompts usually include review criteria such as code style, performance optimization, security, etc., and output format requirements. After analyzing the filtered files, the large language model generates a score and review comments according to the requirements of the prompts. The content includes: code score, potential problems, suggested improvement measures, etc., and organizes the results in Markdown format for easy subsequent display and integration. The Markdown format output helps to structure the display of the review results, such as hierarchical headings, code blocks, and tables, improving readability.

[0132] S150, Score each submitted code to obtain a code score and give review comments on the submitted code.

[0133] In a possible implementation, before scoring each submitted code to obtain a code score and giving a review opinion on the submitted code, the method further includes: collecting open-source codes from the enterprise internal code library of developers; preprocessing the open-source codes to obtain processed codes; sequentially performing length segmentation, string normalization, and index construction on the processed codes to obtain a plurality of code data blocks; generating corresponding vector embeddings for each code data block to obtain first vector embeddings; and storing the first vector embeddings and the code data blocks together in a vector data block.

[0134] Specifically, in an enterprise scenario, there is a large amount of private data within the enterprise, such as internal development documents, middleware SDKs, API documents, etc. The code review model needs to learn and adapt to this type of private data. Traditional model training methods require preparing an expensive GPU machine cluster, with a high training cost and being unable to take effect in real time, unable to meet the need for rapid knowledge update.

[0135] To solve this problem, we use the RAG enhancement method to optimize the large language model for code review. The code review model obtains private knowledge in the retrieval engine in real time through retrieval enhancement technology, enabling the code review model to generate codes or answer questions based on enterprise knowledge. And the enterprise can update and take effect the private knowledge immediately, so that the code review model can always generate or answer questions based on the latest private knowledge. At the same time, the code review model improves the recall rate of knowledge data through technologies such as multi-dimensional recall and refined knowledge segmentation, and selects the most relevant knowledge to introduce into the context of generation / question answering through ranking strategies and relevance screening strategies, further enhancing the practicality and accuracy of the system.

[0136] The system extracts open-source codes from the enterprise internal code library, including codes independently developed and open-sourced by the enterprise and open-source codes obtained from third parties. These codes may be included in a public code management platform. When extracting, it is necessary to ensure the copyright compliance of the codes and screen out the codes related to the current business field to lay a foundation for subsequent analysis.

[0137] The steps for preprocessing the open-source codes are aimed at cleaning and standardizing the code content. Code cleaning mainly cleans the collected open-source codes and the internal code library of the company. The specific cleaning rules include: filtering files with more than 1000 lines, filtering data with functions less than 50 lines, filtering data with a letter ratio of less than 25%, filtering files with a comment ratio greater than 80%, filtering decompiled data, removing sensitive information in the code, filtering files with a large number of comment lines, etc.

[0138] The AST extraction method uses abstract syntax parsing to parse the methods in the class file and constructs a method-level code database by parsing the class unit by unit. The method body data parsed from the abstract syntax tree is written into the method-level code database.

[0139] The processing code is further processed for length segmentation, string normalization, and index construction. Since the code data is too long in its original state to fit into the context window of the model, it needs to be divided into smaller segments for indexing and then placed into the context window of the large model. Specifically, the CharacterTextSplitter, a character-based text splitter, is used to split the code segments in the code file. The specific parameters are: chunksize = 1024 and chunkoverlap = 128. Then, the split code segments are subjected to string normalization, which specifically includes removing extra spaces at both ends of the string using regular expressions, converting all strings to lowercase, and removing special characters in the code segments. Finally, the normalized code segments are used to construct the index. For the storage of specific code data blocks, since cross-text block semantic search needs to be supported, vector embeddings need to be generated for each code data block and then stored together with their vector embeddings. The method used to generate vector embeddings for each code data block is a general text vector model, and the generated vector embeddings are stored in the vector database Weaviate.

[0140] In a possible implementation, each submitted code is scored to obtain a code score, and review comments are given for the submitted code. Specifically, it further includes: for the submitted code, generating a corresponding vector embedding to obtain a second vector embedding; performing a similarity search in the vector database through the vector embedding to calculate the cosine similarity between the second vector embedding and each first vector embedding; querying multiple cosine similarities that exceed the preset similarity to determine the similar vector embeddings corresponding to the multiple cosine similarities that exceed the preset similarity; determining the similar data blocks corresponding to the similar vector embeddings; using the similar data blocks as context information and the submitted code as input to generate review comments; providing background information related to the submitted code in the context information, combining the private domain knowledge of the developer's enterprise, performing semantic understanding and problem identification on the content of the submitted code, including identifying code structure, readability, and naming conventions, and giving optimization suggestions to obtain review comments.

[0141] Specifically, for the submitted code, the same general text vector model as in the preprocessing stage is used to generate the vector embedding of the code. This embedding represents the syntactic structure and semantic features of the submitted code and is used to locate other code segments semantically similar to it in the vector space.

[0142] A similarity search is performed on the second vector embedding of the submitted code to calculate the cosine similarity between it and each first vector embedding stored in the vector database. The cosine similarity reflects the similarity degree between two code segments in the semantic space, and the closer the value is to 1, the more similar the code segments are.

[0143] Filter out the items from the calculated cosine similarities that exceed a preset similarity threshold, such as a threshold of 0.5. These high-similarity items represent the reference code snippets in the vector database that are most relevant to the submitted code. Based on the filtered high-similarity items, determine the corresponding first vector embeddings. These similar vector embeddings store rich context information in the database and can serve as an important reference basis for the current code review.

[0144] Further query the code data blocks associated with the high-similarity vector embeddings. These data blocks are extracted from the enterprise code repository or open-source code and have a high degree of relevance to the submitted code in terms of function, structure, or semantics.

[0145] Provide the filtered similar data blocks as context information to the code review model, and at the same time use the submitted code as the main input. The model combines the context information and the submitted code for comprehensive analysis and generates review opinions for the current code according to the implementation logic and best practices of the existing code.

[0146] Extract the background information directly related to the function or module of the submitted code in the context information, such as algorithm ideas, design goals, or performance optimization measures, and combine enterprise private knowledge, such as coding specifications and business scenarios, to perform semantic understanding and problem identification on the content of the submitted code. This process focuses on evaluating the structural rationality, readability, and naming conventions of the code to help discover potential problems or areas for improvement. Based on the results of the model analysis, combined with the context information and private knowledge, give specific optimization suggestions. This includes adjusting the code logic, improving variable naming, and enhancing code efficiency. Finally, the model outputs the score for the submitted code and detailed review opinions in Markdown format, providing developers with a clear and definite direction for improvement.

[0147] Specifically, use a language model to perform semantic parsing on the submitted code to identify the logical structure of the code, such as loops, conditional judgments, function calls, etc., and potential problems, such as the risk of infinite loops and duplicate code. Combine the context information to determine whether the code conforms to the team standards. For example, check whether it meets the naming specifications, whether the variable names are descriptive, or whether the code is modular enough to improve maintainability.

[0148] For the discovered problems, provide optimization suggestions by combining the context information and private knowledge. For example, modify unclear or verbose variable names, optimize code segments with low performance, and provide implementation logic suggestions that are more in line with the business intent. Finally, generate a clear Markdown-format review report, including code scores, problem lists, specific suggestions, etc.

[0149] In a possible implementation, each submitted code is scored to obtain a code score, and review comments are given for the submitted code, specifically including: determining multiple review levels for preset code review, including function level, file level, and system level; determining the weight values of each similar data block for each review level; obtaining the dynamic weight value of the review level by calculating the average value according to the multiple weight values of each review level; for the submitted code, generating the corresponding level score according to the review criteria corresponding to the review level; calculating the code score based on the dynamic weight value and the level score corresponding to each review level.

[0150] Specifically, when conducting code review, it is first necessary to define multiple levels for review in order to refine and precision the review criteria layer by layer. The review levels disclosed in the embodiments of the present application include: function-level review, which evaluates the quality of each function or method, and focuses on the readability, functionality, and efficiency of the code. File-level review, which reviews the overall structure and specifications of all the code within a file, including the organization and modularity of the code. System-level review, which evaluates the code quality of the entire system or project, and checks the overall architecture design, maintainability, scalability, etc. of the system.

[0151] In some projects, these three review levels of function level, file level, and system level may be the most important and necessary review dimensions, and other levels such as module level and component level are not used as the priority review levels. A typical scenario is when there are relatively independent functional modules in the project, and the development of these modules is relatively independent, and the overall system architecture is relatively stable. In this case, the review focus usually concentrates on the logical correctness of a single function (function level), the standardization and maintainability of the code within a single file (file level), and the structural rationality and cooperation between modules of the entire system (system level). In particular, for the development of large-scale Web applications, developers mainly focus on the implementation of a single interface, the organization and structure of the interface files, and the architecture design of the entire application system, rather than paying too much attention to the review of finer-grained modules or components. In this case, the function-level, file-level, and system-level reviews are sufficient to cover the code quality assurance requirements, so no other additional review levels are set.

[0152] When conducting code review, the scoring criteria for each level are not the same, so it is necessary to allocate appropriate weight values to the review criteria for each level. Similar data blocks such as the complexity of functions and the code duplication rate will have different impacts in different review levels. For example, the function-level review can focus on the performance and logical clarity of the code, while the file-level review may pay more attention to the structure and specifications of the code. According to the influence degree of the data block, a weight value is assigned to each level, so that the influence of different review dimensions on the final score is more accurate.

[0153] For each review level, a dynamic weight value can be calculated based on the specific weight value. This dynamic weight value is not statically fixed, but dynamically adjusted according to the specific code submission. For example, if a submission involves more function-level changes, the function-level weight can be appropriately increased. By calculating the weighted average of each weight value, the final dynamic weight value of each review level is obtained, thereby ensuring the flexibility and accuracy of the review.

[0154] During the review process, hierarchical scores are generated for the standards of each level. For example, for function-level reviews, you can check the execution efficiency, logical clarity, naming conventions, etc. of the function; for file-level reviews, you can check whether the code complies with the overall structural specifications of the project and whether the comments are sufficient; system-level reviews involve the rationality of the architecture design, the degree of coupling between modules, etc. Each review level has specific scoring criteria, and reviewers score the code based on these criteria.

[0155] Finally, the final code score is obtained by comprehensively calculating the score of each review level and its dynamic weight value. Specifically, code score = function-level score × function-level dynamic weight + file-level score × file-level dynamic weight + system-level score × system-level dynamic weight. In this way, the scores of different levels will have a corresponding impact on the final score according to their weights, thereby ensuring the comprehensiveness and accuracy of the review results.

[0156] S160, calculating the average scores of multiple code ratings to obtain a comprehensive score corresponding to the target code.

[0157] After scoring each file in a single submission, the server generates a score list, with each score corresponding to the modification score of the file. All file scores are traversed, added up and divided by the number of files to calculate the average score of all files. This score represents the overall performance of the entire submission in terms of modification quality. The calculated average score is used as the comprehensive score of the target code and recorded in the review report, with corresponding explanations and data support so that developers can understand the overall quality situation.

[0158] In a possible implementation, after scoring the filtered files and providing review opinions, and outputting the scores and review opinions for the filtered files in Markdown format, the method also includes: sorting the file scores obtained by scoring each filtered file in order of high and low scores; determining a preset number of file scores with the lowest scores based on the sorting results; re-reviewing the filtered files corresponding to the preset number of file scores with the lowest scores to obtain secondary review opinions; and outputting the secondary review opinions.

[0159] Specifically, collect the results of scoring all filtered files and sort these scores in descending order. Use a sorting algorithm (such as quicksort or a built-in sorting function) to arrange the score list while preserving the association information between the scores and the corresponding files. According to a preset quantity (such as the 3 files with the lowest scores), select multiple files with the lowest scores from the sorted score list and record the identification information of these files for further operations.

[0160] For the files with the lowest scores, re-call the large language model or other review modules for in-depth review. In the secondary review, combine context information, enterprise private domain knowledge, and problem types to further analyze possible problems in the files. The model outputs more detailed and specific improvement suggestions, such as code logic errors, potential security issues, or unreasonable designs. Output the results of the secondary review in Markdown format, providing improvement suggestions for low-scoring files and attaching specific example code if necessary. The generated content includes: the name of the secondary review file, problem description, optimization suggestions, and comprehensive improvement plan.

[0161] By adopting the technical solution disclosed in the embodiment of the present application, based on the bound trigger mechanism, when a developer submits code in the code management tool, the server can automatically call the background interface. At the same time, the server uses scheduled tasks to centrally process the submission information, decomposing complex code review tasks into fine-grained individual submitted codes for step-by-step analysis, which helps to improve the accuracy of the review. In addition, through intelligent scoring and automatic generation of review opinions, combined with the comprehensive statistics of multiple scores, the server provides fast and standardized review results, avoiding the subjectivity and omission problems of manual review, thus significantly improving the efficiency and quality of code review.

[0162] This embodiment also discloses a system for intelligent code review. Refer to Figure 2 The system includes an acquisition module 201, a processing module 202, and an output module 203, where:

[0163] The acquisition module 201 is used to bind a trigger mechanism based on the project for which the developer needs to conduct code review in the code management tool, so that after the code management tool receives the code pushed to the specified branch by the user, it can automatically call the background interface of the intelligent code review system.

[0164] The acquisition module 201 is used to obtain the submitted local code and automatically call the background interface according to the trigger mechanism, where the local code includes the incremental code or full-scale code submitted by the developer.

[0165] The processing module 202 is used to, if it is determined that the type of the local code is incremental code, filter the local code and write the submission information corresponding to the target code that meets the filtering conditions into the cache to wait for the scheduled task to execute the review task.

[0166] The processing module 202 is configured to retrieve the submission information from the cache when executing a timed task, and split the target code into individual submission codes according to the submission information.

[0167] The processing module 202 is configured to score each submission code to obtain a code score, and give a review opinion on the submission code.

[0168] The output module 203 is configured to calculate the average score of multiple code scores to obtain the comprehensive score corresponding to the target code.

[0169] In a possible implementation, the acquisition module 201 is configured to obtain the source file of the local code if it is determined that the type of the local code is a full-scale code.

[0170] The processing module 202 is configured to filter out the target files that need to be scanned for security from the source file.

[0171] The acquisition module 201 is configured to parse the syntax tree of the target file and recursively obtain the target attributes in the target file.

[0172] The acquisition module 201 is configured to check whether the fully qualified name of the target attribute contains a preset suffix. If it is determined that the fully qualified name contains the preset suffix, a full-scale scan is performed on the target file.

[0173] The processing module 202 is configured to perform a method-level risk review on the target code based on the results of the full-scale scan.

[0174] The output module 203 is configured to output a review result according to the review result of the method-level risk review on the target code. The review result includes that there is a risk or there is no risk.

[0175] In a possible implementation, the processing module 202 is configured to call a relevant interface from a code management tool using the developer's submission ID to obtain the modification content for the target code.

[0176] The processing module 202 is configured to traverse the modification content to find multiple modified files for each submission code.

[0177] The processing module 202 is configured to filter out the modified files that do not require code review by determining whether the file type of the modified file is a file type that requires code review, and obtain the differential files.

[0178] The processing module 202 is configured to filter out the invalid characters in the differential files to obtain the filtered files.

[0179] The processing module 202 is used to score the filtered files and give review opinions, and outputs the scores and review opinions for the filtered files in Markdown format.

[0180] In a possible implementation, the processing module 202 is used to sort the file scores obtained by scoring each filtered file in descending order of scores.

[0181] The processing module 202 is used to determine multiple file scores with the lowest preset number according to the sorting result.

[0182] The processing module 202 is used to re-review the filtered files corresponding to multiple file scores with the lowest preset number to obtain secondary review opinions.

[0183] The output module 203 is used to output the secondary review opinions.

[0184] In a possible implementation, the acquisition module 201 is used to collect open-source codes from the developer's enterprise internal code library.

[0185] The processing module 202 is used to preprocess the open-source codes to obtain processed codes.

[0186] The processing module 202 is used to perform length segmentation, string normalization, and index construction on the processed codes in sequence to obtain multiple code data blocks.

[0187] The processing module 202 is used to generate corresponding vector embeddings for each code data block to obtain the first vector embeddings.

[0188] The processing module 202 is used to store the first vector embeddings and the code data blocks together 305 into vector data blocks.

[0189] In a possible implementation, the processing module 202 is used to generate corresponding vector embeddings for the submitted codes to obtain second vector embeddings.

[0190] The processing module 202 is used to perform similarity search in the vector database through the vector embeddings, and calculate the cosine similarity between the second vector embeddings and each of the first vector embeddings.

[0191] The processing module 202 is used to query multiple cosine similarities that exceed the preset similarity, and determine the similar vector embeddings corresponding to the multiple cosine similarities that exceed the preset similarity.

[0192] The processing module 202 is used to determine the similar data blocks corresponding to the similar vector embeddings.

[0193] The processing module 202 is configured to use the similar data blocks as context information, take the submitted code as input, and generate review opinions.

[0194] The processing module 202 is configured to provide background information related to the submitted code in the context information, combine the private knowledge of the developer's enterprise, perform semantic understanding and problem recognition on the content of the submitted code, including identifying code structure, readability, and naming conventions, and give optimization suggestions to obtain review opinions.

[0195] In a possible implementation manner, the processing module 202 is configured to determine multiple review levels for the preset code review, including function level, file level, and system level.

[0196] The processing module 202 is configured to determine the weight values of each similar data block for each review level.

[0197] The processing module 202 is configured to calculate the average value of the multiple weight values of each review level to obtain the dynamic weight value of the review level.

[0198] The processing module 202 is configured to generate corresponding level scores for the submitted code according to the review criteria corresponding to the review levels.

[0199] The processing module 202 is configured to calculate the code score based on the dynamic weight values and level scores corresponding to each review level.

[0200] It should be noted that: when the system provided in the above embodiments implements its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be elaborated here.

[0201] This embodiment also discloses an electronic device. Refer to Figure 3 , the electronic device may include: at least one processor 301, at least one communication bus 302, a user interface 303, a network interface 304, and at least one memory 305.

[0202] Among them, the communication bus 302 is used to implement connection communication between these components.

[0203] Among them, the user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.

[0204] Among them, the network interface 304 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface).

[0205] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server using various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling the data stored in the memory 305, it performs various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one of the hardware forms of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 301 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.

[0206] Among them, the memory 305 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store the data involved in the above-mentioned various method embodiments. Optionally, the memory 305 may also be at least one storage device located far from the aforementioned processor 301. As a computer storage medium, the memory 305 may include an operating system, a network communication module, a user interface 303 module, and an application program for a method of intelligent code review.

[0207] In the electronic device shown in FIG. 3, the user interface 303 is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor 301 can be used to call the application program stored in the memory 305 that stores a method for intelligent code review. When executed by one or more processors 301, the electronic device is caused to execute the method of one or more of the above embodiments.

[0208] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be adopted in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0209] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0210] In the several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0211] The unit described as a separated component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0212] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0213] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. And the aforementioned memory 305 includes: various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0214] This application also discloses a computer-readable storage medium storing instructions. When executed by one or more processors 301, it causes the electronic device to execute the method as described in one or more of the above embodiments.

[0215] The foregoing are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and the disclosure of the practical truth, those skilled in the art will readily think of other implementation manners of the present disclosure. This application aims to cover any variations, uses, or adaptive changes of the present disclosure, which follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A method for intelligent code review, characterized in that: The method comprises: A project binding trigger mechanism based on the developer's need to conduct code review in the code management tool, so that the code management tool can automatically call the background interface of the intelligent code review system after receiving the code pushed by the user's specified branch; Obtaining the submitted local code, and automatically calling the background interface according to the trigger mechanism, wherein the local code includes the incremental code or the full code submitted by the developer; If it is determined that the type of the local code is the incremental code, the local code is filtered, and the submission information corresponding to the target code that meets the filtering condition is written into the cache to wait for the scheduled task to execute the review task; When executing the scheduled task, the submission information is retrieved from the cache, and the target code is split into individual submission codes according to the submission information; Scoring each submitted code to obtain a code score, and providing review opinions on the submitted code; Counting the average scores of the multiple code scores to obtain a comprehensive score corresponding to the target code; After obtaining the submitted local code and automatically calling the background interface according to the trigger mechanism, the method further includes: If it is determined that the type of the local code is full code, obtaining a source file of the local code; Filtering target files that need to be security scanned from the source files; Parsing the syntax tree of the target file, and recursively obtaining target attributes in the target file; Check whether the fully qualified name of the target attribute contains a preset suffix, and if it is determined that the fully qualified name contains the preset suffix, perform a full scan on the target file; Based on the result of the full scan, a method-level risk review is performed on the target code; According to the review result of the method-level risk review of the target code, the review result is output, and the review result includes whether there is a risk or there is no risk.

2. The method of intelligent code review according to claim 1, characterized in that: After taking out the submission information from the cache when executing the scheduled task and splitting the target code into individual submission codes according to the submission information, the method further includes: Using the developer's submission ID to call a related interface from the code management tool to obtain the modified content for the target code; Traversing the modified contents to search for multiple modified files for each of the submitted codes; By determining whether the file type of the modified file is a file type that requires code review, the modified files that do not require code review are filtered out to obtain a difference file; Filter invalid characters in the difference file to obtain a filtered file; The filter file is scored and review comments are given, and the score and review comments for the filter file are output in Markdown format.

3. The method of intelligent code review according to claim 2, characterized in that: After scoring the filter file and providing review opinions, and outputting the score and review opinions for the filter file in Markdown format, the method further includes: The file scores obtained by scoring each of the filtered files are sorted in order of high and low scores; According to the sorting results, determining scores for a preset number of files with the lowest scores; Re-evaluating the filtered files corresponding to the preset number of file scores with the lowest scores to obtain a second review opinion; Output the secondary review opinion.

4. The method of intelligent code review according to claim 1, characterized in that: Before scoring each submitted code to obtain a code score and providing review opinions for the submitted code, the method further includes: Collecting open source code from the developer's internal code base; Preprocessing the open source code to obtain a processed code; The processing code is sequentially length-segmented, character string-normalized, and indexed to obtain a plurality of code data blocks; Generating a corresponding vector embedding for each of the code data blocks to obtain a first vector embedding; The first vector is embedded and stored together with the code data block into a vector data block.

5. The method of intelligent code review according to claim 4, characterized in that: The step of scoring each submitted code to obtain a code score and providing review opinions for the submitted code specifically includes: For the submitted code, generate a corresponding vector embedding to obtain a second vector embedding; Performing a similarity search in a vector database by embedding the vector, and calculating the cosine similarity between the second vector embedding and each of the first vector embeddings; Querying multiple cosine similarities whose cosine similarities exceed a preset similarity, and determining similarity vector embeddings corresponding to the cosine similarities whose cosine similarities exceed the preset similarity; Determine the similar data block corresponding to the embedding of the similar vector; Using the similar data blocks as context information and the submitted code as input, generating review comments; The context information is provided with background information related to the submitted code, and combined with the private domain knowledge of the developer's enterprise, semantic understanding and problem identification are performed on the content of the submitted code, including identification of code structure, readability and naming conventions, and optimization suggestions are given to obtain the review opinion.

6. The method of intelligent code review according to claim 5, characterized in that: The step of scoring each submitted code to obtain a code score and providing review opinions for the submitted code specifically includes: Determine multiple review levels for preset code reviews, including function level, file level, and system level; Determining a weight value of each of the similar data blocks for each of the review levels; According to the multiple weight values ​​of each review level, an average is calculated to obtain a dynamic weight value of the review level; For the submitted code, generating a corresponding level score according to the review criteria corresponding to the review level; The code score is calculated based on the dynamic weight values ​​corresponding to each review level and the level score.

7. An intelligent code review system, characterized in that: The system comprises an acquisition module (201), a processing module (202) and an output module (203), wherein: The acquisition module (201) is used to bind a trigger mechanism based on the project that the developer needs to conduct code review in the code management tool, so that the code management tool can automatically call the background interface of the intelligent code review system after receiving the code pushed by the user-specified branch; The acquisition module (201) is used to acquire the submitted local code and automatically call the background interface according to the trigger mechanism, wherein the local code includes the incremental code or the full code submitted by the developer; The processing module (202) is used to filter the local code if it is determined that the type of the local code is the incremental code, and write the submission information corresponding to the target code that meets the filtering condition into the cache to wait for the scheduled task to execute the review task; The processing module (202) is used to retrieve the submission information from the cache when executing the scheduled task, and split the target code into individual submission codes according to the submission information; The processing module (202) is used to score each submitted code, obtain a code score, and provide review opinions for the submitted code; The output module (203) is used to calculate the average score of the multiple code scores to obtain the comprehensive score corresponding to the target code; The acquisition module (201) is used to acquire the source file of the local code if it is determined that the type of the local code is full code; The processing module (202) is used to filter out target files that need to be security scanned from the source files; The processing module (202) is used to parse the syntax tree of the target file and recursively obtain the target attributes in the target file; The processing module (202) is used to check whether the fully qualified name of the target attribute contains a preset suffix, and if it is determined that the fully qualified name contains the preset suffix, perform a full scan on the target file; The processing module (202) is used to perform a method-level risk review on the target code based on the result of the full scan; The output module (203) is used to output the review result according to the review result of the method-level risk review of the target code, and the review result includes the existence of risk or the absence of risk.

8. An electronic device, characterized in that: The electronic device comprises a processor (301), a communication bus (302), a user interface (303), a network interface (304) and a memory (305), wherein the memory (305) is used to store instructions, the user interface (303) and the network interface (304) are both used to communicate with other devices, the communication bus (302) is used to realize connection and communication between components in the electronic device, and the processor (301) is used to execute the instructions stored in the memory (305) so that the electronic device executes the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 6 is performed.