Code auditing method, code item storage method and computing equipment

By extracting target search fields and historical data fragments from the code project database and calculating the audit score based on preset rules, the problem of low efficiency in private domain project code auditing is solved, and efficient machine-assisted automated auditing is achieved.

CN120849283APending Publication Date: 2025-10-28XFUSION DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510979735.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

The low efficiency of reviewing code for private domain projects is mainly due to the reviewers' unfamiliarity with the existing codebase, resulting in low efficiency of manual review.

Method used

By obtaining the code to be reviewed and the code description text, extracting the target search field, and using the historical data fragments and preset review rules in the code project database, the review score is automatically calculated to achieve machine-assisted code review.

Benefits of technology

It improves the efficiency and quality of code review, reduces the need for manual understanding of the code base, and implements an automated review process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849283A_ABST
    Figure CN120849283A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a code auditing method, a code item storage method and computing equipment. When a code auditing request instruction for the to-be-audited code is received, obtaining the to-be-audited code and the code description text; and extracting a target retrieval field from the to-be-audited code and the code description text. And obtaining the target historical data fragment from the code item database according to the target retrieval field. And calculating a target auditing score based on the to-be-audited code, the target historical data fragment and a preset auditing rule. And if the target auditing score is greater than or equal to the auditing score threshold, generating first prompt information for representing that the to-be-audited code passes the auditing process. In this way, the to-be-audited codes can be automatically audited in combination with the historical data of the code items, and under the condition that the to-be-audited codes have secrecy, the features of the to-be-audited codes can still be quickly understood and audited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data processing technology, and specifically designs a code review method, a code project storage method, and a computing device. Background Technology

[0002] Private domain project code refers to code written by enterprise or individual developers for specific business or projects, as well as pre-agreed programming methods that developers can use for specific businesses. Private domain project code is usually not publicly disclosed, meaning it has strong confidentiality.

[0003] When business or project requirements change, developers can provide a code modification history, including the modified code and change documentation, which will then be reviewed by code reviewers. Code reviewers can evaluate the modified code based on its operational logic, compatibility, and other aspects. Once the modified code passes the review, it can be merged into the project.

[0004] However, due to the confidential nature of private domain project code, reviewers may not be familiar with the existing codebase and code, so they need to re-understand the codebase and the modified code during the review process, resulting in low review efficiency for the modified code. Summary of the Invention

[0005] This application provides a code review method, a code project storage method, and a computing device to solve the problem of low code review efficiency in code review access control scenarios.

[0006] In a first aspect, embodiments of this application provide a code review method, including:

[0007] In response to a code review request instruction for code to be reviewed, the code to be reviewed and its description text are obtained, with the description text describing the code to be reviewed. A target retrieval field is extracted from the code to be reviewed and the description text. The target retrieval field must at least be related to the code topic of the code to be reviewed. A target historical data fragment is retrieved from the code project database based on the target retrieval field. The target historical data fragment corresponds to the same code project as the code to be reviewed, and the code topic of the target historical data fragment is the same as the code topic of the code to be reviewed. Based on the code to be reviewed, the target historical data fragment, and preset review rules, a target review score is calculated. If the target review score is greater than or equal to a review score threshold, a first prompt message indicating that the code to be reviewed has passed the review process is generated.

[0008] Based on this implementation method, upon receiving a code review request, historical data fragments for evaluating the code to be reviewed can be retrieved from the database using target retrieval fields associated with the code topic. Then, the target review score for the code to be reviewed, the historical data fragments, and preset rules are combined to calculate the target review score. If the target review score exceeds a threshold, a first warning message is generated. This approach, which combines historical data fragments from the code project to which the code belongs, improves the quality of code review. Furthermore, the automated review of code based on machines or models enhances the efficiency of the review process.

[0009] In one feasible implementation, the target retrieval fields are extracted from the code to be reviewed and the code description text, specifically including:

[0010] Extract the first semantic information from the code description text. Determine the code topic information and source location information from the first semantic information. Extract at least one summary field from the code to be reviewed and the code description text based on the code topic information. Extract the source location field corresponding to the source location information from the code description text. Combine the summary field and the source location field into the target retrieval field.

[0011] Based on this implementation method, the retrieval field can be obtained by concatenating the summary field and the original location field. On the one hand, this allows the retrieval field to include summary information from the code to be reviewed and its description text, improving the accuracy of historical data retrieval. On the other hand, including the source location information of the code to be reviewed in the retrieval field helps distinguish the role of the same code in different structures, further enhancing the accuracy of historical data retrieval.

[0012] In another feasible implementation, historical data fragments include historical code fragment vectors, historical code fragments, historical document fragment vectors, and historical document fragments. The code project database includes a vector library, a knowledge graph database, and a document database. Historical data fragments are retrieved from the code repository based on search fields, specifically including:

[0013] The target vector index is obtained by traversing the vector indices in the vector library based on the target retrieval field. The target vector index consists of a summary field and / or a source location field from the target historical data. The target historical code fragment vector and target historical document fragment vector retrieved based on the target vector index are then obtained. Node identifiers are extracted from the target historical code fragment vector or target historical document fragment vector. The target historical code fragment is retrieved from the knowledge graph database based on the node identifiers. The target historical document fragment is retrieved from the document database based on the node identifiers.

[0014] Based on this implementation, the target vector index can be retrieved from the vector library by searching for a field, thereby obtaining the target historical code fragment vector and the target historical document fragment vector in the storage directory of the target vector index. Then, based on the node identifiers in these two vectors, the target historical code fragment and the target historical document fragment can be further obtained. In this way, during the code review process, the code to be reviewed can first be compared and reviewed based on the fragment dimension, and then reviewed based on the more granular code and document data, which can improve the accuracy of the code review.

[0015] In another feasible implementation, before calculating the review score based on the code to be reviewed, historical code fragments, and preset rules, the following is also included:

[0016] Retrieve preset rules corresponding to the code projects from the code project database. Extract second semantic information from the preset rules. The second semantic information is used to characterize the review dimensions and review scoring criteria of the code to be reviewed.

[0017] Based on this implementation, the model used to review code can access a pre-stored rule file in the database before reviewing it, and extract the second semantic information from the rule file to obtain the review dimensions and scoring criteria contained therein. In this way, the pre-stored rule file can guide the model's review direction and scoring criteria, and transform the review process for the code to be reviewed into a non-manual review process, thereby improving the efficiency of the review.

[0018] In another feasible implementation, the review score is calculated based on the code to be reviewed, historical code fragments, and preset rules, specifically including:

[0019] Independent code review information and first review score information are extracted from the second semantic information. The first review score information represents the mapping relationship between the review result of the code to be reviewed and the first review score. Review dimension information includes independent code review information, which instructs the execution of independent tests on the code to be reviewed. A first testing tool is invoked based on the independent code review information. The code to be reviewed and the first test cases based on the independent code review information are input into the first testing tool. The first test cases drive the first testing tool to output the independent review evaluation information of the code to be reviewed. Based on the first review score information, the first review score corresponding to the independent review evaluation information is calculated. The first review score is marked as the target review score for the code to be reviewed.

[0020] Based on this implementation method, the first review dimension information can be extracted from the second semantic information of the preset rule file. Then, according to the review direction indicated by the first review dimension information, a first testing tool is invoked. The first testing tool then tests the execution logic of the code to be reviewed using first test cases, and generates execution logic evaluation information. Finally, the review score corresponding to the execution logic evaluation information is calculated according to the addition and subtraction conditions indicated by the first scoring criteria. This allows for an automated testing process of the execution logic of the code to be reviewed, based on the first review dimension information, the first testing tool, and the first test cases, thereby improving the review efficiency of the code.

[0021] In another feasible implementation, the review score is calculated based on the code to be reviewed, historical code fragments, and preset rules, specifically including:

[0022] Code compatibility review information and second review score information are extracted from the second semantic information. The second review score information is used to characterize the mapping relationship between the review result of the code to be reviewed and the second review score. The review dimension information also includes code compatibility review information, which is used to instruct the code to be reviewed and the target historical data fragment to perform merge testing. The second testing tool is invoked based on the code compatibility review information. The code to be reviewed, the target historical data fragment, and the second test cases of the code to be reviewed are input into the second testing tool. The second test cases are used to drive the second testing tool to output the compatibility evaluation information of the code to be reviewed. Based on the second review score information, the second review score corresponding to the compatibility evaluation information is calculated. The second review score is marked as the target review score of the code to be reviewed.

[0023] Based on this implementation method, second review dimension information can be extracted from the second semantic information of the preset rule file. Then, according to the review direction indicated by the second review dimension information, a second testing tool is invoked. The second testing tool then uses second test cases to test the compatibility between the code to be reviewed and historical code, and generates compatibility evaluation information. In this way, by implementing an automated testing process for the compatibility of the code to be reviewed through the second review dimension information, the second testing tool, and the second test cases, the review efficiency of the code to be reviewed can be improved.

[0024] In another feasible implementation, it also includes:

[0025] If the review score is less than the score threshold, a second prompt message is generated to indicate that the code to be reviewed has failed the review process, and a third prompt message containing code modification suggestions is generated.

[0026] Based on this implementation method, when the calculated target review score is less than the score threshold, a second prompt message can be generated to indicate that the code under review has failed the review. Furthermore, a third prompt message can be generated to provide modification suggestions. In this way, the second and third prompt messages can help programmers quickly identify problems in the code under review and improve modification efficiency.

[0027] In another feasible implementation, when the target review score is greater than or equal to the review score threshold, the method further includes:

[0028] Merge the code awaiting review into the code project.

[0029] Based on this implementation method, when the target review score of the code to be reviewed exceeds the review score threshold, the code to be reviewed can be merged into the code project. On the one hand, this can enrich the code knowledge base, and on the other hand, it can increase or optimize the business or project corresponding to the code project.

[0030] Secondly, embodiments of this application provide a code project storage method, including:

[0031] Retrieve historical code and historical description documents. Extract historical code semantic information and historical document semantic information from the historical code and historical description documents, respectively. Divide the historical code into multiple historical code fragments based on the historical code semantic information. Divide the historical description documents into multiple historical document fragments based on the historical document semantic information. Establish a mapping relationship between historical code fragments and historical document fragments when there is a correlation between their semantic information and document fragment semantic information. Store the historical code fragments, historical document fragments, and mapping relationships in the code project database.

[0032] Based on this implementation method, dividing the historical code and historical description documents in the code project into multiple segments can improve data processing efficiency. Furthermore, by establishing a mapping relationship between historical code and historical description documents, the relationship between the code and the description documents can be fully reflected. This helps the model understand the meaning and association between the code and the description documents when invoked by the model, thereby facilitating the model's review of the code to be reviewed based on the historical data of the code project.

[0033] In one feasible implementation, the code project database includes a knowledge graph database. Storing historical code snippets, historical document snippets, and mapping relationships in the code project database specifically includes:

[0034] Extract semantic information from historical code snippets. Based on this semantic information, extract entity units and their dependencies from the historical code snippets. Create storage nodes and node identifiers to distinguish these nodes in the knowledge graph database. Store the entity units, their dependencies, and the mapping between entity units and document snippets in the storage nodes.

[0035] Based on this implementation, entity units such as functions, classes, and variables in the code can be stored as nodes in a knowledge graph database, and the dependencies of these entity units can also be stored in the nodes, forming a knowledge graph with entity units as nodes and dependencies as edges. The knowledge graph can provide rich dependencies between entity units, facilitating the model's discovery of code relationships and the relationships between entity units within the code.

[0036] In another feasible implementation, it also includes:

[0037] Extract source location semantic information of entity units from historical document semantic information. Source location semantic information is used to characterize the position of entity units within the code project. Store the source location semantic information in a storage node.

[0038] Based on this implementation method, by storing source location semantic information to storage nodes, entity units located in different code segments can be distinguished. Thus, when the model accesses the same entity unit, it can combine the source location information of the entity unit to determine the different roles of the entity unit in different code segments.

[0039] In another feasible implementation, the code project database includes a vector library for storing vector indexes and a directory for storing these indexes. Specifically, storing code snippets, document snippets, and mapping relationships in the database includes:

[0040] Vectorize the code snippets and document snippets separately to obtain code snippet vectors and document snippet vectors. Each code snippet vector and document snippet vector contains at least one node identifier. Construct a vector index based on the information in the storage nodes corresponding to the code snippet vectors and document snippet vectors. Store the code snippet vectors, document snippet vectors, and the mapping relationship between them in the vector index's storage directory.

[0041] Based on this implementation method, code and documents can be stored in a vectorized manner, and a vector index can be built. This allows the model to retrieve historical data corresponding to the code to be reviewed by searching fields and using the vector index, thus improving the efficiency of historical data retrieval.

[0042] In another feasible implementation, historical code is divided into multiple historical code fragments based on historical code semantic information, specifically including:

[0043] Get the code length of the historical code snippet. If the code length exceeds the code length threshold, split the historical code snippet.

[0044] Based on this implementation, when the length of historical code exceeds the length threshold, the length of each historical code segment can be reduced to less than or equal to the length threshold by splitting the historical code into segments. This facilitates the alignment of code vectors and document vectors and helps reduce the data processing time of historical code segments.

[0045] Thirdly, embodiments of this application provide a computing device, including: a processor and a memory for storing first executable instructions and second executable instructions of the processor. The processor is configured to read the first executable instructions from the memory and execute the first executable instructions to implement the code review method of the first aspect. Alternatively, it is configured to read the second executable instructions from the memory and execute the second executable instructions to implement the code item storage method of the second aspect. Attached Figure Description

[0046] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 A flowchart of the code review method provided in this application embodiment;

[0048] Figure 2 A flowchart for extracting retrieval fields provided in this application embodiment;

[0049] Figure 3 This is a schematic diagram of the database composition provided in an embodiment of this application;

[0050] Figure 4 A flowchart illustrating the preset rule extraction process provided in this application embodiment;

[0051] Figure 5 A flowchart illustrating the first code review model provided in this application embodiment;

[0052] Figure 6 A flowchart illustrating the second code review model provided in this application embodiment;

[0053] Figure 7 This application provides a flowchart of the process for handling code that fails the review process in an embodiment of the present application.

[0054] Figure 8 A flowchart of code project storage provided for embodiments of this application;

[0055] Figure 9 The knowledge graph database construction process provided in this application embodiment;

[0056] Figure 10 A flowchart illustrating the storage process of source location semantic information provided in this application embodiment;

[0057] Figure 11 Flowchart of historical code fragment vectors and historical document fragment vectors provided for embodiments of this application;

[0058] Figure 12 A flowchart illustrating the historical code partitioning provided for embodiments of this application. Detailed Implementation

[0059] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0060] Figure 1 A flowchart illustrating the code review method provided in this application embodiment. Figure 1 As shown, this application provides a code review method to improve code review efficiency based on non-human code review. The code review method includes:

[0061] S100: In response to a code review request instruction for code to be reviewed, obtain the code to be reviewed and the code description text.

[0062] In some embodiments, code projects can be stored in a database. For example, computing devices such as servers can implement specific functions or services based on the code projects stored in the database. Within the same code project, constraints such as syntax and coding standards must be followed. When a function or service changes, the modified code can be merged into the code project in the database, allowing for modifications to the service or function as needed.

[0063] Taking the example of a server implementing a specific service based on code projects, a code modification event is generated due to changes in service requirements. Therefore, the server needs to modify the code that has been modified. In this embodiment of the application, the code that has been modified is referred to as the code to be reviewed.

[0064] In some embodiments, after modifying the code according to the changed service requirements, the programmer can submit a record document containing the code to be reviewed to the server, and the server will then perform subsequent review processing based on the code to be reviewed and the code description document in the record document.

[0065] After programmers submit their documentation, the server receives the code review request and extracts semantic information from the code description text. This code description text explains the code's execution logic and functionality, and may also include information such as required interfaces, creation time, author, and source location. Therefore, after extracting the semantic information, the server can categorize the code description text into various document types, such as project documents and user documents. For example, project documents can provide specific content like project descriptions, requirements analysis, and design specifications. User documents can provide specific content such as permission design aspects of the code.

[0066] Understandably, each segmented document can be further subdivided based on its semantic hierarchy; for example, code in different locations could correspond to different requirements analyses. In this way, different types of documents are used to describe the characteristics of the corresponding code, allowing the code review model to better understand the historical data of the code project, thereby achieving better code review results.

[0067] Specifically, semantic information can be used to determine the code topic of the code to be reviewed, and the code topic corresponds to the service or function of the code to be reviewed. Therefore, historical data corresponding to the code to be reviewed can be retrieved from the database based on the semantic information representing the code topic. The code to be reviewed can then be reviewed based on this historical data.

[0068] S200: Extract the target search fields from the code to be reviewed and the code description text.

[0069] In some embodiments, extracting retrieval fields from the code to be reviewed and the code description text based on semantic information is equivalent to extracting retrieval fields from the code to be reviewed and the code description text based on the code topic.

[0070] For example, if the code topic is to modify the user permissions of function A of the target service, based on this code topic, fields related to the user permissions of function A of the service can be extracted from the code description text. Functions / variables / classes related to permission modification can also be extracted from the code to be reviewed, and then the target search fields can be formed based on the extracted fields.

[0071] In this way, the target search field is related to the code topic, that is, the target search field can reflect the code topic of the code to be reviewed, and thus the historical code fragments related to the code to be reviewed can be accurately found based on the target search field.

[0072] S300: Retrieve historical data fragments of the target from the code project database based on the target retrieval field.

[0073] In some embodiments, the database also includes a data index corresponding to historical data fragments. The data index may include important information from the historical data fragments, such as the subject and source location of the historical data fragments. In this way, the server can traverse the data indexes in the database based on the search field, and then, when the search field successfully matches the target data index, retrieve the target historical data fragment stored in the storage directory of the target data index.

[0074] When the server traverses the data index in the database based on the search field, it can simultaneously extract the semantics corresponding to the search field and the data index, and then determine the target data index based on semantic matching.

[0075] Understandably, historical data snippets retrieved based on search fields correspond to the same code projects as the code to be reviewed. This ensures that the historical data snippets provide sufficient reference for the code to be reviewed. Especially when the code project is highly confidential, its coding standards and syntax are pre-defined; therefore, only historical data from the same code project can provide sufficient evidence for the review process.

[0076] Furthermore, based on the extraction method of the target search field, the code topic corresponding to the target historical data fragment is the same as the code topic of the code to be reviewed. Therefore, based on the information contained in the target search field, the historical data fragment corresponding to the code to be reviewed can be accurately obtained, thereby improving the review quality of the code to be reviewed.

[0077] S400: Calculates the target review score based on the code to be reviewed, target historical data fragments, and preset review rules.

[0078] In some embodiments, a code review model can be loaded into the server, and the target review score of the code to be reviewed can be calculated using the code review model. The code review model can be a review model based on a Large Language Model (LLM). The code review model can perform various semantic and text-related tasks such as text generation and semantic understanding, and can inherit various testing tools to evaluate text and code data.

[0079] The preset rules can include the review dimensions and scoring criteria of the code to be reviewed. In this way, the code review model can review the code to be reviewed according to the review dimensions and scoring criteria provided by the preset rules.

[0080] For example, when the audit dimension includes the consistency between the code to be audited and the target historical data, the code audit model can call a consistency detection tool to check the consistency between the code to be audited and the target historical data, and calculate the target audit score according to the scoring criteria.

[0081] For example, when the review dimension includes reviewing the functional implementation of the code to be reviewed, the code review model can call a functional testing tool to check whether the code to be reviewed can implement all the functions described in the code description document, and calculate the target review score according to the scoring criteria.

[0082] In this way, the code review model can automate the review of code to be reviewed, thereby significantly improving the efficiency of the review process. Furthermore, the target review score generated by the code review model is sufficiently objective, which helps to control the code quality of the code to be reviewed and ensures that the code to be reviewed will not negatively impact the project.

[0083] S401: If the target review score is greater than or equal to the review score threshold, generate a first prompt message to indicate that the code to be reviewed has passed the review process.

[0084] In some embodiments, the review score threshold is preset, and the review score threshold can be set in conjunction with the score range corresponding to the scoring criteria. For example, if the score increment / decrement step in the scoring criteria is 0.5 points, then the review score threshold can be set to a single-digit number, such as 3 points. In this way, when the code review model calculates a target review score greater than or equal to 3 points, it can generate a first notification indicating that the code to be reviewed has passed the review process.

[0085] In this way, based on the setting of the review score threshold, the code review model can identify good and bad code, thereby realizing the automated review of the code to be reviewed, which helps to improve the review efficiency of the code to be reviewed.

[0086] Furthermore, if the code to be reviewed passes the code review model, it can be merged into the code project in the database. The server can then update and run the service based on the updated code project to adapt to service changes.

[0087] Figure 2 A flowchart illustrating the extraction of retrieval fields provided in an embodiment of this application. For example... Figure 2 As shown, the steps for extracting the target retrieval fields from the code to be reviewed and the code description text include:

[0088] S201: Extract the first semantic information from the code description text.

[0089] S202: Determine the code topic information and source location information in the first semantic information.

[0090] S203: Extract at least one summary field from the code to be reviewed and the code description text based on the code topic information.

[0091] S204: Extract the source location field corresponding to the source location information from the code description text.

[0092] S205: Combine the summary field and the source location field into the target retrieval field.

[0093] In some embodiments, source location information is used to describe characteristics such as the file origin of the code to be reviewed and its line number within the code project. When the code to be reviewed is written, its source location information is noted in the code description document. Then, during the field extraction phase, the server can extract the source location information from the first semantic information in the code description document. For example, the source location information may indicate that the code to be reviewed is located between lines 100 and 200 in the code project. Alternatively, the source location information may also indicate that the code to be reviewed corresponds to code segment A within the code project.

[0094] In this way, source location information can be extracted from the initial semantic information of the code description document, and the source location field can be determined based on the source location information. Based on the source location field, the functionality of the same function / class / variable in different code structures can be distinguished, thereby improving the accuracy of historical data retrieval.

[0095] Similarly, the server can also extract code topic information from the first semantic information. Code topic information can be used to describe code characteristics such as the functionality and design requirements of the code to be reviewed. Therefore, after extracting the code topic information, the server can extract a summary field from the code to be reviewed and the code description text based on the code topic information.

[0096] The summary field serves to summarize the code to be reviewed and the code description text, and can fully reflect the main content that the code to be reviewed and the code description text want to express.

[0097] Understandably, multiple summary fields can be extracted from the code to be reviewed and the code description text to fully reflect the main content to be expressed by the code to be reviewed and the code description text, so as to construct accurate target retrieval fields.

[0098] In this way, after extracting the source location field and the summary field, the server can combine them to form the target retrieval field. This target retrieval field provides both the main content of the code to be reviewed and its description text, as well as the source location of the code. This facilitates accurate retrieval of historical data from the database and helps distinguish the functions / classes / variables used in different scenarios, thereby improving the efficiency of historical data fragment retrieval.

[0099] Figure 3 This is a schematic diagram of the database composition provided in an embodiment of this application. For example... Figure 3 As shown, historical data fragments include historical code fragment vectors, historical code fragments, historical document fragment vectors, and historical document fragments. The database includes a vector library, a knowledge graph database, and a document database. Specifically, historical code fragment vectors and historical document fragment vectors are stored in the vector library, historical code fragments are stored in the knowledge graph database, and historical document fragments are stored in the document database.

[0100] In this way, when the server retrieves historical data fragments based on the target retrieval field, it can obtain the target historical data fragments from various databases. The steps include:

[0101] S301: Traverse the vector indices in the vector library based on the target retrieval field to obtain the target vector index. The target vector index consists of the summary field and / or source location field from the target historical data.

[0102] S302: Obtain the target historical code fragment vector and the target historical document fragment vector retrieved based on the target vector index.

[0103] S303: Extract node identifiers from the target historical code fragment vector or the target historical document fragment vector.

[0104] S304: Retrieve target historical code fragments from a knowledge graph database based on node identifiers.

[0105] S305: Retrieve target historical document fragments from the document database based on node identifiers.

[0106] In some embodiments, the target vector index consists of a summary field and a source location field extracted from historical code snippet vectors and historical document snippet vectors. That is, the target vector index can summarize feature information such as the topic and source location of historical code snippet vectors and historical document snippet vectors stored in its storage directory. In this way, the server can match the semantics of the target retrieval field with the vector indexes in the vector library to find the target vector index from the vector library.

[0107] Understandably, based on the extraction method of the target retrieval field and the construction foundation of the target vector index, the search efficiency and accuracy of target historical code snippets and target historical document snippets can be effectively improved.

[0108] In some embodiments, entity units and their related information in the target historical code fragment are stored in nodes, and different nodes can be distinguished by different node identifiers. A historical code fragment may include multiple entity units and their related information, and there is a correspondence between historical code fragments and historical code fragment vectors. Therefore, during the transformation from a historical code fragment to a historical code fragment vector, node identifiers can be transformed along with the historical code fragment vector so that the historical code fragment vector also contains node identifiers.

[0109] In this way, after the server finds the target historical code fragment vector, it can also search for the node corresponding to the node identifier in the knowledge graph database based on the mapping relationship provided by the node identifier in the historical code fragment vector, and obtain the entity unit and the dependency relationship of the entity unit from the node.

[0110] Similarly, after the server finds the target historical document fragment vector, it can search the document database for the target historical document fragment corresponding to the node identifier based on the mapping relationship provided by the node identifier.

[0111] In this way, by setting node identifiers, mapping relationships can be established between the data in the vector library, knowledge graph database, and document database. When searching historical data fragments, the server can prioritize searching the data in the vector library, and then search the knowledge graph database and document database based on the node identifiers obtained from the search. Prioritizing the search of data in the vector library effectively improves the search efficiency of historical data fragments, thereby improving the review efficiency of the code to be reviewed. Furthermore, the target historical document fragment vector and the target historical code fragment vector can provide feature information such as topic, coding style, and source location for the code to be reviewed, so that the code review model can use the target historical document fragment vector and the target historical code fragment vector as reference data for reviewing the code to be reviewed.

[0112] Furthermore, the entity units and their dependencies obtained from the knowledge graph database can provide references, inheritance, and calling relationships between functions / variables / classes. This allows the code review model to fully understand the code project in which the code to be reviewed resides, thereby improving the efficiency and quality of the review process.

[0113] Furthermore, historical document data fragments obtained from the document database can provide descriptive information about historical code fragments. When the code review model calls historical document data fragments, it can also extract semantic information from them and understand the historical code fragments based on the semantic information. This can improve the ability to understand the various relationships between the code to be reviewed and historical code fragments, as well as improve the efficiency of understanding.

[0114] In this way, by providing the code review model with various historical data fragments, the code review model can analyze the overall logic of the code to be reviewed, as well as the underlying logic of each entity unit and the relationships between entity units in the code to be reviewed, thus effectively improving the review quality of the code review model.

[0115] Figure 4 A flowchart illustrating the preset rule extraction process provided in this application embodiment. For example... Figure 4 As shown, when the code review model is reviewing the code to be reviewed, prompts and scoring criteria can be input into the model. The model can then execute the review process according to the prompts, which can include instructions on the review dimensions. The model can also combine the scoring criteria to score the review results and ultimately output the review outcome for the code to be reviewed.

[0116] In this way, the code review model can obtain review dimensions and scoring criteria before calculating the review score based on the code to be reviewed, historical code snippets, and preset rules. The steps include:

[0117] S311: Retrieve the preset rules corresponding to the code project from the code project database.

[0118] S312: Extract the second semantic information from the preset rules. The second semantic information is used to characterize the review dimensions and review scoring criteria of the code to be reviewed.

[0119] In some embodiments, preset rules can be data stored in a code review rule base. Preset rules can also be categorized and stored according to different code projects to build a rich code review rule base. Then, when the code review model reviews the code of each code project, it only needs to search for the corresponding preset rule based on the code project.

[0120] Preset rules can include prompts and scoring criteria. The prompt file can provide scoring dimensions, which may include, but are not limited to: consistency of style between the code to be reviewed and historical code, the execution logic of the code to be reviewed, the coding standards of the code to be reviewed, and the compatibility between the code to be reviewed and historical code. By setting scoring dimensions, the code review model can understand the semantics corresponding to the scoring dimensions and conduct test reviews of the code to be reviewed according to the scoring dimensions.

[0121] The scoring criteria document can provide scoring standards. Taking the code review model's calculation of the target review score for the code to be reviewed as an example, the scoring criteria can be to award points when the code meets the review requirements and deduct points when it does not. The review requirements can be further subdivided, and consequently, the step size for awarding and deducting points can also be further subdivided to enrich the review scenarios of the code review model.

[0122] Understandably, the preset rule files can be stored in the database in advance. When the code review model receives the code to be reviewed, it can obtain the preset rule files in the database, extract the semantic information in the preset rule files, and perform review processes such as testing and scoring on the code to be reviewed according to the semantic information.

[0123] The database can be pre-allocated to store preset rule files. Different preset rule files can be set for different code projects to adapt to their content. When the code review model receives the code to be reviewed and the code description document, it can also retrieve the corresponding preset rule file from the database based on the code project pointed to by the code to be reviewed and the code description document to obtain the review dimensions and scoring criteria.

[0124] In this way, the code review model does not need to prepare prompts or scoring criteria files in advance for each review. It can directly call the preset rule file in the database to obtain the review dimensions and scoring criteria corresponding to the code to be reviewed. This can effectively reduce the preparation time of the code review process and improve the review efficiency of the code review model.

[0125] Figure 5 This is a flowchart illustrating the first code review model provided in this application embodiment. Figure 5 As shown, the code review model calculates the review score based on the code to be reviewed, historical code snippets, and preset rules, specifically including the following steps:

[0126] S410: Extract code independent review information and first review score information from the second semantic information.

[0127] S411: Invoke the first testing tool based on the independent code review information.

[0128] S412: Input the first test case of the code to be audited and the independent audit information of the code into the first test tool. The first test case is used to drive the first test tool to output the independent audit evaluation information of the code to be audited.

[0129] S413: Based on the first review scoring information, calculate the first review score corresponding to the independent review evaluation information.

[0130] S414: Mark the first review score as the target review score for the code to be reviewed.

[0131] In some embodiments, first audit scoring information and independent code audit information are extracted from second semantic information. The first audit scoring information is used to characterize a first audit scoring standard, which can be used to characterize the mapping relationship between the audit result of the code to be audited and the first audit score. The independent code audit information is used to characterize an independent code audit standard, which is used to instruct that independent tests be performed on the code to be audited.

[0132] It should be noted that the review of code to be reviewed can include the first review dimension, which means that when testing the code to be reviewed, only the code itself is tested. For example, the execution logic of the code to be reviewed is tested.

[0133] In this way, after obtaining the preset rule file, the code review model can extract the first review score information and the independent code review information from the second semantic information of the preset rule file. The code review model can understand multiple review dimensions from the independent code review information and call the corresponding first test tool to test the code to be reviewed according to these review dimensions.

[0134] Each review dimension can be pre-configured with a corresponding primary testing tool. When the code review model invokes the primary testing tool, it can also obtain the primary test cases corresponding to the independent review dimension of the code. These primary test cases include the input, expected output, and execution conditions for testing the code to be reviewed. In this way, the code review model can use the primary testing tool to test both the code to be reviewed and the primary test cases to obtain test results.

[0135] In some embodiments, the first test case can be automatically generated by the code review model based on the characteristics of the code to be reviewed, such as its running logic and theme. This helps to improve the adaptability of the first test case to the code to be reviewed, thereby improving the review effect of the code to be reviewed.

[0136] For the primary testing tool used to test runtime logic issues, it can compare the actual output generated by executing the code to be reviewed with the expected output. If the actual output and the expected output match, the tool indicates that the runtime logic has passed the test. If the actual output and the expected output do not match, the tool indicates that the runtime logic has failed the test.

[0137] Building upon the improvements to the first testing tool, it can now also provide analysis of the code awaiting review within the test results, specifically evaluation information on its execution logic. For example, it might state, "There is a problem with the execution logic of line 5." This allows the test results to assist programmers in subsequently modifying the code awaiting review.

[0138] In addition, the first testing tool can also be used to test the completeness of the comments or corresponding code description documents of the code to be reviewed, and provide analysis results based on the test results, such as "the code comments are sufficient and the code description documents are detailed".

[0139] In this way, the code review model can combine the test results generated by the first testing tool with the scoring criteria corresponding to the review dimensions to generate a review score. For example, when the evaluation information for the execution logic is "There is a problem with the execution logic of line 5", a scoring operation of "deducting 3 points" is performed.

[0140] Understandably, the code review model generates a review score after testing the code to be reviewed using multiple primary testing tools based on multiple review dimensions. By summing the review scores, a primary review score corresponding to the primary review dimension can be obtained.

[0141] In some embodiments, where only the runtime logic of the code to be reviewed needs to be tested and there is no need to test the association between the code to be reviewed and historical code, the code review model can mark the first review score as the target review score and determine whether the code to be reviewed passes the current code review process based on the target review score.

[0142] In this way, the code review model can automatically review code in the order of calling testing tools, testing the code to be reviewed, and generating the target review score. The computational power of the code review model can effectively improve the review efficiency, and its ability to learn from historical code can improve both review efficiency and the quality of the reviewed code.

[0143] Figure 6 This is a flowchart illustrating the second code review model provided in this application embodiment. Figure 6 As shown, the code review model's process for reviewing code to be reviewed also includes:

[0144] S420: Extract code compatibility review information and second review score information from the second semantic information.

[0145] S421: Input the code to be audited, the target historical data fragment, and the second test case of the code to be audited into the second testing tool.

[0146] S422: Mark the second audit score as the target audit score for the code to be audited.

[0147] In some embodiments, the review criteria also include a second review scoring criterion, which corresponds to the code compatibility review information. The code compatibility review information is used to indicate at least several aspects, including compatibility, style consistency, coupling, and scalability, between the code to be reviewed and the target historical code segment. That is, in testing aspects such as style consistency, coupling, and scalability, the code review model needs to test the code to be reviewed in conjunction with the target historical code segment. These review dimensions are collectively referred to as code compatibility review dimensions in this embodiment.

[0148] In this way, after obtaining the preset rule file, the code review model can extract code compatibility review information from the second semantic information of the preset rule file. The code review model can then call the corresponding second testing tool based on the code compatibility review information, and generate second test cases corresponding to the code to be reviewed based on the code compatibility review information.

[0149] After generating the second test case and calling the second testing tool, the code review model can test the code to be reviewed based on the second test case, the code to be reviewed, and the target historical code snippets.

[0150] For example, when the code compatibility review dimension is to test the compatibility between the code to be reviewed and the target historical code segment, it is necessary to generate test cases for testing compatibility based on the code to be reviewed and the target historical code segment, and then test the test cases based on a second testing tool to obtain the test results.

[0151] In some embodiments, the second testing tool can be a module integrated into the code review model. Taking the code review model's assessment of the adequacy of comments in the code to be reviewed as an example, the second testing tool can first understand and learn the relationship between the target historical code snippets and comments, and based on the relationship between the target historical code snippets and comments, determine the relationship between the code to be reviewed and the comments, and further determine whether the comments in the code to be reviewed are adequacy.

[0152] In some embodiments, the code review model can invoke a second testing tool from an external source based on pre-set access permissions, so that the code to be reviewed can be reviewed by the second testing tool.

[0153] It is understood that the method of generating test cases and the testing tools are not limited in the embodiments of this application. The embodiments of this application aim to integrate the generation or invocation capabilities of testing tools and test cases into the code review model, so that the code review model can perform automated review of the code to be reviewed, thereby improving the efficiency and quality of code review.

[0154] When the code review model includes multiple code compatibility review dimensions in the code compatibility review information, it can obtain multiple review scores based on multiple second test cases and second testing tools corresponding to these dimensions. Furthermore, the code review model can sum these review scores to obtain a second review score. In scenarios where only merging testing of the code to be reviewed with the target historical code snippet is required, the second review score can be marked as the target review score to further determine whether the code to be reviewed passes the test.

[0155] In some embodiments, the review scenario for the code to be reviewed includes independent testing of the code to be reviewed, as well as combined testing of the code to be reviewed and a target historical code segment. Thus, after obtaining the first review score and the second review score of the code to be reviewed, the code review model can sum the first and second review scores, mark the sum as the target review score, and further determine whether the code passes the review based on the target review score.

[0156] In this way, the code to be reviewed can be evaluated using different testing tools based on different review dimensions. Furthermore, the review process can incorporate the modular semantics provided by the target historical code snippets and document snippets, facilitating the review of the code's functional implementation, design requirements, and other characteristics. It can also utilize the entity units and dependencies provided in the target historical code and documents to further evaluate the code's structural hierarchy. This enhances the code review model's adaptability to various types of code projects and improves the overall quality of the review process.

[0157] Figure 7 This is a flowchart illustrating the process when code awaiting review fails the review, as provided in an embodiment of this application. Figure 7 As shown, when the code review model fails the review, that is, when the target review score is less than the score threshold, the code review model generates a second prompt message to represent that the code has failed the review process, and generates a third prompt message containing code modification suggestions.

[0158] In some embodiments, the second prompt message can be used to indicate that the code to be reviewed has failed the review, for example, outputting the prompt text "The code to be reviewed has a score of 1 point and has failed the review".

[0159] The third set of prompts can be used to provide modification suggestions for the code awaiting review. For example, modification suggestions can be output in text form, and these suggestions may include, but are not limited to, the source of any issues in the code and individual performance scores. This allows for sufficient modification suggestions based on the code review model to improve the efficiency of code modification in scenarios where the code fails review.

[0160] Figure 8 A flowchart illustrating the code project storage process provided in this application embodiment. For example... Figure 8 As shown, the code project storage methods include:

[0161] S600: Retrieve historical codes and historical code description text.

[0162] In some embodiments, code project documentation includes historical code and descriptive text for that historical code. The descriptive text can be used to describe information such as design specifications, project descriptions, requirements analyses, and themes related to the historical code. Therefore, the characteristics of the historical code can be understood based on the descriptive text.

[0163] Therefore, upon completion of a code project, both the historical code and its description text need to be stored in a designated code project database to facilitate subsequent code maintenance and iteration. Furthermore, during code iteration, modified code can be tested and reviewed using a code review model. Moreover, based on the pre-established code project database using historical code and its description text, the code review model can utilize these historical code and description files as the basis for review, thereby improving the efficiency and quality of code review.

[0164] S700: Extract semantic information of historical codes and semantic information of historical documents from historical codes and historical description documents, respectively.

[0165] In some embodiments, semantic information in historical descriptive text can be extracted by invoking a semantic model. Semantic information in historical code can also be extracted using a code semantic extraction model, thereby separating entity units such as functions / classes / variables, and the dependencies between these entity units. The code semantic extraction model can be a model used to extract code structural features, such as an abstract syntax tree.

[0166] In this way, the semantic information of the historical code description text can be used to divide the historical code description text into multiple document fragments, thereby reducing the amount of data processing during storage. At the same time, the semantic information can also be used to divide the text into different document fragments corresponding to different thematic content.

[0167] Furthermore, based on semantic information extraction, the detailed structure of historical code can be extracted, such as identifying key nodes like functions, classes, and variables in historical code, as well as determining the dependencies between functions, classes, and variables.

[0168] S800: Based on historical code semantic information, historical code is divided into multiple historical code segments.

[0169] In some embodiments, historical code can be divided into multiple historical code segments based on entity units extracted from historical code and the dependencies of those entity units.

[0170] Understandably, each historical code segment can include a combination of one or more functions, variables, and classes. Given the dependencies between these entity units, entity units with explicit relationships can be grouped into the same historical code segment based on these dependencies.

[0171] For example, if functions, variables, and classes in lines 20 to 30 of the historical code exhibit clear call, reference, and inheritance relationships, then these entity units in lines 20 to 30 can be grouped into the same historical code segment. Entity units with stronger dependencies tend to implement the same functionality. Therefore, this segmentation method based on the dependencies between entity units can divide the historical code into multiple historical code segments according to the functions they correspond to. Different historical code segments can be used to implement different functions.

[0172] For example, historical code snippet A is used to control user login permissions. Historical code snippet B is used to control user access permissions. Therefore, by dividing historical code snippets based on different functions, the implementation function of each historical code snippet can be used as the semantic information corresponding to that historical code snippet.

[0173] This approach allows for the division of historical code into multiple code segments based on their functionalities, building upon the structural characteristics of the historical code. This facilitates the storage of historical code segments according to their functions, and also enables the code review model to search by function and analyze code under review based on its function when accessing historical code segments.

[0174] S900: Based on the semantic information of historical documents, the historical description document is divided into multiple historical document fragments.

[0175] In some embodiments, historical code description documents can be divided into multiple document fragments based on semantic information. For example, historical code description documents can be divided into design requirement document fragments and function implementation document fragments based on design requirement information and function implementation information identified in the semantic information.

[0176] Based on the design requirements document fragments, they can be further divided according to semantic information, for example, dividing the design requirements document fragments into design requirements document fragment A and design requirements document fragment B.

[0177] In this way, historical document fragments can correspond to partial semantic information, namely, document fragment semantic information. Document fragment semantic information can characterize features such as the theme of the corresponding historical document fragment. Therefore, based on document fragment semantic information, a connection can also be established between historical document fragments and historical code fragments.

[0178] Based on the established connection between historical code snippets and historical document snippets, the code review model can simultaneously call the other when calling either one, thereby obtaining sufficient reference for the code to be reviewed, which is conducive to improving the review quality of the code to be reviewed.

[0179] S1000: When there is a correlation between the semantic information of code fragments and the semantic information of document fragments, establish a mapping relationship between historical code fragments and historical document fragments.

[0180] In some embodiments, the semantic information of code segments divided based on functional implementation can be the topic of the code segment. Therefore, when the semantic information of code segments and document segments are the same or related, a mapping relationship between historical code segments and historical document segments can be established.

[0181] Taking historical code fragment A as an example of implementing user login permission control, the semantic information of the corresponding historical document fragment A is used to characterize the function of historical code fragment A as controlling user login permissions. Therefore, when the semantic information corresponding to the historical code fragment successfully matches the semantic information of the historical document fragment, a mapping relationship between historical code fragment A and historical document fragment A can be established.

[0182] Furthermore, historical code snippets can be associated with multiple historical document snippets. For example, a historical code snippet can also be associated with historical document snippets that describe its source location, modification time, and other characteristics. This allows for the establishment of mapping relationships between historical code snippets and multiple historical document snippets. In this way, historical code snippets can be jointly described by multiple historical document snippets. When the code review model calls historical data, it can fully utilize historical document snippets to aid in understanding the historical code snippets, thereby improving review efficiency and quality.

[0183] S1100: Store historical code snippets, historical document snippets, and mapping relationships in the code project database.

[0184] In some embodiments, after dividing historical code segments, historical document segments, and establishing a mapping relationship between them, the historical code segments, historical document segments, and mapping relationship can be stored in a code project database.

[0185] In this way, the code project database stores both the code used for project execution and corresponding code description documents based on different code snippets. This enriches the information in the code project database, allowing for the retrieval of abundant reference data when the code review model calls historical data snippets, which can then be used in the review process of the code to be reviewed.

[0186] Figure 9 The knowledge graph database construction process provided in this application embodiment. For example... Figure 9 As shown, the specific steps for storing historical code snippets, historical document snippets, and mapping relationships to the code project database include:

[0187] S1110: Extract entity units and their dependencies from historical code snippets.

[0188] S1120: Create storage nodes and node identifiers used to distinguish storage nodes in the knowledge graph database.

[0189] S1130: Store entity units, entity unit dependencies, and the mapping relationship between entity units and document fragments to the storage node.

[0190] In some embodiments, a code semantic understanding model, such as an abstract syntax tree (API), can be invoked to extract entity units and their dependencies from historical code snippets. Specifically, an API can parse entity units such as functions, variables, and classes from historical code snippets, as well as the dependencies between these entity units, forming a tree structure. The tree structure contains multiple nodes, with each function, variable, or class corresponding to one node. Traversing these nodes allows for the collection of these nodes, i.e., the collection of entity units within the nodes. Furthermore, based on the semantic understanding capabilities provided by the API, the dependencies between nodes can also be obtained.

[0191] In this way, after obtaining the entity units and their dependencies in the historical code snippets, storage nodes can be created in the knowledge graph database, and the entity units and their dependencies can be stored in the storage nodes.

[0192] Understandably, storage nodes can be logically layered by pre-setting tags and attributes, for example, by logically dividing them into entity unit layers and logical layers. This allows entity units to be stored in the entity unit layer and dependencies in the logical layer based on tags and attributes.

[0193] When creating storage nodes in a knowledge graph database, you can also set node identifiers to distinguish different nodes. For example, a code review model can retrieve historical code data from the node corresponding to that identifier. This helps to prevent data overlap to some extent.

[0194] In this way, by storing the entity units extracted from historical code snippets and the dependencies between them in a knowledge graph database, a knowledge graph corresponding to the historical code is formed. Furthermore, when the code review model calls historical code snippets, it can call each node one by one during the call process and obtain the dependencies of the entity units in each node. This facilitates understanding the code from its underlying logic, thereby improving the quality of code review.

[0195] Furthermore, storage nodes can also store mappings between historical code snippets and historical document snippets. This allows data stored in different databases to remain connected based on the node information stored in each node. Thus, when retrieving one set of data, related data can be retrieved simultaneously, improving the efficiency and comprehensiveness of historical data retrieval.

[0196] Different entity units can correspond to the same historical document fragment, so the mapping relationships in different storage nodes may point to the same historical document fragment. Thus, when retrieving historical document fragments, these entity units can also be retrieved from the knowledge graph database based on the mapping relationships.

[0197] It is understandable that different entity units can correspond to different historical document fragments. The correspondence between entity units and historical document fragments can dynamically adapt to the granularity of historical document division to fully reflect the correspondence between historical code and historical documents.

[0198] In this way, establishing a knowledge graph database facilitates code review models to understand the code from its underlying logic, thereby improving the quality of code review. Furthermore, the node information stored in the database also facilitates the retrieval of data from other databases.

[0199] Figure 10 A flowchart illustrating the storage process of source location semantic information provided in this application embodiment. For example... Figure 10 As shown, when storing data to the storage node, the process also includes storing the source location semantic information of historical code fragments to the storage node. The steps include:

[0200] S1140: Extract source location semantic information of entity units from historical document semantic information.

[0201] S1150: Store the source location semantic information to the storage node.

[0202] In some embodiments, the same entity unit may appear in different code snippets, but these entity units play different roles in different code snippets. Therefore, it is necessary to distinguish these entity units.

[0203] Understandably, historical code description documents can include metadata information such as code source, code modification time, and author. The code source can refer to the code's location within a code project and the specific file it belongs to. Therefore, source location semantic information, representing the code's origin, can be extracted from the semantic information of the historical code description documents. Since the semantic information of the historical code description documents can be represented by vectors, the source location semantic information can also be a vector. Consequently, the source location semantic information can be stored in a storage node.

[0204] In this way, even if the entity units stored in the storage nodes are the same, the code review model can distinguish the meaning of the same entity unit in different code snippets based on the semantic information of the source location when reading the data in these nodes. This can effectively improve the quality of code review. Especially in scenarios with complex code projects, it can solve the problem of difficulty in identifying the repeated use of the same entity unit.

[0205] In this way, the storage nodes include node identifiers used to establish connections between the knowledge graph database and the document database. They also include source location semantic information used to distinguish the functions of the same entity unit in different segments. This helps improve the efficiency and quality of the code review model when acquiring historical data, thereby improving the review efficiency and quality of the code review model.

[0206] Figure 11 A flowchart illustrating the historical code fragment vector and historical document fragment vector provided in the embodiments of this application. (See attached flowchart.) Figure 11 As shown, the process of establishing a vector database includes:

[0207] S1101: Vectorize the code snippet and document snippet respectively to obtain code snippet vector and document snippet vector. Each code snippet vector and document snippet vector must contain at least one node identifier.

[0208] S1102 constructs a vector index based on the information in the storage nodes corresponding to the code fragment vector and document fragment vector.

[0209] S1103: Store the code snippet vector, document snippet vector, and the mapping relationship between the code snippet vector and the document snippet vector in the storage directory of the vector index.

[0210] In some embodiments, encoding code or documents into vectors facilitates data storage, and the vectors still retain the content represented by the code and documents. Furthermore, searching based on vectors can improve modification efficiency during the search process. Therefore, when saving code projects, code snippets and document snippets can also be vectorized, and the vectorized code snippets and document snippets can be stored in a vector library.

[0211] For example, vectorization can be performed on predefined historical code and document fragments to obtain historical code fragment vectors and historical document fragment vectors. Furthermore, a vector index can be built in a vector database, enabling the code review model to search historical data based on the target retrieval field and the vector index.

[0212] Historical code snippet vectors and historical document snippet vectors can be stored together in the vector database or separately in different vector databases. However, the mapping relationship between historical code snippet vectors and historical document snippet vectors is stored together in the vector index storage directory regardless of the storage scenario.

[0213] In addition, when performing vectorization, the node information in the storage nodes of the knowledge graph database can be referenced to perform vectorization on historical code fragments and historical document fragments, so that the historical code fragments and historical document fragments have richer information, which is more convenient for the code review model to obtain.

[0214] For example, after adding node identifiers to the historical code fragment vector, the code review model can obtain at least one node identifier by parsing the vector after retrieving the historical code fragment vector from the vector database, and then obtain the node in the knowledge graph database based on the node identifier.

[0215] For example, after adding source location semantic information to the historical code fragment vector, the code review model can obtain the source location semantic information after parsing the vector, and then distinguish the functions of the same entity unit in different code fragments based on the source location semantic information.

[0216] In this way, by establishing a vector database, the data acquisition efficiency of the code review model can be improved. Furthermore, by adding node information to the vectorized data, the data acquisition efficiency and accuracy of the code review model can be further improved.

[0217] Figure 12 A flowchart illustrating the historical code partitioning provided for embodiments of this application. For example... Figure 12 As shown, in the process of dividing historical code, excessively long historical code is not conducive to storage or performing operations such as vector conversion. Therefore, when dividing historical code into multiple historical code fragments based on its semantic information, the following steps are also included:

[0218] S901: Get the code length of a historical code snippet.

[0219] S902: If the code length exceeds the code length threshold, split the historical code segment.

[0220] In some embodiments, the code used to implement a specific function in a code project may be excessively long, for example, 500 lines of code. 500 lines of code correspond to a large amount of data, which can lead to inefficiencies in data processing processes such as vector conversion. Furthermore, the large amount of data can also reduce storage and retrieval efficiency.

[0221] Therefore, a code length threshold can be set based on the actual data processing capacity, making the code length threshold a judgment condition. In this way, when dividing historical code segments, the code length of the historical code segments can be obtained, and if the code length exceeds the code length threshold, the historical code segments can be further split into sub-segments to avoid the data volume being too large and affecting data processing such as storage and retrieval.

[0222] It is understandable that when splitting historical code fragments, the semantic information provided by the corresponding historical code document fragments can be combined with the semantic information of the historical code fragments to ensure that the split historical code fragments still have complete logical features, so that the subsequent code review model can read the data and understand the functions and other features corresponding to the split historical code fragments.

[0223] This application provides a computing device, which includes a processor and a memory. The memory stores a first executable instruction and a second executable instruction of the processor. The processor reads the first executable instruction from the memory and executes it to implement the code item storage method described in the above embodiments. The processor can also read the second executable instruction from the memory and execute it to implement the code item storage method described in the above embodiments.

[0224] Similar parts between the embodiments provided in this application can be referred to mutually. The specific implementation methods provided above are only a few examples under the overall concept of this application and do not constitute a limitation on the scope of protection of this application. For those skilled in the art, any other implementation methods extended from the solution of this application without creative effort shall fall within the scope of protection of this application.

Claims

1. A code review method, characterized in that, include: In response to a code review request instruction for code to be reviewed, the code to be reviewed and code description text are obtained, wherein the code description text is used to describe the code to be reviewed; Extract target retrieval fields from the code to be reviewed and the code description text; the target retrieval fields are at least related to the code topic of the code to be reviewed. The target historical data fragment is obtained from the code project database according to the target retrieval field; the target historical data fragment corresponds to the same code project as the code to be reviewed, and the code theme corresponding to the target historical data fragment is the same as the code theme of the code to be reviewed; Based on the code to be reviewed, the target historical data fragment, and the preset review rules, the target review score is calculated; If the target review score is greater than or equal to the review score threshold, a first prompt message is generated to indicate that the code to be reviewed has passed the review process.

2. The method according to claim 1, characterized in that, The step of extracting the target retrieval field from the code to be reviewed and the code description text also includes: Extract the first semantic information from the code description text; Determine the code topic information and source location information from the first semantic information; Extract at least one summary field from the code to be reviewed and the code description text based on the code topic information; Extract the source location field corresponding to the source location information from the code description text; The summary field and the source location field are combined to form the target retrieval field.

3. The method according to claim 1 or 2, characterized in that, The historical data fragments include historical code fragment vectors, historical code fragments, historical document fragment vectors, and historical document fragments; the code project database includes a vector library, a knowledge graph database, and a document database; the step of retrieving historical data fragments from the code library based on the search field specifically includes: The target vector index is obtained by traversing the vector indexes in the vector library according to the target retrieval field; the target vector index is composed of the summary field and / or source location field in the target historical data. Obtain the target historical code fragment vector and the target historical document fragment vector retrieved based on the target vector index; Extract node identifiers from the target historical code fragment vector or the target historical document fragment vector; Based on the node identifier, retrieve the target historical code fragment from the knowledge graph database; The target historical document fragment is retrieved from the document database based on the node identifier.

4. The method according to claim 1, characterized in that, Before calculating the review score based on the code to be reviewed, the historical code fragments, and preset rules, the process also includes: Retrieve the preset rules corresponding to the code project from the code project database; The second semantic information is extracted from the preset rules; the second semantic information is used to characterize the review dimension and review scoring criteria of the code to be reviewed.

5. The method according to claim 4, characterized in that, The calculation of the review score based on the code to be reviewed, the historical code fragments, and preset rules specifically includes: Independent code review information and first review score information are extracted from the second semantic information; the first review score information is used to characterize the mapping relationship between the review result of the code to be reviewed and the first review score; the review dimension information includes independent code review information, which is used to indicate that independent testing is performed on the code to be reviewed; The first testing tool is invoked based on the independent review information of the code. The first test case, which is a code to be reviewed and independent review information of the code, is input into the first test tool. The first test case is used to drive the first test tool to output the independent review evaluation information of the code to be reviewed. Based on the first review scoring information, calculate the first review score corresponding to the independent review evaluation information; The first review score is marked as the target review score for the code to be reviewed.

6. The method according to claim 5, characterized in that, The calculation of the review score based on the code to be reviewed, the historical code fragments, and preset rules specifically includes: Extract code compatibility review information and second review score information from the second semantic information; the second review score information is used to characterize the mapping relationship between the review result of the code to be reviewed and the second review score; the review dimension information also includes code compatibility review information, which is used to instruct the code to be reviewed and the target historical data fragment to perform merge testing; and call the second testing tool according to the code compatibility review information; The code to be reviewed, the target historical data fragment, and the second test case of the code to be reviewed are input into the second testing tool. The second test case is used to drive the second testing tool to output the compatibility evaluation information of the code to be reviewed. Based on the second review score information, the second review score corresponding to the compatibility evaluation information is calculated. The second review score is marked as the target review score for the code to be reviewed.

7. A method for storing code projects, characterized in that, include: Retrieve historical code and historical description documents; Semantic information of historical code and semantic information of historical document are extracted from the historical code and the historical description document, respectively. Based on the semantic information of the historical code, the historical code is divided into multiple historical code fragments; Based on the semantic information of the historical documents, the historical description document is divided into multiple historical document fragments; When the semantic information of the code fragment and the semantic information of the document fragment are associated, a mapping relationship between the historical code fragment and the historical document fragment is established; The historical code snippets, the historical document snippets, and the mapping relationships are stored in the code project database.

8. The method according to claim 7, characterized in that, The code project database includes a knowledge graph database; storing the historical code fragments, the historical document fragments, and the mapping relationships in the code project database specifically includes: Extract code semantic information from the historical code fragments; Based on the semantic information of the code, entity units and their dependencies are extracted from the historical code fragments. Create storage nodes and node identifiers for distinguishing the storage nodes in the knowledge graph database; The entity unit, the dependencies of the entity unit, and the mapping relationship between the entity unit and the document fragment are stored in the storage node.

9. The method according to claim 8, characterized in that, Also includes: The source location semantic information of the entity unit is extracted from the semantic information of the historical document; the source location semantic information is used to characterize the position of the entity unit in the code project; The source location semantic information is stored in the storage node.

10. The method according to claim 9, characterized in that, The code project database includes a vector library, which stores vector indexes and a directory for storing vector indexes; storing the code snippets, document snippets, and mapping relationships in the code project database specifically includes: The code snippet and the document snippet are vectorized respectively to obtain a code snippet vector and a document snippet vector; each code snippet vector and document snippet vector contains at least one node identifier; A vector index is constructed based on the information in the storage nodes corresponding to the code fragment vector and the document fragment vector; The code snippet vector, the document snippet vector, and the mapping relationship between the code snippet vector and the document snippet vector are stored in the storage directory of the vector index.

11. A computing device, characterized in that, include: processor: A memory for storing the first and second executable instructions of the processor; The processor is configured to read the first executable instruction from the memory and execute the first executable instruction to implement the code review method according to any one of claims 1-6. Alternatively, it can be used to read the second executable instruction from the memory and execute the second executable instruction to implement the code item storage method according to any one of claims 7-10.