A method for model training, a method for code recognition, and corresponding devices

The generated code critical identification model through model training solves the problem of long-term understanding and review of code in code review, and improves the efficiency and accuracy of code review.

CN116187410BActive Publication Date: 2025-06-27HUAWEI TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111425345.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2025-06-27
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

During code review, it takes a long time to understand and review the code, especially when the project code is large and the staff is large, which leads to inefficient review.

Method used

Through model training methods, a model can be generated that can identify the criticality of method codes in the code. The model optimizes parameters to determine the critical information of the method code by obtaining a combined call graph of multiple training samples, converting it into a collection of path-contexts, and vectorizing it with the call relationship.

Benefits of technology

Improve the efficiency of code review, and automatically identify code criticality, reduce the time it takes for reviewers to understand and review code, and enhance the accuracy of code quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116187410B_ABST
    Figure CN116187410B_ABST
Patent Text Reader

Abstract

The present application discloses a method for model training and a method for code recognition. The path-context obtained by the method code from the project code can be used to train a key model, and then the trained key model can be used to identify the key information of the method code or the key sorting of multiple method codes in the project code to be reviewed, thereby assisting code reviewers in code review. In the solution provided by the present application, since the path-context obtained by the method code has a small granularity, the accuracy of the trained key model is high. Through this key model, the sorting of multiple method codes can be quickly output, thereby improving the speed of code review.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and particularly to a method for model training, a method for code recognition, and corresponding devices. Background Art

[0002] In the software project development process, code review is an essential link. Code review is one of the best practices in software development, which can effectively improve the overall code quality and timely detect possible problems in the code.

[0003] In the code review process, the most time-consuming part is for the reviewer to read and understand the code. In existing experiments, it is statistically shown that programmers need to spend an average of 345 seconds to understand a code submission that on average contains 4 classes. With the rapid growth of the project code volume and the continuous addition of personnel to the project, the time consumption of this part has doubled. There is a pain point in code review: reviewers need to spend a lot of time understanding the code and changes, especially those involving changes to multiple files.

[0004] Therefore, it is meaningful to solve the problem of the time taken to understand the code. Summary of the Invention

[0005] This application provides a method for model training, which is used to obtain a model that can identify the key nature of method codes in the project codes to be reviewed, so that the key nature of method codes in the project codes or the key nature ranking of multiple method codes can be determined through this model, thereby improving the code review efficiency. This application also provides corresponding devices, computer devices, computer-readable storage media, computer program products, etc.

[0006] The first aspect of the present application provides a method for model training, including: obtaining a plurality of training samples, where each training sample is a combined call graph of project codes that have undergone changes. The combined call graph is obtained by combining the code call graph before the project code change and the code call graph after the change. The combined call graph represents multiple method codes included in the project code before and after the change through multiple nodes, and represents the call relationship between two method codes among the multiple method codes through edges. The method code corresponding to each node of the combined call graph also corresponds to first key information; for each training sample, converting each method code among the multiple method codes into a first set of path-contexts, where the first set includes multiple path-contexts, and each path-context represents the path information between any two leaf nodes and the middle part after the method code is converted into an abstract syntax tree; training a first key model according to the first set corresponding to each method code among the multiple method codes and the first key information corresponding to each method code to obtain a second key model; where: The first key model includes a first layer, a second layer, and a third layer. The first layer is used to convert the first set into a first vector representation, the second layer is used to process the first vector representation of each method code in combination with the call relationship between multiple method codes to obtain a second vector representation of each method code, and the third layer is used to determine the second key information of each method code according to the second vector representation of each method code, and supervise the second key information according to the first key information of each node in the combined call graph to optimize the parameters in the first key model; The second key model is used to output the key information of each method code in the target project code to be reviewed or the key sorting information of multiple method codes in the target project code.

[0007] The model training solution of the present application can be completed under an integrated development environment (IDE). An IDE is an application program used to provide a program development environment, generally including tools such as a code editor, a compiler, a debugger, and a graphical user interface.

[0008] In the present application, project code refers to the code written for a project, and method code is the code for various functions written to complete the functions of the project. The method code can be divided into key method code and non-key method code. The key method code refers to the code written to complete the key logic related to computing, and the non-key method code refers to the code written to assist or cooperate with the key method. It can also be said that the key degree of the method code is different. A project code contains multiple method codes.

[0009] In this application, the project code that has undergone changes refers to the modified project code. Once the project code is modified, a code review needs to be performed again. Code review refers to the process of systematically checking the source code during the software development process. The general purpose is to find various defects, including code defects, function implementation issues, coding rationality, performance optimization, etc., to ensure the overall quality of the software.

[0010] In this application, the code call graph refers to the call graph obtained by analyzing the method calls of the project code using the Doxygen tool. The code call graph includes nodes and edges. Each node represents a method code, and the edge represents the call relationship between two method codes with a scheduling relationship. When a project undergoes changes, it usually means that some code is modified, and there will be some code that is not modified. In this way, based on the unmodified code, the code call graph before the change and the code call graph after the change can be combined to obtain a combined call graph. The combined call graph will include the nodes and edges corresponding to all the method codes before the change, as well as the nodes and edges corresponding to the method codes of the modified part after the change.

[0011] In this application, the first key information can be the key value of the method code, and this value can be represented in a normalized form, such as: 1, 0.9, 0.8, or other values, which are not limited in this application. The first key information of the method code corresponding to each node in the combined call graph can be marked by experienced programmers.

[0012] In this application, during the model training process, each training sample can be used as a batch for training.

[0013] In this application, each method code can be converted into an abstract syntax tree. The abstract syntax tree includes a parent node and leaf nodes. The leaf nodes of the upper layer can be used as the parent nodes of the leaf nodes of the lower layer. A path can be formed between any two leaf nodes through their parent nodes. In this way, when there are n leaf nodes, there will be n*(n - 1) / 2 paths, and each pair of leaf nodes and the path between them can be called a path-context. Therefore, each method code can be converted into a first set of path-contexts. The first set contains the path-contexts composed of any two leaf nodes and the intermediate path in the abstract syntax tree converted from the method code.

[0014] In this application, the first key model can perform vectorization processing on the first set, so as to obtain the first vector representation of a single method code. Then, by combining the call relationships between multiple method codes and processing the first vector representation, the second vector representation for finally representing each method code can be obtained. Furthermore, the second key information of each method code can be determined based on the second vector representation, and the form of the second key information can be understood by referring to the first key information.

[0015] In this application, the process of optimizing the parameters in the first key model can be to use the gradient descent algorithm to optimize the parameters, and finally obtain the second key model that can be used for code recognition through the training of multiple training samples.

[0016] As can be seen from the above content of the first aspect, during the model training process, for the method codes in each project code, path-context conversion is performed, and then vectorization processing is performed on the first set of path-context of the method code, thus refining the granularity of the training samples, and the accuracy of the trained model in code recognition will also be higher. In addition, when determining the vector of the method code, the call relationships between different method codes are also combined, considering the relevance between different method codes, further improving the accuracy of model training, and thus further improving the accuracy of the model in code recognition. In this way, the key points of the code can be determined through this model in the code review link, improving the code review efficiency.

[0017] In a possible implementation manner of the first aspect, the above step: for each training sample, converting each method code in multiple method codes into the first set of path-context includes: for each training sample, using the first key model to convert each method code in multiple method codes into the first set of path-context.

[0018] In this possible implementation manner, the process of converting the method code into the first set of path-context can be completed in the first key model. In this way, it is equivalent to integrating more functions into the model, enhancing the capabilities of the model.

[0019] In a possible implementation of the first aspect, the above steps: training the first key model according to the first set corresponding to each method code and the first key information corresponding to each method code in multiple method codes, include: vectorizing each path-context in the first set to obtain a third vector representation of each path-context; determining a first vector representation corresponding to the first set according to the third vector representation of each path-context; determining a second vector representation of each method code according to the first vector representation of each method code and the call relationship between multiple method codes in the combined call graph; determining the second key information of each method code by using the self-attention mechanism according to the second vector representation of each method code, and supervising the second key information according to the first key information of each node in the combined call graph to optimize the parameters in the first key model.

[0020] In this possible implementation, the third vector representation refers to the vector representation for each path-context. The first vector representation can be obtained by summing the third vector representations, or by weighting each third vector representation and then summing, or by performing dimensionality reduction on the third vector representation, weighting the dimensionally reduced third vector representation and then summing. The specific way of obtaining the first vector representation from the third vector representation is not limited in this application. It can be seen from this possible implementation that through three layers of vectorization processing, the accuracy of the second vector representation finally used to determine the second key information of the method code can be higher, thereby improving the accuracy of the model.

[0021] In a possible implementation of the first aspect, the above steps: vectorizing each path-context in the first set to obtain a third vector representation of each path-context, include: performing word segmentation on each of the two leaf nodes of each path-context, and using the average vector representation of multiple sub-words obtained after word segmentation of each leaf node as the vector representation of each leaf node; concatenating the vector representations of the two leaf nodes and the vector representation of the path information between the two leaf nodes to obtain a third vector representation of each path-context.

[0022] In this possible implementation, a leaf node can be understood as a sub-word string, including multiple sub-words. The vector representation of each sub-word can be retrieved from the vocabulary of sub-words. The vector representation of the leaf node can be determined by first summing and then averaging the vector representations of each sub-word in the sub-word string. The vector representation of the path between two leaf nodes can be found in the path vocabulary. In this way, by concatenating the vector representations of the two leaf nodes and the vector representation of the intermediate path, the third vector representation of the path-context can be obtained. In this possible implementation, the granularity of the vector representation used to train the model is refined to the level of sub-words and paths, further improving the accuracy of the model.

[0023] In a possible implementation of the first aspect, the above step: determining the first vector representation corresponding to the first set according to the third vector representation of each path-context includes: performing dimensionality reduction on the third vector representation to obtain the vector representation after dimensionality reduction; summing the products of the vector representation after dimensionality reduction corresponding to each path-context in the first set and the corresponding attention magnitude to obtain the first vector representation corresponding to the first set, where the attention magnitude is determined by the vector representation after dimensionality reduction and the attention vector.

[0024] In this possible implementation, since the third vector representation is obtained by concatenation, its dimension is relatively high. By performing dimensionality reduction, the vector representation after dimensionality reduction of the third vector representation can be obtained. Using the attention magnitude as the weighted weight after dimensionality reduction can increase the proportion of the third vector representation of important path-contexts in the first vector representation. As can be seen from this possible implementation, the attention mechanism is used when determining the first vector representation, which is beneficial to highlighting the key points of some method codes, thereby improving the accuracy of the model.

[0025] In a possible implementation of the first aspect, the above step: converting each method code in multiple method codes into the first set of path-contexts includes: converting each method code in multiple method codes into an abstract syntax tree, where the abstract syntax tree includes multiple leaf nodes, and a path is formed between any two leaf nodes among the multiple leaf nodes and their parent node; determining any two leaf nodes and the path between the two leaf nodes as a path-context, and the first set includes multiple path-contexts obtained by pairwise combination of all the leaf nodes in the abstract syntax tree.

[0026] In this possible implementation, converting path-contexts through an abstract syntax tree can improve the speed of path-context conversion.

[0027] In a possible implementation of the first aspect, determining any two leaf nodes and the path between any two leaf nodes as a path-context includes: performing word segmentation processing on the two leaf nodes in any two leaf nodes and the path between any two leaf nodes according to the naming rule to obtain a path-context including any two leaf nodes and the path between any two leaf nodes.

[0028] In this possible implementation, performing word segmentation processing on the leaf nodes according to the naming rule can substitute the function information and semantic information in the code into the vector representation of the path-context for subsequent model training, which is beneficial to improving the accuracy of the model.

[0029] In a possible implementation of the first aspect, before the above step of obtaining a plurality of training samples, the method further includes: obtaining a plurality of historical project codes that have undergone changes; for each historical project code, obtaining the file content before the change and the file content after the change; performing call analysis on the file content before the change to obtain a first call graph, and performing call analysis on the file content after the change to obtain a second call graph; based on the unchanged content in the first call graph and the second call graph, merging the first call graph and the second call graph to obtain a combined call graph.

[0030] In this possible implementation, the training samples can be the historical project codes that have undergone changes selected from the historical project codes. In this way, a large number of samples can be obtained for training, which can improve the accuracy of the model.

[0031] The second aspect of the present application provides a method for code recognition, including: receiving a target project code to be reviewed, where the target project code has been changed; determining a corresponding combined call graph according to the target project code, and the combined call graph is obtained by combining the call graph before the change of the target project code and the call graph after the change. The combined call graph represents multiple method codes included in the target project code before and after the change through multiple nodes, and represents the call relationship between two method codes among the multiple method codes through edges; converting each method code among the multiple method codes into a first set of path-contexts, where the first set includes multiple path-contexts, and each path-context represents the path information between any two leaf nodes and the path in between after the method code is converted into an abstract syntax tree; determining the criticality information of each method code in the target project code or the criticality sorting information of multiple method codes in the target project code according to the first set corresponding to each method code among the multiple method codes and the target criticality model. The target criticality model includes a first layer, a second layer, and a third layer. The first layer is used to convert the first set into a first vector representation, the second layer is used to process the first vector representation of each method code in combination with the call relationship between multiple method codes to obtain a second vector representation of each method code, and the third layer is used to determine the criticality information of each method code or the criticality sorting information of multiple method codes according to the second vector representation of each method code.

[0032] The code recognition solution of the present application can be completed in an integrated development environment (IDE). An IDE is an application program for providing a program development environment, generally including tools such as a code editor, a compiler, a debugger, and a graphical user interface.

[0033] In the present application, the target project code to be reviewed refers to that the target project code has been modified and needs to be submitted to the reviewer for re-review. To facilitate the reviewer to review the code, the criticality information of each method code in the target project code or the criticality sorting of multiple method codes in the target project code can be determined first through the target criticality model.

[0034] In the present application, the target criticality model can be a second criticality model trained by adopting the above-mentioned first aspect or any possible implementation manner of the first aspect.

[0035] In the present application, for other features that are the same as those in the first aspect, reference can be made to the description in the first aspect for understanding, and details will not be repeated here.

[0036] As can be seen from the content of the second aspect above, for the target project code to be reviewed, first use the target key model to determine the key information of each method code in the target project code or the key sorting of multiple method codes in the target project code. This can assist code reviewers in reviewing the code and improve the efficiency of code review.

[0037] In a possible implementation manner of the second aspect above, the above step: converting each method code in multiple method codes into a first set of path-contexts includes: using the target key model to convert each method code in multiple method codes into a first set of path-contexts.

[0038] In this possible implementation manner, the process of converting the method code into the first set of path-contexts can be completed in the target key model. In this way, it is equivalent to integrating more functions into the model, enhancing the capabilities of the model.

[0039] In a possible implementation manner of the second aspect above, the above step: determining the key information of each method code in the target project code or the key sorting information of multiple method codes in the target project code according to the first set corresponding to each method code in multiple method codes and the target key model includes: vectorizing each path-context in the first set to obtain a third vector representation of each path-context; determining a first vector representation corresponding to the first set according to the third vector representation of each path-context; determining a second vector representation of each method code according to the first vector representation of each method code and the call relationship between multiple method codes in the combined call graph; and determining the key information of each method code or the key sorting information of multiple method codes in the target project code by using the self-attention mechanism according to the second vector representation of each method code.

[0040] In this possible implementation manner, the third vector representation refers to the vector representation of each path-context. The first vector representation can be obtained by summing the third vector representations, or by weighting each third vector representation and then summing, or by performing dimensionality reduction on the third vector representations and then weighting and summing the dimensionality-reduced third vector representations. The specific way of obtaining from the third vector representation to the first vector representation is not limited in this application. As can be seen from this possible implementation manner, through three-layer vectorization processing, the accuracy of the key information of each method code determined or the key sorting information of multiple method codes in the target project code is improved.

[0041] In a possible implementation manner of the second aspect described above, the above step of vectorizing each path-context in the first set to obtain a third vector representation of each path-context includes: performing word segmentation on each of the two leaf nodes of each path-context, and using the average vector representation of the multiple sub-words obtained after word segmentation of each leaf node as the vector representation of each leaf node; concatenating the vector representations of the two leaf nodes respectively and the vector representation of the path information between the two leaf nodes to obtain a third vector representation of each path-context.

[0042] In this possible implementation manner, a leaf node can be understood as a sub-word string including multiple sub-words. The vector representation of each sub-word can be retrieved from the vocabulary of sub-words. The vector representation of a leaf node can be determined by first summing and then averaging the vector representations of each sub-word in the sub-word string. The vector representation of the path between the two leaf nodes can be found in the path vocabulary. In this way, by concatenating the vector representations of the two leaf nodes and the vector representation of the intermediate path, a third vector representation of the path-context can be obtained. In this possible implementation manner, the granularity of the vector representation for determining the key information of the method code is refined to the level of sub-words and paths, further improving the accuracy of the key information of each method code determined or the key sorting information of multiple method codes in the target project code.

[0043] In a possible implementation manner of the second aspect described above, the above step of determining a first vector representation corresponding to the first set according to the third vector representation of each path-context includes: performing dimensionality reduction on the third vector representation to obtain a vector representation after dimensionality reduction; summing the product of the vector representation after dimensionality reduction corresponding to each path-context in the first set and the corresponding attention magnitude, where the attention magnitude is determined by the vector representation after dimensionality reduction and the attention vector, to obtain a first vector representation corresponding to the first set.

[0044] In this possible implementation manner, since the third vector representation is obtained by concatenation, it has a relatively high dimension. By performing dimensionality reduction, a vector representation after dimensionality reduction of the third vector representation can be obtained. Using the attention magnitude as the weighted weight after dimensionality reduction can increase the proportion of the third vector representation of important path-contexts in the first vector representation. It can be seen from this possible implementation manner that the attention mechanism is used when determining the first vector representation, which is beneficial to highlighting the key nature of some method codes, thereby further improving the accuracy of the key information of each method code determined or the key sorting information of multiple method codes in the target project code.

[0045] In a possible implementation of the second aspect described above, the above steps of converting each method code among multiple method codes into a first set of path-contexts include: converting each method code among multiple method codes into an abstract syntax tree, where the abstract syntax tree includes multiple leaf nodes, and a path is formed between any two leaf nodes among the multiple leaf nodes and the root node closest to the two leaf nodes; determining any two leaf nodes and the path between the two leaf nodes as a path-context, and the first set includes multiple path-contexts obtained by pairwise combination of all the leaf nodes in the abstract syntax tree.

[0046] In this possible implementation, converting path-contexts through an abstract syntax tree can improve the speed of path-context conversion.

[0047] In a possible implementation of the second aspect described above, the above steps of determining any two leaf nodes and the path between the two leaf nodes as a path-context include: performing word segmentation processing on the two leaf nodes among any two leaf nodes and the path between the two leaf nodes according to a naming rule to obtain a path-context including any two leaf nodes and the path between the two leaf nodes.

[0048] In this possible implementation, performing word segmentation processing on leaf nodes according to a naming rule can substitute the functional information and semantic information in the code into the vector representation of the path-context for subsequent determination of key information, which is beneficial to improving the accuracy of the key information determined for each method code or the key sorting information of multiple method codes in the target project code.

[0049] In a possible implementation of the second aspect described above, the above steps of determining a corresponding combined call graph according to the target project code include: for the target project code, obtaining the file content before the change and the file content after the change; performing call analysis on the file content before the change to obtain a first call graph, and performing call analysis on the file content after the change to obtain a second call graph; based on the unchanged content in the first call graph and the second call graph, merging the first call graph and the second call graph to obtain a combined call graph.

[0050] The third aspect of the present application provides a model training device, and this model training device has the function of implementing the method of the first aspect or any possible implementation manner of the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. For example: an acquisition unit, a first processing unit, and a second processing unit, and these units can be implemented by one processing unit or multiple processing units.

[0051] The fourth aspect of the present application provides a code recognition device, which has the function of implementing the method according to the second aspect or any possible implementation manner of the second aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, such as: a receiving unit, a first processing unit, a second processing unit, and a third processing unit, and these three processing units can be implemented by one processing unit or multiple processing units.

[0052] The fifth aspect of the present application provides a computer device, which includes at least one processor, a memory, an input / output (I / O) interface, and computer-executable instructions stored in the memory and executable on the processor. When the computer-executable instructions are executed by the processor, the processor executes the method according to the first aspect or any possible implementation manner of the first aspect.

[0053] The sixth aspect of the present application provides a computer device, which includes at least one processor, a memory, an input / output (I / O) interface, and computer-executable instructions stored in the memory and executable on the processor. When the computer-executable instructions are executed by the processor, the processor executes the method according to the second aspect or any possible implementation manner of the second aspect.

[0054] The seventh aspect of the present application provides a computer-readable storage medium storing one or more computer-executable instructions. When the computer-executable instructions are executed by a processor, one or more processors execute the method according to the first aspect or any possible implementation manner of the first aspect.

[0055] The eighth aspect of the present application provides a computer-readable storage medium storing one or more computer-executable instructions. When the computer-executable instructions are executed by a processor, one or more processors execute the method according to the second aspect or any possible implementation manner of the second aspect.

[0056] The ninth aspect of the present application provides a computer program product storing one or more computer-executable instructions. When the computer-executable instructions are executed by one or more processors, one or more processors execute the method according to the first aspect or any possible implementation manner of the first aspect.

[0057] The tenth aspect of the present application provides a computer program product storing one or more computer-executable instructions. When the computer-executable instructions are executed by one or more processors, one or more processors execute the method according to the second aspect or any possible implementation manner of the second aspect.

[0058] The eleventh aspect of this application provides a chip system, which includes at least one processor. The at least one processor is used to support the device for model training to implement the functions involved in the above-mentioned first aspect or any possible implementation manner of the first aspect. In a possible design, the chip system may further include a memory, and the memory is used to store the necessary program instructions and data for the device for handling page faults. The chip system may be composed of chips or may include chips and other discrete devices.

[0059] The twelfth aspect of this application provides a chip system, which includes at least one processor. The at least one processor is used to support the device for code recognition to implement the functions involved in the above-mentioned second aspect or any possible implementation manner of the second aspect. In a possible design, the chip system may further include a memory, and the memory is used to store the necessary program instructions and data for the device for handling page faults. The chip system may be composed of chips or may include chips and other discrete devices. Description of the Drawings

[0060] Figure 1 is a schematic diagram of a scenario for model training and application provided by an embodiment of this application;

[0061] Figure 2 is a schematic diagram of an embodiment of the method for model training provided by an embodiment of this application;

[0062] Figure 3 is a schematic diagram of an example for collecting samples provided by an embodiment of this application;

[0063] Figure 4 is a schematic diagram of another embodiment of the method for model training provided by an embodiment of this application;

[0064] Figure 5 is a schematic diagram of the structure of an abstract syntax tree provided by an embodiment of this application;

[0065] Figure 6 is a schematic diagram of an example of an abstract syntax tree provided by an embodiment of this application;

[0066] Figure 7 is a schematic diagram of an embodiment of the method for model training provided by an embodiment of this application;

[0067] Figure 8 is a schematic diagram of the structure of a key model provided by an embodiment of this application;

[0068] Figure 9 is a schematic diagram of a code review process provided by an embodiment of this application;

[0069] Figure 10It is a schematic diagram of an embodiment of the method for code recognition provided by an embodiment of the present application;

[0070] Figure 11 It is a schematic diagram of another embodiment of the method for code recognition provided by an embodiment of the present application;

[0071] Figure 12 It is a schematic diagram of another scenario for model training and application provided by an embodiment of the present application;

[0072] Figure 13 It is a schematic structural diagram of an apparatus for model training provided by an embodiment of the present application;

[0073] Figure 14 It is a schematic structural diagram of an apparatus for code recognition provided by an embodiment of the present application;

[0074] Figure 15 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0075] The embodiments of the present application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Those of ordinary skill in the art can know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0076] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0077] The embodiments of the present application provide a method for model training, which is used to obtain a model that can identify the key points of method codes in the project codes to be reviewed, so that the key points of method codes in the project codes or the key point ranking of multiple method codes can be determined through this model, improving the code review efficiency. The present application also provides corresponding apparatuses, computer devices, computer-readable storage media, computer program products, etc. The following will be described in detail respectively.

[0078] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce an intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable them to have functions of perception, reasoning, and decision-making.

[0079] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0080] Intelligent manufacturing, intelligent transportation, smart home, intelligent healthcare, intelligent security, autonomous driving, smart city, intelligent terminals, etc.

[0081] Models are usually trained in the computer devices or platforms of the model owners (such as servers, virtual machines (VMs), or containers), and the trained models are stored in the form of model files. When the devices of the model users (such as terminal devices, servers, edge devices, VMs, or containers) need to use the model, the devices of the model users actively load the model files of the model or the devices of the model owners actively send the model files for installing the model to the devices of the model users, so that the model can be applied on the devices of the model users to perform corresponding functions.

[0082] The server refers to a physical machine.

[0083] A terminal device (which can also be referred to as a user equipment (UE)) is a device with wireless transceiver functions. It can be deployed on land, including indoor or outdoor, handheld or vehicle-mounted; it can also be deployed on water (such as on a ship, etc.); it can also be deployed in the air (such as on an airplane, a balloon, a satellite, etc.). The terminal can be a mobile phone, a tablet (pad), a computer with wireless transceiver functions, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc.

[0084] Both a VM or a container can be a virtualized device partitioned in a virtualized manner on the hardware resources of a physical machine.

[0085] In the embodiments of the present application, an initial model designed for code recognition can be placed on a server for training to obtain a target model capable of code recognition. After the target model is installed on the terminal device, code recognition can be performed. This process can refer to Figure 1 the schematic diagram for immediate execution.

[0086] Such as Figure 1 a schematic diagram of a system architecture for model training and application as shown. A computer device is used for model training. An initial model designed for code recognition, that is, a first key model, is installed on this computer device. After the computer device receives a training sample, it trains the first key model to obtain a second key model. The second key model is a model that can be used for code recognition. After the second key model is installed on the terminal device of the code reviewer, when the terminal device receives the project code to be reviewed, it can use the second key model to perform key sorting on the method code in the project code, which can assist the code reviewer to quickly conduct code review.

[0087] The model training solution and the code recognition solution of the present application can be completed under an integrated development environment (IDE). An IDE is an application program used to provide a program development environment, generally including tools such as a code editor, a compiler, a debugger, and a graphical user interface.

[0088] As can be seen from the above introduction, the solution provided by the embodiments of the present application includes two processes: model training and model application, which will be introduced below with reference to the accompanying drawings respectively.

[0089] I. Model training.

[0090] As Figure 2 shown, an embodiment of the method for model training provided by the embodiments of the present application includes:

[0091] 101. The computer device obtains a plurality of training samples.

[0092] Wherein, each training sample is a combined call graph of a project code with changes. The combined call graph is obtained by combining the code call graph before the project code change and the code call graph after the change. The combined call graph represents multiple method codes included in the project code before and after the change through a plurality of nodes, and represents the call relationship between two method codes among the multiple method codes through edges. The method code corresponding to each node of the combined call graph also corresponds to first key information.

[0093] In the embodiments of the present application, the project code refers to the code written for a project, and the method code is the code for various functions written to complete the functions of the project. The method code can be divided into key method codes and non-key method codes. The key method code refers to the code written to complete the key logic related to computing, and the non-key method code refers to the code written to assist or cooperate with the key method. It can also be said that the key degrees of the method codes are different. A project code will include multiple method codes.

[0094] In the embodiments of the present application, the project code with changes refers to the project code that has been modified. Once the project code is modified, code review needs to be performed again. Code review refers to the process of systematically checking the source code during the software development process. The general purpose is to find various defects, including code defects, function implementation problems, coding rationality, performance optimization, etc., to ensure the overall quality of the software.

[0095] In the embodiments of the present application, the code call graph refers to the call graph obtained by using the Doxygen tool to analyze the method calls of the project code. The code call graph includes nodes and edges. Each node represents a method code, and the edge represents the call relationship between two method codes with a scheduling relationship. When a project changes, it usually means that some code is modified, and there will be some code that is not modified. In this way, based on the unmodified code, the code call graph before the change and the code call graph after the change can be combined to obtain a combined call graph. The combined call graph will include the nodes and edges corresponding to all the method codes before the change, as well as the nodes and edges corresponding to the method codes of the modified part after the change.

[0096] In the embodiments of the present application, the first key information may be the value of the criticality of the method code, and this value may be represented in a normalized form, such as: 1, 0.9, 0.8 or other numerical values, which are not limited in the present application. The first key information of the method code corresponding to each node in the combined call graph can be marked by experienced programmers.

[0097] In the embodiments of the present application, during the model training process, each training sample can be used as a batch for training.

[0098] 102. For each training sample, the computer device converts each of the multiple method codes into a first set of path-contexts.

[0099] The first set includes multiple path-contexts, where each path-context represents the path information between any two leaf nodes and the path in between after the method code is converted into an abstract syntax tree.

[0100] In the embodiments of the present application, each method code can be converted into an abstract syntax tree. The abstract syntax tree includes a parent node and leaf nodes. The leaf nodes of the upper layer can be used as the parent nodes of the leaf nodes of the lower layer. A path can be formed between any two leaf nodes through their parent nodes. In this way, when there are n leaf nodes, there will be n*(n - 1) / 2 paths, and each pair of leaf nodes and the path between them can be called a path-context. Therefore, each method code can be converted into a first set of path-contexts. The first set contains the path-contexts composed of any two leaf nodes and the path in between in the abstract syntax tree converted from the method code.

[0101] Optionally, in step 102, the first key model can also be used to convert each of the multiple method codes into a first set of path-contexts. This is equivalent to integrating more functions into the model and improving the capabilities of the model.

[0102] 103. The computer device trains a first criticality model based on the first set corresponding to each method code among multiple method codes and the first criticality information corresponding to each method code to obtain a second criticality model.

[0103] Among them: The first criticality model includes a first layer, a second layer, and a third layer. The first layer is used to convert the first set into a first vector representation. The second layer is used to process the first vector representation of each method code in combination with the call relationship between multiple method codes to obtain a second vector representation of each method code. The third layer is used to determine the second criticality information of each method code according to the second vector representation of each method code, and supervise the second criticality information according to the first criticality information of each node in the combined call graph to optimize the parameters in the first criticality model; the second criticality model is used to output the criticality information of each method code in the target project code to be reviewed or the criticality sorting information of multiple method codes in the target project code.

[0104] In the embodiment of the present application, the first criticality model can perform vectorization processing on the first set, so that a first vector representation of a single method code can be obtained. Then, in combination with the call relationship between multiple method codes, the first vector representation is processed to obtain a second vector representation for finally representing each method code. Furthermore, the second criticality information of each method code can be determined according to the second vector representation. The form of the second criticality information can be understood by referring to the first criticality information.

[0105] In the embodiment of the present application, the process of optimizing the parameters in the first criticality model can be to use the gradient descent algorithm to optimize the parameters, and finally obtain a second criticality model that can be used for code recognition through the training of multiple training samples.

[0106] In the embodiment of the present application, during the model training process, for the method codes in each project code, path-context conversion is performed, and then vectorization processing is performed on the first set of path-context of the method code, which refines the granularity of the training samples. The trained model will also have higher accuracy in code recognition. In addition, when determining the vector of the method code, the call relationship between different method codes is also combined, considering the relevance between different method codes, further improving the training accuracy of the model, and thus further improving the accuracy of the model in code recognition. In this way, the criticality of the code can be determined by this model in the code review link, improving the code review efficiency.

[0107] As can be seen from the above, to train the model, it is necessary to first collect training samples, that is, the combined call graphs of the project codes that have undergone changes, and label the first key information on the nodes corresponding to the method codes in the combined call graph. Then, process the multiple method codes corresponding in the combined call graph to obtain the first set of each method code, and further use this first set to obtain the vector representation for model training to complete the model training process. The following will introduce these several stages separately.

[0108] 1. Collect training samples and label the first key information.

[0109] This process can be referred to Figure 3 for understanding. As Figure 3 shown, this process can include:

[0110] 201. Obtain multiple historical project codes that have undergone changes.

[0111] The historical project codes can be Java project codes on the source website.

[0112] 202. For each historical project code, obtain the file content before the change and the file content after the change.

[0113] 203. Conduct call analysis on the file content before the change to obtain the first call graph.

[0114] The first call graph includes nodes and edges before the change. The nodes represent method codes, and the edges represent the call relationships between two associated method codes.

[0115] In the embodiments of the present application, the nodes usually include node names and node attributes. The node names are usually used to indicate a certain method code, and the node attributes represent the content of the method code. In the present application, no distinction is made between the node names and node attributes. If any one of the node names and node attributes changes, it is considered that the method code has changed. If neither the node name nor the node attribute has changed, it is considered that the method code has not changed.

[0116] 204. Conduct call analysis on the file content after the change to obtain the second call graph.

[0117] The second call graph includes nodes and edges after the change. The nodes represent method codes, and the edges represent the call relationships between two associated method codes.

[0118] 205. Based on the unchanged content in the first call graph and the second call graph, merge the first call graph and the second call graph to obtain the combined call graph.

[0119] The content that has not changed refers to that the content of the method code has not been modified, and the node names and node attributes of the corresponding nodes in the first call graph and the second call graph have not changed.

[0120] As can be seen from Figure 3 , there is a part in the first call graph and the second call graph that has not changed, and a part that has changed. Based on the unchanged part, the first call graph and the second call graph can be combined to obtain Figure 3 the combined call graph shown in

[0121] After obtaining the combined call graph through the above steps, experienced programmers can label the first key information for each node in combination with the functions of each method code.

[0122] Since multiple training samples need to be collected, this process can be understood as a process of continuously adding from an empty set. The following describes this process through a piece of logic.

[0123] Input: The set T of historical project codes;

[0124] Output: The set C of combined call graphs;

[0125] ( indicating that the set C is an empty set in the initial state);

[0126] for t in T do (indicating for a historical project code t in the set T of historical project codes);

[0127] Use the git tool to obtain the set S of commits of the historical project code t at different times;

[0128] for s in S do (indicating for a commit code s in the set S of commits, s is the historical project code committed at a certain time);

[0129] If the content changed by the commit code s is very small, it can be skipped;

[0130] If the content changed by the commit code s meets the processing conditions, extract the file F1 before the change and the file F2 after the change from the commit code s;

[0131] r1 = Doxygen(F1);

[0132] r2 = Doxygen(F2);

[0133] c = Merge(r1, r2);

[0134] add c into C;

[0135] return C。

[0136] In the above logic, the Doxygen function represents the use of the Doxygen tool to analyze the call graph. The Merge function represents the merging of the call graphs of two versions before and after the change with the unchanged function as the common node. c finally represents a directed graph, where each node in the graph corresponds to the method code in the changed file, and the edges in the graph represent the call relationships between the method codes. Each node has an attribute, and the attribute is the content of the method code.

[0137] After obtaining the methods and their call relationships in the code submission, data needs to be sorted out. This process marks the first key information for the method representatives by experienced programmers.

[0138] 2. Process the method codes to obtain the first set of path-contexts of the method codes.

[0139] In the embodiments of the present application, as Figure 4 shown, the process of obtaining the first set of path-contexts from the method codes includes:

[0140] 301. Convert each of the multiple method codes into an abstract syntax tree.

[0141] The abstract syntax tree includes multiple leaf nodes, and a path is formed between any two of the multiple leaf nodes and the parent node of the two leaf nodes.

[0142] As Figure 5 shown, by parsing the method code, the method code can be converted into an abstract syntax tree. The node at the first level in this abstract syntax tree is the root node, and other nodes are all leaf nodes. The node at the upper level is the parent node of the leaf nodes at the lower level. In fact, the abstract syntax tree may also contain non-leaf nodes, and one or more non-leaf nodes may be included in the path formed by two leaf nodes through the parent node.

[0143] In this way, if there are n leaf nodes in the abstract syntax tree, n*(n - 1) / 2 combinations can be obtained by combining them in pairs, and n*(n - 1) / 2 path-contexts can be obtained.

[0144] 302. Determine any two leaf nodes and the path between any two leaf nodes as a path-context.

[0145] The first set includes multiple path-contexts obtained by combining all the leaf nodes in the abstract syntax tree in pairs.

[0146] Step 302 may include: performing word segmentation on any two leaf nodes and two leaf nodes in the path between any two leaf nodes according to the naming rule, so as to obtain a path-context including any two leaf nodes and the path between any two leaf nodes.

[0147] The process of determining the path-context can be understood in combination with the example of the code "numberOfPath = 7". numberOfPath = 7 is an assignment statement, which includes three parts: name (NameExpr), assignment relationship (AssignExpr), and assigned value (IntergerLiteralExpr). The name is numberOfPath, and the value assigned to this name is 7. Therefore, after parsing this assignment statement, an abstract syntax tree as shown in Figure 6 can be obtained. In this abstract syntax tree, numberOfPath and 7 are leaf nodes, and NameExpr, AssignExpr, and IntergerLiteralExpr are non-leaf nodes. In this way, the path-context formed by the two leaf nodes numberOfPath and 7 and the intermediate path can be expressed as:

[0148] "〈numberOfPath,(NameExpr↑AssignExpr↓IntergerLiteralExpr),7〉", where ↑ and ↓ represent the direction of the path.

[0149] However, this representation cannot obtain the semantic information that may be contained in the leaf nodes. For example, the leaf node numberOfPath is actually "number of path". Based on this, in this application, word segmentation is performed on the leaf nodes according to the naming rule, and finally its "path-context" is expressed as:

[0150] "〈(number of path),(NameExpr↑AssignExpr↓IntergerLiteralExpr),(7)〉".

[0151] In the embodiments of this application, by using the abstract syntax tree to transform the path-context, the speed of path-context transformation can be improved. In addition, performing word segmentation on the leaf nodes according to the naming rule can substitute the function information and semantic information in the code into the vector representation of the path-context for subsequent model training, which is beneficial to improving the accuracy of the model.

[0152] 3. Vectorize the first set to obtain a vector representation for model training for model training.

[0153] In the embodiments of the present application, the process of vectorizing and model training for the first combination can be referred to Figure 7 for understanding. As Figure 7 shown, this process includes:

[0154] 401. Vectorize each path-context in the first set to obtain a third vector representation of each path-context.

[0155] In the embodiments of the present application, the third vector representation refers to the vector representation for each path-context.

[0156] In the embodiments of the present application, a path-context includes two leaf nodes and the path between the two leaf nodes. A leaf node can be understood as a sub-word string, including multiple sub-words. The vector representation of each sub-word can be retrieved from the vocabulary of sub-words. The vector representation of a leaf node can be determined by first summing and then averaging the vector representations of each sub-word in the sub-word string. The vector representation of the path between the two leaf nodes can be found in the path vocabulary. In this way, by concatenating the vector representations of the two leaf nodes and the intermediate path, the third vector representation of the path-context can be obtained.

[0157] During the vectorization process, both the sub-words and the paths can be randomly initialized to a d-dimensional vector representation.

[0158] The vocabulary X of the sub-words described above can be expressed as: value_vocab ∈ |X|×d , and the path vocabulary can be expressed as: path_vocab ∈ |P|×d .

[0159] In the embodiments of the present application, taking the leaf node x i including k sub-words (x i,1 , x i,2 , …, x i,k ) as an example, where the vector representation of the j-th sub-word x i,j can be expressed as:

[0160] In this way, the leaf node x i using the mean of k sub-words as its vector representation can be expressed as:

[0161] The vector representation of the path p i can be directly retrieved from the path vocabulary, that is: embedding(p i ) = path_vocab pi ∈ R d .

[0162] In this way, by concatenating the vector representations of two leaf nodes and the vector representation of the path, a third vector representation of the path-context can be obtained. The third vector representation of the i-th path-context can be expressed as: c i ∈R 3d .

[0163] If the i-th path-context is the path-context formed by the s-th leaf node and the t-th leaf node, and the path between the two leaf nodes is path p i , then this path c i can be expressed as:

[0164] c i =embedding(<x s ,p j ,x t )=[embedding(x s );embedding(p j );embedding(x t )].

[0165] 402. Determine the first vector representation corresponding to the first set according to the third vector representation of each path-context.

[0166] This step 402 may include: performing dimensionality reduction on the third vector representation to obtain a vector representation after dimensionality reduction; performing a summation process on the product of the vector representation after dimensionality reduction corresponding to each path-context in the first set and the corresponding attention magnitude, to obtain the first vector representation corresponding to the first set, where the attention magnitude is determined by the vector representation after dimensionality reduction and the attention vector.

[0167] Combined with the derivation process in 401 above, c i ∈R 3d It can be seen that the dimension of this c i is 3d, with a relatively high dimension. A fully connected layer can be used to perform dimensionality reduction on c i , and reduce c i to where W ∈ R d×3d is a trainable parameter.

[0168] In this way, the first vector representation v of the first combination including m path-contexts can be determined through the vector representation after dimensionality reduction and the attention magnitude a i . The first vector representation v can be expressed as: where where represents transpose of Denote the attention vector, where \(u\in\mathbb{R}^d\).

[0169] 403. Determine the second vector representation of each method code according to the first vector representation of each method code and the call relationship between multiple method codes in the combined call graph.

[0170] The above can be the vector representation of a single method code. However, usually there are call relationships between method codes. Therefore, only using the vector representation of a single method code will affect the accuracy of model training. In the embodiments of the present application, in combination with the call relationships between multiple method codes, the second vector representation of the method code is further determined.

[0171] In the embodiments of the present application, the call relationship between method codes is represented by edges in the combined call graph. If there are \(N\) nodes in the combined call graph, then the combined call graph \(A\) is a directed graph, and this directed graph can be represented in the form of an adjacency matrix as \(A\in\mathbb{R}^{}\) N×N In the embodiments of the present application, an undirected graph can be obtained from the directed graph \(A\) The can be expressed as: where \(A^{}\) T represents the transpose of \(A\), and \(I\) N represents the self-loop, which is a diagonal matrix. Then, perform normalization processing on to obtain the normalized matrix where

[0172] The first vector representation of each method code is \(v\), and the attribute matrix of each method code is represented by \(V\) as \(V\in\mathbb{R}^{}\) N×d Then, the second vector representation \(Z\) can be expressed as: where \(W\) (0) and \(W\) (1) \(\in\mathbb{R}^{}\) d×d are the trainable parameters of the model.

[0173] 404. According to the second vector representation of each method code, use the self-attention mechanism to determine the second key information of each method code, and supervise the second key information according to the first key information of each node in the combined call graph to optimize the parameters in the first key model.

[0174] In the embodiments of the present application, the second vector representation \(Z\) can be used to calculate the second key information of the method code, and this calculation process can be understood by the following relational expression: where represents the transpose of \(Z\), and \(b\) represents a trainable parameter. i

[0175] After calculating β above, the first key information l marked by the programmer can be used to supervise β, where l ∈ R N , and the gradient descent method is used to optimize various parameters in the model.

[0176] During the above model training process, if a training sample has m "path-contexts" for a method, it can be represented as m d-dimensional vectors. Calculate the attention for the m d-dimensional vectors and fuse them into a d-dimensional vector, which can be used as the vector representation of the method. Each training sample contains k methods, and the call relationships between the methods form an adjacency matrix of size k*k. The vector representations of the k methods form a k*d attribute matrix. Using a graph convolutional neural network for modeling, a k*d result matrix can be obtained as the final embedding representation of each method. Based on this, a mechanism similar to attention is used to calculate the key information, and the relative key ranking of each method can be obtained. One or several method codes with the greatest key importance can be understood as the key method codes. During training, for a training sample, if the number of method codes is less than k, padding is performed, and if the number of "path-contexts" in a single method code is less than m, padding is also performed. The key information calculated by the final model is supervised by the key method annotations obtained in the data collection phase, the cross-entropy between the two is calculated, and then the gradient descent algorithm is used to optimize the parameters in the model.

[0177] The above introduced the process of model training. Next, the structure of the first key model or the second key model (abbreviation: key model) will be introduced in combination with the above text, as Figure 8 shown. The structure of the key model provided by the embodiments of the present application may include a first layer, a second layer, and a third layer.

[0178] The first layer includes a path-context vectorization layer, a fully connected layer, and a self-attention layer. The path-context vectorization layer can perform the above 401, and the fully connected layer and the self-attention layer can perform the corresponding content introduced in the above 402.

[0179] The second layer includes two graph convolutional layers. Of course, in the embodiments of the present application, it is not limited to two graph convolutional layers, and it can also be one graph convolutional layer or other more graph convolutional layers. The two graph convolutional layers can be used to perform the corresponding content introduced in the above step 403.

[0180] The third layer includes a self-attention layer and a key determination layer. The self-attention layer and the key determination layer can be used to perform the corresponding content introduced in the above step 404.

[0181] The above introduced the process of model training. Next, the process of model application will be introduced in combination with the accompanying drawings.

[0182] II. Model Application.

[0183] In the embodiments of the present application, after the project code developed by the programmer or the modified project code, it needs to be submitted to the configuration library, and then a code review request is initiated. The code reviewer will review the project code in the configuration library, give review opinions, and then the programmer will modify it. Finally, the code will confirm the modification result. After confirming that there is no problem with the modification result, the modified project code can be used.

[0184] Combined with the aforementioned trained target criticality model in the embodiments of the present application, the above process in the embodiments of the present application can be referred to Figure 9 for understanding. As Figure 9 shown, this process includes:

[0185] 501. After the programmer modifies the project code, submit the modified project code to the configuration library.

[0186] 502. Process the project code and use the target criticality model trained by the above model training method to determine the criticality information or criticality ranking of the method code in the project code.

[0187] 503. The reviewer puts forward review opinions according to the criticality information or criticality ranking of the method code.

[0188] 504. The programmer modifies the project code according to the review opinions of the reviewer.

[0189] 505. The reviewer confirms the modification result.

[0190] In the embodiments of the present application, before the reviewer reviews the project code, the project code is first analyzed intelligently to give the criticality information or criticality ranking of the method code in the project code, which improves the efficiency of the reviewer's understanding of the project code, thereby improving the efficiency of code review.

[0191] The above process of the reviewer reviewing the project code can be carried out on the reviewer's terminal device. Among them, the execution of step 502 can be carried out on the terminal device or can be completed by the terminal device in combination with the server. Below, taking the scenario of the terminal device combined with the server as an example, the content executed by the corresponding device is introduced.

[0192] As Figure 10As shown, in the scenario where a terminal device is combined with a server, a client is installed on the reviewer's terminal device. The client includes a code review user interface 601 and a code review plugin 602. The code review user interface 601 includes a code review interface function entry 6011, and the code review plugin 602 includes a code review service call module 6021. A code recognition service module 701 is configured on the server. The code recognition service module 701 includes a code processing module 7011, a first vectorization module 7012, a second vectorization module 7013, and a criticality sorting module 7014.

[0193] Before reviewing the project code, the reviewer can first enter the code review user interface 601 and click the code review interface function entry 6011. The code review interface function entry 6011 will call the code review service call module 6021 in the code review plugin 602. The code review service call module 6021 will call the code recognition service module 701 in the server. Then, the code processing module 7011 will process the project code to be reviewed by the reviewer to obtain a first set of path-contexts of the project code. The first vectorization module 7012 will process the first set to obtain a first vector representation. The second vectorization module 7013 will determine a second vector representation in combination with the first vector representation and the call relationships between different method codes. The criticality sorting module 7014 will determine the criticality information of the method codes according to the second vector representation, and sort the criticality information of multiple method codes, and return a list of criticality sorting to the terminal device.

[0194] Figure 10 Shown is the scenario where the terminal device and the server are combined for code recognition. If the terminal device performs the above functions by itself, the code recognition service module 701 can be configured on the terminal device. The specific execution process is basically the same as the above process and will not be elaborated here.

[0195] The above briefly describes Figure 10 the process of code recognition. Next, the method for code recognition provided by the embodiments of the present application will be described in combination with Figure 11 As shown, an embodiment of the method for code recognition provided by the embodiments of the present application includes: Figure 11

[0196] 801. The computer device receives the target project code to be reviewed.

[0197] The target project code has changed.

[0198] 802. The computer device determines the corresponding combined call graph according to the target project code.

[0199] ​The combined call graph is obtained by combining the call graph of the target project's code before the change and the call graph of the code after the change. The combined call graph represents multiple method codes included in the target project's code before and after the change through multiple nodes, and represents the call relationship between two method codes among the multiple method codes through edges.

[0200] 803. The computer device converts each method code among the multiple method codes into a first set of path-contexts.

[0201] The first set includes multiple path-contexts, where each path-context represents the path information between any two leaf nodes and the path in between after the method code is converted into an abstract syntax tree.

[0202] This step 803 can also be to use the target criticality model to convert each method code among the multiple method codes into a first set of path-contexts.

[0203] 804. The computer device determines the criticality information of each method code in the target project's code or the criticality sorting information of multiple method codes in the target project's code according to the first set corresponding to each method code among the multiple method codes and the target criticality model.

[0204] The target criticality model includes a first layer, a second layer, and a third layer. Among them, the first layer is used to convert the first set into a first vector representation, the second layer is used to process the first vector representation of each method code in combination with the call relationship between multiple method codes to obtain a second vector representation of each method code, and the third layer is used to determine the criticality information of each method code or the criticality sorting information of multiple method codes according to the second vector representation of each method code.

[0205] The target criticality model in the embodiments of this application can be the second criticality model trained in the above-mentioned model training method section.

[0206] For features that are the same as or similar to those involved in model training in model application, reference can be made to the introduction in the foregoing model training section for understanding, and details will not be repeated here.

[0207] In the embodiments of this application, for the target project's code to be reviewed, first use the target criticality model to determine the criticality information of each method code in the target project's code or the criticality sorting of multiple method codes in the target project's code. This can assist code reviewers in reviewing the code and improve the efficiency of code review.

[0208] Optionally, in the embodiments of the present application, step 804 above includes: determining the criticality information of each method code in the target project code or the criticality ranking information of multiple method codes in the target project code according to the first set corresponding to each method code in the multiple method codes and the target criticality model, including: vectorizing each path-context in the first set to obtain a third vector representation of each path-context; determining a first vector representation corresponding to the first set according to the third vector representation of each path-context; determining a second vector representation of each method code according to the first vector representation of each method code and the call relationship between multiple method codes in the combined call graph; and determining the criticality information of each method code or the criticality ranking information of multiple method codes in the target project code by using a self-attention mechanism according to the second vector representation of each method code.

[0209] In the embodiments of the present application, the third vector representation refers to the vector representation for each path-context. The first vector representation can be obtained by summing the third vector representations, or by weighting each third vector representation and then summing, or by performing dimensionality reduction on the third vector representation and then weighting and summing the dimensionally reduced third vector representation. The specific way of obtaining the first vector representation from the third vector representation is not limited in the present application. From this possible implementation, it can be seen that through three-layer vectorization processing, the accuracy of the criticality information of each method code determined or the criticality ranking information of multiple method codes in the target project code is improved.

[0210] Optionally, the above step: vectorizing each path-context in the first set to obtain a third vector representation of each path-context includes: performing word segmentation on each leaf node of the two leaf nodes of each path-context, and using the average vector representation of the multiple sub-words obtained after word segmentation of each leaf node as the vector representation of each leaf node; and concatenating the vector representations of the two leaf nodes and the vector representation of the path information between the two leaf nodes to obtain a third vector representation of each path-context.

[0211] Optionally, the above step: determining a first vector representation corresponding to the first set according to the third vector representation of each path-context includes: performing dimensionality reduction on the third vector representation to obtain a dimensionally reduced vector representation; and summing the product of the dimensionally reduced vector representation corresponding to each path-context in the first set and the corresponding attention magnitude to obtain a first vector representation corresponding to the first set, where the attention magnitude is determined by the dimensionally reduced vector representation and the attention vector.

[0212] Optionally, the above steps of converting each method code in multiple method codes into a first set of path-contexts include: converting each method code in the multiple method codes into an abstract syntax tree, where the abstract syntax tree includes multiple leaf nodes, and a path is formed between any two leaf nodes among the multiple leaf nodes and the root node closest to the two leaf nodes; determining any two leaf nodes and the path between the two leaf nodes as a path-context, and the first set includes multiple path-contexts obtained by pairwise combination of all the leaf nodes in the abstract syntax tree.

[0213] Optionally, the above step of determining any two leaf nodes and the path between the two leaf nodes as a path-context includes: performing word segmentation processing on the two leaf nodes in any two leaf nodes and the path between the two leaf nodes according to the naming rule to obtain a path-context including any two leaf nodes and the path between the two leaf nodes.

[0214] Optionally, the above steps of determining the corresponding combined call graph according to the target project code include: for the target project code, obtaining the file content before the change and the file content after the change; performing call analysis on the file content before the change to obtain a first call graph, and performing call analysis on the file content after the change to obtain a second call graph; based on the unchanged content in the first call graph and the second call graph, merging the first call graph and the second call graph to obtain a combined call graph.

[0215] In the embodiments of the present application, for the process of converting method codes into the first set of path-contexts, the third vector representation, the first vector representation, and the second vector representation, and the obtaining process of the combined call graph, reference can be made to the corresponding content in the foregoing model training method section for understanding, and details are not repeated here.

[0216] Above, for the process of combining model training and model application, reference can also be made to Figure 12 for understanding. As Figure 12As shown, multiple historical project codes are obtained, each historical project code is processed to obtain a combined call graph for each historical project code, and the method code corresponding to each node in the combined call graph is processed to obtain a first set of path-contexts for each method code. The first set is used to train a model to obtain a target criticality model that can be applied. For the target project code to be reviewed, first, the target project code is processed to obtain a first set of each method code during the processing of the target code, and then the first set is input into the target criticality model. The target criticality model uses the aforementioned vectorization processing method for the first set to obtain a first vector representation and a second vector representation, and then obtains the criticality information of each method code. The criticality information of each method code can be sorted, and a list of the criticality information of multiple method codes is output, thereby assisting reviewers in quickly reviewing the code.

[0217] The methods of model training and code recognition are introduced above. Next, the model training device and code recognition device provided by the embodiments of the present application are introduced with reference to the accompanying drawings.

[0218] As Figure 13 shown, an embodiment of the model training device 90 provided by the embodiments of the present application includes:

[0219] An acquisition unit 901, configured to acquire multiple training samples, where each training sample is a combined call graph of a project code that has undergone changes. The combined call graph is obtained by combining the code call graph before the project code change and the code call graph after the change. The combined call graph represents multiple method codes included in the project code before and after the change through multiple nodes, and represents the call relationship between two method codes among the multiple method codes through edges. The method code corresponding to each node of the combined call graph also corresponds to first criticality information. The acquisition unit 901 can be used to execute step 101 in the above method embodiment.

[0220] A first processing unit 902, configured to convert each method code among the multiple method codes into a first set of path-contexts for each training sample acquired by the acquisition unit 901. The first set includes multiple path-contexts, where each path-context represents the path information between any two leaf nodes and the middle thereof after the method code is converted into an abstract syntax tree. The first processing unit 902 can be used to execute step 102 in the above method embodiment.

[0221] A second processing unit 903, configured to train a first criticality model based on a first set corresponding to each method code among multiple method codes obtained by the first processing unit 902 and first criticality information corresponding to each method code, so as to obtain a second criticality model; wherein: the first criticality model includes a first layer, a second layer, and a third layer. The first layer is configured to convert the first set into a first vector representation. The second layer is configured to process the first vector representation of each method code in combination with the call relationship between multiple method codes to obtain a second vector representation of each method code. The third layer is configured to determine second criticality information of each method code according to the second vector representation of each method code, and supervise the second criticality information according to the first criticality information of each node in the combined call graph, so as to optimize the parameters in the first criticality model. The second criticality model is configured to output criticality information of each method code in the target project code to be reviewed or criticality ranking information of multiple method codes in the target project code. This second processing unit 903 may be configured to execute step 103 in the above method embodiment.

[0222] In the embodiment of the present application, during the model training process, for each method code in each project code, a path-context conversion is performed, and then a vectorization process is performed on the first set of path-contexts of the method code, so that the granularity of the training samples is refined, and the accuracy of the trained model in code recognition will be higher. In addition, when determining the vector of the method code, the call relationship between different method codes is also combined, and the relevance between different method codes is considered, further improving the accuracy of model training, and thus further improving the accuracy of the model in code recognition. In this way, in the code review link, the criticality of the code can be determined through this model, improving the code review efficiency.

[0223] Optionally, the first processing unit 902 is configured to, for each training sample, use the first criticality model to convert each method code among multiple method codes into a first set of path-contexts.

[0224] Optionally, the second processing unit 903 is configured to: perform vectorization processing on each path-context in the first set to obtain a third vector representation of each path-context; determine a first vector representation corresponding to the first set according to the third vector representation of each path-context; determine a second vector representation of each method code according to the first vector representation of each method code and the call relationship between multiple method codes in the combined call graph; determine second criticality information of each method code by using a self-attention mechanism according to the second vector representation of each method code, and supervise the second criticality information according to the first criticality information of each node in the combined call graph, so as to optimize the parameters in the first criticality model.

[0225] Optionally, the second processing unit 902 is configured to: perform word segmentation on each of the two leaf nodes of each path-context, and use the average vector representation of the multiple sub-words obtained after word segmentation of each leaf node as the vector representation of each leaf node; splice the vector representations of the two leaf nodes and the vector representation of the path information between the two leaf nodes to obtain a third vector representation of each path-context.

[0226] Optionally, the second processing unit 903 is configured to: perform dimensionality reduction processing on the third vector representation to obtain a vector representation after dimensionality reduction; perform a summation process on the product of the vector representation after dimensionality reduction corresponding to each path-context in the first set and the corresponding attention magnitude, to obtain a first vector representation corresponding to the first set, where the attention magnitude is determined by the vector representation after dimensionality reduction and the attention vector.

[0227] Optionally, the first processing unit 902 is configured to: convert each method code among multiple method codes into an abstract syntax tree, where the abstract syntax tree includes multiple leaf nodes, and any two leaf nodes among the multiple leaf nodes and the parent node between the two leaf nodes form a path; determine any two leaf nodes and the path between the any two leaf nodes as a path-context, and the first set includes multiple path-contexts obtained by pairwise combination of all the leaf nodes in the abstract syntax tree.

[0228] Optionally, the first processing unit 902 is configured to: perform word segmentation on the two leaf nodes in any two leaf nodes and the path between the any two leaf nodes according to a naming rule, to obtain a path-context including the any two leaf nodes and the path between the any two leaf nodes.

[0229] Optionally, the obtaining unit is further configured to: obtain multiple historical project codes that have undergone changes; for each historical project code, obtain the file content before the change and the file content after the change; perform call analysis on the file content before the change to obtain a first call graph, and perform call analysis on the file content after the change to obtain a second call graph; based on the unchanged content in the first call graph and the second call graph, merge the first call graph and the second call graph to obtain a combined call graph.

[0230] The apparatus 90 for model training provided by the embodiments of the present application can be understood by referring to the embodiments in the foregoing method part for model training, and will not be repeated here.

[0231] As Figure 14 shown, an embodiment of the apparatus 100 for code recognition provided by the embodiments of the present application includes:

[0232] A receiving unit 1001 is configured to receive a target project code to be reviewed, where the target project code has undergone changes. The receiving unit 1001 can execute step 801 in the above method embodiment.

[0233] A first processing unit 1002 is configured to determine a corresponding combined call graph according to the target project code received by the receiving unit 1001. The combined call graph is obtained by combining the code call graph before the change of the target project code and the code call graph after the change. The combined call graph represents multiple method codes included in the target project code before and after the change through multiple nodes, and represents the call relationship between two method codes among the multiple method codes through edges. The first processing unit 1002 can execute step 802 in the above method embodiment.

[0234] A second processing unit 1003 is configured to convert each method code among the multiple method codes obtained by the first processing unit 1002 into a first set of path-contexts. The first set includes multiple path-contexts, where each path-context represents any two leaf nodes and the path information in between after the method code is converted into an abstract syntax tree. The second processing unit 1003 can execute step 803 in the above method embodiment.

[0235] A third processing unit 1004 is configured to determine the criticality information of each method code in the target project code or the criticality sorting information of multiple method codes in the target project code according to the first set corresponding to each method code among the multiple method codes obtained by the second processing unit 1003 and the target criticality model. The target criticality model includes a first layer, a second layer, and a third layer. The first layer is configured to convert the first set into a first vector representation. The second layer is configured to process the first vector representation of each method code in combination with the call relationship between multiple method codes to obtain a second vector representation of each method code. The third layer is configured to determine the criticality information of each method code or the criticality sorting information of multiple method codes according to the second vector representation of each method code. The third processing unit 1004 can execute step 804 in the above method embodiment.

[0236] In the embodiment of the present application, for the target project code to be reviewed, the target criticality model is first used to determine the criticality information of each method code in the target project code or the criticality sorting of multiple method codes in the target project code. This can assist code reviewers in reviewing the code and improve the efficiency of code review.

[0237] Optionally, the second processing unit 1003 is configured to use the target criticality model to convert each method code among the multiple method codes into a first set of path-contexts.

[0238] Optionally, the third processing unit 1004 is configured to: vectorize each path-context in the first set to obtain a third vector representation of each path-context; determine a first vector representation corresponding to the first set according to the third vector representation of each path-context; determine a second vector representation of each method code according to the first vector representation of each method code and the call relationship between multiple method codes in the combined call graph; and determine the key information of each method code or the key sorting information of multiple method codes in the target project code by using the self-attention mechanism according to the second vector representation of each method code.

[0239] Optionally, the third processing unit 1004 is configured to: perform word segmentation on each of the two leaf nodes of each path-context, and use the average vector representation of multiple sub-words obtained after word segmentation of each leaf node as the vector representation of each leaf node; splice the vector representations of the two leaf nodes and the vector representation of the path information between the two leaf nodes to obtain a third vector representation of each path-context.

[0240] Optionally, the third processing unit 1004 is configured to: perform dimensionality reduction processing on the third vector representation to obtain a vector representation after dimensionality reduction; perform a summation process on the product of the vector representation after dimensionality reduction corresponding to each path-context in the first set and the corresponding attention magnitude, where the attention magnitude is determined by the vector representation after dimensionality reduction and the attention vector, to obtain a first vector representation corresponding to the first set.

[0241] Optionally, the second processing unit 1003 is configured to: convert each method code in multiple method codes into an abstract syntax tree, where the abstract syntax tree includes multiple leaf nodes, and a path is formed between any two leaf nodes among the multiple leaf nodes and the root node closest to the two leaf nodes; determine any two leaf nodes and the path between the two leaf nodes as a path-context, and the first set includes multiple path-contexts obtained by pairwise combination of all leaf nodes in the abstract syntax tree.

[0242] Optionally, the second processing unit 1003 is configured to: perform word segmentation on the two leaf nodes and the path between the two leaf nodes according to a naming rule to obtain a path-context including the two leaf nodes and the path between the two leaf nodes.

[0243] Optionally, the first processing unit 1002 is configured to: obtain the file content before the change and the file content after the change for the target project code; perform call analysis on the file content before the change to obtain a first call graph, and perform call analysis on the file content after the change to obtain a second call graph; based on the unchanged content in the first call graph and the second call graph, merge the first call graph and the second call graph to obtain a combined call graph.

[0244] The apparatus 100 for code recognition provided by the embodiments of the present application can be understood by referring to the embodiments in the foregoing method section for code recognition, and will not be repeated here.

[0245] Figure 15 As shown, it is a schematic diagram of a possible logical structure of a computer device 110 provided by an embodiment of the present application. The computer device 110 may be a device for model training or a device for code recognition. The computer device 110 includes: a processor 1101, a communication interface 1102, a memory 1103, and a bus 1104. The processor 1101, the communication interface 1102, and the memory 1103 are connected to each other through the bus 1104. In the embodiments of the present application, the processor 1101 is used to control and manage the actions of the computer device 110. For example, the processor 1101 is used to execute Figures 3 to 12 the method embodiments for model training or code recognition processes. The communication interface 1102 is used to support the computer device 110 to communicate. The memory 1103 is used to store the program code and data of the computer device 110.

[0246] Among them, the processor 1101 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 1101 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 1104 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 15 only a thick line is shown in, but it does not mean that there is only one bus or one type of bus.

[0247] In another embodiment of the present application, there is also provided a computer-readable storage medium storing computer-executable instructions. When the processor of the device executes the computer-executable instructions, the device executes the above-mentioned Figures 3 to 8 method for model training, or executes the above-mentioned Figures 9 - 12 method for code recognition.

[0248] In another embodiment of the present application, there is also provided a computer program product including computer-executable instructions stored in a computer-readable storage medium. When the processor of the device executes the computer-executable instructions, the device executes the above-mentioned Figures 3 to 8 method for model training, or executes the above-mentioned Figures 9 - 12 method for code recognition.

[0249] In another embodiment of the present application, there is also provided a chip system including a processor for implementing the above-mentioned Figures 3 to 8 method for model training, or executing the above-mentioned Figures 9 - 12 method for code recognition. In a possible design, the chip system may further include a memory for storing the necessary program instructions and data for the inter-process communication device. The chip system may be composed of chips or may include chips and other discrete devices.

[0250] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0251] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0252] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0253] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0254] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0255] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

Claims

1. A method for model training, characterized in that, Including: Obtain a plurality of training samples, where each training sample is a combined call graph of a project code with changes. The combined call graph is obtained by combining the code call graph before the project code change and the code call graph after the change. The combined call graph represents multiple method codes included in the project code before and after the change through multiple nodes, and represents the call relationship between two method codes among the multiple method codes through edges. The method code corresponding to each node in the combined call graph also corresponds to first key information; For each training sample, convert each method code among the multiple method codes into a first set of path-contexts. The first set includes a plurality of path-contexts, where each path-context represents any two leaf nodes and the path information in between after the method code is converted into an abstract syntax tree; Train the first key model according to the first set corresponding to each method code among the multiple method codes and the first key information corresponding to each method code to obtain a second key model; where: The first key model includes a first layer, a second layer, and a third layer. The first layer is used to convert the first set into a first vector representation. The second layer is used to process the first vector representation of each method code in combination with the call relationship between the multiple method codes to obtain a second vector representation of each method code. The third layer is used to determine the second key information of each method code according to the second vector representation of each method code, and supervise the second key information according to the first key information of each node in the combined call graph to optimize the parameters in the first key model; The second key model is used to output the key information of each method code in the target project code to be reviewed or the key sorting information of multiple method codes in the target project code.

2. The method according to claim 1, characterized in that The step of, for each training sample, converting each method code among the multiple method codes into a first set of path-contexts includes: For each training sample, use the first key model to convert each method code among the multiple method codes into a first set of path-contexts.

3. The method according to claim 1 or 2, characterized in that, The step of training the first key model according to the first set corresponding to each method code among the multiple method codes and the first key information corresponding to each method code includes: Perform vectorization processing on each path-context in the first set to obtain a third vector representation of each path-context; Determine the first vector representation corresponding to the first set according to the third vector representation of each path-context; Determine the second vector representation of each method code according to the first vector representation of each method code and the call relationship between the multiple method codes in the combined call graph; According to the second vector representation of each method code, a self-attention mechanism is used to determine the second key information of each method code, and the second key information is supervised according to the first key information of each node in the combined call graph to optimize the parameters in the first key model.

4. The method according to claim 3, characterized in that, The vectorizing each path-context in the first set to obtain a third vector representation of each path-context includes: Performing word segmentation on each of the two leaf nodes of each path-context, and using the average vector representation of multiple sub-words obtained after word segmentation of each leaf node as the vector representation of each leaf node; Concatenating the vector representations of the two leaf nodes and the vector representation of the path information between the two leaf nodes to obtain a third vector representation of each path-context.

5. The method according to claim 3 or 4, characterized in that, The determining the first vector representation corresponding to the first set according to the third vector representation of each path-context includes: Performing dimensionality reduction processing on the third vector representation to obtain a vector representation after dimensionality reduction; Performing a summation process on the product of the vector representation after dimensionality reduction corresponding to each path-context in the first set and the corresponding attention magnitude to obtain the first vector representation corresponding to the first set, where the attention magnitude is determined by the vector representation after dimensionality reduction and the attention vector.

6. The method according to any one of claims 1-5, characterized in that, The converting each method code in the multiple method codes into a first set of path-contexts includes: Converting each method code in the multiple method codes into an abstract syntax tree, where the abstract syntax tree includes multiple leaf nodes, and a path is formed between any two of the multiple leaf nodes and the parent node of the two leaf nodes; Determining any two leaf nodes and the path between the any two leaf nodes as a path-context, and the first set includes multiple path-contexts obtained by pairwise combination of all leaf nodes in the abstract syntax tree.

7. The method according to claim 6, wherein The determining any two leaf nodes and the path between the any two leaf nodes as a path-context includes: Performing word segmentation on the any two leaf nodes and the path between the any two leaf nodes according to a naming rule to obtain a path-context including the any two leaf nodes and the path between the any two leaf nodes.

8. The method according to any one of claims 1-7, characterized in that, Before obtaining the multiple training samples, the method further includes: Obtaining multiple historical project codes that have undergone changes; For each historical project code, obtaining the file content before the change and the file content after the change; Performing call analysis on the file content before the change to obtain a first call graph, and performing call analysis on the file content after the change to obtain a second call graph; Based on the unchanged content in the first call graph and the second call graph, merging the first call graph and the second call graph to obtain the combined call graph.

9. A method for code recognition, characterized in that, including: Receive a target project code to be reviewed, where the target project code has changed; Determine a corresponding combined call graph according to the target project code. The combined call graph is obtained by combining the code call graph before the change of the target project code and the code call graph after the change. The combined call graph represents multiple method codes included in the target project code before and after the change through multiple nodes, and represents the call relationship between two method codes among the multiple method codes through edges; Convert each method code among the multiple method codes into a first set of path-contexts. The first set includes multiple path-contexts. Each path-context represents the path information between any two leaf nodes and the middle part after the method code is converted into an abstract syntax tree; According to the first set corresponding to each method code among the multiple method codes and a target criticality model, determine the criticality information of each method code in the target project code or the criticality sorting information of multiple method codes in the target project code. The target criticality model includes a first layer, a second layer, and a third layer. The first layer is used to convert the first set into a first vector representation, the second layer is used to process the first vector representation of each method code in combination with the call relationship between the multiple method codes to obtain a second vector representation of each method code, and the third layer is used to determine the criticality information of each method code or the criticality sorting information of the multiple method codes according to the second vector representation of each method code.

10. The method according to claim 9, wherein The step of converting each method code among the multiple method codes into a first set of path-contexts includes: Using the target criticality model, convert each method code among the multiple method codes into a first set of path-contexts.

11. The method according to claim 9 or 10, characterized in that, The step of determining the criticality information of each method code in the target project code or the criticality sorting information of multiple method codes in the target project code according to the first set corresponding to each method code among the multiple method codes and the target criticality model includes: Perform vectorization processing on each path-context in the first set to obtain a third vector representation of each path-context; Determine the first vector representation corresponding to the first set according to the third vector representation of each path-context; Determine the second vector representation of each method code according to the first vector representation of each method code and the call relationship between the multiple method codes in the combined call graph; According to the second vector representation of each method code, use a self-attention mechanism to determine the criticality information of each method code or the criticality sorting information of multiple method codes in the target project code.

12. The method according to claim 11, wherein The step of performing vectorization processing on each path-context in the first set to obtain a third vector representation of each path-context includes: Perform word segmentation on each of the two leaf nodes of each path-context respectively, and use the average vector representation of the multiple sub-words obtained after word segmentation of each leaf node as the vector representation of each leaf node; Concatenate the vector representations of the two leaf nodes respectively and the vector representation of the path information between the two leaf nodes to obtain the third vector representation of each path-context.

13. The method according to claim 11 or 12, characterized in that The determining the first vector representation corresponding to the first set according to the third vector representation of each path-context includes: Perform dimensionality reduction processing on the third vector representation to obtain the vector representation after dimensionality reduction; Sum the products of the vector representations after dimensionality reduction corresponding to each path-context in the first set and the corresponding attention magnitudes, where the attention magnitudes are determined by the vector representations after dimensionality reduction and the attention vectors, to obtain the first vector representation corresponding to the first set.

14. The method according to any one of claims 9-13, characterized in that, The converting each method code in the multiple method codes into a first set of path-contexts includes: Convert each method code in the multiple method codes into an abstract syntax tree, which includes multiple leaf nodes, and a path is formed between any two leaf nodes among the multiple leaf nodes and the root node closest to the two leaf nodes; Determine any two leaf nodes and the path between the two leaf nodes as a path-context, and the first set includes multiple path-contexts obtained by pairwise combination of all the leaf nodes in the abstract syntax tree.

15. The method according to claim 14, wherein The determining any two leaf nodes and the path between the two leaf nodes as a path-context includes: Perform word segmentation on the two leaf nodes and the path between the two leaf nodes according to the naming rules to obtain a path-context including the two leaf nodes and the path between the two leaf nodes.

16. The method according to any one of claims 9-15, characterized in that, The determining the corresponding combined call graph according to the target project code includes: For the target project code, obtain the file content before the change and the file content after the change; Perform call analysis on the file content before the change to obtain a first call graph, and perform call analysis on the file content after the change to obtain a second call graph; Based on the unchanged content in the first call graph and the second call graph, merge the first call graph and the second call graph to obtain the combined call graph.

17. An apparatus for model training, characterized in that, including: An acquisition unit for acquiring a plurality of training samples, where each training sample is a combined call graph of project codes that have undergone changes. The combined call graph is obtained by combining the code call graph before the project code change and the code call graph after the project code change. The combined call graph represents a plurality of method codes included in the project code before and after the change through a plurality of nodes, and represents the call relationship between two method codes among the plurality of method codes through edges. Each method code corresponding to a node in the combined call graph also corresponds to first key information; A first processing unit for, for each training sample acquired by the acquisition unit, converting each method code among the plurality of method codes into a first set of path-contexts. The first set includes a plurality of path-contexts, where each path-context represents any two leaf nodes and the path information in between after the method code is converted into an abstract syntax tree; A second processing unit for training the first key model according to the first set corresponding to each method code among the plurality of method codes obtained by the first processing unit and the first key information corresponding to each method code, to obtain a second key model; where: The first key model includes a first layer, a second layer, and a third layer. The first layer is used to convert the first set into a first vector representation. The second layer is used to process the first vector representation of each method code in combination with the call relationship between the plurality of method codes to obtain a second vector representation of each method code. The third layer is used to determine the second key information of each method code according to the second vector representation of each method code, and supervise the second key information according to the first key information of each node in the combined call graph to optimize the parameters in the first key model; The second key model is used to output the key information of each method code in the target project code to be reviewed or the key sorting information of a plurality of method codes in the target project code.

18. The apparatus according to claim 17, wherein The first processing unit is used to, for each training sample, use the first key model to convert each method code among the plurality of method codes into a first set of path-contexts.

19. The apparatus according to claim 17 or 18, wherein The second processing unit is used to: Perform vectorization processing on each path-context in the first set to obtain a third vector representation of each path-context; Determine the first vector representation corresponding to the first set according to the third vector representation of each path-context; Determine the second vector representation of each method code according to the first vector representation of each method code and the call relationship between the plurality of method codes in the combined call graph; According to the second vector representation of each method code, use a self-attention mechanism to determine the second critical information of each method code, and supervise the second critical information according to the first critical information of each node in the combined call graph to optimize the parameters in the first criticality model.

20. The apparatus according to claim 19, wherein the second processing unit is configured to: perform word segmentation processing on each of the two leaf nodes of each path-context, and use the average vector representation of the multiple sub-words obtained after word segmentation processing of each leaf node as the vector representation of each leaf node; concatenate the vector representations of the two leaf nodes respectively and the vector representation of the path information between the two leaf nodes to obtain a third vector representation of each path-context.

21. The apparatus according to claim 19 or 20, wherein the second processing unit is configured to: perform dimensionality reduction processing on the third vector representation to obtain a vector representation after dimensionality reduction; perform a summation process on the product of the vector representation after dimensionality reduction corresponding to each path-context in the first set and the corresponding attention magnitude, to obtain the first vector representation corresponding to the first set, where the attention magnitude is determined by the vector representation after dimensionality reduction and an attention vector.

22. The apparatus according to any one of claims 17-21, wherein the first processing unit is configured to: convert each method code in the multiple method codes into an abstract syntax tree, the abstract syntax tree includes a plurality of leaf nodes, and a path is formed between any two leaf nodes among the plurality of leaf nodes and the parent node of the two leaf nodes; determine any two leaf nodes and the path between the any two leaf nodes as a path-context, and the first set includes a plurality of path-contexts obtained by pairwise combination of all leaf nodes in the abstract syntax tree.

23. The apparatus according to claim 22, wherein the first processing unit is configured to: perform word segmentation processing on the two leaf nodes in any two leaf nodes and the path between the any two leaf nodes according to a naming rule to obtain a path-context including the any two leaf nodes and the path between the any two leaf nodes.

24. The apparatus according to any one of claims 17-23, wherein the obtaining unit is further configured to: obtain a plurality of historical project codes that have undergone changes; for each historical project code, obtain the file content before the change and the file content after the change; perform call analysis on the file content before the change to obtain a first call graph, and perform call analysis on the file content after the change to obtain a second call graph; merge the first call graph and the second call graph based on the content that has not changed in the first call graph and the second call graph to obtain the combined call graph.

25. A device for code recognition, characterized in that, Comprising: A receiving unit, configured to receive a target project code to be reviewed, where the target project code has changed; A first processing unit, configured to determine a corresponding combined call graph according to the target project code received by the receiving unit, where the combined call graph is obtained by combining a code call graph before the change of the target project code and a code call graph after the change, and the combined call graph represents multiple method codes included in the target project code before and after the change through multiple nodes, and represents the call relationship between two method codes among the multiple method codes through edges; A second processing unit, configured to convert each method code among the multiple method codes obtained by the first processing unit into a first set of path-contexts, where the first set includes multiple path-contexts, and each path-context represents any two leaf nodes and the path information in between after the method code is converted into an abstract syntax tree; A third processing unit, configured to determine the criticality information of each method code in the target project code or the criticality sorting information of multiple method codes in the target project code according to the first set corresponding to each method code among the multiple method codes obtained by the second processing unit and a target criticality model, where the target criticality model includes a first layer, a second layer, and a third layer, and the first layer is configured to convert the first set into a first vector representation, the second layer is configured to process the first vector representation of each method code in combination with the call relationship between the multiple method codes to obtain a second vector representation of each method code, and the third layer is configured to determine the criticality information of each method code or the criticality sorting information of the multiple method codes according to the second vector representation of each method code.

26. The apparatus according to claim 25, wherein the second processing unit is configured to use the target criticality model to convert each method code among the multiple method codes into a first set of path-contexts.

27. The apparatus according to claim 25 or 26, wherein the third processing unit is configured to: vectorize each path-context in the first set to obtain a third vector representation of each path-context; determine the first vector representation corresponding to the first set according to the third vector representation of each path-context; determine the second vector representation of each method code according to the first vector representation of each method code and the call relationship between the multiple method codes in the combined call graph; determine the criticality information of each method code or the criticality sorting information of multiple method codes in the target project code according to the second vector representation of each method code by using a self-attention mechanism.

28. The apparatus according to claim 27, wherein the third processing unit is configured to: Perform word segmentation on each of the two leaf nodes of each path-context respectively, and use the average vector representation of the multiple sub-words obtained after word segmentation of each leaf node as the vector representation of each leaf node; Concatenate the vector representations of the two leaf nodes respectively and the vector representation of the path information between the two leaf nodes to obtain the third vector representation of each path-context.

29. The device according to claim 27 or 28, wherein The third processing unit is configured to: Perform dimensionality reduction processing on the third vector representation to obtain a vector representation after dimensionality reduction; Perform a summation process on the product of the vector representation after dimensionality reduction corresponding to each path-context in the first set and the corresponding attention magnitude, to obtain the first vector representation corresponding to the first set, where the attention magnitude is determined by the vector representation after dimensionality reduction and the attention vector.

30. The device according to any one of claims 25-29, wherein The second processing unit is configured to: Convert each of the multiple method codes into an abstract syntax tree, the abstract syntax tree includes multiple leaf nodes, and a path is formed between any two leaf nodes among the multiple leaf nodes and the root node closest to the two leaf nodes; Determine any two leaf nodes and the path between the any two leaf nodes as a path-context, and the first set includes multiple path-contexts obtained by pairwise combination of all leaf nodes in the abstract syntax tree.

31. The device according to claim 30, wherein The second processing unit is configured to: perform word segmentation on the two leaf nodes in any two leaf nodes and the path between the any two leaf nodes according to a naming rule, to obtain a path-context including the any two leaf nodes and the path between the any two leaf nodes.

32. The device according to any one of claims 25-31, wherein The first processing unit is configured to: For the target project code, obtain the file content before change and the file content after change; Perform call analysis on the file content before change to obtain a first call graph, and perform call analysis on the file content after change to obtain a second call graph; Based on the unchanged content in the first call graph and the second call graph, merge the first call graph and the second call graph to obtain the combined call graph.

33. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by one or more processors, implements the method according to any one of claims 1-8.

34. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by one or more processors, implements the method according to any one of claims 9-16.

35. A computing device, characterized in that, Comprising one or more processors and a computer-readable storage medium storing a computer program; The computer program, when executed by the one or more processors, implements the method according to any one of claims 1-8.

36. A computing device, characterized in that, Comprising one or more processors and a computer-readable storage medium storing a computer program; When the computer program is executed by the one or more processors, it implements the method according to any one of claims 9-16.

37. A chip system, characterized in that, Comprising one or more processors, the one or more processors being invoked to execute the method according to any one of claims 1-8.

38. A chip system, characterized in that, Comprising one or more processors, the one or more processors being invoked to execute the method according to any one of claims 9-16.

39. A computer program product, characterized in that, Comprising a computer program, the computer program being used to implement the method according to any one of claims 1-8 when executed by one or more processors.

40. A computer program product, characterized in that, Comprising a computer program, the computer program being used to implement the method according to any one of claims 9-16 when executed by one or more processors.

Citation Information

Patent Citations

  • Meta-learning-based code self-adaptive generation method

    CN112114791A

  • Code reconstruction method based on programming context information

    CN113190269A