Code quality automatic evaluation method and related equipment
By generating the token sequence of code, converting it into word vectors and constructing a code representation vector, inputting the feedforward neural network model for quality scoring, it solves the problem that code quality evaluation relies on manual reading in the existing technology, and realizes efficient and accurate automatic code quality evaluation.
Patent Information
- Application Number
- CN202411952175.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, code quality evaluation relies on manual reading, is inefficient and inaccurate, and is affected by personal preferences and subjectivity, making it difficult to achieve efficient and accurate automatic code quality evaluation.
By receiving the target code, generating a token sequence, converting it into a word vector, extracting semantic correlation, building a code representation vector, and inputting a feedforward neural network model for quality scoring, realizing automated code quality evaluation.
Accurate and reliable code quality evaluation is achieved, reducing the time and cost of manual review, improving the frequency and timeliness of reviews, and reducing the impact of subjective preferences.
Smart Images

Figure CN120029627A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of electronic technology, and in particular to a method, device, server, computer-readable storage medium, and computer program product for automatic code quality assessment. Background Art
[0002] At present, the software development model is basically a multi-person collaborative development model, using multi-person code submission platforms based on distributed version control systems such as GitHub and GitLab. In the process of code submission, if the developer writes low-quality code that is not readable, scalable, or elegant, it will increase the difficulty of subsequent code reading and modification. In addition, when evaluating the performance of developers, in addition to evaluating the amount of code, evaluating the quality of the code is also particularly important.
[0003] The current code quality evaluation relies on manual code reading to determine whether the code functions as expected, whether it follows programming specifications, and whether it has the characteristics of high-quality code such as readability and scalability. The quality of the code written by the developer is evaluated by manual scoring. However, manual code review is time-consuming and inefficient, which limits the frequency and timeliness of the review and also brings certain manpower costs to the enterprise. In addition, manual code quality evaluation is usually affected by the reviewer's personal preferences, experience, own professional knowledge and interpretation, which is highly subjective, with coarse evaluation granularity, and it is easy to ignore problems that are not easy to detect in the code, resulting in inaccurate evaluation results. Summary of the invention
[0004] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide a method, device, server, computer-readable storage medium and computer program product for automatic code quality assessment, which can automatically achieve accurate and reliable code quality assessment.
[0005] In a first aspect, an embodiment of the present application provides a method for automatically evaluating code quality, comprising:
[0006] Receiving a target code to be evaluated, and generating a token sequence of the target code;
[0007] Convert the token sequence into a word vector;
[0008] Extracting semantic relevance of adjacent tokens in the word vector to obtain a code representation vector;
[0009] The code representation vector is input into a feedforward neural network model, a quality score is performed according to features extracted from the code representation vector, and an output score is obtained as a code quality score.
[0010] In one embodiment, extracting semantic relevance of adjacent tokens in the word vector includes:
[0011] The semantic relevance of adjacent tokens in the word vector is extracted through a recursive autoencoder.
[0012] In one embodiment, converting the token sequence into a word vector includes:
[0013] The word vector generation model is called to convert the token sequence into a word vector.
[0014] In one embodiment, the training method of the word vector generation model includes:
[0015] Get the original code snippet used for pre-training;
[0016] Preprocessing the original code segment;
[0017] Generate tokens for the preprocessed original code segments to obtain token sequence samples;
[0018] Building a vocabulary based on the token sequence sample;
[0019] Call the pre-built word vector generation model, perform word vector conversion training on the token sequence samples according to the vocabulary, and optimize the model parameters according to the training results.
[0020] In one embodiment, preprocessing the original code segment includes:
[0021] Remove the comments in the original code segment;
[0022] Splitting compound words in the original code segment;
[0023] The variable names and function names in the original code segment are uniformly named.
[0024] In one embodiment, calling a word vector generation model to convert the token sequence into a word vector includes:
[0025] The Word2Vec model is called to convert the token sequence into a word vector.
[0026] In a second aspect, an embodiment of the present application provides a device for automatically evaluating code quality, comprising:
[0027] A token generation unit, used for receiving a target code to be evaluated and generating a token sequence of the target code;
[0028] A word vector conversion unit, used to convert the token sequence into a word vector;
[0029] A code representation unit, used for extracting semantic relevance of adjacent tokens in the word vector to obtain a code representation vector;
[0030] The feature scoring unit is used to input the code representation vector into a feedforward neural network model, perform quality scoring according to the features extracted from the code representation vector, and obtain the output score as the code quality score.
[0031] In a third aspect, an embodiment of the present application provides a server, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the embodiment of the present application when executing the program.
[0032] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the embodiment of the present application.
[0033] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the method described in the embodiment of the present application.
[0034] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0036] Figure 1 A schematic diagram of a process flow of a method for automatically evaluating code quality provided by an embodiment of the present application is shown;
[0037] Figure 2 A schematic diagram of the model structure of the evaluation system provided in an embodiment of the present application is shown;
[0038] Figure 3 The following is a structural block diagram of a device for automatically evaluating code quality provided by an embodiment of the present application;
[0039] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing a server of an embodiment of the present application is shown. DETAILED DESCRIPTION
[0040] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the invention are shown in the accompanying drawings.
[0041] It should be noted that, in the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments. Although the embodiments of the present application provide the method operation instruction steps shown in the following embodiments or drawings, more or fewer operation instruction steps may be included in the method based on routine or no creative labor. In the steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiments of the present application. The method may be executed in the order of the methods shown in the embodiments or drawings or in parallel during the actual processing process or when the device is executed.
[0042] Please refer to Figure 1 , Figure 1 FIG. 1 is a flow chart of a method for automatically evaluating code quality provided by an embodiment of the present application. Figure 1 As shown, the method includes:
[0043] S101, receiving the target code to be evaluated and generating a token sequence of the target code;
[0044] Receive the target code to be evaluated and generate the corresponding token sequence. The method of generating the token sequence can refer to the relevant technology and will not be described here.
[0045] In order to avoid the influence of noise interference in the target code on subsequent evaluation, the target code can be further preprocessed before generating a token sequence for the target code. The preprocessing methods include removing comments in the target code, splitting compound words in the target code, and uniformly naming the variable names and function names in the target code. The corresponding preprocessing can be performed according to the code type in the actual application scenario, and is not limited here.
[0046] S102, converting the token sequence into a word vector;
[0047] The specific implementation method of converting the token sequence into a word vector can refer to related technologies. For example, you can call word vector conversion models such as the bag-of-words model and the TF-IDF model, or you can call encoders such as ELMo (Embeddings from LanguageModels). I will not go into details here.
[0048] S103, extracting semantic relevance of adjacent tokens in the word vector to obtain a code representation vector;
[0049] After modeling the word vector and converting the token of the code snippet into a vector, the code snippet can be represented as x 1 ,x 2 ,x 3 …,x n Because similar tokens in a code segment usually have related semantics or connections, the semantic correlation between adjacent code tokens in the word vector is captured to construct a code representation, which converts the original code token from text information to numerical vector information.
[0050] In this method, code representation is applied to the downstream neural network regression model prediction to predict the quality score of the code segment. It can objectively and finely score the code segment with one click, realizing low-cost, high-efficiency and automated intelligent code evaluation.
[0051] It should be noted that the code representation construction and learning method proposed in this method is not only applicable to the code scoring task of the present invention, but also to other programming language processing tasks, such as code clone detection, code defect detection, and automatic code generation.
[0052] Optionally, a recursive autoencoder can be used to extract semantic relevance of adjacent tokens in the word vector to construct a code representation. Other deep learning / machine learning methods can also be used in this step, but because the special structure of the autoencoder neural network can highly abstract the semantic information of the code token, and the recursive structure can fully learn the associated semantic relationship between code sequences, the recursive autoencoder is a model with a better effect in learning code representation. In addition, referring to its performance in natural language processing tasks, this method can give priority to using a recursive autoencoder to construct code representation and fully learn the hidden patterns in the code.
[0053] To deepen the understanding of the implementation of code representation, here is a specific method of extracting semantic relevance of adjacent tokens in word vectors through recursive autoencoders as follows: Figure 2 The figure shows a schematic diagram of the model structure of an evaluation system, wherein the processing of the code in the first part of the figure can refer to the above introduction, and the scoring mechanism in the third part can refer to the introduction of step S104. Here, the implementation process of the second part of the figure is mainly introduced.
[0054] Each time we select two adjacent word vectors x m , x m+1As the input vector of the autoencoder, the encoder maps the input data to a lower-dimensional latent space, and the decoder extracts the latent vector y from this latent space. m Reconstruct the original data x′ m , x′ m+1 By minimizing the difference between the input vector and the reconstructed vector ||x′ m -x m || 2 +||x′ m+1 -x m+1 || 2 , capturing the important features of the code token. The latent vector y encoded by the trained autoencoder model m And the vector x converted from the original code token m+1 , continue to obtain the latent vector y in the above recursive autoencoder method m+1 , perform this operation recursively, and finally obtain a vector y that can represent the entire code n-1 , which is the code representation vector of this code segment.
[0055] S104: Input the code representation vector into a feedforward neural network model, perform a quality score based on features extracted from the code representation vector, and obtain an output score as a code quality score.
[0056] The obtained code representation vector is input into a multi-layer feedforward neural network (FNN). Each layer of the neural network performs a linear transformation on the input data and applies the activation function ReLU (Rectified Linear Unit) to learn and simulate the nonlinear relationship in the input data. Through a series of linear transformations and nonlinear activations, the neural network converts the low-level features of the input into a more abstract high-level representation. The result is sent to the output layer. The output layer uses a specific activation function to limit the output data between 0 and 1, and obtains the output score as the code quality score.
[0057] Furthermore, the output data can be multiplied by 100 to obtain a numerical value, which can be scored in base 100 to obtain a more fine-grained score for predicting code quality, which is not limited here.
[0058] It should be noted that although the operations of the method of the present invention are described in a particular order in the drawings, this does not require or imply that the operations must be performed in this particular order or that all illustrated operations must be performed to achieve desired results.
[0059] Based on the above introduction, the method provided in this embodiment generates a token sequence of the target code to be evaluated, converts it into a word vector, extracts the semantic relevance of adjacent tokens in the word vector, constructs a code representation, converts the original code token from text information to numerical vector information, and then inputs the code representation vector into a feedforward neural network model for regression prediction, and automatically scores the code quality. In this method, code representation is applied to code segment score prediction, making full use of the grammatical and semantic information contained in the code token sequence, mining the hidden representation contained in the code segment, and predicting the quality score of the code segment. It can accurately score with one click in an objective and fine-grained manner, realizing low-cost, high-efficiency, and automated intelligent code evaluation.
[0060] The above embodiment does not limit the implementation method of converting the token sequence to the word vector. In one embodiment, the word vector generation model can be called to convert the token sequence into a word vector. The word vector generation model captures the semantic relationship between words and provides a low-dimensional representation, and has strong generalization ability. Optionally, the Word2Vec model (a word vector model) can be called to convert the token sequence into a word vector. Calling the Word2Vec model can accurately capture the semantic relationship in the token sequence and generate a fixed-length word vector, which is convenient for subsequent unified data processing.
[0061] For the training of the word vector generation model, this embodiment proposes an efficient training method. Specifically, an implementation step is as follows:
[0062] (1) Obtain the original code segment for pre-training;
[0063] (2) Preprocessing the original code segment;
[0064] (3) Generate tokens from the preprocessed original code segment to obtain token sequence samples;
[0065] (4) Build a vocabulary based on token sequence samples;
[0066] (5) Call the pre-built word vector generation model, perform word vector conversion training on the token sequence samples according to the vocabulary, and optimize the model parameters based on the training results.
[0067] The original code snippets used for pre-training can be obtained from public code libraries, projects, and open source software, without limitation. The original code snippets are preprocessed to build a vocabulary, and a word vector generation model (such as the Word2Vec model) is used to convert the code token vocabulary into a vector model. The parameters of the word vector optimization model, i.e., vector size (dimension), window size, minimum word frequency, etc., are set to obtain the optimal vector representation of the vocabulary.
[0068] It should be noted that the original code segments used for pre-training have been scored and annotated in advance. Each code segment is scored by experienced developers. In order to better train the model, the code scoring and annotation dataset used can use the average of the scores of three people as the score label of the code segment quality, thereby eliminating the influence of subjective factors on the code score. These scores are based on factors such as code complexity, readability, and maintainability, and a score is obtained. This score is the label of the code segment. The above dataset is divided into training set, validation set, and test set according to 3:1:1.
[0069] When training the model, minimizing the mean square error between the input vector and the reconstructed vector of the recursive autoencoder and the sum of the mean square errors of the prediction score and the label score of the FNN can be used as the loss function.
[0070] Among them, the specific preprocessing means used in preprocessing the original code segment is not limited in this embodiment. Optionally, a preprocessing means includes: removing comments in the original code segment; splitting compound words in the original code segment; and uniformly naming the variable names and function names in the original code segment.
[0071] Preprocess the original code segment, remove the comments in the code segment to prevent the introduction of noise during word vector training, split compound words, and convert variable names and function names to the same naming style to facilitate subsequent unified recognition and processing. It should be noted that there is no limit on the order in which these three preprocessing methods are executed, and they can be set accordingly according to the needs of the actual scenario.
[0072] This embodiment provides a device for automatically evaluating code quality. Figure 3 , Figure 3 The following is a structural block diagram of the automatic code quality assessment device, which mainly includes:
[0073] The token generation unit 101 is used to receive the target code to be evaluated and generate a token sequence of the target code;
[0074] A word vector conversion unit 102, used to convert a token sequence into a word vector;
[0075] The code representation unit 103 is used to extract semantic relevance between adjacent tokens in the word vector to obtain a code representation vector;
[0076] The feature scoring unit 104 is used to input the code representation vector into the feedforward neural network model, perform quality scoring according to the features extracted from the code representation vector, and obtain the output score as the code quality score.
[0077] It should be understood that the units described in the above device are similar to those in the reference Figure 1 The steps in the method described above correspond to each other. Therefore, the operations and features described above for the method are also applicable to the device and the units contained therein, and will not be repeated here. The device can be pre-implemented in the browser or other security application of the server, or loaded into the browser or its security application of the server by downloading or the like. The corresponding units in the device can cooperate with the units in the server to implement the solution of the embodiment of the present application.
[0078] For the several units mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units described above can be embodied in one unit. On the contrary, the features and functions of one unit described above can be further divided into multiple units to be embodied.
[0079] It should be noted that for details not disclosed in the automatic code quality assessment device in the embodiment of the present application, please refer to the details disclosed in the above embodiments of the present application, which will not be repeated here.
[0080] Reference below Figure 4 , Figure 4 A schematic diagram of the structure of a computer system suitable for implementing a server of an embodiment of the present application is shown.
[0081] like Figure 4 As shown, the computer system includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage part 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation instructions of the system are also stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0082] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed, so that a computer program read therefrom is installed into the storage section 408 as needed.
[0083] In particular, according to an embodiment of the present application, the above reference flow chart Figure 1 The described process can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program includes a program code for executing the method shown in the flow chart. In such an embodiment, the computer program includes a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 409, and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the above-mentioned functions defined in the system of the present application are executed.
[0084] It should be noted that the computer-readable medium shown in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium such as a computer-readable storage medium that can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0085] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operating instructions of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the aforementioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, the boxes represented by two connections can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operating instruction, or can be implemented with a combination of dedicated hardware and computer instructions.
[0086] The units or modules involved in the embodiments described in the present application may be implemented by software or hardware. The units or modules described may also be arranged in a processor. The names of these units or modules do not, in some cases, constitute limitations on the units or modules themselves.
[0087] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the server described in the above embodiment, or may exist independently without being assembled into the server. The above computer-readable storage medium stores one or more programs, and when the above programs are used by one or more processors to execute the automatic code quality assessment method described in the present application.
[0088] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the aforementioned disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.
Claims
1. A method for automatic code quality assessment, characterized in that: include: Receiving a target code to be evaluated, and generating a token sequence of the target code; Convert the token sequence into a word vector; Extracting semantic relevance of adjacent tokens in the word vector to obtain a code representation vector; The code representation vector is input into a feedforward neural network model, a quality score is performed according to features extracted from the code representation vector, and an output score is obtained as a code quality score.
2. The method according to claim 1, characterized in that Extracting semantic relevance of adjacent tokens in the word vector includes: The semantic relevance of adjacent tokens in the word vector is extracted through a recursive autoencoder.
3. The method according to claim 1, characterized in that Convert the token sequence into a word vector, including: The word vector generation model is called to convert the token sequence into a word vector.
4. The method according to claim 3, characterized in that The training method of the word vector generation model includes: Get the original code snippet used for pre-training; Preprocessing the original code segment; Generate tokens for the preprocessed original code segments to obtain token sequence samples; Building a vocabulary based on the token sequence sample; Call the pre-built word vector generation model, perform word vector conversion training on the token sequence samples according to the vocabulary, and optimize the model parameters according to the training results.
5. The method according to claim 4, characterized in that Preprocessing the original code segment includes: Remove the comments in the original code segment; Splitting compound words in the original code segment; The variable names and function names in the original code segment are uniformly named.
6. The method according to claim 3, characterized in that The calling of the word vector generation model to convert the token sequence into a word vector includes: The Word2Vec model is called to convert the token sequence into a word vector.
7. A device for automatically evaluating code quality, characterized in that: include: A token generation unit, used for receiving a target code to be evaluated and generating a token sequence of the target code; A word vector conversion unit, used to convert the token sequence into a word vector; A code representation unit, used for extracting semantic relevance of adjacent tokens in the word vector to obtain a code representation vector; The feature scoring unit is used to input the code representation vector into a feedforward neural network model, perform quality scoring according to the features extracted from the code representation vector, and obtain the output score as the code quality score.
8. A server comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Front-end code review method and device, electronic equipment and storage medium
CN120523708A
A front-end code review method and device, an electronic device and a storage medium
CN120523708B