Code annotation information processing method and device, computer equipment and storage medium
By detecting and completing the missing comment information of complex codes, the ambiguity caused by the missing comment information is solved, the readability and maintenance of the code are improved, and the development efficiency is improved.
Patent Information
- Application Number
- CN202510154683.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-06
AI Technical Summary
The lack of annotation information in complex code leads to ambiguity, reducing the clarity and comprehensibility of the annotations, and affecting the maintenance of the code.
By obtaining the code, evaluating its complexity, obtaining and preprocessing the annotation information, performing syntactic analysis, and determining and completing the missing parts in the annotation information.
Significantly improve the completeness and accuracy of annotation information of complex code, enhance the readability and maintainability of the code, and reduce the time for developers to manually write and revise annotation information.
Smart Images

Figure CN120104173A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of code detection technology, and in particular to a code annotation information processing method, device, computer equipment and storage medium. Background Art
[0002] In the process of software development and maintenance, programmers usually spend nearly half of their time on code understanding activities. Code comments play an important role in this process. As a basic component of software documentation, comments are crucial to the understanding and maintenance of code. They can not only help programmers quickly understand the functions and logic of the code, but also provide valuable references for subsequent code modification and optimization. However, there is a common problem of missing syntactic structure in comment information, which often causes ambiguity and reduces the clarity and comprehensibility of the comments. This ambiguity may lead to programmers misunderstanding of the intent of the code, especially for complex codes, which in turn has a negative impact on code maintenance. At present, there is no technology that can effectively detect and automatically generate missing comment information in complex codes. Summary of the invention
[0003] Based on this, it is necessary to provide a code annotation information processing method, device, computer equipment and storage medium for effectively detecting and automatically generating missing annotation information in complex codes in response to the above technical problems.
[0004] In a first aspect, a method for processing code annotation information is provided, comprising:
[0005] Get the code;
[0006] Evaluate the complexity of the code and determine the complex code based on the evaluation results;
[0007] Acquire first code annotation information corresponding to the complex code, and preprocess the first code annotation information to obtain target code annotation information;
[0008] Performing syntactic analysis on the target code comment information to obtain syntactic analysis results;
[0009] According to the results of syntax analysis, determine and complete the missing information in the target code comment information.
[0010] Optionally, evaluate the complexity of the code, and based on the evaluation result, determine that the complex code includes:
[0011] Get key complexity metrics for your code;
[0012] According to the attributes of the code, the eigenvalues corresponding to the key complexity indicators are determined, and the eigenvalues corresponding to the key complexity indicators are standardized to obtain the target eigenvalues;
[0013] Get the weights corresponding to key complexity indicators;
[0014] Generate target samples based on the weights and target feature values corresponding to key complexity indicators;
[0015] Input the target sample into the pre-built complexity assessment model to obtain the complexity score of the code. The complexity assessment model includes:
[0016]
[0017] in, represents the complexity score, B represents the number of trees in the random forest, Represents the predicted value of the b-th tree for the target sample x;
[0018] Based on the code complexity score, determine the complexity of the code.
[0019] Optionally, the complexity of the code is evaluated, and based on the evaluation result, it is determined that the complex code also includes:
[0020] A preset threshold is obtained, and in response to a complexity score corresponding to the code being greater than the preset threshold, the code is defined as a complex code.
[0021] Optionally, preprocessing the first code annotation information to obtain the target code annotation information includes:
[0022] Based on the code annotation information classification mechanism, redundant code annotation information in the first code annotation information is filtered out to obtain second code annotation information;
[0023] According to the length distribution of the second code annotation information, the code annotation information in the second code annotation information that meets the preset standard length is screened to obtain the target code annotation information.
[0024] Optionally, the target code comment information is subjected to syntactic analysis, and the syntactic analysis results include:
[0025] Obtain the syntactic structure components of the target code comment information;
[0026] According to a preset grammatical pattern and syntactical components of the target code annotation information, the syntactical structure missing type of the target code annotation information is determined, and the grammatical pattern is used to describe the mapping relationship between the syntactical structure missing type and the syntactical components.
[0027] Optionally, based on the syntactic analysis results, determine and complete the missing information in the target code comment information, including:
[0028] Determine the target sub-model in the neural network model according to the missing type of the syntactic structure of the target code annotation information;
[0029] The target sub-model is used to generate missing information corresponding to the missing type of syntactic structure.
[0030] Optionally, determining and completing missing information in the target code comment information based on the syntactic analysis result also includes:
[0031] The target code annotation information is filled with the missing information generated by the target sub-model to generate complete code annotation information.
[0032] In a second aspect, a code annotation information processing device is provided, comprising:
[0033] Data acquisition module, used to acquire code;
[0034] An evaluation module is used to evaluate the complexity of the code and determine the complex code based on the evaluation result;
[0035] A preprocessing module, used for obtaining first code annotation information corresponding to the complex code, and preprocessing the first code annotation information to obtain target code annotation information;
[0036] An analysis module is used to perform syntactic analysis on the target code annotation information to obtain a syntactic analysis result;
[0037] The information processing module is used to determine and complete the missing information in the target code comment information according to the syntactic analysis results.
[0038] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0039] Get the code;
[0040] Evaluate the complexity of the code and determine the complex code based on the evaluation results;
[0041] Acquire first code annotation information corresponding to the complex code, and preprocess the first code annotation information to obtain target code annotation information;
[0042] Performing syntactic analysis on the target code comment information to obtain syntactic analysis results;
[0043] According to the results of syntax analysis, determine and complete the missing information in the target code comment information.
[0044] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0045] Get the code;
[0046] Evaluate the complexity of the code and determine the complex code based on the evaluation results;
[0047] Acquire first code annotation information corresponding to the complex code, and preprocess the first code annotation information to obtain target code annotation information;
[0048] Performing syntactic analysis on the target code comment information to obtain syntactic analysis results;
[0049] According to the results of syntax analysis, determine and complete the missing information in the target code comment information.
[0050] In a fifth aspect, a computer program product is provided, the computer program product comprising a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0051] Get the code;
[0052] Evaluate the complexity of the code and determine the complex code based on the evaluation results;
[0053] Acquire first code annotation information corresponding to the complex code, and preprocess the first code annotation information to obtain target code annotation information;
[0054] Performing syntactic analysis on the target code comment information to obtain syntactic analysis results;
[0055] According to the results of syntax analysis, determine and complete the missing information in the target code comment information.
[0056] The above-mentioned code comment information processing method, device, computer equipment and storage medium, the method includes: obtaining code; evaluating the complexity of the code, and determining the complex code based on the evaluation result; obtaining the first code comment information corresponding to the complex code, and preprocessing the first code comment information to obtain the target code comment information; performing syntactic analysis on the target code comment information to obtain the syntactic analysis result; according to the syntactic analysis result, determining and completing the missing information in the target code comment information. The present application can significantly improve the completeness and accuracy of the comment information of the complex code by accurately detecting and generating the missing parts in the comment information, making the comment content more coherent and clear, and improving the readability of the complex code; the present application reduces the time for developers to manually write and revise the complex code comment information by automatically detecting and generating the comment information, thereby improving the efficiency of the overall development process. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 An application environment diagram of a code comment information processing method in an embodiment;
[0058] Figure 2A schematic diagram of a flow chart of a method for processing code comment information in one embodiment;
[0059] Figure 3 Another schematic diagram of a process of processing code comment information in one embodiment;
[0060] Figure 4 A schematic diagram of an example of a comment syntax analysis tree of a code comment information processing method in one embodiment;
[0061] Figure 5 It is a structural block diagram of a code comment information processing device in one embodiment;
[0062] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0064] It should be understood that in the description of the present application, unless the context clearly requires otherwise, words such as "include", "comprises", and the like throughout the specification should be interpreted as including rather than being exclusive or exhaustive; that is, as including but not limited to.
[0065] It should also be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, "plurality" means two or more.
[0066] It should be noted that the terms "S1", "S2", etc. are only used for the purpose of describing the steps, and do not specifically refer to the order or sequence, nor are they used to limit the present application. They are only for the convenience of describing the method of the present application, and cannot be understood as indicating the order of the steps. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in this field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.
[0067] The code annotation information processing method provided in this application can be applied to Figure 1In the application environment shown, the terminal 102 communicates with the data processing platform set on the server 104 through the network, wherein the terminal 102 can be but not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0068] In one embodiment, Figure 2 As shown, a code annotation information processing method is provided, and the method is applied to Figure 1 The terminal in is used as an example to illustrate, including the following steps:
[0069] S1: Get the code.
[0070] It should be noted that code refers to a set of instructions written by programmers, written in a specific programming language, which can be understood and executed by computers; the codes and their corresponding annotation information in this application are mainly obtained from the open source code hosting platform GitHub, such as Figure 3 As shown, the specific acquisition steps include: using Java Development Tools (JDT) to extract annotation information and corresponding code pairs from multiple GitHub projects, wherein JDT (Java Development Tools) is a set of APIs provided by Eclipse for programming operations on Java resources, such as creating projects, generating Java source code, executing builds, or detecting problems in the code.
[0071] S2: Evaluate the complexity of the code and determine the complex code based on the evaluation results.
[0072] It should be noted that complex code refers to a code block with higher difficulty involving multiple logical relationships; complex code is defined through a multi-indicator comprehensive evaluation method, which can include multiple key indicators such as the cyclomatic complexity of the code, the number of lines of code, the depth and breadth of the AST (abstract syntax tree), etc.
[0073] In some specific embodiments, Figure 3 As shown, the complexity of the code is evaluated, and based on the evaluation results, it is determined that the complex code includes:
[0074] Obtain the key complexity indicators of the code, where the key complexity indicators refer to indicators that can effectively measure the complexity of the code, including: (1) Cyclomatic complexity: Calculate the number of independent paths in the control flow graph of the code to reflect the logical complexity of the code. The increase in cyclomatic complexity usually means that the code is more difficult to test and maintain; (2) Number of lines of code: The number of lines of code is a direct measure of the code size, which usually reflects the complexity of the code implementation. A larger code size often means more branching logic and dependencies; (3) AST (Abstract Syntax Tree, which is a tree-like representation of the abstract syntax structure of the source code) depth and breadth. The AST depth reflects the degree of nesting of the code structure. Deeper code usually means more complex function calls or nested logic. The AST breadth reflects the degree of branching and dispersion of the code. A wider AST structure means more parallel dependencies and functional modules, which increases the difficulty of code reading and maintenance;
[0075] According to the attributes of the code, the eigenvalues corresponding to the key complexity indicators are determined, and the eigenvalues corresponding to the key complexity indicators are standardized to obtain the target eigenvalues, wherein the attributes of the code are the specific cyclomatic complexity, number of lines of code, and AST depth and breadth of the code snippet to be evaluated. The specific eigenvalues corresponding to the code attributes can be determined, such as the number of lines of code is 1000 lines, etc. The data normalization method is Min-Max normalization. Specifically, since the value ranges of various complexity indicators may vary greatly (for example, cyclomatic complexity is usually an integer value, while the number of lines of code may be tens of thousands), in order to unify the influence of various complexity indicators in the fitting process, the original data must be standardized. This application adopts the Min-Max normalization method to linearly convert values in different ranges into a unified interval, that is, 0 to 1, so as to ensure that the weights of various complexity indicators can be reasonably applied in the fitting process. Since the normalization method is a commonly used method, the process of standardizing the data based on the normalization method is not repeated here.
[0076] Obtain the weights corresponding to key complexity indicators. Different complexity indicators have different impacts on code complexity. Therefore, it is necessary to assign appropriate weights to each indicator. You can assign weights to each indicator based on project experience. For example, cyclomatic complexity accounts for 45% of the weight, the number of lines of code accounts for 25% of the weight, AST depth accounts for 20% of the weight, and AST breadth accounts for 10% of the weight. The weight allocation method for each indicator can also include:
[0077] Obtain the indicator weight allocation result in the previous period of the current time, and determine the fluctuation coefficient of the corresponding weight of the target indicator in the indicator weight allocation result. The length of the time period can be set according to actual needs, such as 1 hour, etc. The fluctuation coefficient refers to the number of weight fluctuations at multiple time nodes. As long as the weights of different time nodes change, fluctuations occur. For example, if the weight of the first time node is 25% and the weight of the second time node is 26%, fluctuations occur, and the number of fluctuations + 1, and so on, which will not be repeated;
[0078] In response to the sum of the weight fluctuation coefficients within the time period being greater than a first preset value and the number of indicators that fluctuate within the time period being greater than a second preset value, obtaining the indicator weight allocation results of the target indicator within multiple historical time periods to determine multiple weight values of the target indicator, wherein the first preset value and the second preset value can be set according to actual needs;
[0079] In response to the number of occurrences of the target weight value of the target indicator being greater than the third preset value and the absolute value of the difference between the target weight value and the corresponding weight values of the indicator at two previous and subsequent time points being less than a fourth preset value, marking the target weight value, wherein the third preset value and the fourth preset value can be set according to actual needs;
[0080] In response to the target indicator appearing at the current time node, the marked target weight value is defined as the weight value of the target indicator.
[0081] Among them, through the above rules, the weight of each indicator can be accurately and efficiently allocated, thereby improving the accuracy of code complexity evaluation.
[0082] Generate a target sample based on the weight corresponding to the key complexity index and the target feature value, wherein the target sample is the product of the weight corresponding to the key complexity index and the target feature value;
[0083] Input the target sample into the pre-built complexity assessment model to obtain the complexity score of the code. The complexity assessment model includes:
[0084]
[0085] in, represents the complexity score, that is, the final prediction value of the random forest for the target sample x, which is the average of the prediction values of all trees. B represents the number of trees in the random forest. represents the prediction value of the b-th tree for the target sample x, that is, the probability that the current input is a complex code;
[0086] Based on the complexity score of the code, the complexity of the code is determined, wherein the higher the complexity score of the code, the higher the complexity of the code.
[0087] Specifically, after the data is standardized, all standardized data are fitted to generate a comprehensive complexity score. This application uses random forest regression as a fitting algorithm to construct a corresponding complexity assessment model, and its expression is shown in the above formula. The training process of the complexity assessment model includes: first, the product of the complexity index after standardization and the weight is used as the feature input, and the complexity score of the complex code snippet manually marked by the expert is used as the target value to construct a training set. Then, the complexity assessment model is trained, and cross-validation is used to evaluate the fitting effect of the complexity assessment model to ensure the accuracy of the comprehensive complexity score. Finally, the trained complexity assessment model is used to evaluate the complex code snippet to be evaluated to generate a comprehensive complexity score for each complex code snippet. The higher the score, the higher the complexity.
[0088] In some specific implementations, the complexity of the code is evaluated, and according to the evaluation result, it is determined that the complex code further includes:
[0089] A preset threshold is obtained, and in response to a complexity score corresponding to the code being greater than the preset threshold, the code is defined as a complex code, wherein the preset threshold can be set according to actual needs, such as 0.7.
[0090] Specifically, after the comprehensive complexity score is calculated, a reasonable threshold is determined to screen and determine the complex code. The preferred threshold of the comprehensive complexity score in this application is 0.7, which can ensure that the truly complex code (that is, the code with multiple complexity indicators being high) will be screened out. That is, when the comprehensive complexity score is greater than 0.7, it means that the code snippet meets the definition of complex code. When the comprehensive complexity score is less than or equal to 0.7, it means that the code snippet is relatively simple and does not fall into the category of complex code.
[0091] In the above implementation, by quantitatively analyzing multiple key indicators such as the cyclomatic complexity of the code, the number of lines of code, the depth and breadth of the AST (abstract syntax tree), and combining weight allocation and data fitting to systematically evaluate the code complexity, the code complexity can be defined scientifically and systematically, and ultimately the code fragments that meet the definition of complex code can be screened out, thereby improving the recognition accuracy of complex code.
[0092] S3: Acquire first code annotation information corresponding to the complex code, and preprocess the first code annotation information to obtain target code annotation information.
[0093] It should be noted that the code comment information refers to the comment information located at the head of the code method, usually the first descriptive comment information therein. The first code comment information refers to the comment information corresponding to the complex code that has not been preprocessed; the preprocessing method refers to the data cleaning of the initial code comment information to construct a high-quality data set for further analysis; the target code comment information refers to the preprocessed code comment information.
[0094] In some specific embodiments, Figure 3 As shown, the first code annotation information is preprocessed to obtain the target code annotation information including:
[0095] Based on the code comment information classification mechanism, redundant code comment information in the first code comment information is screened out to obtain second code comment information, wherein the code comment information classification mechanism refers to classifying the validity of the code comment information to obtain two categories: comment information that can effectively describe and explain the code, and irrelevant or redundant comment information. Here, the comments that can effectively describe and explain the code are screened, and those irrelevant or redundant comments are excluded;
[0096] According to the length distribution of the second code annotation information, the code annotation information that meets the preset standard length in the second code annotation information is screened to obtain the target code annotation information. Specifically, the length distribution of the second code annotation information refers to the length distribution of the text. The preset standard length can be set according to actual needs, such as longer length, larger amount of information, etc. Specifically, based on the length distribution of the second code annotation information, the second code annotation information is further refined to remove annotation information that is too short and has limited information content, so as to ensure the validity and representativeness of the target code annotation information.
[0097] In the above implementation, the code annotation information is preprocessed to construct high-quality code annotation information for further analysis, thereby ensuring the validity and representativeness of the code annotation information.
[0098] S4: Perform syntactic analysis on the target code annotation information to obtain a syntactic analysis result.
[0099] It should be noted that syntactic analysis refers to determining the subject, predicate and object structures in the code comment information. The syntactic analysis results mainly include subject structure missing, predicate structure missing, object structure missing and complete syntactic structure. For example, a complete English sentence usually contains a specific structure, as shown in the following expression:
[0100]
[0101]
[0102] Any syntactic structure part of the sentence defined in this application may be missing. Specifically, for a given English sentence St, its syntactic structure may be composed of a subject (Sub), a predicate (Pre), an object (Obj) and a sentence complement (Cl), where the object can be divided into a direct object (D-Obj) and an indirect object (I-Obj). Sub is a person, place or thing that is performing an action, Pre represents the action in the sentence, D-Obj represents receiving an action, and I-Obj represents to whom or for whom an action is being performed. Cl represents renaming or describing the subject or object. This application mainly detects and generates missing syntactic structures of the subject, predicate and object in the annotation information.
[0103] In some specific embodiments, Figure 3 As shown, the target code annotation information is subjected to syntactic analysis, and the syntactic analysis results include:
[0104] The syntactic structure components of the target code annotation information are obtained, where the syntactic structure components refer to the subject structure, predicate structure and object structure. By inputting the target code annotation information into the large model, the corresponding syntactic structure components can be analyzed and identified, such as Figure 4 As shown, it shows examples of various components of the annotated sentences obtained by the large model, including noun phrases (NP), verb phrases (VP), etc. These components provide basic data for further syntactic structure analysis;
[0105] According to the preset grammatical pattern and the syntactic structure components of the target code annotation information, the syntactic structure missing type of the target code annotation information is determined. The grammatical pattern is used to describe the mapping relationship between the syntactic structure missing type and the syntactic structure components. Among them, the missing analysis of the annotation information mainly focuses on three types of syntactic structure missing, namely, subject structure missing, predicate structure missing and object structure missing. For these missing types, this application proposes three grammatical patterns for annotating syntactic structure missing types. The specific contents are shown in Table 1. These grammatical patterns cover three categories: subject structure missing, predicate structure missing and object structure missing. Among them, each missing type has a corresponding grammatical pattern, which is obtained through syntactic analysis results. For example, the first grammatical pattern in Table 1 "<C,In> → Only contains VP” means that if the annotation contains only verb phrases (VP), it indicates that the annotation may have a missing subject, and so on, no further explanation is given.
[0106] Table 1: Types of missing syntactic structures and the meanings of different symbols in each missing type.
[0107]
[0108]
[0109] In the above implementation, the syntactic analysis method proposed in this application can not only effectively identify and process different types of missing annotation information phenomena, but also provide guidance for the subsequent automatic generation of missing information, thereby improving the processing efficiency of missing information in code annotation information.
[0110] S5: According to the syntactic analysis results, determine and complete the missing information in the target code comment information.
[0111] It should be noted that the Bart model is used in this application to automatically generate missing information in the target code annotation information. The Bart model includes three sub-models, which are models for generating information for subject missing, predicate missing and object missing. Among them, the BART model (Bidirectional and Auto-Regressive Transformers) is a deep learning model based on the Transformer architecture.
[0112] In some specific embodiments, Figure 3 As shown, according to the syntactic analysis results, the missing information in the target code comment information is determined and completed, including:
[0113] According to the missing type of the syntactic structure of the target code annotation information, the target sub-model in the neural network model is determined. That is, when the missing type of the syntactic structure is subject structure missing, predicate structure missing and object structure missing, the corresponding target sub-models are the sub-models for subject missing, predicate missing and object missing in the Bart model respectively.
[0114] The target sub-model is used to generate missing information corresponding to the missing type of syntactic structure.
[0115] In some specific implementations, determining and completing missing information in the target code annotation information according to the syntactic analysis result further includes:
[0116] The target code annotation information is filled with the missing information generated by the target sub-model to generate complete code annotation information.
[0117] Specifically, based on the code missing information detection, the missing information is automatically generated by the constructed Bart model. The construction method of the model includes: obtaining a data set for model training and prediction, that is, the code comment information in the previous text, and then building a neural network model (that is, the Bart model) on this data set to predict the missing of the comment syntactic structure at the word level to ensure that the comment can fully reflect the logical structure of the complex code. The steps are as follows:
[0118] (1) First, all annotation information that does not contain missing syntactic structures is extracted from the code base, and the subject, predicate, and object in these annotation information are used as the true value for model training and prediction. These annotation information must accurately describe the function and behavior of the code. For annotation information whose subject or object is personal pronouns or demonstrative pronouns, these pronouns have a low correlation with the specific code content and rarely provide effective information about the code behavior, so they are filtered out. Then, based on the remaining annotation information, three datasets are created for predicting the subject, predicate, and object structures that may be missing in the annotation information.
[0119] (2) Based on the above dataset, the Bart model is trained into three sub-models to predict and generate missing information corresponding to subject missing, predicate missing, and object missing. The input of the model is the annotation information with missing syntactic structure and its corresponding complex code snippet, and the output is the complete annotation after generating the missing information, accurate to the word level. During the model training and prediction process, all parameters use the default values of the Transformers library and Keras library. The specific parameter configuration is as follows: the batch size is 64, the decoding method is beam search, the parameter beam size is set to 5, the learning rate is 5e-5, the maximum input length is 128, the maximum predicted missing text length is 16, and the optimizer is the Adam optimizer.
[0120] (3) Input the target code annotation information into the trained model and output the complete code annotation information.
[0121] In the above implementation, given incomplete comment information with missing information, it is possible to output complete comment information with missing information accurate to the word level. Based on this, the completeness of the comment information of complex codes can be improved, the readability and maintainability of complex codes can be enhanced, and it is helpful for developers to better understand and modify complex codes.
[0122] In the above-mentioned code comment information processing method, the method includes: obtaining code; evaluating the complexity of the code, and determining the complex code according to the evaluation result; obtaining the first code comment information corresponding to the complex code, and preprocessing the first code comment information to obtain the target code comment information; performing syntactic analysis on the target code comment information to obtain the syntactic analysis result; determining and completing the missing information in the target code comment information according to the syntactic analysis result. The beneficial effects of the present application include:
[0123] (1) Improve the completeness and accuracy of annotation information in complex code: By accurately detecting and generating missing parts in annotations, the completeness of annotation information in complex code can be significantly improved. Complete annotation information can more accurately describe the functions and logic of complex code, reducing ambiguity and misunderstanding when developers understand and maintain complex code. In particular, accurate understanding of complex code is crucial in project maintenance.
[0124] (2) Enhance code readability and maintainability: By automatically generating missing parts in the comment information, the comment content is made more coherent and clear, which improves the readability of complex code. Based on this, it can not only help existing developers understand and maintain the code, but also provide code documentation support for future developers, helping to reduce the technical debt of complex code and reduce maintenance costs;
[0125] (3) Improve development efficiency: By automatically detecting and generating annotation information, the time developers spend manually writing and revising complex code annotation information is greatly reduced, improving the efficiency and productivity of the overall development process;
[0126] (4) Enhanced code traceability: Complete and accurate annotation information is crucial for code review, debugging, and auditing. This application includes detailed annotation information in the key parts of the code, clearly recording its functions and logic, making it easier for developers to quickly trace the intent of the code when problems occur, thereby accelerating problem location and resolution.
[0127] It should be understood that although Figure 2-Figure 3 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2-Figure 3 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0128] In one embodiment, Figure 5 As shown, a code annotation information processing device is provided, including: a data acquisition module, an evaluation module, a preprocessing module, an analysis module and an information processing module, wherein:
[0129] Data acquisition module, used to acquire code;
[0130] An evaluation module is used to evaluate the complexity of the code and determine the complex code based on the evaluation result;
[0131] A preprocessing module, used for obtaining first code annotation information corresponding to the complex code, and preprocessing the first code annotation information to obtain target code annotation information;
[0132] An analysis module is used to perform syntactic analysis on the target code annotation information to obtain a syntactic analysis result;
[0133] The information processing module is used to determine and complete the missing information in the target code comment information according to the syntactic analysis results.
[0134] As a preferred implementation, in the embodiment of the present invention, the evaluation module is specifically used for:
[0135] Get key complexity metrics for your code;
[0136] According to the attributes of the code, the eigenvalues corresponding to the key complexity indicators are determined, and the eigenvalues corresponding to the key complexity indicators are standardized to obtain the target eigenvalues;
[0137] Get the weights corresponding to key complexity indicators;
[0138] Generate target samples based on the weights and target feature values corresponding to key complexity indicators;
[0139] Input the target sample into the pre-built complexity assessment model to obtain the complexity score of the code. The complexity assessment model includes:
[0140]
[0141] in, represents the complexity score, B represents the number of trees in the random forest, Represents the predicted value of the b-th tree for the target sample x;
[0142] Based on the code complexity score, determine the complexity of the code.
[0143] As a preferred implementation, in the embodiment of the present invention, the evaluation module is further used for:
[0144] A preset threshold is obtained, and in response to a complexity score corresponding to the code being greater than the preset threshold, the code is defined as a complex code.
[0145] As a preferred implementation, in the embodiment of the present invention, the preprocessing module is specifically used for:
[0146] Based on the code annotation information classification mechanism, redundant code annotation information in the first code annotation information is filtered out to obtain second code annotation information;
[0147] According to the length distribution of the second code annotation information, the code annotation information in the second code annotation information that meets the preset standard length is screened to obtain the target code annotation information.
[0148] As a preferred implementation, in the embodiment of the present invention, the analysis module is specifically used for:
[0149] Obtain the syntactic structure components of the target code comment information;
[0150] According to a preset grammatical pattern and syntactical components of the target code annotation information, the syntactical structure missing type of the target code annotation information is determined, and the grammatical pattern is used to describe the mapping relationship between the syntactical structure missing type and the syntactical components.
[0151] As a preferred implementation, in the embodiment of the present invention, the information processing module is specifically used for:
[0152] Determine the target sub-model in the neural network model according to the missing type of the syntactic structure of the target code annotation information;
[0153] The target sub-model is used to generate missing information corresponding to the missing type of syntactic structure.
[0154] As a preferred implementation, in the embodiment of the present invention, the information processing module is further used for:
[0155] The target code annotation information is filled with the missing information generated by the target sub-model to generate complete code annotation information.
[0156] For the specific definition of the code annotation information processing device, please refer to the definition of the code annotation information processing method above, which will not be repeated here. The various modules in the above-mentioned code annotation information processing device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0157] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 6As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a code annotation information processing method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0158] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0159] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0160] S1: Get code;
[0161] S2: Evaluate the complexity of the code and determine the complex code based on the evaluation results;
[0162] S3: Acquire first code annotation information corresponding to the complex code, and preprocess the first code annotation information to obtain target code annotation information;
[0163] S4: Performing syntactic analysis on the target code comment information to obtain syntactic analysis results;
[0164] S5: According to the syntactic analysis results, determine and complete the missing information in the target code comment information.
[0165] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0166] Get key complexity metrics for your code;
[0167] According to the attributes of the code, the eigenvalues corresponding to the key complexity indicators are determined, and the eigenvalues corresponding to the key complexity indicators are standardized to obtain the target eigenvalues;
[0168] Get the weights corresponding to key complexity indicators;
[0169] Generate target samples based on the weights and target feature values corresponding to key complexity indicators;
[0170] Input the target sample into the pre-built complexity assessment model to obtain the complexity score of the code. The complexity assessment model includes:
[0171]
[0172] in, represents the complexity score, B represents the number of trees in the random forest, Represents the predicted value of the b-th tree for the target sample x;
[0173] Based on the code complexity score, determine the complexity of the code.
[0174] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0175] A preset threshold is obtained, and in response to a complexity score corresponding to the code being greater than the preset threshold, the code is defined as a complex code.
[0176] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0177] Based on the code annotation information classification mechanism, redundant code annotation information in the first code annotation information is filtered out to obtain second code annotation information;
[0178] According to the length distribution of the second code annotation information, the code annotation information in the second code annotation information that meets the preset standard length is screened to obtain the target code annotation information.
[0179] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0180] Obtain the syntactic structure components of the target code comment information;
[0181] According to a preset grammatical pattern and syntactical components of the target code annotation information, the syntactical structure missing type of the target code annotation information is determined, and the grammatical pattern is used to describe the mapping relationship between the syntactical structure missing type and the syntactical components.
[0182] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0183] Determine the target sub-model in the neural network model according to the missing type of the syntactic structure of the target code annotation information;
[0184] The target sub-model is used to generate missing information corresponding to the missing type of syntactic structure.
[0185] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0186] The target code annotation information is filled with the missing information generated by the target sub-model to generate complete code annotation information.
[0187] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0188] S1: Get code;
[0189] S2: Evaluate the complexity of the code and determine the complex code based on the evaluation results;
[0190] S3: Acquire first code annotation information corresponding to the complex code, and preprocess the first code annotation information to obtain target code annotation information;
[0191] S4: Performing syntactic analysis on the target code comment information to obtain syntactic analysis results;
[0192] S5: According to the syntactic analysis results, determine and complete the missing information in the target code comment information.
[0193] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0194] Get key complexity metrics for your code;
[0195] According to the attributes of the code, the eigenvalues corresponding to the key complexity indicators are determined, and the eigenvalues corresponding to the key complexity indicators are standardized to obtain the target eigenvalues;
[0196] Get the weights corresponding to key complexity indicators;
[0197] Generate target samples based on the weights and target feature values corresponding to key complexity indicators;
[0198] Input the target sample into the pre-built complexity assessment model to obtain the complexity score of the code. The complexity assessment model includes:
[0199]
[0200] in, represents the complexity score, B represents the number of trees in the random forest, Represents the predicted value of the b-th tree for the target sample x;
[0201] Based on the code complexity score, determine the complexity of the code.
[0202] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0203] A preset threshold is obtained, and in response to a complexity score corresponding to the code being greater than the preset threshold, the code is defined as a complex code.
[0204] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0205] Based on the code annotation information classification mechanism, redundant code annotation information in the first code annotation information is filtered out to obtain second code annotation information;
[0206] According to the length distribution of the second code annotation information, the code annotation information in the second code annotation information that meets the preset standard length is screened to obtain the target code annotation information.
[0207] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0208] Obtain the syntactic structure components of the target code comment information;
[0209] According to a preset grammatical pattern and syntactical components of the target code annotation information, the syntactical structure missing type of the target code annotation information is determined, and the grammatical pattern is used to describe the mapping relationship between the syntactical structure missing type and the syntactical components.
[0210] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0211] Determine the target sub-model in the neural network model according to the missing type of the syntactic structure of the target code annotation information;
[0212] The target sub-model is used to generate missing information corresponding to the missing type of syntactic structure.
[0213] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0214] The target code annotation information is filled with the missing information generated by the target sub-model to generate complete code annotation information.
[0215] In one embodiment, a computer program product is provided, the computer program product comprising a computer program, the computer program when executed by a processor implements the following steps:
[0216] S1: Get code;
[0217] S2: Evaluate the complexity of the code and determine the complex code based on the evaluation results;
[0218] S3: Acquire first code annotation information corresponding to the complex code, and preprocess the first code annotation information to obtain target code annotation information;
[0219] S4: Performing syntactic analysis on the target code comment information to obtain syntactic analysis results;
[0220] S5: According to the syntactic analysis results, determine and complete the missing information in the target code comment information.
[0221] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0222] Get key complexity metrics for your code;
[0223] According to the attributes of the code, the eigenvalues corresponding to the key complexity indicators are determined, and the eigenvalues corresponding to the key complexity indicators are standardized to obtain the target eigenvalues;
[0224] Get the weights corresponding to key complexity indicators;
[0225] Generate target samples based on the weights and target feature values corresponding to key complexity indicators;
[0226] Input the target sample into the pre-built complexity assessment model to obtain the complexity score of the code. The complexity assessment model includes:
[0227]
[0228] in, represents the complexity score, B represents the number of trees in the random forest, Represents the predicted value of the b-th tree for the target sample x;
[0229] Based on the code complexity score, determine the complexity of the code.
[0230] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0231] A preset threshold is obtained, and in response to a complexity score corresponding to the code being greater than the preset threshold, the code is defined as a complex code.
[0232] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0233] Based on the code annotation information classification mechanism, redundant code annotation information in the first code annotation information is filtered out to obtain second code annotation information;
[0234] According to the length distribution of the second code annotation information, the code annotation information in the second code annotation information that meets the preset standard length is screened to obtain the target code annotation information.
[0235] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0236] Obtain the syntactic structure components of the target code comment information;
[0237] According to a preset grammatical pattern and syntactical components of the target code annotation information, the syntactical structure missing type of the target code annotation information is determined, and the grammatical pattern is used to describe the mapping relationship between the syntactical structure missing type and the syntactical components.
[0238] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0239] Determine the target sub-model in the neural network model according to the missing type of the syntactic structure of the target code annotation information;
[0240] The target sub-model is used to generate missing information corresponding to the missing type of syntactic structure.
[0241] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0242] The target code annotation information is filled with the missing information generated by the target sub-model to generate complete code annotation information.
[0243] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0244] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0245] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application.
Claims
1. A code annotation information processing method, characterized in that: The method comprises: Get the code; Evaluating the complexity of the code and determining a complex code based on the evaluation result; Acquire first code annotation information corresponding to the complex code, and preprocess the first code annotation information to obtain target code annotation information; Performing syntactic analysis on the target code annotation information to obtain a syntactic analysis result; According to the syntactic analysis result, the missing information in the target code annotation information is determined and completed.
2. The code annotation information processing method according to claim 1, characterized in that: The complexity of the code is evaluated, and based on the evaluation result, it is determined that the complex code includes: Obtain key complexity indicators of the code; Determine the characteristic value corresponding to the key complexity index according to the attribute of the code, and perform standardization on the characteristic value corresponding to the key complexity index to obtain a target characteristic value; Obtaining the weight corresponding to the key complexity indicator; Generate a target sample based on the weights and target feature values corresponding to the key complexity indicators; The target sample is input into a pre-built complexity assessment model to obtain a complexity score of the code, wherein the complexity assessment model includes: in, represents the complexity score, B represents the number of trees in the random forest, Represents the predicted value of the b-th tree for the target sample x; Based on the complexity score of the code, the complexity of the code is determined.
3. The code annotation information processing method according to claim 2, characterized in that: The complexity of the code is evaluated, and based on the evaluation result, it is determined that the complex code also includes: A preset threshold is obtained, and in response to a complexity score corresponding to the code being greater than the preset threshold, the code is defined as a complex code.
4. The code annotation information processing method according to claim 1, characterized in that: Preprocessing the first code annotation information to obtain target code annotation information includes: Based on the code annotation information classification mechanism, redundant code annotation information in the first code annotation information is filtered out to obtain second code annotation information; According to the length distribution of the second code annotation information, the code annotation information in the second code annotation information that meets a preset standard length is screened to obtain the target code annotation information.
5. The code annotation information processing method according to claim 1 or 4, characterized in that: The target code annotation information is subjected to syntactic analysis, and the syntactic analysis results obtained include: Obtain the syntactic structure components of the target code comment information; The syntactic structure missing type of the target code annotation information is determined according to a preset grammatical pattern and the syntactic structure components of the target code annotation information, wherein the grammatical pattern is used to describe the mapping relationship between the syntactic structure missing type and the syntactic structure components.
6. The code annotation information processing method according to claim 5, characterized in that: Determining and completing the missing information in the target code annotation information according to the syntactic analysis result includes: Determining a target sub-model in a neural network model according to a type of missing syntax structure of the target code annotation information; The target sub-model is used to generate missing information corresponding to the missing type of the syntactic structure.
7. The code annotation information processing method according to claim 6, characterized in that: Determining and completing the missing information in the target code annotation information according to the syntactic analysis result also includes: The target code annotation information is supplemented with the missing information generated by the target sub-model to generate complete code annotation information.
8. A code annotation information processing device, characterized in that: The device comprises: Data acquisition module, used to acquire code; An evaluation module, used to evaluate the complexity of the code and determine the complex code according to the evaluation result; A preprocessing module, used for obtaining first code annotation information corresponding to the complex code, and preprocessing the first code annotation information to obtain target code annotation information; An analysis module, used for performing syntactic analysis on the target code annotation information to obtain a syntactic analysis result; An information processing module is used to determine and complete the missing information in the target code annotation information according to the syntactic analysis result.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Code generation method and device, storage medium and electronic equipment
CN120447880A