A method for accurate recommendation and review of method names
Through the pre-trained model and prompt learning framework, combined with closed class context information and binary classification model, the problem of difficulty in learning target dissonance and consistency checking of method name detection and recommendation in the prior art is solved, and more efficient method naming recommendation and consistency checking is achieved, which improves the comprehensibility and maintainability of the software.
Patent Information
- Application Number
- CN202310289887.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-03-23
AI Technical Summary
In the prior art, the automatic detection and recommendation models of method names have problems such as learning objective dissonance, underutilization of closed class context information, and difficulty in measuring quality and semantic consistency of the inspection results rely on the generation of new method names, resulting in insufficient software comprehensibility and maintainability.
The pre-trained model and prompt learning framework are adopted, and the method context information of the code file is extracted, the method naming and recommendation is recommended using the CodeT5 model, and the binary classification model is constructed for consistency checking. The inconsistent named data set is generated by combining the difficult negative sample mining method, and the method naming recommendation and consistency checking tasks are coordinated to make full use of the advantages of the pre-trained model.
It improves the accuracy of method naming recommendations and the accuracy of consistency checks, which is significantly better than other models, and improves the comprehensibility and maintainability of the software.
Smart Images

Figure CN116166789B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for accurate recommendation and review of method naming based on a pre-trained model and prompt learning, belonging to the field of computer technology applications. Background Art
[0002] The comprehensibility of software programs is crucial for software upgrading and maintenance. Among them, the method name, as a brief description of the source code, directly affects developers' understanding of the program and development efficiency. However, low-quality method names widely exist in various projects, resulting in developers spending a large amount of time on re-understanding methods and renaming methods in order to improve the readability and maintainability of software. Therefore, automatically detecting low-quality method names and recommending more appropriate method names according to the actual functions of methods, that is, method naming consistency checking and method naming recommendation, have important application value and broad development prospects in practical scenarios.
[0003] Currently, the review and recommendation of method names include many method naming recommendation and method naming consistency checking models, but most models have the following problems:
[0004] 1) All models based on deep learning methods are trained from scratch. While the model is learning the semantic representations of programming languages and natural languages, it also has to learn the relationship between method names and functional implementations. The above two distinct learning objectives will cause learning objective imbalance, reduce training efficiency, and thus lead to sub-optimal results;
[0005] 2) The context information of closed classes is not fully utilized. The datasets of currently commonly used models focus on extracting the context of closed classes mainly on class names and sibling methods, without considering class attribute information, making it difficult for the model to learn the abstract functions of methods;
[0006] 3) The current process of method naming consistency checking is to first generate a new method name, and then judge whether the method name is consistent according to the calculation result of the lexical string similarity between the current method name and the new method name and the selected threshold. This process only considers string similarity, resulting in the result being highly dependent on the quality of the generated new method name, and it is also difficult to measure semantic consistency, affecting the accuracy of the inspection result.
[0007] With the rapid development of computer technology and open-source / industry application software, the demand for software updates and maintenance in various industries is continuously escalating. How to achieve efficient and accurate method naming review and recommendation that can be used in practical scenarios and improve the comprehensibility and maintainability of programs has become an urgent problem in the industry. Summary of the Invention
[0008] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a method name precise recommendation and review method. Different from the existing deep learning-based methods, this method first learns the context representations of programming languages and natural languages through a pre-trained model, and then fully utilizes the capabilities and knowledge of large language models through prompt tuning to detect inconsistent method names and recommend more accurate names. The method name precise recommendation and review method of the present invention adopts a "pre-training, prompting, and prediction" framework, filling the gap between pre-training tasks and downstream naming tasks, and being able to better coordinate the method name recommendation task and the method name consistency check task compared with existing models and methods. In addition, this method uses a prompt-based binary classification model to implement the method name consistency check task, which can measure semantic consistency and avoid the inherent limitations of the method name consistency check model based on generating first and then comparing.
[0009] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0010] A method name precise recommendation and review method, the steps of which include:
[0011] 1) Select multiple code files, each of the code files containing the code and annotation information of a software engineering project; extract the method context information in each of the code files, and use each of the method context information as a training sample to generate a training data set;
[0012] 2) Annotate a method name indicator for each of the training samples, connect each of the training samples with the corresponding method name indicator, and then train a method name recommendation model to predict and output the method names of each of the training samples;
[0013] 3) Use the training data set as a positive sample data set, and construct a negative sample data set with inconsistent naming. Annotate a method name indicator for each training sample in the positive and negative sample data sets, connect each training sample in the positive and negative sample data sets with the corresponding method name indicator to generate a method name consistency check data set; use the method name consistency check data set to train a method name consistency check model;
[0014] 4) For a code source file to be detected, extract the method context information T of a given method name from the code source file and input it into the trained method name consistency check model to determine whether the method in the method context information T is consistent with the given method name. If it is consistent, the consistency check is passed; otherwise, the method name recommendation model generates a new candidate name according to the method context information T.
[0015] Further, each of the method context information includes function body information and enclosing class information; the function body information includes a method name, an identifier name, and a return type, and the enclosing class information includes a class name, class attributes, and sibling methods.
[0016] Further, a negative sample data set with inconsistent naming is constructed by using an edit-based hard negative sample mining method based on the positive sample data set.
[0017] Further, the method naming recommendation model is a pre-trained model CodeT5 that is sensitive to identifiers.
[0018] Further, the method naming consistency check model is a binary classification model.
[0019] A server, comprising a memory and a processor, wherein the memory stores a computer program configured to be executed by the processor, and the computer program includes instructions for performing the steps in the above method.
[0020] A computer-readable storage medium, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the steps of the above method.
[0021] The method naming precise recommendation and review method of the present invention includes the following steps:
[0022] In the first step, a training data set is constructed. First, a series of data cleaning and normalization operations are performed on the four data sets of MNire and Java-small / med / large. Then, method context information at different levels such as function body information (method name, identifier name, return type, etc.) and enclosing class information (class name, class attributes, sibling methods, etc.) is extracted. Then, sub-word sequences are constructed from the program entity names in the extracted method context information to obtain a training data set. The four data sets of MNire and Java-small / med / large are all code files of about fifteen thousand open-source software engineering projects sorted out by well-known international researchers, including code and annotation information, etc. Sub-word sequences corresponding to different types of entity names (for example, identifiers, class names, return types, etc.) in the extracted method context information are constructed, that is, a corresponding sub-word sequence is constructed for each type of program entity name to generate a training data set.
[0023] In the second step, train the method name recommendation model, perform prompt learning tuning on the identifier-aware pre-trained model CodeT5, train the model CodeT5 based on the constructed training dataset, connect all method contexts with indicator words, and output the method name by building a natural language prompt template (i.e., templated prompt learning training data). Among them, the prompt template mainly consists of three parts. The first part is the indicator word, which is placed before the corresponding different contexts as a prompt; the second part is the separator, which is placed after each method context to separate different method contexts; the third part is the prompt word, which is placed at the end of the input template to prompt the tasks performed by the model, such as method name recommendation, method name consistency check, etc.
[0024] In the third step, train the method name consistency check model. First, use the hard negative sample mining method to construct a sufficient dataset of inconsistent naming examples for training, that is, adopt a sampling method based on editing to destroy and reconstruct the original text using consistent naming examples to construct inconsistent naming data, thereby constructing the training set of the method name consistency check model (i.e., the method name consistency check dataset); then combine the method name and context with the prompt template to obtain the method name consistency check dataset, and then use this dataset to train a binary classification model based on prompt learning, that is, the method name consistency check model.
[0025] In the fourth step, after completing the training of the method name recommendation model and the method name consistency check model, build the overall functional application. First, extract the necessary context information from the given Java source file, and then use the method name consistency check model to judge whether the given method names are consistent. If they are consistent, it means that the given method names pass the consistency check; otherwise, the method name recommendation model will generate new candidate names using the previously extracted context to replace the previous inappropriate names.
[0026] In the first step mentioned above, to construct the training dataset, the method process is as follows:
[0027] (1) Download the open-source datasets MNire and Java-small / med / large;
[0028] (2) Perform filtering and cleaning operations on the data:
[0029] 1) Delete empty methods;
[0030] 2) Replace separators;
[0031] 3) Split identifiers;
[0032] 4) Normalize the length;
[0033] (3) Extract context information, including function body context information (method name, identifier name, return type, etc.) and enclosing class context information (class name, class attributes, sibling methods, etc.);
[0034] (4) Construct a dataset based on the sub-word sequences corresponding to the extracted context information;
[0035] In the second step, the training process of the method naming recommendation model is as follows:
[0036] (1) Use the dataset obtained in the first step as the method naming generation dataset (including method context and method name);
[0037] (2) Connect all method contexts with method name indicators to construct templated prompt learning training data;
[0038] (3) According to the data obtained in (2), perform prompt learning tuning on the identifier-aware pre-trained model CodeT5;
[0039] (4) The model generates a predicted method name based on the input method context;
[0040] (5) Use the method name in (1) as the ground truth to evaluate the accuracy of the predicted method name generated by the model.
[0041] In the third step, the training process of the method naming consistency check model is as follows:
[0042] (1) Use the method-based hard negative sample mining method to use the method naming generation dataset as the naming consistent dataset, and generate an inconsistent naming dataset on this basis. The specific operations are as follows:
[0043] 1) Randomly obtain name sub-words in the method naming generation dataset with a certain probability;
[0044] 2) Through editing-based operations, disrupt and reconstruct each sub-word obtained in 1), including operations such as adding, deleting, replacing, and not changing the original sub-word;
[0045] 3) Generate an inconsistent naming dataset according to the results obtained in 2);
[0046] (2) Use the samples in the method naming generation dataset as positive samples and the samples in the inconsistent naming dataset as negative samples to generate a method naming consistency check dataset;
[0047] (3) Train a binary classification model based on prompt learning on the basis of the method naming consistency check dataset.
[0048] In the fourth step, the functional implementation process based on the method naming consistency check model and the method naming recommendation model is as follows: (1) Preprocess the input Java source file;
[0049] (2) Extract the method name and method context;
[0050] (3) Use the method naming consistency check model obtained in the third step to determine whether the corresponding method name is consistent according to the method context. If the predicted output is positive (consistent), the given method ends after passing the consistency check. If the predicted output is negative (inconsistent), go to step (4);
[0051] (4) When the method context and method name do not pass the method naming consistency check, call the method naming recommendation model obtained in the second step, and use the method context extracted in step (2) as the input to generate a new method name.
[0052] Compared with the existing technical solutions, the beneficial effects of the present invention are:
[0053] (1) Use a prompt-based classification model to perform method naming consistency checking, model this task as a binary classification problem, can measure semantic consistency, and avoid the inherent limitations of other method naming consistency check models;
[0054] (2) For the two different learning objectives of method naming recommendation and method naming consistency checking, the model adopts a "pre-training, prompting, and prediction" framework, fills the gap between pre-training tasks and downstream naming tasks, can better coordinate these two tasks than other models, makes full use of the advantages of pre-trained models, and improves training efficiency;
[0055] (3) In the method naming consistency checking task, use an edit-based hard negative sample mining method to generate an inconsistent naming dataset from a consistent naming dataset, effectively solving the problem of insufficient inconsistent naming sample data;
[0056] (4) Make full use of the context information of closed classes, and the class attribute information effectively improves the accuracy of method naming recommendation;
[0057] (5) Significantly outperforms other methods in both method naming recommendation and method naming consistency checking tasks. Description of the Drawings
[0058] Figure 1 It is a flowchart for dataset construction.
[0059] Figure 2 It is a flowchart for training the method naming recommendation model.
[0060] Figure 3It is a flow chart for training a method naming consistency check model.
[0061] Figure 4 It is a functional flow chart based on the method naming consistency check model and the method naming recommendation model.
[0062] Figure 5 It is a general flow chart of the method. Specific implementation manners
[0063] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0064] The basic idea of the present invention is to construct a method naming recommendation dataset using an open-source dataset. Compared with other method naming recommendation task datasets, the present invention extracts the class attribute context information of closed classes when constructing the dataset. The closed class attribute context information is part of the method context, as shown below:
[0065]
[0066]
[0067] True value: getNumParameters
[0068] The method naming recommendation result of the present invention (excluding class attributes): getParametersLength.
[0069] The method naming recommendation result of the present invention (including class attributes): getNumParameters.
[0070] Class attribute context information can effectively improve the accuracy of results; by using the method naming recommendation dataset and combining the difficult negative sample mining method based on editing proposed in the present invention to construct a negative sample dataset with inconsistent naming, a method naming consistency check dataset is generated, which effectively solves the common problem of insufficient number of inconsistent samples in this task and ensures the training effect of the subsequent model; on the basis of the method naming recommendation dataset, the pre-trained model CodeT5 for identifier perception is fine-tuned by prompt learning in combination with a prompt template to generate a method naming recommendation model with high quality and accuracy; the present invention adopts a binary classification model based on prompt learning, and uses the method naming consistency check dataset to obtain a method naming consistency check model. This model corresponds to a binary classification task, which can correctly distinguish the semantics of words, especially for the case where strings are similar but semantics are different, improving the accuracy of this task and avoiding the irrationality and limitations of other method naming consistency check models based on string similarity; the present invention adopts a "pre-training, prompting and prediction" framework to bridge the gap between pre-training tasks and downstream naming tasks, coordinate the method naming recommendation task and the method naming consistency check task, avoid the generation of sub-optimal results, and ensure the quality of inspection results and recommendation results.
[0071] The overall flowchart of the method of the present invention is as Figure 5 shown. First, a method naming recommendation dataset is constructed through operations such as data cleaning and context extraction of the open-source dataset. Among them, data cleaning includes operations such as deleting empty content, replacing delimiters, splitting identifiers, and length standardization, and context extraction includes function body context (method name, return type, identifier, etc.) and closed class context (class name, sibling methods, class attributes, etc.); then, a method naming consistency check dataset is constructed on the basis of the method naming recommendation dataset. The main method is to use the method naming recommendation dataset as the positive sample dataset, and on this basis, use the difficult negative sample mining method based on editing to construct a negative sample dataset with inconsistent naming, and combine the positive and negative sample datasets to generate a method naming consistency check dataset; then, the training of the method naming recommendation model and the method naming consistency check model is carried out respectively: using the method naming recommendation dataset in combination with a prompt template to train the pre-trained model CodeT5 for identifier perception to generate a method naming recommendation model, and using the method naming consistency check dataset to train a binary classification model based on prompt learning to generate a method naming consistency check model; finally, using the above two models, the accurate recommendation and review method of Method is completed.
[0072] The specific steps of the method of the present invention are as follows:
[0073] The first step is to construct a method naming recommendation dataset
[0074] The construction of the method naming recommendation dataset is asFigure 1 As shown in the figure, first download four open-source datasets such as MNire and Java-small / med / large, then perform data cleaning operations to standardize the data format, then extract context information from them, form sub-word sequences from the program entity names in the context, and finally obtain the method naming recommendation dataset.
[0075] The specific steps for constructing the method naming recommendation dataset are as follows:
[0076] (1) Download the open-source datasets MNire and Java-small / med / large;
[0077] (2) Perform filtering and cleaning operations on the data:
[0078] 1) Delete the methods with empty content in the dataset;
[0079] 2) Replace the delimiters in the data;
[0080] 3) Split the identifiers, convert all contexts to lowercase form and cut them into sub-word sequences;
[0081] 4) Standardize the length, where the number of sibling methods and the number of class attributes are both set to 10, and the maximum length of the sub-word sequence is 512;
[0082] (3) Extract the method context information, including the method context information of the function body (method name, identifier name, return type, etc.) and the method context information of the enclosing class (class name, class attributes, sibling methods, etc.);
[0083] (4) Finally, take the method as the granularity, integrate the corresponding context information, and construct the method naming recommendation dataset.
[0084] The second step is to construct the method naming consistency check dataset
[0085] As Figure 3 shown, use the method naming recommendation dataset generated in the first step as the positive sample dataset with consistent naming. On this basis, use the difficult negative sample mining method based on editing to construct the negative sample dataset with inconsistent naming. The combination of the two generates the method naming consistency check dataset.
[0086] The specific steps for constructing the method naming consistency check dataset are as follows:
[0087] (1) Randomly obtain the name sub-words in the method naming generation dataset with a certain probability;
[0088] (2) Perform operations such as adding, deleting, replacing, and not changing on each obtained sub-word;
[0089] (3) Use the sub-words obtained in step (2) to construct a dataset with inconsistent naming;
[0090] (4) Use the dataset generated by method naming as the positive samples, and the dataset with inconsistent naming obtained in step (3) as the negative samples,
[0091] to generate a dataset for checking method naming consistency.
[0092] The third step is to train a method naming recommendation model
[0093] As Figure 2 shown, use the method naming recommendation dataset, connect the method context with the indicator words, and combine the constructed natural language prompt template to construct templated prompt learning training data, and then train the identifier-aware pre-trained model CodeT5 to generate a method naming recommendation model after training.
[0094] The training process of the method naming recommendation model is as follows:
[0095] (1) Extract the method context and method name from the dataset generated by method naming;
[0096] (2) Connect all the method contexts with the indicator words to construct templated prompt learning training data;
[0097] (3) According to the data obtained in (2), perform prompt learning tuning on the identifier-aware pre-trained model CodeT5. The hyperparameters of the model are set as follows: the learning rate is 5e-5, the maximum input length is 512, the maximum output length is 16, the beam search width is 10, and the batch size is 16;
[0098] (4) The model generates a predicted method name based on the input method context;
[0099] (5) Use the method name in (1) as the ground truth to evaluate the accuracy of the predicted method name generated by the model.
[0100] The fourth step is to train a method naming consistency check model
[0101] As Figure 3 shown, use the method naming consistency check dataset to train a binary classification model based on prompt learning to generate a method naming consistency check model, where the hyperparameters of the model are set as follows: the learning rate is 5e-5, the maximum input length is 512, and the batch size is 16.
[0102] The fifth step is to implement accurate recommendation and review methods for Method naming
[0103] After completing the training of the method naming recommendation model and the method naming consistency check model, construct the functional application of the overall Method naming accurate recommendation and review method, such as Figure 5 As shown, first extract the necessary context information from the given Java source file, and then use the method naming consistency check model to determine whether the given method names are consistent. If they are consistent, it means that the given method passes the consistency check. Otherwise, the method naming recommendation model will use the previously extracted context to generate new candidate names to replace the previous inappropriate names.
[0104] The process of the Method naming accurate recommendation and review method is as follows:
[0105] (1) The input Java source file;
[0106] (2) Perform preprocessing operations on the source file, such as deleting empty methods, replacing delimiters, splitting identifiers, and length normalization;
[0107] (3) Extract the method names and method contexts of the preprocessed file;
[0108] (4) Use the method naming consistency check model to determine whether the method context and its method name are consistent. If the model output result is positive, it means they are consistent and the given method passes the consistency check. If the model output is negative, it means the current method name of the method context is inappropriate, and go to step (5);
[0109] (5) When the method context and the method name do not pass the method naming consistency check, call the method naming recommendation model, use the method context extracted in step (3) as the input of this model, and generate a new method name.
[0110] To prove the effectiveness of the method of the present invention, in the comparative experiment of the method naming recommendation task, based on the open-source datasets (MNire and Java-small / med / large), four most used baseline methods (Code2vec, MNire, Cognac, and GTNM) are selected for comparison. The experimental results are shown in Table 1; in the method naming consistency check task, four most advanced baseline methods (DebugMethodName, MNire, DeepName, and Cognac) are selected for comparison. The experimental results are shown in Table 2. In both the method naming recommendation task and the method naming consistency check task, the experimental results of the method of the present invention are significantly better than those of other baseline methods.
[0111] Table 1
[0112]
[0113] Table 2
[0114]
[0115]
[0116] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the principles and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.
Claims
1. A method for accurate recommendation and review of method names, the steps of which include: 1) Select a plurality of code files, each of the code files containing the code and annotation information of a software engineering project; Extract the method context information from each of the code files, and use each of the method context information as a training sample to generate a training data set; 2) Label a method name indicator for each of the training samples, and connect each of the training samples with the corresponding method name indicator and then train a method name recommendation model to predict and output the method names of each of the training samples; the training process of the method name recommendation model is: (1) Use the training data set as a method name generation data set; (2) Connect all the method contexts with the method name indicators to construct templatized prompt learning training data; (3) According to the prompt learning training data obtained in step (2), perform prompt learning tuning on the identifier-aware pre-trained model CodeT5; (4) The pre-trained model CodeT5 generates a predicted method name according to the input method context; (5) Use the method name in step (1) as the ground truth to evaluate the accuracy of the predicted method name generated by the pre-trained model CodeT5; 3) Use the training data set as a positive sample data set, and construct a negative sample data set with inconsistent naming. Label a method name indicator for each training sample in the positive and negative sample data sets, and connect each training sample in the positive and negative sample data sets with the corresponding method name indicator to generate a method name consistency check data set; use the method name consistency check data set to train a method name consistency check model; 4) For a code source file to be detected, extract the method context information T of a given method name from the code source file and input it into the trained method name consistency check model to determine whether the method in the method context information T is consistent with the given method name. If it is consistent, the consistency check is passed; otherwise, the method name recommendation model generates a new candidate name according to the method context information T.
2. The method according to claim 1, characterized in that, Each of the method context information includes function body information and enclosing class information; the function body information includes method name, identifier name, return type, and the enclosing class information includes class name, class attributes, and sibling methods.
3. The method according to claim 1, characterized in that Construct a negative sample data set with inconsistent naming based on the positive sample data set by using an edit-based hard negative sample mining method.
4. The method according to claim 1 or 2 or 3, characterized in that The method name recommendation model is the identifier-aware pre-trained model CodeT5.
5. The method according to claim 1 or 2 or 3, characterized in that The method name consistency check model is a binary classification model.
6. A server, characterized in that, Including a memory and a processor, the memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing each step in any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Object-oriented program method naming odor detection method based on deep learning
CN114398076A
Meta-learning data augmentation framework
US20220351071A1