Code comment updating method and device combining context environment and reconstruction judgment

By combining the context environment and code reconstruction judgment encoding and decoding model, detecting whether code comments need to be updated, solving the problems of misjudgment and poor universality in the existing technology, achieving high-accuracy code comment updates, and solving the OOV problem.

CN115033244BActive Publication Date: 2025-05-16NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210679595.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-16
Publication Date
2025-05-16
Estimated Expiration
2042-06-16

AI Technical Summary

Technical Problem

When detecting whether code comments need to be updated, the prior art lacks consideration of the code context and reconstruction impact, resulting in misjudgment, and is not very universal, making it difficult to solve the OOV problem.

Method used

By obtaining the code training set, establishing the correspondence between method code and annotations, using the encoding and decoding model combined with the context environment and code reconstruction judgment, detecting the annotations that should be updated in the object code, and performing method call detection after the code reconstruction judgment to improve the annotation.

Benefits of technology

Effectively avoid misjudgments, improve the accuracy of code comment updates, solve OOV problems, and ensure the coordination between comment updates and code reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115033244B_ABST
    Figure CN115033244B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for updating code comments that combines context environment and reconstruction judgment, the method comprising: using a code training set to train a coding and decoding model; the coding and decoding model comprises a coding layer, a decoding layer and a collaborative attention layer; the coding layer is connected to the collaborative attention layer so that the output features of each encoder are linked to the output features of other encoders, and the collaborative attention layer outputs global features; using a code reconstruction model to perform code reconstruction judgment on a target code, excluding the code reconstructed part, and then using the coding and decoding model to detect the target code, and querying the comments that should be updated in the target code. Using the above technical solution, according to the data features of the context of the target code, it is comprehensively judged whether the code comments should be updated, while considering the impact of code reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of code analysis, and in particular to a code comment updating method and device combining context environment and reconstruction judgment. Background Art

[0002] Code comments are crucial for software development, maintenance, and understanding. First, code comments are an important tool to ensure that developers can effectively understand the code. Second, inconsistent code comments will inevitably cause ambiguity in understanding, leading to serious software development problems. However, developers often ignore the synergy between code updates and comment updates.

[0003] In the prior art, detection schemes based on inconsistent comments are used to detect whether code comments should be updated. The most typical one is outdated comment detection based on random forest. The problem is that the detection scheme is based on special code types or attributes and takes a certain version of code as input. It does not have good versatility and requires manual summary of rules and features. Manual summary of rules and features is prone to ignore the data features of the code context and the impact of code refactoring on the code, which leads to misjudgment. Summary of the invention

[0004] Purpose of the invention: The present invention provides a code comment updating method and device that combine context environment and reconstruction judgment, aiming to comprehensively judge whether the code comment should be updated based on the data characteristics of the target code context, while considering the impact of code reconstruction to avoid misjudgment; further, the calling function in the updated method code is commented on and updated to improve the global data characteristics of the code and solve the OOV (Out Of Vocabulary) problem.

[0005] Technical solution: The present invention provides a code comment updating method that combines context environment and reconstruction judgment, including: obtaining a code training set, and establishing a correspondence between method code and comments according to method description information in the comments; the code training set includes old method code, old comments, new method code and new comments; processing the code training set to exclude training data in which no method code description information is added or changed; using the code training set to train a coding and decoding model; the coding and decoding model includes a coding layer, a decoding layer and a collaborative attention layer, the coding layer includes a new code encoder, a deleted code encoder, a replaced code encoder and an old comment encoder; the decoding layer includes a new comment decoder; the coding layer is connected to the collaborative attention layer so that the output features of each encoder are linked to the output features of other encoders, and the collaborative attention layer outputs global features; using the code reconstruction model to perform code reconstruction judgment on the target code, excluding the code reconstructed part, and then using the coding and decoding model to detect the target code and query the comments that should be updated in the target code.

[0006] Specifically, the code training set is processed, including deleting inline comments in method codes and deleting labels of comments.

[0007] Specifically, new codes, deleted codes, replaced codes and old comments are obtained by comparing each clause of the old method code and the corresponding new method code; new codes, deleted codes and replaced codes are all changed codes.

[0008] Specifically, the training data that does not add or change the method code description information is excluded, including: splitting the tokens in the old comments. If the token is not in the corresponding method code and is not in the vocabulary, it is considered a spelling error. In the case of a spelling error, if the new comment only corrects the spelling, the corresponding training data is excluded; if only Java annotations are added to the changed code, the corresponding training data is excluded; use NLP grammar detection to exclude old and new comments that do not conform to grammatical rules; split the old comments and the corresponding new comments into tokens respectively, if the tokens are consistent, the corresponding training data is excluded.

[0009] Specifically, the training data without adding or changing the method code description information is excluded, and then the following includes: splitting the comments according to the camel case naming and inserting a connection character between each identifier.

[0010] Specifically, the encoding and decoding model also includes a codeBert model embedding layer and a fully connected layer. New codes, deleted codes, replaced codes and old comments are input into the encoding side codeBert model embedding layer. The output of the encoding side codeBert model embedding layer is connected to the encoding layer. The encoding layer is a bidirectional LSTM. The output of the encoding layer is input into the collaborative attention layer. The new comments are input into the decoding side codeBert model embedding layer. The output of the decoding side codeBert model embedding layer is connected to the decoding layer. The decoding layer is a unidirectional LSTM. The output of the collaborative attention layer and the output of the decoding layer are both input into the fully connected layer, and the fully connected layer outputs weights.

[0011] Specifically, the code reconstruction model is obtained by training the XGboost model using old comments and changed codes as reconstruction training sets.

[0012] Specifically, the target code is reconstructed and judged, which includes: detecting the method code corresponding to the old comment, and if there is a function call therein, adding the comment of the corresponding function to the old comment.

[0013] The present invention also provides a code comment updating device combining context environment and reconstruction judgment, comprising: a data processing unit, a model training unit, a reconstruction unit and a detection unit, wherein: the data processing unit is used to obtain a code training set, and establish a correspondence between method code and annotation according to method description information in the annotation; the code training set includes old method code, old annotation, new method code and new annotation; the code training set is processed to exclude training data in which no method code description information is added or changed; the model training unit is used to train a coding and decoding model using the code training set; the coding and decoding model includes a coding layer, a decoding layer and a collaborative attention layer, the coding layer includes a new code encoder, a deleted code encoder, a replaced code encoder and an old comment encoder; the decoding layer includes a new comment decoder; the coding layer is connected to the collaborative attention layer so that the input features of each encoder are linked to the input features of other encoders, and the collaborative attention layer outputs global features; the reconstruction unit is used to use the code reconstruction model to perform code reconstruction judgment on the target code, excluding the code reconstructed part; the detection unit is used to detect the target code using the coding and decoding model, and query the annotations that should be updated in the target code.

[0014] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: comprehensively judging whether code comments should be updated based on the data features of the target code context; considering the impact of code refactoring to avoid misjudgment; solving the OOV (OutOf Vocabulary) problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A schematic diagram of the flow chart of the code annotation updating method provided by the present invention;

[0016] Figure 2 A schematic diagram of the framework of the encoding and decoding model provided by the present invention. DETAILED DESCRIPTION

[0017] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0018] See also Figure 1 , which is a flow chart of the code annotation updating method provided by the present invention.

[0019] In the embodiment of the present invention, a code training set is obtained, and a correspondence between method codes and annotations is established according to method description information in the annotations.

[0020] In the embodiment of the present invention, commit information can be extracted from the code training set to obtain method codes and comments. The method code includes the old method code (method_old) and the new method code (method_new), and the comment includes the old comment (comment_old) and the new comment (comment_new); there is a corresponding relationship between the old method code and the updated new method code, and there is a corresponding relationship between the old comment and the updated new comment. The code training set is a labeled original data set, which is used to train the model.

[0021] In the embodiment of the present invention, the code training set is processed to exclude training data in which the method code description information is not added or changed.

[0022] In the specific implementation, in order to obtain a high-quality training set, the method codes and comments can be processed, mainly excluding the method codes and corresponding comments that have no substantial changes in content. Whether there are substantial changes in content is mainly reflected in whether the new annotations have added or changed the method code description information compared to the corresponding old annotations. If they have been added or changed, it indicates that there are substantial changes. If old annotations are found during the model detection process, the old annotations should be annotated and updated.

[0023] In the embodiment of the present invention, the processing of the code training set includes deleting inline comments in the method code and deleting tags of the comments.

[0024] In practice, inline comments are relatively simple comments that are placed on the same line as expressions or statements, usually separated by two spaces. The comment part starts with a pound sign and a space. Inline comments cannot fully describe the meaning of method code, and are not helpful for model training and obtaining code context features. The comment label is also not helpful.

[0025] In an embodiment of the present invention, tokens in old comments are split. If the token is not in the corresponding method code and is not in the vocabulary, it is considered a spelling error. In the case of a spelling error, if the new comment only performs spelling correction, the corresponding training data is excluded; if only Java annotations are added to the changed code, the corresponding training data is excluded; NLP (Natural Language Processing) natural grammar detection is used to exclude old comments and new comments that do not conform to grammatical rules; the old comments and the corresponding new comments are split into tokens respectively, and if the tokens are consistent, the corresponding training data is excluded.

[0026] In the embodiment of the present invention, the new code, deleted code, replaced code and old comments are obtained by comparing each clause of the old method code and the corresponding new method code; the new code (code_add), deleted code (code_delete) and replaced code (code_replace) all belong to the changed code (code_change).

[0027] In the specific implementation, the corresponding training data is excluded, including excluding the corresponding old method code, old comments, new method code and new comments; excluding the old comments and new comments that do not conform to the grammatical rules, and the corresponding

[0028] In the specific implementation, Java annotations include @Test, @Override, and @suppressWarnings. If only Java annotations are added, it is considered unnecessary to update the annotations, and such training data is excluded.

[0029] In the specific implementation, the old annotations and the corresponding new annotations are split into tokens respectively. If the tokens of the old annotations and the new annotations are consistent, it is considered that the only change between the old annotations and the new annotations is the text format, and such training data is excluded.

[0030] In the embodiment of the present invention, after excluding part of the training data, the annotations (old annotations and new annotations) are split according to camel case naming, and a connecting character is inserted between each identifier.

[0031] In the specific implementation, the code contains many compound identifiers. In order to prevent the OOV (Out Of Vocabulary) problem from destroying the format of the original comments, the identifiers are split according to camel case naming, and characters (connection characters) are inserted in the middle of each identifier to better indicate whether a token is part of an identifier. For example, for a compound word such as read-only, it becomes readonly after processing, which effectively reduces the OOV problem.

[0032] In the specific implementation, in order to ensure the sufficiency of training data, the data of modified logical statements such as if and for are copied to facilitate better training and learning of the neural network model.

[0033] See also Figure 2 , which is a schematic diagram of the framework of the encoding and decoding model provided by the present invention.

[0034] In the embodiment of the present invention, a code training set is used to train the encoding and decoding model.

[0035] In an embodiment of the present method, the encoding and decoding model includes an encoding layer, a decoding layer and a collaborative attention layer, the encoding layer includes a new code encoder (code_add Encoder), a deletion code encoder (code_delete Encoder), a replacement code encoder (code_replace Encoder) and an old comment encoder (comment_old Encoder); the decoding layer includes a new comment decoder; the encoding layer is connected to the collaborative attention layer so that the output features of each encoder are linked to the output features of other encoders, and the collaborative attention layer outputs global (taking into account context) features.

[0036] In an embodiment of the present invention, the encoding and decoding model (Encoder-Decoder) (Fine-tune model) also includes a codeBert model embedding layer and a fully connected layer. New codes, deleted codes, replaced codes and old comments are input into the encoding side codeBert model embedding layer. The output of the encoding side codeBert model embedding layer is connected to the encoding layer, the encoding layer is a bidirectional LSTM, and the output of the encoding layer is input into the collaborative attention layer; the new comments are input into the decoding side codeBert model embedding layer, the output of the decoding side codeBert model embedding layer is connected to the decoding layer, and the decoding layer is a unidirectional LSTM; the output of the collaborative attention layer and the output of the decoding layer are both input into the fully connected layer, and the fully connected layer outputs the weights.

[0037] In the specific implementation, the codeBert model on the encoding side serves as a unified embedding layer. Each code token and comment token have the same feature vector, and the code and comment tokens have the same feature vector, which solves the OOV problem to a certain extent and makes the tokens in the two have the same semantics. After the embedding layer, each token is vectorized into e i .

[0038] In the specific implementation, after the embedding layer of the codeBert model on the encoding side, it enters the encoding layer. The encoding layer is a two-layer bidirectional LSTM (LSTM, Long Short-Term Memory) to facilitate semantic capture. The specific formula is as follows: i =B i LSTM(h i-1 ,h i+1 , e i ), where h i is the output of the LSTM module in the new code encoder. Similarly, the LSTM outputs of the replacement code encoder, the deletion code encoder, and the old comment encoder are the same as the above formula. The outputs of the corresponding LSTM modules are h i '、h i ” and h i ”'.

[0039] In the specific implementation, the input of the collaborative attention layer is the output of the encoding layer. The output features of each encoder are linked to the output features of other encoders. The output of the collaborative attention layer takes into account the outputs of all encoders, that is, it integrates the contextual information features.

[0040] In the specific implementation, the formula of the collaborative attention layer is as follows: i =H'β i , β i =softmax(H' T ω β h i ), where b i is the output of the encoder, β i is the attention parameter corresponding to the encoder output, and H' is h i 'The matrix formed, ω β Is β i The matrix formed.

[0041] In the specific implementation, the fully connected layer (Dense+softmax) outputs weights. Based on the comparison between the weight value and the threshold, it can be determined whether a specific annotation should be updated. The threshold can be obtained by self-training of the model or set according to the actual scenario.

[0042] In an embodiment of the present invention, a code reconstruction model is used to perform code reconstruction judgment on the target code, and the code reconstructed part is excluded. Then, the encoding and decoding model is used to detect the target code and query the comments that should be updated in the target code.

[0043] In the embodiment of the present invention, the code reconstruction model is obtained by training the XGboost model using old comments and changed codes as reconstruction training sets.

[0044] In the specific implementation, many new comments do not have much significance for the changes to the old comments, and do not change the original semantics. This part of the data will undoubtedly reduce the accuracy of the model. Not every code change requires an annotation update. The most typical case is code refactoring. Therefore, before updating the comments, a code refactoring judgment is added. It is an XGboost model used to classify whether the current code is a code refactoring. In the refactoring training set used for XGboost model training, each training data includes the changed code (code_change) of the refactored code, the old comments, and the label of whether it is refactored. In the comment update detection phase, if the refactoring classifier considers it to be a code refactoring, it ends directly and there is no need to update the comments.

[0045] In the embodiment of the present invention, method call detection is performed after code refactoring judgment, specifically, the method code corresponding to the old comment is detected, and if there is a function call therein, the comment of the corresponding function is added to the old comment.

[0046] In the specific implementation, the update of comments is related to the context of the current method code, so a method call detection step can be added. If a function call is found in the method code, the corresponding function comment is added to the old comment. This can improve the information features of the morning and afternoon, and the update detection results can be more accurate and comprehensive.

[0047] The present invention also provides a code comment updating device combining context environment and reconstruction judgment, comprising: a data processing unit, a model training unit, a reconstruction unit and a detection unit, wherein: the data processing unit is used to obtain a code training set, and establish a correspondence between method code and annotation according to method description information in the annotation; the code training set includes old method code, old annotation, new method code and new annotation; the code training set is processed to exclude training data in which no method code description information is added or changed; the model training unit is used to train a coding and decoding model using the code training set; the coding and decoding model includes a coding layer, a decoding layer and a collaborative attention layer, the coding layer includes a new code encoder, a deleted code encoder, a replaced code encoder and an old comment encoder; the decoding layer includes a new comment decoder; the coding layer is connected to the collaborative attention layer so that the input features of each encoder are linked to the input features of other encoders, and the collaborative attention layer outputs global features; the reconstruction unit is used to use the code reconstruction model to perform code reconstruction judgment on the target code, excluding the code reconstructed part; the detection unit is used to detect the target code using the coding and decoding model, and query the annotations that should be updated in the target code.

[0048] In the embodiment of the present invention, the data processing unit is used to delete inline comments in the method code and delete the tags of the comments.

[0049] In an embodiment of the present invention, the data processing unit is used to obtain new codes, deleted codes, replaced codes and old comments by comparing each clause of the old method code and the corresponding new method code; the new codes, deleted codes and replaced codes all belong to changed codes.

[0050] In an embodiment of the present invention, the data processing unit is used to split tokens in old comments. If the token is not in the corresponding method code and is not in the vocabulary, it is considered a spelling error. In the case of a spelling error, if the new comment only performs spelling correction, the corresponding training data is excluded; if only Java annotations are added to the changed code, the corresponding training data is excluded; use NLP grammar detection to exclude old comments and new comments that do not conform to grammatical rules; split the old comments and the corresponding new comments into tokens respectively, and if the tokens are consistent, the corresponding training data is excluded.

[0051] In the embodiment of the present invention, the data processing unit is used to split the comments according to camel case naming and insert a connecting character between each identifier.

[0052] In an embodiment of the present invention, the encoding and decoding model also includes a codeBert model embedding layer and a fully connected layer. New codes, deleted codes, replaced codes and old comments are input into the encoding side codeBert model embedding layer. The output of the encoding side codeBert model embedding layer is connected to the encoding layer, the encoding layer is a bidirectional LSTM, and the output of the encoding layer is input into the collaborative attention layer; the new comments are input into the decoding side codeBert model embedding layer, the output of the decoding side codeBert model embedding layer is connected to the decoding layer, and the decoding layer is a unidirectional LSTM; the output of the collaborative attention layer and the output of the decoding layer are both input into the fully connected layer, and the fully connected layer outputs the weight.

[0053] In the embodiment of the present invention, the reconstruction unit is used to train the XGboost model using the old comments and the changed code as a reconstruction training set to obtain the code reconstruction model.

[0054] In the embodiment of the present invention, the data processing unit is used to detect the method code corresponding to the old comment after the code reconstruction judgment, and if there is a function call therein, the comment of the corresponding function is added to the old comment.

Claims

1. A code comment updating method combining context environment and refactoring judgment, characterized in that: include: Obtain the code training set, and establish the correspondence between method code and annotation based on the method description information in the annotation; The code training set includes old method codes, old annotations, new method codes and new annotations; Process the code training set and exclude the training data in which no method code description information is added or changed; The encoding and decoding model is trained using the code training set; the encoding and decoding model includes an encoding layer, a decoding layer, a collaborative attention layer, a codeBert model embedding layer and a fully connected layer, the encoding layer includes a new code encoder, a deleted code encoder, a replacement code encoder and an old comment encoder; the decoding layer includes a new comment decoder; the new code, deleted code, replacement code and old comment are input into the encoding side codeBert model embedding layer, the output of the encoding side codeBert model embedding layer is connected to the encoding layer, the encoding layer is a bidirectional LSTM, and the output of the encoding layer is input into the collaborative attention layer; the encoding layer is connected to the collaborative attention layer so that the output features of each encoder are linked to the output features of other encoders, and the collaborative attention layer outputs global features; The new annotation is input into the embedding layer of the codeBert model on the decoding side. The output of the embedding layer of the codeBert model on the decoding side is connected to the decoding layer, which is a unidirectional LSTM. The output of the collaborative attention layer and the output of the decoding layer are both input into the fully connected layer, which outputs the weights. Based on the comparison between the weight value and the threshold, determine whether the annotation should be updated; The code refactoring model is used to judge the code refactoring of the target code, and the code refactoring parts are excluded. Then, the encoding and decoding model is used to detect the target code and query the comments that should be updated in the target code.

2. The code comment updating method combining context environment and reconstruction judgment according to claim 1 is characterized in that: The processing of the code training set includes: Delete the inline comments in the method code and remove the comment tags.

3. The code annotation updating method combining context environment and reconstruction judgment according to claim 1 is characterized in that: New codes, deleted codes, replaced codes and old comments are obtained by comparing each clause of the old method code and the corresponding new method code; new codes, deleted codes and replaced codes are all changed codes.

4. The code comment updating method combining context environment and reconstruction judgment according to claim 3 is characterized in that: The excluding of training data in which no method code description information is added or changed includes: Split the tokens in the old comments. If the token is not in the corresponding method code and is not in the vocabulary, it is considered a spelling error. In the case of a spelling error, if the new comment only has spelling corrections, exclude the corresponding training data; if only Java annotations are added to the changed code, exclude the corresponding training data; use NLP grammar detection to exclude old and new comments that do not conform to grammatical rules; split the old comments and the corresponding new comments into tokens respectively. If the tokens are consistent, exclude the corresponding training data.

5. The code comment updating method combining context environment and reconstruction judgment according to claim 4 is characterized in that: The excluding of training data in which no method code description information is added or changed, then includes: Split comments according to camelCase, inserting hyphens between each identifier.

6. The code comment updating method combining context environment and reconstruction judgment according to claim 3 is characterized in that: The code refactoring model is obtained by training the XGboost model using old comments and changed codes as a refactoring training set.

7. The code comment updating method combining context environment and reconstruction judgment according to claim 6 is characterized in that: The code refactoring model is used to perform code refactoring judgment on the target code, excluding the code refactoring part, and then includes: The method code corresponding to the old comment is detected. If there is a function call, the comment of the corresponding function is added to the old comment.

8. A code comment updating device combining context environment and reconstruction judgment, characterized in that: include: Data processing unit, model training unit, reconstruction unit and detection unit, wherein: The data processing unit is used to obtain a code training set, and establish a correspondence between method codes and comments according to the method description information in the comments; the code training set includes old method codes, old comments, new method codes and new comments; the code training set is processed to exclude training data in which no method code description information is added or changed; The model training unit is used to train the encoding and decoding model using the code training set; the encoding and decoding model includes an encoding layer, a decoding layer, a collaborative attention layer, a codeBert model embedding layer and a fully connected layer, the encoding layer includes a new code encoder, a deleted code encoder, a replacement code encoder and an old comment encoder; the decoding layer includes a new comment decoder; the new code, deleted code, replacement code and old comment are input into the encoding side codeBert model embedding layer, the output of the encoding side codeBert model embedding layer is connected to the encoding layer, the encoding layer is a bidirectional LSTM, and the output of the encoding layer is input into the collaborative attention layer; the encoding layer is connected to the collaborative attention layer so that the input features of each encoder are linked to the input features of other encoders, and the collaborative attention layer outputs global features; the new comment is input into the decoding side codeBert model embedding layer, the output of the decoding side codeBert model embedding layer is connected to the decoding layer, and the decoding layer is a unidirectional LSTM; the output of the collaborative attention layer and the output of the decoding layer are both input into the fully connected layer, and the fully connected layer outputs the weight; according to the comparison between the weight value and the threshold, it is judged whether the annotation should be updated; The refactoring unit is used to use the code refactoring model to perform code refactoring judgment on the target code, and exclude the code refactoring part thereof; The detection unit is used to detect the target code using the encoding and decoding model, and query the annotations that should be updated in the target code.

9. The code comment updating device combining context environment and reconstruction judgment according to claim 8 is characterized in that: The data processing unit is used to obtain the newly added code, deleted code, replaced code and old comments by comparing the old method code with each clause of the corresponding new method code.