Code annotation generation method and device, electronic equipment, storage medium and product
By obtaining the target program code and determining the annotation prompt rules, using a large language model to analyze the code and generate annotation text, the problem of inefficient writing code annotation by manual writing is solved, and fast, efficient and accurate code annotation generation is achieved, which improves the readability of the code.
Patent Information
- Application Number
- CN202510608698.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-13
AI Technical Summary
In the prior art, the writing of code annotations is manual, inefficient, and the writing habits and methods of different programmers lead to understanding burdens and development interference.
By obtaining the target program code and determining the annotation prompt rules, analyzing the code using a large language model, determining the code snippet to be commented, its semantic information and statement types, generating the target annotation text based on the annotation prompt rules and statement types, and inserting it into the code.
It realizes fast and efficient code annotation generation, ensuring the accuracy and practicality of the annotation, and improving the readability of the code.
Smart Images

Figure CN120144167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method for generating code comments, and also relates to a code comment generation device, an electronic device, a non-volatile storage medium, and a computer program product. Background Art
[0002] Inline code comments, as an important part of software documentation, are widely distributed in software projects. They can describe the logic and implementation details of the code in the form of natural language, and are necessary text practices in the software development process to improve the readability of the code in software projects, which is helpful for developers to understand and review the code. Therefore, writing comments has become an important activity for programmers in programming tasks. Currently, the writing of computer program code comments mainly relies on programmers to write manually, with relatively low efficiency; moreover, due to different writing habits and writing methods of different programmers, it is easy to cause an understanding burden and development interference to other programmers.
[0003] Therefore, how to achieve fast and efficient code comments, ensure the accuracy and practicality of the comments, and thus effectively improve the readability of the code is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0004] The object of the present invention is to provide a method for generating code comments, which can achieve fast and efficient code comments, ensure the accuracy and practicality of the comments, and thus effectively improve the readability of the code; another object of the present invention is to provide a code comment generation device, an electronic device, a non-volatile storage medium, and a computer program product, all of which have the above beneficial effects.
[0005] In a first aspect, the present invention provides a method for generating code comments, including: Obtain a target program code and determine a comment hint rule; the comment hint rule includes the correspondence between the code statement type and the code comment type; Use a large language model to analyze the target program code to determine the code snippet to be commented in the target program code, as well as the semantic information and statement type of the code snippet to be commented; According to the comment hint rule and the statement type of the code snippet to be commented, determine the comment type corresponding to the code snippet to be commented; Generate a target comment text corresponding to the code snippet to be commented according to the comment type and semantic information corresponding to the code snippet to be commented, and insert the target comment text into the comment position corresponding to the code snippet to be commented.
[0006] Among them, the large language model is used to analyze the target program code to determine the code snippets to be annotated in the target program code, as well as the semantic information and statement types of the code snippets to be annotated, including: The large language model is used to divide the target program code to obtain multiple target code snippets; According to the code statement types included in the annotation hint rules, the code snippets to be annotated and the statement types of the code snippets to be annotated are screened and determined in each of the target code snippets; Semantic analysis is performed on the code snippets to be annotated to obtain the semantic information of the code snippets to be annotated.
[0007] Among them, using the large language model to divide the target program code to obtain multiple target code snippets includes: The large language model is used to perform a syntax structure analysis on the target program code to determine the syntax structure of the target program code; The target program code is divided according to the syntax structure of the target program code to obtain multiple target code snippets.
[0008] Among them, the annotation hint rules also include the correspondence between the code statement type and the code annotation position; Accordingly, inserting the target annotation text into the corresponding annotation position of the code snippet to be annotated includes: Determining the corresponding annotation position of the code snippet to be annotated according to the annotation hint rules; Inserting the target annotation text into the annotation position.
[0009] Among them, the code statement type includes one or a combination of more of expression statements, variable declaration statements, return statements, loop statements, and conditional judgment statements; The code annotation type includes one or a combination of more of implementation function description type annotations, implementation method description type annotations, and design decision description type annotations.
[0010] Among them, the generation process of the annotation hint rules includes: Collecting a code dataset in a computer program project; Extracting each code annotation pair from the code dataset; Determining the statement type of the code text and the annotation type of the annotation text in the code annotation pair; According to each code annotation pair, the code text corresponding to each statement type, and the annotation text corresponding to each annotation type, determining the association relationship between each statement type and each annotation type; Generate the annotation hint rules according to the association relationships between the respective statement types and the respective annotation types.
[0011] Among them, a code dataset is collected in a computer program project, including: Determine all computer program projects in the target open-source platform; Filter out target computer program projects from all the computer program projects according to a preset filtering rule; Collect the code dataset in the target computer program project.
[0012] Among them, filtering out target computer program projects from all the computer program projects according to a preset filtering rule includes: Determine the project popularity of each of the computer program projects; Select a preset number of the computer program projects with the highest project popularity as the target computer program projects.
[0013] Among them, after selecting a preset number of the computer program projects with the highest project popularity as the target computer program projects, it further includes: Determine the target annotation ratio in each of the target computer program projects; Eliminate the target computer program projects whose target annotation ratio is lower than a preset threshold.
[0014] Among them, after selecting a preset number of the computer program projects with the highest project popularity as the target computer program projects, it further includes: Determine the project type of each of the target computer program projects; Eliminate the target computer program projects whose project type is a test project and / or a documentation project and / or an experimental project.
[0015] Among them, extracting each code annotation pair from the code dataset includes: For each computer program code in the code dataset, perform syntax analysis on the computer program code to obtain the abstract syntax tree corresponding to the computer program code; Determine the annotation nodes in the computer program code according to the abstract syntax tree, and extract the annotation text at the annotation nodes; Use a preset matching rule to determine the code text corresponding to the annotation text; Generate the code annotation pair by using each of the annotation texts and each of the code texts.
[0016] Among them, after determining the comment nodes in the computer program code according to the abstract syntax tree and extracting the comment text at the comment nodes, it further includes: Merging and processing each of the comment texts according to a preset merging rule; Cleaning and processing each of the comment texts according to a preset cleaning rule.
[0017] Among them, the code comment generation method further includes: Determining the positional relationship between each of the comment nodes and the code text corresponding to the comment node; Constructing a correspondence relationship between the code statement type and the code comment position according to each of the positional relationships and the statement types of each of the code texts; Adding the correspondence relationship between the code statement type and the code comment position to the comment hint rule.
[0018] Among them, determining the comment type of the comment text in the code comment pair includes: Performing feature extraction on the comment text in the code comment pair to obtain comment text features; Determining the comment type of the comment text according to the comment text features.
[0019] Among them, performing feature extraction on the comment text in the code comment pair to obtain comment text features includes: Performing syntax analysis on the comment text in the code comment pair by using a preset classification model to obtain each word and phrase in the comment text; Performing feature extraction on each of the word and phrases to obtain each word feature; Generating comment text features of the comment text according to each of the word features.
[0020] Among them, generating the comment hint rule according to the association relationship between each of the statement types and each of the comment types includes: Dividing all the code comment pairs according to the statement type and the comment type to obtain a plurality of code comment pair sets; each code comment pair set corresponds to the same statement type and the same comment type; Evaluating each of the code comment pair sets by using a preset quality analysis tool to obtain each evaluation result; Updating the association relationship between each of the statement types and each of the comment types according to each of the evaluation results to obtain the correspondence relationship between the code statement type and the code comment type, so as to generate the comment hint rule.
[0021] In a second aspect, the present invention further discloses a code comment generation device, including: An acquisition module, configured to acquire target program code and determine annotation hint rules; the annotation hint rules include the correspondence between code statement types and code annotation types; An analysis module, configured to analyze the target program code by using a large language model, determine the code snippet to be annotated in the target program code, and the semantic information and statement type of the code snippet to be annotated; A determination module, configured to determine the annotation type corresponding to the code snippet to be annotated according to the annotation hint rules and the statement type of the code snippet to be annotated; An annotation module, configured to generate a target annotation text corresponding to the code snippet to be annotated according to the annotation type and semantic information corresponding to the code snippet to be annotated, and insert the target annotation text into the annotation position corresponding to the code snippet to be annotated.
[0022] In a third aspect, the present invention further discloses an electronic device, including: A memory, configured to store a computer program; A processor, configured to implement the steps of any of the above code annotation generation methods when executing the computer program.
[0023] In a fourth aspect, the present invention further discloses a non-volatile storage medium, on which a computer program is stored, and the computer program, when executed by a processor, implements the steps of any of the above code annotation generation methods.
[0024] In a fifth aspect, the present invention further discloses a computer program product, including computer programs / instructions, and the computer programs / instructions, when executed by a processor, implement the steps of any of the above code annotation generation methods.
[0025] The present invention provides a code annotation generation method, including: acquiring target program code and determining annotation hint rules; the annotation hint rules include the correspondence between code statement types and code annotation types; analyzing the target program code by using a large language model to determine the code snippet to be annotated in the target program code, and the semantic information and statement type of the code snippet to be annotated; determining the annotation type corresponding to the code snippet to be annotated according to the annotation hint rules and the statement type of the code snippet to be annotated; generating a target annotation text corresponding to the code snippet to be annotated according to the annotation type and semantic information corresponding to the code snippet to be annotated, and inserting the target annotation text into the annotation position corresponding to the code snippet to be annotated.
[0026] Applying the technical solution provided by the present invention, a comment hint rule is created in advance to indicate the correspondence between the code statement type and the code comment type, so that different comment type hints can be given for codes of different statement types; a large language model is created in advance to implement semantic analysis and statement type analysis of the code; thus, for the target program code to be commented, the large language model can be used to analyze and determine each code segment to be commented therein, its semantic information and statement type, and then the comment type corresponding to each code segment to be commented can be determined in combination with the comment hint rule, so that the target comment text corresponding to the code segment to be commented can be generated according to the comment type and semantic information corresponding to the code segment to be commented and inserted into the corresponding position to complete automated code commenting. It can be seen that this technical solution does not require manual writing of comments, is faster and more efficient, can effectively ensure the accuracy and practicality of code comments, and further improve the readability of the code.
[0027] The code comment generation device, electronic device, non-volatile storage medium, and computer program product provided by the present invention also have the above technical effects, and the present invention will not elaborate herein. Brief Description of the Drawings
[0028] In order to more clearly illustrate the prior art and the technical solutions in the embodiments of the present invention, the drawings required for description in the prior art and the embodiments of the present invention will be briefly introduced below. Of course, the following drawings related to the embodiments of the present invention only describe a part of the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings, and the other obtained drawings also fall within the protection scope of the present invention.
[0029] Figure 1 It is a flowchart of a code comment generation method provided by an embodiment of the present invention; Figure 2 It is an overall architecture diagram of a code comment generation system provided by an embodiment of the present invention; Figure 3 It is a structural diagram of a code comment generation device provided by an embodiment of the present invention; Figure 4 It is a structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0030] The core of the present invention is to provide a code comment generation method, which can achieve fast and efficient code commenting, ensure the accuracy and practicality of the comments, and thus effectively improve the readability of the code; another core of the present invention is to provide a code comment generation device, an electronic device, a non-volatile storage medium, and a computer program product, all of which have the above beneficial effects.
[0031] In order to describe the technical solutions in the embodiments of the present invention more clearly and completely, the following will introduce the technical solutions in the embodiments of the present invention in combination with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0032] The embodiments of the present invention provide a method for generating code comments.
[0033] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for generating code comments provided by an embodiment of the present invention. The method for generating code comments may include the following S101 to S104.
[0034] S101: Obtain the target program code and determine the comment hint rule; the comment hint rule includes the correspondence between the code statement type and the code comment type.
[0035] This step aims to achieve the acquisition of the target program code and the determination of the comment hint rule. Among them, the target program code is the computer program code that needs to be commented, and it can be the computer program code in any development project. The comment hint rule is used to represent the correspondence between the code statement type and the code comment type, that is, codes of different statement types can correspond to different types of code comments. For example, for expression statements and variable declaration statements, the comment hint rule tends to suggest generating implementation function description type comments; for return statements and loop statements, the comment hint rule tends to suggest generating design decision description type comments. It should be noted that the comment hint rule is pre-constructed based on the statistics of a large number of data samples and has high accuracy. The following embodiments give several different code statement types and code comment types.
[0036] In an embodiment of the present invention, the code statement type may include one or more combinations of expression statements (Expression statements), variable declaration statements (Variable Declaration statements), return statements (Return statements), loop statements (For statements, While statements), conditional judgment statements (If statements, Switch statements); the code comment type may include one or more combinations of implementation function description type comments (what type comments), implementation method description type comments (how type comments), and design decision description type comments (why type comments).
[0037] It can be understood that What-type comments are mainly used to describe the core functions of code / code snippets, the uses of variables / constants, or the calculation purposes of expressions; for example, to explain the results of a complex calculation or the meaning of a key variable. Why-type comments aim to explain the reasons why the code / code snippet is designed in such a way, the considerations for choosing a specific implementation method, or the purpose of a certain logical branch; for example, to illustrate why a specific if check or loop is needed to process a certain type of data. How-type comments focus on describing the specific implementation details of the code / code snippet, the algorithm steps, or how to correctly use the code / code snippet; for example, to explain the internal working process of a complex algorithm or the calling method of an auxiliary function.
[0038] S102: Analyze the target program code using a large language model to determine the code snippets to be commented in the target program code, as well as the semantic information and statement types of the code snippets to be commented.
[0039] This step aims to implement the analysis of the target program code by the large language model to determine each code snippet to be commented therein, as well as its semantic information and statement type. Specifically, a language model can also be trained using a large-scale data sample and directly called when using this large language model. The target program code is input into the large language model for processing, and the output of the model is each code snippet to be commented in the target program code, as well as the semantic information and statement type of each code snippet to be commented. Among them, the semantic information can be obtained by the large language model through context understanding of the target program code, and the statement type can be obtained by the large language model through syntax structure analysis of the target program code.
[0040] In an embodiment of the present invention, analyzing the target program code using a large language model to determine the code snippets to be commented in the target program code, as well as the semantic information and statement types of the code snippets to be commented, may include: Divide the target program code using a large language model to obtain multiple target code snippets; According to the code statement types included in the annotation hint rules, screen and determine the code snippets to be commented and the statement types of the code snippets to be commented in each target code snippet; Perform semantic analysis on the code snippets to be commented to obtain the semantic information of the code snippets to be commented.
[0041] It can be understood that the target program code generally consists of multiple target code segments. However, in a complete target program code, not all target code segments need to be commented. As mentioned above, the annotation hint rule is used to represent the correspondence between the code statement type and the code annotation type. Essentially, it also hints at the code statement types that generally need to be commented in the computer program code. Based on this, the complete target program code can be first divided into multiple target code segments, and then the large language model can be used to screen out the code segments to be commented based on the code statement types involved in the annotation hint rule, that is, only screen out the target code segments corresponding to the statement types included in the annotation hint rule among all target code segments, and finally use the large language model to perform semantic analysis on each code segment to be commented.
[0042] As mentioned above, the statement type can be obtained by the large language model through syntax structure analysis of the target program code. Therefore, in a possible implementation, using the large language model to divide the target program code to obtain multiple target code segments may include: using the large language model to perform syntax structure analysis on the target program code to determine the syntax structure of the target program code; dividing the target program code according to the syntax structure of the target program code to obtain multiple target code segments.
[0043] S103: Determine the annotation type corresponding to the code segment to be commented according to the annotation hint rule and the statement type of the code segment to be commented.
[0044] The purpose of this step is to determine the annotation type corresponding to the code segment to be commented. Obviously, this annotation type is used to indicate which type of annotation is suitable for inserting into the corresponding code segment to be commented. It can be understood that since the annotation hint rule records the correspondence between the code statement type and the code annotation type, after determining the statement type of the code segment to be commented, the annotation type corresponding to each code segment to be commented can be determined by querying this annotation hint rule. It should be noted that this step can also be implemented based on the large language model, that is, combined with S102, input the target program code and the annotation hint rule into the large language model together, so that the large language model can refer to this annotation hint rule to implement the analysis and processing of the target program code. At this time, the output of the model is the semantic information of the code segment to be commented and its corresponding annotation type.
[0045] S104: Generate the target annotation text corresponding to the code segment to be commented according to the annotation type and semantic information corresponding to the code segment to be commented, and insert the target annotation text into the annotation position corresponding to the code segment to be commented.
[0046] This step aims to generate and insert annotation text. Specifically, for each code snippet to be annotated, after determining its semantic information and the corresponding annotation type, the target annotation text corresponding to it can be generated by combining the semantic information and the annotation type, and the target annotation text can be inserted into the annotation position corresponding to the corresponding code snippet to be annotated. Thus, the automatic annotation of computer program code is completed.
[0047] In an embodiment of the present invention, the annotation hint rule may further include the correspondence between the code statement type and the code annotation position; correspondingly, inserting the target annotation text into the annotation position corresponding to the code snippet to be annotated may include: Determine the annotation position corresponding to the code snippet to be annotated according to the annotation hint rule; Insert the target annotation text into the annotation position.
[0048] Specifically, to achieve the precise insertion of the target annotation text in the target program code, the correspondence between the code statement type and the code annotation position can be added to the annotation hint rule, that is, codes of different statement types can correspond to code annotations at different positions. For example, the annotation position corresponding to an expression declaration statement is usually above the expression declaration statement or may be on the same line; the annotation position corresponding to a conditional statement is usually above the conditional statement. In a possible implementation, the code annotation position may include one or a combination of more of the start position of the code snippet, the middle position of the code snippet, and the end position of the code snippet. Among them, the start position of the code snippet refers to the line before the initial line of the code snippet, the end position of the code snippet refers to the line after the last line of the code snippet, and the middle position of the code snippet refers to the same line of the code snippet.
[0049] It can be seen that the code annotation generation method provided by the embodiment of the present invention pre-creates an annotation hint rule to indicate the correspondence between the code statement type and the code annotation type, so that different annotation type hints can be given for codes of different statement types; a large language model is pre-created to implement semantic analysis and statement type analysis of the code; thus, for the target program code that needs to be annotated, the large language model can be used to analyze and determine each code snippet to be annotated therein, its semantic information and statement type, and then the annotation type corresponding to each code snippet to be annotated can be determined in combination with the annotation hint rule, so that the target annotation text corresponding to it can be generated according to the annotation type and semantic information corresponding to the code snippet to be annotated and inserted into the corresponding position to complete the automated code annotation. It can be seen that this technical solution does not require manual writing of annotations, is more rapid and efficient, can effectively ensure the accuracy and practicality of code annotations, and further improves the readability of the code.
[0050] Based on the above various embodiments, regarding the annotation hint rule, this embodiment provides an implementation method. Specifically, the generation process of the annotation hint rule may include:
[0051] S201: Collect a code data set in a computer program project.
[0052] As described above, the annotation hint rule is pre-constructed based on large-scale data sample statistics. Therefore, when constructing the annotation hint rule, it is necessary to collect large-scale data samples, and this data sample is a set of computer program codes with annotations, that is, the above-mentioned code data set, which can be specifically obtained by collecting in various computer program projects that have been developed or are being developed.
[0053] In a possible implementation manner, collecting a code data set in a computer program project may include: Determine all computer program projects in the target open source platform; Filter out target computer program projects from all computer program projects according to a preset filtering rule; Collect a code data set in the target computer program project.
[0054] This embodiment provides an implementation method for collecting a code data set in a computer program project. It can be understood that a large number of computer program projects generally gather in the open source platform. Therefore, a part of applicable target computer program projects can be filtered out in the specified open source platform to implement the collection of the code data set. Of course, the filtering rule for computer program projects and the specific type of the target open source platform can be set by technicians according to actual needs, and the present invention does not limit this. In a possible implementation manner, the target open source platform can specifically be the GitHub open source platform. The GitHub open source platform is the world's largest open source code hosting platform, which gathers a large number of high-quality computer program projects with active development and standardized maintenance.
[0055] Among them, filtering out target computer program projects from all computer program projects according to a preset filtering rule may include: determining the project popularity of each computer program project; selecting a preset number of computer program projects with the highest project popularity as target computer program projects.
[0056] This embodiment provides a preset screening rule, that is, multiple computer program projects with relatively high project popularity can be selected as target computer program projects in the target open-source platform. It can be understood that computer program projects with relatively high project popularity are generally more popular computer program projects. These projects are usually actively developed and may contain more inline comments. Taking the GitHub open-source platform as an example, the project popularity ranking of computer program projects can be achieved according to the "star (star mark, meaning "favorite") quantity" (the higher the star quantity, the higher the project popularity), and then the screening of target computer program projects can be achieved. Of course, the specific value of the preset quantity does not affect the implementation of this technical solution and can be set according to the actual situation. The present invention does not limit this.
[0057] Furthermore, after selecting a preset number of computer program projects with the highest project popularity as target computer program projects, it may further include: determining the target annotation ratio in each target computer program project; and removing the target computer program projects with a target annotation ratio lower than the preset threshold.
[0058] It can be understood that the annotation hint rule is used to give annotation hints. Therefore, the code data set used to construct the annotation hint rule should be a collection of computer program codes with annotations, and the annotation content should be relatively large to effectively ensure the accuracy of the annotation hint rule. Based on this, for each target computer program project, the target annotation ratio can also be separately counted, and some target computer program projects with a relatively low target annotation ratio can be removed. Similarly, the specific value of the preset threshold does not affect the implementation of this technical solution and can be set according to the actual situation. The present invention does not limit this.
[0059] It should be noted that when used to implement Chinese annotation generation, the above target annotation ratio should be the Chinese annotation ratio; when used to implement English annotation generation, the above target annotation ratio should be the English annotation ratio.
[0060] Furthermore, after selecting a preset number of computer program projects with the highest project popularity as target computer program projects, it may further include: determining the project type of each target computer program project; and removing the target computer program projects with a project type of test projects and / or document projects and / or experimental projects.
[0061] This embodiment aims to exclude toy computer program projects (such as the above-mentioned test projects, document projects, and experimental projects) because toy computer program projects are different from development computer program projects and cannot effectively ensure the accuracy of annotation hint rules. During the implementation process, the determination and exclusion of toy computer program projects can be achieved through keyword matching technology. The keywords that can be used for matching can be "toy", "test", "experiment", "learn", "exercises", etc. Just match the above keywords in the readme files of each target computer program project. To further ensure accuracy, for each target computer program project after excluding toy computer program projects, the readme file and code library can be manually checked to ensure that the target computer program projects used to collect the code dataset are all non-toy computer program projects, that is, the above-mentioned development computer program projects.
[0062] S202: Extract each code-annotation pair from the code dataset.
[0063] This step aims to achieve the extraction of code-annotation pairs, that is, to extract each code-annotation pair from each code data sample in the code dataset. The code-annotation pair is a combination of "code" and "annotation", and the two are in one-to-one correspondence. Its representation form can be <annotation, code>.
[0064] In a possible implementation manner, extracting each code-annotation pair from the code dataset may include: For each computer program code in the code dataset, perform syntax analysis on the computer program code to obtain the abstract syntax tree corresponding to the computer program code; Determine the annotation nodes in the computer program code according to the abstract syntax tree, and extract the annotation text at the annotation nodes; Use the preset matching rules to determine the code text corresponding to the annotation text; Generate code-annotation pairs using each annotation text and each code text.
[0065] This embodiment provides an implementation method for extracting code comment pairs from a code dataset. It can be understood that extracting code comment pairs from a code dataset essentially means extracting the existing annotation text and the corresponding code text in each code data sample (computer program code) in the code dataset. On this basis, for each computer program code in the code dataset, the determination of the annotation node therein can be achieved by constructing its corresponding abstract syntax tree. This annotation node is the annotation position, and thus the extraction of the annotation text can be realized at this annotation node; further, the code text corresponding to the annotation text is matched using a preset matching rule. At this time, the annotation text and the code text form a code comment pair. Among them, the construction of the abstract syntax tree can be implemented based on the JDT tool (Java Development Tools, a Java development tool) provided by the Eclipse platform (a development platform).
[0066] In addition, the preset matching rule can specifically be a rule for the position correspondence relationship between the annotation text and the code text, and the code text corresponding to the annotation text is determined through position matching. As an example above, the annotation position corresponding to an expression statement is usually above the expression statement or may be on the same line; the annotation position corresponding to a conditional statement is usually above the conditional statement. Generally speaking, the positions of the annotation text and its corresponding code text are relatively close. On this basis, the following preset matching rules can be referred to: (1) If the annotation text and the code text are on the same line (such as a simple variable declaration code statement), then select this code text as the corresponding code text of the annotation text; (2) If the annotation text is written before a code block, usually marked with a left parenthesis (for example, the if{...} statement block), then this code block is regarded as the code text corresponding to the annotation text; (3) If the annotation is on different code lines, then select all the code statements at the same level as the annotation text (that is, single-line code statements or code blocks before reaching an empty line or the next annotation text) as the code text corresponding to this annotation text; (4) If the annotation text does not match any of the above three matching rules, for example, the annotation text is written on the last line of the code text, then this annotation text can be ignored.
[0067] Furthermore, after determining the annotation nodes in the computer program code according to the abstract syntax tree and extracting the annotation text at the annotation nodes, it can also include: merging each annotation text according to a preset merging rule; cleaning each annotation text according to a preset cleaning rule.
[0068] First, in actual projects, developers often split a semantically complete comment into multiple lines of " / / " comments. To restore its semantic boundary, after obtaining each comment text, it can also be merged to obtain the complete comment text. In one possible implementation, the merging rules may include: (1) Inline comments are composed of consecutive line comments and there are no blank lines among them; (2) Inline comments are not composed of consecutive line comments, but the corresponding code is composed of multiple single-line codes.
[0069] Second, to effectively improve the quality and usage effect of comment hint rules, after obtaining each comment text, it can also be cleaned to obtain applicable comment text. In one possible implementation, the cleaning rules may include: (1) Clean out empty comments. Such comments do not contain text information related to the code text. Therefore, inline comments containing blank text will be deleted; (2) Clean out the comment text corresponding to the commented code. Such comments are usually discarded code that has been commented out. Therefore, inline comments containing discarded source code will be deleted; (3) Clean out non-English comments / English comments. When used to implement Chinese comment generation, clean out English comments; when used to implement English comment generation, clean out non-English comments. Taking the cleaning of non-English comments as an example, non-English-written comment text can be deleted by checking whether the comment can be encoded using only ASCII characters. Comment text that cannot be encoded will be deleted because it contains characters from other alphabets.
[0070] S203: Determine the statement type of the code text and the comment type of the comment text in the code comment pair.
[0071] This step aims to determine the code text type and the comment text type to facilitate the construction of comment hint rules. As mentioned above, the code statement type can include one or a combination of multiple types such as expression statements, variable declaration statements, return statements, loop statements, and conditional judgment statements; the code comment type can include one or a combination of multiple types such as implementation function description comments, implementation method description comments, and design decision description comments.
[0072] In one possible implementation, determining the comment type of the comment text in the code comment pair may include: Extract features from the comment text in the code comment pair to obtain comment text features; Determine the comment type of the comment text based on the comment text features.
[0073] Among them, feature extraction is performed on the annotation text in the code annotation pair to obtain annotation text features, which may include: performing syntax analysis on the annotation text in the code annotation pair using a preset classification model to obtain each word and phrase in the annotation text; performing feature extraction on each word and phrase to obtain each word feature; and generating annotation text features of the annotation text according to each word feature.
[0074] Furthermore, the preset classification model can specifically adopt a BERT (Bidirectional Encoder Representation from Transformers) model; the extracted feature types may include but are not limited to: the number of tokens TokenNum in the inline annotation, whether the inline annotation only contains punctuation marks OnlySymbol, specific prepositional phrases PrepStr and the quantity PrepNum in the inline annotation, specific conjunctive phrases ConjunStr and the quantity ConjunNum in the inline annotation, words Keywords that may represent feature types in the inline annotation, and the ratio Ratio of the tokens in the inline annotation to the tokens in the corresponding code.
[0075] Specifically, larger TokenNum and Ratio may indicate that the annotation has a higher probability of being an implementation detail or an explanatory intention type, that is, an annotation of the How or Why type; PrepNum and ConjunNum represent the specific relationships (prepositions or conjunctions) in the annotation. For example, the word because may indicate an annotation of the Why type; Keywords represent specific words that may indicate the type. For example, the word via may indicate an annotation of the How type; OnlySymbol indicates that the annotation only contains symbols. For example, an annotation that only contains "———" cannot be classified as the What, How, and Why types.
[0076] S204: Determine the association relationship between each statement type and each annotation type according to each code annotation pair, the code text corresponding to each statement type, and the annotation text corresponding to each annotation type.
[0077] S205: Generate annotation prompt rules according to the association relationship between each statement type and each annotation type.
[0078] The above steps are aimed at determining the association relationship between statement types and comment types, and then generating comment hint rules. It can be imagined that based on the above code comment pairs, the code texts corresponding to each statement type, and the comment texts corresponding to each comment type, the number of code texts corresponding to each statement type and the number of comment texts corresponding to each comment type can be statistically obtained. Then, through comprehensive analysis, the text proportion of each comment type corresponding to each statement type can be obtained. Taking expression statements as an example, according to the total number of comment texts corresponding to them and the comment types to which the comment texts belong, the proportion of What-type comment texts, the proportion of How-type comment texts, the proportion of Why-type comment texts, and the proportion of Other-type comment texts can be obtained respectively. Thus, by referring to the association relationship between each statement type and each comment type, the generation of comment hint rules can be achieved to indicate which statement type of code text is suitable for generating which comment type of comment text.
[0079] In a possible implementation manner, generating comment hint rules according to the association relationship between each statement type and each comment type may include: Dividing all code comment pairs according to statement types and comment types to obtain multiple code comment pair sets; each code comment pair set corresponds to the same statement type and the same comment type; Using a preset quality analysis tool to evaluate each code comment pair set to obtain each evaluation result; Updating the association relationship between each statement type and each comment type according to each evaluation result to obtain the correspondence between code statement types and code comment types, so as to generate comment hint rules.
[0080] To further improve the quality and usage effect of comment hint rules, a quality evaluation method can be used to verify and update the association relationship between statement types and comment types, so as to effectively improve the accuracy of comment hint rules. Among them, the preset quality evaluation tool for implementing quality evaluation can be the PMD tool and the Checkstyle tool (which belong to two different static analysis tools) to effectively reduce the error analysis caused by a single tool. PMD can check Java source code to identify potential problems and provide improvement suggestions; Checkstyle can statically check Java code according to specified coding conventions. Both of these tools can evaluate the input Java source code according to a set of rules and output an error report of possible violations item by item after analysis.
[0081] Furthermore, the construction process of the above comment hint rules may also include: Determining the positional relationship between each comment node and the code text corresponding to the comment node; Construct the correspondence between the code statement type and the code comment position according to the positional relationships and the statement types of the code texts; Add the correspondence between the code statement type and the code comment position to the comment hint rules.
[0082] As described above, in addition to the correspondence between the code statement type and the code comment type, the comment hint rules can also include the correspondence between the code statement type and the code comment position. Based on this, this embodiment provides a method for generating the correspondence between the code statement type and the code comment position. It can be imagined that the correspondence between the code statement type and the code comment position is similar to the correspondence between the code statement type and the code comment type, and both can be realized through big data statistical analysis, which will not be elaborated here.
[0083] This embodiment of the present invention provides another code comment generation method.
[0084] Please refer to Figure 2 , Figure 2 , which is the overall architecture diagram of a code comment generation system provided by this embodiment of the present invention. The code comment generation system mainly includes the following four parts: construction of the data sample set, statistical analysis of the comment position and content, correlation analysis between the comment content and the code quality, and automatic generation of inline comments. Based on Figure 2 the code comment generation system shown, the implementation process of the code comment generation method provided by this embodiment can include:
[0085] I. Construction of the data sample set.
[0086] 1. Project selection.
[0087] To ensure the representativeness and quality of the comment data, the GitHub open-source platform can be selected as the data source, and Java projects that meet the following criteria can be selected as potential experimental projects (target computer program projects):
[0088] (1) Popular projects: Popular projects are usually actively developed and may contain more inline comments; specifically, the popularity of projects on GitHub can be measured by referring to the number of Stars. Rank the Java projects according to the number of Stars, and select a certain number of projects with the most Stars and ranked at the top as experimental projects.
[0089] (2) Projects with English comments: That is, only consider projects with comments written in English; specifically, first check whether the comments are in ASCII encoding, and then calculate the percentage of all ASCII-encoded comments in the project. If the percentage exceeds 90%, this project will be regarded as a project with English comments, otherwise it will be regarded as a non-English comment project and removed from the potential projects.
[0090] (3) Software development projects that are not toys: Potential experimental projects mainly need to be software development projects, rather than document / experiment / test projects; specifically, the heuristic mode can be used first to identify potential toy projects, check whether their readme files contain keywords such as "toy", "test", "experiment", "learn", "exercises", etc., and then manually check the readme files and code libraries of each obtained project to determine whether it is a toy project.
[0091] Based on the above screening rules, 1000 projects are statistically obtained as experimental projects. The statistical information of these projects is shown in Table 1, and Table 1 is a project statistical information table provided by an embodiment of the present invention.
[0092] Table 1 Project Information Statistical Table
[0093]
[0094] As can be seen from Table 1, 10.5% (the ratio of the total number of 599,233 "methods with inline comments" to the total number of 5,702,126 "methods") of the methods contain inline comments, and only 2.3% (the ratio of the total number of 132,896 "methods with both method header comments and inline comments" to the total number of 5,702,126 "methods") of the methods contain both method header comments and inline comments. This means that in 8.2% (the difference between 10.5 and 2.3%) of the methods, developers can only rely on inline comments for code understanding activities. Therefore, it is very meaningful to explore the location, content, and automatic generation method of inline comments.
[0095] 2. Extraction of annotation-code pairs.
[0096] Based on the JDT tool provided by the Eclipse platform, an abstract syntax tree of the source code can be constructed, and it supports identifying annotation nodes (such as LineComment, BlockComment) and their attached statement nodes (such as expressions, declarations, control structures, etc.). Thus, by traversing the abstract syntax tree, the annotation nodes can be automatically located, and their adjacent code statements can be obtained to achieve preliminary <annotation, code> extraction.
[0097] To improve the structural integrity and semantic coherence of the annotations, the following rules can be referred to for merging inline comments and determining the corresponding code of the inline comments:
[0098] First, in actual projects, developers often split a semantically complete comment into multiple lines of " / / " comments. The following merging rules can be adopted: (1) Inline comments are composed of consecutive line comments and there are no blank lines among them; (2) Inline comments are not composed of consecutive line comments, but the corresponding code is composed of multiple single-line codes. Take the following code snippet with comments as an example:
[0099] “public RGBLuminanceSource(int width, int height, int[]pixels){ ... / / In order to measure pure decoding speed, we convert the entireimage to a greyscale array / / up front,which is the same as the Y channel of theYUVLuminanceSource in the real app / / / / Total number of pixels suffices, can ignore shape int size= width * height; luminances=new byte[size]; for(int offset0;offset < size; offset++){ ...} }”。
[0100] According to merging rule (1), lines 3, 4, 5, and 6 will be merged into one comment because there are no blank lines between the code lines; according to merging rule (2), the obtained merged comment will be further merged with lines 7 and 8 to form the final comment because line 7 is blank and the corresponding code is composed of multiple single-line codes (lines 9 and 10).
[0101] Secondly, the merged inline comments need to be further precisely matched to the target code snippets they describe. The following heuristic matching methods can be adopted: (1) If the comment and the code are on the same line (such as a simple variable declaration code statement), then select this code as the corresponding code text for the comment; (2) If the comment is written before a code block, usually marked with a left parenthesis (for example, the if{...} statement block), then this code block is regarded as the corresponding code text for the comment; (3) If the comment is on different code lines, then select all the code statements at the same level as the comment (that is, single-line code statements or code blocks before reaching a blank line or the next comment text) as the corresponding code text for this comment; (4) If the comment does not match any of the above three matching rules, such as the comment is written on the last line of the code, then this comment can be ignored.
[0102] 3. Data cleaning.
[0103] To effectively improve the quality and usage effect of the comment hint rules, the following cleaning rules can be adopted: (1) Clean up empty comments. Such comments do not contain text information related to the code text. Therefore, inline comments containing blank text will be deleted; (2) Clean up the comment text corresponding to commented-out code. Such comments are usually obsolete code that has been commented out. Therefore, inline comments containing obsolete source code will be deleted; (3) Clean up non-English comments. Non-English-written comment text can be deleted by checking whether the comment can be encoded using only ASCII characters. Comment text that cannot be encoded will be deleted because it contains characters from other alphabets.
[0104] The reason for adopting the above cleaning rules is that the comment position and content are the analysis objects of this technical solution. Therefore, although some comments cover special content, the positions where they appear are also worthy of analysis and exploration. For example, a comment containing only "————" may be used to split the code. Therefore, the position where it appears is also meaningful and will not be treated as a noisy comment.
[0105] II. Statistical analysis of comment positions and content.
[0106] Sample partial data and conduct multiple rounds of analysis experiments to classify the annotation positions and contents into different categories. Specifically, the analysis experiments are completed manually and the process is divided into multiple rounds. Specifically, in each round of the experiment, 500 pairs of annotation codes are randomly sampled without replacement from the dataset. In the first round, the positions where annotations appear in the sample data are classified into different categories. When there are disagreements, the experiment participants will discuss until they reach an agreement. Then, each participant maintains the same shared category pool. When conducting a new round of data analysis, an existing category is selected from this pool or a new category is added to the shared category pool in the next round. This analysis process ends when no new categories are added in three consecutive rounds. Finally, the experiment analyzes the annotation categories from two perspectives: code structure and code syntax. After the analysis experiment is completed, a statistical analysis of the occurrence frequencies of different types of annotation positions is conducted.
[0107] Among them, the definition of annotation position categories: (1) The annotation is located in the expression statement of the code, that is, the Expression statement. Such annotation positions are usually above the expression statement or may also be on the same line; (2) The annotation is located in the variable declaration statement of the code, that is, the VariableDeclaration statement. Such annotation positions are usually above the variable declaration statement or may also be on the same line; (3) The annotation is located in the conditional statement of the code, such as the If statement. Such annotation positions are usually above the If statement structure; (4) The annotation is located in the return statement of the code, that is, the Return statement. Such annotation positions are usually above the return statement or may also be on the same line; (5) The annotation is located in the loop statement of the code, such as the For statement. Such annotation positions are usually above the loop statement; (6) The annotation is located in other positions.
[0108] Among them, the definition of annotation content categories: (1) What content type inline annotation: mainly used to describe the core function of the code / code snippet, the purpose of variables / constants, or the calculation purpose of expressions; for example, explaining the result of a complex calculation or the meaning of a key variable; (2) How content type inline annotation: focuses on describing the specific implementation details of the code / code snippet, algorithm steps, or how to correctly use the code / code snippet; for example, explaining the internal working process of a complex algorithm or the calling method of an auxiliary function; (3) Why content type inline annotation: aims to explain the reason why the code / code snippet is designed like this, the considerations for choosing a specific implementation method, or the purpose of a certain logical branch; for example, explaining why a specific if check or loop is needed to process a certain type of data.
[0109] Furthermore, after the sampling and classification are completed, the automatic classification of the annotation content is performed on the entire data set, and the annotation content type with the highest frequency of occurrence in the annotations at different positions is counted. Among them, the StanfordParser tool can be used for feature calculation of the automatic classification, and the BERT model can be used as the classification model. Specifically, the Stanford Parser tool can obtain the syntactic composition of the annotation text, which includes verb phrases (VP), noun phrases (NP), etc. In the classification effect evaluation, the BERT training uses a ten-fold cross-validation method, that is, each time a random piece of data is selected as the test set, and the other data is used as the training set. Five traditional metrics can be selected, namely Precision, Recall, F1-score, Accuracy, and Hamming Loss. Finally, the experiment counts the annotation content type with the highest frequency of occurrence in the annotations at different positions. Please refer to Table 2, which is a classification feature table of inline annotations provided by an embodiment of the present invention:
[0110] Table 2 Classification Feature Table of Inline Annotations
[0111]
[0112] Specifically, larger TokenNum and Ratio may indicate that the annotation has a higher probability of being an implementation detail or an explanatory intention type, that is, an annotation of the How or Why type; PrepNum and ConjunNum represent the specific relationships (prepositions or conjunctions) in the annotation. For example, the word because may indicate an annotation of the Why type; Keywords represent specific words that may indicate the type. For example, the word via may indicate an annotation of the How type; OnlySymbol indicates that the annotation contains only symbols. For example, an annotation that contains only "———" cannot be classified as the What, How, and Why types.
[0113] Finally, in the process of using the BERT model for classification, this technical solution achieved 90.0% Precision, 90.1% Recall, 90.0% F1-score, 90.0% Accuracy, and 0.0921 Hamming Loss in the data set respectively, indicating the effectiveness of the BERT model.
[0114] In addition, the statistical experimental results based on the classification results are shown in Table 3, which is a statistical analysis table of annotation positions and contents provided by an embodiment of the present invention:
[0115] Table 3 Statistical Analysis Table of Annotation Positions and Contents
[0116]
[0117] As can be seen from Table 3: (1) The most frequently occurring type of comment content covered by different code statement positions is the What type of comment, that is, the comments left by developers in the current software project are mostly comments summarizing the code functions, which is especially reflected in the comments corresponding to the Expression statement, VariableDeclaration statement, and For statement. The proportions of their What type comments are 65.1%, 63.9%, and 67.7% respectively. (2) There is a tendency in the comment content covered by different code statement positions. The proportion of What type comments written by developers in the If statement and Return statement is relatively low. Compared with other code statements, the proportions are 55.8% and 53.3% respectively, while the proportions of How and Why type comments are relatively high, which are 23.3% and 20.7% as well as 15.8% and 16.2% respectively. This indicates that developers tend to explain code details and code intentions more when writing comments corresponding to the If statement and Return statement.
[0118] III. Analysis of the Relevance between Comment Content and Code Quality.
[0119] After obtaining the comment position and content, it is also possible to further confirm whether the comment content at the current position will affect the code quality, that is, the type of comment content recommended for different code segments, providing high-quality comment content suggestions for the large model to generate comments. Therefore, in order to measure whether there is an association between different comment contents and code quality metrics, it can be analyzed based on quality analysis tools such as PMD and Checkstyle tools to effectively reduce the error analysis caused by a single tool. Both of them can evaluate the input Java source code according to a set of rules and output error reports of possible violations item by item after analysis. It can be understood that statistically analyzing the relationship between comment content and code quality at different positions is reflected as the difference in quality metric results and provides some insights into the hint word rules. Based on the analysis results of the quality analysis tools shown in Table 4 and Table 5, Table 4 and Table 5 are an analysis result information table of a PMD tool and an analysis result information table of a Checkstyle tool provided by the embodiments of the present invention respectively:
[0120] Table 4 Analysis Result Information Table of PMD Tool
[0121]
[0122] Table 5 Analysis Result Information Table of Checkstyle Tool
[0123]
[0124] The metric shown in Tables 4 and 5 is avg_err, and the calculation formula is: avg_err = data_num / err_num, which is the ratio of the total number of data data_num to the number of errors err_num output by the code quality analysis tool. This means that an error will appear for every number of code data. In theory, the higher the value, the better, indicating fewer errors. Among them, the total number of data refers to the total number of analyzed code snippets, and the number of errors output by the code quality analysis tool refers to the total number of error reports output by the quality analysis tool.
[0125] Based on the analysis of Tables 4 and 5, the comment prompt rules obtained are as follows: (1) What type comments have a positive correlation with Expression statements and VariableDeclaration statements. Therefore, in the subsequent comment generation, it is possible to focus on generating What type comments in these two types of code statements. (2) Why type comments have a positive correlation with Return statements, For statements, and If statements. Therefore, in the subsequent comment generation, it is possible to focus on generating Why type comments in these two types of code statements. (3) When performing the task of automatic comment generation, the generation of How type comments cannot be ignored, that is, all types of code statements may need to describe the implementation details of the code snippet, or describe how to use the corresponding code snippet.
[0126] 4. Automatic generation of inline comments.
[0127] According to the prompt word rules above, you can define the prompt words used for comment generation and insert high-quality comment content into the appropriate code location. The system has a set of built-in comment rules to guide the LLM model (large language model) to generate different types of comments. These rules provide tendency suggestions based on the analysis results of the correlation between comments and code quality. Specifically: 1. What type of comments: For statements such as Expression and VariableDeclaration, the rules tend to suggest generating such comments; 2. Why type comments: For control flow or logic key statements such as Return, For (loop), If (conditional judgment), the rules tend to recommend generating such comments; 3. How type comments: For all types of code statements, the rules will consider suggesting the generation of such comments.
[0128] Finally, the implementation process of the LLM model combining context understanding to generate inline annotation content is as follows:
[0129] 1. Input Processing: Provide the complete code of a function or method body (target program code), along with the guiding natural language prompt rules (annotation prompt rules), as input to the LLM. The role of the prompt is to guide the model to focus on specific code structure features and potential annotation requirements, rather than pre-specifying the annotation location or type.
[0130] 2. Deep Context Understanding and Structure Analysis: (1) Structure Analysis: The LLM first performs syntax and structure analysis on the entire code block to understand the function signature, statement sequence, control flow (loops, conditions), data flow, and the logical hierarchy of the code block. (2) Semantic Understanding: Understand the intention and semantics of the code, analyze the purpose and effect of function calls, understand the internal logic and calculation process of complex expressions or algorithms, and determine the importance of code segments in the overall function and their dependencies on other parts.
[0131] 3. Intelligent Annotation Point Identification: The LLM uses its deep context understanding ability and combines the focus on specific statement types in the experimental results, that is, the model is guided to give priority to those statements that usually contain core logic, define key data, or control the program flow, to autonomously identify the positions in the code that most need annotation. These types include, but are not limited to: Expression statements, especially those with complex calculations or unclear result usage; VariableDeclaration statements, especially those that declare important state variables, configuration parameters, or variables whose lifecycle / scope needs to be explained; Return statements, explaining the meaning of the return value or the reason for specific return conditions; For / While statements: explaining the purpose of the loop, the dataset being processed, or the key logic within the loop; If / Switch statements: elaborating on the reason for setting the condition, the logic of branch processing, or the coverage of specific scenarios.
[0132] 4. Determine the insertion position: Once it is determined that annotation is needed, the LLM refers to the annotation type suggestions (What / Why / How) regarding the statement type in the prompt rules, and combines its in-depth understanding of the specific context of this code point (variable names, operations involved, the logical flow before and after, the overall goal of the code, etc.) to generate high-quality inline annotation text. For example, for a complex If condition, the prompt may suggest a Why-type annotation. The LLM will analyze the specific composition of this condition (variables, comparisons, logical operators) and its role in the entire function logic (e.g., handling a specific error situation or edge case), and then generate an annotation like " / / Check [specific condition] to handle [specific reason or scenario]", rather than simply saying " / / Condition check". Another example, for a variable declaration, the prompt suggests a What-type annotation. The LLM will analyze the actual use and scope of influence of this variable in the subsequent code and generate an annotation such as " / / Store [specific data] for [subsequent key calculation / operation]", rather than just " / / Declare variable".
[0133] As can be seen, the code annotation generation method provided by the embodiments of the present invention pre-creates annotation prompt rules to indicate the correspondence between code statement types and code annotation types, so that different annotation type prompts can be given for codes of different statement types; a large language model is pre-created to implement semantic analysis and statement type analysis of the code; thus, for the target program code that needs to be annotated, the large language model can be used to analyze and determine each code segment to be annotated therein, its semantic information and statement type, and then the annotation type corresponding to each code segment to be annotated can be determined in combination with the annotation prompt rules, so that the target annotation text corresponding to it can be generated according to the annotation type and semantic information corresponding to the code segment to be annotated and inserted into the corresponding position to complete automated code annotation. As can be seen, this technical solution does not require manual writing of annotations, is faster and more efficient, can effectively ensure the accuracy and practicality of code annotations, and further improves the readability of the code.
[0134] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a code annotation generation device provided by the present invention. The code annotation generation device may include: An acquisition module 1, configured to acquire the target program code and determine the annotation prompt rules; the annotation prompt rules include the correspondence between code statement types and code annotation types; An analysis module 2, configured to analyze the target program code by using the large language model to determine the code segments to be annotated in the target program code, as well as the semantic information and statement type of the code segments to be annotated; A determination module 3, configured to determine the annotation type corresponding to the code segment to be annotated according to the annotation prompt rules and the statement type of the code segment to be annotated; The annotation module 4 is used to generate the target annotation text corresponding to the code snippet to be annotated according to the annotation type and semantic information corresponding to the code snippet to be annotated, and insert the target annotation text into the annotation position corresponding to the code snippet to be annotated.
[0135] It can be seen that the code annotation generation device provided by the embodiments of the present invention pre-creates annotation hint rules to indicate the correspondence between code statement types and code annotation types, so that different annotation type hints can be given for codes of different statement types; a large language model is pre-created to implement semantic analysis and statement type analysis of the code; thus, for the target program code to be annotated, the large language model can be used to analyze and determine each code snippet to be annotated therein, its semantic information and statement type, and then combine the annotation hint rules to determine the annotation type corresponding to each code snippet to be annotated, so that the target annotation text corresponding to the code snippet to be annotated can be generated according to the annotation type and semantic information corresponding to the code snippet to be annotated and inserted into the corresponding position to complete automated code annotation. It can be seen that this technical solution does not require manual writing of annotations, is faster and more efficient, can effectively ensure the accuracy and practicality of code annotations, and further improves the readability of the code.
[0136] In an embodiment of the present invention, the above analysis module 2 may include: A division unit for dividing the target program code by using a large language model to obtain a plurality of target code snippets; A screening unit for screening and determining the code snippets to be annotated and the statement types of the code snippets to be annotated in each target code snippet according to the code statement types included in the annotation hint rules; An analysis unit for performing semantic analysis on the code snippets to be annotated to obtain the semantic information of the code snippets to be annotated.
[0137] In an embodiment of the present invention, the above division unit may specifically be used to perform syntax structure analysis on the target program code by using a large language model to determine the syntax structure of the target program code; divide the target program code according to the syntax structure of the target program code to obtain a plurality of target code snippets.
[0138] In an embodiment of the present invention, the annotation hint rules may further include the correspondence between code statement types and code annotation positions; Correspondingly, the above annotation module 4 may specifically be used to determine the annotation position corresponding to the code snippet to be annotated according to the annotation hint rules; insert the target annotation text into the annotation position.
[0139] In one embodiment of the present invention, the code statement types include one or more combinations of expression statements, variable declaration statements, return statements, loop statements, and conditional judgment statements; the code comment types include one or more combinations of implementation function description comments, implementation method description comments, and design decision description comments.
[0140] In one embodiment of the present invention, the code comment generation device may further include: An acquisition module, configured to acquire a code data set in a computer program project; An extraction module, configured to extract each code comment pair from the code data set; A first determination module, configured to determine the statement type of the code text and the comment type of the comment text in the code comment pair; A second determination module, configured to determine the association relationship between each statement type and each comment type according to each code comment pair, the code text corresponding to each statement type, and the comment text corresponding to each comment type; A generation module, configured to generate a comment prompt rule according to the association relationship between each statement type and each comment type.
[0141] In one embodiment of the present invention, the above acquisition module may include: A first determination unit, configured to determine all computer program projects in a target open source platform; A selection unit, configured to screen out target computer program projects from all computer program projects according to a preset screening rule; An acquisition unit, configured to acquire a code data set in the target computer program project.
[0142] In one embodiment of the present invention, the above selection unit may specifically be configured to determine the project popularity of each computer program project; select a preset number of computer program projects with the highest project popularity as target computer program projects.
[0143] In one embodiment of the present invention, after the above selection unit selects a preset number of computer program projects with the highest project popularity as target computer program projects, the selection unit is further configured to determine the target comment ratio in each target computer program project; eliminate the target computer program projects with a target comment ratio lower than a preset threshold.
[0144] In one embodiment of the present invention, after the above selection unit selects a preset number of computer program projects with the highest project popularity as target computer program projects, the selection unit is further configured to determine the project type of each target computer program project; eliminate the target computer program projects with a project type of test projects and / or document projects and / or experimental projects.
[0145] In one embodiment of the present invention, the above extraction module may be specifically configured to, for each computer program code in the code dataset, perform syntax analysis on the computer program code to obtain an abstract syntax tree corresponding to the computer program code; determine comment nodes in the computer program code according to the abstract syntax tree, and extract comment text at the comment nodes; determine code text corresponding to the comment text by using a preset matching rule; and generate code-comment pairs by using each comment text and each code text.
[0146] In one embodiment of the present invention, after determining comment nodes in the computer program code according to the abstract syntax tree and extracting comment text at the comment nodes, the above extraction module may also be used to perform a merging process on each comment text according to a preset merging rule; and perform a cleaning process on each comment text according to a preset cleaning rule.
[0147] In one embodiment of the present invention, the above extraction module may also be used to determine the positional relationship between each comment node and the code text corresponding to the comment node; construct a correspondence between the code statement type and the code comment position according to each positional relationship and the statement type of each code text; and add the correspondence between the code statement type and the code comment position to the comment hint rule.
[0148] In one embodiment of the present invention, the above first determination module may include: An extraction unit, configured to extract features of the comment text in the code-comment pair to obtain comment text features; A second determination unit, configured to determine the comment type of the comment text according to the comment text features.
[0149] In one embodiment of the present invention, the above extraction unit may be specifically configured to perform syntax analysis on the comment text in the code-comment pair by using a preset classification model to obtain each word phrase in the comment text; extract features of each word phrase to obtain each word feature; and generate comment text features of the comment text according to each word feature.
[0150] In one embodiment of the present invention, the above generation module may be specifically configured to divide all code-comment pairs according to the statement type and the comment type to obtain a plurality of code-comment pair sets; each code-comment pair set corresponds to the same statement type and the same comment type; evaluate each code-comment pair set by using a preset quality analysis tool to obtain each evaluation result; and update the association relationship between each statement type and each comment type according to each evaluation result to obtain a correspondence between the code statement type and the code comment type, so as to generate a comment hint rule.
[0151] For the introduction of the device provided in the embodiment of the present invention, please refer to the above method embodiment, and the present invention will not be elaborated herein.
[0152] An embodiment of the present invention provides an electronic device.
[0153] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by the present invention. The electronic device may include: A memory 11 for storing a computer program; A processor 10, which can implement the steps of any of the above code annotation generation methods when executing the computer program.
[0154] As Figure 4 shown, it is a schematic diagram of the composition structure of the electronic device. The electronic device may include: a processor 10, a memory 11, a communication interface 12, and a communication bus 13. The processor 10, the memory 11, and the communication interface 12 all complete communication with each other through the communication bus 13.
[0155] In the embodiment of the present invention, the processor 10 may be a central processing unit (CPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic devices, etc.
[0156] The processor 10 may call the program stored in the memory 11. Specifically, the processor 10 may execute the operations in the embodiment of the code annotation generation method.
[0157] The memory 11 is used to store one or more programs. The program may include program code, and the program code includes computer operation instructions. In the embodiment of the present invention, the memory 11 stores at least programs for implementing the following functions: Obtain the target program code and determine the annotation hint rule; the annotation hint rule includes the correspondence between the code statement type and the code annotation type; Use the large language model to analyze the target program code to determine the code snippet to be annotated in the target program code, as well as the semantic information and statement type of the code snippet to be annotated; According to the annotation hint rule and the statement type of the code snippet to be annotated, determine the annotation type corresponding to the code snippet to be annotated; Generate the target annotation text corresponding to the code snippet to be annotated according to the annotation type and semantic information corresponding to the code snippet to be annotated, and insert the target annotation text into the annotation position corresponding to the code snippet to be annotated.
[0158] In a possible implementation, the memory 11 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function, etc.; the data storage area may store the data created during use.
[0159] In addition, the memory 11 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device or other volatile solid-state storage devices.
[0160] The communication interface 12 may be an interface of the communication module for connecting to other devices or systems.
[0161] Of course, it should be noted that Figure 4 the structure shown does not constitute a limitation on the electronic device in the embodiments of the present invention. In practical applications, the electronic device may include more or fewer components than Figure 4 shown, or combine certain components.
[0162] The embodiments of the present invention provide a non-volatile storage medium.
[0163] The computer program stored on the non-volatile storage medium provided by the embodiments of the present invention can implement the steps of any of the above code annotation generation methods when executed by a processor.
[0164] Among them, the non-volatile storage medium may be any available medium that a computer can store or a data storage device such as a server or a data center integrating one or more available media. For example, it may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape, etc.), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc., which can store computer program code.
[0165] For the introduction of the non-volatile storage medium provided by the embodiments of the present invention, please refer to the above method embodiments, and the present invention will not be elaborated here.
[0166] The embodiments of the present invention provide a computer program product.
[0167] The computer program product provided by the embodiments of the present invention includes computer programs / instructions, and the computer programs / instructions can implement the steps of any of the above code annotation generation methods when executed by a processor.
[0168] Specifically, in the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0169] Among them, the computer program product may include one or more computer programs / instructions. When the computer program / instructions are loaded and executed on a computer, they may wholly or partly generate the processes or functions described in the embodiments of the present invention. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a non-volatile storage medium or transmitted from one non-volatile storage medium to another non-volatile storage medium. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line, etc.) or a wireless manner (such as infrared, wireless, microwave, etc.).
[0170] For the introduction of the computer program product provided in the embodiments of the present invention, please refer to the above method embodiments, and the present invention will not be elaborated herein.
[0171] The embodiments in the specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for the relevant parts.
[0172] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0173] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0174] The above has introduced the technical solution provided by the present invention in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A code comment generation method, characterized in that: include: Obtain the target program code and determine the comment prompt rules; The comment prompt rule includes the corresponding relationship between code statement types and code comment types; Analyzing the target program code using a large language model to determine a code segment to be annotated in the target program code and semantic information and statement type of the code segment to be annotated; Determining the comment type corresponding to the code snippet to be commented on according to the comment prompt rule and the statement type of the code snippet to be commented on; A target annotation text corresponding to the code snippet to be annotated is generated according to the annotation type and semantic information corresponding to the code snippet to be annotated, and the target annotation text is inserted into the annotation position corresponding to the code snippet to be annotated.
2. The code comment generation method according to claim 1, characterized in that: Analyzing the target program code using a large language model to determine the code fragment to be annotated in the target program code and the semantic information and statement type of the code fragment to be annotated, including: Using the large language model to divide the target program code to obtain multiple target code fragments; According to the code statement types included in the comment prompt rules, the code snippets to be annotated and the statement types of the code snippets to be annotated are screened and determined in each of the target code snippets; Perform semantic analysis on the code snippet to be annotated to obtain semantic information of the code snippet to be annotated.
3. The code comment generation method according to claim 2, characterized in that: The target program code is divided by using the large language model to obtain multiple target code fragments, including: Performing grammatical structure analysis on the target program code using the large language model to determine the grammatical structure of the target program code; The target program code is divided according to the grammatical structure of the target program code to obtain a plurality of target code segments.
4. The code comment generation method according to claim 1, characterized in that: The comment prompt rules also include the corresponding relationship between the code statement type and the code comment position; Accordingly, inserting the target comment text into the comment position corresponding to the code snippet to be commented on includes: Determine the comment position corresponding to the code snippet to be commented according to the comment prompt rule; The target annotation text is inserted into the annotation position.
5. The code comment generation method according to claim 1, characterized in that: The code statement type includes one or more combinations of expression statements, variable declaration statements, return statements, loop statements, and conditional judgment statements; The code comment type includes one or more combinations of implementation function description type comments, implementation method description type comments, and design decision description type comments.
6. The code comment generation method according to any one of claims 1 to 5, characterized in that: The generation process of the annotation prompt rule includes: Code data sets are collected in computer program projects; Extracting each code annotation pair from the code data set; Determining a statement type of the code text and a comment type of the comment text in the code comment pair; Determine the association relationship between each of the statement types and each of the comment types according to each of the code comment pairs, the code text corresponding to each of the statement types, and the comment text corresponding to each of the comment types; The annotation prompt rule is generated according to the association relationship between each of the statement types and each of the annotation types.
7. The code comment generation method according to claim 6, characterized in that: Code datasets collected in computer program projects include: Identify all computer program projects in the target open source platform; Screening all the computer program projects according to a preset screening rule to obtain a target computer program project; The code data set is acquired from the target computer program project.
8. The code comment generation method according to claim 7, characterized in that: The target computer program project is obtained by screening all the computer program projects according to the preset screening rules, including: determining the project popularity of each of the computer program projects; A preset number of the computer program projects with the highest project popularity are selected as the target computer program projects.
9. The code comment generation method according to claim 8, characterized in that: After selecting a preset number of the computer program projects with the highest project popularity as the target computer program projects, the method further includes: Determining a target annotation ratio in each of the target computer program projects; Target computer program projects whose target annotation ratio is lower than a preset threshold are eliminated.
10. The code comment generation method according to claim 8, characterized in that: After selecting a preset number of the computer program projects with the highest project popularity as the target computer program projects, the method further includes: determining a project type of each of the target computer program projects; Eliminate target computer program projects whose project types are test projects and / or document projects and / or experimental projects.
11. The code comment generation method according to claim 6, characterized in that: The code annotation pairs are extracted from the code dataset, including: For each computer program code in the code data set, performing syntax analysis on the computer program code to obtain an abstract syntax tree corresponding to the computer program code; Determine a comment node in the computer program code according to the abstract syntax tree, and extract the comment text at the comment node; Determine the code text corresponding to the annotation text using a preset matching rule; The code-annotation pair is generated using each of the annotation texts and each of the code texts.
12. The code comment generation method according to claim 11, characterized in that: Determining a comment node in the computer program code according to the abstract syntax tree, and after extracting the comment text at the comment node, further comprising: Merge the annotation texts according to the preset merging rules; Each of the annotation texts is cleaned according to preset cleaning rules.
13. The code comment generation method according to claim 11, characterized in that: Also includes: Determine the positional relationship between each of the annotation nodes and the code text corresponding to the annotation node; Constructing a correspondence between a code statement type and a code comment position according to each of the position relationships and the statement type of each of the code texts; The correspondence between the code statement type and the code comment position is added to the comment prompt rule.
14. The code comment generation method according to claim 6, characterized in that: Determining the comment type of the comment text in the code comment pair includes: Extracting features of the annotation text in the code annotation pair to obtain annotation text features; The annotation type of the annotation text is determined according to the annotation text feature.
15. The code comment generation method according to claim 14, characterized in that: Feature extraction is performed on the comment text in the code comment pair to obtain comment text features, including: Using a preset classification model to perform grammatical analysis on the annotation text in the code annotation pair to obtain each word phrase in the annotation text; Extracting features from each word phrase to obtain features of each word; Generate annotation text features of the annotation text according to each of the word features.
16. The code comment generation method according to claim 6, characterized in that: Generating the annotation prompt rule according to the association relationship between each of the statement types and each of the annotation types includes: Dividing all the code comment pairs according to the statement type and the comment type to obtain a plurality of code comment pair sets; each code comment pair set corresponds to the same statement type and the same comment type; Using a preset quality analysis tool to evaluate each of the code annotation pairs to obtain each evaluation result; The association relationship between each of the statement types and each of the comment types is updated according to each of the evaluation results, and the corresponding relationship between the code statement type and the code comment type is obtained to generate the comment prompt rule.
17. A code comment generation device, characterized in that: include: An acquisition module is used to acquire target program code and determine comment prompt rules; The comment prompt rule includes the corresponding relationship between code statement types and code comment types; An analysis module, used to analyze the target program code using a large language model, and determine the code fragment to be annotated in the target program code and the semantic information and statement type of the code fragment to be annotated; A determination module, used to determine the comment type corresponding to the code snippet to be annotated according to the comment prompt rule and the statement type of the code snippet to be annotated; The annotation module is used to generate a target annotation text corresponding to the code snippet to be annotated according to the annotation type and semantic information corresponding to the code snippet to be annotated, and insert the target annotation text into the annotation position corresponding to the code snippet to be annotated.
18. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the code annotation generating method as claimed in any one of claims 1 to 16 when executing the computer program.
19. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the code annotation generating method according to any one of claims 1 to 16 are implemented.
20. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the code annotation generation method according to any one of claims 1 to 16 are implemented.
Citation Information
Patent Citations
Program annotation method and device
CN103324513A
Source code annotation automatic generation method
CN110399162A
Annotation positioning method based on program analysis and neural network
CN111104159A
Deep learning-based Java program internal annotation generation method and system
CN113076133A
Code annotation generation method and related device
CN118605936A
Cited By
Code comparison method, system and equipment for open source component and medium
CN120994231A
A method, system, device, and medium for open source component oriented code comparison
CN120994231B