A code comment generation method and device based on a sequence generation adversarial network

By introducing improved algorithms based on sequence generation adversarial networks in code annotation generation, integrating structural features and semantic features, and adding reinforcement learning, the problem of limited code annotation generation in the existing technology is solved, high-quality annotation generation is achieved, ensuring the integrity and consistency of annotations, and thus improving the efficiency and quality of software development.

CN115756475BActive Publication Date: 2025-06-27ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211333572.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2025-06-27
Estimated Expiration
2042-10-28

AI Technical Summary

Technical Problem

The prior art has limited effect in the automatic generation of code comments, and it is difficult to ensure the integrity of comments and consistency with code entities, which leads to more time and energy required for other developers to read code and may cause incorrect code calls or modifications.

Method used

An improved algorithm based on sequence generation adversarial network is adopted to improve the effect of code annotation generation by fusing structural features and semantic features and adding reinforcement learning. Specific methods include building a generative network and a discriminative network, performing pre-training and adversarial training, and optimizing the performance of the generative network using a policy gradient method.

Benefits of technology

It significantly improves the quality and accuracy of code annotation generation, reduces the time and energy of manual maintenance, ensures the integrity of annotations and consistency with code entities, and thus accelerates the software development process and improves the development quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115756475B_ABST
    Figure CN115756475B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and device for generating code comments based on a sequence generative adversarial network. By introducing reinforcement learning, the effect of code comment generation is further improved. Complete code comments can help developers quickly understand the corresponding code content and greatly improve development efficiency. In order to reduce the workload of developers outside development while ensuring the completeness and quality of code comments, the task of code comment generation is proposed. Since the generation effect of current common code comment generation algorithms is limited, in order to further improve the effect, after constructing a generation network based on a graph neural network and a discriminant network based on a convolutional neural network, the present invention introduces a sequence generative adversarial network, and completes the transfer of the loss function gradient between the generation network and the discriminant network through the policy gradient method to establish an adversarial relationship. The present invention goes through three training stages in total and finally obtains a generation network with a better generation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to code comment generation and discrimination, and particularly to a code comment generation method based on an encoder-decoder framework and the fusion of Seq2Seq and Graph2Seq, a code comment discrimination method based on a convolutional neural network, and an improved code comment generation method based on a sequence generative adversarial network. Background Art

[0002] The software development process has become increasingly large and complex and can no longer be completed by a single person at one time. More team cooperation, post-maintenance, and even subsequent extended development are required. Therefore, there is an urgent need for a way to record and facilitate communication about all aspects of the software, that is, software development documents.

[0003] Software development documents cover all aspects of the software development process and are essential materials and guiding guidelines throughout the software development, use, and maintenance processes. Software development documents with standardized content can greatly improve the efficiency of software development and ensure the quality of the software. Software documents are generated and used throughout all stages of software engineering. Generally speaking, the documents generated during the development process will include requirement analysis documents generated during the requirement analysis stage, general design documents, system design documents, and detailed design documents generated during the design stage, software test documents generated during the test stage, and summary report documents such as QA documents and user manuals that may be required after development is completed.

[0004] Software technical documents, as the name implies, focus more on the technical aspects during the software development process, that is, requirement analysis documents, general design documents, system design documents, detailed design documents, software test documents, and other documents related to the development implementation technology, rather than other documents such as deployment documents for system administrators and user manuals for users.

[0005] High-quality software documents can effectively sort out and connect all stages of the entire software engineering, effectively assist the process of the entire software engineering, and help developers and users better grasp the software. High-quality software technical documents can clarify the details of the development design, help developers efficiently and accurately understand important information such as user requirements, the overall and local design of the software, the software change history, and the existing code, enable developers to quickly understand the software, promote the simplification of communication between developers, and thus accelerate the software development process and improve the software development quality.

[0006] Code comments are the most closely related documents to code entities in software technical documents and are the most intuitive documents to assist other developers in quickly understanding program code. Code comments often record the developer's code implementation ideas, implementation methods, usage methods, etc. Complete code comments can help developers quickly master the core content and core intentions of code entities, facilitate developers to quickly call or modify, and are an indispensable part of software technical documents.

[0007] However, the maintenance of code comments also requires developers to spend extra time and effort. Therefore, it often happens that developers forget or neglect to maintain code comments after completing the development of code entities, resulting in the absence or obsolescence of code comments. The absence of code comments will make it lack guidance for other developers to read the code entity and require more time and effort to understand the code; the obsolescence of code comments will lead to incorrect guidance for other developers, the mismatch between the understanding of the code entity and the actual content of the code entity, and then lead to incorrect calls or even incorrect modifications of the code entity, thus entering software defects and even causing more serious problems. Therefore, it is very important to maintain the integrity of code comments and the consistency with the content of code entities.

[0008] In order to reduce manual effort, save labor, and better ensure the integrity of code comments and the consistency with code entities, the research topic of intelligent code comment generation has emerged. As the name implies, automatic code comment generation is to generate corresponding code comments intelligently by a machine based on the code entities implemented by developers. Summary of the Invention

[0009] The purpose of the present invention is to provide an improved algorithm based on a sequence generation adversarial network for the automatic generation of code comments in view of the deficiencies of the prior art, so as to generate better code comments than a single neural sequence network. Current research relying on machine learning can already perform a certain degree of intelligent code comment generation, but there is still much room for improvement in the generation effect. The present invention proposes a method of further improving the generation effect by integrating structural features and semantic features and adding a sequence adversarial network, that is, introducing reinforcement learning to further strengthen the generation ability of a single generation network.

[0010] The present invention is implemented through the following technical solutions:

[0011] According to the first aspect of this specification, there is provided a method for generating code comments based on a sequence generation adversarial network, the method comprising the following steps:

[0012] S1, obtaining code-comment pair data, constructing a data set, dividing it into a training set and a test set according to a certain ratio, the data volume of the training set should be much larger than that of the test set, and constructing a corpus according to the data set;

[0013] S2. Build a generation network to complete the code annotation generation task and perform pre-training for subsequent adversarial training. Specifically, code has its own specific structural features. For example, loops, branches, etc. all have obvious structural features. The control flow graph corresponding to the code is obtained through a static analyzer as the input. A control flow graph encoder is constructed based on a graph neural network to learn the structural information. A code encoder is constructed based on the Transformer model, using the code and annotation text as the input to learn semantic information. The two types of feature information are fused, and a decoder is constructed based on the Transformer model to form a generation network, which obtains the generated annotation sequence from the feature vector. Adding the Graph2Seq model to the traditional Seq2Seq model helps to understand the code and better generate code annotations. In addition, for code with a long sequence, according to the 80 / 20 principle, only 20% of the most important learning features are retained to save computing resources and speed up the training process.

[0014] S3. Establish a discriminant network for discriminating the generated annotations based on a convolutional neural network model and perform pre-training. Specifically, a CNNs model can be used to establish a discriminant network that can complete a binary classification task, which can discriminate whether the input annotation is the true annotation or the generated annotation corresponding to the code, and is used to compete with the generation network in adversarial training.

[0015] S4. Incorporate reinforcement learning, establish an adversarial network between the generation network and the discriminant network based on the sequence generation adversarial network model, and perform adversarial training. Specifically, the policy gradient method is used to complete the process of calculating the loss and gradient of the generation network according to the output result of the discriminant network, establishing an adversarial relationship between the generation network and the discriminant network, enabling the two to further strengthen learning through confrontation, so that the generation network has a better generation effect and the discriminant network has a more accurate judgment result.

[0016] S5. Use BLEU-4 as the evaluation metric to evaluate the generation effect of the generation network, evaluate on the test set, and during the adversarial training process, adjust the parameters of the generation network, discriminant network, and adversarial network according to the evaluation results, control the respective strengths of the generation network and the discriminant network, and find the balance point of the model to achieve the best adversarial effect. In addition, due to the phenomenon that the learning speed of the discriminant network is much faster than that of the generation network found during the evaluation process, it is considered that a discriminant network with a simple structure, poor representation ability, not reaching the best training level, and a discrimination accuracy of about 0.75 should be selected, and a generation network with a complex structure, strong representation ability, and the best training level should be selected to maintain the initial adversarial balance.

[0017] S6. Input the code to be annotated into the trained generation network to obtain the code annotation result.

[0018] Further, in S1, the java data part of the dataset CodeXGLUE, which is specifically cleaned for the code-to-text generation task in the publicly available dataset CodeSearchNet, is adopted. This dataset itself has been divided into a training set containing 164,923 data and a test set containing 10,955 data, and targeted data cleaning is performed according to the characteristics of the code annotation generation task. A corpus of code and annotations is generated based on the cleaned dataset. In addition, <unk> <pad> <bos> <eos>respectively represent the unknown word identifier, the completion identifier, the sequence start identifier, and the sequence end identifier; the cleaning process is specifically as follows: split the camel case nouns and snake case nouns in the code and comments; clean the code samples containing comment content; clean the comment samples containing meaningless repeated special symbols such as *, #, or -; make unified processing for strings and numbers, respectively using <str>And <num>To uniformly identify.

[0019] Furthermore, during the pre-training process of S2, the conversion of code and real annotations from words to tokens is respectively completed, and <bos>and <eos>The corresponding tokens are used to mark the start and end of the sequence; the generation network mainly consists of an encoder and a decoder. To better retain the rich structural information of the code, two types of encoders are used in the encoder part. One is a graph neural network, such as a graph attention network, whose corresponding input is the code control flow graph analyzed by a static analyzer. The other is a Transformer model, whose corresponding input is the token sequence corresponding to the code and the true annotation. After separately learning the structural feature vector and the semantic feature vector, the two are fused to obtain the final feature vector. The decoder also uses a Transformer model, and decodes the learned feature vector to obtain the generated annotation token sequence. The cross-entropy loss function is used to calculate the loss between the generated annotation sequence and the true annotation sequence, and the gradient is passed to optimize the network. After multiple verifications, the generation network obtained by pre-training with AdamW as the optimizer is more conducive to subsequent adversarial training. In addition, it is found that a considerable part of the code data has a sequence length exceeding 200 words. For sequences that are too long, the accuracy will decrease during learning while the resource consumption increases. To improve this problem, according to the 80 / 20 principle, only the most important 20% of the feature vectors of the long sequences are retained, and the 80% with smaller weights are ignored, so as to maximize the retention of features while minimizing resource consumption and accelerating the training speed.

[0020] Further, in the pre-training process of S3, the code is combined and mixed with the true annotation and the generated annotation respectively, and used as the input of the discriminant network. After being processed by the embedding layer, the respective word vectors are obtained. The convolutional layer of the discriminant network uses multiple convolutional kernels of different sizes to extract the features of the code and the two types of annotations in multiple dimensions. Then, the code features are respectively connected with the true annotation features and the generated annotation features as the overall features of the input data, and sent to the ReLU activation function. After activation, a pooling operation is performed, and finally the result of the discriminant network for binary classification is obtained through the connection layer, and a Softmax process is performed on the result. The final output gives the probabilities of the two classes respectively. If the probability of being true is higher, it means that the discriminant network believes that this annotation is the true annotation corresponding to the code. If the probability of being false is higher, it means that the discriminant network believes that this annotation is a false annotation generated by the machine. Finally, the loss of this round is calculated by comparing with the data label, and the gradient is further passed to train the discriminant network.

[0021] Further, the specific step S4 is as follows: In the design of the sequence generation adversarial network, policy gradient means regarding the process of the generation network selecting the next word token of the sequence as a selection strategy. By adjusting this strategy, the generation network can select more tokens that fit the truth. Therefore, a score needs to be given to the result of each selection, and this score comes from the discriminant network. The calculation of the policy gradient J(θ) is expressed as:

[0022]

[0023] In the formula, Y 1:T represents the generated complete annotation sequence, G θ represents the generation network with hyperparameters θ, X represents the input code sequence, CFG represents the input control flow graph, represents the discriminator network with hyperparameters and Y 1:T-1 represents that there are T - 1 tokens in the generated annotation sequence, y T represents the T-th token selected by the generation network in this case, represents the reward given by the discriminator network for the behavior of the generation network selecting y T for this token;

[0024] For the incomplete annotation sequence Y 1:t (t < T) during the generation process, Monte Carlo search is adopted to quickly select words and generate the ungenerated part of the annotation sequence. At the same time, since the sequences generated by Monte Carlo quickly have randomness, multiple sequences will be generated during actual use. N represents the number of sequences, and MC represents the Monte Carlo search process. The calculation process is shown in the following formula:

[0025]

[0026] is the reward score given by the discriminator network to the generation network. Its essence is the probability that the discriminator network believes that the current annotation sequence is the true annotation sequence corresponding to the code sequence; when t < T, this probability is the output obtained by using multiple complete annotation sequences obtained through Monte Carlo search as the input of the discriminator network. Use to represent the output of the discriminator network corresponding to the n-th sequence, and take the average of multiple sequences; when t = T, directly use the complete annotation sequence generated by the generation network and the source code sequence as the input of the discriminator network to obtain the reward score. The specific calculation process is:

[0027]

[0028] Based on the above process, an adversarial relationship can be established between the generation network and the discriminator network. The generation network selects tokens to generate an annotation sequence according to the input code. The discriminator network accepts the code and the generated annotation as inputs and judges whether the annotation sequence is true, and uses the probability of being true as the reward score for the current selection of the generation network. The reward score is passed back to the generation network as the loss guiding gradient descent. After the generation network completes gradient update, it generates a new annotation sequence again and passes it to the discriminator network for judgment;

[0029] The update of the discriminative network is carried out separately after the generative network has been updated for a certain number of rounds. The principle is the same as that of pre-training, except that the generated annotations are generated by the generative network obtained from the latest training.

[0030] Furthermore, the specific steps of step S5 are as follows: During the initial adversarial training, due to the imbalance in the strength of the two adversarial parties, the discrimination task of the discriminative network is much simpler than the generation task of the generative network. After pre-training, it can reach a higher strength, resulting in the generative network being unable to compete with it during adversarial training and thus quickly collapsing. To solve this problem, according to the change in the BLEU-4 score of the annotations generated by the generative network, the generative network and the discriminative network are readjusted and then pre-trained. The optimizer of the generative network is changed to AdamW, and the hyperparameters learning rate and number of training rounds are optimized. The number of training rounds of the discriminative network is reduced so that it does not reach the highest discrimination accuracy but is maintained at a medium strength of about 0.75 to achieve the balance between the two.

[0031] After adjusting the pre-training stage of the generative network and the discriminative network, the relevant training parameters in the adversarial training are adjusted according to the change in the BLEU-4 score. It is found from the change in the BLEU-4 score in the experiment that when the discriminative network is trained more, the generation effect of the generative network improves faster and can quickly reach the best generation effect. However, due to the too strong effect of the discriminative network, it will quickly enter the collapse state after reaching the best, and the generation effect gradually decreases to 0, resulting in poor adversarial training effect. By setting the step size of the generative network to be greater than the step size step of the discriminative network, when the generative network is trained more, although the learning speed decreases because the generative network is more likely to pass the judgment of the discriminative network, it can be stably trained to obtain a stable best generation effect. This guides the setting of the step size step and the number of epochs parameters in the adversarial network.

[0032] Furthermore, the specific steps of step S6 are as follows: The code is first processed by a static analyzer to obtain the corresponding control flow graph, and the control flow graph is used as the input of the control flow graph encoder of the generative network; in the sequence of generated annotations, initially there is only the <bos>The corresponding token. After that, it is the token sequence of the selected words. The code sequence and the current generated annotation sequence are used as the input of the code encoder of the generation network. After being processed by their respective encoders to obtain the structural feature vector and the semantic feature vector, they are fused. The decoder of the trained generation network selects the token of the next word that the annotation should have word by word according to the fused feature vector. The most basic method here is the greedy method, that is, to select the word with the highest probability each time. The beam search method with better effect can also be selected, but the search cost is greater and the time consumption is more. Until the end-of-sequence identifier appears <eos>, generate the generated annotation; convert the words corresponding to the corpus tokens into a complete sequence of English annotations.

[0033] According to the second aspect of the present specification, there is provided a code annotation generation device based on a sequence generative adversarial network, including a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it is used to implement the code annotation generation method based on the sequence generative adversarial network as described in the first aspect.

[0034] The beneficial effect of the present invention is that although there are various intelligent code annotation generation methods currently, their generation effects are limited, and there is currently no attempt to use a sequence adversarial generation network for intelligent code annotation generation. The present invention has made a pioneering attempt in this regard and has confirmed that the generation effect can be further improved on the basis of the existing intelligent code annotation generation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is the framework diagram of the adversarial training algorithm of the present invention;

[0036] Figure 2 is the source code example diagram of the present invention;

[0037] Figure 3 is of the present invention Figure 2 source code conversion control flow diagram example;

[0038] Figure 4 is the process diagram of the policy gradient method of the present invention;

[0039] Figure 5 is the structure diagram of the code annotation generation device based on the sequence generative adversarial network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] The present invention uses the dataset CodeXGLUE provided by the CodeSearchNet team, which is dedicated to code-text generation tasks, and further cleans it for use in code annotation generation tasks. The present invention proposes, for the code annotation generation task, to stack a generative adversarial network on the basis of an existing neural machine translation model, and through the confrontation between the generative network and the discriminative network, guide the learning of the neural machine translation model to obtain better generation effects.

[0042] The present invention includes a total of three networks and three training stages. The algorithm framework for adversarial training is as Figure 1 shown.

[0043] The first network is a generation network, which adopts an encoder-decoder framework and uses two different encoders and a decoder to represent the input code control flow graph as CFG = (V, E, V T , R), where V represents the nodes of the graph, E represents the edges of the graph, V T represents the node type, and R represents the relationship type of the edges. The code sequence is represented as X = (X1, …, x m ), and the annotation sequence as the output target is represented as Y = (y1, …, y n ). We attempt to learn a generation model such that G(CFG, X) = Y.

[0044] The second network is a discriminative network, which adopts the mainstream Convolutional Neural Networks (CNNs), includes an embedding layer and multiple convolutional layers, jointly completes the feature learning of the input data, and finally completes the binary classification discrimination task, that is, whether the input data is a true or false annotation sequence. For the discriminative network, its input data includes both the code sequence and the annotation sequence. Let X represent the code sequence, Y represent the annotation sequence, D represent the discriminative network, and True / False represent whether the annotation sequence is true or false. This learning process can be expressed as D(X, Y) = True / False.

[0045] The last network is a generative adversarial network. Based on the idea of sequence adversarial networks, the present invention realizes the confrontation between the generation network and the discriminative network by adding policy gradient training.

[0046] The three training stages are the pre-training stage of the generation network, the pre-training stage of the discriminative network, and the adversarial training stage of both.

[0047] The specific working process of the present invention is as follows:

[0048] 1. Further cleaning of the CodeXGLUE data.

[0049] 2. Establish a generation network for code annotation generation based on the encoder-decoder framework and perform pre-training.

[0050] 3. Establish a discriminative network for discriminating the generated annotations based on the CNNs convolutional neural network and perform pre-training.

[0051] 4. Add reinforcement learning, establish an adversarial network between the generation network and the discriminative network based on the policy gradient method, and perform adversarial training.

[0052] 5. Evaluate the code comment generation effect of the trained network and adjust the adversarial network according to the results.

[0053] 6. Input the code into the generation network to obtain the generated comments.

[0054] The beneficial effect of the present invention is that although there are various intelligent code comment generation methods currently, their generation effects are limited, and there is no attempt to use the sequence adversarial generation network for intelligent code comment generation. The present invention makes a pioneering attempt in this regard and proves that the generation effect can be further improved on the basis of the existing intelligent code comment generation methods.

[0055] The following further introduces the present invention from six aspects:

[0056] 1. Further cleaning of the CodeXGLUE data.

[0057] The present invention adopts a partial dataset in the CodeSearchNet dataset collected and processed by Husain et al. with the programming language Java for the "code to text" task. For this task, Husain et al. specifically cleaned a dataset named CodeXGLUE. Based on the CodeSearchNet, the following four types of examples were removed from this dataset:

[0058] (1) Code examples that cannot be parsed syntactically;

[0059] (2) Overlong or overly short comment examples that are longer than 256 or shorter than 3 words;

[0060] (3) Comment examples containing special symbols such as <img…> or http: etc.;

[0061] (4) Comment examples in languages other than English;

[0062] However, it was found during the experiment that the CodeXGLUE dataset was still not thoroughly cleaned. Although some abnormal examples that were not conducive to comment generation were removed, the camel-case words, meaningless repeated symbols, numerical strings, etc. contained in the examples would still have a negative impact on the creation of the corpus and the learning of sequence features. Therefore, on the basis of the CodeXGLUE dataset, the present invention further processed the data, mainly including the following aspects:

[0063] (1) Split the camel-case and snake-case words in the code and comments, reduce OOV words, shrink the corpus size, and improve the corpus quality;

[0064] (2) Clean the examples in the code that repeatedly contain comment content. Code that fails to completely separate the comments will affect the learning of code features.

[0065] (3) Clean the samples with comments containing meaningless repeated special symbols such as *, # or -, which are meaningless delimiters without valid information and are interference items in the code comment content;

[0066] (4) For strings and numbers (including scientific notation), different strings and different numbers represent the meanings of strings and numbers respectively in the code. During learning, only the general categories of strings and numbers need to be processed, without caring about their specific contents. Therefore, strings and numbers are uniformly processed, and <str>And <num>to uniformly identify;

[0067] After completing the above data processing, the present invention creates corpora for code and comments respectively based on the cleaned dataset. The size of the corpus for code is 31040, and the size of the corpus for comments is 21239. The words in the corpus are sorted in descending order of usage frequency from high to low, and the first four are set as special symbols <unk> , <pad> , <bos>and <eos>, respectively representing unknown word markers, completion markers, sequence start markers, and sequence end markers, to deal with the appearance of words that are not in the corpus, or to complete the lengths of different sequences and mark the start and end of sequence generation for simultaneous processing as a batch.

[0068] 2. Establish a generative network for code comment generation based on the encoder-decoder framework and perform pre-training.

[0069] There are two common sequence models used for Seq2Seq: the RNNSearch model and the Transformer model. The present invention uses the Transformer model, which is considered to be more effective, as the basic model for semantic feature learning of the generative network. In addition, the generative network model can also use the BERT model, which has been proven to be effective. The basic model for structural feature learning can use various graph neural network models, such as R-GAT or graph Transformer.

[0070] The present invention first obtains the control flow graph of the code based on the static analysis tool before generating the network processing, such as Figure 2 The Java code example shown can be processed by the static analyzer to obtain Figure 3 The code control flow graph shown in the figure is then transformed from words to tokens based on the corpus, and the head and tail are added <bos>And <eos>The corresponding tokens are used to mark the start and end of the sequence. After the control flow graph is processed by the control flow graph encoder based on the graph neural network, a feature vector regarding the structural information is obtained. After the token sequence is processed by the code encoder based on the Transformer model, a feature vector regarding the semantic information is obtained. The two feature vectors are fused and then input into the decoder based on the Transformer model to obtain the predicted generated annotation sequence. The cross-entropy loss function is used to calculate the loss between the generated annotation sequence and the real annotation sequence, and the gradient is passed to optimize the network. After multiple verifications, the generation network obtained by pre-training with AdamW as the optimizer is more conducive to subsequent adversarial training.

[0071] In addition, it is found that a considerable part of the code data has a situation where the sequence length exceeds 200 words. For sequences that are too long, the accuracy will decrease during learning while the resource consumption increases. To improve this problem, according to the 80 / 20 principle, only the most important 20% of the feature vectors of the long sequences are retained, and the 80% with smaller weights is ignored, so as to maximize the retention of features while minimizing resource consumption and accelerating the training speed.

[0072] Settings for the pre-training parameters of the generation network:

[0073] After statistics, most of the code sequence lengths are within 200, and most of the annotation sequence lengths are within 30, as shown in Table 1 and Table 2. Therefore, when reading the data, the maximum lengths of the code sequence and the annotation sequence are limited to 200 and 30 respectively, so as to shorten the sequence lengths to be processed as much as possible and accelerate the training efficiency. Finally, there are 138,230 pieces of data in the train set used for training, and 4,246 pieces of data in the valid set used to verify the training effect. During training, for code sequences or annotation sequences with lengths less than 200 or 30, no special processing is done. For those exceeding 200 or 30, according to the 80 / 20 principle, that is, in any group of things, the most important only accounts for a small part, about 20%, and the remaining about 80% although accounting for the majority is secondary. Therefore, the learning of their feature vectors is processed according to the 80 / 20 rule, that is, the 20% with the largest weights and the most important is retained, and the 80% with smaller weights that can be ignored is discarded to improve the learning efficiency.

[0074] Table 1 Statistics of code sequence lengths

[0075] Sequence length range Number of samples within the range Percentage of the total dataset <100 94191 0.572 <120 107902 0.655 <140 117991 0.717 <160 125795 0.764 <180 131863 0.801 <200 136719 0.831 ≥200 164611 1.0

[0076] Table 2 Statistics of annotation sequence lengths

[0077]

[0078]

[0079] When the generation network is pre-trained, the cross-entropy loss function is used as the loss function, the AdamW optimizer is adopted, the learning rate is set to 0.0001, and the batch size is set to 64.

[0080] 3. Establish a discriminant network for discriminant annotation generation based on the CNNs convolutional neural network and perform pre-training.

[0081] The discriminant network (Discriminator, D) mainly completes the task of judging whether the input data is real or fake. This task can be regarded as a binary classification, that is, classifying the object into the real class or the fake class.

[0082] Since what needs to be judged is whether the current annotation is the real annotation corresponding to the code, the input of the discriminant network requires both the code and the annotation. First, the input code and annotation need to be transformed. After obtaining the token sequence, the word vectors are obtained through the embedding layer, and then they are sent to the convolutional layer to extract features respectively. In the present invention, multiple convolutional kernels of different sizes are adopted to extract the features of the code and the annotation in multiple dimensions. After obtaining the features of the code and the annotation respectively, they are connected as the overall features of the input data and sent to the ReLU activation function. After activation, a pooling operation is performed, and finally the result of the binary classification of the discriminant network is obtained through the connection layer. The result is processed by Softmax once, and the probabilities of the two classes are finally output. If the probability of being true is higher, it means that the discriminant network believes that the annotation is the real annotation corresponding to the code. If the probability of being false is higher, it means that the discriminant network believes that the annotation is a fake annotation generated by the machine.

[0083] Pre-training parameter settings for the discriminant network:

[0084] The discriminant network is set with convolutional kernels of sizes 1, 2, 4, 6, 8, 9, 10, 15, and 20, with 100, 200, 200, 100, 100, 100, 100, 160, and 160 layers respectively, forming a multi-convolutional layer to fully learn the features of the input sequence data. In the binary classification label, the class with subscript 0 represents the probability that the discriminant network believes that the annotation sequence is false, and the class with subscript 1 represents the probability that the discriminant network believes that the annotation sequence is true.

[0085] The cross-entropy loss function is adopted as the loss function of the discriminant network, the Adam optimizer is adopted, the learning rate is set to 0.0001, and the batch size is set to 128. Since the dataset is large, the training of the discriminant network improves rapidly. And for the generative adversarial network, the generation network and the discriminant network need to have similar strengths. Therefore, the discriminant network is only pre-trained for 2 rounds, obtaining the minimum loss of 0.450. At this time, the accuracy Accuracy of the discriminant network is calculated to be 0.782, ensuring that the generation network has a high probability of passing the judgment of the discriminant network.

[0086] 4. Incorporate reinforcement learning, establish an adversarial network between the generator network and the discriminator network based on the policy gradient method, and conduct adversarial training.

[0087] The adversarial nature of the adversarial network means that the generator network aims to generate better, that is, the difference between the generated examples and the real data is minimized, while the discriminator network aims to distinguish more accurately, that is, the difference between the generated examples and the real data is maximized, thus forming an adversarial relationship.

[0088] In the design of the sequence generation adversarial network, the policy gradient means regarding the process of the generator network selecting the next word token of the sequence as a selection strategy, and by adjusting this strategy, the generator network can select tokens that are more in line with the real ones. Therefore, a score needs to be given for the result of each selection, and this score comes from the discriminator network. The process is as Figure 4 shown.

[0089] Use G θ to represent the generator network under the current parameter θ, represent the discriminator network under the current parameter , CFG represents the input control flow graph, X represents the input code sequence, Y represents the generated annotation sequence, T represents the maximum length of the annotation sequence, and the calculation of the policy gradient J(θ) can be expressed as:

[0090]

[0091] In the formula, Y 1:T represents the complete generated annotation sequence, 1:T-1 represents that there are already T - 1 tokens in the generated annotation sequence, y T represents the Tth token selected by the generator network in this case, represents the reward given by the discriminator network for the behavior of the generator network selecting y T this token. The complete formula represents the process of the generator network generating the complete annotation sequence Y 1:T when inputting the code sequence X. During the process, after there are already T - 1 tokens, the behavior of selecting the Tth token as y T , and the reward score given by the discriminator network. This calculation occurs repeatedly during each token selection process of the generated annotation sequence.

[0092] The annotation sequence used in the formula calculation is Y 1:T , that is to say, a complete annotation sequence is required to be input into the discriminator network for evaluation. However, the algorithm hopes to evaluate the selection process of each word during the generation process of the annotation sequence. Therefore, for the incomplete annotation sequence Y during the generation process 1:t (t < T), Monte Carlo search (MC) is adopted to quickly select words and generate the annotation sequence of the ungenerated part. At the same time, since the sequences quickly generated by Monte Carlo are random, multiple sequences will be generated during actual use. Let N represent the number of sequences, and MC represent the Monte Carlo search process. The calculation process is shown in the following formula:

[0093]

[0094] is the reward score given by the discriminative network to the generative network. Its essence is the probability that the discriminative network believes that the current annotation sequence is the true annotation sequence corresponding to the code sequence. When t < T, this probability is the output obtained by using the complete annotation sequence obtained through Monte Carlo search as the input of the discriminative network. Let represent the output of the discriminative network corresponding to the nth sequence, and the average value is taken for multiple sequences. When t = T, the complete annotation sequence and the source code sequence generated by the generative network are directly used as the input of the discriminative network to obtain the reward score. The specific calculation process is as follows:

[0095]

[0096] Based on the above three formulas, an adversarial relationship can be established between the generative network and the discriminative network. The generative network selects tokens to generate the annotation sequence. The discriminative network accepts the generated annotation sequence as the input and judges the probability that this annotation sequence is true as the reward for the current selection of the generative network. The reward is backpropagated to the generative network as the loss guiding gradient descent. After the generative network completes the gradient update, it generates a new annotation sequence again and passes it to the discriminative network for judgment. The update of the discriminative network is carried out separately after the generative network has been updated for a certain number of rounds.

[0097] 5. Evaluate the code annotation generation effect of the trained network, and adjust the adversarial network according to the results.

[0098] For the evaluation metrics of sequence generation effect, BLEU is commonly used. According to different N selected in N-grams, it is commonly divided into BLEU-1, BLEU-2, BLEU-3, and BLEU-4. Since the code annotations generated are all relatively long sentences (more than 5 words), BLEU-4 is adopted as the evaluation metric in the present invention.

[0099] During adversarial training, the generative network and the discriminative network still adopt the network parameters set during pre-training. Before training, the best network model parameters obtained after pre-training of the generative network and the discriminative network are loaded separately.

[0100] However, during the initial adversarial training, it was found that the BLEU-4 score quickly collapsed after a brief increase and dropped all the way to 0. After analysis, the reason was that the intensities of the two adversarial parties were unbalanced. The discrimination task of the discriminative network was much simpler than the generation task of the generative network. Therefore, after pre-training, it could reach a higher intensity. During the adversarial process, the generative network could not compete with it, resulting in the inability to find the direction of reinforcement learning from the adversarial process, and thus quickly collapsed. To solve this problem, according to the change of the BLEU-4 score of the annotations generated by the generative network, the generative network and the discriminative network were readjusted and then pre-trained. The optimizer of the generative network was changed to AdamW, and the hyperparameters learning rate and number of training epochs were optimized. The number of training epochs of the discriminative network was reduced so that it did not reach the highest discrimination accuracy, but was maintained at a medium intensity of about 0.75 to achieve the balance between the two.

[0101] After adjusting the pre-training stage of the generative network and the discriminative network, the relevant training parameters in the adversarial training were adjusted according to the change of the BLEU-4 score. The most important parameters in the adversarial training were the step and epoch settings of the generative network training and the step and epoch settings of the discriminative network training. After certain experiments, it was found that when the step of the discriminative network was greater than the step of the generative network, that is, when the discriminative network was trained more, the generation effect of the generative network improved faster and could quickly reach the best generation effect. However, it would also quickly enter a collapse state after reaching the best due to the too strong effect of the discriminative network, and the generation effect gradually decreased to 0. Therefore, this adversarial training effect was poor. When the step of the generative network was set to be greater than the step of the discriminative network and the generative network was trained more, since the generative network was more likely to pass the judgment of the discriminative network, the learning speed decreased and the improvement was slower. To better achieve the balance between the generative network and the discriminative network and steadily improve the effects of both the generative network and the discriminative network, after certain experiments, it was found that better training effects could be obtained when the step of the generative network was set to 16, the epoch was set to 5, the step of the discriminative network was set to 400, and the epoch was set to 1.

[0102] 6. Input the code into the generative network to obtain the generated annotations.

[0103] The code is first processed by a static analyzer to obtain the corresponding control flow graph, and the control flow graph is used as the input of the control flow graph encoder of the generative network; in the sequence of generated annotations, initially there is only the <bos>The corresponding token. After that, it is the token sequence of the selected words. Using the code sequence and the current generated annotation sequence as the input of the code encoder of the generation network, after being processed by their respective encoders to obtain the structural feature vector and the semantic feature vector and then fusing them; the decoder of the trained generation network selects the token of the next word that the annotation should have word by word according to the fused feature vector. The most basic method here is the greedy method, that is, selecting the word with the highest probability each time. The beam search method with better effect can also be selected, but the search cost is greater and the time consumption is more; until the end-of-sequence identifier appears <eos>, generate generated annotations; convert the words corresponding to the corpus tokens into a complete sequence of English annotations.

[0104] Corresponding to the embodiments of the foregoing code annotation generation method based on a sequence generative adversarial network, the present invention also provides embodiments of a code annotation generation device based on a sequence generative adversarial network.

[0105] See Figure 5 , a code annotation generation device based on a sequence generative adversarial network provided by an embodiment of the present invention includes a memory and one or more processors, and executable code is stored in the memory. When the processor executes the executable code, it is used to implement the code annotation generation method based on a sequence generative adversarial network in the foregoing embodiments.

[0106] Embodiments of the code annotation generation device based on a sequence generative adversarial network of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by a processor of any device with data processing capabilities where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From a hardware level, as Figure 5 shown, it is a hardware structure diagram of any device with data processing capabilities where the code annotation generation device based on a sequence generative adversarial network of the present invention is located. In addition to Figure 5 the shown processor, memory, network interface, and non-volatile memory, any device with data processing capabilities where the device in the embodiment is located usually also includes other hardware according to the actual functions of the any device with data processing capabilities, which will not be elaborated here.

[0107] The implementation processes of the functions and roles of each unit in the above device are specifically detailed in the implementation processes of the corresponding steps in the above method, which will not be elaborated here.

[0108] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial descriptions of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0109] An embodiment of the present invention further provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the method for generating code comments based on a sequence generation adversarial network in the above embodiment is implemented.

[0110] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.

[0111] The above are only the preferred embodiments of one or more embodiments of this specification, and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.< / eos> < / bos> < / eos> < / bos> < / eos> < / bos> < / pad> < / unk> < / num> < / str> < / eos> < / bos> < / eos> < / bos> < / num> < / str> < / eos> < / bos> < / pad> < / unk> is additionally added to the corpus <unk> <pad> <bos> <eos>respectively represent the unknown word identifier, the completion identifier, the sequence start identifier, and the sequence end identifier; the cleaning process is specifically as follows: split the camel case nouns and snake case nouns in the code and comments; clean the code samples containing comment content; clean the comment samples containing meaningless repeated special symbols such as *, #, or -; make unified processing for strings and numbers, respectively using <str>And <num>To uniformly identify.

[0019] Furthermore, during the pre-training process of S2, the conversion of code and real annotations from words to tokens is respectively completed, and <bos>and <eos>The corresponding tokens are used to mark the start and end of the sequence; the generation network mainly consists of an encoder and a decoder. To better retain the rich structural information of the code, two types of encoders are used in the encoder part. One is a graph neural network, such as a graph attention network, whose corresponding input is the code control flow graph analyzed by a static analyzer. The other is a Transformer model, whose corresponding input is the token sequence corresponding to the code and the true annotation. After separately learning the structural feature vector and the semantic feature vector, the two are fused to obtain the final feature vector. The decoder also uses a Transformer model, and decodes the learned feature vector to obtain the generated annotation token sequence. The cross-entropy loss function is used to calculate the loss between the generated annotation sequence and the true annotation sequence, and the gradient is passed to optimize the network. After multiple verifications, the generation network obtained by pre-training with AdamW as the optimizer is more conducive to subsequent adversarial training. In addition, it is found that a considerable part of the code data has a sequence length exceeding 200 words. For sequences that are too long, the accuracy will decrease during learning while the resource consumption increases. To improve this problem, according to the 80 / 20 principle, only the most important 20% of the feature vectors of the long sequences are retained, and the 80% with smaller weights are ignored, so as to maximize the retention of features while minimizing resource consumption and accelerating the training speed.

[0020] Further, in the pre-training process of S3, the code is combined and mixed with the true annotation and the generated annotation respectively, and used as the input of the discriminant network. After being processed by the embedding layer, the respective word vectors are obtained. The convolutional layer of the discriminant network uses multiple convolutional kernels of different sizes to extract the features of the code and the two types of annotations in multiple dimensions. Then, the code features are respectively connected with the true annotation features and the generated annotation features as the overall features of the input data, and sent to the ReLU activation function. After activation, a pooling operation is performed, and finally the result of the discriminant network for binary classification is obtained through the connection layer, and a Softmax process is performed on the result. The final output gives the probabilities of the two classes respectively. If the probability of being true is higher, it means that the discriminant network believes that this annotation is the true annotation corresponding to the code. If the probability of being false is higher, it means that the discriminant network believes that this annotation is a false annotation generated by the machine. Finally, the loss of this round is calculated by comparing with the data label, and the gradient is further passed to train the discriminant network.

[0021] Further, the specific step S4 is as follows: In the design of the sequence generation adversarial network, policy gradient means regarding the process of the generation network selecting the next word token of the sequence as a selection strategy. By adjusting this strategy, the generation network can select more tokens that fit the truth. Therefore, a score needs to be given to the result of each selection, and this score comes from the discriminant network. The calculation of the policy gradient J(θ) is expressed as:

[0022]

[0023] In the formula, Y 1:T represents the generated complete annotation sequence, G θ represents the generation network with hyperparameters θ, X represents the input code sequence, CFG represents the input control flow graph, represents the discriminator network with hyperparameters and Y 1:T-1 represents that there are T - 1 tokens in the generated annotation sequence, y T represents the T-th token selected by the generation network in this case, represents the reward given by the discriminator network for the behavior of the generation network selecting y T for this token;

[0024] For the incomplete annotation sequence Y 1:t (t < T) during the generation process, Monte Carlo search is adopted to quickly select words and generate the ungenerated part of the annotation sequence. At the same time, since the sequences generated by Monte Carlo quickly have randomness, multiple sequences will be generated during actual use. N represents the number of sequences, and MC represents the Monte Carlo search process. The calculation process is shown in the following formula:

[0025]

[0026] is the reward score given by the discriminator network to the generation network. Its essence is the probability that the discriminator network believes that the current annotation sequence is the true annotation sequence corresponding to the code sequence; when t < T, this probability is the output obtained by using multiple complete annotation sequences obtained through Monte Carlo search as the input of the discriminator network. Use to represent the output of the discriminator network corresponding to the n-th sequence, and take the average of multiple sequences; when t = T, directly use the complete annotation sequence generated by the generation network and the source code sequence as the input of the discriminator network to obtain the reward score. The specific calculation process is:

[0027]

[0028] Based on the above process, an adversarial relationship can be established between the generation network and the discriminator network. The generation network selects tokens to generate an annotation sequence according to the input code. The discriminator network accepts the code and the generated annotation as inputs and judges whether the annotation sequence is true, and uses the probability of being true as the reward score for the current selection of the generation network. The reward score is passed back to the generation network as the loss guiding gradient descent. After the generation network completes gradient update, it generates a new annotation sequence again and passes it to the discriminator network for judgment;

[0029] The update of the discriminative network is carried out separately after the generative network has been updated for a certain number of rounds. The principle is the same as that of pre-training, except that the generated annotations are generated by the generative network obtained from the latest training.

[0030] Furthermore, the specific steps of step S5 are as follows: During the initial adversarial training, due to the imbalance in the strength of the two adversarial parties, the discrimination task of the discriminative network is much simpler than the generation task of the generative network. After pre-training, it can reach a higher strength, resulting in the generative network being unable to compete with it during adversarial training and thus quickly collapsing. To solve this problem, according to the change in the BLEU-4 score of the annotations generated by the generative network, the generative network and the discriminative network are readjusted and then pre-trained. The optimizer of the generative network is changed to AdamW, and the hyperparameters learning rate and number of training rounds are optimized. The number of training rounds of the discriminative network is reduced so that it does not reach the highest discrimination accuracy but is maintained at a medium strength of about 0.75 to achieve the balance between the two.

[0031] After adjusting the pre-training stage of the generative network and the discriminative network, the relevant training parameters in the adversarial training are adjusted according to the change in the BLEU-4 score. It is found from the change in the BLEU-4 score in the experiment that when the discriminative network is trained more, the generation effect of the generative network improves faster and can quickly reach the best generation effect. However, due to the too strong effect of the discriminative network, it will quickly enter the collapse state after reaching the best, and the generation effect gradually decreases to 0, resulting in poor adversarial training effect. By setting the step size of the generative network to be greater than the step size step of the discriminative network, when the generative network is trained more, although the learning speed decreases because the generative network is more likely to pass the judgment of the discriminative network, it can be stably trained to obtain a stable best generation effect. This guides the setting of the step size step and the number of epochs parameters in the adversarial network.

[0032] Furthermore, the specific steps of step S6 are as follows: The code is first processed by a static analyzer to obtain the corresponding control flow graph, and the control flow graph is used as the input of the control flow graph encoder of the generative network; in the sequence of generated annotations, initially there is only the <bos>The corresponding token. After that, it is the token sequence of the selected words. The code sequence and the current generated annotation sequence are used as the input of the code encoder of the generation network. After being processed by their respective encoders to obtain the structural feature vector and the semantic feature vector, they are fused. The decoder of the trained generation network selects the token of the next word that the annotation should have word by word according to the fused feature vector. The most basic method here is the greedy method, that is, to select the word with the highest probability each time. The beam search method with better effect can also be selected, but the search cost is greater and the time consumption is more. Until the end-of-sequence identifier appears <eos>, generate the generated annotation; convert the words corresponding to the corpus tokens into a complete sequence of English annotations.

[0033] According to the second aspect of the present specification, there is provided a code annotation generation device based on a sequence generative adversarial network, including a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it is used to implement the code annotation generation method based on the sequence generative adversarial network as described in the first aspect.

[0034] The beneficial effect of the present invention is that although there are various intelligent code annotation generation methods currently, their generation effects are limited, and there is currently no attempt to use a sequence adversarial generation network for intelligent code annotation generation. The present invention has made a pioneering attempt in this regard and has confirmed that the generation effect can be further improved on the basis of the existing intelligent code annotation generation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is the framework diagram of the adversarial training algorithm of the present invention;

[0036] Figure 2 is the source code example diagram of the present invention;

[0037] Figure 3 is of the present invention Figure 2 source code conversion control flow diagram example;

[0038] Figure 4 is the process diagram of the policy gradient method of the present invention;

[0039] Figure 5 is the structure diagram of the code annotation generation device based on the sequence generative adversarial network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] The present invention uses the dataset CodeXGLUE provided by the CodeSearchNet team, which is dedicated to code-text generation tasks, and further cleans it for use in code annotation generation tasks. The present invention proposes, for the code annotation generation task, to stack a generative adversarial network on the basis of an existing neural machine translation model, and through the confrontation between the generative network and the discriminative network, guide the learning of the neural machine translation model to obtain better generation effects.

[0042] The present invention includes a total of three networks and three training stages. The algorithm framework for adversarial training is as Figure 1 shown.

[0043] The first network is a generation network, which adopts an encoder-decoder framework and uses two different encoders and a decoder to represent the input code control flow graph as CFG = (V, E, V T , R), where V represents the nodes of the graph, E represents the edges of the graph, V T represents the node type, and R represents the relationship type of the edges. The code sequence is represented as X = (X1, …, x m ), and the annotation sequence as the output target is represented as Y = (y1, …, y n ). We attempt to learn a generation model such that G(CFG, X) = Y.

[0044] The second network is a discriminative network, which adopts the mainstream Convolutional Neural Networks (CNNs), includes an embedding layer and multiple convolutional layers, jointly completes the feature learning of the input data, and finally completes the binary classification discrimination task, that is, whether the input data is a true or false annotation sequence. For the discriminative network, its input data includes both the code sequence and the annotation sequence. Let X represent the code sequence, Y represent the annotation sequence, D represent the discriminative network, and True / False represent whether the annotation sequence is true or false. This learning process can be expressed as D(X, Y) = True / False.

[0045] The last network is a generative adversarial network. Based on the idea of sequence adversarial networks, the present invention realizes the confrontation between the generation network and the discriminative network by adding policy gradient training.

[0046] The three training stages are the pre-training stage of the generation network, the pre-training stage of the discriminative network, and the adversarial training stage of both.

[0047] The specific working process of the present invention is as follows:

[0048] 1. Further cleaning of the CodeXGLUE data.

[0049] 2. Establish a generation network for code annotation generation based on the encoder-decoder framework and perform pre-training.

[0050] 3. Establish a discriminative network for discriminating the generated annotations based on the CNNs convolutional neural network and perform pre-training.

[0051] 4. Add reinforcement learning, establish an adversarial network between the generation network and the discriminative network based on the policy gradient method, and perform adversarial training.

[0052] 5. Evaluate the code comment generation effect of the trained network and adjust the adversarial network according to the results.

[0053] 6. Input the code into the generation network to obtain the generated comments.

[0054] The beneficial effect of the present invention is that although there are various intelligent code comment generation methods currently, their generation effects are limited, and there is no attempt to use the sequence adversarial generation network for intelligent code comment generation. The present invention makes a pioneering attempt in this regard and proves that the generation effect can be further improved on the basis of the existing intelligent code comment generation methods.

[0055] The following further introduces the present invention from six aspects:

[0056] 1. Further cleaning of the CodeXGLUE data.

[0057] The present invention adopts a partial dataset in the CodeSearchNet dataset collected and processed by Husain et al. with the programming language Java for the "code to text" task. For this task, Husain et al. specifically cleaned a dataset named CodeXGLUE. Based on the CodeSearchNet, the following four types of examples were removed from this dataset:

[0058] (1) Code examples that cannot be parsed syntactically;

[0059] (2) Overlong or overly short comment examples that are longer than 256 or shorter than 3 words;

[0060] (3) Comment examples containing special symbols such as <img…> or http: etc.;

[0061] (4) Comment examples in languages other than English;

[0062] However, it was found during the experiment that the CodeXGLUE dataset was still not thoroughly cleaned. Although some abnormal examples that were not conducive to comment generation were removed, the camel-case words, meaningless repeated symbols, numerical strings, etc. contained in the examples would still have a negative impact on the creation of the corpus and the learning of sequence features. Therefore, on the basis of the CodeXGLUE dataset, the present invention further processed the data, mainly including the following aspects:

[0063] (1) Split the camel-case and snake-case words in the code and comments, reduce OOV words, shrink the corpus size, and improve the corpus quality;

[0064] (2) Clean the examples in the code that repeatedly contain comment content. Code that fails to completely separate the comments will affect the learning of code features.

[0065] (3) Clean the samples with comments containing meaningless repeated special symbols such as *, # or -, which are meaningless delimiters without valid information and are interference items in the code comment content;

[0066] (4) For strings and numbers (including scientific notation), different strings and different numbers represent the meanings of strings and numbers respectively in the code. During learning, only the general categories of strings and numbers need to be processed, without caring about their specific contents. Therefore, strings and numbers are uniformly processed, and <str>And <num>to uniformly identify;

[0067] After completing the above data processing, the present invention creates corpora for code and comments respectively based on the cleaned dataset. The size of the corpus for code is 31040, and the size of the corpus for comments is 21239. The words in the corpus are sorted in descending order of usage frequency from high to low, and the first four are set as special symbols <unk> , <pad> , <bos>and <eos>, respectively representing unknown word markers, completion markers, sequence start markers, and sequence end markers, to deal with the appearance of words that are not in the corpus, or to complete the lengths of different sequences and mark the start and end of sequence generation for simultaneous processing as a batch.

[0068] 2. Establish a generative network for code comment generation based on the encoder-decoder framework and perform pre-training.

[0069] There are two common sequence models used for Seq2Seq: the RNNSearch model and the Transformer model. The present invention uses the Transformer model, which is considered to be more effective, as the basic model for semantic feature learning of the generative network. In addition, the generative network model can also use the BERT model, which has been proven to be effective. The basic model for structural feature learning can use various graph neural network models, such as R-GAT or graph Transformer.

[0070] The present invention first obtains the control flow graph of the code based on the static analysis tool before generating the network processing, such as Figure 2 The Java code example shown can be processed by the static analyzer to obtain Figure 3 The code control flow graph shown in the figure is then transformed from words to tokens based on the corpus, and the head and tail are added <bos>And <eos>The corresponding tokens are used to mark the start and end of the sequence. After the control flow graph is processed by the control flow graph encoder based on the graph neural network, a feature vector regarding the structural information is obtained. After the token sequence is processed by the code encoder based on the Transformer model, a feature vector regarding the semantic information is obtained. The two feature vectors are fused and then input into the decoder based on the Transformer model to obtain the predicted generated annotation sequence. The cross-entropy loss function is used to calculate the loss between the generated annotation sequence and the real annotation sequence, and the gradient is passed to optimize the network. After multiple verifications, the generation network obtained by pre-training with AdamW as the optimizer is more conducive to subsequent adversarial training.

[0071] In addition, it is found that a considerable part of the code data has a situation where the sequence length exceeds 200 words. For sequences that are too long, the accuracy will decrease during learning while the resource consumption increases. To improve this problem, according to the 80 / 20 principle, only the most important 20% of the feature vectors of the long sequences are retained, and the 80% with smaller weights is ignored, so as to maximize the retention of features while minimizing resource consumption and accelerating the training speed.

[0072] Settings for the pre-training parameters of the generation network:

[0073] After statistics, most of the code sequence lengths are within 200, and most of the annotation sequence lengths are within 30, as shown in Table 1 and Table 2. Therefore, when reading the data, the maximum lengths of the code sequence and the annotation sequence are limited to 200 and 30 respectively, so as to shorten the sequence lengths to be processed as much as possible and accelerate the training efficiency. Finally, there are 138,230 pieces of data in the train set used for training, and 4,246 pieces of data in the valid set used to verify the training effect. During training, for code sequences or annotation sequences with lengths less than 200 or 30, no special processing is done. For those exceeding 200 or 30, according to the 80 / 20 principle, that is, in any group of things, the most important only accounts for a small part, about 20%, and the remaining about 80% although accounting for the majority is secondary. Therefore, the learning of their feature vectors is processed according to the 80 / 20 rule, that is, the 20% with the largest weights and the most important is retained, and the 80% with smaller weights that can be ignored is discarded to improve the learning efficiency.

[0074] Table 1 Statistics of code sequence lengths

[0075] Sequence length range Number of samples within the range Percentage of the total dataset <100 94191 0.572 <120 107902 0.655 <140 117991 0.717 <160 125795 0.764 <180 131863 0.801 <200 136719 0.831 ≥200 164611 1.0

[0076] Table 2 Statistics of annotation sequence lengths

[0077]

[0078]

[0079] When the generation network is pre-trained, the cross-entropy loss function is used as the loss function, the AdamW optimizer is adopted, the learning rate is set to 0.0001, and the batch size is set to 64.

[0080] 3. Establish a discriminant network for discriminant annotation generation based on the CNNs convolutional neural network and perform pre-training.

[0081] The discriminant network (Discriminator, D) mainly completes the task of judging whether the input data is real or fake. This task can be regarded as a binary classification, that is, classifying the object into the real class or the fake class.

[0082] Since what needs to be judged is whether the current annotation is the real annotation corresponding to the code, the input of the discriminant network requires both the code and the annotation. First, the input code and annotation need to be transformed. After obtaining the token sequence, the word vectors are obtained through the embedding layer, and then they are sent to the convolutional layer to extract features respectively. In the present invention, multiple convolutional kernels of different sizes are adopted to extract the features of the code and the annotation in multiple dimensions. After obtaining the features of the code and the annotation respectively, they are connected as the overall features of the input data and sent to the ReLU activation function. After activation, a pooling operation is performed, and finally the result of the binary classification of the discriminant network is obtained through the connection layer. The result is processed by Softmax once, and the probabilities of the two classes are finally output. If the probability of being true is higher, it means that the discriminant network believes that the annotation is the real annotation corresponding to the code. If the probability of being false is higher, it means that the discriminant network believes that the annotation is a fake annotation generated by the machine.

[0083] Pre-training parameter settings for the discriminant network:

[0084] The discriminant network is set with convolutional kernels of sizes 1, 2, 4, 6, 8, 9, 10, 15, and 20, with 100, 200, 200, 100, 100, 100, 100, 160, and 160 layers respectively, forming a multi-convolutional layer to fully learn the features of the input sequence data. In the binary classification label, the class with subscript 0 represents the probability that the discriminant network believes that the annotation sequence is false, and the class with subscript 1 represents the probability that the discriminant network believes that the annotation sequence is true.

[0085] The cross-entropy loss function is adopted as the loss function of the discriminant network, the Adam optimizer is adopted, the learning rate is set to 0.0001, and the batch size is set to 128. Since the dataset is large, the training of the discriminant network improves rapidly. And for the generative adversarial network, the generation network and the discriminant network need to have similar strengths. Therefore, the discriminant network is only pre-trained for 2 rounds, obtaining the minimum loss of 0.450. At this time, the accuracy Accuracy of the discriminant network is calculated to be 0.782, ensuring that the generation network has a high probability of passing the judgment of the discriminant network.

[0086] 4. Incorporate reinforcement learning, establish an adversarial network between the generator network and the discriminator network based on the policy gradient method, and conduct adversarial training.

[0087] The adversarial nature of the adversarial network means that the generator network aims to generate better, that is, the difference between the generated examples and the real data is minimized, while the discriminator network aims to distinguish more accurately, that is, the difference between the generated examples and the real data is maximized, thus forming an adversarial relationship.

[0088] In the design of the sequence generation adversarial network, the policy gradient means regarding the process of the generator network selecting the next word token of the sequence as a selection strategy, and by adjusting this strategy, the generator network can select tokens that are more in line with the real ones. Therefore, a score needs to be given for the result of each selection, and this score comes from the discriminator network. The process is as Figure 4 shown.

[0089] Use G θ to represent the generator network under the current parameter θ, represent the discriminator network under the current parameter , CFG represents the input control flow graph, X represents the input code sequence, Y represents the generated annotation sequence, T represents the maximum length of the annotation sequence, and the calculation of the policy gradient J(θ) can be expressed as:

[0090]

[0091] In the formula, Y 1:T represents the complete generated annotation sequence, 1:T-1 represents that there are already T - 1 tokens in the generated annotation sequence, y T represents the Tth token selected by the generator network in this case, represents the reward given by the discriminator network for the behavior of the generator network selecting y T this token. The complete formula represents the process of the generator network generating the complete annotation sequence Y 1:T when inputting the code sequence X. During the process, after there are already T - 1 tokens, the behavior of selecting the Tth token as y T , and the reward score given by the discriminator network. This calculation occurs repeatedly during each token selection process of the generated annotation sequence.

[0092] The annotation sequence used in the formula calculation is Y 1:T , that is to say, a complete annotation sequence is required to be input into the discriminator network for evaluation. However, the algorithm hopes to evaluate the selection process of each word during the generation process of the annotation sequence. Therefore, for the incomplete annotation sequence Y during the generation process 1:t (t < T), Monte Carlo search (MC) is adopted to quickly select words and generate the annotation sequence of the ungenerated part. At the same time, since the sequences quickly generated by Monte Carlo are random, multiple sequences will be generated during actual use. Let N represent the number of sequences, and MC represent the Monte Carlo search process. The calculation process is shown in the following formula:

[0093]

[0094] is the reward score given by the discriminative network to the generative network. Its essence is the probability that the discriminative network believes that the current annotation sequence is the true annotation sequence corresponding to the code sequence. When t < T, this probability is the output obtained by using the complete annotation sequence obtained through Monte Carlo search as the input of the discriminative network. Let represent the output of the discriminative network corresponding to the nth sequence, and the average value is taken for multiple sequences. When t = T, the complete annotation sequence and the source code sequence generated by the generative network are directly used as the input of the discriminative network to obtain the reward score. The specific calculation process is as follows:

[0095]

[0096] Based on the above three formulas, an adversarial relationship can be established between the generative network and the discriminative network. The generative network selects tokens to generate the annotation sequence. The discriminative network accepts the generated annotation sequence as the input and judges the probability that this annotation sequence is true as the reward for the current selection of the generative network. The reward is backpropagated to the generative network as the loss guiding gradient descent. After the generative network completes the gradient update, it generates a new annotation sequence again and passes it to the discriminative network for judgment. The update of the discriminative network is carried out separately after the generative network has been updated for a certain number of rounds.

[0097] 5. Evaluate the code annotation generation effect of the trained network, and adjust the adversarial network according to the results.

[0098] For the evaluation metrics of sequence generation effect, BLEU is commonly used. According to different N selected in N-grams, it is commonly divided into BLEU-1, BLEU-2, BLEU-3, and BLEU-4. Since the code annotations generated are all relatively long sentences (more than 5 words), BLEU-4 is adopted as the evaluation metric in the present invention.

[0099] During adversarial training, the generative network and the discriminative network still adopt the network parameters set during pre-training. Before training, the best network model parameters obtained after pre-training of the generative network and the discriminative network are loaded separately.

[0100] However, during the initial adversarial training, it was found that the BLEU-4 score quickly collapsed after a brief increase and dropped all the way to 0. After analysis, the reason was that the intensities of the two adversarial parties were unbalanced. The discrimination task of the discriminative network was much simpler than the generation task of the generative network. Therefore, after pre-training, it could reach a higher intensity. During the adversarial process, the generative network could not compete with it, resulting in the inability to find the direction of reinforcement learning from the adversarial process, and thus quickly collapsed. To solve this problem, according to the change of the BLEU-4 score of the annotations generated by the generative network, the generative network and the discriminative network were readjusted and then pre-trained. The optimizer of the generative network was changed to AdamW, and the hyperparameters learning rate and number of training epochs were optimized. The number of training epochs of the discriminative network was reduced so that it did not reach the highest discrimination accuracy, but was maintained at a medium intensity of about 0.75 to achieve the balance between the two.

[0101] After adjusting the pre-training stage of the generative network and the discriminative network, the relevant training parameters in the adversarial training were adjusted according to the change of the BLEU-4 score. The most important parameters in the adversarial training were the step and epoch settings of the generative network training and the step and epoch settings of the discriminative network training. After certain experiments, it was found that when the step of the discriminative network was greater than the step of the generative network, that is, when the discriminative network was trained more, the generation effect of the generative network improved faster and could quickly reach the best generation effect. However, it would also quickly enter a collapse state after reaching the best due to the too strong effect of the discriminative network, and the generation effect gradually decreased to 0. Therefore, this adversarial training effect was poor. When the step of the generative network was set to be greater than the step of the discriminative network and the generative network was trained more, since the generative network was more likely to pass the judgment of the discriminative network, the learning speed decreased and the improvement was slower. To better achieve the balance between the generative network and the discriminative network and steadily improve the effects of both the generative network and the discriminative network, after certain experiments, it was found that better training effects could be obtained when the step of the generative network was set to 16, the epoch was set to 5, the step of the discriminative network was set to 400, and the epoch was set to 1.

[0102] 6. Input the code into the generative network to obtain the generated annotations.

[0103] The code is first processed by a static analyzer to obtain the corresponding control flow graph, and the control flow graph is used as the input of the control flow graph encoder of the generative network; in the sequence of generated annotations, initially there is only the <bos>The corresponding token. After that, it is the token sequence of the selected words. Using the code sequence and the current generated annotation sequence as the input of the code encoder of the generation network, after being processed by their respective encoders to obtain the structural feature vector and the semantic feature vector and then fusing them; the decoder of the trained generation network selects the token of the next word that the annotation should have word by word according to the fused feature vector. The most basic method here is the greedy method, that is, selecting the word with the highest probability each time. The beam search method with better effect can also be selected, but the search cost is greater and the time consumption is more; until the end-of-sequence identifier appears <eos>, generate generated annotations; convert the words corresponding to the corpus tokens into a complete sequence of English annotations.

[0104] Corresponding to the embodiments of the foregoing code annotation generation method based on a sequence generative adversarial network, the present invention also provides embodiments of a code annotation generation device based on a sequence generative adversarial network.

[0105] See Figure 5 , a code annotation generation device based on a sequence generative adversarial network provided by an embodiment of the present invention includes a memory and one or more processors, and executable code is stored in the memory. When the processor executes the executable code, it is used to implement the code annotation generation method based on a sequence generative adversarial network in the foregoing embodiments.

[0106] Embodiments of the code annotation generation device based on a sequence generative adversarial network of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by a processor of any device with data processing capabilities where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From a hardware level, as Figure 5 shown, it is a hardware structure diagram of any device with data processing capabilities where the code annotation generation device based on a sequence generative adversarial network of the present invention is located. In addition to Figure 5 the shown processor, memory, network interface, and non-volatile memory, any device with data processing capabilities where the device in the embodiment is located usually also includes other hardware according to the actual functions of the any device with data processing capabilities, which will not be elaborated here.

[0107] The implementation processes of the functions and roles of each unit in the above device are specifically detailed in the implementation processes of the corresponding steps in the above method, which will not be elaborated here.

[0108] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial descriptions of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0109] An embodiment of the present invention further provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the method for generating code comments based on a sequence generation adversarial network in the above embodiment is implemented.

[0110] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.

[0111] The above are only the preferred embodiments of one or more embodiments of this specification, and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.< / eos> < / bos> < / eos> < / bos> < / eos> < / bos> < / pad> < / unk> < / num> < / str> < / eos> < / bos> < / eos> < / bos> < / num> < / str> < / eos> < / bos> < / pad> < / unk>

Claims

1. A code comment generation method based on a sequence generative adversarial network, characterized in that, The method comprises the following steps: S1. Obtain code-comment pair data, construct a data set, divide the data set into a training set and a test set, and construct a corpus according to the data set; S2. Establish a generation network based on the encoder-decoder framework, complete the code comment generation task, and perform pre-training for subsequent adversarial training. Specifically, obtain the control flow graph corresponding to the code through a static analyzer as the input, construct a control flow graph encoder based on a graph neural network to learn structural information, construct a code encoder based on the Transformer model, use the code and the annotation text as the input to learn semantic information, fuse the two feature information, construct a decoder based on the Transformer model to form a generation network, and obtain the generated annotation sequence from the feature vector. For codes with a sequence length exceeding the threshold, according to the 80 / 20 principle, only retain the most important 20% of its learning features to save computing resources and accelerate the training speed; S3. Establish a discriminant network for discriminating the generated annotations based on a convolutional neural network model, and perform pre-training for use in adversarial training against the generation network; S4. Incorporate reinforcement learning, establish an adversarial network between the generation network and the discriminant network based on the sequence generation adversarial network model, and perform adversarial training. Specifically, use the policy gradient method to complete the process of calculating the loss and gradient of the generation network according to the output result of the discriminant network, establish an adversarial relationship between the generation network and the discriminant network, enable the two to further strengthen learning through confrontation, so that the generation network has a better generation effect and the discriminant network has a more accurate judgment result; S5. Use BLEU-4 as an evaluation metric to evaluate the generation effect of the generation network, evaluate on the test set, and during the adversarial training process, adjust the parameters of the generation network, the discriminant network, and the adversarial network according to the evaluation results, control the respective strengths of the generation network and the discriminant network, and find the balance point of the model to achieve the best adversarial effect; S6. Input the code to be annotated into the trained generation network to obtain the code annotation result.

2. The method for generating code comments based on a sequence generation adversarial network according to claim 1, characterized in that, In S1, the CodeXGLUE dataset, which is specifically cleaned for the code-to-text generation task in the publicly available CodeSearchNet dataset, is used. Targeted data cleaning is performed according to the characteristics of the code comment generation task, and a corpus of code and comments is generated based on the cleaned dataset. In addition, <unk> <pad> <bos> <eos>respectively represent the unknown word identifier, the completion identifier, the sequence start identifier, and the sequence end identifier. < / eos> < / bos> < / pad> < / unk> is additionally added to the corpus <unk> <pad> <bos> <eos>respectively represent the unknown word identifier, the completion identifier, the sequence start identifier, and the sequence end identifier. < / eos> < / bos> < / pad> < / unk> 3. The method for generating code comments based on a sequence generative adversarial network according to claim 2, wherein In S1, the cleaning process is specifically as follows: split the camel case nouns and snake case nouns in the code and comments; clean the code samples containing comment content; clean the samples containing meaningless repeated special symbols; make unified processing for strings and numbers, respectively using <str>and <num>for unified identification. < / num> < / str> 4. The method for generating code comments based on a sequence generation adversarial network according to claim 1, wherein During the pre-training process, S2 respectively completes the conversion of code and real comments from words to tokens, and <bos>and <eos>The corresponding tokens are used to identify the start and end of the sequence; the generation network mainly consists of two parts: an encoder and a decoder. The encoder part uses two encoders. One is a graph neural network, and the corresponding input is the code control flow graph analyzed by a static analyzer. The other is a Transformer model, and the corresponding input is the token sequence corresponding to the code and the true annotation. After learning the structural feature vector and the semantic feature vector respectively, the two are fused to obtain the final feature vector. The decoder also uses the Transformer model, and decodes through the learned feature vector to obtain the generated annotation token sequence. Use the cross-entropy loss function to calculate the loss between the generated annotation sequence and the true annotation sequence, and perform gradient transmission to optimize the network. The optimizer uses AdamW.< / eos> < / bos> is added at both the beginning and the end. <bos>and <eos>The corresponding tokens are used to identify the start and end of the sequence; the generation network mainly consists of two parts: an encoder and a decoder. The encoder part uses two encoders. One is a graph neural network, and the corresponding input is the code control flow graph analyzed by a static analyzer. The other is a Transformer model, and the corresponding input is the token sequence corresponding to the code and the true annotation. After learning the structural feature vector and the semantic feature vector respectively, the two are fused to obtain the final feature vector. The decoder also uses the Transformer model, and decodes through the learned feature vector to obtain the generated annotation token sequence. Use the cross-entropy loss function to calculate the loss between the generated annotation sequence and the true annotation sequence, and perform gradient transmission to optimize the network. The optimizer uses AdamW.< / eos> < / bos> 5. The method for generating code comments based on a sequence generation adversarial network according to claim 1, wherein During the pre-training process of S3, the code is combined and mixed with the true annotation and the generated annotation respectively, and then used as the input of the discriminant network. After being processed by the embedding layer, the respective word vectors are obtained. The convolutional layer of the discriminant network uses multiple convolutional kernels of different sizes to extract the features of the code and the two annotations in multiple dimensions. Then, the code features are connected with the true annotation features and the generated annotation features respectively as the overall features of the input data, and are sent to the ReLU activation function. After activation, a pooling operation is performed, and finally, the result of the binary classification of the discriminant network is obtained through the connection layer, and a Softmax operation is performed on the result. Finally, the probabilities of the two classes are output. If the probability of being true is higher, it means that the discriminant network believes that this annotation is the true annotation corresponding to the code. If the probability of being false is higher, it means that the discriminant network believes that this annotation is a false annotation generated by the machine. Finally, the loss of this round is calculated by comparing with the data label, and the gradient is further propagated to train the discriminant network.

6. The method for generating code comments based on a sequence generation adversarial network according to claim 1, wherein Specifically, in S4: In the design of the sequence generation adversarial network, policy gradient means regarding the process of the generation network selecting the next word token of the sequence as a selection strategy. By adjusting this strategy, the generation network can select more tokens that fit the truth. Therefore, a score needs to be given to the result of each selection, and this score comes from the discriminant network. The calculation of the policy gradient J(θ) is expressed as: where Y 1:T represents the generated complete annotation sequence, G θ represents the generation network with hyperparameters θ, X represents the input code sequence, CFG represents the input control flow graph, represents the discriminative network with hyperparameters , Y 1:T-1 represents that there are T - 1 tokens in the generated annotation sequence, y T represents the T-th token selected by the generation network in this case, represents the reward given by the discriminative network for the behavior of the generation network selecting the token y T ; For the incomplete annotation sequence Y during the generation process 1:t (t < T), Monte Carlo search is adopted to quickly select words for the ungenerated part of the annotation sequence. At the same time, since the sequences quickly generated by Monte Carlo are random, multiple sequences will be generated during actual use. Let N represent the number of sequences, and MC represent the Monte Carlo search process. The calculation process is shown in the following formula: is the reward score given by the discriminative network to the generative network. Its essence is the probability that the discriminative network believes that the current annotation sequence is the true annotation sequence corresponding to the code sequence. When t < T, this probability is the output obtained by using multiple complete annotation sequences obtained through Monte Carlo search as the input of the discriminative network. It is denoted by as the output of the discriminative network corresponding to the nth sequence, and the average value is taken for multiple sequences. When t = T, the complete annotation sequence and the source code sequence generated by the generative network are directly used as the input of the discriminative network to obtain the reward score. The specific calculation process is as follows: Based on the above process, an adversarial relationship is established between the generation network and the discriminant network. The generation network selects tokens according to the input code to generate an annotation sequence. The discriminant network accepts the code and the generated annotation as input and judges whether this annotation sequence is true, and uses the probability of being true as the reward score for the current selection of the generation network. The reward score is backpropagated to the generation network as the loss guiding the gradient descent. After the generation network completes the gradient update, it generates a new annotation sequence again and passes it to the discriminant network for judgment. The update of the discriminant network is carried out separately after the generation network has been updated for a certain number of rounds. Its principle is the same as that of pre-training, except that the generated annotation is generated by the latest trained generation network.

7. The method for generating code comments based on a sequence generation adversarial network according to claim 1, characterized in that In S5, due to the phenomenon that the learning speed of the discriminant network is much faster than that of the generation network found during the evaluation process, a discriminant network with a simple structure, poor representation ability, not reaching the best training level, and a discriminant accuracy of about 0.75, and a generation network with a complex structure, strong representation ability, and the best training level are selected to maintain the initial adversarial balance.

8. The method for generating code comments based on a sequence generative adversarial network according to claim 1, wherein Specifically, in S5: During the initial adversarial training, due to the imbalance in the strength of the two adversarial parties, the discrimination task of the discriminant network is much simpler than the generation task of the generation network. After pre-training, it can reach a higher strength, resulting in the generation network being unable to compete with it during the confrontation and thus quickly collapsing. To solve this problem, according to the change in the BLEU-4 score result of the annotation generated by the generation network, the generation network and the discriminant network are readjusted and then pre-trained. The optimizer of the generation network is changed to AdamW, and the hyperparameters learning rate and training rounds are optimized. The training rounds of the discriminant network are reduced, and it is not allowed to reach the highest discriminant accuracy, but maintained at a medium strength of about 0.75 to achieve the balance between the two. After adjusting the pre-training phases of the generation network and the discriminative network, relevant training parameters in the adversarial training are adjusted according to the change in the BLEU-4 score; it is found from the change in the BLEU-4 score in the experiment that when the discriminative network is trained more, the generation effect of the generation network improves faster and can quickly reach the best generation effect. However, due to the too strong effect of the discriminative network, it will quickly enter a collapsed state after reaching the best, and the generation effect gradually decreases to 0, resulting in poor adversarial training effect; while setting the step size of the generation network to be greater than the step size step of the discriminative network, when the generation network is trained more, since the generation network is more likely to pass the judgment of the discriminative network, the learning speed decreases, but it can be stably trained to obtain a stable best generation effect; thus guiding the setting of the step size step and the number of epochs epoch parameters in the adversarial network.

9. The method for generating code comments based on a sequence generative adversarial network according to claim 1, wherein The specific content of step S6 is as follows: The code is first processed by a static analyzer to obtain a corresponding control flow graph, and the control flow graph is used as the input for generating the encoder of the network control flow graph; initially, only the identifier sequence start is included in the generated annotation sequence <bos>The corresponding token. After that, it is the token sequence of the selected words. Using the code sequence and the current generated annotation sequence as the input of the code encoder of the generation network, after being processed by their respective encoders to obtain the structural feature vector and the semantic feature vector, they are fused. The decoder of the trained generation network selects the token of the next word that the annotation should have word by word according to the fused feature vector. The most basic method here is the greedy method, that is, selecting the word with the highest probability each time. Beam search with better effect can also be selected, but the search cost is greater and the time consumption is more. Until the end-of-sequence identifier appears <eos>, generate generated annotations; convert the tokens corresponding to words in the corpus into a complete sequence of English annotations. < / eos> < / bos> 10. A code comment generation device based on a sequence generative adversarial network, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the processor executes the executable code, it is used to implement the code annotation generation method based on the sequence generation adversarial network according to any one of claims 1-9.

Citation Information

Patent Citations

  • Method for generating Web codes based on UI of generative adversarial and convolutional neural networks

    CN110377282A

  • Enhanced code annotation automatic generation method and system

    CN111522581A