Code Summary Generation Method and System Based on KNN Decoding Enhancement

By introducing KNN decoding enhancement and offline databases in deep learning models, the problem of low code summary quality is solved, and higher quality and wider applicable code summary generation is achieved.

CN115237424BActive Publication Date: 2025-07-11DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210922180.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-02
Publication Date
2025-07-11
Estimated Expiration
2042-08-02

AI Technical Summary

Technical Problem

The existing code summary method based on deep learning models generates the summary quality of the summary and is difficult to read.

Method used

In the traditional deep learning-based code summary method, non-parametric KNN decoding enhancement is added. By constructing offline database storage code-summary pairs, using the KNN algorithm to select similar candidate key value pairs from the database for decoding, and combining the output of the Seq2Seq model, a higher-quality code summary is generated.

Benefits of technology

Improves the quality of code summary and generalization of models, reduces migration difficulty, and performs well when applied in multiple programming languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115237424B_ABST
    Figure CN115237424B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for generating code summaries based on KNN decoding enhancement. The encoder in the code summary model is used to perform vector representation on the context semantic information of the code sequence; the decoder in the model decodes the vector and outputs the corresponding predicted summary; at the same time, an offline database storing vector-word data pairs trained based on annotated corpus is innovatively added after the decoder, so that the prediction at each time step of the model performs weighted probability distribution on both the output of the decoder end of the original model and the judgment using the KNN algorithm to refer to the offline database. The present invention uses a non-parametric method to add KNN to the probability output distribution function of the original model, which plays a role in decoding enhancement on the basis of the original output, proves the practicability of the method and system for generating code summaries based on KNN decoding enhancement, improves the quality of the code summaries of the model and the generalization ability of the model, and at the same time reduces the migration difficulty of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of code summarization in software engineering, and particularly relates to a code summarization generation method and system based on KNN decoding enhancement. Background Art

[0002] Software development is an extremely expensive process. Software engineers have to solve the inherent complexity of software while avoiding errors and deliver powerful software products on time. Therefore, there is a continuous demand for innovation in software tools that enhance software reliability and maintainability. People are constantly seeking new methods to reduce software complexity to help engineers build better software. With the rapid development of Internet technology, another valuable resource has emerged: open-source software systems. The code repositories of these systems disclose source code, metadata information such as author information, bug fix records, and review processes. At the same time, the available data scale is huge, with billions of code sequences and millions of metadata instances. The availability of this "big code" has also promoted the emergence of a new data-driven software tool development method.

[0003] Code summarization is a task of writing a short natural language description for code. Its core purpose is to allow developers to understand the code content and purpose, or the use of this code in the whole program, through code summaries (or code annotation documents) without reading the code itself or reading only a small amount of code.

[0004] Currently, existing code summarization methods based on traditional deep learning models still have the defects of low-quality generated summaries and being difficult to read. Summary of the Invention

[0005] To solve the deficiencies of existing methods, the present invention provides a code summarization generation method and system based on KNN decoding enhancement. By adding a non-parametric KNN method to the traditional deep learning-based code summarization method to enhance the decoding part, the performance of the machine's automatic code summarization is improved, and the problem of poor performance of the existing technology in the field of code summarization is overcome.

[0006] The technical solution of the present invention is as follows:

[0007] A code summarization generation method based on KNN decoding enhancement, comprising the following steps:

[0008] Step 1, perform data preprocessing on the code segment to be summarized;

[0009] Step 2: Input the code snippet preprocessed in Step 1 into the encoder of the pre-trained Seq2Seq code summary generation model to obtain its context feature vector as the semantic information representation of the code snippet. Then, by inputting the code context feature vector into the decoder, the hidden layer vector at the current time step is output;

[0010] Step 3: Use the KNN algorithm to select the K candidate key-value pairs most similar to the hidden layer vector at the current time step output in Step 3 from the offline database, count the word frequencies among the current K candidate targets, and obtain the probability distribution of the target words on the vocabulary after normalization;

[0011] Step 4: Map the hidden layer vector output at the current time step by the decoder in Step 2 to a probability distribution on the vocabulary dimension through a linear layer and a Softmax operation, and combine it with the probability distribution on the vocabulary dimension obtained in Step 3 with the hyperparameter weight value to obtain the final prediction probability distribution; select the word corresponding to the highest probability as the prediction result at the current time step;

[0012] Step 5: Concatenate the prediction result in Step 4 to the input end of the decoder in Step 2, continue Steps 2 to 4, gradually generate the predicted words at each time step until the summary length limit is exceeded or the natural termination condition is triggered, and output the code summary of the target code snippet.

[0013] Further, the data preprocessing in Step 1 is specifically as follows: First, textify the code snippet, remove useless symbol information such as comments, carriage returns, and indents in the code. Second, perform word segmentation on the code text to form a text sequence composed of pseudo-English words. At the same time, perform splitting processing on identifiers using camel case naming and snake case naming, so that one identifier is split into multiple readable consecutive words, increasing the readability of the generated summary and the feasibility of expanding the summary candidate vocabulary.

[0014] Further, the pre-training in Step 2 is specifically as follows: Input the preprocessed code snippet and the manually annotated summary based on the code snippet into the decoder of the Seq2Seq code summary generation model to obtain the loss function score of the code summary generation model, and optimize the loss function of the model according to backpropagation, that is, complete the model pre-training.

[0015] Further, the Seq2Seq code summary generation model is specifically a Transformer model.

[0016] Further, the construction method of the offline database in Step 3 is:

[0017] Use the Faiss tool to build a key-value pair index library;

[0018] The code-summary pairs in the dataset used for pre-training are processed as follows:

[0019] Input the code part into the model encoder and the summary part into the model decoder;

[0020] Process the hidden layer vectors output by the last layer of the decoder. Each time step corresponds to a one-dimensional vector, and each one-dimensional vector corresponds to the word at the current position. Using each one-dimensional vector as the key and the word coordinate at the corresponding position as the value, store them in the key-value pair index library to form an offline database. Since this offline database may be large in scale, it may be necessary to calculate and store the data in the local memory for subsequent calculations.

[0021] The offline database contains a series of key-value pairs. The key is the final high-dimensional output vector representation of the decoder part of the Transformer model. Note that do not use the high-dimensional vector after passing through the vocabulary mapping layer as the key because the vocabulary is usually large. Therefore, the offline database constructed in this way has extremely high storage occupancy, and it usually represents a probability distribution and cannot reflect the semantic information contained in the current text; the value uses an integer to represent the token (i.e., label) that should be correctly predicted at the current time step, and the integer is the index of the token in the vocabulary. As shown in formula (1),

[0022]

[0023] where K represents the key, V represents the value, S is the encoder input; T i-1 is the decoder input for the first i - 1 time steps; t i is the word to be predicted at the current time step.

[0024] Through the trained code-summary model, obtain an offline database with key-value pair representations. Assuming the input PL sequence is S and the NL sequence is T, the calculation formula (2) of the final prediction probability distribution in step 4 is:

[0025] P(t i |S,T i-1 )=λP KNN (t i |S,T i-1 )+(1-λ)P c2nl (t i |S,T i-1 ) (2) where, P c2nl is the prediction probability of the Seq2Seq code-summary generation model; λ is a hyperparameter used to control the proportion of the KNN part participating in the prediction; t iRepresents the word to be predicted currently, T i-1 = t1, t2, ..., t i-1 .

[0026] Among them, P c2nl is the prediction probability of the Transformer module, and λ is a hyperparameter used to control the proportion of the KNN part participating in the prediction. T i-1 = t1, t2, …, t i-1 Represents the word to be predicted currently.

[0027] Furthermore, the specific implementation formula for step 3 is:

[0028]

[0029] Among them, P KNN is the word prediction probability distribution obtained using the KNN method on the offline database; t i is the word to be predicted currently; S is the encoder input; T i-1 is the decoder input for the previous i - 1 time steps; N is the offline database set; I is the indicator function; F is the retrieval vector for the current time step; d is the cosine distance calculation function; T is the hyperparameter.

[0030] Furthermore, the value range of the hyperparameter in step 4 is 0.4 to 0.7. The preferred value of the hyperparameter is 0.6.

[0031] The code summary generation system based on KNN decoding enhancement includes:

[0032] A data preprocessing module for preprocessing the code segments to be summarized;

[0033] An encoder module for extracting the context feature vectors of the code text sequence output by the data preprocessing module;

[0034] A decoder module for decoding the context feature vectors output by the encoder module to output the hidden layer vector for the current time step;

[0035] An offline database module for storing vector - word pairs, using the KNN algorithm to select the K candidate key - value pairs most similar to the hidden layer vector output by the decoder module for the current time step, counting the word frequencies among the current K candidate targets, and obtaining the probability distribution of the target words on the vocabulary after normalization;

[0036] A generation module is used to map the hidden layer vector output by the decoder at the current time step to a probability distribution in the vocabulary dimension through a linear layer and a Softmax operation, and combine it with the probability distribution in the vocabulary dimension obtained from the offline database module with a hyperparameter weight value to obtain a final predicted probability distribution; select the word corresponding to the highest probability as the output at the current time step;

[0037] The output result of the generation module is spliced to the input end of the decoder module to complete the entire output of the code summary of the target code segment.

[0038] The beneficial effects of the present invention are as follows. The method of the present invention stores vector data for the code summary model, uses the non-parametric method KNN to join the probability output distribution function of the original model, which plays a role in decoding enhancement on the basis of the original output, not only improves the quality of the code summary of the model, but also can be generalized to the code summary tasks of multiple programming languages. Using this method, not only can a code summary generation framework be built from scratch, but also an offline database can be added only on the basis of a model trained by predecessors, both of which can improve the quality of the generated summary to a certain extent. And while improving the generalization of the model, the migration difficulty of the model is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a schematic diagram of the overall process of the code summary generation method based on KNN decoding enhancement of the present invention.

[0040] Figure 2 is a schematic diagram of the overall architecture of the model in the code summary generation method based on KNN decoding enhancement of the present invention.

[0041] Figure 3 is a schematic diagram of probability prediction with KNN decoding enhancement added in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0043] Embodiment 1

[0044] Adopt as Figure 1 , Figure 2 shown, the code summary generation method based on KNN decoding enhancement, including the following steps:

[0045] S101: Perform data preprocessing on the code samples;

[0046] Among them, in data preprocessing, the code file or code snippet is first texturized, and useless symbol information such as comments, carriage returns, and indents in the code is removed. Then, the code text is tokenized to form a text sequence consisting of a set of quasi-English words. At the same time, in order to improve the readability of the generated abstract and the feasibility of expanding the abstract candidate word list, identifiers using camel case naming and snake case naming are split, and an identifier is divided into multiple readable consecutive words;

[0047] S102: Encode the preprocessed code text sequence based on the encoder part in the code abstract generation model to obtain its context feature vector as the semantic information representation of the code snippet; the decoder part decodes the context feature vector by the encoder part and outputs the hidden layer vector at the current time step.

[0048] Among them, the encoder part adopts the left half of the Transformer model. The preprocessed code text sequence in S101 is input into the encoder through one-hot encoding by the vocabulary, and the position embedding method adopts relative position encoding; this part also provides the encoding part for the input of the target language, that is, the natural language code abstract, in the training and prediction phases. This function is used for calculating the loss function in training and the teacher-forcing inference process in the prediction process;

[0049] Among them, the decoder part adopts the right half of the Transformer model. There are two input ends in the decoder part here. One is the context feature vector of the final output of the encoder, which represents the context semantic feature information of the code snippet to be abstracted; the other is the encoding part for the input of the target language, that is, the natural language code abstract. Further, in the training process, this part is the complete natural language code abstract, that is, the label, for calculating the loss function value, while in the prediction, that is, the process of generating the abstract, this part adopts a step-by-step process for decoding operations. The input of the initial decoding end is " <sos>”, a single program-recognizable symbol, used to represent the start of a sentence.

[0050] S103: In the offline database construction part of the code summary generation method based on KNN decoding enhancement, this part is the core algorithm part of this embodiment. Its objective is to store the hidden layer vector output of the decoder at each time step of the code summary model and the actual word corresponding to this vector in the offline database in the form of key-value pairs by offline storage;

[0051] Among them, in the final code summary generation process, not only the hidden layer vector output of the S102 decoder part is no longer relied on as the only consideration for the predicted target. This method first uses the non-parametric K-Nearest Neighbor (KNN) algorithm to select the K candidate key-value pairs most similar to the hidden layer vector of the final output of the S102 decoder part at the current time step from the offline database, and statistically counts the word frequencies among the current K candidate targets. After normalization, the predicted probability of the target word on the vocabulary can be obtained.

[0052] S104, maps the hidden layer vector output by the S102 decoder at the current time step to a probability distribution on the vocabulary dimension through a linear layer and Softmax operation, and combines it with the probability distribution on the vocabulary dimension obtained in S103 with the hyperparameter weight value to obtain the final predicted probability distribution; selects the word corresponding to the highest probability as the prediction result at the current time step;

[0053] Among them, in this part, the proportion of the S103 offline database in the prediction process is set, and the weight value is set through the hyperparameter λ. The output vector after integration of the foregoing module is subjected to dimensional feature transformation through a linear layer network to obtain the word probability distribution mapped to the vocabulary. The final predicted target can be selected according to different algorithms. For example, beam search selects the N candidates with the highest current probability each time; the greedy algorithm can also be used, which selects the candidate with the highest current probability as the prediction result each time. Here, the greedy algorithm can also be regarded as a special case of beam search when the beam window value is 1.

[0054] S105, splices the prediction result of S104 to the input end of the S102 decoder, continues with S102 to 104, and gradually generates the predicted words at each time step until the summary length limit is exceeded or the natural termination condition is triggered, and outputs the code summary of the target code segment.

[0055] As shown in Figure 2 The overall architecture of the model of the present invention is used to describe the overall method proposed by the present invention in detail.

[0056] In this embodiment, based on the traditional code summarization method of the sequence-to-sequence (Seq2Seq) deep learning model, an offline database storing vector-word data pairs trained based on labeled anticipation is innovatively added after the decoder of Seq2Seq. This enables the model's prediction at each time step to not only depend on the output of the decoder end of the original model, but also refer to the judgment of the offline database using the KNN algorithm, forming a weighted probability distribution between the two, thus improving the performance and generalization ability of the model.

[0057] The model first receives the preprocessed code text sequence as the input part. In the Transfomer encoder, the self-attention mechanism calculates its attention scores, and the calculation method is shown in formula (4):

[0058]

[0059] Among them, Q, K, and V are the matrices after the encoder end projects the input H of the code model respectively, and the calculation method is shown in formula (5):

[0060] Q = HW q , K = HW k , V = HW v (5)

[0061] Among them, W q , W k , W v are the trainable projection weight matrices in the model;

[0062] Since this embodiment uses the relative position encoding method in the position embedding method of the traditional Transformer model, the calculation of the attention scores after adding the relative position embedding is shown in formulas (6), (7), and (8):

[0063]

[0064]

[0065]

[0066] Among them, is calculated through formula (9):

[0067]

[0068] clip(x, k) = max(-k, min(x, k)) (10)

[0069] Define the final output of a single encoder as X hidden , and integrate the above Attention score z i into a matrix representation, defined as X attention . Next, it needs to go through a Feed-Forward network, and the calculation formula is shown in formula (11):

[0070]

[0071] Then, through the residual connection and the layer normalization network, the calculation formula is shown in formula (12):

[0072] X hidden = LayerNorm(X hidden + X attention ) (12)

[0073] The above obtained X hidden is the context feature vector that contains the context semantic information of the code snippet to be summarized and is finally output by the encoder and required by the decoder.

[0074] Furthermore, the model transports X hidden to each layer of the decoder as the input vector of K and V;

[0075] At the same time, the decoder module also accepts the summary text (target) as input. This part calculates the self-attention score of the label text itself to itself through the Masked Multi-Head Attention mechanism. Here, the MASK matrix will cover the weight score of the former to the latter in the text sequence as negative infinity to prevent it from having an impact during the Softmax process;

[0076] After the Masked Multi-Head Attention module, the output vector is used as the input of Q in the next module, while the inputs of K and V come from the output of the last layer of the encoder. The structure of this part is the same as that of the module unit of the encoder, so the specific calculation formula is exactly the same as that described above. The difference is that here, the model calculates the attention score of the target text to the source text, so the obtained matrix is no longer a square matrix.

[0077] In the traditional code summary generation method based on the Seq2Seq model, the above has roughly covered its model architecture content, that is Figure 2 the left half of the Transformer part, and maps the output vector of the last layer of its decoder to the vocabulary dimension for code summary inference work;

[0078] In the inference part, the model adopts a step-by-step generation method based on Teacher-Forcing. First, the input part of the model decoder is modified to " <sos>", this flag represents the start of a sentence, and the input to the encoder is the sequence of code text to be summarized. After both are input into the trained model, the output of the hidden layer vector of the last layer of the decoder can be obtained. The sliced vector of the last time step is taken as the predicted semantic vector for the next word. After mapping it to the vocabulary dimension and performing the Softmax operation, this vector can be regarded as the predicted probability distribution of words on the scale of the vocabulary. The word corresponding to the highest probability is selected as the prediction result for the current time step. The prediction result is concatenated behind the input of the decoder, and the above process is continued to gradually generate the predicted words for each time step, finally completing the generation of the code summary text.

[0079] The code summary generation method based on KNN decoding enhancement provided in this embodiment, the core of which lies in using a parameter-free KNN algorithm for the prediction and judgment of the offline database, including: the construction of the offline database and the use of the offline database;

[0080] Exemplarily, the construction method of the offline database is as follows:

[0081] Use the Faiss tool to build a key-value pair index library;

[0082] The code-summary pairs in the dataset used for pre-training are processed as follows:

[0083] Input the code part into the model encoder and the summary part into the model decoder;

[0084] Process the hidden layer vector output by the last layer of the decoder. Each time step corresponds to a one-dimensional vector, and each one-dimensional vector corresponds to the word at the current position. Since the summary text is concatenated with " <sos>Therefore, an offset (Shift) process is performed between the hidden layer vector and the label text to make their positions corresponding. Taking each dimension vector as the key and the word coordinates at the corresponding position as the value, they are stored in the key-value pair index library to form an offline database. Since the scale of this offline database may be huge, it needs to be calculated and stored in the local memory for subsequent calculations.

[0085] Among them, the offline database contains a series of key-value pairs. The key is the final high-dimensional output vector representation of the decoder part of the Transformer model (in this embodiment, the dimension is set to 768). Do not use the high-dimensional vector after passing through the vocabulary mapping layer as the key because the vocabulary is usually very large. Therefore, the Data Store storage constructed thereby occupies extremely high space, and it usually represents a probability distribution and cannot reflect the semantic information contained in the current text; the value uses an integer to represent the token (i.e., the label) that should be correctly predicted at the current time step, and the integer is the index of the token in the vocabulary, as shown in the following formula:

[0086]

[0087] Adopt as Figure 3 shown, the usage method of the offline database is as follows:

[0088] Through the trained code summary model, an offline candidate word library with key-value pair representations is obtained. Assuming that the input PL sequence is S and the NL sequence is T, the following model probability distribution holds, as shown in the following formula:

[0089] P(t i |S,T i-1 )=λP KNN (t i |S,T i-1 )+(1-λ)P c2nl (t i |S,T i-1 )

[0090] Among them, P c2nl is the prediction probability of the Transformer model; λ is a hyperparameter used to control the proportion of the KNN part participating in the prediction; t i represents the word to be predicted currently, T i-1 =t1,t2,...,t i-1 .

[0091] And the calculation method of P KNN is as shown in the following formula:

[0092]

[0093] The model architecture after adding the KNN decoding enhancement function module is as Figure 2 shown as a whole;

[0094] The inference process of code summary text generation adopts the method of beam search, where the beam window size is set to 4.

[0095] A sample of the finally generated code summary is as follows:

[0096] Code sample:

[0097] def getPath edges pathIndexes loop z path[]for pathIndexIndex inxrange len pathIndexes pathIndex pathIndexes[pathIndexIndex]edge edges[pathIndex]carveIntersection getCarveIntersectionFromEdge edge loop z pathappend carveIntersection return path

[0098] Manual summary (tag):

[0099] get the path from the edge intersections.

[0100] Generated summary:

[0101] get the path component of an unproven vector.

[0102] Evaluation metrics:

[0103] Table 1 Experimental results of the code summary generation method based on KNN decoding enhancement

[0104]

[0105] As shown in Table 1, it can be found that starting from λ greater than or equal to 0.4 until 0.6, the KNN decoding enhancement model can improve the performance of the original Transformer to a certain extent. Until 0.7, the performance no longer improves, but instead shows a slight downward trend in metrics. When λ is equal to 1.0, that is, completely ignoring the final output vector representation of the Transformer's Decoder and only using KNN to output the current word from the candidate word library, the effect is worse than the performance of the original Transformer.

[0106] Example 2

[0107] A second embodiment of the present invention provides a code summary generation system based on KNN decoding enhancement;

[0108] The code summary generation system based on KNN decoding enhancement includes:

[0109] A data preprocessing module, which is used to preprocess the code snippets to be summarised;

[0110] The encoder module is used to extract the contextual semantic feature vector of the code text sequence output by the above module;

[0111] The decoder module is used to decode and output the semantic feature vector generated by the encoder;

[0112] The offline database module is used to store word-vector pairs and provide data support for the decoding enhancement function of this method;

[0113] A generation module, which generates relevant code summaries about code snippets;

[0114] It should be noted that the data preprocessing module, encoder module, decoder module, offline database module and generation module correspond to S101-S105 in Example 1 respectively, and the examples and application backgrounds implemented by the above modules and their corresponding steps are the same, including but not limited to the contents disclosed in Example 1. Furthermore, the modules can be run as part of the system on a computer system with computer instruction execution function.

[0115] Example 3

[0116] A computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.

[0117] Example 4

[0118] An electronic device is provided, comprising a processor, a storage device, and an executable program stored on the storage device and executable on the processor, wherein the executable program can implement a code summary generation method provided by any possible implementation of various possible implementations of the first aspect.

[0119] This text uses the preferred embodiments of the present invention to expound on the principles and implementation schemes of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.< / sos> < / sos> < / sos>

Claims

1. A method for generating code summaries based on enhanced KNN decoding, characterized in that It includes the following steps: Step 1: Perform data preprocessing on the code snippet for which the summary is to be generated; Step 2: Input the code snippet preprocessed in Step 1 into the encoder of the pre-trained Seq2Seq code summary generation model to obtain its context feature vector as the semantic information representation of the code snippet; then, by inputting the code context feature vector into the decoder, output the hidden layer vector at the current time step; Step 3: Use the KNN algorithm to select the K candidate key-value pairs most similar to the hidden layer vector at the current time step output in Step 3 from the offline database, count the word frequencies among the current K candidate targets, and after normalization, obtain the probability distribution of the target words on the vocabulary; Step 4: Map the hidden layer vector output at the current time step of the decoder in Step 2 to the probability distribution on the vocabulary dimension through a linear layer and a Softmax operation, and combine it with the probability distribution on the vocabulary dimension obtained in Step 3 with the hyperparameter weight value to obtain the final predicted probability distribution; Select the word corresponding to the highest probability as the prediction result at the current time step; Step 5: Concatenate the prediction result in Step 4 to the input end of the decoder in Step 2, and continue Steps 2 to 4 to gradually generate the predicted words at each time step until the summary length limit is exceeded or the natural termination condition is triggered, and output the code summary of the target code snippet.

2. The method for generating code summaries based on KNN decoding enhancement according to claim 1, wherein The data preprocessing in Step 1 is specifically as follows: First, textify the code snippet, remove the useless symbol information in the code, and second, perform word segmentation on the code text to form a text sequence composed of a group of pseudo-English words. At the same time, perform splitting processing on the identifiers using camel case naming and snake case naming, so that one identifier is split into multiple readable consecutive words.

3. The method for generating code summaries based on KNN decoding enhancement according to claim 1, wherein The pre-training in Step 2 is specifically as follows: Input the preprocessed code snippet and the manually annotated summary based on the code snippet into the decoder of the Seq2Seq code summary generation model at the same time to obtain the loss function score of the code summary generation model, and optimize the loss function of the model according to backpropagation, that is, complete the model pre-training.

4. The method for generating a code summary based on KNN decoding enhancement according to claim 1, wherein The Seq2Seq code summary generation model is specifically a Transformer model.

5. The method for generating a code summary based on KNN decoding enhancement according to claim 1, wherein The construction method of the offline database in Step 3 is: Use the Faiss tool to build a key-value pair index library; Perform the following processing on the code-summary pairs in the dataset used for pre-training: Input the code part into the model encoder and the summary part into the model decoder; Process the hidden layer vector output by the last layer of the decoder. Each time step corresponds to a one-dimensional vector, and each one-dimensional vector corresponds to the word at the current position. Use each one-dimensional vector as the key and the word coordinate at the corresponding position as the value, and store it in the key-value pair index library to form an offline database.

6. The method for generating a code summary based on KNN decoding enhancement according to claim 1, wherein The specific implementation formula of Step 3 is: Among them, P KNN is the word prediction probability distribution obtained using the KNN method on the offline database; t i is the word to be predicted currently; S is the encoder input; T i-1 is the decoder input at the previous i - 1 time steps; N is the offline database set; I is the indicator function; F is the retrieval vector at the current time step; d is the cosine distance calculation function; T is the hyperparameter.

7. The method for generating code summaries based on KNN decoding enhancement according to claim 1, wherein The calculation formula of the final predicted probability distribution in Step 4 is: P(t i |S,T i-1 ) = λP KNN (t i |S,T i-1 ) + (1 - λ)P c2nl (t i |S,T i-1 ) Among them, P c2nl is the predicted probability of the Seq2Seq code summary generation model; λ is a hyperparameter used to control the proportion of the KNN part participating in the prediction; t i represents the word to be predicted currently, T i-1 = t1, t2,..., t i-1 .

8. The method for generating a code summary based on KNN decoding enhancement according to claim 1, wherein The value range of the hyperparameter in Step 4 is 0.4 to 0.

7.

9. The method for generating a code summary based on KNN decoding enhancement according to claim 1, wherein The hyperparameter in Step 4 takes the value of 0.

6.

10. A code summary generation system based on KNN decoding enhancement, characterized in that, It includes: A data preprocessing module for performing preprocessing work on the code snippet for which the summary is to be generated; An encoder module, which is used to extract the context feature vectors of the code text sequence output by the data preprocessing module; A decoder module, which is used to decode the context feature vectors output by the encoder module to output the hidden layer vector at the current time step; An offline database module, which is used to store vector-word pairs, select the K candidate key-value pairs most similar to the hidden layer vector output by the decoder module at the current time step from the offline database using the KNN algorithm, count the word frequencies among the current K candidate targets, and obtain the probability distribution of the target words on the vocabulary after normalization; A generation module, which is used to map the hidden layer vector output by the decoder at the current time step to a probability distribution on the vocabulary dimension through a linear layer and a Softmax operation, and combine it with the probability distribution on the vocabulary dimension obtained by the offline database module with a hyperparameter weight value to obtain the final predicted probability distribution; Select the word corresponding to the highest probability as the output at the current time step; The output result of the generation module is spliced to the input end of the decoder module to complete the full output of the code summary of the target code segment.

Citation Information

Patent Citations

  • Text abstract automatic generation method based on self-attention network

    CN110209801A

  • Reader preference-based personalized digital book recommendation system and method, computer and storage medium

    CN113590970A