Source code abstract generation method and system based on enhanced prompt learning framework
By adopting an enhanced prompt learning framework in code summary generation, combining code pre-trained models and large language models to generate knowledge and structure prompt vectors, the problems of complex prompt word design and shallow code structure understanding in the existing technology are solved, and more accurate and flexible code summary generation is achieved.
Patent Information
- Application Number
- CN202510525258.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
When generating code summary, the existing technology faces multilingual and diverse code snippets, the prompt word design is complex and has poor flexibility, and the code structure is shallow, making it difficult to generate accurate fine-grained code summary.
Using a method based on an enhanced prompt learning framework, combining code pre-trained models and large language models, knowledge prompt vectors and structural prompt vectors are generated through mapping modules and structural proxy modules, reducing dependence on manual design prompt templates, and generating more accurate code summary.
It realizes flexible adaptation to multiple programming languages and code snippets of different scales, and the generated code summary is more accurate and conforms to the code function, improving the efficiency of code automation understanding and document generation.
Smart Images

Figure CN120045226A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of natural language processing and software engineering, and in particular to a source code summary generation method and system based on an enhanced prompt learning framework. Background Art
[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.
[0003] In the modern software development process, understanding and maintaining source code is a key link. Developers usually need to annotate the code to improve the readability and maintainability of the code. However, manually writing code comments is a time-consuming and easily overlooked task, especially in complex projects or rapid iterative development, where incomplete or inaccurate code comments are very common. As the project is updated and iterated, the original comments often do not match the actual code functions, further increasing the difficulty of understanding and maintaining the code. To solve this problem, many technologies for automatically generating code summaries have gradually been proposed, aiming to help developers quickly understand the functions and logic of the code by automatically generating natural language descriptions.
[0004] In the prior art, patent CN118152287A proposes a rich text code review method based on a large language model. This method matches code snippets in rich text through regular expressions, combines knowledge bases of different programming languages with predefined code review prompt words, and uses the natural language processing capabilities of the large language model to generate review opinions. This method can effectively extract code content from rich text and conduct reviews, but its limitation is that it relies less on code understanding, and mainly identifies code snippets by matching regular expressions and prompt words. Especially when dealing with complex code structures and multi-language scenarios, this method has poor flexibility and needs to design specific prompt templates for different programming languages, which increases the complexity of prompt design and makes it difficult to cope with diverse code formats and syntax.
[0005] In addition, patent CN118550579A proposes a hierarchical code summary generation method based on a large language model thinking chain. The call relationship between code modules is extracted through a static analysis tool, and a heuristic algorithm is used to calculate the weight of each code file to generate a module-level code summary. This method performs well when dealing with large modular software systems and can generate high-level code structure summaries. However, this solution has limitations when dealing with smaller-grained code snippets, especially in the generation of code summaries at the function level or class level, and often fails to accurately capture the detailed functions of the code. This method relies on static analysis and heuristic algorithms. For complex code logic, especially non-modular code, the generated summary is too general and lacks a deep understanding of the internal structure and semantics of the code.
[0006] Another prior art patent CN117873559A proposes a code summary generation method based on a large language model and a static analysis tool. This method generates an abstract syntax tree through lexical and grammatical parsing, and generates a control flow graph to describe the execution flow of the code. Combined with the pseudocode block generation method, the code summary is depth-first traversed to generate a complete code summary. Although this method has certain advantages in processing complex code logic and arithmetic expressions, it relies on static analysis, control flow graphs, and pseudocode block generation, and is not as flexible as the automatic generation method based on prompt learning. In addition, this method has limitations in the number of tokens of the input code and the range of programming languages supported, and has limitations in the generation and understanding of multi-language code.
[0007] In summary, existing code summary generation technologies generally have the following shortcomings when facing multi-language and diverse code snippets: 1. The prompt word design is complex and inflexible. Many methods rely on manually designed prompt templates and are difficult to adapt to different types of programming languages and code structures; 2. Existing methods have a superficial understanding of code structure, especially when generating fine-grained summaries of code snippets, they cannot accurately reflect the function and logic of the code; 3. Some existing technologies are mainly suitable for top-level code summary generation of modular systems. For smaller-grained code snippets, the summaries they generate are too abstract, which is not conducive to developers' in-depth understanding of the specific functions of the code. Summary of the invention
[0008] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a source code summary generation method and system based on an enhanced prompt learning framework, which is used to automatically generate code summaries, and generate more accurate code summaries by combining the structural understanding ability of the code pre-training model with the natural language generation ability of the large language model.
[0009] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: In a first aspect, the present invention provides a method for generating a source code summary based on an enhanced prompt learning framework, comprising: Obtain code snippets, input them into the enhanced prompt learning framework for processing, and generate code summaries in natural language form; The enhanced prompt learning framework includes a code pre-training model, a mapping module, a structure proxy module and a large language model; the code pre-training model receives code snippets and performs feature extraction to obtain code features, and inputs the code features into the mapping module to generate a knowledge prompt vector; at the same time, the code features are input into the structure proxy module to generate a structure prompt vector; the knowledge prompt vector and the structure prompt vector are spliced and input into the large language model as prompt information to generate a code summary in natural language form.
[0010] According to a further technical solution, the mapping module includes a left encoder and a right encoder, the left encoder includes a self-attention layer and a feedback-forward layer connected in sequence, and the right encoder includes a self-attention layer, a cross-attention layer and a feedback-forward layer connected in sequence.
[0011] According to a further technical solution, the specific steps of the mapping module generating the knowledge hint vector are as follows: When training the mapping module, the background knowledge is input into the left encoder for processing to obtain a high-order representation of the background knowledge; the virtual tag is input into the right encoder, and the code feature is input into the cross-attention layer of the right encoder. After the virtual tag and the code feature interact in the cross-attention layer, they are sent to the forward feedback layer to obtain a simulated knowledge representation, and finally a trained mapping module is obtained; The trained mapping module only retains the right encoder, inputs the virtual label into the right encoder, and inputs the code features into the cross-attention layer of the right encoder to obtain the simulated knowledge representation and input it into the projector for feature mapping and dimension alignment to obtain the knowledge prompt vector.
[0012] A further technical solution is that when the mapping module is trained and optimized, the high-order representation of background knowledge and the simulated knowledge representation are used to design the contrast estimation loss, cross entropy loss, and bidirectional matching loss respectively, and the total loss is the sum of the contrast estimation loss, the cross entropy loss, and the bidirectional matching loss.
[0013] A further technical solution is that the specific steps of the structure proxy module generating a structure hint vector through a variational autoencoder are: encoding the code features through the encoder of the variational autoencoder, and optimizing the variational autoencoder using reconstruction loss in the decoder, so that the variational autoencoder captures the structure hint vector.
[0014] A further technical solution is to concatenate the knowledge hint vector with the structure hint vector by inputting the code snippet into the embedding layer of the large language model to obtain text embedding, concatenating and merging the text embedding, the knowledge hint vector and the structure hint vector, and inputting the concatenated code snippet into the large language model.
[0015] A further technical solution is to freeze the large language model for training optimization, and its joint optimization loss is the weighted sum of reconstruction loss, KL loss and difference generation loss.
[0016] In a second aspect, the present invention provides a source code summary generation system based on an enhanced prompt learning framework, comprising: A code acquisition module is configured to: acquire code snippets; A summary generation module is configured to: input the code snippet into an enhanced prompt learning framework for processing, and generate a code summary in a natural language form; The enhanced prompt learning framework includes a code pre-training model, a mapping module, a structure proxy module and a large language model; the code pre-training model receives code snippets and performs feature extraction to obtain code features, and inputs the code features into the mapping module to generate a knowledge prompt vector; at the same time, the code features are input into the structure proxy module to generate a structure prompt vector; the knowledge prompt vector and the structure prompt vector are spliced and input into the large language model as prompt information to generate a code summary in natural language form.
[0017] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in a method for generating a source code summary based on an enhanced prompt learning framework as described in the first aspect.
[0018] In a fourth aspect, the present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a method for generating a source code summary based on an enhanced prompt learning framework as described in the first aspect is implemented.
[0019] One or more of the above technical solutions have the following beneficial effects: By combining the advantages of LLM and CodePLM, this paper proposes an enhanced hint learning framework that can deeply understand the code structure and generate more accurate natural language summaries that conform to the code function. At the same time, this method reduces the dependence on manually designed hint templates, has higher flexibility and adaptability, and is suitable for multiple programming languages and code snippets of different sizes.
[0020] The present invention combines the code pre-training model's understanding of code structure and the natural language generation capability of the large language model, and the generated code summary is more accurate and in line with the actual function. At the same time, by using the mapping module and the structure proxy module to generate knowledge prompt vectors and structure prompt vectors, the reliance on manually designed prompt templates is reduced, the flexibility of prompt generation is enhanced, and code snippets in multiple programming languages are adapted. It supports multiple programming languages, such as Python, Java, C++, etc., and enhances the applicability in different programming environments. In addition, the present invention optimizes the alignment of the prompt vector and the embedding of the large language model, and does not need to waste a lot of time to fine-tune the large language model. In low-resource scenarios (such as insufficient training data or limited computing resources), it can still generate high-quality code summaries that are applicable to multiple programming languages and code structure types (such as functions, classes, control flows, etc.), which improves the efficiency of automatic code understanding and document generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0022] Figure 1 is a flow chart of a method for generating a source code summary according to an embodiment of the present invention; Figure 2 This is a structural diagram of an enhanced prompt learning framework according to an embodiment of the present invention; Figure 3 This is a code summary example of an embodiment of the present invention; Figure 4 The embodiment of the present invention is based on Figure 3 The code snippets in the code summary are generated by different methods. DETAILED DESCRIPTION
[0023] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0024] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0025] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0026] Embodiment 1 like Figure 1 As shown, this embodiment discloses a method for generating a source code summary based on an enhanced prompt learning framework, the method comprising the following steps: Obtain code snippets, input them into the enhanced prompt learning framework for processing, and generate code summaries in natural language form; The enhanced prompt learning framework includes a code pre-training model, a mapping module, a structure proxy module and a large language model; the code pre-training model receives code snippets and performs feature extraction to obtain code features, and inputs the code features into the mapping module to generate a knowledge prompt vector; at the same time, the code features are input into the structure proxy module to generate a structure prompt vector; the knowledge prompt vector and the structure prompt vector are spliced and input into the large language model as prompt information to generate a code summary in natural language form.
[0027] In this embodiment, combining the advantages of large-scale language model (LLM) and code pre-trained model (CodePLM), an enhanced prompt learning framework (EPG4CS framework) is proposed. The EPG4CS framework aims to use CodePLMs to create advanced features of LLM and complete the code summarization task. In the enhanced prompt learning framework, a two-stage independent training strategy is adopted. In the first stage, the parameters of the code pre-trained model (CodePLM) are frozen to ensure stable vector output, and the mapping module is optimized by contrast estimation loss, cross entropy loss and bidirectional matching loss; in the second stage, the parameters of the large language model (LLM) are set to a frozen state, and the mapping module, projector and structure proxy module are optimized.
[0028] EPG4CS includes the following key components: (1) Code pre-trained model (frozen code pre-trained model CodePLMs), which is pre-trained on code language tasks and generates embedding vectors for input code that are useful for downstream tasks; The code pre-training model may be CodeBERT, CodeT5, or other code-based pre-training models, which can extract the grammatical structure and context information of the code snippet.
[0029] (2) Mapper, which aims to align code representation with knowledge information through various pre-training tasks; (3) Structural Agent, which aims to extract the structural features of the code using VAE; (4) Large language model (FrozenLLM), which is responsible for generating accurate code summaries after receiving rich prompt information. The large language model is StarcodeBase, PolyCoder, or other models that support natural language generation.
[0030] In this embodiment, the randomly initialized soft hint method often leads to unstable LLM output and reduced generalization ability. In order to solve this problem, the present invention designs three training losses in the first stage to reduce the difficulty of training in the second stage, including contrast estimation loss, cross entropy loss and bidirectional matching loss, and pre-trains the mapping module Mapper, aiming to allow Mapper to generate the optimal soft hint vector (i.e., knowledge hint vector, hint vector containing code background knowledge) in the second stage.
[0031] like Figure 2As shown on the left, the mapping module (mapper) includes a left encoder (LT) and a right encoder (RT). The left encoder includes a self-attention layer and a feedforward layer connected in sequence, and the right encoder includes a self-attention layer, a cross-attention layer and a feedforward layer connected in sequence. The weights of the self-attention layer and the feedforward layer of the left encoder and the right encoder are shared. The whole process of processing the pre-trained data pair <code snippet C, background knowledge K> can be described as follows: First, the frozen code pre-training model generates a representation from the input (code snippet) (i.e., code features), and then Input to the right encoder. The code snippet contains multiple codes, each of which produces a representation At the same time, the left encoder is used to learn the background knowledge Processing to extract a high-level representation Then, the right encoder integrates the virtual marker and code features , generating a simulated knowledge representation The virtual tokens are a set of randomly initialized vector parameters whose dimensions match the input dimensions of the first layer of the encoder and are used as the input to the first layer of the encoder. The formal form of these steps is as follows:
[0032]
[0033] in, A high-level representation of background knowledge, represents the right encoder, represents the left encoder, represents the model parameters shared between the left encoder and the right encoder, represents simulated knowledge representation, represents a virtual tag, Indicates code characteristics, It represents the cross-attention mechanism, which is used to realize the information interaction between virtual tags and code features using the cross-attention mechanism, so that the virtual tags can effectively extract knowledge that is closely integrated with the logic and function of the code from the code features.
[0034] The steps of the mapping module during training are as follows: the background knowledge is processed through the left encoder to obtain a high-order representation of the background knowledge and capture the deep semantic information therein; the virtual tag is input into the right encoder, and the code features are input into the cross-attention layer of the right encoder at the same time. The virtual tag is also input into the cross-attention layer after being calculated by the self-attention layer. The virtual tag and the code features interact with each other through the cross-attention mechanism, and then the virtual tag is further sent to the forward feedback layer to obtain the simulated knowledge representation; when the mapping module is trained and optimized, the high-order representation of the background knowledge and the simulated knowledge representation are used to design the contrast estimation loss, cross entropy loss and bidirectional matching loss. The final optimization model is the sum of the three losses.
[0035] The steps of the mapping module when applied are as follows: After the training is completed, the trained mapping module is obtained, and only the right encoder part is retained, which is consistent with the training. The virtual tag is input into the right encoder, and the code feature is input into the cross-attention layer of the right encoder. The virtual tag is also input into the cross-attention layer after being calculated by the self-attention layer. At this time, the virtual tag obtains the code information most relevant to the text from the code feature to obtain the simulated knowledge representation; the simulated knowledge representation is input into the projector for feature mapping and dimension alignment to obtain the knowledge prompt vector.
[0036] Next, the design of the pre-training task of the mapping module Mapper in the first phase will be introduced in detail.
[0037] 1) Code knowledge contrastive learning: Contrastive learning enhances the model’s ability to extract key features for distinguishing positive and negative samples by clearly distinguishing code snippets. Based on contrastive learning, the output of the left encoder is maximized. Simulated knowledge representation with right encoder output The mutual information between them is optimized, thereby optimizing the simulation effect of the virtual marker.
[0038] Specifically, first calculate and Select the pair of samples with the highest score (e.g. & ) as positive samples, while other sample pairs with low similarity or irrelevant are regarded as negative samples. In the training optimization process, the noise contrast estimation (InfoNCE) loss function is used, which is expressed as:
[0039] in, represents the comparison estimated loss, represents a high-level representation of the virtual label, A high-level representation of background knowledge, represents the temperature parameter, Represents a similarity function, such as cosine similarity or Euclidean distance.
[0040] 2) Knowledge text generation: To further enhance the text modeling capability of Mapper, text generation loss is introduced. Its training goal is to enable Mapper to receive code features. Generating text with background knowledge Consistent output. In the design of Mapper, since Mapper’s architecture does not allow direct interaction between the frozen code pre-trained model and the knowledge text, the information required to generate the text must first be extracted by the query and then passed to the knowledge text through the self-attention layer. Therefore, the virtual tag Forced from code features Here, a multimodal causal self-attention mask is used to control the interaction between the left encoder and the right encoder, and the [DEC] token is used as a signal to start decoding in the left encoder, so is the text sequence to be generated by Mapper. Finally, the Mapper is optimized by cross entropy loss:
[0041] in, represents the cross entropy loss, represents the number of samples in the collected sample, , Represents the index, Indicates The true probability distribution of knowledge texts, Indicates The model predicts the probability distribution of knowledge text.
[0042] 3) Code-knowledge matching loss: Through the above two tasks, the relationship between virtual tags and background knowledge is aligned. However, due to the modality gap between programming language and natural language, a two-way matching loss is designed to enable Mapper to capture code snippets more accurately. and background knowledge text The specific optimization loss function is described as follows:
[0043] in, represents the bidirectional matching loss, Indicates Code features, Indicates A high-level representation of background knowledge, represents the margin hyperparameter, represents the distance function, Defines the rectifier operation.
[0044] Generate learning from frozen LLM.
[0045] 1) Structural Agent: Although LLMs are widely considered to have excellent performance in understanding and generalization, their decoding-based training model does not fully consider the structural semantics of the code. To address this problem, this paper designs a structural agent module based on a variational autoencoder, which models the structured representation of the code as a probability distribution through variational reasoning. Specifically, the structural agent module obtains code features from the code pre-training model , and then generate latent variables through the encoder of the structural agent module , the decoder from The original code is reconstructed in the latent space to ensure that the reconstructed code retains as much of the original structure and semantic features as possible. The following are the specific implementation steps:
[0046]
[0047]
[0048] in, represents the conditional probability distribution, indicating that given the input Latent variables Distribution of represents normal distribution; Represents latent variables The mean of and input Decide; Represents latent variables The variance of and input Decide; represents the conditional probability distribution, which means that given the latent variable Input Distribution of Represents latent variables The mean of and Decide; Represents latent variables The variance of and input Decide; and Represents the parameters of the encoder and decoder in the structure proxy module; represents the reconstruction loss, represents the KL divergence.
[0049] Compared with previous methods, the structural agent module makes the latent variables Can be trained as latent structured features, using latent variables As a structural hint vector (hint vector containing information about the code structure).
[0050] 2) Fusion embedding generation: In order to provide the necessary background knowledge for LLM during the decoding process, code snippets are extracted from the mapping module Simulation Knowledge Representation , and then input it into the projector to get the knowledge hint vector . After that, replace the code snippet Input into LLM to get text embedding .Will , , These three elements are combined and regularized to solve the negative impact of distribution differences on the model. The specific steps are as follows:
[0051] in, represents the KL loss, represents the KL divergence.
[0052] Knowledge Tips Vector After regularization, we get the regularized knowledge hint vector , After regularization, we get the regularized structure hint vector , , is effectively used to precisely influence the output of LLM. In addition, by comparing the difference between the generated code summary and the actual summary, the loss minimization process is guided, and the loss function is modeled as:
[0053] in, represents the difference generation loss, represents the trainable parameters of the model, represents the number of tokens in the vocabulary, Indicates The probability of predicting the label, Indicates each The probability of the true label is .
[0054] 3) Joint Optimization: In this invention, a two-stage independent training strategy is adopted. In the first stage, we focus on optimizing the left encoder and right encoder in the mapping module and virtual labeling, while freezing CodePLM to ensure stable vector output. The total loss in the first stage includes the sum of contrast estimation loss, cross entropy loss, and bidirectional matching loss. In the second stage, in order to generate code summaries, LLM is set to a frozen state, and the mapping module, projector, and structure proxy module are further optimized. The joint optimization loss function of this stage is defined as follows:
[0055] in, , and Represents the corresponding weight of the loss, which is used to balance the impact of each loss on model training. In this embodiment, they are set to 1.0, 0.5 and 0.1 respectively.
[0056] The present invention significantly improves the accuracy and efficiency of code summary generation by combining code structure information and natural language background knowledge. Especially under low-resource conditions, the enhanced prompt learning method proposed in the present invention can effectively reduce training time and resource consumption. Compared with the existing discrete prompt learning method, the summary generated by the present invention is more accurate and has stronger context relevance, and is suitable for scenarios such as automated code annotation, code understanding and code retrieval.
[0057] Embodiment 2 This embodiment discloses a source code summary generation system based on an enhanced prompt learning framework, including: A code acquisition module is configured to: acquire code snippets; A summary generation module is configured to: input the code snippet into an enhanced prompt learning framework for processing, and generate a code summary in a natural language form; The enhanced prompt learning framework includes a code pre-training model, a mapping module, a structure proxy module and a large language model; the code pre-training model receives code snippets and performs feature extraction to obtain code features, and inputs the code features into the mapping module to generate a knowledge prompt vector; at the same time, the code features are input into the structure proxy module to generate a structure prompt vector; the knowledge prompt vector and the structure prompt vector are spliced and input into the large language model as prompt information to generate a code summary in natural language form.
[0058] Embodiment 3 The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method of embodiment 1 when executing the program.
[0059] Embodiment 4 The purpose of this embodiment is to provide a computer-readable storage medium, a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the method of embodiment 1 are performed.
[0060] The steps involved in the apparatus of the above embodiments 3 and 4 correspond to the method embodiment 1, and the specific implementation method can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0061] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0062] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
[0063] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A source code summary generation method based on an enhanced hint learning framework, characterized in that: include: Obtain code snippets, input them into the enhanced prompt learning framework for processing, and generate code summaries in natural language form; The enhanced prompt learning framework includes a code pre-training model, a mapping module, a structure proxy module and a large language model; the code pre-training model receives code snippets and performs feature extraction to obtain code features, and inputs the code features into the mapping module to generate a knowledge prompt vector; at the same time, the code features are input into the structure proxy module to generate a structure prompt vector; the knowledge prompt vector and the structure prompt vector are spliced and input into the large language model as prompt information to generate a code summary in natural language form.
2. A source code summary generation method based on an enhanced prompt learning framework as claimed in claim 1, characterized in that: The mapping module includes a left encoder and a right encoder, the left encoder includes a self-attention layer and a feedforward layer connected in sequence, and the right encoder includes a self-attention layer, a cross-attention layer and a feedforward layer connected in sequence.
3. A source code summary generation method based on an enhanced hint learning framework as claimed in claim 2, characterized in that: The specific steps of the mapping module generating the knowledge hint vector are: When training the mapping module, the background knowledge is input into the left encoder for processing to obtain a high-order representation of the background knowledge; the virtual tag is input into the right encoder, and the code feature is input into the cross-attention layer of the right encoder. After the virtual tag and the code feature interact in the cross-attention layer, they are sent to the forward feedback layer to obtain a simulated knowledge representation, and finally a trained mapping module is obtained; The trained mapping module only retains the right encoder, inputs the virtual label into the right encoder, and inputs the code features into the cross-attention layer of the right encoder to obtain the simulated knowledge representation and input it into the projector for feature mapping and dimension alignment to obtain the knowledge prompt vector.
4. A source code summary generation method based on an enhanced prompt learning framework as claimed in claim 2, characterized in that: When the mapping module is trained and optimized, the high-order representation of background knowledge and the simulated knowledge representation are used to design the contrast estimation loss, cross entropy loss, and bidirectional matching loss respectively, and the total loss is the sum of the contrast estimation loss, the cross entropy loss, and the bidirectional matching loss.
5. The method for generating source code summaries based on an enhanced hint learning framework according to claim 1, characterized in that: The specific steps of the structure proxy module generating a structure hint vector through a variational autoencoder are as follows: encoding the code features through the encoder of the variational autoencoder, and optimizing the variational autoencoder using reconstruction loss in the decoder so that the variational autoencoder captures the structure hint vector.
6. A source code summary generation method based on an enhanced hint learning framework as claimed in claim 1, characterized in that: The knowledge hint vector and the structure hint vector are concatenated as follows: a code snippet is input into an embedding layer of a large language model to obtain text embedding, and the text embedding, the knowledge hint vector and the structure hint vector are concatenated and merged, and then input into the large language model.
7. A source code summary generation method based on an enhanced hint learning framework as claimed in claim 6, characterized in that: The large language model is frozen for training optimization, and its joint optimization loss is the weighted sum of reconstruction loss, KL loss, and difference generation loss.
8. A source code summary generation system based on an enhanced hint learning framework, characterized in that: include: A code acquisition module is configured to: acquire code snippets; A summary generation module is configured to: input the code snippet into an enhanced prompt learning framework for processing, and generate a code summary in a natural language form; The enhanced prompt learning framework includes a code pre-training model, a mapping module, a structure proxy module and a large language model; the code pre-training model receives code snippets and performs feature extraction to obtain code features, and inputs the code features into the mapping module to generate a knowledge prompt vector; at the same time, the code features are input into the structure proxy module to generate a structure prompt vector; the knowledge prompt vector and the structure prompt vector are spliced and input into the large language model as prompt information to generate a code summary in natural language form.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in a method for generating a source code summary based on an enhanced prompt learning framework as described in any one of claims 1 to 7 are implemented.
10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for generating a source code summary based on an enhanced prompt learning framework according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Label generation method and device of electronic book and electronic equipment
CN114398854A
Few-sample abstract generation method based on pre-training soft hint
CN114647723A
Code abstract automatic generation method based on knowledge enhancement pre-training model
CN117193848A
Code abstract generation method, system and equipment based on improved Transform model
CN118963825A
Program translation through code distillation
WO2024259558A1