Method and system for generating submitted information based on language model and retrieval enhancement

By combining language model and search enhancement methods, using hybrid search technology and pre-trained language model, the problems of insufficient generation quality and poor interpretability of existing submitted information generation methods are solved, and high-quality and consistent submitted information generation are achieved.

CN120010908APending Publication Date: 2025-05-16WUHAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510022975.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing submission information generation methods have problems with insufficient generation quality and poor interpretability, and rely too much on predefined rules or existing submission information databases, making it difficult to generate appropriate descriptions of novel code changes.

Method used

Using a language model and search enhancement method, the most relevant code difference-submit information pairs are obtained from large-scale source databases through mixed search technology, and combined with the code difference of the submitted information to be generated, and the pre-trained language model is used to generate the submitted information.

Benefits of technology

It significantly improves the quality of the submitted information generation, enhances the interpretability and consistency of the generated results, and exceeds the performance of existing benchmark methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010908A_ABST
    Figure CN120010908A_ABST
Patent Text Reader

Abstract

The invention discloses a submitted information generation method and system based on a language model and retrieval enhancement, and belongs to the technical field of text generation, and the method comprises the steps: carrying out retrieval in a source database according to the obtained code difference of to-be-generated submitted information, and obtaining a code difference-submitted information pair most related to the code difference of the to-be-generated submitted information; based on the most relevant code difference-submission information pair and the code difference of to-be-generated submission information, a sequence is combined, and a specific mark or a prompt template is adopted to enhance the combined sequence; and inputting the enhanced sequence into a generator constructed based on a language model to generate submission information. According to the method, retrieval is carried out through a retrieval enhancement method, the feasibility in the submitted information generation task is verified, and in addition, the accuracy and consistency of submitted information generation are remarkably improved in combination with the powerful generation capacity of the pre-training language model or the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of text generation, and in particular relates to a method and system for generating submission information based on a language model and retrieval enhancement (REtrieval-Augmented CommiT message generation framework, referred to as REACT framework). Background Art

[0002] A commit message is a brief text used to describe the content of each code change in a version control system. It is usually used to record the reasons for code changes and the specific modifications, helping developers understand code history, track errors, and collaborate on development. Therefore, commit messages play a key role in the software development and maintenance process, and can improve the readability of the code and the maintainability of the project. A clear and informative commit message can enable other developers to quickly understand the purpose of the change, thereby reducing communication costs and improving collaboration efficiency. However, writing high-quality commit messages is often time-consuming and subjective, so the automatic generation of commit messages has gradually become a research focus.

[0003] Commit message generation (CMG) refers to the use of technical means to automatically generate commit messages to describe the content and purpose of code changes. With the widespread use of version control systems in software development, the importance of commit messages has become increasingly prominent. However, manually writing high-quality commit messages is usually time-consuming, labor-intensive and subjective. Many developers may not be willing to spend time to record in detail or only provide simple, uninformative descriptions, which may lead to difficulties in subsequent maintenance and collaboration. To solve this problem, commit message generation technology has emerged. In recent years, researchers have proposed many methods for automatic generation of commit messages, including three categories: retrieval-based generation methods, machine learning-based generation methods, and hybrid generation methods. In the field of commit message generation, Buse et al. proposed a rule-based method for automatic document generation of program changes. Shen et al. proposed an automatic summary method for what and why information in source code changes. Linares-Vásquez et al. developed a tool called ChangeScribe to automatically generate commit messages. Among the retrieval-based methods, Liu et al. proposed NNGen, which uses information retrieval technology to recommend commit messages from similar code differences. Hoang et al. proposed CC2Vec to generate commit messages by learning representations of code differences. Among learning-based methods, Jiang et al. first proposed CommitGen, which uses an encoder-decoder model of a recurrent neural network for commit message generation. Xu et al. proposed CoDiSum, which uses multi-layer bidirectional gated recurrent units as encoders to better learn representations of code changes. Liu et al. proposed PtrGNCMsg, based on an improved sequence-to-sequence model and a pointer generation network. Dong et al. proposed FIRA, which first used fine-grained graphs to represent code differences. Among hybrid methods, Liu et al. proposed ATOM, which combines abstract syntax trees and hybrid sorting for commit message generation. Wang et al. proposed CoRec, which uses information retrieval and neural machine translation techniques to solve the problems of low-frequency words and exposure bias. Shi et al. proposed RACE, which combines retrieval and generation techniques in a more integrated way.

[0004] In terms of pre-trained language models, Ahmad et al. proposed PLBART, a unified pre-trained model for program understanding and generation. Feng et al. proposed CodeBERT, a pre-trained model for programming and natural language. Guo et al. proposed UniXcoder for unified cross-modal pre-training of code representation. Wang et al. proposed CodeT5, a unified pre-trained encoder-decoder model with identifier-aware pre-training objectives.

[0005] However, the existing methods have at least the following technical problems: Rule-based methods, such as the automatic document generation method proposed by Buse et al. and the ChangeScribe tool developed by Linares-Vásquez et al., rely too much on predefined rules and lack flexibility in the generated results. Retrieval-based methods, such as NNGen proposed by Liu et al. and CC2Vec proposed by Hoang et al., rely heavily on existing commit information repositories and have difficulty generating appropriate descriptions for novel code changes. Learning-based methods, such as CommitGen proposed by Jiang et al. and CoDiSum proposed by Xu et al., introduce deep learning technology, but due to the model training from scratch, they fail to fully utilize the rich knowledge contained in the pre-trained model, and the generation quality still has room for improvement. In addition, existing methods often regard commit information generation as a single sequence-to-sequence conversion task, ignoring the role of retrieval enhancement in improving generation quality.

[0006] Therefore, it is necessary to design a submission information generation method and system based on language model and retrieval enhancement to address the above problems. Summary of the invention

[0007] The purpose of the present invention is to solve the problems of insufficient generation quality and poor interpretability commonly existing in existing submission information generation methods. A method and system for automatically generating submission information based on language model and retrieval enhancement is provided. The most relevant code difference-submission information pairs are obtained from a large-scale source database through hybrid retrieval technology, and are combined with the code differences of the submission information to be generated into a sequence for enhancement. A generator is constructed in combination with a pre-trained language model or a large language model to generate submission information, which significantly improves the generation quality of the submission information.

[0008] According to one aspect of the present specification, a method for generating submission information based on a language model and retrieval enhancement is provided, comprising:

[0009] Searching in a source database according to the obtained code differences of the submission information to be generated, obtaining a code difference-submission information pair most relevant to the code differences of the submission information to be generated;

[0010] Based on the most relevant code difference-commit information pair, the code difference of the commit information to be generated is combined into a sequence, and a specific mark or prompt template is used to enhance the combined sequence;

[0011] The enhanced sequence is input into the generator built based on the language model to generate the submission information.

[0012] Furthermore, obtaining the most relevant code difference-commit information pair to be generated also includes:

[0013] BM25 (Best Match 25) is used to calculate the textual relevance of code differences, and a deep encoder is used to calculate the semantic similarity of code differences;

[0014] The calculation results of BM25 and the deep encoder are weighted fused, and the difference-commit information pair with the highest fusion score is selected as the most relevant code difference-commit information pair; the formula for the weighted fusion is:

[0015]

[0016] in, is the BM25 score weight, is the encoder similarity weight, is the cosine similarity.

[0017] Furthermore, the construction of the generator includes:

[0018] Train based on a pre-trained code language model or a general large language model to obtain a trained generator.

[0019] Furthermore, the training of the generator further includes:

[0020] Fine-tune the pre-trained code language model or the general large language model, and use the optimizer to optimize the parameters of the language model.

[0021] Furthermore, the method also includes: generating submission information using an autoregressive generation method.

[0022] According to one aspect of the present specification, a submission information generation system based on a language model and a retrieval enhancement method is provided, comprising:

[0023] A retrieval module, used for searching in a source database according to the obtained code differences of the submission information to be generated, and obtaining the code difference-submission information pair most relevant to the code differences of the submission information to be generated;

[0024] An enhancement module, configured to combine the most relevant code difference-commit information pair with the code difference of the commit information to be generated into a sequence, and enhance the combined sequence by using a specific mark or prompt template;

[0025] The generation module is used to input the enhanced sequence into the generator built based on the language model to generate submission information.

[0026] According to one aspect of the present specification, there is provided an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of the method for generating submission information based on a language model and a retrieval enhancement method when executing the computer program.

[0027] According to one aspect of the present specification, there is provided a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the submission information generation method based on the language model and the retrieval enhancement method are implemented.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] 1. The present invention enhances the effect of retrieval enhancement on the generation quality through hybrid retrieval technology. The experimental results show that the hybrid retrieval method adopted by the REACT framework has achieved significant improvement effects, verifying the feasibility of the retrieval enhancement method in submitting information generation tasks.

[0030] 2. In order to solve the problem of insufficient performance of existing submission information generation methods, the present invention significantly improves the accuracy and consistency of submission information generation by combining the powerful generation capabilities of pre-trained language models or large language models.

[0031] 3. By integrating the existing language model into the REACT framework, the present invention improves the BLEU score of CodeT5 on the test set by 55%, and the BLEU score of Llama3 by 102%, which greatly enhances the performance of the language model in the submission information generation task, surpassing the existing benchmark method, and has applicability and versatility. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0033] Figure 1 A flowchart of a method for generating submission information based on a language model and a retrieval enhancement method according to an embodiment of the present invention;

[0034] Figure 2 A schematic diagram of generating a REACT enhanced Llama 3 model according to an embodiment of the present invention;

[0035] Figure 3 It is a module diagram of a submission information generation system based on a language model and a retrieval enhancement method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0036] It should be noted that:

[0037] A code discrepancy refers to the differences between two or more code versions;

[0038] Commit information refers to a brief text description of the submitted code differences by the developer during the code submission process;

[0039] Code difference-commit information pair refers to the one-to-one correspondence between code differences and their corresponding commit information;

[0040] A large-scale source database refers to a database containing multiple code difference-commit information pairs;

[0041] The specific tags refer to [QUERY], [DIFF], and [MSG], which are used to identify the code difference for which the commit message is to be generated, the most relevant code difference retrieved, and the corresponding commit message, respectively.

[0042] A hint template is a predefined text snippet or structure that is used to combine the generated code diff with the retrieved submissions and guide the large language model to perform specific actions.

[0043] The technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all of the embodiments. Based on the embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0044] The embodiment of the present invention provides a method for generating submission information based on a language model and retrieval enhancement, such as Figure 1 As shown, it includes: obtaining the code difference of the submission information to be generated, retrieving the code difference-submission information pair that is most relevant to the code difference of the submission information to be generated from the source database; combining the code difference of the submission information to be generated into a sequence based on the most relevant code difference-submission information pair, and enhancing it with a specific tag or prompt template; building a generator based on the language model, inputting the enhanced sequence into the generator, and generating the submission information.

[0045] Specifically, the embodiment of the present invention further provides specific steps for hybrid retrieval based on BM25 and a deep encoder, as follows:

[0046] 1. Collect code difference text and corresponding commit information from public code repositories. Select popular projects (high-star projects) and filter out invalid or low-quality commit information, such as meaningless information such as "init commit" or "update". At the same time, filter out commit information that is too long or too short with a length of less than 8 or greater than 200 to ensure the integrity and high quality of the source database. Extract semantic and structural features based on code difference text. Match the code difference text with the commit information one by one, and use the pre-trained model encoder CodeT5+ to encode the code difference into a fixed-length dense vector, expressed as:

[0047] (1)

[0048] in, It is a difference vector representation (the dimension is fixed to 256), which is used to calculate the semantic similarity of the difference; For code differences.

[0049] 2. Based on the query code difference text, use the BM25 algorithm to build a term frequency-inverse document frequency (TF-IDF) index to capture the correlation at the text level and store the difference dense vector generated by the encoder. Use the BM25 algorithm to calculate the BM25 scores of all code differences in the source database and filter out the candidate set with the most term frequency correlation. The candidate set is defined as The similarity score formula of BM25 is:

[0050] (2)

[0051] in, represents the query (the text of the code diff), represents a document (source database), is the inverse document frequency, represents the term frequency, parameter and Used to control the balance between document length and frequency.

[0052] for For each code difference in , the dense vector representation generated by the deep encoder is used to calculate and query the vector The cosine similarity calculation formula is as follows:

[0053] (3)

[0054] in, is a dense vector representation of the code differences to generate the commit information, A dense vector representing each candidate difference.

[0055] 3. Weighted fusion of BM25 score and cosine similarity score to obtain the final mixed score . Assume that the BM25 score weight is , the encoder similarity weight is , then the mixed fraction formula is:

[0056] (4)

[0057] in, is the BM25 score weight, is the encoder similarity weight, is the cosine similarity.

[0058] According to the mixed score results, the difference-commit information pair with the highest score is selected as the most relevant code difference-commit information pair.

[0059] Specifically, the vectors generated by the deep encoder and the BM25 scores are uniformly stored in the source database to ensure that each record includes a dense vector representation of the code difference and the BM25 score. Persistent storage helps reduce the time consumption of subsequent query operations and improve the efficiency of the entire system.

[0060] Specifically, the embodiment of the present invention further provides specific steps for enhancing the combined sequence, as follows:

[0061] 1. Combine the code difference of the commit message to be generated (Query Diff), the most relevant code difference retrieved (Retrieved Diff), and the most relevant commit message retrieved (Retrieved Message) into an input sequence. In order to allow the generator to distinguish the various parts of the input, specific tags ([QUERY], [DIFF], [MSG]) are introduced, which are used to identify the code difference of the commit message to be generated, the retrieved code difference and its corresponding commit message. This approach ensures that the model can understand the input structure and identify the role of each part. The format of the input is expressed as:

[0062] (5)

[0063] 2. Design prompt word templates for large language models, and fill the combined input sequences with the designed templates so that the model can perform context learning and prompt learning.

[0064] 3. Combine the code differences of the commit information to be generated and the retrieved difference-commit information pairs into a sequence, structure the query differences, retrieved differences and commit information through tags such as [QUERY], [DIFF], [MSG], and select the appropriate enhancement method based on the target model.

[0065] Specifically, the embodiment of the present invention also provides specific steps for constructing a generator based on a language model, as follows:

[0066] 1. Use a pre-trained code language model or a general large language model as the generator. For the code language model, fine-tune the parameters; for the large language model, use the prompt learning method to directly use its context learning ability to complete the generation task. The purpose of fine-tuning is to train the pre-trained model with enhanced input with examples, so that it can better understand the input structure and generate submission information that meets the requirements. Fine-tuning is mainly based on the sequence after input enhancement. And the corresponding target output (i.e. submit information).

[0067] Construct input-output pairs, input : The enhanced input sequence generated by the enhancer, the target output : The actual commit message written by the developer;

[0068] (6)

[0069] Loss function, the model is based on the input Generate target output step by step The probability distribution of , optimizes the model parameters by minimizing the cross entropy loss. The cross entropy loss function formula is:

[0070] (7)

[0071] in, is the enhanced combined data of the input, To submit information, represents the length of the generated sequence, Represents the model at time step Generate the correct word The probability of and previously generated words , Represents the parameters of the model, which are continuously updated through fine-tuning.

[0072] Update the model parameters using the Adam optimizer to minimize the cross entropy loss The update formula of the optimization process is:

[0073] (8)

[0074] in, is the learning rate, is the loss function over parameters Through fine-tuning, the model can more accurately predict the submission information at each time step, improving the generation effect.

[0075] 2. Submission information generation stage: An autoregressive generation method is used, that is, in the process of gradually generating each word, the previously generated word is used as input to predict the next word until a complete submission information is generated.

[0076] Given enhanced input , the model will gradually generate each word of the submission information according to the autoregressive generation method. The generation process is as follows: the initial input is the enhanced sequence In the time steps, the model generates words , its conditional probability is:

[0077] (9)

[0078] in, For the generated word, the model continues to predict the word of the next time step , until a terminator is generated (such as <eos>) or reaches the set maximum length.

[0079] Generate final submission information The joint probability is:

[0080] (10)

[0081] In the submission information generation stage, a beam search strategy is used to form the final output sequence, and the model temperature is set to 0.5.

[0082] The embodiment of the present invention also provides experimental results of a submission information generation method based on a language model and retrieval enhancement, wherein the REACT framework can effectively enhance the generation capability of the existing model, and the BLEU score after integrating CodeT5 into the REACT framework is increased by 55%, and the BLEU score after integrating Llama 3 into the REACT framework is increased by 102%, which greatly exceeds the existing benchmark method and has wide applicability and versatility. The REACT framework can be integrated with different types of code language models (including pre-trained language models and large language models), and the model performance is enhanced by fine-tuning or contextual learning, which verifies the universality of the framework and provides practical value for scenarios within the project. In the case study of the Electron project, REACT can generate content that meets the writing specifications of project-specific submission information by retrieving the historical submission information of the project as an example, which significantly improves the practicality of the generated results and the effectiveness of hybrid retrieval. The experimental results show that the hybrid retrieval method adopted by REACT can effectively retrieve to guide the generation of submission information, which is 60% higher than random retrieval in BLEU score, verifying the feasibility of the retrieval enhancement method in the submission information generation task. The REACT framework proposed in the embodiment of the present invention can not only significantly improve the performance of submitting information generation, but also has good versatility and practical value, and provides an effective solution for the task of submitting information generation.

[0083] The embodiment of the present invention provides an example of using REACT to enhance the performance of Llama 3 model generation, such as Figure 2 As shown in the figure, BLEU is an effective measure of translation accuracy for the accuracy of generated submission information and is widely used in related work. It measures the similarity between the reference submission information and the submission information generated by the method. The larger the BLEU value, the better the generation performance. The expression is as follows:

[0084] (11)

[0085] in, is the length penalty factor, To calculate the matching ratio for each n, is the weight, () is the weighted geometric mean.

[0086] ROUGE-L is a recall evaluation method based on the longest common subsequence, which is specifically used to measure the similarity between the generated text and the reference text. In the submission information generation task, the fluency and completeness of the generated content are evaluated by calculating the longest common subsequence between the generated submission information and the reference submission information. The higher the ROUGE-L value, the better the quality of the generated submission information. The expression is as follows:

[0087] , , (12)

[0088] in, is the longest common subsequence, is the accuracy, is the recall rate, is the harmonic mean of recall and precision, as a parameter.

[0089] METEOR is an alignment-based evaluation metric that takes into account synonym matching and can better capture semantic similarity. In the task of automatically generating submission information, the semantic consistency between the generated submission information and the reference submission information is comprehensively evaluated through word sequence matching, synonym recognition, etc. The higher the METEOR score, the better the semantic accuracy of the generated result. The expression is as follows:

[0090] (13)

[0091] in, is the harmonic mean of precision and recall, For penalty points.

[0092] The automatic generation method of submission information proposed in the present invention is compared with other traditional methods, and the comparison results of the following three evaluation indicators are obtained, as shown in Table 1.

[0093] Table 1 Comparison of evaluation indicators of various methods

[0094]

[0095] As can be seen from Table 1, in terms of various evaluation indicators, the performance of the proposed REACT combined with the pre-trained language model CodeT5 surpasses other traditional methods. As shown in Table 2, the experimental results of combining REACT with different models are compared.

[0096] Table 2 Comparison of evaluation indicators of REACT combined with different models

[0097]

[0098] It can be seen from the results in Table 1 and Table 2 that the method proposed in the embodiment of the present invention significantly improves the performance of submission information generation. The BLEU score after integrating CodeT5 into the REACT framework is improved by 55%. In addition, it has good versatility and can be integrated with different types of code language models.

[0099] The implementation basis of each embodiment of the present invention is to implement programmed processing through a device with a processor function. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this reality, on the basis of the above embodiments, an embodiment of the present invention provides a submission information generation system based on a language model and retrieval enhancement, which is used to execute the submission information generation method based on a language model and retrieval enhancement in the above method embodiment.

[0100] See also Figure 3 The system includes: a retrieval module, which is used to search in a source database according to the obtained code differences of the submission information to be generated, and obtain the code difference-submission information pair that is most relevant to the code difference of the submission information to be generated; an enhancement module, which is used to combine the code differences of the submission information to be generated into a sequence based on the most relevant code difference-submission information pair, and enhance the combined sequence by using a specific tag or prompt template; and a generation module, which is used to input the enhanced sequence into a generator built based on a language model to generate submission information.

[0101] The system for generating submission information based on language model and retrieval enhancement provided by the embodiment of the present invention solves the problem of insufficient performance of the existing submission information generation method. Figure 3 Several modules in it, by combining the powerful generation capabilities of pre-trained language models or large language models, significantly improve the accuracy and consistency of submission information generation, overcome the problems of missing information and insufficient consistency in the traditional model generation process, are suitable for code management scenarios of different projects, and significantly improve development efficiency.

[0102] Based on the same inventive concept as the aforementioned embodiment, an embodiment of the present invention further provides an electronic device, including a memory and a processor, the memory being used to store computer-executable instructions, and the processor being used to execute computer-executable instructions, to implement a submission information generation method based on a language model and retrieval enhancement as proposed in the aforementioned embodiment.

[0103] The embodiment of the present invention also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by the processor, it overcomes the problems of missing information and insufficient consistency in the traditional model generation process, strengthens the effect of retrieval enhancement on the improvement of generation quality, is applicable to code management scenarios of different projects, and significantly improves development efficiency. The storage medium can be any non-volatile storage device such as a hard disk, a solid-state drive, a flash drive, an optical disk, etc., for storing computer program code and necessary data files, and the stored computer program includes: a retrieval module, an enhancement module, and a generation module.

[0104] Finally, it should be pointed out that the above specific embodiments are only representative examples of the present invention. Obviously, the present invention is not limited to the above specific embodiments, and there are many variations. Any simple modification, equivalent changes and modifications made to the above specific embodiments based on the technical essence of the present invention should be considered to belong to the protection scope of the present invention.< / eos>

Claims

1. A submission information generation method based on language model and retrieval enhancement, characterized in that: include: Searching in a source database according to the obtained code differences of the submission information to be generated, obtaining a code difference-submission information pair most relevant to the code differences of the submission information to be generated; Based on the most relevant code difference-commit information pair, the code difference of the commit information to be generated is combined into a sequence, and a specific mark or prompt template is used to enhance the combined sequence; The enhanced sequence is input into the generator built based on the language model to generate the submission information.

2. A method for generating submission information based on language model and retrieval enhancement according to claim 1, characterized in that: The obtaining of the most relevant code difference-commit information pair to be generated also includes: BM25 is used to calculate the textual relevance of code differences, and a deep encoder is used to calculate the semantic similarity of code differences; The calculation results of BM25 and the deep encoder are weighted fused, and the difference-commit information pair with the highest fusion score is selected as the most relevant code difference-commit information pair.

3. The method for generating submission information based on language model and retrieval enhancement according to claim 1, characterized in that: The construction of the generator includes: Train based on a pre-trained code language model or a general large language model to obtain a trained generator.

4. The method for generating submission information based on language model and retrieval enhancement according to claim 3, characterized in that: The training of the generator further includes: Fine-tune the pre-trained code language model or the general large language model, and use the optimizer to optimize the parameters of the language model.

5. The method for generating submission information based on language model and retrieval enhancement according to claim 1, characterized in that: The method further comprises: generating submission information by adopting an autoregressive generation method.

6. A submission information generation system based on language model and retrieval enhancement, characterized in that: include: A retrieval module, used for searching in a source database according to the obtained code differences of the submission information to be generated, and obtaining the code difference-submission information pair most relevant to the code differences of the submission information to be generated; An enhancement module, configured to combine the most relevant code difference-commit information pair with the code difference of the commit information to be generated into a sequence, and enhance the combined sequence by using a specific mark or prompt template; The generation module is used to input the enhanced sequence into the generator built based on the language model to generate submission information.

7. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for generating submitted information based on language model and retrieval enhancement according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for generating submitted information based on a language model and retrieval enhancement according to any one of claims 1 to 5 are implemented.

Citation Information

Cited By

  • Library migration recommendation method based on retrieval enhancement generation

    CN120950480A