Method and device for finely adjusting code model

By fine-tuning the code model, using the source code and patch code in the code library for similarity search and prompt text construction, the problem of insufficient performance of the code model in the existing technology in program repair tasks is solved, and higher inference accuracy and generalization capabilities are achieved.

CN119988198AActive Publication Date: 2025-05-13ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510088841.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The performance and generalization capabilities of existing automatic program repair methods in program repair tasks are still insufficient, and it is difficult to effectively improve the execution efficiency and accuracy of the code model in program repair tasks.

Method used

By fine-tuning the code model, using the source code and patch code in the code base, the characterization vector is generated for similarity search, the target context is obtained, and the patch code is inferred by building a prompt text guide code model, and finally fine-tuning the code model parameters.

Benefits of technology

It improves the code model's understanding of code semantics and structure, enhances the inference accuracy and generalization ability in program repair tasks, and optimizes the stability of the code model in the face of diverse programming languages ​​and code styles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988198A_ABST
    Figure CN119988198A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method for finely adjusting a code model, which is carried out on the basis of a code library, and the code library comprises a plurality of source codes and corresponding patch codes. The method comprises the following steps: acquiring a first representation vector corresponding to any first source code in a plurality of source codes, wherein the first representation vector is determined by performing first coding on a text of the first source code and performing second coding on a first abstract syntax tree corresponding to the first source code; and according to the first representation vector, performing similarity-based retrieval in a code library to obtain a target context. The target context comprises a plurality of other source codes and patch codes corresponding to the source codes, wherein the similarity between the source codes and the first source code meets a first threshold value. And generating a prompt text according to the first source code and the target context, and inputting the code model to obtain an inference patch. According to a first patch code, corresponding to the first source code, of the inference patch, parameters of the code model are finely adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present specification relate to the technical field of large language models and code generation, and more particularly, to a method and apparatus for fine-tuning a code model. Background Art

[0002] With the rapid development of the software industry, in the process of software development and maintenance, program repair (such as bug fixes, code error correction, etc.) takes up more and more time, squeezing the effective development time in software engineering projects and consuming a lot of manpower costs, which has become an increasingly serious problem. In order to cope with the increasingly heavy program repair tasks, the Automatic Program Repair (APR) method came into being, aiming to reduce the time and human resources consumed by program repair tasks in software engineering. Automatic program repair methods mainly include: heuristic-based methods, template-based methods, semantic-based methods, and deep learning-based methods. These methods are either limited by the cumbersome rule design process based on expert experience, or limited by the types of code problems that can be repaired, or limited by model parameters and training data quality. The performance and generalization ability of program repair tasks are still insufficient.

[0003] Therefore, it is hoped that there will be a solution that can improve the execution efficiency and accuracy of the code model in program repair tasks through technical means. Summary of the invention

[0004] One or more embodiments of the present specification describe a method and apparatus for fine-tuning a code model to improve the code model's ability to understand a target source code, thereby better completing a program repair task.

[0005] According to a first aspect, a method for fine-tuning a code model is provided, which is performed based on a code library, wherein the code library includes a plurality of source codes and corresponding patch codes, and the method comprises:

[0006] A first representation vector corresponding to any first source code among the plurality of source codes is obtained, where the first representation vector is determined by performing a first encoding on a text of the first source code and performing a second encoding on a first abstract syntax tree corresponding to the first source code.

[0007] According to the first representation vector, a similarity-based search is performed in the code library to obtain a target context; the target context includes a number of other source codes whose similarity to the first source code meets a first threshold and their corresponding patch codes.

[0008] According to the first source code and the target context, a prompt text is generated, and a code model is input to obtain an inferred patch.

[0009] Fine-tune the parameters of the code model according to the first patch code corresponding to the inferred patch and the first source code.

[0010] According to one implementation, the fine-tuned code model is used to infer the corresponding patch code for the source code to be checked based on the target prompt text; the target prompt text includes the source code to be checked, a number of reference source codes whose similarity with the source code to be checked meets a first threshold and their corresponding patch codes.

[0011] According to one implementation, the first encoding includes: performing text encoding on the first source code using a first text model.

[0012] According to one implementation, the second encoding includes: performing a serialization operation on the first abstract syntax tree to obtain a converted text sequence, wherein the converted text sequence includes node identifiers representing each node of the first abstract syntax tree and structure labels representing the structural relationship between the nodes; and performing text encoding on the converted text sequence using a second text model.

[0013] In one scenario of the above implementation, the structure tag includes a tag pair marked with a node identifier, and the tag pair is used to indicate a child node of a node corresponding to the marked node identifier; and the serialization operation includes:

[0014] During the target traversal of the first abstract syntax tree, for any target node currently visited, its node identifier is extracted, and the node identifier is added to the first label pair range in the current sequence, where the first label pair is marked with the node identifier of the parent node of the target node; the current sequence obtained at the end of the target traversal is used as the converted text sequence.

[0015] According to an implementation, the first characterization vector is a weighted average of a first vector obtained by the first encoding and a second vector obtained by the second encoding.

[0016] According to one implementation, in the prompt text, the other source codes are arranged from high to low according to their similarity to the first source code, and the number of the other source codes is a maximum number such that the prompt text does not exceed the upper limit of the number of input word units.

[0017] According to an implementation, in the prompt text, each source code is marked with a first tag, and each patch code is marked with a second tag.

[0018] According to an implementation manner, the code model is a large language model, and fine-tuning the parameters of the code model includes: fine-tuning all parameters of the large language model.

[0019] According to a second aspect, there is provided an apparatus for fine-tuning a code model, which is run based on a code library, wherein the code library includes a plurality of source codes and corresponding patch codes; the apparatus comprises:

[0020] The encoding module is configured to obtain a first representation vector corresponding to any first source code among the multiple source codes, where the first representation vector is determined by first encoding the text of the first source code and second encoding the first abstract syntax tree corresponding to the first source code.

[0021] The retrieval module is configured to perform a similarity-based search in the code library according to the first representation vector to obtain a target context; the target context includes a number of other source codes whose similarity to the first source code meets a first threshold and their corresponding patch codes.

[0022] The inference module is configured to generate a prompt text according to the first source code and the target context, input a code model, and obtain an inferred patch.

[0023] The fine-tuning module is configured to fine-tune the parameters of the code model according to the first patch code corresponding to the inferred patch and the first source code.

[0024] According to a third aspect, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0025] According to a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein an executable code is stored in the memory, and when the processor executes the executable code, the method described in the first aspect is implemented.

[0026] In summary, in the method and device provided in the embodiments of this specification, a method for fine-tuning the code model is designed, which can retrieve relevant code snippets based on the semantic representation and structural representation of the source code in the code library, thereby improving the accuracy of context code snippet retrieval. By using the context obtained by the retrieval, a prompt text is constructed to guide the code model to infer the patch code of the training source code. By gradually fine-tuning the parameters of the code model, the code model can output an inference result that is close to the sample label.

[0027] During this fine-tuning process, the code model's ability to understand code semantics and code structure can be enhanced, and the inference accuracy of the code model in performing program repair tasks can be improved; the generalization ability of the code model can be optimized, so that the code model can still maintain stable code repair performance when facing a variety of programming languages ​​and code styles. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0029] Figure 1 A schematic diagram of a framework of a method for fine-tuning a code model disclosed in an embodiment of this specification;

[0030] Figure 2 A flow chart of a method for fine-tuning a code model provided according to an embodiment of this specification;

[0031] Figure 3A A schematic diagram of a serialization operation provided in an embodiment of this specification;

[0032] Figure 3B A schematic diagram of a serialization operation provided in an embodiment of this specification;

[0033] Figure 4 A schematic diagram of a device for fine-tuning a code model according to an embodiment of this specification. DETAILED DESCRIPTION

[0034] The solution provided by the embodiments of this specification is described below in conjunction with the accompanying drawings.

[0035] In one or more embodiments of this specification, the Python programming language is used as an example to explain the process of fine-tuning the code model, but this does not limit the application scenarios and technical tools of the embodiments of the present invention. The technical concepts embodied in each embodiment of this specification can be applied to various code models that can perform program repair tasks and the programming language environments they support.

[0036] As mentioned above, in order to improve the R&D efficiency of software engineering projects, the automatic program repair method APR is widely used. Considering the limitations of traditional APR, the Large Language Model (LLM) is used to perform APR tasks. Compared with traditional deep learning models, LLM has shown certain potential in understanding code and generating patches, thanks to the massive code sample training. A large number of relevant experimental results also show that LLM can not only deeply understand the source code, but also automatically generate the corresponding patch code.

[0037] On this basis, the Retrieval-augmented Generation (RAG) technology can help LLM retrieve and integrate relevant code snippets (context) required in the reasoning process, thereby effectively improving LLM's reasoning ability when handling code repair tasks. Specifically, in the process of LLM inferring the patch code of the source code, RAG technology can first retrieve the context snippets related to the source code from the huge external knowledge base and pass them to LLM; then, LLM can use these context snippets to guide its reasoning of the patch code, thereby improving the quality and accuracy of LLM's inferred patch code.

[0038] However, in some related technologies, there are still large errors when using RAG technology to assist LLM in inferring patch codes. The reason is that in the context code snippets used for prompt engineering and model fine-tuning, key code features are usually lacking, which makes it difficult for large language models to accurately capture the deep meaning of source code files. Through research, the inventors found that in the process of retrieving context code snippets, related technologies often ignore the analysis of the code structure and grammatical features of source code files, resulting in low matching between context retrieval results and source code features, or containing a large amount of redundant information, or lacking key context snippets. These factors affect the fine-tuning effect of the large language model and its performance in APR tasks.

[0039] In order to solve the above technical problems, the inventors proposed a method for fine-tuning the code model, which can comprehensively consider the semantic features and structural features of the code during the context retrieval process, retrieve context fragments closely related to the source code to be repaired, so as to assist the code model in inferring the patch code for the source code to be repaired, and continuously fine-tune the code model parameters by comparing with the sample labels of the patch code, so as to gradually improve the code model's ability to understand the source code.

[0040] Figure 1A method framework for fine-tuning a code model in an embodiment is disclosed. In the process of encoding a representation vector, the source code text is first encoded to obtain a coding representation containing semantic information, and the abstract syntax tree (AST) corresponding to the source code is secondly encoded to obtain a coding representation containing code structure information. Based on these coding representations, the representation vector of the source code can be determined accordingly. It is not difficult to understand that the representation vector obtained in this way can not only contain the semantic features of the source code, but also its structural features. Any source code in the code library can be used as the source code to be repaired, and its corresponding patch code can be used as a sample label. Based on the representation vector corresponding to the source code to be repaired, through similarity retrieval, several other source codes related to the source code to be repaired are obtained. These related source codes and their corresponding patch codes together constitute the target context, which is combined with the source code to be repaired as a prompt text, and input into the code model for inference to obtain an inferred patch corresponding to the source code to be repaired. According to the evaluation result (for example, cross entropy loss) between the inferred patch and the sample label, the code model is fine-tuned. Since the extraction and retrieval of code structure features are added in the context retrieval process, the target context can be closer to the characteristics of the source code to be repaired, which promotes the code model to accurately grasp the deep characteristics of the source code and perform reasoning, so as to give full play to the program repair ability of the code model.

[0041] Following the above technical ideas, Figure 2 In FIG. 1 , a flow chart of a method for fine-tuning a code model according to an embodiment of the present specification is shown. It can be understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. Figure 2 In one embodiment, the method is performed based on a code library, and the code library contains multiple source codes and corresponding patch codes; the method includes at least the following steps. S201: Obtain a first representation vector corresponding to any first source code among the multiple source codes, and the first representation vector is determined by first encoding the text of the first source code and second encoding the first abstract syntax tree corresponding to the first source code. S203: According to the first representation vector, perform a similarity-based search in the code library to obtain a target context; the target context includes several other source codes whose similarity to the first source code meets a first threshold and their corresponding patch codes. S205: Generate a prompt text based on the first source code and the target context, input the code model, and obtain an inferred patch. S207: Fine-tune the parameters of the code model based on the inferred patch and the first patch code corresponding to the first source code.

[0042] The specific implementation methods of the above steps will be described in detail below with reference to the accompanying drawings.

[0043] In step S201, a first representation vector corresponding to any first source code among the plurality of source codes is obtained, where the first representation vector is determined by performing a first encoding on a text of the first source code and performing a second encoding on a first abstract syntax tree corresponding to the first source code.

[0044] The multiple source codes contained in the code library can be program codes written in any programming language, for example, C++, Java, Python, etc. In some scenarios, the code library may also contain source codes written in multiple programming languages. In the embodiments of this specification, Python programming language will be used as an example for explanation. However, it should be understood that one or more embodiments in this specification are intended to provide a method for fine-tuning a code model, and are not limited to a specific programming language. All scenarios related to the technical concepts provided by the embodiments of the present invention can apply the methods provided by the embodiments of the present invention.

[0045] First of all, it should be pointed out that there are various code defects in the source code, including but not limited to syntax errors, code dependency problems, logical errors, etc. These defects may cause the source code to be unable to achieve its intended function. The specific types of defects are not limited in the embodiments of this specification. For code defects in the source code, the code library also contains corresponding patch codes. The patch codes can repair the code defects of the source code to ensure that the source code can be interpreted, compiled, and executed. The patch code can be presented in various forms, such as code text, patch package (Patch), configuration statement, etc., which is not specifically limited in the embodiments of this specification.

[0046] In this step, for any source code in the code library (i.e., the first source code), its first characterization vector is obtained. The first characterization vector may be pre-encoded and stored in the code library with the first source code, or may be generated by real-time encoding when the first source code is selected, which is not limited here.

[0047] As mentioned above, the first representation vector can be determined by first encoding the text of the first source code and second encoding the first abstract syntax tree corresponding to the first source code. The result of the first encoding output can represent the code semantic features in the first source code, and the result of the second encoding output can represent the code structural features in the first source code.

[0048] According to one implementation, the first encoding may be to perform text encoding on the first source code using a first text model. The first text model is usually a natural language processing model, which has a strong language understanding capability and can encode the code text of the first source code, extract semantic information from the code, and represent it with a dense vector. For example, the first text model may be a BERT model based on Transformer, or a CodeGPT model pre-trained for code files, and so on.

[0049] The second encoding is performed based on the first abstract syntax tree, and can capture the structured information of the code in the first source code. Specifically, the abstract syntax tree is a data structure that displays each code element and code relationship in the source code file in a tree structure, wherein each node represents a code element in the source code, such as a module, a class declaration, a function definition, etc.; the connection between the nodes reflects the grammatical hierarchy and code relationship between the code elements, such as function calls, variable references, etc. The second encoding converts the grammatical structure and logical relationship in the first source code into a vector representation by encoding the nodes and connecting edges in the first abstract syntax tree, and can capture the potential relationship between the various code elements in the first source code and characterize its deep code structure characteristics. The abstract syntax tree can also be regarded as a special graph structure, and the second encoding can encode the first abstract syntax tree using a graph neural network (GNN). Further, in one embodiment, the second encoding can be encoded using a graph neural network based on an attention mechanism, wherein a higher weight is applied to the key nodes in the first abstract syntax tree, so that these key nodes can be highlighted.

[0050] According to one implementation, the second encoding process may be to first convert the first abstract syntax tree into a text sequence and then perform text encoding on it. In this implementation, the second encoding specifically includes the following processing steps: performing a serialization operation on the first abstract syntax tree to obtain a converted text sequence, which includes node identifiers representing each node of the first abstract syntax tree and structure labels representing the structural relationship between the nodes; and performing text encoding on the converted text sequence using a second text model.

[0051] The conversion text sequence includes not only the node information of the first abstract syntax tree, but also the structural information between the nodes. The serialization operation may be to first instantiate the first abstract syntax tree as a programming object, and then use the object serialization method provided in the programming language (for example, the pickle module in Python) to convert the first abstract syntax tree as a whole into a byte stream as the conversion text sequence.

[0052] According to one implementation, the serialization operation can introduce preset tags in the converted text sequence to mark the parent-child relationship between nodes. As mentioned above, the parent-child relationship between the nodes of the abstract syntax tree (represented as the connection between the nodes) reflects the grammatical hierarchy and code relationship between the code elements. Therefore, it can be understood that marking the parent-child relationship between the abstract syntax tree nodes in the text sequence can achieve the purpose of incorporating the code structure information into the text sequence.

[0053] Specifically, the structural tag includes a tag pair marked with a node identifier, and the tag pair is used to indicate the child node of the node corresponding to the marked node identifier; the serialization operation may include: in the process of target traversal of the first abstract syntax tree, for any target node currently visited, extract its node identifier, add the node identifier to the range of the first tag pair in the current sequence, and the first tag pair is marked with the node identifier of the parent node of the target node; the current sequence obtained at the end of the target traversal is used as the converted text sequence.

[0054] The target traversal can use traversal methods such as breadth-first traversal and depth-first traversal to access each node of the first abstract syntax tree. During the target traversal, for any target node accessed, its node identifier is extracted. The node identifier can be the name of the code element represented by the node, such as a module name, a class name, a function name, etc., or it can be the summary information of the code element represented by the node, such as the definition body of a function, the definition header of a class and its members, etc.

[0055] In order to accurately depict the hierarchical relationship between nodes, when adding the node identifier of the target node to the current sequence, it can be added to the first label pair range corresponding to the node identifier of the parent node of the target node. In this way, the content in each label pair range can reflect the parent-child relationship of the node; in other words, the nested relationship between structural labels can reflect the hierarchical relationship between nodes. It is not difficult to understand that for a single syntax tree, the root node can be used as the node identifier marked by the outermost label pair in the current sequence. Figure 3A This is a schematic diagram of a serialization operation. Each code element of the source code is represented by a node in the abstract syntax tree, and the relationship between the code elements is represented by connecting edges. It should be understood that the abstract syntax tree in the figure is used as an example. For the sake of simplicity, the node identifiers of each node are represented by numbers, and only a three-layer tree is used as an example. In actual applications, the abstract syntax tree usually has a deeper level and more detailed node identifiers. Figure 3A, during the target traversal execution process, any target node visited will be added to the current sequence. In the example, brackets are used as label pairs in the structural labels. Taking any step in the target traversal process as an example, when the target node visited is #1.1.1, in the current sequence, the corresponding label pair of the target node's parent node is (#1.1). Therefore, the node identifier of the target node can be added to the range of the label pair, and the label pair is updated to (#1.1#1.1.1). In this way, after the target traversal execution is completed, the structural relationship between each node can be embedded in the current sequence through the label pair, and the sequence can be used as the conversion text sequence. It should be understood that the target traversal process shown in the accompanying drawings does not represent a specific limitation on the order of node access. In practice, the choice of traversal algorithm will determine the order of target nodes visited. The nodes in the syntax tree may be visited in different orders according to different traversal algorithms, which are not specifically limited here.

[0056] In addition to using self-closing tags such as brackets in the above examples as tag pairs in structural tags, in some practices, double tags can also be used as tag pairs. This type of tag consists of a start tag and a corresponding end tag, and the middle can contain tag content, for example: the start tag <x> , end tag< / x> . Figure 3B A schematic diagram of a serialization operation using a double tag as a tag pair is shown. In this figure, Figure 3A Taking the same abstract syntax tree in the example as an example, by performing target traversal, a current sequence that can be used as the conversion text sequence is generated. In the current sequence, a double label ( <x> ……< / x> ) as a label pair, and the node identifier of the parent node of the target node is marked. After the target traversal is completed, the current sequence shown in the figure can be obtained. The node identifier in any label of the label pair is the parent node, and the content between the start label and the end label of the label pair is the node identifier of each child node of the parent node. In this way, the purpose of expressing the structural relationship between nodes in the converted text sequence can also be achieved.

[0057] It is understandable that in some specific application scenarios, filters can be applied during the execution of the target traversal to filter the nodes that are accessed and have no coding significance, so as to avoid including nodes that are irrelevant to the code structure in the converted text sequence. The filter rules can be flexibly formulated according to the specific application scenario, and will not be specifically expanded here.

[0058] The above is an explanation of the serialization operation of the abstract syntax tree included in the second encoding process. After obtaining the conversion text sequence corresponding to the first abstract syntax tree, text encoding can be performed based on the conversion text sequence to obtain a vector that can characterize the structural features of the first source code. Similar to the aforementioned first encoding method, in the process of implementing the second encoding, a second text model capable of text encoding can be used. The second text model can use the same model as the first text model, or other models can be used, which is not limited here.

[0059] By executing the first encoding and the second encoding, a first vector and a second vector are obtained respectively. On this basis, the first characterization vector can be formed by fusing the two vectors. This process can flexibly adopt different fusion strategies according to specific needs in practice. According to one practice, a principal component analysis strategy can be used to extract the main components in the first vector and the second vector to generate the first characterization vector, thereby retaining the key features in the two vectors. According to another practice, the first vector and the second vector can be connected end to end and spliced ​​in a vector series manner to form the first characterization vector. There are many ways to obtain a fused vector, and the embodiments of this specification do not give examples one by one.

[0060] According to one implementation, the first characterization vector may be a weighted average of the first vector and the second vector. Specifically, a specific weight coefficient is applied to corresponding elements in the first vector and the second vector, and the elements are added to obtain the first characterization vector. In this way, the feature information carried by the two vectors can be more flexibly integrated, and the weight setting can be used to control the proportion of a certain type of feature of the first source code in the first characterization vector.

[0061] According to an implementation manner, the first vector and the second vector may not be fused, that is, the first representation vector may simply include the first vector and the second vector.

[0062] Next, in step S203, a similarity-based search is performed in the code base according to the first representation vector to obtain a target context; the target context includes a number of other source codes whose similarity to the first source code meets a first threshold and their corresponding patch codes.

[0063] In this step, for the first source code, a search is performed in the code library to obtain other source code files related to the first source code, and the search process can be performed based on the first representation vector. Other source codes in the code library have their own corresponding representation vectors for similarity retrieval. These representation vectors can also be determined based on the corresponding source code text and abstract syntax tree according to the calculation method of the first representation vector.

[0064] During the retrieval process, any other source code in the code base is compared with the first source code based on the first characterization vector to calculate the similarity. In the case where the first characterization vector is a single vector obtained by fusing the first vector and the second vector, the vector similarity between the first characterization vector and the characterization vector of some other source code can be calculated. If the vector similarity is higher than a preset similarity threshold (i.e., the first threshold), the other source code and its corresponding patch code are included in the target context. The first threshold defines the minimum similarity level that needs to be achieved between other source codes that can be used as target contexts and the first source code in order to be considered relevant. That is, if the similarity between the characterization vector of a certain source code and the first characterization vector exceeds the first threshold, then it can be considered that this source code is highly related to the first source code in terms of code semantics and code structure, and the source code and its corresponding patch code can be included in the target context.

[0065] If the first characterization vector includes a first vector and a second vector, the first vector and the second vector can be compared separately, and then the comparison results can be integrated. For example, the first similarity between the first vector of the first source code and the first vector of other source codes is calculated, and the first source code set is screened out according to the first similarity, and the second similarity between the second vector of the first source code and the second vector of other source codes is calculated, and the second source code set is screened out according to the second similarity, and the source code and its corresponding patch code in the intersection of the first source code set and the second source code set are included in the target context. Alternatively, the comprehensive similarity between the other source code and the first source code is determined based on the first similarity and the second similarity of the other source code, and several other source codes similar to the first source code are determined based on the comprehensive similarity, and these source codes and their corresponding patch codes are included in the target context.

[0066] In other examples, other similarity retrieval strategies may be flexibly adopted based on the first representation vector, which are not described one by one here.

[0067] After the target context is determined, in step S205, a prompt text may be generated according to the first source code and the target context, and a code model may be input to obtain an inferred patch.

[0068] In this step, the determined target context can be used in combination with the first source code to generate a prompt text for guiding the code model. The purpose of the prompt text is to allow the code model to understand the potential characteristics of the first source code and to generate a patch code that can repair the program defects in the first source code. The target context contained in the prompt text can provide the code model with rich relevant information (including several related source codes and corresponding defect repair patch codes) to facilitate it to accurately infer the patch code applicable to the first source code.

[0069] However, it should be noted that in actual practice, the code base often contains a large number of code fix pairs (i.e., source code and its corresponding patch code), and the content of the fix pairs is usually relatively large (some source code files can even have tens of thousands of lines of code). This may cause the content of the prompt text to be extremely large, containing a large number of tokens. A token is the basic unit for the code model to process text, which can represent a word, punctuation mark, or other text element. Generally speaking, the longer the content of the prompt text, the more tokens it contains, which will increase the amount of information that the code model needs to process during the reasoning process, thereby affecting the analysis efficiency. Therefore, in order to ensure the efficiency of reasoning, the code model usually presets a limit on the number of input tokens.

[0070] Accordingly, according to one implementation, in the prompt text, the other source codes are arranged from high to low according to their similarity to the first source code, and the number of the other source codes is such that the prompt text does not exceed the maximum number of the upper bound of the number of input tokens. The purpose of sorting is to place the other source codes most relevant to the first source code and their corresponding patch codes at the top of the prompt text so that they can be given priority in subsequent analysis or reasoning.

[0071] At the same time, in order to comply with the code model's limit on the number of input tokens (if any), the number of other source codes in the prompt text can be controlled at a critical value, under which the total number of tokens in the prompt text will not exceed the maximum number of input tokens that the code model can accept, and there will be no space to accommodate the next code repair pair. The prompt text constructed according to this implementation method not only ensures the information richness and relevance of the prompt text, but also avoids the risk of the number of tokens exceeding the upper limit of the code model's acceptance capacity due to the length of the text.

[0072] In addition, according to one implementation, in order to enhance the code model's ability to understand the relationship between the source code and the corresponding patch code in the prompt text, each source code can be marked with a first label and each patch code can be marked with a second label in the prompt text. The introduction of this label mechanism can enable the code model to more clearly identify the corresponding relationship between the source code and the patch code, so that in the reasoning process, the relevant information can be accurately located and used to effectively extract the repair mode between the source code and the patch code.

[0073] The specific forms of the first label and the second label can be various, for example, the first label can be [BUG], and the corresponding second label can be [FIX]; the first label can be [source code], and the corresponding second label can be [patch], etc. This specification embodiment does not list this in detail.

[0074] It is not difficult to understand that in the prompt text, for several other source codes related to the first source code, each is marked with the first label, and the corresponding patch code is marked with the second label. The first source code is also marked with the first label, but the content marked with the corresponding second label can be empty, or the corresponding second label is not set, because the patch code marked with the second label is the result that needs to be inferred and output by the code model.

[0075] Next, the code model can perform inference according to the input prompt text and output an inferred patch corresponding to the first source code. In step S207: fine-tune the parameters of the code model according to the inferred patch and the first patch code corresponding to the first source code.

[0076] By comparing the difference between the inferred patch and the first patch code (as the sample label), the code model can capture the deviations in the reasoning process and fine-tune the model parameters to optimize the predictive ability of the code model so that its inferred patch code is more accurate and closer to the sample label.

[0077] In a specific practice, the code model is a large language model, and the parameter fine-tuning may be fine-tuning of all parameters of the large language model. During the parameter fine-tuning process, the large language model can use the feedback obtained from the comparison between the inferred patch and the first patch code to adjust its internal parameters through the back propagation algorithm, including but not limited to the update of weight parameters and bias parameters. Full parameter fine-tuning can update all parameters of the model, which means that both the general feature knowledge learned in the pre-training stage and the knowledge for specific tasks learned in the fine-tuning stage will be integrated into the parameters of the model, which can not only enhance the inference performance of the model on the current automatic program repair task, but also improve the generalization ability of the model when processing similar tasks.

[0078] This iterative parameter fine-tuning can help the code model better adapt to the characteristics of the programming language and learn the characteristics of program defects, so that in the subsequent automatic program repair tasks, it can efficiently and accurately identify and repair various types of code defects. The steps S203-S205 described above are also applicable to the processing of the automatic program repair task by the code model. According to the code semantic features and code structure features of the input source code to be checked, the relevant repair pairs are retrieved in the code library, and screened based on the first threshold. Together with the source code to be checked, the prompt text is formed and input into the code model to infer the patch code. In other words, the fine-tuned code model can be used to infer the corresponding patch code for the source code to be checked according to the target prompt text; the target prompt text includes the source code to be checked, and several reference source codes whose similarity with the source code to be checked meets the first threshold and their corresponding patch codes.

[0079] The above text, based on one or more embodiments, elaborates on a method for fine-tuning a code model. Using the above method provided in the embodiments of this specification, the semantic representation and structural representation of the code can be comprehensively utilized to perform accurate context retrieval, and the context obtained by the retrieval can be used to construct prompt text to guide the code model to infer the patch code. The code model is fine-tuned based on the difference between the inferred patch and the patch sample label. Through iterative parameter fine-tuning, the code model's ability to understand the code semantics and code structure is continuously enhanced, the inference accuracy of the code repair task is improved, and the generalization ability of the code model is optimized.

[0080] In this specification, the word "first" in terms such as first coding and first model, and the corresponding "second", "third" (if any), etc. in the text are merely for the convenience of distinction and description and do not have any limiting meaning.

[0081] The foregoing describes certain embodiments of the present specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in an order different from that in the embodiments, and the desired results may still be achieved. In addition, the processes depicted in the accompanying drawings do not necessarily have to be performed in the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0082] Figure 4 Schematic diagram of a device for fine-tuning a code model according to an embodiment of this specification. The device 400 is deployed in a computing device, which can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities. Figure 2 The method embodiment shown corresponds to the method embodiment, which is based on a code library, and the code library includes multiple source codes and corresponding patch codes; the device 400 includes:

[0083] The encoding module 401 is configured to obtain a first representation vector corresponding to any first source code among the plurality of source codes, where the first representation vector is determined by performing a first encoding on the text of the first source code and performing a second encoding on a first abstract syntax tree corresponding to the first source code;

[0084] The retrieval module 402 is configured to perform a similarity-based search in the code base according to the first representation vector to obtain a target context; the target context includes a number of other source codes whose similarity to the first source code meets a first threshold and their corresponding patch codes;

[0085] The inference module 403 is configured to generate a prompt text according to the first source code and the target context, input a code model, and obtain an inferred patch;

[0086] The fine-tuning module 404 is configured to fine-tune the parameters of the code model according to the first patch code corresponding to the inferred patch and the first source code.

[0087] According to another embodiment of the present specification, there is also provided a computer program product, including a computer program / instruction, which implements the aforementioned combination when executed by a processor. Figure 2 The steps of the method.

[0088] According to another embodiment, the present specification further provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the aforementioned combination is implemented. Figure 2 The steps of the method.

[0089] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the embodiments of the present invention may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0090] The specific implementation methods described above further describe the purpose, technical solutions and beneficial effects of the embodiments of the present invention in detail. It should be understood that the above description is only a specific implementation method of the embodiments of the present invention and is not intended to limit the scope of protection of the present invention. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solution of the present invention shall be included in the scope of protection of the present invention.

Claims

1. A method for fine-tuning a code model based on a code library, wherein the code library includes a plurality of source codes and corresponding patch codes; the method comprises: Obtaining a first representation vector corresponding to any first source code among the plurality of source codes, where the first representation vector is determined by performing a first encoding on a text of the first source code and performing a second encoding on a first abstract syntax tree corresponding to the first source code; According to the first representation vector, perform a similarity-based search in the code library to obtain a target context; The target context includes a plurality of other source codes whose similarity with the first source code meets a first threshold and their corresponding patch codes; Generate a prompt text according to the first source code and target context, input a code model, and obtain an inferred patch; Fine-tune the parameters of the code model according to the first patch code corresponding to the inferred patch and the first source code.

2. The method according to claim 1, wherein: The fine-tuned code model is used to infer corresponding patch codes for the source code to be checked according to the target prompt text; the target prompt text includes the source code to be checked, several reference source codes whose similarity with the source code to be checked meets a first threshold and their corresponding patch codes.

3. The method according to claim 1, wherein: The first encoding includes: The first source code is text-encoded using a first text model.

4. The method according to claim 1, wherein: The second encoding includes: Performing a serialization operation on the first abstract syntax tree to obtain a converted text sequence, wherein the converted text sequence includes a node identifier representing each node of the first abstract syntax tree and a structure label representing a structural relationship between the nodes; The converted text sequence is text encoded using a second text model.

5. The method according to claim 4, wherein: The structural label includes a label pair marked with a node identifier, and the label pair is used to indicate a child node of a node corresponding to the marked node identifier; The serialization operation includes: During the target traversal of the first abstract syntax tree, for any target node currently visited, its node identifier is extracted, and the node identifier is added to the first label pair range in the current sequence, where the first label pair is marked with the node identifier of the parent node of the target node; the current sequence obtained at the end of the target traversal is used as the converted text sequence.

6. The method according to claim 1, wherein: The first characterization vector is a weighted average of the first vector obtained by the first encoding and the second vector obtained by the second encoding.

7. The method according to claim 1, wherein: In the prompt text, the other source codes are arranged from high to low according to their similarity with the first source code, and the number of the other source codes is a maximum number such that the prompt text does not exceed the upper limit of the number of input word units.

8. The method according to claim 1, wherein: In the prompt text, each source code is marked with a first tag, and each patch code is marked with a second tag.

9. The method according to claim 1, wherein: The code model is a large language model, and the fine-tuning of the parameters of the code model includes: fine-tuning all parameters of the large language model.

10. A device for fine-tuning a code model, running based on a code library, the code library comprising a plurality of source codes and corresponding patch codes, the device comprising: an encoding module configured to obtain a first representation vector corresponding to any first source code among the plurality of source codes, wherein the first representation vector is determined by performing a first encoding on a text of the first source code and performing a second encoding on a first abstract syntax tree corresponding to the first source code; A retrieval module is configured to perform a similarity-based search in the code library according to the first representation vector to obtain a target context; The target context includes a plurality of other source codes whose similarity with the first source code meets a first threshold and their corresponding patch codes; An inference module is configured to generate a prompt text according to the first source code and the target context, input a code model, and obtain an inferred patch; The fine-tuning module is configured to fine-tune the parameters of the code model according to the first patch code corresponding to the inferred patch and the first source code.

11. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 9.

12. A computing device comprising a memory and a processor, characterized in that: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Program code similarity measurement method based on abstract syntax tree access context

    CN113434145A

  • Method-level program repairing system and method based on pre-training model

    CN114546828A

  • Program defect automatic repairing method and system based on code language model

    CN116755753A

  • Code repairing system and method based on large language model

    CN116991467A

  • Method for performing efficient parameter fine tuning on code model in combination with abstract syntax tree

    CN118733052A