Model fine tuning method and electronic equipment

By filtering and synthesizing high-quality code problem data, iteratively fine-tuning the task model, solving the problem of uneven quality of the code data generated by the model, and improving the model's learning ability and fine-tuning effect.

CN120354932APending Publication Date: 2025-07-22ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510201125.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, the quality of the code data generated by the model is uneven, resulting in poor model fine-tuning effect.

Method used

The task model of the target task generates multiple first test codes of code problem data. The quality evaluation indicators of the test code determine the problem difficulty level, filter out the code problem data that reaches the preset level, synthesize the second test code, and iteratively fine-tune the task model.

Benefits of technology

The learning ability and fine-tuning effect of the task big model are improved, and the performance of the model is gradually improved through multiple iterations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354932A_ABST
    Figure CN120354932A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model fine tuning method and electronic equipment. The model fine tuning method comprises the steps of performing code generation on code problem data of a target task through a task large model of the target task to obtain a plurality of first test codes corresponding to the code problem data; the target task is a UTDD task or a coding task; testing the first test code, and determining a problem difficulty level of the code problem data according to a test result; in response to the condition that the problem difficulty level reaches a preset level, synthesizing a second test code of the code problem data according to the plurality of first test codes; and performing iterative fine tuning on the task large model according to the code problem data and the second test code to obtain a target large model. According to the method, code data with higher quality can be synthesized, and the model fine tuning efficiency and the model fine tuning effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model fine-tuning method and an electronic device. Background Art

[0002] In recent years, large models have achieved remarkable achievements in various fields such as coding, medicine, law, and finance. In order to further improve the knowledge and generalization ability of large models, a large amount of high-quality domain data is required to optimize the models. In the field of coding, high-quality code data is crucial for training and optimizing the performance of large models, directly affecting the performance and accuracy of the final model coding.

[0003] Currently, methods for obtaining coding-related training data include crawling open-source data, manually generating data, etc. However, these methods have some disadvantages: Although the cost of open-source code data is low, the available data volume is small, the quality is poor, and the open-source data lacks diversity and category imbalance, resulting in the possibility that the large model after training may not cover all problem solutions. The human cost of manually generating data is high and the generation speed is slow, which cannot meet the requirements of rapid iteration and large-scale data use. To solve this problem, a method of synthetic data is adopted to obtain training data. By using the method of synthetic data to obtain training data, the human cost can be reduced, and the types and quantities of fine-tuning data can be increased. However, currently, the code data generated by the model still has the problem of uneven quality and performs poorly in model optimization. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a model fine-tuning method and an electronic device, so as to solve the technical problem that in the related art, the code data generated by the model has poor quality, resulting in uneven quality of the code data based on which the task large model is fine-tuned (or optimized), and further resulting in poor model fine-tuning effects.

[0005] To solve the above technical problem, the embodiments of this application are implemented as follows: On the one hand, the embodiments of this application provide a model fine-tuning method, including: Using the task large model of the target task to generate code for the code problem data of the target task, obtaining a plurality of first test codes corresponding to the code problem data; the target task is a UTDD (Unit Test-Driven Development) task or a coding task; Testing the first test codes, and determining the problem difficulty level of the code problem data according to the test results; In response to the problem difficulty level reaching a preset level, synthesizing a second test code of the code problem data according to the plurality of first test codes; Iteratively fine-tune the task large model according to the code problem data and the second test code to obtain a target large model.

[0006] On the other hand, an embodiment of the present application provides a model fine-tuning device, including: A code generation module, configured to generate codes for the code problem data of the target task through a task large model of the target task to obtain a plurality of first test codes corresponding to the code problem data; the target task is a unit test-driven development UTDD task or a coding task; A determination module, configured to test the first test code and determine the problem difficulty level of the code problem data according to the test result; A code synthesis module, configured to, in response to the problem difficulty level reaching a preset level, synthesize a second test code of the code problem data according to the plurality of first test codes; A model fine-tuning module, configured to iteratively fine-tune the task large model according to the code problem data and the second test code to obtain a target large model.

[0007] On yet another aspect, an embodiment of the present application provides an electronic device, including a processor and a memory electrically connected to the processor, where the memory stores a computer program, and the processor is configured to call and execute the computer program from the memory to implement the above model fine-tuning method.

[0008] On yet another aspect, an embodiment of the present application provides a computer-readable storage medium for storing a computer program, and the computer program can be executed by a processor to implement the above model fine-tuning method.

[0009] On yet another aspect, an embodiment of the present application provides a computer program product, including a computer program, and the computer program is executed by a processor to implement the above model fine-tuning method.

[0010] Adopting the technical solution of the embodiment of the present application, through the task large model of the target task, code generation is performed on the code problem data of the target task to obtain multiple first test codes corresponding to the code problem data, where the target task is a UTDD task or a coding task; the first test codes are tested, and the problem difficulty level of the code problem data is determined according to the test results; in response to the problem difficulty level reaching a preset level, according to the multiple first test codes, a second test code of the code problem data is synthesized, and the task large model is iteratively fine-tuned according to the code problem data and the second test code to obtain a target large model. Since the code problem data with a low difficulty level does not have a learning space to a certain extent for the task large model, therefore, by screening out the code problem data with a relatively high problem difficulty level (such as reaching the preset level), and performing model fine-tuning according to the code problem data with a high problem difficulty level and its corresponding second test code, in fact, the code problem data that does not need to be learned by the task large model is screened out, and only the code problem data with a high problem difficulty level is retained, which is beneficial to improving the learning ability of the task large model. Moreover, this way of screening code problem data for UTDD tasks and coding tasks can gradually improve the quality of the code problem data retained after each iteration through multiple iteration processes, so that the data used for each fine-tuning of the task large model is of better quality than the data used for the previous iteration fine-tuning. By continuously synthesizing higher-quality model fine-tuning data, the learning ability of the model is continuously improved, so as to obtain an extremely performant target large model through multiple iteration fine-tuning. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in one or more embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in one or more embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0012] Figure 1 is a schematic flowchart of a model fine-tuning method according to an embodiment of the present application; Figure 2 is a schematic flowchart of a model fine-tuning method according to another embodiment of the present application; Figure 3 is a schematic principle diagram of a UTDD model fine-tuning method according to an embodiment of the present application; Figure 4 is a schematic principle diagram of a coding model fine-tuning method according to an embodiment of the present application; Figure 5 is a schematic block diagram of a model fine-tuning device according to an embodiment of the present application; Figure 6 It is a schematic block diagram of an electronic device according to an embodiment of the present application. Specific Embodiments

[0013] Embodiments of the present application provide a model fine-tuning method and an electronic device.

[0014] In order to enable those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0015] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein.

[0016] The model fine-tuning method provided by the embodiments of the present application can be executed by an electronic device or by software installed in the electronic device. Specifically, the electronic device can be a terminal device or a server device. Among them, the terminal device can include a smart phone, a laptop computer, a smart wearable device, a vehicle-mounted terminal, etc., and the server device can include an independent physical server, a server cluster composed of multiple servers, or a cloud server capable of performing cloud computing.

[0017] Before introducing the model fine-tuning method provided by the present application, first, the application environment of the embodiments of the present application will be described. The present application is applicable to the research and development coding field, mainly used for writing functional code or test case code according to user requirements, and using the written functional code and / or test case code to fine-tune a task large model. The model fine-tuning method in the embodiments of the present application can support any existing type of coding language, such as languages like C, C++, Java, Golang, Python, etc. Before executing the model fine-tuning method, a compilation environment and a test environment for the corresponding coding language can be set up according to requirements.

[0018] Figure 1 It is a schematic flowchart of a model fine-tuning method according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps S102 to S108: Step S102: Use the task large model of the target task to generate code for the code problem data of the target task, obtaining multiple first test codes corresponding to the code problem data.

[0019] Among them, the target task is a UTDD task or a coding task. The task large model of the target task refers to a large model that can serve the target task. Optionally, if the target task is a UTDD task, the corresponding task large model can be a pre-trained test model that can generate code data for unit testing; if the target task is a coding task, the corresponding task large model can be a pre-trained coding model that can generate code data for coding. Both the test model and the coding model are mature R & D large models and both have the ability to generate code, so the training process is not elaborated here.

[0020] Optionally, before executing step S102, the code problem data of the target task can be automatically generated by a requirements description model matching the target task. Specifically, first obtain the original code data of the code problem data, and then use the requirements description model matching the target task to generate a requirements description statement corresponding to the original code data, which is used to describe the function to be implemented by the original code data. Furthermore, determine the code problem data according to the requirements description statement. After determining the code problem data, the code problem data can be input into the task large model to generate the answer data corresponding to the code problem data, and the first test code corresponding to the code problem data can be extracted according to the answer data. The answer data is the complete code information, which can be used as the first test code, or some useful information can be extracted from the complete code information as the first test code.

[0021] One code problem data can correspond to one or more first test codes. For the same code problem data, the corresponding multiple first test codes can be different in one or more dimensions. For example, the coding languages corresponding to the multiple first test codes are different, the compilation methods are different, etc. The original code data can be one of the multiple first test codes.

[0022] A folder corresponding to the code data can be created, which is used to store data related to the code problem data. After generating the first test code corresponding to the code problem data, store the code problem data and its corresponding first test code in the folder. When executing step S102, obtain the pre-generated code problem data and its corresponding first test code from the folder. In addition, other data related to the first test code can also be stored in the folder, including but not limited to the following data: quality evaluation metrics of the first test code, problem difficulty level of the code problem data corresponding to the first test code, code running report of the first test code, etc.

[0023] Step S104: Test the first test code and determine the problem difficulty level of the code problem data according to the test results.

[0024] Among them, the representation method of the problem difficulty level is not limited. For example, it can be divided into three levels: high, medium, and low, or it can be divided into multiple levels such as the first level, the second level, and the third level. Various forms such as numbers, letters, symbols, etc. can also be used to distinguish different problem difficulty levels, which are not listed one by one in this embodiment.

[0025] The problem difficulty level can characterize the answering ability of the task large model to the code problem data. That is, the higher the problem difficulty level, the lower the accuracy of the answering result (i.e., the answer) of the task large model to the code problem data. The test data used to optimize the task large model is screened by the problem difficulty level, so that the test code with no learning space for the task large model (i.e., the first test code corresponding to the code problem data with a lower difficulty level) is screened out.

[0026] Step S106: In response to the problem difficulty level reaching the preset level, synthesize the second test code of the code problem data according to multiple first test codes.

[0027] In this embodiment, multiple first test codes corresponding to the code problem data can be generated by the task large model. These multiple first test codes may include first test codes whose problem difficulty levels do not reach the preset level. The first test codes whose problem difficulty levels do not reach the preset level are screened out, so as to only retain the first test codes whose problem difficulty levels reach the preset level. The second test code is synthesized using the retained first test codes after screening, so that the task large model only learns the first test codes with a higher difficulty, which is beneficial to improving the learning ability of the task large model.

[0028] The preset level can be set according to requirements. For example, if it is desired that the learning ability of the task large model improves relatively quickly, then the preset level can be set to a higher level. Assuming that the problem difficulty level is divided into three levels: low, medium, and high, and the preset level is high, then the first test codes with a high problem difficulty level are screened out. If it is desired that the learning ability of the task large model improves at a moderate speed, then the preset level can be set to a medium level. Assuming that the problem difficulty level is divided into three levels: low, medium, and high, and the preset level is medium, then the first test codes with a medium and high problem difficulty level are screened out.

[0029] Step S108: Iteratively fine-tune the task large model according to the code problem data and the second test code to obtain the target large model.

[0030] In this embodiment, the iteration termination conditions for iterative fine-tuning of the task large model may include at least one of the following: the number of iteration rounds reaches a preset number, and the quality evaluation index of the answer data of the task large model for the question data meets the preset index requirements. The question data refers to the data for testing the model performance of the task large model, and can be called test question data. The quality evaluation index may include at least one of the following: compilation pass rate, branch coverage rate, and line coverage rate. The preset index requirements may include at least one of the following: the quality evaluation index is among the top X, and the quality evaluation index reaches a preset index value; X is an integer greater than or equal to 1.

[0031] Among them, the quality evaluation index being among the top X may include at least one of the following: the compilation pass rate is among the top X of all compilation pass rates, the branch coverage rate is among the top X of all branch coverage rates, and the line coverage rate is among the top X of all line coverage rates. The quality evaluation index reaching the preset index value may include at least one of the following: the compilation pass rate reaches the preset compilation pass rate threshold, the branch coverage rate reaches the preset branch coverage rate threshold, and the line coverage rate reaches the preset line coverage rate threshold.

[0032] After each iteration is completed, it can be determined whether the task large model meets the iteration termination conditions. Optionally, when the iteration termination condition is that the quality evaluation index of the answer data of the task large model for the question data meets the preset index requirements, the test question data of the target task can be obtained in advance. The test question data can be obtained by crawling open-source data, generating data manually, etc. The test question data of the target task is input into the task large model of the current iteration to obtain the answer data of the test question data; then, according to the code metrics of the answer data, the quality of the answer data is evaluated to obtain the quality evaluation index of the answer data; furthermore, it is determined whether the quality evaluation index of the answer data meets the preset index requirements. In response to the quality evaluation index of the answer data meeting the preset index requirements, the task large model of the current iteration is determined as the target large model, that is, the iteration terminates. In response to the quality evaluation index of the answer data not meeting the preset index requirements, the task large model of the current iteration is continuously fine-tuned until the task large model obtained after iteration meets the iteration termination conditions.

[0033] Adopting the technical solution of the embodiment of the present application, through the task large model of the target task, code generation is performed on the code problem data of the target task to obtain multiple first test codes corresponding to the code problem data, where the target task is a UTDD task or a coding task; the first test codes are tested, and the problem difficulty level of the code problem data is determined according to the test results; in response to the problem difficulty level reaching the preset level, according to the multiple first test codes, a second test code of the code problem data is synthesized, and the task large model is iteratively fine-tuned according to the code problem data and the second test code to obtain the target large model. Since the code problem data with a low difficulty level does not have a learning space to a certain extent for the task large model, therefore, by screening out the code problem data with a relatively high problem difficulty level (such as reaching the preset level) and performing model fine-tuning according to the code problem data with a high problem difficulty level and its corresponding second test code, in fact, the code problem data that does not need to be learned by the task large model is screened out, and only the code problem data with a high problem difficulty level is retained, which is beneficial to improving the learning ability of the task large model. Moreover, this way of screening code problem data for UTDD tasks and coding tasks can gradually improve the quality of the code problem data retained after each iteration through multiple iteration processes, so that the data used for each fine-tuning of the task large model is of better quality than the data used for the previous iterative fine-tuning. By continuously synthesizing higher-quality model fine-tuning data, the learning ability of the model is continuously improved, so as to obtain an extremely performant target large model through multiple iterative fine-tuning.

[0034] In one embodiment, testing the first test code can be performed according to the following steps A1 - A2: Step A1, for each first test code, according to the code dynamic metrics of the first test code in the running state, quality evaluation of the first test code is performed to obtain the quality evaluation metrics of the first test code.

[0035] Among them, the code dynamic metrics belong to one of the code measurement metrics. Generally, the code measurement metrics can include: code static metrics of the code in the non-running state, and code dynamic metrics of the code in the running state. The code static metrics include but are not limited to metrics such as the lexical, syntactic, semantic, and cyclomatic complexity of the code, and the code dynamic metrics include but are not limited to metrics such as the functional test report, execution time, memory peak, concurrent performance, and network latency during the code running process.

[0036] The quality evaluation metrics include at least one of the following: branch coverage rate and line coverage rate. The branch coverage rate refers to the ratio of the number of branches covered during the code execution to the total number of branches, and is a key metric for measuring the effectiveness of the code in covering the possible paths of the program control flow. The line coverage rate is used to measure whether each executable statement in the code has been executed. By running the first test code and based on the performance of various metrics in the running state of the first test code, the branch coverage rate and the line coverage rate can be obtained. Therefore, by evaluating the quality of the first test code according to the code metric corresponding to the first test code, the code quality of the first test code can be comprehensively and accurately evaluated from multiple dimensions (including one or more dimensions of the branch coverage rate and the line coverage rate), thereby providing an accurate data basis for determining the problem difficulty level of the code problem data.

[0037] Step A2: Determine the problem difficulty level of the code problem data according to the quality evaluation metrics of each first test code.

[0038] In addition to the above-described steps A1 - A2, the quality of the first test code can also be evaluated according to the code static metrics of the first test code to obtain the quality evaluation metrics of the first test code, and the quality evaluation metrics include the compilation pass rate. Furthermore, the problem difficulty level of the code problem data is determined according to the quality evaluation result.

[0039] The compilation pass rate refers to the proportion of the code that can be successfully compiled, reflecting the smooth progress of the code during the conversion to an executable program. By compiling the code static metrics of the first test code through a compiler, the compilation pass rate of the first test code can be obtained.

[0040] Optionally, the quality of the first test code can also be evaluated by combining the code static metrics and the code dynamic metrics of the first test code to obtain the quality evaluation metrics of the first test code. In this case, the quality evaluation metrics include the compilation pass rate, the branch coverage rate, and the line coverage rate. Furthermore, the problem difficulty level of the code problem data is determined through multiple quality evaluation metrics.

[0041] When determining the problem difficulty level of the code problem data according to the quality evaluation metrics, the problem difficulty level can be determined according to one or more of the compilation pass rate, the branch coverage rate, and the line coverage rate.

[0042] Optionally, when determining the problem difficulty level according to one quality evaluation metric among the compilation pass rate, the branch coverage rate, and the line coverage rate, the problem difficulty level and each quality evaluation metric are negatively correlated, that is, the higher the compilation pass rate, the branch coverage rate, or the line coverage rate, the lower the corresponding problem difficulty level.

[0043] When determining the problem difficulty level based on multiple quality evaluation metrics (such as two or three) among the compilation pass rate, branch coverage rate, and line coverage rate, first, based on the multiple quality evaluation metrics, determine the overall pass rate of the code problem data, and then determine the problem difficulty level according to the overall pass rate.

[0044] Optionally, when comprehensively determining the problem difficulty level according to the three quality evaluation metrics of the compilation pass rate, branch coverage rate, and line coverage rate, step A2 can be executed as the following steps A21 - A23: Step A21, determine the first test code whose compilation pass rate reaches the preset pass rate threshold.

[0045] Step A22, according to the branch coverage rate and / or line coverage rate of the first test code whose compilation pass rate reaches the preset pass rate threshold, determine the overall pass rate of the code problem data.

[0046] Step A23, determine the problem difficulty level of the code problem data according to the overall pass rate.

[0047] Optionally, when determining the overall pass rate of the code problem data according to multiple quality evaluation metrics, the index weight corresponding to each quality evaluation metric can be determined in advance, and then, perform a weighted sum calculation according to each determined quality evaluation metric and its respective corresponding index weight to obtain the overall pass rate.

[0048] Optionally, according to the overall pass rate of the code problem data, the problem difficulty level of the code problem data can be divided into three levels: low, medium, and high. There is a negative correlation between the overall pass rate and the problem difficulty level, that is, the higher the overall pass rate, the lower the corresponding problem difficulty level.

[0049] In this embodiment, the quality evaluation metrics of the first test code can characterize the ability of the task large model to answer questions to a certain extent. For the UTDD task of the Java language, the problem difficulty level can be divided into three levels: easy (Easy), medium (Medium), and hard (Hard). Among them, the overall pass rate corresponding to the code problem data at the Easy level is greater than or equal to 80%, the overall pass rate corresponding to the code problem data at the Medium level is greater than 10% and less than 80%, and the overall pass rate corresponding to the code problem data at the Hard level is less than or equal to 10%. Of course, the above data are only examples and are not limiting conditions for this application.

[0050] Since the code problem data with a lower problem difficulty level does not have a learning space for the task large model, and the task large model can already answer the code problem data relatively accurately, while the task problem data with a higher problem difficulty level is a weak point for the task large model, that is, the answer results for this part of the code problem data are not accurate enough. Therefore, by determining the problem difficulty level of the code problem data through the quality evaluation index of each first test code, the code problem data that is a blank item or a weak point for the task large model can be accurately screened out, thereby improving the knowledge learning ability of the task large model.

[0051] In one embodiment, after screening out multiple first test codes whose problem difficulty level reaches the preset level, a second test code of the code problem data is synthesized according to the screened multiple first test codes. When synthesizing the second test code, the following steps B1 - B2 can be executed: Step B1, extract the valid test codes from the multiple first test codes according to the quality evaluation index of each screened first test code.

[0052] Step B2, combine the valid test codes to obtain the second test code. Among them, the valid test code can be understood as a useful information segment in the first test code. When combining multiple valid test codes, the multiple valid test codes can be input into the task large model so that the task large model synthesizes the multiple valid test codes into a complete second test code. In this embodiment, the task large model has the function of synthesizing test codes.

[0053] In this embodiment, there are multiple ways to extract the valid test codes from the multiple first test codes. The following exemplarily lists two extraction methods.

[0054] Method 1, extract at least one first test code whose quality evaluation index meets the first index requirement from the multiple first test codes as the valid test code. The first index requirement includes at least one of the following: the quality evaluation index is among the top N, and the quality evaluation index reaches the preset index value; N is an integer greater than or equal to 1.

[0055] Among them, the quality evaluation index being among the top N may include at least one of the following: the compilation pass rate is among the top N of all compilation pass rates, the branch coverage rate is among the top N of all branch coverage rates, and the line coverage rate is among the top N of all line coverage rates. The quality evaluation index reaching the preset index value may include at least one of the following: the compilation pass rate reaches the preset compilation pass rate threshold, the branch coverage rate reaches the preset branch coverage rate threshold, and the line coverage rate reaches the preset line coverage rate threshold.

[0056] Method 2: Using the greedy algorithm to extract effective test codes, that is, selecting the combination of the first test codes with the smallest number from multiple first test codes and ensuring that the combination of the first test codes can achieve the best quality evaluation index. Optionally, the method of using the greedy algorithm to extract effective test codes includes the following steps B11 - B13: Step B11: Select the first test code with the highest quality evaluation index from the multiple first test codes that have not been selected.

[0057] Before the first selection, all the first test codes screened out belong to the first test codes that have not been selected. The selected first test codes can be marked, for example, marked as "selected", or removed from the set of the first test codes that have not been selected.

[0058] Step B12: Combine the selected first test codes and determine the quality evaluation index of the combined code obtained after combination.

[0059] After the first selection of the first test code, since only one first test code is selected, the combined code only includes one first test code, and the quality evaluation index of this first test code is the quality evaluation index of the combined code.

[0060] Step B13: In response to the quality evaluation index of the combined code meeting the second index requirement, determine the selected first test code as the effective test code. The second index requirement includes at least one of the following: the quality evaluation index reaches the preset index value, the quality evaluation index reaches the highest. In response to the quality evaluation index of the combined code not meeting the second index requirement, continue to select the first test code that has not been selected and has the highest quality evaluation index.

[0061] Among them, the quality evaluation index reaching the highest means that the quality evaluation indexes corresponding to other combinations of the first test codes are not higher than the quality evaluation index of the current combined code.

[0062] For code problem data with different problem difficulty levels, the corresponding effective test codes can be extracted in the same or different ways. Taking the UTDD task in the Java language as an example, for code problem data at the Medium level, the effective test codes can be extracted according to Method 1, for example, extracting the first test codes with the quality evaluation indexes in the top N as the effective test codes. For code problem data at the Hard level, the effective test codes can be extracted according to Method 2. First, determine the quality evaluation index of each first test code, and then add the first test codes in order from high to low quality evaluation index without repetition for combination until the quality evaluation index of the combined first test codes (i.e., the combined code) reaches the highest.

[0063] After combining multiple first test codes, there are multiple calculation methods for the quality evaluation index of the combined code. For example, the quality evaluation indexes of multiple first test codes are weighted and summed, or the quality evaluation indexes of multiple first test codes are averaged to obtain the quality evaluation index of the combined code. Since the addition is performed in the order of the quality evaluation indexes of the first test codes from high to low, when the quality evaluation index of the combined code obtained after one addition starts to decline, it is determined that the quality evaluation index of the combined code obtained after the previous addition reaches the highest.

[0064] Figure 2 It is a schematic flowchart of a test code generation method according to an embodiment of the present application. As Figure 2 shown, it includes the following steps S201-S209: Step S201, through the task large model of the target task, generate code for the code problem data of the target task to obtain multiple first test codes corresponding to the code problem data.

[0065] Taking the UTDD task based on the Java programming language as an example, obtaining the first test code essentially means extracting the complete functions in a class from the code problem data and its corresponding complete code information. In specific implementation, the Tree-sitter syntax tree can be used for extraction, and DFS (Depth First Search) can be combined to identify the Methods in the class of the UTDD program (i.e., the test class program) to obtain the complete functions in the class, so as to obtain the first test code. For example, if the task problem data corresponds to n complete code information, the test case functions are extracted from each complete code information using the Tree-sitter syntax tree to obtain test case function 1, test case function 2,..., test case function n, and these n test case functions are the n first test codes corresponding to the code problem data.

[0066] Step S202, determine the compilation pass rate of each first test code.

[0067] Optionally, the compilation pass rate is determined according to the static indexes of the first test code. The static indexes include but are not limited to indexes such as the lexical, syntactic, semantic, and cyclomatic complexity of the code. An existing compiler can be used to compile the first test code to obtain the compilation pass rate. For Java code, the Java compiler can be used to check the static indexes of the first test code.

[0068] Step S203, when the compilation pass rate is greater than or equal to the preset pass rate threshold, determine the code coverage rate of the first test code.

[0069] Among them, the code coverage includes branch coverage / line coverage. The code coverage can be determined according to the dynamic metrics of the first test code, and the dynamic metrics include but are not limited to functional test reports, execution time, peak memory, concurrent performance, network latency, etc. during the code running process. For Java code, the dynamic metric report of the first test code can be obtained through the deployed test framework.

[0070] Optionally, for the first test code with a compilation pass rate less than the preset compilation pass rate threshold, it indicates that the quality of the first test code is poor, and the subsequent steps are no longer executed.

[0071] Step S204, according to the code coverage of each first test code, determine the overall pass rate of the multiple first test codes corresponding to the code problem data.

[0072] Among them, the overall pass rate can be understood as: the proportion of the number of first test codes with a code coverage reaching the preset coverage threshold among all the first test codes corresponding to the code problem data.

[0073] Step S205, according to the corresponding relationship between the overall coverage rate and the problem difficulty level, determine the problem difficulty level of the code problem data.

[0074] The corresponding relationship between the overall coverage rate and the problem difficulty level can be preset. Taking the UTDD task in the Java language as an example, the problem difficulty level can be divided into three levels: easy (Easy), medium (Medium), and hard (Hard). Among them, the overall pass rate corresponding to the code problem data at the Easy level is greater than or equal to 80%, the overall pass rate corresponding to the code problem data at the Medium level is greater than 10% and less than 80%, and the overall pass rate corresponding to the code problem data at the Hard level is less than or equal to 10%.

[0075] Step S206, screen out the target code problem data with the problem difficulty level reaching the preset level.

[0076] Step S207, for each target code problem data, according to the quality evaluation metrics of each first test code corresponding to the target code problem data, extract the effective test codes from the first test codes corresponding to the target code problem data.

[0077] The extraction method of the effective test codes has been described in detail in the above embodiments and will not be repeated here.

[0078] Step S208, combine the effective test codes to obtain the second test code corresponding to the target code problem data.

[0079] Among them, when combining multiple valid test codes, the multiple valid test codes can be concatenated in sequence to form the second test code.

[0080] After generating the second test code corresponding to each target code problem data, the second test code can be applied to the fine-tuning of the task large model.

[0081] It can be seen that in the embodiments of the present application, the target code problem data with the problem difficulty level reaching the preset level is screened out through quality evaluation indicators (such as compilation pass rate and code coverage rate), and then the second test code is synthesized for the target code problem data, so as to synthesize the test code for fine-tuning the task large model with high quality and improve the efficiency of fine-tuning the task large model.

[0082] In one embodiment, after synthesizing the second test code of the code problem data according to multiple first test codes, the task large model can be iteratively fine-tuned directly according to the screened code problem data and its corresponding second test code. It is also possible to first perform a quality evaluation on the second test code, and when the quality evaluation result of the second test code meets the requirements, then iteratively fine-tune the task large model according to the screened code problem data and its corresponding second test code.

[0083] Optionally, the method for performing a quality evaluation on the second test code may include the following steps: First, run the second test code, and determine the quality evaluation indicators corresponding to the second test code according to the running result of the second test code. The quality evaluation indicators include at least one of the following: compilation pass rate, branch coverage rate, line coverage rate.

[0084] Secondly, when the quality evaluation indicators corresponding to the second test code meet the requirements of the third indicator, it is determined that the second test code meets the task code requirements of the target task. The requirements of the third indicator may include at least one of the following: the quality evaluation indicator is in the top Y positions, the quality evaluation indicator reaches a preset indicator value; Y is an integer greater than or equal to 1.

[0085] Still taking the UTDD task in the Java language as an example, if the quality evaluation indicators of the synthesized second test code meet the requirements of the third indicator, for example, the line coverage rate and branch coverage rate of the second test code both exceed 95%, it can be determined that the second test code meets the task code requirements of the UTDD task.

[0086] After the above quality assessment, when performing step S108, in the case where the second test code meets the requirements of the task code, the second code data is applied to iteratively fine-tune the task large model. That is, according to the second test code that meets the requirements of the task code and its corresponding code problem data, the task large model is fine-tuned. By performing a quality assessment on the second test code, it is possible to ensure that the data quality applied to model fine-tuning is better, thereby improving the effect of model fine-tuning.

[0087] Taking the target task as the UTDD task and the task large model as the UTDD model as an example, a model fine-tuning method provided by an embodiment of the present application is described. Figure 3 It is a schematic diagram of a UTDD model fine-tuning method according to an embodiment of the present application. During the UTDD model fine-tuning process, a small amount of high-quality and multi-language UTDD fine-tuning data can be first used as cold start to activate the potential of the UTDD model. The multi-language may include languages such as C, C++, Java, Golang, and Python. The UTDD fine-tuning data includes: the second test code obtained by combining valid test codes extracted from a plurality of first test codes corresponding to the task problem data, and the task problem data corresponding to the second test code.

[0088] As Figure 3 shown, first determine the seed data of the code problem data, that is, the original code data. Use the requirements description model to write detailed requirements description statements for the seed data. These requirements description statements are used to describe the functions to be implemented by the original code data, and the requirements description statements can be stored in the requirements description corpus. Then, the requirements description statements can be obtained from the requirements description corpus, the corresponding code problem data can be determined according to the requirements description statements, and the code problem data is input into the UTDD model to generate a plurality of first test codes for the code problem data. The plurality of first test codes can be different in terms of compilation language, compilation method, etc. The code problem data and its corresponding first test codes can be stored in the test corpus.

[0089] After that, for each code problem data in the test corpus, determine the quality assessment indicators of each first test code corresponding to the code problem data, and determine the problem difficulty level of the code problem data according to the quality assessment indicators. In the case where the problem difficulty level of the code problem data reaches the preset level, according to the problem difficulty level of the code problem data, an appropriate screening strategy is adopted to extract valid test codes from the plurality of first test codes. Among them, how to determine the problem difficulty level of the code problem data and the method of extracting valid test codes from the plurality of first test codes have been described in detail in the above embodiments and will not be repeated here.

[0090] After extracting the effective test code, the effective test code is combined to obtain the second test code. The second test code obtained after combination and its corresponding task problem data are used as the UTDD fine-tuning data for the current iteration. The UTDD model is fine-tuned using the UTDD fine-tuning data to obtain a fine-tuned UTDD model. For the UTDD model fine-tuned in the current iteration, the test problem data of the target task is used to check whether the UTDD model meets the iteration termination condition. The test problem data can be part or all of the task problem data obtained in advance, or other problem data re-obtained. This embodiment does not limit this. The test problem data is input into the task large model of the current iteration to obtain the answer data of the test problem data; and then, according to the code metric index of the answer data, the quality of the answer data is evaluated to obtain the quality evaluation index of the answer data. When the quality evaluation index of the answer data meets the preset index requirements, it is determined that the UTDD model fine-tuned in the current iteration meets the iteration termination condition, and the UTDD model fine-tuned in the current iteration is determined as the target UTDD model. If the UTDD model fine-tuned in the current iteration does not meet the iteration termination condition, the UTDD fine-tuning data obtained after the current iteration is used as the seed data for the next round, and the UTDD model is continued to be fine-tuned in the next round.

[0091] It can be seen that since in the iterative fine-tuning process of each round, sample data that is relatively simple for the current model (i.e., task problem data and its corresponding first test code) is screened out and excluded according to the quality evaluation index of the test code, and only sample data with a higher problem difficulty level for the current model is retained for the next round of iteration, it is possible to ensure that the capabilities of the UTDD model are continuously optimized and improved in multiple iterations.

[0092] In this embodiment, for the UTDD task of each language, the iterative fine-tuning of the UTDD model can be performed respectively according to the Figure 3 steps shown. During the execution of the UTDD iterative task of each language, the UTDD fine-tuning data corresponding to the single language is obtained. After that, the UTDD fine-tuning data of each language can be merged to obtain the UTDD fine-tuning data of multiple languages. Optionally, using the UTDD fine-tuning data of multiple languages to perform iterative fine-tuning on the UTDD model can enable the fine-tuned UTDD model to have the UTDD development ability of multiple languages, and the performance of the UTDD model is better.

[0093] When the target task is a coding task, the same Figure 3 execution principle shown can be adopted to implement the iterative fine-tuning of the coding model. The iterative execution process of the coding task is the same as that of the UTDD task and will not be repeated here.

[0094] In one embodiment, in addition to usingFigure 3 In addition to implementing the execution principle shown, the UTDD fine-tuning data collected during the UTDD iterative task can also be used as the startup data for the coding task. The UTDD fine-tuning data includes: the second test code obtained by combining the valid test codes extracted from the multiple first test codes corresponding to the task problem data of UTDD, and the task problem data corresponding to the second test code.

[0095] For the case where the target task is a coding task and the task large model includes a coding model, when iteratively fine-tuning the task large model according to the code problem data and the second test code to obtain the target large model, the following steps C1 - C4 can also be executed: Step C1, obtain the target code problem data whose problem difficulty level reaches the preset level, and the second test code corresponding to the target code problem data.

[0096] Among them, step C1 can be to obtain the target code problem data and its corresponding second test code during the execution of the UTDD iterative task. After screening out the target code problem data and its corresponding second test code and combining them, it is the UTDD fine-tuning data.

[0097] Step C2, input the target code problem data into the coding model to obtain multiple first function codes corresponding to the target task problem data.

[0098] Step C3, determine the coding fine-tuning data corresponding to the coding model according to the target code problem data and the multiple first function codes. The coding fine-tuning data includes the target code problem data and its corresponding second function code.

[0099] When executing step C3, first, according to the target code problem data and its corresponding multiple first function codes, determine the quality evaluation index of each first function code; the quality evaluation index includes at least one of the following: compilation pass rate, branch coverage rate, line coverage rate. Then, according to the quality evaluation index of each first function code, select the second function code that meets the requirements of the fourth index from the multiple first function codes.

[0100] Among them, the requirements of the fourth index include at least one of the following: the quality evaluation index is in the top M positions, the quality evaluation index reaches the preset index value, and M is an integer greater than or equal to 1.

[0101] Step C4, iteratively fine-tune the coding model according to the coding fine-tuning data.

[0102] Figure 4 is a schematic diagram of a coding model fine-tuning method according to an embodiment of the present application. As Figure 4As shown in the figure, during the fine-tuning process of the encoding model, the UTDD fine-tuning data collected during the UTDD iterative task is used as the starting data for the encoding task.

[0103] First, the UTDD fine-tuning data is input into the encoding model to obtain an encoding corpus. The encoding corpus includes: the target code problem data in the UTDD fine-tuning data and the first functional code corresponding to the target code problem data.

[0104] Secondly, a quality assessment is performed on the target code problem data in the encoding corpus and the first functional code corresponding to the target code problem data, and encoding fine-tuning data is generated according to the quality assessment results. Among them, the quality assessment process may include: determining the quality assessment index of each first functional code, and selecting the second functional code that meets the requirements of the fourth index from multiple first functional codes according to the quality assessment index of each first functional code.

[0105] If there is only one second functional code that meets the requirements of the fourth index, the second functional code and its corresponding target code problem data can be directly determined as the encoding fine-tuning data. If there are multiple second functional codes that meet the requirements of the fourth index, multiple second functional codes can be combined to obtain a target functional code, and the target functional code and its corresponding target code problem data are determined as the encoding fine-tuning data. The encoding model is fine-tuned using the encoding fine-tuning data to obtain a fine-tuned encoding model. When the encoding model reaches the iteration termination condition, the encoding model fine-tuned in the current iteration is determined as the target encoding model. When the encoding model does not reach the iteration termination condition, the encoding fine-tuning data obtained during the current iteration is used as the starting data for the next round, and the encoding model is continuously fine-tuned in the next round.

[0106] It can be seen that in this embodiment, by using the UTDD fine-tuning data as the starting data for the encoding task, since the UTDD fine-tuning data is already high-quality data screened during the execution of the UTDD iterative task, including high-quality task problem data and its corresponding code data, the data quality of the starting data for the encoding task is greatly improved, thereby improving the fine-tuning efficiency of the encoding model and the model optimization effect.

[0107] In summary, specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims may be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing may be advantageous.

[0108] The above is the model fine-tuning method provided by the embodiments of this application. Based on the same idea, the embodiments of this application also provide a model fine-tuning device.

[0109] Figure 5 It is a schematic block diagram of a model fine-tuning device according to an embodiment of this application. As Figure 5 shown, the device includes: A code generation module 51, configured to generate codes for the code problem data of the target task through a task large model of the target task, to obtain a plurality of first test codes corresponding to the code problem data; the target task is a unit test-driven development (UTDD) task or a coding task; A determination module 52, configured to test the first test codes, and determine the problem difficulty level of the code problem data according to the test results; A code synthesis module 53, configured to, in response to the problem difficulty level reaching a preset level, synthesize a second test code for the code problem data according to the plurality of first test codes; A model fine-tuning module 54, configured to iteratively fine-tune the task large model according to the code problem data and the second test code, to obtain a target large model.

[0110] Those skilled in the art should understand that Figure 5 the model fine-tuning device in

[0111] By using the device according to the embodiment of the present application, through the task large model of the target task, code generation is performed on the code problem data of the target task to obtain a plurality of first test codes corresponding to the code problem data, and the target task is a UTDD task or a coding task; the first test codes are tested, and the problem difficulty level of the code problem data is determined according to the test results; in response to the problem difficulty level reaching a preset level, according to the plurality of first test codes, a second test code of the code problem data is synthesized, and the task large model is iteratively fine-tuned according to the code problem data and the second test code to obtain a target large model. Since the code problem data with a low difficulty level has little learning space for the task large model to a certain extent, therefore, by screening out the code problem data with a relatively high problem difficulty level (such as reaching the preset level), and performing model fine-tuning according to the code problem data with a relatively high problem difficulty level and its corresponding second test code, in fact, the code problem data that does not need to be learned by the task large model is screened out, and only the code problem data with a high problem difficulty level is retained, which is beneficial to improving the learning ability of the task large model. Moreover, this way of screening code problem data can, for the UTDD task and the coding task, gradually improve the quality of the code problem data retained after each iteration through multiple iteration processes, so that the data used for each fine-tuning of the task large model is of better quality than the data used for the previous iteration fine-tuning. By continuously synthesizing higher-quality model fine-tuning data and continuously improving the learning ability of the model, an excellent target large model with extremely good performance can be obtained through multiple iteration fine-tuning.

[0112] Based on the same idea, the embodiment of the present application also provides an electronic device, as Figure 6 shown. The electronic device may vary greatly due to configuration or performance differences, and may include one or more processors 601 and a memory 602. One or more application programs or data may be stored in the memory 602. Among them, the memory 602 may be short-term storage or persistent storage. The application programs stored in the memory 602 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the electronic device. Further, the processor 601 may be set to communicate with the memory 602 and execute a series of computer-executable instructions in the memory 602 on the electronic device. The electronic device may further include one or more power supplies 603, one or more wired or wireless network interfaces 604, one or more input / output interfaces 605, and one or more keyboards 606.

[0113] Specifically, in this embodiment, the electronic device includes a memory and one or more programs. One or more of the programs are stored in the memory, and one or more of the programs may include one or more modules. Each module may include a series of computer-executable instructions in the electronic device and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions for: Generate multiple first test codes corresponding to the code problem data of the target task through the task large model of the target task. The target task is a UTDD task or a coding task; Test the first test codes and determine the problem difficulty level of the code problem data according to the test results; In response to the problem difficulty level reaching a preset level, synthesize a second test code for the code problem data according to the multiple first test codes; Iteratively fine-tune the task large model according to the code problem data and the second test code to obtain a target large model.

[0114] Adopting the technical solution of the embodiment of the present application, through the task large model of the target task, generate multiple first test codes corresponding to the code problem data of the target task. The target task is a UTDD task or a coding task; test the first test codes and determine the problem difficulty level of the code problem data according to the test results; in response to the problem difficulty level reaching a preset level, synthesize a second test code for the code problem data according to the multiple first test codes, and iteratively fine-tune the task large model according to the code problem data and the second test code to obtain a target large model. Since the code problem data with a low difficulty level does not have a learning space to a certain extent for the task large model, by screening out the code problem data with a higher problem difficulty level (such as reaching the preset level) and performing model fine-tuning according to the code problem data with a higher problem difficulty level and its corresponding second test code, it actually screens out the code problem data that does not need to be learned by the task large model and only retains the code problem data with a high problem difficulty level, which is beneficial to improving the learning ability of the task large model. Moreover, this way of screening code problem data for UTDD tasks and coding tasks can gradually improve the quality of the code problem data retained after each iteration through multiple iteration processes, making the data used for each fine-tuning of the task large model better in quality than the data used for the previous iteration fine-tuning. By continuously synthesizing higher-quality model fine-tuning data and continuously improving the learning ability of the model, an extremely performant target large model can be obtained through multiple iteration fine-tuning.

[0115] An embodiment of the present application also provides a computer-readable storage medium storing one or more computer programs, where the one or more computer programs include instructions that, when executed by an electronic device including a plurality of application programs, can cause the electronic device to execute each process of the above-described model fine-tuning method embodiment, and are specifically used to execute: Generate code for the code problem data of the target task through the task large model of the target task to obtain a plurality of first test codes corresponding to the code problem data; the target task is a UTDD task or a coding task; Test the first test code, and determine the problem difficulty level of the code problem data according to the test results; In response to the problem difficulty level reaching a preset level, synthesize a second test code for the code problem data according to the plurality of first test codes; Iteratively fine-tune the task large model according to the code problem data and the second test code to obtain a target large model.

[0116] Adopting the technical solution of the embodiment of the present application, generate code for the code problem data of the target task through the task large model of the target task to obtain a plurality of first test codes corresponding to the code problem data, where the target task is a UTDD task or a coding task; test the first test code, and determine the problem difficulty level of the code problem data according to the test results; in response to the problem difficulty level reaching a preset level, synthesize a second test code for the code problem data according to the plurality of first test codes, and iteratively fine-tune the task large model according to the code problem data and the second test code to obtain a target large model. Since the code problem data with a low difficulty level has little learning space for the task large model to a certain extent, by screening out the code problem data with a relatively high problem difficulty level (such as reaching the preset level) and performing model fine-tuning according to the code problem data with a relatively high problem difficulty level and its corresponding second test code, it actually filters out the code problem data that does not require the task large model to learn, and only retains the code problem data with a high problem difficulty level, which is beneficial to improving the learning ability of the task large model. Moreover, this way of screening code problem data for UTDD tasks and coding tasks can gradually improve the quality of the retained code problem data after each iteration through multiple iteration processes, making the data used for each fine-tuning of the task large model of better quality than the data used for the previous iteration fine-tuning. By continuously synthesizing higher-quality model fine-tuning data and continuously improving the learning ability of the model, a target large model with excellent performance can be obtained through multiple iteration fine-tuning.

[0117] An embodiment of the present application provides a computer program product, including a computer program, which is executed by a processor to implement each process of the above-mentioned embodiment of the model fine-tuning method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0118] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0119] For the convenience of description, when describing the above devices, they are divided into various units according to functions and described separately. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0120] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0121] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0122] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes.

[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one Figure 1 one process or more processes and / or boxes Figure 1 step of the functions specified in one box or more boxes.

[0124] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0125] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0126] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0127] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0128] This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0129] Each embodiment in this application is described in a progressive manner. For parts that are the same or similar among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for relevant content.

[0130] The above are only the embodiments of this application and are not intended to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modification, equivalent replacement, improvement, etc. made within the principle of this application shall be included within the scope of the claims of this application.

Claims

1. A method for model fine-tuning, comprising: Generating multiple first test codes corresponding to the code problem data of the target task through a task large model of the target task; The target task is a unit test-driven development (UTDD) task or a coding task; Testing the first test codes and determining the problem difficulty level of the code problem data according to the test results; In response to the problem difficulty level reaching a preset level, synthesizing a second test code of the code problem data according to the multiple first test codes; Iteratively fine-tuning the task large model according to the code problem data and the second test code to obtain a target large model.

2. The method according to claim 1, wherein the testing the first test codes and determining the problem difficulty level of the code problem data according to the test results comprises: For each of the first test codes, performing a quality assessment on the first test code according to the code dynamic metrics of the first test code in the running state to obtain a quality assessment metric of the first test code; the quality assessment metric includes at least one of the following: branch coverage rate, line coverage rate; Determining the problem difficulty level of the code problem data according to the quality assessment metrics of each of the first test codes.

3. The method according to claim 2, wherein the synthesizing a second test code of the code problem data according to the multiple first test codes comprises: Extracting valid test codes from the multiple first test codes according to the quality assessment metrics of each of the first test codes; Combining the valid test codes to obtain the second test code.

4. The method according to claim 3, wherein the extracting valid test codes from the multiple first test codes comprises: Extracting at least one first test code whose quality assessment metric meets the first metric requirement from the multiple first test codes as the valid test code; The first metric requirement includes at least one of the following: the quality assessment metric is among the top N, the quality assessment metric reaches a preset metric value; N is an integer greater than or equal to 1.

5. The method according to claim 3, wherein the extracting valid test codes from the multiple first test codes comprises: Selecting the first test code with the highest quality assessment metric from the multiple unselected first test codes; Combining the selected first test codes and determining the quality assessment metric of the combined code obtained after combination; In response to the quality assessment metric of the combined code meeting the second metric requirement, determining the selected first test codes as the valid test code; The second metric requirement includes at least one of the following: the quality assessment metric reaches a preset metric value, the quality assessment metric is the highest; In response to the quality assessment metric of the combined code not meeting the second metric requirement, continuing to select the unselected first test code with the highest quality assessment metric.

6. According to the method described in claim 1, after synthesizing the second test code for the code problem data based on the multiple first test codes, the method further includes: Running the second test code, and determining the quality evaluation index corresponding to the second test code according to the running result of the second test code; When the quality evaluation index corresponding to the second test code meets the requirements of the third index, determining that the second test code meets the task code requirements of the target task; The iterative fine-tuning of the task large model according to the code problem data and the second test code includes: When the second test code meets the task code requirements, performing iterative fine-tuning on the task large model according to the second code data.

7. According to the method described in claim 1, before generating multiple first test codes corresponding to the code problem data of the target task through the task large model of the target task, the method further includes: Obtaining the original code data of the code problem data; Using a requirement description model matching the target task to generate requirement description statements corresponding to the original code data; The requirement description statements are used to describe the functions to be implemented by the original code data; Determining the code problem data according to the requirement description statements.

8. According to the method described in claim 1, the target task is the coding task; the task large model includes a coding model; The iterative fine-tuning of the task large model according to the code problem data and the second test code to obtain a target large model includes: Obtaining target code problem data whose problem difficulty level reaches the preset level and the second test code corresponding to the target code problem data; Inputting the target code problem data into the coding model to obtain multiple first function codes corresponding to the target task problem data; Determining the coding fine-tuning data corresponding to the coding model according to the target code problem data and the multiple first function codes; The coding fine-tuning data includes the target code problem data and its corresponding second function code; Performing iterative fine-tuning on the coding model according to the coding fine-tuning data.

9. According to the method described in claim 8, the determining the coding fine-tuning data corresponding to the coding model according to the target code problem data and the multiple first function codes includes: Determining the quality evaluation index of each first function code; The quality evaluation index includes at least one of the following: compilation pass rate, branch coverage rate, line coverage rate; Selecting the second function codes that meet the requirements of the fourth index from the multiple first function codes according to the quality evaluation index of each first function code; the requirements of the fourth index include at least one of the following: the quality evaluation index is among the top M, the quality evaluation index reaches a preset index value; M is an integer greater than or equal to 1.

10. The method according to claim 1, wherein iteratively fine-tuning the task large model according to the code problem data and the second test code to obtain a target large model comprises: Inputting the test problem data of the target task into the task large model of the current iteration to obtain answer data for the test problem data; Evaluating the quality of the answer data according to the code metrics of the answer data to obtain a quality evaluation metric for the answer data; the quality evaluation metric includes at least one of the following: compilation pass rate, branch coverage rate, line coverage rate; In response to the quality evaluation metric of the answer data meeting the preset metric requirements, determining the task large model of the current iteration as the target large model; in response to the quality evaluation metric of the answer data not meeting the preset metric requirements, continuing to fine-tune the task large model of the current iteration.

11. An electronic device, comprising a processor and a memory electrically connected to the processor, the memory storing a computer program, the processor being configured to call and execute the computer program from the memory to implement the model fine-tuning method according to any one of claims 1-10.

12. A computer-readable storage medium for storing a computer program that can be executed by a processor to implement the model fine-tuning method according to any one of claims 1-10.

13. A computer program product comprising a computer program that is executed by a processor to implement the model fine-tuning method according to any one of claims 1-10.