Large model training optimization method and device, computer equipment and storage medium
By training the initial scoring model, combining the correspondence between the preset scoring model and the interface calling model, the target interface calling model is determined, which solves the problem of inaccurate results when querying API interface information by large models, and improves the accuracy and training efficiency of query results.
Patent Information
- Application Number
- CN202510561131.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-30
AI Technical Summary
When querying API interface information using a big model, directly entering the prompt word will cause inaccurate results and the matching API interface information cannot be effectively found from the API document.
By obtaining the training data set, including interface call question samples, answer samples and their standard ranking information, the initial scoring model is used to score the answer samples, and the initial scoring model is trained according to the differences, the target scoring model is obtained, and the target interface call model is determined based on the correspondence between the preset scoring model and the interface call model is determined to achieve preference alignment training.
It improves the accuracy of API interface information query results, makes the output results of interface calling models more in line with the API field, reduces the number of model training, improves training efficiency, and reduces the computer video memory usage and running rate.
Smart Images

Figure CN120087436A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, computer device, and storage medium for optimizing large model training. Background Art
[0002] With the rapid development of artificial intelligence technology, significant progress has been made in the field of artificial intelligence large models in natural language processing, image processing, etc. Among them, the API (Application Programming Interface) is a collection of agreements and tools for interaction and communication between different components of a software system. During the development or use of the API, it is necessary to search for detailed information such as the required API configuration parameters from a large amount of API documentation. For developers or users, only the information related to functional requirements cannot find the matching API interface information from the API documentation through a simple search query function. Therefore, by inputting a prompt, the large model can be used to search for API interface information.
[0003] However, directly inputting the prompt for querying API interface information into the large model may result in the situation that the obtained results do not belong to the API field at all, leading to inaccurate query results. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, computer-readable storage medium, and computer program product for optimizing large model training, which can improve the accuracy of query results for API interface information.
[0005] In a first aspect, this application provides a method for optimizing large model training, including:
[0006] Obtain a training data set; the training data set includes multiple training sample pairs, each training sample pair includes an interface call problem sample, multiple interface call answer samples corresponding to the interface call problem sample, and standard ranking information of the multiple interface call answer samples; the multiple interface call answer samples are obtained through an initial interface call model;
[0007] Score the interface call answer samples through an initial scoring model, and obtain predicted ranking information of the interface call answer samples according to the score;
[0008] Train the initial scoring model according to the difference between the predicted ranking information and the standard ranking information to obtain a target scoring model;
[0009] Determine a target interface call model according to the corresponding relationship between a preset scoring model and a preset interface call model and the target scoring model.
[0010] In a second aspect, the present application provides an optimization device for large model training, including:
[0011] An acquisition module, configured to acquire a training data set; the training data set includes a plurality of training sample pairs, each training sample pair includes an interface call problem sample, a plurality of interface call answer samples corresponding to the interface call problem sample, and standard ranking information of the plurality of interface call answer samples; the plurality of interface call answer samples are obtained through an initial interface call model;
[0012] A scoring module, configured to score the interface call answer samples through an initial scoring model, and obtain predicted ranking information of the interface call answer samples according to the scores;
[0013] A training module, configured to train the initial scoring model according to the difference between the predicted ranking information and the standard ranking information, and obtain a target scoring model;
[0014] A determination module, configured to determine a target interface call model according to the corresponding relationship between a preset scoring model and a preset interface call model and the target scoring model.
[0015] In a third aspect, the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above method are implemented.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method are implemented.
[0017] In a fifth aspect, the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps in the above method are implemented.
[0018] The above large model training optimization method, device, computer device, computer-readable storage medium, and computer program product train an initial scoring model according to a training data set to obtain a trained target scoring model, and determine a target interface call model that has been aligned and trained according to the target scoring model and the corresponding relationship between the preset scoring model and the preset interface call model, which can achieve preference alignment training for the initial interface call model, enabling the target interface call model to output results that are more in line with the interface call field and improving the accuracy of the output results. In addition, by training the initial scoring model to obtain the target scoring model and determining the target interface call model according to the corresponding relationship between the preset scoring model and the preset interface call model, that is, only one model needs to be trained to obtain a model with preference alignment optimization, without training multiple models, reducing the number of model trainings, improving the training efficiency of the interface call model, while reducing the computer video memory occupancy and improving the computer operation speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 FIG. is an application environment diagram of a large model training optimization method provided by an embodiment of the present application;
[0020] Figure 2 FIG. is a flowchart of a large model training optimization method provided by an embodiment of the present application;
[0021] Figure 3 FIG. is a flowchart of constructing a corresponding relationship between a preset scoring model and a preset interface call model provided by an embodiment of the present application;
[0022] Figure 4 FIG. is a structural block diagram of a large model training optimization device provided by an embodiment of the present application;
[0023] Figure 5 FIG. is an internal structure diagram of a computer device provided by an embodiment of the present application;
[0024] Figure 6 FIG. is an internal structure diagram of another computer device provided by an embodiment of the present application;
[0025] Figure 7 FIG. is an internal structure diagram of a computer-readable storage medium provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0027] The large model training optimization method provided by the embodiments of the present application can be applied to, for example, Figure 1 the application environment shown. Among them, the terminal 102 communicates with the server 104 through a communication network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0028] As Figure 2 shown, the embodiments of the present application provide a large model training optimization method. Taking the method applied to Figure 1 the terminal 102 or the server 104 in as an example for illustration. It can be understood that the computer device can include at least one of the terminal and the server. The method includes the following steps:
[0029] S202. Obtain a training data set; the training data set includes multiple training sample pairs, each training sample pair includes an interface call problem sample, multiple interface call answer samples corresponding to the interface call problem sample, and standard ranking information of the multiple interface call answer samples; the multiple interface call answer samples are obtained through an initial interface call model.
[0030] Among them, the interface call problem sample refers to a problem sample for calling an API interface. The interface call problem sample can be data input by the user in text form, voice form, or picture form. The interface call problem sample can be a sample input into the initial interface call model in order to obtain an interface call answer, which is equivalent to a prompt word input into the initial interface call model. The interface call problem sample can be, for example, "Please find the API interface configuration function that can implement the XX function", etc. The interface call answer sample corresponds to the interface call problem sample, that is to say, the interface call answer sample is an interface call answer generated based on the interface call problem sample. It is easy to understand that the format of the interface call answer sample can also be text, voice, or picture. The initial interface call model can be any artificial intelligence model. The artificial intelligence model includes large models. The initial interface call model can generate an answer to the input problem.
[0031] Exemplarily, the interface call problem samples can be input into the initial interface call model, and multiple different interface call answer samples can be obtained. Then, the multiple interface call answer samples are ranked through manual or ranking algorithms to obtain the standard ranking information of the multiple interface call answer samples. It is easy to understand that the standard ranking information can be marked on the corresponding interface call answer samples.
[0032] It is easy to understand that the more the number of training sample pairs included in the training dataset and the richer the types, the more beneficial it is for the preference alignment training of the model. The training dataset for final training can be obtained by data augmentation of the manually sorted training dataset.
[0033] S204. Score the interface call answer samples through the initial scoring model, and obtain the predicted ranking information of the interface call answer samples according to the scores.
[0034] Among them, the initial scoring model refers to a model that can evaluate answers, and the initial scoring model can also be an artificial intelligence model. Usually, the initial scoring model and the initial interface call model are not the same artificial intelligence model. The initial scoring model can score the interface call answer samples based on the evaluation method. The predicted ranking information is used to represent the ranking information obtained by scoring based on the initial scoring model.
[0035] Exemplarily, the interface call answer samples can be scored through the initial scoring model, and the multiple interface call answer samples are ranked according to the scores to obtain the predicted ranking information of the interface call answer samples.
[0036] S206. Train the initial scoring model according to the difference between the predicted ranking information and the standard ranking information to obtain the target scoring model.
[0037] Among them, the target scoring model refers to the model obtained by performing preference alignment training on the initial scoring model.
[0038] Exemplarily, the initial scoring model can be trained according to the difference between the predicted ranking information and the standard ranking information, so that when the scoring model can output predicted ranking information approximately consistent with the standard ranking information, the target scoring model is obtained.
[0039] In some embodiments, during the process of training the initial scoring model, the initial scoring model is guided to learn the ranking situation in the standard ranking information, that is, the predicted ranking information corresponding to the interface call answer samples with a relatively higher ranking in the standard ranking information also needs to be greater than the predicted ranking information corresponding to the interface call answer samples with a relatively lower ranking in the standard ranking information, and the probability of being greater needs to be higher than the preset threshold to obtain the target scoring model.
[0040] In some embodiments, in addition to considering the difference between the predicted ranking information and the standard ranking information, the initial scoring model can also be trained in combination with the video memory capacity of the computer to obtain the target scoring model. For example, the target scoring model is obtained when the difference between the predicted ranking information and the standard ranking information is less than the difference threshold and the video memory capacity matching the initial scoring model is less than the capacity threshold.
[0041] S208. Determine the target interface invocation model according to the correspondence between the preset scoring model and the preset interface invocation model and the target scoring model.
[0042] Among them, the target interface invocation model refers to the interface invocation model obtained by performing preference alignment optimization on the initial interface invocation model. It should be noted that the preset scoring model is a pre-selected artificial intelligence model that can score the input answers. The preset interface invocation model is a pre-selected artificial intelligence model that can output interface invocation information.
[0043] Exemplarily, the correspondence between the preset scoring model and the preset interface invocation model can be constructed based on the initial alignment optimization objective function. Among them, the initial alignment optimization objective function is determined by the preset interface invocation model, the preset reference model, and the preset scoring model. For example, the initial alignment optimization objective function can be determined based on the difference between the preset interface invocation model and the preset reference model and the score of the preset scoring model for the interface invocation answer generated by the preset interface invocation model. The optimization objective is to minimize the difference between the preset interface invocation model and the preset reference model, while maximizing the score of the preset scoring model for the interface invocation answer generated by the preset interface invocation model. For example, the objective function of preference alignment optimization using the PPO (Proximal Policy Optimization) algorithm can be used as the initial alignment optimization objective function. The correspondence between the preset scoring model and the preset interface invocation model can be obtained by transforming the initial alignment optimization objective function and working to remove or replace the preset reference model.
[0044] It can be seen that in the embodiments of the present application, by training the initial scoring model according to the training data set, the trained target scoring model is obtained. According to the target scoring model and the corresponding relationship between the preset scoring model and the preset interface call model, the target interface call model after alignment training is determined, which can indirectly implement the preference alignment training for the initial interface call model, so that the interface call model can output results that are more in line with the interface call field, improving the accuracy of the output results. In addition, by training the initial scoring model to obtain the target scoring model and determining the target interface call model according to the corresponding relationship between the preset scoring model and the preset interface call model, that is, only one model needs to be trained to obtain the model optimized by preference alignment, without training multiple models, reducing the number of model trainings, improving the training efficiency of the interface call model, while reducing the computer video memory occupancy and improving the computer operation speed.
[0045] In some embodiments, as Figure 3 shown, before determining the target interface call model according to the corresponding relationship between the preset scoring model and the preset interface call model and the target scoring model, the above method further includes:
[0046] S302. Obtain the initial alignment optimization objective function; wherein, the initial alignment optimization objective function is determined by the preset interface call model, the preset reference model, and the preset scoring model.
[0047] Among them, the initial alignment optimization objective function is the optimization objective function for preferentially aligning and optimizing the preset interface call model to be trained. In the process of optimizing and training the preset interface call model, some auxiliary models can be used to participate in the preferential alignment and optimization of the preset interface call model to be trained. The auxiliary models include the preset reference model and the preset scoring model. Among them, the preset interface call model and the preset reference model can be any different artificial intelligence models. The preset scoring model is a model that can score the results output by the preset interface call model. Exemplarily, in the initial alignment optimization objective function, the KL divergence (Kullback-Leibler Divergence) between the preset interface call model and the preset reference model can be included, which can prevent the output result of the preset interface call model from deviating too far from the output result of the preset reference model during the training of the preset interface call model, resulting in the situation that the model training does not converge.
[0048] In some embodiments, the initial alignment optimization objective function includes the KL divergence between the output result of a preset interface call model and the output result of a preset reference model, and the output result of a preset scoring model. The preset interface call model is the model to be trained, the preset scoring model is the trained model, and the preset reference model is the model used as a reference for training the preset interface call model. For example, the result of the initial alignment optimization objective function is the difference between the output result of the preset scoring model and the KL divergence between the output results of the preset interface call model and the preset reference model. The initial alignment optimization objective function can be shown as the following formula (1).
[0049] Formula (1)
[0050] Wherein, represents the preset interface call model; represents the preset scoring model; represents the preset reference model; x represents the input prompt word, i.e., the interface call question; D represents the set corresponding to x; y represents the interface call answer corresponding to the input x, represents a hyperparameter, represents the KL divergence, and E represents the optimization objective. The first half of the initial alignment optimization objective function is equivalent to maximizing the output value of the preset scoring model, and the second half is to minimize the difference between the preset interface call model and the preset reference model. In other words, when the output value of the preset scoring model is greater than the preset scoring threshold, and the difference between the preset interface call model and the preset reference model is less than the difference threshold, the corresponding optimization objective is achieved, and the optimized interface call model is obtained. It is easy to understand that the initial alignment optimization objective function obtained by transforming formula (1) should also be within the protection scope corresponding to the embodiments of the present application. For example, formula (2) obtained by transforming formula (1) is shown as follows.
[0051] Formula (2)
[0052] The maximization problem in formula (2) can be converted into a minimization problem. Then, formula (2) can be transformed into the following formula (3)
[0053] Formula (3)
[0054] That is to say, the initial alignment optimization objective function can be shown as formula (1), formula (2), or formula (3).
[0055] S304. Substitute the obtained partial function into the initial alignment optimization objective function to obtain the objective function to be optimized; wherein, the partial function is a function with the preset scoring model and the preset reference model as independent variables.
[0056] In practical application scenarios, the role of the partial function is to remove the preset scoring model from the optimization process. That is to say, during the transformation process, the parameters of the preset scoring model will not be changed accordingly.
[0057] Exemplarily, the partial function can be represented by the following formula (4).
[0058] Formula (4)
[0059] Exemplarily, in formula (4) represents all possible y that the preset reference model may generate under the premise of a certain prompt word X. Therefore , and the partial function represented by formula (4) is a function of x, and has nothing to do with . Then, substituting the partial function shown in formula (4) into the initial alignment optimization objective function shown in formula (3), the objective function to be optimized is obtained as shown in the following formula (5).
[0060] Formula (5)
[0061] S306. Determine the explicit solution of the preset interface call model according to the objective function to be optimized.
[0062] Among them, the explicit solution refers to that the solution of a mathematical equation or problem can be directly expressed in the explicit form of a known function, usually manifested as the functional relationship between the dependent variable and the independent variable can be clearly expressed. Determining the explicit solution of the interface call model according to the objective function to be optimized is equivalent to directly calculating the parameters of the preset interface call model according to the preset reference model and the preset scoring model.
[0063] Exemplarily, the explicit solution of the preset interface call model determined according to the objective function to be optimized corresponding to formula (5) is shown in the following formula (6). Among them, in the process of solving the explicit solution, to minimize the expectation in formula (5), since has nothing to do with the optimization of , in the gradient descent algorithm, whether the update of the model is positive or negative will cause the gradient to be updated. Therefore, the smallest update is 0, that is, to make equal to 0, log1 = 0, that is, the explicit solution shown in formula (6) of is obtained.
[0064] Formula (6)
[0065] S308. Determine the corresponding relationship between the preset scoring model and the preset interface call model according to the explicit solution of the preset interface call model.
[0066] After obtaining the explicit solution of the preset interface call model, it is equivalent to obtaining the correspondence between the preset interface call model, the preset scoring model, and the preset reference model. Then, after removing the preset reference model, the correspondence between the preset interface call model and the preset scoring model can be obtained.
[0067] Exemplarily, the maximum likelihood estimation result corresponding to the preset interface call model can be calculated, and by replacing the preset reference model with this maximum likelihood estimation result, the correspondence between the preset scoring model and the preset interface call model can be obtained.
[0068] It can be seen that in this embodiment, through transformation processing using the partial function and the initial alignment optimization objective function to obtain the explicit solution of the preset interface call model, and based on this explicit solution of the preset interface call model to determine the correspondence between the preset scoring model and the preset interface call model, it can ensure that the output of the preset interface call model to be trained is a valid probability distribution, and can improve the accuracy of the correspondence between the preset scoring model and the preset interface call model.
[0069] In some embodiments, determining the correspondence between the preset scoring model and the preset interface call model according to the explicit solution of the preset interface call model includes:
[0070] Converting the explicit solution of the preset interface call model into an expression with the preset scoring model as the dependent variable;
[0071] Generating the maximum likelihood estimation result corresponding to the preset interface call model;
[0072] Replacing the preset reference model in the expression with the maximum likelihood estimation result to obtain the correspondence between the preset scoring model and the preset interface call model.
[0073] Among them, the explicit solution is an expression with the preset interface call model as the dependent variable. By transforming the explicit solution, an expression with the preset scoring model as the dependent variable can be obtained. Exemplarily, transforming formula (6) can obtain an expression with the preset interface call model as the dependent variable as shown in formula (7) below.
[0074] Formula (7)
[0075] It can be seen from formula (7) that it is equivalent to substituting the preset interface call model to be optimized into the optimization objective of the preset scoring model, that is, the training steps of the preset interface call model can be incorporated into the training of the preset scoring model. The maximum likelihood estimation result corresponding to the preset interface call model is generated as shown in formula (8) below.
[0076] Formula (8)
[0077] During the training process, assume that an interface call problem is input into a preset interface call model to be trained, and two interface call answers are generated. The better one is determined manually as y w , and the worse one is y l , and during the optimization process, it satisfies the ranking corresponding to the preset scoring model as , but it doesn't necessarily mean that the likelihood ranking satisfies , therefore, the ranking of the preset scoring model can be replaced by the likelihood ranking, that is, the preset reference model is replaced by the maximum likelihood estimation result, and the following formula (9) is obtained.
[0078] Formula (9)
[0079] Formula (9) represents the corresponding relationship between the preset scoring model and the preset interface call model.
[0080] It can be seen that in this embodiment, by replacing the preset reference model with the maximum likelihood estimation result corresponding to the preset interface call model in the expression with the preset scoring model as the dependent variable, the corresponding relationship between the preset scoring model and the preset interface model is obtained, and the result of the likelihood ranking can be made consistent with the manual ranking, thereby improving the accuracy of the corresponding relationship between the preset scoring model and the preset interface model.
[0081] In some embodiments, the predicted ranking information includes a first predicted ranking and a second predicted ranking; training an initial scoring model according to the difference between the predicted ranking information and the standard ranking information to obtain a target scoring model, including:
[0082] Respectively obtain the first predicted ranking obtained by predicting the first interface call answer sample through the initial scoring model, and the second predicted ranking obtained by predicting the second interface call answer sample; wherein, the ranking of the first interface call answer sample in the standard ranking information is before the ranking of the second interface call answer sample in the standard ranking information;
[0083] Determine the probability that the first predicted ranking is greater than the second predicted ranking;
[0084] When the probability is greater than a preset probability threshold, determine the target scoring model.
[0085] Among them, the initial scoring model can be an artificial intelligence model. The initial scoring model scores the interface call answer samples, and can score according to the occurrence frequency of target terms in the interface call answer samples, where the target terms are used to represent the standard answers. Alternatively, the initial scoring model can calculate the keyword overlap rate between the interface call answer sample and the standard answer, and score the interface call answer sample according to the keyword overlap rate. Alternatively, the initial scoring model can calculate the edit distance (Levenshtein) or cosine similarity between the interface call answer sample and the standard answer, and score the interface call answer sample according to the edit distance or cosine similarity. Alternatively, the initial scoring model can count data such as the browsing frequency, click frequency, and answer accuracy feedback of users for different interface call answer samples, and score the interface call answer samples according to data such as the browsing frequency, click frequency, and answer accuracy feedback of users. The predicted ranking information can be understood as the predicted ranking of the interface call answer samples.
[0086] Exemplarily, the initial scoring model scores the first interface call answer sample to obtain a first answer score, and scores the second interface call answer sample to obtain a second score, ranks the first answer score and the second answer score, and obtains a first predicted ranking corresponding to the first answer score and a second predicted ranking corresponding to the second answer score. The ranking of the first interface call answer sample in the standard ranking information is before the ranking of the second interface call answer sample in the standard ranking information. Therefore, during the training process, efforts should also be made to make the first predicted ranking obtained by the initial scoring model greater than the second predicted ranking. Therefore, calculate the probability that the first predicted ranking is greater than the second predicted ranking. When this probability is greater than a preset probability threshold, the target scoring model is obtained. Among them, the preset probability threshold can be set according to the actual application scenario. For example, the preset probability threshold can be 0.95, 0.96, or 0.99, etc.
[0087] In the actual application scenario, the same interface call question sample may correspond to two or more interface call answer samples. The probability calculation of the predicted ranking can be performed for different interface call answer samples corresponding to the same interface call question sample, or the probability calculation of the predicted ranking can be performed for the interface call answer samples corresponding to multiple interface call question samples. Specifically, it is possible to count the probability that the first predicted ranking is greater than the second predicted ranking among the predicted rankings corresponding to different interface call answer samples of the same interface call question sample. Alternatively, it is also possible to count the probability that the first predicted ranking is greater than the second predicted ranking among the predicted rankings of the interface call answer samples corresponding to different interface call question samples. Until the counted probability is greater than the preset probability threshold, the target scoring model is obtained.
[0088] In one example, the Bradley-Terry model (a statistical model of sports games) can be used for modeling, and formula (10) is obtained.
[0089] Formula (10)
[0090] Among them, formula (10) can calculate y w Greater than y l probability.
[0091] In order to make y w As much as possible greater than y l , for the entire training data set , N represents the number of training samples. The optimization objective of the scoring model can be expressed as follows:
[0092] Formula (11)
[0093] Will The specific form of is formula (10) is substituted into formula (11), and the optimization objective of the scoring model can be obtained as shown in the following formula (12).
[0094] Formula (12)
[0095] The optimization objective of the scoring model is usually described as a loss problem. In formula (12), when the scoring model is applied to y w The ranking is greater than y l When the ranking of is maximized, the corresponding loss is minimized.
[0096] It can be seen that in this embodiment, by determining the probability that the predicted ranking obtained by the initial scoring model for the interface call answer samples that are ranked higher in the standard ranking information is greater than the predicted ranking obtained by predicting the interface call answer samples that are ranked lower in the standard ranking information, when the probability is greater than the preset probability threshold, it means that the predicted ranking obtained by the initial scoring model for the interface call answer samples is relatively consistent with the ranking in the standard ranking information, that is, the ranking of the interface call answer samples by the initial scoring model is consistent with the standard ranking, thereby obtaining a more accurate target scoring model.
[0097] In some embodiments, the above method further comprises:
[0098] Obtaining respectively a first score obtained by scoring the first interface call answer sample through the initial scoring model, and a second score obtained by scoring the second interface call answer sample;
[0099] Determine the training loss based on the difference between the first score and the second score and the score difference term;
[0100] The initial scoring model is trained according to the training loss to obtain the target scoring model.
[0101] Among them, the score difference term is used to reduce the possibility of overfitting in model training. If the difference between the first score and the second score is too large, the model may be biased towards the answer of the first interface call answer sample corresponding to the first score; conversely, if the difference between the first score and the second score is too small, the model may be biased towards the answer of the second interface call answer sample corresponding to the second score, which will affect the generalization performance of the model. Therefore, the overfitting of model training is limited by the score difference term. The score difference term can be a constant greater than zero, usually in the range of (0,1]. The score difference term can be, for example, 0.2 or 0.3.
[0102] Exemplarily, the initial scoring model is trained according to the training loss, that is, the parameters of the initial scoring model are adjusted so that the training loss becomes smaller, until the training loss is less than the loss threshold, and the target scoring model is obtained.
[0103] In some embodiments, the specific form of the initial scoring model (i.e., replacing the preset scoring model with the initial scoring model in formula (9)) can be substituted into formula (12), so as to realize the conversion of the optimization target of the initial scoring model into the optimization target of the initial interface call model, as shown in the following formula (13).
[0104] Formula (13)
[0105] Since the difference between the two classes affects the generalization ability of the model classifier, in the preference alignment training, these two classes are the winning or losing response of a single input, and the score difference term is introduced in the optimization objective of the initial interface call model. , the final optimization target of the initial interface call model can be obtained as shown in the following formula (14).
[0106] Formula (14)
[0107] It can be seen that in this embodiment, by determining the training loss of the initial scoring model based on the difference between the score of the first interface call answer sample that ranks higher in the standard ranking information and the score of the second interface call answer sample that ranks lower in the standard ranking information according to the initial scoring model, as well as the score difference item, it is possible to reduce overfitting in the model training process and improve the generalization ability of the target scoring model.
[0108] In some embodiments, the initial scoring model is trained according to the difference between the predicted ranking information and the standard ranking information to obtain a target scoring model, including:
[0109] Determine the video memory capacity that matches the initial scoring model;
[0110] Determine the training conditions based on the difference between the predicted ranking information and the standard ranking information and the video memory capacity matching the initial scoring model;
[0111] Train the initial scoring model, and obtain the target scoring model when the training conditions are met.
[0112] Among them, the video memory capacity (GPU Memory Capacity) refers to the capacity of the video memory on the computer graphics card (Graphics Processing Unit, GPU), and the size of the video memory capacity determines the ability of the video memory to temporarily store data. The video memory capacity of the graphics card can be, for example, 128MB, 256MB, 512MB, 1024MB, 2GB, 4GB, 8GB or 1TB, etc. The video memory capacity matching the initial scoring model can be understood as the video memory capacity occupied by the initial scoring model during operation.
[0113] Exemplarily, the training conditions include a difference condition and a video memory condition. During the training process of the initial scoring model, the video memory capacity matching the initial scoring model can be calculated in real time. When the difference between the predicted ranking information and the standard ranking information meets the difference condition, and the video memory capacity matching the initial scoring model meets the video memory condition, the target scoring model is obtained. Among them, the difference condition can be, for example, less than or equal to a difference threshold, and the video memory condition can be, for example, less than or equal to a preset video memory capacity. Among them, the difference threshold and the preset video memory capacity can be set according to the actual application scenario. For example, the difference threshold is 0.2, 0.1 or 0.01, etc. The preset video memory capacity can be less than or equal to the maximum video memory capacity of the graphics card in the computer. For example, the preset video memory capacity is 50%, 70% or 80% of the computer graphics card capacity, etc.
[0114] It can be seen that in this embodiment, by determining the training conditions of the initial scoring model based on the difference between the predicted ranking information and the standard ranking information and the video memory capacity matching the initial scoring model, and obtaining the target scoring model when the training conditions are met, it is possible to use the video memory capacity occupied by the initial scoring model as a training condition to participate in the training of the initial scoring model, so that the video memory capacity occupied by the trained model meets the preset conditions, and effectively control the video memory capacity occupied by the trained model.
[0115] In an exemplary application scenario, input the interface call problem sample x into the initial interface call model to obtain two interface call answer samples y w and y l , and it is easy to understand that in other application scenarios, three or more interface call answer samples may also be obtained. This example is described with two as an example. Manually for yw and y l are ranked, and the standard ranking information is y w greater than y l . The initial scoring model is used to score the answer samples y of the two interface calls respectively w and y l to obtain the first score corresponding to y w , and the second score corresponding to y l . According to the first score and the second score, the predicted ranking information of y w and y l can be obtained. According to the difference between the predicted ranking information and the standard ranking information, the initial scoring model is trained until the difference between the predicted ranking information and the standard ranking information is less than the difference threshold, and the target scoring model is obtained. Then, according to the corresponding relationship between the preset scoring model and the preset interface call model and the target scoring model, the target interface call model is determined. Alternatively, according to the difference between the predicted ranking information and the standard ranking information, the initial interface call model can be trained until the difference between the predicted ranking information and the standard ranking information is less than the difference threshold, and the target interface call model is obtained.
[0116] In some embodiments, during the alignment training of the initial interface call model by the PPO algorithm, the coordinated operation of multiple models is required, resulting in problems such as large video memory occupation, many training hyperparameters, and unstable training during the training process. The optimization target of the PPO algorithm can be used as the initial alignment optimization objective function. By transforming the optimization target of the PPO algorithm, a partial function with the preset scoring model and the preset reference model as dependent variables is configured, and the maximum likelihood estimation result corresponding to the preset interface call model is determined. According to the partial function and the maximum likelihood estimation result, the optimization target of the PPO algorithm is transformed to obtain the corresponding relationship between the preset scoring model and the preset interface call model, that is, the relevant parameters of the preset reference model are eliminated. Thus, only the initial scoring model needs to be trained to obtain the target scoring model. According to the corresponding relationship between the preset scoring model and the preset interface call model, the target interface call model corresponding to the target scoring model can be obtained, which can reduce the type of training data, reduce the number of trained models and the number of hyperparameters, reduce the tuning cost, reduce the video memory requirements for multi-model collaboration in PPO, enable the training of the model to be completed with less video memory resources, and improve the computer operation speed.
[0117] When actually using the target interface call model, inputting the interface call problem into the target interface call model can obtain an interface call answer that is more consistent with the API field. The interface call answer can include information such as the configuration function, class, return type, and parameters of the required API interface to be called.
[0118] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0119] Based on the same inventive concept, an embodiment of the present application also provides a large model training optimization device. The implementation solution provided by this device to solve problems is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the large model training optimization device provided below can refer to the limitations on the large model training optimization method in the above text, and will not be repeated here.
[0120] As Figure 4 shown, an embodiment of the present application provides a large model training optimization device 400, including:
[0121] An acquisition module 402, configured to acquire a training data set; the training data set includes multiple training sample pairs, each training sample pair includes an interface call problem sample, multiple interface call answer samples corresponding to the interface call problem sample, and standard ranking information of the multiple interface call answer samples; the multiple interface call answer samples are obtained through an initial interface call model;
[0122] A scoring module 404, configured to score the interface call answer samples through an initial scoring model, and obtain predicted ranking information of the interface call answer samples according to the scoring;
[0123] A training module 406, configured to train the initial scoring model according to the difference between the predicted ranking information and the standard ranking information to obtain a target scoring model;
[0124] A determination module 408, configured to determine a target interface call model according to the correspondence between a preset scoring model and a preset interface call model and the target scoring model.
[0125] In some embodiments, the above device further includes a construction module, and the construction module is configured to: obtain an initial alignment optimization objective function; the initial alignment optimization objective function is determined by calling a model, a preset reference model, and a preset scoring model through a preset interface; substitute the obtained partial function into the initial alignment optimization objective function to obtain an objective function to be optimized; the partial function is a function with the preset scoring model and the preset reference model as independent variables; determine the display solution of the preset interface call model according to the objective function to be optimized; determine the correspondence between the preset scoring model and the preset interface call model according to the display solution of the preset interface call model.
[0126] In some embodiments, in terms of determining the correspondence between the preset scoring model and the preset interface call model according to the display solution of the preset interface call model, the construction module is specifically configured to:
[0127] Convert the display solution of the preset interface call model into an expression with the preset scoring model as the dependent variable;
[0128] Generate the maximum likelihood estimation result corresponding to the preset interface call model;
[0129] Replace the preset reference model in the expression with the maximum likelihood estimation result to obtain the correspondence between the preset scoring model and the preset interface call model.
[0130] In some embodiments, the prediction ranking information includes a first prediction ranking and a second prediction ranking; in terms of training the initial scoring model according to the difference between the prediction ranking information and the standard ranking information to obtain a target scoring model, the training module 406 is specifically configured to:
[0131] Obtain the first prediction ranking obtained by predicting the first interface call answer sample through the initial scoring model and the second prediction ranking obtained by predicting the second interface call answer sample respectively; the ranking of the first interface call answer sample in the standard ranking information is before the ranking of the second interface call answer sample in the standard ranking information;
[0132] Determine the probability that the first prediction ranking is greater than the second prediction ranking;
[0133] When the probability is greater than a preset probability threshold, determine the target scoring model.
[0134] In some embodiments, the training module 406 is further configured to:
[0135] Obtain the first score obtained by scoring the first interface call answer sample through the initial scoring model and the second score obtained by scoring the second interface call answer sample respectively; determine the training loss according to the difference between the first score and the second score and the score difference term; train the initial scoring model according to the training loss to obtain the target scoring model.
[0136] In some embodiments, in training the initial scoring model according to the difference between the predicted ranking information and the standard ranking information to obtain the target scoring model, the training module 406 is specifically configured to:
[0137] Determine the video memory capacity matching the initial scoring model;
[0138] Determine the training conditions according to the difference between the predicted ranking information and the standard ranking information and the video memory capacity matching the initial scoring model;
[0139] Train the initial scoring model, and obtain the target scoring model when the training conditions are met.
[0140] Each module in the above large model training optimization device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of the processor, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0141] In some embodiments, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to large model training optimization. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements the steps in the above large model training optimization method.
[0142] In some embodiments, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements the steps in the above-mentioned large model training optimization method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen; the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0143] Those skilled in the art can understand that Figure 5 or Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0144] In some embodiments, a computer device is provided. The computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the steps in the above-mentioned method embodiments.
[0145] In some embodiments, as Figure 7 shown, an internal structure diagram of a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. When the computer program is executed by the processor, it implements the steps in the above-mentioned method embodiments.
[0146] In some embodiments, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by the processor, it implements the steps in the above-mentioned method embodiments.
[0147] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0148] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0149] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0150] The above embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A large model training optimization method, characterized in that: include: Get the training dataset; The training data set includes a plurality of training sample pairs, each of which includes an interface call question sample, a plurality of interface call answer samples corresponding to the interface call question sample, and standard ranking information of the plurality of interface call answer samples; the plurality of interface call answer samples are obtained through an initial interface call model; Scoring the interface call answer samples using an initial scoring model, and obtaining predicted ranking information of the interface call answer samples according to the scores; Training the initial scoring model according to the difference between the predicted ranking information and the standard ranking information to obtain a target scoring model; The target interface calling model is determined according to the corresponding relationship between the preset scoring model and the preset interface calling model and the target scoring model.
2. The method according to claim 1, characterized in that Before determining the target interface call model according to the correspondence between the preset scoring model and the preset interface call model and the target scoring model, the method further includes: Obtaining an initial alignment optimization objective function; the initial alignment optimization objective function is determined by a preset interface calling model, a preset reference model, and a preset scoring model; Substituting the obtained partial function into the initial alignment optimization objective function to obtain the objective function to be optimized; the partial function is a function with the preset scoring model and the preset reference model as independent variables; Determining a display solution of the preset interface calling model according to the objective function to be optimized; According to the display solution of the preset interface calling model, a corresponding relationship between the preset scoring model and the preset interface calling model is determined.
3. The method according to claim 2, characterized in that The determining, according to the display solution of the preset interface calling model, the corresponding relationship between the preset scoring model and the preset interface calling model comprises: Converting the display solution of the preset interface calling model into an expression with the preset scoring model as a dependent variable; Generate a maximum likelihood estimation result corresponding to the preset interface calling model; The preset reference model in the expression is replaced with the maximum likelihood estimation result to obtain a corresponding relationship between the preset scoring model and the preset interface calling model.
4. The method according to claim 1, characterized in that: The predicted ranking information includes a first predicted ranking and a second predicted ranking; and the training of the initial scoring model according to the difference between the predicted ranking information and the standard ranking information to obtain a target scoring model includes: The first predicted ranking obtained by predicting the first interface call answer sample through the initial scoring model, and the second predicted ranking obtained by predicting the second interface call answer sample are respectively obtained; the ranking of the first interface call answer sample in the standard ranking information is before the ranking of the second interface call answer sample in the standard ranking information; determining a probability that the first predicted rank is greater than the second predicted rank; When the probability is greater than a preset probability threshold, a target scoring model is determined.
5. The method according to claim 4, characterized in that The method further comprises: respectively obtaining a first score obtained by scoring the first interface call answer sample through the initial scoring model, and a second score obtained by scoring the second interface call answer sample; Determining a training loss according to a difference between the first score and the second score and a score difference term; The initial scoring model is trained according to the training loss to obtain a target scoring model.
6. The method according to claim 1, characterized in that The step of training the initial scoring model according to the difference between the predicted ranking information and the standard ranking information to obtain a target scoring model comprises: Determining a video memory capacity that matches the initial scoring model; determining a training condition according to a difference between the predicted ranking information and the standard ranking information and a video memory capacity that matches the initial scoring model; The initial scoring model is trained to obtain a target scoring model when the training conditions are met.
7. A large model training optimization device, characterized in that: include: An acquisition module is used to obtain a training data set; The training data set includes a plurality of training sample pairs, each of which includes an interface call question sample, a plurality of interface call answer samples corresponding to the interface call question sample, and standard ranking information of the plurality of interface call answer samples; the plurality of interface call answer samples are obtained through an initial interface call model; A scoring module, used to score the interface call answer sample through an initial scoring model, and obtain predicted ranking information of the interface call answer sample according to the scoring; A training module, used for training the initial scoring model according to the difference between the predicted ranking information and the standard ranking information to obtain a target scoring model; The determination module is used to determine the target interface calling model according to the corresponding relationship between the preset scoring model and the preset interface calling model and the target scoring model.
8. The device according to claim 7, characterized in that The device also includes a construction module, which is used to: obtain an initial alignment optimization objective function; the initial alignment optimization objective function is determined by a preset interface call model, a preset reference model and a preset scoring model; substitute the obtained partial function into the initial alignment optimization objective function to obtain an objective function to be optimized; the partial function is a function with the preset scoring model and the preset reference model as independent variables; determine a display solution of the preset interface call model according to the objective function to be optimized; According to the display solution of the preset interface calling model, a corresponding relationship between the preset scoring model and the preset interface calling model is determined.
9. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
API (Application Program Interface) recommendation method based on prompt learning and dual-information source fusion
CN117034135A
Method for two-stage fine tuning of large language model agent
CN118350414A
Apparatus and method for training a machine learning model
US12259864B1