Adaptive data matching method and device for instruction supervision fine tuning

By searching and fitting multiple reference language models, a proportion prediction model is generated, and the problem of low accuracy of data proportion setting in instruction supervision and fine-tuning of large language model is solved, and efficient and accurate task ratio is achieved, which is suitable for large language models of multiple scales.

CN120353889APending Publication Date: 2025-07-22PENG CHENG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510287685.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, during the instruction supervision and fine-tuning process of large language models, the proportion of task data mainly depends on experience setting, resulting in low accuracy, and in the case of multi-task mixing, tasks have complex mutual influences, making it difficult to determine the optimal data ratio.

Method used

By acquiring multiple reference language models, performing parameter search and model fitting one by one, generating a proportional prediction model, and using this model to automatically predict the task ratio of the target large language model, reducing the search and training costs of large-scale parameter models.

Benefits of technology

It significantly improves the accuracy and efficiency of instruction supervision fine-tuning, is suitable for large language models of different sizes, is highly versatile and scalable, and reduces training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353889A_ABST
    Figure CN120353889A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a self-adaptive data matching method and device for instruction supervision fine tuning, and relates to the technical field of artificial intelligence. According to the method, for a reference language model, at least one round of parameter search is carried out based on a baseline weight interval of task types to obtain a matching parameter, and a task matching of the reference language model is obtained according to the matching parameter of each task type; and performing model fitting on the model parameter quantity of the reference language model and the corresponding task ratio to obtain a proportion prediction model, and performing prediction by using the proportion prediction model based on the model parameter quantity of the target large language model to obtain a target task ratio. The method comprises the following steps: performing optimal search of task matching on a plurality of models with small parameter quantities, generating a plurality of groups of data pairs of model parameter quantities and task matching, training a proportion prediction model according to the data pairs, and performing automatic prediction on a target large language model with large-scale parameters to obtain a target task proportion with relatively high reliability. And the performance and efficiency of instruction supervision fine tuning are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to an adaptive data ratio method and device for instruction supervised fine-tuning. Background Art

[0002] The training process of large language models is divided into two stages: First, unsupervised pre-training is carried out on a large-scale text dataset to learn the general representation and knowledge of the language; Second, by performing instruction supervised fine-tuning on various task data, the instruction following ability and task processing performance of the model are improved. Instruction supervised fine-tuning designs data samples of various tasks in the form of text generation, and all tasks share the same task form, training objective, and loss function during the fine-tuning process. Different from usual multi-task learning, the instruction supervised fine-tuning of large language models involves a larger number of tasks and higher task diversity, so it is necessary to use as little data as possible for fine-tuning while ensuring performance.

[0003] In related technologies, according to the specific application direction of the large language model, data ratios are set for different tasks during the instruction supervised fine-tuning process, so that the sample data amounts corresponding to different tasks are different, thereby balancing the performance of each task and reducing the training cost. However, the current data ratios of different tasks are mainly set based on experience and completely rely on developers' understanding of the data, resulting in a low accuracy of instruction supervised fine-tuning. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose an adaptive data ratio method, device, equipment, and storage medium for instruction supervised fine-tuning, so as to improve the accuracy and reliability of the instruction supervised fine-tuning results of large language models.

[0005] To achieve the above purpose, the first aspect of the embodiments of this application proposes an adaptive data ratio method for instruction supervised fine-tuning, including:

[0006] Obtain multiple reference language models. The data for instruction supervised fine-tuning of each reference language model includes multiple task types, each task type includes a corresponding baseline weight interval, and at least two of the reference language models have different model parameters;

[0007] Select the reference language models one by one. For the task type, perform at least one round of parameter search based on the baseline weight interval to obtain ratio parameters, and obtain the task ratio of the reference language model according to the ratio parameters corresponding to each task type;

[0008] Obtain the task ratio of each reference language model, and perform model fitting on the model parameters of the reference language model and the corresponding task ratio to obtain a ratio prediction model;

[0009] Based on the number of model parameters of the target large language model, use the ratio prediction model to predict the target task ratio, and the number of model parameters of the reference language models are all smaller than that of the target large language model.

[0010] In some embodiments, obtaining the ratio parameters through at least one round of parameter search based on the baseline weight interval includes:

[0011] In the process of each round of parameter search, obtain the search interval of the current round, and the initial value of the search interval is the baseline weight interval;

[0012] Determine the search strategy of the current round, and search in the search interval according to the search strategy to obtain the interval information of the current round, and use the interval information as the search interval of the next round;

[0013] After meeting the iteration termination condition, obtain the ratio parameters according to the interval information obtained from the last round of parameter search.

[0014] In some embodiments, determining the search strategy of the current round, and searching in the search interval according to the search strategy to obtain the interval information of the current round includes:

[0015] If the current round is less than or equal to the preset number of rounds, the search strategy is preliminary search, otherwise it is progressive search;

[0016] When the search strategy is preliminary search, randomly select a preset number of parameter points from the search interval, obtain the performance prediction results of the parameter points, select candidate points from the parameter points based on the performance prediction results, and use all the candidate points as the interval information;

[0017] When the search strategy is progressive search, select the candidate points from the search interval, take the candidate points as the center, determine at least one new reference point within a preset range, obtain the performance prediction results of the new reference points, and determine the interval information from the new reference points based on the performance prediction results.

[0018] In some embodiments, obtaining the performance prediction results of the parameter points includes:

[0019] Input the parameter points into a pre-trained performance prediction model for data prediction to obtain performance metrics;

[0020] Obtain multiple neighborhood points within the neighborhood of the reference point, obtain the performance metrics corresponding to the neighborhood points, and calculate the uncertainty metric value of the reference point based on the performance metrics;

[0021] Obtain the performance test result based on the performance metric and the uncertainty metric value.

[0022] In some embodiments, the using the interval information as the search interval for the next round includes:

[0023] Obtain a preset exploration ratio, and expand the interval information based on the preset exploration ratio to obtain an expanded interval;

[0024] Use the expanded interval as the search interval.

[0025] In some embodiments, after obtaining the task ratio of each reference language model, the method further includes:

[0026] Obtain the total sample set of the reference language model, obtain the initial sample subset corresponding to each task type based on the initial task ratio of the reference language model and the total sample set, and fine-tune the reference language model based on the initial sample subset to obtain the initial performance metric of each task type;

[0027] For at least one preset detection time, obtain the current sample subset corresponding to each task type based on the task ratio and the total sample set, and fine-tune the reference language model based on the current sample subset to obtain the current performance metric of each task type;

[0028] Obtain the difference between the initial performance metric and the current performance metric to obtain the performance change value corresponding to each task type, and generate a change window based on the maximum and minimum values of the performance change value;

[0029] Use the change window of the previous preset detection time as the performance evaluation window, update the task ratio based on the performance evaluation window, if the performance change value is less than the minimum value of the performance evaluation window, increase the ratio parameter corresponding to the task type, if the performance change value is greater than the maximum value of the performance evaluation window, decrease the ratio parameter corresponding to the task type, and the sum of all ratio parameters is one.

[0030] In some embodiments, after updating the task ratio based on the performance evaluation window, the method further includes:

[0031] Obtain the updated sample subset corresponding to each task type based on the updated task ratio and the total sample set, the updated sample subset includes a plurality of task samples, and each task sample includes a task label;

[0032] For each task sample in each of the updated sample subsets, obtain a first probability that the prediction result of the reference language model is the task label given the task sample, and obtain a second probability that the reference language model generates the task sample. Obtain a first intermediate value according to the ratio of the first probability to the second probability;

[0033] Accumulate all the first intermediate values to obtain a second intermediate value, and obtain the task average following difficulty of the task type according to the ratio of the second intermediate value to the number of samples in the total sample set. The task average following difficulty is used to evaluate the update performance of the task ratio.

[0034] In some embodiments, the model fitting of the model parameter number of the reference language model and the corresponding task ratio to obtain a proportional prediction model includes:

[0035] Obtain the model parameter number of each reference language model, and obtain the model complexity and the task correlation between each task type;

[0036] Generate an initial feature vector according to at least one of the model parameter number, model complexity, and task correlation, and perform scaling adjustment on the initial feature vector to obtain a multi-dimensional feature vector;

[0037] Use the multi-dimensional feature vector and the corresponding task ratio as a fitting data group, and perform model fitting using multiple fitting data groups to obtain the proportional prediction model.

[0038] To achieve the above object, a second aspect of the embodiments of the present application proposes an adaptive data ratio device for instruction supervised fine-tuning, including:

[0039] An acquisition module: used to acquire a plurality of reference language models. Each reference language model includes a plurality of task types, and each task type includes a corresponding baseline weight interval. The model parameter numbers of at least two reference language models are different;

[0040] A parameter search module: used to select the reference language models one by one. For the task type, perform at least one round of parameter search based on the baseline weight interval to obtain ratio parameters, and obtain the task ratio of the reference language model according to the ratio parameters corresponding to each task type;

[0041] A fitting module: used to obtain the task ratio of each reference language model, and perform model fitting on the model parameter number of the reference language model and the corresponding task ratio to obtain a proportional prediction model;

[0042] Prediction module: It is used to predict the target task ratio by using the ratio prediction model based on the number of model parameters of the target large language model, and the number of model parameters of the reference language models is less than that of the target large language model.

[0043] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0044] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a storage medium, which is a storage medium that stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0045] The adaptive data ratio method, device, equipment and storage medium for instruction supervised fine-tuning proposed in the embodiments of the present application obtain multiple reference language models. The instruction supervised fine-tuning data of each reference language model includes multiple task types, and each task type includes a corresponding baseline weight interval. Each reference language model is selected one by one. For a task type, at least one round of parameter search is performed based on the baseline weight interval to obtain ratio parameters, and the task ratio of the reference language model is obtained according to the ratio parameters corresponding to each task type. The task ratios of each reference language model are obtained, and a model fitting is performed on the number of model parameters of the reference language model and the corresponding task ratio to obtain a ratio prediction model. Based on the number of model parameters of the target large language model, the target task ratio is predicted by using the ratio prediction model. The number of model parameters of the reference language models is less than that of the target large language model. In the embodiments of the present application, the optimal search for task ratios of multiple reference language models with small numbers of parameters is performed to generate multiple pairs of data of the number of model parameters and task ratios, and the ratio prediction model is trained with this. Subsequently, the ratio prediction model is used to automatically predict the target large language model with large-scale parameters to obtain a target task ratio with high reliability. Compared with the setting based on empirical values, this embodiment can significantly improve the accuracy of instruction supervised fine-tuning. In addition, since both the task ratio search and the training of the ratio prediction model are performed on small-parameter models, the search and training costs for large-scale parameter models can be greatly reduced, which is applicable to large language models of different scales and has high versatility and scalability. Description of the Drawings

[0046] Figure 1 It is a schematic diagram of generating training data based on a human-machine collaboration method provided by the embodiments of the present application.

[0047] Figure 2It is a flowchart of an adaptive data ratio method for instruction supervised fine-tuning provided by an embodiment of the present application.

[0048] Figure 3 It is a schematic diagram of the parameter search process of a reference language model provided by an embodiment of the present application.

[0049] Figure 4 It is a flowchart of obtaining ratio parameters by performing at least one round of parameter search based on a baseline weight interval provided by an embodiment of the present application.

[0050] Figure 5 It is a flowchart of determining the search strategy for the current round and performing a search in the search interval according to the search strategy to obtain the interval information for the current round provided by an embodiment of the present application.

[0051] Figure 6 It is a flowchart of obtaining the performance prediction result of a parameter point provided by an embodiment of the present application.

[0052] Figure 7 It is a flowchart of task ratio dynamic update provided by an embodiment of the present application.

[0053] Figure 8 It is a flowchart of the calculation process of average following difficulty provided by an embodiment of the present application.

[0054] Figure 9 It is another schematic diagram of task ratio dynamic update provided by an embodiment of the present application.

[0055] Figure 10 It is a flowchart of constructing a ratio prediction model by performing model fitting on the model parameter quantity of a reference language model and the corresponding task ratio provided by an embodiment of the present application.

[0056] Figure 11 It is a structural block diagram of an adaptive data ratio device for instruction supervised fine-tuning provided by another embodiment of the present application.

[0057] Figure 12 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0058] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0059] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.

[0061] First, several terms involved in this application are analyzed as follows:

[0062] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also refers to the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0063] A large language model is a model in the field of Natural Language Processing (NLP). Its training process is divided into two stages: First, it undergoes unsupervised pre-training on a large-scale text dataset to learn the general representation and knowledge of the language; second, it improves the model's instruction-following ability and task processing performance through instruction-supervised fine-tuning on various task data. Instruction-supervised fine-tuning designs data samples of various tasks in the form of text generation. During the fine-tuning process, all tasks share the same task form, training objective, and loss function, namely text generation, next token prediction, and cross-entropy loss.

[0064] Different from traditional multi-task learning, the instruction-supervised fine-tuning of large language models involves a larger number of tasks, such as hundreds or even thousands. Even some large language models need to perform instruction-supervised fine-tuning learning on tens of thousands of sub-tasks, and the task diversity is higher, that is, there may not be a correlation between different tasks. Therefore, on the premise of ensuring performance, it is necessary to use as little data as possible to configure the data ratio of each task to balance the performance of each task and reduce the training cost.

[0065] In the related art, according to the specific application direction of the large language model, the data proportion is set for different tasks during the instruction supervised fine-tuning process, so that the sample data volume corresponding to different tasks is different, thereby balancing the performance of each task and reducing the training cost. However, the current data proportion of different tasks is mainly set based on experience and completely depends on the developer's understanding of the data, resulting in a low accuracy of instruction supervised fine-tuning. For example, in the sample dataset, the translation task accounts for 80%, and the question-and-answer task accounts for 20%. At this time, the large language model may be overly biased towards the translation ability, resulting in a significant decline in the question-and-answer performance. Another example is that when the same data proportion is set for each task in a uniform manner, due to the difficulty differences between tasks, the uniform distribution method is also difficult to achieve the optimal fine-tuning performance.

[0066] In addition, the method of determining the optimal data ratio for each task through grid search does not work well in this scenario. The main reason is that when the number of tasks is large, the search space will expand rapidly. For example, when the number of tasks is k and the search space for each task is n, the total search space will reach n to the power of k, which is difficult to implement in practice. Secondly, the optimal data ratios suitable for large language models with different parameter scales are not the same. If the data ratios are searched for all large language models, especially those with huge parameter scales, it will cost a huge amount.

[0067] At the same time, in the case of multi-task mixing, the mutual influence between tasks also needs to be considered, such as the positive interaction between tasks and the negative interaction between tasks. The positive interaction between tasks means that each task promotes each other during mixed training, and the negative interaction between tasks means that each task repels each other during mixed training. The positive interaction usually can reduce the data volume required for each task, while the negative interaction will increase the data volume required for each task. These effects make it no longer applicable to directly adopt the optimal data ratio during the single-task fine-tuning of each task for mixing.

[0068] Based on this, the embodiments of the present application provide an adaptive data ratio method, device, device and storage medium for instruction supervised fine-tuning. By performing an optimal search for task ratios on multiple reference language models with small parameter quantities, multiple data pairs of model parameter quantities and task ratios are generated, and a ratio prediction model is trained with this. Subsequently, the ratio prediction model is used to automatically predict the large language model with large-scale parameters, and a target task ratio with high reliability is obtained. Compared with the setting based on empirical values, this embodiment can significantly improve the accuracy of instruction supervised fine-tuning. In addition, since both the task ratio search and the training of the ratio prediction model are performed on models with small parameter quantities, the search and training costs for large-scale parameter models can be greatly reduced, which is applicable to large language models of different scales and has high generality and scalability.

[0069] The embodiments of the present application provide an adaptive data ratio method, device, equipment, and storage medium for instruction supervised fine-tuning, which will be specifically described through the following embodiments. First, the adaptive data ratio method for instruction supervised fine-tuning in the embodiments of the present application will be described.

[0070] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machine to have the functions of perception, reasoning, and decision-making.

[0071] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0072] The adaptive data ratio method for instruction supervised fine-tuning provided by the embodiments of the present application relates to the field of artificial intelligence technology. The adaptive data ratio method for instruction supervised fine-tuning provided by the embodiments of the present application can be applied to terminals, can also be applied to server sides, and can also be a computer program running on a terminal or a server side. For example, the computer program can be a native program or software module in an operating system; it can be a local (Native) application (Application, APP), that is, a program that needs to be installed in the operating system to run, such as a client that supports the adaptive data ratio of instruction supervised fine-tuning, that is, a program that only needs to be downloaded to the browser environment to run; it can also be a small program that can be embedded in any APP. All in all, the above computer program can be any form of application program, module, or plugin. Among them, the terminal communicates with the server through a network. This adaptive data ratio method for instruction supervised fine-tuning can be executed by the terminal or the server, or jointly executed by the terminal and the server.

[0073] In some embodiments, the terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart watch, etc. The server may be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; it may also be a service node in a blockchain system, and the service nodes in the blockchain system form a peer-to-peer (P2P) network, and the P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). The terminal and the server can be connected through communication connection methods such as Bluetooth, Universal Serial Bus (USB), or network, and this embodiment does not limit this here.

[0074] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0075] First, describe the training data generation process used in the instruction supervision fine-tuning process of the embodiments of this application.

[0076] In one embodiment, refer to Figure 1 , Figure 1 is a schematic diagram of generating training data based on a human-machine collaboration method provided by the embodiments of this application. Refer to Figure 1 , this process is mainly divided into four steps: data collection stage, clustering and induction stage, annotator construction stage, and prediction and classification stage.

[0077] First, the core objective of the data collection phase is to construct a comprehensive and balanced basic dataset. The comprehensiveness is reflected in that the data covers various task types of large language models, such as question answering, dialogue, translation, classification, code, reasoning, generation and creation, summarization, etc., while the balance requires the data to be evenly distributed among different task types and scenarios. Therefore, when screening data from open databases or self-built databases in the embodiments of the present application, a preliminary evaluation is carried out according to the representativeness, diversity, and quality of the data. For example, when extracting data, means such as keyword screening and topic classification are used to ensure that the data covers a wide range of topic fields, such as news, social media, or academic literature.

[0078] Next, in the clustering and induction phase, clustering, induction, and sampling annotation are performed on the basic dataset. First, an automatic clustering algorithm, such as K-means, DBSCAN, or hierarchical clustering, etc., is used to cluster the data in the preliminarily screened basic dataset. After clustering, all possible task types are generalized by type to obtain multiple clustering clusters, and each task type corresponds to a clustering cluster. Then, multiple data samples are evenly sampled from each clustering cluster, and a task type and a quality level are labeled as tags for each data sample. Here, the quality level uses a 0-10 scale, and the scoring criteria can be language fluency, information accuracy, task relevance, etc. Through this processing phase, not only can potential problems in the data be discovered, but also heuristic rules can be summarized from it for subsequent data filtering and annotator training.

[0079] Next, in the annotator construction phase, a classifier is trained as an annotator through an iterative optimization method, and an optimization algorithm and heuristic rules are designed. First, the data samples and labels obtained by the above clustering are used as the training set to train the annotator. The annotator can be constructed based on a deep learning model (such as BERT, GPT) or a traditional machine learning model (such as SVM, random forest), depending on the data scale and task complexity. In addition, the data samples and labels are filtered according to heuristic rules, for example, deleting low-quality samples or duplicate data. In addition, an optimization algorithm is used to automatically annotate high-quality data. Finally, through sampling testing and manual annotation, the classifier, optimization algorithm, and heuristic rules are continuously iteratively optimized to improve their generalization ability and adaptability. For example, in each iteration, the performance of the annotator is evaluated through cross-validation, and the model parameters or optimization rules are adjusted according to the results.

[0080] Finally, in the prediction and classification stage, based on the trained annotator, label the task types and quality levels for other data in the basic dataset, and screen high-quality data according to the labeling results. The screening criteria include high scores (such as above 8 points) and multi-task coverage (such as applicable to both generation and question-and-answer tasks). Through this process, multi-task sample data that can be used for subsequent instruction supervised fine-tuning is finally obtained. The data ratio can be understood as the proportion of sample data corresponding to different task types in the total multi-task sample data. It can be understood that multi-task sample data can also be obtained through manual annotation, and this embodiment does not limit this.

[0081] Next, an adaptive data ratio method for instruction supervised fine-tuning in an embodiment of the present application is described.

[0082] Figure 2 is an optional flowchart of the adaptive data ratio method for instruction supervised fine-tuning provided by an embodiment of the present application. Figure 2 The method in may include but is not limited to steps 110 to 140. At the same time, it can be understood that this embodiment Figure 2 does not specifically limit the order of steps 110 to 140 in, and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added.

[0083] Step 110: Obtain multiple reference language models.

[0084] In one embodiment, since the purpose of the embodiment of the present application is to automatically infer and determine the optimal data ratio of a larger or different-scale large language model in multi-task learning by using the data ratio obtained from a small-parameter model, it can be regarded as predicting the performance of a larger system by learning the behavior patterns of a smaller-scale system. Therefore, the large language model that needs to perform data ratio prediction is called the target large language model, and a series of small models with different numbers of parameters (for example, from 100K to 100M, and the number of parameters of two consecutive small models differs by a factor of 2) selected are called reference language models.

[0085] It can be understood that the number of parameters of different reference language models is different, and the number of parameters of all reference language models is less than that of the target large language model.

[0086] Among them, the instruction supervised fine-tuning data of the target large language model and each reference language model includes multiple task types, such as question and answer, dialogue, translation, classification, code, reasoning, generation and creation, summary, etc. The task types of different reference language models and the target large language model can be different.

[0087] In one embodiment, first, a baseline weight range for data allocation is set for each task type based on expert experience. That is, if the final data allocation for this task type is within this baseline weight range, it can ensure that the sampling ratio of the total multi-task data samples matches the actual importance and data complexity of this task type. The purpose is to delimit an efficient exploration area for each task type in the possibility space, thereby reducing ineffective searches. It can be understood that the baseline weight range can be set according to empirical values.

[0088] Step 120: Select a reference language model one by one. For the task type, perform at least one round of parameter search based on the baseline weight range to obtain the allocation parameters, and obtain the task allocation of the reference language model according to the allocation parameters corresponding to each task type.

[0089] In one embodiment, after obtaining the reference language model, the behavior pattern between the number of model parameters of this learning model and the task allocation can be based on it. First, it is necessary to determine the optimal task allocation for each reference language model. The task allocation is composed of the allocation parameters of each task type. Specifically, the task allocation corresponding to each reference language model can be determined by means of parameter search in the baseline weight range.

[0090] Refer to Figure 3 , Figure 3 FIG. is a schematic diagram of the parameter search process of the reference language model provided by the embodiment of the present application. First, select a reference language model one by one as the reference language model currently being processed. For each task type, it is necessary to search for its optimal allocation parameters.

[0091] In one embodiment, refer to Figure 4 , Figure 4 FIG. is a flowchart of obtaining allocation parameters by performing at least one round of parameter search based on the baseline weight range provided by the embodiment of the present application, specifically including the following steps:

[0092] Step 410: In the process of each round of parameter search, obtain the search range of the current round.

[0093] In one embodiment, for a certain task type, a multi-round search process needs to be executed. That is to say, taking the baseline weight range as the search space, multiple searches need to be performed to obtain the optimal allocation parameters corresponding to this task type. Among them, the search range is a subset of the parameter space and is usually used to find the optimal parameters. The total number of rounds of parameter search here can be set according to actual needs.

[0094] Taking one round of parameter search as an example, it is necessary to determine the search range for the current round. In the first iteration, the search range is the baseline weight range, that is, the initial value of the search range is the baseline weight range. In other iterations, the search range is the result obtained in the previous iteration. That is to say, the next round of search is further searched based on the search result of the previous round. The following steps will explain this.

[0095] Step 420: Determine the search strategy for the current round, and search in the search range according to the search strategy to obtain the interval information for the current round, and use the interval information as the search range for the next round.

[0096] In one embodiment, referring to Figure 3 , at the beginning of each round of parameter search, it is necessary to adaptively determine the search strategy in the decision center. After having the search strategy, it is possible to search from the search range based on this search strategy. The result obtained from the search is called interval information, and this interval information is the search range for the next round.

[0097] In one embodiment, referring to Figure 5 , Figure 5 is the flowchart for determining the search strategy for the current round and searching in the search range according to the search strategy to obtain the interval information for the current round provided by the embodiment of the present application, which specifically includes the following steps:

[0098] Step 510: If the current round is less than or equal to the preset round, the search strategy is a preliminary search; otherwise, it is a progressive search.

[0099] In one embodiment, intelligent search is performed by using an adaptive hybrid search strategy, which can automatically switch the search mode according to the requirements of different search stages. This embodiment adopts a strategy of first rough search and then fine search. Therefore, if the current round is less than or equal to the preset round, it is determined that it is in the initial stage of the search, and the search strategy is set as a preliminary search at this time. If the current round is greater than the preset round, it is considered that it is necessary to enter the process of detailed search, and the search strategy is set as a progressive search. The preset round here can be set according to actual needs.

[0100] Step 520: When the search strategy is a preliminary search, randomly select a preset number of parameter points from the search range, obtain the performance prediction results of the parameter points, select candidate points from the parameter points based on the performance prediction results, and use all candidate points as the interval information.

[0101] In one embodiment, if the search strategy is a preliminary search, first, a preset number of parameter points are randomly selected from the search interval based on an optimization strategy of random sampling to explore the potential optimal region of the parameter space with limited computing resources. Among them, the method of random sampling can be a uniform distribution or a pseudo-random sequence, and the diversity and coverage of the sampling points are ensured by the way of random sampling. After obtaining the parameter points, the performance prediction results of the parameter points are obtained, and candidate points are screened out from the parameter points through the performance prediction results. Here, the selection strategy of the candidate points can be the points with the best performance, the points with performance higher than a certain threshold. Finally, these candidate points are used as new interval information to update the search interval of the next round to guide the subsequent search process.

[0102] For example, 20 parameter points are randomly selected from the search interval, and the 5 points with the best performance are selected as candidate points through the performance prediction results of these parameter points. These candidate points are used as new interval information for the next round of optimization. By iterating this process, the optimal parameter combination is gradually approximated, thereby achieving efficient parameter optimization.

[0103] In one embodiment, referring to Figure 3 , the performance prediction results at least include performance metrics and uncertainty metric values, and the performance prediction results will in turn affect the adaptive selection of the search strategy. The calculation processes of these two parameters are described in detail below. Referring to Figure 6 , Figure 6 is a flowchart for obtaining the performance prediction results of parameter points provided by an embodiment of the present application, which specifically includes the following steps:

[0104] Step 610: Input the parameter points into a pre-trained performance prediction model for data prediction to obtain performance metrics.

[0105] In one embodiment, the parameter points are essentially the possible data ratios corresponding to the task types. For example, when the reference language model is a multi-task language model supporting generation, question answering, and translation tasks, for the generation task type, its parameter points can be expressed as 40% for the generation task, 30% for the generation task, etc.

[0106] Therefore, this application pre-trains a performance prediction model, which can be a neural network model. The data ratio is the input feature, and the performance is the target label. The performance can be quantitatively judged according to actual needs. For example, in a generation task, BLEU or ROUGE metrics can be used; in a question-answering task, accuracy or F1 score can be used; in a translation task, METEOR or TER metrics can be used. The performance prediction model can be a neural network model, such as a multi-layer perceptron (MLP), a convolutional neural network (CNN), or a Transformer-based model. The specific choice depends on the data scale and task complexity, which is not limited in this embodiment. For example, a three-layer MLP model can be used. Its input layer receives the data ratio, the hidden layer extracts features, and the output layer predicts the performance metrics.

[0107] Finally, the trained performance prediction model can perform data prediction according to the input parameter points to obtain the possible performance metrics corresponding to these parameter points. For example, in a generation task, if the input is that the generation task accounts for 50%, the performance prediction model can predict that the BLEU score of the generation task is 0.85.

[0108] Step 620: Obtain multiple neighborhood points within the neighborhood of the reference point, obtain the performance metrics corresponding to the neighborhood points, and calculate the uncertainty measure value of the reference point based on the performance metrics.

[0109] In one embodiment, in order to evaluate the uncertainty level of the area around the current optimal solution and ensure that premature convergence will not occur due to accidental factors. Premature convergence may cause the optimization algorithm to fall into a local optimum and ignore the global optimum solution. Therefore, in the embodiment of this application, multiple neighborhood points are obtained within the neighborhood of the reference point, and their performance metrics are calculated. The uncertainty measure value of the reference point is calculated based on the performance metrics to quantify the uncertainty of the reference point.

[0110] Among them, the neighborhood size can be determined according to the actual situation. Within this neighborhood range, multiple reference points are selected as neighborhood points. Then, according to the above process, the performance metrics of each neighborhood point are obtained using the performance prediction model, and then the mean and variance of the performance metrics are calculated. The variance is used as the uncertainty measure value.

[0111] Step 630: Obtain the performance test result according to the performance metrics and the uncertainty measure value.

[0112] In one embodiment, the size relationship between the performance metrics and the preset metric value, and the size relationship between the uncertainty measure value and the preset measure value are respectively judged. The performance test result is obtained according to the judgment result. If the performance metrics are within the preset metric threshold and the uncertainty measure value is within the preset measure threshold, the performance test result indicates that this reference point can be used as a candidate point.

[0113] Step 530: When the search strategy is progressive search, select candidate points from the search interval. Taking the candidate points as the center, determine at least one new reference point within a preset range, obtain the performance prediction results of the new reference points, and determine interval information from the new reference points based on the performance prediction results.

[0114] In one embodiment, after the initial stage of the search, the further screening process can be entered. Therefore, when the search strategy is progressive search, the number of candidate points has become smaller at this time. Therefore, select candidate points from the search interval, define its local search range with the candidate points as the center, and determine multiple new reference points within the preset range in the manner of grid search. For example, if the candidate point is 0.01, the corresponding local search range can be [0.008, 0.012]. Subsequently, uniformly select multiple new reference points within the preset range in the manner of grid search. For example, select 5 new reference points: 0.008, 0.009, 0.01, 0.011, and 0.012.

[0115] Next, obtain the performance prediction results of the new reference points. According to the evaluation method of the candidate points, determine from the new reference points based on the performance prediction results that the new reference points with the performance index located within the preset index threshold and the uncertainty measurement value located within the preset measurement value threshold constitute the interval information.

[0116] It can be understood that in the embodiments of the present application, the current search strategy will be determined in each parameter search process, and the search strategy can be adaptively switched according to the iteration rounds. In the initial stage of the search, the system tends to widely cover the entire possible space to discover potential candidate points. As the search progresses, when promising regions are identified, the system will gradually narrow or expand the interval based on the previous search results and turn to a more precise method for local refinement search. If certain specific proportion values perform well, the regions around these values will be explored more intensively; conversely, regions with poor performance will be quickly excluded.

[0117] In one embodiment, in order to avoid falling into local optimality, whether it is the initial search or the progressive search, after obtaining the interval information in each parameter search process, a certain degree of exploratory search needs to be carried out. The purpose of this exploratory search is to further explore the possible better solutions in the parameter space outside the known high-performance regions, so as to ensure the global search ability of the optimization algorithm. In the embodiments of the present application, obtain a preset exploration ratio, which is usually set to a relatively small value (such as 10% or 20%), expand the interval information based on the preset exploration ratio to obtain an expanded interval, and use the expanded interval as the search interval. For example, if the interval information is [0.01, 0.02] and the preset exploration ratio is 10%, the range of the exploratory search can be expanded to [0.009, 0.021], and the obtained search interval is [0.009, 0.021].

[0118] Step 430: After the iteration termination condition is met, obtain the ratio parameters based on the interval information obtained from the parameter search in the last round.

[0119] In one embodiment, after the iteration termination condition is met, the parameter search process ends. The iteration termination condition here can be the total number of rounds, or the stable change of the performance test results, or that each task type reaches the optimal performance of a single task, etc., which is set according to actual requirements. Then, obtain the ratio parameters based on the interval information obtained from the parameter search in the last round. The interval information here can be a single parameter point or an interval of parameter points. If it is an interval, select one parameter point from it as the ratio parameter.

[0120] According to the above process, the ratio parameters for each task type in the reference language model can be obtained and form the task ratio of the reference language model. The sum of all ratio parameters is one. In the whole process, through the adaptive hybrid search method, fine optimization is carried out within the baseline weight interval delimited by experts. Single-task training is performed on each task type to find the optimal ratio parameters, that is, the proportion of data to be sampled, under specific conditions. Using the actual performance feedback of the data to fine-tune the search process ensures the accuracy and adaptability of the search. At the same time, through the dynamically adjusted interval information, the intelligently selected search mode, and the accelerated search of performance prediction, the computational complexity can be significantly reduced and the overall efficiency can be improved. In addition, since the ratio parameters for each task type are optimized separately, it will not bring an exponential increase in the search space.

[0121] In one embodiment, considering that the performance of the reference language model will be gradually optimized during the fine-tuning process, an iterative feedback control system is also used to ensure that the obtained task ratio can be automatically adjusted with the change of the performance of the reference language model, so as to simply and efficiently adaptively update the optimal ratio parameters of each task type.

[0122] In one embodiment, referring to Figure 7 , Figure 7 is the flowchart of the dynamic update of the task ratio provided by the embodiment of the present application, which specifically includes the following steps:

[0123] Step 710: Obtain the total sample set of the reference language model, obtain the initial sample subset corresponding to each task type based on the initial task ratio and the total sample set of the reference language model, and fine-tune the reference language model based on the initial sample subset to obtain the initial performance index of each task type.

[0124] In one embodiment, a total sample set of a reference language model is obtained from the multi-task sample data obtained above. The total sample set contains multiple sample data corresponding to each task type in the reference language model. Suppose an initial proportion is set for each task type in the reference language model, and these initial proportions form an initial task ratio. For example, the generation task accounts for 40%, the question-and-answer task accounts for 30%, and the translation task accounts for 30%. At this time, the initial task ratio is 4:3:3. The initial task ratio can be determined according to empirical values (such as a larger amount of generation task data, a higher weight is assigned) or evenly distributed (such as each task type accounts for 33.3%).

[0125] At this time, according to the initial task ratio of the reference language model and the total sample set, an initial sample subset corresponding to each task type can be obtained. According to the initial task ratio and the total sample set, the initial sample subset corresponding to each task type is extracted. For example, if the generation task sample set in the total sample set contains 1000 samples, the question-and-answer task sample set contains 750 samples, the translation task sample set contains 750 samples, and the initial task ratio is 4:3:3, then 400 samples are selected from the generation task sample set to form the initial sample subset of the generation task, 300 samples are selected from the question-and-answer task sample set to form the initial sample subset of the question-and-answer task, and 300 samples are selected from the translation task sample set to form the initial sample subset of the translation task.

[0126] Next, using the initial sample subset, the reference language model is trained in the way of instruction supervised fine-tuning, and according to the performance evaluation quantization method of the fine-tuned reference language model, the initial performance index of each task type is obtained. The performance evaluation quantization method can be selected according to the task type. For example, the BLEU or ROUGE index is used for the generation task, the accuracy or F1 score is used for the question-and-answer task, and the METEOR or TER index is used for the translation task. For example, the BLEU score of the fine-tuned model on the generation task is 0.85, the F1 score on the question-and-answer task is 0.92, and the METEOR score on the translation task is 0.78.

[0127] Step 720: For at least one preset detection time, based on the task ratio and the total sample set, obtain the current sample subset corresponding to each task type, and fine-tune the reference language model based on the current sample subset to obtain the current performance index of each task type.

[0128] In one embodiment, multiple preset detection times are set, and the prediction detection time is set according to actual needs. At each preset detection time, the current task ratio is obtained, and then according to the task ratio and the total sample set, the current sample subset corresponding to each task type is obtained. According to the calculation method of the above initial performance index, the reference language model is fine-tuned based on the current sample subset to obtain the current performance index of each task type.

[0129] Step 730: Obtain the difference between the initial performance metric and the current performance metric to get the performance change value corresponding to each task type, and generate a change window based on the maximum and minimum values of the performance change values.

[0130] In one embodiment, calculate the difference between the initial performance metric and the current performance metric to get the performance change value corresponding to each task type. At this time, for a reference language model, since it has multiple task types, there will be multiple performance change values ΔP. A change window is formed based on the maximum and minimum values of the performance change values, denoted as [ΔP_min, ΔP_max], where ΔP_max represents the maximum value of the performance change value, and ΔP_min represents the minimum value of the performance change value. This window reflects the change range of the task performance of the current reference language model. The precise capture of the task performance change avoids frequent adjustments caused by minor fluctuations.

[0131] Step 740: Use the change window of the previous preset detection time as the performance evaluation window, and update the task ratio based on the performance evaluation window. If the performance change value is less than the minimum value of the performance evaluation window, increase the ratio parameter corresponding to the task type. If the performance change value is greater than the maximum value of the performance evaluation window, decrease the ratio parameter corresponding to the task type.

[0132] In one embodiment, use the change window of the previous preset detection time as the performance evaluation window to judge the change in the performance of the current reference language model compared to the previous detection. At this time, update the task ratio based on the performance evaluation window. If the performance change value ΔP is less than the minimum value ΔP_min of the performance evaluation window, it means the performance has deteriorated and the data volume of this task type needs to be increased, so increase the ratio parameter corresponding to the task type. If the performance change value ΔP is greater than the maximum value ΔP_max of the performance evaluation window, it means the data volume of this task type needs to be decreased, so decrease the ratio parameter corresponding to the task type. In this embodiment, once it is determined that the task ratio needs to be adjusted, the ratio parameter of the corresponding task type will be appropriately increased or decreased according to the specific situation of the change window to ensure that more resources are allocated to those task types with poor performance. It can be understood that in order to avoid the negative impact brought by the drastic changes during the adjustment process, the adjustment of the ratio parameter is usually implemented in a moving average manner, gradually approaching the optimal value, based on the smoothing process to ensure the effective allocation of resources and prevent the occurrence of overfitting.

[0133] The entire process of the embodiment of this application forms a closed-loop control system, enabling the system to continuously learn and optimize its own parameter configuration during operation, with high flexibility and intelligence.

[0134] It can be understood that the adjustment process of the ratio parameters is based on the task ratio obtained by parameter search, so as to ensure that the adjustment of each task type starts at least at a level close to the optimal sampling magnitude, providing a reasonable starting point for adaptive optimization.

[0135] In one embodiment, in order to evaluate the performance change value of the reference language model before and after adjusting the task ratio, the parameter of task average following difficulty is introduced for quantitative evaluation. Referring to Figure 8 , Figure 8 is the flowchart of the calculation process of the average following difficulty provided by the embodiment of the present application, which specifically includes the following steps:

[0136] Step 810: Obtain an updated sample subset corresponding to each task type based on the updated task ratio and the total sample set.

[0137] In one embodiment, since the task ratio has been updated, an updated sample subset corresponding to each task type is obtained based on the updated task ratio and the total sample set. The updated sample subset includes multiple task samples, and each task sample includes a task label, and the task label is used to indicate the task type corresponding to the task sample.

[0138] Step 820: For each task sample in each updated sample subset, obtain the first probability that the prediction result of the reference language model is the task label under the condition of the given task sample, and obtain the second probability that the reference language model generates the task sample, and obtain a first intermediate value according to the ratio of the first probability and the second probability.

[0139] In one embodiment, the first intermediate value is expressed as:

[0140]

[0141] where x i represents the i-th task sample corresponding to a certain task type, y i represents the task label of the task sample, and P(y i |x i ) represents the first probability, and P(x i ) represents the second probability.

[0142] Step 830: Accumulate all the first intermediate values to obtain a second intermediate value, and obtain the task average following difficulty of the task type according to the ratio of the second intermediate value and the number of samples in the total sample set.

[0143] In one embodiment, the second intermediate value is expressed as:

[0144]

[0145] where N represents the total number of samples in the updated sample subset of a certain task type.

[0146] Therefore, the average following difficulty FD of the tasks corresponding to this task type is expressed as:

[0147]

[0148] The average following difficulty of the tasks represents the difficulty of the task samples. The fine-tuning process is actually a process of reducing the difficulty of the task samples. The lower the difficulty drops, the better the performance. After obtaining the average following difficulty of the tasks corresponding to each task type, it is possible to judge the performance changes of each task type before and after updating the task ratio. Therefore, the average following difficulty of the tasks can be used to evaluate the updated performance of the task ratio. In this embodiment, the average following difficulty of the tasks is used as a unified evaluation index for cross-task performance, aligning the evaluation of all tasks to the same space to ensure the comprehensiveness and accuracy of the evaluation and simplify the evaluation process.

[0149] In one embodiment, referring to Figure 9 , Figure 9 is another schematic diagram of the dynamic update of the task ratio provided by the embodiment of the present application. Figure 9 In it, each time the task ratio is updated, dynamic performance monitoring needs to be carried out based on the average following difficulty of the tasks. When the average following difficulty indicates that the performance is developing well, the previous change window is used as the performance evaluation window, and this is used as a threshold to determine whether to trigger the adjustment of the task ratio, avoiding frequent and unnecessary changes caused by small fluctuations in the model performance and ensuring the stability and efficiency of the learning process.

[0150] The task ratio is adjusted and updated based on the performance evaluation window, and the reference language model is fine-tuned again according to the updated task ratio, and the new performance changes are continuously monitored and fed back into the next evaluation to form a self-learning and optimizing learning loop. As time goes by, the change window mechanism can continuously accumulate experience and become more intelligent in identifying which changes are real performance transformations and which are just random noise, thereby improving the decision-making accuracy.

[0151] Through this dynamic adjustment process, performance bottleneck tasks can be efficiently identified and preferentially processed, achieving on-demand allocation of resources, avoiding overfitting while promoting the balanced improvement of overall performance. Especially in the case of limited resources, through intelligent data allocation strategies, the effect of even surpassing the training with all data can be achieved, greatly saving computing resources and time costs. If the model performance of a certain task type declines compared to the previous stage, the system will automatically increase its data proportion, aiming to improve the performance of this task type through more targeted training; conversely, if the performance improves, its data proportion will be appropriately reduced, and more learning resources will be directed to other task types that may be in the bottleneck period. Experimental results show that even when only 80% of the total data volume is sampled, the task proportion method of the embodiments of the present application can still surpass the performance of full data sampling, improving the training efficiency and the comprehensive performance of the model.

[0152] Step 130: Obtain the task ratio of each reference language model, and perform model fitting on the model parameter quantity of the reference language model and the corresponding task ratio to obtain a ratio prediction model.

[0153] In one embodiment, after obtaining the task ratios of all reference language models in the above manner, behavior pattern learning can be carried out. Refer to Figure 10 , Figure 10 is the flowchart for constructing a ratio prediction model by performing model fitting on the model parameter quantity of the reference language model and the corresponding task ratio provided by the embodiments of the present application, which specifically includes the following steps:

[0154] Step 1010: Obtain the model parameter quantity of each reference language model, and obtain the model complexity and the task correlation between various task types.

[0155] In one embodiment, the model parameter quantity refers to the total number of all trainable parameters in the reference language model, usually expressed in millions (M) or billions (B), and the model parameter quantity can be directly queried from the model architecture or obtained using tools provided by the deep learning framework.

[0156] The model complexity is an indicator used to measure the computing and storage requirements of the reference language model, and can be obtained by analyzing parameters such as the amount of computation, memory occupancy, or training time through analysis tools, so as to quantitatively determine the model complexity. Among them, the amount of computation represents the number of floating-point operations required for the reference language model to complete one forward propagation, the memory occupancy represents the size of the video memory or memory occupied during model training or inference, and the training time represents the time required for the model to complete one training.

[0157] Task relevance refers to the degree of mutual influence between different task types, which can be quantified by comparing the performance of joint training. Specifically, compare the performance of jointly training multiple task types with that of training them separately, and calculate the performance gain or loss. For example, when jointly training the generation and translation tasks, whether the BLEU score of the generation task is significantly higher than that when training the generation task alone.

[0158] Step 1020: Generate an initial feature vector based on at least one of the model parameter quantity, model complexity, and task relevance, and perform scaling adjustment on the initial feature vector to obtain a multi-dimensional feature vector.

[0159] In one embodiment, an initial feature vector is generated based on at least one of the model parameter quantity, model complexity, and task relevance. That is to say, the model parameter quantity is essential, while the model complexity and task relevance can be selected alternatively. Next, the model parameter quantity, and at least one of the model complexity and task relevance are concatenated to generate an initial feature vector. Considering the potential differences in the scales of different reference language models, a scaling factor is introduced to perform scaling adjustment on the initial feature vector to obtain a multi-dimensional feature vector. Among them, the scaling factor can be set according to the actual situation, and the influence brought by the change in the size of the reference language model is corrected through the scaling adjustment process to ensure a smooth transition from small-scale models to large-scale models.

[0160] Step 1030: Use the multi-dimensional feature vector and the corresponding task ratio as a fitting data group, and perform model fitting using multiple fitting data groups to obtain a ratio prediction model.

[0161] In one embodiment, the multi-dimensional feature vector and the corresponding task ratio are used as a fitting data group. At this time, the multi-dimensional feature vector is used as the independent variable, and the task ratio is used as the dependent variable for non-linear function fitting. During the fitting process, a polynomial function model or a machine learning algorithm suitable for processing high-dimensional inputs and capturing complex relationships, such as a neural network, can be selected to fit the non-linear function to obtain a ratio prediction model.

[0162] It can be understood that the input of the ratio prediction model includes the multi-dimensional feature vector after feature engineering including the model parameter quantity, and the output is the task ratio containing the optimal data ratio of each task type.

[0163] Step 140: Based on the model parameter quantity of the target large language model, use the ratio prediction model to predict and obtain the target task ratio.

[0164] In one embodiment, after obtaining the fitted ratio prediction model, an input vector consistent with the multi-dimensional feature vector is generated according to the number of model parameters of the target large language model, and is input into the ratio prediction model for reverse data prediction, so as to obtain the target task ratio corresponding to the target large language model. Among them, the number of model parameters of the reference language models is less than that of the target large language model.

[0165] It can be seen that in the embodiment of the present application, behavior pattern recognition is performed on the reference language models with small parameter scales, and then a ratio prediction model is constructed by fitting based on these data, avoiding the time-consuming and resource-intensive experimental process for each large-scale large language model one by one, and greatly improving the efficiency of instruction supervised fine-tuning. In particular, the introduced scaling factor can ensure a smooth transition from small models to large models, enhancing the accuracy and reliability of the prediction model. Once the ratio prediction model is trained and verified, it can be applied to large language models of larger or different scales, and automatically infer the optimal task ratio, with great flexibility and applicability. It is not only applicable to the current multi-task learning system, but also can be extended to a wider range of fields, such as transfer learning, continuous learning, etc., providing a new idea for quickly determining the optimal configuration in different application scenarios.

[0166] The technical solution provided by the embodiment of the present application includes obtaining multiple reference language models. The instruction supervised fine-tuning data of each reference language model includes multiple task types, and each task type includes a corresponding baseline weight interval. Each reference language model is selected one by one. For each task type, at least one round of parameter search is performed based on the baseline weight interval to obtain the ratio parameters, and the task ratio of the reference language model is obtained according to the ratio parameters corresponding to each task type. The task ratios of each reference language model are obtained, and the ratio prediction model is obtained by fitting the number of model parameters of the reference language model and the corresponding task ratio. Based on the number of model parameters of the target large language model, the target task ratio is predicted by using the ratio prediction model. The number of model parameters of the reference language models is less than that of the target large language model. In the embodiment of the present application, the optimal search for the task ratio of multiple reference language models with small parameter scales is performed to generate multiple pairs of data of the number of model parameters and the task ratio, and the ratio prediction model is trained with this. Subsequently, the ratio prediction model is used to automatically predict the target large language model with large-scale parameters, obtaining a target task ratio with high reliability, improving the performance of each task while effectively reducing the search and training costs. Compared with the setting based on empirical values, this embodiment can significantly improve the accuracy of instruction supervised fine-tuning. In addition, since both the task ratio search and the training of the ratio prediction model are performed on small-parameter models, the search and training costs for large-scale parameter models can be greatly reduced, and it is applicable to large language models of different scales, with high generality and scalability.

[0167] The embodiment of the present application further provides an adaptive data ratio device for instruction supervision fine-tuning, which can implement the above-mentioned adaptive data ratio method for instruction supervision fine-tuning. Refer to Figure 11 , the device includes:

[0168] An acquisition module 1110: configured to acquire a plurality of reference language models, each reference language model includes a plurality of task types, each task type includes a corresponding baseline weight interval, and the model parameters of at least two reference language models are different.

[0169] A parameter search module 1120: configured to select a reference language model one by one. For a task type, perform at least one round of parameter search based on the baseline weight interval to obtain a ratio parameter, and obtain the task ratio of the reference language model according to the ratio parameter corresponding to each task type.

[0170] A fitting module 1130: configured to obtain the task ratio of each reference language model, and perform model fitting on the model parameters of the reference language model and the corresponding task ratio to obtain a ratio prediction model.

[0171] A prediction module 1140: configured to predict the target task ratio by using the ratio prediction model based on the model parameters of the target large language model. The model parameters of the reference language models are all smaller than the model parameters of the target large language model.

[0172] The specific implementation manner of the adaptive data ratio device for instruction supervision fine-tuning in this embodiment is basically the same as the specific implementation manner of the above-mentioned adaptive data ratio method for instruction supervision fine-tuning, and will not be elaborated here.

[0173] The embodiment of the present application further provides an electronic device, including:

[0174] At least one memory;

[0175] At least one processor;

[0176] At least one program;

[0177] The program is stored in the memory, and the processor executes the at least one program to implement the above-mentioned adaptive data ratio method for instruction supervision fine-tuning of the present application. This electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (Personal Digital Assistant, PDA), an in-vehicle computer, etc.

[0178] Please refer to Figure 12 , Figure 12 illustrates the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0179] The processor 1201 can be implemented in the form of a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0180] The memory 1202 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1202 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1202, and are called by the processor 1201 to execute the adaptive data ratio method for instruction supervision and fine-tuning of the present application;

[0181] The input / output interface 1203 is used to implement information input and output;

[0182] The communication interface 1204 is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0183] The bus 1205 transmits information between various components of the device (such as the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204);

[0184] Among them, the processor 1201, the memory 1202, the input / output interface 1203, and the communication interface 1204 achieve communication connections with each other inside the device through the bus 1205.

[0185] The embodiments of the present application also provide a storage medium, which is a storage medium that stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned adaptive data ratio method for instruction supervision and fine-tuning.

[0186] As a non-transitory storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0187] The adaptive data ratio method, device, equipment, and storage medium for instruction supervised fine-tuning proposed in the embodiments of the present application obtain multiple reference language models. The instruction supervised fine-tuning data of each reference language model includes multiple task types, and each task type includes a corresponding baseline weight interval. Each reference language model is selected one by one. For a task type, at least one round of parameter search is performed based on the baseline weight interval to obtain a ratio parameter. The task ratio of the reference language model is obtained according to the ratio parameter corresponding to each task type. The task ratio of each reference language model is obtained, and a model fitting is performed on the model parameter quantity and the corresponding task ratio of the reference language model to obtain a ratio prediction model. Based on the model parameter quantity of the target large language model, the ratio prediction model is used for prediction to obtain the target task ratio. The model parameter quantities of the reference language models are all smaller than the model parameter quantity of the target large language model. In the embodiments of the present application, through an optimal search for the task ratio of multiple reference language models with small parameter quantities, multiple data pairs of model parameter quantity and task ratio are generated, and the ratio prediction model is trained with these data pairs. Subsequently, the ratio prediction model is used to automatically predict the target large language model with large-scale parameters to obtain a target task ratio with high reliability. Compared with the setting based on empirical values, this embodiment can significantly improve the accuracy of instruction supervised fine-tuning. In addition, since both the task ratio search and the training of the ratio prediction model are performed on small-parameter models, the search and training costs for large-scale parameter models can be greatly reduced, which is applicable to large language models of different scales and has high generality and scalability.

[0188] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0189] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0190] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0191] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0192] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0193] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0194] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0195] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0196] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0197] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store programs.

[0198] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. An adaptive data ratio method for instruction supervised fine-tuning, characterized in that Including: Obtain multiple reference language models. The data for the instruction supervised fine-tuning of each reference language model includes multiple task types. Each task type includes a corresponding baseline weight interval, and at least two of the reference language models have different model parameter quantities. Select the reference language models one by one. For each task type, perform at least one round of parameter search based on the baseline weight interval to obtain the ratio parameters, and obtain the task ratio of the reference language model according to the ratio parameters corresponding to each task type. Obtain the task ratio of each reference language model, and perform model fitting on the model parameter quantity of the reference language model and the corresponding task ratio to obtain a proportional prediction model. Based on the model parameter quantity of the target large language model, use the proportional prediction model to perform prediction to obtain the target task ratio. The model parameter quantities of the reference language models are all smaller than the model parameter quantity of the target large language model.

2. The adaptive data ratio matching method for instruction supervised fine-tuning according to claim 1, wherein The performing at least one round of parameter search based on the baseline weight interval to obtain the ratio parameters includes: In the process of each round of parameter search, obtain the search interval of the current round. The initial value of the search interval is the baseline weight interval. Determine the search strategy of the current round, and perform search in the search interval according to the search strategy to obtain the interval information of the current round, and use the interval information as the search interval of the next round. When the iteration termination condition is satisfied, obtain the ratio parameters according to the interval information obtained from the last round of parameter search.

3. The adaptive data ratio adjustment method for instruction supervision fine-tuning according to claim 2, wherein The determining the search strategy of the current round, and performing search in the search interval according to the search strategy to obtain the interval information of the current round includes: If the current round is less than or equal to the preset number of rounds, the search strategy is preliminary search; otherwise, it is progressive search. When the search strategy is preliminary search, randomly select a preset number of parameter points from the search interval, obtain the performance prediction results of the parameter points, select candidate points from the parameter points based on the performance prediction results, and use all the candidate points as the interval information. When the search strategy is progressive search, select the candidate points from the search interval, determine at least one new reference point within a preset range centered on the candidate points, obtain the performance prediction results of the new reference points, and determine the interval information from the new reference points based on the performance prediction results.

4. The adaptive data ratio adjustment method for instruction supervision fine-tuning according to claim 3, wherein, The obtaining the performance prediction results of the parameter points includes: Input the parameter points into a pre-trained performance prediction model for data prediction to obtain performance metrics. Obtain multiple neighborhood points within the neighborhood of the reference point, obtain the performance metrics corresponding to the neighborhood points, and calculate the uncertainty metric value of the reference point based on the performance metrics. Obtain the performance test results according to the performance metrics and the uncertainty metric value.

5. The adaptive data ratio adjustment method for instruction supervised fine-tuning according to claim 2, wherein, The using the interval information as the search interval of the next round includes: Obtain a preset exploration ratio, expand the interval information based on the preset exploration ratio to obtain an expanded interval. Use the expanded interval as the search interval.

6. The adaptive data ratio adjustment method for instruction supervision fine-tuning according to claim 1, characterized in that After obtaining the task ratio of each of the reference language models, the method further includes: Obtaining the total sample set of the reference language model, obtaining an initial sample subset corresponding to each task type based on the initial task ratio of the reference language model and the total sample set, and fine-tuning the reference language model based on the initial sample subset to obtain an initial performance metric for each task type; For at least one preset detection time, obtaining a current sample subset corresponding to each task type based on the task ratio and the total sample set, and fine-tuning the reference language model based on the current sample subset to obtain a current performance metric for each task type; Obtaining the difference between the initial performance metric and the current performance metric to obtain a performance change value corresponding to each task type, and generating a change window based on the maximum and minimum values of the performance change value; Using the change window of the previous preset detection time as a performance evaluation window, updating the task ratio based on the performance evaluation window. If the performance change value is less than the minimum value of the performance evaluation window, increasing the ratio parameter corresponding to the task type. If the performance change value is greater than the maximum value of the performance evaluation window, decreasing the ratio parameter corresponding to the task type. The sum of all the ratio parameters is one.

7. The adaptive data ratio adjustment method for instruction supervision and fine-tuning according to claim 6, characterized in that, After updating the task ratio based on the performance evaluation window, the method further includes: Obtaining an updated sample subset corresponding to each task type based on the updated task ratio and the total sample set. The updated sample subset includes a plurality of task samples, and each task sample includes a task label; For each task sample in the updated sample subset, obtaining a first probability that the prediction result of the reference language model is the task label under the condition of the given task sample, and obtaining a second probability that the reference language model generates the task sample, and obtaining a first intermediate value according to the ratio of the first probability and the second probability; Accumulating all the first intermediate values to obtain a second intermediate value, and obtaining the task average following difficulty of the task type according to the ratio of the second intermediate value and the number of samples in the total sample set. The task average following difficulty is used to evaluate the update performance of the task ratio.

8. The adaptive data ratio adjustment method for instruction supervision fine-tuning according to claim 1, wherein, The model fitting to obtain a ratio prediction model for the model parameter amount of the reference language model and the corresponding task ratio includes: Obtaining the model parameter amount of each reference language model, and obtaining the model complexity and the task correlation between each task type; Generating an initial feature vector according to at least one of the model parameter amount, the model complexity, and the task correlation, and performing scaling adjustment on the initial feature vector to obtain a multi-dimensional feature vector; Using the multi-dimensional feature vector and the corresponding task ratio as a fitting data group, and performing model fitting using a plurality of fitting data groups to obtain the ratio prediction model.

9. An adaptive data ratio device for instruction supervised fine-tuning, characterized in that, Including: Acquisition module: configured to acquire a plurality of reference language models, each of the reference language models including a plurality of task types, each task type including a corresponding baseline weight range, and at least two of the reference language models having different model parameter amounts; Parameter search module: configured to sequentially select the reference language models, and for each task type, perform at least one round of parameter search based on the baseline weight range to obtain ratio parameters, and obtain the task ratio of the reference language model according to the ratio parameters corresponding to each task type; Fitting module: configured to obtain the task ratio of each reference language model, and perform model fitting on the model parameter amount of the reference language model and the corresponding task ratio to obtain a ratio prediction model; Prediction module: configured to predict a target task ratio by using the ratio prediction model based on the model parameter amount of the target large language model, and the model parameter amounts of the reference language models are all smaller than the model parameter amount of the target large language model.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the adaptive data ratio method for instruction supervised fine-tuning according to any one of claims 1 to 8 is implemented.

11. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the adaptive data ratio method for instruction supervised fine-tuning according to any one of claims 1 to 8 is implemented.