Task model training method and device, electronic equipment and readable storage medium
By dynamically adjusting the loss slope and threshold in large language models and optimizing the weight coefficient, the problem of task conflict and uneven resource allocation in multi-task learning is solved, and the accuracy and stability of model training is improved.
Patent Information
- Application Number
- CN202510643894.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-26
AI Technical Summary
In the multi-task learning process of large language models, there are problems such as task conflict, catastrophic forgetting and uneven resource allocation, resulting in a degradation in the performance of the model in practical applications.
By determining the loss slope and loss threshold of the task, dynamically adjusting the weight coefficient, optimizing the loss value of the task model, dynamic adjustment and optimization of the task model are achieved.
It improves the accuracy and stability of task model training and improves the performance of the model in practical applications.
Smart Images

Figure CN120541573A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a task model training method and device, electronic equipment, computer-readable storage medium, and computer program product. Background Art
[0002] With the widespread application of large language models (LLM) in various technical fields, the demand for training large language models is also changing. Multi-Task Learning (MTL) of large language models aims to optimize multiple tasks simultaneously by sharing model parameters to improve the generalization ability and training efficiency of the model. In the scenario of multi-task learning, if the large language model needs to be applicable to tasks of multiple types at the same time, the large language model needs to be trained separately or simultaneously on datasets of multiple task types. However, in the process of training large language models, problems such as task conflicts, catastrophic forgetting, and uneven resource allocation often occur, which reduce the performance of large language models in practical applications. Summary of the Invention
[0003] The present disclosure provides a task model training method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0004] In a first aspect, the present disclosure provides a task model training method, comprising:
[0005] Determining, based on sample training data of a first task, a plurality of first loss values of a task model to be trained for the first task, wherein the first task is any one of a plurality of tasks used to train the task model;
[0006] determining a loss slope and a loss threshold of the first task based on a plurality of the first loss values;
[0007] Determining a weight coefficient of the first task according to the loss slope and the loss threshold;
[0008] determining a target loss value of the first task according to the weight coefficient and the second loss value, wherein the second loss value is one of the plurality of first loss values;
[0009] A total loss value of the task model is determined according to target loss values of the multiple tasks, and the task model is trained based on the total loss value.
[0010] In a second aspect, the present disclosure provides a task model training device, comprising:
[0011] A first determining module is configured to determine a plurality of first loss values of a task model to be trained for a first task based on sample training data of the first task, wherein the first task is any one of a plurality of tasks used to train the task model;
[0012] a second determining module, configured to determine a loss slope and a loss threshold of the first task based on a plurality of the first loss values;
[0013] a third determining module, configured to determine a weight coefficient of the first task according to the loss slope and the loss threshold;
[0014] a fourth determining module configured to determine a target loss value of the first task according to the weight coefficient and a second loss value, wherein the second loss value is one of the plurality of first loss values;
[0015] A fifth determination module is configured to determine a total loss value of the task model according to the target loss values of the multiple tasks, and train the task model based on the total loss value.
[0016] In a third aspect, the present disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor so that the at least one processor can perform the above-mentioned task model training method.
[0017] In a fourth aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned task model training method when executed by a processor.
[0018] In a fifth aspect, the present disclosure provides a computer program product comprising a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned task model training method.
[0019] The task model training method provided by the embodiment of the present disclosure can obtain multiple first loss values of the task model for the first task based on the sample training data of the first task, and determine the loss slope and loss threshold of the first task based on the multiple first loss values. Thus, the weight coefficient of the task model for the first task can be further determined based on the determined loss slope and loss threshold, and the weighted target loss value of the first task can be determined based on the weight coefficient, so as to dynamically adjust the weight coefficient of the first task and improve stability. Then, based on the target loss values of multiple tasks, the total loss value used to train the task model is determined, and the task model is continued to be trained based on the total loss value, thereby improving the accuracy and stability of the task model training and improving the performance of the trained task model in practical applications.
[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent to those skilled in the art by describing detailed example embodiments with reference to the accompanying drawings. In the accompanying drawings:
[0022] Figure 1 is a flowchart of a task model training method provided by an embodiment of the present disclosure;
[0023] Figure 2 is a schematic diagram of a total loss value-based training task model provided by an embodiment of the present disclosure;
[0024] Figure 3 is a processing flow chart of a task model training method provided by an embodiment of the present disclosure;
[0025] Figure 4 is a block diagram of a task model training device provided by an embodiment of the present disclosure;
[0026] Figure 5 It is a block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] To enable those skilled in the art to better understand the technical solutions of the present disclosure, exemplary embodiments of the present disclosure are described below in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0028] In the absence of conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.
[0029] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0030] The terms used herein are only used to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "made of" are used in this specification, the presence of the features, wholes, steps, operations, elements and / or components is specified, but the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof is not excluded. Similar words such as "connected" or "connected" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.
[0031] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined as such herein.
[0032] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution complies with relevant national laws and regulations (for example, the "Information Security Technology Personal Information Security Specification", etc.). For example: corresponding prescribed measures are taken to control access to personal information; the display of personal information is subject to prescribed restrictions; the purpose of using personal information does not exceed the scope of direct or reasonable connection; when using personal information, clear identity reference is eliminated to avoid precise positioning of specific individuals.
[0033] With the widespread application of large language models in various technical fields, the requirements for training large language models are also constantly changing. Multi-task learning for large language models aims to improve the model's generalization ability and training efficiency by simultaneously optimizing multiple tasks by sharing model parameters. In the multi-task learning scenario, if a large language model needs to be applicable to tasks of multiple types simultaneously, it must be trained separately or simultaneously on datasets of multiple task types. However, during the training process of large language models, problems such as task conflicts, catastrophic forgetting, and uneven resource allocation often arise, reducing the performance of large language models in practical applications.
[0034] Based on this, the embodiments of the present disclosure provide a task model training method, a task model training device, an electronic device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
[0035] The task model training method according to the embodiment of the present disclosure can be performed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server can be an independent physical server, a server cluster composed of multiple physical servers, or a cloud server capable of cloud computing. The method can be implemented by a processor calling computer-readable program instructions stored in a memory.
[0036] See also Figure 1 , Figure 1 A flowchart of a task model training method provided according to an embodiment of the present disclosure is shown, which specifically includes the following steps:
[0037] Step 102: Based on sample training data of a first task, determine a plurality of first loss values of a task model to be trained for the first task.
[0038] The first task is any one of multiple tasks used to train the task model. In practical applications, multiple tasks that the task model needs to perform during application can be determined based on actual needs, and the task model can be trained based on the determined multiple tasks. Multiple tasks include, but are not limited to, text classification tasks, text recognition tasks, topic classification tasks, object detection tasks, image processing tasks, and video processing tasks.
[0039] For each of the multiple tasks, the task model is trained based on its corresponding sample training data to obtain the first loss value of the task model at each training round. The sample training data refers to the data used to train the task model for each task, and different tasks have different corresponding sample training data. In actual applications, the sample training data of each task can be input into the task model to obtain the output of the task model, and the first loss value of each task can be calculated based on the output.
[0040] Step 104: Determine a loss slope and a loss threshold of the first task based on a plurality of the first loss values.
[0041] In current multi-task learning scenarios, models are usually trained by setting a fixed loss threshold. Although this training method can enable the trained model to have a certain task learning ability, the model cannot adapt to the magnitude differences and dynamic changes of the loss values of different tasks. As a result, in the multi-task learning scenario of the model, it causes the model to misjudge the execution status of the model for each task.
[0042] In the embodiments provided herein, after obtaining multiple first loss values corresponding to multiple tasks, the loss slope and loss threshold for each task can be determined based on the multiple first loss values of each task, thereby dynamically adjusting the loss threshold and eliminating the impact of differences in the magnitude of loss values between different tasks. Furthermore, the overall direction and rate of change of the loss for each task can be determined based on the determined loss slope. The specific implementation methods for determining the loss slope and loss threshold are explained below.
[0043] In a specific embodiment provided by the present disclosure, determining the loss slope of the first task based on the plurality of first loss values includes:
[0044] determining a plurality of third loss values from the plurality of first loss values based on a preset number of steps;
[0045] A loss slope of the first task is determined according to the plurality of third loss values.
[0046] The preset number of steps refers to the pre-set number of training rounds used to determine the third loss value from multiple first loss values. The preset number of steps can be set according to actual needs and is not limited in this embodiment of the present disclosure. The third loss value is the loss value selected from the multiple first loss values based on the preset number of steps, and the number of third loss values is the same as the preset number of steps. The loss slope refers to the slope of the straight line fitted based on the multiple third loss values.
[0047] Specifically, after obtaining multiple first loss values for each task (hereinafter referred to as the "first task"), based on a pre-set number of steps, a corresponding number of third loss values are determined from the multiple first loss values of the first task, and a straight line is fitted based on the determined third loss values, and the slope of the fitted straight line is calculated, which is the loss slope of the first task.
[0048] For example, when the number of steps is preset to k, the third loss value determined from the multiple first loss values of task i can be expressed as Among them, L is the loss value, t is the training round of the task model, that is, the number of training steps. Using linear regression, based on The fitting line L = a*d+b, a is the slope, which can be written as It is used to represent the loss slope of task i in the tth training round, reflecting the overall change direction and rate of the loss of task i.
[0049] In a specific embodiment provided by the present disclosure, the second loss value is one of a plurality of the third loss values;
[0050] Determining a loss threshold of the first task based on the plurality of first loss values includes:
[0051] Calculate, based on an exponential moving average loss algorithm, a historical first loss threshold of a previous training round to which the second loss value belongs and the second loss value to obtain a first loss threshold for the first task;
[0052] Calculating the plurality of third loss values based on a quantile loss algorithm to obtain a second loss threshold for the first task;
[0053] A loss threshold for the first task is determined according to the first loss threshold and the second loss threshold.
[0054] Among them, the exponential moving average loss algorithm, specifically the EMA (Exponential Moving Average) algorithm, can be used to smooth the changing trend of the loss. The first loss threshold refers to the loss threshold calculated based on the exponential moving average loss algorithm. The historical first loss threshold refers to the first loss threshold of the previous training round of the training round to which the second loss value of the first task belongs. For example, the first loss threshold of the currently calculated task i is Then the historical first loss threshold is The second loss value at this time is
[0055] The quantile loss algorithm, specifically the 90th percentile loss algorithm, can be used to focus on the error at the 90th percentile of the data distribution, rather than the average error, thereby reducing the impact of outliers. The second loss threshold is the loss threshold calculated based on the quantile loss algorithm. The loss threshold is the final loss threshold determined.
[0056] Specifically, to prevent misjudgment of the model state due to differences in the magnitude of the loss, the exponential moving average loss algorithm and the quantile loss algorithm can be combined to calculate the first loss threshold and the second loss threshold for the first task, respectively. The final loss threshold for the first task can be determined based on the first and second loss thresholds. The calculation method for the first loss threshold can be found in the following formula 1:
[0057]
[0058] in, is the first loss threshold of task i in the tth training round, is the historical first loss threshold of task i in the t-1th training round, is the second loss value of task i in the tth training round, and γ is the smoothing factor. The value of γ ranges from 0 to 1 and can be set according to actual needs, for example, 0.9.
[0059] The calculation method of the second loss threshold can be referred to the following formula 2:
[0060]
[0061] in, is the second loss threshold of task i in the tth training round, are k third loss values determined from multiple first loss values.
[0062] After obtaining the first loss threshold and the second loss threshold of the first task, the maximum value of the first loss threshold and the second loss threshold is determined as the loss threshold of the first task, as shown in the following formula 3:
[0063]
[0064] That is, the loss threshold of task i in the tth training round.
[0065] It should be noted that the preset number of steps can be less than or equal to the number of first loss values. If the preset number of steps is less than the number of first loss values, it is necessary to calculate the corresponding number of loss slopes and loss thresholds based on the number of times multiple third loss values can be obtained from the first loss value, and then execute the subsequent process. If the preset number of steps is equal to the number of first loss values, it means that the multiple loss values are all third loss values, and it is only necessary to calculate the loss slope and loss threshold once, and then execute the subsequent process.
[0066] For example, when the preset number of steps is 5 and the number of first loss values is 10, that is, when the preset number of steps is less than the number of first loss values, the multiple third loss values that can be obtained from the 10 first loss values are respectively Accordingly, it is necessary to calculate the corresponding loss slope and loss threshold according to the obtained third loss value; when the preset number of steps is 10 and the number of first loss values is also 10, that is, when the preset number of steps is equal to the number of first loss values, multiple third loss values can be obtained from the 10 first loss values. Accordingly, we only need to calculate Corresponding loss slope and loss threshold.
[0067] In an embodiment of the present disclosure, a preset number of steps is set to determine multiple third loss values, and the loss slope of the first task is calculated based on the multiple third loss values. The change direction and rate of the loss of the first task can be reflected according to the loss slope; the exponential moving average loss algorithm and the quantile loss algorithm are combined to calculate the first loss threshold and the second loss threshold of the first task respectively, and then the loss threshold of the first task is determined based on the first loss threshold and the second loss threshold, so as to realize dynamic adjustment of the loss threshold and avoid differences in the magnitude of the loss.
[0068] Step 106: Determine a weight coefficient of the first task according to the loss slope and the loss threshold.
[0069] After determining the loss slope and loss threshold of the first task, the weight coefficient of the first task can be determined based on the loss slope and loss threshold. Based on the weight coefficient, the task model can be adjusted in the subsequent process to adjust the loss value of executing the first task to achieve optimization of the loss value.
[0070] In a specific embodiment provided by the present disclosure, determining the weight coefficient of the first task according to the loss slope and the loss threshold includes:
[0071] determining a state response coefficient of the task model performing the first task according to the loss slope, the loss threshold, and the second loss value;
[0072] adjusting a weight response coefficient of the first task based on the state response coefficient to obtain an adjusted weight response coefficient;
[0073] A weight coefficient of the first task is determined according to the adjusted weight response coefficient.
[0074] The state response coefficient is used to characterize the responsiveness of the task model when executing the first task and is also used to adjust the weight response of the task model. The weight response coefficient is used to characterize the weight response of the task model when executing the first task and is also used to adjust the weight coefficient of the task model. The weight coefficient is used to adjust the loss value determined by the task model when executing the first task.
[0075] Specifically, after determining the loss slope and loss threshold for the first task, the state response coefficient of the task model when executing the first task is determined based on the loss slope, loss threshold, and second loss value. The weight response coefficient of the first task is then adjusted based on the state response coefficient to obtain an adjusted weight response coefficient. The weight coefficient from the previous training round is then updated based on the adjusted weight response coefficient to obtain the current weight coefficient for the first task.
[0076] Furthermore, after determining the state response coefficient of the task model when executing the first task, adjusting the weight response coefficient of the first task based on the state response coefficient can be achieved based on the following formula 4:
[0077]
[0078] in, is the weight response coefficient of task i in the t+1th training round, is the weight response coefficient of task i in the tth training round, is the state response coefficient.
[0079] The weight coefficient of the first task is determined based on the adjusted weight response coefficient, which can be implemented based on the following formula 5:
[0080]
[0081] in, is the weight coefficient of task i in the t+1th training round, is the weight coefficient of task i in the tth training round, T i is the dynamic temperature of task i.
[0082] The disclosed embodiments can determine the state response coefficient of the task model when executing the first task based on the loss slope, loss threshold and second loss value of the first task, adjust the weight response coefficient of the first task model based on the state response coefficient when the task model executes the first task, and obtain the current adjusted weight response coefficient. Thus, the weight coefficient of the first task is determined based on the weight response coefficient, which is conducive to dynamically adjusting the loss value of the first task in the subsequent process.
[0083] Furthermore, after determining the loss slope and loss threshold of the first task, the state of the task model when executing the first task can also be known, so that the training process of the task model can be intervened and adjusted in advance according to the state of the task model when executing the first task, as follows:
[0084] In a specific embodiment provided by the present disclosure, determining a state response coefficient of the task model performing the first task according to the loss slope, the loss threshold, and the second loss value includes:
[0085] Based on the preset loss coefficient and the loss threshold, determining an adjusted loss threshold and obtaining a trend significance threshold;
[0086] Determine a state of the task model when performing the first task and a state response coefficient corresponding to the state based on the loss slope, the trend significance threshold, the second loss value, the loss threshold, and the adjusted loss threshold.
[0087] The preset loss coefficient refers to a pre-set loss coefficient used to adjust the loss threshold, for example, 0.7, 0.8, etc. The trend significance threshold can be used to filter out the interference of small fluctuations on state classification, and is usually set to 0.05.
[0088] Specifically, a preset loss coefficient is multiplied by a loss threshold, and the product of the preset loss coefficient and the loss threshold is determined as the adjusted loss threshold to obtain a trend significance threshold. Based on the loss slope, the trend significance threshold, the second loss value, the loss threshold, and the adjusted loss threshold, the state of the task model when executing the first task is determined, and the corresponding state response coefficient is adjusted based on this state.
[0089] The disclosed embodiments achieve the goal of filtering noise through a trend significance threshold, reducing the possibility of misjudging the state of the task model when performing the first task, determining the state of the task model when performing the first task based on the loss slope, the trend significance threshold, the second loss value, the loss threshold and the adjusted loss threshold, improving the accuracy of determining the state of the task model when performing the first task, achieving accurate classification of the state, and dynamically adjusting the state response coefficient based on the determined state.
[0090] Furthermore, since the preset loss coefficient has a value range between 0 and 1, the adjusted loss threshold is less than the loss threshold. Based on this, in a specific embodiment provided by the present disclosure, the state of the task model when executing the first task is determined based on the loss slope, the trend significance threshold, the second loss value, the loss threshold, and the adjusted loss threshold, including:
[0091] When the loss slope is less than the negative value of the trend significance threshold and the second loss value is less than the adjusted loss threshold, determining that the state of the task model when performing the first task is a continuous progress state;
[0092] When the loss slope is less than the negative value of the trend significance threshold and the second loss value is within an interval formed by the adjusted loss threshold and the loss threshold, determining that the state of the task model when performing the first task is a learning state;
[0093] When the absolute value of the loss slope is less than the trend significance threshold, and the second loss value is less than the loss threshold, determining that the state of the task model when performing the first task is a short-term fluctuation state;
[0094] When the absolute value of the loss slope is less than the trend significance threshold and the second loss value is greater than or equal to the loss threshold, determining that the state of the task model when executing the first task is a platform warning state;
[0095] When the loss slope is greater than the trend significance threshold and the second loss value is less than the loss threshold, determining that the state of the task model when performing the first task is a potential forgetting state;
[0096] When the loss slope is greater than the trend significance threshold and the second loss value is greater than or equal to the loss threshold, it is determined that the state of the task model when performing the first task is a degraded state.
[0097] Specifically, when the loss slope is less than the negative value of the trend significance threshold, and the second loss value is less than the adjusted loss threshold, that is, and (δ is the trend significance threshold, and 0.8 is the preset loss coefficient), which means that the loss of the first task decreases rapidly and is significantly lower than the adjusted loss threshold. At this time, it is determined that the state of the task model when executing the first task is a continuous improvement state, and the state response coefficient can be 0.7.
[0098] When the loss slope is less than the negative value of the trend significance threshold, and the second loss value is within the interval formed by the adjusted loss threshold and the loss threshold, that is, and The loss of the first task decreases steadily but has not fully converged. At this time, it is determined that the state of the task model when performing the first task is a learning state, and the state response coefficient can be 1.
[0099] When the absolute value of the loss slope is less than the trend significance threshold, and the second loss value is less than the loss threshold, that is, and It means there is no obvious trend and it may be affected by noise. In this case, the state of the task model when executing the first task is determined to be a short-term fluctuation state, and the state response coefficient can be 1.
[0100] When the absolute value of the loss slope is less than the trend significance threshold, and the second loss value is greater than or equal to the loss threshold, that is, and The loss of the first task is stagnant and higher than the loss threshold, indicating possible degradation. At this time, the state of the task model when executing the first task is determined to be the platform warning state, and the state response coefficient can be 1.2.
[0101] When the loss slope is greater than the trend significance threshold and the second loss value is less than the loss threshold, that is, and This indicates that the loss of the first task unexpectedly increases, but does not exceed the loss threshold, and may begin to be forgotten. At this time, it is determined that the state of the task model when performing the first task is a potential forgetting state, and the state response coefficient can be 1.5.
[0102] When the loss slope is greater than the trend significance threshold, and the second loss value is greater than or equal to the loss threshold, that is, and The loss of the first task continues to deteriorate and exceeds the loss threshold, and the first task fails. At this time, it is determined that the state of the task model when executing the first task is a degraded state, and the state response coefficient can be 2.
[0103] The disclosed embodiment achieves the determination of the state of the task model when performing the first task based on the loss slope, the trend significance threshold, the second loss value, the loss threshold and the adjusted loss threshold, thereby achieving accurate classification of the state and dynamically adjusting the state response coefficient based on the determined state, thereby enabling early intervention and adjustment of the training process of the task model based on the state of the task model when performing the first task, thereby improving training efficiency.
[0104] Step 108: Determine a target loss value for the first task based on the weight coefficient and the second loss value, wherein the second loss value is one of the plurality of first loss values.
[0105] The target loss value refers to the loss value adjusted based on the weight coefficient. After determining the weight coefficient of the first task, the second loss value is weighted according to the weight coefficient of the first task and the second loss value to obtain the target loss value of the first task.
[0106] Step 110: Determine a total loss value of the task model according to the target loss values of the multiple tasks, and train the task model based on the total loss value.
[0107] Based on the above method, the target loss value of each task can be obtained. The total loss value of the task model can be obtained by adding up the target loss values of each task. For details, please refer to the astaxanthin formula 6:
[0108]
[0109] in, is the total loss value of the task model. After obtaining the total loss value, the model parameters of the task model are adjusted based on the total loss value, and the task model is trained continuously until the model training stopping condition is reached, thereby obtaining a trained task model. The model training stopping condition may be when the total loss value reaches a set threshold, or when the number of training rounds for the task model reaches a preset number. This disclosure does not limit the model training stopping condition for the task model.
[0110] See also Figure 2 , Figure 2 FIG. 1 shows a schematic diagram of a total loss value training task model provided according to an embodiment of the present disclosure. Figure 2 As shown, the multiple tasks used to train the task model include task 1, task 2, ... task i. For each task, the first loss value of each training round in the process of training the task model is determined respectively, and the loss slope and loss threshold of each task are calculated according to the multiple first loss values. The weight coefficient of each task is determined according to the calculated loss slope and loss threshold, and finally, each loss value is adjusted according to the weight coefficient to obtain the target loss value. The target loss value of each task is combined to obtain the total loss value, and the task model is trained based on the total loss value. After the training is completed, the trained task model can be obtained. The specific implementation methods for determining the loss slope, loss threshold, weight coefficient, target loss value and total loss value can be found in the above description, and the embodiments of the present disclosure will not be repeated here.
[0111] The disclosed embodiment achieves the goal of improving the accuracy of task model training by separately determining the target loss value for each task, combining the target loss value of each task to obtain a total loss value, and training the task model based on the total loss value.
[0112] After obtaining the trained task model, it can be used in practical applications. The following is a brief description of the application process of the task model.
[0113] In a specific embodiment provided by the present disclosure, after training the task model based on the total loss value, the method further includes:
[0114] Get the pending data of the pending tasks;
[0115] The data to be processed is input into the trained task model to obtain the task processing result of the task to be processed.
[0116] The "unprocessed task" refers to a task that needs to be processed based on the task model, and the type of the unprocessed task is one of the multiple tasks in the task model training process. The "unprocessed data" is the data used to process the unprocessed task. The task processing result is the result obtained by processing the unprocessed task. For example, if the unprocessed task is a text classification task, the unprocessed data is the text to be classified, and the task processing result is the text classification result after the unprocessed text is classified; if the unprocessed task is an image processing task, the unprocessed data is the image to be processed, and the task processing result is the image processing result after processing the unprocessed image, etc.
[0117] Specifically, in practical applications, the data to be processed of the task to be processed can be obtained in response to the task to be processed, and the data to be processed can be input into the trained task model to obtain the task processing result of the task to be processed.
[0118] The task model training method provided by the present disclosure includes: determining multiple first loss values of the task model to be trained for the first task based on sample training data of the first task, wherein the first task is any one of the multiple tasks used to train the task model; determining the loss slope and loss threshold of the first task based on the multiple first loss values; determining the weight coefficient of the first task according to the loss slope and loss threshold; determining the target loss value of the first task according to the weight coefficient and the second loss value, wherein the second loss value is one of the multiple first loss values; determining the total loss value of the task model according to the target loss values of the multiple tasks, and training the task model based on the total loss value.
[0119] The disclosed embodiment achieves the following: based on the sample training data of the first task, multiple first loss values of the task model for the first task are obtained, and the loss slope and loss threshold of the first task are determined based on the multiple first loss values. Thus, based on the determined loss slope and loss threshold, the weight coefficient of the task model for the first task can be further determined, and the weighted target loss value of the first task can be determined based on the weight coefficient, so as to dynamically adjust the weight coefficient of the first task and improve stability. Then, based on the target loss values of the multiple tasks, the total loss value used to train the task model is determined, and the task model is continued to be trained based on the total loss value, thereby improving the accuracy and stability of the task model training and improving the performance of the trained task model in practical applications.
[0120] The following combination Figure 3 , further explains the task model training method provided in the embodiment of the present disclosure. Figure 3 A processing flow chart of a task model training method provided according to an embodiment of the present disclosure is shown as follows: Figure 3 As shown, the LLM model can be fine-tuned using a multi-task training set, and during the fine-tuning training of the model, the model is trained using a multi-task validation set. Specifically, in each step of the model fine-tuning training (that is, each training round), the first loss value of the model for each task validation set is calculated. According to the multiple first loss values corresponding to each task, the loss slope and loss threshold of each task are calculated every k steps. Specifically, a linear regression method can be used to fit a straight line based on the multiple first loss values to calculate the loss slope of each task, and the exponential sliding average loss algorithm is used to calculate the first loss threshold of each task. The 90% quantile loss algorithm is used to calculate the second loss threshold of each task, and the maximum of the first loss threshold and the second loss threshold is determined as the loss threshold of each task ( Figure 3 not shown).
[0121] By judging whether the second loss value (specifically one of the multiple first loss values) is lower than the loss threshold and whether the loss slope is lower than the trend significance threshold, the state of the LLM model when executing each task is judged. The state response coefficient corresponding to each task is matched by the state of the LLM model, and the weight coefficient of different tasks is adjusted according to the state response coefficient, thereby adjusting the loss value of each task according to the weight coefficient to obtain the target loss value of each task. Finally, the target loss value of each task is added up to obtain the total loss value of multiple tasks, and the LLM model is trained based on the total loss value ( Figure 3 (not shown) until the trained LLM model is obtained and the model training is completed.
[0122] The disclosed embodiments achieve dynamic adjustment of the weight coefficient of each task, determine the total loss value used to train the LLM model according to the target loss value of each task, and train the LLM model based on the total loss value, thereby improving the accuracy and stability of LLM model training and improving the performance of the trained LLM model in practical applications.
[0123] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0124] In addition, the present disclosure also provides a task model training device, an electronic device, a computer-readable storage medium and a computer program product, all of which can be used to implement any task model training method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method section and will not be repeated here.
[0125] Figure 4 A block diagram of a task model training device provided according to an embodiment of the present disclosure is shown. Figure 4 , an embodiment of the present disclosure provides a task model training device, the task model training device comprising:
[0126] A first determining module 402 is configured to determine a plurality of first loss values of a task model to be trained for a first task based on sample training data of the first task, wherein the first task is any one of a plurality of tasks used to train the task model;
[0127] A second determining module 404 is configured to determine a loss slope and a loss threshold of the first task based on a plurality of the first loss values;
[0128] A third determination module 406 is configured to determine a weight coefficient of the first task according to the loss slope and the loss threshold;
[0129] A fourth determining module 408 is configured to determine a target loss value of the first task according to the weight coefficient and a second loss value, wherein the second loss value is one of the plurality of first loss values;
[0130] The fifth determination module 410 is configured to determine a total loss value of the task model according to the target loss values of the multiple tasks, and train the task model based on the total loss value.
[0131] Optionally, the second determining module 404 is further configured to:
[0132] determining a plurality of third loss values from the plurality of first loss values based on a preset number of steps;
[0133] A loss slope of the first task is determined according to the plurality of third loss values.
[0134] Optionally, the second loss value is one of a plurality of the third loss values;
[0135] The second determining module 404 is further configured to:
[0136] Calculate, based on an exponential moving average loss algorithm, a historical first loss threshold of a previous training round to which the second loss value belongs and the second loss value to obtain a first loss threshold for the first task;
[0137] Calculating the plurality of third loss values based on a quantile loss algorithm to obtain a second loss threshold for the first task;
[0138] A loss threshold for the first task is determined according to the first loss threshold and the second loss threshold.
[0139] Optionally, the third determining module 406 is further configured to:
[0140] determining a state response coefficient of the task model performing the first task according to the loss slope, the loss threshold, and the second loss value;
[0141] adjusting a weight response coefficient of the first task based on the state response coefficient to obtain an adjusted weight response coefficient;
[0142] A weight coefficient of the first task is determined according to the adjusted weight response coefficient.
[0143] Optionally, the third determining module 406 is further configured to:
[0144] Based on the preset loss coefficient and the loss threshold, determining an adjusted loss threshold and obtaining a trend significance threshold;
[0145] Determine a state of the task model when performing the first task and a state response coefficient corresponding to the state based on the loss slope, the trend significance threshold, the second loss value, the loss threshold, and the adjusted loss threshold.
[0146] Optionally, the adjusted loss threshold is less than the loss threshold;
[0147] The third determining module 406 is further configured to:
[0148] When the loss slope is less than the negative value of the trend significance threshold and the second loss value is less than the adjusted loss threshold, determining that the state of the task model when performing the first task is a continuous progress state;
[0149] When the loss slope is less than the negative value of the trend significance threshold and the second loss value is within an interval formed by the adjusted loss threshold and the loss threshold, determining that the state of the task model when performing the first task is a learning state;
[0150] When the absolute value of the loss slope is less than the trend significance threshold, and the second loss value is less than the loss threshold, determining that the state of the task model when performing the first task is a short-term fluctuation state;
[0151] When the absolute value of the loss slope is less than the trend significance threshold and the second loss value is greater than or equal to the loss threshold, determining that the state of the task model when executing the first task is a platform warning state;
[0152] When the loss slope is greater than the trend significance threshold and the second loss value is less than the loss threshold, determining that the state of the task model when performing the first task is a potential forgetting state;
[0153] When the loss slope is greater than the trend significance threshold and the second loss value is greater than or equal to the loss threshold, it is determined that the state of the task model when performing the first task is a degraded state.
[0154] Optionally, the device further includes a model application module configured to:
[0155] Get the pending data of the pending tasks;
[0156] The data to be processed is input into the trained task model to obtain the task processing result of the task to be processed.
[0157] The task model training device provided by the present disclosure includes: a first determination module, configured to determine multiple first loss values of the task model to be trained for the first task based on sample training data of the first task, wherein the first task is any one of the multiple tasks used to train the task model; a second determination module, configured to determine the loss slope and loss threshold of the first task based on the multiple first loss values; a third determination module, configured to determine the weight coefficient of the first task according to the loss slope and loss threshold; a fourth determination module, configured to determine the target loss value of the first task according to the weight coefficient and the second loss value, wherein the second loss value is one of the multiple first loss values; a fifth determination module, configured to determine the total loss value of the task model according to the target loss values of the multiple tasks, and train the task model based on the total loss value.
[0158] The disclosed embodiment achieves the following: based on the sample training data of the first task, multiple first loss values of the task model for the first task are obtained, and the loss slope and loss threshold of the first task are determined based on the multiple first loss values. Thus, based on the determined loss slope and loss threshold, the weight coefficient of the task model for the first task can be further determined, and the weighted target loss value of the first task can be determined based on the weight coefficient, so as to dynamically adjust the weight coefficient of the first task and improve stability. Then, based on the target loss values of the multiple tasks, the total loss value used to train the task model is determined, and the task model is continued to be trained based on the total loss value, thereby improving the accuracy and stability of the task model training and improving the performance of the trained task model in practical applications.
[0159] Each module in the task model training device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0160] Figure 5 A block diagram of an electronic device provided in an embodiment of the present disclosure.
[0161] See also Figure 5 , an embodiment of the present disclosure provides an electronic device 500, which includes: at least one processor 501; at least one memory 502, and one or more I / O interfaces 503, connected between the processor 501 and the memory 502; wherein the memory 502 stores one or more computer programs that can be executed by the at least one processor 501, and the one or more computer programs are executed by the at least one processor 501 to enable the at least one processor 501 to perform the above-mentioned task model training method.
[0162] Each module in the above-mentioned electronic device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0163] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor / processing core, implements the task model training method described above. The computer-readable storage medium may be volatile or non-volatile.
[0164] An embodiment of the present disclosure also provides a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned task model training method.
[0165] It will be understood by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable storage medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium).
[0166] As is well known to those skilled in the art, the term computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information (such as computer-readable program instructions, data structures, program modules or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technology, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those skilled in the art, communication media typically contains computer-readable program instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0167] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0168] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0169] The computer program product described herein may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).
[0170] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0171] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0172] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0173] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0174] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.
Claims
1. A task model training method, characterized in that: include: Determining, based on sample training data of a first task, a plurality of first loss values of a task model to be trained for the first task, wherein the first task is any one of a plurality of tasks used to train the task model; determining a loss slope and a loss threshold of the first task based on a plurality of the first loss values; Determining a weight coefficient of the first task according to the loss slope and the loss threshold; determining a target loss value of the first task according to the weight coefficient and the second loss value, wherein the second loss value is one of the plurality of first loss values; A total loss value of the task model is determined according to target loss values of the multiple tasks, and the task model is trained based on the total loss value.
2. The method according to claim 1, wherein Determining a loss slope of the first task based on the plurality of first loss values includes: determining a plurality of third loss values from the plurality of first loss values based on a preset number of steps; A loss slope of the first task is determined according to the plurality of third loss values.
3. The method according to claim 2, wherein The second loss value is one of a plurality of third loss values; Determining a loss threshold of the first task based on the plurality of first loss values includes: Calculate, based on an exponential moving average loss algorithm, a historical first loss threshold of a previous training round to which the second loss value belongs and the second loss value to obtain a first loss threshold for the first task; Calculating the plurality of third loss values based on a quantile loss algorithm to obtain a second loss threshold for the first task; A loss threshold for the first task is determined according to the first loss threshold and the second loss threshold.
4. The method according to claim 1, wherein Determining a weight coefficient of the first task according to the loss slope and the loss threshold includes: determining a state response coefficient of the task model performing the first task according to the loss slope, the loss threshold, and the second loss value; adjusting a weight response coefficient of the first task based on the state response coefficient to obtain an adjusted weight response coefficient; A weight coefficient of the first task is determined according to the adjusted weight response coefficient.
5. The method according to claim 4, wherein Determining a state response coefficient of the task model performing the first task according to the loss slope, the loss threshold, and the second loss value includes: Based on the preset loss coefficient and the loss threshold, determining an adjusted loss threshold and obtaining a trend significance threshold; Determine a state of the task model when performing the first task and a state response coefficient corresponding to the state based on the loss slope, the trend significance threshold, the second loss value, the loss threshold, and the adjusted loss threshold.
6. The method according to claim 5, wherein the adjusted loss threshold is less than the loss threshold; Determining a state of the task model when executing the first task according to the loss slope, the trend significance threshold, the second loss value, the loss threshold, and the adjusted loss threshold includes: When the loss slope is less than the negative value of the trend significance threshold and the second loss value is less than the adjusted loss threshold, determining that the state of the task model when performing the first task is a continuous progress state; When the loss slope is less than the negative value of the trend significance threshold and the second loss value is within an interval formed by the adjusted loss threshold and the loss threshold, determining that the state of the task model when performing the first task is a learning state; When the absolute value of the loss slope is less than the trend significance threshold, and the second loss value is less than the loss threshold, determining that the state of the task model when performing the first task is a short-term fluctuation state; When the absolute value of the loss slope is less than the trend significance threshold and the second loss value is greater than or equal to the loss threshold, determining that the state of the task model when executing the first task is a platform warning state; When the loss slope is greater than the trend significance threshold and the second loss value is less than the loss threshold, determining that the state of the task model when performing the first task is a potential forgetting state; When the loss slope is greater than the trend significance threshold and the second loss value is greater than or equal to the loss threshold, it is determined that the state of the task model when performing the first task is a degraded state.
7. The method according to any one of claims 1 to 6, wherein: After training the task model based on the total loss value, the method further includes: Get the pending data of the pending tasks; The data to be processed is input into the trained task model to obtain the task processing result of the task to be processed.
8. A task model training device, characterized in that: include: A first determining module is configured to determine a plurality of first loss values of a task model to be trained for a first task based on sample training data of the first task, wherein the first task is any one of a plurality of tasks used to train the task model; a second determining module, configured to determine a loss slope and a loss threshold of the first task based on a plurality of the first loss values; a third determining module, configured to determine a weight coefficient of the first task according to the loss slope and the loss threshold; a fourth determining module configured to determine a target loss value of the first task according to the weight coefficient and a second loss value, wherein the second loss value is one of the plurality of first loss values; A fifth determination module is configured to determine a total loss value of the task model according to the target loss values of the multiple tasks, and train the task model based on the total loss value.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 7.