Multi-task model training method and device, equipment and medium

By analyzing the task characteristics in multi-task learning and dynamically adjusting the model structure, the problems of fixed structure, inter-task interference and resource waste in the existing multi-task learning methods are solved, and more efficient and accurate multi-task training effects are achieved.

CN119940480APending Publication Date: 2025-05-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510091888.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing multi-task learning methods have problems such as fixed model structure, inter-task interference and waste of computing resources. They cannot dynamically adjust the model structure and resource allocation, which affects training efficiency and task accuracy.

Method used

By performing task feature analysis on each target task participating in multi-task learning, we can obtain task complexity, data distribution characteristics and task type, dynamically adjust the number of neurons and layers of the network layer of the multi-task model, and perform corresponding training processes to ensure that each task obtains appropriate training resources.

Benefits of technology

It avoids mutual interference between tasks, improves the training efficiency and accuracy of each task, dynamically optimizes the model structure and resource allocation, and improves the overall performance of multi-task learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940480A_ABST
    Figure CN119940480A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-task model training method and device, equipment and a medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the task feature analysis of each target task, and obtaining the task complexity, data distribution features and task types; if the target tasks with the same task type exist, marking the target tasks belonging to the same type as similar tasks, setting a first neuron number of each network layer and a first layer number of the network layers according to task complexity and data distribution characteristics of each similar task, and executing a first training process for training the multi-task model; if the target tasks of the same task type do not exist, setting the number of second neurons of each network layer and the number of second layers of the network layers according to the task complexity and data distribution characteristics of each target task, and executing a second training process for training the multi-task model; the obtained current multi-task model meeting the training ending condition is the trained multi-task model. And the task training efficiency and precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multi-task model training method, device, equipment and medium. Background Art

[0002] Multi-Task Learning (MTL) is a technique for training multiple related tasks by sharing representations and resources, aiming to improve training efficiency and enhance the performance of each task. Traditional multi-task learning methods usually share the model layers of multiple tasks, assuming that these tasks have certain similarities and can effectively cooperate under the same model framework. However, in reality, factors such as feature differences between tasks, different data distributions, and differences in task complexity often lead to interference between models, thereby affecting the accuracy of each task. Currently, some methods attempt to improve computational efficiency by sharing the underlying representation of the network and separating task-specific parts to avoid interference between tasks. However, these methods still have the following problems:

[0003] Fixed model structure: Existing multi-task learning methods usually fix the model structure at the beginning, that is, which layers are shared by each task and the number of network layers used by each task. This fixed structure design cannot be dynamically adjusted according to the characteristics of different tasks and changes in tasks during training, thus affecting training efficiency and task accuracy.

[0004] Interference between tasks: Although sharing network layers can improve computational efficiency, for scenarios where there are large differences between tasks, simple sharing strategies often lead to cross-contamination of information, affecting the performance of each task. For example, when speech recognition and text classification tasks share the underlying feature extraction layer, the task-specific information of a specific task may be blurred, resulting in a decrease in classification accuracy.

[0005] Waste of computing resources: Due to the large differences in the computational complexity and data volume of different tasks, existing methods are often unable to dynamically adjust the allocation of computing resources during training. For example, for a task with high complexity, if a fixed structure is used, the task may not obtain sufficient training resources, while for simpler tasks, computing resources may be wasted. Summary of the invention

[0006] In view of this, the purpose of the present invention is to provide a multi-task model training method, device, equipment and medium, which can avoid mutual interference between tasks and improve the training efficiency and accuracy of each task. The specific scheme is as follows:

[0007] In a first aspect, the present application discloses a multi-task model training method, comprising:

[0008] Performing task feature analysis on each target task involved in multi-task learning to obtain the task complexity of each target task, the data distribution characteristics of each target task, and the task type of each target task;

[0009] If there are target tasks of the same type among the target tasks, the target tasks of the same type are marked as similar tasks, and the number of first neurons in each network layer and the number of first layers in the network layer of the multi-task model are set according to the task complexity and data distribution characteristics of each similar task, and then a first training process of training the multi-task model using the similar tasks is executed;

[0010] If there is no target task of the same type among the target tasks, the number of second neurons in each network layer and the number of second layers in the network layer of the multi-task model are set according to the task complexity and data distribution characteristics of each target task, and then a second training process of training the multi-task model using the target tasks is executed;

[0011] The current multi-task model that reaches the training end condition is obtained as the trained multi-task model.

[0012] Optionally, the multi-task model training method further includes:

[0013] A multi-task model is constructed including a first network layer, a second network layer, and a third network layer; wherein the first network layer is a network layer shared by similar tasks, the second network layer is a network layer independently established for similar tasks, and the third network layer is a network layer independently established for dissimilar tasks, and the number of neurons in each network layer and the number of network layers of the multi-task model are not fixed.

[0014] Optionally, the multi-task model training method further includes:

[0015] When the task type of the similar task is a text task type, the first network layer of the multi-task model is set to include a shared input layer and a shared common feature extraction layer, and the second network layer of the multi-task model is set to include a separate fully connected layer corresponding to each similar task.

[0016] Optionally, the multi-task model training method further includes:

[0017] If there is no target task of the same type among the target tasks, the third network layer of the multi-task model is set to include a separate input layer, a separate feature extraction layer and a separate fully connected layer corresponding to each target task.

[0018] Optionally, in the multi-task model training method, the first number of neurons and the first number of layers of the first network layer and the second network layer are both smaller than the second number of neurons and the second number of layers of the third network layer.

[0019] Optionally, executing a training process for training the multi-task model includes:

[0020] Using each of the target tasks carrying the real task execution results to train the multi-task model including the number of neurons and the number of network layers, and determining the task accuracy of the current multi-task model according to the difference between the execution result of the output of the current multi-task model and the real task execution result;

[0021] When the task accuracy is lower than the preset accuracy, the number of neurons in the network layer corresponding to the current target task and the number of layers of the network layer are increased, and the process jumps to the step of training the multi-task model including the number of neurons and the number of layers of the network layer using the target tasks carrying the real task execution results until the training end condition is reached, and then the current multi-task model is output.

[0022] Optionally, after determining the task accuracy of the current multi-task model according to the difference between the execution result of the output of the current multi-task model and the actual task execution result, the method further includes:

[0023] When the task accuracy is higher than the preset accuracy, the number of neurons in the network layer corresponding to the current target task and the number of layers in the network layer are reduced, and the process jumps to the step of training the multi-task model including the number of neurons and the number of layers in the network layer using the target tasks carrying the real task execution results until the training end condition is reached, and then the current multi-task model is output.

[0024] In a second aspect, the present application discloses a multi-task model training device, comprising:

[0025] A feature analysis module is used to perform task feature analysis on each target task involved in multi-task learning to obtain the task complexity of each target task, the data distribution characteristics of each target task, and the task type of each target task;

[0026] A first training module is used for marking target tasks of the same type as similar tasks if there are target tasks of the same type among the target tasks, and setting the number of first neurons in each network layer and the number of first layers in the network layer of the multi-task model according to the task complexity and data distribution characteristics of each similar task, and then executing a first training process of training the multi-task model using the similar tasks;

[0027] A second training module is used for setting the second number of neurons in each network layer and the second number of layers in the network layer of the multi-task model according to the task complexity and data distribution characteristics of each target task if there is no target task of the same type among the target tasks, and then executing a second training process of training the multi-task model using the target tasks;

[0028] The model output module is used to obtain the current multi-task model that meets the training end condition as the trained multi-task model.

[0029] In a third aspect, the present application discloses an electronic device, comprising:

[0030] Memory, used to store computer programs;

[0031] A processor is used to execute the computer program to implement the steps of the multi-task model training method disclosed above.

[0032] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the multi-task model training method disclosed above are implemented.

[0033] It can be seen that the present application discloses a multi-task model training method, including: performing task feature analysis on each target task participating in multi-task learning to obtain the task complexity of each target task, the data distribution characteristics of each target task, and the task type of each target task; if there are target tasks of the same type among the target tasks, the target tasks of the same type are marked as similar tasks, and the first number of neurons in each network layer and the first number of layers of the network layer of the multi-task model are set according to the task complexity and data distribution characteristics of each similar task, and then the first training process of training the multi-task model using the similar tasks is executed; if there are no target tasks of the same type among the target tasks, the second number of neurons in each network layer and the second number of layers of the network layer of the multi-task model are set according to the task complexity and data distribution characteristics of each target task, and then the second training process of training the multi-task model using the target tasks is executed; the current multi-task model that meets the training end condition is obtained as the trained multi-task model. It can be seen that by dynamically optimizing the model structure of the multi-task model according to the complexity of the target task, data characteristics and similarities between tasks, executing the corresponding training process, ensuring that different tasks obtain appropriate training and network model resources, the overall performance of multi-task learning can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0035] Figure 1 A flow chart of a multi-task model training method disclosed in this application;

[0036] Figure 2 A schematic diagram of training resource allocation disclosed in this application;

[0037] Figure 3 This is a schematic diagram of the structure of a multi-task model training device disclosed in this application;

[0038] Figure 4 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0040] Multi-task learning is a technique for training multiple related tasks by sharing representations and resources, aiming to improve training efficiency and enhance the performance of each task. Traditional multi-task learning methods usually share the model layers of multiple tasks, assuming that these tasks have certain similarities and can effectively cooperate under the same model framework. However, in reality, factors such as feature differences between tasks, different data distributions, and differences in task complexity often lead to interference between models between multiple tasks, thereby affecting the accuracy of each task. Currently, some methods attempt to improve computational efficiency by sharing the underlying representation of the network and separating task-specific parts to avoid interference between tasks. However, these methods still have the following problems:

[0041] Fixed model structure: Existing multi-task learning methods usually fix the model structure at the beginning, that is, which layers are shared by each task and the number of network layers used by each task. This fixed structure design cannot be dynamically adjusted according to the characteristics of different tasks and changes in tasks during training, thus affecting training efficiency and task accuracy.

[0042] Interference between tasks: Although sharing network layers can improve computational efficiency, for scenarios where there are large differences between tasks, simple sharing strategies often lead to cross-contamination of information, affecting the performance of each task. For example, when speech recognition and text classification tasks share the underlying feature extraction layer, the task-specific information of a specific task may be blurred, resulting in a decrease in classification accuracy.

[0043] Waste of computing resources: Due to the large differences in the computational complexity and data volume of different tasks, existing methods are often unable to dynamically adjust the allocation of computing resources during training. For example, for a task with high complexity, if a fixed structure is used, the task may not obtain sufficient training resources, while for simpler tasks, computing resources may be wasted.

[0044] To this end, the present invention also discloses a multi-task model training scheme, which can avoid mutual interference between tasks and improve the training efficiency and accuracy of each task.

[0045] Reference Figure 1 As shown, an embodiment of the present invention discloses a multi-task model training method, comprising:

[0046] Step S11: performing task feature analysis on each target task involved in multi-task learning to obtain the task complexity of each target task, the data distribution characteristics of each target task, and the task type of each target task.

[0047] In this embodiment, in order to better train the multi-task model, task feature analysis is performed on each target task participating in multi-task learning to obtain the specific task features after feature analysis, and then determine the appropriate network structure of the multi-task model according to the specific task features. Specifically:

[0048] Based on task objectives and rule evaluation, task feature analysis is performed on each target task of multi-task learning to obtain the task complexity of each target task. Task complexity is an indicator to measure the difficulty of a target task. Task complexity involves the amount of model calculation required to complete the task, the depth and breadth of understanding of data features. Furthermore, the steps to obtain the task complexity of each target task are as follows: First, clarify the goal of the target task, and determine the task complexity based on the difficulty of achieving the task goal.

[0049] For example, in natural language processing, if the target task is a text classification task, then the goal of the task is to classify the text into predefined categories, such as classifying news into sports, entertainment, finance, and other categories. This goal is relatively clear and simple, because it is mainly a coarse-grained division of text topics, so the initial judgment complexity is low. The semantic role labeling task requires analysis of the semantic role played by each word in the sentence, such as subject, predicate, object, time adverbial, etc. This involves an in-depth understanding of the grammatical and semantic structure of the sentence, as well as the complex relationship between different words, so it can be inferred that the task complexity is high; the complexity is judged according to the operation rules of the task. Taking the text generation task as an example, generating text that conforms to grammatical and semantic rules according to given conditions (such as theme, style, etc.) involves multiple complex links such as vocabulary selection, sentence construction, and semantic coherence, and the complexity is high; while the simple text classification task only needs to match the keywords in the text with the predefined category labels, and the complexity is low.

[0050] In addition, if most studies in the relevant literature show that a certain task (such as named entity recognition) usually requires a complex model structure and a large amount of computing resources to achieve good performance, then it can be preliminarily judged that the task is highly complex.

[0051] Perform task feature analysis on each target task involved in multi-task learning to obtain the data distribution characteristics of each target task. Specifically, calculate the basic statistics of the task data in the target task, such as mean, variance, median, etc. In image data, if the mean and variance of pixel values ​​are significantly different between different categories, this indicates different data distributions. For example, in the cat and dog classification task, the image of a cat may be lighter in color overall, with a lower pixel mean and a smaller variance; while the image of a dog may be richer in color, with a larger pixel mean and variance. For text data, the frequency distribution of vocabulary can be statistically analyzed. If there are significant differences in vocabulary frequencies between different categories of text (such as technology and entertainment), this reflects the data distribution characteristics. For example, the frequency of technical vocabulary is higher in technology texts, while the frequency of celebrity names and film and television vocabulary is higher in entertainment texts. For low-dimensional data (such as simple two-dimensional or three-dimensional data), scatter plots, bar graphs, etc. can be used for visualization. For example, in a simple two-class data set, use the two features of the data as the horizontal and vertical coordinates to draw a scatter plot, observe the distribution of the two types of data points, whether they are clustered together or clearly separated, so as to understand whether the data distribution is concentrated or discrete. For high-dimensional data (such as word vector representation of text or feature vector of image), dimensionality reduction techniques (such as principal component analysis PCA) can be used to reduce the data to two or three dimensions and then visualize it. This helps to intuitively see the distribution of data in low-dimensional space, determine whether there is clustering phenomenon, whether the data has obvious boundaries, etc.

[0052] Perform task feature analysis on each target task involved in multi-task learning to obtain the task type of each target task. Specifically, perform similarity analysis between tasks and compare the data types of tasks, such as determining whether the target tasks are all text data, image data, or other types of data; if they are text data, further compare the format and source of the text. For example, are they all news texts or multiple text types such as academic papers and novels? At the task goal level, clarify the output type of the task. If they are all classification tasks, compare the category properties of the classification. For example, are they all binary classification tasks or multi-classification tasks, and whether there is a logical connection between the categories of the classification. Taking natural language processing as an example, although text sentiment classification (positive / negative) and text topic classification (sports / entertainment / technology, etc.) are both classification tasks, they have different category properties. Consider the semantic similarity of task goals. For example, machine translation and text rewriting tasks have certain similarities in semantics, both of which perform some form of transformation on the text content; while text classification and text generation tasks have large differences in goals.

[0053] Step S12: If there are target tasks of the same type among the target tasks, the target tasks of the same type are marked as similar tasks, and the number of first neurons in each network layer and the number of first layers in the network layer of the multi-task model are set according to the task complexity and data distribution characteristics of each similar task, and then the first training process of training the multi-task model using the similar tasks is executed.

[0054] In this embodiment, before training, a multi-task model including a first network layer, a second network layer, and a third network layer is constructed; wherein the first network layer is a network layer shared by similar tasks, the second network layer is a network layer established separately for similar tasks, and the third network layer is a network layer established separately for dissimilar tasks, and the number of neurons in each network layer of the multi-task model and the number of layers of the network layer are not fixed. It can be understood that the multi-task model structure design is specifically designed as follows: for similar tasks with similar data characteristics and requirements, some network layers are shared, such as the underlying feature extraction layer, to reduce redundant calculations of the model and improve training efficiency. Task-specific layer design: for parts with large differences between tasks, such as task-specific feature extraction layers or decision layers, a separation mechanism is adopted to design specific network layers for each target task separately to ensure that each target task can benefit from the most appropriate structure. Dynamic adjustment of the number of layers and width: dynamically adjust the number of layers and width of the model according to the complexity of the task and the change in the accuracy of the task during training. For example, if the accuracy of the NER task is found to be low during training, the network depth and width of the task can be automatically increased.

[0055] In this embodiment, when the task type of the similar task is a text task type, the first network layer of the multi-task model is set to include a shared input layer and a shared common feature extraction layer, and the second network layer of the multi-task model is set to include a separate fully connected layer corresponding to each similar task. It can be understood that if each target task is a text classification task and a sentiment analysis task, since in the feature analysis part, it is determined that the task types of both are text task types, the two are marked as similar tasks. In this way, when the task type of the similar task is a text task type, the multi-task model is used as a classification model and an analysis model to construct the model architecture. Specifically, in the shared layer design, the common layer of the two is designed as a shared input layer and a shared underlying feature extraction layer (shared common feature extraction layer), and for the non-common parts of different tasks, it is still necessary to provide a separate fully connected layer, a separate feature extraction layer, etc. for the text classification task and the sentiment analysis task, and the number of neurons in the network layer and the number of network layers are set according to the task complexity and data distribution characteristics of the corresponding tasks.

[0056] Specifically, for similar tasks such as text classification and sentiment analysis, the input layer of the multi-task model: Since both data types are text, the input layer can be shared. The shared input layer converts the text data of the task into a vector form that the multi-task model can process, for example, using word embedding technology to map each word to a low-dimensional vector. Shared underlying feature extraction layer: Some basic feature extraction layers can be shared, for example, using the first few layers of the convolutional neural network, this part of the network layer can extract common features of the text, such as lexical and syntactic features. These common features are important for text classification and sentiment analysis. The task-specific layer is designed as follows: Text classification task-specific layer: After the shared underlying feature extraction layer, some separate fully connected layers dedicated to text classification are added. These separate fully connected layers further learn features related to the text topic based on the extracted common features to determine the category to which the text belongs. For example, the features are combined and transformed through the fully connected layer to output the probability that the text belongs to each topic category. Sentiment analysis task-specific layer: Also after the shared layer, a fully connected layer is designed separately for the sentiment analysis task. These layers focus on learning sentiment-related features, such as the sentiment tendency of words, the sentiment intensity of sentences, etc., to output the sentiment category probability of the text.

[0057] In this embodiment, the multi-task model including the number of neurons and the number of network layers is trained using each of the target tasks that carry the actual task execution results, and the task accuracy of the current multi-task model is determined based on the difference between the execution result of the output of the current multi-task model and the actual task execution result; when the task accuracy is lower than the preset accuracy, the number of neurons in the network layer and the number of network layers corresponding to the current target task are increased, and the step of training the multi-task model including the number of neurons and the number of network layers using each of the target tasks that carry the actual task execution results is jumped to until the training end condition is reached, and the current multi-task model is output. It can be understood that the steps for training a multi-task model for similar tasks are as follows:

[0058] Collect a large amount of text data and annotate the category labels for text classification and the sentiment labels for sentiment analysis. Preprocess the text data, including operations such as word segmentation, removal of stop words, and conversion of text into fixed-length sequences, to meet the input requirements of the model. Divide the data set into training, validation, and test sets. Initialize the parameters of the shared layer and the task-specific layer, such as using random initialization or pre-trained parameters. Input the training set data into the model and calculate the loss functions for the text classification task and the sentiment analysis task respectively. For example, the cross entropy loss function can be used for the text classification task, and the cross entropy loss function can also be used for the sentiment analysis task. Calculate the weighted sum of the two task loss functions as the total loss function of the entire multi-task model. The weights of the two task loss functions can be adjusted according to factors such as the importance of the task or the amount of data. Use optimization algorithms such as stochastic gradient descent, Adam (Adaptive Moment Estimation), etc. Update the parameters of the multi-task model according to the total loss function, including the parameters of the shared layer and the task-specific layer. During the training process, the parameters of the shared layer will be updated under the influence of both tasks, while the parameters of the task-specific layer will be updated mainly under the influence of their respective tasks. During the training process, the validation set data is regularly used to evaluate the performance of the model on text classification and sentiment analysis tasks, such as accuracy, recall and other indicators. According to the evaluation results of the validation set, the training parameters (such as learning rate) or model structure (such as increasing or decreasing the number of neurons in certain layers) are adjusted to prevent overfitting or underfitting. Then the test set data is used to evaluate the trained multi-task model, and the accuracy, recall, F1 value and other indicators of the text classification task and sentiment analysis task on the test set are calculated respectively to comprehensively evaluate the performance of the model on these two tasks, that is, to determine the task accuracy of the current multi-task model.

[0059] In this embodiment, after determining the task accuracy of the current multi-task model according to the difference between the execution result of the output of the current multi-task model and the actual task execution result, it also includes: when the task accuracy is higher than the preset accuracy, reducing the number of neurons in the network layer corresponding to the current target task and the number of layers in the network layer, and jumping to the step of training the multi-task model including the number of neurons and the number of layers in the network layer using each of the target tasks carrying the actual task execution results until the training end condition is reached, and then outputting the current multi-task model. It can be understood that if the task accuracy of the current multi-task model is too high, in order not to waste computing resources during the training process, it is necessary to reduce the number of neurons in the network layer corresponding to the current target task and the number of layers in the network layer, that is, to reduce the width and depth of the model. Specifically, in multi-task learning, each target task shares part of the network layer and computing resources, but dynamically adjusts the computing resources according to the accuracy of each task. For example, for tasks with lower task accuracy, increase computing resources (such as increasing the training batch, the number of layers or the number of neurons), and for tasks with higher accuracy, reduce the allocation of computing resources. Figure 2 As shown in the figure, the shared network layer can be used as a shared layer for Task 1, Task 2, and Task 3. In this way, during the model training process, on the sharable network, these target tasks are uniformly processed through the shared network layer.

[0060] Step S13: If there is no target task of the same type among the target tasks, the number of second neurons in each network layer and the number of second layers in the network layer of the multi-task model are set according to the task complexity and data distribution characteristics of each target task, and then the second training process of training the multi-task model using the target tasks is executed.

[0061] In this embodiment, if there is no target task of the same type in each target task, the third network layer of the multi-task model is set to include a separate input layer, a separate feature extraction layer and a separate fully connected layer corresponding to each target task. It is understandable that if the target tasks participating in multi-task learning are not of the same task type, a shared network layer cannot be set. Taking the text classification task and machine translation task in natural language processing as examples, their data processing methods and target outputs are completely different. Text classification is mainly to judge the category to which the text belongs, while machine translation is to convert the text of one language into the text of another language. Due to this huge difference, it is difficult for them to share parts such as the input layer and the underlying feature extraction layer. Therefore, it is necessary to set a separate input layer, a separate feature extraction layer and a separate fully connected layer for each target task, and then set the corresponding number of network layer neurons and the number of network layers according to the task complexity and data distribution characteristics.

[0062] The purpose of this is to avoid interference between different tasks. Because if the network layer is forced to be shared, it will cause cross-contamination of information. For example, when sharing the underlying feature extraction layer, the information of one task may be mistakenly passed to another task, affecting the performance of each task, causing the task-specific information of a specific task to be blurred, thereby reducing the accuracy of each task. Therefore, in the absence of similarity between tasks, a separate network layer design is used to ensure the learning effect of each task.

[0063] In this embodiment, the first number of neurons and the first number of layers of the first network layer and the second network layer are smaller than the second number of neurons and the second number of layers of the third network layer. It can be understood that, according to the similar tasks mentioned above, dissimilar tasks have higher task complexity. Therefore, the number of neurons and the number of layers of the first network layer and the second network layer corresponding to similar tasks are smaller than the number of neurons and the number of layers of the third network layer. For example: NER tasks require more contextual information due to the characteristics of the task, so they require additional specific network layers and more neurons and layers.

[0064] In this embodiment, the multi-task model including the number of neurons and the number of network layers is trained using each of the target tasks carrying the real task execution results, and the task accuracy of the current multi-task model is determined based on the difference between the execution result of the output of the current multi-task model and the real task execution result; when the task accuracy is lower than the preset accuracy, the number of neurons in the network layer corresponding to the current target task and the number of network layers are increased, and the step of training the multi-task model including the number of neurons and the number of network layers using each of the target tasks carrying the real task execution results is jumped to until the training end condition is reached, and the current multi-task model is output. It can be understood that, like training similar tasks, the multi-task model with the model network layer, the number of network layers, and the number of neurons on the network layer is trained, and its task accuracy is obtained, and when the task accuracy is too low, its training resources (the number of neurons and the number of network layers) are increased until the training end condition is reached, and the current multi-task model is output.

[0065] In this embodiment, after determining the task accuracy of the current multi-task model based on the difference between the execution result of the output of the current multi-task model and the actual task execution result, it also includes: when the task accuracy is higher than the preset accuracy, reducing the number of neurons in the network layer corresponding to the current target task and the number of layers in the network layer, and jumping to the step of training the multi-task model including the number of neurons and the number of layers in the network layer using each of the target tasks carrying the actual task execution results until the training end condition is reached, and then outputting the current multi-task model. It can be understood that when the task accuracy is too high, in order not to waste computing resources during the training process, it is necessary to reduce the number of neurons in the network layer corresponding to the current target task and the number of layers in the network layer, that is, to reduce the width and depth of the model.

[0066] Step S14: obtaining the current multi-task model that has reached the training end condition as the trained multi-task model.

[0067] In this embodiment, during the completion of multiple rounds of training iterations, the performance of the current multi-task model on the training set and the validation set is continuously monitored. When the model meets the pre-set specific conditions in the performance evaluation index, it can be determined that the training end condition has been reached. These conditions usually include but are not limited to the task accuracy reaching a predetermined threshold, for example: the accuracy of the text classification task is stably maintained at more than 90%, and the fluctuation is very small in multiple consecutive training rounds; and the loss function value converges to a minimum value, such as after multiple rounds of optimization, the loss function value has stabilized at a lower level, and there is no obvious decline in subsequent training. When these conditions are met, the multi-task model at this time is obtained and used as the multi-task model after training. It can be seen that the model integrates effective learning results for multiple tasks, and can efficiently and accurately process different types of tasks in actual application scenarios, providing reliable support for subsequent actual task execution.

[0068] It can be seen that the present application discloses a multi-task model training method, including: performing task feature analysis on each target task participating in multi-task learning to obtain the task complexity of each target task, the data distribution characteristics of each target task, and the task type of each target task; if there are target tasks of the same type among the target tasks, the target tasks of the same type are marked as similar tasks, and the first number of neurons in each network layer and the first number of layers of the network layer of the multi-task model are set according to the task complexity and data distribution characteristics of each similar task, and then the first training process of training the multi-task model using the similar tasks is executed; if there are no target tasks of the same type among the target tasks, the second number of neurons in each network layer and the second number of layers of the network layer of the multi-task model are set according to the task complexity and data distribution characteristics of each target task, and then the second training process of training the multi-task model using the target tasks is executed; the current multi-task model that meets the training end condition is obtained as the trained multi-task model. It can be seen that by dynamically optimizing the model structure of the multi-task model according to the complexity of the target task, data characteristics and similarities between tasks, executing the corresponding training process, ensuring that different tasks obtain appropriate training and network model resources, the overall performance of multi-task learning can be improved.

[0069] Reference Figure 3 As shown, the present invention also discloses a multi-task model training device, including:

[0070] A feature analysis module 11 is used to perform task feature analysis on each target task involved in multi-task learning to obtain the task complexity of each target task, the data distribution characteristics of each target task, and the task type of each target task;

[0071] The first training module 12 is used for marking the target tasks of the same type as similar tasks if there are target tasks of the same type among the target tasks, and setting the number of first neurons in each network layer and the number of first layers in the network layer of the multi-task model according to the task complexity and data distribution characteristics of each similar task, and then executing a first training process of training the multi-task model using the similar tasks;

[0072] A second training module 13 is used for setting the second number of neurons in each network layer and the second number of layers in the network layer of the multi-task model according to the task complexity and data distribution characteristics of each target task if there is no target task of the same type among the target tasks, and then executing a second training process of training the multi-task model using the target tasks;

[0073] The model output module 14 is used to obtain the current multi-task model that meets the training end condition as the trained multi-task model.

[0074] It can be seen that the present application discloses a task feature analysis of each target task participating in multi-task learning to obtain the task complexity of each target task, the data distribution characteristics of each target task, and the task type of each target task; if there are target tasks of the same type among the target tasks, the target tasks of the same type are marked as similar tasks, and the first number of neurons in each network layer and the first number of layers of the network layer of the multi-task model are set according to the task complexity and data distribution characteristics of each similar task, and then the first training process of training the multi-task model using the similar tasks is executed; if there are no target tasks of the same type among the target tasks, the second number of neurons in each network layer and the second number of layers of the network layer of the multi-task model are set according to the task complexity and data distribution characteristics of each target task, and then the second training process of training the multi-task model using the target tasks is executed; the current multi-task model that meets the training end condition is obtained as the trained multi-task model. It can be seen that by dynamically optimizing the model structure of the multi-task model according to the complexity of the target task, data characteristics and similarities between tasks, executing the corresponding training process, ensuring that different tasks obtain appropriate training and network model resources, the overall performance of multi-task learning can be improved.

[0075] Furthermore, the present application also discloses an electronic device. Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.

[0076] Figure 4 A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the multi-task model training method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0077] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0078] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0079] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0080] Among them, the operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20 to realize the operation and processing of the processor 21 on the massive data 223 in the memory 22, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the multi-task model training method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks. In addition to data transmitted from an external device received by the electronic device, the data 223 can also include data collected by its own input and output interface 25, etc.

[0081] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the multi-task model training method disclosed above is implemented. For the specific steps of the method, reference can be made to the corresponding contents disclosed in the aforementioned embodiments, and no further description will be given here.

[0082] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0083] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. The steps of the method or algorithm described in conjunction with the embodiments disclosed herein can be implemented directly with hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory RAM (Random Access Memory), memory, read-only memory ROM (Read Only Memory), electrically programmable EPROM (Electrically Programmable Read Only Memory), electrically erasable programmable EEPROM (ElectricErasable Programmable Read Only Memory), register, hard disk, removable disk, CD-ROM (CompactDisc-Read Only Memory), or any other form of storage medium known in the technical field.

[0084] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0085] The scheme provided by the present invention is introduced in detail above. Specific examples are used in this article to illustrate the principle and implementation mode of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation mode and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A multi-task model training method, characterized in that: include: Performing task feature analysis on each target task involved in multi-task learning to obtain the task complexity of each target task, the data distribution characteristics of each target task, and the task type of each target task; If there are target tasks of the same type among the target tasks, the target tasks of the same type are marked as similar tasks, and the number of first neurons in each network layer and the number of first layers in the network layer of the multi-task model are set according to the task complexity and data distribution characteristics of each similar task, and then a first training process of training the multi-task model using the similar tasks is executed; If there is no target task of the same type among the target tasks, the number of second neurons in each network layer and the number of second layers in the network layer of the multi-task model are set according to the task complexity and data distribution characteristics of each target task, and then a second training process of training the multi-task model using the target tasks is executed; The current multi-task model that reaches the training end condition is obtained as the trained multi-task model.

2. The multi-task model training method according to claim 1, characterized in that: Also includes: A multi-task model is constructed including a first network layer, a second network layer, and a third network layer; wherein the first network layer is a network layer shared by similar tasks, the second network layer is a network layer independently established for similar tasks, and the third network layer is a network layer independently established for dissimilar tasks, and the number of neurons in each network layer and the number of network layers of the multi-task model are not fixed.

3. The multi-task model training method according to claim 2, characterized in that: Also includes: When the task type of the similar task is a text task type, the first network layer of the multi-task model is set to include a shared input layer and a shared common feature extraction layer, and the second network layer of the multi-task model is set to include a separate fully connected layer corresponding to each similar task.

4. The multi-task model training method according to claim 2, characterized in that: Also includes: If there is no target task of the same type among the target tasks, the third network layer of the multi-task model is set to include a separate input layer, a separate feature extraction layer and a separate fully connected layer corresponding to each target task.

5. The multi-task model training method according to claim 2, characterized in that: The first number of neurons and the first number of layers in the first network layer and the second network layer are both smaller than the second number of neurons and the second number of layers in the third network layer.

6. The multi-task model training method according to any one of claims 1 to 5, characterized in that: The training process of executing the multi-task model includes: Using each of the target tasks carrying the real task execution results to train the multi-task model including the number of neurons and the number of network layers, and determining the task accuracy of the current multi-task model according to the difference between the execution result of the output of the current multi-task model and the real task execution result; When the task accuracy is lower than the preset accuracy, the number of neurons in the network layer corresponding to the current target task and the number of layers of the network layer are increased, and the process jumps to the step of training the multi-task model including the number of neurons and the number of layers of the network layer using the target tasks carrying the real task execution results until the training end condition is reached, and then the current multi-task model is output.

7. The multi-task model training method according to claim 6, characterized in that: After determining the task accuracy of the current multi-task model according to the difference between the execution result of the output of the current multi-task model and the actual task execution result, the method further includes: When the task accuracy is higher than the preset accuracy, the number of neurons in the network layer corresponding to the current target task and the number of layers in the network layer are reduced, and the process jumps to the step of training the multi-task model including the number of neurons and the number of layers in the network layer using the target tasks carrying the real task execution results until the training end condition is reached, and then the current multi-task model is output.

8. A multi-task model training device, characterized in that: include: A feature analysis module is used to perform task feature analysis on each target task involved in multi-task learning to obtain the task complexity of each target task, the data distribution characteristics of each target task, and the task type of each target task; A first training module is used for marking target tasks of the same type as similar tasks if there are target tasks of the same type among the target tasks, and setting the number of first neurons in each network layer and the number of first layers in the network layer of the multi-task model according to the task complexity and data distribution characteristics of each similar task, and then executing a first training process of training the multi-task model using the similar tasks; A second training module is used for setting the second number of neurons in each network layer and the second number of layers in the network layer of the multi-task model according to the task complexity and data distribution characteristics of each target task if there is no target task of the same type among the target tasks, and then executing a second training process of training the multi-task model using the target tasks; The model output module is used to obtain the current multi-task model that meets the training end condition as the trained multi-task model.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the multi-task model training method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, the steps of the multi-task model training method as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Model training method and system based on artificial intelligence, medium and product

    CN121979693A

  • An artificial intelligence-based model training method, system, medium and product

    CN121979693B