Multitask learning resource optimization method based on lottery hypothesis
Through the multi-task learning resource optimization method based on the lottery assumption, specific parts of the all-around expert layer are activated and optimized, and the problem of difficult to balance resource allocation and feature extraction in the existing technology is solved, and the flexibility of optimization of computing resources and information sharing among tasks is realized, and the overall performance and efficiency of multi-task learning is improved.
Patent Information
- Application Number
- CN202510054498.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-14
AI Technical Summary
When existing multi-task learning methods deal with tasks with large differences in task correlation or different complexity, it is difficult to effectively balance the resource allocation and feature extraction requirements between tasks, resulting in waste of computing resources and interference between tasks, affecting the overall performance of the model.
Using a multi-task learning resource optimization method based on lottery assumptions, the specific parts of the all-around expert layer are efficiently activated and optimized through mask selection and lottery assumption principles, making the network more sparse and achieving task-specific advanced feature extraction and computing resource optimization.
Effectively prune neural network model parameters, reduce computing costs and storage overhead, improve model efficiency, enhance the flexibility of information sharing among tasks, reduce resource waste and redundant calculations, and improve the overall performance and efficiency of multi-task learning.
Smart Images

Figure CN119990232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-task learning and neural network optimization, and in particular to a multi-task learning resource optimization method based on lottery hypothesis. Background Art
[0002] The lottery ticket hypothesis is a pruning method used in the field of deep learning to compress neural network models. The lottery ticket hypothesis assumes that a relatively complex deep neural network has a relatively optimized sparse subnetwork structure (called winning tickets), which can be applied to model compression. Compared with the original network, the number of parameters and complexity of the sparse subnetwork are much lower, but the inference accuracy is basically the same, thereby reducing the computational cost and storage overhead of the model.
[0003] At present, multi-task learning is widely used in scenarios where multiple related or unrelated tasks need to be processed simultaneously, and has become an important technical means to improve model performance and optimize the utilization of computing resources. The core idea of multi-task learning is to improve learning efficiency and enhance the generalization ability of the model by sharing relevant information between tasks. Traditional multi-task learning methods mainly include hard parameter sharing, soft parameter sharing, expert sharing, and sparse sharing. These methods have achieved information sharing and optimization of computing resources between tasks to a certain extent, but when dealing with tasks with large differences in task relevance or different complexities, it is still difficult to effectively balance the resource allocation and feature extraction requirements between tasks. These methods may lead to waste of computing resources and interference between tasks, thus affecting the overall performance of the model.
[0004] In recent years, the Multi-gate Mixture-of-Experts (MMoE) model has received widespread attention as an improved multi-task learning method. The MMoE model introduces multiple expert networks and independent task gating networks, so that each task can select and combine different expert outputs according to its own needs. Specifically, each task in the MMoE model has an independent gating network, which determines the weight of each expert based on the input features of the task, thereby dynamically selecting and activating experts suitable for the current task. This mechanism enhances the flexibility of the model to a certain extent and avoids direct conflicts and information interference between tasks.
[0005] Although the multi-gate mixture of experts (MMoE) model achieves task-specific expert selection by introducing independent task gating networks, it still faces the problems of high consumption of computing resources, redundant calculations and lack of flexibility. When processing multiple tasks, the MMoE model needs to calculate the outputs of multiple experts for each task at the same time, which leads to resource waste and increased computational overhead. In addition, the fixed expert network and task gating network structure are difficult to adapt to task changes and complex task requirements, limiting the dynamic adjustment ability of the model. Especially when the task relevance is low, the existing MMoE model still has deficiencies in feature sharing between tasks and cannot effectively utilize shared information, resulting in the overall efficiency and performance not being optimal. Therefore, there is an urgent need for a mechanism that can dynamically adjust resource allocation and expert selection to improve the flexibility and resource utilization efficiency of multi-task learning, flexibly select and activate different parts of the expert network according to task requirements, reduce the waste of computing resources, and enhance information sharing between tasks, thereby improving the overall performance and efficiency of multi-task learning. Summary of the invention
[0006] The technical problem to be solved by the present invention is to provide a multi-task learning resource optimization method based on the lottery hypothesis in view of the deficiencies of the above-mentioned prior art. The method efficiently activates and optimizes specific parts of the universal expert layer through mask selection and the lottery hypothesis principle, making the network more sparse, and realizing task-specific advanced feature extraction and computing resource optimization.
[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is: a multi-task learning resource optimization method based on lottery hypothesis, comprising the following steps:
[0008] Step 1: Obtain the original data set for each task and divide it into training set and test set to ensure that each task has an independent data set for testing and evaluating the neural network model;
[0009] Select the original data sets of task a and task b and divide them into the training set of task a and test set And the training set of task b and test set
[0010] The data in the original data set includes images, text, audio and video;
[0011] Step 2: Use the embedding layer to convert the input data of task a and task b into a low-dimensional continuous vector representation, that is, generate the embedding vectors of the input data of task a and task b; use the softmax function as a router, and calculate the routing probability data of task a and task b according to the generated embedding vectors; initialize the common mask matrix of task a and task b, and randomly initialize the universal expert matrix. The specific method is:
[0012] S1: Use the embedding layer to convert the input data of each task into a low-dimensional continuous vector representation; convert the input data X of task a into (a) and the input data X of task b (b) They are converted into low-dimensional continuous vector representations through the embedding layer. The specific conversion formula is:
[0013] E (a) =Embed(X (a) )
[0014] E (b) =Embed(X (b) )
[0015] Among them, E (a) For input data X (a) The embedding vector, E (b) For input data X (b) Embedding vector, Embed is the embedding layer;
[0016] S2: Design the router to initialize the routing probability vectors of task a and task b; use the softmax function as the router, according to the embedding vector E (a) and the embedding vector E (b) Initialize the routing probability vectors for task a and task b and in are the n expert coefficients of task a, They are the n expert coefficients of task b, and the specific calculation formula is:
[0017] α a =soft max x(E (a) )
[0018] α b =softmax(E (b) )
[0019] S3: Initialize the common mask matrices M1, M2, M3...M for tasks a and b n , each element of the mask matrix is initialized to 1; a universal expert matrix is randomly initialized for subsequent training, and the mask matrices M1, M2, M3...M n Select and activate a specific part of the parameters in the Universal Expert Matrix;
[0020] Step 3: randomly initialize a set of neural network model parameters E; jointly train the initialized n mask matrices using the training sets of task a and task b, and after the training of the training set is completed, use iterative amplitude pruning based on the lottery hypothesis principle to remove p% of the parameters after the absolute value of the neural network model parameters E, and then update the mask matrix; use the updated mask matrix to select and activate specific part parameters in the universal expert matrix, including the following steps:
[0021] Step 3.1: Use the training set of task a and task b and Jointly train the initialized n mask matrices; train the training sets of task a and task b and Divide into small batches of data, forward propagate a small batch of data of task a and a small batch of data of task b respectively, and obtain the loss function L of task a and task b a and L b , and then the loss function L of task a and task b is a and L b Add together to get the loss function L total , the specific formula is:
[0022] L total =L a +L b
[0023] Backpropagate the loss function L using the gradient descent algorithm total , that is, the loss function L total Minimize and update the neural network model parameters E;
[0024] Step 3.2: After training the entire original data set according to step 3.1, perform iterative amplitude pruning according to the principle of the lottery hypothesis, set the pruning ratio p, prune the parameters of the neural network model after the absolute value of p%, and convert the mask matrices M1, M2, M3...M n The corresponding mask in is set to 0; re-divide the small batch data and train the task a and task b sets and Retrain, perform iterative amplitude pruning after each training until the set sparsity is reached, stop training, and use mask matrices M1, M2, M3...M n Training completed;
[0025] Step 4: Use the training set of task a and the training set of task b for cross-training, update the parameters of the universal expert matrix after each round of training, until the parameters of the universal expert matrix converge, stop training, and obtain the final universal expert matrix, including the following steps:
[0026] Step 4.1: Use the training set of task a and task b and Perform cross-training; train task a and task b and Divide into multiple small batches of data, forward propagate a small batch of data of task a, and obtain the loss function L a ', and then back propagation is performed, using the gradient descent method to reduce the loss function L a 'Minimize, at this time the universal expert matrix parameters are also updated; then use a small batch of task b data for forward propagation to obtain the loss function L b ', use the gradient descent method to reduce the loss function L b 'Minimize, back propagate to minimize the loss function L b ', then update the universal expert matrix parameters;
[0027] Step 4.2: Then, cross-train the small batch data of task a and task b according to step 4.1. After each training, the parameters of the universal expert matrix will be updated until the parameters of the universal expert matrix converge. Then, the training is stopped to obtain the final universal expert matrix E. all , training completed;
[0028] Step 5: Use the test set of task a and the test set of task b The data in the initialization task a and task b probability vector α a and α b Multiply it with the trained sparse mask matrix and universal expert matrix, and finally generate the output result of the final task through the tower layer structure, which depends on the specific task;
[0029] Use the routing probability vector α of task a and task b a and α b As the input of the training neural network model, each element in the vector is respectively matched with the trained sparse mask matrix M'1, M'2, M'3...M' n And the all-round expert matrix E all Multiply, that is, only activate specific part of the parameters of the universal expert matrix, and then add the results, that is, only activate the sparse mask matrix M'1, M'2, M'3...M' n The universal expert matrix parameter corresponding to the position with a median value of 1 is used to obtain the output result Y of the entire neural network model. a and Y b :
[0030]
[0031] Finally, Y a and Yb Through the tower structure, the relevant features of task a and task b are learned.
[0032] The beneficial effects of adopting the above technical solution are: the multi-task learning resource optimization method based on lottery hypothesis provided by the present invention has the following advantages compared with the prior art:
[0033] (1) The sparse expert hybrid method based on the lottery hypothesis is used to effectively prune the parameters of the neural network model, thereby significantly reducing the computational cost and storage overhead of the neural network model; the sparse expert network structure retains a high degree of sensitivity to task-critical information, while improving the efficiency of the model while ensuring the accuracy of reasoning;
[0034] (2) The softmax-based router and mask matrix are introduced to achieve dynamic resource allocation, so that each task can select and activate the appropriate expert network part according to the characteristics of its input data; this mechanism enhances the flexibility of the multi-task learning model for information sharing between different tasks and effectively reduces the waste of computing resources and redundant calculations;
[0035] (3) By optimizing the universal expert matrix through cross-training and iterative amplitude pruning, the generalization ability of the model in handling complex tasks and different data distributions is further improved;
[0036] (4) The task-specific tower layer independently processes and optimizes the features of each task, reducing the mutual interference between tasks and improving the accuracy of feature extraction and the efficiency of task processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A flow chart of a multi-task learning resource optimization method based on lottery hypothesis provided by an embodiment of the present invention;
[0038] Figure 2 An architecture diagram of a multi-task learning resource optimization method based on lottery hypothesis provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0039] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0040] In this embodiment, a multi-task learning resource optimization method based on the lottery hypothesis is provided. Figure 1-2 As shown, the following steps are included:
[0041] Step 1: Obtain the original data set for each task and divide it into training set and test set to ensure that each task has an independent data set for testing and evaluating the neural network model;
[0042] Select the original data sets of task a and task b and divide them into the training set of task a and test set And the training set of task b and test set
[0043] The data in the original data set includes multiple forms such as images, texts, audios and videos;
[0044] In this embodiment, task a is an image detection task, task b is an object classification task, and the input data of task a and task b are both images;
[0045] Step 2: Use the embedding layer to convert the input data of task a and task b into a low-dimensional continuous vector representation, that is, generate the embedding vectors of the input data of task a and task b; use the softmax function as a router, and calculate the routing probability data of task a and task b according to the generated embedding vectors; initialize the common mask matrix of task a and task b (all set to 1), and randomly initialize the universal expert matrix. The specific method is:
[0046] S1: Use the embedding layer to convert the input data of each task into a low-dimensional continuous vector representation; convert the input data X of task a into (a) and the input data X of task b (b) They are converted into low-dimensional continuous vector representations through the embedding layer. The specific conversion formula is:
[0047] E (a) =Embed(X (a) )
[0048] E (b) =Embed(X (b) )
[0049] Among them, E (a) For input data X (a) The embedding vector, E (b) For input data X (b) Embedding vector, Embed is the embedding layer;
[0050] S2: Design the router to initialize the routing probability vectors of task a and task b; use the softmax function as the router, according to the embedding vector E (a) and the embedding vector E (b) Initialize the routing probability vectors for task a and task b and in are the n expert coefficients of task a, They are the n expert coefficients of task b, and the specific calculation formula is:
[0051] α a =softmax(E (a) )
[0052] α b =softmax(E (b) )
[0053] S3: Initialize the common mask matrices M1, M2, M3...M for tasks a and b n (As many mask matrices as there are elements in the routing probability vector), each element of the mask matrix is initialized to 1; a universal expert matrix is randomly initialized for subsequent training, and the mask matrices M1, M2, M3...M n Select and activate a specific part of the parameters in the Universal Expert Matrix;
[0054] Step 3: Randomly initialize a set of neural network model parameters E; use the training sets of task a and task b to jointly train the initialized n mask matrices. After the training of the training set is completed, iterative amplitude pruning is used based on the lottery hypothesis principle to remove p% of the parameters (unimportant parameters) after the absolute value of the neural network model parameters E, and then the mask matrix is updated; after A rounds of iterative amplitude pruning, after reaching the target sparsity, the mask matrix will gradually become sparse, retaining the most important parameter parts for task a and task b, including the following steps:
[0055] Step 3.1: Use the training set of task a and task b and Jointly train the initialized n mask matrices; train the training sets of task a and task b and Divide into multiple small batches of data, forward propagate a small batch of data of task a and a small batch of data of task b respectively, and obtain the loss function L of task a and task b a and L b , and then the loss function L of task a and task b is a and L b Add together to get the loss function L total , the specific formula is:
[0056] L total =L a +L b
[0057] Backpropagate the loss function L using the gradient descent algorithm total , that is, the loss function L total Minimize and update the neural network model parameters E;
[0058] In this embodiment, the cross entropy function is used to calculate the loss function;
[0059] Step 3.2: After training the entire original data set (epoch) according to step 3.1, perform iterative amplitude pruning according to the principle of the lottery hypothesis, set an appropriate pruning ratio p, prune the parameters of the neural network model after the absolute value of p%, and convert the mask matrices M1, M2, M3...M n The corresponding mask in is set to 0; re-divide the mini-batch data and train the task a and task b sets and Retrain, perform iterative amplitude pruning after each training until the set sparsity is reached, stop training, and use mask matrices M1, M2, M3...M n Training completed;
[0060] Step 4: Use the training set of task a and the training set of task b for cross-training, update the parameters of the universal expert matrix after each round of training, until the parameters of the universal expert matrix converge, stop training, and obtain the final universal expert matrix, including the following steps:
[0061] Step 4.1: Use the training set of task a and task b and Perform cross-training; train task a and task b and Divide into multiple small batches of data, forward propagate a small batch of data of task a, and obtain the loss function L a ', and then back propagation is performed, using the gradient descent method to reduce the loss function L a 'Minimize, at this time the universal expert matrix parameters are also updated; then use a small batch of task b data for forward propagation to obtain the loss function L b ', use the gradient descent method to reduce the loss function L b 'Minimize, back propagate to minimize the loss function L b ', then update the universal expert matrix parameters;
[0062] Step 4.2: Then, cross-train the small batch data of task a and task b according to step 4.1. After each training, the parameters of the universal expert matrix will be updated until the parameters of the universal expert matrix converge. Then, the training is stopped to obtain the final universal expert matrix E. all , training completed;
[0063] Step 5: Use the test set of task a and the test set of task b The data in the initialization task a and task b probability vector α a and α bMultiplying the trained sparse mask matrix and the universal expert matrix, that is, only activating certain parameters of the universal expert matrix, and finally generating the output result of the final task through the tower structure, which depends on the specific task;
[0064] Use the routing probability vector α of task a and task b a and α b As the input of the training neural network model, each element in the vector is respectively matched with the trained sparse mask matrix M'1, M'2, M'3...M' n And the all-round expert matrix E all Multiply and add the results, that is, only activate the sparse mask matrix M'1, M'2, M'3...M' n The universal expert matrix parameter corresponding to the position with a median value of 1 is used to obtain the output result Y of the entire neural network model. a and Y b :
[0065]
[0066] Finally, Y a and Y b Through the tower structure, the relevant features of task a and task b are learned;
[0067] The tower structure is a network structure specialized for each task, allowing each task to learn specific features related to itself to adapt to the needs of the specific task.
[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A multi-task learning resource optimization method based on lottery hypothesis, characterized by: The following steps are involved: Step 1: Obtain the original data set for each task and divide it into training set and test set to ensure that each task has an independent data set for testing and evaluating the neural network model; Step 2: Use the embedding layer to convert the input data of task a and task b into a low-dimensional continuous vector representation, that is, generate the embedding vectors of the input data of task a and task b; use the softmax function as a router to calculate the routing probability data of task a and task b based on the generated embedding vectors; initialize the common mask matrix of task a and task b, and randomly initialize the universal expert matrix; Step 3: Randomly initialize a set of neural network model parameters E; use the training sets of task a and task b to jointly train the initialized n mask matrices, and after the training of the training set is completed, use iterative amplitude pruning based on the lottery hypothesis principle to remove the parameters after p% of the absolute value of the neural network model parameters E, and then update the mask matrix; use the updated mask matrix to select and activate specific part parameters in the universal expert matrix; Step 4: Use the training set of task a and the training set of task b for cross-training, update the parameters of the universal expert matrix after each round of training, until the parameters of the universal expert matrix converge, stop training, and obtain the final universal expert matrix; Step 5: Use the test set of task a and the test set of task b The data in the initialization task a and task b probability vector α a and α b It is multiplied with the trained sparse mask matrix and the universal expert matrix, and finally the output result of the final task is generated through the tower layer structure, which depends on the specific task.
2. The multi-task learning resource optimization method based on lottery hypothesis according to claim 1 is characterized in that: The specific method of step 1 is: Select the original data sets of task a and task b and divide them into the training set of task a and test set And the training set of task b and test set The data in the original data set includes images, texts, audios and videos.
3. The multi-task learning resource optimization method based on lottery hypothesis according to claim 2 is characterized in that: Step 2: S1: Use the embedding layer to convert the input data of each task into a low-dimensional continuous vector representation; convert the input data X of task a into (a) and the input data X of task b (b) They are converted into low-dimensional continuous vector representations through the embedding layer; S2: Design the router to initialize the routing probability vectors of task a and task b; use the softmax function as the router, according to the embedding vector E (a) and the embedding vector E (b) Initialize the routing probability vectors for task a and task b and in are the n expert coefficients of task a, are the n expert coefficients of task b respectively; S3: Initialize the common mask matrices M1, M2, M3...M for tasks a and b n , each element of the mask matrix is initialized to 1; a universal expert matrix is randomly initialized for subsequent training, and the mask matrices M1, M2, M3...M n Select and activate a specific section of parameters in the All-round Expert Matrix.
4. The multi-task learning resource optimization method based on lottery hypothesis according to claim 3 is characterized in that: The input data X of task a (a) and the input data X of task b (b) The specific conversion formula is converted into low-dimensional continuous vector representation through the embedding layer: E (a) =Embed(X (a) ) E (b) =Embed(X (b) ) Among them, E (a) For input data X (a) The embedding vector, E (b) For input data X (b) The embedding vector of , Embed is the embedding layer.
5. The multi-task learning resource optimization method based on lottery hypothesis according to claim 4 is characterized in that: The softmax function is used as a router, according to the embedding vector E (a) and the embedding vector E (b) Initialize the routing probability vectors for task a and task b and The specific calculation formula is: a a =softmax(E (a) ) a b =softmax(E (b) )。 6. The multi-task learning resource optimization method based on lottery hypothesis according to claim 5 is characterized in that: Step 3 includes the following steps: Step 3.1: Use the training set of task a and task b and Jointly train the initialized n mask matrices; train the training sets of task a and task b and Divide into small batches of data, forward propagate a small batch of data of task a and a small batch of data of task b respectively, and obtain the loss function L of task a and task b a and L b , and then the loss function L of task a and task b is a and L b Add together to get the loss function L total , the specific formula is: L total =L a +L b Backpropagate the loss function L using the gradient descent algorithm total , that is, the loss function L total Minimize and update the neural network model parameters E; Step 3.2: After training the entire original data set according to step 3.1, perform iterative amplitude pruning according to the principle of the lottery hypothesis, set the pruning ratio p, prune the parameters of the neural network model after the absolute value of p%, and convert the mask matrices M1, M2, M3...M n The corresponding mask in is set to 0; re-divide the small batch data and train the task a and task b sets and Retrain, perform iterative amplitude pruning after each training until the set sparsity is reached, stop training, and use mask matrices M1, M2, M3...M n Training completed.
7. The multi-task learning resource optimization method based on lottery hypothesis according to claim 6 is characterized in that: Step 4 includes the following steps: Step 4.1: Use the training set of task a and task b and Perform cross-training; train task a and task b and Divide into multiple small batches of data, forward propagate a small batch of data of task a, and obtain the loss function L a ', and then back propagation is performed, using the gradient descent method to reduce the loss function L a 'Minimize, at this time the universal expert matrix parameters are also updated; then use a small batch of task b data for forward propagation to obtain the loss function L b ', use the gradient descent method to reduce the loss function L b 'Minimize, back propagate to minimize the loss function L b ', then update the universal expert matrix parameters; Step 4.2: Then, cross-train the small batch data of task a and task b according to step 4.
1. After each training, the parameters of the universal expert matrix will be updated until the parameters of the universal expert matrix converge. Then, the training is stopped to obtain the final universal expert matrix E. all , training completed.
8. The multi-task learning resource optimization method based on lottery hypothesis according to claim 7 is characterized in that: The specific method of step 5 is: Use the routing probability vector α of task a and task b a and α b As the input of the training neural network model, each element in the vector is respectively matched with the trained sparse mask matrix M'1, M'2, M'3...M' n And the all-round expert matrix E all Multiply, that is, only activate specific part of the parameters of the universal expert matrix, and then add the results, that is, only activate the sparse mask matrix M'1, M'2, M'3...M' n The universal expert matrix parameter corresponding to the position with a median value of 1 is used to obtain the output result Y of the entire neural network model. a and Y b : Finally, Y a and Y b Through the tower structure, the relevant features of task a and task b are learned.
Citation Information
Patent Citations
Description mining system and method based on multi-task sparse shared learning
CN113641819A
Personalized collaborative learning method and device based on neural network model pruning
CN114418085A
Building edge optimization method based on multi-task learning and dual lottery hypothesis
CN116052006A
Hybrid expert reinforcement learning method and system
WO2020155994A1