Parameter isolation continuous learning method based on random neural network and neuron mask
By introducing a neuron mask mechanism into the neural network and combining the parameter isolation method of random neural networks, the problem that neural networks are difficult to maintain the performance of old tasks in multi-task scenarios is solved, and the effect of neural networks completing more tasks without affecting the backbone network parameters is achieved, and the network reuse capability is improved.
Patent Information
- Application Number
- CN202311560810.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-23
AI Technical Summary
Disaster forgetting makes it difficult for neural networks to maintain the performance of old tasks in multi-task scenarios, resulting in manual fine-tuning of parameters or repeated training during training, which is time-consuming and labor-intensive, hindering the development of neural networks.
The parameter isolation continuous learning method based on random neural networks and neuron masks is adopted. Through the neuron mask mechanism, the neural network is no longer limited to a single classification task without affecting the parameters of the backbone network, and the one-to-many effect is achieved, effectively saving the mask storage space and improving the reuse capability of the backbone network.
It enables the neural network to complete more tasks without affecting the parameters of the backbone network, saves the mask storage space, improves the reuse capability of the backbone network, and solves the problem of catastrophic forgetting.
Smart Images

Figure BDA0004562350900000021 
Figure BDA0004562350900000031 
Figure BDA0004562350900000032
Abstract
Description
Technical Field
[0001] The present invention belongs to the problem of continuous learning in computer vision classification tasks and relates to a parameter isolation method. Background Art
[0002] Catastrophic forgetting is a common problem in neural networks, which is manifested in the devastating decline of the performance of neural networks on past tasks in multi-task scenarios as new tasks are trained. Catastrophic forgetting forces the use of manual fine-tuning of parameters or repeated training of past data sets to maintain the performance of neural networks on old tasks during the training of new tasks, which will consume a lot of time and manpower costs and hinder the development of neural networks. Solving catastrophic forgetting is the goal of continuous learning. The parameter isolation method allows different parts of the neural network to be responsible for different tasks, avoiding mutual interference between network parameters, and indirectly achieving the effect of continuous learning. How to enable neural networks to complete more tasks under limited network capacity is one of the main goals of the parameter isolation method. The neuron mask mechanism proposed in this patent allows a neural network to no longer be limited to a single classification task without affecting the parameters of the backbone network, achieving a one-to-many effect. In addition, the use of the mask mechanism at the neuron level effectively saves the storage space of the mask and improves the reuse capacity of the backbone network. Summary of the invention
[0003] The purpose of the present invention is to provide a parameter isolation continuous learning method based on random neural network and neuron mask to address the deficiencies of the prior art. This method adopts the neuron mask mechanism to make a neural network no longer limited to a single classification task without affecting the parameters of the backbone network, thus achieving a one-to-many effect. In addition, it effectively saves the storage space of the mask and improves the reuse capacity of the backbone network.
[0004] The technical solution for achieving the purpose of the present invention is:
[0005] The parameter isolation continuous learning method based on random neural network and neuron mask is different from the existing technology in that it includes the following steps:
[0006] 1) Import a computer vision classification dataset and divide it into multiple subsets;
[0007] 2) Randomly initialize a neural network, freeze all parameters, and obtain a random feature extractor module and a random classifier module;
[0008] 3) According to the output neuron dimension of each layer of the neural network, a neural network mask is randomly initialized to obtain a neuron mask module;
[0009] 4) Take a subset from the data set divided in step 1) to train the neural network, and update the parameters of the neuron mask through the back propagation algorithm until the task training is completed;
[0010] 5) Save the neuron mask of the current task, randomly initialize a new neural network mask, and train the neural network by taking another subset from the dataset without repetition;
[0011] 6) Repeat steps 4) and 5) until all subsets are trained.
[0012] The random feature extractor module described in step 2) is: the backbone network part of any neural network architecture in the field of computer vision; different from the common application method of the backbone network, the random feature extractor module freezes all the parameters of the backbone network after randomly initializing the backbone network, so the parameters of the backbone network are not iteratively updated during the neural network training process; assuming that the backbone network has L layers, the forward propagation process is as follows:
[0013] h L =σW L σ(W L-1 …σ(W 1 X)))
[0014] Where σ(·) is the ReLU activation function, W 1 is the weight of the first layer of the backbone network, W L is the weight of the Lth layer of the backbone network, and so on.
[0015] The random classifier module described in step 2) is: the classifier part of any neural network architecture in the computer vision classification task; different from the common application method of the classifier, the random classifier module freezes all the parameters of the classifier after randomly initializing the classifier, so the parameters of the classifier are not iteratively updated during the neural network training process; assuming that the classifier weight is θ, the forward propagation process is as follows:
[0016] y=θh L
[0017] where h L is the output of the backbone network, and y is the output of the classifier.
[0018] The neuron mask module in step 3) is a learnable parameter module for processing the output neuron mask of each layer; assuming that the neuron mask is M = {m L ,m L-1 ,...,m 1}, where m L The output neuron mask corresponding to the Lth layer of the backbone network, and so on; the use of neuron mask includes the following steps:
[0019] Inference stage: The operation mechanism of neuron mask is as follows:
[0020]
[0021] where h i is the output feature of the i-th layer of the backbone network, h′ i is the output feature after mask processing, Represents element-wise multiplication, m′ i is the neuron mask of the i-th layer of the backbone network after discretization, and its elements are either 0 or 1;
[0022] Training phase: The neuron mask is updated by back-propagation and then discretized:
[0023]
[0024]
[0025] Where η is the learning rate, is m i The gradient of i is the original mask, m i ′ is the mask actually used in the inference stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is an overall flow chart in the embodiment.
[0027] Figure 2 Schematic diagram of neuron mask in the embodiment. DETAILED DESCRIPTION
[0028] The content of the present invention is further described below in conjunction with the drawings and embodiments, but the present invention is not limited thereto.
[0029] Example:
[0030] Reference Figure 1 ,The parameter isolation continuous learning method based on random neural network and neuron mask, which is different from the existing technology, includes the following steps:
[0031] 1) Import a computer vision classification dataset and divide it into multiple subsets;
[0032] 2) Randomly initialize a neural network, freeze all parameters, and obtain a random feature extractor module and a random classifier module;
[0033] 3) According to the output neuron dimension of each layer of the neural network, a neural network mask is randomly initialized to obtain a neuron mask module;
[0034] 4) Take a subset from the data set divided in step 1) to train the neural network, and update the parameters of the neuron mask through the back propagation algorithm until the task training is completed;
[0035] 5) Save the neuron mask of the current task, randomly initialize a new neural network mask, and train the neural network by taking another subset from the dataset without repetition;
[0036] 6) Repeat steps 4) and 5) until all subsets are trained.
[0037] The random feature extractor module described in step 2) is: the backbone network part of any neural network architecture in the field of computer vision; different from the common application method of the backbone network, the random feature extractor module freezes all the parameters of the backbone network after randomly initializing the backbone network, so the parameters of the backbone network are not iteratively updated during the neural network training process; assuming that the backbone network has L layers, the forward propagation process is as follows:
[0038] h L =σ(W L σ(W L-1 …σ(W 1 X)))
[0039] Where σ(·) is the ReLU activation function, W 1 is the weight of the first layer of the backbone network, W L is the weight of the Lth layer of the backbone network, and so on.
[0040] The random classifier module described in step 2) is: the classifier part of any neural network architecture in the computer vision classification task; different from the common application method of the classifier, the random classifier module freezes all the parameters of the classifier after randomly initializing the classifier, so the parameters of the classifier are not iteratively updated during the neural network training process; assuming that the classifier weight is θ, the forward propagation process is as follows:
[0041] y=θh L
[0042] where h L is the output of the backbone network, and y is the output of the classifier.
[0043] Reference Figure 2 The neuron mask module in step 3) is: a learnable parameter module for processing the output neuron mask of each layer; assuming that the neuron mask is M = {m L ,m L-1 ,...,m 1}, where m L The output neuron mask corresponding to the Lth layer of the backbone network, and so on; the use of neuron mask includes the following steps:
[0044] Inference stage: The operation mechanism of neuron mask is as follows:
[0045]
[0046] where h i is the output feature of the i-th layer of the backbone network, h′ i is the output feature after mask processing, Represents element-wise multiplication, m′ i is the neuron mask of the i-th layer of the backbone network after discretization, and its elements are either 0 or 1;
[0047] Training phase: The neuron mask is updated by back-propagation and then discretized:
[0048]
[0049]
[0050] Where η is the learning rate, is m i The gradient of i is the original mask, m i ′ is the mask actually used in the inference stage.
Claims
1. Parameter isolation continuous learning method based on random neural network and neuron mask, It is characterized in that The steps include: 1-1) Import a computer vision classification dataset and divide it into multiple subsets; 1-2) Randomly initialize a neural network, freeze all parameters, and obtain a random feature extractor module and a random classifier module; 1-3) According to the output neuron dimension of each layer of the neural network, a neural network mask is randomly initialized to obtain a neuron mask module; 1-4) Take a subset from the data set divided in step 1-1) to train the neural network, and update the parameters of the neuron mask through the back propagation algorithm until the task training is completed; 1-5) Save the neuron mask of the current task, randomly initialize a new neural network mask, and train the neural network by taking another subset from the dataset without repetition; 1-6) Repeat steps 1-4) and 1-5) until all subsets are trained.
2. The parameter isolation continuous learning method based on random neural network and neuron mask according to claim 1, It is characterized in that The random feature extractor module described in step 1-2) is: the backbone network part of any neural network architecture in the field of computer vision; different from the common application method of the backbone network, the random feature extractor module freezes all the parameters of the backbone network after randomly initializing the backbone network, so the parameters of the backbone network are not iteratively updated during the neural network training process; assuming that the backbone network has L layers, the forward propagation process is as follows: h L =σ(W L σ(W L-1 …σ(W 1 X))) Where σ(·) is the ReLU activation function, W 1 is the weight of the first layer of the backbone network, W L is the weight of the Lth layer of the backbone network, and so on.
3. The parameter isolation continuous learning method based on random neural network and neuron mask according to claim 1, It is characterized in that The random classifier module described in step 1-2) is: the classifier part of any neural network architecture in the computer vision classification task; different from the common application method of the classifier, the random classifier module freezes all the parameters of the classifier after randomly initializing the classifier, so the parameters of the classifier are not iteratively updated during the neural network training process; assuming that the classifier weight is θ, the forward propagation process is as follows: and=θh L where h L is the output of the backbone network, and y is the output of the classifier.
4. The parameter isolation continuous learning method based on random neural network and neuron mask according to claim 1, It is characterized in that The neuron mask module described in step 1-3) is: a learnable parameter module for processing the output neuron mask of each layer; assuming that the neuron mask is M = {m L ,m L-1 ,...,m 1 }, where m L The output neuron mask corresponding to the Lth layer of the backbone network, and so on; the use of neuron mask includes the following steps: 4-1) In the inference stage, the operation mechanism of the neuron mask is as follows: where h i is the output feature of the i-th layer of the backbone network, h′ i is the output feature after mask processing, Represents element-wise multiplication, m′ i is the neuron mask of the i-th layer of the backbone network after discretization, and its elements are either 0 or 1; 4-2) During the training phase, the neuron mask is updated through back-propagation and then discretized: Where η is the learning rate, is m i The gradient of i is the original mask, m i ′ is the mask actually used in the inference stage.