Image classification continual learning method based on information bottleneck
Patent Information
- Application Number
- CN202410305768.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-18
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-03-18
AI Technical Summary
因此,维持这些不重要的权重或神经元将构建冗余子网络,导致网络容量的过度消耗和性能的下降
[0031]1)本发明通过引入变分参数,可以有效解决灾难性遗忘问题,并减少子网络的冗余,提高了模型的效率和泛化能力;
Smart Images

Figure CN118313438B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image classification technology, and more specifically, relates to a continuous learning method for image classification based on information bottlenecks. Background Technology
[0002] Currently, deep neural networks exhibit excellent performance in image classification. However, when faced with new tasks, they often forget previously acquired task-specific information, overwriting old task knowledge and leading to a "catastrophic decline" in performance on older tasks—a phenomenon known as catastrophic forgetting. To address this issue, continuous learning, also known as lifelong learning, has been proposed. This method retains previously acquired knowledge while acquiring new information, enabling the network to evolve towards human learning patterns. In recent years, many effective continuous learning methods have emerged. These methods can be conceptually categorized into regularization-based methods, memory-based methods, and architecture-based methods.
[0003] While existing methods have achieved significant success, catastrophic forgetting remains far from being resolved. In the realm of architecture-based methods, parameter isolation is arguably the most effective approach, involving compressing or pruning networks to construct a dedicated subnetwork for each task, mitigating interference with older tasks. These methods select weights or neurons for subnetwork construction based on roughly the same fundamental criterion: their size. However, other work has demonstrated that the size of weights or neurons is not necessarily related to their importance. Therefore, maintaining these less important weights or neurons will result in redundant subnetworks, leading to excessive network capacity consumption and performance degradation. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a continuous learning method for image classification based on information bottleneck. This method introduces variational parameters to solve catastrophic forgetting and optimizes the image classification model based on information bottleneck to reduce sub-network redundancy, thereby improving the performance of continuous learning for image classification.
[0005] To achieve the above-mentioned objectives, the image classification continuous learning method based on information bottlenecks of the present invention includes the following steps:
[0006] S1: Obtain a training sample set A for N tasks according to actual needs. n n = 1, 2, ..., N;
[0007] S2: Determine the image classification model according to actual needs and initialize its parameters, and initialize the fusion parameter mask matrix of each intermediate layer in the image classification model. Where i = 1, 2, ..., L, L represents the number of intermediate layers in the image classification model, and K i-1 K i Ki represents the number of neurons in the intermediate layers i-1 and i in the image classification model, and K0 represents the number of neurons in the input layer.
[0008] S3: Set the task number n = 1;
[0009] S4: Initialize the variational parameter matrix for each intermediate layer of the image classification model corresponding to task n. and i = 1, 2, ..., L, initialization formula is:
[0010] μ n,i =μ n-1,i ⊙M all,i +μ random ⊙(1-M all,i )
[0011] σ n,i =σ n-1,i ⊙M all,i +σ random ⊙(1-M all,i )
[0012] Where, μ n-1,i σ n-1,i Let μ represent the variational parameter matrix of intermediate layer i learned from task n-1, respectively. 0,i =0, σ 0,i In the normal distribution N(σ) init_mean ,σ init_var σ was obtained from sampling in ) init_mean σ init_var These represent the preset variational parameters σ. n,i The mean matrix and variance matrix of the initial values, μ random σ random Represents the generated random matrix;
[0013] S5: Use the training sample set A corresponding to task n. n The image classification model is trained, and the features h of each layer during the training process are... i The following calculations were performed using the reparameterized formula:
[0014] h i =f i (h i-1 ,(μ n,i +ε i ⊙σ n,i )⊙W i )
[0015] Among them, f i() represents the mapping function of the intermediate layer i. Let N(0,1) represent the matrix sampled from the standard normal distribution N(0,1). This represents the weight parameter matrix from intermediate layer i-1 to intermediate layer i in the image classification model;
[0016] Calculate the loss function for the training samples, and based on this, calculate the weight parameter W for each intermediate layer in the image classification model. i Variational parameter μ n,i and σ n,i The gradient is then applied to the weight parameter W. i Variational parameter μ n,i and σ n,i Update as follows; the loss function LOSS is calculated as follows:
[0017] The Lagrangian function Lag is calculated for each intermediate layer i using the following formula. i :
[0018] Lag i =γ i ×I(h i h i-1 )-I(h i ;y)
[0019] Where, γ i This indicates the compression ratio of the intermediate layer i;
[0020] Then, the loss function LOSS is calculated using the following formula:
[0021]
[0022] Where λ represents the preset weights, and loss represents the classification loss of the image classification model;
[0023] S6: The parameter mask matrix M corresponding to each intermediate layer in task n is calculated using the following formula. n,i :
[0024]
[0025] Among them, M n,i (p,q) represents the parameter mask matrix M. n,i The mask corresponding to the weight parameters from the p-th neuron in the i-1th intermediate layer to the q-th neuron in the i-th intermediate layer, where p = 1, 2, ..., K. i-1 q = 1, 2, ..., K i μ n,i (p,q), σ n,i (p, q) represent the current variational parameter matrix μ, respectively. n,i and σ n,i The corresponding element in;
[0026] S7: Determine if n < N. If yes, proceed to step S8; otherwise, proceed to step S9.
[0027] S8: Update the fusion parameter mask matrix M all,i =M n,i ||M all,i Let n = n + 1, and return to step S4;
[0028] S9: For each task n, based on the parameter mask matrix M of each of its intermediate layers... n,i The image classification model is pruned to obtain the information bottleneck mask subnetwork corresponding to task n, which is used for the actual execution of task n.
[0029] This invention presents a continuous learning method for image classification based on information bottlenecks. For each task, two variational parameter matrices are set for each intermediate layer of the image classification model. These variational parameter matrices are initialized based on the fusion parameter mask matrix of each intermediate layer. Then, the image classification model is trained using training sample sets from different tasks. During training, the features of each intermediate layer are calculated using a reparameterization formula. The overall loss function during training is calculated by combining the Lagrange function and the conventional classification loss. Subsequently, the parameter mask matrix of each intermediate layer is determined using the variational parameter matrices. Finally, the image classification model is pruned according to the parameter mask matrix of each intermediate layer corresponding to the task, resulting in the information bottleneck mask sub-network for that task.
[0030] The present invention has the following beneficial effects:
[0031] 1) By introducing variational parameters, this invention can effectively solve the catastrophic forgetting problem, reduce the redundancy of subnetworks, and improve the efficiency and generalization ability of the model.
[0032] 2) This invention can effectively enhance the model's adaptability to learning new tasks, enabling the model to retain its knowledge of previous tasks when processing continuous tasks;
[0033] 3) This invention also improves the interpretability of the model because the parameter mask matrix can intuitively show how the weights of each layer are selected and pruned, thereby constructing a non-redundant information bottleneck mask sub-network. Attached Figure Description
[0034] Figure 1 This is a flowchart illustrating a specific implementation of the image classification continuous learning method based on information bottlenecks according to the present invention.
[0035] Figure 2 This is an example diagram of continuous learning for image classification;
[0036] Figure 3This is a comparison chart of the number of parameters in the present invention and the comparative method in Example 1. Detailed Implementation
[0037] The specific embodiments of the present invention will now be described with reference to the accompanying drawings to enable those skilled in the art to better understand the invention. It should be particularly noted that in the following description, detailed descriptions of known functions and designs that might obscure the main content of the invention will be omitted here.
[0038] Figure 1 This is a flowchart illustrating a specific implementation of the image classification continuous learning method based on information bottlenecks according to the present invention. Figure 1 As shown, the specific steps of the image classification continuous learning method based on information bottleneck of the present invention include:
[0039] S101: Obtain the training sample set:
[0040] Obtain a training sample set A for N tasks based on actual needs. n , n=1,2,…,N.
[0041] S102: Initialize the image classification model:
[0042] Determine the image classification model based on actual needs and initialize its parameters. The parameter initialization of the image classification model can be randomized or reused from pre-trained parameters. Initialize the fusion parameter mask matrix for each intermediate layer in the image classification model. Where i = 1, 2, ..., L, L represents the number of intermediate layers in the image classification model, and K i-1 K i K represents the number of neurons in the intermediate layers i-1 and i in the image classification model, and K0 represents the number of neurons in the input layer.
[0043] S103: Set the task number n = 1.
[0044] S104: Initialize variational parameters:
[0045] To determine the parameter mask matrix for each intermediate layer in different tasks, this invention introduces two variational parameter matrices for each layer of the original image classification model, reparameterizing the network weights of the original image classification model. Therefore, it is necessary to first initialize the variational parameter matrix corresponding to each intermediate layer of the image classification model for task n. and i = 1, 2, ..., L.
[0046] To address the knowledge transfer problem that may result from the freezing of subnetworks for old tasks, this invention, before training each new task, integrates the variational parameters of previous tasks and the fusion parameter matrix to generate the variational parameter matrix μ.n,i and σ n,i The initialization of is calculated using the following formula:
[0047] μ n,i =μ n-1,i ⊙M all,i +μ random ⊙(1-M all,i )
[0048] σ n,i =σ n-1,i ⊙M all,i +σ random ⊙(1-M all,i )
[0049] Where, μ n-1,i σ n-1,i Let μ represent the variational parameter matrix of intermediate layer i learned from task n-1, respectively. 0,i =0, σ 0,i In the normal distribution N(σ) init_mean ,σ init_var σ was obtained from sampling in ) init_mean σ init_var These represent the preset variational parameters σ. n,i The mean matrix and variance matrix of the initial values, μ random σ random This represents the generated random matrix.
[0050] The above methods can effectively solve the catastrophic forgetting problem in the continuous learning process of image classification tasks. That is, when learning new image classification task knowledge, the previously learned image classification knowledge is retained. This is achieved by retaining the weights that are important to previous tasks and adjusting the network weights in real time when training new tasks, thus avoiding conflicts between new and old knowledge.
[0051] S105: Training the image classification model:
[0052] Using the training sample set A corresponding to task n n The image classification model is trained, and the features h of each layer during the training process are... i The following calculations were performed using the reparameterized formula:
[0053] h i =f i (h i-1 ,(μ n,i +ε i ⊙σ n,i )⊙W i )
[0054] Among them, f i () represents the mapping function of the intermediate layer i. Let N(0,1) represent the matrix sampled from the standard normal distribution N(0,1). This represents the weight parameter matrix from intermediate layer i-1 to intermediate layer i in the image classification model.
[0055] Calculate the loss function for the training samples, and based on this, calculate the weight parameter W for each intermediate layer in the image classification model. i Variational parameter μ n,i and σ n,i The gradient is then applied to the weight parameter W. i Variational parameter μ n,i and σ n,i Update.
[0056] The loss function is crucial for model training. To improve training performance, this invention calculates the loss function based on the information bottleneck principle. The information bottleneck is a concept in information theory that posits that deep learning networks essentially squeeze information out of a bottleneck, removing noisy input data containing irrelevant details and retaining only the features most relevant to the general concept. In this invention, the information bottleneck principle is utilized, employing the Lagrange function to minimize the mutual information I(h) between features from adjacent layers. i h i-1 Simultaneously maximize the mutual information I(h) between features and labels y. i ;y), that is, the Lagrangian function Lagrangian of each intermediate layer i is calculated using the following formula. i :
[0057] Lag i =γ i ×I(h i h i-1 )-I(h i ;y)
[0058] Where, γ i This represents the compression ratio of the intermediate layer i, i.e., retaining more feature information (γ). i (smaller) or compress more knowledge (γ) i The larger (the larger).
[0059] In practical applications, variational inference can be used to calculate the upper bound of the Lagrange function, thereby approximating the Lagrange function. The specific calculation process will not be elaborated here.
[0060] Therefore, the formula for calculating the loss function LOSS during image classification model training in this invention is as follows:
[0061]
[0062] Where λ represents the preset weights, and loss represents the classification loss of the image classification model. In this embodiment, the classification loss is the cross-entropy loss.
[0063] By using the above loss function, effective information can be focused on specific weights, while ineffective weights can be masked, thereby effectively extracting and compressing feature information to form optimized model parameters for each classification task.
[0064] In this embodiment, to reduce the impact on previous tasks, the weight parameter W is adjusted. i The gradient is constrained to freeze weights that are valuable for previous tasks. The formula for the constraint is as follows:
[0065]
[0066] Among them, ▽W i , These represent the weight parameters W before and after the restriction process, respectively. i The gradient.
[0067] Furthermore, in this embodiment, considering the changes in feature distribution during image classification model training, the compression ratio γ can be periodically and dynamically adjusted. i To achieve more accurate compression, thereby maintaining the flexibility and efficiency of the network structure in learning new tasks, the compression ratio γ i The specific adjustment method is as follows: the image classification model is updated E times per iteration, where E is set according to the actual situation. Then, singular value decomposition (SVD) is used to analyze the features h of the current intermediate layer i. i The singular values are obtained by decomposition. All singular values are sorted in descending order, and the top v values are selected. i There are singular values that satisfy v. i The ratio of the sum of a singular value to the sum of all singular values is greater than a preset threshold and v i Minimize, then let the compression ratio γ i =v i / k i .
[0068] S106: Calculate the parameter mask matrix:
[0069] The parameter mask matrix M corresponding to each intermediate layer in task n is calculated using the following formula. n,i :
[0070]
[0071] Among them, M n,i (p,q) represents the parameter mask matrix M. n,iThe mask corresponding to the weight parameters from the p-th neuron in the i-1th intermediate layer to the q-th neuron in the i-th intermediate layer, where p = 1, 2, ..., K. i-1 q = 1, 2, ..., K i μ n,i (p,q), σ n,i (p, q) represent the current variational parameter matrix μ, respectively. n,i and σ n,i The corresponding element in.
[0072] S107: Determine if n < N. If yes, proceed to step S108; otherwise, proceed to step S109.
[0073] S108: Update parameters:
[0074] Update the fusion parameter mask matrix M all,i =M n,i ||M all,i Let n = n + 1, and return to step S104.
[0075] S109: Determine the information bottleneck mask subnetwork:
[0076] For each task n, based on the parameter mask matrix M of each of its intermediate layers n,i The image classification model is pruned to obtain the information bottleneck mask subnetwork corresponding to task n, which is used for the actual execution of task n.
[0077] Figure 2 This is an example diagram of continuous learning for image classification. For example... Figure 2 As shown, in the continuous learning process of image classification in this invention, the weights that are important to previous tasks are retained by freezing parameters, and the network weights are adjusted when training new tasks, thereby solving the catastrophic forgetting problem in the continuous learning process.
[0078] To better illustrate the technical effects of the present invention, specific examples are used to experimentally verify the present invention.
[0079] Example 1
[0080] In this embodiment, the experimental conditions are set as follows: System: Ubuntu 20.04, Software: Python 3.6, Processor: Intel(R) Xeon(R) CPU E5-2678 v3@2.50GHz×2, Memory: 256GB.
[0081] Based on the image classification task, this paper compares existing parameter isolation methods with the method of this invention by counting the masked weights for each layer of the first subnet (for the first task). In this embodiment, the best current parameter isolation method, WSN (Winning SubNetworks), is used as the comparison method. With WSN, weight selection involves a uniform distribution based on the top k percentile weight ranking scores of all layers (set to 50% in this experiment). In contrast, the method of this invention allows for dynamic pruning of irrelevant weights and corresponding adjustment of the selection percentage.
[0082] Figure 3 This is a comparison chart of the number of parameters in the present invention and the comparative method in Example 1. From... Figure 3 As can be seen, due to the significant increase in the number of parameters, WSN retains a limited number of weights in the lower layers but exhibits a rich number of weights in the higher layers. Compared to WSN, this invention uses more weights in the lower layers but significantly reduces the number of weights required in the higher layers. This phenomenon can be attributed to the fact that this invention needs to capture sufficient visual information in the lower layers to allow it to propagate through the network and provide sufficient semantic information for the higher layers. Therefore, this invention achieves a 90% reduction in total weights and better performance, exceeding WSN by 0.6%. The results provide empirical support for the conjecture that addressing information bottlenecks can effectively reduce subnet redundancy.
[0083] Example 2
[0084] The experimental conditions set in this embodiment are as follows: System: Ubuntu 20.04, Software: Python 3.6, Processor: Intel(R) Xeon(R) CPU E5-2678 v3@2.50GHz×2, Memory: 256GB.
[0085] To more accurately evaluate the effectiveness of this invention in a continuous learning environment, this embodiment performs sequential task learning for image classification on the CIFAR-100 dataset. Using AlexNet as the original model, this invention is compared with 14 existing methods. Table 1 is a comparison table of the image classification performance of this invention and the compared methods in Example 2.
[0086]
[0087] Table 1
[0088] As shown in Table 1, this invention achieved a best average accuracy of 82.69%. Notably, since the weights learned from previous tasks have been frozen, this invention proves to be a forgetting-free model, resulting in 0% backward transfer. Furthermore, by re-initializing the variational parameters, this invention demonstrates the ability to transfer acquired knowledge to facilitate learning of new tasks, achieving a forward transfer rate of 3.52%, which is superior to previous parameter isolation methods.
[0089] Then, to further verify the behavior of the present invention on deeper networks, experiments were conducted on the original ResNet18 model. Table 2 is a comparison table of the image classification performance of the present invention and the comparison method in Example 2 on three different datasets.
[0090]
[0091]
[0092] Table 2
[0093] As shown in Table 2, the average accuracy of this invention consistently remains at the best level on CIFAR-100 (WSN), TinyImageNet (WSN), and MiniImageNet (der++), improving upon state-of-the-art methods by 1.68%, 3.02%, and 0.78%, respectively. Furthermore, on all datasets, this invention even outperforms Multi-task, which is considered a strong baseline for continuous learning. The possible reasons are multifaceted:
[0094] (1) Compact subnetworks perform better than over-parameterized dense neural networks. Further analysis of the results of this invention and Multi-task for each task reveals that this invention consistently outperforms Multi-task.
[0095] (2) Multi-task lacks knowledge transfer capability. Since it trains a new network for each task, there is no interaction between individual networks, which ultimately limits the transferability of knowledge.
[0096] In contrast, this invention utilizes the reinitialization of variational parameters to selectively retain valuable knowledge and achieve efficient knowledge transfer. This is evident in the FWT results, where the invention achieves significant improvements of 2.98%, 2.40%, and 1.22% on three different datasets, respectively.
[0097] Through the above data analysis and comparison, it can be seen that the present invention is superior to other parameter isolation methods. These results verify the effectiveness and superiority of the present invention.
[0098] Although the illustrative specific embodiments of the present invention have been described above to enable those skilled in the art to understand the invention, it should be understood that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of the present invention are protected.
Claims
1. A continuous learning method for image classification based on information bottleneck, characterized in that, Includes the following steps: S1: Obtain a training sample set A for N tasks according to actual needs. n n = 1, 2, ..., N; S2: Determine the image classification model according to actual needs and initialize its parameters, and initialize the fusion parameter mask matrix of each intermediate layer in the image classification model. Where i = 1, 2, ..., L, L represents the number of intermediate layers in the image classification model, and K i-1 K i Ki represents the number of neurons in the intermediate layers i-1 and i in the image classification model, and K0 represents the number of neurons in the input layer. S3: Set the task number n = 1; S4: Initialize the variational parameter matrix for each intermediate layer of the image classification model corresponding to task n. and The initialization formula is: m n,i =μ n-1,i ⊙M all,i +m random ⊙(1-M all,i ) s n,i =s n-1,i ⊙M all,i +s random ⊙(1-M all,i ) Where, μ n-1,i σ n-1,i Let μ represent the variational parameter matrix of intermediate layer i learned from task n-1, respectively. 0,i =0, σ 0,i In the normal distribution N(σ) init_mean ,σ init_var σ was obtained from sampling in ) init_mean σ init_var These represent the preset variational parameters σ. n,i The mean matrix and variance matrix of the initial values, μ random σ random Represents the generated random matrix; S5: Use the training sample set A corresponding to task n. n The image classification model is trained, and the features h of each layer during the training process are... i The following calculations were performed using the reparameterized formula: h i =f i (h i-1 ,(m n,i +e i ⊙s n,i )⊙W i ) Among them, f i () represents the mapping function of the intermediate layer i. Let N(0,1) represent the matrix sampled from the standard normal distribution N(0,1). This represents the weight parameter matrix from intermediate layer i-1 to intermediate layer i in the image classification model; Calculate the loss function for the training samples, and based on this, calculate the weight parameter W for each intermediate layer in the image classification model. i Variational parameter μ n,i and σ n,i The gradient is then applied to the weight parameter W. i Variational parameter μ n,i and σ n,i Update as follows; the loss function LOSS is calculated as follows: The Lagrangian function Lag is calculated for each intermediate layer i using the following formula. i : Lag i =γ i ×I(h i ;h i-1 )-I(h i ;y) Where, γ i This indicates the compression ratio of the intermediate layer i; Then, the loss function LOSS is calculated using the following formula: Where λ represents the preset weights, and loss represents the classification loss of the image classification model; S6: The parameter mask matrix M corresponding to each intermediate layer in task n is calculated using the following formula. n,i : Among them, M n,i (p,q) represents the parameter mask matrix M. n,i The mask corresponding to the weight parameters from the p-th neuron in the i-1th intermediate layer to the q-th neuron in the i-th intermediate layer, where p = 1, 2, ..., K. i-1 q = 1, 2, ..., K i μ n,i (p,q), σ n,i (p, q) represent the current variational parameter matrix μ, respectively. n,i and σ n,i The corresponding element in; S7: Determine if n < N. If yes, proceed to step S8; otherwise, proceed to step S9. S8: Update the fusion parameter mask matrix M all,i =M n,i ||M all,i Let n = n + 1, and return to step S4; S9: For each task n, based on the parameter mask matrix M of each of its intermediate layers... n,i The image classification model is pruned to obtain the information bottleneck mask subnetwork corresponding to task n, which is used for the actual execution of task n.
2. The image classification continuous learning method according to claim 1, characterized in that, In step S5, gradient constraint processing is applied during image classification model training. The specific formula is as follows: in, These represent the weight parameters W before and after the restriction process, respectively. i The gradient.
3. The image classification continuous learning method according to claim 1, characterized in that, In step S5, the compression ratio γ of the intermediate layer i during the image classification model training process i Periodic updates are performed, specifically as follows: the image classification model is updated E times per iteration, where E is set according to the actual situation. Then, singular value decomposition is used to update the features h of the current intermediate layer i. i The singular values are obtained by decomposition. All singular values are sorted in descending order, and the top v values are selected. i There are singular values that satisfy v. i The ratio of the sum of a singular value to the sum of all singular values is greater than a preset threshold and v i Minimize, then let the compression ratio γ i =v i / k i .
Citation Information
Patent Citations
Image classification method based on federal knowledge distillation and ensemble learning
CN117523291A
Party-specific environmental interface with artificial intelligence (AI)
US20200082293A1