Dense knowledge distillation method in continuous learning scene
By building a dense task pool, adaptive weighted tasks in continuous learning scenarios, and using random selection strategies for knowledge distillation, the high complexity and forgetting problems in traditional methods are solved, and the stability and adaptability of the model in multi-task learning are achieved.
Patent Information
- Application Number
- CN202510267815.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the continuous learning scenario, the traditional knowledge distillation method has increased exponentially due to the exponential increase in time complexity, which makes it complex and difficult to choose the optimal old task combination, and cannot effectively balance the learning of new tasks and the retention of old tasks knowledge, which can easily lead to catastrophic forgetting.
An intensive knowledge distillation method in continuous learning scenarios is adopted, including building a intensive task pool, adaptive weighted tasks, knowledge distillation through random selection strategies, and learning is guided by KL divergence.
Effectively prevent the model from forgetting old tasks, promote the generalization ability of new task learning, enhance the stability and adaptability of the model in multi-task learning, and significantly reduce the risk of catastrophic forgetting.
Smart Images

Figure CN120197679A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and more particularly to a dense knowledge distillation method in a continuous learning scenario. Background Art
[0002] Due to the catastrophic forgetting that easily occurs when traditional models are trained on a fixed dataset, that is, covering or losing the knowledge learned previously when learning new tasks, this phenomenon is particularly important in real-world applications such as autonomous driving, medical diagnosis, and intelligent assistants. In such scenarios, the model needs to be able to adapt to dynamically changing data and tasks. Therefore, researchers are constantly exploring effective methods to solve the problems of catastrophic forgetting and knowledge retention in order to achieve a more intelligent autonomous learning system.
[0003] Continuous learning is a machine learning method aimed at enabling a model to gradually learn new tasks in a changing environment while retaining the knowledge of previous tasks. Research directions include knowledge distillation, regularization techniques, memory enhancement, and the flexibility of the model architecture. In the prior art, by combining different old tasks in the task pool to guide the model, tracking the learning progress of the model on different old tasks, and trying to combine multiple old tasks, enumerating all possible combinations of old tasks to guide the current task learning, the time complexity increases exponentially, aiming to mitigate the catastrophic forgetting problem in continuous learning and improve the comprehensive performance and adaptability of the model.
[0004] However, how to select the most effective combination of old tasks remains a difficult problem. Traditional methods usually rely on enumerating all different combinations of tasks in the task pool for knowledge distillation. In the case of a large task pool scale, it will make the selection of the optimal combination complex and difficult. At the same time, as the task pool scale expands, the number of combinations increases rapidly, and the computational overhead rises significantly. Although attempts are made to reduce the computational complexity through random combinations, the knowledge of all tasks is often not fully utilized, resulting in weak distillation effects, failing to effectively balance the learning of new tasks and the retention of old task knowledge, and unable to effectively reduce the forgetting of the model, leading to the model being prone to catastrophic forgetting.
[0005] Therefore, a dense knowledge distillation method in a continuous learning scenario is provided to solve the above problems. Summary of the Invention
[0006] The object of the present invention is to provide a dense knowledge distillation method in a continuous learning scenario, which can effectively prevent the model from forgetting old tasks, while promoting the generalization ability of the new task learning process, enhancing the stability and adaptability of the model in multi-task learning, and thus further improving the overall performance and efficiency of the model in the face of multi-task learning.
[0007] To achieve the above object, the present invention provides a dense knowledge distillation method in a continuous learning scenario, including the following steps:
[0008] S1: Construct a dense task pool P t ;
[0009] S2: Adaptively weight the tasks in the dense task pool P t ;
[0010] S3: Randomly select a task group from the dense task pool P t for knowledge distillation by a random selection strategy.
[0011] Preferably, step S1 specifically includes the following steps:
[0012] S11: When a new task T t is added to the learning process, the dense task pool P t obtains the set of current tasks and learned tasks;
[0013] S12: Add the new task T t to the current tasks to expand the content of the dense task pool P t ;
[0014] S13: Construct the dense task pool P t as {T {0} , T {1} , T {2} , T {0,1} , T {0,2} , T {1,2} , T {0,1,2}}.
[0015] Preferably, step S2 specifically includes the following steps:
[0016] S21: Calculate the number of classes C t of each task, and determine the weighting value of the task according to the number of classes C t ;
[0017] S22: Determine the weighting coefficient λ t according to the similarity between the new task and the old tasks.
[0018] Preferably, in step S22, the weighting coefficient λ t is set as:
[0019] λ t = λ base * r(|p|, |C t |) * s(p, C t )
[0020] where p represents the dense task pool P tsubtask, |p| represents the number of old categories in task pool p, |C t | represents the number of categories of the t-th new task, and r(|p|, |C t |) represents the proportion, and s(p, C t ) represents the similarity.
[0021] Preferably, the proportion r(|p|, |C t |) and the similarity s(p, C t ) are respectively set as:
[0022]
[0023] Among them, represents the average feature vector of all categories in task pool p, represents C t the average feature vector of all new categories in.
[0024] Preferably, step S3 specifically includes the following steps:
[0025] S31: In each training step of each training cycle, randomly select a task group from the dense task pool P t for knowledge distillation;
[0026] S32: Compare the output logits of the randomly selected task group with the output logits of the current model, and guide the model to perform knowledge distillation through KL divergence .
[0027] Preferably, in step S32, the output logits of the current model are set as:
[0028] z t = f t (x t )
[0029] Among them, x t represents the input of the current model, z t represents the output of the current model, and f t represents the calculation process of the model.
[0030] Preferably, in step S32, the KL divergence is set as:
[0031]
[0032] Among them, represents the output logits of the current model at the t-th step under task pool p, represents the output logits of the current model at the (t - 1)-th step under task pool p.
[0033] Therefore, the present invention adopts the above-mentioned intensive knowledge distillation method in a continuous learning scenario, and has the following beneficial effects:
[0034] (1) By combining intensive distillation and adaptive weighting, the present invention combines with existing technologies such as knowledge distillation, replay, and weight regularization to improve the stability and generalization ability of multi-task learning;
[0035] (2) The present invention introduces a task group random selection strategy. By randomly selecting task groups for distillation in each optimization step, the computational overhead is greatly reduced, and the computational efficiency is improved while ensuring the distillation effect;
[0036] (3) The present invention densely groups the output logits of multiple historical tasks. Each task group represents a knowledge set of different historical tasks, ensuring that the model can fully share the knowledge between multiple tasks and improving the accumulation and transfer effect of knowledge;
[0037] (4) The present invention combines intensive task group distillation with a random selection strategy to ensure that the model effectively retains the knowledge of old tasks while learning new tasks, thereby significantly reducing the risk of catastrophic forgetting and effectively decoupling the relationship between old tasks and new tasks.
[0038] Next, through the accompanying drawings and embodiments, the method solution of the present invention will be further described in detail. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flowchart of an intensive knowledge distillation method in a continuous learning scenario of the present invention;
[0040] Figure 2 is a flowchart of constructing an intensive task pool of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] The method solution of the present invention will be further described below through the accompanying drawings and embodiments.
[0042] Unless otherwise defined, the method terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the art to which the present invention belongs.
[0043] In the present invention, words such as "including" or "comprising" mean that the elements before this word cover the elements listed after this word, and it does not exclude the possibility of also covering other elements. The orientation or positional relationship indicated by terms such as "inside", "outside", "above", "below", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly. In the present invention, unless otherwise clearly specified and defined, terms such as "attached" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be directly connected, or indirectly connected through an intermediate medium. It can be the internal communication of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0044] Embodiment
[0045] As Figure 1 and Figure 2 shown, the present invention provides a dense knowledge distillation method in a continuous learning scenario, including the following steps:
[0046] S1: Construct a dense task pool P t , the construction of the dense task pool P t is based on the record of the learned tasks, and each task represents the learning objective and learning content of the model at a certain stage;
[0047] Step S1 specifically includes the following steps:
[0048] S11: When a new task T t is added to the learning process, the dense task pool P t obtains the set of the current task and the learned tasks;
[0049] S12: Add the new task T t to the current task to expand the content of the dense task pool P t ;
[0050] S13: Construct the dense task pool P t as {T {0} , T {1} , T {2} , T {0,1} , T {0,2} , T {1,2} , T {0,1,2}}.
[0051] S2: For the dense task pool P tAdaptive weighting is performed on the tasks in
[0052] Step S2 specifically includes the following steps:
[0053] S21: Calculate the number of categories C of each task t , and determine the weighting value of the task according to the number of categories C t ;
[0054] S22: Determine the weighting coefficient λ according to the similarity between the new task and the old task t , so that similar tasks obtain a higher weight during distillation, thereby reducing the knowledge interference of the new task on the old task.
[0055] In step S22, the weighting coefficient λ t is set as:
[0056] λ t = λ base * r(|p|, |C t |) * s(p, C t )
[0057] where p represents a subtask of the dense task pool P t , |p| represents the number of old categories of the task pool p, |C t | represents the number of categories of the t-th new task, r(|p|, |C t |) represents the ratio, and s(p, C t ) represents the similarity.
[0058] The ratio r(|p|, |C t |) and the similarity s(p, C t ) are respectively set as:
[0059]
[0060] where represents the average feature vector of all categories in the task pool p, represents the average feature vector of all new categories in C t ;
[0061] S3: Randomly select a task group from the dense task pool P t for knowledge distillation through a random selection strategy, which effectively reduces the computational complexity of task combination.
[0062] Step S3 specifically includes the following steps:
[0063] S31: In each training step of each training cycle, randomly select a task group from the dense task pool P t for knowledge distillation, ensuring that within each training cycle, the dense task pool Pt All combinations of tasks in
[0064] are learned in different training steps instead of calculating all task groups each time, thus avoiding high computational overhead;
[0065] In step S32, the output logits of the randomly selected task group are compared with the output logits of the current model, and the knowledge distillation of the model is guided through the KL divergence
[0066] z t = f t (x t )
[0067] where x t represents the input of the current model, z t represents the output of the current model, and f t represents the calculation process of the model.
[0068] In step S32, the KL divergence is set as:
[0069]
[0070] where represents the output logits of the current model at the t-th step under the task pool p, represents the output logits of the current model at the (t - 1)-th step under the task pool p.
[0071] Comparative experiments are carried out based on the CIFAR100 dataset using different knowledge distillation methods, and the experimental results are shown in Table 1;
[0072] Table 1: Comparative experimental results
[0073] Knowledge distillation method Accuracy on CIFAR100 dataset Knowledge distillation 66.65% Random dense knowledge distillation 67.51%
[0074] As shown in Table 1, the accuracy rate of the ordinary knowledge distillation method is 66.65%, while the accuracy rate of the random dense knowledge distillation method adopted in the technical solution of this embodiment is 67.51%. Therefore, the random dense knowledge distillation method in this embodiment can improve the computational efficiency while ensuring the distillation effect.
[0075] Therefore, the present invention adopts the above-mentioned dense knowledge distillation method in a continuous learning scenario, effectively preventing the model from forgetting old tasks, promoting the generalization ability of the new task learning process at the same time, enhancing the stability and adaptability of the model in multi-task learning, and further improving the overall performance and efficiency of the model when facing multi-task learning.
[0076] Finally, it should be noted that the above embodiments are only used to illustrate the method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the method of the present invention, and these modifications or equivalent replacements cannot make the modified method deviate from the spirit and scope of the method of the present invention.
Claims
1. A dense knowledge distillation method in a continuous learning scenario, characterized in that: The following steps are involved: S1: Building a dense task pool P t ; S2: for dense task pool P t Adaptively weight the tasks in S3: Randomly select a strategy from the dense task pool P t A task group is randomly selected for knowledge distillation.
2. The dense knowledge distillation method in a continuous learning scenario according to claim 1, characterized in that: Step S1 specifically includes the following steps: S11: When the new task T t is added to the learning process, the dense task pool P t Get the current task and the set of learned tasks; S12: New task T t Add to the current task and expand the intensive task pool P t content; S13: Building a dense task pool P t For {T {0} , T {1} , T {2} , T {0,1} , T {0,2} , T {1,2} , T {0,1,2} }.
3. The dense knowledge distillation method in a continuous learning scenario according to claim 1, characterized in that: Step S2 specifically includes the following steps: S21: Calculate the number of categories C for each task t , according to the number of categories C t Determine the weighted value of the task; S22: Determine the weighting coefficient λ based on the similarity between the new task and the old task t .
4. The dense knowledge distillation method in a continuous learning scenario according to claim 3, characterized in that: In step S22, the weighting coefficient λ t Set to: l t =λ base *r(|p|,|C t |*s(p,C t ),p∈P t Among them, p represents the dense task pool P t subtask, |p| represents the number of old categories in task pool p, |C t | represents the number of categories of the tth new task, r(|p|,|C t |) represents the ratio, s(p,C t ) indicates similarity.
5. The dense knowledge distillation method in a continuous learning scenario according to claim 4, characterized in that: Ratio r(|p|,|C t |) and similarity s(p,C t ) are set to: in, represents the average feature vector of all categories in the task pool p, Represents C t The average feature vector of all new categories in .
6. The dense knowledge distillation method in a continuous learning scenario according to claim 1, characterized in that: Step S3 specifically includes the following steps: S31: In each training step of each training cycle, the dense task pool P t Randomly select a task group for knowledge distillation; S32: Compare the output logits of the randomly selected task group with the output logits of the current model, using KL divergence Guide the model to perform knowledge distillation.
7. The dense knowledge distillation method in a continuous learning scenario according to claim 6, characterized in that: In step S32, the output logits of the current model are set to: z t =f t (x t ) Among them, x t Represents the input of the current model, z t represents the output of the current model, f t Represents the calculation process of the model.
8. The dense knowledge distillation method in a continuous learning scenario according to claim 6, characterized in that: In step S32, KL divergence Set to: in, Indicates that the current model outputs logits at the tth step under task pool p, Indicates that the current model outputs logits at the t-1th step under task pool p.