Image target detection incremental learning method based on maximum cosine loss
By introducing huge cosine loss and knowledge distillation technology in incremental learning of image object detection, the problems of catastrophic forgetting and poor feature distribution constraints in incremental learning are solved, and the model's continuous memory of old knowledge and new tasks are realized, reducing storage and computing costs.
Patent Information
- Application Number
- CN202510194672.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-24
AI Technical Summary
The existing incremental learning methods have only limited effects in alleviating catastrophic forgetting, and there are problems such as unbalanced performance of new and old tasks, poor feature distribution constraints, and high storage and computing costs.
The incremental learning method of image object detection based on great cosine loss is adopted. The regression bounding box, projection angle and classification characteristics of the positive sample are distilled by introducing knowledge distillation technology, and pseudo-label generation and sample part playback techniques are used in the incremental stage to ensure that the learning of the new task does not lose old task knowledge.
Effectively slows down catastrophic forgetting, ensures the model's continuous memory of old knowledge and adapts to new tasks, reduces storage and computing costs, and improves the separability of feature distribution.
Smart Images

Figure CN120198713A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of object detection, and specifically to an incremental learning method for image object detection based on maximum cosine loss. Background Art
[0002] For incremental learning, the commonly used method is to fine-tune the model. However, due to the imbalance between new and old categories, it often leads to catastrophic forgetting of the model, that is, the model will quickly forget the knowledge learned before while learning new knowledge.
[0003] In current research work, the methods to alleviate catastrophic forgetting in class-incremental object detection can be roughly divided into three categories: methods based on parameter freezing and fine-tuning, methods based on knowledge distillation, and methods based on sample memory replay. The methods based on parameter freezing and fine-tuning maintain the performance of old tasks by freezing some model parameters, but this usually leads to insufficient adaptability of the model to new tasks. The methods based on knowledge distillation use the output or intermediate features of the teacher model with frozen parameters as labels to supervise the student model by guiding the training of the student model with the teacher model. However, since the distillation loss only constrains the model output, it cannot fundamentally optimize the feature representation. The methods based on sample memory replay store some old task data or their features for replay training. Although it can alleviate forgetting, its storage cost is high. In addition, due to data privacy issues, the original samples of old tasks are often difficult to obtain.
[0004] Although the above methods alleviate the catastrophic forgetting problem to a certain extent, they still have some deficiencies. For example, the control of the balance between new and old tasks is insufficient, which may lead to a seesaw effect between the performance of new and old tasks. There is a lack of refined constraints on the feature distribution, resulting in poor separability of features in the high-dimensional space. The storage cost and computational complexity are high, making it difficult to meet the requirements of real-time and resource-constrained scenarios. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides an incremental learning method for image object detection based on maximum cosine loss, which solves the problem that the existing incremental learning methods can only alleviate catastrophic forgetting to a certain extent but still have some defects.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] An incremental learning method for image object detection based on maximum cosine loss, the incremental learning method specifically includes the following steps:
[0008] S1. Obtain initial images as sample features, divide them into several training tasks, divide the training tasks into a training set and a test set, and divide the sample features in each training task into several categories;
[0009] S2. Construct an object detection model including a feature extraction module, a classification head, and a regression module, and set the loss function as the maximum cosine similarity loss function;
[0010] S3. Perform incremental learning on the trained object detection model according to several training tasks in the training set.
[0011] Preferably, in step S2, the expression of the maximum cosine similarity loss function is:
[0012]
[0013] In the above formula, L i is the loss of the i-th sample, n is the total number of divided categories, is the cosine value of the correct category corresponding to the i-th sample feature, cosθ ij (j≠y i ) is the cosine distance of the wrong category corresponding to the i-th sample feature, is the normalized cosine distance.
[0014] Preferably, in step S3, it specifically includes the following steps:
[0015] S31. In each incremental learning stage, use the sample replay algorithm to obtain some samples from the previous incremental learning stage to obtain old samples;
[0016] S32. Obtain each training task in turn, and use the sample features in this training task as new samples;
[0017] S33. Generate pseudo-labels for the new samples based on the teacher model;
[0018] S34. Input the old samples and the new samples with pseudo-labels into the teacher model and the student model, and perform knowledge distillation on the features of the teacher model and the student model;
[0019] S35. Through the backpropagation algorithm, update the parameters of the object detection model based on knowledge distillation;
[0020] S36. Repeat steps S31 - S35 until the training of all training tasks in the training set is completed, and test through the test set to obtain the final object detection model.
[0021] Preferably, in step S33, it specifically includes the following steps:
[0022] S331. Input the new samples into the trained teacher model to obtain a number of regression boxes and object categories; the regression boxes represent the positions and sizes of the detected target objects, and each category corresponds to a confidence value, indicating the possibility that the object belongs to a certain category.
[0023] S332. Select a number of regression boxes with the highest confidence values and obtain the final boxes through the NMS algorithm;
[0024] S333. Generate pseudo-labels according to the object categories corresponding to the final boxes.
[0025] Preferably, in step S34, it specifically includes the following steps:
[0026] S341. Input the old samples and the new samples with pseudo-labels into the teacher model and the student model;
[0027] S342. Perform knowledge distillation on the positive samples in the old samples generated by the teacher model. The positive samples include regression boxes and projection angles, and the regression boxes, projection angles, and category outputs of the teacher model are used as guidance for the student model;
[0028] S343. Perform knowledge distillation on the category outputs of the new samples on the old model.
[0029] Preferably, in steps S341 and S342, the knowledge distillation includes projection angle distillation, regression box distillation, feature distillation, and prediction distillation.
[0030] Preferably, the calculation formula for the projection angle distillation is:
[0031]
[0032] In the above formula, represents the projection angle distillation loss, and cosγ is calculated through the inner product of the feature vector and the classification hyperplane weight vector.
[0033] Preferably, the calculation formula for the regression box distillation is:
[0034]
[0035] In the above formula, represents the regression box distillation loss, N pos is the number of positive samples, is the square of the Euclidean distance, and b is the coordinate of the regression box.
[0036] Preferably, the calculation formula for the feature distillation is:
[0037]
[0038] In the above formula, represents the feature distillation loss, and are the normalized feature vectors extracted from the samples on the old model and the new model respectively, and <·,·> represents the inner product of two vectors.
[0039] Preferably, the calculation formula of the prediction distillation is:
[0040]
[0041] In the above formula, represents the prediction distillation loss, where represents the number of new samples in the current network, represents the output of the current network on the old classes, represents the output of the previous-stage teacher network on the old classes.
[0042] Compared with the prior art, the present invention provides an incremental learning method for image object detection based on maximum cosine loss, which has the following beneficial effects:
[0043] 1. By introducing the knowledge distillation technology, the present invention distills the regression bounding boxes, projection angles, and classification features of positive samples to ensure that the learning of new tasks does not lose the knowledge of old tasks, thereby effectively alleviating catastrophic forgetting.
[0044] 2. The present invention also distills the class outputs of new samples on the old model to ensure that the old-class outputs of new samples on the old and new models are consistent, thereby ensuring the continuous memory of the model for old knowledge and the adaptability to new tasks.
[0045] 3. In order to further alleviate catastrophic forgetting and retain old knowledge, the present invention also introduces the pseudo-label generation and sample partial replay technologies. In the incremental stage, since the old-class samples have no labels, the model will generate pseudo-labels for the old samples without labels to ensure that these old samples can be correctly guided during training, thereby avoiding the negative impact of unlabeled old samples on model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0047] Figure 1 is a schematic diagram of the architecture of the object detection model of the present invention;
[0048] Figure 2 is a schematic diagram of the classification head of the object detection model of the present invention;
[0049] Figure 3 It is a schematic diagram of the overall framework of the incremental learning model of the present invention;
[0050] Figure 4 It is a schematic diagram of feature distillation of the present invention. Specific embodiments
[0051] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Thereby, a full understanding of how the present application uses technical means to solve technical problems and achieve the realization process of technical effects can be obtained and implemented accordingly.
[0052] Those of ordinary skill in the art can understand that all or part of the steps in implementing the following embodiments can be completed by instructing relevant hardware through a program. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0053] In order to solve the problem that existing incremental learning methods can only alleviate catastrophic forgetting to a certain extent but still have some defects, the present invention provides an image object detection incremental learning method based on maximum cosine loss. As Figure 3 shown, it is a network model that can support object detection and incremental learning. The learning method specifically includes the following steps:
[0054] S1. Obtain initial images as sample features. The initial images come from PASCAL VOC and MS COCO, and are divided into several training tasks. The training tasks are divided into a training set and a test set, and the sample features in each training task are divided into several categories. Specifically, for a dataset containing M (M = 80) categories, the categories are equally divided into T (T = 8) tasks, where the number of categories in each task is N = M / (T), and there are a total of T tasks.
[0055] S2. Build an object detection model including a feature extraction module, a classification head, and a regression module, and set the loss function as the maximum cosine similarity loss function. Through step S2, the object detection model is initially trained. During the initial training process, set the optimizer and learning rate strategy, and optimize the object detection classification task through the maximum cosine similarity loss function (CosMax-Loss); the expression of the maximum cosine similarity loss function for the i-th sample is:
[0056]
[0057] In the above formula, L i is the loss of the i-th sample, n is the total number of divided categories, is the cosine value of the i-th correct category y corresponding to the i-th sample feature i of cosθ ij (j≠y i ) is the cosine distance of the j-th wrong category corresponding to the i-th sample feature, is the normalized cosine distance, n represents the number of all categories. The CosMax-Loss loss function is optimized by calculating the cosine distance between the sample feature and the correct category weight, ensuring the maximization of the cosine distance between the feature and the category, and the minimization of the cosine distance of the non-label category. When the CosMax-Loss loss function gradually converges, the cosine distance of the wrong category will tend to 0, that is, orthogonal between the feature and the non-label category weight.
[0058] S3. Perform incremental learning on the object detection model trained in step S2 according to several training tasks in the training set. The following further describes step S3 in detail. In step S3, it specifically includes the following steps:
[0059] S31. In each incremental learning stage, use the sample replay algorithm to obtain some samples of the previous incremental learning stage to obtain old samples;
[0060] S32. Obtain each training task in turn, and use the sample features in this training task as new samples;
[0061] S33. Generate pseudo-labels for the new samples based on the teacher model. In step S33, the generation of pseudo-labels specifically includes the following steps:
[0062] S331. Input the new samples into the trained teacher model to obtain several regression boxes and object categories. The regression box represents the position and size of the detected target object, and each category corresponds to a confidence value, indicating the possibility that the object belongs to a certain category.
[0063] S332. Select several regression boxes with the highest confidence values, and obtain the final boxes through the NMS algorithm. The NMS algorithm can effectively filter out those regression boxes with too high overlap, ensuring that the finally retained regression boxes have high accuracy.
[0064] S333. Generate pseudo-labels according to the object categories corresponding to the final boxes.
[0065] The above steps will traverse all data samples until pseudo-labels are generated for all new samples and are ready for subsequent training.
[0066] S34. Input the old samples and the new samples with pseudo - labels into the teacher model and the student model, and perform knowledge distillation on the features of the teacher model and the student model. To improve the effect of knowledge distillation, the following centralized knowledge distillation algorithms are designed to better improve the optimization effect of the object detection model. In step S34, it specifically includes the following steps:
[0067] S341. Input the old samples and the new samples with pseudo - labels into the teacher model and the student model to ensure that the knowledge of the old tasks is not lost during the learning of new tasks;
[0068] S342. Perform knowledge distillation on the positive samples in the old samples generated by the teacher model. The positive samples include the regression box and the projection angle. The regression box, projection angle, and class output of the teacher model are used as the guidance for the student model;
[0069] S343. Perform knowledge distillation on the class output of the new samples on the old model to ensure that the output of the new samples for the old classes is consistent on the old and new models. Knowledge distillation includes projection angle distillation, regression box distillation, feature distillation, and prediction distillation. In the incremental learning scenario, knowledge distillation is a very effective strategy, which can significantly alleviate the catastrophic forgetting problem by increasing the constraints of the new model on the previous knowledge. Therefore, in order to achieve seamless knowledge transfer between the old and new tasks, the present invention introduces feature distillation and prediction distillation between the old and new models.
[0070] At the feature level of the positive samples, the new model needs to learn the feature direction extracted by the old model to ensure that the positive samples in the feature space can correctly distinguish the foreground from the background. The projection angle is an important representation of the positive sample features, and its cosine value is used to measure the projection size of the positive sample on the classification hyperplane. Therefore, the present invention designs the projection angle distillation, and the calculation formula is:
[0071]
[0072] In the above formula, represents the projection angle distillation loss, N pos is the number of positive samples, cosγ is calculated by the inner product of the feature vector and the classification hyperplane weight vector, and respectively represent the projection angles of the positive sample features on the class hyperplane in the current stage and the previous stage, and the value range is between - 1 and 1. Projection angle distillation optimizes the classification accuracy of the new model in the feature space by constraining the projection angle consistency of the positive samples between the old and new models. The distillation mechanism of the projection angle aims to ensure that the projection angles of the positive samples of the old and new models are consistent.
[0073] To retain the knowledge of the old categories as much as possible, the present invention also introduces regression box distillation to perform more refined knowledge transfer for the positive samples in the old categories, mainly including the distillation of regression boxes and projection angles. This design can help the new model better adapt to the feature distribution of the old categories and avoid catastrophic forgetting of the old tasks.
[0074] For the positive samples in the object detection task, the predicted values of the regression boxes are important information for describing the position and size of the objects. In the incremental learning stage, the new model needs to learn the knowledge of the regression boxes of the old model for the positive samples through distillation, so as to ensure that the new model can maintain accurate localization ability for the old categories.
[0075] In the specific implementation, the regression box distillation loss is realized by calculating the difference in the regression parameters of the predicted boxes between the new and old models. The calculation formula of the regression box distillation is:
[0076]
[0077] In the above formula, represents the regression box distillation loss, N pos is the number of positive samples, is the square of the Euclidean distance, b is the coordinate of the regression box, and respectively represent the coordinate values of the regression boxes of the positive samples in the current stage and the previous stage. Its size is 4, including the coordinates (x1, y1) of the upper left corner of the bounding box and the coordinates (x2, y2) of the lower right corner. The regression box distillation calculates the difference in the predicted values of the regression boxes between the new and old models and optimizes it through the loss function to ensure the localization accuracy of the new model for the old tasks.
[0078] Although the prediction distillation can numerically constrain the outputs of the new and old models, if there is a lack of constraint on the feature direction, the new model may be completely opposite to the old model in terms of feature distribution. If only the prediction results are constrained, it may cause the Student model to have a feature distribution completely opposite to that of the Teacher model, that is, only numerically constraining will lose important feature direction information. As Figure 4 shown, the features can fit towards the class center from different directions of the class center, and some of these directions may very likely be wrong directions, completely opposite to the feature direction of the previous stage model. However, imposing strong constraints on the features from the feature level can make the features extracted by the previous stage model and the features extracted by the current stage model have similar directions. Thus, this situation is avoided.
[0079] To avoid this situation, the present invention designs the feature distillation loss Strongly constrain the feature distributions of the new and old models at the feature level. This method calculates the cosine similarity between the feature vectors extracted by the new and old models to ensure that the feature directions of the new model are the same as those of the old model.
[0080] The calculation formula for feature distillation is:
[0081]
[0082] In the above formula, represents the feature distillation loss, and are the normalized feature vectors extracted from the sample on the old model and the new model respectively. The old model and the new model represent the object detection models optimized in this stage and the previous stage respectively. <·,·> represents the inner product of two vectors. Feature distillation calculates the consistency between the new and old models in the feature space through cosine similarity to ensure that the feature directions extracted by the new model are similar to those of the old model, so as to effectively retain the feature representation of the old task. It strongly constrains the feature distributions of the new and old models at the feature level. This method calculates the cosine similarity between the feature vectors extracted by the new and old models to ensure that the feature directions of the new model are the same as those of the old model. The feature direction can reflect its relationship with the class to a certain extent. Therefore, by constraining the cosine angle to make the new model similar to the old model in terms of feature direction, the previous knowledge can be effectively retained.
[0083] Prediction distillation can make the outputs of new samples on the new and old models for old classes consistent, so as to maintain the model's memory of the old task. During the incremental learning process, new samples usually have lower probabilities for old classes. Although these probabilities are low, they contain rich information that can help the new model better adapt to the old task when learning new tasks. Therefore, the prediction distillation loss When training the second network and subsequent networks, calculate the mean square error between the outputs of new samples on the new network for old classes and their outputs at the same positions on the old network. The calculation formula is:
[0084]
[0085] In the above formula, represents the prediction distillation loss, where represents the number of new samples in the current network, represents the output of the current network for old classes, and represent the cosine values of the outputs of the model in the current stage for old classes and the cosine values of the outputs of the model in the previous stage for old classes respectively, and represents the outputs of the teacher network in the previous stage and the current stage on the old categories. Let \(t\) denote the training at the \(t\)-th stage. This loss function only takes effect for \(t > 1\) because there is no teacher network for the first network to perform distillation. With this constraint, the new model can better learn the information of the old categories, ensuring the consistency between the new and old models on the old tasks.
[0086] S35. Update the parameters of the object detection model through the backpropagation algorithm and based on knowledge distillation;
[0087] S36. Repeat steps S31 - S35 until the training of all training tasks in the training set is completed, and test through the test set to obtain the final object detection model.
[0088] The above embodiments have introduced the present invention in detail. Specific examples are used in this article to elaborate on the principle and implementation of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An incremental learning method for image target detection based on maximum cosine loss, characterized in that: The incremental learning method specifically includes the following steps: S1. Obtain an initial image as a sample feature, and divide it into several training tasks, and divide the training tasks into a training set and a test set, and divide the sample features in each training task into several categories; S2. Build an object detection model including a feature extraction module, a classification head and a regression module, and set the loss function to the maximum cosine similarity loss function; S3. Perform incremental learning on the trained target detection model according to several training tasks in the training set.
2. The incremental learning method according to claim 1, characterized in that: In step S2, the expression of the maximum cosine similarity loss function is: In the above formula, L i is the loss of the i-th sample, n is the total number of divided categories, is the cosine value of the correct category corresponding to the i-th sample feature, cosθ ij (j≠y i ) is the cosine distance of the error category corresponding to the i-th sample feature, is the normalized cosine distance.
3. The incremental learning method according to claim 1, characterized in that: In step S3, the following steps are specifically included: S31, in each incremental learning stage, using a sample replay algorithm to obtain part of the samples in the previous incremental learning stage to obtain old samples; S32, obtaining each training task in turn, and using the sample features in the training task as new samples; S33, generate pseudo labels for new samples based on the teacher model; S34, input the old samples and the new samples with pseudo labels into the teacher model and the student model, and perform knowledge distillation on the features of the teacher model and the student model; S35, updating the parameters of the target detection model through the back propagation algorithm and based on knowledge distillation; S36. Repeat steps S31 to S35 until all training tasks in the training set are completed, and then test the test set to obtain the final target detection model.
4. The incremental learning method according to claim 3, characterized in that: In step S33, the following steps are specifically included: S331. Input the new sample into the trained teacher model to obtain several regression boxes and object categories; the regression box represents the position and size of the detected target object, and each category corresponds to a confidence value, indicating the possibility that the object belongs to a certain category. S332, select several regression frames with the highest confidence values, and obtain the final frame through the NMS algorithm; S333. Generate a pseudo label according to the object category corresponding to the final frame.
5. The incremental learning method according to claim 3, characterized in that: In step S34, the following steps are specifically included: S341, inputting old samples and new samples with pseudo labels into the teacher model and the student model; S342, performing knowledge distillation on positive samples in the old samples generated by the teacher model, wherein the positive samples include a regression box and a projection angle, and the regression box, projection angle, and category output of the teacher model serve as guidance for the student model; S343. Perform knowledge distillation on the category output of the new sample on the old model.
6. The incremental learning method according to claim 5, characterized in that: In step S341 and step S342, the knowledge distillation includes projection angle distillation, regression box distillation, feature distillation and prediction distillation.
7. The incremental learning method according to claim 6, characterized in that: The calculation formula of the projection angle distillation is: In the above formula, represents the projection angle distillation loss, and cosγ is calculated by the inner product of the feature vector and the classification hyperplane weight vector.
8. The incremental learning method according to claim 6, characterized in that: The calculation formula of the regression frame distillation is: In the above formula, represents the regression box distillation loss, N pos is the number of positive samples, is the square of the Euclidean distance, and b is the coordinate of the regression box.
9. The incremental learning method according to claim 6, characterized in that: The calculation formula of the characteristic distillation is: In the above formula, represents the feature distillation loss, and are the normalized feature vectors extracted from the samples on the old model and the new model respectively, and <·,·> represents the inner product of the two vectors.
10. The incremental learning method according to claim 6, characterized in that: The calculation formula for predicting distillation is: In the above formula, represents the predicted distillation loss, where Represents the number of new samples in the current network, represents the output of the current network on the old category, Represents the output of the teacher network in the previous stage on the old categories.