End-to-end incremental target detection method through internal feature alignment and task decoupling

Through the end-to-end incremental object detection method of intrinsic feature alignment and task decoupling, the problem of inconsistent optimization within classification branches in the prior art is solved, and higher recognition accuracy and new and old categories detection performance are achieved.

CN120047715APending Publication Date: 2025-05-27HEBI POWER SUPPLY OF HENAN ELECTRIC POWERCORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411962188.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The distillation loss of the existing incremental object detection method in the classification branch is inconsistent with the classification loss of the new category, resulting in optimization problems, and the response characteristic deviation may introduce noise interference, affecting detection performance.

Method used

The end-to-end incremental object detection method of intrinsic feature alignment and task decoupling is adopted. By building a teacher network and a student network, the task alignment learning module, the task decoupling expansion strategy module and the knowledge distillation module are integrated to achieve feature alignment and knowledge transfer. The divided and contagious distillation strategy is used for classification and regression response.

Benefits of technology

Improve the network recognition accuracy, ensure that the classification and regression response of teacher detectors are consistent in the feature space, promote the learning of new categories, and do not forget old knowledge while taking into account the performance of new tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047715A_ABST
    Figure CN120047715A_ABST
Patent Text Reader

Abstract

The invention discloses an end-to-end incremental target detection method through internal feature alignment and task decoupling. The method comprises the following steps: fusing a task alignment learning module in a teacher network; a task decoupling extension strategy module, a task alignment learning module and a knowledge distillation module are fused in the student network; the task alignment learning module combines distribution focus loss, classification loss and boundary box regression loss as a final optimization target of the teacher network or the student network; when a new task is processed, a category information header output by a student network firstly passes through a task decoupling network expansion module to decouple new category information and old category information; for a new category, the student network directly outputs a position information head and a category information head with information of a new target; and for the old category, the student network distillates the category information header and the position information header of the teacher network through knowledge, so that the output position information header and the category information header acquire the information of the old target. According to the invention, the network identification precision can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of incremental object detection, and particularly relates to an end-to-end incremental object detection method through intrinsic feature alignment and task decoupling. Background Art

[0002] Conventional object detection methods are usually set in the closed-world assumption, trained on specific datasets, learn a fixed number of object categories, and are applied to specific scenarios. With the development of the information age, the speed of product replacement has accelerated, and traditional object detection algorithms are difficult to be flexibly applied in actual scenarios. Incremental Object Detection (IOD) solves this problem by incrementally training an object detector with instances from new classes while retaining the knowledge obtained from previously learned classes. In the past, the expansion of classification networks with incremental tasks mainly used the minimum unfolding method, which only modified the convolutional kernels of the last convolutional layer to increase the output channels while sharing the previous convolutional layers among various tasks. However, this method has an optimization problem that the distillation loss of the basic classes in the classification branch is inconsistent with the classification loss of the new classes.

[0003] The latest technology ERD selects the maximum component of the regression response distribution as the index for selecting the regression distillation region in the knowledge distillation strategy. However, it is found that the response feature deviation between the classification and regression branches may introduce noise interference in the subsequent stages of incremental learning, and there are performance limitations caused by the inconsistency between the training stage and the testing stage. Therefore, the method of ERD is an obvious compromise solution. In addition, for the distillation method of responses in incremental object detection, considering the distillation of localization knowledge and category knowledge mixed together, it is difficult to judge whether the mixed knowledge at each position is beneficial to the detection performance and which regions are beneficial to the transfer of a certain type of knowledge. Summary of the Invention

[0004] The present invention discloses an end-to-end incremental object detection method through intrinsic feature alignment and task decoupling, which can effectively improve the recognition accuracy of the network.

[0005] The present invention discloses an end-to-end incremental object detection method through intrinsic feature alignment and task decoupling, and the method comprises the following steps:

[0006] S1, constructing an old-task training dataset and a new-task training dataset for incremental learning of object detection, where the data in the old-task training dataset and the new-task training dataset have no intersection; the new-task training dataset includes old classes and new classes, wherein the training data related to the new classes have labels, and the training data related to the old classes are not set with labels;

[0007] S2. Construct a teacher network, integrate a task alignment learning module into the teacher network, and train the old task training dataset to obtain a teacher model;

[0008] S3. Construct a student network, integrate a task decoupling and extension strategy module, a task alignment learning module, and a knowledge distillation module into the student network. The new task training dataset outputs a location information head and a category information head through the student network, which are used for location prediction and category prediction respectively, and the location information head output by the student network only involves the generation of the bounding box; the knowledge distillation module transfers the old knowledge learned by the teacher model to the student model, enabling the student model to have the ability to recognize both new and old targets;

[0009] The task alignment learning module introduces a distribution focal loss on the predicted distribution and the ground truth label, and combines the distribution focal loss, the classification loss, and the bounding box regression loss as the final optimization objective of the teacher network or the student network;

[0010] S4. When processing a new task, the category information head output by the student network first passes through the task decoupling network expansion module to decouple the new category and old category information; for the new category, the student network directly outputs the location information head and the category information head, and the location information head and the category information head have the information of the new target; for the old category, the student network distills the category information head and the location information head of the teacher network, so that the output location information head and category information head obtain the information of the old target.

[0011] As a preferred example, the method further includes the following steps:

[0012] A spatio-temporal interaction module based on matrix operations uses matrix multiplication to process the feature maps of the teacher network and the student network; the spatio-temporal interaction module includes three parts: an input dimension conversion unit, a spatio-temporal interaction unit, and an excitation weighting unit;

[0013] Among them, the input dimension conversion unit is used to change the feature map with an input dimension of [C×T×H×W] into a feature map with a dimension of [CT×HW] through matrix conversion operations, as the first operand of the matrix multiplication; the spatio-temporal interaction unit performs a series of dimension rearrangement reshape operations and dimension-invariant 1×1 convolution operations on the feature map with an input dimension of [C×T×H×W], making it into a feature map with a dimension of [HW×1×1], as the second operand of the matrix multiplication; the excitation weighting unit performs matrix multiplication on the first operand and the second operand, and then weights the channel weights to the original features channel by channel through multiplication to perform the original feature recalibration in the channel dimension.

[0014] As a preferred example, the task decoupling and extension strategy module includes multiple decoupled classification heads;

[0015] When the features extracted by the front-end module are input, the decoupled classification head generates classification responses for the features of each incremental task. The classification responses of the features of each incremental task perform a convolution operation on the previous incremental task, and finally these responses are concatenated to make it consistent with the output in the minimum expansion method. Each incremental task maintains a separate classification network;

[0016] Current new category task conv t The final output response of the classification head in is expressed as:

[0017] C stu = concat[conv 1 (feats), conv 2 (feats)... conv t (feats)]

[0018] In the formula, conv i represents the convolutional layer of the object category classification head in task i; concat represents the operation of concatenating multiple tensors conv 1 , conv 2 , conv i along a specific dimension to form a larger tensor; feats represents the network feature information.

[0019] As a preferred example, the knowledge distillation module takes the entire classification response as the classification distillation area; at the same time, an adaptive statistical threshold is constructed based on the classification confidence to screen the regression distillation area, and after correction and screening, the high-certainty response area on the regression branch is determined for regression distillation to adapt to the different characteristics of the classification and regression responses.

[0020] As a preferred example, the task alignment learning modules of the student network and the teacher network operate independently;

[0021] The task alignment learning modules of the student network and the teacher network respectively perform a high-order combination of the intersection IoU between the predicted class information and location information of the new task training data set output by the student network, the predicted class information and location information of the old task training data set output by the teacher network, and the class and location of the real data to quantify the task sample alignment. After the task sample alignment, the class loss, the generalized intersection over union GIoU loss, and the distribution focal loss Distribution Focal Loss are calculated.

[0022] As a preferred example, the overall loss function L total of the teacher network and the student network is:

[0023] L total = αLmodel_cls +L model_reg +β Ldist_cls +L dist_reg

[0024] Among them, the loss term L model is the classification and localization loss for the student network, used to calculate the loss of the student network for detecting new targets; L dist represents the knowledge distillation loss; the parameter α is a hyperparameter used to balance the detector loss weight, and the parameter β is a hyperparameter used to balance the hierarchical distillation loss weight. The hyperparameters α and β are used to improve the robustness of the detection performance of new and old categories; the subscripts reg and cls correspond to the regression and classification regions respectively.

[0025] The beneficial effects of the present invention are as follows:

[0026] First, the end-to-end incremental object detection method of the present invention through intrinsic feature alignment and task decoupling integrates a task alignment learning mechanism to align features, so that the classification and regression responses of the teacher detector are consistent in the feature space. This consistency can not only provide a more robust teacher model for the subsequent incremental stage to promote effective knowledge distillation, but also facilitate the learning of new categories.

[0027] Second, the end-to-end incremental object detection method of the present invention through intrinsic feature alignment and task decoupling adopts a divide-and-conquer distillation strategy for classification knowledge and localization knowledge in the response. Different from the feature-based method, the response-based method can provide the inference information of the teacher detector. By extracting incremental knowledge from the responses of classification knowledge and localization knowledge, a powerful and efficient student object detector can be gradually learned.

[0028] Third, for the overall architecture of the network, the end-to-end incremental object detection method of the present invention through intrinsic feature alignment and task decoupling designs a loss function. While taking into account the performance of the student network in identifying new tasks, the function also ensures that old knowledge is not forgotten. At the same time, for the divide-and-conquer distillation strategy, the target loss function is redesigned, the network model is optimized, and the detection performance is improved.

[0029] Fourth, in order to better improve the incremental learning task of the network, the end-to-end incremental object detection method of the present invention optimizes the object detection and recognition algorithm. By adopting the integrated task alignment learning module and task decoupling extension strategy module, while ensuring the inference speed of the network, a dedicated knowledge distillation module is synchronously customized and designed to focus on key regions, which can effectively improve the recognition accuracy of the network. Description of the Drawings

[0030] Figure 1It is the principle framework diagram of the end-to-end incremental object detection method through intrinsic feature alignment and task decoupling of the present invention;

[0031] Figure 2 It is the schematic diagram of the response-based knowledge distillation module;

[0032] Figure 3 It is the structural schematic diagram of the Spatial-Temporal (ST) module;

[0033] Figure 4 It is the schematic diagram of the task decoupling extension strategy module;

[0034] Figure 5 It is the schematic diagram of the customized knowledge distillation module;

[0035] Figure 6 It is the schematic diagram of the Task Alignment Learning (TAL) module;

[0036] Figure 7 It is the flow chart of the end-to-end incremental object detection method through intrinsic feature alignment and task decoupling of the present invention. Specific embodiments

[0037] The following embodiments can enable those skilled in the art to understand the present invention more comprehensively, but do not limit the present invention in any way.

[0038] See Figure 7 , the present invention discloses an end-to-end incremental object detection method through intrinsic feature alignment and task decoupling, and the method includes the following steps:

[0039] S1. Construct an old task training data set and a new task training data set for incremental learning of object detection. There is no intersection between the data in the old task training data set and the new task training data set; the new task training data set includes old categories and new categories. Among them, the training data related to the new categories has labels, and the training data related to the old categories is not labeled;

[0040] S2. Construct a teacher network, integrate a task alignment learning module in the teacher network, and train the old task training data set to obtain a teacher model;

[0041] S3. Construct a student network, integrate a task decoupling extension strategy module, a task alignment learning module, and a knowledge distillation module in the student network. The new task training data set outputs a location information head and a category information head through the student network, which are used for location prediction and category prediction respectively. And the location information head output by the student network only involves the generation of the bounding box; the knowledge distillation module transfers the old knowledge learned by the teacher model to the student model, enabling the student model to have the ability to identify new and old targets at the same time;

[0042] The task alignment learning module introduces a distribution focal loss on the predicted distribution and the ground truth labels, and combines the distribution focal loss, classification loss, and bounding box regression loss as the final optimization objective for the teacher network or the student network;

[0043] S4. When processing a new task, the class information head output by the student network first passes through the task decoupling network expansion module to decouple the new class and old class information; for the new class, the student network directly outputs the location information head and the class information head, and the location information head and the class information head have the information of the new target; for the old class, the student network obtains the class information head and the location information head of the old target by distilling the knowledge of the teacher network, so that the output location information head and class information head obtain the information of the old target.

[0044] As Figure 1 shown, the present invention proposes an end-to-end incremental object detection method through intrinsic feature alignment and task decoupling, which is mainly applied to the task of incremental object detection. Traditional object detection algorithms are often limited to training on specific datasets, learning a fixed number of classes for specific scenarios. For example, if only the person class is trained on an object detection model, when there are new classes to be detected, traditional object detection needs to collect pictures of relevant classes and label a large amount of data, and then retrain the labeled data of the two classes. This greatly limits the development and application of object detection technology. Moreover, in real scenarios, videos and pictures contain a large number of unknown object classes, and the types of objects are numerous, far exceeding the number of types defined by humans. In this case, it is difficult to keep up by continuously labeling data and training models. Incremental object detection solves this problem. The end-to-end incremental object detection method of the present invention designates the previously learned detector as the teacher detector (teacher network), denoted as tea, and the detector that performs incremental learning on the new task is called the student detector (student network), denoted as stu. During the incremental learning process of the new class, the student detector uses the task decoupling expansion module we proposed to expand its network structure to adapt to the base class and the new class. This decoupling method can alleviate the performance constraints caused by balancing different optimization objectives in the classification branch. Subsequently, the teacher detector and the student detector perform inference on randomly sampled samples from the new incremental task dataset. The predicted responses of the classification head and the regression head are denoted as Cla and Reg respectively. Then our learning alignment strategy is applied to the predicted responses of the two detectors. In addition, the integrated task alignment learning mechanism ensures effective learning of new classes from GT (boxes generated by annotations) annotations by aligning the inconsistent training and testing phases in the dense detector. These alignment operations improve the incremental learning efficiency of the detector we constructed in the learning alignment strategy.

[0045] (I) Knowledge distillation module

[0046] AsFigure 2 The following is the schematic diagram of the response-based knowledge distillation module. Response-based knowledge generally refers to the output of the last layer of the teacher model. The idea comes from directly mimicking the final output of the teacher model. This type of knowledge is intuitive and effective and has been widely applied in different tasks and application scenarios. Its formal definition is as follows:

[0047] L ResD (p(z t ,T),p(z s ,T))=L R (p(z t ,T),p(z s ,T))

[0048] where z t and z s are the Logits of the teacher model and the student model respectively.

[0049] The most well-known response-based knowledge in the field of image classification is the soft target. The soft target is the probability that the input belongs to a certain class, and its calculation method is as follows:

[0050]

[0051] Hinton believes that the soft target contains the implicit knowledge of the teacher model, and the corresponding distillation loss can be expressed as:

[0052] L ResD (p(z t ,T),p(z s ,T))=L R (p(z t ,T),p(z s ,T)).

[0053] As Figure 5 shown in the figure, the following is the schematic diagram of the principle of the knowledge distillation module customized for incremental object detection in the present invention. First, the selection of the distillation region is discussed in this embodiment. Considering optimizing the entire response of the classification branch during the training process, the entire classification response is selected as the classification distillation region, denoted as R cls ; on the contrary, since only a small part of the positive sample regions in the regression branch are optimized, a series of filtering operations are designed in this embodiment to select high-certainty regions for regression distillation, denoted as R regThis method enables the knowledge distillation module of this embodiment to effectively adapt to the different characteristics of classification and regression responses. Different from non-incremental detectors, incremental detectors lack GT markers of basic classes as references for selecting highly deterministic regression distillation regions. To solve this problem, it is assumed that classification and regression responses show good consistency in the feature space. Subsequently, an adaptive statistical threshold is constructed based on classification confidence to screen the regression distillation region. After correction and further screening, a highly deterministic response region on the regression branch is finally determined.

[0054] In the appendix Figure 5 For class information, the class output information of the teacher network and the class information of the student network are input into the knowledge distillation module. The class information of the student network learns knowledge by distilling the class output information of the teacher network. The classification knowledge distillation loss is expressed as:

[0055]

[0056] Here, P tea is the confidence of the classification response in the teacher detector, P stu is the confidence of the classification response in the student detector, M represents the number of predicted samples in the classification region cls, K represents the number of all detectable classes in the teacher detector, S tea represents applying the Sigmoid function transformation to the last layer features of the teacher detector, S stu represents applying the Sigmoid function transformation to the last layer features of the student detector, represents taking the minimum value.

[0057] For regression information, after correction and further screening, a highly deterministic response region on the regression branch is finally determined. For the class information of the teacher network, μ and β are first obtained through confidence, θ is obtained by weighted summation of μ and β, and then the preliminary candidate regions are screened by comparing P tea with θ, and finally the regression distillation region is filtered out through NMS.

[0058] After inputting the regression output information of the teacher network and the regression information of the student network into the knowledge distillation module, the regression information of the student network learns knowledge by distilling the regression output information of the teacher network. The regression knowledge distillation loss is expressed as:

[0059]

[0060] In the formula, N represents the number of regression regions reg of predicted samples, B i represents the predicted bounding box corresponding to the i th sample in the regression response.

[0061] (2) Spatial-Temporal Interaction Module for Matrix Operations In the present invention, a spatial-temporal interaction module for matrix operations is inserted between two feature layers of the backbone networks of the student network and the teacher network. Its function is to multiply the original feature channel weights channel by channel and weight them to the original features, completing the recalibration of the original features in the channel dimension to enhance effective information and suppress useless information.

[0062] As Figure 3 shown, in this embodiment, based on the spatial-temporal interaction module for matrix operations, matrix multiplication is used to process the feature maps of the teacher network and the student network. The spatial-temporal interaction module for matrix operations of the present invention is a plug-and-play attention module. Feature maps in neural networks are essentially matrix data, and matrix multiplication can be used to process feature maps. This patent incorporates a spatial-temporal interaction module, which can effectively improve the recognition accuracy of the network for key regions. The spatial-temporal interaction module is mainly divided into three parts: an input dimension conversion unit, a spatial-temporal interaction unit, and an excitation weighting unit. In Figure 3 , the input dimension is [C×T×H×W], where C represents the number of channels, T represents the value of the image sequence, H represents the height, and W represents the width. In the input dimension conversion unit, the module input [C×T×H×W] is transformed into [CT×HW] through a simple matrix conversion operation, thus obtaining the first operand of matrix multiplication. The input of the spatial-temporal interaction unit is the same as that of the input dimension conversion unit. The purpose of the spatial-temporal interaction unit is to obtain the second operand of matrix multiplication. Specifically, the spatial-temporal interaction unit performs a series of dimension rearrangement reshape operations and 1×1 convolution operations with unchanged dimensions on the feature map with an input dimension of [C×T×H×W], making it a feature map with a dimension of [HW×1×1] as the second operand of matrix multiplication. Considering that if the input is an RGB sequence, then extracting the correlation information between the input features during the dimension transformation process will improve the overall result. Therefore, a Reshape-Conv composite operation is used to achieve this goal. The excitation weighting unit multiplies the outputs of the above two parts to obtain the input of the third part. The third part uses an excitation weighting operation, whose function is to multiply the channel weights channel by channel and weight them to the original features, completing the recalibration of the original features in the channel dimension.

[0063] (3) Task Decoupling and Expansion Strategy Module

[0064] When facing new incremental tasks, existing methods usually adopt the minimum expansion mode to expand the classification head. For incremental tasks, this only involves expanding the last convolutional layer of the classification head, resulting in significant performance limitations. This embodiment proposes a task decoupling expansion strategy module to overcome this limitation. In this embodiment, each incremental task maintains a separate classification network. When the features extracted by the input front-end module (such as FPN) are input, the decoupled classification head generates classification responses for each incremental task. Then these responses are concatenated to make the output consistent with that in the minimum expansion method. This method alleviates the trade-off of the shared convolutional layer in balancing fundamentally different optimization objectives in the classification branches.

[0065] When processing new tasks, according to the output category information, the student network first passes through the task decoupling network expansion module to decouple the new category and old category information; for the new category, the student network detects the new category according to the new category information header; for the old category, the student network transfers knowledge to the old category information header of the student network by distilling the category information of the teacher network according to the decoupled old category information header.

[0066] As Figure 4 shown, the task decoupling expansion strategy module includes multiple decoupled classification heads; when the features extracted by the front-end module are input, the decoupled classification heads generate classification responses for the features of each incremental task, and the classification responses of the features of each incremental task perform a convolutional operation on the previous incremental task, and finally these responses are concatenated to make the output consistent with that in the minimum expansion method, and each incremental task maintains a separate classification network;

[0067] The final output response of the classification head in the current new category task conv t is expressed as:

[0068] C stu = concat[conv 1 (feats), conv 2 (feats)... conv t (feats)]

[0069] In the formula, conv i represents the convolutional layer of the object category classification head in task i; concat represents the operation of concatenating multiple tensors conv 1 , conv 2 , conv i along a specific dimension to form a larger tensor; feats represents the network feature information.

[0070] (IV) Task alignment learning module

[0071] In this embodiment, the task alignment learning modules of the student network and the teacher network operate independently; the task alignment learning modules of the student network and the teacher network respectively perform high-order combination of the class information prediction and location information prediction of the new task training data set output by the student network, and the class information prediction and location information prediction of the old task training data set output by the teacher network with the intersection IoU between the class and location of the real data to quantify the task sample alignment. After the task sample alignment, the class loss, the generalized intersection over union GIoU loss, and the distribution focal loss are calculated.

[0072] Task Alignment Learning (TAL) is a widely adopted method for aligning the classification and regression feature spaces of responses in various non-incremental dense detectors. To evaluate its effectiveness in incremental object detection, it is first integrated into the constructed incremental detector. TAL uses a high-order combination t of the classification score and the intersection (IoU) between the predicted bounding box and the ground truth to quantify the task alignment, and then uses this alignment score (t) to guide the training sample assignment and weight the classification and localization losses. To enable the task alignment to adapt to the technical solution of this embodiment, we modified the task alignment learning module to adapt to our incremental detector. In this embodiment, in addition to the traditional classification loss L cls and the bounding box regression loss L box explicitly, a distribution focal loss L distribution is also introduced on the predicted distribution and the ground truth label to enhance the learning of localization knowledge. This embodiment combines these three components as the final optimization objective for the new class, denoted as L detector .

[0073] Specifically, for the task alignment learning modules of the student network and the teacher network of the present invention, the class information prediction and location information prediction of the new task training data set output by the student network pass through the task head alignment module. The processing flow of the task head alignment module is as follows: the class information prediction and location information prediction are first mixed and pass through a series of task interaction feature stacks of multiple convolutional layers, and multi-level features and multi-scale effective receptive fields are provided for the two tasks of class and location. Formally, X fpn represents the FPN feature. The feature extractor uses N consecutive convolutional layers with activation functions to calculate the task interaction features:

[0074]

[0075] where conv k and δ refer to the k-th convolutional layer and the ReLU function respectively. In addition, in the task head alignment module, the present invention also inserts a deformable convolutional module, which can automatically focus on different features beneficial to classification and regression. One of the learnable convolutions predicts the offset p nto adaptively select the functional space of the current task. For the feature map F at each position pi, there is:

[0076]

[0077] where p n represents n points of the feature grid output by the TAP module, w(p n ) represents the convolution weight coefficient at the corresponding point position, Δp n represents the offset of this point, and x(.) represents the original feature.

[0078] After the feature map is processed by the task head alignment and deformable convolution module, the class branch and the regression branch are output. For the class branch, the class L1 loss is calculated; for the regression branch, the Generalized Intersection over Union (GIoU) loss and the Distribution Focal Loss are used.

[0079] This embodiment proposes an end-to-end incremental object detection method based on response-based knowledge distillation. Considering the catastrophic forgetting problem of the network, knowledge technology is adopted to transfer the knowledge of the teacher model to the student model, improving the learning efficiency of the network and the model generalization ability; considering that the distributions of different knowledge are different, it may not be beneficial to transfer classification knowledge and localization knowledge simultaneously at the same position, and a divide-and-conquer strategy is adopted, in which different distillation methods are used for classification knowledge and localization knowledge, thus greatly improving the knowledge transfer performance of the network.

[0080] For the incremental learning task, the present invention designs a loss function, and calculates the classification and localization losses L total :

[0081] L total = αL model_cls + L model_reg + βL dis_cls + L dist_reg

[0082] where the loss term L model is the classification and localization loss for the student network, used to calculate the loss of the student network for detecting new targets; L dist represents the knowledge distillation loss; the parameter α is a hyperparameter used to balance the loss weight of the detector, the parameter β is a hyperparameter used to balance the hierarchical distillation loss weight, and the hyperparameters α and β are used to improve the robustness of the detection performance of new and old classes; the subscripts reg and cls correspond to the regression and classification regions respectively.

[0083] Example

[0084] In this example, training samples CoCo2017 are collected starting from the incremental detection task, with a total of 80 classes. An old-task training dataset and a new-task training dataset are constructed, where there are 40 classes for the old task and 40 classes for the new task, and there is no overlap between the classes of the old and new tasks. Such a design ensures the effectiveness of the incremental learning process. Subsequently, a teacher network is constructed, and a Task Alignment Learning (TAL) module is integrated into the teacher network. Then, the training samples of the old task are input into the teacher network for systematic learning and training to obtain the teacher model. Next, a student network is constructed. To enable the student model to better learn the old knowledge of the teacher model, a Task Decoupling and Expansion Strategy module and a Task Alignment Learning (TAL) module are integrated into the student network, and a customized knowledge distillation module is also used. Subsequently, the training samples of the new task are input into the student network to enable the network to have the performance of recognizing new knowledge. At the same time, to avoid catastrophic forgetting of the network, a divide-and-conquer strategy is used, and different distillation losses are used for the location information and class information respectively, and the performance of the teacher model in recognizing the old task is transferred to the student model, enabling the student model to recognize both new and old tasks simultaneously, thereby achieving the effect of incremental learning.

[0085] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. An end-to-end incremental object detection method via intrinsic feature alignment and task decoupling, characterized in that The method comprises the following steps: S1, construct an old task training dataset and a new task training dataset for incremental learning of target detection, where the data in the old task training dataset and the new task training dataset have no intersection; the new task training dataset includes old categories and new categories, where the training data related to the new category has a label, and the training data related to the old category has no label; S2, build a teacher network, integrate the task alignment learning module into the teacher network, and train the old task training data set to obtain the teacher model; S3, construct a student network, integrate the task decoupling extension strategy module, task alignment learning module and knowledge distillation module in the student network, and output the position information header and category information header of the new task training data set through the student network, which are used for position prediction and category prediction respectively, and the position information header output by the student network is only designed to generate the bounding box; the knowledge distillation module transfers the old knowledge learned by the teacher model to the student model, so that the student model has the ability to recognize both new and old targets; The task alignment learning module introduces distribution focus loss on the predicted distribution and the ground truth label, and combines the distribution focus loss, classification loss and bounding box regression loss as the final optimization target of the teacher network or the student network; S4, when processing a new task, the category information header output by the student network first passes through the task decoupling network expansion module to decouple the new category and old category information; for the new category, the student network directly outputs the position information header and the category information header, which have the information of the new target; for the old category, the student network distills the category information header and the position information header of the teacher network through knowledge, so that the output position information header and the category information header obtain the information of the old target.

2. The end-to-end incremental object detection method through intrinsic feature alignment and task decoupling according to claim 1, characterized in that The method further comprises the following steps: A spatiotemporal interaction module of matrix operation is inserted between two feature layers of the backbone network of the student network and the teacher network, and the original feature channel weights are weighted to the original features channel by channel through multiplication to complete the recalibration of the original features in the channel dimension; specifically, the spatiotemporal interaction module of matrix operation uses matrix multiplication to process the feature graphs of the teacher network and the student network; the spatiotemporal interaction module includes three parts: an input dimension conversion unit, a spatiotemporal interaction unit and an excitation weighting unit; Among them, the input dimension conversion unit is used to convert the feature map with input dimension [C×T×H×W] into a feature map with dimension [CT×HW] through matrix conversion operation, which is used as the first operand of matrix multiplication; the spatiotemporal interaction unit performs a series of dimension reshape operations and dimension-invariant 1×1 convolution operations on the feature map with input dimension [C×T×H×W] to convert it into a feature map with dimension [HW×1×1], which is used as the second operand of matrix multiplication; the excitation weighting unit performs matrix multiplication operation on the first operand and the second operand, and then weights the channel weights to the original features channel by channel through multiplication to perform recalibration of the original features in the channel dimension.

3. The end-to-end incremental object detection method through intrinsic feature alignment and task decoupling according to claim 1, characterized in that: When the features extracted by the front-end neck network module are input, the task decoupling network expansion module generates new category information headers and old category information headers for the features of the new category and the old category respectively; for the new category, the student network detects the new category based on the new category information header; for the old category, the student network transfers knowledge to the old category information header of the student network through knowledge distillation of the category information of the teacher network based on the decoupled old category information header; the new category information header and the old category information header contain multiple layers of features respectively, and the response of each layer of features can perform convolution operations on the features of the previous layer, and finally connect these responses to keep them consistent with the output in the minimum expansion method, and the new category information header and the old category information header each maintain a separate category network; The final output response of the new category information detection head is expressed as: C stu =concat[conv 1 (feats),conv 2 (feats)...conv t (feats)] In the formula, conv i Represents the convolution operation on the i-layer object category classification head; concat represents multiple tensors conv 1 ,conv 2 ,conv i Connect along a specific dimension to form a larger tensor operation; feats represents network feature information.

4. The end-to-end incremental object detection method through intrinsic feature alignment and task decoupling according to claim 1, characterized in that The knowledge distillation module takes the entire classification response as the classification distillation region, and the classification knowledge distillation loss is expressed as: Where P tea is the confidence of the classification response in the teacher detector, M represents the number of predicted samples in the classification region cls, K represents the number of all detectable categories in the teacher model, and S tea Indicates that the last layer of features of the teacher model is transformed using the Sigmoid function, S stu Indicates that the last layer of features of the student model is transformed using the Sigmoid function. Indicates seeking the minimum value; At the same time, the knowledge distillation module first obtains μ, β through confidence for the category information of the teacher network, where μ is obtained by Mean(P tea ) is obtained by averaging, and β is the standard deviation StandardDeviation (P tea ) is obtained, and then μ, β are weighted summed to obtain θ, and P is compared tea The initial candidate regions are selected with θ, and finally the regression distillation regions are filtered out through NMS. After the regression output information of the teacher network and the regression information of the student network are input into the knowledge distillation module, the regression information of the student network learns knowledge through the regression output information of the teacher network through knowledge distillation. The regression knowledge distillation loss is expressed as: Where N represents the number of predicted sample regression regions reg, B i Indicates the regression response i th The predicted bounding box corresponding to the sample. The superscripts tea and stu represent the teacher model and the student model respectively.

5. The end-to-end incremental object detection method through intrinsic feature alignment and task decoupling according to claim 1, characterized in that: The task alignment learning modules of the student network and the teacher network operate independently of each other; the task alignment learning modules of the student network and the teacher network respectively perform a high-order combination of the intersection IoU between the category information prediction and position information prediction of the new task training data set output by the student network, the category information prediction and position information prediction of the old task training data set output by the teacher network and the category and position of the real data to quantify the task sample alignment, and then calculate the category loss, generalized intersection-over-union (GIoU) ​​loss and distribution focal loss (Distribution FocalLoss) after the task samples are aligned; Specifically, the task alignment learning module includes a task head alignment module and a deformable convolution module; wherein, after the category information prediction and the position information prediction enter the task head alignment module, they are mixed through a series of task interaction feature stacks of multiple convolutional layers, providing multi-level features and multi-scale effective receptive fields for the two tasks of category and position, and the deformable convolution module automatically focuses on different features that are beneficial to classification and regression; wherein the convolution layer of the task head alignment module predicts the offset p at each position n To adaptively select the function space of the current task; for each position p i The feature map F(p i ),have: Where p n represents the n points of the feature grid output by the TAP module, w(p n ) represents the convolution weight coefficient of the corresponding point position, Δp n represents the offset of the point, x(.) represents the original feature; R represents the regular grid; After the feature map is processed by the task head alignment module and the deformable convolution module, the category branch and the regression branch are output; for the category branch, the category L1 loss is calculated; for the regression branch, the generalized intersection-over-union (GIoU) ​​loss and the distribution focal loss (Distribution FocalLoss) are used.

6. The end-to-end incremental object detection method through intrinsic feature alignment and task decoupling according to claim 1, characterized in that: The overall loss function L of the teacher network and the student network total for: L total =a model_cls +L model_reg +βL dist_cls +L dist_reg Among them, the loss term L model is the classification and positioning loss for the student network, which is used to calculate the loss of the student network to detect new targets; L dist It is expressed as knowledge distillation loss; parameter α is a hyperparameter used to balance the detector loss weight, parameter β is a hyperparameter used to balance the graded distillation loss weight, and hyperparameters α and β are used to improve the robustness of new and old category detection performance; subscripts reg and cls correspond to regression and classification areas, respectively.

Citation Information

Cited By

  • Student model training method and device based on gradient decoupling

    CN121390201A