An intelligent target detection method based on incremental learning and catastrophic forgetting prevention

CN122821098APending Publication Date: 2026-09-25XIDIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611037822.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

(1)现有技术在处理新数据时,模型常常会忘记之前学习到的重要特征,导致在旧数据集上的性能急剧下降,影响整体检测准确率

Benefits of technology

(1)本发明通过引入知识蒸馏机制,确保教师网络能够有效传递旧知识给学生网络,同时引入正则化约束,在训练过程中动态平衡新旧数据,从而减轻灾难性遗忘,使得学生网络在引入新数据后,仍能保持对旧数据的高准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821098A_ABST
    Figure CN122821098A_ABST
Patent Text Reader

Abstract

The application provides an intelligent target detection method based on incremental learning and catastrophic forgetting prevention, which greatly improves the flexibility of target detection by providing multiple basic target detection algorithms selected by users, and only the student network participates in the training of the current multi-modal data in the incremental process, thereby reducing the computing cost and training time. Moreover, the application combines the regularization protection mechanism with knowledge distillation to provide double protection for the training process of the student network, effectively prevents catastrophic forgetting, and ensures the balance of new and old tasks. In addition, the application tracks the performance of the student network in real time to adjust the training strategy, further improving the adaptability and stability of the student network. The application also has the ability to process multi-modal data, enhancing the generalization performance of the student network, and is particularly suitable for target detection tasks in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of target detection technology, specifically relating to an intelligent target detection method based on incremental learning and catastrophic forgetting protection. Background Technology

[0002] Object detection technology is an important branch of computer vision, widely used in security monitoring, autonomous driving, and medical image analysis. With the rapid development of deep learning, object detection algorithms based on Convolutional Neural Networks (CNNs), such as YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), and R-CNN (Region-based Convolutional Neural Network), have made significant progress. These algorithms have driven the widespread application of intelligent vision systems by achieving fast and high-precision object detection. However, in dynamic environments, data distribution and categories may change over time, and traditional models often cannot adapt to new data distributions, leading to performance degradation on new data. Incremental learning methods have emerged to address this, aiming to allow models to retain old knowledge while continuously receiving new data.

[0003] Patent CN202210199571.7 addresses the incremental learning problem in image object detection through deep supervised self-distillation, particularly targeting the catastrophic forgetting phenomenon commonly seen with new data. This method constructs two parallel neural networks (a teacher model and a student model). During training, the teacher model learns features from old data and transfers this knowledge to the student model, while the student model simultaneously learns features from both new and old data. Its core lies in designing a deep supervised distillation loss function, including feature distillation loss, output distillation loss, and ground truth loss. This allows the model to adaptively balance features from both new and old data during updates, improving generalization ability and preventing forgetting.

[0004] In the current field of object detection, although existing technologies attempt to address the problems in incremental learning, issues such as catastrophic forgetting, insufficient feature sharing, and low training efficiency still exist. These shortcomings not only affect model performance but also limit the application of object detection algorithms in diverse and dynamic environments. Specifically, they manifest as follows: (1) When processing new data, existing technologies often forget the important features learned previously, resulting in a sharp decline in performance on old datasets and affecting the overall detection accuracy.

[0005] (2) Existing technologies usually focus on processing a single modality and lack effective strategies to process data from different modalities simultaneously, which limits the detection capability in complex scenarios (such as low light or smoke environments).

[0006] (3) Many existing methods require retraining the entire model, especially when new data is introduced. This retraining process is not only time-consuming, but also consumes a lot of computing resources. Frequent retraining and the need for large-scale datasets also lead to high costs for existing solutions, which are difficult to afford in practical applications and limit their promotion.

[0007] (4) Existing systems lack the ability to dynamically select the best algorithm based on the characteristics of the input data (such as lighting conditions, target type, etc.). Most detection systems use fixed algorithms and do not consider the differences in the characteristics of the input data, resulting in insufficient detection accuracy in complex scenes. Summary of the Invention

[0008] To address the aforementioned problems in existing technologies, this invention provides an intelligent target detection method based on incremental learning and catastrophic forgetting protection. The technical problem to be solved by this invention is achieved through the following technical solution: An intelligent target detection method based on incremental learning and catastrophic forgetting protection includes: S100: Acquire historical multimodal data and preprocess it to obtain preprocessed data; S200: The user-selected basic target detection model is trained using preprocessed data to obtain the trained basic target detection model as the teacher network; S300, the weights of the teacher network are used as the initial weights of the student network, and the current multimodal data is obtained; the model architectures of the teacher network and the student network are similar; S400 utilizes the output distillation loss and feature distillation loss between the teacher network and the student network, and designs the total loss function of the student network through EWC regularization; S500, the current multimodal data is input into the student network for training, and the total loss function is used to guide the learning process of the student network, resulting in the target detection result output by the student network at the end of the last training.

[0009] Beneficial effects: (1) This invention introduces a knowledge distillation mechanism to ensure that the teacher network can effectively transfer old knowledge to the student network. At the same time, it introduces regularization constraints to dynamically balance the old and new data during the training process, thereby reducing catastrophic forgetting and enabling the student network to maintain a high accuracy rate for the old data after the introduction of new data.

[0010] (2) This invention provides an intelligent target detection method based on incremental learning and catastrophic forgetting protection, which can process image data of multiple modalities (such as visible light and infrared images) and perform flexible model optimization and target detection tasks according to the characteristics of the data modalities.

[0011] (3) By designing a mechanism that trains only the student network, the present invention avoids retraining the teacher network, thereby significantly shortening the training time, reducing computational requirements, making the incremental learning process more efficient, thereby reducing the overall implementation cost, and making the method of the present invention more feasible in practical applications.

[0012] (4) This invention provides a dynamic selection mechanism for different types of target detection algorithms, which allows users to flexibly select target detection algorithms according to specific task requirements. This dynamic selection mechanism improves the adaptive ability of the detected target and can select the optimal model according to the actual scenario, thereby improving detection efficiency and accuracy.

[0013] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating an intelligent target detection method based on incremental learning and catastrophic forgetting protection provided by the present invention. Figure 2 This is an overall schematic diagram of the intelligent target detection method based on incremental learning and catastrophic forgetting protection provided by the present invention; Figure 3 This is a detailed schematic diagram of the intelligent target detection method based on incremental learning and catastrophic forgetting protection provided by the present invention; Figure 4 This is a schematic diagram illustrating the selection of a basic target detection model based on the actual application scenario provided by the present invention; Figure 5 This is a schematic diagram illustrating the detection performance of the student network when the regularization coefficient takes different values, as provided by the present invention. Figure 6 This is a schematic diagram of knowledge distillation provided by the present invention; Figure 7 This is a schematic diagram illustrating the process of dynamically adjusting student network performance provided by the present invention. Detailed Implementation

[0015] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0016] Combining as 1 to Figure 4 This invention provides an intelligent target detection method based on incremental learning and catastrophic forgetting protection, comprising: S100: Acquire historical multimodal data and preprocess it to obtain preprocessed data; In one specific embodiment of the present invention, S100 includes: S110, acquire historical multimodal data containing multiple visible light and infrared images; S120, the historical multimodal data is sequentially denoised, aligned and normalized to obtain normalized multimodal data; S130, Use a feature analysis tool to perform data feature analysis on the normalized multimodal data to obtain statistical features under different modes; refer to Figure 2 This step utilizes feature analysis tools to perform data feature analysis on the normalized multimodal data, such as illumination condition detection, target basic feature analysis, and overall image quality analysis, to obtain the statistical characteristics of the multimodal data under different modalities. Examples include pixel value distribution, texture information, and color channel distribution.

[0017] S140, the normalized multimodal data is grouped using the statistical features to obtain visible light image group and infrared image group, and the visible light image group and infrared image group are used as preprocessed data.

[0018] This step utilizes statistical features to distinguish between visible light and infrared images, and ultimately outputs feature analysis results for multimodal data, including intermodal differences, data grouping, and quality, providing a basis for users to dynamically select the most suitable basic target detection model.

[0019] S200: The user-selected basic target detection model is trained using preprocessed data to obtain the trained basic target detection model as the teacher network; In an optional embodiment of the present invention, S200 includes: S210, receives the basic target detection model selected by the user in the visual interface according to the target detection task; Combination Figure 2 and Figure 3 The basic object detection model of this invention includes various object detection algorithms, such as YOLOv5, SSD, Cascade R-CNN, and Faster R-CNN. Users can manually select YOLOv5, SSD, Cascade R-CNN, Faster R-CNN, or other classic object detection algorithms according to their specific needs. The framework supports flexible algorithm replacement. For example, YOLOv5 can be selected for fast detection in real-time applications, while Cascade R-CNN can be switched for tasks requiring high accuracy. This allows users to quickly switch algorithm modules between different tasks to meet different needs. Figure 4 This diagram illustrates how users can choose the appropriate algorithm for practical applications. In security monitoring scenarios, where high target detection accuracy is required, the Faster R-CNN detection algorithm can be selected. In drone aerial photography scenarios, where real-time target detection is crucial, the YOLOv5 detection algorithm can be chosen. In intelligent traffic monitoring scenarios, strong real-time multi-scale target detection capabilities are required, so the SSD detection algorithm can be selected. This invention allows for the adaptive selection of suitable target detection algorithms based on actual needs; the specific applicable detection algorithm is not limited herein.

[0020] S220, The preprocessed data is input into the basic target detection model for training, and the training process is optimized using cross-entropy loss and regression loss until the training ends, resulting in a trained basic target detection model. S230, the trained basic target detection model is used as the teacher network.

[0021] refer to Figure 3 After selecting a basic target detection model, this invention begins the training process of the basic target detection model. After the training is completed, a teacher network is obtained, and then incremental learning is performed.

[0022] S300, the weights of the teacher network are used as the initial weights of the student network, and the current multimodal data is obtained; the model architectures of the teacher network and the student network are similar, and the student network is relatively lightweight; The teacher network shares the same backbone as the student network, which includes a feature extraction layer responsible for extracting features from the input data. The teacher and student networks differ in their classification, regression, detection, and output layers. The student network has fewer classification, regression, detection, and output layers than the teacher network, retaining only the key layers for new tasks.

[0023] The current multimodal data can be infrared images and / or visible light images under different conditions.

[0024] Combination Figure 2 and Figure 3 In the incremental learning process, it is necessary to ensure that the student network can learn new knowledge without forgetting old knowledge. Therefore, the weights of the teacher network are used as the initial weights of the student network. The weights of the teacher network do not participate in backpropagation and updates. The student network extracts features of the current multimodal data based on the feature extraction capabilities of the teacher network.

[0025] S400 utilizes the output distillation loss and feature distillation loss between the teacher network and the student network, and designs the total loss function of the student network through EWC regularization; Combination Figure 2 and Figure 3 This invention designs a knowledge distillation mechanism and a regularization protection mechanism to improve the performance of the student network. When introducing current multimodal data, the knowledge distillation mechanism transfers knowledge of old tasks from the teacher network to the student network, enabling the student network to learn new tasks without forgetting old ones. Simultaneously, regularization (EWC) is employed to constrain changes in key model weights, preventing the learning of new data from interfering with the knowledge of old data, thereby protecting the model weights and avoiding catastrophic forgetting.

[0026] In one specific embodiment of the present invention, S400 includes: S410, determine the predicted outputs and feature representations of the teacher network and the student network respectively; S420, by utilizing the difference between the predicted outputs of the teacher network and the student network, designs an output distillation loss function; In training the incremental branch, i.e. training the student network, this invention uses the output of the teacher network (old knowledge) as a distillation signal, combined with the output of the student network, to calculate the difference between the teacher network and the student network through feature distillation loss and output distillation loss. This difference is then added as a constraint to the total loss of incremental learning, thereby guiding the learning of the incremental branch and ensuring that the learning of new knowledge does not forget old knowledge.

[0027] Combination Figure 2 and Figure 3 The teacher network uses the old dataset for forward propagation to generate predicted outputs. and feature representation The student network uses the same input image to generate predicted output. and feature representation The output distillation loss function typically uses the Kullback-Leibler divergence (KL divergence) to measure the difference between two output distributions. The output distillation loss function is expressed as: ; In the formula, This indicates the output distillation loss function. It is the Kullback-Leibler divergence. It is the probability distribution of the predicted output of the teacher network. It is the probability distribution of the predicted output of the student network. It is the predicted output of the teacher network. It is the predicted output of the student network. It is the input for both teacher and student networks. This is a temperature parameter used to adjust the smoothness of the softmax function.

[0028] S430, utilizing the differences in feature representations between the teacher network and the student network, designs a feature distillation loss function; whereby the feature distillation loss function is expressed by the formula: ; In the formula, Represents the characteristic distillation loss function. This represents the feature representation of the penultimate layer of the teacher network. This represents the feature representation of the penultimate layer of the student network. This indicates the calculation of the square value of the L2 norm.

[0029] Before calculating the feature distillation loss, the features of the teacher network are first size aligned so that the feature representations of the teacher network have the same size and number of channels as the feature representations of the student network.

[0030] S440, the EWC regularization loss is designed using the EWC regularization method, and the total loss function of the student network is designed by combining the output distillation loss function and the feature distillation loss function.

[0031] During the training of the incremental branch, the importance of each layer's weights to the old task is calculated using the EWC (Elastic Weight Consolidation) regularization method. The idea behind EWC is to use the Fisher information matrix to estimate the importance of each parameter and impose strong regularization constraints on important parameters. The EWC regularization loss is expressed by the formula: ; In the formula, This represents the EWC regularization loss. This is the regularization coefficient, used to control the strength of regularization. These are parameters in the Fisher information matrix. The value represents the parameter. The importance of old knowledge It is the first The parameters of the student network were trained. It is the first Historical parameters of the student network during the training session; Figure 5 This demonstrates how, using a general model as an example, different... The different effects of student network training under numerical conditions ( When the regularization coefficient is too small, overfitting occurs, and while the accuracy of the training set increases, the accuracy of the test set decreases. When the regularization constraint is strengthened, the generalization ability of the student network is significantly enhanced, verifying the important role of regularization loss in enhancing model robustness and preventing overfitting. The total loss function of the student network can be designed through weighted combination, and the total loss function of the student network is expressed by the formula: ; In the formula, , , The weights are used to control the importance of output distillation loss, feature distillation loss, and EWC regularization loss in the total loss function, respectively. These three loss functions work together to ensure that the student network inherits the knowledge of the teacher network while preventing the forgetting of knowledge from old tasks during incremental learning.

[0032] S500, the current multimodal data is input into the student network for training, and the total loss function is used to guide the learning process of the student network, resulting in the target detection result output by the student network at the end of the last training.

[0033] refer to Figure 6 The student network is trained using the current multimodal data. Newly added branches are primarily responsible for processing the current multimodal data, and only the weights of these new branches are optimized. The weights of the teacher network remain fixed to avoid disrupting existing knowledge. Knowledge is transferred from the teacher network to the student network, and distillation and regularization optimizations are performed during the transfer process.

[0034] The student network is trained specifically for the current multimodal data, while the teacher network remains unchanged. The incremental branch is trained only on the new task, avoiding the need to retrain the teacher network, thus significantly reducing computational costs and training time, while ensuring the student network's rapid adaptability to the current multimodal data and stability on older tasks.

[0035] Following S500, the intelligent target detection method based on incremental learning and catastrophic forgetting protection further includes: calculating a detection index using the target detection results, and dynamically adjusting the weights, regularization coefficients, and learning rates of the output distillation loss and feature distillation loss according to the detection index.

[0036] During training and inference, this invention can monitor the performance of student networks on both old and new data. Combined with... Figure 2 , Figure 3 and Figure 7This invention evaluates the detection performance of a student network during incremental learning by calculating metrics such as F1 score and mAP (mean Average Precision). If the student network's performance on prior knowledge significantly declines, it is protected and optimized by adjusting the distillation loss weights or regularization coefficients. Based on the student network's detection performance, the training strategy is dynamically adjusted, including adjusting the incremental learning rate, regularization loss, and loss function, to optimize the student network's performance.

[0037] refer to Figure 2 The student network outputs object detection results, such as detected object categories, bounding boxes, and confidence scores, for practical applications (e.g., object detection, image classification). This invention can summarize the performance metrics of the student network on both new and old tasks. It generates reports from the object detection results of the student network, using metrics such as F1 score and mAP, to evaluate the overall performance of the student network.

[0038] This invention provides users with a variety of basic object detection algorithms to choose from, significantly improving the flexibility of object detection compared to a single algorithm. Furthermore, when introducing current multimodal data, it trains only the student network, thereby reducing computational costs and training time. Moreover, this invention combines regularization protection mechanisms with knowledge distillation, providing dual protection for the student network training process, effectively preventing catastrophic forgetting and ensuring a balance between new and old tasks. In addition, this invention tracks the performance of the student network in real time to adjust the training strategy, further improving the adaptability and stability of the student network. This invention also has the ability to process multimodal data, enhancing the generalization performance of the student network, making it particularly suitable for object detection tasks in complex environments.

[0039] Although this application has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, the disclosure, and the appended claims in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality.

[0040] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. An intelligent target detection method based on incremental learning and catastrophic forgetting protection, characterized in that, include: S100: Acquire historical multimodal data and preprocess it to obtain preprocessed data; S200: The user-selected basic target detection model is trained using preprocessed data to obtain the trained basic target detection model as the teacher network; S300, the weights of the teacher network are used as the initial weights of the student network, and the current multimodal data is obtained; the model architectures of the teacher network and the student network are similar; S400 utilizes the output distillation loss and feature distillation loss between the teacher network and the student network, and designs the total loss function of the student network through EWC regularization; S500, the current multimodal data is input into the student network for training, and the total loss function is used to guide the learning process of the student network, resulting in the target detection result output by the student network at the end of the last training.

2. The intelligent target detection method based on incremental learning and catastrophic forgetting protection according to claim 1, characterized in that, S100 includes: S110, acquire historical multimodal data containing multiple visible light and infrared images; S120, the historical multimodal data is sequentially denoised, aligned and normalized to obtain normalized multimodal data; S130, Use a feature analysis tool to perform data feature analysis on the normalized multimodal data to obtain statistical features under different modes; S140, the normalized multimodal data is grouped using the statistical features to obtain visible light image group and infrared image group, and the visible light image group and infrared image group are used as preprocessed data.

3. The intelligent target detection method based on incremental learning and catastrophic forgetting protection according to claim 1, characterized in that, S200 includes: S210, receives the basic target detection model selected by the user in the visual interface according to the target detection task; S220, The preprocessed data is input into the basic target detection model for training, and the training process is optimized using cross-entropy loss and regression loss until the training ends, resulting in a trained basic target detection model. S230, the trained basic target detection model is used as the teacher network.

4. The intelligent target detection method based on incremental learning and catastrophic forgetting protection according to claim 1 or 3, characterized in that, The basic object detection models include YOLOv5, SSD, R-CNN, and Faster R-CNN.

5. The intelligent target detection method based on incremental learning and catastrophic forgetting protection according to claim 1, characterized in that, The S400 includes: S410, determine the predicted outputs and feature representations of the teacher network and the student network respectively; S420, by utilizing the difference between the predicted outputs of the teacher network and the student network, an output distillation loss function is designed; S430, utilizing the differences in feature representations between the teacher network and the student network, designs a feature distillation loss function; S440, the EWC regularization loss is designed using the EWC regularization method, and the total loss function of the student network is designed by combining the output distillation loss function and the feature distillation loss function.

6. The intelligent target detection method based on incremental learning and catastrophic forgetting protection according to claim 5, characterized in that, The output distillation loss function in S420 is expressed by the following formula: ; In the formula, This indicates the output distillation loss function. It is the Kullback-Leibler divergence. It is the probability distribution of the predicted output of the teacher network. It is the probability distribution of the predicted output of the student network. It is the predicted output of the teacher network. It is the predicted output of the student network. It is the input for both teacher and student networks. This refers to the temperature parameter.

7. The intelligent target detection method based on incremental learning and catastrophic forgetting protection according to claim 6, characterized in that, The characteristic distillation loss function in S430 is expressed by the following formula: ; In the formula, Represents the characteristic distillation loss function. This represents the feature representation of the penultimate layer of the teacher network. This represents the feature representation of the penultimate layer of the student network. This indicates the calculation of the square value of the L2 norm.

8. The intelligent target detection method based on incremental learning and catastrophic forgetting protection according to claim 7, characterized in that, The EWC regularization loss in S440 is expressed by the following formula: ; In the formula, This represents the EWC regularization loss. This is the regularization coefficient, used to control the strength of regularization. These are parameters in the Fisher information matrix. The value represents the parameter. The importance of old knowledge It is the first The parameters of the student network were trained. It is the first Historical parameters of the student network during each training session.

9. The intelligent target detection method based on incremental learning and catastrophic forgetting protection according to claim 8, characterized in that, The total loss function of the S440 student network is expressed by the formula: ; In the formula, , , The weighting coefficients control the importance of output distillation loss, feature distillation loss, and EWC regularization loss in the total loss function, respectively.

10. The intelligent target detection method based on incremental learning and catastrophic forgetting protection according to claim 9, characterized in that, Following S500, the intelligent target detection method based on incremental learning and catastrophic forgetting protection also includes: The detection index is calculated using the target detection results, and the weights, regularization coefficients, and learning rates of the output distillation loss and feature distillation loss are dynamically adjusted based on the detection index.

Citation Information

Patent Citations

  • A method for image object detection based on deep supervised self-distillation

    CN114863248B