Class image incremental learning method and system based on diffusion model

Through the class image incremental learning method based on diffusion model, high-quality pseudo-samples are generated and category distinction is optimized, which solves the catastrophic forgetting and imbalance problems in class incremental learning, and improves the classification performance and adaptability of the model.

CN120495739APending Publication Date: 2025-08-15SHANGHAI LINGANG JINGHONG SECURITY TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510566984.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing class incremental learning methods have obvious bottlenecks in catastrophic forgetting, generation sample stability and classifier feature consistency, especially when dealing with grayscale fluctuations or lighting uneven scenarios, the stability of classification results decreases.

Method used

The class image incremental learning method based on diffusion model is adopted, pseudo-samples are generated through the Unet backbone network and the PEDCC layer, and the classification head is combined for classification. The diffusion model is used for forward diffusion and reverse denoising processes to generate high-quality pseudo-samples, and the category distinction is optimized through POD Loss and NaC Loss to achieve incremental learning.

Benefits of technology

Effectively solve the problems of category imbalance and catastrophic forgetting, improve the classification performance and generalization capabilities of the model, maintain the recognition ability of old categories, adapt to the learning of new categories, and improve the performance of the system in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495739A_ABST
    Figure CN120495739A_ABST
Patent Text Reader

Abstract

The invention relates to the field of class image incremental learning, and provides a class image incremental learning method and system based on a diffusion model, and the method comprises the steps: obtaining a training data set; real images of all image categories in the initial task data set are used as input data of a network architecture, image features of the real images are extracted through a Unet backbone network, and category discrimination of the image features in forward diffusion and reverse denoising processes in a diffusion model is optimized according to parameters set by a PEDCC layer; generating pseudo sample classifications to form an effective pseudo sample set; and taking the first image category in the incremental task data set as a data set of the incremental tasks, combining the data set of the incremental tasks with the pseudo samples in the effective pseudo sample set to form a new task data set, and inputting the new task data set into a network architecture to carry out incremental learning playback training until all the incremental tasks are trained. The problems of class imbalance and disastrous forgetting are effectively solved, and meanwhile the classification performance of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of incremental learning of category images, and specifically to a method and system for incremental learning of category images based on a diffusion model, which is mainly used to integrate image generation and image classification functions, and use generated images to generate playback, thereby improving the incremental learning ability of category images. Background Art

[0002] With the widespread application of artificial intelligence in computer vision, image classification models are being widely deployed in real-world scenarios such as security surveillance, medical imaging, and autonomous driving. Real-world tasks often require continuous scalability, meaning that models must continuously learn new categories while maintaining the ability to recognize previously learned categories, creating the need for incremental learning. The core challenge of incremental learning is catastrophic forgetting, whereby the model loses its ability to recognize old categories when learning new ones, resulting in overall performance degradation.

[0003] How to alleviate catastrophic forgetting is the research focus of incremental learning. In current research work, the methods to alleviate catastrophic forgetting in incremental learning can be roughly divided into three categories: (1) Constraint methods based on regularization, which impose constraints on important parameters in the old task by adding regularization terms in the learning of the new task, thereby limiting the forgetting of old knowledge. However, in practical applications, it is difficult to design a reasonable metric to estimate the importance of each parameter in the neural network; (2) Bias correction methods, which correct the bias between the new and old categories according to some bias correction strategies to avoid the incremental learning network being biased towards the category of the new task, but new biases may be introduced during the correction process; (3) Replay methods, which can be divided into three categories according to the content of the replay: example replay, generated replay, and feature replay, that is, by replaying the examples of the old task or generating data or features to memorize the information of the old category. Among them, the incremental learning method based on example replay selects some representative old samples from the original dataset of the old task as memory samples according to some sample screening strategies, and trains them together with the new task dataset. However, in some scenarios, due to data privacy issues, this type of method will not be able to obtain the original data of the old task. Quasi-incremental learning methods based on feature replay require storing features extracted from previous tasks. While storing features consumes less memory than storing samples, this approach suffers from poor model flexibility when the feature extraction model is fixed, and feature drift occurs when the feature extraction model is updated. Methods based on generative replay avoid the problem of old data being unavailable by training a generative model. However, the quality of the generated samples is directly affected by the generative model's capabilities. Sequentially updating the generative model can lead to catastrophic forgetting, which can increase forgetfulness in the incremental classifier model and result in poor performance.

[0004] Convolutional neural networks (CNNs), the foundational architecture for image classification, have achieved excellent performance in various fields. However, they are inherently sensitive to differences in the modulus values of different image channels, such as RGB. This leads to reduced classification stability when dealing with scenes with fluctuating grayscale or uneven lighting. For example, in handwritten text recognition, the classifier can misclassify the same character due to varying shades of handwriting.

[0005] In summary, existing incremental learning methods still have obvious bottlenecks in alleviating catastrophic forgetting, generating sample stability, and classifier feature consistency. There is an urgent need for a category image incremental learning method and system based on a diffusion model to achieve a more stable and efficient image category incremental learning process. Summary of the Invention

[0006] In response to the defects in the prior art, the purpose of the present invention is to propose a category image incremental learning method and system based on a diffusion model to solve the problems of category imbalance and catastrophic forgetting in the traditional image incremental learning method proposed in the above background technology.

[0007] To achieve the above objectives, the present invention is implemented through the following technologies:

[0008] The present invention provides a method for incremental learning of category images based on a diffusion model, comprising:

[0009] Step S1: Obtain a training dataset and divide it into an initial task dataset and an incremental task dataset;

[0010] Step S2: Construct a network architecture for incremental learning of category images. The network architecture includes a Unet backbone network, a PEDCC layer, a diffusion model, and a classification head.

[0011] Step S3: In the basic training phase, real images of all image categories in the initial task dataset are used as input data for the network architecture. The image features of the real images are extracted through the Unet backbone network and transmitted to the PEDCC layer. The image features are optimized according to the parameters set by the PEDCC layer during the forward diffusion and reverse denoising processes in the diffusion model to generate pseudo samples of the real images. Finally, the pseudo samples are classified using the classification head and the corresponding classification results are output. This also includes the process of screening out pseudo samples whose classification confidence exceeds the confidence threshold to form a valid pseudo sample set for the incremental learning replay training phase.

[0012] Step S4: In the incremental learning replay training phase, based on the trained network architecture, the real images of the Tth image category in the incremental task dataset are used as the dataset of the incremental task, and are merged with the pseudo samples in the valid pseudo sample set to form a new task dataset, and ρ is a natural number greater than 0. The new task dataset is input into the network architecture for incremental learning replay training, and the classification results and corresponding pseudo samples corresponding to the new task dataset are output; wherein, the pseudo samples of the real images whose classification confidence exceeds the classification confidence threshold are screened out and added to the valid pseudo sample set;

[0013] Step S5: Merge the valid pseudo sample set and the data set of the ρ+1th incremental task in the incremental task data set into a new task data set, and perform incremental learning replay training until all incremental tasks in the incremental task data set are trained.

[0014] Furthermore, step S1 includes:

[0015] Step S11: obtaining a training dataset containing multiple image categories, where the image categories include digital categories and general object categories;

[0016] Step S12: Divide the training dataset into an initial task dataset and an incremental task dataset according to a preset ratio, wherein each image category in the incremental task dataset is used as an incremental task for incremental learning playback training;

[0017] Step S13: Perform image preprocessing on the real images in the initial task dataset and the incremental task dataset. The image preprocessing includes converting the real images into tensor format and performing image enhancement and pixel normalization.

[0018] Furthermore, step S3 includes:

[0019] Step S31: Input the initial task dataset into the network architecture, extract image features through the Unet backbone network, and generate corresponding class centroids according to the parameters set in the PEDCC class center file in the PEDCC layer to optimize the class discrimination of image features in the feature space;

[0020] Step S32: Based on the image features optimized by the PEDCC layer, in the forward diffusion process of the diffusion model, the image features are forward diffused and added as noise, converted into a noise distribution, and the original image features are restored through the reverse denoising process to generate corresponding pseudo samples;

[0021] Step S33: input the generated pseudo samples into the classification head for classification, and output the category prediction result of each pseudo sample;

[0022] Step S34: Calculate the classification confidence of the pseudo samples, screen valid pseudo samples according to the set confidence threshold, and form a valid pseudo sample set for the incremental learning replay training phase.

[0023] Furthermore, step S31 includes:

[0024] Input the initial task dataset containing multiple categories of images into the network architecture, and extract the image features of the real images in the category images through the Unet backbone network;

[0025] In the PEDCC layer, the class centroids are uniformly generated in the feature space according to the PEDCC algorithm, and the feature characteristics are optimized so that the image features of different categories are reasonably distributed in the feature space and the class discrimination is improved;

[0026] The image features optimized by the PEDCC layer are output to the diffusion model to generate pseudo samples.

[0027] Furthermore, in step S32, in the forward diffusion process of the diffusion model, the image features are forward diffused and added as noise to convert into a noise distribution, including:

[0028] Get the image feature x0 and use it as the initial state input in the diffusion model;

[0029] The time step t randomly sampled from the uniform distribution U(0,T) is used to determine the intensity of the noise added during the forward diffusion process;

[0030] Update the image features according to the diffusion equation dx = f(x, t)dt + g(t)dw, and pass the drift coefficient f(x, t) and diffusion coefficient g (t) Adjust the image feature changes, x is the state at time step t during the forward diffusion process, and w is the standard Wiener process;

[0031] By formula Add noise ∈ to the original image feature x0 and convert it into a noise distribution image x t and through α t Control weight impact;

[0032] By minimizing the loss function The value of optimizes the forward diffusion process, where s θ (x t ,t) is the model output, Part of the noise.

[0033] Furthermore, in step S32, the original image features are restored through an inverse denoising process to generate corresponding pseudo samples, including:

[0034] Sample from the noise distribution image to obtain the initial noise sample image xt , used to perform the reverse process through the diffusion model to remove noise and restore image features;

[0035] The label y of the image category corresponding to the real image in the initial task dataset is used as a conditional input to guide the generation of corresponding pseudo samples in the reverse denoising process;

[0036] According to the reverse denoising formula p θ (X t-1 |x t ,y)=N(X t·1 ;μ θ (x t ,t,y),Σ θ (x t ,t)), calculate the image x t Restore to state X t-1 process, where μ θ (x t ,t,y) is the predicted mean, Σ θ (x t ,t) is the covariance matrix of the noise, N is the number of samples in the data set;

[0037] The generated image X t-1 Pseudo samples as output.

[0038] Furthermore, in step S3, pseudo samples whose classification confidence exceeds the confidence threshold are screened out to form a valid pseudo sample set, including:

[0039] The classification head maps the pseudo samples to the category label space for classification, calculates the classification loss of each pseudo sample through the cross entropy loss function, and outputs the classification confidence;

[0040] Determine whether the classification confidence exceeds the confidence threshold;

[0041] When the classification confidence of the pseudo sample is lower than the confidence threshold, the pseudo sample is discarded;

[0042] When the classification confidence of the pseudo sample is not lower than the confidence threshold, it is added to the valid pseudo sample set.

[0043] Furthermore, it also includes: in the network architecture, the class centroid is loaded into the weight of the classification head, and kept fixed in the basic training stage and the incremental learning replay training stage, and the distance between the class centroids is maximized through POD Loss to optimize the separation of image categories; at the same time, the cosine distance between image features and class centroids is minimized through NaC Loss to ensure the aggregation of pseudo samples of the same image category in the feature space.

[0044] Based on the same inventive concept, the present invention provides a diffusion model-based incremental learning system for category images, which adopts the aforementioned incremental learning method for category images and includes:

[0045] The data processing module is used to obtain the training data set and divide it into the initial task data set and the incremental task data set;

[0046] The network architecture module is used to build the network architecture for incremental learning of categorical images. The network architecture includes the Unet backbone network, the PEDCC layer, the diffusion model, and the classification head.

[0047] The basic training module is used to use real images of all image categories in the initial task dataset as input data for the network architecture, extract image features of the real images through the Unet backbone network, and transmit them to the PEDCC layer. According to the parameters set by the PEDCC layer, the category discrimination of the image features during the forward diffusion and reverse denoising processes in the diffusion model is optimized, and pseudo samples of the real images are generated. Finally, the pseudo samples are classified using the classification head and the corresponding classification results are output. This also includes the process of screening out pseudo samples whose classification confidence exceeds the confidence threshold to form a valid pseudo sample set for the incremental learning replay training phase.

[0048] The incremental replay training module is used to take the real images of the Tth image category in the incremental task dataset as the dataset of the incremental task, and merge them with the pseudo samples in the valid pseudo sample set into a new task dataset, and ρ is a natural number greater than 0, and input them into the network architecture for incremental learning replay training, and output the classification results and corresponding pseudo samples corresponding to the new task dataset; which also includes screening out pseudo samples of real images whose classification confidence exceeds the classification confidence threshold, and adding them to the valid pseudo sample set; merging the valid pseudo sample set with the data set of the ρ+1th incremental task in the incremental task dataset into a new task dataset, and performing incremental learning replay training until all incremental tasks in the incremental task dataset are trained.

[0049] Furthermore, the basic training module includes,

[0050] The feature category optimization unit is used to input the initial task dataset into the network architecture, extract image features through the Unet backbone network, and generate the corresponding class centroid according to the parameters set in the PEDCC class center file in the PEDCC layer to optimize the class discrimination of image features in the feature space;

[0051] The diffusion training unit is used to add the image features optimized by the PEDCC layer as noise in the forward diffusion process of the diffusion model, convert it into a noise distribution, and restore the original image features through the reverse denoising process to generate the corresponding pseudo samples;

[0052] The classification prediction unit is used to input the generated pseudo samples into the classification head for classification and output the category prediction result of each pseudo sample;

[0053] The pseudo sample generation unit is used to calculate the classification confidence of the pseudo samples, screen the valid pseudo samples according to the set confidence threshold, and form a valid pseudo sample set for the incremental learning replay training phase.

[0054] Compared with the prior art, the present invention has at least one of the following technical effects:

[0055] (1) The present invention adopts a score-based diffusion model for incremental class training, which effectively solves the problems of class imbalance and catastrophic forgetting. At the same time, by generating pseudo-sample replays, it ensures that the number of generated samples is consistent with the number of real samples, avoiding the impact of class imbalance. The diffusion model can optimize the memory of old categories in incremental learning and reduce forgetting. At the same time, the classification network helps generate higher quality old class samples, thereby improving the classification performance of the model.

[0056] (2) The incremental learning method for category images based on the diffusion model provided by the present invention utilizes the combination of the diffusion model and the classification network, generates pseudo samples and performs incremental learning, and significantly improves the efficiency and performance of category image learning. This method effectively improves the category separation by optimizing the category discrimination of image features in the feature space, and avoids the catastrophic forgetting problem caused by forgetting old tasks in traditional incremental learning. In addition, the PEDCC layer and POD Loss loss function are used to optimize the inter-class distance, and combined with NaC Loss to maximize the intra-class compactness, the quality of pseudo samples is further improved, thereby enhancing the generalization ability and robustness of the model. Through the iterative incremental training process, it is possible to continuously learn new categories while maintaining a good memory of old categories without saving the original data, solving the problems of category imbalance and low quality of pseudo samples, thereby improving the performance of the system in practical applications. This method not only improves the classification accuracy in the incremental learning process, but also has strong adaptability and flexibility in multi-task training, and has high practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments:

[0058] Figure 1 It is a flowchart of the steps of the incremental learning method of category images based on the diffusion model of the present invention;

[0059] Figure 2 It is a flowchart of the incremental learning task training in the incremental learning method of category images based on the diffusion model of the present invention. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solution and advantages of the embodiments of this application clearer, the following will describe in detail a category image incremental learning method based on a diffusion model provided by the present invention in combination with the drawings in the embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and provides a detailed implementation method and specific operation process.

[0061] First embodiment

[0062] Regarding the practical application of catastrophic forgetting, especially in incremental learning scenarios, the following are some real-world examples and problems:

[0063] ① In intelligent security surveillance, cameras often need to detect and classify a constantly changing array of object categories. For example, initially, they only need to recognize "people" and "cars," but as the monitoring environment becomes more complex, they gradually need to recognize more object categories, such as "animals," "packages," and "abnormal behavior." Traditional image classification methods struggle to adapt to constantly changing environments because models tend to "forget" old object categories during training, leading to a loss of ability to recognize old categories as new ones are learned. Specifically, as new tasks, such as newly added object categories, are introduced, the network may overfit the features of the new task and neglect its memory of the old tasks, resulting in "catastrophic forgetting." For example, when "package" recognition is added as an incremental task, the model may lose its ability to recognize the "person" category.

[0064] ② In medical image classification, new disease categories are constantly being identified as medical technology advances. For example, an initial model might only detect "lung cancer" and "breast cancer," but as new medical research progresses, it needs to be able to identify diseases such as "stomach cancer" or "brain tumor." To this end, the model needs to perform incremental learning to adapt to new disease categories. Similarly, when adding new tasks, such as "stomach cancer" or "brain tumor," it is easy to forget previously learned tasks, such as "lung cancer" or "breast cancer." In this case, the model's accuracy decreases and its ability to recognize old disease categories is lost.

[0065] ③ In autonomous driving systems, object detection is a core task. As system applications continue to expand, new object categories (such as "bicycles," "pedestrians," and "trucks") are constantly being introduced to improve the vehicle's perception capabilities. In detecting new categories, existing detection models (such as "vehicles" and "traffic lights") must be retained without losing recognition capabilities. Therefore, in incremental learning, the addition of new categories can cause existing detection models to "forget" their ability to detect old categories, leading to misjudgments when the autonomous driving system encounters old objects. For example, when recognizing "trucks," the model may lose its ability to identify "traffic lights."

[0066] To address the above three situations, a diffusion model-based incremental learning method for categorical images is proposed. By using classification guidance within the Unet backbone network to generate images with similar labels, and then using the generated images of the same category as pseudo-image labels for incremental learning of the network, the network's incremental learning performance can be effectively improved. For classification, a classification head can be added to the feature extraction portion of the backbone network to simultaneously address both image generation and image classification, without relying on the independent classification discriminant network commonly found in existing technologies. First, the traditional Unet is enhanced to serve as the backbone network for the diffusion model. Channel normalization and the POD-Loss loss function can be used to further improve the classification performance of the classification network compared to the original Unet. The entire network is then trained to generate categorical images. Finally, the images generated by the trained model are used as pseudo-samples for incremental learning of the next category. This process is repeated until the network is trained on all categorical images. The present invention retains the advanced classification network and diffusion model while integrating image classification and generation functions, thereby realizing a category incremental learning network that completely retains no pseudo samples, eliminating the redundancy of the traditional simultaneous operation of two separate classification and generation networks, and experimentally verifying that the present invention is ahead of other existing public technical methods in the research on category image incremental learning.

[0067] The purpose of this invention is to address the shortcomings of existing technologies and provide a diffusion model-based incremental learning method for categorical images, which has both generation and classification capabilities. First, an incremental pattern generation model based on the currently mainstream diffusion model is developed, which has good generation capabilities. Specific implementation methods of this invention are as follows:

[0068] like Figure 1 As shown, the present invention provides a method for incremental learning of category images based on a diffusion model, comprising:

[0069] Step S1: Obtain a training dataset and divide it into an initial task dataset and an incremental task dataset;

[0070] Step S2: Construct a network architecture for incremental learning of category images. The network architecture includes a Unet backbone network, a PEDCC layer, a diffusion model, and a classification head.

[0071] It should be noted that the basic training network is constructed and the training process is set up as follows:

[0072] First, image category information is obtained from the initial task dataset. The PEDCC algorithm is used to generate a predefined class centroid file containing evenly distributed class centroids in the feature space. These class centroids are then loaded into the PEDCC layer for subsequent image feature optimization. In the feature extraction portion of the Unet backbone network, a PEDCC layer is added and its parameters are configured. The PEDCC layer parameters include the locations of the class centroids and parameters related to the diffusion model. These parameters are kept fixed during training to avoid perturbing the structure of the feature space. Furthermore, the optimizer and learning rate strategy are set. Training is performed using the AdamW optimizer, which combines adaptive learning rates and weight decay to stabilize the network learning process. A warm-up learning rate strategy is also implemented, slowly increasing the learning rate at the beginning of training to avoid gradient explosion caused by excessively high learning rates at the beginning of training. After the network is set up, the number of iterations is set and the forward denoising process begins. Training is then completed, and the training data is fed into the network architecture for the forward diffusion process. During this process, noise is gradually added to the image features, and the features are then restored using the diffusion model for reverse denoising, generating pseudo samples. In each training session, as the diffusion process iterates, the model gradually optimizes the generated pseudo samples and classifies them. Combined with the optimized loss function, the network architecture learns the relationship between image features and class centroids, thereby continuously improving pseudo sample generation and classification capabilities in an incremental learning process.

[0073] PEDCC is the abbreviation of Predefined Evenly-Distribution Class Centroids. Class centroids are used to represent the center points of different categories. The core goal of PEDCC technology is to generate multiple class centroids through uniform distribution in high-dimensional space, so that these class centroids are well spaced and evenly distributed in the feature space, thereby optimizing the category discrimination of image features.

[0074] PEDCC refers to a set of k points uniformly distributed on a unit hypersphere in n-dimensional space. Suppose we need to generate k uniformly distributed class centroids on a hypersphere in n-dimensional space. The specific steps are as follows:

[0075] 1. Input the number of centroids k and the spatial dimension n

[0076] 2. Get the class centroid u i ∈N(0,I),i=1,…,k

[0077] Get (n-k+2)-dimensional positive and negative basic points and

[0078] for j=1,...,k-2do

[0079] Get the (n-k+2+j)-dimensional positive basic point

[0080] Expand the obtained points by 1 dimension

[0081] Calculate the new point

[0082] Returns the centroid of the class U=(u1,…,u k )

[0083] in represents the centroid of the i-th class obtained in the j-th iteration, represents the centroid of the class after one-dimensional expansion in the jth iteration. Assuming that k uniformly distributed points (k <= n+1) need to be generated on the hypersphere of n-dimensional space, the first step directly generates the (n-k+2)-dimensional positive and negative basic points. In the second step, the two basic points obtained in the previous step are expanded by one dimension and assigned a value of 0, giving a positive basic point in (n-k+3) dimensions. The algorithm then calculates two new points on the (n-k+3)-dimensional hypersphere. At this point, there are three uniformly distributed points in the (n-k+3)-dimensional space. Repeat the second step until k uniformly distributed points are obtained on the hypersphere of n-dimensional space.

[0084] Step S3: In the basic training phase, real images of all image categories in the initial task dataset are used as input data for the network architecture. The image features of the real images are extracted through the Unet backbone network and transmitted to the PEDCC layer. The image features are optimized according to the parameters set by the PEDCC layer during the forward diffusion and reverse denoising processes in the diffusion model to generate pseudo samples of the real images. Finally, the pseudo samples are classified using the classification head and the corresponding classification results are output. This also includes the process of screening out pseudo samples whose classification confidence exceeds the confidence threshold to form a valid pseudo sample set for the incremental learning replay training phase.

[0085] Specifically, the goal of training the forward diffusion process is to distribute the data Converted to standard Gaussian distribution This process is implemented through stochastic differential equations (SDE), which are:

[0086] dx=f(x,t)dt+g (t) dw;

[0087] Where: x is the state of the diffusion process at time t; f(x,t) is the drift coefficient, which represents the deterministic change g of the data at time t (t) is the diffusion coefficient, which indicates the intensity of the noise; w is the standard Wiener process, i.e., Brownian motion.

[0088] In actual training, the loss function is often further simplified by the reparameterization technique: by replacing x t Directly use x0 and ∈ to represent; time t is randomly sampled from the uniform distribution U(0,T).

[0089] This facilitates the generation of high-quality data during reverse SDE sampling.

[0090] The reverse denoising process and the classifier are used to determine whether the classification confidence of the generated sample is greater than or equal to the set confidence threshold, including:

[0091] Given a label, guide the diffusion model to generate the next moment sample x towards a specific category t The reverse process starts from x t Generate x step by step t-1 , conditional generation introduces the category label y as a condition:

[0092] p θ (X t·1 |x t ,y)=N(X t·1 ;μ θ (x t ,t,y),Σ θ (x t ,t));

[0093] Where: μ θ (x t ,t,y) is the mean value predicted by the model, which depends on the current state x t , time step t, and condition y; Σ θ (x t ,t) is the covariance matrix of the noise (usually set to be fixed or time-dependent).

[0094] The classifier discriminates the convolution operation, with input x, output y, and convolution kernel f. Channel modulus normalization normalizes the output value of each channel to a fixed range through the unitization operation, thereby reducing the deviation introduced by scale differences. The output after unitization mainly depends on the relative angle cosθ between the input and the convolution kernel. i , rather than its absolute size.

[0095] y i = <x·f i >=||x||·||f i ||·cosθ i ;

[0096]

[0097] Loss function: Cross-Entropy Loss is a loss function widely used in classification tasks. The formula is as follows, where y ij is the predicted probability value of the i-th sample in the j-th category, is the one-hot label of the i-th sample in the j-th category. However, the softmax loss function adopts an inter-class competition mechanism, which only cares about the accuracy of the predicted probability of the correct label, ignores the difference of incorrect labels, and cannot guarantee the compactness within the class and the alienation between classes. At the same time, there are fewer constraints after the feature extraction layer. From the perspective of features, we hope to obtain features with the largest inter-class separation and the best intra-class compactness. It is a better method to constrain the features directly and uniquely. The POD Loss loss function has these characteristics. On the one hand, it uses a predefined uniformly distributed class centroid PEDCC and loads it into the weight of the classification layer in the network. At the same time, the weight is fixed during the training phase to maximize the inter-class distance. On the other hand, the cosine distance between the sample features and PEDCC is limited by the norm-adaptive cosine loss (NaC Loss) to minimize the intra-class distance. Among them, N represents the number of samples, cosNaC is the number of samples, and N is the number of samples. is the improved norm-adaptive cosine similarity.

[0098]

[0099] The generation process of applying the classifier to the diffusion model is as follows: In the generation process of the score-based diffusion model, a number of points are sampled from the random Gaussian space as the sample x at time t. t ,As the reverse denoising process continues, the statistical features of the noise at different moments are calculated, and x-1 that is closer to the real sample is generated, and then discriminated by the classifier to obtain effective pseudo samples with the same label.

[0100] This application adopts a class-incremental mode to train a score-based diffusion model and includes a classifier, and uses the generated pseudo samples to replay the memory to keep the number of generated samples consistent with the real samples, thereby avoiding the problem of class imbalance.

[0101] This method uses a diffusion model as a generative model. The diffusion model is based on score generative modeling and learns incrementally to reduce forgetting during training. At the same time, the classification network is used to guide the diffusion model to generate higher-quality old class samples for playback, thereby remembering the information of old categories, which can effectively alleviate catastrophic forgetting and improve classification performance.

[0102] Step S4: In the incremental learning replay training phase, based on the trained network architecture, the real images of the first image category in the incremental task dataset are used as the dataset of the incremental task, and are merged with the pseudo samples in the valid pseudo sample set to form a new task dataset, and ρ is a natural number greater than 0. The new task dataset is input into the network architecture for incremental learning replay training, and the classification results and corresponding pseudo samples corresponding to the new task dataset are output; wherein, the pseudo samples of the real images whose classification confidence exceeds the classification confidence threshold are screened out and added to the valid pseudo sample set;

[0103] Step S5: Merge the valid pseudo sample set and the data set of the ρ+1th incremental task in the incremental task data set into a new task data set, where ρ is a natural number greater than 0, and perform incremental learning replay training until all incremental tasks in the incremental task data set are trained.

[0104] Specifically, before sampling pseudo samples for inference, the categories and number of pseudo samples required are set. Once this is done, random sampling is performed in Gaussian space, fed into the network architecture, and the statistical characteristics of the noise at each moment are calculated. This gradually completes the reverse denoising process, yielding a number of generated samples. Before denoising, a confidence threshold should be set for the generated samples. If the classification confidence of the generated samples is greater than or equal to the confidence threshold, the generated samples are of high quality; if it is less than the confidence threshold, the generated samples are of low quality and should be discarded. Finally, the same number of generated samples are used together with the dataset from the ρ+1 class to train the next task, achieving a generative replay effect until all tasks are trained.

[0105] Furthermore, step S1 includes:

[0106] Step S11: obtaining a training dataset containing multiple image categories, where the image categories include digital categories and general object categories;

[0107] Step S12: Divide the training dataset into an initial task dataset and an incremental task dataset according to a preset ratio, wherein each image category in the incremental task dataset is used as an incremental task for incremental learning playback training;

[0108] Step S13: Perform image preprocessing on the real images in the initial task dataset and the incremental task dataset. The image preprocessing includes converting the real images into tensor format and performing image enhancement and pixel normalization.

[0109] It should be specifically explained that, taking the CIFAR10 dataset as an example, the dataset contains M (M=10) categories. Half of the categories are used as the initial tasks, and then the other half of the categories are divided into T (T=5) tasks, where each task contains M / (2T) categories, for a total of T+1 tasks. Prepare experimental data, convert images into tensors, and perform image enhancement and normalization.

[0110] Furthermore, step S3 includes:

[0111] Step S31: Input the initial task dataset into the network architecture, extract image features through the Unet backbone network, and generate corresponding class centroids according to the parameters set in the PEDCC class center file in the PEDCC layer to optimize the class discrimination of image features in the feature space;

[0112] Step S32: Based on the image features optimized by the PEDCC layer, in the forward diffusion process of the diffusion model, the image features are forward diffused and added as noise, converted into a noise distribution, and the original image features are restored through the reverse denoising process to generate corresponding pseudo samples;

[0113] Step S33: input the generated pseudo samples into the classification head for classification, and output the category prediction result of each pseudo sample;

[0114] Step S34: Calculate the classification confidence of the pseudo samples, screen valid pseudo samples according to the set confidence threshold, and form a valid pseudo sample set for the incremental learning replay training phase.

[0115] Furthermore, step S31 includes:

[0116] Input the initial task dataset containing multiple categories of images into the network architecture, and extract the image features of the real images in the category images through the Unet backbone network;

[0117] In the PEDCC layer, the class centroids are uniformly generated in the feature space according to the PEDCC algorithm, and the feature characteristics are optimized so that the image features of different categories are reasonably distributed in the feature space and the class discrimination is improved;

[0118] The image features optimized by the PEDCC layer are output to the diffusion model to generate pseudo samples.

[0119] Furthermore, in step S32, in the forward diffusion process of the diffusion model, the image features are forward diffused and added as noise to convert into a noise distribution, including:

[0120] Get the image feature x0 and use it as the initial state input in the diffusion model;

[0121] The time step t randomly sampled from the uniform distribution U(0,T) is used to determine the intensity of the noise added during the forward diffusion process;

[0122] Update the image features according to the diffusion equation dx = f(x, t)dt + g(t)dw, and pass the drift coefficient f(x, t) and diffusion coefficient g (t) Adjust the image feature changes, x is the state at time step t during the forward diffusion process, and w is the standard Wiener process;

[0123] By formula Add noise ∈ to the original image feature x0 and convert it into a noise distribution image x t and through α t Control weight impact;

[0124] By minimizing the loss function The value of optimizes the forward diffusion process, where s θ (x t ,t) is the model output, Part of the noise.

[0125] Furthermore, in step S32, the original image features are restored through an inverse denoising process to generate corresponding pseudo samples, including:

[0126] Sample from the noise distribution image to obtain the initial noise sample image x t , used to perform the reverse process through the diffusion model to remove noise and restore image features;

[0127] The label y of the image category corresponding to the real image in the initial task dataset is used as a conditional input to guide the generation of corresponding pseudo samples in the reverse denoising process;

[0128] According to the reverse denoising formula p θ (X t-1 |x t ,y)=N(X t·1 ;μ θ (x t ,t,y),Σ θ (x t ,t)), calculate the image x t Restore to state X t-1 process, where μ θ (x t ,t,y) is the predicted mean, Σ θ (x t ,t) is the covariance matrix of the noise, N is the number of samples in the data set;

[0129] The generated image X t-1 Pseudo samples as output.

[0130] Furthermore, in step S3, pseudo samples whose classification confidence exceeds the confidence threshold are screened out to form a valid pseudo sample set, including:

[0131] The classification head maps the pseudo samples to the category label space for classification, calculates the classification loss of each pseudo sample through the cross entropy loss function, and outputs the classification confidence;

[0132] Determine whether the classification confidence exceeds the confidence threshold;

[0133] When the classification confidence of the pseudo sample is lower than the confidence threshold, the pseudo sample is discarded;

[0134] When the classification confidence of the pseudo sample is not lower than the confidence threshold, it is added to the valid pseudo sample set.

[0135] Furthermore, it also includes: in the network architecture, the class centroid is loaded into the weight of the classification head, and kept fixed in the basic training stage and the incremental learning replay training stage, and the distance between the class centroids is maximized through POD Loss to optimize the separation of image categories; at the same time, the cosine distance between image features and class centroids is minimized through NaC Loss to ensure the aggregation of pseudo samples of the same image category in the feature space.

[0136] It should be specifically noted that the neural architecture of this application uses Unet as the backbone network of the diffusion model. The feature extraction part and channel modulus normalization operation can be used, and the softmax loss function is replaced by POD Loss to improve the classification performance of the classification network.

[0137] The following is a specific practical process. The example datasets come from MNIST and CIFAR10. The datasets need to be divided. Specifically, for a dataset containing M (M=10) categories, half of the categories are used as initial tasks. Then, the other half of the categories are divided into T (T=5) tasks, where each task contains M / (2T) categories, for a total of T+1 tasks.

[0138] 1. The experimental environment is built with PyTorch and trained on a single NVIDIA GeForce RTX 3090 GPU running on Linux CentOS 7.9 system.

[0139] 2. Generate a predefined PEDCC center point file and train the set deep neural network on two datasets (taking CIFAR10 as an example) for the initial task to obtain the basic network.

[0140] 3. According to Figure 2Process, input the first T categories of real samples, use the basic network to generate the first T categories of pseudo sample images, and obtain the confidence corresponding to the pseudo sample images through cross entropy.

[0141] 4. Set a confidence threshold y0. Pseudo sample images above the threshold y0 are regarded as original true samples, and pseudo sample images below the threshold y0 are discarded, while keeping the number of samples no less than the original samples.

[0142] 5. Use the first T pseudo samples and the real samples of the T+1 class as the data set for the next task, and perform the next stage of category incremental learning.

[0143] 6. Increase T by 1 and repeat steps 3 to 5 until all categories of samples are trained. Then, obtain the incremental learning performance of the final model by testing the test set samples.

[0144] Second embodiment

[0145] Based on the same inventive concept, the present invention also provides a diffusion model-based incremental learning system for category images, which adopts the aforementioned incremental learning method for category images and includes:

[0146] The data processing module is used to obtain the training data set and divide it into the initial task data set and the incremental task data set;

[0147] The network architecture module is used to build the network architecture for incremental learning of categorical images. The network architecture includes the Unet backbone network, the PEDCC layer, the diffusion model, and the classification head.

[0148] The basic training module is used to use real images of all image categories in the initial task dataset as input data for the network architecture, extract image features of the real images through the Unet backbone network, and transmit them to the PEDCC layer. According to the parameters set by the PEDCC layer, the category discrimination of the image features during the forward diffusion and reverse denoising processes in the diffusion model is optimized, and pseudo samples of the real images are generated. Finally, the pseudo samples are classified using the classification head and the corresponding classification results are output. This also includes the process of screening out pseudo samples whose classification confidence exceeds the confidence threshold to form a valid pseudo sample set for the incremental learning replay training phase.

[0149] The incremental replay training module is used to take the real images of the first image category in the incremental task data set as the data set of the incremental task, and merge them with the pseudo samples in the valid pseudo sample set into a new task data set, and ρ is a natural number greater than 0, and input them into the network architecture for incremental learning replay training, and output the classification results and corresponding pseudo samples corresponding to the new task data set; which also includes screening out pseudo samples of real images whose classification confidence exceeds the classification confidence threshold, and adding them to the valid pseudo sample set; merging the valid pseudo sample set with the data set of the ρ+1th incremental task in the incremental task data set into a new task data set, and ρ is a natural number greater than 0, and performing incremental learning replay training until all incremental tasks in the incremental task data set are trained.

[0150] Furthermore, the basic training module includes,

[0151] The feature category optimization unit is used to input the initial task dataset into the network architecture, extract image features through the Unet backbone network, and generate the corresponding class centroid according to the parameters set in the PEDCC class center file in the PEDCC layer to optimize the class discrimination of image features in the feature space;

[0152] The diffusion training unit is used to add the image features optimized by the PEDCC layer as noise in the forward diffusion process of the diffusion model, convert it into a noise distribution, and restore the original image features through the reverse denoising process to generate the corresponding pseudo samples;

[0153] The classification prediction unit is used to input the generated pseudo samples into the classification head for classification and output the category prediction result of each pseudo sample;

[0154] The pseudo sample generation unit is used to calculate the classification confidence of the pseudo samples, screen the valid pseudo samples according to the set confidence threshold, and form a valid pseudo sample set for the incremental learning replay training phase.

[0155] Although the present invention has been disclosed above in terms of preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art may make possible changes and modifications to the technical solutions of the present invention by using the methods and technical contents disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall fall within the scope of protection of the technical solutions of the present invention.

Claims

1. A method for incremental learning of category images based on a diffusion model, characterized in that: include: Step S1: Obtain a training dataset and divide it into an initial task dataset and an incremental task dataset; Step S2: constructing a network architecture for incremental learning of category images, wherein the network architecture includes a Unet backbone network, a PEDCC layer, a diffusion model, and a classification head; Step S3: In the basic training phase, the real images of all image categories in the initial task data set are used as input data of the network architecture, the image features of the real images are extracted through the Unet backbone network and transmitted to the PEDCC layer, the category discrimination of the image features during the forward diffusion and reverse denoising processes in the diffusion model is optimized according to the parameters set by the PEDCC layer, pseudo samples of the real images are generated, and finally the pseudo samples are classified using the classification head, and the corresponding classification results are output; wherein, the process also includes screening out the pseudo samples whose classification confidence exceeds the confidence threshold to form a valid pseudo sample set for the incremental learning replay training phase; Step S4: In the incremental learning replay training phase, based on the trained network architecture, the real image of the T-th image category in the incremental task data set is used as the data set of the incremental task, and is merged with the pseudo samples in the valid pseudo sample set to form a new task data set, and ρ is a natural number greater than 0, and is input into the network architecture for incremental learning replay training, and the classification result corresponding to the new task data set and the corresponding pseudo samples are output; wherein, the method further includes screening out the pseudo samples of the real images whose classification confidence exceeds the classification confidence threshold, and adding them to the valid pseudo sample set; Step S5: Merge the valid pseudo sample set and the data set of the ρ+1th incremental task in the incremental task data set into the new task data set, and perform the incremental learning replay training until all the incremental tasks in the incremental task data set are trained.

2. The incremental learning method for category images according to claim 1, characterized in that Step S1 includes: Step S11: obtaining the training data set containing a plurality of image categories, wherein the image categories include digital categories and general object categories; Step S12: dividing the training dataset into the initial task dataset and the incremental task dataset according to a preset ratio, wherein each image category in the incremental task dataset serves as an incremental task for the incremental learning playback training; Step S13: performing image preprocessing on the real images in the initial task dataset and the incremental task dataset, wherein the image preprocessing includes converting the real images into a tensor format, and performing image enhancement and pixel normalization.

3. The incremental learning method for category images according to claim 2, characterized in that Step S3 includes: Step S31: Input the initial task dataset into the network architecture, extract the image features through the Unet backbone network, and generate corresponding class centroids according to the parameters set in the PEDCC class center file in the PEDCC layer to optimize the class discrimination of the image features in the feature space; Step S32: Based on the image features optimized by the PEDCC layer, in the forward diffusion process of the diffusion model, the image features are forward diffused and added as noise to convert them into noise distribution, and the original image features are restored through a reverse denoising process to generate the corresponding pseudo samples; Step S33: inputting the generated pseudo samples into the classification head for classification, and outputting the category prediction result of each pseudo sample; Step S34: Calculate the classification confidence of the pseudo samples, screen the valid pseudo samples according to the set confidence threshold, and form the valid pseudo sample set for the incremental learning replay training phase.

4. The incremental learning method for category images according to claim 3, characterized in that The step S31 includes: Inputting the initial task dataset containing multiple category images into the network architecture, and extracting the image features of the real images in the category images through the Unet backbone network; In the PEDCC layer, the class centroids are uniformly generated in the feature space according to the PEDCC algorithm, and the feature features are optimized so that the image features of different categories are reasonably distributed in the feature space, thereby improving the class distinction; The image features optimized by the PEDCC layer are output to the diffusion model to generate the pseudo samples.

5. The incremental learning method for category images according to claim 4, characterized in that: In step S32, in the forward diffusion process of the diffusion model, the image features are forward diffused and added as noise to convert into a noise distribution, including: Obtaining the image feature x0 and using it as an initial state input in the diffusion model; A time step t randomly sampled from a uniform distribution U(0,T) is used to determine the intensity of the noise added during the forward diffusion process; According to the diffusion equation dx=f(x,t)dt+g (t) dw updates the image features and uses the drift coefficient f(x,t) and diffusion coefficient g (t) Adjust the image feature change, x is the state of the forward diffusion process at the time step t, and w is the standard Wiener process; By formula Add noise ∈ to the original image feature x0 and convert it into the image x of the noise distribution t and through α t Control weight impact; By minimizing the loss function The value of optimizes the forward diffusion process, where s θ (x t ,t) is the model output, Part of the noise.

6. The incremental learning method for category images according to claim 5, characterized in that: In step S32, the original image features are restored through an inverse denoising process to generate the corresponding pseudo samples, including: Sampling from the image of the noise distribution, obtaining the image x of the initial noise sample t , used to perform a reverse process through the diffusion model to remove the noise and restore the image features; The label y of the image category corresponding to the real image in the initial task dataset is used as a conditional input to guide the generation of the corresponding pseudo sample in the reverse denoising process; According to the reverse denoising formula p θ (X t-1 |x t ,y)=N(X t·1 ;μ θ (x t ,t,y),Σ θ (x t ,t)), calculate the image x t Restore to state X t-1 process, in which the μ θ (x t ,t,y) is the predicted mean, Σ θ (x t , t) is the covariance matrix of the noise, and N is the number of samples in the data set; The generated image X t-1 Pseudo samples as output.

7. The incremental learning method for category images according to claim 6, characterized in that: In step S3, the pseudo samples whose classification confidence exceeds the confidence threshold are screened out to form a valid pseudo sample set, including: The classification head maps the pseudo samples to the category label space for classification, calculates the classification loss of each pseudo sample through a cross entropy loss function, and outputs the classification confidence; Determining whether the classification confidence exceeds the confidence threshold; When the classification confidence of the pseudo sample is lower than the confidence threshold, discard the pseudo sample; When the classification confidence of the pseudo sample is not lower than the confidence threshold, the pseudo sample is added to the valid pseudo sample set.

8. The incremental learning method for category images according to claim 7, characterized in that: Also includes: In the network architecture, the class centroids are loaded into the weights of the classification head and kept fixed in the basic training phase and the incremental learning replay training phase. The distance between the class centroids is maximized by POD Loss to optimize the separation of the image categories. At the same time, NaC Loss is used to minimize the cosine distance between the image features and the class centroids to ensure the aggregation of the pseudo samples of the same image category in the feature space.

9. A diffusion model-based incremental learning system for category images, using the incremental learning method for category images according to any one of claims 1 to 8, characterized in that: include: The data processing module is used to obtain the training data set and divide it into the initial task data set and the incremental task data set; A network architecture module is used to build a network architecture for incremental learning of categorical images, which includes a Unet backbone network, a PEDCC layer, a diffusion model, and a classification head; A basic training module is configured to use real images of all image categories in the initial task dataset as input data for the network architecture, extract image features of the real images through the Unet backbone network, and transmit them to the PEDCC layer. The module optimizes the category discrimination of the image features during forward diffusion and reverse denoising in the diffusion model according to the parameters set by the PEDCC layer, generates pseudo samples of the real images, and finally classifies the pseudo samples using the classification head to output corresponding classification results. The module also includes screening out pseudo samples whose classification confidence exceeds a confidence threshold to form a valid pseudo sample set for use in the incremental learning replay training phase. An incremental replay training module is used to use the real image of the Tth image category in the incremental task data set as the data set of the incremental task, and merge it with the pseudo samples in the valid pseudo sample set into a new task data set, and ρ is a natural number greater than 0, input it into the network architecture for incremental learning replay training, and output the classification results corresponding to the new task data set and the corresponding pseudo samples; wherein, it also includes screening out the pseudo samples of the real images whose classification confidence exceeds the classification confidence threshold, and adding them to the valid pseudo sample set; merging the valid pseudo sample set with the data set of the ρ+1th incremental task in the incremental task data set into the new task data set, and performing the incremental learning replay training until all the incremental tasks in the incremental task data set are trained.

10. The incremental learning system for category images according to claim 9, characterized in that: The basic training module includes: a feature class optimization unit, configured to input the initial task dataset into the network architecture, extract the image features through the Unet backbone network, and generate corresponding class centroids based on the parameters set in the PEDCC class center file in the PEDCC layer, so as to optimize the class discrimination of the image features in the feature space; a diffusion training unit configured to add the image features optimized by the PEDCC layer as noise in the forward diffusion process of the diffusion model, convert the image features into a noise distribution, and restore the original image features through a reverse denoising process to generate the corresponding pseudo samples; a classification prediction unit, configured to input the generated pseudo samples into the classification head for classification, and output the category prediction result of each pseudo sample; The pseudo sample generation unit is used to calculate the classification confidence of the pseudo sample, screen the valid pseudo samples according to the set confidence threshold, and form the valid pseudo sample set for the incremental learning replay training phase.