Diffusion noise enhancement method for long-tail remote sensing image

Through the diffusion model, the foreground and background of remote sensing images are generated and mixed, and combined with the comparison learning method, the problem of inconsistent distribution of generated images and real images in the long-tail distribution of remote sensing images is solved, and the recognition ability of tail categories and the classification accuracy of the model are improved.

CN120339746APending Publication Date: 2025-07-18NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510387862.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When the prior art processes the long-tail distribution of remote sensing images, there is a problem of inconsistency between the generated image and the real image, resulting in poor tail category recognition capabilities, and traditional data enhancement methods are difficult to significantly improve the diversity and quality of tail samples.

Method used

A diffusion model is used to generate new samples, and the foreground of the original image is mixed with the background of the generated image through the DiffCam-Mix module. The high-quality generated samples are screened in combination with the CLIP model. The SimSiam comparison learning is used to calibrate the sample distribution in the feature space to improve the consistency between the generated samples and the original samples.

Benefits of technology

Effectively expand the number of samples in tail category, improve the model's ability to identify tail categories, and improve the overall performance and generalization ability of remote sensing image classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339746A_ABST
    Figure CN120339746A_ABST
Patent Text Reader

Abstract

The invention provides a diffusion noise enhancement method for a long-tail remote sensing image. According to the method, before training, a specially designed condition prompt set and an original training set are used for jointly guiding a diffusion model to generate a new sample, and a CLIP model is used for screening. In addition, the invention provides two effective generation data utilization strategies, and optimization is respectively carried out at the end of the first-stage training and in the second-stage training process. Firstly, a DiffCam-Mix module is designed, the module uses a Grad-CAM + + class activation mapping method to extract a background of a generated image and a foreground of an original image, the background and the foreground are mixed, and mixed data with real features of the original image and diversity of the generated image are constructed. And secondly, in a training process, introducing a comparative learning method based on cosine similarity, so that the mixed data and the corresponding original data are kept consistent in a feature space, thereby further calibrating data distribution. The method is suitable for a long-tail distribution scene in a remote sensing image classification task, the number of tail category samples can be effectively increased, the classification precision of the model is improved, and the recognition capability of long-tail data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image recognition technology in deep learning. Specifically for the long-tail remote sensing image classification task, it particularly relates to a data augmentation method based on diffusion noise to optimize the long-tail data distribution and improve the classification performance of the model. Background Art

[0002] Remote sensing image classification is mainly used to identify and classify land cover types in images obtained from satellite or aerial photography, which is of great significance in fields such as environmental monitoring, urban planning, and agriculture. In recent years, deep convolutional neural networks have been widely applied to remote sensing image classification tasks. These networks can automatically learn and extract complex features of remote sensing images, thus greatly improving the performance of image classification tasks. Although traditional convolutional neural networks have shown great advantages in remote sensing image classification, the distribution of remote sensing images in the real world often shows imbalance, that is, the long-tail distribution of data: a small number of classes have a large number of samples, namely the head classes; while the vast majority of classes have a small number of samples, namely the tail classes. In this case, traditional deep convolutional neural networks may learn representations biased towards the head classes, resulting in poor recognition of the tail classes.

[0003] Data augmentation is an effective method to alleviate the long-tail problem. By increasing the number of samples and sample diversity of the tail classes, it improves the data distribution, enabling the model to learn the features of each class more evenly, thereby improving the classification performance. Traditional data augmentation methods mainly expand the dataset by applying geometric transformations, color adjustments, noise addition, etc. to the original images. Although these methods can increase sample diversity to a certain extent, the additional information they provide mainly focuses on the surface changes of image features and lacks the expansion of deeper information. In addition, when dealing with extreme long-tail distributions, these methods are difficult to significantly improve the diversity and quality of tail class samples. As an outstanding generative model, diffusion models have performed well in image generation tasks in recent years. Currently, some scholars have tried to directly use the images generated by diffusion models for data augmentation. However, when dealing with remote sensing data, the effects of these methods are relatively limited, mainly because there is a problem of inconsistent distribution between the generated images and the real images.

[0004] From the perspective of data augmentation, the present invention aims to utilize the rich external knowledge of the diffusion model to expand the diversity of the original dataset. Considering the distribution shift between the generated samples and the original samples, we not only use the new samples generated by the diffusion model but also combine these new samples with the original samples to improve their generalization ability. In addition, we reduce the distance between the mixed samples and their corresponding original samples in the feature space to further calibrate the distribution of different samples. The present invention can be applied to the long-tail distribution scenario of remote sensing datasets, effectively expanding the number of samples in the tail category, improving the classification accuracy of the model, and enhancing the recognition ability of long-tail data. Summary of the Invention

[0005] Object of the Invention: The present invention aims to optimize the distribution characteristics of long-tail remote sensing data by introducing a data augmentation method based on diffusion noise, improve the recognition ability of the model for tail classes, and thus enhance the overall performance and generalization ability of remote sensing image classification.

[0006] The technical solution is as follows:

[0007] 1. A diffusion noise enhancement method for long-tail remote sensing images, the method comprising:

[0008] S1. Input the original remote sensing dataset I with a long-tail distribution o ;

[0009] S2. Define the conditional prompt set P, input the original dataset I o and the prompt set P into the data generation module, and obtain the generated image dataset I' after CLIP filtering g ;

[0010] S3. Perform class-balanced sampling on the original dataset I o and use the cross-entropy loss function for the first-stage training of the model;

[0011] S4. Input the original dataset I o and the generated image dataset I' g into the DiffCam-Mix module, use the model trained in the first stage to calculate the foreground masks and background masks of the images in I o and I' g respectively, and use the masks to mix the foreground of the original image with the background of the generated image to obtain the mixed sample dataset I m ;

[0012] S5. Use the original dataset I o and the mixed sample dataset I m as inputs for the second-stage training of the model.

[0013] Furthermore, for the generated image dataset I' in step S2 g, the present invention uses the pre-trained diffusion model InstructPix2Pix as the generation model and designs a prompt set P = {Snowy, Sunset, Autumn, Rainy, Winter, Sunny, Summer, Cloudy, Spring, Foggy} to guide the diffusion model to generate images of different styles. For

[0014] each input image x i ∈ I o , the diffusion model generates a generated image x' of the corresponding style according to the prompt P j ∈ P. x ij is the original image corresponding to x' i , and their labels are kept consistent. Repeat the above generation process to finally obtain the generated dataset I ij . Next, use the pre-trained CLIP model to filter the generated image dataset I g . Define the template T g of class y as "A photo of the class [y]", where [y] represents the name of class y. Using the image encoder and text encoder of CLIP, y calculate the image embedding of the generated image x'

[0015] ij and the text embedding of its corresponding class template T y y , and then calculate the cosine similarity score between them

[0016]

[0017] For class y, select K-N y samples with the highest similarity scores (where K is the maximum number of samples for each class, and N y is the number of original samples of class y), filter the remaining samples, and finally obtain the high-quality generated dataset I' g .

[0018] To accurately extract the foreground part of the original image and the background part of the generated image, the present invention divides the training of the model into two stages. The model used here is a deep convolutional neural network, including ResNet-32, ResNet-50, etc. The first stage is the first n epochs, using the original data to train the feature extractor and classifier through a class-balanced sampler. The loss function used here is the cross-entropy loss function L o , and then use the obtained network for the class activation map calculation in step 4.

[0019] For the DiffCam-Mix module in step S4, the present invention uses Grad-CAM++ to extract the key parts (foreground) of the original image and mixes them with the non-key parts (background) of the generated image. For each image, through forward and backward propagation, the gradient of the target class c with respect to the feature map is obtained, denoted as where Y c is the output score of the target class c, and A k is the activation feature map of the target layer. For each activation map A k , the gradient weight is calculated as follows:

[0020]

[0021] where (m, n) and (a, b) are the same iterators on the entire activation map A k . and represent the second-order partial derivative and the third-order partial derivative respectively. is multiplied by the gradient after Relu activation, and the sum over all positions (m, n) is taken to obtain the final weight

[0022]

[0023] The calculated weight is multiplied by the feature map and summed, and then the class activation map CAM is obtained through the ReLU layer:

[0024]

[0025] Normalization and resizing are applied for further processing to match the size of the input image to obtain the final class activation map CAM resized . According to the following formula:

[0026] M = (1 - CAM resized ) 2 ,

[0027] the background mask of each generated image can be obtained, where M ∈ {0, 1} W×H . The mixing operation is defined as:

[0028]

[0029] where x i (x i ∈ I o ) is the original image sample, and x′ i (x′ i ∈ I′ g ) is the corresponding generated image sample. is based on the above

[0030] the above-mentioned x i and x' i to generate the mixed data. λ is a hyperparameter ranging from 0 to 1, which is used to control the mixing ratio of the foreground and the background, and ⊙ is the element-wise multiplication. Since this operation is performed between two samples belonging to the same class, the label of the mixed data will not change. The mixed sample dataset I obtained by this module m and the original sample dataset I o are together used for the second-stage training of the model in step 5.

[0031] For step S5, use I m and I o to fine-tune the model after the first-stage training in step S3 is completed. In this stage, SimSiam contrastive learning is used. The original sample x i and the corresponding mixed sample x' i of these two views are regarded as a positive sample pair, which are processed by an encoder network composed of a feature extractor (from the first stage) and a projection head. The encoder shares weights between these two views. In addition, a prediction head is included, which is used to transform the output of one view and match it with the other view. Both the projection head and the prediction head are composed of MLP, and the prediction head is located after the projection head. During the second-stage training, only the predictor branch can be updated. After extracting features from the positive sample pair, they are fed into the projection head to obtain and and then input into the prediction head to obtain and The goal of this step is to minimize the negative cosine similarity between the positive sample pairs, which can be expressed as:

[0032]

[0033] The final loss function of the second-stage training process is:

[0034]

[0035] where β is a hyperparameter used to control the contrastive learning loss. L o and L m correspond to the cross-entropy loss functions for the original data and the mixed data training respectively.

[0036] Advantages of the present invention: The present invention provides a diffusion noise enhancement method for long-tailed remote sensing images, which can effectively alleviate the data imbalance problem in remote sensing image classification and improve the recognition ability of the model for tail class samples. Compared with traditional data enhancement methods (such as Mixup, CutMix, etc.), the present invention uses a diffusion model to generate more diverse samples, expands the feature space of tail classes, and thus enhances the generalization ability of the model. In addition, by generating mixed samples through the DiffCam-Mix module and calibrating the mixed samples through contrastive learning, the present invention can ensure the consistency between the generated samples and the original samples in the feature space. Brief Description of the Drawings

[0037] Figure 1 is the flowchart of the implementation of the method of the present invention;

[0038] Figure 2 is the construction process of the mixed sample;

[0039] Figure 3 is the contrastive learning process in the two-stage training. Detailed Embodiments

[0040] The present invention provides a data enhancement method based on diffusion noise to solve the long-tailed remote sensing data classification problem. The specific implementation steps of the present invention are further described below with reference to the accompanying drawings of the specification. As Figure 1 shown, a diffusion noise enhancement method for long-tailed remote sensing images includes the following steps:

[0041] Step 1: Input the original remote sensing data set I with a long-tailed distribution o ;

[0042] Step 2: Define the conditional prompt set P. The original data set I o and the prompt set P are input into the data generation module. After being filtered by CLIP, the generated image data set I' g is obtained;

[0043] Step 3: Perform class-balanced sampling on the original data set I o and use the cross-entropy loss function for the first-stage training of the model;

[0044] Step 4: Input the original data set I o and the generated image data set I' g into the DiffCam-Mix module. Use the model trained in the first stage to calculate the foreground masks and background masks of the images in I o and I' g respectively, and use the masks to mix the foreground of the original image with the background of the generated image to obtain the mixed sample data set I m ;

[0045] Step 5: Original dataset I o and Hybrid sample dataset I m are used as inputs for the second-stage training of the model.

[0046] The long-tailed remote sensing datasets input in Step 1 include SIRI-WHU-LT, PatternNet-LT, RSI-CB256-LT, etc. For the above datasets, use the formula to construct them into long-tailed datasets, where y is the class index, N y is the number of samples contained in class y, N max is the number of samples in the class with the largest number of samples, μ is the imbalance factor, and C is the total number of classes. As the class index increases, the corresponding number of samples decreases. In Step 2, the pre-trained diffusion model InstructPix2Pix is used as the generation model, and a prompt set P = {Snowy, Sunset, Autumn, Rainy, Winter, Sunny, Summer, Cloudy, Spring, Foggy} is designed to guide the diffusion model to generate images of different styles. For each input image x i ∈ I o , the diffusion model generates a generated image x j of the corresponding style according to the prompt P ij . Keep the label of x ij consistent with the corresponding x i . Repeat the above generation process to finally obtain the generated dataset I g . Next, use the pre-trained CLIP model to filter I g . Define the template T y of class y as "A photo of the class [y]", where [y] represents the name of class y. Using the image encoder and text encoder of CLIP, calculate the image embedding of x′ ij and the text embedding of its corresponding class template T y , and then calculate the cosine similarity score between them

[0047]

[0048] For class y, select K - N y samples with the highest similarity scores (here K is the maximum number of samples for each class, and N y is the original number of samples of class y), filter the remaining samples, and finally obtain the high-quality generated dataset I′ g .

[0049] Step 3 uses the original data after class-balanced sampling and the cross-entropy loss function for the first-stage training of the model. Step 4 For each image, calculate the gradient of the target class c with respect to the feature map, denoted as where Y c is the output score of the target class c, and A k is the activation feature map of the target layer. For each activation map A k , calculate its gradient weight using the following formula

[0050]

[0051] where (m,n) and (a,b) are the same iterators on the entire activation map A k . and represent the second-order partial derivative and the third-order partial derivative respectively. Then obtain the final weight through the following formula

[0052]

[0053] Multiply the calculated weight by the feature map and sum, and then obtain the class activation map CAM through the ReLU layer:

[0054]

[0055] Apply normalization and size adjustment for further processing to match the size of the input image to obtain the final class activation map CAM resized . According to the following formula:

[0056] M = (1 - CAM resized ) 2 ,

[0057] the background mask of each generated image can be obtained, where M ∈ {0,1} W×H . Figure 2 Shows the construction process of the mixed samples, and defines the mixing operation as:

[0058]

[0059] where x i (x i ∈ I o ) is the original image sample, and x′ i (x′ i ∈ i′ g ) is the corresponding generated image sample. is based on the above x i and x′ iThe generated mixed data. λ is a hyperparameter ranging from 0 to 1, which is used to control the mixing ratio of the foreground and the background and is set to 0.5 in the present invention. ⊙ is the element-wise multiplication.

[0060] Step 5 uses SimSiam contrastive learning, as Figure 3 shown, regarding the original sample x i and the corresponding mixed sample x i of these two views as a positive sample pair. After extracting features from the positive sample pair, they are fed into the projection head to obtain and respectively, and then fed into the prediction head to obtain and respectively. The goal of this step is to minimize the negative cosine similarity between the positive sample pairs, that is:

[0061]

[0062] The final loss function of the second-stage training process is:

[0063]

[0064] where β is a hyperparameter used to control the contrastive learning loss, set to 1 on the SIRI-WHU-LT dataset, and set to 10 in PatternNet-LT and RSI-CB256-LT. L o and L m correspond to the cross-entropy loss functions for training with the original data and the mixed data respectively.

[0065] Through the above steps, the present invention constructs an efficient data augmentation method for long-tailed remote sensing data, which can effectively improve the problem of unbalanced data distribution and enhance the model's classification ability for tail classes. Experimental results show that compared with other methods, the method of the present invention has achieved significant performance improvements on the three long-tailed remote sensing datasets SIRI-WHU-LT, PatternNet-LT, and RSI-CB256-LT.

[0066] Table 1 Accuracy (%) of the dataset SIRI-WHU-LT on the network ResNet-32.

[0067]

[0068]

[0069] Table 1 shows the accuracy comparison of the present invention with some common long-tail methods and data augmentation methods such as LDAM, CB, BBN, BKD, Mixup, Cutmix, DIFFUSEMIX, etc. on the long-tail remote sensing dataset SIRI-WHU-LT. SIRI-WHU-LT is a long-tail dataset constructed using SIRI-WHU. 40 samples are selected for the test set for each category, and the number of samples N of the class with the largest number of samples max = 140. It can be seen that compared with other methods, the present invention achieves the highest accuracy on the SIRI-WHU-LT dataset with three imbalance factors on the network ResNet-32, increasing by 5.00% - 27.50%, 3.06% - 37.64%, and 5.28% - 31.11% respectively, fully demonstrating the effectiveness of the present invention. Even when the data is extremely imbalanced, the present invention can maintain relatively stable performance.

[0070] Table 2 Accuracy (%) of the dataset PatternNet-LT on the network ResNet-50.

[0071]

[0072] Table 2 shows the accuracy of some other long-tail methods, data augmentation methods, and the present invention on the PatternNet-LT dataset, using the network ResNet-50. Regarding the construction of PatternNet-LT, 100 samples are selected for the test set for each category, and the remaining are constructed as the long-tail dataset, where the number of samples N of the class with the largest number of samples max = 700. Compared with the comparison methods, the accuracy of the present invention is improved by 0.42% - 2.31%, 0.89% - 8.26%, and 1.08% - 11.60% respectively under the conditions of three imbalance factors. In contrast, even in the most extreme imbalance situation, the present invention still shows superiority and achieves better results.

[0073] Table 3 Accuracy (%) of the dataset RSI-CB256-LT on the network ResNet-50.

[0074]

[0075] Table 3 shows the accuracy comparison of the present invention with some common long-tail methods and data augmentation methods on RSI-CB256-LT, using the ResNet-50 network. Regarding the construction of RSI-CB256-LT, 50 samples are selected for the test set for each category, and the remaining samples are constructed as the long-tail dataset, where the number of samples N of the class with the largest number of samples max= 100. Compared with other methods, the precision of the present invention is optimal in all cases, with improvements of 1.14% - 12.46%, 2.63% - 24.51%, and 1.71% - 23.03% respectively for the three imbalance factors.

Claims

1. A diffusion noise enhancement method for long-tail remote sensing images, characterized in that, The method includes: S1. Input the original remote sensing dataset I with a long-tailed distribution o ; S2. Define the conditional prompt set P and the original dataset I o and the input data generation module of the prompt set P, and obtain the generated image dataset I' after CLIP filtering g ; Among them: the conditional prompt set P = {Snowy, Sunset, Autumn, Rainy, Winter, Sunny, Summer, Cloudy, Spring, Foggy}, and the prompt set is defined as P = {P1, P2, …, P m}, the original image and the conditional prompt set are input into the diffusion model together. For each input image x i ∈I o , the diffusion model generates a generated image x′ j corresponding to the style according to the prompt P ij , x i is the original image corresponding to x′ ij . The label of the generated sample is consistent with the label of its corresponding original sample. Repeat the above generation process to finally obtain the generated image dataset I g ; S3. Perform class-balanced sampling on the original dataset I o and use the cross-entropy loss function for the first-stage training of the model; S4. Input the original dataset I o and the generated image dataset I' g into the DiffCam-Mix module, and use the model trained in one stage to calculate the foreground masks and background masks of the images in I o and I' g respectively. Then use the masks to mix the foreground of the original image with the background of the generated image to obtain the mixed sample dataset I m ; S5. Original dataset I o and mixed sample dataset I m are used as inputs for the second stage of model training to achieve diffusion noise enhancement for long-tail remote sensing images.

2. The diffusion noise enhancement method for long-tail remote sensing images according to claim 1, characterized in that: The data generation module described in step S2 uses the pre-trained diffusion model InstructPix2Pix as the generation model, and then uses the pre-trained CLIP model to filter the generated image dataset I g The specific process is as follows: Define the template T for class y y as "A photo of the class [y]", where [y] represents the name of class y. Using the image encoder and text encoder of CLIP, calculate and generate the image x' ij of the image embedding and its corresponding class template T y of the text embedding, and then calculate the cosine similarity score between them For class y, select K - N y samples with the highest cosine similarity scores, where K is the maximum number of samples per class and N y is the original number of samples of class y. Filter the remaining samples to finally obtain the high-quality generated dataset I'. g .

3. The diffusion noise enhancement method for long-tail remote sensing images according to claim 1, characterized in that: Step S4 generates the hybrid sample dataset I m The steps are as follows: Performing a data mixing operation on the generated data and the corresponding original data using the model trained in the first stage and the Grad-CAM++ class activation mapping method; For each image, through forward and backward propagation, the gradient of the target class c with respect to the feature map is obtained, denoted as where Y c is the output score of the target class c, and A k is the activation feature map of the target layer; For each activation map A k , the gradient weight is calculated as follows: where (m,n) and (a,b) are the same iterators throughout the activation map A k above, and represent the second-order partial derivative and the third-order partial derivative respectively, multiply the gradient after Relu activation, sum over all positions (m,n), and obtain the final weight Multiplying and summing the calculated weights with the feature map, and then obtaining the class activation map CAM through the ReLU layer: Further processing is performed using normalization and resizing to match the size of the input image to obtain the final Class Activation Map (CAM). resized , according to the following formula: M = (1 - CAM resized ) 2 , Obtain the background mask for each generated image, where M ∈ {0, 1} W×H ; Define the mixing operation as: where x i (x i ∈I o ) is the original image sample, x′ i (x′ i ∈I′ g ) is the corresponding generated image sample, is the mixed data generated based on the above x i and x′ i . λ is a hyperparameter ranging from 0 to 1 and is used to control the mixing ratio of the foreground and background, is used to calculate the background part of x′ i , is used to calculate the foreground part of x i , and ⊙ is element-wise multiplication; Since this operation is performed between two samples belonging to the same category, the label of the mixed data will not change.

4. The diffusion noise enhancement method for long-tail remote sensing images according to claim 1, characterized in that: The two-stage training process of the model in step S5 is as follows: After obtaining the mixed data, use these samples and the original samples for the second-stage training to fine-tune the network; according to the idea of SimSiam, the original sample x i and the corresponding mixed sample are regarded as two different augmented views, and they are used as a positive sample pair. After extracting features from the positive sample pair, they are fed into the projection head to obtain and respectively, and then fed into the prediction head to obtain and respectively. The goal of this step is to minimize the negative cosine similarity between the positive sample pairs, which is expressed as: The final loss function of the second-stage training process is: where β is a hyperparameter used to control the contrastive learning loss, L o and L m correspond to the cross-entropy loss functions for training on the original data and the mixed data, respectively.

Citation Information

Cited By

  • Deep forgery detection classification and positioning method based on noise inconsistency

    CN117079354A

  • A deepfake detection classification and localization method based on noise inconsistency

    CN117079354B