Small sample image generation method and system based on feature difference dynamic guidance

Through the small sample image generation method based on the diffusion model, the loss function is constructed using the feature difference between the latent variable and the reference image, the problem of insufficient image generation quality and diversity in the prior art is solved, and high-quality and diverse image generation is achieved.

CN120014081AActive Publication Date: 2025-05-16SOUTH CHINA UNIV OF TECH

Patent Information

Application Number
CN202411532682.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2025-05-16
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

The existing small sample image generation methods have shortcomings in generating similar images with high visual quality and rich diversity. In particular, GAN models are prone to aliasing artifacts, and the overall blurring of the generated images is strong, semantic convergence, and the quality is poor.

Method used

Using a small sample image generation method based on the diffusion model, the reference image is noise-added and iteratively denoised, and the loss function is constructed using the feature difference between the latent variable and the reference image, and the diffusion model is trained to generate high-quality and diverse images.

Benefits of technology

It realizes the generation of similar images with high visual quality and rich diversity, avoids the common aliasing artifact problems in GAN models, and the generated image details are rich and the overall quality is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014081A_ABST
    Figure CN120014081A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample image generation method and system based on feature difference dynamic guidance. The method comprises the following steps: acquiring image data and category labels thereof, and establishing a training set; the method comprises the following steps: constructing a diffusion model, in the diffusion model, continuously adding noise to an input reference image to obtain an image after the noise is added for t steps, which is equivalent to isotropic Gaussian noise, and iteratively executing multiple single-step denoising reduction on the image after the noise is added for t steps to obtain a predicted image; constructing a loss function according to the feature difference of the predicted image and the reference image, and training a diffusion model by using data in the training set to obtain a trained diffusion model; and taking the image data of the unknown category which does not appear in the training set as a reference image, inputting the reference image into the trained diffusion model, and generating a small sample image. The diversity of the generated image is improved, and the visual quality of the image is improved by fusing the information of the reference image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image generation, and in particular to a method and system for generating a small sample image based on dynamic guidance of feature differences. Background Art

[0002] The widespread application of deep learning has made daily life more convenient and efficient, and has greatly improved people's work efficiency. However, the problem that still exists when applying deep learning to real-world scenarios is that the model requires a large amount of training data. In order to alleviate the limitation of limited data on the application of deep learning, the small sample image generation task was proposed. Small sample image generation refers to the task of synthesizing rich and diverse new samples of the target category through a deep learning model using limited data samples of the target category. Small sample image generation can be regarded as a way of data augmentation, which expands the existing data set and assists the application and promotion of deep learning models in real life and production fields.

[0003] Existing methods for generating small sample images can be mainly divided into three categories: optimization-based methods, conversion-based methods, and fusion-based methods. Most of these methods are implemented based on Generative Adversarial Networks (GAN), and methods based on diffusion models have also been gradually proposed in recent years. FIGR pioneered the use of optimization-based meta-learning Reptile in the training of generative adversarial network models to generate small sample images, but the computational cost of training the model is very high. DAGAN generates novel samples by injecting random noise into the representation of a single input image, but the generated images lack diversity. WaveGAN (WaveGAN: Frequency-aware GAN for High-Fidelity Few-shot Image Generation) decomposes the encoded features into multiple frequency components, feeds the high-frequency components to the decoder through high-frequency jump connections to alleviate the difficulty of the generator in synthesizing fine details, and uses low-frequency jump connections to retain perceptible basic information, but it is prone to aliasing artifacts. FSDM uses ViT to aggregate image block information in units of image sets, and fuses image block information and latent variables through cross-attention to adapt to the generation process conditioned on a set of images of a given category. However, the generated images have a strong overall blur and similar semantics, and the generation quality is poor. Summary of the invention

[0004] The purpose of the present invention is to provide a method and system for generating small sample images based on dynamic guidance of feature differences, so as to generate similar images with high visual quality and rich diversity for a small number of samples.

[0005] This paper studies the small sample image generation task based on the diffusion model. Compared with GAN, the diffusion model has a certain training target and better distribution coverage, and has great potential in the small sample image generation task.

[0006] The present invention makes full use of the information of the reference image and reasonably guides the generation process of the diffusion model. The generation process provides partial context information of the unknown category image for the latent variable, and the auxiliary diffusion model controls the details of the generated image. By reducing the distance between the generated image estimated by the latent variable and the reference image in the feature space, the noise distribution removed in each step is adjusted. The generated image has greater flexibility in variation, which is conducive to obtaining diverse and high-quality generated images, and can avoid the common aliasing artifact problem in the GAN-based model.

[0007] The purpose of the present invention is achieved by at least one of the following technical solutions.

[0008] A method for generating a small sample image based on dynamic guidance of feature difference comprises the following steps:

[0009] S1, obtain image data and its category labels, and establish a training set;

[0010] S2. Construct a diffusion model. In the diffusion model, noise is continuously added to the input reference image to obtain the noise-added image. t The image after the step is equivalent to an isotropic Gaussian noise, and then the noise is added t The image after the step is iteratively performed multiple times of single-step denoising and restoration to obtain the predicted image;

[0011] S3, constructing a loss function based on the feature difference between the predicted image and the reference image, and training the diffusion model using the data in the training set to obtain a trained diffusion model;

[0012] S4. Using image data of unknown categories that do not appear in the training set as reference images, inputting them into the trained diffusion model to generate a small sample image.

[0013] Furthermore, in step S1, an image acquisition device (such as a high-definition camera) is used to image the target object and put the image into a folder of a corresponding category, and the name of the folder is used as an image label to complete the establishment of a training set.

[0014] Furthermore, in step S2, in the diffusion model, noise is continuously added to the input known image, and the expression is:

[0015]

[0016] Among them, X0 is the input reference image, X t is the latent variable, which is the image after adding noise for t steps, t = 1, 2, 3, ..., T, T is the total number of steps for adding noise, αt =1-β t , β t is the variance used in the tth noise addition process, ∈~N(0,1), ∈ is Gaussian noise and has the same dimension as the original image X0.

[0017] Furthermore, in step S2, the existing deep neural network Unet (U-Net: Convolutional Networks for Biomedical Image Segmentation) is used to predict the noise removed in a single step, and the latent variable X t The expression of single-step denoising is:

[0018]

[0019] Among them, the latent variable X t-1 is the latent variable X t The result of single denoising, ∈ θ (X t , t) is the noise distribution predicted by the deep neural network Unet, δ t is the latent variable X t The standard deviation of , Z~N(0,I).

[0020] Furthermore, in each iterative denoising step of the diffusion model, the latent variable is fused with the context information of the input reference image, which specifically includes the following steps:

[0021] S2.1. Latent variable X t After single-step denoising, the latent variable X′ is obtained t-1 ;

[0022] S2.2, add noise to the reference image Y for t-1 steps to obtain Y t-1 ;

[0023] S2.3, the fusion expression of the latent variable and the context information of the reference image is as follows:

[0024] X″ t-1 =M·Y t-1 +(1-M)·X′ t-1

[0025] Where M is the mask, X″ t-1 is the result of fusion of latent variables and reference image.

[0026] Furthermore, in the diffusion model, the latent variable X t Iterate denoising to get the current latent variable X t Prediction image after denoising The expression is as follows:

[0027]

[0028] Furthermore, in step S3, during the training of the diffusion model, for a single reference image in the training set of the input diffusion model, the output prediction image is calculated. The distance D between the feature representation of the reference image feature , the expression is as follows:

[0029]

[0030] in, Represents the predicted image The feature representation of the kth layer after N layers of feature extractors, k = 1, 2, ... N, represents the feature representation of the kth layer after the reference image passes through the feature extractor; D k Represents the predicted image The distance between the feature representation of the reference image output at the kth layer of the feature extractor; D feature It is the sum of the distances between the feature representations output by each layer of the N-layer feature extractor.

[0031] Furthermore, during the training of the diffusion model, the loss function is constructed to minimize the predicted image The sum of the distances D between the feature representations output by each layer of the feature extractor of the reference image after passing through N layers feature .

[0032] The difference between the feature representation of the predicted image and the reference image output by each layer of the feature extractor varies. If the difference between the feature representations of a certain layer is the largest, D feature The feature difference of this layer will dominate, and the noise distribution predicted by Unet will be adjusted mainly by reducing the feature difference of this layer.

[0033] Furthermore, the N-layer feature extractor includes Resnet or VGG.

[0034] A small sample image generation system based on a diffusion model, comprising:

[0035] The first module is used to obtain image data and its category labels and establish a training set;

[0036] The second module is used to build a diffusion model. In the diffusion model, noise is continuously added to the input reference image to obtain the noise-added t The image after the step is equivalent to an isotropic Gaussian noise, and then the noise is added t The image after the step is iteratively performed multiple times of single-step denoising and restoration to obtain the predicted image;

[0037] The third module is used to construct a loss function based on the feature difference between the predicted image and the reference image, and train the diffusion model using the data in the training set to obtain a trained diffusion model;

[0038] The fourth module is used to use image data of unknown categories that do not appear in the training set as reference images, input them into the trained diffusion model, and generate a small sample image.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] This paper studies the small sample image generation task based on the diffusion model. Compared with GAN, the diffusion model has a certain training target and better distribution coverage, and has great potential in the small sample image generation task.

[0041] The present invention makes full use of the information of the reference image and reasonably guides the generation process of the diffusion model. The generation process provides partial context information of the unknown category image for the latent variable, and the auxiliary diffusion model controls the details of the generated image. By reducing the distance between the generated image estimated by the latent variable and the reference image in the feature space, the noise distribution removed in each step is adjusted. The generated image has greater flexibility in variation, which is conducive to obtaining diverse and high-quality generated images, and can avoid the common aliasing artifact problem in the GAN-based model. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The following description combines specific illustrations to explain the technical solution so as to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and similar generalized embodiments made by ordinary technicians in this field without creative work are all within the scope of protection of the present invention.

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0044] Figure 1 It is an implementation flow chart of a method for generating a small sample image based on dynamic guidance of feature difference in an embodiment of the present invention;

[0045] Figure 2 Schematic diagram of the overall structure of a method for generating a small sample image based on dynamic guidance of feature differences in an embodiment of the present invention;

[0046] Figure 3 Schematic diagram of the fusion of latent variables and context information of a reference image in an embodiment of the present invention;

[0047] Figure 4 Schematic diagram of calculating the characteristic distance between a latent variable prediction image and a reference image in an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The following description combines specific illustrations to explain the technical solution so as to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and similar generalized embodiments made by ordinary technicians in this field without creative work are all within the scope of protection of the present invention.

[0049] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. The singular forms of "a", "the" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0050] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish similar objects and are not necessarily used to describe the order or sequence of features described in one or more embodiments of this specification. In addition, the terms "having" and "including" are similarly expressed to indicate a non-exclusive scope. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to the contents listed in detail, but may include inherent contents related to these steps or modules that are not listed.

[0051] Example:

[0052] A small sample image generation method based on dynamic guidance of feature difference, such as Figure 1 and Figure 2 As shown, the following steps are included:

[0053] S1, obtain image data and its category labels, and establish a training set;

[0054] In one embodiment, the target object is imaged by an image acquisition device (such as a high-definition camera) and the image is placed in a folder of a corresponding category, and the name of the folder is used as an image label to complete the establishment of a training set.

[0055] S2. Construct a diffusion model. In the diffusion model, noise is continuously added to the input reference image to obtain an image after t steps of noise addition, which is equivalent to an isotropic Gaussian noise. Then, multiple single-step denoising and restoration are performed iteratively on the image after t steps of noise addition to obtain a predicted image.

[0056] In the diffusion model, noise is continuously added to the input known image, and the expression is:

[0057]

[0058] Among them, X0 is the input reference image, X t is the latent variable, which is the image after adding noise for t steps, t = 1, 2, 3, ..., T, T is the total number of steps for adding noise, α t =1-β t , β t is the variance used in the tth noise addition process, ∈~N(0,1), ∈ is Gaussian noise and has the same dimension as the original image X0.

[0059] The data sample X0 gradually loses its distinguishable features as the step size t increases. Finally, when T→∞, X T It is equivalent to an isotropic Gaussian distribution.

[0060] In one embodiment, the existing deep neural network Unet (U-Net: Convolutional Networks for Biomedical Image Segmentation) is used to predict the noise removed in a single step, such as Figure 3 As shown, for the latent variable X t The expression of single-step denoising is:

[0061]

[0062] Among them, the latent variable X t-1 is the latent variable X t The result of single denoising, ∈ θ (X t , t) is the noise distribution predicted by the deep neural network Unet, δ t is the latent variable X t The standard deviation of , Z~N(0,I).

[0063] In each iterative denoising step of the diffusion model, the latent variable is fused with the contextual information of the input reference image, which specifically includes the following steps:

[0064] S2.1. Latent variable X t After single-step denoising, the latent variable X′ is obtained t-1 ;

[0065] S2.2, add noise to the reference image Y for t-1 steps to obtain Y t-1 ;

[0066] S2.3, the fusion expression of the latent variable and the context information of the reference image is as follows:

[0067] X″ t-1 =M·Y t-1 +(1-M)·X′ t-1

[0068] Where M is the mask, X″ t-1 is the result of fusion of latent variables and reference image.

[0069] In the diffusion model, the latent variable X t Iterate denoising to get the current latent variable X t Prediction image after denoising The expression is as follows:

[0070]

[0071] S3. In one embodiment, Figure 4 As shown, the loss function is constructed by using the feature difference between the predicted image and the reference image, and the diffusion model is trained using the data in the training set to obtain a trained diffusion model;

[0072] During the training of the diffusion model, for a single reference image in the training set of the input diffusion model, the output predicted image is calculated The distance D between the feature representation of the reference image feature , the expression is as follows:

[0073]

[0074] in, Represents the predicted image The feature representation of the kth layer after N layers of feature extractors, k = 1, 2, ... N, represents the feature representation of the kth layer after the reference image passes through the feature extractor; D k Represents the predicted image The distance between the feature representation of the reference image output at the kth layer of the feature extractor; D feature It is the sum of the distances between the feature representations output by each layer of the N-layer feature extractor.

[0075] During the training of the diffusion model, the loss function is constructed to minimize the predicted image The sum of the distances between the feature representations output by each layer of the feature extractor of the reference image after passing through N layers D feature .

[0076] The difference between the feature representation of the predicted image and the reference image output by each layer of the feature extractor varies. If the difference between the feature representations of a certain layer is the largest, D feature The feature difference of this layer will dominate, and the noise distribution predicted by Unet will be adjusted mainly by reducing the feature difference of this layer.

[0077] In one embodiment, the N-layer feature extractor is Resnet.

[0078] In one embodiment, the N-layer feature extractor is VGG.

[0079] S4. Using image data of unknown categories that do not appear in the training set as reference images, inputting them into the trained diffusion model to generate a small sample image.

[0080] In one embodiment, a small sample image generation system based on a diffusion model includes:

[0081] The first module is used to obtain image data and its category labels and establish a training set;

[0082] The second module is used to build a diffusion model. In the diffusion model, noise is continuously added to the input reference image to obtain the noise-added t The image after the step is equivalent to an isotropic Gaussian noise, and then the noise is added t The image after the step is iteratively performed multiple times of single-step denoising and restoration to obtain the predicted image;

[0083] The third module is used to construct a loss function based on the feature difference between the predicted image and the reference image, and train the diffusion model using the data in the training set to obtain a trained diffusion model;

[0084] The fourth module is used to use image data of unknown categories that do not appear in the training set as reference images, input them into the trained diffusion model, and generate a small sample image.

[0085] The preferred embodiments of the present application disclosed above are only used to help understand the present invention and its core ideas. For those skilled in the art, according to the ideas of the present invention, there will be changes in specific application scenarios and implementation operations, and this specification should not be interpreted as limiting the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A method for generating small sample images based on dynamic guidance of feature difference, characterized in that: The following steps are involved: S1, obtain image data and its category labels, and establish a training set; S2. Construct a diffusion model. In the diffusion model, noise is continuously added to the input reference image to obtain an image after t steps of noise addition, which is equivalent to an isotropic Gaussian noise. Then, multiple single-step denoising and restoration are performed iteratively on the image after t steps of noise addition to obtain a predicted image. S3, constructing a loss function based on the feature difference between the predicted image and the reference image, and training the diffusion model using the data in the training set to obtain a trained diffusion model; S4. Using image data of unknown categories that do not appear in the training set as reference images, inputting them into the trained diffusion model to generate a small sample image.

2. The method for generating small sample images based on dynamic guidance of feature differences according to claim 1, characterized in that: In step S1, the target object is imaged by an image acquisition device and the image is placed in a folder of the corresponding category, and the name of the folder is used as the image label to complete the establishment of the training set.

3. The method for generating small sample images based on dynamic guidance of feature difference according to claim 1, characterized in that: In step S2, in the diffusion model, noise is continuously added to the input known image, and the expression is: Among them, X0 is the input reference image, X t is the latent variable, which is the image after adding noise for t steps, t = 1, 2, 3, ..., T, T is the total number of steps for adding noise, α t =1-β t , β t is the variance used in the tth noise addition process, is Gaussian noise and has the same dimension as the original image X0.

4. The method for generating small sample images based on dynamic guidance of feature differences according to claim 3, characterized in that: In step S2, the existing deep neural network Unet (U-Net: Convolutional Networks for Biomedical Image Segmentation) is used to predict the noise removed in a single step, and the latent variable X t The expression of single-step denoising is: Among them, the latent variable X t-1 is the latent variable X t The result of single denoising, ∈ θ (X t , t) is the noise distribution predicted by the deep neural network Unet, δ t is the latent variable X t The standard deviation of , Z~N(0,1).

5. The method for generating small sample images based on dynamic guidance of feature differences according to claim 4, characterized in that: In each iterative denoising step of the diffusion model, the latent variable is fused with the contextual information of the input reference image, which specifically includes the following steps: S2.

1. Latent variable X t After single-step denoising, the latent variable X′ is obtained t-1 ; S2.2, add noise to the reference image Y for t-1 steps to obtain Y t-1 ; S2.3, the fusion expression of the latent variable and the context information of the reference image is as follows: X″ t-1 =M·Y t-1 +(1-M)·X′ t-1 Where M is the mask, X″ t-1 is the result of fusion of latent variables and reference image.

6. The method for generating small sample images based on dynamic guidance of feature differences according to claim 5, characterized in that: In the diffusion model, the latent variable X t Iterate denoising to get the current latent variable X t Prediction image after denoising The expression is as follows:

7. The method for generating small sample images based on dynamic guidance of feature differences according to claim 6, characterized in that: In step S3, during the training of the diffusion model, for a single reference image in the training set of the input diffusion model, the output predicted image is calculated. The distance D between the feature representation of the reference image feature , the expression is as follows: in, Represents the predicted image The feature representation of the kth layer after N layers of feature extractors, k = 1, 2, ... N, represents the feature representation of the kth layer after the reference image passes through the feature extractor; D k Represents the predicted image The distance between the feature representation of the reference image output at the kth layer of the feature extractor; D feature It is the sum of the distances between the feature representations output by each layer of the N-layer feature extractor.

8. The method for generating small sample images based on dynamic guidance of feature differences according to claim 7, characterized in that: During the training of the diffusion model, the loss function is constructed to minimize the predicted image The sum of the distances D between the feature representations output by each layer of the feature extractor of the reference image after passing through N layers feature .

9. The method for generating small sample images based on dynamic guidance of feature difference according to claim 7, characterized in that: The N-layer feature extractor includes Resnet or VGG.

10. A small sample image generation system based on a diffusion model based on a small sample image generation method based on dynamic guidance of feature difference according to any one of 1 to 9, characterized in that: include: The first module is used to obtain image data and its category labels and establish a training set; The second module is used to build a diffusion model. In the diffusion model, noise is continuously added to the input reference image to obtain an image after t steps of noise addition, which is equivalent to an isotropic Gaussian noise. Then, multiple single-step denoising and restoration are performed iteratively on the image after t steps of noise addition to obtain a predicted image. The third module is used to construct a loss function based on the feature difference between the predicted image and the reference image, and train the diffusion model using the data in the training set to obtain a trained diffusion model; The fourth module is used to use image data of unknown categories that do not appear in the training set as reference images, input them into the trained diffusion model, and generate a small sample image.

Citation Information

Patent Citations

  • Small sample image generation method and system based on diffusion model

    CN116957964A

  • SAR image generation method based on de-noising diffusion probability model

    CN118230191A

  • Image generation method and system based on limited data set

    CN118379594A

Cited By

  • Image generation method and device and electronic equipment

    CN121437671A