Generative artificial intelligence image data enhancement method based on diffusion model
By generating high-quality and diverse image data through diffusion models and specific strategies, the problems of difficult image data acquisition and insufficient diversity are solved, achieving the generation of realistic and diverse images and improving the applicability and robustness of image analysis tasks.
Patent Information
- Application Number
- CN202511121420.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-12-02
AI Technical Summary
Existing image data acquisition is difficult, and the generated new samples lack diversity, making it difficult to meet the needs of complex image analysis tasks. Existing data augmentation methods are also insufficient.
By employing a diffusion model combined with specific enhancement strategies, image features are extracted through a cross-attention mechanism and a self-supervised model to generate high-quality and diverse image data. A time-aware strategy is used to process dynamic scenes, and structural consistency and high-frequency detail loss functions are designed to optimize image generation.
The generated images exhibit excellent realism and diversity, adapting to different scenario requirements, enabling personalized data augmentation, and improving the algorithm's versatility and robustness.
Smart Images

Figure CN121053484A_ABST
Abstract
Description
Technical Field
[0001] This invention discloses an image data enhancement method based on a diffusion model and generative artificial intelligence, which relates to the technical field of image data enhancement. Background Technology
[0002] In numerous image-related fields, the demand for high-quality, diverse image data is extremely urgent. However, the image data actually acquired often faces many challenges. For example, in the field of medical imaging, the number of images acquired is limited due to factors such as the imaging principles of equipment, the complexity of human physiological structures, and radiation dose limitations. Furthermore, some images suffer from quality issues such as noise interference and low contrast. The acquired images cannot cover all possible target states and environmental changes, resulting in a severe lack of data diversity. Existing data augmentation methods, such as simple image flipping, scaling, and cropping, can expand the dataset to some extent, but the newly generated samples lack sufficient diversity, making it difficult to meet the increasingly complex image analysis tasks. Summary of the Invention
[0003] This invention addresses the problems of existing technologies by providing an image data enhancement method based on a diffusion model and generative artificial intelligence. By combining the generative capabilities of the diffusion model with specific enhancement strategies, it generates high-quality and diverse image data, ultimately achieving the generation, synthesis, and data enhancement of scarce samples and long-tailed data under zero-sample settings. This optimizes sample balance, improves the algorithm's generality and robustness, thereby increasing detection accuracy and shortening the data acquisition cycle.
[0004] The specific solution proposed in this invention is as follows:
[0005] This invention provides an image data augmentation method based on a diffusion model and generative artificial intelligence, comprising:
[0006] Step 1: Feature extraction from the object image: Extracting the object image's identity features and detailed features.
[0007] Step 2: Inject dual-path identity features and detail features into the image enhancement model to achieve the image enhancement process. Identity features are injected into each layer of the UNet network through a cross-attention mechanism. In the encoder stage, identity features interact with image features through attention to generate global identity features of the object. In the decoder stage, identity constraints are further strengthened to avoid feature drift during the generation of global identity features of the object, and the global identity features are output. Detail features are then concatenated and fused with global identity features to obtain an enhanced image that maintains texture realism in different scenes. For continuous changes in object pose and lighting in dynamic scenes, a time-aware sampling strategy is adopted. Multiple frame sequences of the same object are extracted from the video, and video frames of the object are collected by dynamically adjusting the time step sampling density. The sampled video frames are mixed with static object image data and input into the image enhancement model. During training, the image enhancement model learns both the static features and dynamic evolution patterns of the object, enabling the model to adapt to dynamic scenes and generate enhanced images with temporal coherence and scene fusion.
[0008] Furthermore, in step 1 of the generative artificial intelligence image data augmentation method based on a diffusion model, the backbone network for extracting identity features from object images uses a self-supervised model DINO-V2. The object image is then processed by removing the background, center-aligning it, and inputting it into the self-supervised model DINO-V2 for global and local identity feature extraction.
[0009] To extract detailed features from an image, the background is removed. A high-frequency map is extracted from the image using a high-pass filter Sobel operator. The high-frequency map is then spatially concatenated with the base background map. The concatenated composite feature is input into a pre-trained detail feature extraction network to extract hierarchical detail feature maps containing different scales.
[0010] Furthermore, in step 2 of the image data augmentation method based on a diffusion model for generative artificial intelligence, when training the model for image augmentation, a training dataset is constructed. The training dataset includes the input target image, the scene image and its corresponding location information, and the output image.
[0011] Furthermore, by employing adaptive time-step sampling, the model prioritizes sampling earlier time steps for video data and later time steps for image data, enabling it to better learn the features of different types of data.
[0012] Furthermore, in step 2 of the generative artificial intelligence image data augmentation method based on diffusion model, a loss function for the model used for image augmentation is designed. The structural consistency loss L1 is used to ensure that the overall structure of the generated object matches the target object; the high-frequency detail loss L1 is used to ensure that the high-frequency details of the generated object are consistent with the target object; and the adversarial loss GAN is used to enhance the authenticity of the generated result.
[0013] The present invention also provides an image data enhancement device based on a diffusion model and generative artificial intelligence, comprising a feature extraction module and an image enhancement module.
[0014] The feature extraction module extracts features from object images: it extracts the object's identity features and detailed features.
[0015] The image enhancement module injects dual-path identity features and detail features into the image enhancement model to achieve the image enhancement process. Identity features are injected into each layer of the UNet network through a cross-attention mechanism. In the encoder stage, identity features interact with image features through attention to generate global identity features of the object. In the decoder stage, identity constraints are further strengthened to avoid feature drift during the generation of global identity features of the object, and the global identity features are output. Detail features are then concatenated and fused with global identity features to obtain an enhanced image that maintains the texture realism of the object in different scenes. For continuous changes in object pose and lighting in dynamic scenes, a time-aware sampling strategy is adopted. Multiple frame sequences of the same object are extracted from the video, and video frames of the object are acquired by dynamically adjusting the sampling density of the time step. The sampled video frames are mixed with static object image data and input into the image enhancement model. During training, the image enhancement model learns both the static features and dynamic evolution rules of the object, enabling the model to adapt to dynamic scenes and generate enhanced images with temporal coherence and scene fusion.
[0016] Furthermore, the backbone network of the feature extraction module of the generative artificial intelligence image data enhancement device based on a diffusion model, which extracts identity features from object images, adopts the self-supervised model DINO-V2. The image image is then processed by removing the background, center-aligning it, and inputting it into the self-supervised model DINO-V2 for global and local identity feature extraction.
[0017] To extract detailed features from an image, the background is removed. A high-frequency map is extracted from the image using a high-pass filter Sobel operator. The high-frequency map is then spatially concatenated with the base background map. The concatenated composite feature is input into a pre-trained detail feature extraction network to extract hierarchical detail feature maps containing different scales.
[0018] Furthermore, in the image enhancement module of the generative artificial intelligence image data enhancement device based on a diffusion model, when training the model for image enhancement, a training dataset is constructed. The training dataset includes the input target image, scene image and corresponding location information, and output image.
[0019] Furthermore, by employing adaptive time-step sampling, the model prioritizes sampling earlier time steps for video data and later time steps for image data, enabling it to better learn the features of different types of data.
[0020] Furthermore, the image enhancement module of the generative artificial intelligence image data enhancement device based on the diffusion model is designed with a loss function for the image enhancement model. The loss function uses structural consistency loss L1 to ensure that the overall structure of the generated object matches the target object; high-frequency detail loss L1 is used to ensure that the high-frequency details of the generated object are consistent with the target object; and adversarial loss GAN is used to improve the authenticity of the generated result.
[0021] The advantages of this invention are:
[0022] 1. Excellent enhancement effect: Compared with existing image enhancement methods, the enhanced images generated based on the model used for image enhancement have higher realism and diversity. They can generate new image content that conforms to the laws of real scenes, rather than simply transforming the original image, effectively making up for the shortcomings of traditional methods.
[0023] 2. High flexibility: Through innovative mechanisms such as condition constraints and dynamic noise scheduling, it can flexibly control the features and styles of enhanced images according to different application scenarios and task requirements, realize personalized and targeted data enhancement, and improve the applicability of the technology.
[0024] 3. Easy to expand: The core framework of this technology has good scalability. It can quickly adapt to new application scenarios by adjusting the input conditions of the model and network parameters according to new image types and enhancement needs, without the need for large-scale modifications to the overall architecture. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation
[0026] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.
[0027] Example 1
[0028] This invention provides an image data augmentation method based on a diffusion model and generative artificial intelligence, comprising:
[0029] Step 1: Feature extraction from object images: Extracting identity features and detail features from object images. The backbone network for extracting identity features uses the self-supervised model DINO-V2. Background is removed from the object image, center alignment is performed, and the image is then input into the self-supervised model DINO-V2 for global and local identity feature extraction.
[0030] To extract detailed features from an image, the background is removed. A high-frequency map is extracted from the image using a high-pass filter Sobel operator. The high-frequency map is then spatially concatenated with the base background map. The concatenated composite feature is input into a pre-trained detail feature extraction network to extract hierarchical detail feature maps containing different scales.
[0031] Step 2: Inject dual-path identity features and detail features into the image enhancement model to achieve the image enhancement process. Identity features are injected into each layer of the UNet network through a cross-attention mechanism. In the encoder stage, identity features interact with image features through attention to generate global identity features of the object. In the decoder stage, identity constraints are further strengthened to avoid feature drift during the generation of global identity features of the object, and the global identity features are output. Detail features are then concatenated and fused with global identity features to obtain an enhanced image that maintains texture realism in different scenes. For continuous changes in object pose and lighting in dynamic scenes, a time-aware sampling strategy is adopted. Multiple frame sequences of the same object are extracted from the video, and video frames of the object are collected by dynamically adjusting the time step sampling density. The sampled video frames are mixed with static object image data and input into the image enhancement model. During training, the image enhancement model learns both the static features and dynamic evolution patterns of the object, enabling the model to adapt to dynamic scenes and generate enhanced images with temporal coherence and scene fusion.
[0032] When training the model for image enhancement, a training dataset is constructed. The training dataset includes the input target image, scene image and corresponding location information, and output image.
[0033] Furthermore, by employing adaptive time-step sampling, the model prioritizes sampling earlier time steps for video data and later time steps for image data, enabling it to better learn the features of different types of data.
[0034] Furthermore, a loss function for the image enhancement model is designed, in which structural consistency loss L1 is used to ensure that the overall structure of the generated object matches the target object; high-frequency detail loss L1 is used to ensure that the high-frequency details of the generated object are consistent with the target object; and adversarial loss GAN is used to improve the realism of the generated result.
[0035] This invention enables lightweight deployment and generates a lightweight architecture. The lightweight architecture achieves seamless integration of any target object with a new scene without requiring additional model fine-tuning. The specific process is as follows: After inputting the target object image and the scene image to be fused, the system automatically locates the target object's position within the scene and generates a bounding box. This bounding box is then expanded into a square to ensure the target object has sufficient spatial dimensions during generation. The expanded square region is then fed into the diffusion model. The model, by utilizing pre-learned dual-feature fusion capabilities, maintains the target object's identity and detailed features while adjusting the object's lighting, shadows, and other attributes according to the scene environment, ultimately generating a naturally blended result, such as precisely integrating furniture into interior scenes of different styles.
[0036] Lightweight deployment enables shape control: allowing users to intuitively adjust the shape of objects through simple interactions to meet personalized needs. Users only need to draw a rough mask, such as outlining the direction of folds in clothing or the bending angle of robot joints, and the model can recognize the shape constraint information contained in the mask and adjust it in combination with the identity and detailed features of the target object. During the generation process, the diffusion model performs fine optimization of the local structure of the object based on the spatial distribution of the mask, so that the posture of the generated object highly matches the mask drawn by the user, such as changing the bending shape of mechanical parts or the drape of clothing according to the hand-drawn outline.
[0037] It also enables multi-object synthesis: it can simultaneously handle the generation and interaction control of multiple objects, increasing the complexity and richness of scene construction. Each object is assigned independent identity and detail features, and a cross-attention mechanism is used to achieve feature isolation and association between objects. During generation, the model ensures the uniqueness of each object's identity based on its identity features, controls the object's texture, color, and other attributes using detail features, and adjusts the relative posture of objects according to the user-defined positional relationships. This enables the natural generation of multi-objective interactive scenes such as "multi-person collaborative work" and "multi-furniture combination placement," while maintaining good integration between objects and the scene, as well as between objects themselves.
[0038] Example 2
[0039] The present invention also provides an image data enhancement device based on a diffusion model and generative artificial intelligence, comprising a feature extraction module and an image enhancement module.
[0040] The feature extraction module extracts features from object images: it extracts the object's identity features and detailed features.
[0041] The image enhancement module injects dual-path identity features and detail features into the image enhancement model to achieve the image enhancement process. Identity features are injected into each layer of the UNet network through a cross-attention mechanism. In the encoder stage, identity features interact with image features through attention to generate global identity features of the object. In the decoder stage, identity constraints are further strengthened to avoid feature drift during the generation of global identity features of the object, and the global identity features are output. Detail features are then concatenated and fused with global identity features to obtain an enhanced image that maintains the texture realism of the object in different scenes. For continuous changes in object pose and lighting in dynamic scenes, a time-aware sampling strategy is adopted. Multiple frame sequences of the same object are extracted from the video, and video frames of the object are acquired by dynamically adjusting the sampling density of the time step. The sampled video frames are mixed with static object image data and input into the image enhancement model. During training, the image enhancement model learns both the static features and dynamic evolution rules of the object, enabling the model to adapt to dynamic scenes and generate enhanced images with temporal coherence and scene fusion.
[0042] The information interaction and execution process between the modules in the above-mentioned device are based on the same concept as the method embodiment of the present invention, and the specific details can be found in the description in the method embodiment of the present invention, and will not be repeated here.
[0043] Similarly, the advantages of the device of the present invention are:
[0044] 1. Excellent enhancement effect: Compared with existing image enhancement methods, the enhanced images generated based on the model used for image enhancement have higher realism and diversity. They can generate new image content that conforms to the laws of real scenes, rather than simply transforming the original image, effectively making up for the shortcomings of traditional methods.
[0045] 2. High flexibility: Through innovative mechanisms such as condition constraints and dynamic noise scheduling, it can flexibly control the features and styles of enhanced images according to different application scenarios and task requirements, realize personalized and targeted data enhancement, and improve the applicability of the technology.
[0046] 3. Easy to expand: The core framework of this technology has good scalability. It can quickly adapt to new application scenarios by adjusting the input conditions of the model and network parameters according to new image types and enhancement needs, without the need for large-scale modifications to the overall architecture.
[0047] It should be noted that not all steps and modules in the above processes and device structures are mandatory; some steps or modules can be omitted as needed. The execution order of each step is not fixed and can be adjusted as required. The system structure described in the above embodiments can be a physical structure or a logical structure. That is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or they may be jointly implemented by certain components in multiple independent devices.
[0048] The above-described embodiments are merely preferred embodiments provided to fully illustrate the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are all within the scope of protection of the present invention. The scope of protection of the present invention is defined by the claims.
Claims
1. An image data augmentation method based on a diffusion model and generative artificial intelligence, characterized by: include: Step 1: Feature extraction from the object image: Extracting the object image's identity features and detailed features. Step 2: Inject dual-path identity features and detail features into the image enhancement model to achieve the image enhancement process. Identity features are injected into each layer of the UNet network through a cross-attention mechanism. In the encoder stage, identity features interact with image features through attention to generate global identity features of the object. In the decoder stage, identity constraints are further strengthened to avoid feature drift during the generation of global identity features of the object, and the global identity features are output. Detail features are then concatenated and fused with global identity features to obtain an enhanced image that maintains texture realism in different scenes. For continuous changes in object pose and lighting in dynamic scenes, a time-aware sampling strategy is adopted. Multiple frame sequences of the same object are extracted from the video, and video frames of the object are collected by dynamically adjusting the time step sampling density. The sampled video frames are mixed with static object image data and input into the image enhancement model. During training, the image enhancement model learns both the static features and dynamic evolution patterns of the object, enabling the model to adapt to dynamic scenes and generate enhanced images with temporal coherence and scene fusion.
2. The image data enhancement method based on a diffusion model and generative artificial intelligence according to claim 1, characterized in that: Step 1 uses the self-supervised model DINO-V2 to extract identity features from object images. The object images are then processed by removing the background, centering them, and inputting them into the DINO-V2 model for global and local identity feature extraction. To extract detailed features from an image, the background is removed. A high-frequency map is extracted from the image using a high-pass filter Sobel operator. The high-frequency map is then spatially concatenated with the base background map. The concatenated composite feature is input into a pre-trained detail feature extraction network to extract hierarchical detail feature maps containing different scales.
3. The image data enhancement method based on a diffusion model and generative artificial intelligence according to claim 1, characterized in that: In step 2, when training the model for image enhancement, a training dataset is constructed. The training dataset includes the input target image, scene image and corresponding location information, and output image. Furthermore, by employing adaptive time-step sampling, the model prioritizes sampling earlier time steps for video data and later time steps for image data, enabling it to better learn the features of different types of data.
4. The image data enhancement method based on a diffusion model and generative artificial intelligence according to claim 1, characterized in that: In step 2, a loss function for the image enhancement model is designed. The structural consistency loss L1 is used to ensure that the overall structure of the generated object matches the target object; the high-frequency detail loss L1 is used to ensure that the high-frequency details of the generated object are consistent with the target object; and the adversarial loss GAN is used to improve the realism of the generated result.
5. An image data enhancement device based on a diffusion model and generative artificial intelligence, characterized in that: Including feature extraction Modules and image enhancement modules, The feature extraction module extracts features from object images: it extracts the object's identity features and detailed features. The image enhancement module injects dual-path identity features and detail features into the image enhancement model to achieve the image enhancement process. Identity features are injected into each layer of the UNet network through a cross-attention mechanism. In the encoder stage, identity features interact with image features through attention to generate global identity features of the object. In the decoder stage, identity constraints are further strengthened to avoid feature drift during the generation of global identity features of the object, and the global identity features are output. Detail features are then concatenated and fused with global identity features to obtain an enhanced image that maintains the texture realism of the object in different scenes. For continuous changes in object pose and lighting in dynamic scenes, a time-aware sampling strategy is adopted. Multiple frame sequences of the same object are extracted from the video, and video frames of the object are acquired by dynamically adjusting the sampling density of the time step. The sampled video frames are mixed with static object image data and input into the image enhancement model. During training, the image enhancement model learns both the static features and dynamic evolution rules of the object, enabling the model to adapt to dynamic scenes and generate enhanced images with temporal coherence and scene fusion.
6. The image data enhancement device based on a diffusion model and generative artificial intelligence according to claim 5, characterized in that feature extraction... The backbone network for extracting identity features from object images uses the self-supervised model DINO-V2. Background removal and center alignment are performed on the object images before they are input into the DINO-V2 model for global and local identity feature extraction. To extract detailed features from an image, the background is removed. A high-frequency map is extracted from the image using a high-pass filter Sobel operator. The high-frequency map is then spatially concatenated with the base background map. The concatenated composite feature is input into a pre-trained detail feature extraction network to extract hierarchical detail feature maps containing different scales.
7. The image data enhancement device based on a diffusion model and generative artificial intelligence according to claim 5, characterized in that: When training the image enhancement model, the image enhancement module constructs a training dataset. This dataset includes the input target image, scene image and corresponding location information, and the output image. Furthermore, by employing adaptive time-step sampling, the model prioritizes sampling earlier time steps for video data and later time steps for image data, enabling it to better learn the features of different types of data.
8. The image data enhancement device based on a diffusion model and generative artificial intelligence according to claim 5, characterized in that: The image enhancement module is designed with loss functions for the image enhancement model. It employs structural consistency loss L1 to ensure that the overall structure of the generated object matches the target object; high-frequency detail loss L1 is used to ensure that the high-frequency details of the generated object are consistent with the target object; and adversarial loss GAN is used to enhance the realism of the generated result.