Remote sensing image semantic segmentation data enhancement method and system based on controllable diffusion model

By using a controllable diffusion model-based approach, combined with a visual-language model and HED edge detection to generate visual prior features of remote sensing images, the problem of insufficient semantic consistency and pixel-level annotation in remote sensing image semantic segmentation is solved. This achieves high-quality data augmentation, improves remote sensing image segmentation performance and dataset diversity.

CN120997839APending Publication Date: 2025-11-21WUHAN UNIV

Patent Information

Application Number
CN202511041835.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing remote sensing image semantic segmentation methods suffer from insufficient semantic consistency and pixel-level annotation during data augmentation. Traditional methods cannot generate images that conform to the real scene, and the generation model lacks explicit control over the land cover categories, resulting in a deviation between the semantic distribution of the generated samples and the real data.

Method used

A method based on a controllable diffusion model is adopted, which generates text descriptions and semantic segmentation masks through a vision-language model, generates visual prior features by combining the HED edge detection algorithm, performs multi-scale controllable generation using a diffusion generation model, and generates high-resolution remote sensing images through a text-visual cross-attention mechanism, ensuring semantic consistency and pixel-level annotation of the generated images.

Benefits of technology

The generated remote sensing images significantly improve the segmentation performance and robustness of the model in semantic segmentation tasks, enhance the diversity of the dataset, reduce data annotation costs, and provide high-quality training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997839A_ABST
    Figure CN120997839A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image semantic segmentation data enhancement method based on a controllable diffusion model, and the method comprises the steps: generating a text description based on an original remote sensing image through a vision-language model, and constructing a text prompt in combination with a class label in a semantic segmentation mask; the method comprises the following steps of: generating a refined edge structure diagram through a Holistic Nested Edge Detection (HED) edge detection algorithm on the basis of an original remote sensing image, and coding the refined edge structure diagram into a visual priori feature in combination with a semantic segmentation mask; and embedding the visual priori into a diffusion generation model through a control adapter, and generating a remote sensing image in combination with text prompt. According to the method, the bottleneck of the prior art is broken through through a semantic-guided diffusion process, multi-scale controllable generation and visual priori embedding, the denoising characteristic of the diffusion model can generate a high-fidelity image, and a conditional constraint mechanism can ensure semantic consistency and physical rationality of an enhanced sample; and a better data support is provided for a remote sensing semantic segmentation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of data augmentation and image generation technology, specifically to a method and system for semantic segmentation data augmentation of remote sensing images based on a controllable diffusion model. Background Technology

[0002] Remote sensing images play a crucial role in various fields such as environmental monitoring, urban planning, and disaster management. With the continuous development of remote sensing technology, acquiring high-resolution remote sensing images globally has become easier. These images are of great significance for monitoring climate change, managing urban expansion, and assessing natural disasters. Semantic segmentation is one of the core tasks in remote sensing image analysis, aiming to assign each pixel in an image to a specific land cover category, such as buildings, roads, water bodies, and forests. However, semantic segmentation of remote sensing images faces significant challenges, especially due to the scarcity of labeled data and the high cost of high-quality labeling, which have become key factors restricting the improvement of model performance and generalization ability.

[0003] Traditional data augmentation methods, such as geometric transformations (rotation, flipping, cropping, etc.) and color transformations (brightness and contrast adjustments), can effectively increase the diversity of training data and alleviate model overfitting. However, they are essentially simple operations on the original image, and the generated new samples are highly dependent on the content of the original image, failing to introduce entirely new land cover distributions, textures, or scenes. Therefore, these traditional methods have significant limitations in enhancing the semantic diversity of images and enriching pixel-level annotations.

[0004] In recent years, generative model-based data augmentation methods have played a significant role in enhancing image diversity. However, they have struggled to effectively address the consistency issue between generated images and semantic labels in semantic segmentation tasks for remote sensing images. While these methods can generate high-quality images, they often lack pixel-level annotation information and fail to achieve a high degree of semantic consistency with the ground truth labeled images, thus limiting their application in semantic segmentation tasks. Therefore, how to combine generative models with the specific needs of semantic segmentation to generate semantically consistent remote sensing images with pixel-level annotations remains a crucial challenge in remote sensing image semantic segmentation research.

[0005] There are generally three existing methods for data augmentation in the field of remote sensing semantic segmentation: 1. Traditional methods based on basic data augmentation techniques The main principle is to alter the spatial layout and shape of the input image through geometric transformations (such as rotation, scaling, translation, and flipping), and to change the pixel intensity distribution of the input image through radiation domain transformations (such as grayscale adjustment and color space conversion), or to generate new samples by image blending (such as linear interpolation to fuse two input images). These methods expand the dataset through simple parameterization operations, aiming to improve the model's robustness to geometric deformation and illumination changes.

[0006] Method limitations: Geometric transformations only change the spatial structure of an image through global affine operations, failing to generate realistic local deformations that reflect the physical characteristics of different land cover categories in remote sensing scenes (such as the rigid structure of buildings and the topological connectivity of roads). More complex non-rigid geometric transformations, such as random elastic deformation, may lead to road network breaks or building distortions, disrupting the semantic consistency of the scene. Radiometric transformations (such as color jitter and contrast stretching) struggle to reproduce complex spectral distortions during remote sensing imaging (such as atmospheric scattering and multi-sensor differences), especially offering limited enhancement of multi-band correlation in hyperspectral images. Hybrid image techniques (such as MixUp) that fuse different samples through linear interpolation may disrupt the semantic coherence of local areas (such as the disordered overlay of water bodies and farmland), blurring the correspondence between labels and pixels and exacerbating model misjudgments of category boundaries.

[0007] 2. Methods based on imaging simulation systems The main principle is to generate precisely labeled synthetic remote sensing data through a physics engine or virtual environment, or to align the style (such as texture and color) of simulated images with real images using neural style transfer to compensate for the deficiencies of real labeled data. Its core lies in reducing the distributional differences between simulated and real data through domain adaptation techniques.

[0008] Method limitations: Physical simulation systems rely on preset imaging parameters (such as lighting models and sensor noise), making it difficult to accurately reproduce the complex, dynamically changing conditions in real scenes (such as cloud cover and seasonal changes), leading to systematic deviations in the semantic distribution between synthetic and real data. Simulated vegetation areas may exhibit unrealistic colors due to simplification of the spectral reflectance model, or building shadows may not match the actual distribution due to geometric projection errors. Neural style transfer techniques lack constraints on high-level semantic consistency, potentially causing misalignment between labels and content in the transferred image (e.g., transferring the "forest" style from simulated data to a real "farmland" area results in a mismatch between spectral features and labels). High-fidelity simulation requires constructing complex 3D scene models and rendering multi-view images, with computational costs far exceeding conventional augmentation requirements.

[0009] 3. Generative methods based on deep learning Main principle: Generative Adversarial Networks (GANs) and their variants (such as cGAN and Pix2Pix) are used to automatically learn the distribution of remote sensing data. Realistic images are synthesized through adversarial training between the generator and discriminator, or the latent feature space is reconstructed through autoencoders to implicitly enhance data diversity. Such methods aim to reduce the domain gap between real and generated images through a data-driven approach.

[0010] Method limitations: The mode collapse problem during GAN training can easily lead to insufficient diversity in generated images, such as repeatedly generating farmland or broken road structures with similar textures. This is especially true when processing high-resolution remote sensing imagery, where the fidelity of geometric details (such as building edges and road linear features) in the generated samples significantly decreases. The generation method lacks explicit control mechanisms for specific land cover categories, and the coupling between the generation process and semantic labels is weak, potentially producing semantically incorrect samples (e.g., generating pixels with vegetation spectral characteristics in a "water" labeled region). The discriminator's evaluation of generated images focuses on low-level statistical features (such as color distribution and noise patterns) rather than high-level semantic plausibility, which may generate contradictory samples that conform to the data distribution but are impossible in real-world scenarios (e.g., irregular mixing of desert and forest). Summary of the Invention

[0011] To overcome the shortcomings of the existing technologies, this invention provides a remote sensing image semantic segmentation data enhancement method and system based on a controllable diffusion model. Through semantically guided diffusion process, multi-scale controllable generation, and visual prior embedding, it breaks through the bottleneck of the existing technologies. The denoising characteristics of the diffusion model can generate high-fidelity images, while the conditional constraint mechanism can ensure the semantic consistency and physical rationality of the enhanced samples, providing better data support for remote sensing semantic segmentation tasks.

[0012] According to one aspect of the present invention, a method for semantic segmentation data augmentation of remote sensing images based on a controllable diffusion model is provided, comprising: Based on the original remote sensing images, a visual-language model is used to generate text descriptions, and text prompts are constructed by combining category labels in a semantic segmentation mask. Based on the original remote sensing image, a refined edge structure map is generated by the Holistically-Nested Edge Detection (HED) edge detection algorithm, and then combined with semantic segmentation mask encoding as visual prior features; The visual prior is embedded into a diffusion generation model through a control adapter and combined with text prompts to generate remote sensing images.

[0013] As a further technical solution, after generating the remote sensing image, the following is also included: Establish quality assessment and screening criteria for generated data, and perform post-processing optimization on the generated remote sensing images.

[0014] As a further technical solution, after acquiring the original remote sensing image, it also includes: The original remote sensing images are parsed using a visual-language model to generate text descriptions. Analyze the semantic segmentation mask of the original remote sensing image to extract the land cover category labels; The text description and category labels are combined to form a structured multi-level text prompt.

[0015] As a further technical solution, the visual prior is embedded into a diffusion generation model through a control adapter, including: Based on the visual prior of the composition, multi-scale visual features are extracted through the feature extraction layer; The extracted multi-scale visual features are fused, and the fused features are gradually injected into the diffusion generation model through a zero convolutional network. The weight distribution of the noise prediction network is dynamically adjusted by feature inverse normalization.

[0016] As a further technical solution, the method also includes: The diffusion generation model based on the encoder-decoder architecture receives multimodal conditions, aligns semantic and spatial structure constraints through a text-visual cross-attention mechanism, and performs iterative denoising in the latent space to generate high-resolution images. The multimodal conditions include text cue embedding and visual priors.

[0017] As a further technical solution, the diffusion process of the diffusion generation model is achieved through the following iterative denoising process: , in, The image features represent the current time step t. These are image features from the previous time step. It is a text prompt embedding. It is a visual prior control feature that contains information from the semantic segmentation mask and the HED edge map.

[0018] According to one aspect of the present invention, a remote sensing image semantic segmentation data enhancement system based on a controllable diffusion model is provided, comprising: The text prompt building module is used to generate text descriptions based on the original remote sensing images using a visual-language model, and to build text prompts by combining category labels in the semantic segmentation mask; The visual prior construction module is used to generate a refined edge structure map based on the original remote sensing image using the HED edge detection algorithm, and combine it with semantic segmentation mask encoding to form visual prior features; A controllable image generation module is used to embed the visual prior into a diffusion generation model through a control adapter and combine it with text prompts to generate remote sensing images.

[0019] As a further technical solution, it also includes: The post-processing optimization module is used to establish quality assessment and screening criteria for generated data and to perform post-processing optimization on the generated remote sensing images.

[0020] According to one aspect of the present invention, an electronic device is provided, including a memory and a processor, the memory storing program instructions executable by the processor, the processor invoking the program instructions to execute the remote sensing image semantic segmentation data enhancement method based on a controllable diffusion model.

[0021] According to one aspect of the present invention, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium storing computer instructions that cause the computer to execute the remote sensing image semantic segmentation data enhancement method based on a controllable diffusion model.

[0022] The image enhancement method provided by this invention has significant beneficial effects, especially in improving the accuracy and robustness of semantic segmentation of remote sensing images, enhancing the diversity of datasets, and reducing data annotation costs. Compared with existing technologies, the beneficial effects of this invention are as follows: 1. Improved Segmentation Performance: Experimental results show that the augmented data generated using the method of this invention can significantly improve the segmentation performance of the semantic segmentation model on multiple types of land features. Due to the improved semantic consistency and pixel-level label accuracy of the generated images, the trained model has significantly improved the classification accuracy of various land features (such as buildings, roads, forests, etc.).

[0023] 2. Enhanced Data Diversity: By generating semantically consistent and stylistically diverse images, this invention effectively increases the diversity of training data. Enhanced images of different styles provide the model with more training samples, thereby improving the model's generalization ability and preventing overfitting in practical applications.

[0024] 3. Reduced annotation costs: The enhanced images generated by this invention can effectively expand the training set and reduce the need for manual annotation. By generating high-quality images with pixel-level labels, this invention not only reduces the time and cost of data annotation but also increases the size and diversity of the dataset, providing richer input data for model training. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a schematic diagram of the data enhancement method for semantic segmentation of remote sensing images based on a controllable diffusion model, provided in an embodiment of the present invention.

[0027] Figure 2 This is a diagram showing the result of remote sensing image generation provided in an embodiment of the present invention.

[0028] Figure 3 A comparison chart showing the improvement in semantic segmentation accuracy based on original data and generated data, provided for embodiments of the present invention.

[0029] Figure 4 This is a framework diagram of a remote sensing image semantic segmentation data enhancement method based on a controllable diffusion model, provided in an embodiment of the present invention.

[0030] Figure 5 A schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0031] The terms “comprising” and “having”, and any variations thereof, in the specification, claims, and accompanying drawings of this invention are intended to cover a non-exclusive inclusion, such as a process, method, system, product, or apparatus that includes a series of steps or units, not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In addition, the technical features of the various embodiments or individual embodiments provided by the present invention can be arbitrarily combined to form new technical solutions. Such combinations are not bound by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0033] To address the limitations of existing data augmentation methods in the field of remote sensing semantic segmentation, this invention designs a multimodal data augmentation framework driven by a controllable diffusion model. This framework solves the coupling problem between semantic consistency maintenance, edge detail preservation, and annotation inheritance in existing remote sensing image generation methods. It constructs a multi-source conditional injection mechanism based on semantic segmentation masks, edge maps, and enhanced text prompts to generate diverse remote sensing image samples with pixel-level annotation inheritance. This effectively overcomes the limitations of traditional data augmentation methods, such as a single semantic structure and missing annotations in the generation model. It also enriches the scene coverage and object morphology diversity of the training dataset, improving the recognition accuracy and generalization performance of semantic segmentation models for multi-scale ground objects in complex geographical scenes.

[0034] Please see Figure 1 This invention proposes a remote sensing image semantic segmentation data enhancement method based on a controllable diffusion model. First, based on the original remote sensing image, a visual-language model is used to generate a text description, and a text prompt is constructed by combining the category labels in the semantic segmentation mask. Next, based on the original remote sensing image, a refined edge structure map is generated using the HED edge detection algorithm, and encoded into visual prior features by combining the semantic segmentation mask. Finally, the visual prior is embedded into the diffusion generation model through a control adapter, and combined with the text prompt to generate the remote sensing image.

[0035] The method described in this embodiment of the invention, by integrating text prompts, visual priors, and a controllable diffusion model, can effectively generate semantically consistent remote sensing images with pixel-level labels, thereby providing high-quality training data for semantic segmentation tasks of remote sensing images. Specifically, it includes the following steps: Step 1, Text Hint Construction. This step designs a two-stage text generation mechanism based on the BLIP-2 visual-language model and semantic label fusion to address the problem of insufficient semantic coverage in generated hints. By parsing the segmentation mask of the original image to extract land cover category labels, explicit category description phrases are constructed and concatenated with the scene description text generated by the BLIP-2 model to form a structured multi-level text hint. This constrains the global semantic distribution and local category attributes of the generated image, ensuring accurate generation control of the diffusion model for multi-category targets.

[0036] Specifically, the BLIP-2 vision-language model is used to generate text descriptions for remote sensing images. Simultaneously, to improve the semantic consistency of the generated images, category information from the semantic segmentation mask is combined to construct text prompts. The formula for constructing text prompts is as follows:

[0037] in, It is a class hint that includes all semantic categories in the image, such as "Buildings; Bareland; Road", while This is the image description text generated by the BLIP-2 model. The final text prompt. This is a combination of the two parts mentioned above, which ensures that the generated image is semantically consistent with the original remote sensing image.

[0038] Step 2, Visual Prior Integration. This step constructs a multi-source visual conditional injection module based on semantic segmentation masks and HED edge maps, overcoming the challenge of aligning the generated image structure with the annotations. The HED network is used to extract edge structure priors from the original image, which are then combined with a semantic mask encoder to generate multi-scale feature maps. The visual priors are embedded into a diffusion model through a control adapter, and the global distribution constraints of the semantic mask are progressively fused with the detailed features of the HED edges using zero-convolutional layers. This achieves coordinated optimization of pixel-level annotation inheritance, edge sharpness, and scene complexity in the generated image.

[0039] Specifically, to enhance the structural and semantic consistency of the image generation process, this embodiment uses HED edge maps and semantic segmentation masks as visual priors. These visual priors are integrated into each stage of the generation process through a control adapter, ensuring that the generated image maintains consistency with the real image in terms of structural detail and semantic information. The specific integration process can be represented as follows:

[0040] in This is the noise characteristic of the m-th layer in the diffusion model. The control features are extracted at the m-th layer. and These are normalized parameters learned through a multilayer perceptron. These control features are progressively injected into the generation process of the diffusion model. Through a multi-scale feature fusion strategy, visual prior information is effectively passed to each layer of the model, ensuring that the generated images are consistent in detail and semantics.

[0041] Step 3, Controlled Diffusion Generation. This step designs a controlled generation process based on Stable Diffusion 1.5 to address the problem of poor adaptability of traditional generation models to remote sensing scene structures. During the denoising process of the diffusion model, semantic embeddings of text prompts and multi-scale features of visual priors (covering pixel-level spatial constraints of segmentation masks) are injected simultaneously. The weights of the noise prediction network are dynamically modulated through the Feature Denormalization (FDN) module to achieve high-resolution remote sensing image generation under multiple conditions of text-visual-spatial coupling, ensuring that the generated samples have texture diversity on the basis of annotation inheritance.

[0042] Specifically, in the generation stage, an image is generated using a method based on the Stable Diffusion model. The diffusion process is accomplished through iterative denoising, and the specific update process is as follows:

[0043] in, The image features represent the current time step t. These are image features from the previous time step. It is a text prompt embedding. These are visual prior control features, containing information from the semantic segmentation mask and the HED edge map. The subscript represents the denoising function of the diffusion model. This represents the optimization parameters. At each time step, the diffusion model progressively generates remote sensing images that meet semantic and spatial structure requirements through denoising and adjustment of control features.

[0044] Step 4, Post-processing Optimization. This step establishes quality assessment and screening criteria for the generated data, addressing issues of generation noise and annotation offset. It also introduces an expert verification process, manually reviewing the generated results for complex scenarios (such as dense building clusters and intersecting roads) to ensure the reliability of the augmented data's annotations and construct a high-quality augmented dataset that can be directly used for training semantic segmentation models.

[0045] To ensure the quality of the generated images, the images generated in this embodiment will be optimized through a post-processing step. During post-processing, a manual inspection criterion Q is used to filter out images with high noise or poor quality. Specifically, the optimized image... This can be expressed by the following formula:

[0046] in, This is the optimized image. It generates images. This indicates a post-processing operation, where Q is a human quality standard. This process ensures that the final output image is suitable for subsequent semantic segmentation training, removing unqualified or excessively noisy images.

[0047] Through the above steps, the method of this embodiment of the invention achieves the generation of semantically consistent and pixel-level accurately labeled remote sensing images. The resulting remote sensing image is shown in the figure below. Figure 2 As shown.

[0048] Figure 2This figure demonstrates the generation results of the method of this invention under multi-condition coupling of text prompts, semantic segmentation masks, and edge maps. The first column on the left of the figure represents pixel-level semantic constraints; the second column is a grayscale edge map, preserving structural details such as road outlines and building boundaries; the third column is text prompts, providing multi-level text descriptions that clearly define the generation target and attribute requirements; the right-hand column of generated images shows the model output results, which strictly follow the spatial distribution of the semantic mask and the semantic description of the text prompts. For example, in the "field road" scene, the generated image accurately aligns with the road's central axis and renders the texture of farmland covered with green vegetation; the rightmost column shows examples of original satellite images, covering three typical scenes: urban, suburban, and farmland. This figure verifies the model's ability to generate high-fidelity, highly consistent remote sensing images under multi-modal input collaboration, providing a reliable sample source for semantic segmentation data augmentation.

[0049] The method of this invention significantly improves the accuracy and robustness of semantic segmentation tasks for remote sensing images, as shown in the visualization results. Figure 3 As shown, this demonstrates that the method of the present invention can effectively expand the diversity of training datasets.

[0050] Figure 3 This paper demonstrates the performance optimization results of applying the high-quality augmented data generated by the method of this invention to a semantic segmentation task. The three horizontal sets of examples in the figure correspond to three typical geographical regions: Chiangmai, Aachen, and Les Cayes. The four vertical columns present the original RGB image, ground truth annotations, baseline model predictions, and the segmentation results trained with augmented data. The comparison shows that the augmented prediction results significantly outperform the baseline model in terms of road continuity, building outline accuracy, and distinction between vegetation and farmland boundaries, verifying the role of the generated data in improving the segmentation accuracy of complex scenes.

[0051] The implementation of the various embodiments of the present invention is based on programmed processing by a device with processor functionality. Therefore, in practical engineering, the technical solutions and functions of the various embodiments of the present invention are encapsulated into various modules. Based on this reality, and building upon the above embodiments, the embodiments of the present invention provide a remote sensing image semantic segmentation data enhancement system based on a controllable diffusion model. This system is used to execute a remote sensing image semantic segmentation data enhancement method based on a controllable diffusion model from the above method embodiments.

[0052] See Figure 4The system includes: a text prompt construction module, used to generate text descriptions based on the original remote sensing image using a visual-language model, and construct text prompts by combining category labels in a semantic segmentation mask; a visual prior construction module, used to generate a refined edge structure map based on the original remote sensing image using the HED edge detection algorithm, and encode it into visual prior features by combining it with a semantic segmentation mask; and a controllable image generation module, used to embed the visual priors into a diffusion generation model through a control adapter, and generate remote sensing images by combining them with text prompts.

[0053] This invention provides a remote sensing image semantic segmentation data augmentation system based on a controllable diffusion model. Addressing the limitations of existing data augmentation methods in the field of remote sensing semantic segmentation, this system employs... Figure 4 The framework in this paper breaks through the bottleneck of existing technologies through semantically guided diffusion process, multi-scale controllable generation and visual prior embedding. The denoising characteristics of the diffusion model can generate high-fidelity images, while the condition constraint mechanism can ensure the semantic consistency and physical rationality of enhanced samples, providing better data support for remote sensing semantic segmentation tasks.

[0054] Figure 4 This paper presents a framework diagram of the remote sensing image semantic segmentation data enhancement system based on a controllable diffusion model proposed in this invention. The diagram is divided into three functional areas from left to right, corresponding to the input processing, multimodal conditional fusion, and output generation stages, respectively. The left input area displays the original remote sensing image sample, which is parsed by the BLIP-2 vision-language model to generate scene description text. Simultaneously, category labels from the semantic mask are extracted, and the two are concatenated to form a structured multi-level text prompt. The middle area is the core conditional generation module. It generates a refined edge structure map from the original image using a HED edge detection network, which is then encoded into visual prior features using the semantic mask. After extracting multi-scale visual features through a feature extraction layer, a control adapter module is used to achieve feature fusion. During this process, features are progressively injected into the diffusion model through a zero-convolutional network, and feature denormalization dynamically adjusts the weight distribution of the noise prediction network, ensuring that the global constraints of the semantic mask and the detailed features of the edge structure work synergistically during the generation process. The right side represents the controllable generation module. Based on an encoder-decoder architecture, the diffusion generation model receives multimodal conditional inputs (textual description embeddings and visual prior features). It aligns semantic and spatial structural constraints through a text-visual cross-attention mechanism, performing iterative denoising within the latent space to generate high-resolution images. The data flow diagram clearly illustrates the end-to-end process from original image input, text and visual prior construction, feature fusion to generated image output. The generated results strictly align with the semantic information of the mask annotations and the structural information of the edge map, and possess diverse texture details.

[0055] It should be noted that the system embodiments provided by the present invention are used not only to implement the methods in the above method embodiments, but also to implement the methods in other method embodiments provided by the present invention. The only difference is that corresponding functional modules are set. The principle is basically the same as that of the above system embodiments provided by the present invention. As long as those skilled in the art can improve the modules in the above system embodiments by referring to the specific technical solutions in other method embodiments and combining technical features to obtain corresponding technical means and technical solutions composed of these technical means, on the basis of the above system embodiments, and on the premise of ensuring the practicality of the technical solutions, they can obtain corresponding system-like embodiments for implementing the methods in other method-like embodiments.

[0056] The method in this embodiment of the invention is implemented using an electronic device; therefore, it is necessary to introduce the relevant electronic device. For this purpose, embodiments of the present invention provide an electronic device, such as... Figure 5 As shown, the electronic device includes: at least one processor, a communication interface, at least one memory, and a communication bus, wherein the at least one processor, the communication interface, and the at least one memory communicate with each other via the communication bus. The at least one processor invokes logical instructions stored in the at least one memory to execute all or part of the steps of the methods provided in the foregoing method embodiments.

[0057] Furthermore, when the logical instructions in at least one of the aforementioned memories are implemented as software functional units and sold or used as independent products, they are stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, is embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (a personal computer, server, or network device) to execute all or part of the steps of the methods described in the various method embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks—various media for storing program code.

[0058] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, located in one place, or distributed across multiple network units. The purpose of this embodiment is achieved by selecting some or all of the modules according to actual needs. Those skilled in the art will understand and implement this without any inventive effort.

[0059] Based on the same inventive concept as the foregoing embodiments, this embodiment of the invention also provides a non-transitory computer-readable storage medium storing computer instructions that cause the computer to execute the remote sensing image semantic segmentation data enhancement method based on the controllable diffusion model. By integrating text prompts, visual priors, and the controllable diffusion model, it can effectively generate semantically consistent remote sensing images with pixel-level labels, thereby providing high-quality training data for the semantic segmentation task of remote sensing images.

[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A remote sensing image semantic segmentation data enhancement method based on a controllable diffusion model, characterized in that, The method comprises the following steps: Based on the original remote sensing image, a text description is generated using a visual-linguistic model, and a text prompt is constructed by combining the class labels in the semantic segmentation mask; Based on the original remote sensing image, a fine edge structure map is generated by the HED edge detection algorithm, and a visual prior feature is encoded by combining the semantic segmentation mask; The visual prior is embedded into the diffusion generation model through a control adapter, and a semantically consistent remote sensing image is generated by combining the text prompt.

2. The method of claim 1, wherein, After generating the remote sensing image, the method further comprises the following steps: Establishing generation data quality evaluation and screening standards to post-process and optimize the generated remote sensing image.

3. The method of claim 1, wherein, After obtaining the original remote sensing image, the method further comprises the following steps: The original remote sensing image is parsed to generate a text description using a visual-linguistic model; The semantic segmentation mask of the original remote sensing image is parsed to extract the class labels of the ground objects; The text description and the class labels are spliced to form a structured multi-level text prompt.

4. The method of claim 1, wherein, The visual prior is embedded into the diffusion generation model through a control adapter, which comprises the following steps: Based on the constructed visual prior, multi-scale visual features are extracted through a feature extraction layer; The extracted multi-scale visual features are fused, and the fused features are gradually injected into the diffusion generation model through a zero convolution network, and the weight distribution of the noise prediction network is dynamically adjusted using feature de-normalization.

5. The method of claim 4, wherein, The method further comprises the following steps: The diffusion generation model based on the encoder-decoder architecture receives multi-modal conditions, aligns the semantic and spatial structure constraints through a text-visual cross-attention mechanism, and generates a high-resolution image through iterative denoising in the latent space, wherein the multi-modal conditions include text prompt embedding and visual prior.

6. The method of claim 1, wherein, The diffusion process of the diffusion generation model is realized through the following iterative denoising process: , wherein, represents the image feature of the current time step t, is the image feature of the previous time step, is the text prompt embedding, is the control feature of the visual prior, containing the information of the semantic segmentation mask and the HED edge map.

7. A remote sensing image semantic segmentation data augmentation system based on a controllable diffusion model, characterized in that, The method comprises the following steps: A text prompt construction module is used to generate a text description based on the original remote sensing image using a visual-linguistic model, and a text prompt is constructed by combining the class labels in the semantic segmentation mask; A visual prior construction module is used to generate a fine edge structure map based on the original remote sensing image using the HED edge detection algorithm, and a visual prior feature is encoded by combining the semantic segmentation mask; A controllable image generation module is used to embed the visual prior into the diffusion generation model through a control adapter, and generate a remote sensing image by combining the text prompt.

8. The system according to claim 7, wherein, The method further comprises the following steps: A post-processing optimization module is used to establish generation data quality evaluation and screening standards to post-process and optimize the generated remote sensing image.

9. An electronic device, comprising: The memory stores program instructions that are executed by the processor, and the processor invokes the program instructions to execute the remote sensing image semantic segmentation data enhancement method based on the controllable diffusion model according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, comprising: The non-transitory computer-readable storage medium stores computer instructions that cause the computer to execute the remote sensing image semantic segmentation data enhancement method based on the controllable diffusion model according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image description generation method and device, electronic equipment and storage medium

    CN118736243A

  • Remote sensing image generation method and device based on multi-condition controllable diffusion model

    CN118982597A

  • Remote sensing image data augmentation method and device based on diffusion generation model

    CN119723247A

  • AI generative model training method, image generation method and electronic equipment

    CN119963941A

  • Training and deployment of image generation models

    US11995803B1

Cited By

  • Semantic-driven image reconstruction method and device based on edge and color assistance

    CN121239857A

  • Small-sample target remote sensing image generation method based on structure perception and detail enhancement

    CN122133737A

  • Few-shot Target Remote Sensing Image Generation Method Based on Structure Awareness and Detail Enhancement

    CN122133737B