A Multi-modal Sample Generation Method Based on a Geographical Information Conditional Diffusion Model

Through the multimodal sample generation method of the geographic information conditional diffusion model, the problem of lack of text description of remote sensing image annotation data is solved, low-cost and efficient data set generation is achieved, and the adaptability and accuracy of the model is improved.

CN119180998BActive Publication Date: 2025-07-08CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411001511.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2025-07-08
Estimated Expiration
2044-07-25

AI Technical Summary

Technical Problem

The data samples of remote sensing image labeling lack text description information, and traditional image enhancement technology cannot fully simulate the sample diversity of remote sensing images, resulting in high cost and inefficiency in data set construction.

Method used

Using a geographic information conditional diffusion model, multi-source images and semantic labels are channel-stitched through joint input modules, and noise injection-noise reduction process and data decomposition are combined with data chimerization modules to generate multi-modal samples.

Benefits of technology

Diversity and richness of sample data are generated, which reduces the cost of remote sensing image annotation, improves the quality and efficiency of the data set, and enhances the generalization ability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119180998B_ABST
    Figure CN119180998B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-modal sample generation method based on a geographic information conditional diffusion model, which relates to the technical field of remote sensing image generation. The method includes: obtaining a remote sensing multi-modal semantic segmentation data set, including multi-source images and semantic labels corresponding to the multi-source images; inputting the remote sensing multi-modal semantic segmentation data set into a geographic information conditional diffusion model for training to obtain a plurality of multi-modal samples; wherein, the geographic information conditional diffusion model includes a joint input module and a data embedding module; through the joint input module, the multi-source images and semantic labels in the remote sensing multi-modal semantic segmentation data set are subjected to channel splicing processing to obtain target image label data; through the data embedding module, the target image label data is subjected to a noise injection-denoising process and data decomposition processing to obtain a plurality of multi-modal samples. The present invention solves the problem that the labeled data samples of remote sensing images usually lack text description information, and realizes the generation of data samples at low cost and high efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image generation, and in particular, to a multi-modal sample generation method, device, electronic device and storage medium based on a geographic information conditional diffusion model. Background Art

[0002] As an important part of the earth's resources, mines are crucial for the development of social economy. However, mining activities may be accompanied by environmental changes and ecosystem damage. Therefore, monitoring and managing the land cover changes in mining areas is crucial for ecological, environmental and social development. Remote sensing images are one of the important data resources for recording surface information. Through deep learning methods, remote sensing images can be automatically interpreted and analyzed, and are applicable to applications such as scene classification, semantic segmentation and instance segmentation, providing an effective tool for mine monitoring and management.

[0003] However, constructing a pre-trained model for remote sensing images requires a large amount of labeled data sets, and labeling remote sensing images is an expensive and time-consuming task. Using traditional image enhancement techniques may not be able to fully simulate the sample diversity of remote sensing images. Therefore, a common solution is to use data augmentation to increase sample diversity and make full use of existing labels. However, traditional image enhancement techniques (such as flipping, rotating and scaling, etc.) are usually not sufficient to simulate the sample diversity of satellite images. Summary of the Invention

[0004] The problem solved by the present invention is how to solve the problem that the labeled data samples of remote sensing images usually lack text description information.

[0005] To solve the above problems, the present invention provides a multi-modal sample generation method, device, electronic device and storage medium based on a geographic information conditional diffusion model.

[0006] In a first aspect, the present invention provides a multi-modal sample generation method based on a geographic information conditional diffusion model, including:

[0007] Obtain a remote sensing multi-modal semantic segmentation data set, where the remote sensing multi-modal semantic segmentation data includes multi-source images and semantic labels corresponding to the multi-source images;

[0008] Input the remote sensing multi-modal semantic segmentation data set into a geographic information conditional diffusion model for training to obtain a plurality of multi-modal samples;

[0009] Wherein, the geographic information conditional diffusion model includes a joint input module and a data embedding module;

[0010] Through the joint input module, perform channel splicing processing on the multi-source images and semantic labels in the remote sensing multi-modal semantic segmentation data set to obtain target image label data;

[0011] Through the data embedding module, perform noise injection - noise reduction process and data decomposition processing on the target image label data to obtain a plurality of the multi - modal samples.

[0012] Optionally, the obtaining of the remote sensing multi - modal semantic segmentation data set includes:

[0013] Obtain the image of the current study area and the fine - classification annotation data of the mines in the current study area;

[0014] According to the image of the current study area and the fine - classification annotation data of the mines, obtain a plurality of the multi - source images and a plurality of the semantic labels;

[0015] According to a plurality of the multi - source images and a plurality of the semantic labels, obtain the remote sensing multi - modal semantic segmentation data set.

[0016] Optionally, the process of performing channel splicing on the multi - source images and semantic labels in the remote sensing multi - modal semantic segmentation data set through the joint input module to obtain the target image label data includes:

[0017] Perform normalization processing on the multi - source images to obtain normalized images;

[0018] Map the discrete label values corresponding to the semantic labels into binary vector representations to obtain binary labels;

[0019] Perform splicing processing on the normalized images and the binary labels in the channel dimension to obtain the target image label data.

[0020] Optionally, the process of performing noise injection - noise reduction process and data decomposition processing on the target image label data through the data embedding module to obtain a plurality of the multi - modal samples includes:

[0021] Obtain the geographical information data and terrain surface data of the current study area;

[0022] Perform noise injection - noise reduction process on the target image label data through the geographical information data and the terrain surface data to obtain image - label joint data;

[0023] Perform data decomposition on the image - label joint data to obtain a plurality of the multi - modal samples.

[0024] Optionally, the obtaining of the geographical information data and terrain surface data of the current study area includes:

[0025] Obtain the basic digital elevation model data of the current study area, and obtain the terrain surface data according to the basic digital elevation model data, where the terrain surface data includes slope data, aspect data, and mountain shadow data;

[0026] According to the remote sensing multi-modal semantic segmentation data set, obtain geographical metadata and time step data, linearly embed the geographical metadata and the time step data respectively, and add them to obtain the geographical information data.

[0027] Optionally, the process of injecting noise and denoising the target image label data through the geographical information data and the terrain surface data to obtain the image-label joint data includes:

[0028] Use the geographical information data as conditional information, and use the terrain surface data as a supplement to the conditional information to perform a noise injection-denoising process on the target image label data to obtain the image-label joint data.

[0029] Optionally, the process of decomposing the image-label joint data to obtain multiple multi-modal samples includes:

[0030] Perform inverse binary coding on the latter C bands in the image-label joint data and restore them to obtain multiple multi-modal samples, where C represents the number of categories of the image-label joint data, and \(\log_2\) represents the logarithmic function with base 2.

[0031] In a second aspect, the present invention provides a multi-modal sample generation device based on a geographical information conditional diffusion model, including:

[0032] An acquisition unit for acquiring a remote sensing multi-modal semantic segmentation data set, where the remote sensing multi-modal semantic segmentation data includes multi-source images and semantic labels corresponding to the multi-source images;

[0033] A training unit for inputting the remote sensing multi-modal semantic segmentation data set into a geographical information conditional diffusion model for training to obtain multiple multi-modal samples;

[0034] Wherein, the geographical information conditional diffusion model includes a joint input module and a data embedding module;

[0035] Through the joint input module, perform channel splicing on the multi-source images and semantic labels in the remote sensing multi-modal semantic segmentation data set to obtain target image label data;

[0036] Through the data embedding module, perform a noise injection-denoising process and data decomposition process on the target image label data to obtain multiple multi-modal samples.

[0037] In a third aspect, the present invention provides an electronic device, including a memory and a processor;

[0038] The memory is used to store a computer program;

[0039] The processor is configured to, when executing the computer program, implement the multi-modal sample generation method based on a geographic information conditional diffusion model as described in the first aspect.

[0040] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the multi-modal sample generation method based on a geographic information conditional diffusion model as described in the first aspect is implemented.

[0041] The beneficial effects of the multi-modal sample generation method, device, electronic device, and storage medium based on the geographic information conditional diffusion model of the present invention are as follows: First, the present invention obtains a data set containing multiple remote sensing image modalities and their corresponding semantic labels. It can be understood that the images can be obtained by different sensors, such as optical images, hyperspectral data, radar data, etc., and the semantic labels correspond to the ground object category information of each pixel. The obtained remote sensing multi-modal semantic segmentation data set is input into the geographic information conditional diffusion model for training. The model includes a joint input module and a data embedding module. The geographic information conditional diffusion model is designed to generate multiple multi-modal samples using the geographic information conditional diffusion algorithm. Among them, the joint input module is responsible for performing channel splicing processing on the multi-source images and semantic labels in the remote sensing multi-modal semantic segmentation data set, integrating the information of different input sources together to obtain target image label data, so that the model can better understand the image features and ground object category information. The data embedding module is responsible for performing a noise injection-denoising process and data decomposition processing on the target image label data to generate multiple multi-modal samples. The noise injection-denoising process can help reduce the interference information in the data, making the samples clearer and more accurate. Data decomposition helps extract different features from the original data to generate diverse sample data. Through the method of the present invention, the geographic information conditional diffusion model can learn the ground object category information from the remote sensing multi-modal semantic segmentation data set, and combine the geographic information conditional diffusion algorithm to generate multiple multi-modal samples. Taking the multi-source images and semantic labels as inputs and using the joint input module and the data embedding module for processing, more diverse and rich sample data can be generated, providing more comprehensive information and more choices for remote sensing image annotation, solving the problem that remote sensing image annotation data samples usually lack text description information, and realizing the generation of data samples with low cost and high efficiency. Description of the Drawings

[0042] Figure 1 It is a flowchart showing a multi-modal sample generation method based on a geographic information conditional diffusion model according to an embodiment of the present invention;

[0043] Figure 2 A structural schematic diagram of a geographic information condition diffusion model according to an embodiment of the present invention;

[0044] Figure 3 A structural schematic diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0045] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments described herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0046] It should be understood that the various steps described in the method embodiments of the present invention can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.

[0047] The term "including" and its variants used herein are open-ended, that is, "including but not limited to"; the term "based on" is "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiment". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules, or units.

[0048] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0049] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0050] As Figure 1 shown, a multi-modal sample generation method based on a geographic information condition diffusion model provided by an embodiment of the present invention includes:

[0051] Step S1, obtain a remote sensing multimodal semantic segmentation dataset, where the remote sensing multimodal semantic segmentation data includes multi-source images and semantic labels corresponding to the multi-source images.

[0052] Specifically, obtain a dataset containing multiple remote sensing image modalities and their corresponding semantic labels. It can be understood that the images can be obtained by different sensors, such as optical images, hyperspectral data, radar data, etc., and the semantic labels correspond to the ground object category information of each pixel.

[0053] Step S2, input the remote sensing multimodal semantic segmentation dataset into a geographic information conditional diffusion model for training to obtain multiple multimodal samples.

[0054] Specifically, in this embodiment, through the geographic information conditional diffusion model, channel splicing processing can be performed using multi-source images and semantic labels to generate multiple multimodal samples, thereby enriching data diversity, improving the richness and diversity of samples, and enhancing the generalization ability of the model.

[0055] Step S3, where the geographic information conditional diffusion model includes a joint input module and a data embedding module;

[0056] Perform channel splicing processing on the multi-source images and semantic labels in the remote sensing multimodal semantic segmentation dataset through the joint input module to obtain target image label data;

[0057] Perform a noise injection - denoising process and data decomposition processing on the target image label data through the data embedding module to obtain multiple of the multimodal samples.

[0058] Specifically, the data embedding module in this embodiment can perform a noise injection - noise reduction process and data decomposition processing on the target image label data to optimize the data quality, improve the usability and effectiveness of the samples. This embodiment can better integrate and process multi-modal data according to geographical information conditions, improve the model's understanding and expression ability of geographical information and surface features, and provide a more effective solution for information processing in a specific geographical environment. The multi-modal sample generation method based on the geographical information condition diffusion model in this embodiment can effectively improve data diversity, optimize data quality, and improve the model's understanding and utilization ability of geographical information and multi-modal data, making the model more adaptable to information processing and analysis in a specific geographical environment. Among them, during the process of generating images, adding random noise to the model input helps to improve the model's robustness and generalization ability, helps the model better capture the subtle features in the data, and makes the model more interpretable. It can be understood that in the actual application process, the noise addition methods include but are not limited to introducing random perturbations in the input data, such as Gaussian noise, impulse noise, or masked noise, etc. The purpose is to enable the model to better learn the distribution of the data, thereby generating more diverse and realistic images. After generating the images, it is necessary to perform denoising processing on the generated images to reduce the impact of the noise introduced during the generation process on the image quality, improve the clarity and quality of the generated images, and make them more in line with the visual characteristics of the real world. It can be understood that the denoising methods include but are not limited to using neural network models, filters, or denoising algorithms. The above methods can effectively remove the noise in the images and improve the visual quality of the images.

[0059] Specifically, the joint input module of this embodiment can perform channel splicing processing on the multi-source images and semantic labels in the remote sensing multi-modal semantic segmentation dataset, thereby realizing the integration of multi-source data, enabling the model to obtain rich information from different perspectives, and improving the variation and diversity of the data. By performing a noise injection - noise reduction process and data decomposition processing on the target image label data through the data embedding module, the robustness of the data can be enhanced, the interference of noise on model modeling can be reduced, and the generalization ability of the model can be improved. Through this embodiment, samples with various diversity characteristics can be generated, better simulating the diversity of actual remote sensing images, improving the diversity and complexity of the training data, and providing better conditions for the training of deep learning models.

[0060] In this embodiment, a dataset containing multiple remote sensing image modalities and their corresponding semantic labels is first obtained. It can be understood that the images can be acquired by different sensors, such as optical images, hyperspectral data, radar data, etc., and the semantic labels correspond to the land cover class information of each pixel. The obtained remote sensing multi-modal semantic segmentation dataset is input into a geographic information conditional diffusion model for training. This model includes a joint input module and a data embedding module. The geographic information conditional diffusion model aims to generate multiple multi-modal samples using the geographic information conditional diffusion algorithm. Among them, the joint input module is responsible for performing channel splicing on the multi-source images and semantic labels in the remote sensing multi-modal semantic segmentation dataset, integrating the information from different input sources together to obtain target image label data, enabling the model to better understand the image features and land cover class information. The data embedding module is responsible for performing a noise injection-denoising process and data decomposition on the target image label data to generate multiple multi-modal samples. The noise injection-denoising process can help reduce the interference information in the data, making the samples clearer and more accurate. Data decomposition helps extract different features from the original data to generate diverse sample data. Through the method of this embodiment, the geographic information conditional diffusion model can learn the land cover class information from the remote sensing multi-modal semantic segmentation dataset, and combine the geographic information conditional diffusion algorithm to generate multiple multi-modal samples. Taking the multi-source images and semantic labels as inputs, and using the joint input module and the data embedding module for processing, more diverse and rich sample data can be generated, providing more comprehensive information and more choices for remote sensing image annotation, solving the problem that the remote sensing image annotation data samples usually lack text description information, and realizing the generation of data samples with low cost and high efficiency.

[0061] Optionally, the obtaining of the remote sensing multi-modal semantic segmentation dataset includes:

[0062] Obtain the image of the current study area and the fine classification annotation data of mines in the current study area;

[0063] According to the image of the current study area and the fine classification annotation data of mines, obtain multiple of the multi-source images and multiple of the semantic labels;

[0064] According to multiple of the multi-source images and multiple of the semantic labels, obtain the remote sensing multi-modal semantic segmentation dataset.

[0065] Specifically, in this embodiment, by obtaining the current study area image and the mine fine classification annotation data, multi-source images and semantic labels are obtained, enabling the model to fully understand the information of different spectral bands. At the same time, combined with the semantic information of the ground objects, it provides richer data features for the model, including spectral, spatial, and semantic information, which helps to improve the recognition and interpretation ability of the deep learning model for ground objects. By using the multi-modal semantic segmentation data set in this embodiment, the information of ground objects, including their spatial distribution and semantic information, can be better mined, providing more accurate data support for the segmentation and recognition tasks of mine ground objects and improving the precise classification and recognition ability of mine ground objects. The method of this embodiment can handle the complex and diverse features of remote sensing data, providing more comprehensive and diverse data features for the deep learning model, overcoming the limitations of traditional single-band or specific images, making better use of multi-modal information for ground object recognition and segmentation, and improving the performance of the deep learning model. Thus, it provides a data basis for generating more effective sample data.

[0066] Optionally, the process of performing channel splicing on the multi-source images and semantic labels in the remote sensing multi-modal semantic segmentation data set through the joint input module to obtain target image label data includes:

[0067] Performing normalization processing on the multi-source images to obtain normalized images;

[0068] Mapping the discrete label values corresponding to the semantic labels into binary vector representations to obtain binary labels;

[0069] Performing splicing processing on the normalized images and the binary labels in the channel dimension to obtain target image label data.

[0070] Specifically, in this embodiment, by standardizing the range of image data, that is, performing standardization processing on multi-source images to obtain a standardized image with pixel values in the range of [0, 1], and mapping semantic labels to binary vectors, the problem of lack of descriptive information in remote sensing image annotation data is overcome, and the processing ability of the deep learning model for remote sensing data is improved. Among them, standardizing the range of image data helps to improve the stability of training, mitigate the influence of different data ranges, enables the model to more easily capture the main features of the data, increases the learning speed, and thus improves the sample generation efficiency; mapping semantic labels to binary vectors enables the model to better understand the specific ground object information at the pixel level, provides richer information, and improves the recognition and segmentation accuracy of the model for ground objects; through channel splicing processing, the standardized image and the binary label are spliced in the channel dimension to provide richer training data and improve the generalization ability and accuracy of the deep learning model. The method of this embodiment improves the processing effect of the deep learning model on remote sensing images by performing standardization processing and vectorization processing on remote sensing images and semantic labels, and reduces the problem of lack of descriptive information in annotation data.

[0071] Optionally, the process of injecting noise - denoising and data decomposition on the target image label data by the data chimeric module to obtain a plurality of the multimodal samples includes:

[0072] Obtain the geographic information data and terrain surface data of the current study area;

[0073] Perform noise injection - denoising process on the target image label data through the geographic information data and the terrain surface data to obtain image - label joint data;

[0074] Decompose the image - label joint data to obtain a plurality of the multimodal samples.

[0075] Specifically, the noise injection - noise reduction process in this embodiment can remove noise and interference in the target image label data, improving the accuracy and reliability of the data; the data decomposition process can decompose the image - label joint data into multiple multimodal samples, effectively enriching the diversity and complexity of the data, providing more data selection and application possibilities. By processing the target image label data with geographic information data and terrain surface data, the relevance and practicality of the data can be enhanced, making the data more applicable and valuable. Multimodal samples can provide more comprehensive data information, which can enhance the robustness and generalization ability of the model, making the model more suitable for different application scenarios and environmental conditions. In this embodiment, the data chimeric module performs noise injection - noise reduction process and data decomposition process on the target image label data, which can effectively improve the data quality, enrich the data samples, enhance the data value, and strengthen the model robustness, thereby achieving efficient and accurate generation of labeled sample data. This embodiment has significant advantages and application prospects.

[0076] Optionally, the obtaining of the geographic information data and terrain surface data of the current study area includes:

[0077] Obtain the basic digital elevation model data of the current study area, and obtain the terrain surface data according to the basic digital elevation model data, where the terrain surface data includes slope data, aspect data, and hillshade data;

[0078] According to the remote sensing multimodal semantic segmentation dataset, obtain geographic metadata and time - step data, linearly embed the geographic metadata and the time - step data respectively, and add them to obtain the geographic information data.

[0079] Specifically, in this embodiment, obtaining terrain surface data including slope data, aspect data, and hillshade data by obtaining basic digital elevation model data enriches the dimension of geographic information, provides more terrain features, and enhances the diversity and expression ability of the data. Obtaining geographic metadata and time - step data using the remote sensing multimodal semantic segmentation dataset, linearly embedding them respectively and adding them to obtain geographic information data improves the relevance of the data, making the geographic information data more practical and reliable. By obtaining and integrating geographic information data and terrain surface data, the quality and integrity of the data can be improved, the value of the data can be increased, and more abundant information and possibilities can be provided for subsequent data processing and analysis. Obtaining rich geographic information data and terrain surface data helps to improve the expression ability of the model, enabling the model to better capture geographic spatial features and change laws, and enhancing the prediction and analysis ability of the model.

[0080] Optionally, the process of injecting noise and denoising the target image label data by using the geographic information data and the terrain surface data to obtain the image-label joint data includes:

[0081] Taking the geographic information data as conditional information and the terrain surface data as a supplement to the conditional information, performing a noise injection-denoising process on the target image label data to obtain the image-label joint data.

[0082] Specifically, the geographic information data and the terrain surface data provide richer background information and context, which can help remove noise and interference in the target image label data, improving the effect and accuracy of the noise injection-denoising process; taking the geographic information data as conditional information and the terrain surface data as a supplement to the conditional information can enhance the relevance between the target image label data and the geographical environment, helping to remove noise inconsistent with the geographic information and reducing mislabeling and irrelevant information in the data; during the process of removing noise, the supplement of the geographic information data and the terrain surface data can help retain useful information in the target image label data, especially information related to the geographical environment, improving the fidelity and practicality of the data after the noise injection-denoising process; by adding conditional constraints of the geographic information data and the terrain surface data, the model can be trained to better understand and utilize geographical environment information, improving the robustness and generalization ability of the model.

[0083] Optionally, the process of decomposing the image-label joint data to obtain multiple multimodal samples includes:

[0084] Performing inverse binary coding and restoration on the latter C bands in the image-label joint data to obtain multiple multimodal samples, where C represents the number of categories of the image-label joint data, and log₂ represents the logarithmic function with base 2.

[0085] Specifically, by performing inverse binary coding and restoration on the image-label joint data, multiple multimodal samples with different bands can be obtained. These samples have different characteristics, enriching the diversity of the data, which helps to improve the expression ability and application value of the data; after obtaining multiple multimodal samples, it is possible to better understand and interpret data features and changes, explore the internal laws and characteristics of the data, and improve the interpretability of the data and the understanding of the application field. Multiple multimodal samples can be used for model training, which can increase the amount of training data for the model, improve the robustness and generalization ability of the model, making the model more suitable for different application scenarios. After obtaining multiple multimodal samples, the effective utilization rate of the data can be improved, meeting the requirements under different application scenarios, and enhancing the flexibility and applicability of the data.

[0086] In one embodiment, the multi-source image and the semantic label encoded in binary are jointly input into the encoder and concatenated along the channels before input. Since the label of semantic segmentation in this embodiment is also in the tif format, it is directly concatenated with the multi-source image along the channels. (For example, the number of channels of the original multi-source image is 4, and the number of channels of the semantic label is generally 1. However, in this embodiment, the label is mapped into a vector by binary encoding. If there are 8 categories, the label is mapped into a data format with 3 channels. Thus, the number of channels after concatenation is 4 + 3 = 7, and the diffused generated image data also has 7 channels. The first 4 channels are reserved as the generated image, and the last three channels are inverse binary encoded to restore to a 1-channel label. In this way, both the image and the label are generated simultaneously. After being encoded by the encoder, it is added to the noise in vector form, and then input into the noise injection-denoising processor for noise injection-denoising process. Combining Figure 2 As shown, in the upper part of the noise injection-denoising processor, the input is shown. The geospatial metadata of the satellite image (such as longitude, latitude, year, etc.) is used as conditional information to guide the generation of the image. The geospatial metadata and the time step are linearly embedded respectively and then added together. Among them, the symbol in the upper left corner represents "vector addition". The information such as longitude and latitude obtained from the remote sensing image metadata is in character form, which will be encoded into vector form through word embedding, and then each vector is added together; the time step is essential in the diffusion model and is also encoded into vector form and added to the encoded metadata to jointly guide the diffusion step; in the lower part of the noise injection-denoising processor, the input is shown. The terrain surface data extracted from the basic DEM data (basic digital elevation model data) is used as a supplement to the conditional information. The image-label joint data generated by performing the noise injection-denoising process multiple times is then decomposed, and the last bands are inverse binary encoded (C represents the number of categories) for restoration, so as to obtain the generated multi-source image-semantic label pair (multi-modal sample). Among them, Figure 2 The lower right corner indicates the loop execution of the noise injection-denoising process (the noise injection-denoising process is carried out step by step according to the setting of the time step, rather than processed all at once. For example, if the time step is set to 1000, the noise injection-denoising process will be executed 1000 times to ensure the quality and diversity of the generated image).

[0087] A multi-modal sample generation device based on a geospatial information conditional diffusion model provided by an embodiment of the present invention includes:

[0088] An acquisition unit, configured to acquire a remote sensing multi-modal semantic segmentation data set, where the remote sensing multi-modal semantic segmentation data includes a multi-source image and a semantic label corresponding to the multi-source image;

[0089] A training unit, configured to input the remote sensing multi-modal semantic segmentation data set into a geospatial information conditional diffusion model for training to obtain a plurality of multi-modal samples;

[0090] Among them, the geographic information conditional diffusion model includes a joint input module and a data embedding module;

[0091] The multi-source images and semantic labels in the remote sensing multi-modal semantic segmentation dataset are subjected to channel splicing processing through the joint input module to obtain target image label data;

[0092] The target image label data is subjected to a noise injection-denoising process and data decomposition processing through the data embedding module to obtain a plurality of the multi-modal samples.

[0093] The multi-modal sample generation device based on the geographic information conditional diffusion model in this embodiment is used to implement the multi-modal sample generation method based on the geographic information conditional diffusion model as described above. Its advantages compared with the prior art are the same as those of the multi-modal sample generation method based on the geographic information conditional diffusion model compared with the prior art, and will not be elaborated here.

[0094] As Figure 3 shown, an electronic device 300 provided by an embodiment of the present invention includes a memory 310 and a processor 320; the memory 310 is used to store a computer program; the processor 320 is used to implement the multi-modal sample generation method based on the geographic information conditional diffusion model as described above when executing the computer program.

[0095] Or, an electronic device 300 includes a memory 310 and a processor 320 coupled to the memory 310; the memory 310 is configured to store a computer program; the processor 320 is configured to perform the following operations when executing the computer program:

[0096] Obtain a remote sensing multi-modal semantic segmentation dataset, where the remote sensing multi-modal semantic segmentation data includes multi-source images and semantic labels corresponding to the multi-source images;

[0097] Input the remote sensing multi-modal semantic segmentation dataset into the geographic information conditional diffusion model for training to obtain a plurality of multi-modal samples;

[0098] Among them, the geographic information conditional diffusion model includes a joint input module and a data embedding module;

[0099] The multi-source images and semantic labels in the remote sensing multi-modal semantic segmentation dataset are subjected to channel splicing processing through the joint input module to obtain target image label data;

[0100] The target image label data is subjected to a noise injection-denoising process and data decomposition processing through the data embedding module to obtain a plurality of the multi-modal samples.

[0101] A computer-readable storage medium provided by an embodiment of the present invention, on which a computer program is stored. When the computer program is executed by a processor, the multi-modal sample generation method based on a geographic information conditional diffusion model as described above is implemented.

[0102] Or, a non-volatile computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor performs the following operations:

[0103] Obtain a remote sensing multi-modal semantic segmentation data set, where the remote sensing multi-modal semantic segmentation data includes multi-source images and semantic labels corresponding to the multi-source images;

[0104] Input the remote sensing multi-modal semantic segmentation data set into a geographic information conditional diffusion model for training to obtain a plurality of multi-modal samples;

[0105] Wherein, the geographic information conditional diffusion model includes a joint input module and a data chimeric module;

[0106] Perform channel splicing processing on the multi-source images and semantic labels in the remote sensing multi-modal semantic segmentation data set through the joint input module to obtain target image label data;

[0107] Perform a noise injection-denoising process and data decomposition processing on the target image label data through the data chimeric module to obtain a plurality of the multi-modal samples.

[0108] Now, an electronic device 300 that can be used as a server or a client of the present invention will be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device 300 is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device 300 can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0109] The electronic device 300 includes a computing unit, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The computing unit, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0110] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc. In the present application, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention. In addition, the functional units in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0111] Although the present invention is disclosed as above, the scope of protection of the present invention is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will all fall within the scope of protection of the present invention.

Claims

1. A multi-modal sample generation method based on a geographical information conditional diffusion model, characterized in that, including: Obtain a remote sensing multimodal semantic segmentation dataset, where the remote sensing multimodal semantic segmentation data includes multi-source images and semantic labels corresponding to the multi-source images; Input the remote sensing multimodal semantic segmentation dataset into a geographic information conditional diffusion model to obtain multiple multimodal samples; Among them, the geographic information conditional diffusion model includes a joint input module and a data embedding module; Through the joint input module, perform channel splicing processing on the multi-source images and semantic labels in the remote sensing multimodal semantic segmentation dataset to obtain target image label data; Through the data embedding module, perform a noise injection-denoising process and data decomposition processing on the target image label data to obtain multiple of the multimodal samples, including: obtaining geographic information data and terrain surface data of the current study area; Perform a noise injection-denoising process on the target image label data through the geographic information data and the terrain surface data to obtain image-label joint data; In the image-label joint data, reverse binary encoding is performed on the last several bands and then restoration is carried out to obtain a plurality of the multimodal samples, where C represents the number of categories of the image-label joint data, represents the logarithmic function with base 2.

2. The multimodal sample generation method based on a geographic information conditional diffusion model according to claim 1, wherein The obtaining of the remote sensing multimodal semantic segmentation dataset includes: Obtain an image of the current study area and mine fine classification annotation data of the current study area; According to the image of the current study area and the mine fine classification annotation data, obtain multiple of the multi-source images and multiple of the semantic labels; According to multiple of the multi-source images and multiple of the semantic labels, obtain the remote sensing multimodal semantic segmentation dataset.

3. The multimodal sample generation method based on the geographical information conditional diffusion model according to claim 2, wherein The performing, through the joint input module, of channel splicing processing on the multi-source images and semantic labels in the remote sensing multimodal semantic segmentation dataset to obtain target image label data includes: Perform normalization processing on the multi-source images to obtain normalized images; Map the discrete label values corresponding to the semantic labels to binary vector representations to obtain binary labels; Perform splicing processing on the normalized images and the binary labels in the channel dimension to obtain target image label data.

4. The multimodal sample generation method based on a geographical information condition diffusion model according to claim 1, wherein The obtaining of the geographic information data and terrain surface data of the current study area includes: Obtain basic digital elevation model data of the current study area, and obtain the terrain surface data according to the basic digital elevation model data, where the terrain surface data includes slope data, aspect data, and mountain shadow data; According to the remote sensing multimodal semantic segmentation dataset, obtain geographic metadata and time step data, perform linear embedding on the geographic metadata and the time step data respectively, and add them to obtain the geographic information data.

5. The multimodal sample generation method based on a geographical information condition diffusion model according to claim 1, wherein The performing, through the geographic information data and the terrain surface data, of a noise injection-denoising process on the target image label data to obtain image-label joint data includes: Use the geographic information data as conditional information, and use the terrain surface data as a supplement to the conditional information, and perform a noise injection-denoising process on the target image label data to obtain image-label joint data.

6. A multi-modal sample generation device based on a geographical information conditional diffusion model, characterized in that, including: An acquisition unit for acquiring a remote sensing multimodal semantic segmentation dataset, where the remote sensing multimodal semantic segmentation data includes multi-source images and semantic labels corresponding to the multi-source images; A training unit for inputting the remote sensing multi-modal semantic segmentation dataset into a geographic information conditional diffusion model for training to obtain a plurality of multi-modal samples; Wherein, the geographic information conditional diffusion model includes a joint input module and a data embedding module; The multi-source images and semantic labels in the remote sensing multi-modal semantic segmentation dataset are subjected to channel splicing processing through the joint input module to obtain target image label data; The target image label data is subjected to a noise injection-denoising process and data decomposition processing through the data embedding module to obtain a plurality of the multi-modal samples, including: obtaining geographic information data and terrain surface data of the current study area; The target image label data is subjected to a noise injection-denoising process through the geographic information data and the terrain surface data to obtain image-label joint data; In the image-label joint data, the last several bands are subjected to inverse binary coding and restored to obtain a plurality of the multimodal samples, where C represents the number of categories of the image-label joint data, represents the logarithmic function with base 2.

7. An electronic device, characterized in that, Including a memory and a processor; The memory is used for storing a computer program; The processor is used for implementing the multi-modal sample generation method based on the geographic information conditional diffusion model according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by the processor, the multi-modal sample generation method based on the geographic information conditional diffusion model according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Urban road network layout design method based on conditional diffusion model

    CN116451398A

  • Image generation model training method and system, image generation method and system and electronic equipment

    CN117541883A