Remote sensing image generation method and system based on mixed semantic embedding
By adopting the method of hybrid semantic embedding and semantic refinement network in remote sensing image generation, the problem of insufficient semantic controllability and diversity in the prior art is solved, high-quality, semantic consistent remote sensing image generation is achieved, and the performance of downstream tasks is significantly improved.
Patent Information
- Application Number
- CN202510062409.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to maintain semantic controllability and diversity when generating remote sensing images, especially when dealing with remote sensing targets with semantic blur, complex occlusion and irregular spatial distribution, the generated images are insufficient semantic consistency and fine-grained controllability.
Using a method based on mixed semantic embedding, a mixed semantic embedding is generated by calculating the combination of the geometric information spatial descriptor and semantic mask of the input image, the mixed semantic embedding is generated, and the generation network is guided to generate remote sensing images, and local, fine-grained semantic feedback and semantic mask consistency feedback are provided through the semantic refinement network.
It realizes the generation of high-quality, semantically consistent remote sensing images, while maintaining generation diversity, significantly improving the performance of downstream tasks, especially in background generation tasks.
Smart Images

Figure CN119991844A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and pattern recognition, and more particularly to a remote sensing image generation method and system based on hybrid semantic embedding. Background Art
[0002] Image generation diversity is divided into global diversity and semantic diversity, and many existing methods focus on global diversity rather than semantic diversity. They use variational autoencoder architectures to constrain generation. Regarding semantic controllability, many recent studies have emphasized the quality aspect of image generation. These methods convert semantic masks into a general image-image framework by directly inputting them into the encoder-decoder network or using spatially adaptive normalization. However, these methods ignore finer semantic layouts, resulting in reduced semantic consistency and limited fine-grained controllability.
[0003] Since the above methods are mainly designed for natural images or everyday objects, their performance is greatly reduced when applied to remote sensing images. Remote sensing objects exhibit a large amount of semantic ambiguity, complex occlusion, and irregular spatial distribution in instances. Semantic ambiguity refers to the extensive feature overlap of remote sensing objects of different categories (such as grassland and forest) at the semantic level. Complex mutual occlusion describes the spatial overlap of remote sensing objects within the same category or between different categories, which brings challenges to the model training process. The irregular spatial distribution within instances is related to the different shapes and geometric arrangements of remote sensing objects (such as buildings). These inherent characteristics can lead to problems such as semantic confusion and geometric pattern collapse, resulting in poor semantic controllability and diversity. Therefore, addressing these issues is crucial to achieve a proper balance between controllability and diversity. Summary of the invention
[0004] In view of this, the present invention provides a remote sensing image generation method and system based on hybrid semantic embedding, which comprehensively considers the semantic controllability and diversity of the generated images, and can excellently maintain the generation diversity while generating high-quality, semantically consistent images.
[0005] In order to achieve the above object, the present invention adopts the following technical solution:
[0006] A remote sensing image generation method based on hybrid semantic embedding, comprising:
[0007] Compute the geometric information space descriptor of the input image based on the semantic embedding;
[0008] Combining the geometric information space descriptor and the semantic mask to generate a hybrid semantic embedding, and generating a remote sensing image based on the hybrid semantic embedding through a generative network guided by the hybrid semantic embedding;
[0009] A semantic refinement network is used to provide local, fine-grained semantic feedback and semantic mask consistency feedback to the generated remote sensing image.
[0010] Preferably, calculating the geometric information space descriptor of the input image according to the semantic embedding specifically includes:
[0011] Establish a polar coordinate system for each pixel in the instance;
[0012] According to certain polar angles and radii, the polar coordinate system established for each pixel is spatially binned;
[0013] For each block of a single pixel, determine whether the edge of the instance is within the block. If so, count the number of pixels; if not, the default value is 0;
[0014] The number of pixels in each block is used as the numerical feature of the block, and the numerical features of all blocks of the pixel are used as the high-dimensional features of the pixel;
[0015] Normalize the high-dimensional features to obtain the geometric information space descriptor.
[0016] Preferably, generating a remote sensing image through a generative network guided by hybrid semantic embedding and based on the hybrid semantic embedding comprises:
[0017] Decouple hybrid semantic embedding into semantic masks and feature descriptors;
[0018] Decouple the input latent variables and extract features related to geometric information;
[0019] The decoupled features are combined with the feature descriptors to calculate the multi-scale modulation parameters through the hybrid semantic feature modulation module;
[0020] The features processed by the hybrid semantic feature modulation module are convolved to generate a multi-scale feature map, and the multi-scale feature map is upsampled to obtain a new feature map;
[0021] When the dimension of the new feature map is consistent with the dimension of the real image, the generated remote sensing image is obtained.
[0022] Preferably, the decoupled features are combined with the feature descriptors to calculate the multi-scale modulation parameters through a hybrid semantic feature modulation module, specifically including:
[0023] The feature descriptor and the decoupled feature are respectively dimensional adjusted through the convolutional network to obtain the adjusted descriptor and the adjusted feature;
[0024] Perform SC convolution according to the adjusted features to obtain the convolution feature map;
[0025] Perform Triplet attention operation based on the adjusted features to obtain an attention feature map;
[0026] Performing depth convolution on the convolution feature map and the adjusted descriptor to obtain the first deep feature;
[0027] Perform deep convolution on the first deep feature and the attention feature map to obtain the second deep feature;
[0028] The feature descriptor, the decoupled features and the second deep features are concatenated, and the multi-scale modulation parameters are calculated through a two-layer convolutional network.
[0029] Preferably, providing local, fine-grained semantic feedback and semantic mask consistency feedback to the generated remote sensing image through a semantic refinement network specifically includes:
[0030] Segment the generated remote sensing image to generate a segmentation mask of the synthetic image;
[0031] Segment the real image and generate a segmentation mask of the real image;
[0032] Compute the loss function between the segmentation mask of the synthetic image and the segmentation mask of the real image.
[0033] Preferably, the loss function of the semantic refinement network is:
[0034]
[0035] Among them, e represents the current training round, Γ is a hyperparameter, represents the loss function of the real image, Corresponding to the loss function for generating remote sensing images, is the loss function of the real image and the generated remote sensing image.
[0036]
[0037] Among them, K represents the total number of semantic categories, I represents the real image, and C k represents the prediction result of the semantic refinement network for the kth class, m j,k is a semantic mask, It means to find the expectation, M represents the number of samples in the training batch;
[0038]
[0039] Among them, G(.) represents the image generated by the generator;
[0040]
[0041] Among them, Pre(I j,k) represents the prediction result of the semantic refinement network for the real image, Pre(z j,k ) represents the prediction result of the semantic refinement network for the generated image.
[0042] Preferably, the total loss function of the model is:
[0043]
[0044] Among them, λ adv , fm , perc and λ ref is the weight coefficient of each loss term, is the constraint loss function, To combat losses, is the feature matching loss, is the perceived loss, is the semantic refinement loss.
[0045] Preferably, the constraint loss calculation formula is:
[0046]
[0047] Among them, I and S represent the real image and semantic label respectively, z is the hidden variable of the input generator, G(.) represents the graph generated by the generator, and D(.) represents the discriminant result of the discriminator;
[0048] The feature matching loss calculation formula is:
[0049]
[0050] Among them, N i Characteristic D i The number of elements in (I,S), D i (.) represents the discriminator’s discrimination result for the i-th sample;
[0051] The perceptual loss calculation formula is:
[0052]
[0053] Among them, Φ represents the features extracted from the pre-trained VGG-19 model, Represents the generated image, and the subscript h represents the hth layer of the network;
[0054] The adversarial loss calculation formula is:
[0055]
[0056] A remote sensing image generation system based on hybrid semantic embedding, comprising:
[0057] Geometric information space description module: calculates the geometric information space descriptor of the input image based on semantic embedding;
[0058] A hybrid semantic embedding guided generative network module: combining the geometric information space descriptor and the semantic mask to generate a hybrid semantic embedding, and generating a remote sensing image according to the hybrid semantic embedding through a generative network guided by the hybrid semantic embedding;
[0059] Semantic refinement network module: providing local, fine-grained semantic feedback and semantic mask consistency feedback to the generated remote sensing image through a semantic refinement network.
[0060] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a remote sensing image generation method and system based on hybrid semantic embedding, and proposes new ideas in semantic controllability, diversity and extensibility. The present invention proposes a geometric information representation method for processing partial-level semantic representation. By introducing a hybrid semantic feature modulation block, accurate alignment between the generated image and the annotation can be achieved, and global and local semantic information can be fully integrated. A new loss function is used to alleviate the semantic confusion problem, enhance the robustness of the model, and ensure reliable content generation. Compared with existing remote sensing image generation methods, the method proposed in the present invention performs well in visual quality and has made significant progress in improving the performance of downstream tasks. Especially in background generation tasks, the proposed model shows extremely high efficiency. In general, the present invention establishes a new benchmark for remote sensing image generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0062] Figure 1 The present invention provides a flow chart of a remote sensing image generation method based on hybrid semantic embedding.
[0063] Figure 2 A schematic diagram of a remote sensing image generation system based on hybrid semantic embedding provided by the present invention.
[0064] Figure 3 A schematic diagram of the structure of the hybrid semantic feature modulation module provided by the present invention.
[0065] Figure 4 Diagram of the encoder and decoder infrastructure of the semantic refinement network provided by the present invention.
[0066] Figure 5 This is a comparison chart of the effects of generating images using different image generation algorithms provided by the present invention. DETAILED DESCRIPTION
[0067] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0068] The embodiment of the present invention discloses a remote sensing image generation method based on hybrid semantic embedding, such as Figure 1 As shown, including:
[0069] Compute the geometric information space descriptor of the input image based on the semantic embedding;
[0070] Combining the geometric information space descriptor and the semantic mask to generate a hybrid semantic embedding, and generating a remote sensing image based on the hybrid semantic embedding through a generative network guided by the hybrid semantic embedding;
[0071] A semantic refinement network is used to provide local, fine-grained semantic feedback and semantic mask consistency feedback to the generated remote sensing image.
[0072] The goal of this invention is to generate a photo-realistic image I of height H and width W consistent with a given semantic mask S, where S contains C class labels. Each pixel in S corresponds to a specific semantic category k from a set of predefined categories 1,2,…,C, representing the expected semantics of the corresponding position in I.
[0073] The specific implementation process of each step is introduced below.
[0074] The present invention proposes a geometric-informed spatial descriptor (GSD) that captures the positional features of each pixel to describe the geometry and features of remote sensing objects. In addition, the present invention extends the semantic input to include one-hot encoding and geometric-informed spatial descriptors, where one-hot encoding is the default embedding method, thereby creating a global-local semantic embedding distribution called hybrid semantic embedding. In this way, clues from the object component-level layout can be effectively utilized to enhance semantic controllability.
[0075] Specifically, the calculation process of the geometric information space description is as follows:
[0076] 1) Establish a polar coordinate system for each pixel in the instance;
[0077] 2) According to a certain polar angle and radius, the coordinate system established for each pixel is spatially binned;
[0078] 3) For each block of a single pixel, determine whether the edge of the instance is within the block. If so, count the number of pixels; if not, default to 0;
[0079] 4) The number of pixels in each block is used as the numerical feature of the block, and the numerical features of all blocks of the pixel are used as the high-dimensional features of the pixel;
[0080] 5) Normalize the high-dimensional features to obtain the geometric information space descriptor.
[0081] The geometric information space descriptor and the semantic mask are combined to generate a hybrid semantic embedding which is input into the generative network guided by the hybrid semantic embedding.
[0082] The hybrid semantic embedding-guided generative network consists of multiple hybrid semantic feature modulated residual blocks with upsampling operations, such as Figure 1 As shown in the multiple generated feature maps in . During the generation process, the geometric information space description features are calculated according to the corresponding semantic embeddings, and the multi-scale modulation parameters are extracted from the hybrid semantic feature block modulation at each layer. Then, the Hybrid Semantic Feature Modulation (HSFM) module uses the modulation parameters to control the generated content through semantic adaptive normalization.
[0083] Specifically, the hybrid semantic embedding is decoupled into semantic masks and feature descriptors;
[0084] Decouple the input latent variables and extract features related to geometric information;
[0085] The decoupled features are combined with the feature descriptors to calculate the multi-scale modulation parameters through the hybrid semantic feature modulation module;
[0086] The features processed by the hybrid semantic feature modulation module are convolved to generate a multi-scale feature map, and the multi-scale feature map is upsampled to obtain a new feature map;
[0087] The above steps are repeated until the dimension of the new feature map is consistent with the dimension of the real image, and the generated remote sensing image is obtained.
[0088] The GSD feature of the present invention supplements the object-level geometric shape and spatial information for the semantic layout, while the semantic layout introduces paired global semantic information for the GSD feature. HSFM adaptively adjusts parameters according to the different shapes of different semantic object categories, thereby effectively guiding the generation of semantic images.
[0089] like Figure 3 As shown, the decoupled features are combined with the feature descriptors and the multi-scale modulation parameters are calculated through the hybrid semantic feature modulation module, which specifically includes:
[0090] The feature descriptor and the decoupled feature are respectively dimensional adjusted through the convolutional network to obtain the adjusted descriptor and the adjusted feature;
[0091] Perform SC convolution according to the adjusted features to obtain a convolution feature map;
[0092] According to the adjusted features, a Triplet attention operation is performed to obtain an attention feature map;
[0093] Performing depth convolution on the convolution feature map and the adjusted descriptor to obtain the first deep feature;
[0094] Perform deep convolution on the first deep feature and the attention feature map to obtain the second deep feature;
[0095] The feature descriptor, the decoupled features and the second deep features are concatenated, and the parameters γ and β are learned through two layers of convolutional networks respectively, and the multi-scale modulation parameters are calculated by the formula
[0096] z′=z·r+β
[0097] Where z is the input and z' is the output.
[0098] In the HSFM module, the hybrid semantic features are first decoupled into semantic masks and GSD features, which are scaled to a uniform size. These features are then fed into two separate convolutional layers to predict two sets of semantically adaptive 3×3 convolution kernels. The features obtained after the semantic layout convolution are then passed through Triplet attention and parameter-free convolution to reduce spatial and channel redundancy within the convolutional neural network features, thereby compressing the CNN model and improving its performance, and obtaining two sets of features. These two sets of features are fused with the previously convolved GSD features through deep convolution and then concatenated. Finally, feature modulation is performed using the hybrid semantic features and adaptive normalized modulation parameters.
[0099] By using the HSFM block, semantic and spatial location information can be effectively integrated. The large amount of prior knowledge contained in the hybrid semantic embedding greatly improves the semantic consistency and fine-grained generation quality of the generated images guided by the hybrid semantic embedding, especially in remote sensing targets, thus achieving excellent performance.
[0100] In the traditional generative adversarial network (GAN) training framework, the generator network and the discriminator network compete with each other. The generator generates images, while the discriminator tries to distinguish between real and fake images. However, this requires the discriminator to simultaneously evaluate the fidelity of the image and the consistency of the semantic mask, which greatly increases the complexity of its learning and may prevent the generator from receiving fine-grained semantic feedback. To address this challenge, this paper introduces a novel semantic refinement network (SRN) in the traditional training framework.
[0101] During training, the discriminator provides global fidelity feedback, while the SRN provides local, fine-grained semantic and mask consistency feedback. The encoder and decoder infrastructure is as follows: Figure 4 shown.
[0102] The encoder and decoder form a semantic refinement network, which is used to generate prediction results. Based on the generated prediction results and through the semantic refinement loss function, local and fine-grained semantic feedback is provided to the generated remote sensing image; in the generative adversarial network architecture, the discriminator provides semantic mask consistency feedback through adversarial loss.
[0103] SRN is designed as a semantic segmentation network with an encoder-decoder architecture that classifies each pixel of the input image. A pixel-wise cross entropy loss is calculated between the output of SRN and the ground truth segmentation mask. SRN is trained jointly with the generator and the discriminator. When given real images, SRN trains itself and generates alignment information. Conversely, for fake images, SRN takes the segmentation results and registration data as fine-grained semantic feedback. Loss and Activated after 80 training epochs.
[0104] In order to ensure a high degree of semantic consistency between the generated image and the given semantic mask, this paper proposes a semantic refinement loss to optimize the model training process by using To provide local, fine-grained semantic feedback, as shown in the formula:
[0105]
[0106] Among them, e represents the current training round, Γ is a hyperparameter, represents the loss function of the real image, Corresponding to the loss function for generating remote sensing images, is the loss function of the real image and the generated remote sensing image. The definition of is shown in the formula:
[0107]
[0108] Among them, K represents the total number of semantic categories, I represents the real image, and C k represents the prediction result of the semantic refinement network for the kth class, m j,k is a semantic mask, It means to find the expectation, and M represents the number of samples in the training batch. The definition of is as follows:
[0109]
[0110] Among them, G(.) represents the image generated by the generator, as shown in the formula:
[0111]
[0112] Where Pre(I j,k ) represents the prediction result of the semantic refinement network for the real image, Pre(z j,k ) represents the prediction result of the semantic refinement network for the generated image.
[0113] although and There are similarities in form, but they each produce different effects. The segmentation results of the real image are aligned with the given mask, thus facilitating the training of the semantic refinement network. The segmentation results generated by the semantic refinement network for the generated image are fed back to the generator through back-propagation to guide image generation. The output of the generator is further refined by minimizing the difference between the segmentation results of the generated images and the real images, thereby improving the semantic consistency and quality of the generated images.
[0114] The training strategy of the remote sensing image generation model based on hybrid semantic embedding in the present invention is:
[0115] The generator is trained using a constrained loss function Fighting Losses Feature matching loss Perceived loss and the proposed semantic refinement loss Training is performed. The shape of the total loss is shown in the formula:
[0116]
[0117] Among them, λ adv , fm , perc and λ ref is the weight coefficient of each loss term. Adversarial learning can effectively maintain the consistency of the generated image and the real image distribution. The present invention uses the adversarial loss function to constrain image generation. The constrained loss function is expressed as follows:
[0118]
[0119] Using feature matching loss To strengthen supervision and stabilize the training process. This loss encourages the features of the generated image to be closer to the features of the real image in the feature space of the discriminator D. Feature matching loss The definition is as follows:
[0120]
[0121] The present invention uses a pre-trained VGG-19 model to extract features from the real image I and the generated image respectively. As shown below, the perceptual loss is calculated in the multi-scale feature space:
[0122]
[0123] Among them, Φ represents the features extracted from the pre-trained VGG-19 model, Represents the generated image, and the subscript h represents the hth layer of the network;
[0124] The adversarial loss calculation formula is:
[0125]
[0126] The embodiment of the present invention discloses a remote sensing image generation system based on hybrid semantic embedding, such as Figure 2 As shown, the remote sensing image generation model based on hybrid semantic embedding constructed by the present invention consists of three main parts: a geometric information space description module, a hybrid semantic embedding guided generation network module and a semantic refinement network module.
[0127] Geometric information space description module: calculates the geometric information space descriptor of the input image based on semantic embedding;
[0128] A hybrid semantic embedding guided generative network module: combining the geometric information space descriptor and the semantic mask to generate a hybrid semantic embedding, and generating a remote sensing image according to the hybrid semantic embedding through a generative network guided by the hybrid semantic embedding;
[0129] Semantic refinement network module: providing local, fine-grained semantic feedback and semantic mask consistency feedback to the generated remote sensing image through a semantic refinement network.
[0130] The specific implementation process and method of each module of the system of the present invention are the same and will not be repeated here.
[0131] Embodiment 1:
[0132] Image generation on satellite remote sensing datasets:
[0133] The following is a comparison of the performance of the method of the present invention and other methods on GID-15, and the results are shown in Table 1 and Figure 5 As shown. In the semantic controllability experiment, several state-of-the-art methods are compared from two dimensions: generation quality (mainly measured by FID) and semantic consistency (assessed by mIoU and accuracy). Compared with other generation methods, the present invention significantly outperforms other methods in terms of FID, mIoU and accuracy. This shows that the present invention has obvious advantages in both the quality and semantic consistency of generated images. Diversity experiments show that the method of the present invention achieves suboptimal performance in the LPIPS and mCSD indicators, and achieves optimal performance in the mOCD indicator. This shows that the method of the present invention can maintain excellent generation diversity while generating high-quality and semantically consistent images, which is also confirmed by the diversity indicators.
[0134] Table 1 Comparison of different image generation algorithms - Visualization effect of satellite remote sensing image dataset
[0135]
[0136] LPIPS of 0 means that the algorithm does not support diversity generation; - means that the algorithm does not support this indicator test.
[0137] Embodiment 2:
[0138] Image generation on aerial airborne remote sensing datasets:
[0139] Using the ISPRS aerial remote sensing dataset, it is verified that the images generated by the present invention can improve the performance of downstream tasks (image segmentation). Table 2 introduces the comparison of different methods in downstream task enhancement. The results show that the method of the present invention performs well in accuracy, scalability and stability, and can effectively support downstream tasks in the field of remote sensing while ensuring semantic controllability and diversity.
[0140] Table 2 Comparison of the improvement of downstream tasks by different image generation algorithms - Visualization effect of aerial remote sensing image dataset
[0141]
[0142]
[0143] Source Only means that no generated images are used as a means of enhancement.
[0144] The hybrid semantic embedding guided generative adversarial network of the present invention utilizes hierarchical information from a single source. In the process of feature description, the present invention proposes a hybrid semantic embedding method, which can coordinate fine-grained local semantic layout to describe the geometric structure of remote sensing objects without the need for additional information. From the perspective of feature modeling, a novel loss function is proposed by constructing a semantic refinement network to ensure fine-grained semantic feedback. The proposed method can alleviate semantic confusion and prevent geometric pattern collapse. Compared with existing remote sensing image generation methods, the model constructed by the present invention comprehensively considers the semantic controllability and diversity of generated images, and achieves leading performance in two different remote sensing image scenarios.
[0145] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0146] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A remote sensing image generation method based on hybrid semantic embedding, characterized in that: include: Compute the geometric information space descriptor of the input image based on the semantic embedding; Combining the geometric information space descriptor and the semantic mask to generate a hybrid semantic embedding, and generating a remote sensing image based on the hybrid semantic embedding through a generative network guided by the hybrid semantic embedding; A semantic refinement network is used to provide local, fine-grained semantic feedback and semantic mask consistency feedback to the generated remote sensing image.
2. The remote sensing image generation method based on hybrid semantic embedding according to claim 1, characterized in that: The geometric information space descriptor of the input image is calculated based on the semantic embedding, including: Establish a polar coordinate system for each pixel in the instance; According to certain polar angles and radii, the polar coordinate system established for each pixel is spatially binned; For each block of a single pixel, determine whether the edge of the instance is within the block. If so, count the number of pixels; if not, the default value is 0; The number of pixels in each block is used as the numerical feature of the block, and the numerical features of all blocks of the pixel are used as the high-dimensional features of the pixel; Normalize the high-dimensional features to obtain the geometric information space descriptor.
3. The remote sensing image generation method based on hybrid semantic embedding according to claim 1, characterized in that: A generative network guided by hybrid semantic embedding is used to generate remote sensing images based on hybrid semantic embedding, including: Decouple hybrid semantic embedding into semantic masks and feature descriptors; Decouple the input latent variables and extract features related to geometric information; The decoupled features are combined with the feature descriptors to calculate the multi-scale modulation parameters through the hybrid semantic feature modulation module; The features processed by the hybrid semantic feature modulation module are convolved to generate a multi-scale feature map, and the multi-scale feature map is upsampled to obtain a new feature map; When the dimension of the new feature map is consistent with the dimension of the real image, the generated remote sensing image is obtained.
4. The remote sensing image generation method based on hybrid semantic embedding according to claim 3 is characterized in that: After combining the decoupled features with the feature descriptors, the multi-scale modulation parameters are calculated through the hybrid semantic feature modulation module, including: The feature descriptor and the decoupled feature are respectively dimensional adjusted through the convolutional network to obtain the adjusted descriptor and the adjusted feature; Perform SC convolution according to the adjusted features to obtain the convolution feature map; Perform Triplet attention operation based on the adjusted features to obtain an attention feature map; Performing depth convolution on the convolution feature map and the adjusted descriptor to obtain the first deep feature; Perform deep convolution on the first deep feature and the attention feature map to obtain the second deep feature; The feature descriptor, the decoupled features and the second deep features are concatenated, and the multi-scale modulation parameters are calculated through a two-layer convolutional network.
5. The remote sensing image generation method based on hybrid semantic embedding according to claim 1, characterized in that: Providing local, fine-grained semantic feedback and semantic mask consistency feedback to the generated remote sensing image through a semantic refinement network, specifically including: Segment the generated remote sensing image to generate a segmentation mask of the synthetic image; Segment the real image and generate a segmentation mask of the real image; Compute the loss function between the segmentation mask of the synthetic image and the segmentation mask of the real image.
6. The remote sensing image generation method based on hybrid semantic embedding according to claim 5, characterized in that: The loss function of the semantic refinement network is: Among them, e represents the current training round, Γ is a hyperparameter, represents the loss function of the real image, Corresponding to the loss function for generating remote sensing images, is the loss function of the real image and the generated remote sensing image. Among them, K represents the total number of semantic categories, I represents the real image, and C k represents the prediction result of the semantic refinement network for the kth class, m j,k is a semantic mask, It means to find the expectation, M represents the number of samples in the training batch; Among them, G(.) represents the image generated by the generator; Among them, Pre(I j,k ) represents the prediction result of the semantic refinement network for the real image, Pre(z j,k ) represents the prediction result of the semantic refinement network for the generated image.
7. The remote sensing image generation method based on hybrid semantic embedding according to claim 6, characterized in that: The total loss function of the model is: Among them, λ adv , fm , perc and λ ref is the weight coefficient of each loss term, is the constraint loss function, To combat losses, is the feature matching loss, is the perceived loss, is the semantic refinement loss.
8. The remote sensing image generation method based on hybrid semantic embedding according to claim 7, characterized in that: The constraint loss calculation formula is: Among them, I and S represent the real image and semantic label respectively, z is the hidden variable of the input generator, G(.) represents the graph generated by the generator, and D(.) represents the discriminant result of the discriminator; The feature matching loss calculation formula is: Among them, N i Characteristic D i The number of elements in (I,S), D i (.) represents the discriminator’s discrimination result for the i-th sample; The perceptual loss calculation formula is: Among them, Φ represents the features extracted from the pre-trained VGG-19 model, Represents the generated image, and the subscript h represents the hth layer of the network; The adversarial loss calculation formula is:
9. A remote sensing image generation system based on hybrid semantic embedding, used to implement a remote sensing image generation method based on hybrid semantic embedding as described in any one of claims 1 to 8, characterized in that: include: Geometric information space description module: calculates the geometric information space descriptor of the input image based on semantic embedding; A hybrid semantic embedding guided generative network module: combining the geometric information space descriptor and the semantic mask to generate a hybrid semantic embedding, and generating a remote sensing image according to the hybrid semantic embedding through a generative network guided by the hybrid semantic embedding; Semantic refinement network module: providing local, fine-grained semantic feedback and semantic mask consistency feedback to the generated remote sensing image through a semantic refinement network.