A remote sensing information driven remote sensing image cascaded generation method

By using a cascaded generation method driven by remote sensing information and guiding image generation with remote sensing metadata, the problems of lack of information, poor feature preservation, and poor consistency in remote sensing image generation are solved. This enables the generation of high-quality, large-size, and multimodal remote sensing images, meeting the application requirements of real-time and large-scene coverage.

CN122336036APending Publication Date: 2026-07-03WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN UNIV
Filing Date
2026-04-10
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing remote sensing image generation methods lack effective remote sensing information guidance mechanisms, suffer from poor feature preservation during cascade generation, exhibit poor consistency between generated images and metadata, and have low efficiency in generating large-size images, making it difficult to meet the requirements of high quality, real-time processing, and large-scene coverage.

Method used

A remote sensing information-driven cascade generation method is adopted. By constructing a multi-layer cascade generator and a remote sensing feature fusion module, combined with a diffusion model and attention mechanism, the method uses remote sensing metadata information to guide image generation. The remote sensing consistency constraint module ensures the geometric and spectral consistency of the generated results and supports dynamic window sampling and multimodal remote sensing information fusion.

Benefits of technology

The generated remote sensing images are consistent with the actual remote sensing data in terms of geographical location and imaging parameters, maintaining high quality and high fidelity. They can generate large-size continuous borderless remote sensing images, adapt to various remote sensing data types, and meet the needs of real-time monitoring and large-scene coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122336036A_ABST
    Figure CN122336036A_ABST
Patent Text Reader

Abstract

This invention discloses a remote sensing information-driven cascaded generation method for remote sensing images, belonging to the fields of remote sensing image processing and artificial intelligence. The method constructs a remote sensing information encoder to encode remote sensing metadata information into conditional vectors, guiding a multi-level cascaded generator to progressively generate remote sensing images from low resolution to high resolution. During the generation process, a remote sensing feature fusion module integrates geographic and spectral features into the image generation, and a remote sensing consistency constraint module is applied to verify geometric and spectral consistency. This invention can generate high-quality remote sensing images that conform to actual geographic and spectral characteristics based on remote sensing metadata information, and has significant application value in scenarios such as data augmentation, simulation, missing data completion, and emergency response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to remote sensing image processing and artificial intelligence, and more particularly to a method for cascading remote sensing image generation driven by remote sensing information. Background Technology

[0002] Remote sensing imagery consists of images of the Earth's surface acquired by sensors mounted on platforms such as satellites and aircraft. It plays a crucial role in fields such as agricultural monitoring, urban planning, environmental assessment, and disaster early warning. However, the acquisition of remote sensing images is affected by various factors, including sensor imaging conditions, weather conditions, and orbital limitations, leading to missing or poor-quality image data for certain areas or time periods. Furthermore, acquiring high-resolution remote sensing images is costly and has a long imaging cycle, making it difficult to meet the needs of real-time monitoring and rapid response.

[0003] To address the aforementioned issues, remote sensing image generation technology has received widespread attention in recent years. Traditional remote sensing image generation methods are mainly based on physical modeling and procedural generation. These methods require a large amount of prior knowledge and complex parameter settings, resulting in poor flexibility and difficulty in generating diverse remote sensing scenes. With the development of deep learning technology, remote sensing image generation methods based on generative adversarial networks (GANs) and diffusion models have made significant progress. However, existing remote sensing image generation methods still suffer from the following technical problems: First, there is a lack of effective remote sensing information guidance mechanisms. Most existing methods generate remote sensing images based solely on text descriptions or random noise, failing to fully utilize specialized knowledge in the field, such as geographic coordinates, imaging parameters, and spectral characteristics. This results in generated images lacking realism and failing to meet the needs of practical applications.

[0004] Second, there is the issue of feature preservation during the cascade generation process. Cascade generation methods achieve high-quality image generation by progressively generating images from low to high resolution, but problems such as feature distortion and texture discontinuity can easily occur during the cascading process. Especially in remote sensing images, the scale range of ground features is large, from macroscopic topography to microscopic building details, all of which need to be accurately generated, which places higher demands on the feature preservation capabilities of cascade generation.

[0005] Third, there is the issue of consistency between the generated image and remote sensing metadata. Remote sensing images have clearly defined geographic locations and imaging parameters, and the generated image needs to be consistent with this metadata information. Existing methods lack effective consistency verification mechanisms, resulting in deviations between the generated image and the actual remote sensing data in terms of geographic location, spectral characteristics, etc., which affects the usability of the generated results.

[0006] Fourth, there is the efficiency issue in generating large-size remote sensing images. Remote sensing images are typically large in size, with a single image potentially reaching tens of thousands of pixels in resolution. Directly generating large-size images using existing generative models is limited by computational resources, while using block-based generation can easily result in stitching artifacts, making it difficult to generate continuous, boundless large-size remote sensing images.

[0007] Therefore, a new method is needed that can effectively utilize remote sensing information, maintain consistency of cascaded generation features, ensure consistency between the generated results and remote sensing metadata, and efficiently generate large-size remote sensing images. Summary of the Invention

[0008] Purpose of the invention: The purpose of this invention is to provide a remote sensing information-driven cascaded generation method for remote sensing images, which fully utilizes remote sensing metadata to guide the image generation process and achieves high-quality, high-fidelity remote sensing image generation.

[0009] Technical solution: A method for cascaded generation of remote sensing images driven by remote sensing information, comprising the following steps: Step S1: Obtain remote sensing metadata information, which includes at least one of geographic coordinates, imaging time, sensor type, spectral range, imaging resolution, and imaging angle; Step S2: Construct a remote sensing information encoder to encode the remote sensing metadata information into a conditional vector of a multi-layer cascaded generator; Step S3: Construct a multi-level cascaded generator based on the diffusion model. The multi-level cascaded generator includes a primary generator and at least one secondary generator. Each generator is responsible for generating images at different resolutions. Step S4: Using a sampling strategy guided by remote sensing information, the conditional vector is used to guide the multi-level cascaded generator to generate an initial low-resolution remote sensing image; Step S5: Integrate the geographic features and spectral features from the remote sensing metadata into the generated image features using the remote sensing feature fusion module; Step S6: The secondary generator is used sequentially to upsample and enhance the details of the low-resolution image to generate a remote sensing image of the target resolution; Step S7: Apply the remote sensing consistency constraint module to verify the geometric and spectral consistency of the generated remote sensing images; Step S8: Output a high-resolution remote sensing image that satisfies the consistency of remote sensing features; The remote sensing information encoder adopts a network structure that combines a multilayer perceptron and an attention mechanism. The multi-level cascaded generator adopts a U-Net-based architecture. The remote sensing feature fusion module integrates remote sensing features into the image generation process using feature recalibration and attention weighting. The remote sensing consistency constraint module uses georegistration error and spectral angle similarity as consistency evaluation indicators.

[0010] Furthermore, the remote sensing information encoder includes a geographic coordinate encoding submodule, an imaging parameter encoding submodule, and a fusion submodule; The geographic coordinate encoding submodule encodes latitude and longitude coordinates using sine and cosine functions to form a 64-dimensional geographic feature vector. The imaging parameter encoding submodule converts imaging time, sensor type, spectral range, imaging resolution, and imaging angle into corresponding feature vectors through an embedding layer; The fusion submodule uses a multi-head attention mechanism to fuse the geographic feature vector and the imaging parameter feature vector, and outputs the condition vector.

[0011] Furthermore, the multi-stage cascade generator adopts a three-stage cascade structure; The primary generator produces 64×64 resolution remote sensing images, with this stage focusing on capturing the macroscopic layout of the scene and the main land cover categories; The first-level generator upsamples the 64×64 resolution image to 256×256 resolution, focusing on generating medium-scale ground features and texture details; The second-stage generator upsamples the 256×256 resolution image to 1024×1024 resolution, generating high-resolution fine textures and edge details; Each generator performs a diffusion denoising process guided by a conditional vector, using either the DDPM or DDIM sampling algorithm.

[0012] Furthermore, the remote sensing feature fusion module includes a geographic feature extraction submodule and a spectral feature injection submodule; The geographic feature extraction submodule extracts the topographic features, land cover distribution features, and climate features of the corresponding region from a pre-built geographic feature database based on geographic coordinate information; The spectral feature injection submodule generates corresponding spectral response curve features based on sensor type and spectral range information, and injects the spectral features into the feature map of the generator through a channel attention mechanism; The remote sensing feature fusion module performs remote sensing feature fusion at each stage of cascade generation to ensure that the generated image is consistent with the remote sensing metadata in both spatial and spectral dimensions.

[0013] Furthermore, the remote sensing consistency constraint module includes a geometric consistency verification submodule and a spectral consistency verification submodule; The geometric consistency verification submodule calculates the geographic registration error between the generated image and the reference geographic information system data. When the registration error exceeds a preset threshold, the generated image is geometrically corrected. The spectral consistency verification submodule calculates the difference between the spectral angle similarity of the generated image and the statistical distribution of the real remote sensing image. When the difference exceeds a preset threshold, the spectral feature injection weight is adjusted. The remote sensing consistency constraint module performs consistency optimization in the final stage of cascade generation. The optimization objective function is Ltotal = Lrecon + αLgeo + βLspectral, where Lrecon is the reconstruction loss, Lgeo is the geometric consistency loss, Lspectral is the spectral consistency loss, and α and β are balance coefficients.

[0014] Furthermore, the remote sensing information-guided sampling strategy includes conditional enhancement sampling and frequency-aware sampling; The conditional augmentation sampling randomly discards some remote sensing metadata information with a certain probability during the training process, thereby improving the model's robustness to missing conditions. The frequency-aware sampling divides the sampling process into multiple stages, with each stage focusing on image features in different frequency ranges. The low-frequency stage uses a lower conditional guidance strength, while the high-frequency stage uses a higher conditional guidance strength. The sampling strategy employs a classifier-free guided method, which controls the generation process by weighting the combination of unconditional and conditional noise predictions.

[0015] Furthermore, the method also includes a cascaded training strategy; The cascaded training strategy adopts a bottom-up training sequence, first training the primary generator to enable it to generate high-quality 64×64 resolution images. Then, with the parameters of the primary generator fixed, the first primary generator is trained so that it can generate 256×256 resolution images based on the output of the primary generator. Finally, fix the parameters of the primary generator and the first-level generator, and train the second-level generator so that it can generate a 1024×1024 resolution image based on the output of the first-level generator. The training loss for each generator includes reconstruction loss, adversarial loss, and remote sensing consistency loss.

[0016] Furthermore, the method also includes a dynamic window sampling strategy for generating remote sensing images of arbitrary size; The dynamic window sampling strategy divides the target large image into multiple overlapping image blocks, and the size of each image block is the training size of the generator. The width of the overlapping region is half the size of the image patch to ensure the continuity of adjacent image patches at the boundary; For each image patch, the same initial noise and conditional vector are used for generation to ensure the consistency of the generated content in overlapping areas; All image patches are stitched together according to the average pixel value of the overlapping areas to generate the final borderless large-size remote sensing image.

[0017] Furthermore, the method also includes multimodal remote sensing information fusion; The multimodal remote sensing information fusion supports the generation of various data types, including optical remote sensing images, radar remote sensing images, and hyperspectral remote sensing images; For optical remote sensing images, RGB three-channel or multispectral channel generation methods are used; For radar remote sensing images, a joint generation of amplitude and phase features is introduced, and a complex convolutional network is used to process the radar data. For hyperspectral remote sensing images, a spectral cascade generation strategy is adopted, first generating images with sparse bands, and then gradually generating images with complete spectral bands. Multimodal remote sensing information fusion is achieved through unified conditional vector representation and targeted network architecture design.

[0018] Beneficial effects: This invention constructs a remote sensing information encoder to encode remote sensing metadata information into conditional vectors, effectively utilizing professional knowledge in the field of remote sensing. This ensures that the generated images are consistent with the actual remote sensing data in terms of geographical location, imaging parameters, etc., thereby improving the authenticity and usability of the generated results.

[0019] This invention employs a multi-level cascaded generator architecture to generate remote sensing images step by step from low resolution to high resolution. Remote sensing feature information is incorporated into each generation stage, effectively maintaining feature consistency in the cascaded generation process and avoiding problems such as feature distortion and texture discontinuity.

[0020] This invention proposes a remote sensing feature fusion module that organically integrates geographic and spectral features into the image generation process, resulting in images with high fidelity in both spatial and spectral dimensions, accurately reflecting the distribution and characteristics of actual ground features.

[0021] This invention applies a remote sensing consistency constraint module to perform geometric and spectral consistency checks on the generated images, ensuring a high degree of consistency between the generated results and the remote sensing metadata information, thereby improving the reliability of the generated results.

[0022] This invention supports a dynamic window sampling strategy, which can generate remote sensing images of any size. Through consistency processing of overlapping areas, it achieves the generation of continuous, boundaryless, large-size images, meeting the application needs of large-scene coverage in the field of remote sensing.

[0023] This invention supports multimodal remote sensing information fusion and can generate remote sensing images of various types, such as optical, radar, and hyperspectral images, with strong versatility and adaptability. Attached Figure Description

[0024] Figure 1 This is an overall flowchart of the method provided in the embodiments of the present invention; Figure 2 This is a schematic diagram of the structure of a remote sensing information encoder; Figure 3 This is a schematic diagram of the architecture of a multi-level cascaded generator; Figure 4 This is a schematic diagram of the working principle of the remote sensing feature fusion module; Figure 5 This is a schematic diagram of the remote sensing consistency constraint module. Detailed Implementation

[0025] To make the technical solution of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Example 1

[0026] like Figure 1 As shown, this embodiment provides a remote sensing information-driven cascaded generation method for remote sensing images, specifically including the following steps: Step S1: Obtain remote sensing metadata information. Remote sensing metadata information includes at least one of the following: geographic coordinates (longitude, latitude, elevation), imaging time (year, month, day, hour, minute, second), sensor type (optical sensor, radar sensor, hyperspectral sensor, etc.), spectral range (visible light band, infrared band, microwave band, etc.), imaging resolution (ground sampling distance, unit: meters / pixel), and imaging angle (solar altitude angle, sensor elevation angle, azimuth angle, etc.). This metadata information can be obtained from sensor parameter files, geographic information system databases, or historical remote sensing image metadata.

[0027] Step S2: Construct the remote sensing information encoder. For example... Figure 2As shown, the remote sensing information encoder includes a geographic coordinate encoding submodule, an imaging parameter encoding submodule, and a fusion submodule. The geographic coordinate encoding submodule encodes latitude and longitude coordinates using sine and cosine functions to form a 64-dimensional geographic feature vector. Specifically, for longitude λ and latitude φ, the encoded vector is [sin(λ), cos(λ), sin(2λ), cos(2λ), ..., sin(32λ), cos(32λ), sin(φ), cos(φ), ..., sin(32φ), cos(32φ)]. The imaging parameter encoding submodule converts imaging time, sensor type, spectral range, imaging resolution, and imaging angle into corresponding feature vectors through an embedding layer. For example, imaging time is converted into a 32-dimensional feature vector through a temporal embedding layer, and sensor type is converted into a 16-dimensional feature vector through a category embedding layer. The fusion submodule uses a multi-head attention mechanism to fuse the geographic feature vector and the imaging parameter feature vector, outputting a 128-dimensional conditional vector. The multi-head attention mechanism consists of 8 attention heads, each with a dimension of 16.

[0028] Step S3: Construct a multi-level cascade generator based on the diffusion model. For example... Figure 3 As shown, the multi-level cascaded generator adopts a three-stage cascaded structure, including a primary generator, a first-stage generator, and a second-stage generator. Each generator uses a U-Net-based architecture, consisting of an encoder and a decoder. The encoder performs four downsampling operations, and the decoder performs four upsampling operations, with skip connections between layers of the same resolution. The primary generator takes random noise and a conditional vector as input and outputs a 64×64 resolution remote sensing image, focusing on capturing the macroscopic layout of the scene and major land cover categories. The first-stage generator takes a 64×64 resolution image and a conditional vector as input and outputs a 256×256 resolution image, focusing on generating mesoscale land cover features and texture details. The second-stage generator takes a 256×256 resolution image and a conditional vector as input and outputs a 1024×1024 resolution image, generating high-resolution fine textures and edge details. Each generator performs a diffusion denoising process guided by the conditional vector, using the DDPM sampling algorithm with 1000 sampling steps.

[0029] Step S4: Employ a remote sensing-guided sampling strategy. This strategy includes conditional augmentation sampling and frequency-aware sampling. Conditional augmentation sampling randomly discards some remote sensing metadata information with a 20% probability during training, for example, retaining only geographic coordinates while discarding imaging time, thus improving the model's robustness to missing conditions. Frequency-aware sampling divides the sampling process into multiple stages, each focusing on image features within a different frequency range. In the first 500 sampling steps, a lower conditional guidance strength (e.g., guidance coefficient of 3.0) is used to focus on generating the low-frequency structure of the image; in the latter 500 sampling steps, a higher conditional guidance strength (e.g., guidance coefficient of 7.5) is used to focus on generating the high-frequency details of the image. The sampling strategy employs a classifier-free guidance method, controlling the generation process through a weighted combination of unconditional and conditional noise predictions: ε̃=εuncond+w·(εcond-εuncond), where εuncond is the unconditional noise prediction, εcond is the conditional noise prediction, and w is the guidance coefficient.

[0030] Step S5: The geographic and spectral features from the remote sensing metadata are integrated into the generated image features using the remote sensing feature fusion module. For example... Figure 4 As shown, the remote sensing feature fusion module includes a geographic feature extraction submodule and a spectral feature injection submodule. The geographic feature extraction submodule extracts topographic features (such as elevation, slope, and aspect), land cover distribution features (such as land use type and vegetation cover), and climate features (such as average annual precipitation and average annual temperature) of the corresponding region from a pre-built geographic feature library based on geographic coordinate information. The geographic feature library is built based on global geographic information system data and enables fast querying through spatial indexing. The spectral feature injection submodule generates corresponding spectral response curve features based on sensor type and spectral range information, and injects the spectral features into the generator's feature map through a channel attention mechanism. The spectral response curve features include the spectral reflectance features of typical land features in different bands. The remote sensing feature fusion module performs remote sensing feature fusion at each stage of the cascaded generation process, ensuring that the generated image is consistent with the remote sensing metadata in both spatial and spectral dimensions.

[0031] Step S6: The low-resolution image is sequentially upsampled and enhanced for detail using a secondary generator to generate a remote sensing image at the target resolution. The cascaded generation process follows a bottom-up order: first, a 64×64 resolution image is generated using a primary generator; then, a 256×256 resolution image is generated using a first-stage generator; and finally, a 1024×1024 resolution image is generated using a second-stage generator. In each upsampling step, the secondary generator receives the low-resolution image and conditional vector generated in the previous stage as input and generates a high-resolution image through a diffusion denoising process. The upsampling factor is 4x, implemented using sub-pixel convolution or transposed convolution.

[0032] Step S7: Apply the remote sensing consistency constraint module to verify the geometric and spectral consistency of the generated remote sensing images. For example... Figure 5 As shown, the remote sensing consistency constraint module includes a geometric consistency verification submodule and a spectral consistency verification submodule. The geometric consistency verification submodule calculates the georegistration error between the generated image and the reference geographic information system data, specifically including positional and orientation errors. When the registration error exceeds a preset threshold (e.g., positional error exceeding 5 meters, orientation error exceeding 3 degrees), geometric correction is performed on the generated image, using methods including spatial transformation and affine transformation. The spectral consistency verification submodule calculates the difference between the spectral angular similarity of the generated image and the statistical distribution of the real remote sensing image. Spectral angular similarity is defined as the cosine of the angle between two spectral vectors. When the difference exceeds a preset threshold (e.g., spectral angle exceeding 10 degrees), the spectral feature injection weights are adjusted, and feature fusion and image generation are performed again. The remote sensing consistency constraint module performs consistency optimization in the final stage of cascade generation. The optimization objective function is Ltotal = Lrecon + αLgeo + βLspectral, where Lrecon is the reconstruction loss, using the mean square error; Lgeo is the geometric consistency loss, using the square of the georegistration error; Lspectral is the spectral consistency loss, using 1 - spectral angle similarity; α and β are balance coefficients, set to 0.5 and 0.3, respectively.

[0033] Step S8: Output a high-resolution remote sensing image that satisfies the consistency of remote sensing features. The generated remote sensing image is saved in a standard format (such as GeoTIFF), and the corresponding remote sensing metadata information is embedded in the image file to ensure the integrity and accuracy of the image's geolocation and imaging parameter information.

[0034] This embodiment also includes a cascaded training strategy. The cascaded training strategy employs a bottom-up training sequence, first training a primary generator to generate high-quality 64×64 resolution images. The training dataset includes 500,000 real 64×64 resolution remote sensing images and their corresponding remote sensing metadata. Training losses include reconstruction loss (mean squared error), adversarial loss (using a discriminator to distinguish generated images from real images), and remote sensing consistency loss (geometric and spectral consistency loss). The total loss function is Ltotal1 = LMSE + λ1Ladv + λ2Lconsistency, where λ1 = 0.5 and λ2 = 0.3. Training uses the Adam optimizer with a learning rate of 0.0002 and a training period of 200 epochs. Then, with the parameters of the primary generator fixed, a first-stage generator is trained to generate 256×256 resolution images based on the output of the primary generator. The training dataset includes 300,000 real 256×256 resolution remote sensing images and their corresponding remote sensing metadata, where the low-resolution input is obtained by downsampling from real images. The training losses also include reconstruction loss, adversarial loss, and remote sensing consistency loss. Finally, with the parameters of the primary generator and the first-stage generator fixed, the second-stage generator is trained to generate 1024×1024 resolution images based on the output of the first-stage generator. The training dataset includes 100,000 real remote sensing images at 1024×1024 resolution and their corresponding remote sensing metadata.

[0035] This embodiment also includes a dynamic window sampling strategy for generating remote sensing images of arbitrary sizes. The dynamic window sampling strategy divides the large target image into multiple overlapping image patches, each patch being 1024×1024 pixels in size, which is the training size of the generator. The width of the overlapping region is 512 pixels, half the size of the image patch, ensuring the continuity of adjacent image patches at the boundaries. For each image patch, the same initial noise and conditional vector are used for generation, ensuring the consistency of the generated content in the overlapping region. Specifically, for an image patch Pi (i=1, 2, ..., N), its top-left corner coordinates are (xi, yi), and the overlapping region is Pi∩P_{i+1}. During generation, the same random seed is used to initialize the noise in the overlapping region, and the same location encoding (using the actual geographic coordinates of the image patch) is used. All image patches are stitched together according to the average pixel value of the overlapping region. For pixels in the overlapping region, the average value of corresponding pixels in adjacent image patches is taken. Finally, the stitched image undergoes edge smoothing processing to eliminate possible stitching artifacts.

[0036] This embodiment also includes multimodal remote sensing information fusion. Multimodal remote sensing information fusion supports the generation of various data types, including optical remote sensing images, radar remote sensing images, and hyperspectral remote sensing images. For optical remote sensing images, a three-channel RGB or multispectral channel generation method is used, resulting in a three-channel RGB image or a multispectral image (such as Landsat's seven bands). For radar remote sensing images, joint generation of amplitude and phase features is introduced, using a complex convolutional network to process radar data, generating a complex image containing amplitude and phase maps. For hyperspectral remote sensing images, a spectral cascade generation strategy is adopted, first generating images with sparse bands (e.g., sampling once every 10 bands), and then gradually generating images with complete spectral bands (e.g., 200 bands). Multimodal remote sensing information fusion is achieved through a unified conditional vector representation and targeted network architecture design. Different modalities of images use different encoder and decoder structures, but share the core design of the remote sensing information encoder and feature fusion module. Example 2

[0037] This embodiment provides a specific application example of the method of the present invention. Taking the generation of a 0.5-meter resolution optical remote sensing image of a certain urban area as an example, the specific implementation steps are as follows: The remote sensing metadata information of the target area was obtained as follows: geographic coordinates (116.4°E, 39.9°N), imaging time was 10:00 AM on March 20, 2025, sensor type was high-resolution optical sensor, spectral range was visible light band (450-700nm), imaging resolution was 0.5 m / pixel, imaging angle was solar altitude angle of 45°, and sensor elevation angle was 90° (vertically downward).

[0038] Remote sensing metadata is input into the remote sensing information encoder to generate a 128-dimensional conditional vector. Geographic coordinates are encoded as 64-dimensional vectors, and imaging parameters are encoded as 64-dimensional vectors. These are then fused into a 128-dimensional conditional vector through multi-head attention.

[0039] Initialize with random noise distributed according to a standard normal distribution N(0,1). Input the conditional vector along with the random noise into the primary generator to produce a low-resolution image of 64×64 pixels. This image shows the macroscopic layout of the target area, including the city outline, main roads, large buildings, etc.

[0040] The 64×64 resolution image and conditional vector are input into the first-level generator to generate a 256×256 resolution image. This image displays mesoscale features, including street networks, building clusters, and green space distribution.

[0041] During the generation process, the remote sensing feature fusion module extracts the features of flat terrain and dense buildings in the region from the geographic feature library based on the geographic coordinates, generates corresponding spectral response features based on optical sensors and visible light bands, and injects these features into the feature map of the generator through channel attention.

[0042] The 256×256 resolution image and conditional vectors are input into the second-stage generator to produce a 1024×1024 resolution remote sensing image at 0.5 meters per pixel. This image displays high-resolution details, including building outlines, road textures, and vehicle shapes.

[0043] The remote sensing consistency constraint module verifies the generated images. Geometric consistency verification shows that the registration error between the generated image and the reference map is 2.3 meters, less than the 5-meter threshold, meeting the requirements. Spectral consistency verification shows that the spectral angular similarity of the generated image is 0.92, greater than the 0.95 threshold, slightly below the requirement. Therefore, the spectral feature injection weights are adjusted, and the image is regenerated. The second generated image has a spectral angular similarity of 0.96, meeting the requirements.

[0044] The final output is a 1024×1024 resolution remote sensing image. This image performs well in terms of visual quality, ground feature accuracy, and spectral fidelity, and can meet the needs of practical applications. Example 3

[0045] This embodiment provides comparative experimental results between the method of the present invention and existing methods.

[0046] Experimental setup: The method of this invention was compared with three existing methods on the same dataset: StableDiffusion (a general image generation method), MetaEarth (a global remote sensing image generation method), and GAN-based method (a remote sensing image generation method based on generative adversarial networks). The dataset contains 10,000 remote sensing images and their corresponding remote sensing metadata, covering various geographical regions and imaging conditions.

[0047] Evaluation indicators include: FID (Fréchet Inception Distance): Measures the difference in distribution between the generated image and the real image in the feature space; the smaller the value, the better.

[0048] LPIPS (Learned Perceptual Image Patch Similarity): Measures the perceptual similarity between a generated image and a real image; the lower the better.

[0049] Geometric consistency: The registration error between the generated image and the reference map; the smaller the better.

[0050] Spectral consistency: The similarity of the spectral angles between the generated image and the real image; the higher the better.

[0051] The experimental results are shown in the table below: The experimental results show that the method of this invention significantly outperforms existing methods in all evaluation metrics. In particular, the performance improvement of the method is most pronounced in geometric consistency and spectral consistency metrics, which fully demonstrates the effectiveness of remote sensing information guidance and consistency constraints.

[0052] Furthermore, the method of this invention also has advantages in generation efficiency. Generating a 1024×1024 resolution image requires 12.5 seconds from Stable Diffusion, 8.3 seconds from MetaEarth, and 6.7 seconds from GAN-based methods, while the method of this invention only requires 4.2 seconds. This is mainly due to the optimized design of the cascaded generation architecture, which avoids the computational burden of directly generating high-resolution images.

[0053] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A method for cascaded generation of remote sensing images driven by remote sensing information, characterized in that, Includes the following steps: S1. Obtain remote sensing metadata information, which includes at least one of geographic coordinates, imaging time, sensor type, spectral range, imaging resolution, and imaging angle; S2. Construct a remote sensing information encoder to encode the remote sensing metadata information into a conditional vector of a multi-layer cascaded generator; S3. Construct a multi-level cascaded generator based on the diffusion model. The multi-level cascaded generator includes a primary generator and at least one secondary generator. Each generator is responsible for generating images at different resolutions. S4. Using a sampling strategy guided by remote sensing information, the conditional vector is used to guide the multi-level cascaded generator to generate an initial low-resolution remote sensing image. S5. The geographic features and spectral features in the remote sensing metadata are integrated into the generated image features through the remote sensing feature fusion module; S6. The secondary generator is used sequentially to upsample and enhance the details of the low-resolution image to generate a remote sensing image of the target resolution. S7. Apply the remote sensing consistency constraint module to perform geometric and spectral consistency checks on the generated remote sensing images; S8. Output high-resolution remote sensing images that satisfy the consistency of remote sensing features; The remote sensing information encoder adopts a network structure that combines a multilayer perceptron and an attention mechanism. The multi-level cascaded generator adopts a U-Net-based architecture. The remote sensing feature fusion module integrates remote sensing features into the image generation process using feature recalibration and attention weighting. The remote sensing consistency constraint module uses georegistration error and spectral angle similarity as consistency evaluation indicators.

2. The remote sensing information-driven cascaded generation method for remote sensing images according to claim 1, characterized in that, The remote sensing information encoder includes a geographic coordinate encoding submodule, an imaging parameter encoding submodule, and a fusion submodule; The geographic coordinate encoding submodule encodes latitude and longitude coordinates using sine and cosine functions to form a 64-dimensional geographic feature vector. The imaging parameter encoding submodule converts imaging time, sensor type, spectral range, imaging resolution, and imaging angle into corresponding feature vectors through an embedding layer; The fusion submodule uses a multi-head attention mechanism to fuse the geographic feature vector and the imaging parameter feature vector, and outputs the condition vector.

3. The remote sensing information-driven cascaded generation method for remote sensing images according to claim 1, characterized in that, The multi-stage cascade generator adopts a three-stage cascade structure; The primary generator produces 64×64 resolution remote sensing images, with this stage focusing on capturing the macroscopic layout of the scene and the main land cover categories; The first-level generator upsamples the 64×64 resolution image to 256×256 resolution, focusing on generating medium-scale ground features and texture details; The second-stage generator upsamples the 256×256 resolution image to 1024×1024 resolution, generating high-resolution fine textures and edge details; Each generator performs a diffusion denoising process guided by a conditional vector, using either the DDPM or DDIM sampling algorithm.

4. The remote sensing information-driven cascaded generation method for remote sensing images according to claim 1, characterized in that, The remote sensing feature fusion module includes a geographic feature extraction submodule and a spectral feature injection submodule; The geographic feature extraction submodule extracts the topographic features, land cover distribution features, and climate features of the corresponding region from a pre-built geographic feature database based on geographic coordinate information; The spectral feature injection submodule generates corresponding spectral response curve features based on sensor type and spectral range information, and injects the spectral features into the feature map of the generator through a channel attention mechanism; The remote sensing feature fusion module performs remote sensing feature fusion at each stage of cascade generation to ensure that the generated image is consistent with the remote sensing metadata in both spatial and spectral dimensions.

5. The remote sensing information-driven cascaded generation method for remote sensing images according to claim 1, characterized in that, The remote sensing consistency constraint module includes a geometric consistency verification submodule and a spectral consistency verification submodule; The geometric consistency verification submodule calculates the geographic registration error between the generated image and the reference geographic information system data. When the registration error exceeds a preset threshold, the generated image is geometrically corrected. The spectral consistency verification submodule calculates the difference between the spectral angle similarity of the generated image and the statistical distribution of the real remote sensing image. When the difference exceeds a preset threshold, the spectral feature injection weight is adjusted. The remote sensing consistency constraint module performs consistency optimization in the final stage of cascade generation. The optimization objective function is Ltotal = Lrecon + αLgeo + βLspectral, where Lrecon is the reconstruction loss, Lgeo is the geometric consistency loss, Lspectral is the spectral consistency loss, and α and β are balance coefficients.

6. The remote sensing information-driven cascaded generation method for remote sensing images according to claim 1, characterized in that, The remote sensing information-guided sampling strategies include conditional enhancement sampling and frequency-aware sampling; The conditional augmentation sampling randomly discards some remote sensing metadata information with a certain probability during the training process, thereby improving the model's robustness to missing conditions. The frequency-aware sampling divides the sampling process into multiple stages, with each stage focusing on image features in different frequency ranges. The low-frequency stage uses a lower conditional guidance strength, while the high-frequency stage uses a higher conditional guidance strength. The sampling strategy employs a classifier-free guided method, which controls the generation process by weighting the combination of unconditional and conditional noise predictions.

7. The remote sensing information-driven cascaded generation method for remote sensing images according to claim 1, characterized in that, The method also includes a cascaded training strategy; The cascaded training strategy adopts a bottom-up training sequence, first training the primary generator to enable it to generate high-quality 64×64 resolution images. Then, with the parameters of the primary generator fixed, the first primary generator is trained so that it can generate 256×256 resolution images based on the output of the primary generator. Finally, fix the parameters of the primary generator and the first-level generator, and train the second-level generator so that it can generate a 1024×1024 resolution image based on the output of the first-level generator. The training loss for each generator includes reconstruction loss, adversarial loss, and remote sensing consistency loss.

8. The remote sensing information-driven cascaded generation method for remote sensing images according to claim 1, characterized in that, The method also includes a dynamic window sampling strategy for generating remote sensing images of arbitrary size; The dynamic window sampling strategy divides the target large image into multiple overlapping image blocks, and the size of each image block is the training size of the generator. The width of the overlapping region is half the size of the image patch to ensure the continuity of adjacent image patches at the boundary; For each image patch, the same initial noise and conditional vector are used for generation to ensure the consistency of the generated content in overlapping areas; All image patches are stitched together according to the average pixel value of the overlapping areas to generate the final borderless large-size remote sensing image.

9. The remote sensing information-driven cascaded generation method for remote sensing images according to claim 1, characterized in that, The method also includes multimodal remote sensing information fusion; The multimodal remote sensing information fusion supports the generation of various data types, including optical remote sensing images, radar remote sensing images, and hyperspectral remote sensing images; For optical remote sensing images, RGB three-channel or multispectral channel generation methods are used; For radar remote sensing images, a joint generation of amplitude and phase features is introduced, and a complex convolutional network is used to process the radar data. For hyperspectral remote sensing images, a spectral cascade generation strategy is adopted, first generating images with sparse bands, and then gradually generating images with complete spectral bands. Multimodal remote sensing information fusion is achieved through unified conditional vector representation and targeted network architecture design.