Multi-modal satellite image fusion method based on end-to-end architecture

Through the multimodal satellite image fusion method with an end-to-end architecture, multi-time phase and multi-source satellite images are directly processed to generate high-quality super-resolution images, solving the problems of complex processes and data quality loss in traditional methods, and achieving efficient data processing and storage optimization.

CN120387938APending Publication Date: 2025-07-29SCHOOL OF SOFTWARE ZHEJIANG UNIV (NINGBO) MANAGEMENT CENT (NINGBO SOFTWARE EDUCATION CENT) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510445238.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing satellite image fusion technology separates and processes multi-phase and multi-source data, resulting in cumbersome processes, complex data screening, and traditional compression methods sacrifice data quality to reduce data volume.

Method used

Using an end-to-end architecture, multi-time phase multi-channel low-resolution images and high-resolution single-channel images are directly processed through multi-time phase image fusion, multi-source image fusion and image combination steps. Using feature extraction and adaptive fusion strategies, super-resolution multi-channel images with high spatial resolution and high spectral resolution are generated.

Benefits of technology

Simplify the data processing process, reduce calculation and storage costs, and at the same time reduce the amount of data while ensuring data quality and improve data utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387938A_ABST
    Figure CN120387938A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal satellite image fusion method based on an end-to-end architecture, and belongs to the field of artificial intelligence multi-modal data fusion. For multi-temporal and multi-source satellite image data, the method comprises the steps of multi-temporal image fusion, multi-source image fusion and fused image combination. In the multi-temporal image fusion step, the low-resolution multi-temporal multi-channel image is subjected to feature alignment, sampling and fusion, and a super-resolution multi-channel image is generated. In the multi-source image fusion step, the high-resolution single-channel image and the multi-temporal fusion image are combined to generate a super-resolution multi-channel image. And in the fused image combination step, pixel-level channel-by-channel addition is carried out on the fused result, and a final super-resolution multi-channel image is formed. According to the method, the problems of multi-temporal and multi-source fusion separation and high extra data screening requirement in traditional satellite image fusion are solved, meanwhile, the storage space occupation is reduced, and the data transmission pressure of the satellite Internet of Things is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence multi-modal data fusion and computer vision. Specifically, it relates to a multi-temporal and multi-source satellite image fusion method based on an end-to-end architecture. Background Art

[0002] With the development of low-earth orbit (LEO) small satellite technology, the satellite Internet of Things gradually constructs a constellation system, enabling multi-temporal image data to be acquired at the same location within a short period. Since the observation period of a single small satellite is relatively short, multiple satellites can image the same target area within a short time, thus forming multi-temporal satellite remote sensing data at the same location. However, due to the limitations of the volume, weight, and energy consumption of small satellites, the optical imaging devices carried by them usually have a relatively low spatial resolution. And limited by the physical characteristics of the sensors, satellite sensors cannot simultaneously obtain images with high spatial resolution and high spectral resolution. Currently, satellite remote sensing imaging usually uses two types of sensors: one can acquire single-channel panchromatic (Pan) images with high spatial resolution, and the other can acquire multi-spectral (MS) images with high spectral resolution but low spatial resolution. Therefore, within the same observation area, Pan-MS images can be obtained simultaneously.

[0003] Due to the complementary characteristics of different imaging devices in terms of spatial and spectral resolution, how to effectively fuse multi-temporal and multi-source remote sensing data to improve the image quality while reducing the data storage and transmission costs is an urgent problem to be solved in the application of the satellite Internet of Things. In addition, ground base stations need to process a large amount of satellite remote sensing data. Limited by the storage and transmission capabilities, the problem of data backlog is becoming increasingly prominent. Therefore, the development of efficient multi-modal satellite image fusion technology is of great significance for improving the data utilization rate and application value of the satellite Internet of Things.

[0004] Current satellite image fusion technologies still face challenges in practical applications. First, the fusion process is separated. Existing technologies separately process multi-temporal fusion and multi-source fusion, resulting in a cumbersome data processing process and limited fusion effects. Second, data screening is complex. Due to the huge amount of multi-temporal data, existing methods are non-end-to-end models and rely on complex preprocessing and screening steps, increasing the computational and storage costs. Third, the effect of compressing the data volume is limited. Traditional methods mainly rely on image compression and selection to reduce the data volume. The above methods will sacrifice part of the data quality while reducing the data volume and cannot reduce the data volume while ensuring the data quality. Summary of the Invention

[0005] Aiming at the above technical deficiencies, the present invention proposes a multi-modal satellite image fusion method based on an end-to-end architecture, aiming to efficiently process multi-temporal and multi-source satellite image data.

[0006] To achieve the above object, the present invention includes the following steps:

[0007] A multi-temporal image fusion step is used to align, sample, and fuse the input low-resolution multi-temporal multi-channel images in feature space to generate a first single super-resolution multi-channel image;

[0008] a multi-source image fusion step for fusing the input high-resolution single single-channel image with the first single super-resolution multi-channel image to generate a second single super-resolution multi-channel image;

[0009] The fusion image combination step is used to combine the first single super-resolution multi-channel image and the second single super-resolution multi-channel image to generate a final third single super-resolution multi-channel image.

[0010] The technical effects achieved by the present invention are as follows:

[0011] (1) End-to-end fusion simplifies the data processing process: Traditional methods usually process multi-temporal fusion and multi-source fusion separately, resulting in a complex data processing process and limited fusion effect. The present invention adopts an end-to-end method, directly inputs multi-temporal multi-channel low-resolution images, and fuses high-resolution single-channel images, achieving overall optimization from data input to the final high-quality fused image. It avoids the tedious independent processing steps of splitting the multi-temporal image fusion and multi-source image fusion work, and performing quality screening of the input images before multi-temporal image fusion, thereby improving the efficiency of data processing and the fusion effect.

[0012] (2) Reduced data screening and preprocessing costs: Existing methods rely on complex data screening and preprocessing steps to reduce the impact of abnormal images in multi-temporal images on the fusion results. However, these non-end-to-end processes increase computational and storage costs. The present invention introduces sampling technology to screen the feature space and directly performs adaptive optimization on multi-temporal data during the fusion process, eliminating the need for additional data screening steps, thereby effectively reducing computational complexity and storage requirements.

[0013] (3) Reducing data volume while ensuring data quality: Traditional methods mainly rely on image compression and selection to reduce data volume, but such methods usually sacrifice some data quality, resulting in information loss in the fused image. The present invention fully utilizes the information redundancy of multi-temporal and multi-source data during the fusion process. Through feature extraction and adaptive fusion strategies, it can not only reduce data storage and transmission overhead, but also ensure high spatial resolution and high spectral quality of the final fused image, thereby effectively improving data utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0015] Figure 1 It is the network architecture diagram of a multi-modal satellite image fusion method based on an end-to-end architecture according to the present invention;

[0016] Figure 2 It is the network structure diagram of the multi-temporal image fusion step in a multi-modal satellite image fusion method based on an end-to-end architecture according to the present invention;

[0017] Figure 3 It is the network structure diagram of the multi-source image fusion step in a multi-modal satellite image fusion method based on an end-to-end architecture according to the present invention;

[0018] Reference numerals: 100: multi-temporal image fusion step; 110: multi-source image fusion step; 120: image combination step. Detailed implementation manners

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0020] As Figure 1 shown, a multi-modal satellite image fusion method based on an end-to-end architecture, which is carried out based on a multi-modal satellite image fusion model of an end-to-end architecture. During the multi-modal satellite image fusion process, it includes a multi-temporal image fusion step, a multi-source image fusion step, and an image combination step. The specific process includes the following steps:

[0021] Step S1, data acquisition and preprocessing: Collect multiple paired image data for training. Each paired image data refers to "input as T low-resolution multi-temporal multi-channel images (low spatial resolution but high spectral resolution) for the same location and a single high-resolution single-channel image (high spatial resolution but low spectral resolution); the output reference is a single high-resolution multi-channel image for the same location. (low spatial resolution but high spectral resolution) and a single high-resolution single-channel image (high spatial resolution but low spectral resolution); the output reference is a single high-resolution multi-channel image for the same location (High spatial resolution and high spectral resolution), and perform cropping and data format conversion to meet the input requirements of the model architecture;

[0022] In this application, "high resolution" refers to an image with relatively high spatial resolution (such as a panchromatic image), while "super resolution" refers to enhancing the spatial resolution of a low-resolution image through an algorithm to make it reach or approach the high-resolution level;

[0023] Step S2, multi-temporal data fusion, see Figure 2 , including the following steps:

[0024] Step S2.1, in order to establish the connection between multiple multi-temporal multi-channel images, calculate the average value of the input low-resolution multi-temporal multi-channel images at the pixel level for each channel as the reference image and perform channel-wise splicing with the input multiple images Use a convolutional neural network with shared parameters to calculate the deep features of each input image The convolutional neural network with shared parameters can adapt to any number of input images and is a lightweight model;

[0025]

[0026] Step S2.2, use a convolutional neural network with shared parameters to further extract the deep feature space of each input image In order to sample the deep features beneficial to fusion, use the Krum algorithm to calculate the reference features Calculate the Euclidean distance between other features and and select N corresponding features for selecting fusion features;

[0027]

[0028]

[0029] Step S2.3, use a neural network to fuse the sampled features to form fusion features

[0030]

[0031] Step S2.4, in order to convert the fusion features into a super-resolution multi-channel image, use a convolutional neural network to adjust the number of channels of the fusion features Use sub-pixel convolution PixelShuffle with an amplification factor of r to generate a single super-resolution multi-channel image i.e., the first single - sheet super - resolution multi - channel image;

[0032]

[0033] Step S3, the multi - source data fusion step, see Figure 3 , the input high - resolution single - channel single - sheet image and the super - resolution multi - channel image SR formed by the multi - temporal image fusion step MISR are stitched in the channel dimension through a neural network to fuse them into a super - resolution multi - channel image i.e., the second single - sheet super - resolution multi - channel image;

[0034]

[0035] Step S4, the fused image combination step, includes the following steps: To make greater use of the respective advantages of the super - resolution image formed by the multi - temporal fusion step and the super - resolution image formed by the multi - source fusion step, the super - resolution multi - channel image SR generated by the multi - temporal image fusion step MISR and the super - resolution multi - channel image SR generated by the multi - source image fusion step Sharpening are added channel - by - channel at the pixel level to generate the final single - sheet super - resolution multi - channel image i.e., the third single - sheet super - resolution multi - channel image; The finally generated super - resolution multi - channel image has both high spatial resolution and high spectral resolution;

[0036]

[0037] Step S5, perform model training, calculate the error between SR and HR target , and use the backpropagation gradient descent algorithm to adjust the model parameters proposed by the present invention batch by batch.

Claims

1. A multi-modal satellite image fusion method based on an end-to-end architecture, characterized in that, This method is based on a multi-modal satellite image fusion model with an end-to-end architecture. During the multi-modal satellite image fusion process, it includes the following steps: The multi-temporal image fusion step is used to align, sample, and fuse the feature spaces of multiple input low-resolution multi-temporal multi-channel images to generate a first single super-resolution multi-channel image; The multi-source image fusion step is used to fuse the input high-resolution single-channel image with the first single super-resolution multi-channel image to generate a second single super-resolution multi-channel image; The fused image combination step is used to combine the first single super-resolution multi-channel image and the second single super-resolution multi-channel image to generate a final third single super-resolution multi-channel image.

2. A multi-modal satellite image fusion method based on an end-to-end architecture according to claim 1, wherein: The multi-temporal image fusion step includes: Using a convolutional neural network with shared parameters as an encoder to perform deep feature extraction on the input multiple images; Intelligently sampling the deep features to select the deep features of the images beneficial for fusion; Fusing the selected multiple deep features to form a fused feature; Using a convolutional neural network as a decoder to calculate the fused feature to form a first single super-resolution multi-channel image.

3. A multi-modal satellite image fusion method based on an end-to-end architecture according to claim 1, wherein: The multi-source image fusion step includes: splicing the input high-resolution single-channel image and the first single super-resolution multi-channel image formed in the multi-temporal image fusion step in the channel dimension, and fusing them through a neural network to form a second super-resolution multi-channel image.

4. A multi-modal satellite image fusion method based on an end-to-end architecture according to claim 1, wherein: The fused image combination step includes: adding the first single super-resolution multi-channel image and the second single super-resolution multi-channel image channel by channel at the pixel level to generate a final third single super-resolution multi-channel image.

5. A multi-modal satellite image fusion method based on an end-to-end architecture according to claim 2, wherein: The step of performing deep feature extraction on multiple input images includes: calculating the average value of multiple multi-temporal multi-channel input low-resolution images channel by channel at the pixel level as the reference image LR ref , concatenating it with the multiple input images in the channel dimension, and using a convolutional neural network with shared parameters to calculate the deep features of each input image 6. A multi-modal satellite image fusion method based on an end-to-end architecture according to claim 5, wherein: The step of intelligently sampling deep features includes: extracting the deep feature space of each input image using a convolutional neural network with shared parameters Calculating reference features based on the deep feature space using the Krum algorithm Calculating the deep features and the reference features The Euclidean distance of, and selecting N corresponding features 7. A multi-modal satellite image fusion method based on an end-to-end architecture according to claim 6, wherein: The step of fusing the selected multiple deep features to form a fused feature includes: using a neural network to fuse the features to form a fused feature F fusion .

8. A multi-modal satellite image fusion method based on an end-to-end architecture according to claim 7, wherein: The computing step of using a convolutional neural network as a decoder for the fused features includes: using a convolutional neural network to adjust the number of channels F of the fused features decoder , using sub-pixel convolution to generate a first single super-resolution multi-channel image SR MISR .

9. A multi-modal satellite image fusion method based on an end-to-end architecture according to claim 3, wherein: The multi-source image fusion steps are as follows: The input single-channel high-resolution image HR pan is concatenated with the first single-image super-resolution multi-channel image SR MISR in the channel dimension [HR pan , SR MISR , and is fused through a neural network to form the second single-image super-resolution multi-channel image SR Sharpening .

10. A multimodal satellite image fusion method based on an end-to-end architecture according to claim 1, characterized in that: It also includes training a multi-modal satellite image fusion model, calculating the error between the third single-image super-resolution multi-channel image SR and the reference single high-resolution multi-channel image HR target and using the backpropagation gradient descent algorithm to adjust the model parameters batch by batch.