An image fusion method, device, storage medium and electronic equipment

CN117456323BActive Publication Date: 2026-09-25ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311350663.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-17
Publication Date
2026-09-25
Estimated Expiration
2043-10-17

AI Technical Summary

Technical Problem

[0003]然而,遥感设备的性能限制导致高分辨率多光谱(High-resolution multi-spectral,简称HSMS)图像不能直接通过拍摄的方式获取,卫星传感器通常只能捕获覆盖同一地域范围的全色(panchromatic,简称PAN)图像和与全色图像对应的低分辨率多光谱(Low-resolution multi-spectral,简称LSMS)图像

Benefits of technology

[0061]本说明书提供的图像融合的方法,在图像融合时,采集目标区域的全色图像以及低分辨率多光谱图像。将全色图像调整为与低分辨率多光谱图像的分辨率相同的第一图像。将第一图像输入预先训练的图像融合模型的特征提取子网,确定全色特征,以及将低分辨率多光谱图像输入特征提取子网,确定光谱特征。确定全色特征对应的频域特征,以及确定光谱特征对应的频域特征。将全色特征对应的频域特征以及光谱特征对应的频域特征输入图像融合模型的特征融合子网,确定第一融合特征,并根据第一融合特征,确定第一融合图像。再将第一融合图像调整为与全色图像的分辨率相同的第二图像。将全色图像输入特征提取子网,确定第一特征,以及将第二图像输入特征提取子网,确定第二特征。确定第一特征对应的频域特征,以及确定第二特征对应的频域特征。再将第一特征对应的频域特征以及第二特征对应的频域特征输入特征融合子网,确定第二融合特征。根据第二融合特征,确定目标区域的目标融合图像。采样低高互补的阶段融合方式,以及以融合图像的频域特征为主进行图像融合,可以避免目标融合图像出现光谱扭曲问题,使得目标融合图像更加准确。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117456323B_ABST
    Figure CN117456323B_ABST
Patent Text Reader

Abstract

The specification discloses a method and device for image fusion, a storage medium and an electronic device, comprising: inputting a first image and a low-resolution multispectral image into a feature extraction subnetwork of a pre-trained image fusion model respectively to determine a panchromatic feature and a spectral feature; inputting frequency domain features of the panchromatic feature and frequency domain features of the spectral feature into a feature fusion subnetwork of the image fusion model to determine a first fusion feature; determining a first fusion image according to the first fusion feature; determining a first feature and a second feature through the feature extraction subnetwork according to a panchromatic image and a second image; inputting frequency domain features of the first feature and frequency domain features of the second feature into the feature fusion subnetwork to determine a second fusion feature; and determining a target fusion image according to the second fusion feature. The image fusion is mainly based on the fusion of frequency domain features and supplemented by the fusion of spatial domain features, thereby improving the accuracy of the target fusion image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a method, apparatus, storage medium and electronic device for image fusion. Background Technology

[0002] With the continuous development of science and technology and the rapid development of remote sensing technology, multispectral images have provided rich data support for work such as ground feature analysis, ground feature classification, and image interpretation. Multispectral images are also widely used in fields such as agricultural production, mineral monitoring, and environmental protection.

[0003] However, the performance limitations of remote sensing equipment mean that high-resolution multispectral (HSMS) images cannot be directly acquired through imaging. Satellite sensors typically can only capture panchromatic (PAN) images covering the same geographical area and their corresponding low-resolution multispectral (LSMS) images. Since panchromatic images are high-resolution images, they can be fused with low-resolution multispectral images to generate a high-resolution multispectral image with high resolution in both the spatial and spectral domains. Therefore, how to fuse panchromatic and low-resolution multispectral images to obtain a high-resolution multispectral image is a very important issue.

[0004] Based on this, this specification provides a method for image fusion. Summary of the Invention

[0005] This specification provides a method, apparatus, storage medium, and electronic device for image fusion, in order to partially solve the aforementioned problems existing in the prior art.

[0006] The following technical solution is adopted in this specification:

[0007] This specification provides a method for image fusion, including:

[0008] Acquire panchromatic and low-resolution multispectral images of the target area;

[0009] The panchromatic image is adjusted to a first image with the same resolution as the low-resolution multispectral image;

[0010] The first image is input into the feature extraction subnet of a pre-trained image fusion model to determine panchromatic features, and the low-resolution multispectral image is input into the feature extraction subnet to determine spectral features;

[0011] Determine the frequency domain features corresponding to the panchromatic features, and determine the frequency domain features corresponding to the spectral features;

[0012] The frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features are input into the feature fusion subnet of the image fusion model to determine the first fusion feature, and the first fusion image is determined based on the first fusion feature.

[0013] The first fused image is adjusted to a second image with the same resolution as the panchromatic image;

[0014] The panchromatic image is input into the feature extraction sub-network to determine the first feature, and the second image is input into the feature extraction sub-network to determine the second feature;

[0015] Determine the frequency domain feature corresponding to the first feature, and determine the frequency domain feature corresponding to the second feature;

[0016] The frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature are input into the feature fusion subnet to determine the second fusion feature;

[0017] Based on the second fusion feature, the target fused image of the target region is determined.

[0018] Optionally, the frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features are input into the feature fusion subnetwork of the image fusion model to determine the first fusion feature, specifically including:

[0019] The panchromatic features and the spectral features are fused to determine the first spatial feature;

[0020] Determine the frequency domain features corresponding to the first spatial domain features;

[0021] The first spatial feature, the frequency domain feature corresponding to the panchromatic feature, the frequency domain feature corresponding to the spectral feature, and the frequency domain feature corresponding to the first spatial feature are input into the feature fusion subnet of the image fusion model to determine the first fusion feature.

[0022] Optionally, the feature fusion subnet includes a first fusion layer and a second fusion layer;

[0023] The first spatial feature, the frequency domain feature corresponding to the panchromatic feature, the frequency domain feature corresponding to the spectral feature, and the frequency domain feature corresponding to the first spatial feature are input into the feature fusion subnet of the image fusion model to determine the first fusion feature, specifically including:

[0024] The frequency domain features corresponding to the panchromatic features, the frequency domain features corresponding to the spectral features, and the frequency domain features corresponding to the first spatial features are input into the first fusion layer of the feature fusion subnet of the image fusion model to determine the first frequency domain features.

[0025] The first frequency domain feature and the first spatial domain feature are input into the second fusion layer of the feature fusion subnet of the image fusion model to determine the first fusion feature.

[0026] Optionally, the frequency domain features corresponding to the panchromatic features, the frequency domain features corresponding to the spectral features, and the frequency domain features corresponding to the first spatial features are input into the first fusion layer of the feature fusion subnet of the image fusion model to determine the first frequency domain features, specifically including:

[0027] The high-frequency component corresponding to the panchromatic feature is determined to be a panchromatic high-frequency component, and the low-frequency component corresponding to the panchromatic feature is determined to be a panchromatic low-frequency component; the high-frequency component corresponding to the spectral feature is determined to be a spectral high-frequency component, and the low-frequency component corresponding to the spectral feature is determined to be a spectral low-frequency component; the high-frequency component corresponding to the first spatial feature is determined to be a first spatial high-frequency component, and the low-frequency component corresponding to the first spatial feature is determined to be a first spatial low-frequency component.

[0028] The panchromatic high-frequency component, the panchromatic low-frequency component, the spectral high-frequency component, the spectral low-frequency component, the first spatial high-frequency component, and the first spatial low-frequency component are input into the first fusion layer of the feature fusion subnet of the image fusion model. The first fusion layer fuses the panchromatic high-frequency component, the spectral high-frequency component, and the first spatial high-frequency component to determine the first high-frequency component, and fuses the panchromatic low-frequency component, the spectral low-frequency component, and the first spatial low-frequency component to determine the first low-frequency component.

[0029] The first frequency domain feature is determined based on the first high-frequency component and the first low-frequency component.

[0030] Optionally, the image fusion model further includes a cross-attention fusion layer;

[0031] The panchromatic features and the spectral features are fused to determine the first spatial feature, specifically including:

[0032] The panchromatic features and the spectral features are input into the cross-attention fusion layer of the image fusion model to determine the first spatial features.

[0033] Optionally, an image fusion model is pre-trained, specifically including:

[0034] The panchromatic image of a specified area collected in history is identified as the first sample, and the low-resolution multispectral image corresponding to the first sample is identified as the second sample.

[0035] Adjust the first sample to a first image with the same resolution as the second sample;

[0036] The first image is input into the feature extraction subnet of the image fusion model to be trained to determine the panchromatic features, and the second sample is input into the feature extraction subnet to determine the spectral features;

[0037] Determine the frequency domain features corresponding to the panchromatic features, and determine the frequency domain features corresponding to the spectral features;

[0038] The frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features are input into the feature fusion subnetwork of the image fusion model to be trained, the first fusion feature is determined, and the first fusion image is determined based on the first fusion feature.

[0039] The first fused image is adjusted to a second image with the same resolution as the first sample;

[0040] The first sample is input into the feature extraction subnetwork to determine the first feature, and the second image is input into the feature extraction subnetwork to determine the second feature;

[0041] Determine the frequency domain feature corresponding to the first feature, and determine the frequency domain feature corresponding to the second feature;

[0042] The frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature are input into the feature fusion subnet to determine the second fusion feature, and the second fusion image is determined based on the second fusion feature.

[0043] The reference image of the specified region is determined as the first annotation, and the reference image is adjusted to the same resolution as the second sample and used as the second annotation;

[0044] The image fusion model to be trained is trained based on the first fused image, the second fused image, the first annotation, and the second annotation.

[0045] Optionally, the image fusion model to be trained is trained based on the first fused image, the second fused image, the first annotation, and the second annotation, specifically including:

[0046] The image fusion model to be trained is trained with the goal of minimizing the difference between the first fused image and the second annotation, and minimizing the difference between the second fused image and the first annotation.

[0047] This specification provides an image fusion apparatus, comprising:

[0048] The acquisition module is used to acquire panchromatic images and low-resolution multispectral images of the target area;

[0049] A first processing module is used to adjust the panchromatic image to a first image with the same resolution as the low-resolution multispectral image;

[0050] The first spatial feature extraction module is used to input the first image into the feature extraction subnet of a pre-trained image fusion model to determine panchromatic features, and to input the low-resolution multispectral image into the feature extraction subnet to determine spectral features;

[0051] The first frequency domain feature extraction module is used to determine the frequency domain features corresponding to the panchromatic features and to determine the frequency domain features corresponding to the spectral features.

[0052] The first feature fusion module is used to input the frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features into the feature fusion subnet of the image fusion model to determine the first fusion feature, and to determine the first fusion image based on the first fusion feature;

[0053] The second processing module is used to adjust the first fused image into a second image with the same resolution as the panchromatic image;

[0054] The second spatial feature extraction module is used to input the panchromatic image into the feature extraction sub-network to determine the first feature, and to input the second image into the feature extraction sub-network to determine the second feature;

[0055] The second frequency domain feature extraction module is used to determine the frequency domain feature corresponding to the first feature and to determine the frequency domain feature corresponding to the second feature.

[0056] The second feature fusion module is used to input the frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature into the feature fusion subnet to determine the second fused feature.

[0057] The determining module is used to determine the target fused image of the target region based on the second fusion feature.

[0058] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described image fusion method.

[0059] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described image fusion method.

[0060] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:

[0061] The image fusion method provided in this specification involves acquiring a panchromatic image and a low-resolution multispectral image of the target region during image fusion. The panchromatic image is adjusted to a first image with the same resolution as the low-resolution multispectral image. The first image is input into the feature extraction subnetwork of a pre-trained image fusion model to determine panchromatic features, and the low-resolution multispectral image is input into the feature extraction subnetwork to determine spectral features. Frequency domain features corresponding to the panchromatic features and spectral features are determined. The frequency domain features corresponding to the panchromatic and spectral features are input into the feature fusion subnetwork of the image fusion model to determine a first fusion feature, and a first fused image is determined based on the first fusion feature. The first fused image is then adjusted to a second image with the same resolution as the panchromatic image. The panchromatic image is input into the feature extraction subnetwork to determine a first feature, and the second image is input into the feature extraction subnetwork to determine a second feature. The frequency domain features corresponding to the first and second features are determined. The frequency domain features corresponding to the first and second features are then input into the feature fusion subnetwork to determine a second fusion feature. The target fused image of the target region is determined based on the second fusion feature. The phased fusion method of sampling low and high complementarity, and the image fusion based on the frequency domain features of the fused image, can avoid the spectral distortion problem in the target fused image, making the target fused image more accurate. Attached Figure Description

[0062] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:

[0063] Figure 1 This is a flowchart illustrating an image fusion method provided in this specification;

[0064] Figure 2 This is a schematic diagram of a cross-attention fusion layer provided in this specification;

[0065] Figure 3 This is a schematic diagram of the structure of a spatial attention layer provided in this specification;

[0066] Figure 4 This is a schematic diagram of the structure of a channel attention layer provided in this specification;

[0067] Figure 5 This is a schematic diagram of the structure of a feature extraction subnetwork provided in this specification;

[0068] Figure 6 This is a schematic diagram of the structure of an image fusion model provided in this specification;

[0069] Figure 7 This is a schematic diagram of an image fusion apparatus provided in this specification;

[0070] Figure 8 This specification provides a corresponding Figure 1 A schematic diagram of the structure of an electronic device. Detailed Implementation

[0071] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.

[0072] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0073] Figure 1 This is a flowchart illustrating an image fusion method provided in this specification, including the following steps:

[0074] S100: Acquires panchromatic images and low-resolution multispectral images of the target area.

[0075] Currently, when performing tasks such as feature classification and target detection in a specific area, it is necessary to first acquire images of that area to perform these tasks. However, existing remote sensing equipment cannot directly capture high-resolution multispectral images; it can only capture high-resolution panchromatic images and low-resolution multispectral images of the area. While the panchromatic image has high resolution, it lacks spectral information, and although the low-resolution multispectral image contains spectral information, its low resolution results in an unclear image. To better perform feature classification and target detection based on the images of the area, it is necessary to generate a high-resolution multispectral image based on the high-resolution panchromatic image and the low-resolution multispectral image.

[0076] Therefore, in this specification, the device used for image fusion can acquire panchromatic images and low-resolution multispectral images of the target area. The device used for image fusion can be a server or an electronic device such as a desktop computer or laptop. For ease of description, the image fusion method provided in this specification will be described below using a server as the execution entity. The target area is the area where a high-resolution multispectral image needs to be generated, or it can be the area where target detection, land cover classification, or other tasks are required. This target area can be a residential area, a street, or the sky, etc.

[0077] Specifically, the server can determine the panchromatic image and low-resolution multispectral image of the target area from images of the target area acquired by an image acquisition device, which can be a satellite sensor. The server can also acquire the panchromatic image and low-resolution multispectral image of the target area in response to a user's service request. This service request can be a target detection task request, a land cover classification task request, or any other request that requires image-based services for a region; this specification does not specifically limit this. The service request includes information about the target area, which at least includes the location of the target area.

[0078] S102: Adjust the panchromatic image to a first image with the same resolution as the low-resolution multispectral image.

[0079] Because the resolution of the panchromatic image differs from that of the low-resolution multispectral image (the panchromatic image has a higher resolution), to better fuse the panchromatic and low-resolution multispectral images, the panchromatic image can first be adjusted to have the same resolution as the low-resolution multispectral image. Then, image fusion is performed based on the adjusted image and the low-resolution multispectral image. Based on this, the server can adjust the panchromatic image to obtain a first image with the same resolution as the low-resolution multispectral image. Specifically, the server can downsample the panchromatic image to obtain a first image with the same resolution as the low-resolution multispectral image.

[0080] S104: Input the first image into the feature extraction subnet of the pre-trained image fusion model to determine panchromatic features, and input the low-resolution multispectral image into the feature extraction subnet to determine spectral features.

[0081] The server can input the first image into the feature extraction subnet of a pre-trained image fusion model to determine panchromatic features, and input the low-resolution multispectral image into the feature extraction subnet to determine spectral features. The image fusion model is a pre-trained model that includes a feature extraction subnet and a feature fusion subnet. The feature extraction subnet is used to extract the spatial features of the image (i.e., features in the image space). Both the panchromatic and spectral features mentioned above are feature maps.

[0082] S106: Determine the frequency domain features corresponding to the panchromatic features, and determine the frequency domain features corresponding to the spectral features.

[0083] Since panchromatic and spectral features are both spatial features, image fusion based on spatial features is prone to problems such as blurred edges and artifacts. Therefore, the server can first convert the spatial features into frequency domain features, and then perform image fusion based on the frequency domain features to avoid problems such as blurred edges and artifacts in the fused image.

[0084] Based on this, the server can determine the frequency domain features corresponding to panchromatic features and the frequency domain features corresponding to spectral features. Specifically, when determining the frequency domain features corresponding to panchromatic features, the server can perform a two-dimensional discrete Fourier transform on the panchromatic features to generate the corresponding frequency domain features. Similarly, when determining the frequency domain features corresponding to spectral features, the server can perform a two-dimensional discrete Fourier transform on the spectral features to generate the corresponding frequency domain features. The transformation can be performed using the following formula:

[0085]

[0086] in, This represents the frequency domain feature, and the dimension of this frequency domain feature is... Img is a spatial feature with dimensions H×W. This spatial feature can be a panchromatic feature or a spectral feature.

[0087] Additionally, the server can input panchromatic features or spectral features into the first transformation layer to determine the frequency domain features corresponding to the panchromatic features, or the frequency domain features corresponding to the spectral features. The first transformation layer converts spatial domain features into frequency domain features. This first transformation layer can be a pre-trained network layer or any existing first transformation layer. Of course, it can also be a network layer in the image fusion model, a network layer trained together with the image fusion model; therefore, the image fusion model also includes a first transformation layer.

[0088] S108: Input the frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features into the feature fusion subnet of the image fusion model to determine the first fusion feature, and determine the first fusion image based on the first fusion feature.

[0089] The server can input the frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features into the feature fusion subnet of the image fusion model to determine the first fusion feature, and then determine the first fused image based on the first fusion feature. Specifically, the server can input the frequency domain features corresponding to the panchromatic features, the frequency domain features corresponding to the spectral features, the panchromatic features, and the spectral features into the feature fusion subnet of the image fusion model to determine the first fusion feature, and then determine the first fused image based on the first fusion feature.

[0090] In the feature fusion subnetwork, in addition to fusing frequency domain features, spatial domain features can also be fused to ensure that the spectral information of the image is not lost. Furthermore, when fusing frequency and spatial features, the frequency domain features are generally first converted into their corresponding spatial domain features, and then the converted features are fused with the original spatial domain features. Therefore, the server inputs the frequency domain features corresponding to the panchromatic features, the frequency domain features corresponding to the spectral features, the panchromatic features, and the spectral features into the feature fusion subnetwork. The feature fusion subnetwork first converts the frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features respectively, and then fuses the converted features, panchromatic features, and spectral features to determine the first fused feature, which better represents the spatial and spectral information. The conversion of frequency domain features to spatial features can be performed using inverse Fourier transform, or any other existing method; this specification does not specifically limit this. For ease of explanation, the following descriptions of converting frequency domain features to spatial features all use inverse Fourier transform as an example.

[0091] Furthermore, the server can first fuse frequency domain features, and then fuse the fused frequency domain features with spatial domain features. Therefore, the feature fusion subnet can include a first fusion layer and a second fusion layer. The first fusion layer is used to fuse frequency domain features, and the second fusion layer is used to fuse frequency domain features and spatial domain features. Thus, the server can input the frequency domain features corresponding to panchromatic features and the frequency domain features corresponding to spectral features into the first fusion layer of the feature fusion subnet of the image fusion model to determine the first frequency domain feature. Then, the first frequency domain feature, panchromatic feature, and spectral feature are input into the second fusion layer of the feature fusion subnet of the image fusion model to determine the first fused feature. Specifically, when fusing the first frequency domain feature, panchromatic feature, and spectral feature, the second fusion layer first transforms the first frequency domain feature, and then fuses the transformed features, panchromatic feature, and spectral feature.

[0092] To better integrate frequency and spatial features, the server can first fuse panchromatic and spectral features to determine the first spatial feature. Then, the first spatial feature, the corresponding frequency features of the panchromatic feature, and the corresponding frequency features of the spectral feature are input into the feature fusion subnet of the image fusion model to determine the first fused feature. Alternatively, the server can first fuse panchromatic and spectral features to determine the first spatial feature. The corresponding frequency features of the panchromatic feature and the corresponding frequency features of the spectral feature are then input into the first fusion layer of the feature fusion subnet of the image fusion model to determine the first frequency feature. Finally, the first frequency feature and the first spatial feature are input into the second fusion layer of the feature fusion subnet of the image fusion model to determine the first fused feature.

[0093] In addition, when determining the first fusion feature, besides fusing the frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features, the frequency domain features corresponding to the first spatial features can also be fused. Therefore, the server can also fuse the panchromatic features and the spectral features to determine the first spatial feature. Then, the frequency domain features corresponding to the first spatial feature are determined, and the first spatial feature, the frequency domain features corresponding to the panchromatic features, the frequency domain features corresponding to the spectral features, and the frequency domain features corresponding to the first spatial feature are input into the feature fusion subnet of the image fusion model to determine the first fusion feature. The process of determining the frequency domain features corresponding to the first spatial feature is similar to the process of determining the frequency domain features corresponding to the panchromatic features or the spectral features, and will not be elaborated further here.

[0094] Alternatively, the server can fuse panchromatic and spectral features to determine the first spatial feature. Then, the frequency domain feature corresponding to the first spatial feature is determined. Next, the frequency domain features corresponding to the panchromatic, spectral, and first spatial features are input into the first fusion layer of the feature fusion subnet of the image fusion model to determine the first frequency domain feature. Finally, the first frequency domain feature and the first spatial feature are input into the second fusion layer of the feature fusion subnet of the image fusion model to determine the first fused feature.

[0095] In this specification, when fusing panchromatic and spectral features to determine the first spatial feature, the server directly fuses the panchromatic and spectral features to determine the first spatial feature. Alternatively, the server can input the panchromatic and spectral features into the cross-attention fusion layer of the image fusion model to determine the first spatial feature. The image fusion model also includes a cross-attention fusion layer, which is used to fuse panchromatic and spectral features in both spatial and channel dimensions.

[0096] When determining the first fused image based on the first fusion feature, the server can input the first fusion feature into the output layer to determine the first fused image. This output layer is used to convert image features into an image. This output layer can be any existing network layer, or it can be a network layer in the image fusion model, pre-trained together with the image fusion model. Therefore, the image fusion model can also include an output layer. Of course, the server can also use any existing algorithm to determine the first fused image based on the first fusion feature; this specification does not impose specific limitations.

[0097] S110: Adjust the first fused image to a second image with the same resolution as the panchromatic image.

[0098] Since step S102 involves downsampling the panchromatic image to obtain an image with the same resolution as the low-resolution multispectral image, subsequent steps S104-S108 determine the first fused image based on the first image (i.e., the downsampled panchromatic image) and the low-resolution multispectral image. The resolution of this first fused image is consistent with that of the low-resolution multispectral image. Therefore, to obtain a high-resolution multispectral image, the first fused image needs to be adjusted to have the same resolution as the panchromatic image. Then, image fusion is performed based on the adjusted image and the panchromatic image to obtain the high-resolution multispectral image. Based on this, the server can adjust the first fused image to a second image with the same resolution as the panchromatic image. Specifically, the server can upsample the first fused image to obtain a second image with the same resolution as the panchromatic image.

[0099] S112: Input the panchromatic image into the feature extraction sub-network to determine the first feature, and input the second image into the feature extraction sub-network to determine the second feature.

[0100] The server can input a panchromatic image into the feature extraction subnet to determine the first feature, and input a second image into the feature extraction subnet to determine the second feature. Both the first and second features are feature maps.

[0101] S114: Determine the frequency domain feature corresponding to the first feature, and determine the frequency domain feature corresponding to the second feature.

[0102] The server can determine the frequency domain feature corresponding to the first feature and the frequency domain feature corresponding to the second feature. The process of determining the frequency domain feature corresponding to the first feature or the second feature is similar to the process of determining the frequency domain feature corresponding to the panchromatic feature in step S106 above, except that the panchromatic feature is replaced by the first feature or the second feature, which will not be described again here.

[0103] S116: Input the frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature into the feature fusion subnet to determine the second fusion feature.

[0104] The server can input the frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature into the feature fusion subnet to determine the second fusion feature. Specifically, the server can input the frequency domain features corresponding to the first feature, the frequency domain features corresponding to the second feature, the first feature, and the second feature into the feature fusion subnet to determine the second fusion feature. The specific process is similar to the process in step S108 above, where the frequency domain features corresponding to the panchromatic feature, the frequency domain features corresponding to the spectral feature, the panchromatic feature, and the spectral feature are input into the feature fusion subnet of the image fusion model to determine the first fusion feature, and will not be repeated here.

[0105] Alternatively, the server can first fuse frequency domain features, and then fuse the fused frequency domain features with spatial domain features. Therefore, the server can input the frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature into the first fusion layer of the feature fusion subnet of the image fusion model to determine the second frequency domain feature. Then, the second frequency domain feature, the first feature, and the second feature are input into the second fusion layer of the feature fusion subnet of the image fusion model to determine the second fused feature.

[0106] To better integrate frequency and spatial features, the server can first fuse the first and second features to determine the second spatial feature. Then, it inputs the second spatial feature, the corresponding frequency features of the first feature, and the corresponding frequency features of the second feature into the feature fusion subnet of the image fusion model to determine the second fused feature. Alternatively, the server can first fuse the first and second features to determine the second spatial feature. Then, it inputs the corresponding frequency features of the first and second features into the first fusion layer of the feature fusion subnet of the image fusion model to determine the second frequency feature. Finally, it inputs the second frequency feature and the second spatial feature into the second fusion layer of the feature fusion subnet of the image fusion model to determine the second fused feature.

[0107] Furthermore, when determining the second fusion feature, in addition to fusing the frequency domain features corresponding to the first feature and the second feature, the frequency domain feature corresponding to the second spatial feature can also be fused. Therefore, the server can also fuse the first feature and the second feature to determine the second spatial feature. Then, the frequency domain feature corresponding to the second spatial feature is determined, and the second spatial feature, the frequency domain feature corresponding to the first feature, the frequency domain feature corresponding to the second feature, and the frequency domain feature corresponding to the second spatial feature are input into the feature fusion subnet of the image fusion model to determine the second fusion feature. The process of determining the frequency domain feature corresponding to the second spatial feature is similar to the process of determining the frequency domain feature corresponding to the panchromatic feature or spectral feature, and will not be elaborated further here.

[0108] Alternatively, the server can fuse the first and second features to determine the second spatial feature. Then, it determines the frequency domain feature corresponding to the second spatial feature. Next, it inputs the frequency domain features corresponding to the first, second, and second spatial features into the first fusion layer of the feature fusion subnet of the image fusion model to determine the second frequency domain feature. Finally, it inputs the second frequency domain feature and the second spatial feature into the second fusion layer of the feature fusion subnet of the image fusion model to determine the second fused feature.

[0109] In this specification, the process of fusing the first feature and the second feature to determine the second spatial feature is similar to the process of fusing the panchromatic feature and the spectral feature to determine the first spatial feature in step S108, except that the panchromatic feature in step S108 is replaced with the first feature and the spectral feature is replaced with the second feature, which will not be described again here.

[0110] S118: Determine the target fused image of the target region based on the second fusion feature.

[0111] The server can determine the target fused image of the target region based on the second fusion feature. The target fused image is a high-resolution multispectral image of the target region. When determining the target fused image based on the second fusion feature, the server can input the second fusion feature into the output layer to determine the second fused image. This output layer is used to convert image features into an image. This output layer can be any existing network layer, or it can be a network layer in the image fusion model, pre-trained together with the image fusion model; therefore, the image fusion model can also include an output layer. Of course, the server can also use any existing algorithm based on the second fusion feature to determine the second fused image; this specification does not impose specific limitations.

[0112] As can be seen from the above method, in this application, during image fusion, the server first acquires a panchromatic image and a low-resolution multispectral image of the target region. Then, the panchromatic image is adjusted to a first image with the same resolution as the low-resolution multispectral image. The first image is then input into the feature extraction subnet of a pre-trained image fusion model to determine panchromatic features, and the low-resolution multispectral image is input into the feature extraction subnet to determine spectral features. Next, the frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features are determined. The frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features are input into the feature fusion subnet of the image fusion model to determine the first fusion feature, and based on the first fusion feature, the first fusion image is determined. Then, the first fusion image is adjusted to a second image with the same resolution as the panchromatic image. The panchromatic image is input into the feature extraction subnet to determine the first feature, and the second image is input into the feature extraction subnet to determine the second feature. The frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature are then determined. Finally, the frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature are input into the feature fusion subnet to determine the second fusion feature. Next, based on the second fusion feature, the target fused image of the target region is determined. This is achieved by first downsampling the high-resolution panchromatic image and then fusing it with the low-resolution multispectral image using an image fusion model to determine the first fused image. Then, the first fused image is upsampled and fused again with the high-resolution panchromatic image using the same image fusion model to determine the final target fused image. This phased fusion method, which uses complementary low-to-high sampling, avoids spectral distortion in the target fused image, resulting in a more accurate target fused image.

[0113] Furthermore, when performing image fusion based on the image fusion model, the frequency domain features of the fused image are the primary focus, with the spatial domain features as a secondary consideration. The fused image is then obtained based on these fused features. This approach avoids issues such as edge blurring and artifacts in the fused image while preserving more spatial and spectral information, resulting in a more accurate fused image.

[0114] In this specification, to achieve finer-grained fusion of frequency domain features, the frequency domain features can be decomposed into finer-grained high-frequency and low-frequency components. Frequency domain feature fusion is then performed at this fine-grained dimension of high-frequency and low-frequency components, resulting in better fusion, improved fusion quality, and a more accurate fused image. The aforementioned high-frequency components refer to features in regions of image brightness, grayscale, and other intensity variations. The aforementioned low-frequency components refer to features in regions of image brightness, grayscale, and other intensity variations that are gradual.

[0115] Based on this, in step S108 above, the frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features are input into the first fusion layer of the feature fusion subnet of the image fusion model. When determining the first frequency domain feature, the server can determine that the high-frequency component corresponding to the panchromatic feature is the panchromatic high-frequency component, and the low-frequency component corresponding to the panchromatic feature is the panchromatic low-frequency component. Similarly, the server determines that the high-frequency component corresponding to the spectral feature is the spectral high-frequency component, and the low-frequency component corresponding to the spectral feature is the spectral low-frequency component. Then, the panchromatic high-frequency component, panchromatic low-frequency component, spectral high-frequency component, and spectral low-frequency component are input into the first fusion layer of the feature fusion subnet of the image fusion model. Through the first fusion layer, the panchromatic high-frequency component and the spectral high-frequency component are fused to determine the first high-frequency component, and the panchromatic low-frequency component and the spectral low-frequency component are fused to determine the first low-frequency component. Finally, based on the first high-frequency component and the first low-frequency component, the first frequency domain feature is determined.

[0116] Based on this, in step S108 above, the frequency domain features corresponding to the panchromatic features, the frequency domain features corresponding to the spectral features, and the frequency domain features corresponding to the first spatial features are input into the first fusion layer of the feature fusion subnet of the image fusion model. When determining the first frequency domain feature, the server can determine that the high-frequency component corresponding to the panchromatic feature is the panchromatic high-frequency component, and the low-frequency component corresponding to the panchromatic feature is the panchromatic low-frequency component. Simultaneously, the server determines that the high-frequency component corresponding to the spectral features is the spectral high-frequency component, and the low-frequency component corresponding to the spectral features is the spectral low-frequency component. Furthermore, the server determines that the high-frequency component corresponding to the first spatial feature is the first spatial high-frequency component, and the low-frequency component corresponding to the first spatial feature is the first spatial low-frequency component. Subsequently, the panchromatic high-frequency component, panchromatic low-frequency component, spectral high-frequency component, spectral low-frequency component, first spatial high-frequency component, and first spatial low-frequency component are input into the first fusion layer of the feature fusion subnet of the image fusion model. Through the first fusion layer, the panchromatic high-frequency component, spectral high-frequency component, and first spatial high-frequency component are fused to determine the first high-frequency component, and the panchromatic low-frequency component, spectral low-frequency component, and first spatial low-frequency component are fused to determine the first low-frequency component. Then, based on the first high-frequency component and the first low-frequency component, the first frequency domain feature is determined.

[0117] In this specification, when determining the high-frequency and low-frequency components corresponding to the panchromatic features, spectral features, or first spatial features, the server needs to first determine the frequency domain features corresponding to the panchromatic features, spectral features, or first spatial features, and then determine the high-frequency and low-frequency components based on the determined frequency domain features. The following formula can be used for calculation:

[0118]

[0119]

[0120] in, This represents the high-frequency components after Fourier shift and masking calculations. This represents the high-pass filter mask. This indicates a shift of the frequency domain features. Represents frequency domain characteristics, This indicates a conversion from the spatial domain to the frequency domain. The frequency domain feature can be a panchromatic feature, a spectral feature, or a frequency domain feature corresponding to the first spatial domain feature. This represents the low-frequency components after Fourier shift and masking calculations. This represents a low-pass filter mask.

[0121] The above-mentioned input of panchromatic high-frequency components, panchromatic low-frequency components, spectral high-frequency components, spectral low-frequency components, first spatial high-frequency components, and first spatial low-frequency components into the first fusion layer of the feature fusion subnet of the image fusion model, to fuse the panchromatic high-frequency components, spectral high-frequency components, and first spatial high-frequency components through the first fusion layer to determine the first high-frequency component, and to fuse the panchromatic low-frequency components, spectral low-frequency components, and first spatial low-frequency components to determine the first low-frequency component, the server can input the panchromatic high-frequency components, spectral high-frequency components, and first spatial high-frequency components into the residual inversion layer to determine the first output result, and input the panchromatic low-frequency components, spectral low-frequency components, and first spatial low-frequency components into the residual inversion layer to determine the second output result. The first output result is then input into the first fusion layer of the feature fusion subnet of the image fusion model to determine the first high-frequency component, and the second output result is input into the first fusion layer of the feature fusion subnet of the image fusion model to determine the first low-frequency component.

[0122] The residual inversion layer can be any existing network layer, or it can be a network layer in the image fusion model, and it is trained together with the image fusion model. Therefore, the image fusion model also includes a residual inversion layer. The process of determining the first high-frequency component and the first low-frequency component can be calculated using the following formula:

[0123]

[0124]

[0125] in, Indicates the first high-frequency component. This indicates the first fusion layer. Indicates the residual inversion layer. Represents the panchromatic high-frequency components. Represents the high-frequency components of the spectrum. This represents the high-frequency components in the first spatial domain. Indicates the first low-frequency component. Indicates the full-color low-frequency component. Represents the low-frequency components of the spectrum. This represents the low-frequency component in the first spatial domain.

[0126] In this specification, in step S108 above, when the first frequency domain feature, panchromatic feature, and spectral feature are input into the second fusion layer of the feature fusion subnet of the image fusion model to determine the first fusion feature, the server needs to first convert the first frequency domain feature, and then input the converted feature, panchromatic feature, and spectral feature into the second fusion layer of the feature fusion subnet of the image fusion model to determine the first fusion feature. The converted feature is a spatial domain feature.

[0127] In addition, in step S108 above, the first frequency domain feature and the first spatial domain feature are input into the second fusion layer of the feature fusion subnet of the image fusion model to determine the first fusion feature. The server needs to first convert the first frequency domain feature, and then input the converted feature and the first spatial domain feature into the second fusion layer of the feature fusion subnet of the image fusion model to determine the first fusion feature. Specifically, the following formula can be used for calculation:

[0128]

[0129] in, Indicates the first high-frequency component. Indicates the first low-frequency component. This indicates an inverse shift of the frequency domain features. This indicates a conversion from the frequency domain to the spatial domain. denoted by , where SF1 represents the first spatial domain feature, and ⊙ represents the feature superposition calculation. This represents the first fusion feature.

[0130] In step S116 above, the frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature are input into the first fusion layer of the feature fusion subnet of the image fusion model. When determining the second frequency domain feature, the server can determine the high-frequency component and the low-frequency component corresponding to the first feature. Then, the high-frequency component, the low-frequency component, the high-frequency component, and the low-frequency component of the second feature are input into the first fusion layer of the feature fusion subnet of the image fusion model. Through the first fusion layer, the high-frequency component and the high-frequency component of the first feature are fused to determine the second high-frequency component, and the low-frequency component and the low-frequency component of the first feature are fused to determine the second low-frequency component. Finally, based on the second high-frequency component and the second low-frequency component, the second frequency domain feature is determined.

[0131] Based on this, in step S116 above, the frequency domain features corresponding to the first feature, the frequency domain features corresponding to the second feature, and the frequency domain features corresponding to the second spatial feature are input into the first fusion layer of the feature fusion subnet of the image fusion model. When determining the second frequency domain feature, the server can determine the high-frequency and low-frequency components corresponding to the first feature. The high-frequency and low-frequency components corresponding to the second feature are also determined, as are the high-frequency and low-frequency components corresponding to the second spatial feature. Then, the high-frequency and low-frequency components corresponding to the first feature, the high-frequency and low-frequency components corresponding to the second feature, the high-frequency and low-frequency components corresponding to the second spatial feature, and the first fusion layer of the feature fusion subnet of the image fusion model are input to determine the second high-frequency component. The first fusion layer then fuses the high-frequency components corresponding to the first feature, the high-frequency components corresponding to the second feature, and the high-frequency components corresponding to the second spatial feature to determine the second high-frequency component, and fuses the low-frequency components corresponding to the first feature, the low-frequency components corresponding to the first feature, and the low-frequency components corresponding to the second spatial feature to determine the second low-frequency component. Finally, the second frequency domain feature is determined based on the second high-frequency and second low-frequency components.

[0132] In this specification, the process of determining the high-frequency and low-frequency components corresponding to the first feature, the second feature, or the second spatial feature is similar to the process of determining the high-frequency and low-frequency components corresponding to the panchromatic feature, the spectral feature, or the first spatial feature, and will not be repeated here.

[0133] The above-mentioned high-frequency component, low-frequency component, high-frequency component, low-frequency component, high-frequency component, and low-frequency component of the second spatial feature corresponding to the first feature fusion subnet of the image fusion model are input into the first fusion layer of the feature fusion subnet of the image fusion model. Through the first fusion layer, the high-frequency component, low-frequency component, and high-frequency component of the second spatial feature are fused to determine the second high-frequency component. Similarly, when determining the second low-frequency component, the server can input the high-frequency component, low-frequency component, and high-frequency component of the second spatial feature into the residual inversion layer to determine the third output result, and input the low-frequency component, low-frequency component, and low-frequency component of the second spatial feature into the residual inversion layer to determine the fourth output result. Then, the first output result is input into the first fusion layer of the feature fusion subnet of the image fusion model to determine the first high-frequency component, and the second output result is input into the first fusion layer of the feature fusion subnet of the image fusion model to determine the first low-frequency component.

[0134] The process of determining the second high-frequency component and the second low-frequency component can be calculated using the following formula:

[0135]

[0136]

[0137] in, Indicates the second high-frequency component. This represents the first fusion layer. Indicates the residual inversion layer. This represents the high-frequency component corresponding to the first feature. This represents the high-frequency component corresponding to the second feature. This represents the high-frequency component corresponding to the second spatial feature. This represents the second low-frequency component. This represents the low-frequency component corresponding to the first feature. This represents the low-frequency component corresponding to the second feature. This represents the low-frequency component corresponding to the second spatial feature.

[0138] In this specification, when the second frequency domain feature, the first feature, and the second feature are input into the second fusion layer of the feature fusion subnet of the image fusion model in step S116 above to determine the second fusion feature, the server needs to first convert the second frequency domain feature, and then input the converted feature, the first feature, and the second feature into the second fusion layer of the feature fusion subnet of the image fusion model to determine the second fusion feature.

[0139] In addition, in step S116 above, the second frequency domain feature and the second spatial domain feature are input into the second fusion layer of the feature fusion subnet of the image fusion model. When determining the second fusion feature, the server needs to first convert the second frequency domain feature, and then input the converted feature and the second spatial domain feature into the second fusion layer of the feature fusion subnet of the image fusion model to determine the second fusion feature. Specifically, the following formula can be used for calculation:

[0140]

[0141]

[0142] in, Indicates the second high-frequency component. This represents the second low-frequency component. denoted by , represents the feature after frequency domain feature transformation, SF2 represents the second spatial domain feature, and ⊙ represents feature superposition calculation. This indicates the second fusion feature.

[0143] In this specification, because different feature channels have varying importance, and features at different locations within the same feature channel also have different importance, the server can input the raw features into a cross-attention fusion layer to perform attention weighting in both spatial and channel dimensions. Based on this, as... Figure 2 As shown, Figure 2 This is a schematic diagram of a cross-attention layer fusion layer provided in this specification. Figure 2 In the image fusion model, circles containing the character "X" represent product operations, circles containing the character "+" represent summation operations, and circles containing the character "C" represent concatenation operations. Since the process of inputting panchromatic and spectral features into the cross-attention fusion layer of the image fusion model in step S108 to determine the first spatial domain feature is similar to the process of fusing the first and second features in step S116 to determine the second spatial domain feature, the following combination... Figure 2 The process of determining the first spatial domain feature by inputting panchromatic and spectral features into the cross-attention fusion layer of the image fusion model will be used as an example for illustration.

[0144] Specifically, the cross-attention fusion layer can include a spatial attention layer and a channel attention layer. When inputting panchromatic and spectral features into the cross-attention fusion layer of the image fusion model to determine the first spatial feature, the server can input the panchromatic features into the spatial attention layer to obtain the first attention feature. The first attention feature is then multiplied by the panchromatic feature, and the product is summed with the panchromatic feature. The summed feature is then input into the channel attention layer to obtain the second attention feature. The server also inputs spectral features into the spatial attention layer to obtain the third attention feature. The third attention feature is then multiplied by the spectral feature, and the product is summed with the spectral feature. The summed feature is then input into the channel attention layer to obtain the fourth attention feature. Next, the second and fourth attention features are concatenated, and the concatenated feature is input into the spatial attention layer to obtain the fifth attention feature. The fifth attention feature is then multiplied by the fourth attention feature to obtain the product result. The product result is summed with the second attention feature to obtain the sum result. The sum result is then input into the channel attention layer to obtain the first spatial feature.

[0145] Among them, such as Figure 3 As shown, Figure 3 This is a schematic diagram of the structure of a spatial attention layer provided in this specification. Figure 3The presence of an "X" character in the circle represents a product operation, while the presence of a "C" character represents a concatenation operation. The spatial attention layer can include a first network layer, a first average pooling layer, a first max pooling layer, a second network layer, and a third network layer. Specifically, taking the input of panchromatic features into the spatial attention layer to obtain the first attention feature as an example, the server inputs the panchromatic features into the first network layer, then inputs the output of the first network layer into the first average pooling layer and the first max pooling layer respectively, and then concatenates the outputs of the first average pooling layer and the first max pooling layer to obtain the concatenated result. The concatenated result is input into the second network layer, and the output of the second network layer is multiplied by the panchromatic feature to obtain the first result. The first result is then input into the third network layer to obtain the first attention feature output by the third network layer.

[0146] The aforementioned channel attention layer may include a fourth network layer, a second average pooling layer, a second max pooling layer, a multilayer perceptron, a fifth network layer, and a sixth network layer, such as... Figure 4 As shown, Figure 4 This is a schematic diagram of the structure of a channel attention layer provided in this specification. Figure 4 The circle containing the character "X" represents the product operation. Specifically, taking the summed features input into the channel attention layer to obtain the second attention feature as an example, the server inputs the summed features into the fourth network layer. The output of the fourth network layer is then input into the second average pooling layer and the second max pooling layer, respectively. The outputs of the second average pooling layer and the second max pooling layer are then input into the multilayer perceptron to obtain the output. This output is then input into the fifth network layer, where its output is multiplied by the summed features to obtain the second result. Finally, the second result is input into the sixth network layer to obtain the second attention feature output by the sixth network layer.

[0147] In this specification, to fully explore the intrinsic relationships between features and capture more feature information, the aforementioned feature extraction subnetwork is a multi-layer deep feature extraction subnetwork. This feature extraction subnetwork includes a high-level feature extraction layer, a mid-level feature extraction layer, a low-level feature extraction layer, and a convolutional output layer. Each of the high-level, mid-level, and low-level feature extraction layers includes three convolutional groups, and there is a one-to-one correspondence between convolutional groups in any feature extraction layer and convolutional groups in other feature extraction layers. By cascading and fusing the three feature extraction layers, the feature extraction subnetwork can be viewed as a multi-level detail injection subnetwork, enabling centralized processing of usable feature information from panchromatic and low-resolution multispectral images. It can also filter out features that have a beneficial impact on the fused image, thereby suppressing redundant information in the features.

[0148] Each convolutional group in the high-level feature extraction layers includes a convolutional layer with a 3×3 kernel. Each convolutional group in the mid-level feature extraction layers includes two convolutional layers with a 3×3 kernel. Each convolutional group in the low-level feature extraction layers includes one convolutional layer with a 1×1 kernel and two convolutional layers with a 3×3 kernel. The convolutional output layer includes a convolutional layer with a 1×1 kernel, such as... Figure 5 As shown, Figure 5 This is a schematic diagram of the structure of a feature extraction subnet provided in this specification. Figure 5 The circle containing the character "C" represents a concatenation operation, and the circle containing the character "+" also represents a concatenation operation.

[0149] Combination Figure 5 Specifically, taking step S104 above, where the first image is input into the feature extraction subnet of a pre-trained image fusion model to determine panchromatic features, as an example, for each convolutional group of the low-level feature extraction layer, the server can input the first image into the convolutional group to obtain the first convolution result. The convolutional group of the corresponding intermediate feature extraction layer is determined as the intermediate convolutional group. The first convolution result output by this convolutional group is input into the corresponding intermediate convolutional group to obtain the second convolution result. The second convolution result is then summed with the first convolution result to obtain the summed result. The convolutional group of the corresponding high-level feature extraction layer is determined as the high-level convolutional group. The summed result is input into the high-level convolutional group to obtain the third convolution result. The third convolution results output by each convolutional group in the high-level feature extraction layer are determined, and the third convolution results are concatenated to obtain the concatenated result. The concatenated result is then input into the convolutional output layer to obtain the panchromatic features.

[0150] To ensure that the number of channels for both panchromatic and spectral features is the same, facilitating subsequent image fusion, in step S104, the server can preprocess the first image and the low-resolution multispectral image to guarantee that the number of channels in the processed first image and the processed low-resolution multispectral image is the same. Therefore, the server can preprocess the first image to obtain a processed first image with a specified number of channels, and preprocess the low-resolution multispectral image to obtain a processed spectral image with a specified number of channels. Then, the server inputs the processed first image into the feature extraction subnet of the pre-trained image fusion model to determine the panchromatic features, and inputs the processed multispectral image into the feature extraction subnet to determine the spectral features.

[0151] The specified number of channels can be a pre-set number, such as 1. When preprocessing the first image to obtain a processed first image with the specified number of channels, the first image can be first input into a convolutional layer with a 3×3 kernel to obtain the output result, and then the output result can be input into a convolutional layer with a 1×1 kernel to obtain the processed first image. Similarly, when preprocessing a low-resolution multispectral image to obtain a processed spectral image with the specified number of channels, the low-resolution multispectral image can be input into a convolutional layer with a 3×3 kernel to obtain the processed spectral image.

[0152] Furthermore, to ensure that the number of channels for the first feature and the second feature are the same, facilitating subsequent image fusion, the server can preprocess both the panchromatic image and the second image to guarantee that the number of channels in the processed panchromatic image and the processed second image are the same. Therefore, in step S112 above, the server can preprocess the panchromatic image to obtain a processed panchromatic image with a specified number of channels, and preprocess the second image to obtain a processed second image with a specified number of channels. Then, the server inputs the processed panchromatic image into the feature extraction subnet to determine the first feature, and inputs the processed second image into the feature extraction subnet to determine the second feature. The specific process is similar to that in step S104 above, except that the first image in step S104 is replaced with a panchromatic image, and the low-resolution multispectral image is replaced with the second image; therefore, it will not be elaborated further here.

[0153] In this specification, after obtaining the target fusion image of the target region, when classifying objects in the target region, the objects in the target region can be classified based on the target fusion image to determine the category of each object in the target region. Of course, when performing target detection on objects in the target region, target detection can also be performed based on the target fusion image to determine whether a target object exists in the target region.

[0154] In this specification, the image fusion model may include a feature extraction subnetwork, a cross-attention fusion layer, a residual inversion layer, a feature fusion subnetwork, and an output layer. The feature extraction subnetwork may include a high-level feature extraction layer, a mid-level feature extraction layer, a low-level feature extraction layer, and a convolutional output layer. The cross-attention fusion layer may include a spatial attention layer and a channel attention layer. The feature fusion subnetwork may include a first fusion layer and a second fusion layer, with a specific structure as shown below. Figure 6 As shown, Figure 6This is a schematic diagram of the image fusion model provided in this specification. The following description uses the process of determining the first fused image based on a first image and a low-resolution multispectral image as an example. Specifically, the server can input the first image and the low-resolution multispectral image into the feature extraction subnetwork to determine the panchromatic features corresponding to the first image and the spectral features corresponding to the low-resolution multispectral image. Then, the panchromatic features and spectral features are input into the cross-attention fusion layer to determine the first spatial features. Next, the frequency domain features corresponding to the panchromatic features, spectral features, and the first spatial features are determined. Based on the determined frequency domain features, the high-frequency components and low-frequency components corresponding to the panchromatic features, spectral features, and the first spatial features are determined. The determined high-frequency components and low-frequency components are input into the first fusion layer to determine the first high-frequency component and the first low-frequency component. Then, based on the first high-frequency component and the first low-frequency component, the first frequency domain feature is determined. The first frequency domain feature and the first spatial feature are input into the second fusion layer to determine the first fused feature. Finally, the first fused feature is input into the output layer to determine the first fused image.

[0155] In this specification, the image fusion model described above may include a feature extraction subnetwork and a feature fusion subnetwork. During pre-training of the image fusion model, the server can determine a historically acquired panchromatic image of a specified region as the first sample, and determine the corresponding low-resolution multispectral image as the second sample. Then, the first sample is adjusted to have the same resolution as the second sample. The first image is then input into the feature extraction subnetwork of the image fusion model to be trained to determine panchromatic features, and the second sample is input into the feature extraction subnetwork to determine spectral features. Next, the frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features are determined. The frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features are then input into the feature fusion subnetwork of the image fusion model to be trained to determine the first fusion feature, and based on the first fusion feature, a first fused image is determined. The first fused image is then adjusted to have the same resolution as the first sample. The first sample is input into the feature extraction subnetwork to determine the first feature, and the second image is input into the feature extraction subnetwork to determine the second feature. Finally, the frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature are determined. The frequency domain features corresponding to the first feature and the second feature are then input into the feature fusion subnet to determine the second fusion feature, and based on the second fusion feature, the second fusion image is determined. Next, a reference image for a specified region is determined as the first annotation, and the reference image is adjusted to have the same resolution as the second sample and used as the second annotation. The image fusion model to be trained is then trained based on the first fusion image, the second fusion image, the first annotation, and the second annotation.

[0156] The reference image is a high-resolution multispectral image of a specified region, with the same resolution as the panchromatic image of that region. The process of training the image fusion model can be divided into two main stages. The first stage includes determining the first image based on the first sample, and obtaining the first fused image based on the first image and the second sample. The second stage includes determining the second image based on the first fused image, and obtaining the second fused image based on the second image and the first sample.

[0157] The specific process of the first stage is similar to steps S102 to S108 above, except that steps S102 to S108 are the application process of the image fusion model, while the first stage is the training process of the image fusion model. Therefore, it will not be described in detail here. Similarly, the specific process of the second stage is similar to steps S110 to S118 above, except that steps S110 to S118 are the application process of the image fusion model, while the second stage is the training process of the image fusion model. Therefore, it will not be described in detail here.

[0158] When training the image fusion model to be trained based on the first fused image, the second fused image, the first annotation, and the second annotation, the server can train the image fusion model to be trained with the minimum difference between the first fused image and the second annotation and the minimum difference between the second fused image and the first annotation as training objectives.

[0159] In this specification, to achieve more fine-grained fusion of frequency domain features, the frequency domain features can be decomposed into finer-grained high-frequency and low-frequency components, and then fusion of frequency domain features is performed at this fine-grained dimension of high-frequency and low-frequency components. Therefore, when training the image fusion model to be trained based on the first fused image, the second fused image, the first annotation, and the second annotation, the server can also determine the high-frequency and low-frequency components corresponding to the first annotation, and the high-frequency and low-frequency components corresponding to the second annotation. Then, based on the first high-frequency component and the high-frequency component corresponding to the second annotation, a first high-frequency loss is determined. Based on the first low-frequency component and the low-frequency component corresponding to the second annotation, a first low-frequency loss is determined. Based on the first fused image and the second annotation, a first image loss is determined. The first high-frequency loss, the first low-frequency loss, and the first image loss are used as the first loss. Simultaneously, based on the second high-frequency component and the high-frequency component corresponding to the first annotation, a second high-frequency loss is determined. Based on the second low-frequency component and the low-frequency component corresponding to the first annotation, a second low-frequency loss is determined. Based on the second fused image and the first annotation, a second image loss is determined. The second high-frequency loss, the second low-frequency loss, and the second image loss are used as the second loss. Then, the image fusion model to be trained is trained based on the first loss and the second loss.

[0160] The method for determining the high-frequency and low-frequency components corresponding to the first or second annotation is similar to the method for determining the high-frequency and low-frequency components in step S108 above, and will not be repeated here. The first high-frequency component can be obtained by fusing the high-frequency components corresponding to the first image and the high-frequency components corresponding to the second sample, or by fusing the high-frequency components corresponding to the first image, the high-frequency components corresponding to the second sample, and the high-frequency components corresponding to the first spatial feature. Correspondingly, the first low-frequency component, the second high-frequency component, and the second low-frequency component are also obtained by fusing the low-frequency or high-frequency components corresponding to the relevant images, and will not be repeated here.

[0161] The first loss in the first stage mentioned above can be calculated using the following formula:

[0162]

[0163]

[0164]

[0165] in, Fuse represents the first image loss. img This represents the first fused image. Indicates the second annotation. This represents the loss in the first frequency domain. This indicates the high-frequency component corresponding to the second annotation. This indicates the low-frequency component corresponding to the second annotation. Indicates the first high-frequency loss. This indicates the first low-frequency loss. This indicates the first loss.

[0166] The second loss in the second stage mentioned above can be calculated using the following formula:

[0167]

[0168]

[0169]

[0170] in, Indicates the loss of the second image. Represents the second fused image, GT h Indicates the first annotation. Indicates the loss in the first frequency domain. This indicates the high-frequency component corresponding to the first annotation. This indicates the low-frequency component corresponding to the first annotation. Indicates the second high-frequency loss. This indicates the second low-frequency loss. This indicates the second loss.

[0171] The above describes one or more implementations of the methods described in this specification. Based on the same idea, this specification also provides corresponding image fusion apparatus, such as... Figure 7 As shown.

[0172] Figure 7 This is a schematic diagram of an image fusion apparatus provided in this specification, including:

[0173] Acquisition module 200 is used to acquire panchromatic images and low-resolution multispectral images of the target area;

[0174] The first processing module 202 is used to adjust the panchromatic image to a first image with the same resolution as the low-resolution multispectral image;

[0175] The first spatial feature extraction module 204 is used to input the first image into the feature extraction subnet of a pre-trained image fusion model to determine panchromatic features, and to input the low-resolution multispectral image into the feature extraction subnet to determine spectral features;

[0176] The first frequency domain feature extraction module 206 is used to determine the frequency domain features corresponding to the panchromatic features and to determine the frequency domain features corresponding to the spectral features;

[0177] The first feature fusion module 208 is used to input the frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features into the feature fusion subnet of the image fusion model to determine the first fusion feature, and to determine the first fusion image based on the first fusion feature;

[0178] The second processing module 210 is used to adjust the first fused image into a second image with the same resolution as the panchromatic image;

[0179] The second spatial feature extraction module 212 is used to input the panchromatic image into the feature extraction sub-network to determine the first feature, and to input the second image into the feature extraction sub-network to determine the second feature;

[0180] The second frequency domain feature extraction module 214 is used to determine the frequency domain feature corresponding to the first feature and to determine the frequency domain feature corresponding to the second feature.

[0181] The second feature fusion module 216 is used to input the frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature into the feature fusion subnet to determine the second fused feature.

[0182] The determining module 218 is used to determine the target fused image of the target region based on the second fusion feature.

[0183] Optionally, the first feature fusion module 208 is specifically used to fuse the panchromatic features and the spectral features to determine a first spatial feature; determine the frequency domain features corresponding to the first spatial feature; and input the first spatial feature, the frequency domain features corresponding to the panchromatic features, the frequency domain features corresponding to the spectral features, and the frequency domain features corresponding to the first spatial feature into the feature fusion subnet of the image fusion model to determine a first fused feature.

[0184] Optionally, the feature fusion subnet includes a first fusion layer and a second fusion layer;

[0185] The first feature fusion module 208 is specifically used to input the frequency domain features corresponding to the panchromatic features, the frequency domain features corresponding to the spectral features, and the frequency domain features corresponding to the first spatial features into the first fusion layer of the feature fusion subnet of the image fusion model to determine the first frequency domain features; and to input the first frequency domain features and the first spatial features into the second fusion layer of the feature fusion subnet of the image fusion model to determine the first fusion features.

[0186] Optionally, the first feature fusion module 208 is specifically configured to: determine that the high-frequency component corresponding to the panchromatic feature is a panchromatic high-frequency component, and determine that the low-frequency component corresponding to the panchromatic feature is a panchromatic low-frequency component; determine that the high-frequency component corresponding to the spectral feature is a spectral high-frequency component, and determine that the low-frequency component corresponding to the spectral feature is a spectral low-frequency component; determine that the high-frequency component corresponding to the first spatial feature is a first spatial high-frequency component, and determine that the low-frequency component corresponding to the first spatial feature is a first spatial low-frequency component; input the panchromatic high-frequency component, the panchromatic low-frequency component, the spectral high-frequency component, the spectral low-frequency component, the first spatial high-frequency component, and the first spatial low-frequency component into the first fusion layer of the feature fusion subnet of the image fusion model, so that the first fusion layer fuses the panchromatic high-frequency component, the spectral high-frequency component, and the first spatial high-frequency component to determine the first high-frequency component, and fuses the panchromatic low-frequency component, the spectral low-frequency component, and the first spatial low-frequency component to determine the first low-frequency component; and determine the first frequency domain feature based on the first high-frequency component and the first low-frequency component.

[0187] Optionally, the image fusion model further includes a cross-attention fusion layer;

[0188] The first feature fusion module 208 is specifically used to input the panchromatic features and the spectral features into the cross-attention fusion layer of the image fusion model to determine the first spatial features.

[0189] Optionally, the device further includes:

[0190] Training module 220 is used to determine a panchromatic image of a specified area collected historically as a first sample, and to determine a low-resolution multispectral image corresponding to the first sample as a second sample; adjust the first sample to a first image with the same resolution as the second sample; input the first image into the feature extraction subnet of the image fusion model to be trained to determine panchromatic features, and input the second sample into the feature extraction subnet to determine spectral features; determine the frequency domain features corresponding to the panchromatic features, and determine the frequency domain features corresponding to the spectral features; input the frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features into the feature fusion subnet of the image fusion model to be trained to determine a first fusion feature, and determine a first fusion image based on the first fusion feature; adjust the first fusion image to a first image with the same resolution as the second sample. The first image is given a second image with the same resolution as the first sample; the first sample is input into the feature extraction subnetwork to determine a first feature, and the second image is input into the feature extraction subnetwork to determine a second feature; the frequency domain feature corresponding to the first feature and the frequency domain feature corresponding to the second feature are determined; the frequency domain feature corresponding to the first feature and the frequency domain feature corresponding to the second feature are input into the feature fusion subnetwork to determine a second fusion feature, and a second fusion image is determined based on the second fusion feature; a reference image of the specified region is determined as a first annotation, and the reference image is adjusted to have the same resolution as the second sample and used as a second annotation; the image fusion model to be trained is trained based on the first fusion image, the second fusion image, the first annotation, and the second annotation.

[0191] Optionally, the training module 220 is specifically used to train the image fusion model to be trained with the goal of minimizing the difference between the first fused image and the second annotation and minimizing the difference between the second fused image and the first annotation.

[0192] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 This provides an image fusion method.

[0193] This instruction manual also provides Figure 8 One of the corresponding Figure 1 A schematic diagram of the structure of an electronic device. (e.g.) Figure 8As shown, at the hardware level, this electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above. Figure 1 The image fusion method described above.

[0194] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0195] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0196] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0197] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0198] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware.

[0199] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0200] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0201] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0202] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0203] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0204] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0205] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0206] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0207] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0208] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0209] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0210] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.

Claims

1. A method for image fusion, characterized in that, include: Acquire panchromatic and low-resolution multispectral images of the target area; The panchromatic image is adjusted to a first image with the same resolution as the low-resolution multispectral image; The first image is input into the feature extraction subnet of a pre-trained image fusion model to determine panchromatic features, and the low-resolution multispectral image is input into the feature extraction subnet to determine spectral features; Determine the frequency domain features corresponding to the panchromatic features, and determine the frequency domain features corresponding to the spectral features; The frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features are input into the feature fusion subnet of the image fusion model to determine the first fusion feature, and the first fusion image is determined based on the first fusion feature. The first fused image is adjusted to a second image with the same resolution as the panchromatic image; The panchromatic image is input into the feature extraction sub-network to determine the first feature, and the second image is input into the feature extraction sub-network to determine the second feature; Determine the frequency domain feature corresponding to the first feature, and determine the frequency domain feature corresponding to the second feature; The frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature are input into the feature fusion subnet to determine the second fusion feature; Based on the second fusion feature, the target fused image of the target region is determined.

2. The method as described in claim 1, characterized in that, The frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features are input into the feature fusion subnetwork of the image fusion model to determine the first fusion feature, which specifically includes: The panchromatic features and the spectral features are fused to determine the first spatial feature; Determine the frequency domain features corresponding to the first spatial domain features; The first spatial feature, the frequency domain feature corresponding to the panchromatic feature, the frequency domain feature corresponding to the spectral feature, and the frequency domain feature corresponding to the first spatial feature are input into the feature fusion subnet of the image fusion model to determine the first fusion feature.

3. The method as described in claim 2, characterized in that, The feature fusion subnet includes a first fusion layer and a second fusion layer; The first spatial feature, the frequency domain feature corresponding to the panchromatic feature, the frequency domain feature corresponding to the spectral feature, and the frequency domain feature corresponding to the first spatial feature are input into the feature fusion subnet of the image fusion model to determine the first fusion feature, specifically including: The frequency domain features corresponding to the panchromatic features, the frequency domain features corresponding to the spectral features, and the frequency domain features corresponding to the first spatial features are input into the first fusion layer of the feature fusion subnet of the image fusion model to determine the first frequency domain features. The first frequency domain feature and the first spatial domain feature are input into the second fusion layer of the feature fusion subnet of the image fusion model to determine the first fusion feature.

4. The method as described in claim 3, characterized in that, The frequency domain features corresponding to the panchromatic features, the frequency domain features corresponding to the spectral features, and the frequency domain features corresponding to the first spatial features are input into the first fusion layer of the feature fusion subnet of the image fusion model to determine the first frequency domain features, specifically including: The high-frequency component corresponding to the panchromatic feature is determined to be a panchromatic high-frequency component, and the low-frequency component corresponding to the panchromatic feature is determined to be a panchromatic low-frequency component; the high-frequency component corresponding to the spectral feature is determined to be a spectral high-frequency component, and the low-frequency component corresponding to the spectral feature is determined to be a spectral low-frequency component; the high-frequency component corresponding to the first spatial feature is determined to be a first spatial high-frequency component, and the low-frequency component corresponding to the first spatial feature is determined to be a first spatial low-frequency component. The panchromatic high-frequency component, the panchromatic low-frequency component, the spectral high-frequency component, the spectral low-frequency component, the first spatial high-frequency component, and the first spatial low-frequency component are input into the first fusion layer of the feature fusion subnet of the image fusion model. The first fusion layer fuses the panchromatic high-frequency component, the spectral high-frequency component, and the first spatial high-frequency component to determine the first high-frequency component, and fuses the panchromatic low-frequency component, the spectral low-frequency component, and the first spatial low-frequency component to determine the first low-frequency component. The first frequency domain feature is determined based on the first high-frequency component and the first low-frequency component.

5. The method as described in claim 2, characterized in that, The image fusion model also includes a cross-attention fusion layer; The panchromatic features and the spectral features are fused to determine the first spatial feature, specifically including: The panchromatic features and the spectral features are input into the cross-attention fusion layer of the image fusion model to determine the first spatial features.

6. The method as described in claim 1, characterized in that, Pre-trained image fusion models include: The panchromatic image of a specified area collected in history is identified as the first sample, and the low-resolution multispectral image corresponding to the first sample is identified as the second sample. Adjust the first sample to a first image with the same resolution as the second sample; The first image is input into the feature extraction subnet of the image fusion model to be trained to determine the panchromatic features, and the second sample is input into the feature extraction subnet to determine the spectral features; Determine the frequency domain features corresponding to the panchromatic features, and determine the frequency domain features corresponding to the spectral features; The frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features are input into the feature fusion subnetwork of the image fusion model to be trained, the first fusion feature is determined, and the first fusion image is determined based on the first fusion feature. The first fused image is adjusted to a second image with the same resolution as the first sample; The first sample is input into the feature extraction subnetwork to determine the first feature, and the second image is input into the feature extraction subnetwork to determine the second feature; Determine the frequency domain feature corresponding to the first feature, and determine the frequency domain feature corresponding to the second feature; The frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature are input into the feature fusion subnet to determine the second fusion feature, and the second fusion image is determined based on the second fusion feature. The reference image of the specified region is determined as the first annotation, and the reference image is adjusted to the same resolution as the second sample and used as the second annotation; The image fusion model to be trained is trained based on the first fused image, the second fused image, the first annotation, and the second annotation.

7. The method as described in claim 6, characterized in that, The image fusion model to be trained is trained based on the first fused image, the second fused image, the first annotation, and the second annotation, specifically including: The image fusion model to be trained is trained with the goal of minimizing the difference between the first fused image and the second annotation, and minimizing the difference between the second fused image and the first annotation.

8. An image fusion apparatus, characterized in that, include: The acquisition module is used to acquire panchromatic images and low-resolution multispectral images of the target area; A first processing module is used to adjust the panchromatic image to a first image with the same resolution as the low-resolution multispectral image; The first spatial feature extraction module is used to input the first image into the feature extraction subnet of a pre-trained image fusion model to determine panchromatic features, and to input the low-resolution multispectral image into the feature extraction subnet to determine spectral features; The first frequency domain feature extraction module is used to determine the frequency domain features corresponding to the panchromatic features and to determine the frequency domain features corresponding to the spectral features. The first feature fusion module is used to input the frequency domain features corresponding to the panchromatic features and the frequency domain features corresponding to the spectral features into the feature fusion subnet of the image fusion model to determine the first fusion feature, and to determine the first fusion image based on the first fusion feature; The second processing module is used to adjust the first fused image into a second image with the same resolution as the panchromatic image; The second spatial feature extraction module is used to input the panchromatic image into the feature extraction sub-network to determine the first feature, and to input the second image into the feature extraction sub-network to determine the second feature; The second frequency domain feature extraction module is used to determine the frequency domain feature corresponding to the first feature and to determine the frequency domain feature corresponding to the second feature. The second feature fusion module is used to input the frequency domain features corresponding to the first feature and the frequency domain features corresponding to the second feature into the feature fusion subnet to determine the second fused feature. The determining module is used to determine the target fused image of the target region based on the second fusion feature.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 7.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Remote sensing image fusion method based on multi-redundancy dictionary and sparse reconstruction

    CN104794681A

  • Double-branch panchromatic sharpening method based on attention mechanism

    CN114511470A