Multispectral and panchromatic image fusion method, device and equipment based on semantic information
By introducing multi-scale interaction and feature extraction of semantic information in the fusion of multi-spectral and full-color images, the problem that the difference in land objects types in the prior art is not considered is solved, and the high-quality image fusion effect is achieved, and the effectiveness of the model in practical applications is improved.
Patent Information
- Application Number
- CN202510151050.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-07-08
AI Technical Summary
The existing multi-spectral and full-color image fusion method does not consider the differences in spatial and spectral weights between different landform types, and focuses on advanced features of images at the semantic level, which makes it difficult to generalize the effects of the algorithm in cross-landform applications, and in practice the model is difficult to use effectively.
By obtaining the image dataset and target text description, multi-spectral images, full-color images and semantic features are extracted, multi-scale feature extraction is used for image fusion network, and semantic features are interacted with semantic features and the features to be fused at each scale. Encoder-decoder structure and TCB module are used for feature extraction and integration, combined with Transformer and CNN branches, MLP is used to generate adjustment vectors for semantic interaction, add jump connections and residual connections, and train the network using the average absolute value error loss function.
High-quality multi-spectral image fusion is achieved, which improves the generalization ability in cross-terrestrial applications, outputs high-resolution and rich spectral information images, maintains the characteristics of the land and significantly improves spatial details.
Smart Images

Figure CN120279362A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of computer vision and deep learning, and particularly relates to a multi-spectral and panchromatic image fusion method, apparatus, and device based on semantic information. Background Art
[0002] With the rapid development of remote sensing technology, the demand for high-resolution multi-spectral images is increasing day by day. High-quality multi-spectral images can be widely applied to downstream tasks such as remote sensing land cover change detection, semantic segmentation, and remote sensing interpretation. However, restricted by sensors, existing sensors often can only obtain low-spatial-resolution multi-spectral images and high-spatial-resolution panchromatic images separately. Therefore, multi-spectral and panchromatic image fusion hopes to inject the spatial details of the panchromatic image into the multi-spectral image to obtain a multi-spectral image with both high spatial resolution and high spectral resolution. High-resolution multi-spectral images play an important role in fields such as construction and agriculture.
[0003] Existing multi-spectral and panchromatic image fusion algorithms can be divided into two categories: traditional algorithms and deep learning algorithms. Traditional image fusion methods can be divided into methods based on component replacement, methods based on multi-resolution analysis, and variational optimization methods. Among them, the component replacement methods use the panchromatic image to replace the spatial components of the multi-spectral image. This type of method mainly includes methods based on IHS (intensity-hue-saturation), methods based on PCA (principal component analysis), methods based on GS (Gram-Schmidt), and methods based on brovey transform, etc. This type of method has low computational complexity but is prone to spectral distortion. The multi-resolution analysis methods decompose the source images into LFC (Low-Frequency Component) and HFC (High-Frequency Component) at different scales through various transforms and then fuse them. This type of method can better ensure the consistency of spectral distribution but is prone to spatial distortion. The variational optimization methods construct an energy function by observing the image fusion task or using the sparse representation method, and minimize the energy function through an iterative solution method to obtain the fusion result. This type of method introduces too many artificial priors that may not be accurate and requires a large amount of computational resources.
[0004] In recent years, with the development of deep learning, many deep learning-based image fusion methods have been widely studied and mainly fall into two categories. The first category simulates data through the Wald protocol to generate training data pairs, and uses network architectures such as CNN (Convolutional Neural Network) and Transformer for feature extraction and fusion of the data to be fused. The network training is guided by minimizing the gap between the fusion result and the ground truth. For example, image fusion is achieved by designing a deep residual network architecture, and the aggregation of high-frequency information is realized by designing a Transformer network architecture based on window attention, achieving good fusion results. This type of method usually has good effects on simulation datasets, but due to the gap between simulation data and real data, as well as the resolution and degradation process between different satellites, the above algorithms lack the generalization ability in real scenarios.
[0005] Therefore, blind multi-spectral and panchromatic image fusion research for unknown degradation and unknown satellites has emerged. Variational inference is used to model the physical degradation of multi-spectral and panchromatic image fusion, realizing high-quality fusion under unknown degradation conditions. This type of method improves the generalization ability in real scenarios.
[0006] However, the existing multi-spectral and panchromatic image fusion methods still have the following problems: (1) The above methods are still pixel-to-pixel image enhancement, without considering the differences in spatial and spectral weights between different land cover types and paying attention to the high-level features of the image at the semantic level, resulting in difficult generalization of the algorithm in cross-land cover applications. (2) The fixed model parameters lack the interaction between users and the fusion model, making it difficult to effectively use the model in practice. (3) The semantic information based on multi-modal features for the gain of image fusion has not been explored yet. Summary of the Invention
[0007] This application provides a multi-spectral and panchromatic image fusion method, device, and equipment based on semantic information to solve the problems in the prior art that do not consider the differences in spatial and spectral weights between different land cover types and pay attention to the high-level features of the image at the semantic level, resulting in difficult generalization of the algorithm in cross-land cover applications and difficult effective use of the model in practice.
[0008] The first aspect of the present application provides a multi - spectral and pan - chromatic image fusion method based on semantic information, including the following steps: obtaining an image dataset and the target text description of the image dataset; extracting the semantic features of the multi - spectral image, the pan - chromatic image and the target text description in the image dataset; inputting the multi - spectral image, the pan - chromatic image and the semantic features into an image fusion network, and the image fusion network outputs an image fusion result, where the image fusion network performs multi - scale image feature extraction on the multi - spectral image and the pan - chromatic image, and at each scale, interacts the semantic features with the image features to be fused, and uses semantics to guide image fusion.
[0009] Optionally, extracting the semantic features of the target text description of the image dataset includes: constructing a text description dictionary for different ground objects; using the CLIP image - text large model to construct a semantic feature extraction model, and driving the semantic feature extraction model based on the text description dictionary; identifying the ground object type in the target text description, inputting the ground object type into the semantic feature extraction model, and the semantic feature extraction model outputs the semantic features of the target text description.
[0010] Optionally, the image fusion network includes an encoder and a decoder with multiple layers of down - sampling and up - sampling operations, uses the encoder to extract image features at multiple scales, and uses the decoder to perform feature decoding and result reconstruction, where the encoder and the decoder are composed of TCB modules, and in the encoder stage, the MTI module is used to interact the semantic features with the image features to be fused, and local and global features are extracted and integrated through the TCB module.
[0011] Optionally, the MTI module extracts a weight vector and a bias vector through a multi - layer perceptron, and realizes the interaction between the semantic features and the image features to be fused according to the weight vector and the bias vector.
[0012] Optionally, the interaction formula between the semantic features and the image features to be fused is:
[0013]
[0014] where, the image feature vector after semantic interaction; is the input image feature vector of the j - th layer of the encoder; α is the weight vector; β is the bias vector.
[0015] Optionally, a skip connection is added between the encoder and the decoder of the same layer in the image fusion network, and at the input and output ends of the image fusion network, a residual connection is added to the input of the multi - spectral image.
[0016] Optionally, the TCB module includes a Transformer branch and a CNN branch, the Transformer branch extracts global features, and the CNN branch extracts local features.
[0017] Optionally, before inputting the multi-spectral image, the panchromatic image, and the semantic feature input image into the image fusion network, it further includes: obtaining a training data set and a mean absolute error loss function; training the image fusion network based on the training data set and the mean absolute error loss function, where the mean absolute error loss function is:
[0018]
[0019] where, I f and I gt are respectively the fused high-resolution multi-spectral image and the multi-spectral image ground truth.
[0020] The second aspect of the present application provides a multi-spectral and panchromatic image fusion device based on semantic information, including: an acquisition module for acquiring an image data set and a target text description of the image data set; an extraction module for extracting semantic features of the multi-spectral image, the panchromatic image, and the target text description in the image data set; a fusion module for inputting the multi-spectral image, the panchromatic image, and the semantic features into an image fusion network, and the image fusion network outputs an image fusion result, where the image fusion network performs multi-scale image feature extraction on the multi-spectral image and the panchromatic image, and at each scale, the semantic features are interacted with the image features to be fused, and semantic guidance is used for image fusion.
[0021] The third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the multi-spectral and panchromatic image fusion method based on semantic information in the first aspect.
[0022] Therefore, the present application has the following beneficial effects:
[0023] In the embodiment of the present application, by acquiring an image data set and a target text description of the image data set, and extracting semantic features of the multi-spectral image, the panchromatic image, and the target text description in the image data set, and inputting the three into an image fusion network, the image fusion network performs multi-scale image feature extraction on the multi-spectral image and the panchromatic image, and at each scale, the semantic features are interacted with the image features to be fused, and semantic guidance is used for image fusion to obtain an image fusion result, introducing a new perspective to multi-spectral image fusion, realizing image fusion operations under multi-modal interaction through the guidance of semantic information, and being able to focus on the main ground objects in the image to output high-quality multi-spectral images. Thus, it solves the problems that the prior art does not consider the differences in spatial and spectral weights between different ground object types and focuses on the high-level features of the image at the semantic level, resulting in difficult generalization of the algorithm in cross-ground object applications and ineffective use of the model in practice.
[0024] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Brief Description of the Drawings
[0025] The above-mentioned and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of embodiments in conjunction with the drawings, where:
[0026] Figure 1 is a flowchart of a multi-spectral and panchromatic image fusion method based on semantic information according to an embodiment of the present application;
[0027] Figure 2 is a structural diagram of a semantic-guided multi-spectral and panchromatic image fusion network according to an embodiment of the present application;
[0028] Figure 3 is a test input image of a simulation and real dataset according to an embodiment of the present application;
[0029] Figure 4 is an image fusion result diagram of a simulation dataset according to an embodiment of the present application;
[0030] Figure 5 is a comparison diagram of the image fusion results of a simulation and real dataset and the results of existing advanced methods according to an embodiment of the present application;
[0031] Figure 6 is an example diagram of a multi-spectral and panchromatic image fusion device based on semantic information according to an embodiment of the present application;
[0032] Figure 7 is a schematic structural diagram of an electronic device according to an embodiment of the present application. Detailed Description of the Embodiments
[0033] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.
[0034] The following describes a multi - spectral and panchromatic image fusion method, device, and equipment based on semantic information according to the embodiments of the present application. Aiming at the problems in the prior art mentioned in the above - mentioned background technology that the existing technology does not consider the differences in spatial and spectral weights between different land - cover types and focuses on the high - level features of images at the semantic level, resulting in difficult generalization of the algorithm in cross - land - cover applications and ineffective use of the model in practice, etc., the present application provides a multi - spectral and panchromatic image fusion method based on semantic information. In this method, by obtaining an image dataset and the target text description of the image dataset, and extracting the semantic features of the multi - spectral image, panchromatic image, and target text description in the image dataset, and inputting the three into an image fusion network, the image fusion network performs multi - scale image feature extraction on the multi - spectral image and the panchromatic image, and at each scale, the semantic features are interacted with the image features to be fused, and semantic guidance is used for image fusion to obtain an image fusion result. The introduction of multi - spectral image fusion brings a new perspective, and through the guidance of semantic information, image fusion operations under multi - modal interaction can be realized, and the attention to the main land - cover in the image can be achieved to output a high - quality multi - spectral image. Thus, the problems in the prior art that do not consider the differences in spatial and spectral weights between different land - cover types and focus on the high - level features of images at the semantic level, resulting in difficult generalization of the algorithm in cross - land - cover applications and ineffective use of the model in practice, etc., are solved.
[0035] Specifically, Figure 1 FIG. is a schematic flowchart of a multi - spectral and panchromatic image fusion method based on semantic information provided by an embodiment of the present application.
[0036] As Figure 1 shown, the multi - spectral and panchromatic image fusion method based on semantic information includes the following steps:
[0037] In step S101, an image dataset and the target text description of the image dataset are obtained.
[0038] It can be understood that in the embodiments of the present application, an image dataset and the target text description of the image dataset can be obtained according to the input image.
[0039] In step S102, the semantic features of the multi - spectral image, panchromatic image, and target text description in the image dataset are extracted.
[0040] Among them, the multi - spectral image refers to an image obtained in multiple electromagnetic wave bands (such as visible light, near - infrared, etc.); the panchromatic image is a high - resolution grayscale image that covers the entire visible spectral range but does not distinguish colors.
[0041] It can be understood that in the embodiments of the present application, the multi - spectral image and panchromatic image in the image dataset can be identified and extracted, and the semantic features of the target text description of the image dataset can be extracted.
[0042] In the embodiment of the present application, extracting the semantic features of the target text description of the image dataset includes: constructing a text description dictionary for different ground objects; using the CLIP large text-image model to construct a semantic feature extraction model, and driving the semantic feature extraction model based on the text description dictionary; identifying the ground object type in the target text description, inputting the ground object type into the semantic feature extraction model, and the semantic feature extraction model outputs the semantic features of the target text description.
[0043] Among them, the ground object type refers to different types of ground coverings or objects, such as forests, buildings, water bodies, etc.; the CLIP large text-image model is a multi-modal deep learning model that can jointly understand text and images and establish the connection between the two by learning a large number of text-image pairs.
[0044] It can be understood that in the embodiment of the present application, a text description dictionary containing detailed text descriptions of various ground objects is first created, and the CLIP large text-image model is used to construct a semantic feature extraction model, which can be driven by the constructed text description dictionary. When there is a target text description, the ground object type in the description can be identified, and the ground object type is input into the semantic feature extraction model, and the semantic feature extraction model outputs the semantic features of the target text description.
[0045] In step S103, the multi-spectral image, the panchromatic image, and the semantic features are input into the image fusion network, and the image fusion network outputs an image fusion result. Among them, the image fusion network performs multi-scale image feature extraction on the multi-spectral image and the panchromatic image, and at each scale, the semantic features are interacted with the image features to be fused, and semantic-guided image fusion is utilized.
[0046] Among them, multi-scale image feature extraction refers to the technology of analyzing image features at different resolutions or scales; an adjustment vector is generated using an MLP (Multilayer Perceptron), and the adjustment vector is used to interact the semantic features with the image features to be fused at each scale.
[0047] It can be understood that in the embodiment of the present application, the multi-spectral image, the panchromatic image, and the corresponding semantic features are input into the image fusion network together. At each scale, the network uses an MLP module to generate an adjustment vector, and then these vectors are applied to the image features to be fused to achieve the interaction between semantic information and image features, and semantic-guided image fusion is utilized. After such semantic-guided processing, the fusion network can output a multi-spectral image with high resolution and rich spectral information, which can significantly improve the spatial details while maintaining the original ground object characteristics.
[0048] In the embodiments of the present application, the image fusion network includes an encoder and a decoder with multiple layers of downsampling and upsampling operations. The encoder is used to extract image features at multiple scales, and the decoder is used for feature decoding and result reconstruction. Among them, the encoder and the decoder are composed of TCB modules. In the encoder stage, the MTI module is used for the interaction between semantic features and the image features to be fused, and local features and global features are extracted and integrated through the TCB module.
[0049] Among them, the MTI module is a multi-scale text interaction module; the TCB module is a module that combines a transformer and a convolutional neural network, which will be described in detail below and will not be elaborated here.
[0050] It can be understood that the embodiments of the present application adopt an encoder-decoder structure. In the encoder stage, the MTI module is used for the interaction between semantic features and the image features to be fused, effectively combining semantic guidance and image features at different scales. Then, the TCB module is used for local feature and global feature extraction and integration, and then the decoder is used for feature decoding and result reconstruction.
[0051] In the embodiments of the present application, the MTI module extracts a weight vector and a bias vector through a multi-layer perceptron, and realizes the interaction between semantic features and the image features to be fused according to the weight vector and the bias vector.
[0052] Among them, the generation process of the weight vector and the bias vector is α,β = MLP(f se ), f se is the semantic feature.
[0053] It can be understood that the embodiments of the present application can use the MLP in the MTI module to extract the weight vector and the bias vector through the semantic feature, and realize the interaction between the semantic feature and the image features to be fused according to the weight vector and the bias vector.
[0054] In the embodiments of the present application, the interaction formula between semantic features and the image features to be fused is:
[0055]
[0056] Among them, the image feature vector after semantic interaction; is the input image feature vector of the j-th layer of the encoder; α is the weight vector; β is the bias vector.
[0057] In the embodiments of the present application, a skip connection is added between the encoder and the decoder of the same layer of the image fusion network, and at the input and output ends of the image fusion network, a residual connection is added to the input of the multi-spectral image.
[0058] Among them, a skip connection means directly passing information from one layer of the network to another later layer without passing through each intermediate layer; a residual connection is a special form of skip connection that allows the network to learn the residual mapping between the input and output instead of directly fitting the mapping from the input to the output, which improves the convergence speed of the network.
[0059] It can be understood that in the embodiment of the present application, skip connections are added between the encoder and decoder at the same level, which can directly transfer the features extracted by a certain layer of the encoder to the corresponding layer of the decoder, ensuring the full utilization of features; a residual connection for the multispectral image is also introduced between the input end and the output end of the network, which improves the convergence speed of the network.
[0060] In the embodiment of the present application, the TCB module includes a Transformer branch and a CNN branch. The Transformer branch extracts global features, and the CNN branch extracts local features.
[0061] Among them, the Transformer branch is the transformer branch. Based on the transformer architecture, it can capture long-range dependencies and global context information and can effectively extract global features; the CNN branch is the convolutional neural network branch. Based on the traditional convolutional neural network architecture, it can extract local features.
[0062] It can be understood that the TCB module in the embodiment of the present application is a hybrid module that combines the advantages of the transformer and the convolutional neural network. It is internally divided into a Transformer branch and a CNN branch. Through the transformer architecture of the Transformer branch, global features can be effectively extracted; through the convolutional neural network architecture of the CNN branch, local features can be extracted.
[0063] In the embodiment of the present application, before inputting the multispectral image, the panchromatic image, and the semantic features into the image fusion network, it further includes: obtaining a training data set and a mean absolute error loss function; training the image fusion network based on the training data set and the mean absolute error loss function, where the mean absolute error loss function is:
[0064]
[0065] where, I f and I gt are respectively the fused high-resolution multispectral image and the multispectral image ground truth.
[0066] It can be understood that in the embodiments of the present application, a loss function for calculating the mean absolute error is used with the fused high-resolution multispectral image and the ground truth of the multispectral image. The loss function is used as the L1 norm, and the semantic-guided image fusion network is trained in combination with the loss function to constrain the difference between the fusion result and the true result.
[0067] According to the multispectral and panchromatic image fusion method based on semantic information proposed in the embodiments of the present application, by obtaining an image data set and the target text description of the image data set, and extracting the semantic features of the multispectral image, the panchromatic image and the target text description in the image data set, and inputting the three into the image fusion network. The multispectral image and the panchromatic image are subjected to multi-scale image feature extraction through the image fusion network, and the semantic features are interacted with the image features to be fused at each scale, and semantic-guided image fusion is used to obtain the image fusion result. The multispectral image fusion is introduced from a new perspective, and the image fusion operation under multi-modal interaction is realized through the guidance of semantic information, and the attention to the main ground objects in the image can be realized to output a high-quality multispectral image.
[0068] The multispectral and panchromatic image fusion method based on semantic information will be further described below through a specific embodiment. The embodiments of the present application propose a semantic-guided interactive multispectral and panchromatic image fusion network SEPan. As Figure 2 shown, first in terms of text feature introduction, a ground object type text library is constructed to provide text information according to the input image, and the CLIP model is used to extract semantic features. A multi-scale text feature injection paradigm based on the CLIP model is proposed, which can inject text information at different feature scales during the image feature extraction stage to achieve interactive feature attention. In terms of network design, in order to give full play to the guiding role of semantic information on the image fusion result, an encoder-decoder architecture with downsampling and upsampling operations is adopted, and semantic information guidance is added at different scales of the encoder. The Transformer-CNN module (TCB) is used as the feature extraction module of the encoder and the decoder to improve the feature extraction and representation ability of the network. Finally, a loss function suitable for the network is designed to guide network training, and the training is carried out on the GaoFen-2 simulation data set. Finally, a semantic information-guided multispectral and panchromatic image fusion model is obtained for image fusion. The method includes the following steps:
[0069] Step 1: Use the CLIP large image-text model to construct a semantic feature extraction module for generating corresponding semantic features according to the text description of the input image, which is used to adjust the fusion result. The above semantic feature extraction module is driven by a self-built text information library, and the above text information library contains four common ground object categories in remote sensing data and descriptions of the four ground object categories.
[0070] Assume the input image dataset is as follows:
[0071] D = {(I gt , I p , I lms ) i | i = 1, …, N}
[0072] Where: N represents the number of image pairs in the dataset. In this embodiment, N is 2700. I gt , I p , I lms respectively represent the high - resolution multispectral image, the high - resolution panchromatic image, and the low - resolution multispectral image.
[0073] In the classic multispectral and panchromatic image fusion task, it is hoped to fuse the panchromatic image (I p ) and the low - resolution multispectral image (I lms ) to finally obtain a multispectral (I f ) image with both high spatial resolution and spectral resolution. In the design of the network (F ps ), it usually focuses on how to extract and integrate the spatial and spectral information of the panchromatic image and the multispectral image to obtain the fusion result. Briefly, it can be shown as the formula:
[0074] I f = F ps (I p , I lms )
[0075] In this process, the network only focuses on how to obtain a relatively good overall fusion result, without considering the uniqueness of the spatial and spectral weights of different ground objects in the image, and the network parameters are fixed. Therefore, once the spatial - spectral weight differences in different regions of the image are large, the performance of the network is often poor. Therefore, it is hoped to use the semantic - guided method. By constructing a text description dictionary (T dic ) for different ground objects, integrate the ground object types with the image features to be fused in the semantic modality, and finally obtain a fusion result with ground object attention characteristics. This fusion process is expressed as follows:
[0076] I f = F ps-text (I p , I lms , T dic )
[0077] Where T dicIt is a text description built for four typical remote sensing ground object types. For example, when the ground object type is a building, the indexed semantic information is "This is a panchromatic sharpening task, and we are processing images of buildings". In addition to looking up in the dictionary, this description can also be designed by users themselves to guide the addition of semantics to the ground object categories of interest in the image. F ps-text This is the semantic-guided multi-spectral image fusion network SEPan proposed by this method.
[0078] Furthermore, after constructing the text information dictionary, it is necessary to perform feature encoding on the text information to obtain an accurate representation of the text information. Due to its contrastive learning training mechanism, the CLIP model has good effects on the construction of semantic features and the matching of text features and image features. Therefore, the text encoder (G()) of the CLIP model is selected to extract the semantic features of the given text. In actual operation, the pre-trained weights of the CLIP model are loaded and the gradients of the text encoder are frozen during training. Assume the given text is T dic , and the obtained semantic feature is f se , then the process of extracting text features is:
[0079] f se = G frozen (T dic )
[0080] After extracting the text features, it is necessary to make the semantic features interact with the text features in an appropriate way. By constructing a multi-layer perceptron (MLP), the feature vectors α and β used to adjust the image features are extracted. Among them, α is used as the weight vector and β is used as the bias vector, and their generation process is:
[0081] α,β = MLP(f se )
[0082] According to the obtained feature vectors α and β, the MTI module is proposed for the interaction between semantic features and image features. Specifically, the adjustment method of the image features to be fused is:
[0083]
[0084] where is the input image feature vector of the j-th layer of the encoder, is the image feature vector after semantic interaction.
[0085] Step 2: Design a multi-scale semantic information injection module (MTI) based on semantic features. This module uses a multi-layer perceptron (MLP) to generate adjustment vectors and uses the adjustment vectors to interact with the features of the images to be fused extracted by the encoder of the image fusion network SEPan at multiple scales.
[0086] The MTI module contains an MLP module to generate weight and bias vectors based on semantic information. The weighted and biased vectors semantically guide the input features through multiplication and addition at various scales of the network encoder.
[0087] Step 3: Construct the image fusion network SEPan. The SEPan network selects ResUNet as the basic architecture of the network and extracts image features at multiple scales through an encoder-decoder architecture with 3 layers of downsampling and upsampling operations. In the encoder stage, the features input to each layer first undergo semantic information interaction through the MTI module, and then local and global feature extraction and integration are performed through 2 cascaded TCB modules. To ensure the full utilization of features, skip connections are added between the encoder and decoder at the same layer. At the same time, to improve the spectral fidelity characteristics of the fusion result, residual connections are added to the input and output of the network for the low-resolution multispectral image input.
[0088] The input low-resolution multispectral image is first spatially upsampled to the size corresponding to the output. In this embodiment, the upsampling factor is 4 times, and it is concatenated with the high-resolution panchromatic image at the channel level. The concatenated image first undergoes a 1×1 convolution operation to obtain the shallow features of the image, and then the obtained shallow features are input into the image encoder. Each layer of the image encoder first performs semantic information interaction through the MTI module, and the interacted image features are used by 2 cascaded TCB modules to extract local and global features of the image to be fused. Subsequently, the features are reduced to a smaller dimension through downsampling operations and enter the next level of the encoder. At the input and output of each layer of the encoder, residual connections are added to improve the convergence speed of the network. After the input of the decoder of the network, through a 1×1 convolution operation and two TCB modules again, a 3×3 convolution operation is performed to obtain the final network output.
[0089] Furthermore, the above TCB module evenly divides the input features into two parts and inputs them into the Transformer branch and the CNN branch respectively to extract global and local features.
[0090] Step 4: Train the semantic-guided image fusion network in combination with the loss function, and use the trained network to achieve targeted image fusion for specific ground objects. The loss function adopted is the L1 norm, which is used to constrain the difference between the fusion result and the real result.
[0091] The mean absolute error loss function is:
[0092]
[0093] where: I f and I gtThey are the fused high-resolution multispectral image and the ground truth multispectral image respectively.
[0094] During network training, the Adam optimizer was used, and the parameters of the optimizer were fixed as β1 = 0.9 and β2 = 0.999. The initial learning rate was 3×10 -4 , and 64 groups of data were taken for training each time during training, and a total of 1500 rounds of training were performed. During training, the low-resolution multispectral images in the training set were cropped to a size of 128×128×4, and the panchromatic images were cropped to a size of 512×512×1.
[0095] This embodiment was trained on the GaoFen-2 simulation dataset, and the semantic-guided multispectral image fusion model obtained from the training was tested on the corresponding simulation and real test sets. The input picture examples of the two test sets are as Figure 3 shown. The final processing results on the two datasets are as Figure 4 shown.
[0096] Based on the image fusion results obtained from the above steps, in order to compare with other methods, ADKNet, Hyper-DSNet, MSDDN, and BimPan were selected as comparison methods to compare with the method, and the obtained results are as Figure 5 shown.
[0097] In order to quantitatively evaluate the results of image fusion, the peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), spectral angle similarity (SAM), average global error (ERGAS), and image overall quality evaluation index (Q) were introduced as evaluation indicators to measure the effect of image fusion. Among them, the larger the values of PSNR, SSIM, and Q, the better the fusion effect, and the smaller the values of the SAM and ERGAS indicators, the better the fusion result. 100 pairs of images were randomly selected as test data on the GaoFen-2 simulation dataset, and the test results are shown in Table 1. Among them, Table 1 is a comparison table of evaluation indicators for different methods.
[0098] Table 1
[0099] Method PSNR↑ SSIM↑ ERGAS↓ SAM↓ Q↑ ADKNet 39.454 0.962 0.517 0.045 0.984 Hyper-DSNet 35.661 0.923 0.755 0.052 0.959 MSDDN 41.482 0.976 0.412 0.036 0.990 BimPAN 40.665 0.972 0.446 0.035 0.987 Proposed 43.003 0.981 0.352 0.030 0.993
[0100] In addition, 20 pairs of images were selected as test data on the GaoFen-2 real dataset, and the test results are shown in Table 2. Table 2 is a comparison table of evaluation indicators for 20 pairs of images in the real dataset for different methods.
[0101] Table 2
[0102] Methods D_lamda↓ D_s↓ QNR↑ ADKNet 0.015 0.016 0.969 Hyper-DSNet 0.017 0.022 0.962 MSDDN 0.014 0.023 0.964 BimPAN 0.034 0.077 0.892 Proposed 0.013 0.016 0.971
[0103] Quantitative Index Results: The fusion results obtained by the method proposed in this embodiment are overall better than the existing methods on both the simulation dataset and the real dataset, and there is a significant improvement in the visual effect of fusion.
[0104] Next, a multi-spectral and panchromatic image fusion device based on semantic information proposed according to an embodiment of the present application will be described with reference to the accompanying drawings.
[0105] Figure 5 It is a block diagram of a multi-spectral and panchromatic image fusion device based on semantic information according to an embodiment of the present application.
[0106] As Figure 5 shown, the multi-spectral and panchromatic image fusion device 10 based on semantic information includes: an acquisition module 201, an extraction module 202, and a fusion module 203.
[0107] Among them, the acquisition module 201 is used to acquire an image dataset and a target text description of the image dataset; the extraction module 202 is used to extract semantic features of the multi-spectral image, the panchromatic image, and the target text description in the image dataset; the fusion module 203 is used to input the multi-spectral image, the panchromatic image, and the semantic features into an image fusion network, and the image fusion network outputs an image fusion result. Among them, the image fusion network performs multi-scale image feature extraction on the multi-spectral image and the panchromatic image, and at each scale, it interacts the semantic features with the image features to be fused, and uses semantics to guide image fusion.
[0108] In the embodiment of the present application, the extraction module 202 is further used to: construct a text description dictionary for different ground objects; use the CLIP image-text large model to construct a semantic feature extraction model, and drive the semantic feature extraction model based on the text description dictionary; identify the ground object type in the target text description, and input the ground object type into the semantic feature extraction model, and the semantic feature extraction model outputs the semantic features of the target text description.
[0109] In the embodiment of the present application, the image fusion network includes an encoder and a decoder with multi-layer downsampling and upsampling operations. The encoder is used to extract image features at multiple scales, and the decoder is used to decode features and reconstruct results. Among them, the encoder and the decoder are composed of TCB modules. In the encoder stage, the MTI module is used to interact the semantic features with the image features to be fused, and local features and global features are extracted and integrated through the TCB module.
[0110] In the embodiment of the present application, the MTI module extracts a weight vector and a bias vector through a multi-layer perceptron, and realizes the interaction between the semantic features and the image features to be fused according to the weight vector and the bias vector.
[0111] In the embodiment of the present application, the interaction formula between the semantic features and the image features to be fused is:
[0112]
[0113] Among them, the image feature vector after semantic interaction; is the input image feature vector of the j-th layer of the encoder; α is the weight vector; β is the bias vector.
[0114] In the embodiment of the present application, the fusion module 203 is further configured to: add a skip connection between the encoder and the decoder of the same layer of the image fusion network, and add a residual connection to the input of the multi-spectral image at the input and output ends of the image fusion network.
[0115] In the embodiment of the present application, the TCB module includes a Transformer branch and a CNN branch. The Transformer branch extracts global features, and the CNN branch extracts local features.
[0116] In the embodiment of the present application, a training module is further included. Among them, the training module is further configured to: before inputting the multi-spectral image, the panchromatic image, and the semantic features into the image fusion network, obtain a training data set and an average absolute error loss function; train the image fusion network based on the training data set and the average absolute error loss function, where the average absolute error loss function is:
[0117]
[0118] Among them, I f and I gt are the fused high-resolution multi-spectral image and the multi-spectral image ground truth respectively.
[0119] It should be noted that the foregoing explanation of the embodiment of the multi-spectral and panchromatic image fusion method based on semantic information also applies to the multi-spectral and panchromatic image fusion device based on semantic information in this embodiment, and will not be elaborated here.
[0120] The multi-spectral and panchromatic image fusion device based on semantic information proposed according to the embodiment of the present application obtains an image data set and the target text description of the image data set, extracts the semantic features of the multi-spectral image, the panchromatic image, and the target text description in the image data set, inputs the three into the image fusion network, performs multi-scale image feature extraction on the multi-spectral image and the panchromatic image through the image fusion network, interacts the semantic features with the image features to be fused at each scale, uses semantic guidance for image fusion, obtains an image fusion result, introduces a new perspective into the multi-spectral image fusion, and realizes the image fusion operation under multi-modal interaction through the guidance of semantic information, and can realize the attention to the main ground objects in the image to output a high-quality multi-spectral image.
[0121] Figure 7Schematic diagram of the structure of the electronic device provided by the embodiment of the present application. The electronic device may include:
[0122] A memory 301, a processor 302, and a computer program stored on the memory 301 and executable on the processor 302.
[0123] When the processor 302 executes the program, it implements the multi-spectral and panchromatic image fusion method based on semantic information provided in the above embodiment.
[0124] Furthermore, the electronic device further includes:
[0125] A communication interface 303 for communication between the memory 301 and the processor 302.
[0126] The memory 301 is used to store a computer program executable on the processor 302.
[0127] The memory 301 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0128] If the memory 301, the processor 302, and the communication interface 303 are implemented independently, the communication interface 303, the memory 301, and the processor 302 can be interconnected through a bus and communicate with each other. The bus may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0129] Optionally, in a specific implementation, if the memory 301, the processor 302, and the communication interface 303 are integrated on a single chip, the memory 301, the processor 302, and the communication interface 303 can communicate with each other through an internal interface.
[0130] The processor 302 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiment of the present application.
[0131] In the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0132] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0133] Any process or method description shown in a flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or N executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a manner that is not in the order shown or discussed, including in a substantially simultaneous manner according to the involved functions or in a reverse order, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0134] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware as in another embodiment, any one of the following techniques well known in the art or a combination of them can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, etc.
[0135] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods for implementing the above embodiments can be completed by instructing relevant hardware through a program, and the above program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0136] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limitations on the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A multispectral and panchromatic image fusion method based on semantic information, characterized in that, It includes the following steps: Obtain an image dataset and the target text description of the image dataset; Extract the semantic features of the multi-spectral images, panchromatic images, and target text description in the image dataset; Input the multi-spectral image, the panchromatic image, and the semantic features into an image fusion network, and the image fusion network outputs an image fusion result. Among them, the image fusion network performs multi-scale image feature extraction on the multi-spectral image and the panchromatic image, and at each scale, it interacts the semantic features with the image features to be fused, and uses semantics to guide image fusion.
2. The multispectral and panchromatic image fusion method based on semantic information according to claim 1, wherein The extraction of the semantic features of the target text description of the image dataset includes: Construct a text description dictionary for different ground objects; Use the CLIP image-text large model to construct a semantic feature extraction model, and drive the semantic feature extraction model based on the text description dictionary; Identify the ground object type in the target text description, input the ground object type into the semantic feature extraction model, and the semantic feature extraction model outputs the semantic features of the target text description.
3. The multispectral and panchromatic image fusion method based on semantic information according to claim 1, wherein The image fusion network includes an encoder and a decoder with multiple layers of downsampling and upsampling operations. The encoder is used to extract image features at multiple scales, and the decoder is used for feature decoding and result reconstruction. Among them, the encoder and the decoder are composed of TCB modules. In the encoder stage, the MTI module is used to interact the semantic features with the image features to be fused, and local features and global features are extracted and integrated through the TCB module.
4. The multi-spectral and panchromatic image fusion method based on semantic information according to claim 3, characterized in that, The MTI module extracts a weight vector and a bias vector through a multi-layer perceptron, and realizes the interaction between the semantic features and the image features to be fused according to the weight vector and the bias vector.
5. The multispectral and panchromatic image fusion method based on semantic information according to claim 4, wherein The interaction formula between the semantic features and the image features to be fused is: Among them, The image feature vector after semantic interaction; Is the input image feature vector of the j-th layer of the encoder; α is the weight vector; β is the bias vector.
6. The multispectral and panchromatic image fusion method based on semantic information according to claim 3, wherein Add a skip connection between the encoder and the decoder in the same layer of the image fusion network, and at the input and output ends of the image fusion network, add a residual connection to the input of the multi-spectral image.
7. The method for fusing multi-spectral and panchromatic images based on semantic information according to claim 3, wherein The TCB module includes a Transformer branch and a CNN branch. The Transformer branch extracts global features, and the CNN branch extracts local features.
8. The multispectral and panchromatic image fusion method based on semantic information according to claim 1, characterized in that, Before inputting the multi-spectral image, the panchromatic image, and the semantic features into the image fusion network, it also includes: Obtain a training dataset and a mean absolute error loss function; Train the image fusion network based on the training dataset and the mean absolute error loss function, where the mean absolute error loss function is: Among them, I f and I ft are respectively the fused high-resolution multispectral image and the ground truth of the multispectral image.
9. A multispectral and panchromatic image fusion device based on semantic information, characterized in that, It includes: An acquisition module for obtaining an image dataset and the target text description of the image dataset; An extraction module for extracting the semantic features of the multi-spectral images, panchromatic images, and target text description in the image dataset; A fusion module for inputting the multi-spectral image, the panchromatic image, and the semantic features into an image fusion network, and the image fusion network outputs an image fusion result. Among them, the image fusion network performs multi-scale image feature extraction on the multi-spectral image and the panchromatic image, and at each scale, it interacts the semantic features with the image features to be fused, and uses semantics to guide image fusion.
10. An electronic device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the multi-spectral and panchromatic image fusion method based on semantic information according to any one of claims 1-8.
Citation Information
Cited By
Task prompt perception industrial defect detection model fine tuning method
CN121391828A