A high-resolution multispectral video imaging method and apparatus
By integrating panchromatic video and low spatial resolution multispectral video into a multispectral image fusion network, the constraint between spatial and spectral resolution of multispectral images is solved, enabling the acquisition of high-resolution multispectral images and improving detection efficiency and image quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN UNIV
- Filing Date
- 2024-07-17
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies struggle to improve the spatial resolution of multispectral images while preserving spectral information, resulting in low detection efficiency and poor image quality.
By fusing panchromatic video and low spatial resolution multispectral video acquired from different sensors, a multispectral image fusion network is used to perform hierarchical analysis of spectral and spatial differences, extract spectral and spatial prior knowledge, and reconstruct high-resolution multispectral images using a decoder.
This approach improves the spatial resolution of multispectral images while preserving spectral information, thereby increasing detection efficiency and enhancing image quality.
Smart Images

Figure CN119090733B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and specifically to a high-resolution multispectral video imaging method and apparatus. Background Technology
[0002] Multispectral images contain rich spectral information, thus finding wide application in classification, target detection, environmental monitoring, and agricultural monitoring. However, due to hardware limitations, a single sensor cannot simultaneously achieve high spectral and high spatial resolution. Therefore, in multispectral imaging, spatial resolution is often sacrificed to acquire spectral information. Existing spectral imaging technologies are constrained by the relationship between spatial and spectral resolution, making it challenging to directly acquire high spatial resolution multispectral images, which reduces the practical application value of multispectral images. To address these issues, one feasible approach is to design advanced sensors, but this involves technical challenges and high costs. Another feasible approach is to fuse images acquired by different sensors to achieve information complementarity, thereby effectively improving detection efficiency and image quality. This fusion method provides a feasible solution to overcome the limitations of multispectral images, improving spatial resolution while preserving spectral information. However, how to specifically fuse images acquired by different sensors to achieve information complementarity and effectively improve detection efficiency and image quality has become a key technical problem that urgently needs to be solved. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a high-resolution multispectral video imaging method and apparatus to address the above-mentioned problems in the prior art. The present invention aims to achieve the acquisition of high-resolution multispectral video by fusing panchromatic video and low spatial resolution multispectral video obtained from different sensors of the same scene, thereby improving detection efficiency and image quality.
[0004] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0005] A high-resolution multispectral video imaging method includes inputting a low-resolution multispectral image from a multispectral image video and a high-resolution panchromatic image from a panchromatic image video into a pre-trained multispectral image fusion network to reconstruct a final high-resolution multispectral image. The multispectral image fusion network reconstructs the final high-resolution multispectral image through the following steps:
[0006] S101 will convert low-resolution multispectral images The input embedding layer yields the embedded low-resolution multispectral image. ; to convert high-resolution panchromatic images The input embedding layer yields an embedded high-resolution panchromatic image. ;
[0007] S102, embeds a high-resolution panchromatic image By combining low-resolution multispectral images conduct The system analyzes spectral and spatial differences at each stage to extract spectral and spatial prior knowledge. The output features are downsampled at each stage of spectral and spatial difference analysis, and the final downsampled stage yields the final output features. ; embed a high-resolution panchromatic image By combining low-resolution multispectral images conduct The system analyzes spatial and spectral differences at each stage to extract spectral and spatial prior knowledge. The output features are downsampled at each stage of the spectral and spatial difference analysis, and the final downsampled feature is used to obtain the output features. ;
[0008] S103, features and characteristics The initial fusion features are obtained through fusion. ;
[0009] S104, initial fusion features By combining spectral and spatial prior experience, the final high-resolution multispectral image is reconstructed using a decoder.
[0010] Optionally, in step S102, the embedded high-resolution panchromatic image By combining low-resolution multispectral images Multi-level analysis of spectral and spatial differences to extract spectral and spatial prior knowledge is performed through the spectral difference analysis module. Implemented; the spectral difference analysis module Includes a first linear normalization layer (LN) connected in sequence, and a spectral low-rank cross-attention module. The system comprises a first overlay module, a second linear normalization layer (LN), a feedforward network (FFN), and a second overlay module, wherein the output of the first linear normalization layer (LN) also serves as the input to the first overlay module, and the output of the functional layer also serves as the input to the second overlay module; the spectral low-rank cross-attention module. The function expression for the functional layer is:
[0011] ,
[0012] ,
[0013] In the above formula, For spectral low-rank cross-attention modules Output characteristics Indicates a stacking operation. For the number of attention heads, For the result of cross-attention of the j-th attention head, For learnable parameters, For location embedding functions, location embedding functions It consists of two 3×3 convolutional layers and a GELU activation layer; Let j be the value of the attention head. Let be the attention score of the j-th attention head, and we have:
[0014] ,
[0015] In the above formula, The softmax activation function is used. The key is obtained by linearly mapping the input features. The query is obtained by linearly mapping the input features; the value obtained by linearly mapping the input features is... ,in For the high-resolution panchromatic image embedded in the previous stage, For low-resolution multispectral images The image obtained by changing the size, The number of stages. , and These are learnable parameters.
[0016] Optionally, in step S102, the embedded high-resolution panchromatic image By combining low-resolution multispectral images Multi-level analysis of spatial and spectral differences to extract spectral and spatial prior knowledge is performed through the spatial difference analysis module. Implemented; the spatial difference analysis module Includes a first linear normalization layer (LN) connected in sequence, and a spatially low-rank cross-attention module. The system comprises a first overlay module, a second linear normalization layer (LN), a feedforward network (FFN), and a second overlay module, wherein the output of the first linear normalization layer (LN) also serves as the input to the first overlay module, and the output of the functional layer also serves as the input to the second overlay module; the spatial low-rank cross-attention module. The function expression is:
[0017] ,
[0018] In the above formula, For spatial low-rank cross-attention modules Output characteristics Indicates a stacking operation. For the i-th channel and Features of spatial cross-attention output for and Features of spatial cross-attention output Spatial difference analysis module for the current stage The output results, For location embedding functions, location embedding functions It consists of two 3×3 convolutional layers and a GELU activation layer. Represents the first of N encoding stages Each coding stage This represents the number of spectral bands in the image; and we have:
[0019]
[0020]
[0021] In the above formula, for and Output of spatial cross attention for and Output of spatial cross attention and These are learnable parameters. From and The attention score obtained for the j-th head. for and The attention score obtained for the j-th head. This is a subspace of features extracted from low-resolution multispectral images by the previous network. For the feature subspace of the embedded high-resolution panchromatic image, These are the spatial coefficients of the features extracted from the low-resolution multispectral image by the previous network. Let be the spatial coefficients of the features of the embedded high-resolution panchromatic image, where:
[0022] ,
[0023] ,
[0024] ,
[0025] ,
[0026] In the above formula, This refers to the feature extracted from the i-th channel of the low-resolution multispectral image by the previous network. For the features of the i-th channel of the embedded high-resolution panchromatic image, , , These are learnable parameters.
[0027] Optionally, in step S103, the features and characteristics The initial fusion features are obtained through fusion. The function expression is:
[0028] ,
[0029] In the above formula, This represents the bottleneck structure module. This indicates a stacking operation.
[0030] Optionally, the bottleneck structure module consists of a convolutional module with a kernel size of 1×1 and a spectral-spatial prior-guided fusion module. Composition, the spectral-spatial prior-guided fusion module It includes a first linear normalization layer (LN) connected in sequence, and a spectral-spatial prior-guided self-attention module. The system comprises a first stacking module, a second linear normalization layer (LN), a feedforward network (FFN), and a second stacking module, and the spectral-spatial prior-guided self-attention module. The function expression is:
[0031] ,
[0032] In the above formula, Self-attention module guided by spectral-spatial priors The output characteristics, Indicates a stacking operation. For the number of attention heads, For the result of cross-attention of the j-th attention head, For learnable parameters, For location embedding functions, location embedding functions It consists of two 3×3 convolutional layers and a GELU activation layer; For spatial spectrum prior experience, and we have:
[0033] ,
[0034] In the above formula, For the fusion module guided by the nth-level spectral-spatial prior using an attention mechanism Input features The extracted value, and These represent the embedded low-resolution multispectral images. and embedded high-resolution panchromatic images The spectral prior experience and spatial prior experience extracted from them.
[0035] Optionally, in step S104, the initial fusion features are... Combining spectral and spatial prior experience, the final high-resolution multispectral image is reconstructed using a decoder, including:
[0036] S201, initial fusion features pass Level spectral-spatial prior-guided fusion module Decoding features are obtained by combining spectral prior experience and spatial prior experience. Furthermore, each level of spectral-spatial prior-guided fusion module... The input includes a channel upsampling module. And stacking operations, and each level of spectral-spatial prior-guided fusion module. The functional expression for the output feature is:
[0037] ,
[0038] In the above formula, For the nth level spectral-spatial prior-guided fusion module The output characteristics, This represents the nth-order spectral-spatial prior-guided fusion module. , For the nth level spectral-spatial prior-guided fusion module The input features, and and These represent the embedded low-resolution multispectral images. and embedded high-resolution panchromatic images The extracted spectral and spatial prior experiences are:
[0039] ,
[0040] ,
[0041] In the above formula, Indicates a stacking operation. Indicates the upsampling module of the nth channel. The output characteristics, This indicates the upsampling module for the (n+1)th channel. The output characteristics, Indicates the upsampling module of the nth channel. , For the N-n+1th stage The module's output, ;
[0042] S202, the final high-resolution multispectral image is reconstructed according to the following formula:
[0043] ,
[0044] In the above formula, For the final high-resolution multispectral image, For convolution operations, Indicates a stacking operation. For the initial embedding of low-resolution multispectral images, This is a high-resolution panchromatic image that is being embedded for the first time.
[0045] Furthermore, the present invention also provides a high-resolution multispectral video imaging device, including a microprocessor and a memory interconnected thereto, the microprocessor being programmed or configured to execute the high-resolution multispectral video imaging method.
[0046] Optionally, the microprocessor is also connected to an optical and sensing module, which includes an objective lens, a beam splitter, an eyepiece, a panchromatic imaging sensor, and a multispectral coding sensor. The eyepiece includes a first eyepiece and a second eyepiece. External light enters through the objective lens and is split into two beams by the beam splitter. One beam passes through the first eyepiece and enters the panchromatic imaging sensor, while the other beam passes through the second eyepiece and enters the multispectral coding sensor. The panchromatic imaging sensor and the multispectral coding sensor are respectively connected to the microprocessor.
[0047] Furthermore, the present invention also provides a computer-readable storage medium storing a computer program or instructions that are programmed or configured to execute the high-resolution multispectral video imaging method by a processor.
[0048] In addition, the present invention also provides a computer program product, including a computer program or instructions that are programmed or configured to execute the high-resolution multispectral video imaging method via a processor.
[0049] Compared with the prior art, the present invention has the following main advantages:
[0050] 1. The high-resolution multispectral video imaging method of the present invention includes processing low-resolution multispectral images... and high-resolution panchromatic images Acquire embedded low-resolution multispectral images and high-resolution panchromatic images ;Will By combining conduct Level analysis of spectral and spatial differences yields output characteristics ;Will By combining conduct Level analysis of spatial and spectral differences to extract output features ; Features and characteristics The initial fusion features are obtained through fusion. Then the initial fusion features The high-resolution multispectral image is reconstructed using a decoder. The high-resolution multispectral video imaging method of this invention achieves the acquisition of high-resolution multispectral video by fusing panchromatic video and low spatial resolution multispectral video obtained from different sensors of the same scene. It can achieve information complementarity by fusing images obtained from different sensors, thereby effectively improving detection efficiency and image quality.
[0051] 2. The high-resolution multispectral video imaging method of the present invention is applicable to the fusion of various types of panchromatic and multispectral image data, and has a wide range of applications. Attached Figure Description
[0052] Figure 1 This is a schematic diagram of the network structure of the multispectral image fusion network in an embodiment of the present invention.
[0053] Figure 2 The spectral difference analysis module in this embodiment of the invention A schematic diagram of the network structure.
[0054] Figure 3 This is a schematic diagram of the network structure of the feedforward network (FFN) in an embodiment of the present invention.
[0055] Figure 4 Spatial difference analysis module in this embodiment of the invention A schematic diagram of the network structure.
[0056] Figure 5 This is a spatial low-rank cross-attention module in an embodiment of the present invention. A schematic diagram of the network structure.
[0057] Figure 6 This is the spectral-spatial prior-guided fusion module in this embodiment of the invention. A schematic diagram of the network structure.
[0058] Figure 7 This is a spectral-spatial prior-guided self-attention module in an embodiment of the present invention. A schematic diagram of the network structure.
[0059] Figure 8 This is a spectral angle error diagram of different methods in the embodiments of the present invention.
[0060] Figure 9 This is a spectral angle error diagram of the reconstructed scene using different methods in the embodiments of the present invention.
[0061] Figure 10 This is a diagram showing the spectral angle error of reconstruction using different methods in the embodiments of the present invention.
[0062] Figure 11 This is a spectral angle error diagram for predicting HR-HSI degraded HSI using different fusion methods in embodiments of the present invention.
[0063] Figure 12 The classification results before and after fusion are shown in the embodiments of the present invention.
[0064] Figure 13 This is a schematic diagram of the optical and sensing module structure of the high-resolution multispectral video imaging device in an embodiment of the present invention. Detailed Implementation
[0065] like Figure 1 As shown, this embodiment provides a high-resolution multispectral video imaging method, which includes inputting a low-resolution multispectral image from a multispectral image video and a high-resolution panchromatic image from a panchromatic image video into a pre-trained multispectral image fusion network to reconstruct the final high-resolution multispectral image. The multispectral image fusion network reconstructs the final high-resolution multispectral image through the following steps:
[0066] S101 will convert low-resolution multispectral images The input embedding layer yields the embedded low-resolution multispectral image. ; to convert high-resolution panchromatic images The input embedding layer yields an embedded high-resolution panchromatic image. ;
[0067] S102, embeds a high-resolution panchromatic image By combining low-resolution multispectral images conduct The system analyzes spectral and spatial differences at each stage to extract spectral and spatial prior knowledge. The output features are downsampled at each stage of spectral and spatial difference analysis, and the final downsampled stage yields the final output features. ; embed a high-resolution panchromatic image By combining low-resolution multispectral images conduct The system analyzes spatial and spectral differences at each stage to extract spectral and spatial prior knowledge. The output features are downsampled at each stage of the spectral and spatial difference analysis, and the final downsampled feature is used to obtain the output features. ;
[0068] S103, features and characteristics The initial fusion features are obtained through fusion. ;
[0069] S104, initial fusion features By combining spectral and spatial prior experience, the final high-resolution multispectral image is reconstructed using a decoder.
[0070] like Figure 1 As shown, the low-resolution multispectral image in this embodiment Image frames in low-resolution multispectral imaging (LR-MSI) video frames, high-resolution panchromatic images. This refers to an image frame within a high-resolution panchromatic image (HR-PAN) video frame. Alternatively, a separate image frame can be used instead of a video frame; the principle is similar, so it will not be elaborated further below. Low-resolution multispectral image (LR-MSI) and high-resolution panchromatic image (HR-PAN) Inputting each image into the embedding layer yields the embedded low-resolution multispectral image. Embedded high-resolution panchromatic image .in For low-resolution multispectral images, For high-resolution panchromatic images, and Indicates the number of spectral bands in the image. and It represents the spatial scale information of the image. and The function represents the embedding layer, where " "" indicates "much smaller," for example, a value less than 1 / 10 can represent "much smaller." The function expression for the embedding layer is:
[0071] ,
[0072] ,
[0073] In this embodiment, the embedding layer functions all consist of N pairs of bilinear interpolation layers and 3×3 convolutional layers, where N is the total order of the encoding stage. Represents the first of N encoding stages Each coding stage .
[0074] In step S102 of this embodiment, the embedded high-resolution panchromatic image By combining low-resolution multispectral images Multi-level analysis of spectral and spatial differences to extract spectral and spatial prior knowledge is performed through the spectral difference analysis module. Implementation; such as Figure 2 As shown, the spectral difference analysis module in this embodiment Includes a first linear normalization layer (LN) connected in sequence, and a spectral low-rank cross-attention module. The system comprises a first overlay module, a second linear normalization layer (LN), a feedforward network (FFN), and a second overlay module, wherein the output of the first linear normalization layer (LN) also serves as the input to the first overlay module, and the output of the functional layer also serves as the input to the second overlay module; the spectral low-rank cross-attention module. The function expression for the functional layer is:
[0075] ,
[0076] ,
[0077] In the above formula, For spectral low-rank cross-attention modules Output characteristics Indicates a stacking operation. For the number of attention heads, For the result of cross-attention of the j-th attention head, For learnable parameters, For location embedding functions, location embedding functions It consists of two 3×3 convolutional layers and a GELU activation layer; Let j be the value of the attention head. Let be the attention score of the j-th attention head, and we have:
[0078] ,
[0079] In the above formula, The softmax activation function is used. The key is obtained by linearly mapping the input features. The query is obtained by linearly mapping the input features; the value obtained by linearly mapping the input features is... ,in For the high-resolution panchromatic image embedded in the previous stage, For low-resolution multispectral images The image obtained by changing the size, The width of the image. The height of the image. This represents the number of channels in the image. The number of stages. , and These are learnable parameters.
[0080] like Figure 3 As shown, the feedforward network FFN consists of a series of interconnected 1×1 convolutional layers, LeakReLU activation functions, 1×1 convolutional layers, LeakReLU activation functions, and 1×1 convolutional layers.
[0081] In step S102 of this embodiment, the embedded high-resolution panchromatic image By combining low-resolution multispectral images Multi-level analysis of spatial and spectral differences to extract spectral and spatial prior knowledge is performed through the spatial difference analysis module. Implementation; such as Figure 4 As shown, the spatial difference analysis module in this embodiment Includes a first linear normalization layer (LN) connected in sequence, and a spatially low-rank cross-attention module. The system consists of a first overlay module, a second linear normalization layer (LN), a feedforward network (FFN), and a second overlay module, wherein the output of the first linear normalization layer (LN) is also used as the input of the first overlay module, and the output of the functional layer is also used as the input of the second overlay module; for example... Figure 5 As shown, the spatial low-rank cross-attention module The function expression is:
[0082] ,
[0083] In the above formula, For spatial low-rank cross-attention modules Output characteristics Indicates a stacking operation. For the i-th channel and Features of spatial cross-attention output for and Features of spatial cross-attention output Spatial difference analysis module for the current stage The output results, For location embedding functions, location embedding functions It consists of two 3×3 convolutional layers and a GELU activation layer. Represents the first of N encoding stages Each coding stage This represents the number of spectral bands in the image; and we have:
[0084]
[0085]
[0086] In the above formula, for and Output of spatial cross attention for and Output of spatial cross attention and These are learnable parameters. Here is a hyperparameter used to represent the number of fundamental elements in a subspace. From and The attention score obtained for the j-th head. for and The attention score obtained for the j-th head. This is a subspace of features extracted from low-resolution multispectral images by the previous network. For the feature subspace of the embedded high-resolution panchromatic image, These are the spatial coefficients of the features extracted from the low-resolution multispectral image by the previous network. Let be the spatial coefficients of the features of the embedded high-resolution panchromatic image, where:
[0087] ,
[0088] ,
[0089] ,
[0090] ,
[0091] In the above formula, This refers to the feature extracted from the i-th channel of the low-resolution multispectral image by the previous network. For the features of the i-th channel of the embedded high-resolution panchromatic image, , , These are learnable parameters.
[0092] The encoder output is input to... and In the module ( Module, The modules are all Transformer structures composed of attention modules and feedforward networks (FFNs). The spectral and spatial differences between multispectral and panchromatic images are analyzed, and spectral and spatial prior experiences are extracted. and For the first Phase and The output corresponding to the module, and For the previous stage and Module output results, when hour, .
[0093] ,
[0094] ,
[0095] exist In the module, The purpose of this module is to analyze the differences between PAN and HR-MSI, in which the first In each attention module, the embedded low-resolution multispectral image is... Size changed to Next, the features extracted from MSI by the previous network are first linearly mapped to obtain the query (Q), key (K), and value (V). , and These are the learnable parameters in the linear projection layer. Depending on the different stages... , and The data is split into J heads, and then the similarity score of each head is calculated:
[0096] ,
[0097] ,
[0098] ,
[0099] ,
[0100] in Representing the Attention score per head ,(·) T This indicates the transpose operation.
[0101] Based on this, Representing the The value of each head The result of cross-attention is calculated as follows: :
[0102] ,
[0103] ,
[0104] in, Representing the Size . yes The module's output. These are learnable parameters. It is a positional embedding function, consisting of two 3×3 convolutional layers and GELU activation. In the module, From embedded images What is obtained can be regarded as prior experience of the spectrum. and From The calculated value contains information about PAN. Since MSI contains true spectral information, The module can obtain spectral priors from MSI, enabling the model to measure the spectral differences between PAN and MSI.
[0105] exist In this module, to ensure the model can capture global spatial information while reducing computational cost, we introduce matrix factorization during the attention calculation process. As shown in Equation (11), we factorize the two-dimensional matrix... Decomposed into subspace coefficients and spatial coefficient We use deep learning methods to implement matrix factorization, defining four parameter matrices to decompose the MSI features extracted from the previous network and embedding them into the PAN:
[0106] ,
[0107] ,
[0108] ,
[0109] ,
[0110] ,
[0111] in , , These are learnable parameters. and These are features The subspace and subspace coefficients. and These are features The subspace and subspace coefficients. It is a hyperparameter representing the number of basis elements in the subspace. Setting it to 1 allows for faster and more flexible matrix factorization using deep learning, and also enhances the expressive power of the model.
[0112] Attention score of the j-th head from and The attention score obtained for the j-th head. From and It was obtained from this module. Q is used in this module. S and Q A From respectively and Calculated. [K] S V S ] and [K A V A From respectively and This was calculated. Next, all the heads are connected as shown in the following formula:
[0113] ,
[0114] ,
[0115] in, and yes[ , ]and[ , The result obtained from cross-attention calculation. and These are learnable parameters. Subsequently, and Multiply the results and embed the location information into the final attention result. In the following formula:
[0116] ,
[0117] in, and The position embedding functions in the module are the same.
[0118] Finally, the outputs of the two modules are input into the downsampling module as shown in the following equation:
[0119] ,
[0120] ,
[0121] in, and Both are single 4×4 convolutional layers, which can achieve spatial downsampling and channel upsampling with a sampling factor of 2. Then, the final outputs of the two branches need to be processed. and Initial fusion is performed. In step S103 of this embodiment, the features are... and characteristics The initial fusion features are obtained through fusion. The function expression is:
[0122] ,
[0123] In the above formula, This represents the bottleneck structure module. This represents a stacking operation (along the channel dimension and pointwise convolutional layers). In this embodiment, the bottleneck structure module consists of a convolutional module with a 1×1 kernel and a spectral-spatial prior-guided fusion module. composition.
[0124] like Figure 6 As shown, the spectral-spatial prior-guided fusion module It includes a first linear normalization layer (LN) connected in sequence, and a spectral-spatial prior-guided self-attention module. The first stacking module, the second linear normalization layer LN, the feedforward network FFN, and the second stacking module are as follows: Figure 7 As shown, the spectral-spatial prior-guided self-attention module The function expression is:
[0125] ,
[0126] In the above formula, Self-attention module guided by spectral-spatial priors The output characteristics, Indicates a stacking operation. For the number of attention heads, For the result of cross-attention of the j-th attention head, For learnable parameters, For location embedding functions, location embedding functions It consists of two 3×3 convolutional layers and a GELU activation layer; For spatial spectrum prior experience, and we have:
[0127] ,
[0128] In the above formula, For the fusion module guided by the nth-level spectral-spatial prior using an attention mechanism Input features The extracted value, and These represent the embedded low-resolution multispectral images. and embedded high-resolution panchromatic images The spectral prior experience and spatial prior experience extracted are shown in the figure as Spe and Spa, respectively.
[0129] In step S104 of this embodiment, the initial fusion features are... Combining spectral and spatial prior experience, the final high-resolution multispectral image is reconstructed using a decoder, including:
[0130] S201, initial fusion features pass Level spectral-spatial prior-guided fusion module Decoding features are obtained by combining spectral prior experience and spatial prior experience. Furthermore, each level of spectral-spatial prior-guided fusion module... The input includes a channel upsampling module. And stacking operations, and each level of spectral-spatial prior-guided fusion module. The functional expression for the output feature is:
[0131] ,
[0132] In the above formula, For the nth level spectral-spatial prior-guided fusion module The output characteristics, This represents the nth-order spectral-spatial prior-guided fusion module. , For the nth level spectral-spatial prior-guided fusion module The input features, and and These represent the embedded low-resolution multispectral images. and embedded high-resolution panchromatic images The extracted spectral and spatial prior experiences are:
[0133] ,
[0134] ,
[0135] In the above formula, Indicates a stacking operation. Indicates the upsampling module of the nth channel. The output characteristics, Indicates the upsampling module of the nth channel. , For the N-n+1th stage The module's output, ;
[0136] S202, the final high-resolution multispectral image is reconstructed according to the following formula:
[0137] ,
[0138] In the above formula, For the final high-resolution multispectral image, For convolution operations, Indicates a stacking operation. For the initial embedding of low-resolution multispectral images, This is a high-resolution panchromatic image that is being embedded for the first time.
[0139] To enhance the transmission of shallow feature information, this embodiment also fuses the corresponding encoded features before downsampling, wherein the upsampling module... ( This is used to perform Pixel-Shuffle (PS). In this embodiment, a 5×5 convolutional layer is specifically used to implement channel upsampling of PS. It is the first indivual The function of the module (a Transformer structure consisting of attention blocks and feedforward networks (FFNs) in the decoder) . and They represent from and The prior experience extracted from it, when hour, , This is the corresponding output.
[0140] ,
[0141] ,
[0142] ,
[0143] exist In the module, based on fusion features Through calculation , , and Then, spatial spectrum prior experience was introduced. ,in and They represent from and Prior experience extracted from it.
[0144] ,
[0145] Next, embedded multispectral images were introduced. and embedded full-color images This includes spectral prior experience and spatial prior experience, with the spatial spectral prior guidance mechanism driving... The module pays more attention to the spatial spectral priors hidden in PAN and MSI:
[0146] , ,
[0147] ,
[0148] in and These are spatial masks and spectral masks. (·)and (·) represents the functions for the spatial attention module and the spectral attention module. Finally, self-attention is performed. Result calculation:
[0149] ,
[0150] ,
[0151] Finally, a global residual connection is introduced to reconstruct the final high-resolution multispectral image:
[0152] ,
[0153] In the above formula, For the final high-resolution multispectral image, For convolution operations (3×3 convolutional layer). Indicates a stacking operation. For the initial embedding of low-resolution multispectral images, This is a high-resolution panchromatic image that is being embedded for the first time.
[0154] To verify the high-resolution multispectral video imaging method of this embodiment, the high-resolution multispectral video imaging method (Ours) of this embodiment is compared with the existing methods using NSSR, LTTR, SSRnet, Fformer and SSTFUNET models. For details of NSSR, please refer to: Dong W, Fu F, Shi G, et al. Hyperspectral image super-resolution via non-negative structured sparse representation[J]. IEEE Transactions on Image Processing, 2016, 25(5): 2337-2352.; for details of LTTR, please refer to: Dian R, Li S, Fang L. Learning a low tensor-train rank representation for hyperspectral image super-resolution[J]. IEEE transactions on neural networks and learning systems, 2019, 30(9): 2672-2683.; for details of SSRnet, please refer to: Zhang X, Huang W, Wang Q, et al. SSR-NET: Spatial–spectral reconstruction network for hyperspectral and multispectral image fusion[J]. IEEE Transactions onGeoscience and Remote Sensing, 2020, 59(7): 5953-5965.; For details on Fformer, see the literature: Hu JF, Huang TZ, Deng LJ, et al. Fusformer: A transformer-based fusion network for hyperspectral image super-resolution[J]. IEEE Geoscience and RemoteSensing Letters, 2022, 19: 1-5.; For details of SSTF-Unet, please refer to the literature: Liu H, Feng C, Dian R, etal.Sstf-unet: Spatial–spectral transformer-based u-net for high-resolution hyperspectral image acquisition[J]. IEEE Transactions on Neural Networks and Learning Systems, 2023. The dataset used is CAVE, KAIST, and ICVL. Some test results are shown below. Figure 8 , Figure 9 , Figure 10 , Figure 11 as well as Figure 12 As shown, where Figure 8 This is a spectral angle error diagram of the reconstructed jelly bean MS (row 1), real apple (row 2), and fake chili MS (row 3) images of three test images in the CAVE dataset using different methods in this embodiment. Figure 9 This is a spectral angle error plot showing the reconstructed reflectance of scene 21 (row 1), scene 22 (row 2), and scene 26 (row 3) of three test images on the KAIST dataset using different methods in this embodiment. Figure 10 This is a spectral angle error map of the reconstruction of three test images bgu_0403-1511 (row 1), paper_0503-1330 (row 2), and sami_0331-1019 (row 3) on the ICVL dataset using different methods in this embodiment. Figure 11 In this embodiment, (a) to (g) are spectral angle error diagrams of HR-HSI degradation HSI predicted by different fusion methods on the Gaofen-5 and Gaofen-1 datasets. Figure 12 The following are the classification results before and after fusion in this embodiment, where (a) is the classification result of LR-HSI, (b) is the classification result of predicted HR-HSI, and (c) is the reference classification result. Figures 8-12 The effectiveness of the high-resolution multispectral video imaging method in this embodiment has been verified. Furthermore, the statistical results obtained in this embodiment are shown in Tables 1 to 6.
[0155] Table 1: Quantitative metrics for testing methods on the CAVE, KAIST, and ICVL datasets
[0156]
[0157] The best results in Table 1 are marked in bold. As shown in Table 1, our LRTN achieves the best results in all three metrics: PSNR (Peak Signal-to-Noise Ratio), ERGAS (Root Mean Square Error / Signal-to-Noise Ratio), and SAM (Spectral Angle Mapper). In the UIQI (Global Image Quality Index) test on the CAVE dataset, LRTN ranks second. For the KAIST and ICVL datasets, LRTN demonstrates the best performance across all evaluation metrics. Furthermore, SSTF-Unet's performance is very close to LRTN. From the table, we can see that deep learning-based methods have significant advantages over traditional methods. Among all methods, although SSRnet has a significant speed advantage, its performance in fusion quality is limited, falling short of other deep learning-based methods. In summary, our method LRTN achieves the best fusion results while maintaining fast processing speed, exhibiting superior fusion quality compared to other methods.
[0158] Table 2: External validation experimental results of deep learning-based methods
[0159]
[0160] The best results in Table 2 are marked in bold. As shown in Table 2, all methods were trained on the KAIST dataset and tested on the CAVE dataset. From the results, we can see that although our method LRTN has a slightly lower SAM (Spectral Angle Mapper) than SSTF-Unet in this case, LRTN exhibits the best results on all other evaluation metrics. This indicates that LRTN performs best on most evaluation metrics.
[0161] Table 3: Runtime of the method tested on the Gaofen5 and Gaofen1 datasets, with the best results marked in bold.
[0162]
[0163] The best results in Table 3 are marked in bold. As shown in Table 3, our method LRTN ranks second among all compared methods, second only to SSRnet, which maintains the fastest reconstruction speed. Overall, although our method is slightly slower than SSRnet in fusion speed, it achieves the highest level of fusion quality.
[0164] Table 4: Ablation experiments of SEDL, SADL, and SSGF modules on the CAVE and KAIST datasets.
[0165]
[0166] The best results in Table 4 are marked in bold. As shown in Table 4, when we tried replacing key modules with residual blocks or replacing prior knowledge in SSGF, all evaluation metrics on both datasets decreased. For example, for the CAVE dataset, PSNR (Peak Signal-to-Noise Ratio) decreased by 0.22dB, 0.37dB, 0.90dB, and 0.16dB, respectively. Furthermore, while the "No-MF" model performed similarly to LRTN, its parameters increased by 71.26Mb. Considering the computational complexity analysis, we can conclude that fully utilizing the inherent properties of HSI can reduce model parameters and computational complexity while maintaining fusion accuracy.
[0167] Table 5: Ablation experiments of hyperparameters N and L on the CAVE and KAIST datasets
[0168]
[0169] The best results in Table 5 are marked in bold. As shown in Table 5, LRTN performs best on the CAVE and KAIST datasets when N = 2. We set the hyperparameter L to three different values: L = 1, 2, and 3. The results show that LRTN's performance hardly changes in these three cases, indicating that our method is not sensitive to the setting of hyperparameter L and has good robustness. When choosing hyperparameter N, we selected N = 2 based on the best performance results. Regarding hyperparameter L, although the performance of LRTN does not change much with L = 1, 2, and 3, we chose L = 1 to minimize computational cost, thus reducing computation while maintaining performance.
[0170] Table 6: Classification results of LR-HSI and predicted HSI, with the best results marked in bold.
[0171]
[0172] The best results in Table 6 are marked in bold. As shown in Table 6, most classification results for high-resolution (HR-HSI) data are superior to those for low-resolution (LR-HSI) data. Furthermore, the average accuracy of the high-resolution (HR-HSI) predictions is 6.8% higher than that of the low-resolution (LR-HSI) predictions, demonstrating that the high-resolution multispectral video imaging method in this embodiment has higher accuracy in processing high-resolution data.
[0173] In summary, the high-resolution multispectral video imaging method of this embodiment includes processing low-resolution multispectral images... and high-resolution panchromatic images Acquire embedded low-resolution multispectral images and high-resolution panchromatic images ;Will By combining conduct Level analysis of spectral and spatial differences yields output characteristics ;Will By combining conduct Level analysis of spatial and spectral differences to extract output features ; Features and characteristics The initial fusion features are obtained through fusion. Then the initial fusion features The final high-resolution multispectral image is reconstructed using a decoder. This embodiment of the high-resolution multispectral video imaging method achieves high-resolution multispectral video acquisition by fusing panchromatic videos and low spatial resolution multispectral videos obtained from different sensors of the same scene. This allows for the fusion of images from different sensors to achieve information complementarity, thereby effectively improving detection efficiency and image quality. This high-resolution multispectral video imaging method is applicable to the fusion of various types of panchromatic and multispectral image data, and has a wide range of applications.
[0174] Furthermore, this embodiment also provides a high-resolution multispectral video imaging device, including a microprocessor and a memory interconnected, wherein the microprocessor is programmed or configured to execute the high-resolution multispectral video imaging method described above in this embodiment. The microprocessor in this embodiment is also connected to optical and sensing modules. For example... Figure 13As shown, the optical and sensing module in this embodiment includes an objective lens, a beam splitter, an eyepiece, a panchromatic imaging sensor (PAN imaging sensor), and a multispectral coding sensor (MSI coding sensor). The eyepiece includes a first eyepiece and a second eyepiece. External light entering through the objective lens is split into two beams by the beam splitter. One beam passes through the first eyepiece and enters the panchromatic imaging sensor, while the other beam passes through the second eyepiece and enters the multispectral coding sensor. The panchromatic imaging sensor and the multispectral coding sensor are respectively connected to a microprocessor. This embodiment establishes a data acquisition system and proposes a hardware-software combined method to achieve effective acquisition of high-resolution multispectral video. This embodiment establishes an innovative multispectral imaging model. Utilizing the proposed method, it fully leverages spectral and spatial priors to guide feature extraction and fusion. This method promotes the model to pay more attention to the missing information in different images and extracts feature information more efficiently from panchromatic and multispectral images. This embodiment designs a low-rank Transformer model, combining Transformer modules with matrix factorization theory, significantly reducing the number of model parameters and computational cost. A spatial spectral prior-guided fusion module is designed in the feature fusion stage, improving the spatial and spectral fidelity of the image and enhancing the network fusion performance. This embodiment is applicable to the fusion of various types of panchromatic and multispectral image data, possessing a wide range of applications. This embodiment discloses a high-resolution multispectral video imaging device, which includes a hardware system for high-resolution multispectral computational imaging, mainly composed of a beam splitter prism, a panchromatic (PAN) image imaging sensor, and a multispectral (MSI) coded sensor. By employing the MSI and panchromatic image fusion imaging method, this invention fuses a high-spectral-resolution, low-spatial-resolution multispectral image with a high-spatial-resolution, low-spectral-resolution panchromatic image, thereby achieving the acquisition of high-resolution multispectral video. This embodiment has several advantages, such as high temporal and spatial resolution, low cost, high signal-to-noise ratio, and high computational efficiency, overcoming the problem of mutual constraints between spatial and spectral resolution in existing multispectral cameras. This device enables the acquisition of high-quality multispectral video images more effectively, providing superior solutions for various application scenarios.
[0175] Furthermore, this embodiment also provides a computer-readable storage medium storing a computer program or instructions programmed or configured to execute the high-resolution multispectral video imaging method via a processor. Additionally, this embodiment also provides a computer program product including a computer program or instructions programmed or configured to execute the high-resolution multispectral video imaging method via a processor.
[0176] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0177] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A high-resolution multispectral video imaging method, characterized in that, The method involves inputting low-resolution multispectral images from multispectral image videos and high-resolution panchromatic images from panchromatic image videos into a pre-trained multispectral image fusion network to reconstruct the final high-resolution multispectral image. The reconstruction of the final high-resolution multispectral image by the multispectral image fusion network includes the following steps: S101 will convert low-resolution multispectral images The input embedding layer yields the embedded low-resolution multispectral image. ; to convert high-resolution panchromatic images The input embedding layer yields an embedded high-resolution panchromatic image. ; S102, via the spectral difference analysis module Embedded high-resolution panchromatic image Combined with embedded low-resolution multispectral images conduct The system performs level-by-level analysis of spectral and spatial differences, and at each level, the output features are downsampled and output, with the final downsampling level yielding the final output features. ; Spatial difference analysis module Embedded low-resolution multispectral images Combined with embedded high-resolution panchromatic images conduct The system analyzes spatial and spectral differences at each stage, and the output features are downsampled during each stage of spectral and spatial difference analysis. The final downsampled feature is then used to obtain the output features. ; S103, features and characteristics The initial fusion features are obtained through fusion. ; S104, initial fusion features By combining spectral and spatial prior experience, a decoder is used to reconstruct the final high-resolution multispectral image, including: S201, initial fusion features pass Level spectral-spatial prior-guided fusion module Decoding features are obtained by combining spectral prior experience and spatial prior experience. Furthermore, each level of spectral-spatial prior-guided fusion module... The input includes a channel upsampling module. And stacking operations, and each level of spectral-spatial prior-guided fusion module. The functional expression for the output feature is: , In the above formula, For the nth level spectral-spatial prior-guided fusion module The output characteristics, This represents the nth-order spectral-spatial prior-guided fusion module. , For the nth level spectral-spatial prior-guided fusion module The input features, and and These represent the embedded low-resolution multispectral images. and embedded high-resolution panchromatic images The extracted spectral and spatial prior experiences are: , , In the above formula, Indicates a stacking operation. Indicates the upsampling module of the nth channel. The output characteristics, Indicates the upsampling module of the nth channel. , For the N-n+1th stage The module's output, ; S202, the final high-resolution multispectral image is reconstructed according to the following formula: , In the above formula, For the final high-resolution multispectral image, For convolution operations, Indicates a stacking operation. For the initial embedding of low-resolution multispectral images, This is a high-resolution panchromatic image that is being embedded for the first time.
2. The high-resolution multispectral video imaging method according to claim 1, characterized in that, The spectral difference analysis module Includes a first linear normalization layer (LN) connected in sequence, and a spectral low-rank cross-attention module. The system comprises a first overlay module, a second linear normalization layer (LN), a feedforward network (FFN), and a second overlay module, wherein the output of the first linear normalization layer (LN) also serves as the input to the first overlay module, and the output of the functional layer also serves as the input to the second overlay module; the spectral low-rank cross-attention module. The function expression for the functional layer is: , , In the above formula, For spectral low-rank cross-attention modules Output characteristics Indicates a stacking operation. For the number of attention heads, For the result of cross-attention of the j-th attention head, For learnable parameters, For location embedding functions, location embedding functions It consists of two 3×3 convolutional layers and a GELU activation layer; Let j be the value of the attention head. Let be the attention score of the j-th attention head, and we have: , In the above formula, The softmax activation function is used. The key is obtained by linearly mapping the input features. The query is obtained by linearly mapping the input features; the value obtained by linearly mapping the input features is... ,in For the high-resolution panchromatic image embedded in the previous stage, For low-resolution multispectral images The image obtained by changing the size, The number of stages. , and These are learnable parameters.
3. The high-resolution multispectral video imaging method according to claim 1, characterized in that, The spatial difference analysis module Includes a first linear normalization layer (LN) connected in sequence, and a spatially low-rank cross-attention module. The system comprises a first overlay module, a second linear normalization layer (LN), a feedforward network (FFN), and a second overlay module, wherein the output of the first linear normalization layer (LN) also serves as the input to the first overlay module, and the output of the functional layer also serves as the input to the second overlay module; the spatial low-rank cross-attention module. The function expression is: , In the above formula, For spatial low-rank cross-attention modules Output characteristics Indicates a stacking operation. For the i-th channel and Features of spatial cross-attention output for and Features of spatial cross-attention output Spatial difference analysis module for the current stage The output results, For location embedding functions, location embedding functions It consists of two 3×3 convolutional layers and a GELU activation layer. Represents the first of N encoding stages Each coding stage This represents the number of spectral bands in the image; and we have: In the above formula, for and Output of spatial cross attention for and Output of spatial cross attention and These are learnable parameters. From and The attention score obtained for the j-th head. for and The attention score obtained for the j-th head. This is a subspace of features extracted from low-resolution multispectral images by the previous network. For the feature subspace of the embedded high-resolution panchromatic image, These are the spatial coefficients of the features extracted from the low-resolution multispectral image by the previous network. Let be the spatial coefficients of the features of the embedded high-resolution panchromatic image, where: , , , , In the above formula, This refers to the feature extracted from the i-th channel of the low-resolution multispectral image by the previous network. For the features of the i-th channel of the embedded high-resolution panchromatic image, , , These are learnable parameters.
4. The high-resolution multispectral video imaging method according to claim 1, characterized in that, In step S103, the features and characteristics The initial fusion features are obtained through fusion. The function expression is: , In the above formula, This represents the bottleneck structure module. This indicates a stacking operation.
5. The high-resolution multispectral video imaging method according to claim 4, characterized in that, The bottleneck structure module consists of a convolutional module with a kernel size of 1×1 and a spectral-spatial prior-guided fusion module. Composition, the spectral-spatial prior-guided fusion module It includes a first linear normalization layer (LN) connected in sequence, and a spectral-spatial prior-guided self-attention module. The system comprises a first stacking module, a second linear normalization layer (LN), a feedforward network (FFN), and a second stacking module, and the spectral-spatial prior-guided self-attention module. The function expression is: , In the above formula, Self-attention module guided by spectral-spatial priors The output characteristics, Indicates a stacking operation. For the number of attention heads, For the result of cross-attention of the j-th attention head, For learnable parameters, For location embedding functions, location embedding functions It consists of two 3×3 convolutional layers and a GELU activation layer; For spatial spectrum prior experience, and we have: , In the above formula, For the fusion module guided by the nth-level spectral-spatial prior using an attention mechanism Input features The extracted value.
6. A high-resolution multispectral video imaging device, comprising a microprocessor and a memory interconnected, characterized in that, The microprocessor is programmed or configured to perform the high-resolution multispectral video imaging method according to any one of claims 1 to 5.
7. The high-resolution multispectral video imaging device according to claim 6, characterized in that, The microprocessor is also connected to an optical and sensing module, which includes an objective lens, a beam splitter, an eyepiece, a panchromatic imaging sensor, and a multispectral coding sensor. The eyepiece includes a first eyepiece and a second eyepiece. External light enters through the objective lens and is split into two beams by the beam splitter. One beam passes through the first eyepiece and enters the panchromatic imaging sensor, while the other beam passes through the second eyepiece and enters the multispectral coding sensor. The panchromatic imaging sensor and the multispectral coding sensor are respectively connected to the microprocessor.
8. A computer-readable storage medium storing a computer program or instructions, characterized in that, The computer program or instructions are programmed or configured to execute the high-resolution multispectral video imaging method of any one of claims 1 to 5 via a processor.
9. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instructions are programmed or configured to execute the high-resolution multispectral video imaging method of any one of claims 1 to 5 via a processor.
Citation Information
Patent Citations
Image fusion method and device, storage medium and electronic equipment
CN117456323A
Stream-based dual-condition guided remote sensing image panchromatic sharpening method, system and equipment
CN118229579A