Spectral Reconstruction Method, Apparatus, Electronic Device, and Storage Medium
Through the combined structure of dense residual attention blocks and multi-scale residual attention blocks, combined with adaptive multi-scale blocks, the problems of low spectral reconstruction efficiency and accuracy are solved, and efficient and accurate multi-spectral image reconstruction is achieved.
Patent Information
- Application Number
- CN202210444253.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-04-25
AI Technical Summary
The existing spectral reconstruction technology has low efficiency and poor accuracy, making it difficult to meet the practical application needs in many fields.
The combined structure of dense residual attention block DRAB and multi-scale residual attention block MRAB is used for feature matching and fusion, and the spectrum reconstruction is combined with adaptive multi-scale block AMB. The multi-spectral image is reconstructed through shallow feature extraction, feature matching and feature fusion.
It improves the accuracy and efficiency of spectral reconstruction, can better reconstruct the multispectral image of the target image, and enhances the learning ability and adaptability of the network.
Smart Images

Figure CN117011534B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data transmission, and particularly relates to a spectral reconstruction method, apparatus, electronic device, and storage medium. Background Art
[0002] As the fingerprint of an object, the spectrum can not only accurately characterize colors, but also be used for material characterization, and has been widely applied in industries such as textile, garment manufacturing, pigments, inks, printing, imaging, cultural heritage protection, medical diagnosis, remote sensing, food quality inspection, etc. In view of the actual application requirements of the spectrum in multiple fields, spectral reconstruction technology provides an effective way to obtain the spectral information of an object. However, in related technologies, the efficiency of spectral reconstruction technology is low and the accuracy is poor. Summary of the Invention
[0003] To overcome the problems existing in related technologies, embodiments of the present disclosure provide a spectral reconstruction method, apparatus, electronic device, and storage medium to solve the defects in related technologies.
[0004] According to a first aspect of embodiments of the present disclosure, a spectral reconstruction method is provided, including:
[0005] Performing shallow feature extraction processing on a target image to obtain shallow features of the target image;
[0006] Inputting the shallow features into a feature matching module for feature matching to obtain a feature matching result, where the feature matching module includes a plurality of densely connected residual attention blocks DRAB connected in sequence, the DRAB includes a plurality of multi-scale residual attention blocks MRAB connected in sequence, the input and output of each MRAB except the last MRAB are cascaded and used as the input of the next MRAB, and the sum of the result obtained by cascading the input and output of the last MRAB and the input of the first MRAB is used as the output of the DRAB;
[0007] Performing feature fusion processing on the feature matching result to obtain a feature fusion result;
[0008] Reconstructing a multi-spectral image corresponding to the target image according to the feature fusion result.
[0009] In one embodiment, within the DRAB, the input of each MRAB except the first MRAB is input to the corresponding MRAB after channel compression; and / or,
[0010] Within the DRAB, the sum of the result obtained by cascading the input and output of the last MRAB and the input of the first MRAB after channel compression is used as the output of the DRAB.
[0011] In one embodiment, the MRAB includes a channel attention block CA, a double downsampling spatial attention block DDSA, and three multi-scale residual blocks MRB. Among them, the input of the first MRB serves as the input of the second MRB, the CA, and the DDSA respectively. The concatenation result of the product of the output of the second MRB and the output of the CA, and the sum of the output of the second MRB and the output of the DDSA serves as the input of the third MRB. The output of the third MRB serves as the output of the MRAB.
[0012] In one embodiment, the input of the third MRB is input to the third MRB after channel compression.
[0013] In one embodiment, the MRB includes multiple parallel paths, an activation layer, and at least one convolutional layer. Among them, each path includes a different number of convolutional layers. The result obtained by concatenating the outputs of the multiple paths serves as the input of the activation layer. The output of the activation layer serves as the input of the at least one convolutional layer. The at least one convolutional layer is used for feature fusion and channel compression. The sum of the output of the at least one convolutional layer and the input of the MRB serves as the output of the MRB.
[0014] In one embodiment, the CA includes a pooling layer, two convolutional layers, an activation layer, and a normalization layer. Among them, the output of the pooling layer serves as the input of the first convolutional layer. The first convolutional layer is used for channel compression. The output of the first convolutional layer serves as the input of the activation layer. The output of the activation layer serves as the input of the second convolutional layer. The second convolutional layer is used for channel expansion. The output of the second convolutional layer serves as the input of the normalization layer. The output of the normalization layer serves as the output of the CA.
[0015] In one embodiment, the DDSA includes multiple downsampling paths and a convolutional layer. Among them, at least one path in the multiple downsampling paths performs downsampling in both channel and spatial dimensions simultaneously. The result obtained by multiplying the outputs of the multiple downsampling paths in a certain order serves as the input of the convolutional layer. The convolutional layer is used for channel dimension expansion. The output of the convolutional layer serves as the output of the DDSA.
[0016] In one embodiment, the DDSA further includes a normalization layer, and the multiple downsampling paths include a first downsampling path, a second downsampling path, and a third downsampling path. Among them, the first downsampling path is used for downsampling in the channel dimension, and the second downsampling path and the third downsampling path are used for downsampling in both the channel and spatial dimensions. The product of the output of the first downsampling path and the output of the second downsampling path serves as the input to the normalization layer, and the product of the output of the normalization layer and the output of the third downsampling path serves as the input to the convolutional layer.
[0017] In one embodiment, the feature matching module further includes a plurality of densely connected residual blocks DRB connected in sequence, and the output of the last DRB serves as the input to the first DRAB. Among them, the DRB includes a plurality of multi-scale residual blocks MRB connected in sequence.
[0018] In one embodiment, the performing feature fusion processing on the feature matching result to obtain a feature fusion result includes:
[0019] Cascading the output of each DRAB and the output of each DRB and then performing channel compression to obtain a channel compression result with the same number of channels as the shallow features;
[0020] Determining the sum of the channel compression result and the shallow features as the feature fusion result.
[0021] In one embodiment, the reconstructing the multi-spectral image corresponding to the target image according to the feature fusion result includes:
[0022] Inputting the feature fusion result into an adaptive multi-scale block AMB after being processed by an activation function, and performing channel compression on the output of the AMB to obtain the multi-spectral image corresponding to the target image. Among them, the AMB includes a plurality of feature processing paths, each feature processing path includes a convolutional layer, an activation layer, and a convolutional layer connected in sequence. The convolutional kernel sizes of the two convolutional layers in the same feature processing path are the same, and the convolutional kernel sizes of the convolutional layers in different feature processing paths are different. Each feature processing path has a corresponding weight, and the outputs of the multiple feature processing paths are weighted and summed according to the corresponding weights to obtain the output of the AMB.
[0023] According to a second aspect of the embodiments of the present disclosure, there is provided a spectral reconstruction apparatus, including:
[0024] An extraction module, configured to perform shallow feature extraction processing on a target image to obtain shallow features of the target image;
[0025] A matching module, configured to input the shallow features into a feature matching module for feature matching to obtain a feature matching result. The feature matching module includes a plurality of densely connected residual attention blocks (DRABs) connected in sequence. Each DRAB includes a plurality of multi-scale residual attention blocks (MRABs) connected in sequence. The input and output of each MRAB except the last one are cascaded and used as the input of the next MRAB. The result obtained by cascading the input and output of the last MRAB, and the sum of the input of the first MRAB are used as the output of the DRAB.
[0026] A fusion module, configured to perform feature fusion processing on the feature matching result to obtain a feature fusion result.
[0027] A reconstruction module, configured to reconstruct the multi-spectral image corresponding to the target image according to the feature fusion result.
[0028] In one embodiment, within the DRAB, the input of each MRAB except the first one is input to the corresponding MRAB after channel compression; and / or,
[0029] Within the DRAB, the result obtained by cascading the input and output of the last MRAB is compressed in channels, and the sum of the result and the input of the first MRAB is used as the output of the DRAB.
[0030] In one embodiment, the MRAB includes a pyramid channel attention block (CA), a double down-sampling spatial attention block (DDSA), and three multi-scale residual blocks (MRBs). The input of the first MRB is respectively used as the input of the second MRB, the CA, and the DDSA. The product of the output of the second MRB and the output of the CA, and the concatenation result of the sum of the output of the second MRB and the output of the DDSA are used as the input of the third MRB. The output of the third MRB is used as the output of the MRAB.
[0031] In one embodiment, the input of the third MRB is input to the third MRB after channel compression.
[0032] In one embodiment, the MRB includes a plurality of parallel paths, an activation layer, and at least one convolutional layer. Each path includes a different number of convolutional layers. The result obtained by cascading the outputs of the plurality of paths is used as the input of the activation layer. The output of the activation layer is used as the input of the at least one convolutional layer. The at least one convolutional layer is used for feature fusion and channel compression. The sum of the output of the at least one convolutional layer and the input of the MRB is used as the output of the MRB.
[0033] In one embodiment, the CA includes a pooling layer, two convolutional layers, an activation layer, and a normalization layer. Among them, the output of the pooling layer serves as the input of the first convolutional layer, the first convolutional layer is used for channel compression, the output of the first convolutional layer serves as the input of the activation layer, the output of the activation layer serves as the input of the second convolutional layer, the second convolutional layer is used for channel expansion, the output of the second convolutional layer serves as the input of the normalization layer, and the output of the normalization layer serves as the output of the CA.
[0034] In one embodiment, the DDSA includes a plurality of downsampling paths and a convolutional layer. Among them, at least one of the plurality of downsampling paths performs downsampling in both channels and space simultaneously. The result obtained by multiplying the outputs of the plurality of downsampling paths in a certain order serves as the input of the convolutional layer, the convolutional layer is used for channel dimension expansion, and the output of the convolutional layer serves as the output of the DDSA.
[0035] In one embodiment, the DDSA further includes a normalization layer. The plurality of downsampling paths include a first downsampling path, a second downsampling path, and a third downsampling path. Among them, the first downsampling path is used for downsampling in the channel dimension, the second downsampling path and the third downsampling path are used for downsampling in both channels and space. The product of the output of the first downsampling path and the output of the second downsampling path serves as the input of the normalization layer, and the product of the output of the normalization layer and the output of the third downsampling path serves as the input of the convolutional layer.
[0036] In one embodiment, the feature matching module further includes a plurality of densely connected residual blocks DRB connected in sequence. The output of the last DRB serves as the input of the first DRAB. Among them, the DRB includes a plurality of multi-scale residual blocks MRB connected in sequence.
[0037] In one embodiment, the fusion module is specifically configured to:
[0038] Concatenate the outputs of each DRAB and the outputs of each DRB and then perform channel compression to obtain a channel compression result with the same number of channels as the shallow features;
[0039] Determine the sum of the channel compression result and the shallow features as the feature fusion result.
[0040] In one embodiment, the reconstruction module is specifically configured to:
[0041] The result of the feature fusion is processed by an activation function and then input into an Adaptive Multi-Scale Block (AMB). The output of the AMB is subjected to channel compression to obtain the multi-spectral image corresponding to the target image. The AMB includes multiple feature processing paths. Each feature processing path includes a convolutional layer, an activation layer, and a convolutional layer connected in sequence. The convolutional kernel sizes of the two convolutional layers in the same feature processing path are the same, and the convolutional kernel sizes of the convolutional layers in different feature processing paths are different. Each feature processing path has a corresponding weight. The outputs of the multiple feature processing paths are weighted and summed according to the corresponding weights to obtain the output of the AMB.
[0042] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided. The electronic device includes a memory and a processor. The memory is used to store computer instructions that can be run on the processor, and the processor is used to, when executing the computer instructions, perform the spectral reconstruction method according to the first aspect.
[0043] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect is implemented.
[0044] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0045] The present disclosure first performs shallow feature extraction processing on a target image to obtain the shallow features of the target image, then inputs the shallow features into a feature matching module for feature matching to obtain a feature matching result, then performs feature fusion processing on the feature matching result to obtain a feature fusion result, and finally reconstructs the multi-spectral image corresponding to the target image according to the feature fusion result. Since the feature matching module includes multiple Dense Residual Attention Blocks (DRABs) connected in sequence, and each DRAB includes multiple Multi-Scale Residual Attention Blocks (MRABs) connected in sequence, the input and output of each MRAB except the last one are cascaded and used as the input of the next MRAB. The result obtained by cascading the input and output of the last MRAB and the input of the first MRAB is used as the output of the DRAB. Therefore, different levels of features can be reused without affecting the depth of the network, thereby improving the accuracy and efficiency of spectral reconstruction. Description of the Drawings
[0046] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments that conform to the present invention, and are used together with the specification to explain the principles of the present invention.
[0047] Figure 1It is a flowchart of a spectral reconstruction method shown in an exemplary embodiment of the present disclosure;
[0048] Figure 2 It is a network structure diagram of a dense residual attention block DRAB shown in an exemplary embodiment of the present disclosure;
[0049] Figure 3 It is a network structure diagram of an adaptive multi-scale block AMB shown in an exemplary embodiment of the present disclosure;
[0050] Figure 4 It is a network structure diagram of a multi-scale residual attention block MRAB shown in an exemplary embodiment of the present disclosure;
[0051] Figure 5 It is a network structure diagram of a multi-scale residual block MRB shown in an exemplary embodiment of the present disclosure;
[0052] Figure 6 It is a network structure diagram of a channel attention block CA shown in an exemplary embodiment of the present disclosure;
[0053] Figure 7 It is a network structure diagram of a double downsampling spatial attention block DDSA shown in an exemplary embodiment of the present disclosure;
[0054] Figure 8 It is a network structure diagram of a double attention dense residual network DRN-DA shown in an exemplary embodiment of the present disclosure;
[0055] Figure 9 It is a structural schematic diagram of a spectral reconstruction device shown in an exemplary embodiment of the present disclosure;
[0056] Figure 10 It is a structural block diagram of an electronic device shown in an exemplary embodiment of the present disclosure. Detailed implementation manners
[0057] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0058] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms "a", "the", and "said" used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0059] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0060] In a first aspect, at least one embodiment of the present disclosure provides a spectral reconstruction method. Please refer to the attached Figure 1 , which shows the flow of the method, including step S101 and step S104.
[0061] Among them, this method is used to perform spectral reconstruction on a target image, that is, to reconstruct a multi-spectral image for the target image. Exemplarily, this method can be executed using a pre-trained neural network. In the case of multiple target images, this method can be used for batch processing to reconstruct corresponding multi-spectral images for each target image.
[0062] In step S101, perform shallow feature extraction processing on the target image to obtain the shallow features of the target image.
[0063] This step can be executed using a feature extraction module, that is, input the target image into the feature extraction module for feature extraction processing to obtain the shallow features of the target image. Exemplarily, the feature extraction module can include two 3*3 convolutional layers to extract shallow features F0, and at the same time insert a PReLU (Parametric Rectified Linear Unit) activation function between them to introduce more non-linear features and improve the generalization ability of the network.
[0064] In step S102, input the shallow features into a feature matching module for feature matching to obtain a feature matching result.
[0065] Among them, the feature matching module includes a plurality of (for example, 2) densely residual attention blocks DRAB (Densely Residual Attention Block) connected in sequence. The DRAB includes a plurality of multi-scale residual attention blocks MRAB (Multi-scale Residual Attention Block) connected in sequence. The input and output of each MRAB except the last one are cascaded and used as the input of the next MRAB. The sum of the result obtained by cascading the input and output of the last MRAB and the input of the first MRAB is used as the output of the DRAB. That is to say, lightweight dense residual connections are adopted among the multiple MRABs in the DRAB. The output of the previous MRAB is only cascaded with the output of the next adjacent MRAB, rather than cascaded with the output of each subsequent MRAB as in the related art, thereby reducing the network width and increasing the network depth.
[0066] Further, within the DRAB, the input of each MRAB except the first one is input to the corresponding MRAB after channel compression; and / or, within the DRAB, the sum of the result obtained by cascading the input and output of the last MRAB and the input of the first MRAB after channel compression is used as the output of the DRAB.
[0067] It can be understood that a cascading layer can be set at the corresponding position within the feature matching module for cascading, that is, splicing on the channels. A 3*3 convolutional layer can be set at the corresponding position within the feature matching module for channel compression.
[0068] In a possible embodiment, the network structure of the DRAB is as shown in the appendix Figure 2 shown. From the appendix Figure 2It can be seen that the DRAB includes 3 MRABs, and one cascade layer 201 is provided after each MRAB, and one 3*3 convolutional layer 203 is provided after each cascade layer 201. Among them, the input of the DRAB serves as the input of the first MRAB, and the input of the first MRAB is directly input into the first MRAB. Then, the first cascade layer 201 cascades the input and output of the first MRAB and uses the result as the input of the second MRAB. The input of the second MRAB needs to be channel-compressed by the first 3*3 convolutional layer 202 and then input into the second MRAB. Then, the second cascade layer 201 cascades the input and output of the second MRAB and uses the result as the input of the third MRAB. The input of the third MRAB needs to be channel-compressed by the second 3*3 convolutional layer 202 and then input into the third MRAB. Then, the second cascade layer 201 cascades the input and output of the second MRAB, and after channel compression by the third 3*3 convolutional layer 202, the result is summed with the input of the DRAB to obtain the output of the DRAB.
[0069] In this embodiment, the three MRABs in the DRAB and subsequent cascading and compression can be respectively expressed as:
[0070] x1 = Conv([f MRAB (x0), x0])
[0071] x2 = Conv([f MRAB (x1), f MRAB (x0), x0])
[0072] x3 = Conv([f MRAB (x2), f MRAB (x1), f MRAB (x0), x0])
[0073] Among them, X0 is the input of the first MRAB, X1 is the result after channel compression of the input of the second MRAB, X2 is the result after channel compression of the input of the third MRAB, and X3 is the result obtained by cascading the input and output of the third MRAB, the result after channel compression; f MRAB (·) and Conv(·) are respectively the functions of the MRAB and the 3*3 convolutional layer 202; [·] represents the function of the cascade layer 201, that is, the splicing of features on the channel.
[0074] In addition, the feature matching module may further include multiple (for example, 2) densely connected residual blocks DRB connected in sequence. The input of the last DRB serves as the input of the first DRAB. Among them, the DRB includes multiple multi-scale residual blocks MRB connected in sequence.
[0075] In step S103, perform feature fusion processing on the feature matching result to obtain a feature fusion result.
[0076] In this step, when the feature matching module includes multiple DRABs, the outputs of each DRAB can be cascaded first and then channel compression is performed to obtain a channel compression result with the same number of channels as the shallow feature; then, the sum of the channel compression result and the result of the convolution process on the shallow feature is determined as the feature fusion result. When the feature matching module includes multiple DRBs and multiple DRABs, the outputs of each DRAB and the outputs of each DRB can be cascaded first and then channel compression is performed to obtain a channel compression result with the same number of channels as the shallow feature; then, the sum of the channel compression result and the result of the convolution process on the shallow feature is determined as the feature fusion result.
[0077] This step adopts an adaptive fusion strategy to adaptively reuse the intermediate layer features to improve the learning ability of the network. Exemplarily, an Adaptive Fusion (AF) module can be set up to execute this step. For example, the adaptive fusion module includes a concatenation layer, a 3*3 convolution layer, and a 5*5 convolution layer. Among them, the concatenation layer is used to concatenate the outputs of each DRAB and the outputs of each DRB, that is, after adaptively weighting the features of each layer, they are added on the channels; the 3*3 convolution layer is used to perform channel compression on the concatenation result of the above concatenation layer to reduce the dimension and restore it to the original number of channels; the 5*5 convolution layer is used to perform convolution processing on the shallow features. Using the adaptive fusion module to execute this step can reduce data redundancy and improve the accuracy of the fusion result.
[0078] In step S104, reconstruct the multi-spectral image corresponding to the target image according to the feature fusion result.
[0079] In this step, an activation layer (such as a PReLU activation function layer) and an Adaptive Multi-Scale Block (AMB) can be preset in advance. Then, the feature fusion result is processed by the activation layer and input into the AMB, and the output of the AMB is channel-compressed to obtain the multi-spectral image corresponding to the target image. Among them, the AMB includes multiple feature processing paths. Each feature processing path includes a convolutional layer, an activation layer, and a convolutional layer connected in sequence. The convolutional kernel sizes of the two convolutional layers in the same feature processing path are the same, and the convolutional kernel sizes of the convolutional layers in different feature processing paths are different. Each feature processing path has a corresponding weight. After the outputs of the multiple feature processing paths are weighted and summed according to the corresponding weights, the output of the AMB is obtained. The AMB module processes the feature fusion result using multiple different receptive fields and performs adaptive weighted summation through three separate trainable weights, learning features with a larger adaptive receptive field, so as to reconstruct the multi-spectral image accordingly.
[0080] Exemplarily, the network structure of the AMB is as shown in the appendix Figure 3 Please refer to the appendix Figure 3 . The AMB includes three feature processing paths. The first feature processing path consists of two 3×3 convolutional layers (Conv3×3) and a PReLU activation layer inserted between the two 3×3 convolutional layers. The second feature processing path consists of two 5×5 convolutional layers (Conv5×5) and a PReLU activation layer inserted between the two 5×5 convolutional layers. The third feature processing path consists of two 7×7 convolutional layers (Conv7×7) and a PReLU activation layer inserted between the two 7×7 convolutional layers. Based on this, the processing process of the AMB can be expressed by the following formula:
[0081] h = w1×Conv 3×3 (δ(Conv 3×3 (h0)))+w2×Conv 5×5 (δ(Conv 5×5 (h0)))+w3×Conv 7×7 (δ(Conv 7×7 (h0)))
[0082] Among them, h0 and h are the input and output of the AMB respectively; Conv 3×3 (·), Conv 5×5 (·), Conv 7×7 (·) are 3×3, 5×5, and 7×7 convolutional layers respectively; δ(·) represents the PReLU activation function; w1, w2, and w3 are the weights of the three feature processing paths respectively.
[0083] The present disclosure first performs shallow feature extraction on a target image to obtain the shallow features of the target image, then inputs the shallow features into a feature matching module for feature matching to obtain a feature matching result, then performs feature fusion processing on the feature matching result to obtain a feature fusion result, and finally reconstructs the multi-spectral image corresponding to the target image according to the feature fusion result. Since the feature matching module includes a plurality of densely connected residual attention blocks DRAB connected in sequence, and each DRAB includes a plurality of multi-scale residual attention blocks MRAB connected in sequence, the input and output of each MRAB except the last one are cascaded and used as the input of the next MRAB, and the result obtained by cascading the input and output of the last MRAB and the sum of the input of the first MRAB are used as the output of the DRAB. Therefore, different levels of features can be reused without affecting the depth of the network, thereby improving the accuracy and efficiency of spectral reconstruction.
[0084] In some embodiments of the present disclosure, the MRAB includes a channel attention block CA, a double downsampling spatial attention block DDSA, and three multi-scale residual blocks MRB. Among them, the input of the first MRB is used as the input of the second MRB, the CA, and the DDSA respectively. The product of the output of the second MRB and the output of the CA, and the concatenated result of the sum of the output of the second MRB and the output of the DDSA are used as the input of the third MRB, and the output of the third MRB is used as the output of the MRAB.
[0085] Reference can be made to the attached Figure 4 which shows the above network structure of the MRAB, and it is provided with a concatenation layer 401 for concatenating the product of the output of the second MRB and the output of the CA, and the output of the second MRB and the output of the DDSA; in addition, it is also provided with a convolutional layer 402 for compressing the channels of the input of the third MRB and then inputting it into the third MRB. Based on this, the process of the MRAB processing features can be expressed by the following formula:
[0086] y1 = f MRB (y0)
[0087] y2 = f MRB (y1)
[0088]
[0089] In the formula, y0, y1, y2, and y3 are the input of the MARB and the outputs of the first, second, and third MRBs respectively; f MRB (·), f CA (·), f DDSA(·) and Conv(·) are functions of the MRB, CA, DDSA, and 3*3 convolutional layer respectively; They represent element-wise multiplication and element-wise addition respectively; [·] represents concatenation, that is, the concatenation of features on the channels.
[0090] In a possible embodiment, the MRB includes multiple parallel paths, an activation layer, and at least one convolutional layer. Among them, each path includes a different number of convolutional layers (and the convolutional kernel sizes of each convolutional layer involved in multiple first convolutional channels are equal, for example, 3*3). The input of the MRB is used as the input of each path respectively. The result obtained by concatenating the outputs of multiple paths is used as the input of the activation layer. The output of the activation layer is used as the input of the at least one convolutional layer. The at least one convolutional layer is used for feature fusion and channel compression. The sum of the output of the at least one convolutional layer and the input of the MRB is used as the output of the MRB.
[0091] Please refer to the appendix Figure 5 , which exemplarily shows the network structure of the above MRB. It can be seen from the appendix Figure 5 that the MRB is provided with 3 paths, a PReLU activation layer, a 1*1 convolutional layer (Conv1*1), and a 3*3 convolutional layer (Conv3*3). Among them, the first path consists of 1 3*3 convolutional layer (Conv3*3), the second path consists of 2 3*3 convolutional layers (Conv3*3), and the third path consists of 3 3*3 convolutional layers (Conv3*3). The sum of the input of the MRB and the output of the last 3*3 convolutional layer is used as the output of the MRB; The MRB is also provided with a concatenation layer connected to each path respectively, which is used to concatenate each path and input the concatenation result into the PReLU activation layer.
[0092] Different numbers of convolutional layers can achieve different receptive field effects. For example, the above 2 3*3 convolutional layers can achieve the receptive field effect of a 5*5 convolutional layer, but the number of parameters is much less than that of a 5*5 convolutional layer. The above 3 3*3 convolutional layers can achieve the receptive field effect of a 7*7 convolutional layer, but the number of parameters is much less than that of a 7*7 convolutional layer. Multiple paths of the MRB can replace the 3*3 convolutional layer in the traditional residual layer to extract multi-level features at different scales to achieve the purpose of improving the feature extraction ability, and further promote information transmission through local residual learning.
[0093] In a possible embodiment, the CA includes a pooling layer, two convolutional layers, an activation layer, and a normalization layer. Among them, the output of the pooling layer serves as the input to the first convolutional layer, the first convolutional layer is used for channel compression, the output of the first convolutional layer serves as the input to the activation layer, the output of the activation layer serves as the input to the second convolutional layer, the second convolutional layer is used for channel expansion, the output of the second convolutional layer serves as the input to the normalization layer, and the output of the normalization layer serves as the output of the CA.
[0094] Please refer to the appendix Figure 6 , which exemplarily shows the network structure of the above CA. As can be seen from the appendix Figure 6 , the CA includes a pooling layer (Global pooling), which can transform the input feature u of size C×H×W into the channel global statistic value q of C×1×1, that is:
[0095]
[0096] In the formula, u c (i,j) is the value of the c-th channel feature of u at the position (i,j), and q c is the value of the c-th channel after average pooling.
[0097] As can be seen from the appendix Figure 6 , the first convolutional layer is a 1*1 convolutional layer (Conv1*1), which can process the global statistic value q into a feature of (C / r)*1*1. Then, this feature passes through the PReLU activation layer.
[0098] As can be seen from the appendix Figure 6 , the activation layer of the CA is a PReLU activation layer, which can increase the nonlinearity.
[0099] As can be seen from the appendix Figure 6 , the second convolutional layer is a 1*1 convolutional layer (Conv1*1), which can expand the dimension of the (C / r)*1*1 feature output by the activation layer into a C*1*1 feature.
[0100] As can be seen from the appendix Figure 6 , the normalization layer in the CA is a Sigmoid function layer, and the output of the Sigmoid function layer is the output v of the CA, that is:
[0101] v = σ(Conv2(δ(Conv1(q))))
[0102] In the formula, Conv1(·) and Conv2(·) are the downsampling convolution of the first 1*1 convolutional layer and the upsampling convolution of the second convolutional layer respectively, δ(·) represents the PReLU activation function, and σ(·) represents the Sigmoid function.
[0103] In a possible embodiment, the DDSA includes a plurality of downsampling paths and a convolutional layer. Among them, at least one of the plurality of downsampling paths is used for simultaneous downsampling in both the channel and spatial dimensions. The result obtained by sequentially multiplying the outputs of the plurality of downsampling paths in a certain order is used as the input of the convolutional layer. The convolutional layer is used for expanding the channel dimension, and the output of the convolutional layer is used as the output of the DDSA. Exemplarily, the DDSA further includes a normalization layer. The plurality of downsampling paths include a first downsampling path, a second downsampling path, and a third downsampling path. Among them, the first downsampling path is used for downsampling in the channel dimension, and the second downsampling path and the third downsampling path are used for downsampling in both the channel and spatial dimensions. The product of the output of the first downsampling path and the output of the second downsampling path is used as the input of the normalization layer, and the product of the output of the normalization layer and the output of the third downsampling path is used as the input of the convolutional layer.
[0104] It can be understood that before multiplying the output of the first downsampling path and the output of the second downsampling path, the output of the first downsampling path can be reshaped to transform it from a three-dimensional tensor into a two-dimensional matrix, and then transposed. At the same time, the output of the second downsampling path can be reshaped to transform it from a three-dimensional tensor into a two-dimensional matrix, and then the two are multiplied. Before multiplying the output of the normalization layer and the output of the third downsampling path, the output of the normalization layer can be transposed, and then the two are multiplied. At the same time, before inputting the product of the output of the normalization layer and the output of the third downsampling path into the convolutional layer for expanding the channel dimension, this product can be reshaped to transform it from a two-dimensional matrix into a three-dimensional tensor.
[0105] Please refer to the attached Figure 7 , which exemplarily shows the network structure of the DDSA. As can be seen from the attached Figure 7 , the first downsampling path, the second downsampling path, and the third downsampling path are each composed of a 1*1 convolutional layer (Conv1*1). Suppose the input feature map g0 ∈ R C×H×W . It is respectively input into three downsampling paths composed of 1*1 convolutional layers. The first downsampling path performs channel downsampling, and the second downsampling path and the third downsampling path simultaneously achieve dual downsampling in both the channel and spatial dimensions. The downsampling ratios in the channel and spatial dimensions are s (for example, 8) and t (for example, 4) respectively. Three feature maps g1 ∈ R (C / s)×H×W , g2 ∈ R (C / s)×(H / t)×(W / t) , g3 ∈ R (C / s)×(H / t)×(W / t) are obtained according to the following formula:
[0106] g1 = Conv1(g0)
[0107] g2 = Conv2(g0)
[0108] g3 = Conv3(g0)
[0109] Among them, Conv1(·) is the sampling function of the first downsampling path, Conv2(·) is the sampling function of the second downsampling path, and Conv3(·) is the sampling function of the third downsampling path.
[0110] From Figure 7 it can be seen that the normalization layer of the DDSA is a Softmax function layer. The output g1 of the first downsampling path undergoes Reshape and transpose (Transpose), and after the output g2 of the second downsampling path undergoes Reshape, the spatial attention matrix can be obtained through matrix multiplication, and then the spatial attention matrix is input into the Softmax function layer for normalization processing to obtain the normalized spatial attention matrix B expressed by the following formula:
[0111]
[0112] In the formula, τ(·) represents the Softmax activation function.
[0113] From Figure 7 it can be seen that by multiplying the output g3 of the third downsampling path with the transposed spatial attention matrix B and performing the Reshape operation to restore the feature map to its original size, g4 expressed by the following formula is obtained:
[0114] g4 = g3 × B T
[0115] From Figure 7 it can be seen that the convolutional layer for expanding the channel dimension is a 1*1 convolutional layer (Conv1*1), and the above g4 passes through this 1*1 convolutional layer to expand the channel dimension to the original dimension.
[0116] In this embodiment, based on the classical non-local self-attention, a double downsampling strategy is adopted to form a double downsampling non-local attention mechanism, which greatly reduces the computational cost, can extract spatial attention information at the same time, and can also utilize the dependence between long-distance pixels to extract significant features.
[0117] In some embodiments of the present disclosure, it is pre-constructed and trained as Figure 8The shown Dual Attention Dense Residual Network DRN-DA (Densely Residual Network with Dual Attention) is then used to execute this method, that is, the target image is input into the DRN-DA, and the DRN-DA outputs the multi-spectral image corresponding to the target image. The image input into the network is I RGB ∈R N×3×H×W , and the image output from the network is I MSI ∈R N×31×H×W , where H and W respectively represent the height and width of the image, and 3 and 31 are both the number of channels of the image.
[0118] It can be seen from Figure 8 that the DRN-DA includes a Shallow Feature Extraction module, a Feature Mapping module, an Adaptive Fusion module, and a Reconstruction module. The Shallow Feature Extraction module includes two 3×3 convolutional layers 801 and a PReLU activation layer 802 disposed therebetween. Among them, the two 3×3 convolutional layers 801 are used to extract shallow features F0, and the PReLU (Parametric Rectified Linear Unit) activation function can introduce more non-linear features and improve the generalization ability of the network. The Feature Mapping module includes two DRBs 803 and two DRABs 804, where the DRBs 803 and DRABs 804 can be the structures introduced in any of the above embodiments. The Adaptive Fusion module includes a concatenation layer 805, a 3×3 convolutional layer 801, and a 5×5 convolutional layer 806. Among them, the concatenation layer 805 is used to concatenate the outputs of the two DRBs 803 and the outputs of the two DRABs 804, the 3×3 convolutional layer 801 is used to compress the channels of the concatenation result of the concatenation layer 805, the 5×5 convolutional layer 806 is used to perform convolutional processing on the shallow features F0, and the sum of the output of the 3×3 convolutional layer 801 and the output of the 5×5 convolutional layer 806 is used as the output of this Adaptive Fusion module. The Reconstruction module includes a PReLU activation layer 802, an AMB, and a 3×3 convolutional layer 801. The PReLU activation layer 802 processes the output of the Adaptive Fusion module and then inputs it into the AMB. The output of the AMB is reconstructed into a multi-spectral image after channel compression by the 3×3 convolutional layer 801. Among them, the network structure of the AMB can be the network structure of the AMB introduced in any of the above embodiments. Additionally,
[0119] The hardware platform running DRN-DA can have an Intel E5 dual-core processor, a GPU with two NVIDIA RTX 2080s, and 32G of memory. PyTorch 1.7.0 is selected as the framework, and relevant processing is carried out with the Ubuntu 18.04.5 Linux operating system.
[0120] The size of the training images can be uniformly adjusted to 64×64, and during the testing process, the entire image is directly input into the trained network for testing. During training, the batch size is set to 32. The learning rate is initialized to 10 -4 , and the polynomial decay mode is adopted with a decay exponent of 1.5. After 100 epochs, the learning rate drops to 0. The default parameters of the Adam (Adaptive Moment Estimation) optimizer, β1 = 0.9, β2 = 0.999, and ε = 10 -8 are used to guide the gradient descent.
[0121] In this embodiment, DRN-DA is used to execute the spectral reconstruction method, and compared with various spectral reconstruction methods in the related art, both the efficiency and accuracy are significantly improved.
[0122] According to the second aspect of the embodiments of the present disclosure, a spectral reconstruction device is provided. Please refer to the attached Figure 9 , and the device includes:
[0123] An extraction module 901, configured to perform shallow feature extraction processing on a target image to obtain shallow features of the target image;
[0124] A matching module 902, configured to input the shallow features into a feature matching module for feature matching to obtain a feature matching result. The feature matching module includes a plurality of densely connected residual attention blocks DRAB connected in sequence. Each DRAB includes a plurality of multi-scale residual attention blocks MRAB connected in sequence. The input and output of each MRAB except the last one are cascaded and used as the input of the next MRAB. The result obtained by cascading the input and output of the last MRAB and the sum of the input of the first MRAB are used as the output of the DRAB;
[0125] A fusion module 903, configured to perform feature fusion processing on the feature matching result to obtain a feature fusion result;
[0126] A reconstruction module 904, configured to reconstruct the multi-spectral image corresponding to the target image according to the feature fusion result.
[0127] In some embodiments of the present disclosure, within the DRAB, the input of each MRAB except the first MRAB is input to the corresponding MRAB after channel compression; and / or,
[0128] Within the DRAB, the sum of the result obtained by cascading the input and output of the last MRAB and the input of the first MRAB after channel compression is used as the output of the DRAB.
[0129] In some embodiments of the present disclosure, the MRAB includes a channel attention block CA, a double downsampling spatial attention block DDSA, and three multi-scale residual blocks MRB. Among them, the input of the first MRB serves as the input of the second MRB, the CA, and the DDSA respectively. The product of the output of the second MRB and the output of the CA, and the concatenated result of the sum of the output of the second MRB and the output of the DDSA serve as the input of the third MRB, and the output of the third MRB serves as the output of the MRAB.
[0130] In some embodiments of the present disclosure, the input of the third MRB is input to the third MRB after channel compression.
[0131] In some embodiments of the present disclosure, the MRB includes a plurality of parallel paths, an activation layer, and at least one convolutional layer. Among them, each path includes a different number of convolutional layers. The result obtained by concatenating the outputs of the plurality of paths serves as the input of the activation layer. The output of the activation layer serves as the input of the at least one convolutional layer. The at least one convolutional layer is used for feature fusion and channel compression. The sum of the output of the at least one convolutional layer and the input of the MRB serves as the output of the MRB.
[0132] In some embodiments of the present disclosure, the CA includes a pooling layer, two convolutional layers, an activation layer, and a normalization layer. Among them, the output of the pooling layer serves as the input of the first convolutional layer. The first convolutional layer is used for channel compression. The output of the first convolutional layer serves as the input of the activation layer. The output of the activation layer serves as the input of the second convolutional layer. The second convolutional layer is used for channel expansion. The output of the second convolutional layer serves as the input of the normalization layer. The output of the normalization layer serves as the output of the CA.
[0133] In some embodiments of the present disclosure, the DDSA includes a plurality of downsampling paths and a convolutional layer. Among them, at least one of the plurality of downsampling paths performs downsampling in both the channel and spatial dimensions simultaneously. The result obtained by multiplying the outputs of the plurality of downsampling paths in a certain order is used as the input of the convolutional layer. The convolutional layer is used for channel dimension expansion, and the output of the convolutional layer is used as the output of the DDSA.
[0134] In some embodiments of the present disclosure, the DDSA further includes a normalization layer. The plurality of downsampling paths include a first downsampling path, a second downsampling path, and a third downsampling path. Among them, the first downsampling path is used for downsampling in the channel dimension, and the second downsampling path and the third downsampling path are used for downsampling in both the channel and spatial dimensions. The product of the output of the first downsampling path and the output of the second downsampling path is used as the input of the normalization layer, and the product of the output of the normalization layer and the output of the third downsampling path is used as the input of the convolutional layer.
[0135] In some embodiments of the present disclosure, the feature matching module further includes a plurality of densely connected residual blocks DRB connected in sequence. The output of the last DRB is used as the input of the first DRAB. Among them, the DRB includes a plurality of multi-scale residual blocks MRB connected in sequence.
[0136] In some embodiments of the present disclosure, the fusion module is specifically configured to:
[0137] Concatenate the outputs of each DRAB and the outputs of each DRB, and then perform channel compression to obtain a channel compression result with the same number of channels as the shallow features;
[0138] Determine the sum of the channel compression result and the shallow features as the feature fusion result.
[0139] In some embodiments of the present disclosure, the reconstruction module is specifically configured to:
[0140] Input the feature fusion result after being processed by an activation function into an adaptive multi-scale block AMB, and perform channel compression on the output of the AMB to obtain the multi-spectral image corresponding to the target image. Among them, the AMB includes a plurality of feature processing paths. Each feature processing path includes a convolutional layer, an activation layer, and a convolutional layer connected in sequence. The convolutional kernel sizes of the two convolutional layers in the same feature processing path are the same, and the convolutional kernel sizes of the convolutional layers in different feature processing paths are different. Each feature processing path has a corresponding weight. After the outputs of the plurality of feature processing paths are weighted and summed according to the corresponding weights, the output of the AMB is obtained.
[0141] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments of the method in the first aspect, and will not be elaborated here.
[0142] According to a third aspect of the embodiments of the present disclosure, please refer to the appended Figure 10 , which exemplarily shows a block diagram of an electronic device. For example, the device 1000 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0143] Referring to Figure 10 , the device 1000 may include one or more of the following components: a processing component 1002, a memory 1004, a power component 1006, a multimedia component 1008, an audio component 1010, an input / output (I / O) interface 1012, a sensor component 1014, and a communication component 1016.
[0144] The processing component 1002 generally controls the overall operation of the device 1000, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing element 1002 may include one or more processors 1020 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 1002 may include one or more modules to facilitate the interaction between the processing component 1002 and other components. For example, the processing component 1002 may include a multimedia module to facilitate the interaction between the multimedia component 1008 and the processing component 1002.
[0145] The memory 1004 is configured to store various types of data to support the operation of the device 1000. Examples of such data include instructions for any application or method operating on the device 1000, contact data, phone book data, messages, pictures, videos, etc. The memory 1004 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0146] The power component 1006 provides power to various components of the device 1000. The power component 1006 may include a power management system, one or more power supplies, and other components associated with rebuilding, managing, and distributing power to the device 1000.
[0147] The multimedia component 1008 includes a screen that provides an output interface between the device 1000 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 1008 includes a front camera and / or a rear camera. When the device 1000 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0148] The audio component 1010 is configured to output and / or input audio signals. For example, the audio component 1010 includes a microphone (MIC) that is configured to receive external audio signals when the device 1000 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1004 or transmitted via the communication component 1016. In some embodiments, the audio component 1010 further includes a speaker for outputting audio signals.
[0149] The I / O interface 1012 provides an interface between the processing component 1002 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0150] The sensor component 1014 includes one or more sensors for providing an assessment of the various aspects of the state of the device 1000. For example, the sensor component 1014 can detect the on / off state of the device 1000, the relative positioning of components, such as the display and the keypad of the device 1000. The sensor component 1014 can also detect a change in the position of the device 1000 or a component of the device 1000, the presence or absence of user contact with the device 1000, the orientation or acceleration / deceleration of the device 1000, and the temperature change of the device 1000. The sensor component 1014 can also include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 1014 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 1014 can also include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0151] The communication component 1016 is configured to facilitate communication between the device 1000 and other devices in a wired or wireless manner. The device 1000 can access a communication standard-based wireless network, such as WiFi, 2G or 3G, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 1016 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1016 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0152] In an exemplary embodiment, the device 1000 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the power supply method of the above electronic device.
[0153] In a fourth aspect, in an exemplary embodiment of the present disclosure, there is also provided a non-transitory computer-readable storage medium including instructions, such as a memory 1004 including instructions, and the above instructions can be executed by a processor 1020 of the device 1000 to complete the power supply method of the above electronic device. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0154] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0155] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A spectral reconstruction method, characterized in that Including: Performing shallow feature extraction processing on a target image to obtain shallow features of the target image; Inputting the shallow features into a feature matching module for feature matching to obtain a feature matching result, where the feature matching module includes a plurality of densely connected residual attention blocks (DRABs) connected in sequence, and each DRAB includes a plurality of multi-scale residual attention blocks (MRABs) connected in sequence. The input and output of each MRAB except the last one are cascaded and used as the input of the next MRAB, and the result obtained by cascading the input and output of the last MRAB and the sum of the input of the first MRAB are used as the output of the DRAB; Performing feature fusion processing on the feature matching result to obtain a feature fusion result; Reconstructing a multi-spectral image corresponding to the target image according to the feature fusion result.
2. The spectral reconstruction method according to claim 1, wherein Within the DRAB, the input of each MRAB except the first one is input into the corresponding MRAB after channel compression; And / or Within the DRAB, the result obtained by cascading the input and output of the last MRAB is channel-compressed and the sum of it and the input of the first MRAB are used as the output of the DRAB.
3. The spectral reconstruction method according to claim 2, characterized in that The MRAB includes a channel attention block (CA), a double down-sampling spatial attention block (DDSA), and three multi-scale residual blocks (MRBs). The input of the first MRB is respectively used as the input of the second MRB, the CA, and the DDSA. The product of the output of the second MRB and the output of the CA, and the concatenated result of the sum of the output of the second MRB and the output of the DDSA are used as the input of the third MRB, and the output of the third MRB is used as the output of the MRAB.
4. The spectral reconstruction method according to claim 3, wherein The input of the third MRB is input into the third MRB after channel compression.
5. The spectral reconstruction method according to claim 3, wherein The MRB includes a plurality of parallel paths, an activation layer, and at least one convolutional layer. Each path includes a different number of convolutional layers. The result obtained by cascading the outputs of the plurality of paths is used as the input of the activation layer, the output of the activation layer is used as the input of the at least one convolutional layer, the at least one convolutional layer is used for feature fusion and channel compression, and the sum of the output of the at least one convolutional layer and the input of the MRB is used as the output of the MRB.
6. The spectral reconstruction method according to claim 3, wherein The CA includes a pooling layer, two convolutional layers, an activation layer, and a normalization layer. The output of the pooling layer is used as the input of the first convolutional layer, the first convolutional layer is used for channel compression, the output of the first convolutional layer is used as the input of the activation layer, the output of the activation layer is used as the input of the second convolutional layer, the second convolutional layer is used for channel expansion, the output of the second convolutional layer is used as the input of the normalization layer, and the output of the normalization layer is used as the output of the CA.
7. The spectral reconstruction method according to claim 3, wherein The DDSA includes a plurality of downsampling paths and a convolutional layer. Among them, at least one of the plurality of downsampling paths performs downsampling in both channels and space simultaneously. The results obtained by multiplying the outputs of the plurality of downsampling paths in a certain order are used as the input of the convolutional layer. The convolutional layer is used for channel dimension expansion, and the output of the convolutional layer is used as the output of the DDSA.
8. The spectral reconstruction method according to claim 7, wherein The DDSA further includes a normalization layer. The plurality of downsampling paths include a first downsampling path, a second downsampling path, and a third downsampling path. Among them, the first downsampling path is used for downsampling in the channel dimension, and the second downsampling path and the third downsampling path are used for downsampling in both channels and space. The product of the output of the first downsampling path and the output of the second downsampling path is used as the input of the normalization layer. The product of the output of the normalization layer and the output of the third downsampling path is used as the input of the convolutional layer.
9. The spectral reconstruction method according to claim 1, wherein The feature matching module further includes a plurality of densely connected residual blocks DRBs connected in sequence. The output of the last DRB is used as the input of the first DRAB. Among them, each DRB includes a plurality of multi-scale residual blocks MRBs connected in sequence.
10. The spectral reconstruction method according to claim 9, wherein Performing feature fusion processing on the feature matching result to obtain a feature fusion result, including: Concatenating the outputs of each DRAB and the outputs of each DRB and then performing channel compression to obtain a channel compression result with the same number of channels as the shallow layer features; Determining the sum of the channel compression result and the result of the shallow layer features after convolutional processing as the feature fusion result.
11. The spectral reconstruction method according to claim 1, characterized in that Reconstructing the multi-spectral image corresponding to the target image according to the feature fusion result, including: Inputting the feature fusion result after being processed by an activation function into an adaptive multi-scale block AMB, and performing channel compression on the output of the AMB to obtain the multi-spectral image corresponding to the target image. Among them, the AMB includes a plurality of feature processing paths. Each feature processing path includes a convolutional layer, an activation layer, and a convolutional layer connected in sequence. The convolutional kernel sizes of the two convolutional layers in the same feature processing path are the same, and the convolutional kernel sizes of the convolutional layers in different feature processing paths are different. Each feature processing path has a corresponding weight. After the outputs of the plurality of feature processing paths are weighted and summed according to the corresponding weights, the output of the AMB is obtained.
12. A spectral reconstruction device, characterized in that, Including: An extraction module for performing shallow layer feature extraction processing on a target image to obtain the shallow layer features of the target image; A matching module, configured to input the shallow features into a feature matching module for feature matching to obtain a feature matching result, wherein the feature matching module includes a plurality of densely connected residual attention blocks (DRABs) connected in sequence, and each DRAB includes a plurality of multi-scale residual attention blocks (MRABs) connected in sequence. The input and output of each MRAB except the last one are cascaded and used as the input of the next MRAB. The sum of the result obtained by cascading the input and output of the last MRAB and the input of the first MRAB is used as the output of the DRAB; A fusion module, configured to perform feature fusion processing on the feature matching result to obtain a feature fusion result; A reconstruction module, configured to reconstruct the multi-spectral image corresponding to the target image according to the feature fusion result.
13. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory is configured to store computer instructions that can be run on the processor, and the processor is configured to perform the spectral reconstruction method according to any one of claims 1 to 10 when executing the computer instructions.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 10.