A multi-spectral image demosaicking method based on convolutional neural network
By constructing a mosaic adaptive attention convolution neural network, the problems of spatial resolution loss and artifacts in multispectral image demosaics are solved, and more efficient feature extraction and reconstruction effects are achieved, chessboard effect is reduced, and the retention of image subject information is enhanced.
Patent Information
- Application Number
- CN202310027162.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-01-09
AI Technical Summary
The existing multispectral image demosaic method has spatial resolution loss and artifacts during reconstruction, and the convolutional neural network cannot effectively utilize the relationship between the various segments of the multispectral Raw image during feature extraction, resulting in poor reconstruction effect.
Mosaic adaptive attention convolution neural network is constructed, including mosaic convolution module, mosaic feature encoding module, mosaic feature decoding module and spatial attention module. The convolution kernel weight sharing strategy based on filters and the dense residual attention module are adopted, and the Charbonnier regression loss function is used for training to reduce the checkerboard effect and subject information loss of the feature map.
Without losing spatial information, the reconstruction quality of multi-spectral images is improved, the checkerboard effect is reduced, the ability to characterize the content of the image subject is enhanced, and the reconstruction effect is improved.
Smart Images

Figure CN116029930B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multispectral image, and particularly relates to a multispectral image demosaicking method based on convolutional neural network. Background Art
[0002] Multispectral images have more spectral bands than traditional RGB color images and possess more spectral information, and are widely used in remote sensing image processing, medical image analysis, food quality detection, authenticity target detection and other aspects. Existing multispectral imaging technologies mainly include spatial scanning, spectral scanning, snapshot spectral imaging, etc., all of which sacrifice spatial resolution or temporal resolution in exchange for spectral resolution. Existing snapshot spectral imaging based on multispectral filter array (MSFA) (as shown in Figure 1 ), (as shown in Figure 3 ) sacrifices spatial resolution to improve spectral resolution. As the spectral resolution increases, the spatial resolution will decrease accordingly. To obtain a complete multispectral image, demosaicking processing is required. Usually, interpolation method is used for demosaicking. Due to the sparse sampling of snapshot multispectral images in space, there are serious artifacts and checkerboard effects in the demosaicked images.
[0003] Currently, there are certain deficiencies in the snapshot multispectral image demosaicking methods based on deep convolutional neural network. These methods separate each spectral band of the snapshot multispectral image and perform reconstruction through the multispectral image with low spatial resolution, which will lead to the loss of spatial information and limit the reconstruction effect. Since the sampling pixel positions of each filter in the snapshot spectral image are different, there is a certain degree of offset in the spatial information corresponding to each spectral band. Artifacts will appear in the multispectral images reconstructed by these methods. Tewodros Amberbir Habtegebrial et al. proposed a multispectral Raw image demosaicking method based on deep residual network in their paper "Deep Convolutional Networks ForSnapshot Hypercpectral Demosaicking". This method separates and recombines the pixels corresponding to each spectral band in the multispectral Raw image to generate a multispectral image with low spatial resolution, and reconstructs the multispectral image with low spatial resolution through a deep residual network to generate a hyperspectral image with high spatial resolution.
[0004] However, in the prior art, each spectral band of the multi-spectral Raw image is separated, which reduces the spatial resolution of the image and loses some spatial information in the image. In addition, the spectral bands of the multi-spectral Raw image are not completely isolated from each other, and the relationship between the spectral bands cannot be extracted after separation, which affects the spectral accuracy of the reconstructed image. When using a convolutional neural network to extract features from a multi-spectral Raw image in the prior art, the convolutional kernels share the same weights, resulting in the hybridization of spectral information after feature extraction. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a multi-spectral image demosaicing method based on a convolutional neural network. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0006] The present invention provides a multi-spectral image demosaicing method based on a convolutional neural network, including:
[0007] Obtain a multi-spectral Raw image training set;
[0008] Construct a mosaic adaptive attention convolutional neural network;
[0009] Use the multi-spectral Raw image training set to train the mosaic adaptive attention convolutional neural network to obtain a trained mosaic adaptive attention convolutional neural network;
[0010] Input the multi-spectral Raw image to be measured into the trained mosaic adaptive attention convolutional neural network for demosaicing reconstruction to obtain a corresponding demosaiced reconstructed multi-spectral image;
[0011] Wherein, the mosaic adaptive attention convolutional neural network includes a mosaic convolution module, a mosaic feature encoding module, a mosaic feature decoding module, and a plurality of spatial attention modules, wherein,
[0012] The mosaic convolution module, the mosaic feature encoding module, and the mosaic feature decoding module are cascaded;
[0013] Both the mosaic feature encoding module and the mosaic feature decoding module include a plurality of cascaded dense residual attention modules;
[0014] The plurality of spatial attention modules are correspondingly connected between the dense residual attention modules of the mosaic feature encoding module and the dense residual attention modules of the mosaic feature decoding module.
[0015] In an embodiment of the present invention, obtaining a multi-spectral Raw image training set includes:
[0016] Performing simulation sampling on multiple original multi-spectral images using a multi-spectral filter array to obtain multiple multi-spectral Raw images as the training set of the multi-spectral Raw images;
[0017] Among them, the multi-spectral filter array includes c spectral filters arranged in a spatial pattern of m*n, where c is the number of spectral bands of the multi-spectral filter array, and both m and n are integers greater than zero.
[0018] In an embodiment of the present invention, in the mosaic feature encoding module, the input and output feature map channels of each dense residual attention module are different;
[0019] In the mosaic feature decoding module, the input and output feature map channels of each dense residual attention module are different.
[0020] In an embodiment of the present invention, the dense residual attention module includes a first convolutional unit, a first concatenate layer, a second convolutional unit, a second concatenate layer, a third convolutional unit, a mosaic channel attention layer, and a fusion layer cascaded in sequence, where,
[0021] The first convolutional unit, the second convolutional unit, and the third convolutional unit each include a cascaded convolutional layer and a first activation function layer. Among them, the convolutional kernel size of the convolutional layer is 3×3. The number of convolutional kernels of the convolutional layer of the first convolutional unit and the second convolutional unit is one-fourth of the input channels of this dense residual attention module, and the number of convolutional kernels of the convolutional layer of the third convolutional unit is the same as the output channels of this dense residual attention module; the activation function of the first activation function layer is the PReLU activation function;
[0022] The first concatenate layer is used to concatenate the input feature of the dense residual attention module with the output feature of the first convolutional unit;
[0023] The second concatenate layer is used to concatenate the input feature of the dense residual attention module, the output feature of the first convolutional unit, and the output feature of the second convolutional unit;
[0024] The mosaic channel attention layer is used to aggregate the spectra at the corresponding positions of each spectral filter in the output feature of the third convolutional unit, calculate weights for each pixel of the aggregated feature, and weight the output feature of the third convolutional unit using the calculated weights;
[0025] The fusion layer is used to fuse the input feature of the dense residual attention module with the output feature of the mosaic channel attention layer.
[0026] In one embodiment of the present invention, the spatial attention module includes a spatial feature aggregation layer, a spatial feature screening layer, a third concatenate layer, a mosaic convolution module, a second activation function layer, and a feature attention layer, where,
[0027] The spatial feature aggregation layer is used to perform average weighting on the pixels of all channels at each spatial position of the input features of the spatial attention module to achieve feature aggregation;
[0028] The spatial feature screening layer is used to screen the maximum value of the pixels of all channels at each spatial position of the input features of the spatial attention module;
[0029] The third concatenate layer is used to concatenate the output features of the spatial feature aggregation layer and the output features of the spatial feature screening layer, and the concatenated feature map is input into the mosaic convolution module;
[0030] The mosaic convolution module, the second activation function layer, and the feature attention layer are cascaded in sequence;
[0031] The activation function of the second activation function layer is the Sigmoid activation function.
[0032] In one embodiment of the present invention, the mosaic convolution module includes: a multi-core convolution layer, a feature channel screening layer, and a feature channel fusion layer; where,
[0033] The number of convolution kernels of the multi-core convolution layer is the same as the number of spectral bands of the multispectral filter array, and each convolution kernel of the multi-core convolution layer performs a convolution operation on the input feature map to obtain a corresponding mosaic feature map;
[0034] The feature channel screening layer is used to perform a filtering operation on the multiple mosaic feature maps generated by the multi-core convolution layer with a filter based on the spatial position, where the number of filters is the same as the number of convolution kernels of the multi-core convolution layer, and each filter responds to the spatial position of the filter corresponding to one spectral band in the multispectral filter array;
[0035] The feature channel fusion layer is used to add and fuse all the feature maps generated by the feature channel screening layer.
[0036] In one embodiment of the present invention, the mosaic adaptive attention convolutional neural network is trained using the multispectral Raw image training set to obtain a trained mosaic adaptive attention convolutional neural network, including:
[0037] Input the multi - spectral Raw image training set and the original multi - spectral image into the mosaic adaptive attention convolutional neural network to obtain the reconstructed multi - spectral image corresponding to the multi - spectral Raw image;
[0038] Use the Charbonnier regression loss function to calculate the regression loss between the reconstructed multi - spectral image and the original multi - spectral image;
[0039] According to the regression loss, use the adaptive moment estimation gradient descent algorithm to train the mosaic adaptive attention convolutional neural network for multiple rounds until the regression loss converges, and obtain the trained mosaic adaptive attention convolutional neural network.
[0040] In an embodiment of the present invention, the Charbonnier regression loss function is:
[0041]
[0042] where L represents the regression loss between the reconstructed multi - spectral image and the original multi - spectral image, represents the i - th original multi - spectral image, y i represents the i - th reconstructed multi - spectral image, M represents the total number of images in the multi - spectral Raw image training set, and ε represents the stable bias parameter.
[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0044] 1. The multi - spectral image demosaicking method based on a convolutional neural network of the present invention uses the constructed mosaic adaptive attention convolutional neural network to achieve demosaicking reconstruction of the input multi - spectral Raw image. In this mosaic adaptive attention convolutional neural network, there is a mosaic convolution module with a filter - based convolutional kernel weight sharing strategy for multi - spectral Raw images. This mosaic convolution module uses only one specific convolutional kernel weight for the pixel positions corresponding to the same filter in the MSFA, and different convolutional kernel weights are used for the pixels sampled by different filters, which can extract features from the entire multi - spectral Raw image without spectral band separation. It can better extract the spectral features at each pixel while not losing spatial information;
[0045] 2. In the proposed mosaic adaptive attention convolutional neural network of the multi - spectral image demosaicking method based on a convolutional neural network of the present invention, there is a dense residual attention module for feature aggregation adaptively based on the spatial position of the filter. After feature extraction, it can perform feature fusion on the pixel points in the feature map according to the corresponding positions of different filters in the MSFA (multi - spectral filter array), and generate a mosaic channel attention map in this way, making the network have stronger learning ability and reducing the checkerboard effect of the feature map;
[0046] 3. In the multi-spectral image demosaicking method based on a convolutional neural network of the present invention, in the proposed mosaic adaptive attention convolutional neural network, a spatial attention module for mosaic features is provided. This spatial attention module is mainly used to highlight the main content of the multi-spectral Raw image, enhance the neural network's representation ability of the main content in the mosaic feature map, and reduce the loss of the main information of the image before and after demosaicking.
[0047] 4. In the multi-spectral image demosaicking method based on a convolutional neural network of the present invention, the network parameters of the proposed mosaic adaptive attention convolutional neural network can be adjusted according to the number and arrangement of the filter in the MSFA used to generate the multi-spectral Raw image.
[0048] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the drawings, details are described as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a schematic structural diagram of a multi-spectral filter array (MSFA) provided by an embodiment of the present invention;
[0050] Figure 2 is a schematic diagram of a multi-spectral Raw image provided by an embodiment of the present invention;
[0051] Figure 3 is a schematic diagram of snapshot spectral imaging provided by an embodiment of the present invention;
[0052] Figure 4 is a flowchart of a multi-spectral image demosaicking method based on a convolutional neural network provided by an embodiment of the present invention;
[0053] Figure 5 is a schematic structural diagram of a mosaic adaptive attention convolutional neural network provided by an embodiment of the present invention;
[0054] Figure 6 is a schematic diagram of a mosaic convolution module provided by an embodiment of the present invention;
[0055] Figure 7 is a schematic diagram of a dense residual attention module provided by an embodiment of the present invention;
[0056] Figure 8 is a schematic diagram of a spatial attention module provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following provides a detailed description of a multi-spectral image demosaicing method based on a convolutional neural network proposed according to the present invention in combination with the accompanying drawings and specific embodiments.
[0058] Regarding the foregoing and other technical contents, features, and effects of the present invention, they can be clearly presented in the following detailed description in conjunction with the accompanying drawings. Through the description of the specific embodiments, a more in-depth and specific understanding of the technical means and effects adopted by the present invention to achieve the intended purpose can be obtained. However, the accompanying drawings are only provided for reference and explanation, and are not used to limit the technical solutions of the present invention.
[0059] Embodiment 1
[0060] Please refer to Figure 4 and Figure 5 , Figure 4 which is a flowchart of a multi-spectral image demosaicing method based on a convolutional neural network provided by an embodiment of the present invention; Figure 5 which is a schematic structural diagram of a mosaic adaptive attention convolutional neural network provided by an embodiment of the present invention. As shown in the figure, the multi-spectral image demosaicing method based on a convolutional neural network in this embodiment includes:
[0061] Step 1: Obtain a multi-spectral Raw image training set;
[0062] In an alternative embodiment, multiple original multi-spectral images are simulated and sampled using a multi-spectral filter array to obtain multiple multi-spectral Raw images as the multi-spectral Raw image training set. A schematic diagram of the multi-spectral Raw image is as shown in Figure 2 .
[0063] Among them, the original multi-spectral image is a three-dimensional data S ∈ R a×b×c , and each band in the multi-spectral image corresponds to a two-dimensional matrix S i ∈ R a×b , where ∈ represents the membership symbol, R represents the real number field symbol, a represents the width of the multi-spectral image, b represents the height of the multi-spectral image, c represents the number of spectral bands of the multi-spectral image, i represents the serial number of the spectral band in the multi-spectral image, and i = 1, 2,..., c.
[0064] As shown in Figure 1 which is a schematic structural diagram of a multi-spectral filter array (MSFA) provided by an embodiment of the present invention. In this embodiment, the multi-spectral filter array includes c spectral filters with a spatial arrangement of m * n, where c is the number of spectral bands of the multi-spectral filter array, and both m and n are integers greater than zero.
[0065] In a multi-spectral Raw image, its spatial size is a×b, which is the same as the original multi-spectral image, the number of channels is 1, and spatially, k pixels are grouped together. The information on each pixel is obtained by filtering and sampling through the corresponding spectral filter. The pixels obtained by sampling through each spectral filter are arranged periodically in the multi-spectral Raw image. Optionally, in this example, c = 9, m = 3, and n = 3.
[0066] In this embodiment, among the multiple multi-spectral Raw images obtained by sampling, 80% of the multi-spectral Raw images are selected as the training sample set, and the original multi-spectral images of each RAW image are used as the ground truth for evaluation and comparison. The remaining 20% of the multi-spectral Raw images and their corresponding original multi-spectral images are used as the test sample set for subsequent testing.
[0067] Step 2: Construct a mosaic adaptive attention convolutional neural network;
[0068] As Figure 5 shown, the mosaic adaptive attention convolutional neural network of this embodiment includes a mosaic convolution module, a mosaic feature encoding module, a mosaic feature decoding module, and multiple spatial attention modules. Among them, the mosaic convolution module, the mosaic feature encoding module, and the mosaic feature decoding module are cascaded. Both the mosaic feature encoding module and the mosaic feature decoding module include multiple cascaded dense residual attention modules. Among them, in the mosaic feature encoding module, the number of input and output feature map channels of each dense residual attention module is different; in the mosaic feature decoding module, the number of input and output feature map channels of each dense residual attention module is different. The multiple spatial attention modules are correspondingly connected between the dense residual attention modules of the mosaic feature encoding module and the dense residual attention modules of the mosaic feature decoding module.
[0069] Please refer to Figure 6 shown, a schematic diagram of a mosaic convolution module provided by an embodiment of the present invention. In an optional embodiment, the mosaic convolution module includes: a multi-core convolution layer, a feature channel screening layer, and a feature channel fusion layer.
[0070] Among them, the number of convolution kernels of the multi-core convolution layer is the same as the number of spectral bands of the multi-spectral filter array. Each convolution kernel of the multi-core convolution layer performs a convolution operation on the input feature map to obtain the corresponding mosaic feature map. The feature channel screening layer is used to perform a filtering operation on the multiple mosaic feature maps generated by the multi-core convolution layer with a filter based on the spatial position. The number of filters is the same as the number of convolution kernels of the multi-core convolution layer. Each filter responds to the spatial position of the filter corresponding to one spectral band in the multi-spectral filter array. The feature channel fusion layer is used to add and fuse all the feature maps generated by the feature channel screening layer.
[0071] In this embodiment, the multi-core convolutional layer has 9 convolutional kernels, the size of each of the 9 convolutional kernels is set to 3×3, and the number of each is set to 1. Each convolutional kernel is used to perform a convolution operation on the input feature map to generate 9 mosaic feature maps, and the number of filters in the feature channel screening layer is 9.
[0072] In this embodiment, the mosaic convolution module is designed based on the filter-based convolutional kernel weight sharing strategy for multi-spectral Raw images. This mosaic convolution module only uses a specific convolutional kernel weight for the pixel positions corresponding to the same filter in the MSFA, and the convolutional kernel weights used for the pixels sampled by different filters are different, enabling feature extraction for the entire multi-spectral Raw image without the need for spectral segment separation. It can better extract the spectral features at each pixel while not losing spatial information.
[0073] In an alternative embodiment, the mosaic feature encoding module includes four cascaded dense residual attention modules, namely the first dense residual attention module, the second dense residual attention module, the third dense residual attention module, and the fourth dense residual attention module.
[0074] Optionally, the input channel number of the first dense residual attention module is 32, and the output channel number is 32; the input channel number of the second dense residual attention module is 32, and the output channel number is 64; the input channel number of the third dense residual attention module is 64, and the output channel number is 128; the input channel number of the fourth dense residual attention module is 128, and the output channel number is 128.
[0075] In an alternative embodiment, the mosaic feature decoding module includes four cascaded dense residual attention modules, namely the fifth dense residual attention module, the sixth dense residual attention module, the seventh dense residual attention module, and the eighth dense residual attention module.
[0076] Optionally, the input channel number of the fifth dense residual attention module is 128, and the output channel number is 128; the input channel number of the sixth dense residual attention module is 128, and the output channel number is 64; the input channel number of the seventh dense residual attention module is 64, and the output channel number is 64; the input channel number of the eighth dense residual attention module is 64, and the output channel number is 32.
[0077] In an alternative embodiment, the mosaic adaptive attention convolutional neural network includes four spatial attention modules. Among them, one spatial attention module is used to connect between the first dense residual attention module and the eighth dense residual attention module, one spatial attention module is used to connect between the second dense residual attention module and the seventh dense residual attention module, one spatial attention module is used to connect between the third dense residual attention module and the sixth dense residual attention module, and one spatial attention module is used to connect between the fourth dense residual attention module and the fifth dense residual attention module.
[0078] Please refer to Figure 7 the schematic diagram of a dense residual attention module provided by the embodiment of the present invention shown in the figure. In an alternative embodiment, the dense residual attention module includes a first convolutional unit, a first concatenate layer, a second convolutional unit, a second concatenate layer, a third convolutional unit, a mosaic channel attention layer, and a fusion layer cascaded in sequence.
[0079] Among them, the first convolutional unit, the second convolutional unit, and the third convolutional unit all include a cascaded convolutional layer and a first activation function layer. Among them, the convolution kernel size of the convolutional layer is 3×3. The number of convolution kernels of the convolutional layer of the first convolutional unit and the second convolutional unit is one-fourth of the number of input channels of this dense residual attention module, and the number of convolution kernels of the convolutional layer of the third convolutional unit is the same as the number of output channels of this dense residual attention module; the activation function of the first activation function layer is the PReLU activation function.
[0080] The first concatenate layer is used to concatenate the input feature of the dense residual attention module and the output feature of the first convolutional unit; the second concatenate layer is used to concatenate the input feature of the dense residual attention module, the output feature of the first convolutional unit, and the output feature of the second convolutional unit.
[0081] The mosaic channel attention layer is used to aggregate the spectra at the corresponding positions of each spectral filter in the output feature of the third convolutional unit, calculate the weights for each pixel of the aggregated feature, and use the calculated weights to weight the output feature of the third convolutional unit. The fusion layer is used to fuse the input feature of the dense residual attention module and the output feature of the mosaic channel attention layer.
[0082] Optionally, the mosaic channel attention layer uses adaptive average pooling to perform feature aggregation in the spatial dimension of the feature map and uses the Sigmoid activation function.
[0083] In this embodiment, by setting a dense residual attention module that adaptively aggregates features based on the spatial position of the filter, after feature extraction, the pixel points in the feature map can be feature-fused according to the corresponding positions of different filters in the MSFA (multi-spectral filter array), and a mosaic channel attention map is generated thereby, enabling the network to have stronger learning ability and reducing the checkerboard effect of the feature map.
[0084] It should be noted that a dense residual attention module is used to build the mosaic feature encoding module and the mosaic feature decoding module. Here, the dense residual attention module is used for neural network feature extraction. For the dense residual attention module, in other alternative embodiments, only convolutional layers and activation function layers can be used without using the concatenate layer and the fusion layer, or only convolutional layers, activation function layers, and the concatenate layer can be used without using the fusion layer, or only convolutional layers, activation function layers, and the fusion layer can be used without using the concatenate layer, and all can complete the function of neural network feature extraction.
[0085] Please refer to Figure 8 the schematic diagram of a spatial attention module provided by the embodiment of the present invention shown in
[0086] Among them, the spatial feature aggregation layer is used to perform average weighting on the pixels of all channels at each spatial position of the input feature of the spatial attention module to achieve feature aggregation; the spatial feature screening layer is used to screen the maximum value of the pixels of all channels at each spatial position of the input feature of the spatial attention module; the third concatenate layer is used to concatenate the output feature of the spatial feature aggregation layer and the output feature of the spatial feature screening layer, and the concatenated feature map is input to the mosaic convolution module. The mosaic convolution module, the second activation function layer, and the feature attention layer are cascaded in sequence.
[0087] Optionally, the activation function of the second activation function layer is the Sigmoid activation function.
[0088] In this embodiment, the spatial attention module is set for mosaic features, mainly used to highlight the main content of the multi-spectral Raw image, enhance the representation ability of the neural network for the main content in the mosaic feature map, and can reduce the loss of the main information of the image before and after demosaicing.
[0089] It should be noted that in this embodiment, the mosaic convolution module is used in both the backbone structure and the spatial attention module of the mosaic adaptive attention convolutional neural network. The mosaic convolution module can be applied not only to the multi-spectral Raw image but also to the mosaic feature map. Since the mosaic feature map still has a mosaic arrangement, the mosaic convolution can still be used.
[0090] The network parameters of the mosaic adaptive attention convolutional neural network in this embodiment can be adjusted according to the number and arrangement of the filters in the MSFA used to generate the multi-spectral Raw image.
[0091] Step 3: Use the multi-spectral Raw image training set to train the mosaic adaptive attention convolutional neural network to obtain a trained mosaic adaptive attention convolutional neural network;
[0092] Among them, the specific training process of the mosaic adaptive attention convolutional neural network is as follows:
[0093] Input the multi-spectral Raw image training set and the original multi-spectral image into the mosaic adaptive attention convolutional neural network to obtain the reconstructed multi-spectral image corresponding to the multi-spectral Raw image;
[0094] Use the Charbonnier regression loss function to calculate the regression loss between the reconstructed multi-spectral image and the original multi-spectral image;
[0095] In the image reconstruction task, using the regression loss function as the optimization target, when the traditional L1 and L2 loss functions are relatively similar between the reconstructed image and the real image matrix, it will cause the optimization target to approach 0, making it difficult to continue optimizing, which limits the reconstruction performance of the convolutional neural network. To further improve the reconstruction performance of the neural network, this example uses the Charbonnier regression loss function, and its formula is as follows:
[0096]
[0097] Among them, L represents the regression loss between the reconstructed multi-spectral image and the original multi-spectral image, represents the i-th original multi-spectral image, y i represents the i-th reconstructed multi-spectral image, M represents the total number of images in the multi-spectral Raw image training set, and ε represents the stable bias parameter. In this embodiment, ε = 0.001.
[0098] According to the regression loss, use the adaptive moment estimation gradient descent algorithm to train the mosaic adaptive attention convolutional neural network for multiple rounds until the regression loss converges to obtain a trained mosaic adaptive attention convolutional neural network.
[0099] Step 4: Input the multi-spectral Raw image to be measured into the trained mosaic adaptive attention convolutional neural network for demosaicking reconstruction to obtain the corresponding demosaicked and reconstructed multi-spectral image.
[0100] The multi-spectral image demosaicking method based on convolutional neural network in this embodiment uses the constructed mosaic adaptive attention convolutional neural network to achieve demosaicking reconstruction of the input multi-spectral Raw image, without separating each spectral band in the multi-spectral Raw image, reducing the loss of spatial information during feature extraction, and this embodiment can perform feature extraction on the entire multi-spectral Raw image and extract the spectral connections between different spectra.
[0101] Mosaic channel attention is added to the mosaic adaptive attention convolutional neural network in this embodiment, making the neural network have stronger learning ability and reducing the checkerboard effect of the feature map. Mosaic spatial attention enhances the representation ability of the neural network for the main content in the mosaic feature map, reducing the loss of the main information of the image before and after demosaicking.
[0102] The peak signal-to-noise ratio (PSNR), structural similarity (SSIM), spectral angle matching (SAM) and other index effects of the multi-spectral image generated after demosaicking using the method of this embodiment are better than those of existing methods.
[0103] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the article or device including the element. "Connection" or "connected" and other similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The orientation or positional relationship indicated by "up", "down", "left", "right", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of the present invention.
[0104] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A multi-spectral image demosaicing method based on a convolutional neural network, characterized in that, Including: Obtaining a multi-spectral Raw image training set; Constructing a mosaic adaptive attention convolutional neural network; Training the mosaic adaptive attention convolutional neural network with the multi-spectral Raw image training set to obtain a trained mosaic adaptive attention convolutional neural network; Inputting the multi-spectral Raw image to be measured into the trained mosaic adaptive attention convolutional neural network for demosaicing reconstruction to obtain the corresponding demosaiced reconstructed multi-spectral image; Among them, the mosaic adaptive attention convolutional neural network includes a mosaic convolution module, a mosaic feature encoding module, a mosaic feature decoding module, and a plurality of spatial attention modules, where The mosaic convolution module, the mosaic feature encoding module, and the mosaic feature decoding module are cascaded; Both the mosaic feature encoding module and the mosaic feature decoding module include a plurality of cascaded dense residual attention modules; The plurality of spatial attention modules are correspondingly connected between the dense residual attention modules of the mosaic feature encoding module and the dense residual attention modules of the mosaic feature decoding module.
2. The multispectral image demosaicking method based on a convolutional neural network according to claim 1, wherein Obtaining a multi-spectral Raw image training set includes: Performing simulation sampling on a plurality of original multi-spectral images using a multi-spectral filter array to obtain a plurality of multi-spectral Raw images as the multi-spectral Raw image training set; Among them, the multi-spectral filter array includes c spectral filters with a spatial arrangement of m*n, c is the number of spectral bands of the multi-spectral filter array, and both m and n are integers greater than zero.
3. The multispectral image demosaicing method based on a convolutional neural network according to claim 1, wherein In the mosaic feature encoding module, the input and output feature map channel numbers of each dense residual attention module are different; In the mosaic feature decoding module, the input and output feature map channel numbers of each dense residual attention module are different.
4. The multi - spectral image demosaicking method based on a convolutional neural network according to claim 3, wherein The dense residual attention module includes a first convolutional unit, a first concatenate layer, a second convolutional unit, a second concatenate layer, a third convolutional unit, a mosaic channel attention layer, and a fusion layer cascaded in sequence, where The first convolutional unit, the second convolutional unit, and the third convolutional unit all include a cascaded convolutional layer and a first activation function layer, where the convolutional kernel size of the convolutional layer is 3×3, the number of convolutional kernels of the convolutional layer of the first convolutional unit and the second convolutional unit is one-fourth of the input channels of this dense residual attention module, and the number of convolutional kernels of the convolutional layer of the third convolutional unit is the same as the output channels of this dense residual attention module; the activation function of the first activation function layer is the PReLU activation function; The first concatenate layer is used to concatenate the input feature of the dense residual attention module and the output feature of the first convolutional unit; The second concatenate layer is used to concatenate the input feature of the dense residual attention module, the output feature of the first convolutional unit, and the output feature of the second convolutional unit; The mosaic channel attention layer is used to aggregate the spectra at the corresponding positions of each spectral filter in the output features of the third convolutional unit, calculate weights for each pixel in the aggregated features, and use the calculated weights to weight the output features of the third convolutional unit; The fusion layer is used to fuse the input features of the dense residual attention module and the output features of the mosaic channel attention layer.
5. The multispectral image demosaicking method based on a convolutional neural network according to claim 1, characterized in that, The spatial attention module includes a spatial feature aggregation layer, a spatial feature screening layer, a third concatenate layer, a mosaic convolution module, a second activation function layer, and a feature attention layer, where The spatial feature aggregation layer is used to perform average weighting on the pixels of all channels at each spatial position of the input features of the spatial attention module to achieve feature aggregation; The spatial feature screening layer is used to screen the maximum values of the pixels of all channels at each spatial position of the input features of the spatial attention module; The third concatenate layer is used to concatenate the output features of the spatial feature aggregation layer and the output features of the spatial feature screening layer, and the concatenated feature map is input into the mosaic convolution module; The mosaic convolution module, the second activation function layer, and the feature attention layer are cascaded in sequence; The activation function of the second activation function layer is the Sigmoid activation function.
6. The multi-spectral image demosaicing method based on a convolutional neural network according to claim 5, characterized in that The mosaic convolution module includes: a multi-core convolution layer, a feature channel screening layer, and a feature channel fusion layer; where The number of convolution kernels of the multi-core convolution layer is the same as the number of spectral bands of the multi-spectral filter array. Each convolution kernel of the multi-core convolution layer performs a convolution operation on the input feature map to obtain a corresponding mosaic feature map; The feature channel screening layer is used to perform a filtering operation on the multiple mosaic feature maps generated by the multi-core convolution layer with a filter based on spatial position. The number of filters is the same as the number of convolution kernels of the multi-core convolution layer, and each filter responds to the spatial position of the filter corresponding to one spectral band in the multi-spectral filter array; The feature channel fusion layer is used to add and fuse all the feature maps generated by the feature channel screening layer.
7. The multi-spectral image demosaicking method based on a convolutional neural network according to claim 2, wherein Training the mosaic adaptive attention convolutional neural network using the multi-spectral Raw image training set to obtain a trained mosaic adaptive attention convolutional neural network, including: Inputting the multi-spectral Raw image training set and the original multi-spectral image into the mosaic adaptive attention convolutional neural network to obtain a reconstructed multi-spectral image corresponding to the multi-spectral Raw image; Using the Charbonnier regression loss function to calculate the regression loss between the reconstructed multi-spectral image and the original multi-spectral image; According to the regression loss, using the adaptive moment estimation gradient descent algorithm to perform multiple rounds of training on the mosaic adaptive attention convolutional neural network until the regression loss converges to obtain a trained mosaic adaptive attention convolutional neural network.
8. The multi-spectral image demosaicking method based on a convolutional neural network according to claim 7, wherein The Charbonnier regression loss function is: Among them, L represents the regression loss between the reconstructed multi-spectral image and the original multi-spectral image, represents the i-th original multi-spectral image, y i represents the i-th reconstructed multi-spectral image, M represents the total number of images in the multi-spectral Raw image training set, and ε represents the stable bias parameter.
Citation Information
Patent Citations
Residual neural network based on hole convolution and two-stage image demosaicing method
CN111696036A
Joint denoising and demosaicing method and system based on distributed learning
CN113658060A