Multi-spectral image demosaicing method based on generative adversarial model
Through the multispectral image demosaic method based on the generative adversarial model, the problem of failure to fully consider the deep semantic information of the image in the prior art is solved, and a more complete and accurate reconstruction of spectral information is achieved.
Patent Information
- Application Number
- CN202510417727.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-27
AI Technical Summary
The existing multispectral image demosaic algorithm fails to fully consider the deep semantic information of the image, resulting in incomplete reconstructed spectral information.
The multispectral image demosaic method based on the generative adversarial model is adopted. The unsupervised learning algorithm of the generation adversarial network is used to complete the corresponding training using specific training data. On the basis of the original feature gap between the original calculated image pixel values, the depth feature extraction network calculates the difference between the predicted image and the original image is added.
The optimization direction of the network is improved and the completeness and accuracy of the reconstructed spectral information is ensured.
Smart Images

Figure CN120219155A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multispectral image processing, and particularly to a multispectral image demosaicking method based on a generative adversarial model. Background Art
[0002] Compared with traditional RGB color images, multispectral images contain both spatial information and richer spectral information, and play an important role in scenarios such as target detection, recognition, classification, and tracking. Due to the structure of the sensor, traditional multispectral imagers require multiple exposures or scans, and are not suitable for capturing the spectral-spatial information of moving objects from dynamic scenes. Based on this, snapshot spectral cameras that obtain target multispectral information through a single exposure have emerged. However, the multispectral images output by snapshot spectral cameras are of low resolution. To obtain high-resolution multispectral images, design algorithms need to be used to perform super-resolution restoration on the images. Since this process reconstructs the multispectral mosaic images output by the snapshot camera to obtain a high-resolution multispectral image composed of multiple single-spectral images combined, this process is also called demosaicking.
[0003] Currently, the mainstream demosaicking algorithms are divided into traditional methods and deep learning-based methods. Traditional demosaicking algorithms are based on the principle of compressive sensing, which transforms the demosaicking problem of mosaic images into a compressive sensing sparse signal recovery problem and a method based on an improved guided filter for demosaicking multispectral images; deep learning-based methods mainly use residual stacking and dense networks to supplement semantic and texture information in the deep features of images, and perform feature gap processing on the reconstructed results and the real results to optimize the feature extraction network. In the above methods, when calculating the feature gap, only the corresponding pixel values between the generated image and the real image are used for feature gap calculation, without considering the deep semantic information of the image, resulting in incomplete spectral information reconstructed. Summary of the Invention
[0004] In view of this, the present invention aims to provide a multispectral image demosaicking method based on a generative adversarial model to solve the problem that the existing methods do not consider the deep semantic information of the image, resulting in incomplete spectral information reconstructed. The present invention completes corresponding training using specific training data through the unsupervised learning algorithm of the generative adversarial network, and on the basis of the original calculation of the feature gap between the pixel values of the image, adds the calculation of the difference between the predicted image and the original image by the deep feature extraction network, so that the network has a more accurate optimization direction.
[0005] To achieve the above object, the technical solution of the present invention is realized as follows: A multispectral image demosaicking method based on a generative adversarial model specifically includes the following steps: S1: Construct a single - spectral image dataset based on the original multi - spectral image dataset, and use the single - spectral image dataset to construct an initial multi - spectral mosaic image dataset; S2: Construct a generative adversarial network, and use the initial multi - spectral mosaic image dataset to train the generative adversarial network to obtain a generative adversarial model; The generative adversarial model includes a generator, a discriminator, and a feature loss network; S3: Input the multi - spectral mosaic image to be reconstructed into the generator of the generative adversarial model for processing to obtain a reconstructed single - spectral image dataset.
[0006] Further, in step S1, the specific process of obtaining the single - spectral image dataset is as follows: Use a snapshot spectral camera to take multiple pictures of the target to obtain the original multi - spectral image dataset; Split the original multi - spectral image dataset into single - spectral images with a single filtering channel to obtain the single - spectral image dataset.
[0007] Further, in step S2, use the generator to process the initial multi - spectral mosaic image dataset to obtain a predicted single - spectral image dataset; The feature loss network calculates the deep semantic feature gap between the single - spectral images included in the single - spectral image dataset and the corresponding predicted single - spectral images in sequence, and uses the deep semantic feature gap, the feature gap of the generator, and the feature gap of the discriminator to train the generative adversarial network.
[0008] Further, the generator includes a multi - scale feature extraction module, a first residual feature extraction module, a second residual feature extraction module, a third residual feature extraction module, a fourth residual feature extraction module, a fifth residual feature extraction module, a sixth residual feature extraction module, and a feature encoding module. Among them, each initial multi - spectral mosaic image included in the initial multi - spectral mosaic image dataset is processed by the multi - scale feature extraction module and the first residual feature extraction module in sequence to obtain a first feature map. After the first feature map is processed by the second residual feature extraction module, a second feature map is obtained. After the second feature map is processed by the third residual feature extraction module, a third feature map is obtained. After the third feature map is processed by the fourth residual feature extraction module, a fourth feature map is obtained. After the fourth feature map is processed by the feature encoding module, a fifth feature map is obtained. The third feature map and the fifth feature map are input into the fifth residual feature extraction module for processing to obtain a sixth feature map. The second feature map and the sixth feature map are input into the sixth residual feature extraction module for processing to obtain a predicted single - spectral image dataset.
[0009] Further, the multi-scale feature extraction module includes a first 2D convolutional module, a second 2D convolutional module, a third 2D convolutional module, and a fourth 2D convolutional module. The feature A1 input to the multi-scale feature extraction module is processed by the first 2D convolutional module to obtain a feature map A2. After the feature map A2 is processed by the second 2D convolutional module, the third 2D convolutional module, and the fourth 2D convolutional module respectively, the processing results of the three are added together to obtain the output feature of the multi-scale feature extraction module.
[0010] Further, the network structures of the first residual feature extraction module, the second residual feature extraction module, the third residual feature extraction module, the fourth residual feature extraction module, the fifth residual feature extraction module, and the sixth residual feature extraction module are the same. Among them, the first residual feature extraction module includes a fifth 2D convolutional module, a sixth 2D convolutional module, a seventh 2D convolutional module, an eighth 2D convolutional module, and a ninth 2D convolutional module. The feature map B1 input to the first residual feature extraction module is processed by the fifth 2D convolutional module to obtain a feature map B2. After the feature map B2 is processed by the sixth 2D convolutional module and the seventh 2D convolutional module respectively, a feature map B3 and a feature map B4 are obtained correspondingly. After the feature map B4 is processed by the eighth 2D convolutional module, a feature map B5 is obtained. The feature map B3 and the feature map B5 are added together to obtain a feature map B6. The feature map B1 is input to the ninth 2D convolutional module for processing to obtain a feature map B7. The feature map B6 and the feature map B7 are concatenated to obtain a feature map B8.
[0011] Further, the feature encoding module includes an encoding module, a decoding module, and a linear module. The feature map C1 input to the feature encoding module is processed by the encoding module to obtain a feature map C2. The feature map C1 and the feature map C2 are input to the decoding module for processing to obtain a feature map C3. The feature map C3 is input to the linear module for processing to obtain the output feature of the feature encoding module.
[0012] Further, the discriminator includes a lightweight feature extraction module, a global adaptive average pooling module, and a Sigmoid function module. Among them, each single-spectral image included in the predicted single-spectral image dataset is sequentially input to the lightweight feature extraction module for processing to obtain a seventh feature map. The seventh feature map is processed by the global adaptive average pooling module and the Sigmoid function module to obtain a score.
[0013] Further, the lightweight feature extraction module includes a tenth 2D convolution module, an eleventh 2D convolution module, and a depthwise separable convolution module. After the feature map E1 input to the lightweight feature extraction module is processed by the tenth 2D convolution module, a feature map E2 is obtained. After the feature map E2 is sequentially processed by the depthwise separable convolution module and the eleventh 2D convolution module, a feature map E3 is obtained. The feature map E2 and the feature map E3 are added together to obtain the output feature of the lightweight feature extraction module.
[0014] Further, the feature loss network includes a first feature extraction module and a second feature extraction module. The network structures of the first feature extraction module and the second feature extraction module are the same, including a 3D convolution module, a twelfth 2D convolution module, a thirteenth 2D convolution module, a fourteenth 2D convolution module, a first FC module, a Dropout module, and a second FC module. The feature F1 input to the first feature extraction module or the second feature extraction module is sequentially processed by the 3D convolution module, the twelfth 2D convolution module, the thirteenth 2D convolution module, the fourteenth 2D convolution module, the first FC module, the Dropout module, and the second FC module to obtain an output feature. The feature gap between the output features of the first feature extraction module and the second feature extraction module is used as the deep semantic feature gap.
[0015] Compared with the prior art, the present invention can achieve the following beneficial effects: For the multi-spectral image demosaicking method based on the generative adversarial model of the present invention, through the unsupervised learning algorithm of the generative adversarial network and using specific training data, the corresponding training is completed. On the basis of calculating the feature gap between the original image pixel values, the difference between the predicted image and the original image is calculated by the deep feature extraction network, so that the network has a more accurate optimization direction. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 It is a schematic flowchart of the multi-spectral image demosaicking method based on the generative adversarial model according to the embodiment of the present invention; Figure 2 It is a schematic network structure diagram of the generator according to the embodiment of the present invention; Figure 3 It is a schematic network structure diagram of the multi-scale feature extraction module according to the embodiment of the present invention; Figure 4 It is a schematic network structure diagram of the first residual feature extraction module according to the embodiment of the present invention; Figure 5 Schematic diagram of the network structure of the feature encoding module described in the embodiments of the present invention Figure 6 Schematic diagram of the network structures of the encoding module and the decoding module described in the embodiments of the present invention Figure 7 Schematic diagram of the network structure of the discriminator described in the embodiments of the present invention Figure 8 Schematic diagram of the network structure of the feature loss network described in the embodiments of the present invention Figure 9 Schematic diagram of the network training of the generative adversarial network described in the embodiments of the present invention Detailed implementation manners
[0017] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention.
[0018] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0019] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0020] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0021] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0022] As Figure 1 shown, the multi-spectral image demosaicing method based on a generative adversarial model of the present invention specifically includes the following steps: S1: Construct a single-spectral image dataset based on the original multi-spectral image dataset, and use the single-spectral image dataset to construct an initial multi-spectral mosaic image dataset; S2: Construct a generative adversarial network, and use the initial multi-spectral mosaic image dataset to train the generative adversarial network to obtain a generative adversarial model; the generative adversarial model includes a generator, a discriminator, and a feature loss network; S3: Input the multi-spectral mosaic image to be reconstructed into the generator of the generative adversarial model for processing to obtain a reconstructed single-spectral image dataset.
[0023] The present invention generates a sparse multi-spectral mosaic image from the original multi-spectral image for training according to the required filter array requirements; then constructs a generative adversarial network, which mainly includes a multi-scale feature extraction module and a CSP-ized residual feature extraction module with a richer gradient flow, and a discriminator network for improving the performance of the generator network. The generative adversarial network is an end-to-end network, that is, there is no need to manually interpolate the multi-spectral mosaic image, and directly input a single-channel image to generate a high-resolution multi-spectral image composed of multiple single-spectral images combined. Here, the discriminator and the feature loss network are used to assist in training the generator.
[0024] In some embodiments, in step S1, the specific process of obtaining the single-spectral image dataset is as follows: Use a snapshot spectral camera to take multiple shots of the target to obtain the original multi-spectral image dataset; split the original multi-spectral image dataset into single-spectral images with a single filter channel to obtain the single-spectral image dataset.
[0025] It should be noted that taking a 3×3 square filter array period as an example, the filter array consists of 9 filter channels, and each filter channel transmits light of one band. The spectral information of the mosaic spectral image output by the snapshot spectral camera in one exposure is determined by its corresponding filter channel.
[0026] Further, the original multi-spectral image is spectrally split, and value-taking operations are performed on the pixel values of each spectral band image. In each filter array period, the number of spectral bands has the following coordinate relationship with the coordinates of the pixel values to be taken out: ; ; wherein, i (i = 1....3 2 ) represents the number of spectral bands, x represents the abscissa of the pixel to be taken out, and y represents the ordinate of the pixel to be taken out. represents the modulo operation, and / / represents the integer division operation.
[0027] In some embodiments, in step S2, a generator is used to process the initial multi-spectral mosaic image dataset to obtain a predicted single-spectral image dataset; The feature loss network sequentially calculates the deep semantic feature gaps between the single-spectral images included in the single-spectral image dataset and the corresponding predicted single-spectral images, and uses the deep semantic feature gaps, the feature gaps of the generator, and the feature gaps of the discriminator to train the generative adversarial network.
[0028] In some embodiments, as Figure 2 shown, the generator includes a multi-scale feature extraction module, a first residual feature extraction module, a second residual feature extraction module, a third residual feature extraction module, a fourth residual feature extraction module, a fifth residual feature extraction module, a sixth residual feature extraction module, and a feature encoding module. Among them, each initial multi-spectral mosaic image included in the initial multi-spectral mosaic image dataset is sequentially processed by the multi-scale feature extraction module and the first residual feature extraction module to obtain a first feature map. After the first feature map is processed by the second residual feature extraction module, a second feature map is obtained. After the second feature map is processed by the third residual feature extraction module, a third feature map is obtained. After the third feature map is processed by the fourth residual feature extraction module, a fourth feature map is obtained. After the fourth feature map is processed by the feature encoding module, a fifth feature map is obtained. The third feature map and the fifth feature map are input into the fifth residual feature extraction module for processing to obtain a sixth feature map. The second feature map and the sixth feature map are input into the sixth residual feature extraction module for processing to obtain the predicted single-spectral image dataset.
[0029] In some embodiments, as Figure 3As shown, the multi-scale feature extraction module includes a first 2D convolution module, a second 2D convolution module, a third 2D convolution module, and a fourth 2D convolution module. The feature A1 input to the multi-scale feature extraction module is processed by the first 2D convolution module to obtain a feature map A2. After the feature map A2 is processed by the second 2D convolution module, the third 2D convolution module, and the fourth 2D convolution module respectively, the processing results of the three are added together to obtain the output feature of the multi-scale feature extraction module.
[0030] It should be noted that the convolution kernel of the first 2D convolution module is 1×1, and the convolution kernels of the second 2D convolution module, the third 2D convolution module, and the fourth 2D convolution module are 7×7, 5×5, and 3×3 respectively. The processing process of the multi-scale feature extraction module is as follows: , ; where add represents adding the corresponding elements of the features, represents a 7×7 convolution, represents a 5×5 convolution, represents a 3×3 convolution, represents a 1×1 convolution, x represents the input feature of the multi-scale feature extraction module, represents the output feature of the multi-scale feature extraction module.
[0031] In some embodiments, as Figure 4 shown, the network structures of the first residual feature extraction module, the second residual feature extraction module, the third residual feature extraction module, the fourth residual feature extraction module, the fifth residual feature extraction module, and the sixth residual feature extraction module are the same. Among them, the first residual feature extraction module includes a fifth 2D convolution module, a sixth 2D convolution module, a seventh 2D convolution module, an eighth 2D convolution module, and a ninth 2D convolution module. The feature map B1 input to the first residual feature extraction module is processed by the fifth 2D convolution module to obtain a feature map B2. After the feature map B2 is processed by the sixth 2D convolution module and the seventh 2D convolution module respectively, the feature map B3 and the feature map B4 are obtained correspondingly. After the feature map B4 is processed by the eighth 2D convolution module, the feature map B5 is obtained. The feature map B3 and the feature map B5 are added together to obtain the feature map B6. The feature map B1 is input to the ninth 2D convolution module for processing to obtain the feature map B7. The feature map B6 and the feature map B7 are cascaded to obtain the feature map B8.
[0032] It should be noted that the convolution kernels of the fifth 2D convolution module, the sixth 2D convolution module, the seventh 2D convolution module, and the ninth 2D convolution module are 1×1, and the convolution kernel of the eighth 2D convolution module is 3×3. The calculation formula of the first residual feature extraction module is:
[0033] Among them, represents feature concatenation in the channel direction, and add represents element-wise addition of features. represents a 1×1 convolution. represents a 3×3 convolution. represents the output feature of the multi-scale feature extraction module. represents the output feature of the first residual feature extraction module after CSP transformation.
[0034] In some embodiments, as Figure 5 shown, the feature encoding module includes an encoding module, a decoding module, and a linear module. After the feature map C1 input to the feature encoding module is processed by the encoding module, the feature map C2 is obtained. The feature map C1 and the feature map C2 are input to the decoding module for processing to obtain the feature map C3. The feature map C3 is input to the linear module for processing to obtain the output feature of the feature encoding module.
[0035] It should be noted that the calculation formula of the feature encoding module is: ; Among them, Encoder represents the encoding module, Decoder represents the decoding module, and FC represents linear processing (here referring to the row module). represents the feature output by the fourth residual feature extraction module after CSP transformation in the previous step. is the output feature of the feature encoding module.
[0036] Furthermore, as Figure 6 shown, the encoding module first performs a dimension transformation on the input feature T through a Reshape operation, converting the three-dimensional matrix into a one-dimensional vector, and then uses a linear layer to perform two-dimensionality reduction on the one-dimensional vector to compress the information and obtain the output vector E. Then, the output vector E and the input feature T are input into the decoding module together. The decoding module first converts the input feature T into a one-dimensional vector through a Reshape operation, performs one-dimensionality reduction and compression after converting it into a vector, and outputs the vector t; then uses a linear layer to perform dimensionality increase on the output vector E and numerically add it to the vector t to obtain the vector d; then, uses a linear layer to perform dimensionality increase on the vector d and performs a dimension transformation through a Reshape operation to obtain the output matrix D with position information.
[0037] In some embodiments, as Figure 7As shown in the figure, the discriminator includes a lightweight feature extraction module, a global adaptive average pooling module, and a Sigmoid function module. Among them, each single-spectral image included in the predicted single-spectral image dataset is sequentially input into the lightweight feature extraction module for processing to obtain a seventh feature map. The seventh feature map is processed by the global adaptive average pooling module and the Sigmoid function module to obtain a score.
[0038] In some embodiments, the lightweight feature extraction module includes a tenth 2D convolution module, an eleventh 2D convolution module, and a depthwise separable convolution module. After the feature map E1 input to the lightweight feature extraction module is processed by the tenth 2D convolution module, a feature map E2 is obtained. The feature map E2 is sequentially processed by the depthwise separable convolution module and the eleventh 2D convolution module to obtain a feature map E3. The feature map E2 and the feature map E3 are added together to obtain the output feature of the lightweight feature extraction module.
[0039] It should be noted that the convolution kernels of the tenth 2D convolution module and the eleventh 2D convolution module are 1×1. The calculation formula of the discriminator is: ; ; ; Among them, represents the probability function, represents the global adaptive average pooling operation, represents the 1×1 convolution, represents the 3×3 depthwise separable convolution, Y represents the input single-spectral image, represents the input predicted single-spectral image, represents the feature gap (score) between the single-spectral image and the predicted single-spectral image.
[0040] In some embodiments, as Figure 8 shown, the feature loss network includes a first feature extraction module and a second feature extraction module. The network structures of the first feature extraction module and the second feature extraction module are the same, including a 3D convolution module, a twelfth 2D convolution module, a thirteenth 2D convolution module, a fourteenth 2D convolution module, a first FC module, a Dropout module, and a second FC module. The feature F1 input to the first feature extraction module or the second feature extraction module is sequentially processed by the 3D convolution module, the twelfth 2D convolution module, the thirteenth 2D convolution module, the fourteenth 2D convolution module, the first FC module, the Dropout module, and the second FC module to obtain an output feature. The feature gap between the output features of the first feature extraction module and the second feature extraction module is used as the deep semantic feature gap.
[0041] It should be noted that the calculation formula of the feature loss network is as follows: ; Among them, is the output feature of the feature loss network, FC represents linear processing, represents a 3×3 convolution, represents a dropout operation, is a reshape operation.
[0042] Furthermore, the generative adversarial network is trained using the total feature loss, and the total feature loss Loss is represented by the following formula: ; Among them, represents the feature gap of the generator, represents the feature gap of the discriminator, represents the deep semantic feature gap.
[0043] Furthermore, as Figure 9 shown, during the process of training the generative adversarial network, the feature gaps calculated by the generator, discriminator, and feature loss network are backpropagated to themselves to optimize the weight information and obtain the generative adversarial model.
[0044] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the disclosure of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved. This is not limited herein.
[0045] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multispectral image demosaicing method based on a generative adversarial model, characterized by: The specific steps include: S1: construct a single-spectrum image dataset based on the original multispectral image dataset, and use the single-spectrum image dataset to construct an initial multispectral mosaic image dataset; S2: Construct a generative adversarial network and use the initial multispectral mosaic image dataset to train the generative adversarial network to obtain a generative adversarial model; The generative adversarial model includes a generator, a discriminator and a feature loss network; S3: Input the multispectral mosaic image to be reconstructed into the generator of the generative adversarial model for processing to obtain the reconstructed single-spectrum image dataset.
2. The multispectral image demosaicing method based on the generative adversarial model according to claim 1, characterized in that: In step S1, the specific process of obtaining a single-spectrum image dataset is as follows: using a snapshot spectral camera to shoot the target at least twice to obtain an original multi-spectral image dataset; splitting the original multi-spectral image dataset into single-spectrum images with a single filtering channel to obtain a single-spectrum image dataset.
3. The multispectral image demosaicing method based on the generative adversarial model according to claim 1, characterized in that: In step S2, the initial multispectral mosaic image dataset is processed by the generator to obtain a predicted single-spectrum image dataset; The feature loss network sequentially calculates the deep semantic feature gap between the single-spectrum image contained in the single-spectrum image dataset and the corresponding predicted single-spectrum image, and trains the generative adversarial network using the deep semantic feature gap, the feature gap of the generator, and the feature gap of the discriminator.
4. The multispectral image demosaicing method based on the generative adversarial model according to claim 1, characterized in that: The generator includes a multi-scale feature extraction module, a first residual feature extraction module, a second residual feature extraction module, a third residual feature extraction module, a fourth residual feature extraction module, a fifth residual feature extraction module, a sixth residual feature extraction module and a feature encoding module, wherein each initial multi-spectral mosaic image contained in the initial multi-spectral mosaic image data set is processed by the multi-scale feature extraction module and the first residual feature extraction module in sequence to obtain a first feature map, the first feature map is processed by the second residual feature extraction module to obtain a second feature map, the second feature map is processed by the third residual feature extraction module to obtain a third feature map, the third feature map is processed by the fourth residual feature extraction module to obtain a fourth feature map, the fourth feature map is processed by the feature encoding module to obtain a fifth feature map, the third feature map and the fifth feature map are input into the fifth residual feature extraction module for processing to obtain a sixth feature map, the second feature map and the sixth feature map are input into the sixth residual feature extraction module for processing to obtain a predicted single-spectrum image data set.
5. The multispectral image demosaicing method based on the generative adversarial model according to claim 4, characterized in that: The multi-scale feature extraction module includes a first 2D convolution module, a second 2D convolution module, a third 2D convolution module and a fourth 2D convolution module. The feature A1 input to the multi-scale feature extraction module is processed by the first 2D convolution module to obtain a feature map A2. After the feature map A2 is processed by the second 2D convolution module, the third 2D convolution module and the fourth 2D convolution module respectively, the processing results of the three are added to obtain the output features of the multi-scale feature extraction module.
6. The multispectral image demosaicing method based on a generative adversarial model according to claim 4, characterized in that: The network structures of the first residual feature extraction module, the second residual feature extraction module, the third residual feature extraction module, the fourth residual feature extraction module, the fifth residual feature extraction module and the sixth residual feature extraction module are the same, wherein the first residual feature extraction module includes a fifth 2d convolution module, a sixth 2d convolution module, a seventh 2d convolution module, an eighth 2d convolution module and a ninth 2d convolution module, the feature map B1 input to the first residual feature extraction module is processed by the fifth 2d convolution module to obtain a feature map B2, the feature map B2 is processed by the sixth 2d convolution module and the seventh 2d convolution module respectively to obtain feature maps B3 and B4, the feature map B4 is processed by the eighth 2d convolution module to obtain feature map B5, the feature map B3 and the feature map B5 are added to obtain feature map B6, the feature map B1 is input to the ninth 2d convolution module for processing to obtain feature map B7, the feature map B6 and the feature map B7 are cascaded to obtain feature map B8.
7. The multispectral image demosaicing method based on a generative adversarial model according to claim 4, characterized in that: The feature encoding module includes an encoding module, a decoding module and a linear module. The feature map C1 input to the feature encoding module is processed by the encoding module to obtain the feature map C2. The feature map C1 and the feature map C2 are input to the decoding module for processing to obtain the feature map C3. The feature map C3 is input to the linear module for processing to obtain the output features of the feature encoding module.
8. The multispectral image demosaicing method based on a generative adversarial model according to claim 1, characterized in that: The discriminator includes a lightweight feature extraction module, a global adaptive average pooling module, and a Sigmoid function module, wherein each single-spectrum image contained in the predicted single-spectrum image data set is sequentially input into the lightweight feature extraction module for processing to obtain a seventh feature map, and the seventh feature map is processed by the global adaptive average pooling module and the Sigmoid function module to obtain a score.
9. The multispectral image demosaicing method based on the generative adversarial model according to claim 8, characterized in that: The lightweight feature extraction module includes a tenth 2D convolution module, an eleventh 2D convolution module and a depth-separable convolution module. The feature map E1 input to the lightweight feature extraction module is processed by the tenth 2D convolution module to obtain a feature map E2. The feature map E2 is processed by the depth-separable convolution module and the eleventh 2D convolution module in turn to obtain a feature map E3. The feature map E2 and the feature map E3 are added to obtain the output features of the lightweight feature extraction module.
10. The multispectral image demosaicing method based on a generative adversarial model according to claim 1, characterized in that: The feature loss network includes a first feature extraction module and a second feature extraction module. The first feature extraction module and the second feature extraction module have the same network structure, including a 3D convolution module, a twelfth 2D convolution module, a thirteenth 2D convolution module, a fourteenth 2D convolution module, a first FC module, a Dropout module and a second FC module. The feature F1 input to the first feature extraction module or the second feature extraction module is sequentially passed through the 3D convolution module, the twelfth 2D convolution module, the thirteenth 2D convolution module, the fourteenth 2D convolution module, the first FC module, the Dropout module and the second FC module to obtain the output feature, and the feature gap between the output features of the first feature extraction module and the second feature extraction module is used as the deep semantic feature gap.