A snapshot compressive reconstruction and deep neural network based ultrafast imaging method

By combining snapshot compression reconstruction technology and deep neural networks, the response speed limitation of traditional camera systems in capturing rapidly changing light signals is solved, achieving ultrafast imaging with high spectral and temporal resolution and improving image quality.

CN119693481BActive Publication Date: 2025-11-04NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411746739.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-11-04
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

Traditional camera systems have limitations in response speed when capturing rapidly changing light signals, making it difficult to achieve extremely high temporal and spectral resolutions and effectively record rapid phenomena within a very short time.

Method used

A snapshot compression reconstruction technique and a deep neural network-based approach are employed. By training the compression reconstruction network and the progressive deep learning network on a synthetic simulation dataset, three-dimensional hyperspectral video data is reconstructed. The hierarchical abstraction capability of deep learning is used to extract spectral features and generate four-dimensional hyperspectral video.

Benefits of technology

It significantly improves the imaging quality and temporal resolution of ultrafast imaging, surpassing the performance of traditional methods, and can capture high spectral resolution images in a very short time, thus promoting the development of the field of ultrafast imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693481B_ABST
    Figure CN119693481B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on snapshot compression imaging and deep neural network's ultrafast imaging method.The specific steps are as follows: synthetic simulation dataset;Two-dimensional compressed data of ultrafast laser pulse used for detecting ultrafast scene is collected;End-to-end compressed reconstruction network based on low-rank decomposition and channel attention mechanism is constructed;End-to-end compressed reconstruction network is trained;Two-dimensional compressed data is reconstructed into three-dimensional hyperspectral video data using the trained reconstruction network;Progressive deep neural network is constructed and trained;Using the trained progressive deep neural network, three-dimensional hyperspectral video data is converted into four-dimensional hyperspectral video data containing hyperspectral information in each frame.Compared with the current leading snapshot compression reconstruction technology, the peak signal-to-noise ratio (PSNR) of the ultrafast imaging method proposed in the application is about 1dB higher in three-dimensional hyperspectral video reconstruction;Excellent results can also be achieved in expanding four-dimensional hyperspectral video generation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer vision, and in particular to a superfast imaging method based on snapshot compressive reconstruction technology and deep neural network. BACKGROUND

[0002] Superfast imaging technology plays a crucial role in analyzing transient processes. This technology has the ability to capture phenomena that occur and change rapidly within extremely short time scales, providing a unique perspective for studying and analyzing these rapidly changing processes. However, traditional camera systems, especially those using complementary metal-oxide-semiconductor (CMOS) or charge-coupled device (CCD) sensors, have inherent limitations in capturing rapidly changing light signals. These sensors cannot achieve extremely high levels of electronic response speed, which directly limits their ability to accurately capture rapidly changing light signals within extremely short time periods.

[0003] In addition, the readout circuit design of traditional cameras has bottlenecks. These circuits cannot process and store image data at femtosecond (1 femtosecond equals 10^-15 seconds) or picosecond (1 picosecond equals 10^-12 seconds) levels. Due to this speed limitation, traditional cameras cannot achieve extremely high temporal resolution, i.e., they cannot capture consecutive image frames within extremely short time intervals. This is a significant obstacle for processes occurring on the picosecond or femtosecond time scale. For example, in chemical reactions, biological processes, or material science, many key events occur within such short time periods, and traditional cameras cannot provide sufficient resolution to detail these processes.

[0004] Snapshot compressive imaging technology, as an emerging imaging method, can capture high-dimensional information such as spectral, temporal, or spatial data in a single exposure. The basic principle of this technology is to compress the high-dimensional spectral data cube through an imaging sensor into two-dimensional measurement data, and then use a compression reconstruction algorithm to recover the desired information. The advantage of snapshot compressive imaging technology is that it does not require multiple exposures, making it particularly suitable for capturing dynamic or transient phenomena. It has been widely applied in research work in the fields of biological imaging, remote sensing detection, and ultrafast physical processes.

[0005] In the field of ultrafast imaging, the attempt to introduce the concept of snapshot compressive imaging into the work of the compressive ultrafast spectral time (CUST) photography system compresses time or spectral information into a single two-dimensional image through spatial encoding, and uses the compression sensing algorithm TwIST to reconstruct multiple time or spectral resolution images. The reconstructed images are not limited by the speed of the imaging sensor, thus achieving extremely high temporal resolution and spectral resolution. However, this work only verifies the feasibility of the idea, and the imaging quality of ultrafast phenomena still needs to be further improved.

[0006] In recent years, artificial neural networks have evolved to the stage of deep learning. Deep learning adopts the concept of hierarchical abstraction, learning low-level concepts to obtain high-level concepts. This hierarchical structure is usually constructed by layer-by-layer greedy training algorithm to extract features useful for machine learning. Deep learning is a series of algorithms designed to use processing layers composed of complex structures or multiple nonlinear transformations to abstract data at a high level. With its powerful expression ability, deep learning has achieved outstanding results in various machine learning tasks and has shown the potential to surpass other methods in compression reconstruction. SUMMARY

[0007] In order to optimize the imaging quality of the existing ultrafast imaging system mentioned above, the purpose of the present application is to propose an ultrafast imaging method based on snapshot compression reconstruction technology and deep neural network.

[0008] To achieve the above purpose, the technical solution adopted by the present application is as follows:

[0009] An ultrafast imaging method based on snapshot compression imaging and deep neural network, comprising the following steps:

[0010] S1, synthesizing a four-dimensional hyperspectral video data set of simulation;

[0011] S2, collecting two-dimensional compressed data of an ultrafast laser pulse used to detect an ultrafast scene;

[0012] S3, constructing an end-to-end compression reconstruction network based on a deep unfolding architecture;

[0013] S4, training the compression reconstruction network using the data set synthesized in step S1;

[0014] S5, using the trained compression reconstruction network to reconstruct the two-dimensional compressed data collected in step S2 into three-dimensional hyperspectral video data;

[0015] S6, constructing a progressive deep learning network based on a U-Net architecture for extracting spectral features from hyperspectral images and generating corresponding new hyperspectral images;

[0016] S7, training the progressive deep learning network using the data set synthesized in step S1;

[0017] S8, using the three-dimensional hyperspectral video data obtained in step S5 and the trained progressive deep learning network to produce four-dimensional hyperspectral video data with each frame being a hyperspectral image.

[0018] In view of the fact that the compression reconstruction problem essentially belongs to the category of ill-posed problems, and the deep learning technology has shown significant superiority in the field of compression reconstruction, the present application combines the deep unfolding framework and the encoder-decoder network structure in deep learning, and proposes a novel and effective ultrafast imaging method, which uses a synthetic simulation data set to train the proposed compression reconstruction network and the progressive deep neural network, and can recover and reconstruct three-dimensional hyperspectral video data from two-dimensional compressed measurement data, and has more excellent performance in subjective and objective evaluation compared with other existing advanced snapshot compression reconstruction technologies; meanwhile, the three-dimensional hyperspectral video data obtained by reconstruction is also excellent in the aspect of four-dimensional hyperspectral video expansion generation. The method of the present application has far-reaching significance for promoting the development of the field of ultrafast imaging, and shows great potential and broad prospects beyond traditional methods. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly explain the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings involved in the embodiment or the prior art description. Obviously, the drawings in the following description are only several embodiments of the present application. Based on these drawings, other possible drawings can be deduced by those skilled in the art without creative labor.

[0020] Figure 1 The present application is a general flowchart of the method.

[0021] Figure 2 The present application is a flowchart of generating three-dimensional hyperspectral video data in the method.

[0022] Figure 3 The present application is a schematic diagram of the snapshot compression imaging system in the embodiment.

[0023] Figure 4 The present application is a simplified schematic diagram of the synthetic four-dimensional hyperspectral data set in the embodiment.

[0024] Figure 5 The present application is a network structure diagram of the reconstruction network based on deep unfolding in the embodiment.

[0025] Figure 6 The present application is a structure diagram of the high-order tensor decomposition module in the embodiment.

[0026] Figure 7 The present application is a structure diagram of the residual tensor learning module in the embodiment.

[0027] Figure 8 The present application is a structure diagram of the low-rank feature fusion module in the embodiment.

[0028] Figure 9A structural diagram of a residual dense connection module in an embodiment of the present application.

[0029] Figure 10 A network structural diagram of a progressive deep neural network based on U-net in an embodiment of the present application.

[0030] Figure 11 A method schematic diagram for generating four-dimensional hyperspectral video data based on three-dimensional hyperspectral video data in an embodiment of the present application. DETAILED DESCRIPTION

[0031] To make the purpose, technical solutions and advantages of the present application clearer, the following will make a further detailed description of the implementation method of the present application in combination with the drawings.

[0032] A method of ultrafast imaging (e.g., the time resolution of imaging is less than 1 picosecond) based on snapshot compression reconstruction and deep neural network in the present embodiment, the steps are as follows:

[0033] S1, synthesizing a simulation data set for training, verification and testing;

[0034] S2, combining pulse shaping technology, spectral encoding method, spatial modulation means, dispersion processing and space-time integration technology, realizing the acquisition of two-dimensional compressed data of the ultrafast laser pulse for detecting the ultrafast scene;

[0035] S3, based on the deep unfolding framework, constructing an end-to-end compression reconstruction network combining low-rank decomposition and channel attention mechanism;

[0036] S4, training the end-to-end compression reconstruction network;

[0037] S5, using the trained reconstruction network to reconstruct the two-dimensional compressed data acquired in step S2 into three-dimensional hyperspectral video data;

[0038] S6, constructing a progressive deep neural network based on U-Net;

[0039] S7, training the progressive deep neural network;

[0040] S8, based on the three-dimensional hyperspectral video data and the progressive deep neural network, generating four-dimensional hyperspectral video data containing hyperspectral information in each frame.

[0041] The selection of the data set in the step S1 has a great influence on the training of the entire neural network, and the DAVIS2017 data set is used in the embodiment. The data set is a public video data set provided by an official of a video segmentation image competition. A plurality of frames of video frames in RGB format are selected from the data set as input data, and are input into a neural network model (such as MST++) for generating a hyperspectral image from an RGB image. Four-dimensional hyperspectral video data sets are obtained through the processing of the model, as simulation data sets, and the four dimensions represent the spectral channel, time, height and width respectively, aiming to simulate the data acquisition result of an ultrafast phenomenon in the real world. Figure 4 A schematic diagram of a scene with 4 spectral channels and 4 time dimensions in a data set is given.

[0042] The flow chart of the step S2 is shown in Figure 2 . Specifically, the ultrafast laser pulse used to detect the ultrafast scene needs to be shaped and stretched first. The purpose of this process is to effectively encode the time information in the subsequent steps. After the pulse detects the ultrafast phenomenon, the hyperspectral data I(x, y, λ) is obtained. By forming a correspondence between the different wavelengths of the ultrafast pulse and the time, a three-dimensional hyperspectral video data I(x, y, t) containing spatial, temporal and hyperspectral information is obtained.

[0043] Subsequently, the three-dimensional hyperspectral video data is spatially modulated using an encoding aperture. By introducing additional spatial information through a specific encoding mode, the encoding is introduced for the subsequent two-dimensional measurement process. The design of the encoding mode is that the value of each pixel point is randomly 0 or 1, so that the same designed encoding mode can be used during training and verification. Assuming where H represents the height dimension, W represents the width dimension, C represents the channel dimension, and X is the three-dimensional hyperspectral data cube input into the snapshot compressive imaging system. The schematic diagram of the system is shown in Figure 3 , where the upper part is the imaging principle of the system, and the lower part is the three-dimensional hyperspectral data cube input into the snapshot compressive imaging system. The schematic diagram of the system is shown in Figure 3 , where the upper part is the imaging principle of the system, and the lower part is the optical implementation of the system. Using represents a two-dimensional entity mask, where represents the real number field. The three-dimensional hyperspectral data cube is first modulated by the mask. In the embodiment, an encoding grating or a spatial light modulator is used to realize this step. For any channel m, this process can be represented by the following expression:

[0044] X'(:,:,m) = M ® X(:,:,m) (1)

[0045] where X' represents the modulated data block, X(:,:,m) represents the mth channel of the input data, and ® represents element-wise multiplication.

[0046] The modulated hyperspectral data cube is dispersed using a dispersive prism, and then captured using a camera, as shown in Fig. 1. Figure 3 As shown in the upper half, the element-wise summation of the spatially shifted spectral data cube between channels is performed to obtain two-dimensional measurement data Y containing time information:

[0047]

[0048] where (u, v) represents the coordinates on the camera detection plane, d i represents the number of pixels offset in the ith channel during dispersion (in this embodiment, d i = 2i-2), represents the noise present during measurement, is the two-dimensional measurement data collected. Define is the imaging matrix, where D m is a diagonal matrix with the diagonal elements being the vectorized coded aperture M(:,:,m), and m is a positive integer between 1 and n. By matrix-vectorizing (i.e., concatenating each column of the matrix into a vector) the other three variables Y, X' and N in equation (2) into y, x and n, respectively, equation (2) can be converted into the following form:

[0049] y = Φx + n (3)

[0050] The reconstruction problem can be described as solving for the reconstruction result x given the two-dimensional measurement value y and the imaging matrix Φ. To solve the ill-posed problem of equation (3) in the presence of noise, additional regularization is usually required to achieve accurate and stable solutions. The estimate of the variable x can be derived by solving equation (4) to solve the compression reconstruction problem:

[0051]

[0052] where Ω(x) is a regularization term to restrict the range of the solution, and λ is a parameter of the regularization term. The present application uses a generalized alternating projection algorithm to solve this problem, the principle of which is to decompose this optimization problem into two sub-problems by introducing an auxiliary variable v of the same dimension as the variable x to solve iteratively. In the kth stage of the iteration process (where k = 1...K, K is the maximum number of iterations), the first sub-problem is to update x in the (k+1)th stage by calculating the auxiliary variable v in the kth stage (k) on a linear manifold y = the Euclidean projection on Φx (k+1)

[0053] x (k+1) = v (k) + Φ T (ΦΦ T ) -1 (y-Φv k ) (5)

[0054] The goal of the second sub-problem is to train an effective denoiser (k+1) to make x (k+1) closer to the target result, and update x (k+1) using the denoiser to obtain v in :

[0055]

[0056] Therefore, in step S3, the present application proposes an end-to-end compression reconstruction network based on the deep unfolding framework combining low-rank decomposition and channel attention mechanism based on formula (5) and formula (6), and the specific structure is as follows Figure 5 ​As shown, K (K=9 in this embodiment) modules implementing the generalized alternating projection algorithm are connected in series, corresponding to the K stages in the iterative solving method of step S5. Equation (5) shows that both the two-dimensional compressed measurement data and the mask used for spatial modulation in the measurement process are inputs to the compression reconstruction network. The denoiser used in equation (6) is the LRUnet denoising network designed in this application. This denoising network is also an encoder-decoder structure. In the encoding stage, three high-order tensor decomposition modules are used to enhance the feature extraction capability of the shallow layers of the network. Before the input data enters the high-order tensor decomposition module, a mapping layer with the structure of "convolutional layer-activation function-convolutional layer" (hereinafter referred to as "double-layer convolutional module") is used to map the features of the input degraded video data. The kernel size of the convolutional layer in all double-layer convolutional modules in the entire LRUnet denoising network is set to 3x3. In order to further expand the receptive field of the network, a downsampling operation is introduced after each high-order tensor decomposition module (except the last module). In the decoding stage, two upsampling modules and three double-layer convolutional modules are used to gradually restore the data to its original resolution and effectively suppress noise interference in the process. The last high-order tensor decomposition module of the encoder is connected to the first upsampling module of the decoder. A residual dense connection module is designed between the encoding and decoding stages. By fusing deep and shallow features, efficient multi-level feature fusion and full use of global information are achieved. The output of each upsampling is concatenated with the output of the corresponding residual dense connection module and then input to the double-layer convolutional module to obtain the output result for the next upsampling. Finally, the output result of the second-to-last double-layer convolutional module is input to another double-layer convolutional module, and then the residual connection with the input is performed to generate the reconstructed image. The design methods of the high-order tensor decomposition module and the residual dense connection module will be described in detail below.

[0057] Figure 8 The architecture of the low-rank feature fusion module is shown. This module uses the structure of "global average pooling + convolutional layer + Sigmoid activation function" in the channel, height, and width dimensions to extract features and generate three one-dimensional vectors. Then, the three vectors are re-aggregated using the Kronecker product to obtain the output features. Figure 7 The architecture of the residual tensor learning module is shown. This module uses a block residual structure to capture different frequency features in the image. The input of the first residual tensor learning module is the input of the high-order tensor decomposition module, and the input of each subsequent residual tensor learning module is the output of the previous module. Specifically, first, a low-rank fusion module is used to generate a rank-1 tensor Ox from the input features. Then, the difference between the difference result and the output of another low-rank fusion module is added (this represents the residual structure), thereby obtaining the output of the residual tensor learning module. As shown in Figure 6As shown, the architecture of the higher-order tensor decomposition module includes r residual tensor learning modules, whose function is to transform the input features into r rank-1 tensors. These tensors are concatenated and then passed through a 3×3 convolutional layer to extract a 3D attention map. Subsequently, this attention map is subjected to a Hadamard product operation with the input features, and the result is added to the original input features. Finally, the output features are obtained through a convolutional layer with a 3×3 kernel and an activation function.

[0058] The structure of the residual dense connection module is as follows Figure 9 As shown, it mainly consists of two parallel branches. The first branch contains three 3×3 convolutional layers, which are specifically used for deep feature extraction. Input feature F in It is the input of the first convolutional layer; the input feature F in The output of the second convolutional layer is concatenated with the output F1 of the first convolutional layer; the input feature F in The outputs F1 and F2 of the first and second convolutional layers are concatenated and used as the input to the third convolutional layer. The outputs F1, F2, and F3 of the three convolutional layers are then concatenated and fed into a 1×1 convolutional layer to further fuse these deep features, thereby generating a richer feature representation. The second branch primarily focuses on extracting global contextual information. The input data for this module is first processed by a global pooling layer, then sequentially passed through two fully connected layers, with a ReLU activation function applied after each layer, finally outputting the processing result of this branch. The sum of the outputs of these two branches is the final output of the entire module.

[0059] In step S4, the Adam adaptive gradient optimization algorithm is used to train the end-to-end compressed reconstruction network. The loss function is defined as the mean square error between the network output reconstruction result and the true value. During network training and validation, the mask is a randomly generated 256×256 matrix containing only 0s and 1s; the images in the synthesized four-dimensional hyperspectral video dataset in step S1 are randomly cropped into image patches with a width and height of 256 and fewer than 31 channels (because the maximum number of channels for generating a hyperspectral image from an RGB image is 31). Figure 4 Taking a simplified diagram as an example, the diagram shows a four-dimensional hyperspectral video dataset with four spectral channels and a time dimension of four. Four images with the same coordinates on the time axis (images on the diagonal passing through the origin) are selected and sequentially stitched together to form an image patch with four channels. Adjusting the number of channels only requires modifying the number of diagonal images stitched together during this operation; in this embodiment, the number of channels is selected as eight. These image patches are used to train the network end-to-end. After the loss function reaches convergence, the trained model is saved as an end-to-end compressed reconstruction network.

[0060] In step S5, the dataset synthesized in step S1 is processed according to the procedure in step S2 to obtain compressed two-dimensional measurement data, which, along with a mask, is fed into the trained end-to-end compression and reconstruction network to output three-dimensional hyperspectral video data I(x,y,t). The reconstruction level achieved using the reconstruction network proposed in step S3 demonstrates excellent peak signal-to-noise ratio (PSNR), approximately 1 dB higher than other advanced snapshot compression and reconstruction methods. This result demonstrates that the method proposed in this invention can more effectively preserve the details of the original signal during image reconstruction, and that the theoretical basis for the finding that structural information in noisy images is mainly contained within low-rank components, while noise components are mainly distributed in high-rank parts, is reasonable.

[0061] In step S6, this invention designs a progressive deep neural network based on U-net, which takes a pair of images and a task number as input and outputs a pair of images. Its structure is as follows: Figure 10 As shown, the network first performs initial feature mapping through a 3×3 convolutional layer, resulting in an input feature with 32 channels. The core structure of each layer in the progressive deep neural network is the standard U-net architecture, which performs two downsampling and upsampling operations during both the encoding and decoding stages. All U-net layers in the network share the same parameter set. Each layer's U-net receives the concatenation of the output and input features of the previous layer's U-net as input; for the first layer's U-net, the input is the concatenation of two identical input features to meet the network's specific requirements for input data size. Furthermore, the task number, another input to the network, determines which layer's U-net output is selected as the final output. The output of the selected layer's U-net undergoes feature mapping through a 3×3 convolutional layer, remapping its channel count to 2, which becomes the network's final output. The parameters of the convolutional layer corresponding to the output feature mapping for each task number are different.

[0062] In step S7, the Adam adaptive gradient optimization algorithm is used to train the progressive deep neural network. The loss function is defined as the mean squared error between the network output and the true value. Image pairs are selected from the four-dimensional hyperspectral video dataset synthesized in step S1. Figure 4 For example, let's define the horizontal axis t in the coordinate system of the graph as representing time, and the vertical axis λ as representing the spectral wavelength. i ,λ j ) indicates that at t i The spectral wavelength corresponding to the time is λ jThe image is used as the input to the network. Two frames are randomly selected and stitched together, serving as the images of two vertices on one diagonal of a square. This stitched image is then used as the network input; that is, two images are selected whose absolute difference in time axis coordinates is equal to the absolute difference in spectral wavelength axis coordinates. The two vertices on the other diagonal of the square are used as the ground truth values ​​corresponding to the network output. In each training iteration, the task number is defined as the absolute difference in time axis coordinates between the two input images. The method for selecting the training set can be expressed as: the input image is the coordinate (t... i ,λ j ) and (t m ,λ n Two frames of images are output, where i, j, m, and n are all positive integers greater than 0 and less than the maximum value of the time axis coordinates, and satisfy |im|=|jn| and are not 0. This equation value is defined as the task number, and the true value corresponding to the output image is the coordinate (t). i ,λ n ) and (t m ,λ j The two frames of images are used. Only the parameters of the output convolutional layer corresponding to the task number of the current input data are updated. The task numbers of all training data in the same batch are equal until the loss function converges, and then the training model is saved.

[0063] In step S8, the three-dimensional hyperspectral video data I(x,y,t) reconstructed in step S5 needs to be expanded into a four-dimensional hyperspectral video I(x,y,t,λ) where each frame contains hyperspectral information. The expansion method is as follows: Figure 11 As shown in the diagram. This illustration assumes the reconstructed 3D hyperspectral video data has a size of 4×256×256, and the resulting 4D hyperspectral video, obtained through expansion, would have a size of 4×4×256×256. Figure 11 In this coordinate system, the horizontal axis t represents time, and the vertical axis λ represents the spectral wavelength. i ,λ j Then it means t i At time λ, the corresponding spectral wavelength is j The images. According to the rules followed in step S4 when formulating the training set for the reconstruction network, in this coordinate system, the four images of the reconstructed 4×256×256 three-dimensional hyperspectral video data should be located on the same diagonal line as the time axis coordinate and the spectral channel coordinate (i.e., the image coordinates (t...). i ,λ j (satisfying i = j), such as Figure 11 As shown in the top left corner. Figure 11The operations in the first row represent that the images at (t1, λ1) and (t3, λ3) are selected for splicing, and then input into the trained progressive deep neural network in step S7, according to the definition of the task number in step S7, the task number at this time is |1-3| = 2, and the two channels of the output result of the network at this time represent the images at (t1, λ3) and (t3, λ1) positions. After the same operation is performed on the other 5 pairs of mutually different images on the diagonal line with the same horizontal and vertical axis coordinate numbers, a four-dimensional hyperspectral video data with a size of 4*4*256*256 can be completed, and each frame of the video is a multi-channel hyperspectral image. It can be easily calculated that a three-dimensional hyperspectral video data set with m spectral channels needs to be operated 0.5*m(m-1) times to expand to a four-dimensional hyperspectral video data with m spectral channels and m time channels. The average peak signal-to-noise ratio (PSNR) calculated between the finally output four-dimensional hyperspectral video data and the true value can reach up to 39dB, and it can be seen that the effect of the expansion method of the application is relatively good.

Claims

1. A method for ultrafast imaging based on snapshot compressive imaging and deep neural network, characterized in that, The method comprises the following steps: S1, synthesizing a simulated four-dimensional hyperspectral video dataset; S2, collecting two-dimensional compressed data of an ultrafast laser pulse used for detecting an ultrafast scene; S3, constructing an end-to-end compression reconstruction network based on a deep unfolding architecture; the compression reconstruction network comprises a plurality of modules for implementing a generalized alternating projection algorithm and connected in series, each module comprising a denoising network, the denoising network comprising an encoder and a decoder, the encoder comprising a double-layer convolution module, three high-order tensor decomposition modules and two down-sampling modules, the high-order tensor decomposition modules being used for converting input features into tensors with a rank of 1; the decoder comprising two up-sampling modules and three double-layer convolution modules, and a residual dense connection module being arranged between the encoder and the decoder for performing a jump connection; S4, training the compression reconstruction network using the dataset synthesized in step S1; S5, reconstructing the two-dimensional compressed data collected in step S2 into three-dimensional hyperspectral video data using the trained compression reconstruction network; S6, constructing a progressive deep learning network based on a U-Net architecture, which is used for extracting spectral features from a hyperspectral image and generating a corresponding new hyperspectral image; the progressive deep learning network takes an image pair and a task sequence number as input, and the core structure of each layer of the progressive deep learning network is a U-Net architecture, which performs down-sampling and up-sampling operations twice in the encoding and decoding stages; all the U-Nets in all the layers of the progressive deep learning network share the same parameter set, and each U-Net receives the spliced data of the output of the previous U-Net and the input features as input; for the first U-Net, the input is the splicing of two identical input features; the task sequence number is the absolute value of the coordinate difference value of the two input images on the time axis, the progressive deep learning network selects the output of the corresponding U-Net according to the input task sequence number, and performs feature mapping through a convolution layer to obtain the final output; the parameters of the output feature mapping convolution layer corresponding to each task sequence number are different; S7, training the progressive deep learning network using the dataset synthesized in step S1; S8, producing four-dimensional hyperspectral video data in which each frame is a hyperspectral image using the three-dimensional hyperspectral video data obtained in step S5 and the trained progressive deep learning network. 2.The ultrafast imaging method based on snapshot compressive imaging and deep neural network according to claim 1, wherein, In step S1, the synthesized simulated four-dimensional hyperspectral video dataset specifically comprises: selecting a plurality of consecutive video frames in RGB format from a video dataset as input data, inputting the input data into a neural network model for generating a hyperspectral image from an RGB image, and obtaining a four-dimensional hyperspectral video dataset after processing by the model, wherein the four dimensions represent a spectral channel, a time, a height and a width, respectively. 3.The ultrafast imaging method based on snapshot compressive imaging and deep neural network according to claim 1, wherein, In step S2, the collection of two-dimensional compressed data of an ultrafast laser pulse used for detecting an ultrafast scene specifically comprises: S21, shaping and stretching the ultrafast laser pulse used for detecting the ultrafast scene to obtain three-dimensional hyperspectral video data; S22, spatially modulating the three-dimensional hyperspectral video data using an encoding aperture; S23, dispersing the modulated data using a dispersion prism; S24, capturing the two-dimensional compressed data using a collection device. 4.The ultrafast imaging method based on snapshot compressive imaging and deep neural network of claim 1, wherein, The high-order tensor decomposition module includes r residual tensor learning modules, the input of the first residual tensor learning module is the input of the high-order tensor decomposition module, and the input of each subsequent residual tensor learning module is the output of the previous module; the residual tensor learning module adopts a block residual structure to capture different frequency features in the image.

5. The ultrafast imaging method based on snapshot compressive imaging and deep neural network according to claim 4, characterized in that, The residual tensor learning module includes two low-rank fusion modules, a first low-rank fusion module is first used to generate a tensor with a rank of 1 from the input feature, then a difference operation is performed with the input feature, and the difference result is added to the output of the second low-rank fusion module, thereby obtaining the output of the residual tensor learning module.

6. The ultrafast imaging method based on snapshot compressive imaging and deep neural network according to claim 5, characterized in that, The low-rank fusion module adopts a structure of global average pooling + convolution layer + Sigmoid activation function. 7.The ultrafast imaging method based on snapshot compressive imaging and deep neural network according to claim 1, wherein, The residual dense connection module includes two parallel branches: the first branch includes three 3x3 convolutional layers for extracting deep features, wherein the output F1 of the first convolutional layer is concatenated with the input feature F in as the output of the second convolutional layer; the input feature F in , the output F1 of the first convolutional layer and the output F2 of the second convolutional layer are concatenated as the input of the third convolutional layer, the outputs F1, F2 and F3 of the three convolutional layers are concatenated and input to a 1x1 convolutional layer for further fusing the deep features; the second branch is used for extracting global context information, the input data is first processed by a global pooling layer, then sequentially passes through two fully connected layers, and a ReLU activation function is applied after each layer, and then the processing result of the second branch is output; the output results of the two branches are added to obtain the final output of the entire residual dense connection module. 8.The ultrafast imaging method based on snapshot compressive imaging and deep neural network of claim 1, wherein, In step S7, two frames are randomly selected from the data set synthesized in step S1 as the images of the two vertices of a diagonal of a square for splicing, serving as the input of the progressive deep learning network, that is, two images are selected, the absolute value of the time axis coordinate difference and the absolute value of the spectral wavelength axis coordinate difference are equal, and the two vertices on the other diagonal of the square are taken as the real values corresponding to the network output results.

Citation Information

Patent Citations

  • Polarization and hyper-spectral compression imaging method and system

    CN102155992A

  • Hyperspectral image compression method and device based on deep learning

    CN110348487A