Medium-wave infrared spectral imaging metasurface design and spectral image reconstruction synchronization realization method
Through the neural network optimization design of the metasurface unit, the problems of large size and spectral reconstruction distortion of the medium-wave infrared spectral imaging system are solved, and spectral imaging with high spectral resolution and high signal-to-noise ratio is achieved, which is suitable for multi-scenario applications.
Patent Information
- Application Number
- CN202510853216.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-10
AI Technical Summary
Traditional medium-wave infrared spectral imaging systems are large in size, high in cost, and difficult to achieve both miniaturization and high spectral resolution. In addition, spectral reconstruction is easily distorted by environmental interference.
The structural parameters of the metasurface unit are correlated with the spectral response, and the neural network optimization design is used to achieve synchronous collaborative optimization of spectral imaging and image reconstruction, and forward prediction, spatial spectrum encoding and decoding networks are used to efficiently process spectral data.
It improves the accuracy and stability of spectral reconstruction, enhances the design freedom, and can achieve high spectral resolution and high signal-to-noise ratio in a compact structure, adapting to multi-scenario applications.
Smart Images

Figure CN120764341A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of micro-nano optical design technology, and in particular relates to a method for synchronously realizing medium-wave infrared spectral imaging metasurface design and spectral image reconstruction. Background Art
[0002] Spectral imaging in the mid-wave infrared band can obtain spatial and spectral information of target objects over a longer wavelength range. Traditional infrared spectral imaging systems usually use spectroscopic elements (such as gratings, prisms) or Fourier transform interferometers to achieve multi-spectral or hyperspectral data acquisition, but the system is large in size, complex in structure and high in cost, which is not conducive to portability and integration. In addition, in order to further meet the application requirements of on-site real-time detection and multi-scene adaptation, the resolution and sensitivity required by the system are also constantly improving. Traditional solutions are difficult to achieve both miniaturization, low cost and high consistency in production while ensuring detection performance.
[0003] In recent years, with the continuous advancement of micro-nanofabrication processes, spectral imaging solutions based on chip-scale technology have experienced rapid development. Mid-wave infrared spectral imaging chips, in particular, integrate spectroscopic elements, detectors, and signal processing modules, enabling efficient spectral data acquisition in a compact design. However, due to the unique requirements of MWIR optical devices in terms of substrate materials, nanofabrication precision, and detection sensitivity, achieving high spectral resolution and high signal-to-noise ratio at the chip scale has become a key technical bottleneck that urgently needs to be addressed.
[0004] In terms of spectral reconstruction, since medium-wave infrared imaging is easily affected by factors such as environmental background thermal radiation, detector noise, and stray light from the optical system when acquiring signals from target objects, the final image spectral information is distorted or degraded in both the spatial and spectral domains. Summary of the Invention
[0005] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a method for synchronously implementing the design of medium-wave infrared spectral imaging metasurface and spectral image reconstruction, which correlates the spectral response of the metasurface unit with the selection of its structural parameters, improves the accuracy and stability of spectral reconstruction by jointly optimizing the detection and imaging processes, and synchronously realizes the optimized design of the metasurface structure, thereby increasing the degree of freedom of design.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] A method for synchronously implementing medium-wave infrared spectral imaging metasurface design and spectral image reconstruction includes the following steps:
[0008] Step 1: Using the structural parameter-spectral response data of the metasurface unit to train a forward prediction network, the forward prediction network takes the structural parameter as input and the spectral response as output; the metasurface unit is a micro-nano structure;
[0009] Step 2: Arrange several metasurface units with different spectral response characteristics on a two-dimensional plane, use a spatial spectrum encoding network to encode the input hyperspectral image through convolution operations to generate a measurement map, and use a spatial spectrum decoding network to recover the target hyperspectral data from the measurement map;
[0010] Step 3: Combine the forward prediction network, spatial spectrum encoding network, and spatial spectrum decoding network for collaborative optimization training. The structural parameters of each metasurface unit are trainable variables, which are automatically iterated and updated through training. The required multi-channel spectral information is obtained in one shot, and the trained spatial spectrum decoding network is used to reconstruct the hyperspectral image from the measurement image.
[0011] In one embodiment, the structural parameter of the metasurface unit, spectral response data, is obtained by:
[0012] Determine the number of spectral channels C of the transmission spectrum and set M structural parameters {α1, α2, ..., α M}, use simulation software to calculate the spectral responses corresponding to different combinations of structural parameter values to form a structural parameter-spectral response database.
[0013] In one embodiment, the material of the metasurface unit is silicon, and the metasurface unit is a circular ring structure. The structural parameters include three parts, namely the outer diameter, inner diameter and height of the ring; the metasurface unit period is set to 3 microns, the outer diameter range is set to 0.8-1.5μm, the inner diameter range is 0-0.7μm, the height range is 0.3-1μm, and the spectral reconstruction range is the mid-infrared band of 3.7-4.8 microns.
[0014] In one embodiment, the input layer dimension of the forward prediction network is equal to the number of structural parameters M, and the output layer dimension is the number of spectral channels C of the transmission spectrum; the structural parameter-spectral response mapping data in the database is used as training samples to update the network weights.
[0015] In one embodiment, the method of arranging a plurality of metasurface units with different spectral response characteristics on a two-dimensional plane is as follows:
[0016] Select n 2A metasurface array constitutes a coding unit, each of the metasurface arrays is composed of several metasurface units with the same spectral response characteristics arranged periodically on a two-dimensional plane. The spectral response characteristics of the metasurface units that make up different metasurface arrays are different. A single coding unit covers n×n detector pixels, and each detector pixel corresponds to a metasurface array; several of the coding units are repeatedly arranged in the row and column directions until the entire detector surface is covered, thereby obtaining a complete coded metasurface.
[0017] In one embodiment, the convolution kernel weights of the convolution operation of the spatial spectrum coding network are provided by the forward prediction network, the number of convolution kernels is consistent with the number of the metasurface arrays, and for the input hyperspectral image, the convolution kernel slides on the two-dimensional plane in an incompletely overlapping manner, and the output intermediate feature maps are spatially spliced to generate a measurement map.
[0018] In one embodiment, the size of each convolution kernel is C×1×1, and the starting positions of each convolution kernel are distributed in n 2 In each area, the convolution kernel is calculated in a sliding manner with a step size of n to implement a local convolution operation; the calculation results of each convolution kernel in a region are weighted summed to obtain the output of the region; the outputs of all regions are spliced according to the original spatial positions to obtain the final output image, namely the measurement map.
[0019] In one embodiment, the outputs of all regions are spliced together according to their original spatial positions, and the implementation method is as follows:
[0020] During the stitching process, the convolution result of each region is inserted into the corresponding image position to form a new feature map, namely the measurement map, which has the same size as the original measurement map. Figure 1 Sample.
[0021] In one embodiment, in step 3, during collaborative optimization network training, the structural parameters of the metasurface units in each metasurface array and the trainable parameters in the spatial spectrum decoding network are synchronously updated; and the matching design of the metasurface structure and the decoding algorithm is simultaneously completed in an end-to-end training process.
[0022] In one embodiment, in step 3, the forward prediction network is used alone after training to quickly evaluate and screen the hypersurface structure under new requirements.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] 1. It can increase design freedom. By co-optimizing metasurface structure design and spectral image reconstruction within the same framework, this invention eliminates the need for a single, fixed simulation design process. Instead, multiple structural parameters (such as geometry, material thickness, and periodic arrangement) are set within the neural network and the network is enabled to learn and adjust autonomously. This approach significantly expands the number of design parameter combinations available, allowing for a larger search space to find the optimal solution that meets specified optical performance and imaging requirements.
[0025] 2. It can improve the reconstruction accuracy of spectral images. The spectral encoding and decoding of the present invention, leveraging the powerful nonlinear fitting capabilities of deep neural networks, can better exploit the correlation between spatial and spectral dimensions. In particular, during the training process of the spatial-spectral encoding and decoding network, the mapping relationship between each metasurface array and the spectral channel is automatically learned, maximizing the information provided by each channel to achieve high-fidelity reconstruction of the spectral image.
[0026] 3. The design results can be adjusted according to the spectral reconstruction network. Under the collaborative design concept of the present invention, the parameter optimization of the metasurface structure and the parameter training of the back-end spectral decoding network are carried out simultaneously. When the network recognizes that the current design scheme cannot meet the high-precision reconstruction requirements, the network will update the structural parameters through backpropagation so that the final metasurface design is deeply matched with the reconstruction network. This dynamic update mechanism avoids the "split" problem that may occur in traditional unidirectional designs, and ensures that the structural design and algorithm decoding remain consistent in target performance.
[0027] 4. The optimality of the structural design can be evaluated from the perspective of images. Unlike traditional methods that only measure structural performance from optical parameters (such as transmittance, dispersion curve, etc.), the present invention uses the quality of the final spectral image as the evaluation criterion, allowing the network to be directly optimized for the target application scenario. During the training process, the design effect can be measured by comparing the difference between the reconstructed image and the real image, so as to more intuitively grasp the performance of the design results in actual imaging or recognition applications. At the same time, it can also flexibly incorporate noise factors and environmental interference conditions to achieve a comprehensive evaluation of optimality in real usage scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is a flow chart of the method for simultaneously implementing spectral imaging metasurface design and spectral image reconstruction according to the present invention.
[0029] Figure 2 Schematic diagram of a forward prediction network in an embodiment of the present invention.
[0030] Figure 3 Schematic diagram of the spatial spectrum coding method in an embodiment of the present invention.
[0031] Figure 4 This is a diagram of the working principle of the spatial spectrum coding network in an embodiment of the present invention. A 4×4 spectrum image is used for demonstration, and each number represents a pixel. DETAILED DESCRIPTION
[0032] In order to make the content of the present invention clearer, the present invention is further described in detail below with reference to specific embodiments and drawings, but the present invention is not limited thereto.
[0033] The method of forward designing the spectral response of a metasurface unit using simulation software is relatively mature, but it is time-consuming and difficult to obtain a metasurface structural unit that perfectly matches an excellent measurement matrix. To this end, the present invention provides a method for synchronously implementing the design of a medium-wave infrared spectral imaging metasurface and spectral image reconstruction, which adopts the correlation mechanism between the structural characteristics of the metasurface micro-nano structural unit and the light field control mode, and combines the space-spectrum encoding and decoding mechanism to realize the collaborative optimization design of the metasurface. It combines micro-nano optical design ideas and algorithms to jointly optimize the detection system and the imaging process, and makes full use of prior information to improve the accuracy and stability of spectral reconstruction. On this basis, through reasonable optical parameter settings, chip structure design and signal processing algorithms, the high performance of the medium-wave infrared spectral imaging chip can be achieved in a smaller volume and power consumption, providing a feasible solution for the new generation of infrared spectral imaging technology.
[0034] Specifically, the present invention uses the spectral response data of different metasurface units within a micro-nanostructure as parameters for the convolutional layer of a neural network, equating the actual image acquisition process to a convolution operation. By training the neural network, a combination of spectral modulation curves suitable for the target system is obtained, while also generating a corresponding spectral reconstruction network. In this invention, the spectral reconstruction network consists of a spatial spectrum encoding network and a spatial spectrum decoding network.
[0035] In addition, the pre-trained high-precision forward prediction network can realize the mapping from the metasurface micro-nanostructure to the transmission spectrum; in the collaborative optimization process, the spectral response data output by it is updated as the convolution layer parameters, so that the structural parameters and the convolution kernel weights are linked to each other, so that the design of the metasurface structure and the optimization of the spectral image reconstruction algorithm can be completed simultaneously in the same training. Since the metasurface design and the spatial spectrum encoding and decoding process are collaboratively iterated under the same neural network framework, the optimal metasurface structure combination that closely matches the spatial spectrum encoding method can be obtained. The present invention can be widely used in various scenarios such as spectral recognition, spectral detection, and spectral imaging based on spatial spectrum coding.
[0036] The specific solution of the present invention is as follows Figure 1 As shown, it mainly includes the following steps:
[0037] Step 1: Build the database.
[0038] The database of the present invention is the structural parameters of the metasurface unit and its spectral response, that is, the structural parameter-spectral response data. The specific construction method is: determine the number of spectral channels C of the transmission spectrum, set M structural parameters {α1, α2, ..., α M}, determine the range λ of spectral reconstruction min ,λ max , then the spectral resolution is (λ max -λ min ) / C, use simulation software to calculate the spectral responses corresponding to different combinations of structural parameter values, and form a structural parameter-spectral response database to provide data support for subsequent training.
[0039] For example, by changing the value of α1 while keeping the values of other structural parameters unchanged, multiple sets of data can be obtained. Alternatively, by keeping the value of α1 unchanged and changing the values of other structural parameters, multiple sets of data can be obtained. In the present invention, structural parameters include but are not limited to the outer diameter, inner diameter, height, period distribution, and material refractive index of the metasurface, and can be expanded to meet the needs of different mid-infrared bands.
[0040] In the present invention, the super surface unit is a micro-nano structure.
[0041] Step 2: Build and train the forward prediction network.
[0042] The forward prediction network takes structural parameters as input and spectral response as output. In practical applications, the forward prediction network is a fully connected neural network. Specifically, a multi-layer perceptron or a convolutional neural network can be used. It is trained by minimizing the loss function between the predicted output and the actual spectral response of the database. Regularization and other strategies can be introduced to improve prediction accuracy and prevent overfitting.
[0043] The forward prediction network has an input layer dimension of M, corresponding to the M structural parameters controlling the micro-nanostructure of the metasurface unit. The output layer dimension of C corresponds to the response values of the metasurface unit in C spectral channels. The combination of these dimensions forms the spectral response curve. Using the structural parameter-spectral response mapping data in the database as training samples, the network weights are updated using the mean squared error (MSE) to accurately predict the spectral response of the metasurface based on its structural parameters. After training, the forward prediction network can quickly output the corresponding spectral response by inputting the metasurface structural parameters.
[0044] The structure of the forward prediction network in the embodiment of the present invention is as follows Figure 2 As shown, it consists of several fully connected layers, where α i Represents the i-th parameter of the metasurface unit, I j represents the spectral response value of the metasurface in the jth spectral channel.
[0045] In the forward prediction network training process, the present application adopts a mean squared error (MSE) loss function. The loss function compares the difference between the predicted spectral response value of the neural network and the true spectral response, calculates the squared error between each group of predicted values and target values, and takes the average of the errors on all output channels as the overall loss value. The smaller the loss value, the closer the prediction result of the model to the true response.
[0046] Step 3, determine the spectral coding mode.
[0047] The present application can adopt various coding modes. Taking one of them as an example, since the metasurface structure is integrated on the camera detector, the light after the metasurface coding is collected by the detector. Therefore, the present application introduces a two-dimensional plane of the detector, as shown in Figure 3 n 2 super surface arrays with different spectral response characteristics are selected to form a coding unit, each super surface array is composed of a plurality of periodic arrangements in a two-dimensional plane, and n can be adjusted as needed, for example, it can be 2, 3, 4, etc. In the figure, n = 2, and the four arrays are super surface array one 101, super surface array two 102, super surface array three 103, and super surface array four 104. The super surface array has different spectral response characteristics, which means that the super surface units that make up it have different spectral response characteristics, while in a single super surface array, the spectral response characteristics of the super surface units that make up it are the same, wherein the spectral response characteristics are embodied by the structural parameters of the super surface units. Among them, a single coding unit covers n x n detector pixels, as shown in the black thick line box in Figure 3 , and each detector pixel corresponds to a super surface array. When the overall pixel of the detector is H x W size (H = 6 and W = 6 are shown in the figure), only a plurality of coding units need to be arranged repeatedly in the row and column directions until the entire detector surface is covered, thereby obtaining a complete coding metasurface structure.
[0048] In this way, the incident light on each detector pixel point is modulated in space and spectral dimensions, and the subsequent combination of the spectral decoding network can efficiently reconstruct the hyperspectral image of the target from the original measurement map.
[0049] Further, the spatial distribution of the super surface unit of the present application can be flexibly adjusted, including different arrangement modes, different element spacings and periods, in order to balance the imaging field of view size, spectral bandwidth and manufacturing process limitations.
[0050] Step 4, construct the spectral coding network and decoding network as the spectral reconstruction network.
[0051] According to the above encoding mode, the input hyperspectral image is encoded by a convolution operation of the spatial-spectral encoding network to generate a single-channel measurement map, and the target hyperspectral data is recovered from the measurement map by using a spatial-spectral decoding network to realize spectral reconstruction.
[0052] The input of the spatial-spectral encoding network is a hyperspectral image, and the image size is HxWxC, wherein H represents the image height, W represents the image width, and C represents the number of image channels. After convolution, a measurement map is output, and the measurement map size is HxWx1. The measurement map is input into the spatial-spectral decoding network, and the final decoding output is a hyperspectral image with a size of HxWxC.
[0053] Further, a plurality of convolution kernels K i (i=1,2...n 2 ) are provided in the spatial-spectral encoding network, the number of convolution kernels is the same as the number of super surface arrays, and the size of each convolution kernel is Cx1x1. According to a predetermined starting position, a local convolution operation is performed on the spatial dimension of the input image. For the input hyperspectral image, the convolution kernel slides in an incomplete overlapping manner on a two-dimensional plane, and the output intermediate feature map is spliced in space to generate a single-channel measurement map. Specifically, the position of the convolution kernel is determined by a predetermined offset, and the starting position is distributed on the spatial positions of (1, 1), (1, 2),..., (n, n-1), (n, n), that is, the corresponding region of the n 2 super surface array. Different convolution kernels operate at different spatial starting positions: within each region, the convolution kernel is calculated by sliding with a step size of n to realize local convolution operation. Each convolution kernel performs weighted summation on the image channels of the region to obtain the output at the corresponding position. The outputs of all regions are spliced according to the original spatial position to obtain the final output image, that is, the measurement map. Specifically, in the splicing process, the convolution result of each region is inserted into the corresponding image position to form a new feature map, and the size of the feature map is the same as the original measurement Figure 1 sample.
[0054] Wherein, the convolution kernel weight of the convolution operation of the spatial-spectral encoding network is provided by the forward prediction network: the structural parameters of the super surface unit {alpha i1 , alpha i2 ,..., alpha iM}(i=1,2,...n 2 ) are taken as trainable parameters in the network, and the bias term of the parameter is set to 0. The parameter is taken as the input of the pre-trained forward prediction network, and the spectral response curve of the n 2 super surface can be obtained through the network, which is taken as the convolution kernel of the spatial-spectral encoding network to realize the spatial-spectral encoding of the input image, wherein alpha iM is the Mth structural parameter of the super surface unit in the ith super surface array.
[0055] The input of the spatial spectral decoding network is a measurement map with dimensions of H × W × 1. The output is a reconstructed hyperspectral image. This network can employ an end-to-end convolutional neural network or an attention mechanism network. By comparing with the true hyperspectral image, the decoding weights are iteratively updated using gradient descent and backpropagation algorithms. In this embodiment, a UNet network is used to learn the nonlinear mapping relationship between the measurement map and the true hyperspectral image. Furthermore, during the iterative process, different noise or interference models can be introduced into the spatial spectral decoding network to improve environmental robustness.
[0056] Step 5: Collaborative training.
[0057] The present invention combines the forward prediction network, the space spectrum encoding network, and the space spectrum decoding network to obtain a collaborative optimization network, such as Figure 1 As shown, it is then trained through collaborative optimization. In the collaborative optimization network, the structural parameters of each hypersurface unit {α1,α2,...,α M} are regarded as trainable variables. The forward prediction network uses these structural parameters to output the spectral response curve and connects it with the back-end decoding network in the spatial spectrum encoding process. The collaborative design network is trained using a known hyperspectral image dataset. By comparing the differences between the reconstructed spectrum output by the network and the real image, the structural parameters of the metasurface units in each metasurface array and the trainable parameters in the spatial spectrum decoding network are synchronously updated. The matching design of the metasurface structure and the decoding algorithm is completed simultaneously in a single end-to-end training process. The required multi-channel spectral information is obtained in a single shot, and the trained spatial spectrum decoding network is then used to reconstruct the hyperspectral image from the measurement image. The forward prediction network can also be used alone to quickly evaluate and screen metasurface structures under new requirements, reducing the high computational cost of multiple simulations.
[0058] During this process, the metasurface structure and decoding algorithm are continuously optimized through end-to-end back propagation based on the modulation characteristics of different spectral channels and the adaptability constraints of the manufacturing process, so that the resulting chip has high-precision spectral imaging capabilities in the mid-infrared band.
[0059] Furthermore, the present invention adopts a comprehensive loss function containing two constraints when performing end-to-end training on the network to ensure both the accuracy of spectral reconstruction and the manufacturability of the metasurface unit structure parameters. Specifically, the loss function can be expressed as:
[0060] L = a × L recon +b×L param
[0061] Among them, L reconis the image reconstruction loss, which measures the difference between the reconstructed image and the true hyperspectral image, ensuring the overall reconstruction accuracy in the spectral and spatial domains; param This is a structural parameter constraint penalty, used to penalize situations where the metasurface's geometric or material parameters exceed preset tolerances or violate processing requirements. By setting appropriate weight coefficients a and b, the resulting metasurface unit structural parameters can be made to conform to pre-set physical or manufacturing constraints while maintaining spectral reconstruction accuracy.
[0062] During the actual training process, the network will continuously update the weights of each layer including the hypersurface structure parameters through back propagation based on the comprehensive loss function. When the training is completed, the parameters {α i1 ,α i2 ,...,α iM} is the metasurface unit structure parameter in the i-th metasurface array optimized for the spectral dataset and the spatial spectrum decoding network. After the training is completed, n 2 Group of metasurface structure parameters {α i1 ,α i2 ,...,α iM}, i=1,2,...n 2 By inputting this set of parameters into the forward prediction network, the spectral response curve of the metasurface unit in the i-th metasurface array can be optimized. Thus, in a single end-to-end collaborative training, the simultaneous optimization of the micro-nanostructure design and the spectral image decoding algorithm is completed, ensuring high-quality spectral reconstruction while meeting the constraints in the actual manufacturing and use of metasurfaces.
[0063] Finally, the spatial-spectral decoding network within the collaborative design network is independently extracted and used as a reconstruction model for the actual spectral imaging system. When the original measurement map is acquired through the designed metasurface, it is fed into the spatial-spectral decoding network to restore the target hyperspectral image with the pre-designed number of spectral channels, providing support for spectral recognition and detection.
[0064] According to the method of the present invention, its design process is not only limited to the mid-wave infrared band, but can also be extended to other bands such as visible light, near-infrared or terahertz. By changing the material and structural parameters and reshaping the forward prediction network and spatial spectrum coding network, spatial spectrum coding imaging can be achieved in different spectral ranges.
[0065] In a specific embodiment of the present invention, the method for synchronously implementing the metasurface design and spectral image reconstruction is as follows:
[0066] (1) The super surface material is determined to be silicon, and the substrate is sapphire. The super surface unit is determined to be a circular ring structure, and the structure parameters include three parts, namely the outer diameter, the inner diameter and the height of the circular ring. The spectral reconstruction range of the mid-infrared waveband is determined to be 3.7-4.8 microns, and the number of spectral channels is determined to be 33.
[0067] (2) Obtain a plurality of spectral response data sets of the super surface unit using simulation software. The super surface unit period is set to 3 microns, the outer diameter range is set to 0.8-1.5 microns, the inner diameter range is set to 0-0.7 microns, and the height range is set to 0.3-1 microns. Each parameter is set to have 15 equally spaced sampling points, and a total of 3375 groups of data are generated.
[0068] (3) Construct a forward prediction network, set the number of input units of the network to 3, which corresponds to the three parameters of the super surface structure, and set the number of output units to 33, which corresponds to the number of spectral channels. The data set obtained by the simulation software is used to train the forward prediction network, so as to achieve the purpose of predicting the spectral response of the super surface structure according to the super surface structure.
[0069] (4) Confirm that the number of super surface structure units is 4, and a 2x2 periodic unit array is formed to encode the hyperspectral image. The structure parameters of the i-th super surface unit {α i1 ,α i2 ,α i3}(i=1,2,3,4) are input into the trained forward prediction network to obtain the spectral response curve.
[0070] (5) Construct a spectral-spatial encoding network, and use the spectral response curve of the super surface as the parameter K i (i=1,2,3,4) of the convolution layer in the network. The dimension of K is 33x1x1, and a local convolution operation is performed on the spatial dimension of the input image according to the predetermined starting position. Figure 4 The basic operation of the convolution kernel acting on a 33x4x4 image is shown, specifically, the position of the convolution kernel is determined by the predetermined offset, and the four convolution kernels operate at different spatial positions: the first convolution kernel acts on the region with a starting position of (1,1). The second convolution kernel acts on the region with a starting position of (1,2). The third convolution kernel acts on the region with a starting position of (2,1). The fourth convolution kernel acts on the region with a starting position of (2,2). In each region, the convolution kernel calculates by sliding with a step size of 2. Each convolution kernel performs weighted summation on the image channels in the region to obtain the output at the corresponding position. For the spectral image of 33x128x128 in this example, the principle is the same, that is, the convolution operation is performed according to the corresponding starting position and step size, thereby realizing the spectral-spatial encoding of the input image.
[0071] (6) Stitching results: For each convolution kernel output, stitch it back to the final original measurement map according to its original spatial position. During the stitching process, the convolution result of each region is inserted into the corresponding image position to form the original measurement map. The original measurement dimension is 1×128×128.
[0072] (7) Construct a spatial spectrum decoding network, take the original measurement image as input, and output the reconstructed hyperspectral image.
[0073] (8) Construct a collaborative optimization network to combine the above-mentioned space spectrum coding network, space spectrum encoding network and space spectrum decoding network, such as Figure 1 As shown. Use appropriate spectral images to train the network and continuously update the network parameters {α i1 ,α i2 ,α i3}, when the training is finished, the parameters {α i1 ,α i2 ,α i3 By inputting this set of parameters into the forward prediction network, the optimized metasurface spectral response curve can be obtained.
[0074] (9) The metasurface filter array produced using the parameters and corresponding materials is placed in front of the detector. The light emitted by the object to be measured passes through the metasurface array and is imaged onto the detector to obtain the original measurement image. The original measurement image is input into the trained spatial spectrum decoding network to reconstruct the spectral image. This device can be used for computational spectral imaging in the mid-infrared band.
Claims
1. A method for synchronously implementing the design of a medium-wave infrared spectral imaging metasurface and spectral image reconstruction, characterized in that: The steps include: Step 1: Using the structural parameter-spectral response data of the metasurface unit to train a forward prediction network, the forward prediction network takes the structural parameter as input and the spectral response as output; the metasurface unit is a micro-nano structure; Step 2: Arrange several metasurface units with different spectral response characteristics on a two-dimensional plane, use a spatial spectrum encoding network to encode the input hyperspectral image through convolution operations to generate a measurement map, and use a spatial spectrum decoding network to recover the target hyperspectral data from the measurement map; Step 3: Combine the forward prediction network, spatial spectrum encoding network, and spatial spectrum decoding network for collaborative optimization training. The structural parameters of each metasurface unit are trainable variables, which are automatically iterated and updated through training. The required multi-channel spectral information is obtained in one shot, and the trained spatial spectrum decoding network is used to reconstruct the hyperspectral image from the measurement image.
2. The method for synchronously implementing the design of a medium-wave infrared spectral imaging metasurface and spectral image reconstruction according to claim 1 is characterized in that: The structural parameter of the metasurface unit, spectral response data, is obtained by: Determine the number of spectral channels C of the transmission spectrum and set M structural parameters {α1, α2, ..., α M }, use simulation software to calculate the spectral responses corresponding to different combinations of structural parameter values to form a structural parameter-spectral response database.
3. The method for synchronously implementing the design of a medium-wave infrared spectral imaging metasurface and spectral image reconstruction according to claim 1 or 2, characterized in that: The material of the metasurface unit is silicon, and the metasurface unit has a circular ring structure. The structural parameters include three parts, namely the outer diameter, inner diameter and height of the ring; the metasurface unit period is set to 3 microns, the outer diameter range is set to 0.8-1.5μm, the inner diameter range is 0-0.7μm, the height range is 0.3-1μm, and the spectral reconstruction range is the mid-infrared band of 3.7-4.8 microns.
4. The method for synchronously implementing the design of a medium-wave infrared spectral imaging metasurface and spectral image reconstruction according to claim 1 is characterized in that: The input layer dimension of the forward prediction network is equal to the number M of structural parameters, and the output layer dimension is the number C of spectral channels of the transmission spectrum; the structural parameter-spectral response mapping data in the database is used as training samples to update the network weights.
5. The method for synchronously implementing the design of a medium-wave infrared spectral imaging metasurface and spectral image reconstruction according to claim 1 is characterized in that: The method of arranging a plurality of metasurface units with different spectral response characteristics on a two-dimensional plane is as follows: Select n 2 A metasurface array constitutes a coding unit, each of the metasurface arrays is composed of several metasurface units with the same spectral response characteristics arranged periodically on a two-dimensional plane. The spectral response characteristics of the metasurface units that make up different metasurface arrays are different. A single coding unit covers n×n detector pixels, and each detector pixel corresponds to a metasurface array; several of the coding units are repeatedly arranged in the row and column directions until the entire detector surface is covered, thereby obtaining a complete coded metasurface.
6. The method for synchronously implementing the design of a medium-wave infrared spectral imaging metasurface and spectral image reconstruction according to claim 5, characterized in that: The convolution kernel weights of the convolution operation of the spatial spectrum coding network are provided by the forward prediction network. The number of convolution kernels is consistent with the number of the metasurface arrays. For the input hyperspectral image, the convolution kernel slides on the two-dimensional plane in an incompletely overlapping manner, and the output intermediate feature maps are spatially spliced to generate a measurement map.
7. The method for synchronously implementing the design of a medium-wave infrared spectral imaging metasurface and spectral image reconstruction according to claim 6, characterized in that: The size of each convolution kernel is C×1×1, and the starting positions of each convolution kernel are distributed in n 2 In each area, the convolution kernel is calculated in a sliding manner with a step size of n to implement a local convolution operation; the calculation results of each convolution kernel in a region are weighted summed to obtain the output of the region; the outputs of all regions are spliced according to the original spatial positions to obtain the final output image, namely the measurement map.
8. The method for synchronously implementing the design of a medium-wave infrared spectral imaging metasurface and spectral image reconstruction according to claim 7, characterized in that: The output of all regions is spliced according to the original spatial position. The implementation method is as follows: During the stitching process, the convolution result of each region is inserted into the corresponding image position to form a new feature map, namely the measurement map.
9. The method for synchronously implementing the design of a medium-wave infrared spectral imaging metasurface and spectral image reconstruction according to claim 1, characterized in that: In step 3, during collaborative optimization network training, the structural parameters of the metasurface units in each metasurface array and the trainable parameters in the spatial spectrum decoding network are synchronously updated; and the matching design of the metasurface structure and the decoding algorithm is simultaneously completed in an end-to-end training process.
10. The method for synchronously implementing the design of a medium-wave infrared spectral imaging metasurface and spectral image reconstruction according to claim 1, characterized in that: In step 3, the forward prediction network is used alone after training to quickly evaluate and screen the hypersurface structure under new requirements.
Citation Information
Cited By
Metasurface structure for realizing wide-spectrum measurement base and design method
CN121677935A