Physically driven mid-infrared spectral encoding and reconstruction method
By employing a physics-driven end-to-end deep learning network architecture and a dual-branch spectral reconstruction method, the accuracy and speed issues in mid-wave infrared spectral reconstruction are addressed, achieving high-precision and rapid reconstruction of mid-wave infrared spectra, which is suitable for gas monitoring and on-site analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI
- Filing Date
- 2026-05-09
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies for mid-wave infrared spectral reconstruction suffer from low accuracy, slow speed, and weak decoupling capabilities. In particular, they are difficult to achieve high-precision gas monitoring when there is a lack of large-scale, high-quality training datasets and when the characteristics of mid-wave infrared spectral signals are not adapted.
A physics-driven end-to-end deep learning network architecture is adopted, which combines the Lambert-Beer law, spectral logarithmic transform and signal physical decomposition of mid-wave infrared spectroscopy to design a dual-branch spectral reconstruction network. Through attention-enhanced gas absorption feature extraction module and negative logarithmic transform, targeted reconstruction of smooth background and sharp gas features is achieved.
It achieves high-precision and rapid reconstruction of mid-wave infrared spectra, meets the requirements of qualitative and quantitative gas analysis, improves the model's ability to reconstruct the absorption spectra of complex mixed gases, and enhances the real-time performance of infrared spectral imaging.
Smart Images

Figure CN122492841A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mid-wave infrared spectral imaging technology, and particularly relates to a physical-driven mid-wave infrared spectral encoding and reconstruction method. Background Technology
[0002] Mid-wave infrared spectroscopy imaging technology can simultaneously acquire image information and pixel-level spectral information of the target scene in the molecular fingerprint spectral region. With the unique advantage of gas characteristic absorption peaks corresponding to molecular vibration modes, it can achieve accurate qualitative and quantitative analysis of gases, becoming a core technical means for regional scale gas emission monitoring.
[0003] Currently, research on spectral reconstruction based on spectral encoding and decoding is mostly concentrated in the visible light band. Migrating mature visible light band spectral reconstruction algorithms to mid-infrared gas monitoring faces significant technical bottlenecks: First, data-driven algorithms based on deep learning heavily rely on large-scale, high-quality training datasets. However, the mid-infrared band, as a core application band for gas monitoring, currently lacks a sufficiently large number of publicly available spectral data cubes, making it difficult to support effective training of deep learning models. Second, the signal composition of mid-infrared gas absorption spectra has significant unique characteristics, essentially a combination of smooth background radiation spectra and sharp gas characteristic absorption spectra. Under the constraints of high-sparseness compressed measurement, reconstruction algorithms adapted to the visible light band are not designed for this signal characteristic, failing to achieve high-precision reconstruction of sharp substance fingerprint features in the mid-infrared spectrum. Consequently, the accuracy of existing deep learning methods in reconstructing mid-infrared spectra falls far short of the actual needs of gas monitoring, becoming a key issue restricting the further development and practical application of mid-infrared compressed sensing spectral imaging technology.
[0004] Therefore, there is an urgent need to develop a spectral compression coding and reconstruction method that adapts to the characteristics of mid-wave infrared signals, break through the reconstruction accuracy bottleneck of existing technologies, achieve high-precision and rapid reconstruction of sharp gas absorption spectra in mid-wave infrared, and meet the practical application needs of on-site gas emission monitoring. Summary of the Invention
[0005] In view of this, the present invention aims to provide a physical-driven mid-wave infrared spectral encoding and reconstruction method to solve the problems of low reconstruction accuracy of sharp gas absorption spectra in mid-wave infrared and weak decoupling ability of mixed components in the existing technology. The present invention achieves high-precision and rapid reconstruction of mid-wave infrared gas absorption spectra.
[0006] To achieve the above objectives, the technical solution created by this invention is implemented as follows: A physical-driven method for mid-wave infrared spectral encoding and reconstruction includes the following steps: S1: Obtain the mid-wave infrared spectral signal of the target scene and construct a forward modeling network. Input the thickness of each film layer of each filter contained in the filter array into the forward modeling network for processing to obtain the spectral transmittance matrix. S2: Multiply the mid-wave infrared spectral signal with the spectral transmittance matrix to obtain the compressed measurement result; S3: Input the compressed measurement results into the dual-branch spectral reconstruction network for processing to obtain the smooth background logarithmic spectral components and characteristic absorption absorbance components; S4: The logarithmic spectral components of the smooth background and the characteristic absorption absorbance components are superimposed to obtain the logarithmic spectrum. The logarithmic spectrum is then processed by a post-processing network to output the reconstructed spectral signal.
[0007] Furthermore, in step S3, the bi-branch spectral reconstruction network includes a smooth background logarithmic spectral reconstruction branch and a concentration range product prediction network.
[0008] Furthermore, the forward modeling network, the smooth background log-spectral reconstruction branch, and the post-processing network are all pre-trained fully connected networks.
[0009] Furthermore, the concentration range product prediction network includes an initial convolutional block, a multi-channel self-attention module, a first residual block, a second residual block, a global average pooling module, and a classifier. The compressed measurement results are processed sequentially by the initial convolutional block, the multi-channel self-attention module, the first residual block, the second residual block, the global average pooling module, and the classifier to obtain the concentration range product containing various gas components. The concentration range product of various gas components is multiplied by the transposed gas absorbance matrix to obtain the characteristic absorption absorbance components.
[0010] Furthermore, in the initial convolutional block, the input features are sequentially processed by connected one-dimensional convolution, batch normalization, LeakyReLU activation, and max pooling to obtain the output features.
[0011] Furthermore, the first residual block and the second residual block have the same structure. The first residual block includes two cascaded convolutional blocks, an efficient channel attention module, and a LeakyRelu activation module. The feature A1 input to the first residual block is sequentially processed by the two cascaded convolutional blocks and the efficient channel attention module for local feature extraction and channel adaptive enhancement to obtain feature A2. Feature A1 and feature A2 are then connected by a skip connection and input to the LeakyRelu activation module for processing to obtain feature A3.
[0012] Furthermore, in the classifier, the input features are sequentially processed through a linear fully connected layer and activated by LeakyReLU to obtain the output features.
[0013] Furthermore, the logarithmic spectrum is transformed by inverse negative exponential transform before being input into the post-processing network.
[0014] Compared with the prior art, the present invention can achieve the following beneficial effects: (1) The physical-driven mid-wave infrared spectral encoding and reconstruction method described in this invention adopts a physical-driven end-to-end deep learning network architecture to achieve accurate reconstruction of mid-wave infrared spectra. The core physical mechanisms of mid-wave infrared spectral imaging, such as the Lambert-Beer law, spectral logarithmic transformation, and signal physical decomposition, are deeply integrated into the network architecture design, rather than simply superimposing physical constraints. This achieves a combination of physical-driven interpretability and data-driven generalization capability, specifically adapted to the reconstruction task of mid-wave infrared gas absorption spectra. Based on the physical composition characteristics of mid-wave infrared spectra—"smooth background logarithmic spectrum + sharp gas characteristic absorbance"—a dual-branch mid-wave infrared spectral reconstruction network is designed to perform targeted reconstruction of the two types of components separately.
[0015] (2) The physical-driven mid-wave infrared spectral encoding and reconstruction method described in this invention constructs an attention-enhanced gas absorption feature extraction module (multi-channel self-attention module, first residual block, and second residual block) in the concentration range product prediction network. By integrating the multi-channel self-attention module and embedding a one-dimensional residual block with efficient channel attention (ECA), the global dependency capture and channel-specific feature enhancement of short-sequence spectral compression measurement data are achieved, effectively decoupling the absorption signals of different components in the mixed gas system.
[0016] (3) The physical-driven mid-wave infrared spectral encoding and reconstruction method described in this invention proposes a method for converting and processing physical quantities in mid-wave infrared spectroscopy. It utilizes negative logarithmic transformation to convert the nonlinear Hadamard product operation in the gas modulation process into a linear addition operation of the logarithmic spectrum, solving the technical challenge of backpropagation of neural network gradients. Simultaneously, it combines the Lambert-Beer law to achieve a linear solution for the absorbance of multi-component mixed gases, reducing the difficulty of model training. The smooth background logarithmic spectral components and characteristic absorption absorbance components from the dual-branch output are linearly added, and the logarithmic spectrum is converted into the original spectral vector through negative exponential inverse transformation. Finally, a post-processing network is used to complete the fine optimization of the reconstructed spectrum.
[0017] (4) The physical-driven mid-wave infrared spectral encoding and reconstruction method described in this invention provides a simple, low-cost, and lightweight mid-wave infrared snapshot spectral imaging scheme. It fundamentally solves the core problems of low accuracy, slow speed, and weak decoupling ability in existing mid-wave infrared spectral reconstruction technologies. Based on the characteristics of mid-wave infrared spectral data, a targeted mid-wave infrared spectral reconstruction network, namely a dual-branch spectral reconstruction network, is designed to achieve accurate mid-wave infrared spectral reconstruction, meeting the requirements of qualitative and quantitative gas analysis. An attention-enhanced one-dimensional residual feature extraction module, namely the first residual block and the second residual block, is constructed to capture the global dependency relationship of short-sequence compressed measurement results and enhance channel-specific features, effectively decoupling and identifying the absorption signals of different components in a mixed gas system, and improving the model's ability to reconstruct the absorption spectra of complex mixed gases. This invention utilizes the parallel computing characteristics of graphics cards to achieve fast, scene-level spectral reconstruction, improving the real-time performance of infrared spectral imaging. Attached Figure Description
[0018] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments and descriptions of the invention are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings: Figure 1 A schematic flowchart of the physical-driven mid-wave infrared spectral encoding and reconstruction method described in the embodiments of the present invention; Figure 2 A schematic diagram of the structure of the mid-wave infrared spectral imaging system based on spectral compression coding as described in the embodiment of the present invention; Figure 3 A schematic diagram of the processing flow of mid-wave infrared spectral signals described in the embodiments of the present invention; Figure 4 A schematic diagram of the processing flow of the physical-driven mid-wave infrared spectral encoding and reconstruction method described in the embodiments of the present invention; Figure 5 A schematic diagram of the concentration range product prediction network described in an embodiment of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not constitute a limitation thereof.
[0020] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.
[0021] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are used only for the convenience of describing this invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on this invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.
[0022] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0023] The invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0024] The core of this invention is to complete the overall technical solution design based on the technical requirement of high-precision reconstruction of sharp gas absorption spectra in the mid-wave infrared.
[0025] like Figure 1 As shown, this invention proposes a physical-driven mid-wave infrared spectral encoding and reconstruction method, which specifically includes the following steps: S1: Obtain the mid-wave infrared spectral signal of the target scene and construct a forward modeling network. Input the thickness of each film layer of each filter contained in the filter array into the forward modeling network for processing to obtain the spectral transmittance matrix. Each filter has the same number of optical thin film layers.
[0026] S2: Multiply the mid-wave infrared spectral signal with the spectral transmittance matrix to obtain the compressed measurement result, simulating the actual spectral measurement process; S3: Input the compressed measurement results into the dual-branch spectral reconstruction network for processing to obtain the smooth background logarithmic spectral components and characteristic absorption absorbance components; S4: The logarithmic spectral components of the smooth background and the characteristic absorption absorbance components are superimposed to obtain the logarithmic spectrum. The logarithmic spectrum is then processed by a post-processing network to output the reconstructed spectral signal.
[0027] It should be noted that, Figure 2 This demonstrates the process of cube compression of scene spectral data using a filter array. The optical system includes an imaging lens, a filter array for multispectral detection, and a mid-infrared detector. The optical system images the spectral data cube of the target scene, and the projection of the spectral data cube is obtained on the mid-infrared detector. In this process, such as... Figure 3 As shown, the filter array modulates and encodes the spectral information of the target scene, thereby acquiring the encoded spectral image at the mid-wave infrared detector. For the encoded spectral image, different pixels represent the modulation and integration results of different filters. Using the light intensity information at the corresponding pixels of each filter in the encoded spectral image as the compressed measurement result, the reconstructed spectral data cube can be obtained by applying the physical-driven mid-wave infrared spectral encoding and reconstruction method proposed in this invention.
[0028] Mid-wave infrared spectral signals can essentially be decomposed into a combination of smooth background signals and sharp absorption signals. According to the Lambert-Beer law, for a gas mixture containing multiple components, it is only necessary to solve for the concentration-range product of each gas component and multiply it by the gas absorbance matrix at unit concentration and unit optical path length. The gas absorbance matrix is a pre-defined fixed matrix. Absorbance data for relevant gases can be obtained from the HITRAN on the web website. The gas absorbance matrix is an m×N matrix, where m is the number of gas types and N is the number of spectral channels. Multiplying the transpose of the gas absorbance matrix by the concentration-range product prediction network's output (m×1) yields an N×1 vector, which represents the total absorbance of the gas mixture system (i.e., the characteristic absorption absorbance component), with dimensions consistent with the logarithmic spectral components of the smooth background. Furthermore, in neural networks, element-wise vector multiplication is detrimental to gradient backpropagation. To address this problem, this invention utilizes the properties of logarithmic transformation to convert nonlinear element-wise multiplication into linear addition, thereby achieving the fusion of different types of features and creating favorable conditions for gradient backpropagation in neural networks.
[0029] Furthermore, assuming the gas contains carbon dioxide, carbon monoxide, nitrous oxide, and sulfur dioxide, then the gas composition m=4, the number of spectral channels is taken as mid-wave infrared discrete spectral channels N=6, and the gas absorbance matrix K: ; in, Let be the standard absorbance of the i-th gas per unit concentration and per unit optical path in the j-th spectral channel, where i ranges from 1 to 4 and j ranges from 1 to 6.
[0030] First, the mid-wave infrared spectral signal is multiplied by the spectral transmittance matrix corresponding to the filter array to obtain the corresponding compressed measurement result. This compressed measurement result is input into a two-branch spectral reconstruction network, which reconstructs the logarithmic spectral components of the smoothed background and the characteristic absorption and absorbance components through a smoothed background logarithmic spectral reconstruction branch and a concentration range product prediction network, respectively. Then, through a two-branch information fusion, logarithmic spectrum to spectral physical quantity conversion, and post-processing network (fully connected network), the reconstructed spectral signal, i.e., the reconstructed mid-wave infrared spectral signal, is finally output.
[0031] Furthermore, in step S3, the bi-branch spectral reconstruction network includes a smooth background logarithmic spectral reconstruction branch and a concentration range product prediction network.
[0032] It should be noted that, in order to achieve high-accuracy reconstruction of mid-wave infrared gas absorption spectral signals, such as Figure 4 As shown, the dual-branch spectral reconstruction network is a one-dimensional spectral signal reconstruction neural network. Its core design principle is to prioritize the physical links and signal characteristics of mid-wave infrared spectral imaging over simply adding physical constraints. By deeply integrating the core physical mechanisms into the network architecture, this network is specifically adapted for mid-wave infrared gas absorption spectral reconstruction tasks with smooth backgrounds and sharp absorption characteristics.
[0033] First, a forward modeling network is used to construct the mapping relationship between the thickness of each layer of the filter and its spectral transmittance, thus solving for the spectral transmittance matrix of the filter array. The mid-wave infrared spectral signal is multiplied by the spectral transmittance matrix of the filter array to obtain the corresponding compressed measurement result. This compressed measurement result is input into a dual-branch spectral reconstruction network, where the smooth background logarithmic spectral components and the characteristic absorption absorbance components are reconstructed through a smooth background reconstruction branch and a characteristic absorption reconstruction branch (i.e., the branch containing the concentration-range-product prediction network). The smooth background absorbance reconstruction branch is a fully connected network composed of multiple fully connected layers; the concentration-range-product prediction network outputs the concentration-range-product of various gas components, which, when multiplied by the transposed gas absorbance matrix, yields the characteristic absorption reconstruction result, i.e., the characteristic absorption absorbance component. The dual-branch reconstruction results are superimposed through information fusion, and then processed through a physical quantity conversion and fully connected post-processing network to finally output the reconstructed mid-wave infrared spectral signal.
[0034] In some embodiments, the forward modeling network, the smooth background logarithmic spectral reconstruction branch, and the post-processing network are all fully connected networks.
[0035] Among them, the forward modeling network needs to be pre-trained.
[0036] In some embodiments, the concentration range product prediction network comprises an initial convolutional block, a multi-channel self-attention module, a first residual block, a second residual block, a global average pooling module, and a classifier. The compressed measurement results are processed sequentially by the initial convolutional block, the multi-channel self-attention module, the first residual block, the second residual block, the global average pooling module, and the classifier to obtain the concentration range product containing various gas components. The concentration range product of various gas components is multiplied by the gas absorbance matrix to obtain the characteristic absorption absorbance components.
[0037] In some embodiments, in the initial convolutional block, the input features are sequentially subjected to connected one-dimensional convolution, batch normalization, LeakyReLU activation, and max pooling operations to obtain the output features.
[0038] In some embodiments, the first residual block and the second residual block have the same structure. The first residual block includes two cascaded convolutional blocks, an efficient channel attention module, and a LeakyRelu activation module. The feature A1 input to the first residual block is sequentially processed by the two cascaded convolutional blocks and the efficient channel attention module for local feature extraction and channel adaptive enhancement to obtain feature A2. Feature A1 and feature A2 are then skip-connected and input to the LeakyRelu activation module for processing to obtain feature A3.
[0039] It should be noted that the two cascaded convolutional blocks have the same structure. The features input to the convolutional blocks are sequentially processed by one-dimensional convolution, batch normalization, and LeakyReLU activation to obtain the output features.
[0040] In some embodiments, in the classifier, the input features are sequentially passed through a linear fully connected layer and activated by LeakyReLU to obtain the output features.
[0041] The feature absorption reconstruction branch, specifically the branch containing the concentration-range-length product prediction network, is the core branch for achieving high-precision reconstruction of gas feature absorption. Its core module is an attention-enhanced one-dimensional residual concentration-range-length product prediction network, such as... Figure 5As shown, the concentration range product prediction network can accurately extract the characteristic absorption information of gas components from the physical input encoded by the filter and output the concentration range product of each gas. Its main components include an initial convolutional block, a multi-channel self-attention block, two attention-enhanced one-dimensional residual blocks, two one-dimensional residual blocks (i.e., the first residual block and the second residual block, which have different numbers of channels), and a classifier. The input data is the compressed measurement result of the mid-wave infrared spectral signal through the filter array. The initial convolutional block transforms the compressed measurement result to the target number of channels. The multi-channel self-attention module is used to enhance the feature recognition capability of the concentration range product prediction network. The first residual block is used to strengthen the channel-specific expression of spectral features and improve the targeting and effectiveness of feature extraction. The second residual block has a different number of channels than the first residual block. The global average pooling module is used to implement global average pooling. The classifier is used to output the concentration range product of various gas components.
[0042] To enhance the channel-specific expression of spectral features, this invention designs two attention-enhanced one-dimensional residual blocks as the core building blocks of a one-dimensional concentration range product prediction network. Both the first and second residual blocks embed efficient channel attention (ECA) modules on the basis of traditional residual blocks, realizing integrated learning of local feature extraction and channel adaptive enhancement, and improving the targeting and effectiveness of feature extraction.
[0043] Furthermore, spectral encoding can be achieved not only through thin-film filters or filter arrays, but also through new technologies such as metasurfaces, photonic crystals, and quantum dots. Additionally, attention modules such as Convolutional Block Attention Modules (CBAM), Squeezed and Excited Attention Modules (SENet), and simplified versions of self-attention mechanisms (such as single-head attention and linear attention) can replace multi-head channel self-attention modules and efficient channel attention modules.
[0044] In some embodiments, the logarithmic spectrum is transformed by an inverse negative exponential transform before being input into the post-processing network. To verify the technical effect of the present invention, this embodiment conducted mid-wave infrared gas absorption spectrum reconstruction in the 3.7~4.8μm band with 45 channels. The specific implementation details are as follows: (a) Basic parameter settings for the embodiment 1. Spectral parameters: target wavelength 3.7~4.8μm, number of spectral channels 45, spectral resolution 25nm.
[0045] 2. Parameters of the filter array: The filter array includes 9 broadband filter units arranged in a 3×3 pattern. Each broadband filter unit has 8 optical thin films with a thickness range of 250~650nm. The structure is an alternating structure of high and low refractive index films. The materials are magnesium fluoride (low refractive index) and titanium dioxide (high refractive index), and the substrate is sapphire.
[0046] 3. Gas parameters: The target gases are nitrous oxide, carbon monoxide, sulfur dioxide, and carbon dioxide, covering a volume fraction of 0.1% to 100%. Two scenarios are set up: pure target gas absorption and target gas-air mixed absorption (0.6m optical path of air column and 0.3m optical path of gas cell).
[0047] (II) Implementation Steps 1. Pre-training of the filter forward modeling network: Based on the transfer matrix theory, 3 million sets of filter structure-transmittance data are generated to train a six-layer fully connected network, i.e., to train the forward modeling network.
[0048] 2. Construction of the mid-wave infrared spectroscopy dataset: 110,000 smooth Gaussian background spectra in the 3.7~4.8μm band were generated. The gas absorption spectra were obtained by multiplying them with the gas transmittance. The training set and validation set were divided at a ratio of 10:1. The test set was obtained by multiplying the blackbody radiation spectrum in the 403~503K band with the gas transmittance.
[0049] 3. For example Figure 3 The overall model within the dashed box shown is jointly trained. The overall model includes a dual-branch spectral reconstruction network and a post-processing network. A pre-trained forward modeling network is loaded, and the dual-branch spectral reconstruction network is built. The smooth background logarithmic spectral reconstruction branch is set as a network consisting of 8 fully connected layers. The two residual blocks in the concentration range product prediction network are set to 64 channels and 128 channels respectively. The post-processing network is a three-layer fully connected network. The joint optimization of the filter array and the dual-branch spectral reconstruction network is completed to obtain the optimal model and the structure of the filter array with the best performance.
[0050] The loss function consists of two parts: the mean square error term for spectral reconstruction and the filter structure regularization term. If there are a total of... Each filter has [a certain feature]. Layered optical thin films, and the loss function used to jointly train the overall model. for: ; in, The mean square error of the spectral reconstruction; These are the filter structure regularization parameters, set here to... ; For filter structure regularization function; For the first The first filter The thickness of the optical thin film, where N is the number of spectral channels. The input network consists of mid-infrared spectral signals. This is the reconstructed mid-wave infrared spectral signal.
[0051] To ensure the stability and gradient continuity of model training, the filter structure regularization function... The soft-plus function, which uses a continuous gradient, is used for construction. Its expression is: ; in, The thickness of the optical thin film. and These are the preset minimum and maximum film thicknesses, used to constrain the film thickness within a reasonable process range; The softplus function has the following expression: ; in, Here, the shape parameter is used to control the smoothness of the function. Set as This value ensures that the function has an appropriate degree of smoothness at the film thickness boundary, balancing gradient continuity and the accuracy of thickness constraints. is the independent variable.
[0052] When the film thickness is Within this range, the regularization term has a small value and a weak impact on the loss function. Once the thickness exceeds this range, the regularization term increases rapidly, thus creating a strong constraint on model parameter updates. Compared to hard threshold regularization, soft threshold regularization maintains gradient continuity at the thickness boundary, effectively avoiding training loss oscillations caused by abrupt gradient changes, while also providing more precise constraints near the boundary values.
[0053] 4. Conduct spectral reconstruction simulation and experiments: Input the real spectra of the test set into the trained array, such as... Figure 4 The entire model shown simulates the compression encoding and reconstruction process, analyzing the reconstruction effect under different absorption system types and temperatures. An experimental system was built, and the system's spectral response was calibrated using the least squares method. Actual gas absorption spectrum compression measurement data were collected to complete the actual absorption spectrum reconstruction.
[0054] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0055] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A physical-driven method for encoding and reconstructing mid-wave infrared spectra, characterized in that: Specifically, the steps include the following: S1: Obtain the mid-wave infrared spectral signal of the target scene and construct a forward modeling network. Input the thickness of each film layer of each filter contained in the filter array into the forward modeling network for processing to obtain the spectral transmittance matrix. S2: Multiply the mid-wave infrared spectral signal with the spectral transmittance matrix to obtain the compressed measurement result; S3: Input the compressed measurement results into the dual-branch spectral reconstruction network for processing to obtain the smooth background logarithmic spectral components and characteristic absorption absorbance components; S4: The logarithmic spectral components of the smooth background and the characteristic absorption absorbance components are superimposed to obtain the logarithmic spectrum. The logarithmic spectrum is then processed by a post-processing network to output the reconstructed spectral signal.
2. The physical-driven mid-wave infrared spectral encoding and reconstruction method according to claim 1, characterized in that: In step S3, the dual-branch spectral reconstruction network includes a smooth background logarithmic spectral reconstruction branch and a concentration range product prediction network.
3. The physical-driven mid-wave infrared spectral encoding and reconstruction method according to claim 2, characterized in that: The forward modeling network, the smooth background logarithmic spectral reconstruction branch, and the post-processing network are all fully connected networks.
4. The physical-driven mid-wave infrared spectral encoding and reconstruction method according to claim 1, characterized in that: The concentration range product prediction network includes an initial convolutional block, a multi-channel self-attention module, a first residual block, a second residual block, a global average pooling module, and a classifier. The compressed measurement results are processed sequentially by the initial convolutional block, the multi-channel self-attention module, the first residual block, the second residual block, the global average pooling module, and the classifier to obtain the concentration range product containing various gas components. The concentration range product of various gas components is multiplied by the transposed gas absorbance matrix to obtain the characteristic absorption absorbance components.
5. The physical-driven mid-wave infrared spectral encoding and reconstruction method according to claim 4, characterized in that: In the initial convolutional block, the input features are sequentially processed by connected one-dimensional convolution, batch normalization, LeakyReLU activation, and max pooling to obtain the output features.
6. The physical-driven mid-wave infrared spectral encoding and reconstruction method according to claim 4, characterized in that: The first residual block and the second residual block have the same structure. The first residual block includes two cascaded convolutional blocks, an efficient channel attention module, and a LeakyRelu activation module. The feature A1 input to the first residual block is processed by the two cascaded convolutional blocks and the efficient channel attention module in sequence for local feature extraction and channel adaptive enhancement to obtain feature A2. Feature A1 and feature A2 are then connected by a skip connection and input to the LeakyRelu activation module for processing to obtain feature A3.
7. The physical-driven mid-wave infrared spectral encoding and reconstruction method according to claim 4, characterized in that: In the classifier, the input features are sequentially processed through a linear fully connected layer and activated by LeakyReLU to obtain the output features.
8. The physical-driven mid-wave infrared spectral encoding and reconstruction method according to claim 4, characterized in that: Before the logarithmic spectrum is input into the post-processing network, it is transformed by inverse negative exponential transform.