A TDLAS Gas Concentration Detection Method and System Based on Dual-View Spectral Attention Mechanism
The TDLAS gas concentration detection method based on the dual-view spectral attention mechanism solves the difficulties of gas concentration detection under macroscopic sawtooth waves and photoelectric noise in the existing technology, and realizes high-precision and robust gas concentration measurement, which is applicable to a variety of gases and complex background conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
- Filing Date
- 2026-05-29
- Publication Date
- 2026-06-30
AI Technical Summary
Existing TDLAS gas concentration detection methods struggle to effectively extract weak gas absorption when faced with macroscopic sawtooth waves and photoelectric noise. Furthermore, concentration inversion methods based on absorbance are highly dependent on baseline extraction and are prone to introducing artificial noise. Traditional deep learning methods ignore the multiplicative coupling between optical background and gas absorption, resulting in feature aliasing and a lack of physical interpretability.
A TDLAS gas concentration detection method based on a dual-view spectral attention mechanism is adopted. By training the model, amplitude normalization, local window tokenization, physical prior correlation matrix enhancement, and cross-attention fusion are performed on the transmission spectrum to generate a three-dimensional spectral feature cube. Concentration regression is then performed using a multilayer perceptron to achieve gas concentration detection.
It improves the accuracy and robustness of gas concentration detection, reduces reliance on manual baseline extraction and complex preprocessing, enhances the ability to identify weak absorption features, and is suitable for concentration measurement under different gases and complex background conditions.
Smart Images

Figure CN122310474A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of laser spectroscopy detection and gas analysis technology, and in particular to a TDLAS gas concentration detection method and system based on a dual-view spectral attention mechanism. Background Technology
[0002] Quantitative determination of gas concentrations is crucial for atmospheric environmental monitoring, air quality assessment, pollution control, and industrial process optimization. These gases typically exhibit low and drastically varying concentrations, with significant spatial differences, posing a significant challenge to analytical techniques. Gas concentration determination based on laser absorption spectroscopy (LAS), relying on absorbance, offers advantages such as high sensitivity, selectivity, stability, and low cost, making it outstanding in continuous, real-time, and in-situ detection and meeting diverse application requirements. Tunable diode laser absorption spectroscopy (TDLAS), with its high selectivity, high sensitivity, and rapid response, has become the most advantageous spectroscopic technique for accurate measurement of atmospheric composition.
[0003] However, due to optical interference noise, instability of lasers, detectors, and other electronic equipment, measurements are subject to electronic noise such as jitter and drift. Furthermore, in real-world environments, baseline fluctuations caused by temperature-induced laser power variations, optical interference fringes, and beam deflection in turbulent media after superimposing the original transmission spectrum with a laser scanning sawtooth wave make spectral quantification difficult. Traditional methods require aligning the measured sensor spectrum with a standard spectrum to correct baseline drift and signal noise, followed by concentration inversion based on a physical model. However, even with environmental temperature and pressure corrections, this approach offers limited improvement in measurement accuracy. Many existing quantitative methods heavily rely on absorbance data, requiring baseline feature engineering from the transmission spectrum to absorbance. Common baseline extraction methods obtain the background spectrum or fit the non-absorption region by measuring a non-absorption reference gas cell. However, in real-world environments, the baseline is a nonlinear superposition of multiple physical processes with different time and frequency domain characteristics. Static methods cannot separate these features and are prone to introducing artificial distortions and errors. Yuntao Peng et al. proposed a physically-related dual-wavelength dynamic baseline extraction strategy. This strategy introduces a reference gas with a non-absorption region and derives an unknown baseline containing the target gas by combining the baseline of the reference gas, the laser's own optical power fluctuation signal, and the target gas's optical power fluctuation signal. However, these baseline extraction methods all rely on additional optical paths or cumbersome signal derivation. With the development of machine learning and deep learning, more and more optical sensors are being combined with artificial intelligence, providing solutions to the limitations of traditional signal processing. By processing and analyzing data from gas sensors, artificial intelligence algorithms can detect minute signals and distinguish them from noise features that traditional methods cannot detect, thereby improving sensitivity and selectivity.
[0004] Traditional regression aims to build models based on mathematical methods and predict continuous results based on a set of features or inputs. Deep learning regression extends this approach, with its main advantage being its ability to automatically extract complex patterns and features from data. Feedforward neural networks are suitable for input data with fixed dimensions and no time dependence; convolutional neural networks are suitable for signal or image data; and recurrent neural networks are suitable for time series analysis or speech processing. Due to its powerful ability to process complex and high-dimensional data, deep learning regression has become an effective tool for achieving high-precision continuous numerical prediction tasks.
[0005] Existing deep learning methods applied to TDLAS systems mainly develop in two directions. First, as pre-processing denoising: Rongqi Xu et al., addressing the shortcomings of transmission spectra obtained from direct absorption spectroscopy analysis in noisy environments, proposed a residual network filter based on deep learning algorithms. Utilizing the residual network filter's ability to extract spectral features, they established a mapping relationship between the input noisy spectrum and the denoised spectrum and successfully integrated it into a common methane sensor. Gas concentration inversion demonstrated its strong generalization ability and stability. Peng Zhao et al., addressing the temporal correlation of harmonic signals, introduced a method combining Long Short-Term Memory (LSTM) and Denoising Autoencoder (DAE), verifying its effectiveness in harmonic signal denoising through simulation and real-world experiments. However, the denoising process, without sensing the downstream regression target, may eliminate trace absorption along with noise, leading to irreversible feature loss. Second, end-to-end regression: Yanbin Ren et al., addressing the impact of absorbance spectral noise and background normalization on gas concentration inversion, constructed a dual convolutional neural network and a baseline normalization algorithm to learn good feature representations from the data, overcoming the limitations of traditional regression methods. Yinsong Wang et al. used a moving average filtering algorithm to denoise the second harmonic signal, decomposing each harmonic into several data segments and taking their average, peak, and valley values as inputs to CNN and LSTM neural networks. They established a nonlinear fitting relationship between gas input characteristics and concentration. Furthermore, they introduced an attention mechanism to optimize the neural network parameters, thereby improving inversion accuracy. Aoxue Cai et al. proposed a Fourier kernel convolutional neural network for high-precision gas concentration inversion from raw wavelength-modulated spectra, eliminating the need for additional preprocessing steps. In addition, Linbo Tian et al. proposed using one-dimensional convolution and a multilayer perceptron to directly learn concentration mappings from transmission spectra to measure methane and acetylene concentrations, and verified its robustness in dealing with laser aging and circuit fluctuations. DongqiYu et al. introduced the DBN method for gas concentration inversion after preprocessing the acquired absorption spectra using empirical wavelet transform and principal component analysis, significantly improving inversion accuracy. While the aforementioned methods have yielded significant results, simply treating the absorption spectrum as a flat one-dimensional time series or one-dimensional image for deep learning mapping is insufficient to address the multiplicative coupling between macroscopic sawtooth wave intensity and interference fringes and microscopic gas absorption. Existing deep learning regression quantification is highly dependent on the data distribution in the training set and cannot effectively handle the dynamic changes in the amplitude and slope of the sawtooth wave intensity of the laser incident light due to various factors. Furthermore, traditional attention mechanisms tend to focus on baselines with large amplitude variations in transmission spectrum modeling, thus losing their ability to represent the actual absorption region. Summary of the Invention
[0006] To address the challenges of extracting weak gas absorption from macroscopic sawtooth waves and photoelectric noise in existing technologies, the high dependence of absorbance-based concentration inversion methods on baseline extraction and the susceptibility to artificial noise, and the fact that existing end-to-end deep learning methods treat spectral data as a one-dimensional sequence and ignore the multiplicative coupling between optical background and gas absorption, leading to feature aliasing and a lack of physical interpretability, this invention provides a TDLAS gas concentration detection method and system based on a dual-view spectral attention mechanism.
[0007] To achieve the above objectives, the present invention adopts the following technical solution, including: A TDLAS gas concentration detection method based on a dual-view spectral attention mechanism is proposed. The gas concentration is detected using a trained TDLAS gas concentration detection model, which is trained as follows: S11, Obtain a sample set from the TDLAS device. The sample includes the original transmission spectrum and the gas concentration labels corresponding to the spectrum. S12, normalize the amplitude of the original transmission spectra in the sample set to obtain a normalized spectral sample set; S13, perform learnable local window tokenization on the original transmission spectra in the normalized spectral sample set to obtain the spectral token sequence. S14, inject position encoding information into the spectral token sequence to generate a three-dimensional spectral feature cube; S15, use the physical prior correlation matrix to perform intraspectral attention enhancement on the three-dimensional spectral feature cube to generate spectral enhancement features; S16: Based on the wavelength feature planes of the three-dimensional spectral feature cube, construct the query, construct the key and value with the typical feature spectrum of the corresponding band, and complete the feature matching and aggregation through cross attention to obtain the cross-spectral fusion feature; S17, the spectral enhancement features and cross-spectral fusion features are fused through gated residuals to generate the final fused features; S18, learnable weighted aggregation and multilayer perceptron mapping are performed on the final fused features to obtain the network regression output, and the network regression output is mapped to the concentration label scale to obtain the gas concentration value; S19. Based on the gas concentration values obtained from the model output and the corresponding gas concentration labels, a regression loss function is constructed, and the model parameters are iteratively updated using the Adam optimizer combined with an exponential learning rate decay strategy.
[0008] Preferably, the gas concentration is detected using the trained TDLAS gas concentration detection model. The specific detection process is as follows: S21, Obtain the raw transmission spectrum from the TDLAS device; S22, Based on the model parameters determined during the training phase, the amplitude of the original transmission spectrum is normalized to obtain the normalized test spectrum; S23, perform learnable local window tokenization on the normalized spectrum to be measured to obtain the spectrum token sequence; S24, inject position encoding information into the spectrum token sequence to be measured to generate a three-dimensional spectral feature cube; S25, use the physical prior correlation matrix to perform intraspectral attention enhancement on the three-dimensional spectral feature cube to generate spectral enhancement features; S26, perform inter-spectral cross-attention matching on the three-dimensional spectral feature cube based on typical feature spectra to obtain cross-spectral fusion features; S27, the spectral enhancement features and cross-spectral fusion features are fused through gated residuals to generate the final fused features; S28, the final fused features are subjected to learnable weighted aggregation and multilayer perceptron mapping to obtain network regression output, and the network regression output is mapped to the concentration label scale to obtain the concentration value of the gas to be measured.
[0009] Preferably, the amplitude normalization process in step S12 uses maximum and minimum value normalization to reduce the amplitude scale differences caused by different bands, different acquisition batches and laser output power fluctuations, so as to obtain a normalized spectral sample set.
[0010] Preferably, the specific steps of step S13 are as follows: The original transmission spectrum is segmented along the wavelength dimension according to a preset local window, and the continuous spectrum is divided into multiple spectral segments. Each spectral segment is used as an independent spectral token to form an initial spectral token sequence. The initial spectral token sequence is subjected to a learnable embedding transformation. At the same time, a local sliding window is used to perform feature fusion on adjacent tokens in the wavelength dimension, aggregating the spectral information of adjacent wavelengths to capture the local spectral structure, and finally obtaining a regular spectral token sequence.
[0011] Preferably, the specific operation steps of step S15 are as follows: From the three-dimensional spectral feature cube Take out the first one Original transmission spectrum Perform linear projection transformations on the samples to generate the query matrix required by the attention mechanism. Key matrix AND-value matrix , ,in, The number of original transmission spectra. Indicates the sampling length. The token feature dimension mapped to each wavelength point. For the projection dimension of the attention mechanism; Introducing the correlation matrix guided by physical priors ; Generate spectral enhancement features : ; in, For normalized exponential functions, These are the learnable weights.
[0012] Preferably, the specific operation steps of step S16 are as follows: Obtaining a 3D spectral feature cube Obtain typical characteristic spectra and generate a typical characteristic spectrum cube. ,in, The number of original transmission spectra. The number of typical characteristic spectra, Indicates the sampling length. The token feature dimension mapped to each wavelength point; Take a three-dimensional spectral feature cube A certain wavelength point Plane , Generate a query matrix with cross-attention. : ; in, For the projection dimension of the attention mechanism, The linear projection weights of the query matrix; Using typical characteristic spectral cubes in the corresponding Band matrix generation key matrix Sum matrix ; No. The output characteristics of the cross-attention mechanism in each band are as follows: , : ; in, For the three-dimensional spectral feature cube A query matrix for each band. The typical characteristic spectral cube Transpose of the key matrix of each band The typical characteristic spectral cube Value matrix of each band; The output features of all bands are stacked sequentially along the wavelength sequence dimension, and the cross-spectral fusion features are finally reconstructed. : ; in, This is a tensor splicing function.
[0013] Preferably, step S17 uses two learnable gating parameters to control the contributions of the spectral enhancement feature and the cross-spectral fusion feature, respectively, to generate the final fusion feature. : ; in, It is a three-dimensional spectral feature cube. These are the gating parameters for spectral attention branches guided by physical priors. These are typical characteristic spectral cross-fusion branch gating parameters. For spectral enhancement features, This is a cross-spectral fusion feature.
[0014] Preferably, the specific operation steps of step S18 are as follows: The final fused features are flattened to obtain a one-dimensional fused feature vector. One-dimensional fused features are input into a multilayer perceptron, and feature transformation and nonlinear mapping are performed sequentially through multiple fully connected layers and nonlinear activation functions, and the network regression output is output. A scale mapping strategy is used to map the network regression output to a scale range consistent with the concentration labels in the training set, thereby obtaining the gas concentration values.
[0015] Preferably, the specific operation steps of step S19 are as follows: Perform an inverse label transformation on the gas concentration value output in step S18 to restore it to a numerical scale consistent with the corresponding gas concentration label, and obtain the generated gas concentration label. A loss function is constructed based on generated gas concentration labels and real gas concentration labels. The Adam optimizer is used in conjunction with an exponential learning rate decay strategy to complete the iterative update of model parameters.
[0016] A TDLAS gas concentration detection system based on a dual-view spectral attention mechanism is provided, applicable to the aforementioned TDLAS gas concentration detection method based on a dual-view spectral attention mechanism. The gas concentration detection system includes: Data acquisition and preprocessing module: Acquires the raw transmission spectrum from the TDLAS device, normalizes the raw transmission spectrum, and generates a normalized spectrum; Spectral sequence encoding module: The normalized spectrum is tokenized using a learnable local window to obtain a spectral token sequence, and positional encoding information is injected into the spectral token sequence to generate a three-dimensional spectral feature cube; Physically Prior-Guided Feature Enhancement Module: This module intervenes in intraspectral attention allocation through a correlation matrix guided by physical priors, generating spectral enhanced features. Typical Feature Spectral Cross-fusion Module: Based on the wavelength feature plane of the three-dimensional spectral feature cube, a query is constructed. Keys and values are constructed with the typical feature spectral features of the corresponding band. Feature matching and aggregation are completed through cross attention to obtain cross-spectral fusion features. Dual-gated feature fusion module: fuses spectral enhancement features and cross-spectral fusion features to generate the final fused feature; Regression Head: Performs learnable linear transformation and weighted aggregation on the final fused features to obtain a spectral representation distributed along the wavelength scanning direction, and inputs the spectral representation into the multilayer perceptron regression network to obtain the network regression output; Learnable label processing module: Maps the network regression output to the concentration label scale to obtain the gas concentration detection value.
[0017] The advantages of this invention are: 1. This invention directly encodes the features of the original TDLAS transmission spectrum through a learnable local window tokenization method, reducing the reliance on manual baseline extraction and complex preprocessing steps; 2. This invention constrains the intraspectral attention by introducing a physical prior correlation matrix, enabling the model to focus on gas absorption peaks and related bands, and suppressing the effects of baseline drift, photoelectric noise and edge artifacts on concentration inversion; 3. This invention achieves adaptive matching between the spectrum to be measured and the reference spectrum through a typical characteristic spectrum cross-attention mechanism, thereby improving the ability to identify weak absorption features in complex backgrounds. 4. This invention adaptively adjusts the contribution of physical prior enhancement features and typical feature spectral cross-fusion features through a dual-gated residual fusion mechanism, thereby achieving collaborative modeling of local absorption information and global spectral benchmark. 5. This invention utilizes learnable label scale mapping to convert network output into a true concentration scale, thereby enabling the detection results to have better stability, robustness, and physical interpretability.
[0018] 6. This invention can improve the accuracy, generalization ability and noise resistance of TDLAS gas concentration detection, and is suitable for concentration measurement scenarios under different gases, different wavebands and complex background conditions. Attached Figure Description
[0019] Figure 1This is a schematic diagram of a TDLAS gas concentration detection system based on a dual-view spectral attention mechanism.
[0020] Figure 2 A schematic diagram illustrating the fusion of a physical prior-guided feature enhancement module and a typical feature spectrum cross-fusion module.
[0021] Figure 3 The flowchart shows the method for training the TDLAS gas concentration detection model.
[0022] Figure 4 This is a schematic diagram of the correlation matrix guided by the physical priors of methane gas.
[0023] Figure 5 A schematic diagram of the correlation matrix guided by the physical priors of nitrous oxide gas.
[0024] Figure 6 This is a comparison chart of the sensitivity of the correlation matrix parameters under different adjustment factors.
[0025] Figure 7 This is a schematic diagram of the TDLAS trace gas detection experimental platform.
[0026] Figure 8 This is a schematic diagram illustrating the experimental results of different token generation strategies on the methane dataset.
[0027] Figure 9 This diagram illustrates the experimental results of different token generation strategies on the nitrous oxide dataset.
[0028] Figure 10 This diagram illustrates the parameter sensitivity of different attention modules to the model on the methane dataset.
[0029] Figure 11 This diagram illustrates the parameter sensitivity of different attention modules to the model on the nitrous oxide dataset.
[0030] Figure 12 A weighted graph of the attention matrix for features enhanced by physical priors in the nitrous oxide dataset.
[0031] Figure 13 This is a weighted graph of the attention matrix for the cross-fusion of typical feature spectra in the nitrous oxide dataset.
[0032] Figure 14 This is a schematic diagram showing the original spectrum of the nitrous oxide dataset and the attention confidence weights. Detailed Implementation
[0033] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0034] Example 1 The relevant technologies involved in this invention are described below: 1. TDLAS signal model and its nonlinear coupling The data generated by the TDLAS signal model is based on Beer–Lambert's law, which describes the physical relationship between the absorbance of electromagnetic radiation and the object being analyzed, namely: ; in, Indicates absorbance. Indicates light transmittance. The absorption coefficient is... The thickness of the absorption layer, This represents the concentration of the absorbing substance. The single absorption of a particular component satisfies the following formula: ; in, These represent the transmitted light intensity and the incident light intensity, respectively. The concentration of the absorbed gas, It is the effective optical distance. For absorption cross section, , For wave number, This is the strength proportionality coefficient. The integral intensity of the spectral line is divided by the line intensity. For the sake of light transmittance The composite coefficient, This is a spectral line shape function. Therefore, the problem is modeled as follows: assuming a spectral dataset... The Middle The transmission spectrum is , , The number of transmission spectra. This represents the sampling length, which physically corresponds to wavelength. The actual incident light intensity includes physical perturbations, and Beer-Lambert's law becomes: ; in, For the first A transmission spectrum at wavelength The light intensity at that location, For the first A transmission spectrum at wavelength The actual background baseline at the location, For the first A transmission spectrum at wavelength Noise at that location For the first The tag to be detected by the transmission spectrum, i.e., the first... The detection concentration of each transmission spectrum.
[0035] The formula shows that gas concentration is coupled in an exponential nonlinearity and has multiplicative baseline interference. When the network input is a transmission spectrum, in order to detect concentration... It needs to implicitly complete two tasks internally: first, to find the non-absorption region to estimate the background baseline, and second, to find the absorption valley to deduce the concentration.
[0036] 2. Standard Scaling Dot Product Attention Model From a mathematical perspective, the attention mechanism is a general inter-set information routing function. It maps a set of queries to a set of key-value pairs, thereby generating highly context-sensitive output. (Query matrix) key value Generated via linear mapping: ; ; ; in, Input the feature matrix into the model. This is a learnable weight matrix.
[0037] The standard scaled dot product attention model can be represented as: ; Wherein, scaling factor To ensure the softmax function operates within a relatively flat region and avoid gradient vanishing, it's crucial to ensure it operates within a relatively flat region. However, applying unconstrained attention modules directly to raw TDLAS data will result in the variance-driven disaster. In the raw TDLAS signal, the background baseline contributes most of the signal's amplitude and global variance, and the exponential amplification effect of the softmax function causes weak gas concentration information to be filtered out by the network during the attention weight allocation stage. Traditional self-attention mechanisms, when faced with spectral data where effective information is extremely sparse, force each token to calculate an attention score with all tokens, leading to the introduction of high-frequency noise from the background region into the attention matrix.
[0038] Please see Figure 1 and Figure 2 , Figure 1This is a block diagram of a TDLAS gas concentration detection system based on a dual-view spectral attention mechanism as described in this embodiment. Figure 2 The diagram illustrates the integration of a physical prior-guided feature enhancement module and a typical feature spectrum cross-fusion module. The gas concentration detection system includes: a data acquisition and preprocessing module, a spectral sequence encoding module, a physical prior-guided feature enhancement module, a typical feature spectrum cross-fusion module, a dual-gated feature fusion module, a regression head, a learnable label processing module, and a model parameter update module. The entire system's operation is divided into two parts: the training process and the detection process.
[0039] During training: Data acquisition and preprocessing module: used to acquire the raw transmission spectrum from the TDLAS device and the gas concentration label corresponding to the spectrum, perform amplitude normalization on the raw transmission spectrum, and generate a normalized spectral dataset. Spectral sequence encoding module: The original transmission spectra in the normalized spectral dataset are tokenized using learnable local windows to obtain a spectral token sequence, and positional encoding information is injected into the spectral token sequence to generate a three-dimensional spectral feature cube. Physically Prior-Guided Feature Enhancement Module: This module intervenes in intraspectral attention allocation through a correlation matrix guided by physical priors, generating spectral enhanced features. Typical Feature Spectrum Cross-fusion Module: Based on the wavelength feature plane of the three-dimensional spectral feature cube, a query is constructed, and the key and value are constructed with the typical feature spectrum of the corresponding band. The feature matching and aggregation are completed through cross attention to obtain the cross-spectral fusion feature. Dual-gated feature fusion module: fuses spectral enhancement features and cross-spectral fusion features to generate the final fused feature; Regression Head: Performs learnable linear transformation and weighted aggregation on the final fused features to obtain a spectral representation distributed along the wavelength scanning direction, and inputs the spectral representation into the multilayer perceptron regression network to obtain the network regression output; Learnable label processing module: Maps the network regression output to a concentration label scale to obtain gas concentration values; Model parameter update module: Based on the gas concentration values obtained from the model output and the corresponding concentration labels, a regression loss function is constructed, and the model parameters are iteratively updated using the Adam optimizer combined with an exponential learning rate decay strategy.
[0040] During the testing process: Data acquisition and preprocessing module: Acquires the raw transmission spectrum from the TDLAS device, normalizes the raw transmission spectrum, and generates a normalized spectrum; Spectral sequence encoding module: The normalized spectrum is tokenized using a learnable local window to obtain a spectral token sequence, and positional encoding information is injected into the spectral token sequence to generate a three-dimensional spectral feature cube; Physically Prior-Guided Feature Enhancement Module: This module intervenes in intraspectral attention allocation through a correlation matrix guided by physical priors, generating spectral enhanced features. Typical Feature Spectrum Cross-fusion Module: Based on the wavelength feature plane of the three-dimensional spectral feature cube, a query is constructed, and the key and value are constructed with the typical feature spectrum of the corresponding band. The feature matching and aggregation are completed through cross attention to obtain the cross-spectral fusion feature. Dual-gated feature fusion module: fuses spectral enhancement features and cross-spectral fusion features to generate the final fused feature; Regression Head: Performs learnable linear transformation and weighted aggregation on the final fused features to obtain a spectral representation distributed along the wavelength scanning direction, and inputs the spectral representation into the multilayer perceptron regression network to obtain the network regression output; Learnable label processing module: Maps the network regression output to the concentration label scale to obtain the gas concentration detection value.
[0041] Example 2 This embodiment 2 proposes a TDLAS gas concentration detection method based on a dual-view spectral attention mechanism. The method trains a TDLAS gas concentration detection model incorporating the dual-view spectral attention mechanism, and then uses the trained model to detect gas concentration. Figure 3 As shown, the training method is as follows: S11, Obtain a sample set from the TDLAS device. The sample includes the original transmission spectrum and the gas concentration labels corresponding to the spectrum. Specifically, obtain the sample set from the TDLAS device. , For the number of samples, Indicates the sampling length, where one sample is... , The sample includes the original transmission spectra and the gas concentration labels corresponding to the spectra, so the number of original transmission spectra is also [number missing]. .
[0042] Subsequently, the original spectra were divided and labeled, and strict out-of-distribution (OOD) data partitioning was performed. Training and test sets with non-overlapping concentrations were constructed. The TDLAS gas concentration detection model, which incorporates a dual-view spectral attention mechanism, was trained using samples from the training set. The generalization ability and concentration detection accuracy of the trained model were objectively evaluated using samples from the test set.
[0043] S12, normalize the amplitude of the original transmission spectra in the sample set to obtain a normalized spectral sample set; Specifically, normalization parameters are determined based on the amplitude range of the original transmission spectra in the training set, and each original transmission spectrum is normalized to obtain a normalized spectral sample set. In this embodiment, the amplitude normalization preprocessing uses maximum-minimum value normalization to reduce amplitude scale differences caused by different wavebands, different acquisition batches, and laser output power fluctuations, thus obtaining a normalized spectral sample set. The path integral concentration label maintains the scale of the true physical quantity and is used to construct the loss function with the gas concentration detection values output by the model.
[0044] S13, perform learnable local window tokenization on the original transmission spectra in the normalized spectral sample set to obtain the spectral token sequence. Specifically, the original transmission spectrum is segmented in the wavelength dimension according to a preset local window, dividing the continuous spectrum into multiple spectral segments. Each spectral segment is used as an independent spectral token to form an initial spectral token sequence. The initial spectral token sequence is subjected to a learnable embedding transformation, and at the same time, a local sliding window is used in the wavelength dimension to perform feature fusion on adjacent tokens, aggregating the spectral information of adjacent wavelengths to capture the local spectral structure, and finally obtaining a regular spectral token sequence.
[0045] S14, inject position encoding information into the spectral token sequence to generate a three-dimensional spectral feature cube; Specifically, positional encoding information is added to each spectral token, enabling the model to perceive the order and positional relationship of the spectra in the wavelength dimension, and generating a three-dimensional spectral feature cube. (Model input feature matrix) The number of samples, i.e. The number of original transmission spectra. Indicates the sampling length. The token feature dimension mapped to each wavelength point; S15, through the intervention of the physical prior-guided correlation matrix in the attention allocation of the three-dimensional spectral feature cube, spectral enhancement features are generated; From the three-dimensional spectral feature cube Take out the first one sample The samples are subjected to linear projection transformations to generate the query matrices required by the attention mechanism. Key matrix AND-value matrix , ,in The projection dimension representing the attention mechanism is set in this embodiment. .
[0046] Since a small portion of the absorption information in the TDLAS gas concentration detection model is submerged in noise, resulting in the problem that a small portion of the absorption caused by broadening is not noticed, this invention introduces a physical prior-guided correlation matrix into the in-sample attention mechanism to intervene in the in-sample attention allocation.
[0047] The correlation matrix The two-dimensional mask matrix is constructed by the outer product of the nonlinear weight vector generated from the second derivative of the average spectrum of the training set, and then generated by the Hadamard product of the Spearman correlation matrix of the second derivatives of all training samples. The specific steps are as follows: Average spectrum of the training set The second derivative is used to highlight curvature changes near the absorption peak; Gaussian smoothing is used to suppress high-frequency noise. The first and last 10% of the spectrum are weakened to avoid sensor edge artifacts interfering with attention allocation. By using a soft thresholding function to reduce the weight of the background region and by using nonlinear transformation to enhance the response of the weak absorption region, a one-dimensional spectral weight vector is obtained. The one-dimensional spectral weight vector is multiplied by an outer product to generate a two-dimensional mask matrix. This two-dimensional mask matrix is then multiplied by a Hadamard product with the Spearman correlation matrix corresponding to the second derivative of the training samples. Finally, a physical prior matrix that can simultaneously characterize the saliency of spectral absorption position and the correlation between bands is obtained.
[0048] Specifically, the average spectrum of the original training set Calculate the second derivative and obtain the sequence using Gaussian smoothing. To remove sensor edge artifacts, the bands immediately before and after the sensor are ignored, and a soft thresholding function is used to reduce background interference. To increase the weights for weak absorption, the weights are set... This allows even weak absorption to receive appropriate attention.
[0049] The final physics-prior-guided feature enhancement attention mechanism generates spectral enhancement features. : ; in, For normalized exponential functions, These are learnable weights used to control the influence of the physical prior guidance matrix on the attention mechanism of the relevant band. The correlation matrix is guided by physical priors. For the purpose of inquiry and evidence collection, This is the transpose of the key matrix. For the projection dimension of the attention mechanism, It is a value matrix.
[0050] Please see Figure 4 and Figure 5 , Figure 4 This is a schematic diagram of the correlation matrix guided by the physical priors of methane gas. Figure 5 This diagram illustrates the physical prior guidance correlation matrix for nitrous oxide, showing the physical prior guidance matrices for different gas formations. The dark blue areas, with weights approaching 0, indicate that correlation noise has been suppressed. The bright areas retain significant absorption characteristics, and the red areas represent the strongest absorption positions for that gas. To investigate the weight of this correlation matrix in the model, different adjustment factors were designed. Experiments were conducted separately, with the model structure and hyperparameters fixed, and only statistical analysis was performed. Parameter sensitivity to different datasets. For example... Figure 6 As shown, on the methane dataset, At that time, RMSE reached its lowest value of 4.7575, indicating that the physical prior matrix only needs to provide a small amount of weight assistance to effectively guide the attention mechanism to focus on a specific absorption band; while on the nitrous oxide dataset, The lowest RMSE was 4.6109. Different gases exhibit varying degrees of dependence on intraspectral band correlation, demonstrating the necessity of the proposed physical prior guiding matrix for improving the model's expressive power.
[0051] S16: Based on the wavelength feature planes of the three-dimensional spectral feature cube, a query is constructed. Keys and values are built using typical feature spectra of the corresponding bands. Feature matching and aggregation are completed through cross-attention to obtain cross-spectral fusion features. The specific operation steps are as follows: Obtain the three-dimensional spectral feature cube generated in step S14 ; Obtain typical characteristic spectra, and obtain the typical characteristic spectrum cube through steps S11-S14. ,in The number of typical characteristic spectra, Indicates the sampling length. Each wavelength point is mapped to a token feature dimension that is consistent with the input spectral feature cube, ensuring feature space alignment.
[0052] In this embodiment, the typical characteristic spectra can be derived from simulated transmission spectra generated in the Hisran database under thermodynamic conditions consistent with the target environment, or from high signal-to-noise ratio samples obtained in experiments.
[0053] Extract the three-dimensional spectral feature cube A certain wavelength point Plane , The plane is transformed by linear projection to generate a query matrix for cross-attention. The linear projection process is represented as: ; in, The linear projection weights of the query matrix.
[0054] Using typical characteristic spectral cubes Matrix generation key in the corresponding band Sum .
[0055] For each input spectrum, the query vector is compared with the key vector of a typical feature spectrum. The most relevant spectral pattern is adaptively matched, and the corresponding value vectors are weighted and aggregated to extract effective feature information. The output characteristics of the cross-attention mechanism in each band are as follows: , : ; in, For the three-dimensional spectral feature cube A query matrix for each band. The typical characteristic spectral cube The key matrix of each band, The typical characteristic spectral cube The value matrix of each band.
[0056] The output features of all bands are stacked sequentially along the wavelength sequence dimension, and the cross-spectral fusion features are finally reconstructed. : .
[0057] S17, the spectral enhancement features and cross-spectral fusion features are fused through gated residuals to generate the final fused features; Specifically, by controlling the contributions of physical prior enhancement features and typical spectral cross-fusion features through two learnable gating parameters, an adaptive balance between local absorption information and global spectral benchmark is achieved. That is, the spectral enhancement features and cross-spectral fusion features are fused to generate the final fused feature. ; in, It is a three-dimensional spectral feature cube. Gating parameters for spectral attention branches guided by physical priors, and These are typical characteristic spectral cross-fusion branch gating parameters.
[0058] S18. The final fused features are subjected to learnable weighted aggregation and multilayer perceptron mapping to obtain the network regression output. The network regression output is then mapped to the concentration label scale to obtain the gas concentration value.
[0059] The specific operation process is as follows: The final fused features are flattened to convert the high-dimensional multidimensional fused features into one-dimensional feature vectors, eliminating dimensional differences and facilitating the input and computation of the subsequent perceptron model. The flattened one-dimensional fused features are input into a multilayer perceptron (MLP), and feature transformation and nonlinear mapping are performed sequentially through multiple fully connected layers and nonlinear activation functions (such as ReLU and Sigmoid functions) to explore the deep correlation between fused features and gas concentration. Finally, the network regression output is output through a linear output layer. By employing a scale mapping strategy, the network regression output of the multilayer perceptron is mapped to a scale range consistent with the concentration labels in the training set (i.e., matching the numerical range and dimensions of the concentration labels), eliminating the scale bias between the network output and the actual concentration labels, and finally obtaining accurate and interpretable gas concentration values.
[0060] S19. Based on the gas concentration values obtained from the model output and the corresponding gas concentration labels, a regression loss function is constructed, and the model parameters are iteratively updated using the Adam optimizer combined with an exponential learning rate decay strategy. Specifically, the gas concentration value output in step S18 is subjected to inverse label transformation to restore it to a numerical scale consistent with the real label, thereby obtaining the generated gas concentration label. A loss function is constructed based on the generated gas concentration labels and the real gas concentration labels. The Adam optimizer is used in conjunction with an exponential learning rate decay strategy to complete the iterative update of the model parameters. Meanwhile, an early stopping mechanism is introduced, using the validation set loss as the monitoring metric. When the validation set performance fails to improve for several consecutive iterations, training is terminated early, effectively suppressing model overfitting and improving the model's generalization ability.
[0061] The trained TDLAS gas concentration detection model was used to detect the gas concentration. The specific detection process is as follows: S21, Obtain the raw transmission spectrum from the TDLAS device; S22, Based on the model parameters determined during the training phase, the amplitude of the original transmission spectrum is normalized to obtain the normalized test spectrum; S23, perform learnable local window tokenization on the normalized spectrum to be measured to obtain the spectrum token sequence; S24, inject position encoding information into the spectral token sequence to be tested to generate a spectral feature cube with position information; S25, use the physical prior correlation matrix to perform intraspectral attention enhancement on the spectral feature cube to be measured to generate spectral enhancement features; S26, perform inter-spectral cross-attention matching on the three-dimensional spectral feature cube based on typical feature spectra to obtain cross-spectral fusion features; S27, the spectral enhancement features and cross-spectral fusion features are fused through gated residuals to generate the final fused features; S28, the final fused features are subjected to learnable weighted aggregation and multilayer perceptron mapping to obtain network regression output, and the network regression output is mapped to the concentration label scale to obtain the concentration value of the gas to be measured.
[0062] Example 3 The spectral data collected in this experiment were all from the TDLAS trace gas detection experimental platform, the schematic diagram of which is shown below. Figure 7 As shown in the diagram. Solid arrows represent the electrical signal control and data acquisition links; dashed arrows represent the laser transmission path; and dotted-dashed arrows represent the gas path. In this experimental platform, the absorption gas cell is 25m long; the laser is tuned collaboratively by the current control module and the temperature control module, using a sawtooth wave as the wavelength scanning signal. The gas mixing platform consists of two Sevenstar Mass Flow Meters with ranges of 500 sccm and 1000 sccm, respectively. These are mainly used to prepare gradient concentrations of target gases (methane and nitrous oxide), providing samples of target gases at different concentrations for the experiment. Nitrogen is used as a dilution gas to prepare the gradient gases.
[0063] Based on the above system, we collected methane gas concentrations ranging from 10.5 to 40 ppm. Since the actual target is the integral concentration along the entire path, the label is set to path integral concentration ppm. To overcome the accuracy limitations of the gas distribution system, this paper simulates 10,000 training samples as an auxiliary dataset based on the characteristics of the laser and a forward model of the HITRAN database. The simulation uses transmission spectral data with a signal-to-noise ratio of 50 dB, and follows... Figure 7 The experimental platform shown has completed the simulation parameter configuration. In this embodiment, the laser parameters are specifically set as follows: center current 178mA, power tuning factor 0.12563mW / mA, and wavelength tuning factor 0.01175nm / mA. Due to the overlapping of multiple methane absorption peaks near 1651nm, a selective absorption line stronger than 1e-30 is chosen. The absorption lines are simulated to ensure that weak spectral lines are also included in the calculation.
[0064] To verify the generality of the proposed model, a supplementary dataset of nitrous oxide gas directly absorbed by a mid-infrared laser was also collected. The concentration of nitrous oxide in this dataset ranges from 1 to 10 ppm, and the gas cell length is 25 m.
[0065] The experimental environment details are shown in Table 1. The model is based on the PyTorch 1.12.1 deep learning framework and runs on a Windows system with a Core i9-10900X processor and 32GB of memory. This system uses two NVIDIA GeForce RTX 2080 Ti GPUs with a total of 24GB of video memory to accelerate tensor computation.
[0066] Table 1: Experimental Environment
[0067] To evaluate the performance of the regression model, this paper uses RMSE, MAE, and R... 2 As an evaluation indicator. 2 It is the square of the correlation coefficient between the true value and the model's detected value, representing the percentage of variation in the outcome variable that the detected variable can explain. It ranges from 0 to 1, with a higher value indicating better model stability.
[0068] The root mean square error (RMSE) is the square root of the mean squared error. It measures the average error between the detected value and the true value, and its unit is consistent with the target variable. It is commonly used as a metric for the detection results of machine learning models; the lower the RMSE, the better the model. ; in, For the first The true concentration label of each sample For the first The mean absolute error of the detection concentration label for each sample is the arithmetic mean of the absolute values of the deviations of all detected values from the true values.
[0069] Compared to RMSE, MAE (Mean Absolute Error) is less sensitive to outliers, treating all errors equally and exhibiting a "linear" penalty mechanism. ; The coefficient of determination is calculated based on the difference between the true value, the model's test value, and the data mean. For variations that the model failed to explain, The total variation of the data itself, For all real tags The average of. The formula is: .
[0070] Example 4 To verify the superiority of the proposed method, comparative experiments were conducted with CNN, LSTM, and transformer architectures.
[0071] For Convolutional Neural Networks (CNNs), a 1D-CNN was constructed to evaluate the performance of local feature extraction on the TDLAS spectral dataset. The CNN consists of two cascaded convolutional blocks, with the number of channels increased from 16 to 32. Kernel=5 was set to capture the structure of local absorption peaks in the spectrum, preventing the kernel from being too large and introducing excessive background noise. Stride=2 was used for smooth downsampling during feature extraction, reducing the computational cost of the model. The extracted spectral features were then processed through global average pooling and input to a fully connected layer, outputting the final detection path integral concentration.
[0072] For LSTM, a bidirectional long short-term memory network BiLSTM with a hidden layer size of 64 is used to capture complete contextual information on both sides of the absorption peak. Then, the hidden states of forward and backward propagation are concatenated and fused in the feature dimension as a global feature representation. Finally, the feature vector is fed into a fully connected layer for dimensionality reduction mapping to output the final detection result.
[0073] For transformers, set , , Two encoder layers are stacked to extract diverse representation subspaces. The input sequence is linearly mapped and fused with positional encoding before being fed into the network. The feature output is globally pooled along the sequence dimension and then passed through a fully connected layer to output the detection results. The experimental results are shown in Table 2.
[0074] Table 2. Comparison of detection performance of different models on different spectral datasets
[0075] For the nitrous oxide dataset, Gaussian white noise and random drift are added to the original data to simulate random and systematic errors for data augmentation. CNN local convolutional kernels are prone to local overfitting under high-frequency noise interference; baseline drift disrupts sequence morphology, causing LSTM to fail to establish correct long-range context mapping; and without physical constraints, the transformer's purely data-driven model is easily disrupted by cluttered noise. The model proposed in this paper maintains low RMSE and high R-value on noisy datasets. 2 It is worth noting that although the Transformer exhibits a small detection variance on this dataset, its absolute error RMSE reaches 7.0371 ppm. m, by comparing the extreme values, the detection accuracy of the upper boundary of the error fluctuation of the model proposed in this paper (3.3850+1.2922=4.6772) still exceeds that of the lower boundary of the transformer (7.0371-0.5360=6.5011).
[0076] The above results demonstrate that the proposed model effectively focuses on the global spectral morphology, filtering out high-frequency noise while maintaining high detection accuracy. For the methane dataset, local convolutional kernels of CNNs struggle to extract effective features from complex background interference, resulting in the worst detection performance among all methods. The global self-attention mechanism performs well in larger and more complex methane datasets because it can observe the entire spectrum, while the proposed method remains the best among all baselines due to its injection of a physical correlation matrix to specifically focus on specific bands.
[0077] In summary, the proposed model demonstrates good results for both data-enhanced nitrous oxide datasets in the mid-infrared band and methane datasets with lower signal-to-noise ratios in the near-infrared band. The guidance of the physical prior matrix and the dual-viewpoint attention mechanism effectively compensate for the shortcomings of traditional deep learning in quantifying weak signal path integrals, providing a general solution for TDLAS quantification.
[0078] Example 5 To verify the impact and specific contribution of the two basic modules in the proposed dual-view spectral attention model on regression performance, a systematic ablation experiment was designed. A set of non-optimal hyperparameters was set to differentiate the contributions of architecture and hyperparameter configuration to the model, while keeping all variables consistent. The experiment was repeated five times with randomized reported mean ± standard deviation. The original backbone network without attention was used as the baseline in the ablation experiment. The input spectrum was only represented by a tokenizer and encoded at position, followed by feature aggregation via a feedforward neural network. Based on this, the effects of the physically prior-guided feature enhancement module and the typical spectral cross-fusion module on the model, as well as the impact of the dual-view spectral attention mechanism on the entire model, were verified. As shown in Table 3, the dual-view attention mechanism reduced the RMSE of methane from 10.89 to 6.82 and the RMSE of nitrous oxide from 7.05 to 4.61, and R... 2 The value was significantly increased to 0.9648, indicating that the basic network lacking the attention module cannot effectively cope with complex spectral signals.
[0079] For the methane dataset, using only the PFE module (RMSE 6.89) and only the CPA module (RMSE 6.95) both resulted in an error reduction of over 35% compared to the baseline, indicating that both modules have a positive impact on the model. PFE effectively injects spectral internal correlations by incorporating a physical prior matrix to lock the local absorption curvature boundary, while CPA provides macroscopic baseline alignment for real signal reconstruction by focusing on typical characteristic spectra. For the nitrous oxide dataset, using only the PFE module (variance ±2.21) and the CPA module (variance ±2.24) reduced the RMSE, but resulted in extremely high variance. When the two modules were combined, the model's detection RMSE decreased and the variance converged to ±0.86. This shows that the macroscopic baseline provided by CPA effectively prevents PFE from over-focusing on local noise, while the physical prior constraints provided by PFE smooth out the macroscopic fluctuations caused by typical characteristic spectra in CPA. These experiments demonstrate the necessity of the proposed dual-view spectral attention mechanism based on a dual-gated residual parallel fusion architecture.
[0080] Table 3. Impact of the two attention mechanisms on model performance
[0081] Example 6 To verify the superiority of the proposed learnable local window tokenizer generation strategy, this section compares the impact of fixed local windows and learnable local windows, along with their token sizes, on the model's regression performance. Assume the... Original spectral sample sequence At wavelength The feature at a given location includes the spectral intensity within the neighboring window, let the window radius be... Then the local context centered on it contains common... A fixed local window directly captures the wavelength before and after the point. The light intensity at a given point is used as the token feature for that wavelength. A learnable local window linearly projects the light intensity information within the local window into a comparable space to obtain the first... The original spectral sample sequence at wavelength Post-projection features at: ; in, For parameters The time series model, This represents the original light intensity information vector within the local window. , For learnable linear projection weights, This is a learnable bias term.
[0082] The experiments maintained a fixed implementation structure, differing only in the tokenizer generation method and token size. RMSE and R2 were used to evaluate regression performance. All experiments used the same training strategy, the same optimizer Adam, the same learning rate and batch size, the same number of training epochs (300), and the same early stopping strategy; a fixed random seed was used, and the experiments were repeated 5 times, reporting mean ± standard deviation. Experimental results for different token generation strategies are shown below. Figure 8 and Figure 9 As shown.
[0083] Learnable local windows demonstrate a significant advantage in regression performance, indicating that introducing a learnable mechanism can effectively capture local feature structures. Due to the inhomogeneity of gas absorption lines and the differences in the broadening effect of each absorption line, the degree to which a given wavelength is influenced by its surrounding neighborhood varies. Fixed local windows assume that the contribution of each wavelength within a local window to the token representation is constant, while learnable local windows learn a general representation of a local spectral shape to obtain more discriminative token representations for the absorption characteristics of different gases. For methane, the RMSE first decreases and then rebounds as the token window radius increases. The minimum value was reached. This indicates that within a reasonable range, a larger token size helps the model gather richer contextual information to enhance its robustness; however, an excessively large window radius may introduce redundant information and noise, leading to a distraction in the model's attention. For the nitrous oxide regression task, the model generally exhibits robustness to the token radius, and a lower RMSE can be obtained with a larger token size. This is because the data acquired by the mid-infrared instrument has a higher signal-to-noise ratio, resulting in smoother data with lower noise. The feature distribution of nitrous oxide can be characterized through a wider field of view and is less susceptible to interference from background noise.
[0084] Example 7 To further verify the weight relationship between the two attention modules, this invention conducts ablation experiments with controlled variables. To improve the model's expressive power, the two attention modules are decoupled during training. Two learnable parameters are used. Independent optimization, among which Used to control the weights of the CPA attention module. Used to control the weights of the PFE attention module.
[0085] First fix Other model results and parameters, settings Conduct experiments separately. For the methane dataset, such as... Figure 10 As shown by the Chinese block line, when The presence of high RMSE and error bars indicates that guided models lacking sufficient weights for typical characteristic spectra struggle to extract reliable feature information in complex backgrounds. However, as... The optimal value is reached when the value is increased to 0.7, indicating that the introduction of the CPA module can quickly establish a reliable macroscopic baseline for the feature space; for the nitrous oxide dataset, such as Figure 11 As shown by the Chinese block line, When the error bar converges to its minimum value, it indicates that the CPA module provides macroscopic benchmark alignment for the model using typical characteristic spectra. Similarly, fix the following: For methane data, such as Figure 10 As shown by the dotted line in the middle, with As the value increased from 0.1 to 0.7, the error continued to increase until... At this time, the error reaches its minimum, indicating that the attention mechanism can only effectively capture details beneficial to detection when PFE is given a high weight. However, for nitrous oxide data, such as... Figure 11 As shown by the dotted line, the error increases with... The gradual increase in indicates that excessive internal spectral attention mechanisms may lead to overfitting of the model to local noise.
[0086] Experiments show that the dual-view decoupling mechanism proposed in this paper enables the network to dynamically allocate module weights based on the absorption intensity and signal-to-noise ratio of the target gas band, demonstrating the high flexibility of this method in TDLAS regression tasks.
[0087] Example 8 To visualize the internal mechanism of the attention module proposed in this paper, firstly... Figure 12 The weight map of the feature enhancement attention matrix for the nitrous oxide dataset is guided by physical priors. There are obvious vertical stripes in the figure, indicating that all Q tokens in the sequence are jointly focusing on some specific K tokens. The network implicitly models the background baseline containing system noise through these global attention mechanisms. Figure 13This is a weighted graph of the attention matrix for the cross-fusion of typical spectral features from the nitrous oxide dataset. During the retrieval phase, to prevent the neural network from taking shortcuts and implicitly accessing other samples in the same batch, leading to model learning failure, the weights of the remaining samples in the current batch are forcibly set to 0. The model only searches for samples from the reference sample library that have the most similar absorption features to the current sample and assigns them higher weights. Within the baseline band, there is no gas absorption, and different legitimate reference samples show almost no difference; the Softmax mechanism tends to uniformly diffuse the attention weights to stabilize the noise floor. However, in the core absorption region, the color bars of different reference samples show differences, and the probability of the absorption region collapses. For this sample, reference sample number 66 clearly has a larger weight than number 33, meaning that the 66th sample resonates highly with the current input sample. This local resonance provides a basis for the accurate localization of weak signals. A sample is randomly selected from the test set... Figure 14 As shown, after downsampling with stride=5, the original absorption is almost completely submerged in the original sawtooth wave signal. By extracting... Figure 13 The confidence of the aligned tokens is plotted using the row maximum values of the weight matrix, as shown by the purple dashed line. The model locks in weak absorbing pits without polynomial preprocessing.
[0088] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A TDLAS gas concentration detection method based on a dual-view spectral attention mechanism, characterized in that, The gas concentration is detected using the trained TDLAS gas concentration detection model. The model is trained as follows: S11, Obtain a sample set from the TDLAS device. The sample includes the original transmission spectrum and the gas concentration labels corresponding to the spectrum. S12, normalize the amplitude of the original transmission spectra in the sample set to obtain a normalized spectral sample set; S13, perform learnable local window tokenization on the original transmission spectra in the normalized spectral sample set to obtain the spectral token sequence. S14, inject position encoding information into the spectral token sequence to generate a three-dimensional spectral feature cube; S15, use the physical prior correlation matrix to perform intraspectral attention enhancement on the three-dimensional spectral feature cube to generate spectral enhancement features; S16: Based on the wavelength feature planes of the three-dimensional spectral feature cube, construct the query, construct the key and value with the typical feature spectrum of the corresponding band, and complete the feature matching and aggregation through cross attention to obtain the cross-spectral fusion feature; S17, the spectral enhancement features and cross-spectral fusion features are fused through gated residuals to generate the final fused features; S18, learnable weighted aggregation and multilayer perceptron mapping are performed on the final fused features to obtain the network regression output, and the network regression output is mapped to the concentration label scale to obtain the gas concentration value; S19. Based on the gas concentration values obtained from the model output and the corresponding gas concentration labels, a regression loss function is constructed, and the model parameters are iteratively updated using the Adam optimizer combined with an exponential learning rate decay strategy.
2. The TDLAS gas concentration detection method based on the dual-view spectral attention mechanism according to claim 1, characterized in that, The trained TDLAS gas concentration detection model was used to detect the gas concentration. The specific detection process is as follows: S21, Obtain the raw transmission spectrum from the TDLAS device; S22, Based on the model parameters determined during the training phase, the amplitude of the original transmission spectrum is normalized to obtain the normalized test spectrum; S23, perform learnable local window tokenization on the normalized spectrum to be measured to obtain the spectrum token sequence; S24, inject position encoding information into the spectrum token sequence to be measured to generate a three-dimensional spectral feature cube; S25, use the physical prior correlation matrix to perform intraspectral attention enhancement on the three-dimensional spectral feature cube to generate spectral enhancement features; S26, perform inter-spectral cross-attention matching on the three-dimensional spectral feature cube based on typical feature spectra to obtain cross-spectral fusion features; S27, the spectral enhancement features and cross-spectral fusion features are fused through gated residuals to generate the final fused features; S28, the final fused features are subjected to learnable weighted aggregation and multilayer perceptron mapping to obtain network regression output, and the network regression output is mapped to the concentration label scale to obtain the concentration value of the gas to be measured.
3. The TDLAS gas concentration detection method based on the dual-view spectral attention mechanism according to claim 1, characterized in that, The amplitude normalization process described in step S12 uses maximum and minimum value normalization to reduce the amplitude scale differences caused by different bands, different acquisition batches, and laser output power fluctuations, thereby obtaining a normalized spectral sample set.
4. The TDLAS gas concentration detection method based on the dual-view spectral attention mechanism according to claim 1, characterized in that, The specific steps for step S13 are as follows: The original transmission spectrum is segmented along the wavelength dimension according to a preset local window, and the continuous spectrum is divided into multiple spectral segments. Each spectral segment is used as an independent spectral token to form an initial spectral token sequence. The initial spectral token sequence is subjected to a learnable embedding transformation. At the same time, a local sliding window is used to perform feature fusion on adjacent tokens in the wavelength dimension, aggregating the spectral information of adjacent wavelengths to capture the local spectral structure, and finally obtaining a regular spectral token sequence.
5. The TDLAS gas concentration detection method based on the dual-view spectral attention mechanism according to claim 1, characterized in that, The specific steps for step S15 are as follows: from a three-dimensional spectral feature cube the first original transmission spectrum , respectively, linear projection transformation is performed on the sample to generate the query matrix required by the attention mechanism , the key matrix and the value matrix , , wherein is the number of original transmission spectra, indicates the sampling length, is the token feature dimension mapped by each wavelength point, is the projection dimension of the attention mechanism; Introducing the correlation matrix guided by physical priors ; Generate spectral enhancement features : ; in, For normalized exponential functions, These are the learnable weights.
6. The TDLAS gas concentration detection method based on a dual-view spectral attention mechanism according to claim 1, characterized in that, The specific steps for step S16 are as follows: Obtaining a 3D spectral feature cube Obtain typical characteristic spectra and generate a typical characteristic spectrum cube. ,in, The number of original transmission spectra. The number of typical characteristic spectra, Indicates the sampling length. The token feature dimension mapped to each wavelength point; Take a three-dimensional spectral feature cube A certain wavelength point Plane , Generate a query matrix with cross-attention. : ; in, For the projection dimension of the attention mechanism, The linear projection weights of the query matrix; Using typical characteristic spectral cubes in the corresponding Band matrix generation key matrix Sum matrix ; No. The output characteristics of the cross-attention mechanism in each band are as follows: , : ; in, For the three-dimensional spectral feature cube A query matrix for each band. The typical characteristic spectral cube Transpose of the key matrix of each band The typical characteristic spectral cube Value matrix of each band; The output features of all bands are stacked sequentially along the wavelength sequence dimension, and the cross-spectral fusion features are finally reconstructed. : ; in, This is a tensor splicing function.
7. The TDLAS gas concentration detection method based on a dual-view spectral attention mechanism according to claim 1, characterized in that, Step S17 uses two learnable gating parameters to control the contributions of spectral enhancement features and cross-spectral fusion features respectively, generating the final fusion feature. : ; in, It is a three-dimensional spectral feature cube. These are the gating parameters for spectral attention branches guided by physical priors. These are typical characteristic spectral cross-fusion branch gating parameters. For spectral enhancement features, This is a cross-spectral fusion feature.
8. The TDLAS gas concentration detection method based on a dual-view spectral attention mechanism according to claim 1, characterized in that, The specific steps for step S18 are as follows: The final fused features are flattened to obtain a one-dimensional fused feature vector. One-dimensional fused features are input into a multilayer perceptron, and feature transformation and nonlinear mapping are performed sequentially through multiple fully connected layers and nonlinear activation functions, and the network regression output is output. A scale mapping strategy is used to map the network regression output to a scale range consistent with the concentration labels in the training set, thereby obtaining the gas concentration values.
9. The TDLAS gas concentration detection method based on a dual-view spectral attention mechanism according to claim 1, characterized in that, The specific steps of step S19 are as follows: Perform an inverse label transformation on the gas concentration value output in step S18 to restore it to a numerical scale consistent with the corresponding gas concentration label, and obtain the generated gas concentration label. A loss function is constructed based on generated gas concentration labels and real gas concentration labels. The Adam optimizer is used in conjunction with an exponential learning rate decay strategy to complete the iterative update of model parameters.
10. A TDLAS gas concentration detection system based on a dual-view spectral attention mechanism, characterized in that, A TDLAS gas concentration detection method based on a dual-view spectral attention mechanism, applicable to any one of claims 1-9, comprises a gas concentration detection system including: Data acquisition and preprocessing module: Acquires the raw transmission spectrum from the TDLAS device, normalizes the raw transmission spectrum, and generates a normalized spectrum; Spectral sequence encoding module: The normalized spectrum is tokenized using a learnable local window to obtain a spectral token sequence, and positional encoding information is injected into the spectral token sequence to generate a three-dimensional spectral feature cube; Physically Prior-Guided Feature Enhancement Module: This module intervenes in intraspectral attention allocation through a correlation matrix guided by physical priors, generating spectral enhanced features. Typical Feature Spectral Cross-fusion Module: Based on the wavelength feature plane of the three-dimensional spectral feature cube, a query is constructed. Keys and values are constructed with the typical feature spectral features of the corresponding band. Feature matching and aggregation are completed through cross attention to obtain cross-spectral fusion features. Dual-gated feature fusion module: fuses spectral enhancement features and cross-spectral fusion features to generate the final fused feature; Regression Head: Performs learnable linear transformation and weighted aggregation on the final fused features to obtain a spectral representation distributed along the wavelength scanning direction, and inputs the spectral representation into the multilayer perceptron regression network to obtain the network regression output; Learnable label processing module: Maps the network regression output to the concentration label scale to obtain the gas concentration detection value.