Method for detecting hazardous chemical leaks based on near infrared spectroscopy

By combining near-infrared spectroscopy technology with ICA and Bayesian models, the problem of difficulty in achieving real-time monitoring and distinguishing multiple chemical substances in traditional detection methods has been solved, and rapid and accurate detection of hazardous chemical leaks has been achieved.

CN119757271BActive Publication Date: 2025-10-10GUANGZHOU THINKER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411952402.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-10-10
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

Traditional methods for detecting leaks of hazardous chemicals have difficulty achieving efficient data analysis and real-time monitoring, and have difficulty distinguishing mixtures of multiple chemicals.

Method used

Near-infrared spectroscopy technology is combined with independent component analysis (ICA) and Bayesian model to achieve classification detection of various chemical substances through spectral acquisition, data preprocessing, spectral feature extraction and Bayesian model construction.

Benefits of technology

It realizes real-time monitoring and accurate classification of hazardous chemical leaks, can effectively separate overlapping absorption peaks, and improves the accuracy of identifying complex mixtures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119757271B_ABST
    Figure CN119757271B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of near-infrared spectrum analysis, and particularly relates to a dangerous chemical leakage detection method based on near-infrared spectrum, comprising the following steps: S1) spectrum collection: configuring a detector and a near-infrared spectrometer to detect the corresponding absorption spectrum of the leakage gas or liquid; S2) data preprocessing: correcting the absorption spectrum by using a baseline spectrum to eliminate the influence of light source unevenness or stray light of the light path; S3) spectrum feature extraction: adopting principal component analysis or independent component analysis to reduce the high-dimensional spectrum data of wavelength points to a low-dimensional feature space, retaining main information to extract spectrum feature parameters; S4) constructing a Bayesian model; and S5) model application. The near-infrared spectrum technology has the characteristics of rapid detection, can realize real-time monitoring of the dangerous chemical leakage, and can effectively separate overlapping absorption peaks through the ICA method, improving the accuracy of complex mixture identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of near-infrared spectroscopy analysis, and in particular relates to a dangerous chemical leakage detection method based on near-infrared spectroscopy. Background Art

[0002] Hazardous chemical leaks can cause serious environmental pollution, casualties, and property damage. Rapid chemical identification is a prerequisite for effective regulation. Developing a fast, accurate, contactless, and effective identification method for multiple chemicals is crucial for effective chemical regulation. Traditional hazardous chemical leak detection methods, such as gas chromatography-mass spectrometry (GC-MS) and gas sensor monitoring, have the following shortcomings:

[0003] 1. Traditional single sensors have difficulty distinguishing mixtures of multiple chemicals, and their accuracy is limited.

[0004] Second, data analysis time for technologies such as GC-MS is long, making it difficult to achieve efficient data analysis and real-time monitoring.

[0005] Near-infrared spectroscopy (NIRS) is a technique based on the vibrational spectrum of molecules. When near-infrared light strikes a substance, the functional groups within the molecules absorb specific wavelengths of near-infrared light, producing a characteristic absorption spectrum. Combined with advanced data processing and model building methods, near-infrared spectroscopy can overcome the shortcomings of traditional methods and achieve efficient and accurate detection of hazardous chemical leaks. Summary of the Invention

[0006] To address the technical challenges of traditional technologies, which struggle to achieve efficient data analysis and real-time monitoring, as well as the difficulty distinguishing between mixtures of multiple chemicals, this paper provides a hazardous chemical leak detection method based on near-infrared spectroscopy. This method combines ICA with a Bayesian model to achieve classified detection of multiple chemical mixtures.

[0007] The purpose of the present invention can be achieved through the following technical solutions:

[0008] The method for detecting leaks of hazardous chemicals based on near-infrared spectroscopy includes the following steps:

[0009] S1) Spectral acquisition:

[0010] Use broadband near-infrared light source and light intensity stabilizer to ensure the stability of absorption spectrum;

[0011] Configure detectors and near-infrared spectrometers to detect the corresponding absorption spectra of leaked gas or liquid;

[0012] Set the intermittent continuous scanning mode and record the absorption spectrum by adjusting the acquisition frequency and scanning resolution;

[0013] Preferably, in the step S1) of spectrum acquisition, the detectors selected are InGaAs detectors and PbS detectors; the near-infrared spectrometers selected are array-type NIR near-infrared spectrometers or FT-NIR near-infrared spectrometers.

[0014] S2) Data preprocessing: Use the baseline spectrum to correct the absorption spectrum to eliminate the influence of light source inhomogeneity or stray light in the optical path;

[0015] The denoising algorithm is used to smooth the absorption spectrum and improve the signal-to-noise ratio;

[0016] Preferably, the specific steps of correcting the absorption spectrum in the data preprocessing step S2) include:

[0017] The background spectrum is collected in the absence of gas leakage, which is the baseline spectrum;

[0018] The collected absorption spectrum was subtracted from the baseline spectrum to obtain the net absorption spectrum;

[0019] The baseline drift trend was fitted using a polynomial fitting algorithm and subtracted to correct the net absorbance spectrum.

[0020] Preferably, the data preprocessing in step S2) further comprises: normalizing all spectral data to the same scale using a standardization or normalization method to eliminate the influence of amplitude differences between different spectra.

[0021] S3) Spectral feature extraction:

[0022] S31) Wavelength selection: recursive feature elimination method is used to select the wavelength point with the highest discrimination between different chemical categories in the spectrum as the feature;

[0023] S32) Feature extraction: principal component analysis (PCA) or independent component analysis (ICA) is used to reduce the high-dimensional spectral data of the wavelength points to a low-dimensional feature space, retaining the main information to extract spectral feature parameters;

[0024] The spectral characteristic parameters include peak position, peak intensity, peak area and half-peak width; wherein,

[0025] Peak position : Find the wavelength point with the highest absorbance in each independent signal;

[0026] Peak intensity : Calculate the maximum absorbance;

[0027] Peak area A: integration of the area under the absorption peak of the signal;

[0028] Half-peak width F: Calculate the width at which the light intensity at the peak drops to half.

[0029] Preferably, the specific steps of wavelength selection in step S31) are as follows:

[0030] Select a classifier or regression model to evaluate the importance of wavelength;

[0031] Train all wavelength points and calculate the weight or contribution score of each wavelength;

[0032] Remove the wavelength point with the lowest contribution score, retrain the model, and update the weights;

[0033] Continue iterating until the optimal set of wavelengths remains;

[0034] According to the contribution score curve, determine the wavelength point with the highest discrimination for chemical categories .

[0035] Preferably, the specific steps of extracting spectral characteristic parameters using independent component analysis method include:

[0036] The ICA algorithm is used to decompose the spectral data, and the relevant parameters of the ICA algorithm are set, including the number of iterations and the convergence threshold;

[0037] ;

[0038] Where X is the spectral data; S is the independent signal source in the spectrum, and each column represents an independent component; A is the mixing matrix, which describes how the independent components are mixed to form the spectral data;

[0039] Use non-Gaussian measure functions to separate independent signal sources;

[0040] Extract and calculate peak position, peak intensity, peak area, and half-peak width from independent signal sources.

[0041] S4) Build a Bayesian model:

[0042] Define prior distribution: Initialize prior probability based on historical spectral data and the distribution of known chemical classes ;

[0043] Constructing a likelihood function: The spectral characteristic parameters (such as peak position, peak intensity, and peak area) of the labeled chemical categories are combined into a spectral feature vector set, which is used as the observation data of the Bayesian model;

[0044] For each chemical class, the Euclidean distance is used to calculate the degree of match between the observed data and the characteristic reference standard; the characteristic reference standard is the statistical mean or median of the spectral feature vector set;

[0045] The matching degree is used as the likelihood function of the observed data ;

[0046] Posterior probability calculation: Calculate the posterior probability of the sample belonging to each chemical class according to Bayes' theorem; its expression is:

[0047]

[0048] where, is the chemical class; X is the feature vector of the spectrum; is the prior probability of the chemical class; is the likelihood function; is the evidence factor.

[0049] S5) Model application: Input the newly collected and extracted spectral feature vector into the trained Bayesian model, and output the posterior probability of belonging to a certain chemical class;

[0050] Select the class with the maximum posterior probability as the detection result.

[0051] Preferably, in the step S4) of constructing the Bayesian model, the spectral distribution of each chemical class is fitted with a Gaussian Mixture Distribution (GMM), and the corresponding probability density function is calculated; and the probability density function is used as the likelihood function of the observed data.

[0052] Preferably, the step of calculating the probability density function comprises:

[0053] For each chemical class, use Gaussian Mixture Distribution (GMM) to model the distribution of its spectral feature vectors;

[0054] Determine an initial number of mixture components K (i.e. the number of Gaussian distributions), and subsequently optimize it through Akaike Information Criterion (AIC) or Bayesian Information Criterion (BIC) to ensure the balance between model complexity and data fitting ability;

[0055] Use the Expectation Maximization (EM) algorithm to fit the Gaussian mixture model parameters, including the mean vector, covariance matrix and mixing coefficient of each spectral feature vector;

[0056] Use the EM algorithm to iterate the Gaussian mixture model:

[0057] E step: Calculate the posterior probability of each Gaussian component given the current parameters;

[0058] M step: Update the parameters to maximize the log-likelihood function of the model until convergence;

[0059] For each chemical class, its probability density function is represented as:

[0060] ​​

[0061] wherein, is the multi-dimensional Gaussian distribution corresponding to the feature vector X; is the mean vector of each chemical category ; is the covariance matrix of chemical category ; is the mixing coefficient of chemical category .

[0062] Preferably, further comprising: setting a confidence threshold for the posterior probability output by the Bayesian model; if the maximum posterior probability is lower than the confidence threshold, triggering an alarm mechanism to prompt the uncertainty of the model.

[0063] Advantages of the present application:

[0064] 1. Near-infrared spectroscopy technology has the characteristics of rapid detection, which can realize real-time monitoring of dangerous chemical leakage; and through the ICA method, it can effectively separate overlapping absorption peaks and improve the accuracy of complex mixture recognition.

[0065] 2. By constructing a Bayesian probability model; this scheme first defines the prior probability, then constructs the likelihood function (using the Euclidean distance method or the more optimal Gaussian mixture model method), and finally calculates the posterior probability of each chemical category according to Bayes' theorem, realizing the classification detection of multiple mixed chemical categories. BRIEF DESCRIPTION OF DRAWINGS

[0066] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0067] Figure 1 The step flow chart of the dangerous chemical leakage detection method based on near-infrared spectroscopy of the present application.

[0068] Figure 2 The flow chart of step S4) in the dangerous chemical leakage detection method based on near-infrared spectroscopy of the present application. DETAILED DESCRIPTION

[0069] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0070] See also Figure 1-Figure 2 As shown, the method for detecting leakage of hazardous chemicals based on near-infrared spectroscopy includes the following steps:

[0071] S1) Spectral acquisition:

[0072] Use broadband near-infrared light source and light intensity stabilizer to ensure the stability of absorption spectrum;

[0073] Configure detectors and near-infrared spectrometers to detect the corresponding absorption spectra of leaked gas or liquid;

[0074] Set the intermittent continuous scanning mode and record the absorption spectrum by adjusting the acquisition frequency and scanning resolution;

[0075] Furthermore, in the step S1) of spectrum acquisition, the detectors selected are InGaAs detectors and PbS detectors; the near-infrared spectrometer selects an array NIR near-infrared spectrometer or an FT-NIR near-infrared spectrometer.

[0076] In the specific implementation process, FT-NIR spectrometers have higher resolution and accuracy, but are more expensive; array NIR spectrometers have faster scanning speeds and are more suitable for real-time monitoring. In actual applications, optimization and verification must be carried out based on factors such as the absorption spectrum characteristics of specific chemicals, environmental conditions, and detection sensitivity requirements.

[0077] Acquisition frequency: 1-10 Hz is usually sufficient to monitor most chemical leaks. However, for highly toxic, rapidly diffusing gases such as ammonia or hydrogen sulfide, a higher acquisition frequency (10-50 Hz) may be required for a quick response.

[0078] Scanning resolution: 2-5nm resolution is sufficient for most chemicals. For ammonia and hydrogen sulfide, due to their relatively narrow absorption peaks, a higher scanning resolution (1-2nm) may be required.

[0079] As shown in Table 1 below, the embodiments of the present invention only list several common categories of hazardous chemicals and provide some preliminary suggestions for selecting spectrometer parameters. In actual applications, adjustments and optimizations need to be made according to specific circumstances.

[0080]

[0081] S2) Data preprocessing: Use the baseline spectrum to correct the absorption spectrum to eliminate the influence of light source inhomogeneity or stray light in the optical path;

[0082] The denoising algorithm is used to smooth the absorption spectrum and improve the signal-to-noise ratio;

[0083] Further, the specific step of correcting the absorption spectrum in the step S2) data preprocessing includes:

[0084] Collect the background spectrum in the state of no gas leakage, which is the baseline spectrum;

[0085] Subtract the baseline spectrum from the collected absorption spectrum to obtain the net absorption spectrum;

[0086] Fit the baseline drift trend using a polynomial fitting algorithm, and subtract the trend to correct the net absorption spectrum.

[0087] Further, the step S2) data preprocessing further includes: using standardization or normalization method to normalize all spectral data to the same scale, so as to eliminate the influence of amplitude difference between different spectra.

[0088] Specifically, the purpose of correcting the absorption spectrum is to eliminate the influence of baseline drift and light source non-uniformity in the spectral data, and to ensure the accuracy of the spectral signal. The baseline spectrum can be collected multiple times, and the average value is taken to reduce the influence of random noise. The number of collection times depends on the noise level and the required accuracy.

[0089] The subtraction operation of subtracting the baseline spectrum can effectively remove most of the system errors caused by the instrument and the light source. It is necessary to ensure that the absorption spectrum and the baseline spectrum have the same wavelength range and the number of sampling points.

[0090] In addition, if the baseline has a significant drift trend, simple subtraction correction may not be accurate enough. Therefore, a polynomial fitting algorithm is also used to fit the trend of baseline drift. The order of the polynomial needs to be selected according to the complexity of the baseline drift. Too high order may cause overfitting. Subtract the fitting curve from the baseline spectrum to obtain the corrected baseline spectrum. Subtract the corrected baseline spectrum from the absorption spectrum to obtain the final corrected net absorption spectrum.

[0091] As an optional implementation, the purpose of normalization is to eliminate the influence of amplitude difference between different spectra, so that the spectral data has statistical property, thereby improving the generalization ability of the model. Select appropriate baseline correction, denoising and normalization methods, and optimize related parameters; ensure the accuracy and effectiveness of the final model.

[0092] S3) Spectral feature extraction:

[0093] S31) Wavelength selection: recursive feature elimination method is used to select the wavelength points with the highest discrimination for different chemical categories in the spectrum as features;

[0094] S32) Feature extraction: Principal Component Analysis (PCA) or Independent Component Analysis (ICA) is used to reduce the high-dimensional spectral data of wavelength points to a low-dimensional feature space, retaining the main information to extract spectral feature parameters;

[0095] The spectral feature parameters include peak position, peak intensity, peak area and half-peak width; wherein,

[0096] Peak position : Find the wavelength point with the highest absorbance in each independent signal;

[0097] Peak intensity : Calculate the maximum value of absorbance;

[0098] Peak area A: The area under the absorption peak of the signal is integrated;

[0099] Half-peak width F: Calculate the width of the light intensity at half the peak value.

[0100] Specifically, the step S3) spectral feature extraction step aims to extract feature parameters that can effectively distinguish different chemical categories from the pre-processed high-dimensional spectral data.

[0101] The principle of recursive feature elimination method in step S31) wavelength selection is to iteratively remove the features with the smallest contribution to model prediction until the pre-set number of features is reached or the pre-set condition is met. In this way, redundant and irrelevant features can be effectively removed, improving the generalization ability and efficiency of the model, while reducing the complexity of the model. The specific implementation process is as follows:

[0102] Model selection: Select a suitable classification model as the base model for feature selection, such as Support Vector Machine (SVM), Random Forest (RF) or Logistic Regression, etc.

[0103] Feature importance evaluation: Use the selected base model to train the spectral data of the complete wavelength range. Most machine learning models provide feature importance evaluation functions, such as weight coefficients of SVM, feature importance scores of RF, etc. These indicators reflect the contribution of each wavelength to the classification result.

[0104] Iterative elimination: According to the feature importance indicators, iteratively remove the wavelength points with the lowest importance. Each iteration re-trains the base model and re-evaluates the feature importance of the remaining wavelength points. This process continues until the pre-set number of features is reached or the feature importance indicators meet certain conditions.

[0105] Feature subset selection: The final retained wavelength point set is the selected feature subset. This subset contains the wavelength points with the highest discrimination for different chemical categories.

[0106] Step S32) In feature extraction, the principle of Principal Component Analysis (PCA) is to project high-dimensional data into a lower-dimensional space, maximizing the variance information in the data. The principle of Independent Component Analysis (ICA) is to find a set of statistically independent signal sources, which usually represent different chemical components. Whether using PCA or ICA, after dimensionality reduction, further processing of low-dimensional data is needed to extract specific feature parameters, such as peak position, peak intensity, peak area, and half-peak width. These feature parameters need to be interpreted in combination with specific chemical knowledge and spectroscopy knowledge. Some image processing or signal processing techniques can be used to identify and quantify these features. Among them,

[0107] Peak position : Find the wavelength point with the highest absorbance in each independent signal;

[0108] Peak intensity : Calculate the maximum value of absorbance

[0109] Specifically, absorbance is usually defined by the Lambert-Beer law as: ;

[0110] In the formula: is the background light intensity; is the transmitted light intensity.

[0111] Peak area A: The area integral under the absorption peak of the signal; the calculation formula of peak area A is: ;

[0112] In the formula: and are the starting wavelength and ending wavelength of the absorption peak integral, respectively, determining the integral range; represents the absorbance at wavelength ; represents the small wavelength increment of the integral variable, used for continuous integration.

[0113] Half-peak width F: Calculate the width where the light intensity drops to half at the peak. The specific calculation method is:

[0114] Find the peak height: determine the maximum intensity value at the peak .

[0115] Determine the half-peak height: calculate the half-peak height, which is half of the peak height .

[0116] Find the two wavelength points corresponding to the half-peak height: on both sides of the peak, find the two wavelength points ( ) whose intensity value is equal to or closest to the half-peak height. Interpolation methods can be used to more accurately determine these two wavelength points. For example, linear interpolation.

[0117] Calculate the half-peak width: The half-peak width F is equal to the difference between the two wavelength points: .

[0118] As an optional implementation, spectrum analysis software (such as Origin, MATLAB) provides the function of directly calculating and visualizing the half-peak width.

[0119] Furthermore, the specific steps of wavelength selection in step S31) are as follows:

[0120] Select a classifier or regression model to evaluate the importance of wavelength;

[0121] Train all wavelength points and calculate the weight or contribution score of each wavelength;

[0122] Remove the wavelength point with the lowest contribution score, retrain the model, and update the weights;

[0123] Continue iterating until the optimal set of wavelengths remains;

[0124] According to the contribution score curve, determine the wavelength point with the highest discrimination for chemical categories .

[0125] Furthermore, the specific steps of extracting spectral characteristic parameters using independent component analysis method include:

[0126] The ICA algorithm is used to decompose the spectral data, and the relevant parameters of the ICA algorithm are set, including the number of iterations and the convergence threshold;

[0127] ;

[0128] Where X is the spectral data; S is the independent signal source in the spectrum, and each column represents an independent component; A is the mixing matrix, which describes how the independent components are mixed to form the spectral data;

[0129] Use non-Gaussian measure functions to separate independent signal sources;

[0130] Extract and calculate peak position, peak intensity, peak area, and half-peak width from independent signal sources.

[0131] As an optional implementation, the independent component analysis (ICA) method is used to extract spectral feature parameters, which has a significant advantage especially when dealing with complex spectral data, i.e., multiple gases or liquid mixtures exist simultaneously. Because ICA can separate overlapping absorption peaks and extract independent signal sources, which is difficult to achieve for traditional univariate analysis methods. ICA is a blind source separation technique, which means it does not need to know the specific form of each source signal in the mixed signal in advance, but only needs to assume that they are statistically independent. This is very useful in hazardous chemical leakage scenarios, because the types and concentrations of leaked substances are usually unknown. This is difficult to achieve for traditional peak area or peak height calculation methods, because these methods are difficult to handle overlapping peaks.

[0132] In the spectrum of a mixture, the absorption peaks of different substances often overlap, making it difficult to distinguish individual components. ICA can effectively separate these overlapping absorption peaks and identify individual components by assuming that the mixed signal is linearly mixed by multiple statistically independent source signals. ICA provides an effective method to separate overlapping absorption peaks and extract independent signal sources.

[0133] S4) Constructing Bayesian model:

[0134] Define prior distribution: initialize prior probability based on historical spectral data and distribution of known chemical categories ;

[0135] In the specific implementation process, if there is a lack of prior knowledge, a uniform prior probability can be assumed, i.e., the prior probability of all chemicals is the same. The frequency of each chemical appearing can also be estimated using historical data statistics to estimate the prior probability. The prior probability reflects the estimate of the probability of each chemical category appearing without observation data.

[0136] Construct likelihood function: form a spectral feature vector set by spectral feature parameters (such as peak position, peak intensity and peak area, etc.) of labeled chemical categories, and use it as observation data of Bayesian model;

[0137] For each chemical category, use the Euclidean distance to calculate the matching degree of the observation data and the characteristic reference standard; the characteristic reference standard is the statistical mean or median of the spectral feature vector set;

[0138] Take the matching degree as the likelihood function of the observation data ;

[0139] In the specific implementation process, the likelihood function describes the likelihood of the observation data given the chemical category The probability of observing the spectral feature vector X under the condition of the sample belonging to the chemical class c is calculated. The spectral feature parameters (e.g. peak position, peak intensity, and peak area, etc.) of the labeled chemical classes are grouped into a spectral feature vector set as the feature reference standard. The Euclidean distance between the observation data and the feature reference standard (e.g. mean or median) is calculated. The smaller the distance, the higher the matching degree, and the higher the likelihood function value.

[0140] Posterior probability calculation: According to the Bayes formula, the posterior probability of the sample belonging to each chemical class is calculated; its expression is:

[0141] ;

[0142] In the formula, c is the chemical class; X is the feature vector of the spectrum; P(c) is the prior probability of the chemical class; L(X|c) is the likelihood function; E is the evidence factor.

[0143] S5) Model application: input the new spectral feature vector collected and extracted in real time into the trained Bayesian model, and output the posterior probability belonging to a certain chemical class;

[0144] The class with the maximum posterior probability is selected as the detection result.

[0145] Specifically, the principle of the Bayesian model is to combine the prior probability and the likelihood function to calculate the posterior probability using the Bayes theorem. The prior probability reflects the prior estimate of the occurrence of each chemical class, and the likelihood function reflects the degree of matching between the observation data and each chemical class. The posterior probability considers both the prior probability and the likelihood function, and represents the updated estimate of the probability of the occurrence of each chemical class after observing the data.

[0146] Further, in the step S4) of constructing the Bayesian model, the Gaussian Mixture Distribution (GMM) is used to fit the spectral distribution of each chemical class, and the corresponding probability density function is calculated; and the probability density function is used as the likelihood function of the observation data.

[0147] The calculation steps of the probability density function include:

[0148] For each chemical class, the Gaussian Mixture Distribution (GMM) is used to model the distribution of its spectral feature vector;

[0149] An initial number of mixture components K (i.e. the number of Gaussian distributions) is determined, and the Akaike Information Criterion (AIC) or the Bayesian Information Criterion (BIC) is used for subsequent optimization to ensure the balance between model complexity and data fitting ability;

[0150] The expectation maximization (EM) algorithm was used to fit the Gaussian mixture model parameters, including the mean vector, covariance matrix, and mixing coefficients of each spectral feature vector;

[0151] Use the EM algorithm to iterate the Gaussian mixture model:

[0152] Step E: Calculate the posterior probability of each Gaussian component given the current parameters;

[0153] Step M: Update the parameters to maximize the log-likelihood function of the model until convergence;

[0154] For each chemical category, its probability density function is expressed as:

[0155] ;

[0156] Where, is the multidimensional Gaussian distribution corresponding to the eigenvector X; Each chemical category The mean vector of ; Chemical Category The covariance matrix of Chemical Category The mixing coefficient.

[0157] Specifically, in constructing a likelihood function, the present invention also provides a Gaussian mixture model (GMM) method. This method assumes that the spectral feature vectors of each chemical class follow a Gaussian mixture distribution. The EM algorithm is used to fit the GMM parameters, including the mean vector, covariance matrix, and mixing coefficients of each Gaussian component. The fitted GMM is then used to calculate the probability density function of the observed data, which serves as the likelihood function.

[0158] Another notable innovation of this invention is that, to improve the reliability of the Bayesian classification model, ICA is combined with a Gaussian mixture model (GMM). Using the multimodal distribution constructed by the GMM, the spectral signature distribution of each chemical category is fitted to generate the observation probability of each chemical category, i.e., the probability density function. This indirectly identifies the spectral signals of multiple chemical leak sources, making this method suitable for multi-chemical leak detection in complex scenarios, thereby enhancing the robustness of the model classification.

[0159] Furthermore, it also includes: setting a credibility threshold for the posterior probability output by the Bayesian model; if the maximum posterior probability is lower than the credibility threshold, triggering an alarm mechanism to prompt the uncertainty of the model.

[0160] This method for detecting hazardous chemical leaks based on near-infrared spectroscopy utilizes independent component analysis (ICA) to separate independent components within spectral signals. This method is particularly suitable for feature decomposition of spectral data from complex mixtures. The feature extraction results from absorption spectra provide high-quality input data for chemical leak detection. Bayesian inference is employed, combined with a Gaussian mixture model (GMM) to fit the spectral distributions of multiple chemical categories, thereby improving classification accuracy, adaptability, and real-time responsiveness.

[0161] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the equipment and tool devices described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0162] In the several embodiments provided in this application, it should be understood that the disclosed methods and processes can be implemented in other ways. For example, the module embodiments described above are merely illustrative. If the calculation method or steps are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0163] Throughout the specification, references to terms such as "as an alternative embodiment," "example," and "particularly" indicate that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0164] The above content is merely an example and explanation of the structure of the present invention. Those skilled in the art may make various modifications or additions to the described specific embodiments or replace them in a similar manner. As long as they do not deviate from the structure of the invention or exceed the scope defined by the claims, they should all fall within the scope of protection of the present invention.

Claims

1. A method for detecting hazardous chemical leaks based on near-infrared spectroscopy, characterized by: The following steps are involved: S1) Spectral acquisition: Use a broadband near-infrared light source and a light intensity stabilizer to ensure the stability of the absorption spectrum. Configure a detector and near-infrared spectrometer to detect the corresponding absorption spectrum of the leaked gas or liquid. Set an intermittent continuous scanning mode and record the absorption spectrum by adjusting the acquisition frequency and scanning resolution. S2) Data preprocessing: The absorption spectrum is corrected using the baseline spectrum to eliminate the influence of light source inhomogeneity or stray light in the optical path; a denoising algorithm is used to smooth the absorption spectrum curve to improve the signal-to-noise ratio; S3) Spectral feature extraction: S31) Wavelength selection: using recursive feature elimination to select the wavelength point with the highest discrimination between different chemical categories in the spectrum as a feature; S32) Feature extraction: Using independent component analysis to reduce the high-dimensional spectral data of the wavelength points to a low-dimensional feature space, retaining the main information to extract spectral feature parameters; The spectral characteristic parameters include peak position, peak intensity, peak area and half-peak width; S4) Building a Bayesian model: Based on the historical spectral data and the distribution of known chemical classes, a Bayesian model is built to calculate the posterior probability that the sample belongs to each chemical class; S5) Model application: The new spectral feature vectors collected and extracted in real time are input into the trained Bayesian model to output the posterior probability of belonging to a certain chemical category; the category with the largest posterior probability is selected as the detection result; The specific steps of extracting spectral characteristic parameters using independent component analysis method include: The ICA algorithm is used to decompose the spectral data, and the relevant parameters of the ICA algorithm are set, including the number of iterations and the convergence threshold; X = A·S; Where X is the spectral data; S is the independent signal source in the spectrum, and each column represents an independent component; A is the mixing matrix, which describes how the independent components are mixed to form the spectral data; Use non-Gaussian measure functions to separate independent signal sources; Extract and calculate peak position, peak intensity, peak area and half-peak width from independent signal sources; In step S4), in the process of constructing the Bayesian model, the spectral distribution of each chemical category is fitted with a Gaussian mixture distribution, and the corresponding probability density function is calculated; and the probability density function is used as the likelihood function of the observed data; The calculation steps of the probability density function include: For each chemical class, the distribution of its spectral feature vectors was modeled using a Gaussian mixture distribution; Determine an initial number of mixture components K and perform subsequent optimization using the Akaike Information Criterion or the Bayesian Information Criterion to ensure a balance between model complexity and data fitting ability; The Gaussian mixture model parameters, including the mean vector, covariance matrix, and mixing coefficients of each spectral feature vector, were fitted using the expectation maximization algorithm. Use the EM algorithm to iterate the Gaussian mixture model: Step E: Calculate the posterior probability of each Gaussian component given the current parameters; Step M: Update the parameters to maximize the log-likelihood function of the model until convergence; For each chemical category, its probability density function is expressed as: Where, N(X;μ ik ,Σ ik ) is the multidimensional Gaussian distribution corresponding to the eigenvector X; μ ik Each chemical category C i The mean vector of Σ ik It is a chemical category C i The covariance matrix of ik It is a chemical category C i The mixing coefficient.

2. The method for detecting dangerous chemical leaks based on near-infrared spectroscopy according to claim 1, characterized in that: In the spectrum acquisition step S1), the detector is an InGaAs detector or a PbS detector; and the near-infrared spectrometer is an array-type NIR near-infrared spectrometer or a FT-NIR near-infrared spectrometer.

3. The method for detecting dangerous chemical leaks based on near-infrared spectroscopy according to claim 1, characterized in that: The specific steps of correcting the absorption spectrum in the data preprocessing step S2) include: The background spectrum is collected in the absence of gas leakage, which is the baseline spectrum; The collected absorption spectrum was subtracted from the baseline spectrum to obtain the net absorption spectrum; The baseline drift trend was fitted using a polynomial fitting algorithm and subtracted to correct the net absorbance spectrum.

4. The method for detecting dangerous chemical leaks based on near-infrared spectroscopy according to claim 1, characterized in that: The data preprocessing in step S2) further includes: normalizing all spectral data to the same scale using a standardization or normalization method to eliminate the influence of amplitude differences between different spectra.

5. The method for detecting dangerous chemical leaks based on near-infrared spectroscopy according to claim 1, characterized in that: Among the spectral characteristic parameters, the peak position λ i : Find the wavelength point with the highest absorbance in each independent signal; Peak intensity I: calculate the maximum value of absorbance; Peak area A: integration of the area under the absorption peak of the signal; Half-peak width F: Calculate the width at which the light intensity at the peak drops to half.

6. The method for detecting dangerous chemical leaks based on near-infrared spectroscopy according to claim 1, characterized in that: The specific steps of wavelength selection in step S31) are as follows: Select a classifier or regression model to evaluate the importance of wavelength; Train all wavelength points and calculate the weight or contribution score of each wavelength; Remove the wavelength point with the lowest contribution score, retrain the model, and update the weights; Continue iterating until the optimal set of wavelengths remains; According to the contribution score curve, determine the wavelength point λ that has the highest discrimination for chemical categories i .

7. The method for detecting dangerous chemical leaks based on near-infrared spectroscopy according to claim 1, characterized in that: The specific process of constructing the Bayesian model in step S4) includes: Define the prior distribution: Based on the historical spectral data and the distribution of known chemical categories, initialize the prior probability P(C i ); Constructing the likelihood function: The spectral characteristic parameters of the labeled chemical categories are combined into a spectral characteristic vector set, which is used as the observation data of the Bayesian model; For each chemical class, the Euclidean distance is used to calculate the degree of match between the observed data and the characteristic reference standard; the characteristic reference standard is the statistical mean or median of the spectral feature vector set; The matching degree is used as the likelihood function P(X|C i ); Posterior probability calculation: According to the Bayesian formula, the posterior probability of the sample belonging to each chemical category is calculated; its expression is: Where C i is the chemical category; X is the characteristic vector of the spectrum; P(C i ) is the prior probability of the chemical category; P(X|C i ) is the likelihood function; P(X) is the evidence factor.

Citation Information

Patent Citations

  • Classification method combining Gaussian regression mixture model and MRF hyperspectral function data

    CN116343032A

  • Pollution detection method for food

    CN118010649A