VOCs data digital intelligent quality control auditing and processing method and system and electronic equipment
By employing multi-level anomaly analysis algorithms and equipment correction operations, the problems of numerous stray peaks and excessive noise in spectral data during volatile organic compound (VOC) detection have been resolved, enabling automatic correction of equipment anomalies and ensuring data accuracy.
Patent Information
- Application Number
- CN202511390497.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2025-12-26
AI Technical Summary
During the detection and review of volatile organic compounds (VOCs), the spectral data generated by gas chromatographs and gas chromatography-mass spectrometry systems under long-term monitoring have many cluttered peaks, baseline drift and excessive noise, resulting in low accuracy in identifying abnormal data and inability to effectively correct equipment malfunctions.
A multi-level anomaly analysis algorithm and equipment correction operation are adopted, including data preprocessing, peak analysis processing, substance matching, chemical peak concentration determination, multi-level anomaly data analysis and preset anomaly correction model. The multi-level anomaly data analysis algorithm is used to perform anomaly analysis on the chemical peak concentration set and generate anomaly correction results.
It enables automatic correction of equipment anomalies, ensuring data accuracy and reliability, and avoiding equipment failures and environmental interference caused by abnormal data.
Smart Images

Figure CN121208232A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of computer technology, specifically to a method, system, and electronic device for intelligent quality control auditing and processing of VOCs data. Background Technology
[0002] Volatile organic compound (VOCs) detection and auditing refers to the determination and auditing of the types, concentrations, and emissions of VOCs in environmental media (air, water, soil) using chemical or physical methods. It is commonly used in environmental assessment, health protection, and compliance management, and is a key technology system for environmental quality management. Currently, the typical approach to VOCs detection and auditing is to use gas chromatography (GC) and gas chromatography-mass spectrometry (GC-MS) for qualitative and quantitative monitoring, and to audit the data through offline VOCs spectral data transfer, without timely pollution source tracing and secondary sampling.
[0003] However, when using the above methods to detect and audit volatile organic compounds, the following technical problems often arise:
[0004] In industrial-grade monitoring, gas chromatographs and chromatography-mass spectrometry (GC-MS) instruments often operate under continuous, long-term, and multi-species monitoring conditions. This results in numerous cluttered peaks in the spectral data, accompanied by baseline drift, peak drift, and excessive noise. Consequently, the accuracy of anomaly identification is low, data review efficiency is inefficient, and ultimately, it becomes impossible to correct equipment malfunctions based on the anomaly analysis results.
[0005] The information disclosed in this background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0007] Some embodiments of this disclosure propose a VOCs data processing method, system, electronic device, and computer-readable medium based on a multi-level anomaly analysis algorithm and device correction operations to solve one or more of the technical problems mentioned in the background section above.
[0008] In a first aspect, some embodiments of this disclosure provide a method for intelligent quality control auditing and processing of VOCs data. This method includes: preprocessing the acquired volatile organic compound (VOC) spectral data to obtain preprocessed spectral data; performing peak analysis on the preprocessed spectral data to obtain a chemical peak feature dataset; performing substance matching on the chemical peak feature dataset to obtain substance matching results; determining the concentration of each chemical peak feature data in the chemical peak feature dataset based on the substance matching results using a pre-calibrated calibration curve to obtain chemical peak concentrations, which serve as a chemical peak concentration set; performing anomaly analysis on the chemical peak concentration set based on a multi-level anomaly data analysis algorithm to obtain anomaly analysis results, wherein the anomaly analysis results include statistical anomaly detection results, temporal mutation detection results, and multi-component logic verification results; and performing equipment anomaly correction on the anomaly analysis results based on a preset anomaly correction model to obtain anomaly correction results.
[0009] Secondly, some embodiments of this disclosure provide a VOCs data intelligent quality control audit and processing system. The system includes: a preprocessing unit configured to preprocess the acquired volatile organic compound spectral data to obtain preprocessed spectral data; a peak analysis unit configured to perform peak analysis processing on the preprocessed spectral data to obtain a chemical peak feature dataset; a substance matching unit configured to perform substance matching on the chemical peak feature dataset to obtain substance matching results; a determination unit configured to determine the concentration of each chemical peak feature data in the chemical peak feature dataset based on the substance matching results using a pre-calibrated calibration curve to obtain chemical peak concentrations as a chemical peak concentration set; an anomaly analysis unit configured to perform anomaly analysis on the chemical peak concentration set based on a multi-level anomaly data analysis algorithm to obtain anomaly analysis results, wherein the anomaly analysis results include statistical anomaly detection results, temporal mutation detection results, and multi-component logic verification results; and an anomaly correction unit configured to perform equipment anomaly correction on the anomaly analysis results based on a preset anomaly correction model to obtain anomaly correction results.
[0010] Thirdly, some embodiments of this disclosure provide an electronic device for intelligent quality control auditing and processing of VOCs data, including: one or more processors; and a storage device storing one or more programs thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the implementations of the first aspect above.
[0011] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.
[0012] The above-described embodiments of this disclosure have the following beneficial effects: The VOCs data intelligent quality control audit and processing method of some embodiments of this disclosure avoids the inability to correct equipment anomalies based on abnormal analysis results. Specifically, the reason for the inability to correct equipment anomalies based on abnormal analysis results is that in industrial monitoring, gas chromatographs and chromatography-mass spectrometry instruments often need to operate under continuous, long-term, and multi-species monitoring conditions, resulting in numerous cluttered peaks in the generated spectral data, accompanied by baseline drift, peak drift, and excessive noise. This leads to low accuracy in identifying abnormal data and low efficiency in data auditing. Based on this, the VOCs data intelligent quality control audit and processing method of some embodiments of this disclosure first preprocesses the acquired volatile organic compound spectral data to obtain preprocessed spectral data. This eliminates high-frequency noise and instrument drift interference, providing high-quality input for subsequent analysis. Second, peak analysis processing is performed on the preprocessed spectral data to obtain a chemical peak feature dataset. This accurately identifies and segments the chemical peak features (position, area, and type) in the spectrum, establishing a data foundation for qualitative and quantitative analysis of substances. Then, substance matching is performed on the aforementioned chemical peak feature dataset to obtain substance matching results. The chemical peak features are then compared with a standard mass spectrometry library to determine the target substance type, solving the problem of identifying "what kind of pollutant it is." Next, based on the substance matching results, the concentration of each chemical peak feature data in the aforementioned chemical peak feature dataset is determined using a pre-calibrated calibration curve, obtaining the chemical peak concentration set. This establishes a quantitative relationship regarding "how much pollutant is present." Next, based on a multi-level anomaly data analysis algorithm, anomaly analysis is performed on the aforementioned chemical peak concentration set to obtain anomaly analysis results. These results include statistical anomaly detection results, temporal mutation detection results, and multi-component logical verification results. Thus, using statistical, temporal, and logical verification three-level detection, equipment malfunctions, environmental interference, and actual pollution events in the concentration data are identified. Then, based on a preset anomaly correction model, equipment anomaly correction is performed on the aforementioned anomaly analysis results to obtain anomaly correction results. This automatically triggers calibration, retesting, or compensation operations for different anomaly types, ensuring data accuracy and reliability. This implementation method avoids the inability to perform equipment anomaly correction based on anomaly analysis results. Attached Figure Description
[0013] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0014] Figure 1 This is a flowchart of some embodiments of the VOCs data intelligent quality control audit and processing method disclosed herein;
[0015] Figure 2 This is a structural schematic diagram of some embodiments of the VOCs data intelligent quality control audit and processing system according to the present disclosure;
[0016] Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0020] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0021] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0022] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0023] Figure 1 A flowchart 100 is shown, illustrating some embodiments of the VOCs data intelligent quality control audit and processing method according to this disclosure. The VOCs data intelligent quality control audit and processing method includes the following steps:
[0024] Step 101: Perform data preprocessing on the acquired volatile organic compound spectral data to obtain preprocessed spectral data.
[0025] In some embodiments, the executing entity (e.g., a server) of the VOCs data intelligent quality control audit and processing method can preprocess the acquired volatile organic compound (VOC) spectral data to obtain preprocessed spectral data. The acquired VOC spectral data includes: raw VOC spectral data, the data acquisition time, and the coordinates corresponding to the acquisition time. The raw VOC spectral data is the initial detection result generated by sampling and analyzing VOCs emitted from the environment or pollution sources using specialized instruments; its essence is a spectrum reflecting the distribution characteristics of compounds in the dimensions of time and signal intensity.
[0026] In practice, the aforementioned implementing entities can perform data preprocessing on the acquired volatile organic compound (VOC) spectral data through the following steps to obtain preprocessed spectral data:
[0027] Step one involves removing high-frequency noise from the aforementioned volatile organic compound (VOC) spectral data to obtain a preliminarily denoised spectral signal. In practice, the execution entity can first decompose the VOC spectral data using the Mallat algorithm to obtain wavelet decomposition results. Then, a low-pass filter is used to filter the high-frequency noise in the wavelet decomposition results, resulting in a spectral signal with high-frequency noise removed, which serves as the preliminarily denoised spectral signal.
[0028] Step two involves reconstructing the spectral signal after initial denoising using a reconstruction algorithm to preserve the effective low-frequency signal, thus obtaining the denoised spectral data. For example, the reconstruction algorithm could be a Daubechies wavelet.
[0029] Step three involves baseline calibration of the denoised spectral data to obtain preprocessed spectral data. In practice, firstly, the execution entity can smooth the denoised spectral data using Fourier transform to obtain a smoothed spectral baseline. Then, an SG (Savitzky-Golay) filter is used to perform baseline calibration on the smoothed spectral baseline to obtain preprocessed spectral data.
[0030] Step 102: Perform peak analysis on the preprocessed spectral data to obtain a chemical peak feature dataset.
[0031] In some embodiments, the execution entity may perform peak analysis on the preprocessed spectral data to obtain a chemical peak feature dataset. This chemical peak feature dataset includes a target type chemical peak set and a peak shape parameter set. The peak shape parameters in the peak shape parameter set may include peak height, peak center, peak width, peak area, and the starting point t of the chemical peak. L (Starting point) and ending point t R (Peak point).
[0032] In some optional implementations of certain embodiments, the aforementioned execution entity may obtain the chemical peak feature dataset through the following steps:
[0033] Step 1: Identify the chemical peak types of the preprocessed spectral data to obtain the target type chemical peak set.
[0034] In practice, the aforementioned executing entity can identify the chemical peak types of the preprocessed spectral data through the following sub-steps to obtain the target type of chemical peak set:
[0035] Sub-step one: Apply the continuous wavelet transform (CWT) to the preprocessed spectral data to generate a time-frequency matrix:
[0036]
[0037] Where W(a,b) represents the time-frequency matrix. a represents the scaling parameter (controlling frequency resolution). b represents the translation parameter (time positioning). t represents time. s(t) represents the preprocessed spectral data. Show the wavelet function after scaling and translation Complex Conjugate.
[0038] Sub-step two involves mapping the aforementioned time-frequency matrix W(a, b) to a grayscale image I(x, y). Here, the horizontal axis x represents the retention time (x is directly proportional to b), and the vertical axis y represents the scale parameter (y is directly proportional to ln a). The pixel value I(x, y) is calculated using the normalized energy density |W(a, b)|. 2 Indicates. | | 2 It represents the square of the modulus.
[0039] Sub-step three involves inputting the grayscale image I(x,y) into a preset recognition model for peak type identification to obtain the target type chemical peaks. In practice, the execution entity can input the grayscale image I(x,y) into the preset recognition model for peak type identification to obtain the target type chemical peak set. The preset recognition model can be a model used for image classification and segmentation. For example, the preset recognition model can be a ResNet-34 model.
[0040] Step two involves locating the peak boundaries of each target-type chemical peak in the aforementioned target-type chemical peak set to generate position information for each peak, thus obtaining a position information set. In practice, the execution entity can use a local minima-base point moving peak-shaving algorithm to locate the peak boundaries of each target-type chemical peak in the aforementioned target-type chemical peak set, thereby locating the starting point t of the target chemical peak. L(Starting point) and ending point t R (Peak landing point), obtaining the position information set. The aforementioned local minimum-base point moving peak-shaving algorithm can be an algorithm that uses local minima as base points to identify at least one valley point (local minimum) between peaks in the chromatogram, obtaining a valley point group. Each valley point in the valley point group represents the boundary region between adjacent peaks. In a simple case, the peak between two valley points is defined as an independent peak, with the valley point serving as the boundary. In response to baseline drift or peak asymmetry, a straight line is drawn between adjacent valley points as the baseline, and the starting point t of the chemical peak is determined. L (Starting point) and ending point t R The (peak point) is defined as the intersection of the chromatographic curve and the baseline mentioned above. For incompletely separated peaks, the valley point may be higher than the true baseline. Moving the baseline (i.e., adjusting the valley point position) allows for a more accurate determination of the separation line. Searching to the left of the valley point (in the direction of decreasing time), find the point where the first derivative of the spectrum of the target type of chemical peak changes from negative to positive (or reaches a threshold close to zero), and use this as the starting point t. L (Peak point). Searching to the right from the valley point (in the direction of increasing time), find the point where the first derivative of the spectrum of the target type of chemical peak changes from positive to negative (or reaches a threshold close to zero), and take this as the ending point t. R (Peak point).
[0041] Step 3: Based on the aforementioned location information set, extract peak shape parameters for each target type chemical peak in the aforementioned target type chemical peak set to obtain a chemical peak feature dataset. In practice, the aforementioned execution entity can use the Gaussian iteration method and the aforementioned location information set to extract peak shape parameters for each target type chemical peak in the aforementioned target type chemical peak set to obtain a chemical peak feature dataset. Peak shape parameter extraction can be achieved using the following formula:
[0042] First, the target type chemical peaks are considered to have a Gaussian distribution:
[0043]
[0044] Where t represents time, μ represents the peak center, σ represents the standard deviation, and H represents the peak height.
[0045] Secondly, the parameters H, μ, and σ are solved by least squares fitting to satisfy:
[0046]
[0047] Where t represents time, μ represents the peak center, σ represents the standard deviation, and H represents the peak height. L Indicates the starting point (peak point) of the target chemical peak. t R This indicates the endpoint (fall point) of the target chemical peak. s(t) represents the spectral data of the target chemical peak.
[0048] Finally, the peak width w and peak area A are determined:
[0049] w = 2.355σ.
[0050]
[0051] Where H represents peak height, w represents peak width, A represents peak area, and σ represents standard deviation.
[0052] Step 103: Perform substance matching on the chemical peak feature dataset to obtain the substance matching results.
[0053] In some embodiments, the aforementioned execution entity may perform substance matching on the aforementioned chemical peak feature dataset to obtain substance matching results.
[0054] In practice, the aforementioned implementing entity can perform substance matching on the chemical peak feature dataset through the following steps to obtain the substance matching results:
[0055] Step one: Input the aforementioned chemical peak feature dataset into the time-series prediction model to obtain the predicted retention time and confidence level. Here, the predicted retention time represents the starting point t of the chemical peak. L The time required to reach the peak center μ. The structure of the above time series prediction model may include:
[0056] Input feature processing layer: This layer converts the aforementioned chemical peak feature dataset into numerical vectors and maps each feature to the [0,1] interval through data normalization. In practice, the execution entity can map each feature to the [0,1] interval by dividing the peak center μ by the total chromatographic duration. Then, it takes the logarithm of the peak area A (e.g., log...). 10 A). Perform Min-Max normalization on the peak width w. Apply a positional encoding with twice the weight to the peak center μ to strengthen the dominant role of the peak center μ:
[0057]
[0058] Where i represents the dimension index, d represents the model embedding dimension, μ represents the peak center, and PE() represents the position encoding.
[0059] Transformer encoder: Used to remove the text embedding layer of the traditional Transformer and replace it with numerical features directly connected to the input layer. Only the encoder structure is retained (no decoder). A multi-head self-attention mechanism is employed. Residual connections and layer normalization are used.
[0060] Output layer: Used to output prediction retention time and confidence level.
[0061] Step two: Based on the sequential decision-making mathematical model, the aforementioned prediction retention time, and the aforementioned confidence level, chemical peak-substance matching is performed on the aforementioned chemical peak feature dataset to obtain the optimal substance matching relationship. In practice, the aforementioned executing entity can use the sequential decision-making mathematical model for matching to obtain the optimal substance matching relationship. The aforementioned sequential decision-making mathematical model can be a decision model divided according to the number of answers obtained according to the decision requirements and their interrelationships. For example, the aforementioned sequential decision-making mathematical model can be a Markov decision process optimized using Q-learning. The aforementioned optimal substance matching relationship includes the aforementioned chemical peak feature dataset and the names of substances that match the aforementioned chemical peak feature dataset.
[0062] Step 3: Based on the mass spectrometry library, query the optimal substance matching relationships to obtain the substance matching results. In practice, the executing entity can initiate a query for the optimal substance matching relationships based on the mass spectrometry library (such as the NIST mass spectrometry database) to obtain the chemical peak characteristics of the substance names in the mass spectrometry library. Mass spectrometry similarity is determined using a weighted cosine similarity algorithm.
[0063]
[0064] Where Sim represents mass spectrometry similarity. E and F represent the two chemical peak vectors to be compared. n represents the number of features in the aforementioned chemical peak vectors. i represents the index of the aforementioned chemical peak vector. E i F i Let w represent the i-th eigenvalue of the chemical peak vector. i The weights representing the i-th feature need to be predefined (e.g., w). i (It can be 1 or 2). δ represents the summation operator. As an example, if Sim > 0.85, the two chemical peak vectors can be considered to be the same, resulting in a substance matching result.
[0065] Step 104: Based on the substance matching results, use the pre-calibrated calibration curve to determine the concentration of each chemical peak feature data in the chemical peak feature dataset, and obtain the chemical peak concentration as the chemical peak concentration set.
[0066] In some embodiments, the execution entity may determine the concentration of each chemical peak feature data in the chemical peak feature dataset based on the substance matching results and using a pre-calibrated calibration curve to obtain the chemical peak concentration as a chemical peak concentration set.
[0067] In practice, the aforementioned implementing entity can use a pre-calibrated calibration curve to determine the concentration of each chemical peak characteristic data in the aforementioned chemical peak characteristic dataset. The pre-calibrated calibration curve is a quantitative relationship model between the instrument response signal and the concentration of the analyte, established experimentally. The calibration curve can be expressed as:
[0068] C = p·A 2 +q·A+r.
[0069] Where C represents the concentration of the analyte, A represents the peak area of the analyte, p represents the empirical coefficient of the quadratic term, q represents the empirical coefficient of the linear term, and r represents the constant empirical coefficient. The empirical coefficients p, q, and r can be determined given the concentration C and peak area A of the analyte and used as known data.
[0070] Step 105: Based on the multi-level anomaly data analysis algorithm, perform anomaly analysis on the concentration of each chemical peak in the chemical peak concentration set to obtain the anomaly analysis results.
[0071] In some embodiments, the aforementioned execution entity may perform anomaly analysis on the aforementioned chemical peak concentration set based on a multi-level anomaly data analysis algorithm to obtain anomaly analysis results. These anomaly analysis results include statistical anomaly detection results, temporal abrupt change detection results, and multi-component logic verification results.
[0072] In practice, the aforementioned implementing entity can perform anomaly analysis on the concentration of each chemical peak in the chemical peak concentration set based on a multi-level anomaly data analysis algorithm through the following steps to obtain the anomaly analysis results:
[0073] Step one involves performing statistical anomaly detection on the concentration of each chemical peak in the aforementioned chemical peak concentration set, obtaining the statistical anomaly detection results. These anomaly analysis results include statistical anomaly detection results, temporal abrupt change detection results, and multi-component logic verification results. In practice, the executing entity can perform a Grubbs test on each chemical peak concentration in the aforementioned chemical peak concentration set to generate a Grubbs test value for each chemical peak concentration, which serves as the statistical anomaly detection result.
[0074] Step 2: Perform time-series abrupt change detection on the concentration of each chemical peak in the above chemical peak concentration set to obtain the time-series abrupt change detection results.
[0075] In practice, the aforementioned executing entity can perform temporal abrupt change detection on the concentration of each chemical peak in the aforementioned chemical peak concentration set through the following sub-steps to obtain the temporal abrupt change detection results. These results may include: cumulative offset and T-statistic.
[0076] Sub-step one: Based on the Cumulative Sum Control Chart (CUSUM), perform cumulative offset detection on the concentration of each chemical peak in the above chemical peak concentration set to generate the cumulative offset of each chemical peak concentration.
[0077] Sub-step two involves performing a mean difference significance test on the concentration of each chemical peak in the above chemical peak concentration set based on the sliding T-test, in order to generate a T-statistic for each chemical peak concentration.
[0078] Step 3: Perform multi-component logic verification on the concentration of each chemical peak in the above chemical peak concentration set to obtain the multi-component logic verification results. These results include homologous ratios and physical constraint detection results.
[0079] In practice, the aforementioned executing entity can obtain multi-component logic verification results through the following sub-steps:
[0080] Sub-step one involves performing a ratio test on substances within the same chemical category (based on molecular structural similarity, such as benzene series compounds: benzene, toluene, ethylbenzene, xylene) within each chemical peak concentration in the aforementioned concentration set, to generate the homologous ratio for each chemical peak concentration. In practice, the ratio test for substances within the same chemical category is based on environmental chemical mechanisms and industry experience, and is performed for targeted verification of specific related substance groups. For example, the industry reference range for toluene / benzene is 0.8 ± 0.2 (typical value in the petrochemical industry).
[0081] Sub-step two involves performing physical constraint verification on the concentration of each chemical peak in the aforementioned chemical peak concentration set to generate a physical constraint detection result for each chemical peak concentration. In practice, the aforementioned physical constraint verification may include: comparing the total mass of the chemicals before the chemical reaction with the total mass of the products after the chemical reaction based on the law of conservation of mass, and comparing the concentration ratio of gases and liquids based on Henry's Law. The aforementioned physical constraint detection results may include normal or abnormal results.
[0082] Step 106: Based on the preset anomaly correction model, perform equipment anomaly correction on the anomaly analysis results to obtain the anomaly correction results.
[0083] In some embodiments, the aforementioned execution entity may perform equipment anomaly correction on the aforementioned anomaly analysis results based on a preset anomaly correction model to obtain anomaly correction results.
[0084] In practice, the aforementioned implementing entity can perform equipment anomaly correction based on a preset anomaly correction model through the following steps to obtain the anomaly correction result:
[0085] Step 1: Based on a pre-defined anomaly correction model, the anomaly analysis results, acquired environmental parameters, and device status are encoded to generate corresponding encoded results. The pre-defined anomaly correction model can be a model used to correct the anomaly analysis results, and may include an input layer, a feature fusion layer, a temporal feature extraction layer, a bidirectional LSTM network, an attention mechanism layer, a decision output layer, and a post-processing layer. The input layer encodes the anomaly analysis results, acquired environmental parameters, and device status to generate corresponding encoded results. The feature fusion layer concatenates the encoded results to obtain a fused feature tensor. The temporal feature extraction layer reshapes the fused feature tensor into a temporal sequence to generate a temporal sequence. The bidirectional LSTM network (BiLSTM) performs forward and backward LSTM processing on the temporal sequence to obtain a context-aware feature representation for each time step, serving as the temporal-aware feature representation. The aforementioned attention mechanism layer can be used to weight the temporal-aware feature representation based on feature importance and output weighted features. The aforementioned decision output layer can be used to perform decision fusion on the weighted features and obtain the probability distribution of the correction operation. The aforementioned post-processing layer can be used to generate an anomaly correction result set based on the probability distribution of the correction operation. The acquired environmental parameters may include temperature, humidity, wind direction, and air pressure. The aforementioned equipment status may include operating status and equipment model.
[0086] In practice, the aforementioned executing entity can perform encoding through the following sub-steps to generate corresponding encoding results. These encoding results can include statistical anomaly characteristics, temporal mutation characteristics, multi-component logical characteristics, environmental parameter characteristics, and equipment status characteristics.
[0087] Sub-step one involves normalizing the above statistical anomaly detection result set to obtain statistical anomaly features. In practice, the executing entity can perform Min-Max normalization on the above statistical anomaly detection results to obtain statistical anomaly features.
[0088] Sub-step two involves performing differential standardization on the aforementioned time-series abrupt change detection results to obtain the time-series abrupt change features. In practice, the executing entity can first perform logarithmic compression on the cumulative offset in the aforementioned time-series abrupt change detection results to generate a logarithmically compressed result. Then, the executing entity can perform a hyperbolic tangent transform on the T statistic in the aforementioned time-series abrupt change detection results using the hyperbolic tangent function to obtain the transformed T statistic value. The aforementioned logarithmic compression result can be obtained using the following formula:
[0089] Y = ln(|X|+1)·sign(X).
[0090] Where X represents the cumulative offset. Y represents the logarithmic compression result. sign() returns 1 when the input parameter is greater than 0, -1 when it is less than 0, and 0 when it is equal to 0.
[0091] Sub-step three involves performing a logarithmic transformation on the ratios of the above multi-component logic verification results to obtain the multi-component logic features. In practice, the executing entity can perform a logarithmic transformation on the homologous ratios in the above multi-component logic verification results. The physical constraint detection results in the above multi-component logic verification results are represented as Boolean codes (normal: 0, abnormal: 1).
[0092] Sub-step four involves encoding the acquired environmental parameters to obtain environmental parameter features. These features include temperature encoding, humidity encoding, wind direction encoding, and air pressure encoding.
[0093] In practice, the aforementioned executing entity can first encode the temperature among the environmental parameters using a radial basis function (RBF) kernel to generate a temperature code. Then, it can perform Min-Max normalization on the humidity among the environmental parameters to generate a humidity code. Next, it can take an angle vector (e.g., [cosθ, sinθ]) on the wind direction θ among the environmental parameters to generate a wind direction code. Finally, it can perform differential encoding on the air pressure among the environmental parameters to generate an air pressure code.
[0094]
[0095] Where P′ represents the air pressure code. P represents air pressure.
[0096] Finally, the temperature code, humidity code, wind direction code, and air pressure code mentioned above were determined as environmental parameter characteristics.
[0097] Sub-step five involves encoding the acquired device status to obtain device status characteristics. These characteristics may include the aforementioned operating status code and the aforementioned device model code.
[0098] In practice, the aforementioned executing entity can first perform one-hot encoding on the operating states of the aforementioned equipment to generate operating state codes (e.g., [1,0,0]: normal, [0,1,0]: calibration, [0,0,1]: fault). Then, it can perform model embedding on the equipment model in the aforementioned equipment states to generate equipment model codes. Model embedding is a technique for mapping high-dimensional discrete data (such as text, images, and category labels) to a low-dimensional continuous vector space, capturing semantic or feature information between data through vectors.
[0099] Step two involves feature concatenation of the above encoding results to obtain a fused feature tensor. In practice, the aforementioned executing entity can use statistical anomaly features, temporal mutation features, multi-component logical features, environmental parameter features, and equipment status features from the above encoding results to determine the fused feature tensor through feature concatenation.
[0100] Step 3: Reshape the aforementioned fused feature tensor into a temporal sequence to generate a temporal sequence. In practice, the executing entity can adjust the tensor dimension order of the fused feature tensor from [batch size (number of samples), time step (sequence length), feature dimension] to [time step (sequence length), batch size (number of samples), feature dimension]. Then, it is segmented into multiple independent segments with the structure [batch size (number of samples), feature dimension] according to the time step dimension to obtain the temporal sequence.
[0101] Step four: Based on a bidirectional LSTM network, the aforementioned time-series sequence is processed using forward LSTM and backward LSTM to obtain a context-aware feature representation for each time step, which serves as the time-series-aware feature representation. In practice, the execution entity can input the aforementioned time-series sequence into a Long Short-Term Memory (LSTM) network from start to finish for forward LSTM processing to obtain forward-aware features, and then input the aforementioned time-series sequence into a Long Short-Term Memory (LSTM) network from end to start for backward LSTM processing to obtain backward-aware features. Finally, the aforementioned forward-aware features and backward-aware features are concatenated to generate a context-aware feature representation for each time step, which serves as the time-series-aware feature representation. The aforementioned bidirectional Long Short-Term Memory (BiLSTM) network is a sequence coding architecture that combines bidirectional processing and long short-term memory units.
[0102] Step 5: Weight the time-series-aware feature representations based on feature importance and output the weighted features. In practice, firstly, the executing agent can determine the weight of each dimension of the context-aware feature representation at each time step based on the Multi-Head Latent Attention (MLA) mechanism. Secondly, the weights of each dimension of the context-aware feature representation at each time step are standardized using a normalized exponential function (softmax) to generate normalized weights for each dimension of the context-aware feature representation at each time step. Finally, each dimension feature is multiplied by its corresponding normalized weight, and then summed to obtain the weighted fusion of the context-aware feature representations at each time step, which serves as the weighted feature.
[0103] Step six involves performing decision fusion on the aforementioned weighted features to obtain the probability distribution of the correction operation. In practice, the executing entity can fuse the weighted features into a decision vector based on a fully connected layer, and generate the probability distribution of the correction operation using a normalized exponential function (softmax). The probability distribution of the correction operation can be P = [P0, P1, P2, P3].
[0104] Step 7: Generate an abnormal correction result set based on the probability distribution of the above correction operations. The abnormal correction results include:
[0105] Sub-step one, in response to P1>0.7, indicates sampling interference, interrupts the current data stream, and initiates the cleaning process: nitrogen flushing of the injection line (60 seconds), and vacuum reset of the detection chamber.
[0106] Sub-step two: A response of P1 < 0.7 and P2 > 0.8 indicates a device malfunction. Corrective actions are performed based on the malfunction type. For example: Malfunction type: Sensor drift; Corrective action: Switch to backup channel. Malfunction type: Calibration failure; Corrective action: Initiate five-point calibration. Malfunction type: Signal loss; Corrective action: Switch communication protocol.
[0107] Sub-step three, responding to P1 < 0.7, P2 < 0.8, and P3 > 0.6, indicates a system drift, and temperature compensation is applied to the chemical peak concentration:
[0108] C comp =C×[1-α(ΔT)+β(ΔT)] 2 ].
[0109] Where C represents the concentration of the analyte. compThis indicates the concentration of the analyte after temperature compensation. α represents the linear compensation coefficient (e.g., benzene: 0.015 / ℃). β represents the nonlinear compensation coefficient (e.g., benzene: 0.0002 / ℃²). ΔT represents the difference between the actual operating ambient temperature and the reference temperature during instrument calibration, in ℃ (degrees Celsius).
[0110] Sub-step four: In response to max(P0, P1, P2, P3) < 0.6, indicating a true concentration change, a pollution event warning is generated. This pollution event warning can be a report reflecting a true concentration change, and can be a text message in JSON format. The pollution event warning includes a true concentration change identifier, the substance name, and its concentration.
[0111] Steps one through seven and their related contents, as an inventive point of this disclosure, solve the technical problem that "in the process of volatile organic compound (VOCs) detection, when data distortion, equipment failure, and environmental interference coexist, the existing technology has difficulty distinguishing between true concentration anomalies and interference factors such as equipment failure, sampling interference, or system drift, resulting in the inability to perform relevant corrections based on the anomaly analysis results." The factors that prevent relevant corrections based on the anomaly analysis results are often as follows: data distortion, equipment failure, and environmental interference occur during the VOCs detection process. Solving these factors can avoid the inability to perform equipment cleaning, equipment failure correction, temperature compensation, and pollution event early warning based on the anomaly analysis results. To achieve this effect, firstly, based on a preset anomaly correction model, the above-mentioned anomaly analysis results, the acquired environmental parameters, and the equipment status are encoded, and the corresponding encoded results are output. This allows multi-source heterogeneous data to be converted into consistent, computer-readable data. Secondly, the above-mentioned encoded results are feature-concatenated to obtain a fused feature tensor. This allows multi-source data to be fused for multi-source analysis by the above-mentioned preset anomaly correction model. Third, the aforementioned fusion feature tensor is reshaped into a temporal sequence to generate a temporal sequence. This generates multi-source fusion data for time-series analysis. Fourth, the temporal sequence is processed using a bidirectional LSTM network with forward and backward LSTM processing, outputting a context-aware feature representation for each time step as a temporal-aware feature representation. This extracts high-dimensional temporal-aware features from the multi-source fusion data. Fifth, the temporal-aware feature representation is weighted by feature importance, and weighted features are output. This allows for selective information focus through dynamic weight allocation. Sixth, the weighted features are fused for decision processing, and the probability distribution of correction operations is output. This achieves a quantitative expression of multi-dimensional anomaly decisions. Seventh, anomaly correction results are generated based on the probability distribution of the correction operations. This transforms probabilistic decisions into executable instructions. Ultimately, this achieves the goal of avoiding the inability to perform equipment cleaning, equipment fault correction, temperature compensation, and pollution event early warning based on anomaly analysis results.
[0112] In some optional implementations of the embodiments, the aforementioned execution entity can trace the location of the above-mentioned anomaly correction results based on the acquired meteorological data and geographic information data, obtain the location of volatile organic compound emission hotspots, and control a preset robot to the location of the above-mentioned volatile organic compound emission hotspots to perform data sampling, obtain data sampling results, and use the data sampling results to perform concentration detection and anomaly correction, and generate a pollution event report.
[0113] In practice, the aforementioned implementing entities can use the following steps to trace the location of volatile organic compound (VOC) emission hotspots based on the acquired meteorological and geographic information data. They can then control a pre-set robot to these hotspots to collect data, obtain the sampling results, and use this data for concentration detection and anomaly correction to generate a pollution event report.
[0114] Step 1: In response to the above-mentioned anomaly correction results, a pollution event warning is generated. The chemical peak feature data in the chemical peak feature dataset corresponding to the substance name in the pollution event warning is determined. The target type chemical peak of the above-mentioned chemical peak feature data is time-aligned with the acquired meteorological data to generate time-aligned concentration time series and time-aligned meteorological data.
[0115] In practice, the aforementioned implementing entities can use interpolation to time-align the acquired meteorological data with the target type chemical peaks in the aforementioned chemical peak characteristic data, thereby generating time-aligned concentration time series and time-aligned meteorological data. The acquired meteorological data can be high-resolution gridded data (e.g., 1km × 1km, 1 hour) from on-site meteorological sensors, local weather stations, and WRF (Weather Research and Forecasting) model outputs. The acquired meteorological data can include wind direction, wind speed, temperature, air pressure, humidity, atmospheric stability (Pasquale stability class, or directly using the Moning-Obukhov length output from WRF), and planetary boundary layer height.
[0116] Step two: Based on the aforementioned chemical peak characteristic data, generate grid data. In practice, the executing entity can use the coordinates corresponding to the aforementioned chemical peak characteristic data as the center of the grid to construct the grid data. The resolution of the grid data can be 1km × 1km or 100m × 100m. The aforementioned coordinates can be the starting point t of the chemical peak in the peak shape parameter of the aforementioned chemical peak characteristic data. L The coordinates of the corresponding data collection.
[0117] Step three involves mapping the time-aligned concentration time series, the aforementioned time-aligned meteorological data, and the acquired geographic information data to the aforementioned grid data to generate gridded fused data. In practice, firstly, the executing entity can project the aforementioned time-aligned meteorological data and the acquired geographic information data to the same coordinate system as the aforementioned grid through geographic coordinate transformation, obtaining coordinate-transformed meteorological data and coordinate-transformed geographic information data. Then, the coordinate-transformed meteorological data is spatially interpolated using the inverse distance weighting method to generate time-space aligned meteorological data. Finally, spatial overlay analysis is used to fuse the aforementioned time-aligned concentration time series, the aforementioned time-space aligned meteorological data, and the aforementioned coordinate-transformed geographic information data into gridded fused data. The aforementioned spatial overlay analysis can be an overlay analysis of raster data.
[0118] Step four: Based on the forward diffusion model, in response to the release of pollutants of unit intensity from each grid cell in the above-mentioned gridded fused data, determine the unit concentration generated at the center of the grid to generate a concentration contribution matrix. The aforementioned forward diffusion model is a model used to predict the distribution of atmospheric pollutant concentrations. For example, the aforementioned forward diffusion model can be AERMOD or CALPUFF. The aforementioned unit intensity can be 1 g / s. The aforementioned concentration contribution matrix can be composed of rows m: pollution source grid cells (total number of grids), columns n: receptor sites (sites in the above-mentioned anomaly correction results), and elements d... mn The matrix consisting of the concentration of pollutants at point n when grid m releases pollutants of unit intensity.
[0119] Step 5: Based on the concentration contribution matrix described above, construct the inversion equation. This inversion equation includes the concentration contribution matrix, the source intensity vector, and the noise vector. The inversion equation can be expressed as: Observed concentration vector = Concentration contribution matrix × Source intensity vector + Noise vector. The observed concentration vector can be the concentration from the anomaly correction result. The noise vector is determined by the instrument accuracy (e.g., ±0.1 ppm).
[0120] Step Six: Based on the regression method, determine the source strength vector in the above inversion equation to generate an emission heat map. In practice, the execution entity can determine the source strength vector in the above inversion equation based on the regression method. Then, the source strength vector is subjected to zeroing of negative values and Min-Max normalization to obtain the processed source strength vector. Next, the source strength values in the processed source strength vector are bound to the corresponding grid center coordinates as thermal values to obtain the emission heat map. The regression method can be the least squares method or Tikhonov regularization.
[0121] Step seven involves performing spatial clustering analysis on the aforementioned emission hotspot map to generate the grid range of emission hotspots with the highest contribution. In practice, the implementing entity can use DBSCAN clustering or K-Means clustering to perform spatial clustering analysis on the emission hotspot map to generate the grid range of emission hotspots with the highest contribution.
[0122] Step eight involves overlaying the aforementioned gridded fusion data with the grid range of the emission hotspot with the highest contribution to generate the location of volatile organic compound (VOC) emission hotspots. In practice, the implementing entity can use raster algebra operations to overlay the aforementioned gridded fusion data with the grid range of the emission hotspot with the highest contribution to generate the location of VOC emission hotspots.
[0123] Step 9: Control the preset robot to locate the volatile organic compound emission hotspot to collect data, obtain the data sampling results, use the data sampling results to detect concentration and correct anomalies, and generate a pollution event report.
[0124] In practice, the aforementioned implementing entity can control a pre-set robot to locate a volatile organic compound (VOC) emission hotspot for data sampling, obtain data sampling results, and perform concentration detection and anomaly correction on the data sampling results based on steps 101-106 above, generating a pollution event report. The pre-set robot can include a mobile platform system and a sampling system. The pre-set robot is a machine used to navigate to the VOC emission hotspot for VOC data sampling (e.g., an unmanned vehicle with a navigation system and vacuum pump). The pollution event report can be a report reflecting actual concentration changes and can be a JSON-formatted text report. The pollution event report includes an identifier of the actual concentration change, substance name, concentration, time, and coordinates.
[0125] Steps one through nine and their related contents, as an inventive point of this disclosure, solve the technical problem of "the lack of source tracing of VOCs detection results based on the source tracing model, making it difficult to sample and analyze pollution sources and generate pollution event reports based on the pollution source analysis results." The factors that make it difficult to sample and analyze pollution sources and generate pollution event reports based on the pollution source analysis results are often as follows: the lack of source tracing of VOCs detection results based on the source tracing model. If these factors are resolved, pollution source sampling and analysis can be achieved, avoiding the inability to generate pollution event reports based on the pollution source analysis results. To achieve this effect, firstly, the acquired meteorological data is time-aligned with the concentration time series in the above-mentioned anomaly correction results to generate time-aligned meteorological data and time-aligned concentration time series. This synchronizes the meteorological data with the pollutant concentration time series, eliminating time offset errors. Secondly, a grid is generated centered on the location information in the above-mentioned anomaly correction results. This constructs a spatial analysis grid centered on the monitoring points, providing a standardized spatial framework for multi-source data fusion. Third, the time-aligned meteorological data and acquired geographic information data are mapped to the aforementioned grid to generate gridded fused data. This associates meteorological elements and geographic features with grid cells, generating a gridded fused dataset containing environmental background. Fourth, based on a forward diffusion model, the unit concentration generated by the location information in the above anomaly correction results when each grid cell releases a unit intensity of pollutant in the above gridded fused data is calculated to generate a concentration contribution matrix. Thus, the impact of each grid unit emission on the monitoring point is calculated using the forward diffusion model, establishing a quantitative pollution source-receptor relationship model. Fifth, based on the above concentration contribution matrix, an inversion equation is constructed, which includes the concentration contribution matrix, source strength vector, and noise. This transforms the pollution source tracing problem into a mathematical inversion problem, providing a theoretical framework for emission intensity calculation. Sixth, based on regression methods, the source strength vector in the above inversion equation is determined to generate an emission hotspot map. This quantifies the pollution emission intensity of each grid, generating a visualized emission hotspot map. Seventh, spatial cluster analysis is performed on the above emission hotspot map to generate the range of emission hotspot grids with the highest contribution. Therefore, high-density clusters in the emission hotspot map are identified, and the core pollution source range with the greatest contribution is accurately located. Eighth, the acquired geographic information data is overlaid with the aforementioned emission hotspot grid range with the highest contribution to generate volatile organic compound (VOC) emission hotspot locations. Thus, by integrating geographic information and hotspot grids, precise emission coordinates with environmental attributes are output. Ninth, a pre-set robot is controlled to locate the VOC emission hotspot to collect data. The collected data is used to perform concentration detection and anomaly correction on the collected data, generating a pollution event report. This achieves the sampling, analysis, and correction of pollution sources, and the generation of pollution event reports.Ultimately, this enabled the sampling and analysis of pollution sources, thus avoiding the inability to generate pollution incident reports based on the analysis results.
[0126] The above-described embodiments of this disclosure have the following beneficial effects: The VOCs data intelligent quality control audit and processing method of some embodiments of this disclosure avoids the inability to correct equipment anomalies based on abnormal analysis results. Specifically, the reason for the inability to correct equipment anomalies based on abnormal analysis results is that in industrial monitoring, gas chromatographs and chromatography-mass spectrometry instruments often need to operate under continuous, long-term, and multi-species monitoring conditions, resulting in numerous cluttered peaks in the generated spectral data, accompanied by baseline drift, peak drift, and excessive noise. This leads to low accuracy in identifying abnormal data and low efficiency in data auditing. Based on this, the VOCs data intelligent quality control audit and processing method of some embodiments of this disclosure first preprocesses the acquired volatile organic compound spectral data to obtain preprocessed spectral data. This eliminates high-frequency noise and instrument drift interference, providing high-quality input for subsequent analysis. Second, peak analysis processing is performed on the preprocessed spectral data to obtain a chemical peak feature dataset. This accurately identifies and segments the chemical peak features (position, area, and type) in the spectrum, establishing a data foundation for qualitative and quantitative analysis of substances. Then, substance matching is performed on the aforementioned chemical peak feature dataset to obtain substance matching results. The chemical peak features are then compared with a standard mass spectrometry library to determine the target substance type, solving the problem of identifying "what kind of pollutant it is." Next, based on the substance matching results, the concentration of each chemical peak feature data in the aforementioned chemical peak feature dataset is determined using a pre-calibrated calibration curve, obtaining the chemical peak concentration set. This establishes a quantitative relationship regarding "how much pollutant is present." Next, based on a multi-level anomaly data analysis algorithm, anomaly analysis is performed on the aforementioned chemical peak concentration set to obtain anomaly analysis results. These results include statistical anomaly detection results, temporal mutation detection results, and multi-component logical verification results. Thus, using statistical, temporal, and logical verification three-level detection, equipment malfunctions, environmental interference, and actual pollution events in the concentration data are identified. Then, based on a preset anomaly correction model, equipment anomaly correction is performed on the aforementioned anomaly analysis results to obtain anomaly correction results. This automatically triggers calibration, retesting, or compensation operations for different anomaly types, ensuring data accuracy and reliability. This implementation method avoids the inability to perform equipment anomaly correction based on anomaly analysis results.
[0127] Continue to refer to Figure 2 As a response to the above Figure 1 The implementation of the method shown in this disclosure provides some embodiments of a VOCs data intelligent quality control audit and processing system. These system embodiments are similar to... Figure 1Corresponding to the method embodiments shown, the system can be specifically applied to various electronic devices.
[0128] like Figure 2 As shown, the VOCs data intelligent quality control audit and processing system 200 in some embodiments includes: a preprocessing unit 201, a peak analysis unit 202, a substance matching unit 203, a determination unit 204, an anomaly analysis unit 205, and an anomaly correction unit 206. The preprocessing unit 201 is configured to preprocess the acquired volatile organic compound (VOC) spectral data to obtain preprocessed spectral data; the peak analysis unit 202 is configured to perform peak analysis on the preprocessed spectral data to obtain a chemical peak feature dataset; the substance matching unit 203 is configured to perform substance matching on the chemical peak feature dataset to obtain substance matching results; the determination unit 204 is configured to determine the concentration of each chemical peak feature data in the chemical peak feature dataset based on the substance matching results and using a pre-calibrated calibration curve to obtain chemical peak concentrations, which serve as a chemical peak concentration set; the anomaly analysis unit 205 is configured to perform anomaly analysis on the chemical peak concentration set based on a multi-level anomaly data analysis algorithm to obtain anomaly analysis results, wherein the anomaly analysis results include statistical anomaly detection results, temporal mutation detection results, and multi-component logic verification results; and the anomaly correction unit 206 is configured to perform equipment anomaly correction on the anomaly analysis results based on a preset anomaly correction model to obtain anomaly correction results.
[0129] It is understandable that the units described in system 200 are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to system 200 and the units contained therein, and will not be repeated here.
[0130] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0131] like Figure 3As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0132] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0133] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0134] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0135] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0136] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: preprocess the acquired volatile organic compound spectral data to obtain preprocessed spectral data; perform peak analysis processing on the preprocessed spectral data to obtain a chemical peak feature dataset; perform substance matching on the chemical peak feature dataset to obtain substance matching results; based on the substance matching results, determine the concentration of each chemical peak feature data in the chemical peak feature dataset using a pre-calibrated calibration curve to obtain chemical peak concentrations, as a chemical peak concentration set; perform anomaly analysis on the chemical peak concentration set based on a multi-level anomaly data analysis algorithm to obtain anomaly analysis results, wherein the anomaly analysis results include statistical anomaly detection results, temporal mutation detection results, and multi-component logic verification results; and perform device anomaly correction on the anomaly analysis results based on a preset anomaly correction model to obtain anomaly correction results.
[0137] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0139] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a preprocessing unit, a peak analysis unit, a substance matching unit, a determination unit, an anomaly analysis unit, and an anomaly correction unit. The names of these units do not necessarily limit the specific unit; for example, a preprocessing unit may also be described as "a unit that performs data preprocessing on acquired volatile organic compound spectral data to obtain preprocessed spectral data."
[0140] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0141] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A method for intelligent quality control auditing and processing of VOCs data, comprising: The acquired volatile organic compound spectral data were preprocessed to obtain preprocessed spectral data. Peak analysis is performed on the preprocessed spectral data to obtain a chemical peak feature dataset; Substance matching is performed on the chemical peak feature dataset to obtain substance matching results; Based on the substance matching results, the concentration of each chemical peak feature data in the chemical peak feature dataset is determined using a pre-calibrated calibration curve to obtain the chemical peak concentration, which is then used as the chemical peak concentration set. Based on a multi-level anomaly data analysis algorithm, anomaly analysis is performed on the chemical peak concentration set to obtain anomaly analysis results, wherein the anomaly analysis results include statistical anomaly detection results, temporal mutation detection results, and multi-component logic verification results; Based on the preset anomaly correction model, the anomaly analysis results are used to correct equipment anomalies, and the anomaly correction results are obtained.
2. The method according to claim 1, wherein, The process of preprocessing the acquired volatile organic compound spectral data to obtain preprocessed spectral data includes: High-frequency noise was removed from the volatile organic compound spectral data to obtain a preliminarily denoised spectral signal; Wavelet reconstruction is performed on the initially denoised spectral signal to obtain reconstructed spectral data, which is used as the denoised spectral data. Baseline calibration is performed on the denoised spectral data to obtain preprocessed spectral data.
3. The method according to claim 1, wherein, The preprocessed spectral data is subjected to peak analysis to obtain a chemical peak feature dataset, including: Chemical peak type identification is performed on the preprocessed spectral data to obtain multiple target type chemical peaks, which are used as a target type chemical peak set; Peak boundary localization is performed on each target type chemical peak in the target type chemical peak set to generate position information for each target type chemical peak in the target type chemical peak set, thus obtaining a position information set; Based on the location information set, peak shape parameters are extracted for each target type chemical peak in the target type chemical peak set to obtain a peak shape parameter set. The target type chemical peak set and the peak shape parameter set are used as a chemical peak feature dataset.
4. The method according to claim 1, wherein, The process of performing substance matching on the chemical peak feature dataset to obtain substance matching results includes: The chemical peak feature dataset is input into the time series prediction model to obtain the predicted retention time and confidence level; Based on the sequential decision mathematical model, the predicted retention time, and the confidence level, chemical peak-substance matching is performed on the chemical peak feature dataset to obtain the optimal substance matching relationship. Based on the mass spectrometry library, the optimal substance matching relationship is queried to obtain the substance matching result.
5. The method according to claim 1, wherein, The multi-level anomaly data analysis algorithm performs anomaly analysis on the concentration of each chemical peak in the chemical peak concentration set, and obtains the anomaly analysis results, including: Statistical anomaly detection is performed on the concentration of each chemical peak in the chemical peak concentration set to obtain the statistical anomaly detection results; A time-series abrupt change detection is performed on the concentration of each chemical peak in the chemical peak concentration set to obtain the time-series abrupt change detection results. Perform multi-component logic verification on the concentration of each chemical peak in the chemical peak concentration set to obtain the multi-component logic verification result; The results of statistical anomaly detection, temporal mutation detection, and multi-component logic verification are identified as anomaly analysis results.
6. A VOCs data intelligent quality control audit and processing system, comprising: The preprocessing unit is configured to preprocess the acquired volatile organic compound spectral data to obtain preprocessed spectral data. The peak analysis unit is configured to perform peak analysis processing on the preprocessed spectral data to obtain a chemical peak feature dataset. The substance matching unit is configured to perform substance matching on the chemical peak feature dataset to obtain substance matching results; The determining unit is configured to determine the concentration of each chemical peak feature data in the chemical peak feature dataset based on the substance matching results using a pre-calibrated calibration curve, thereby obtaining the chemical peak concentration as a chemical peak concentration set; An anomaly analysis unit is configured to perform anomaly analysis on the chemical peak concentration set based on a multi-level anomaly data analysis algorithm to obtain anomaly analysis results, wherein the anomaly analysis results include statistical anomaly detection results, time-series abrupt change detection results, and multi-component logic verification results; An anomaly correction unit is configured to perform equipment anomaly correction based on a preset anomaly correction model on the anomaly analysis results, and obtain anomaly correction results.
7. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 5.
8. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 5.