A method and system for predicting chromatographic peak retention time of volatile organic compounds

By combining a dual-modal deep learning model with gas chromatography and historical data, the problem of insufficient information fusion in VOCs monitoring is solved, high-precision retention time prediction and automated analysis are achieved, adapting to complex monitoring conditions, reducing manual analysis, and supporting research on ozone formation mechanisms.

CN119673311BActive Publication Date: 2025-09-26HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411860637.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-09-26
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

Existing VOCs monitoring methods fail to effectively integrate multiple information for monitoring, resulting in poor monitoring results. In particular, in complex, continuous, and long-term monitoring, there are problems such as noise, baseline fluctuation, retention time drift, and difficulty in identifying weak peaks.

Method used

A dual-modal deep learning model is used to combine gas chromatography data and historical retention time data. The retention time of volatile organic compounds is predicted through image processing layer and sequence processing layer, and correction is performed using the derivative ascent algorithm to achieve multi-source information fusion and high-precision prediction.

Benefits of technology

It significantly improves the accuracy and robustness of retention time prediction, can identify weak chromatographic peaks, reduce the burden of manual analysis, improve data consistency and monitoring efficiency, adapt to the data characteristics of different monitoring sites, and support long-term VOCs pollution mechanism research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119673311B_ABST
    Figure CN119673311B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of organic compound monitoring, and specifically discloses a method and system for predicting the retention time of chromatographic peaks of volatile organic compounds. The method includes: obtaining gas chromatography data and historical data of organic compounds at the current moment; inputting the gas chromatography data and historical data into a prediction model to predict the retention time of each chromatographic peak of the organic compound at the current moment; the prediction model includes: a first and a second model; the first model is used to obtain a first predicted value of the retention time based on the gas chromatography data; the second model is used to obtain a second predicted value based on the historical data and the first predicted value, the second predicted value being the retention time predicted by the prediction model, or the second model is used to obtain a second predicted value of the retention time based on the historical data; the retention time predicted by the prediction model is obtained by fusing the two predicted values. Through the present application, the prediction accuracy of the retention time of the chromatographic peaks of organic compounds is significantly improved by making full use of multi-source information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of organic compound monitoring, and more specifically, to a method and system for predicting the retention time of chromatographic peaks of volatile organic compounds. Background Art

[0002] Volatile organic compounds (VOCs) are important precursors of ozone and some secondary organic aerosols. Some of these compounds are highly toxic to the environment and ecosystems, making VOCs a key focus of atmospheric environmental monitoring. Accurate and continuous long-term monitoring of VOCs is crucial for research on ozone depletion, climate change, air quality monitoring, identification of air pollution sources, and public health early warning systems. Gas chromatography (GC) and gas chromatography-mass spectrometry (GC-MS) are commonly used instruments for monitoring atmospheric VOCs. Monitored substances are characterized by chromatographic peaks in a GC spectrum, where the retention time (RT) indicates the substance type and the peak area reflects the concentration. Therefore, chromatographic peak location (i.e., determining the retention time) is a key step in GC spectrum analysis.

[0003] However, accurate manual or automated analysis of complex, continuous, long-term, industrial-scale VOCs monitoring data faces significant challenges: (i) The complexity of environmental samples leads to the mixing of non-target and target substances, resulting in interfering or overlapping peaks. (ii) Due to long-term continuous operation, noise, baseline fluctuation or drift, and significant retention time shifts may occur. (iii) Furthermore, large differences in substance concentrations can result in extremely weak target peaks that are difficult to identify by visual inspection. The vast majority of these weak peaks are attributed to olefin compounds such as trans-2-butene, 1-butene, isobutene, and cis-2-butene. These compounds have exceptionally high ozone formation potentials (OFPs), indicating their significant contribution to ozone formation. Inaccurate analysis of these key substances can adversely affect studies of ozone formation mechanisms and the causes of photochemical pollution. (iv) Due to the high chromatographic complexity caused by these factors, more than 6,000 monitoring stations in China rely primarily on manual analysis, resulting in significant human resource costs. Furthermore, manual analysis relies primarily on personal experience, which can lead to differences in analytical results between sites due to personnel differences, thus affecting data consistency.

[0004] Therefore, the study of accurate and automated identification of chromatograms has attracted widespread attention. Existing studies have shown that machine learning, deep learning, or integrated models can be used to accurately predict the retention time of chromatographic peaks of monitored substances. However, the vast majority of studies treat each analysis as an isolated random event and do not consider the temporal connection between samples. In addition, although existing general chromatography analysis toolkits have demonstrated effective analytical capabilities in most laboratory analysis environments, these models receive information in some isolated form, such as processing only image or string information, or only analyzing information related to molecular structure. This shows that the model lacks the ability to use multimodal information for comprehensive analysis and cannot adapt to application scenarios that require diverse information input. Summary of the Invention

[0005] In response to the defects of the existing technology, the purpose of this application is to provide a method and system for predicting the retention time of volatile organic compound chromatographic peaks, aiming to solve the problems that the existing VOCs monitoring methods are single, fail to effectively integrate multiple information for monitoring, and have poor monitoring effects.

[0006] To achieve the above objectives, in a first aspect, the present application provides a method for predicting the retention time of chromatographic peaks of volatile organic compounds, comprising:

[0007] Obtaining gas chromatography data of organic compounds at the current moment and historical data of organic compounds; the historical data includes: retention time data of each substance chromatographic peak in the gas chromatography measured in the past time period;

[0008] The gas chromatography data and historical data are input into a prediction model to predict the retention time of each chromatographic peak of the organic compound substance at the current moment; the prediction model includes: a first model and a second model; the first model is used to obtain a first predicted value of the retention time based on the gas chromatography data; the second model is used to obtain a second predicted value of the retention time based on the historical data and the first predicted value, and the second predicted value is the retention time predicted by the prediction model; or the second model is used to obtain a second predicted value of the retention time based on the historical data, and the retention time predicted by the prediction model is obtained by fusing the first predicted value and the second predicted value.

[0009] It should be noted that the present application comprehensively considers the gas chromatography data at the current moment and the historical data of retention time when predicting the retention time, fully refers to the current chromatographic characteristics and the historical migration characteristics of the material chromatographic peak, and can effectively integrate the current chromatographic information and the historical migration characteristics of the retention time of the material chromatographic peak, thereby achieving a more accurate retention time prediction; the above-mentioned scheme of the present application can predict the weak material chromatographic peaks that cannot be predicted by the existing retention time prediction relying on a single information. It can be seen that the retention time prediction method provided by the scheme of the present application is highly accurate and reliable.

[0010] In one possible implementation, the first model predicts a first predicted value of the retention time based on gas chromatography data, including:

[0011] The first model predicts the retention time of each substance chromatographic peak based on the gas chromatography data and the derivative data of the gas chromatography data corresponding to the spectrum to obtain a first prediction value; preferably, the derivative includes: a first-order derivative and a second-order derivative.

[0012] It is understood that in this application, when predicting retention times based on gas chromatography, not only the chromatographic data is considered, but also the derivatives at various locations in the chromatogram are calculated. Retention time prediction is performed by combining the chromatographic data and the derivatives (first-order and second-order derivatives) characteristics. Typically, the first-order derivative at the top of a chromatographic peak of a substance is 0, and the second-order derivative is less than 0.

[0013] In a possible implementation, the method further includes:

[0014] Determining the first derivative of the retention time of the substance chromatographic peak predicted by the prediction model on the spectrum corresponding to the gas chromatography data;

[0015] The retention time predicted by the prediction model is corrected according to the first-order derivative until the corresponding first-order derivative is 0, or until the corresponding first-order derivative is 0 and the second-order derivative is less than 0.

[0016] Furthermore, theoretically, the predicted retention time should correspond to the peak top of the substance's chromatographic peak, but there may be some deviation. Therefore, the first-order derivative and / or second-order derivative can be used to verify and adjust the predicted retention time to ensure that the final retention time corresponds to the peak top of the substance's chromatographic peak.

[0017] In a possible implementation, the first model adopts a one-dimensional convolution kernel.

[0018] Among them, since gas chromatography data is one-dimensional data, the first model uses a one-dimensional convolution kernel to adapt to the input data dimension to achieve retention time prediction based on gas chromatography data.

[0019] In one possible implementation, the prediction model is trained as a whole; or

[0020] When the second model obtains the second predicted value of the retention time based on the historical data, the first model and the second model in the prediction model are trained separately, and then the prediction model is trained; when the second model obtains the second predicted value of the retention time based on the historical data and the first predicted value, the first model in the prediction model is trained separately, and then the prediction model is trained.

[0021] It should be noted that the prediction models obtained by using different training methods have been verified to have performance that can meet the expected use.

[0022] In a possible implementation, after the prediction model is trained using the training data of the first group of organic compounds:

[0023] If the prediction model is used to predict the second group of organic compounds, when the second model obtains the second predicted value of the retention time based on the historical data, a small amount of training data of the second group of organic compounds is used to retrain the first model and the second model in the trained prediction model to update the prediction model, and then N dimension adjustment modules are added to the updated prediction model so that the adjusted prediction model can be used to predict the second group of organic compounds; the number and / or type of material spectra included in the first group of organic compounds and the second group of organic compounds are different, N is an integer greater than or equal to 0, and when the number of the first group of organic compounds and the second group of organic compounds is the same, N is 0;

[0024] If the prediction model is used to predict the second group of organic compounds, when the second model obtains the second predicted value of the retention time based on the historical data and the first predicted value, a small amount of training data of the second group of organic compounds is used to retrain the first model in the trained prediction model to update the prediction model, and then N dimension adjustment modules are added to the updated prediction model so that the adjusted prediction model can be used to predict the second group of organic compounds; the number and / or types of material spectra included in the first group of organic compounds and the second group of organic compounds are different, N is an integer greater than or equal to 0, and when the number of the first group of organic compounds and the second group of organic compounds is the same, N is 0.

[0025] It can be understood that the prediction model provided in this application takes into account the type and / or quantity of organic compound combinations at different sites, and only makes a small adaptive migration of the prediction model, that is, it can be effectively migrated to the new site based on the original prediction model, thereby improving the migration performance of the prediction model and increasing the practicality of the prediction model.

[0026] In one possible implementation, the first model is a convolutional neural network; and / or the second model is a recurrent neural network or an attention mechanism-type neural network.

[0027] In a second aspect, the present application provides a system for predicting the retention time of chromatographic peaks of organic compounds, comprising:

[0028] A data acquisition module is used to acquire gas chromatography data of organic compounds at the current moment and historical data of organic compounds; the historical data includes: retention time data of each substance chromatographic peak in the gas chromatography measured in the past time period;

[0029] A retention time prediction module is used to input the gas chromatography data and historical data into a prediction model to predict the retention time of each chromatographic peak of the organic compound substance at the current moment; the prediction model includes: a first model and a second model; the first model is used to obtain a first predicted value of the retention time based on the gas chromatography data; the second model is used to obtain a second predicted value of the retention time based on the historical data and the first predicted value, and the second predicted value is the retention time predicted by the prediction model; or the second model is used to obtain the second predicted value of the retention time based on the historical data, and the retention time predicted by the prediction model is obtained by fusing the first predicted value and the second predicted value.

[0030] In one possible implementation, the first model used in the retention time prediction module predicts the retention time of each substance chromatographic peak based on the gas chromatography data and the derivative data of the spectrum corresponding to the gas chromatography data to obtain a first prediction value; preferably, the derivative includes: a first-order derivative and a second-order derivative.

[0031] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation of the first aspect.

[0032] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.

[0033] In a fifth aspect, the present application provides a computer program product, which, when executed on a processor, enables the processor to execute the method described in the first aspect or any possible implementation of the first aspect.

[0034] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.

[0035] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the existing technologies:

[0036] This application provides a method, system, and device for predicting the retention times of volatile organic compounds. The proposed prediction model can simultaneously process real-time chromatograms and historical retention time series, fully utilizing multi-source information. It also uses the derivative uphill method to adjust the retention times predicted using multi-source information, significantly improving prediction accuracy. In experiments at a typical monitoring station, the prediction model provided by this application achieved a mean absolute error (MAE) of 0.0144 minutes in retention time localization, which is 2.76 to 38.19 times smaller than traditional machine learning or deep learning models.

[0037] This application provides a method, system, and device for predicting the retention times of volatile organic compounds (VOCs). These methods are capable of effectively adapting to various complex and abnormal chromatograms, including baseline fluctuations, baseline drift, and retention time drift, demonstrating strong robustness. This addresses the challenges posed by instrument operating state changes during long-term continuous monitoring.

[0038] This application provides a method, system, and device for predicting the retention times of volatile organic compounds. These methods leverage multi-source information and employ the derivative uphill method to adjust the retention times predicted using this information. This method enables accurate identification of weak chromatographic peaks that are typically difficult to discern visually, particularly for olefin compounds with high ozone-forming potential, such as trans-2-butene, 1-butene, isobutene, and cis-2-butene. This provides important support for in-depth research into the mechanisms of ozone formation and contributes to a better understanding and prediction of photochemical pollution.

[0039] This application provides a method, system, and device for predicting the retention time of volatile organic compounds (VOCs). These methods exhibit excellent transferability and can be fine-tuned to adapt to the data characteristics of different monitoring sites. In cross-migration validation across multiple monitoring sites, the model demonstrated strong transferability. This suggests that this method has broad industrial application potential and can be scaled up for use in different regions and at different types of monitoring sites.

[0040] This application provides a method, system, and device for predicting the retention time of volatile organic compounds (VOCs). This method enables automated, high-precision analysis of VOCs monitoring data, significantly improving monitoring efficiency and reducing the burden of manual analysis. This not only saves significant human resource costs but also improves the consistency and reliability of data analysis, providing more accurate data support for long-term VOC pollution mechanism research. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a flow chart of a method for predicting chromatographic peak retention time of volatile organic compounds provided in an embodiment of the present application;

[0042] Figure 2 This is an example diagram of environmental monitoring data provided by an embodiment of the present application;

[0043] FIG3( a ) is a schematic diagram of the architecture of the first prediction model provided in an embodiment of the present application;

[0044] FIG3( b ) is a schematic diagram of the architecture of the second prediction model provided in an embodiment of the present application;

[0045] FIG3( c ) is a schematic diagram of the architecture of the third prediction model provided in an embodiment of the present application;

[0046] FIG3( d ) is a schematic diagram of the architecture of the fourth prediction model provided in an embodiment of the present application;

[0047] Figure 4 This is a schematic diagram of the correction of retention time based on the derivative rise method provided in the embodiment of the present application;

[0048] Figure 5 Examples of the retention time values ​​of real-world samples predicted by the embodiments of the present application: (a) samples of weak peaks and overlapping peaks at site D; (b) samples of baseline drift, weak peaks, and overlapping peaks at site C; (c) samples of baseline jitter and weak peaks at site W; and (d) samples of peak drift, weak peaks, and overlapping peaks at site D.

[0049] Figure 6 This is a schematic diagram of the prediction model migration provided by the embodiment of the present application;

[0050] Figure 7 This is an architecture diagram of a volatile organic compound chromatographic peak retention time prediction system provided in an embodiment of the present application;

[0051] Figure 8 This is an architectural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0053] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.

[0054] In the specification and claims herein, the terms "first" and "second" are used to distinguish between different objects, rather than to describe a specific order of objects. For example, "first model" and "second model" are used to distinguish between different models, rather than to describe a specific order of models.

[0055] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0056] First, the technical terms involved in the embodiments of this application are introduced.

[0057] (1) Retention time

[0058] Retention time refers to the time from the start of sample injection to the appearance of the chromatographic peak. In gas chromatography analysis, retention time is an important basis for substance identification. Each substance has its own characteristic retention time under specific chromatographic conditions.

[0059] (2) First model

[0060] In this application, the first model is more vividly named "Image Processing Layers" (IPLs) based on its function. In this application, the image processing layer refers to the neural network layer used to process and analyze gas chromatograms. In this application, IPLs are primarily used to extract feature information from chromatograms.

[0061] (3) Second model

[0062] In the present embodiment, the first model is more vividly named "Sequence Processing Layers" (SPLs) based on its function. Sequence Processing Layers refer to the neural network layers used to process and analyze historical retention time series. In this application, SPLs are primarily used to capture the temporal dynamics of retention time.

[0063] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.

[0064] Figure 1 : is a flow chart of a method for predicting chromatographic peak retention time of volatile organic compounds provided in an embodiment of the present application; Figure 1 As shown, the following steps are included:

[0065] Step S101, obtaining gas chromatography data of organic compounds at the current moment and historical data of organic compounds; the historical data includes: retention time data of each substance chromatographic peak in the gas chromatography measured in the past time period;

[0066] Step S102: input the gas chromatography data and historical data into a prediction model to predict the retention time of each chromatographic peak of the organic compound at the current moment; the prediction model includes: a first model and a second model; the first model is used to obtain a first predicted value of the retention time based on the gas chromatography data; the second model is used to obtain a second predicted value of the retention time based on the historical data and the first predicted value, and the second predicted value is the retention time predicted by the prediction model; or the second model is used to obtain a second predicted value of the retention time based on the historical data, and the retention time predicted by the prediction model is obtained by fusing the first predicted value and the second predicted value.

[0067] It should be noted that at each monitoring station, the original spectra of the volatile organic compounds monitored are referenced to Figure 2 shown. Figure 2 The medium blue labels are artificially calibrated substances; among them, 1 is ethane, 2 is ethylene, 3 is methane, 4 is propylene, 5 is isobutane, 6 is n-butane, 7 is acetylene, 8 is trans-2-butene, 9 is 1-butene, 10 isobutene, 11 is cis-2-butene, 12 is cyclopentane, 13 isopentane, and 14 is n-pentane.

[0068] For further example, the above prediction model is trained as a whole. In this case, the architecture of the prediction model can be seen in Figures 3(b) and 3(c); or when the second model obtains the second predicted value of the retention time based on historical data, the first model and the second model in the prediction model are trained separately, and then the prediction model is trained, as shown in Figure 3(a); when the second model obtains the second predicted value of the retention time based on the historical data and the first predicted value, the first model in the prediction model is trained separately, and then the prediction model is trained, as shown in Figure 3(d).

[0069] 3(b) and 3(d) correspond to the first model obtaining a first predicted value of the retention time based on gas chromatography data, and the second model obtaining a second predicted value of the retention time based on historical data and the first predicted value. The second predicted value corresponds to the retention time predicted by the prediction model. 3(a) and 3(c) correspond to the first model obtaining a first predicted value of the retention time based on gas chromatography data, and the second model obtaining a second predicted value of the retention time based on historical data. The retention time predicted by the prediction model is obtained by fusing the first predicted value and the second predicted value.

[0070] Those skilled in the art will appreciate that the prediction model can be referred to as a bimodal deep learning model based on its functionality; the first module can be referred to as image processing layers (IPLs), and the second module can be referred to as sequence processing layers (SPLs). Therefore, in Figures 3(a) to 3(d), the first model and the second model are represented as IPLs and SPLs, respectively.

[0071] It should be noted that the prediction models obtained using different training methods (models corresponding to Figures 3(a) to 3(d)) have been verified to have the performance they need to meet the expected requirements. Specific performance indicators: Mean Absolute Error (MAE) for retention time positioning can be found in Table 1:

[0072] Table 1

[0073]

[0074] As shown in Table 1, it can be seen that the architectures of the four prediction models given in Figures 3(a) to 3(d) above have small MAEs for retention time prediction, which can meet expectations and have good application prospects.

[0075] For further example, since the gas chromatography data is one-dimensional data, the first model can use a one-dimensional convolution kernel to adapt to the input data dimension to achieve retention time prediction based on the gas chromatography data.

[0076] Further preferably, the first model predicts the retention time of each substance chromatographic peak based on gas chromatography data and derivative data of the corresponding gas chromatography data spectrum to obtain a first predicted value; preferably, the derivatives include: first-order derivatives and second-order derivatives. It is understood that when predicting retention times based on gas chromatography, not only the chromatographic spectrum data but also the derivatives at various locations in the spectrum are considered. Combining the characteristics of the spectrum data and derivatives (first-order derivatives and second-order derivatives) to predict retention times can improve prediction accuracy. Typically, the first-order derivative at the peak top of a substance chromatographic peak is 0, and the second-order derivative is less than 0. It can be seen that the first-order derivative and second-order derivative can reflect the characteristics of a certain substance chromatographic peak.

[0077] The following example conducts a comparative experiment on the MAE of the first model inputting different data types for retention time prediction. The results are shown in Table 2:

[0078] Table 2

[0079]

[0080] It can be seen that when the first model selects the residual network (ResidualNetwork, ResNet), and the input data of the first model includes graph data, first-order derivatives, and second-order derivatives, the MAE of the first model is the lowest.

[0081] It can be understood that, by comparing the data in Tables 1 and 2, the lowest MAE for retention time prediction using the single first model is 0.0668, while the MAE for retention time prediction using the multimodal prediction model is far less than half of 0.0668. This indirectly proves that the prediction accuracy of retention time prediction using the multimodal prediction model provided by this application is much higher than that of a single-modal prediction method.

[0082] More preferably, the above Figure 1 The provided method also includes the following steps: determining the first-order derivative of the retention time of the substance chromatographic peak predicted by the prediction model on the corresponding spectrum of the gas chromatography data; and correcting the retention time predicted by the prediction model according to the first-order derivative until the corresponding first-order derivative is 0, or until the corresponding first-order derivative is 0 and the second-order derivative is less than 0.

[0083] It should be noted that, in theory, the predicted retention time should be the peak top of the corresponding substance chromatographic peak, but there may be a certain deviation. Therefore, the first-order derivative and / or second-order derivative can be used to search, verify and adjust, see Figure 4 As shown, it is ensured that the final retention time corresponds to the peak top of the substance chromatographic peak.

[0084] Specifically, as an example, the above-mentioned first model can be a convolutional neural network; and / or the above-mentioned second model can be a recurrent neural network or an attention mechanism-type neural network.

[0085] In one embodiment of the present application, the image processing layer adopts a convolutional network structure for one-dimensional sequences, which is used to specifically process gas chromatography spectra.

[0086] In one embodiment of the present application, the sequence processing layer adopts a recurrent neural network structure or a neural network structure based on an attention mechanism to specifically process historical retention time series.

[0087] In one embodiment of the present application, the calculation results of the sequence processing layer and the image processing layer will be aggregated and output after calculation by the fully connected layer, and the output is the final output result of the dual-modal deep learning model.

[0088] In one embodiment of the present application, the gas chromatogram is preprocessed, and the preprocessing includes smoothing and derivative generation. The smoothing process uses wavelet analysis and Savitzky-Golay (SG) smoothing algorithm to effectively reduce the noise in the spectrum. The derivative generation calculates the first and second order derivatives of the original spectrum as additional channels and the original spectrum. Figure 1 The image processing layer is input, which can enhance the characteristics of the chromatographic peaks and help the model locate the peak positions more accurately.

[0089] In one embodiment of the present application, a derivative-ascent algorithm is also used to correct the retention times predicted by the dual-modal deep learning model. This algorithm exploits the geometric characteristics of retention time positions in the chromatogram, namely, the retention time must correspond to the top of the chromatographic peak, where the first-order derivative of the peak top is zero and the second-order derivative is less than zero. In this way, the accuracy of the prediction can be further improved, especially for weak or overlapping chromatographic peaks.

[0090] In one embodiment of the present application, the historical retention time series includes data of a plurality of historical time steps. This length is optimized through experiments to provide sufficient historical information to capture temporal dynamic characteristics while avoiding the introduction of excessive redundant information that would increase computational complexity.

[0091] In a more specific embodiment, the method for predicting the retention time of volatile organic compounds provided in the present application includes the following steps: obtaining a gas chromatogram and preprocessing it to form multi-channel data features, and simultaneously obtaining a historical retention time series, and inputting these two types of data into a dual-modal deep learning model. Within the dual-modal deep learning model, the multi-channel data features are processed by image processing layers (IPLs), and the historical retention time series are processed by sequence processing layers (SPLs). The processed features are combined within the model, and the predicted retention time is output after passing through the fully connected layer. Finally, the prediction result can be further corrected by applying the derivative ascent algorithm to obtain the final retention time prediction value. This method can achieve high-precision automated analysis of VOCs, significantly improving analysis efficiency and accuracy.

[0092] Because gas chromatography data has the dual characteristics of images and time series, this application uses a dual-modal deep learning approach to process image features and sequence features separately within the model. This enables high-precision analysis of complex industrial-grade continuous monitoring data, effectively addressing the difficulties traditional methods face in processing abnormal chromatograms and weak chromatographic peaks. This method is widely applicable to VOCs analysis in fields such as environmental monitoring and industrial process control.

[0093] For example, specifically, this embodiment may include the following steps:

[0094] S1. Data acquisition and preprocessing

[0095] S11. Obtain a gas chromatogram.

[0096] S12. Obtaining a historical retention time series: For each preset moment, extract the historical retention time data of a number of time steps before the moment to form a historical retention time series.

[0097] S13. Preprocess the gas chromatogram. Preprocessing includes smoothing and derivative generation. Smoothing uses wavelet analysis or the SG smoothing algorithm to effectively reduce noise in the spectrum. Derivative generation calculates the first and second derivatives of the original spectrum. The original spectrum, first derivative, and second derivative are used as multi-channel data features.

[0098] S2. Model structure

[0099] S21, Image Processing Layers (IPLs). IPLs use a convolutional neural network structure for one-dimensional sequences to process multi-channel data features.

[0100] S22, Sequence Processing Layers (SPLs). SPLs use a GRU network structure to process historically preserved time series. Sequence processing layers use a recurrent neural network structure or an attention-based neural network structure specifically for processing historically preserved time series.

[0101] S23. Build a bimodal deep learning model, ResGRU. This model integrates IPLs and SPLs, enabling simultaneous processing of image and sequence features. The model first processes the input multi-channel data features and the historical retention time series using IPLs and SPLs, respectively. It then combines these processed features and outputs the final retention time prediction through a fully connected layer.

[0102] S3. Model training and prediction

[0103] S31. Train the ResGRU model using a training dataset. The training data includes preprocessed gas chromatograms (multi-channel data features), historical retention time series, and corresponding true retention time labels.

[0104] S32. Use the trained ResGRU model to predict retention times for new gas chromatograms (after preprocessing) and historical retention time series.

[0105] S4. Prediction result correction

[0106] S41, applying a derivative rise algorithm to correct the predicted retention time obtained in S32. The algorithm utilizes the geometric characteristics of the RT position in the chromatogram, i.e., the RT must correspond to the top of the chromatographic peak, where the first derivative of the peak top is zero and the second derivative is less than zero.

[0107] S42. Iteratively adjust the predicted point position until a peak top position that meets the conditions is found, thereby obtaining a final corrected retention time prediction value.

[0108] S5. Model evaluation and optimization

[0109] S51. Use the mean absolute error (MAE) as the evaluation indicator to calculate the error between the predicted value and the true value.

[0110] S52. Continuously optimize model performance by adjusting model structure, hyperparameters, etc. until satisfactory prediction accuracy is achieved.

[0111] Preferably, the derivative ascent algorithm described in step S41 is performed as follows: (1) Obtain the RT value initially predicted by the ResGRU model. (2) Calculate the first-order derivative of the point. (3) If the first-order derivative is positive, move the predicted point to the right; if it is negative, move it to the left. (4) Repeat steps (2) and (3) until a point is found where the first-order derivative is zero. (5) Check whether the second-order derivative of the point is less than zero. If so, confirm it as a peak; if not, continue searching.

[0112] It is understandable that the specific structural parameters and training parameters of the model can be adjusted and optimized according to the actual application scenario. Those skilled in the art can design the corresponding parameters according to actual needs.

[0113] The method of the present application can not only accurately predict the retention time of normal chromatograms, but also, through experiments, it has been found that the prediction model provided by the present application can effectively process abnormal chromatograms due to its high prediction ability. Figure 5 As shown, Figure 5 (a) is a sample of weak peaks and overlapping peaks at site D. Figure 5 For the following situations, see (b) for baseline drift, (c) for baseline fluctuation, (d) for retention time drift, and so on. Figure 5 From (a) to (d), it can be seen that the prediction model provided by this application has a strong ability to distinguish chromatograms under unstable and harsh conditions, and can identify chromatographic peaks that are difficult for the human eye to identify under distorted spectral conditions. In particular, for weak chromatographic peaks, such as olefin compounds such as trans-2-butene, 1-butene, isobutylene and cis-2-butene, the method provided by this application shows excellent recognition ability. These compounds have high ozone generation potential and are of great significance to the study of photochemical pollution. Among them, Figure 5 The substances are: (i) isobutane; (ii) n-butane; (iii) acetylene; (iv) trans-2-butene; (v) 1-butene; (vi) isobutene; (vii) cyclopentane; (viii) cis-2-butene.

[0114] Furthermore, the methods provided by the embodiments of this application have excellent transferability. Through fine-tuning, a model trained at one monitoring station can be quickly adapted to other monitoring stations, significantly reducing the workload of training individual models for each station. This opens up the possibility for large-scale industrial applications of VOCs monitoring.

[0115] Different monitoring sites may have different types and numbers of monitored organic compound combinations. For example, sites H, W, D, and C are shown in Table 3:

[0116] Table 3

[0117]

[0118] It can be seen that there are inconsistencies in the types of substances monitored between sites, and the order of peak elution of the same substance is also different. This may be due to differences in the chromatographic columns used in different sites, differences in the chromatographic method parameter settings, and differences in instrument brands.

[0119] In one possible implementation, after the prediction model is trained using the training data of the first set of organic compounds:

[0120] If the prediction model is used to predict the second group of organic compounds, when the second model obtains the second predicted value of the retention time based on the historical data, a small amount of training data of the second group of organic compounds is used to retrain the first model and the second model in the trained prediction model to update the prediction model, and then N dimension adjustment modules are added to the updated prediction model so that the adjusted prediction model can be used to predict the second group of organic compounds; the number and / or type of material spectra included in the first group of organic compounds and the second group of organic compounds are different, N is an integer greater than or equal to 0, and when the number of the first group of organic compounds and the second group of organic compounds is the same, N is 0;

[0121] If the prediction model is used to predict the second group of organic compounds, when the second model obtains the second predicted value of the retention time based on the historical data and the first predicted value, a small amount of training data of the second group of organic compounds is used to retrain the first model in the trained prediction model to update the prediction model, and then N dimension adjustment modules are added to the updated prediction model so that the adjusted prediction model can be used to predict the second group of organic compounds; the number and / or types of material spectra included in the first group of organic compounds and the second group of organic compounds are different, N is an integer greater than or equal to 0, and when the number of the first group of organic compounds and the second group of organic compounds is the same, N is 0.

[0122] It should be noted that the key means of migrating the above-mentioned prediction model between different sites (different groups of organic compounds) include: 1. Using a small amount of training data to train the first model and / or the second model to adapt it to the new types of material chromatographic peaks; 2. Using a dimensionality adjustment module to adjust the dimensions of the input and / or output of the first model and / or the second model and the data fusion module. If the original first model and / or the second model outputs the predicted retention time results of the material chromatographic peaks in N, the number of output results is N. If the number of organic compound material chromatographic peaks at the new site is M, when M is not equal to N, the output results need to be adjusted. The above-mentioned dimensionality adjustment model is used to adjust the number of output results. This can be achieved through a fully connected layer.

[0123] See also Figure 6 As shown, the prediction model corresponding to the first group of organic compounds before migration can be understood as the source domain model. When migrating to the second group of organic compounds for monitoring, the prediction model used for the second group of organic compounds can be understood as the target domain model. Taking the prediction model provided in Figure 3(b) as an example, after training the first module using a small number of gas chromatograms of the second group of organic compounds, a fully connected layer is added to the output of the first module after the prediction of the first module. Then, before the prediction results of the first module and the historical retention time are input into the second module, another fully connected layer is added. Finally, a fully connected layer is added to the output of the second module to achieve the migration of the prediction model.

[0124] It can be understood that the prediction model provided in this application takes into account the type and / or quantity of organic compound combinations at different sites, and only makes a small adaptive migration of the prediction model, that is, it can be effectively migrated to the new site based on the original prediction model, thereby improving the migration performance of the prediction model and increasing the practicality of the prediction model.

[0125] Figure 6 The source domain data dimension (N, 50, 14) indicates that the source site needs to monitor 14 types of data, and the historical data of the previous 50 hours is used to predict the historical data of the next hour. The target domain data dimension (N, 50, 13) indicates that the target site needs to monitor 13 types of data, and the historical data of the previous 50 hours is used to predict the historical data of the next hour. N represents the number of training samples input into the model during training. The fully connected layer at the output end of the first model is used to convert the 14-dimensional data it outputs into 13 dimensions. The fully connected layer before inputting into the second model is used to convert the 13-dimensional data into 14 dimensions. The fully connected layer at the output end of the second model is used to convert the 14-dimensional data into 13 dimensions.

[0126] In summary, the method proposed in this application not only improves the accuracy and efficiency of VOCs monitoring but also provides more precise data support for long-term VOC pollution mechanism research, with broad industrial application prospects. By accurately identifying and quantifying weak chromatographic peaks, especially those of olefin compounds with high ozone generation potential, this method provides important support for in-depth research on ozone formation mechanisms and the causes of photochemical pollution.

[0127] In a specific embodiment, the method provided by the present application is used to continuously monitor site H. The monitoring station is classified as a residential area monitoring station, with a monitoring frequency of once per hour and a continuous monitoring period of four months.

[0128] For each preset moment of data, this application obtains the following two types of information: (1) gas chromatogram (2) historical retention time series; preprocessing of the gas chromatogram includes smoothing and derivative generation. Smoothing uses wavelet analysis and SG smoothing algorithm to effectively reduce the noise in the spectrum. Derivative generation calculates the first-order and second-order derivatives of the original spectrum. The original spectrum, first-order derivative, and second-order derivative are used as multi-channel data features.

[0129] For the historical retention time series, this application extracts the historical retention time data of the 50 time steps before each preset moment.

[0130] This case uses a bimodal deep learning model (abbreviated as ResGRU), whose structure is shown in Figure 3(b). The model consists of the following main components: (1) Image Processing Layers (IPLs): Using the ResNet network structure, it processes multi-channel data features. (2) Sequence Processing Layers (SPLs): Using the GRU network structure, it processes historical retention time series. (3) Fully Connected Layers: This layer combines the outputs of the IPLs and SPLs to generate the final retention time prediction.

[0131] During model training, the preprocessed multi-channel data features and the historically retained time series are fed into the ResGRU model. 90% of the training data is used for training, and 10% for validation. The mean absolute error (MAE) is used as the loss function. Training is terminated when the loss function no longer changes significantly.

[0132] The trained ResGRU model is used to predict the validation set data. The prediction results are further corrected by applying the derivative ascent algorithm.

[0133] On the validation set of site H, the ResGRU model achieves a mean absolute error (MAE) of 0.0144 minutes in retention time localization, as shown in Table 4, significantly outperforming traditional methods.

[0134] Table 4

[0135]

[0136] To verify the migration capability of the model, this application applies the trained model to the monitoring station data of site W, site D and site C. The migration process is as follows Figure 6 As shown, it mainly includes the following steps:

[0137] (1) Adjust the ResGRU model structure so that the outputs of the ResNet and GRU layers are processed through dense layers to adapt to the number of material types at different sites.

[0138] (2) Fix the ResNet layer parameters and only fine-tune the GRU layer and dense layer.

[0139] (3) Use limited data (perhaps 160) of the target site for fine-tuning training.

[0140] Verification showed that the migrated models maintained good performance even when the amount of training data was reduced to 160. For most sites (H, W, and D), while the predictive performance of the migrated models fell short of the locally trained models, they still met acceptable engineering standards. In particular, for site C, all migrated models outperformed the locally trained models, demonstrating the significant potential of transfer learning for specific sites, as shown in Table 5.

[0141] Table 5

[0142]

[0143] Figure 7 This is a diagram of the system architecture for predicting volatile organic compound chromatographic peak retention time provided by an embodiment of the present application, such as Figure 7 Shown, including:

[0144] The data acquisition module 710 is used to acquire the gas chromatography data of the organic compound at the current moment and the historical data of the organic compound; the historical data includes: the retention time data of each substance chromatographic peak in the gas chromatography measured in the past time period;

[0145] The retention time prediction module 720 is used to input the gas chromatography data and historical data into a prediction model to predict the retention time of each chromatographic peak of the organic compound at the current moment; the prediction model includes: a first model and a second model; the first model is used to obtain a first prediction value of the retention time based on the gas chromatography data; the second model is used to obtain a second prediction value of the retention time based on the historical data and the first prediction value, and the second prediction value is the retention time predicted by the prediction model; or the second model is used to obtain a second prediction value of the retention time based on the historical data, and the retention time predicted by the prediction model is obtained by fusing the first prediction value and the second prediction value.

[0146] It should be understood that the above-mentioned system is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the system are similar to those described in the above-mentioned method. The working process of the system can refer to the corresponding process in the above-mentioned method and will not be repeated here.

[0147] Based on the method in the above embodiment, the embodiment of the present application provides an electronic device, such as Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the method in the above embodiment.

[0148] Further preferably, the electronic device may be an industrial-grade computer or a dedicated chromatography data analysis device.

[0149] In addition, the logic instructions in the aforementioned memory 830 can be implemented in the form of a software functional unit and, when sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.

[0150] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.

[0151] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.

[0152] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0153] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC.

[0154] The above embodiments can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions. When loaded and executed on a computer, the computer program instructions fully or partially produce the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)).

[0155] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

[0156] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A method for predicting the retention time of volatile organic compound chromatographic peaks, characterized in that: include: Obtaining gas chromatography data of organic compounds at the current moment and historical data of organic compounds; the historical data includes: retention time data of each substance chromatographic peak in the gas chromatography measured in the past time period; The gas chromatography data and historical data are input into a prediction model to predict the retention time of each chromatographic peak of the organic compound at the current moment; the prediction model includes: a first model and a second model; the first model is used to obtain a first predicted value of the retention time based on the gas chromatography data; the second model is used to obtain a second predicted value of the retention time based on the historical data and the first predicted value, the second predicted value being the retention time predicted by the prediction model; or the second model is used to obtain a second predicted value of the retention time based on the historical data, the retention time predicted by the prediction model being obtained by fusing the first predicted value and the second predicted value; The first model predicts a first predicted value of retention time based on gas chromatography data, comprising: The first model predicts the retention time of each substance chromatographic peak based on the gas chromatography data and the derivative data of the spectrum corresponding to the gas chromatography data to obtain a first predicted value; the derivative includes: a first-order derivative and a second-order derivative; Also includes: Determining the first derivative of the retention time of the substance chromatographic peak predicted by the prediction model on the spectrum corresponding to the gas chromatography data; The retention time predicted by the prediction model is corrected according to the first-order derivative until the corresponding first-order derivative is 0, or until the corresponding first-order derivative is 0 and the second-order derivative is less than 0.

2. The method according to claim 1, characterized in that The first model uses a one-dimensional convolution kernel.

3. The method according to claim 1, characterized in that The prediction model is trained as a whole; or When the second model obtains the second predicted value of the retention time based on the historical data, the first model and the second model in the prediction model are trained separately, and then the prediction model is trained; when the second model obtains the second predicted value of the retention time based on the historical data and the first predicted value, the first model in the prediction model is trained separately, and then the prediction model is trained.

4. The method according to claim 1, wherein After the prediction model is trained using the training data of the first group of organic compounds: If the prediction model is used to predict the second group of organic compounds, when the second model obtains the second predicted value of the retention time based on the historical data, a small amount of training data of the second group of organic compounds is used to retrain the first model and the second model in the trained prediction model to update the prediction model, and then N dimension adjustment modules are added to the updated prediction model so that the adjusted prediction model can be used to predict the second group of organic compounds; the number and / or type of material spectra included in the first group of organic compounds and the second group of organic compounds are different, N is an integer greater than or equal to 0, and when the number of the first group of organic compounds and the second group of organic compounds is the same, N is 0; If the prediction model is used to predict the second group of organic compounds, when the second model obtains the second predicted value of the retention time based on the historical data and the first predicted value, a small amount of training data of the second group of organic compounds is used to retrain the first model in the trained prediction model to update the prediction model, and then N dimension adjustment modules are added to the updated prediction model so that the adjusted prediction model can be used to predict the second group of organic compounds; the number and / or types of material spectra included in the first group of organic compounds and the second group of organic compounds are different, N is an integer greater than or equal to 0, and when the number of the first group of organic compounds and the second group of organic compounds is the same, N is 0.

5. The method according to any one of claims 1 to 4, characterized in that The first model is a convolutional neural network; the second model is a recurrent neural network or an attention mechanism neural network.

6. A system for predicting the retention time of volatile organic compound chromatographic peaks, characterized in that: include: A data acquisition module is used to acquire gas chromatography data of organic compounds at the current moment and historical data of organic compounds; the historical data includes: retention time data of each substance chromatographic peak in the gas chromatography measured in the past time period; A retention time prediction module, configured to input the gas chromatography data and historical data into a prediction model to predict the retention time of each chromatographic peak of the organic compound at the current moment; the prediction model comprising: a first model and a second model; the first model being configured to obtain a first predicted value of the retention time based on the gas chromatography data; the second model being configured to obtain a second predicted value of the retention time based on the historical data and the first predicted value, the second predicted value being the retention time predicted by the prediction model; or the second model being configured to obtain a second predicted value of the retention time based on the historical data, the retention time predicted by the prediction model being obtained by fusing the first predicted value and the second predicted value; The first model used in the retention time prediction module predicts the retention time of each substance chromatographic peak based on the gas chromatography data and the derivative data of the spectrum corresponding to the gas chromatography data to obtain a first predicted value; the derivative includes: a first-order derivative and a second-order derivative; Also includes: Determining the first derivative of the retention time of the substance chromatographic peak predicted by the prediction model on the spectrum corresponding to the gas chromatography data; The retention time predicted by the prediction model is corrected according to the first-order derivative until the corresponding first-order derivative is 0, or until the corresponding first-order derivative is 0 and the second-order derivative is less than 0.

7. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is configured to execute the program stored in the memory, and when the program stored in the memory is executed, the processor is configured to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Transformer oil chromatographic data forecasting method based on D-S evidence theory

    CN103592374A

  • Method for predicting retention time of compound in gas chromatographic analysis method

    CN111879871A