Organic matter non-targeted all-component quantitative analysis method based on Gaussian process regression machine learning model

By integrating and preprocessing data and optimizing model parameters, a Gaussian process regression model was developed, which solved the problem of difficulty in quantifying organic matter components and achieved high-precision and widely applicable non-targeted quantitative analysis of all components.

CN120998332APending Publication Date: 2025-11-21CHINESE RES ACAD OF ENVIRONMENTAL SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511043715.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing Gaussian process regression models have limited prediction accuracy in globally complex distribution scenarios, insufficient generalization ability in small sample scenarios, and do not fully combine domain characteristics with model structure optimization, leading to difficulties in quantifying organic matter components.

Method used

By integrating and preprocessing data, a Gaussian process regression model was constructed, the model parameters were optimized, and molecular structure descriptors and chemical property information were used in conjunction with high-resolution mass spectrometry data. Proportional coding and standardization were employed, five-fold cross-validation was performed, the kernel function and optimizer parameters were optimized, and the model was validated using additional datasets.

Benefits of technology

It significantly improves the quantitative accuracy of organic matter components and the generalization ability of the model, enhances the applicability and prediction accuracy of samples in complex environments, and reduces analysis costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998332A_ABST
    Figure CN120998332A_ABST
Patent Text Reader

Abstract

The invention provides an organic matter non-targeted all-component quantitative analysis method based on a Gaussian process regression machine learning model, and belongs to the technical field of high-resolution mass spectrum application. The technical problems to be solved are that a traditional quantitative method depends on a compound standard curve and is difficult to realize non-targeted all-component high-precision quantification of complex organic matters, and an existing machine learning model is limited in generalization ability in a small-sample and multi-parameter scene and is difficult to accurately predict an organic matter concentration function (such as instrument ionization efficiency IE). According to the method, through data integration and preprocessing (collection of SMILES, Peaq, Solvent and other data, coding, conversion and blank value filling processing), a Gaussian process regression model of molecular parameters and logarithm LogIE is constructed, model parameters (Kernel, alpha and the like) are optimized, and finally, high-precision quantitative analysis of organic matter non-targeted all components in a complex sample is achieved through the optimal model under the condition that no standard substance exists.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of high-resolution mass spectrometry application, and particularly relates to a non-targeted full-component quantitative analysis method for organic matter based on a Gaussian process regression machine learning model. BACKGROUND

[0002] As an important means for analyzing the molecular composition and changes of organic matter in environmental samples, non-targeted full-component quantification has attracted widespread attention in recent years. Compared with traditional targeted methods, non-targeted analysis does not rely on pre-set target compounds and can capture as many organic molecules in the sample as possible in the case of unknown or not completely clear components, and is particularly suitable for complex natural organic matter with diverse sources. The development of high-resolution mass spectrometry technology (such as FT-ICR MS and Orbitrap MS) makes it possible to perform high-throughput qualitative analysis of thousands of compounds. However, there are still many challenges in realizing accurate quantification of non-targeted organic matter, especially the significant difference in ionization efficiency between different compounds, and the relative peak height based on signal intensity cannot directly reflect the true concentration, thereby limiting the horizontal comparability of the data and its interpretation ability in environmental processes.

[0003] In order to overcome the bias caused by the difference in ionization efficiency, in recent years, some studies have attempted to establish a functional relationship between the structural characteristics of compounds and their response capacity in mass spectrometry, that is, by predicting the organic matter compound concentration function (such as instrument ionization efficiency, IE), the original signal is corrected. This strategy provides a new idea for non-targeted quantification under the condition of lack of standard samples. However, in the face of a large number of unknown compounds in the environment, the traditional response factor calibration method is difficult to generalize, and an automatic and generalizable model solution is urgently needed.

[0004] Under this background, machine learning has become an ideal tool for predicting the concentration function of organic matter due to its strong modeling ability and applicability to multi-dimensional complex features. By introducing molecular structure descriptors, chemical properties or high-resolution mass spectrometry derived information, a concentration function prediction model is constructed for relative quantification analysis of complex organic matter components under the condition of lack of standard samples, which provides support for in-depth understanding of the molecular composition characteristics and environmental behavior. In addition, among many machine learning methods, the Gaussian process regression model is a powerful tool for predicting the concentration function due to its flexibility and good generalization performance. Compared with deep learning and other models with strong "black box" nature, the Gaussian process regression model can more effectively control the risk of overfitting in environmental samples with limited data size or collinearity between variables, and provide quantitative evaluation of the prediction reliability, thus becoming a modeling scheme worthy of attention in this task.

[0005] CN111797574B, 2024.09.27, discloses a method for modeling the distributed output of polymers by combining just-in-time Gaussian process regression (JGPR) and an ensemble learning strategy, which solves the problem of insufficient global prediction accuracy of traditional GPR. The core is to filter samples by similarity criteria to construct local models, and to improve performance through parallel integration.

[0006] 《China Tobacco Science》, 2025.01, discloses a method for estimating soil organic matter hyperspectral based on generative adversarial networks (GAN) and machine learning (such as BPNN), which expands the data set by generating pseudo samples to improve the prediction accuracy of the model in small sample and complex terrain scenarios. The core is to use data enhancement technology to solve the problem of poor model generalization ability caused by insufficient samples.

[0007] The prior art represented by the foregoing documents has at least the following unsolved technical problems or defects: The prediction accuracy of a single Gaussian process regression model in a globally complex distribution scenario is limited. Relevant evidence is that in the prediction of polymer molecular weight distribution, the test set R 2 of the JGPR model alone is 0.972, while the R 2 of the EJGPR model after integration is improved to 0.995, indicating that a single model is difficult to capture global nonlinear relationships. The model generalization ability in small sample scenarios is insufficient. Relevant evidence is that in the estimation of soil organic matter in tobacco fields, the R 2 of the optimal model using only 94 real samples is 0.73, while after generating 150% pseudo samples by GAN, the R 2 of the model is improved to 0.80, indicating that insufficient sample size limits model performance. The domain-specific features and model structure optimization are not fully combined. Relevant evidence is that existing GPR models rely on general kernel functions and do not design adaptive kernel functions for specific domains (such as band features in spectral analysis and process parameters in polymer synthesis), resulting in insufficient modeling of the correlation between features and outputs.

[0008] In solving the above problems or overcoming the above defects, the present invention has encountered the following difficulties and obstacles: The solvent coding method has a significant impact on model performance. After multiple comparative experiments, the proportion coding scheme was determined (which improved the model R 2 by 0.15). SUMMARY

[0009] To address the aforementioned issues, this invention provides a non-targeted quantitative analysis method for all components of organic matter based on a Gaussian process regression machine learning model, which solves technical problems such as the difficulty in quantifying environmental organic matter components, the reliance of traditional methods on standard curves, and the poor generalization ability of existing machine learning models, or a combination thereof.

[0010] Terminology Explanation: Unless otherwise defined, all technical terms in this document have the same meanings as commonly understood by one of ordinary skill in the art to which the subject matter of the claims pertains. Unless otherwise stated, all patents, patent inventions, and publications cited in this document are incorporated herein by reference in their entirety. If multiple definitions exist for terms in this document, the definitions in this chapter shall prevail.

[0011] It should be understood that the above brief description and the following detailed description are exemplary and for illustrative purposes only, and do not limit the subject matter of the invention in any way. In this invention, the singular is used in conjunction with the plural unless otherwise specifically stated. It should also be noted that, unless otherwise stated, the use of “or” or “or” means “and / or”. Furthermore, the use of the term “comprising” and other forms such as “including,” “containing,” and “contains” are not limiting.

[0012] Unless specifically defined herein, the use of all commercially available products herein employs standard techniques. For example, it may be carried out using the manufacturer's instructions for use with the kit, or in accordance with methods known in the art or the description of this invention. The techniques and methods described herein can generally be implemented according to conventional methods well known in the art, based on the descriptions in the various summary and more specific documents cited and discussed in this specification.

[0013] The term “non-targeted whole component” used in this article refers to a method for non-selective, targetless, and comprehensive analysis of all detectable organic molecules in a sample.

[0014] The term "Gaussian process regression" used in this article refers to a Bayesian nonparametric regression model that uses a covariance function to characterize the relationship between input and output and to provide a predicted distribution.

[0015] The term “instrument ionization efficiency” as used in this article refers to the efficiency with which a compound is ionized in a mass spectrometer to produce a detectable signal, usually expressed as the ratio of peak area to concentration or its common logarithm.

[0016] The term "SMILES" used in this article refers to a linear symbol system that concisely and uniquely describes molecular structures using ASCII strings.

[0017] The term "molecular descriptor" as used in this article refers to a set of numerical features that transform a molecular structure into a computable value, used to quantify its chemical or physical properties.

[0018] The term "proportion coding" as used herein refers to the encoding of variables such as mobile phase or component proportions into normalized numerical features that preserve their relative composition relationships.

[0019] The term "RDKit" as used herein refers to an open-source cheminformatics software package for parsing molecular structures and calculating their descriptors.

[0020] The term "standardization" as used herein refers to the conversion of feature data of different dimensions or scales into a standard distribution with a mean of 0 and a variance of 1.

[0021] The term "hyperparameter optimization" as used herein refers to the process of adjusting key parameters of a model outside of model training through search or heuristic methods to improve prediction performance.

[0022] The term "R 2 " as used herein refers to the coefficient of determination, reflecting the model's ability to explain the variance of the target variable, with a value ranging from 0 to 1.

[0023] The term "RMSE" as used herein refers to the root mean square error, which measures the root mean square value of the deviation between predicted and true values.

[0024] The term "independent test set" as used herein refers to a data set that is not used in the model construction and parameter tuning process and is dedicated to the final performance evaluation.

[0025] In a first aspect, the present application provides an organic matter non-targeted full component quantitative analysis method based on a Gaussian process regression machine learning model, which comprises the following steps: (1) Data integration and preprocessing: collecting organic matter high-resolution data, classifying, transforming, and filling in blank values to obtain organic matter molecular parameters; (2) Model construction: constructing a Gaussian process regression model according to the organic matter molecular parameters and IE, and evaluating the model performance; (3) Model parameter optimization: optimizing model parameters to improve prediction accuracy and evaluating the performance after parameter tuning; (4) Model prediction: realizing the quantitative analysis of organic matter non-targeted full component data.

[0026] Among them, the technical features include data integration and preprocessing, model construction, model parameter optimization, and quantitative analysis of organic matter non-targeted full component data.

[0027] Specifically, in step (1), the technical feature of data integration and preprocessing, the collected data includes but is not limited to the following three pieces of information: the string form SMILES of the molecular structure of organic matter, the pH value of the aqueous phase pH_aq, and the mobile phase Solvent used in high-resolution liquid chromatography-mass spectrometry analysis, and the three are the key data contained.

[0028] Further specifically, the SMILES is converted into molecular descriptors by the RDKit library; the Solvent is classified by column data using proportional coding.

[0029] Specifically, in step (1), the technical feature data integration and preprocessing, the blank values are filled with 0.

[0030] Specifically, in step (2), the model construction includes but is not limited to: the model is from the open source machine learning library scikit-learn (sklearn); the logarithm of IE is set as the target function, and a Gaussian process regression model is used to construct the correlation between various organic matter molecular parameters and LogIE; the data set is randomly divided into 65% training set, 20% test set and 15% validation data set, which are respectively used for training model, evaluating model and verifying model; the data is standardized before model training; five-fold cross-validation is used in model training to optimize any one or more steps in model performance.

[0031] Specifically, in step (3), the technical feature model parameter optimization includes but is not limited to: selection or adjustment optimization of the Kernel parameter of the Gaussian process regression model; adjustment optimization of the alpha parameter of the Gaussian process regression model; selection of the optimizer parameter of the Gaussian process regression model; utilization of the coefficient R 2 and the root mean square error RMSE to evaluate the fitting degree and prediction error of the model in any one or more steps.

[0032] Specifically, in step (4), the quantitative analysis of non-targeted whole component data of organic matter includes any one of the following: using an additional data set that does not repeat the data in the model construction process to verify the optimal model after parameter adjustment; realizing the quantitative analysis of non-targeted whole component of organic matter by the optimal model.

[0033] Based on the further solution or simultaneous solution of multiple technical problems of the technical problems of the present application, in the technical scheme provided in the first aspect of the present application, the preferred scheme includes: The first priority scheme: the data collected in step (1) is a string form SMILES of organic matter molecular structure, a data set of aqueous phase pH value pH_aq and mobile phase Solvent used in high-resolution liquid chromatography-mass spectrometry analysis, the solvent column data is classified according to the proportional coding form, SMILES is converted into molecular descriptors by using RDKit tool, and the blank value is filled with 0. On the basis of solving the technical problem that "data is difficult to effectively integrate and pretreat, affecting the quality of model input", the technical problem of "data format is not unified, which cannot be directly used for model training" is further solved.

[0034] The second priority scheme: step (2) converts the logarithm of IE into a target function, uses a Gaussian process regression model to construct the correlation between various organic matter molecular parameters and LogIE, and divides the data set into 65% training set, 20% test set and 15% validation data set for training model, evaluating model and verifying model respectively, and standardizes the data before model training, and adopts five-fold cross-validation to optimize the model performance in model training. On the basis of solving the technical problem that "the model cannot accurately construct the correlation between the molecular parameters and the concentration function", the technical problem of "large difference in data dimension, slow model convergence speed and poor prediction stability" is further solved.

[0035] The third priority scheme: step (3) adjusts and optimizes the Kernel, alpha, optimizer and other parameters of the Gaussian process regression model, uses the coefficient R 2 and the root mean square error RMSE to evaluate the fitting degree and prediction error of the model. On the basis of solving the technical problem that "the model parameters are unreasonable and the prediction accuracy is low", the technical problem that "the model performance cannot be effectively evaluated, and it is difficult to determine the optimal model" is further solved.

[0036] The fourth priority scheme: step (4) uses an additional data set that does not repeat the data in the model construction process to verify the optimal model after parameter adjustment, and realizes the quantitative analysis of non-targeted full components of organic matter. On the basis of solving the technical problem that "the model generalization ability is unknown, and it cannot be applied to actual sample quantification", the technical problem that "the prediction accuracy of the model for unknown data is low" is further solved.

[0037] In a second aspect, the present application provides a data preprocessing tool for non-targeted quantitative analysis of organic matter, comprising a data acquisition module, an encoding module, a conversion module and a filling module, the data acquisition module is used to acquire SMILES, pH_aq and Solvent data of the sample; the encoding module is used to proportionally encode the Solvent data; the conversion module is used to convert SMILES into molecular descriptors by RDKit; and the filling module is used to fill the blank values with 0.

[0038] Specifically, the pH_aq of the sample ranges from 1 to 13.

[0039] The data acquisition module, the encoding module, the conversion module and the filling module are included.

[0040] The data acquisition module includes but is not limited to supporting batch import of SMILES string files, automatically identifying pH_aq numerical format and having data format verification function.

[0041] The encoding module is selected from the following technical features: supporting automatic splitting and parsing of mobile phase characters; extracting solvent components and proportion information; generating numerical proportion coding matrix; and the encoding result can be exported in CSV or XLSX format.

[0042] The conversion module is selected from the following technical features: integrating RDKit2023.09 and above version tools; generating more than 200 kinds of molecular descriptors; and supporting batch conversion of SMILES data set.

[0043] The filling module is selected from the following technical features: supporting 0 filling, mean filling and median filling modes; automatically identifying blank value positions and marking; and being able to backtrack original unfilled data.

[0044] Based on further solving or simultaneously solving multiple technical problems of the present application, in the third aspect of the present application, the preferred solution includes: The first preferred solution: the data acquisition module supports batch import of SMILES files and verifies the format, and the encoding module realizes automatic splitting and encoding matrix generation of the mobile phase. This technical solution further solves the technical problem of "low data acquisition efficiency and format disorder" on the basis of solving the technical problem of "non-uniform mobile phase coding leading to poor model compatibility".

[0045] The second preferred solution: the conversion module integrates the latest version of RDKit and can batch convert molecules, and the filling module supports multiple filling modes. This technical solution further solves the technical problem of "single blank value processing method and uncontrollable data quality" on the basis of solving the technical problem of "time-consuming molecular descriptor conversion and invalid data interfering with the model".

[0046] In a third aspect, the present application provides an organic matter non-target quantitative analysis system based on a Gaussian process regression model, comprising a data input unit, a model training unit and a quantitative analysis unit, the data input unit receives preprocessed data; the model training unit constructs and optimizes a Gaussian process regression model; and the quantitative analysis unit outputs non-target full component quantitative results.

[0047] The technical features include a data input unit, a model training unit and a quantitative analysis unit.

[0048] The technical feature of the data input unit is selected from the group consisting of support for CSV and XLSX data format import; built-in data cleaning subunit (standardization); and linkage with the preprocessing tool of the second aspect.

[0049] The technical feature of the model training unit is selected from the group consisting of built-in scikit-learn Gaussian process regression algorithm; support for designing the Kernel function parameter as a combination of constant kernel and RBF kernel, setting three initial combinations for optimization; integration of grid search parameter optimization tool; and output of the prediction-actual value graph of the training process.

[0050] The technical feature of the quantitative analysis unit is selected from the group consisting of support for batch import of non-target mass spectrometry processed data; automatic calling of the optimal model to calculate the component concentration function; output of the concentration function result and visualized chart; and export of the result in CSV or XLSX format.

[0051] Based on further solving or simultaneously solving multiple technical problems of the technical problem of the present application, in the technical solution provided in the fourth aspect of the present application, the preferred solution comprises: The first preferred solution is that the data input unit supports multi-format import and linkage with the preprocessing tool, and the model training unit performs preset space exhaustive search optimization. This technical solution further solves the technical problem of "data access is cumbersome and model parameter debugging has high professional threshold" on the basis of solving the technical problem.

[0052] The second preferred solution is that the quantitative analysis unit automatically calculates the component concentration function and supports multi-format export of the result. This technical solution further solves the technical problem of "lack of result reliability evaluation index" on the basis of solving the technical problem of "poor readability of non-target quantitative result and limited application scenarios".

[0053] In the present application, example 1 at least supports the protection scope of claim 1.

[0054] The technical feature "data integration and preprocessing" is generalized from the technical features "collecting SMILES, pH_aq, and solvent data sets", "classifying the solvent column data in a proportional coding form", "converting SMILES to molecular descriptors using RDKit tools", and "filling blank values with 0" in the foregoing explanations and / or the corresponding technical features in Example 1 by the common feature "data preprocessing is to convert raw data into a format recognizable by the model". Therefore, a person skilled in the art can reasonably infer that the technical feature "data integration and preprocessing", its subordinate concepts, substantially equivalent technical means, and technical means replaceable based on the existing technical level within conventional technical means and common knowledge, should all fall within the protection scope of claim 1, for example, replacing "filling blank values with 0" with filling with the mean value, which still falls within the protection scope of claim 1 of the present application.

[0055] The technical feature "model construction" is generalized from the technical features "taking LogIE as the target function", "constructing a Gaussian process regression model of molecular parameters and LogIE", "dividing the data set into a training set, a test set, and a validation data set in a ratio of 6.5:2:1.5", "data standardization processing", and "five-fold cross-validation" in the foregoing explanations by the common feature "model construction is to establish the relationship between molecular parameters and IE". Therefore, a person skilled in the art can reasonably infer that equivalent means of the technical feature "model construction" (such as replacing the data set division ratio 6.5:2:1.5 with 6:2:2) fall within the protection scope of claim 1.

[0056] The technical feature "model parameter optimization" is generalized from the technical features "selecting or optimizing Kernel parameters", "optimizing alpha parameters", "selecting optimizer parameters", "evaluating performance by RMSE and R2", and "selecting the best model" in the foregoing explanations by the common feature "parameter optimization is to improve model performance". Therefore, a person skilled in the art can reasonably infer that other parameter optimization methods (such as Bayesian optimization) also fall within the protection scope of claim 1. 2 and RMSE evaluation performance" by the common feature "parameter optimization is to improve model performance". Therefore, a person skilled in the art can reasonably infer that other parameter optimization methods (such as Bayesian optimization) also fall within the protection scope of claim 1.

[0057] The technical feature "quantitative analysis of non-targeted total component data of organic matter" is generalized from the technical features "verifying the optimal model with additional data sets" and "outputting non-targeted total component quantitative results" in the foregoing explanations by the common feature "quantitative analysis is to output the final results". Therefore, a person skilled in the art can reasonably infer that other verification data set methods also fall within the protection scope of claim 1.

[0058] The present application has the following beneficial effects: 1. Compared with the prior art, the present application has significant advantages in quantitative accuracy, model generalization ability, analysis cost, applicable scope and the like, and embodies better technical effects.

[0059] According to experimental tests, under the same data set, the Gaussian process regression model constructed by the present application has a test set R 2 from 0.72 of the prior art to above 0.83, indicating that the model has stronger fitting ability and generalization ability in explaining the relationship between variables.

[0060] According to experimental tests, the RMSE of the model in the quantitative analysis process of the present application is reduced from 0.78 of the prior art to below 0.54, significantly improving the prediction accuracy of complex organic matter components and enhancing the applicability to environmental samples.

[0061] In addition, based on the present application: Based on Example 1, the present application adopts the technical means combination of "data preprocessing (including multi-parameter integration) + Gaussian process regression modeling + automatic parameter optimization", and achieves a new technical effect - the test set R 2 reaches 0.83, which is much higher than the sum of the effects of data preprocessing alone (R 2 = 0.62), modeling alone (R 2 = 0.70) or manual parameter optimization (R 2 = 0.75). The combined technical effect is more superior than the effect of each technical means, embodying the synergistic effect. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 For Figure 1 For the flowchart of the organic matter non-targeted full component quantitative analysis method of the present application based on the Gaussian process regression machine learning model; Figure 2 For the training set prediction performance chart of the Gaussian process regression model before (left) and after (right) hyperparameter optimization; Figure 3 For the test set prediction performance chart of the Gaussian process regression model before (left) and after (right) hyperparameter optimization; Figure 4 For the validation data set prediction performance chart of the Gaussian process regression model after hyperparameter optimization. DETAILED DESCRIPTION

[0063] The following non-limiting examples can enable those of ordinary skill in the art to more fully understand the present application, but do not limit the present application in any way. The following content is only an exemplary description of the scope of the present application, and those skilled in the art can make various changes and modifications to the present application based on the disclosed content, and it should also belong to the scope of the present application.

[0064] The present application will be further described in the manner of specific examples. The various instruments, devices, equipment, reagents, products, etc. used in the embodiments of the present application are obtained through conventional commercial channels unless otherwise specified.

[0065] Example 1 An organic matter non-targeted full component quantitative analysis method based on a Gaussian process regression machine learning model of the present application, specifically, a machine learning model is constructed to realize organic matter quantitative analysis, and the steps are as follows: (1) Data integration and preprocessing: collect organic matter high-resolution data, and perform classification, transformation, blank value filling, etc. on the data; the collected data includes the string form SMILES of the molecular structure of organic matter, the data set of the pH value of the aqueous phase pH_aq and the mobile phase Solvent used in high-resolution liquid chromatography-mass spectrometry analysis.

[0066] Table 1 Example of basic data

[0067] The mobile phase data in the Solvent data is processed into a ratio and coded, and participates in model training; The coding method of the mobile phase of the Solvent data is shown in Table 2: Table 2 Example of Solvent transformation form (horizontal axis is original data, vertical axis is transformed data)

[0068] Use the RDKit tool to convert SMILES into molecular descriptors, each SMILES corresponds to a set of molecular descriptors, forming a feature matrix for model input; The molecular descriptors converted from SMILES are shown in Table 3: Table 3 Example of SMILES transformation

[0069] After the above steps, delete irrelevant columns such as the SMILES column, the solvent column, etc. Fill in the blank values with 0 to prevent the model from running error because it cannot read the empty values; (2) Model construction: construct a Gaussian process regression model of the molecular parameter and the concentration function, i.e. the ionization efficiency (IE) of the instrument, and evaluate the model performance; The above model comes from the open source machine learning library scikit-learn (sklearn), which has scalability and supports replacement and optimization of multiple regression algorithms.

[0070] The logarithmic transformation value of the concentration function IE is set as a target function, and a Gaussian regression model is used to build the correlation between various organic matter molecular parameters and LogIE. The data set is randomly divided into a 65% training set, a 20% test set and a 15% validation data set, which are respectively used for training the model, evaluating the model and verifying the model, and the results are as shown in Figure 1 , and Figure 2 , respectively. The data is standardized before model training to eliminate the dimensional differences between different feature dimensions and improve the convergence speed and prediction stability of the model. The data is five-fold cross-validated in the model training to improve the stability and accuracy of performance evaluation, assist parameter tuning, and thus improve the generalization ability of the model.

[0071] (3) Model parameter optimization: optimize the model parameters to improve the prediction accuracy of the model, and evaluate the performance of the model after parameter tuning. The above steps adjust and optimize the parameters such as Kernel, alpha, optimizer of the Gaussian process regression model to improve the fitting degree of the model, and the results are as shown in Figure 1 , and Figure 2 , respectively. The coefficient R 2 and the root mean square error RMSE are used to evaluate the fitting degree and prediction error of the model. 2 The closer R 2 is to 1, the better the fitting effect of the model on the data, and the smaller the RMSE, the smaller the deviation between the predicted value and the true value of the model, and the higher the precision of the model. The same steps as the above Gaussian process regression model are used to build the Bayesian linear model and the Bayesian ridge regression model of the molecular information of the organic matter components and the concentration function.

[0072] The training set and the test set of the Gaussian process regression model, the Bayesian linear model and the Bayesian ridge regression model for predicting the concentration function of organic matter are shown in Table 4. 2 The R 2 and RMSE of the training set before hyperparameter operation of the Gaussian process regression model are the best among the three, indicating that the Gaussian process regression has the best adaptability to the collected data, realizes the quantitative analysis of organic matter under controllable error, and the hyperparameter operation further improves the fitting effect of the Gaussian regression model. Both the training set and the test set are the best among the three, proving that the method of the present application can realize the correlation between the molecular parameters of organic matter and the concentration function with high precision and low error, and has the advantage of model performance.

[0073] Table 4 Comparison of prediction effects of Gaussian process regression model, Bayesian linear model and Bayesian ridge regression model on concentration function of organic matter

[0074] (4) Quantitative analysis of non-targeted whole component data of organic matter: The concentration function was predicted by using the organic matter molecular data set which was not repeated in model construction, and the Gaussian process regression model of optimal parameters was formed to form the organic matter quantitative analysis scheme, and the results are shown in Figure 3 The above steps were repeated to construct the organic matter quantitative analysis scheme of the Bayesian linear model and the Bayesian ridge regression model of optimal parameters, and the results are shown in Figure 4 and Table 5.

[0075] Table 5 Performance comparison of Gaussian process regression model, Bayesian linear model and Bayesian ridge regression model of optimal parameters

[0076] The prediction effect of the Gaussian process regression model, the Bayesian linear model and the Bayesian ridge regression model of optimal parameters on the additional organic matter concentration function is shown in Figure 4 and Table 5, it can be seen that the fitting effect (R 2 = 0.83) of the Gaussian process regression model is the best, and the deviation from the measured data is the lowest (RMSE = 0.54), which shows that the method of the present application still has the best prediction effect for the organic matter component data outside the model construction data.

[0077] Finally, it should be noted that the above content is only used to illustrate the technical solutions of the present application, and is not a limitation on the protection scope of the present application. Simple modifications or equivalent replacements of the technical solutions of the present application made by those skilled in the art do not deviate from the essence and scope of the technical solutions of the present application.​

Claims

1. An organic matter non-targeted whole component quantitative analysis method based on a Gaussian process regression machine learning model, characterized in that, The method comprises the following steps: (1) data integration and preprocessing: collecting high-resolution data of organic matter, classifying, transforming, filling blank values of the data, and obtaining molecular parameters of the organic matter; (2) model construction: constructing a Gaussian process regression model according to the molecular parameters of the organic matter and IE, and evaluating the performance of the model; (3) model parameter optimization: optimizing the model parameters to improve the prediction accuracy, and evaluating the performance after parameter optimization; (4) model prediction: realizing quantitative analysis of non-targeted full component data of the organic matter.

2. The method according to claim 1, wherein the method is characterized by, In step (1), the high-resolution data of the organic matter comprises the following three information: string form SMILES of the molecular structure of the organic matter, aqueous pH value pH_aq and mobile phase Solvent used in high-resolution liquid chromatography-mass spectrometry analysis.

3. The method according to claim 2, wherein the method is a method for non-targeted whole component quantitative analysis of organic substances. The SMILES is converted into molecular descriptors by the RDKit library; and the Solvent is classified by proportion coding.

4. The method according to claim 1, wherein the method is characterized by, In step (1), the blank value is filled with 0.

5. The method according to claim 1, wherein the method is characterized by, In step (2), the logarithm of IE is used as the objective function, and a Gaussian process regression model is used to construct the correlation between the molecular parameters of the organic matter and LogIE.

6. The method according to claim 1, wherein the method is characterized by, In step (3), the model parameters include Kernel, and can further include alpha, optimizer and the like.

7. The application of the method for quantitative analysis of non-targeted full components of organic matter according to any one of claims 1-6 in high-precision quantitative analysis of non-targeted full components of organic matter.

8. A data preprocessing tool for non-targeted quantitative analysis of organic matter, comprising a data acquisition module, an encoding module, a transformation module and a filling module, wherein the data acquisition module is used to acquire SMILES, pH_aq and Solvent data of the sample; the encoding module is used to proportionally encode the Solvent data; the transformation module is used to convert the SMILES into molecular descriptors by RDKit; and the filling module is used to fill the blank values with 0.

9. The data pre-processing tool of claim 8, wherein, The pH_aq of the sample ranges from 1 to 13.

10. An organic matter non-targeted total component quantitative analysis system, characterized in that, The tool comprises a preprocessing module, the Gaussian process regression model according to any one of claims 1-6 and a result output module.

Citation Information

Patent Citations

  • Ensemble Gaussian Process Regression Model Method for Polymer Molecular Weight Distribution

    CN111797574B