Liquid chromatographic analysis aqueous phase type prediction method based on machine learning, medium and program product
Through machine learning-based methods, quantum chemistry software is used to extract the molecular structure characteristics of organic matter and build a prediction model, which solves the problem of water phase type prediction in liquid chromatography analysis, and achieves efficient and accurate prediction, which is suitable for the rapid analysis needs of multiple fields.
Patent Information
- Application Number
- CN202510268396.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
In the liquid chromatography analysis of organic matter, it is difficult to accurately predict the type of water phase, which leads to the time-consuming and inaccurate method development, which cannot meet the rapid analysis needs of large quantities of substances.
Using a machine learning-based method, the molecular structural characteristics of organic matter are extracted through quantum chemistry software, and a prediction model is constructed in combination with machine learning algorithms to predict the water phase type. Specific steps include data collection and classification, structural feature extraction, data processing and model construction, model verification and optimization, and prediction application.
It achieves efficient and accurate prediction of the water phase type, significantly accelerates the development process of liquid chromatography analysis methods, improves prediction accuracy and efficiency, and is suitable for environmental testing, medicine, materials and other fields.
Smart Images

Figure CN120216953A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting the type of aqueous phase, and in particular to a method, medium, and program product for predicting the type of aqueous phase in liquid chromatography analysis based on machine learning. Background Art
[0002] During the development process of liquid chromatography analysis methods for organic substances, it is necessary to determine the type of aqueous phase, whether to use pure water as the mobile phase, and whether to add buffer salts. For substances without literature references, a traditional procedure is to prepare various aqueous solutions with gradually increasing complexity of components as the mobile phase, such as starting from a pure water phase and, if a satisfactory chromatogram cannot be obtained, then using an aqueous solution with added buffer salts; another method is to calculate or experimentally determine the acid dissociation constant pKa of the organic substance and judge the proportion of different forms of the organic substance at neutral pH. If the proportion of the ionic form is relatively high, it indicates that buffer salts may need to be added to the aqueous phase.
[0003] However, both of these methods have limitations of long time consumption or insufficient accuracy. The first method requires more attempts and takes a long time; for the second method, the calculation model of pKa is not completely accurate, and experimentally determining pKa requires a long period. It can be seen that the existing methods at present cannot meet the requirements of method development for a large number of substances.
[0004] In recent years, machine learning technology has gradually been applied to the field of analytical chemistry. For example, CN114255932A discloses a method for predicting the reactivity of new pollutants based on machine learning, which constructs a regression model through multi-stage feature enhancement analysis (MFEA) to predict the reaction rate constants (k values) of sulfate radicals and carbonate radicals. Although this technology shows certain advantages in reactivity prediction (R in the test set 2 > 0.86), it focuses on the continuous value regression prediction of reaction rates, while the type of aqueous phase is a discrete classification problem (Y / N), and its two-stage modeling (classification + regression) framework cannot be directly applied; it relies on static feature selection of molecular fingerprints (MACCS / ECFPs) and does not optimize molecular descriptors (such as polarity parameters, topological indices) for chromatographic analysis requirements.
[0005] Therefore, the existing technology still lacks a dedicated prediction method for the chromatographic aqueous phase type classification scenario, and there is an urgent need for an efficient solution that integrates dynamic feature optimization, data balance processing, and classification verification mechanisms. Summary of the Invention
[0006] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a method, medium, and program product for predicting the type of aqueous phase in liquid chromatography analysis based on machine learning, which can predict whether buffer salts need to be added to the aqueous phase and is conducive to accelerating the development process of liquid chromatography analysis methods for a large number of substances.
[0007] In addition, the solution of the present invention is simple and easy to use, easy to get started, has high accuracy, and has a wide range of applications. It can be used in the development process of organic liquid chromatography analysis methods in disciplines such as environmental detection, environmental ecology, medicine, and materials.
[0008] The object of the present invention can be achieved by the following technical solutions:
[0009] In the first aspect of the present invention, a prediction method for the aqueous phase type of liquid chromatography analysis based on machine learning is provided, including the following steps:
[0010] S1. Data collection and classification: Obtain the aqueous phase type data of the organic liquid chromatography analysis method, and divide the data into category Y that does not require buffer salts and category N that requires buffer salts according to whether buffer salts are added.
[0011] S2. Structural feature extraction: Obtain the chemical structural formula information of the organic matter, and calculate a feature data set including 1D and 2D molecular descriptors and molecular fingerprints through quantum chemistry software.
[0012] S3. Data processing and model construction: Divide the feature data set into a training set and a test set, preprocess the training set to optimize the feature quality, and construct an aqueous phase type prediction model through a machine learning algorithm.
[0013] S4. Model verification and optimization: Double-verify the model performance through the Kappa values of the training set and the test set.
[0014] S5. Prediction application: Extract the chemical structure features of the organic matter to be measured and input them into the optimized model in S4 to predict the aqueous phase type required for its liquid chromatography analysis.
[0015] Further, in S2, the quantum chemistry software is PaDEL-Descriptor, and the output molecular descriptors include at least three combinations of atomic type, bond type, topological parameters, and electrical parameters.
[0016] Further, in S3, the preprocessing includes at least one of the following operations:
[0017] Eliminate variables with more than 95% of the same data;
[0018] Eliminate variables with null values;
[0019] Eliminate variables with a Pearson correlation coefficient > 0.9;
[0020] In S3, the data division ratio is training set: test set = 8:2, and the division process uses a random sampling method.
[0021] Further, in S3, data balancing processing is also included:
[0022] When the number of samples of class N in the training set exceeds that of class Y, the SMOTE or ADASYN oversampling algorithm is used to enhance the samples of class Y;
[0023] The oversampling multiple is dynamically set according to the original data distribution, and the difference in the number of samples between the two classes after enhancement is controlled within ±10%;
[0024] The oversampling multiple is calculated by the following formula: Enhancement multiple = Number of samples of class N × Balance coefficient / Number of samples of class Y,
[0025] where the balance coefficient ranges from 0.8 to 1.2, and the final total number of samples does not exceed 5 times the original data volume.
[0026] Further, in S3, the machine learning algorithm is selected from at least one of the following: XGBoost gradient boosting tree, random forest, support vector machine, deep neural network, K-nearest neighbor algorithm.
[0027] Further, when the XGBoost algorithm is selected, its hyperparameter combination satisfies the following conditions:
[0028] Range of number of trees: 500 - 3000 trees;
[0029] Range of learning rate: 0.01 - 0.1;
[0030] Range of maximum tree depth: 1 - 5 layers;
[0031] The loss function selects the binary classification logistic regression function.
[0032] Further, in S4, the model validation and optimization include the following mechanisms:
[0033] When the Kappa value of the test set ≤ 0.75, adjust the type of machine learning algorithm;
[0034] When the Kappa value of the training set ≤ 0.75, adjust the hyperparameter combination of the selected algorithm.
[0035] Further, in S5, it also includes the post-processing process of the prediction result:
[0036] When the prediction confidence level output by the model is lower than 85%, the following operations are automatically performed:
[0037] a) Generate a structural similarity analysis report of the organic matter;
[0038] b) Provide at least two alternative aqueous phase scheme suggestions;
[0039] c) Mark the identification for manual review in the output result.
[0040] In a second aspect of the present invention, there is provided a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to execute the structural apparent disease detection data standardization batch processing method as described above.
[0041] In a third aspect of the present invention, there is provided a computer program product, including a computer program which, when executed by a processor, implements the liquid chromatography analysis aqueous phase type prediction method based on machine learning as described above.
[0042] The core principle of the present invention lies in: extracting the organic molecule structure characteristics (1D / 2D descriptors and molecular fingerprints) through quantum chemical calculations, and combining machine learning algorithms to construct a prediction model, replacing the traditional trial-and-error experiments or pKa determinations in a data-driven manner. Specifically, it is realized through the following technical chain - collecting historical chromatographic data and classifying and labeling → quantifying molecular characteristics → optimizing the feature set → balancing the data distribution → constructing and validating a classification model (Kappa value > 0.75), and finally quickly predicting the aqueous phase type (requiring buffer salt / pure water) of new substances based on the model, achieving a double improvement in the development efficiency and prediction accuracy of chromatographic analysis methods.
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] 1) The present invention applies machine learning methods to the prediction of the aqueous phase type in the organic liquid chromatography analysis method, which can provide an efficient and relatively accurate prediction tool, and improve the development efficiency of chromatographic analysis methods for a large number of organic substances.
[0045] 2) The present invention is simple and easy to use, and can be used in the development process of organic liquid chromatography analysis methods in disciplines such as environmental detection, environmental ecology, medicine, and materials.
[0046] 3) The present invention has low requirements for computer hardware, and can be realized on common personal computers, with strong practicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a concise flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] Overall, the method for predicting the aqueous phase type of the liquid chromatography analysis method based on machine learning in the present invention includes the following steps:
[0049] 1) Step 1: Collect the aqueous phase type data in the literature, and divide the data into two categories according to whether buffer salt is added to the aqueous phase, denoted as Y (no need to add buffer salt) and N (need to add buffer salt);
[0050] 2) Step 2: Search for or draw the chemical structural formula of the organic substance in Step 1, and save it as a mol file;
[0051] 3) Step 3: Use the PaDEL-Descriptor quantum chemistry software. Take all mol files as input, select 1D&2D and Fingerprints, and calculate molecular descriptors and molecular fingerprints. Randomly divide the obtained organic compounds into a training set and a test set in a ratio of 8:2. Preprocess the molecular descriptors and molecular fingerprints of the training set: i) Remove variables with more than 95% of the data being the same; ii) Remove variables with null values; iii) Remove variables with a Pearson correlation coefficient > 0.9.
[0052] 4) Step 4: Since in the collected data, the category classified as N is more than the category classified as Y, and the data volume is unbalanced. Use an oversampling method for the training set, such as the Smote or Adasyn method, to enhance the data of category Y.
[0053] 5) Step 5: Use machine learning methods such as support vector machines, gradient boosting, neural networks, random forests, and K-nearest neighbors. Take the classification in Step 1 as the dependent variable for the training set, and the molecular descriptors and molecular fingerprints obtained after preprocessing in Step 3 as the independent variables to establish a binary classification model.
[0054] 6) Step 6: Use this model to calculate the classification of the test set and calculate the Kappa values of the training set and the test set. If the Kappa value > 0.75, then the established model can be used for predicting the aqueous phase type of the new organic compound liquid chromatography analysis method. Otherwise, adjust the machine learning method and the hyperparameter values until the Kappa values of the training set and the test set are both > 0.75.
[0055] 7) Step 7: Prediction of the aqueous phase type of the new organic compound liquid chromatography analysis method: Perform the prediction of the new organic compound according to Step 2, Step 3, and Step 6.
[0056] The data processing, modeling, and prediction processes in Step 3, Step 4, Step 5, and Step 7 can be implemented using programming languages such as R and Python.
[0057] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. Features such as component models, material names, connection structures, control methods, algorithms, etc. that are not clearly stated in this technical solution are regarded as common technical features disclosed in the prior art.
[0058] Example 1
[0059] This example provides a modeling process using the liquid chromatography analysis method in the Chinese Pharmacopoeia as the data set. A simple schematic diagram of the process is as Figure 1 shown.
[0060] Step 1: Obtain the data of the aqueous phase types in the organic liquid chromatography analysis methods in the Chinese Pharmacopoeia. A total of 656 pieces of data were collected this time. According to whether buffer salts are added to the aqueous phase, the data were divided into two categories, denoted as Y (no need to add buffer salts) and N (need to add buffer salts), with 105 and 551 pieces of data in categories Y and N respectively;
[0061] Step 2: Obtain the chemical structural formulas of the organic compounds in Step 1 and save them as mol files;
[0062] Step 3: Use the PaDEL-Descriptor quantum chemistry software. Take all mol files as inputs, select 1D, 2D, and Fingerprints, and calculate the molecular descriptors and molecular fingerprints; randomly divide the obtained organic compounds into a training set and a test set according to an 8:2 ratio (i.e., 525:131), and preprocess the molecular descriptors and molecular fingerprints of the training set: i) Remove variables with more than 95% of the data being the same, ii) Remove variables with null values, iii) Remove variables with a Pearson correlation coefficient > 0.9;
[0063] Step 4: Since in the collected data, the category classified as N is more than the category classified as Y, and the data volume is unbalanced, use the Smote oversampling method for the training set to enhance the data of category Y, with the multiple set to 2; after enhancement, there are 882 pieces of data in both category Y and category N;
[0064] Step 5: Use the Extreme Gradient Boosting method XGBoost. Take the classification in Step 1 as the dependent variable for the training set, and the molecular descriptors and molecular fingerprints obtained after preprocessing in Step 3 as the independent variables to establish a binary classification model;
[0065] Step 6: Use this model to calculate the classification of the test set and calculate the Kappa values of the training set and the test set; if the Kappa value > 0.75, the established model can be used for predicting the aqueous phase type of new organic liquid chromatography analysis methods; otherwise, adjust the machine learning method and the hyperparameter values until the Kappa values of the training set and the test set are both > 0.75. After repeated training, the finally obtained hyperparameter values are: use the binary classification logistic regression loss function, 2000 trees, learning rate 0.06, minimum weight of leaf nodes 1, and maximum tree depth 1;
[0066] Check the prediction accuracy of the model for the training set and the test set. See Table 1, and the prediction effect is good.
[0067] Table 1 XGBoost model characterization
[0068] Dataset TP TN FP FN Ac% Se% Sp% F1-score Kappa Training set 882 882 0 0 100 100 100 1 1 Test set 108 16 5 2 94.66 98.18 76.19 0.9686 0.7893
[0069] Note: TP True Positive; TN True Negative; FP False Positive; FN False Negative; positive is the buffer salt solution, negative is pure water; Ac Accuracy; Se Sensitivity; Sp Specificity; F1-score F1 score.
[0070] Step 7: Use the established model to predict the aqueous phase type of liquid chromatography analysis methods for environmental organic pollutants DDT (CAS: 3416-05-5), phenanthrene (CAS: 85-01-8), and perfluorooctanoic acid (CAS: 335-67-1) not included in the dataset. Obtain the mol file according to Step 2, calculate and preprocess the molecular descriptors and molecular fingerprints according to Step 3, and then use the model established in Step 6 for prediction. It can be obtained that the classifications of DDT and phenanthrene are Y, that is, no buffer salt needs to be added in the aqueous phase, and pure water can be used as the liquid chromatography mobile phase; while the classification of perfluorooctanoic acid is N, and a buffer salt needs to be added. The prediction results of the three substances are consistent with the actual detection methods in the literature.
[0071] Thus, the solution of this embodiment has the following technical advantages:
[0072] Through the intelligent prediction of the aqueous phase type by the machine learning model, the defects of repeated trial and error or relying on time-consuming pKa determination in the traditional method are avoided. The embodiment shows that the model can complete the prediction in a single calculation (for example, the prediction time for substances such as DDT < 5 seconds). Compared with the traditional method (usually taking several hours to several days), the efficiency is increased by more than 200 times, meeting the rapid analysis requirements of a large number of substances (≥100 substances / batch).
[0073] Based on the model optimization strategy of the dual verification mechanism (the Kappa values of the training set and the test set > 0.75), the reliability of the prediction is effectively guaranteed. In the embodiment, the accuracy rate of the test set reaches 94.66% (Kappa = 0.7893), the specificity (Sp) is 76.19%, and the sensitivity (Se) is 98.18%, which is significantly better than the traditional empirical judgment method (the accuracy rate is about 60 - 70%), and can reduce the number of invalid experiments by about 30%.
[0074] Adopt the quantitative analysis of molecular descriptors and fingerprint features, combined with the SMOTE oversampling technique (the balance degree reaches 100% after enhancement of category N in the embodiment), to solve the strong dependence on expert experience in the traditional method. Even for new substances without literature records, accurate prediction can still be achieved.
[0075] The model construction and prediction process can be completed on conventional computer hardware (such as a 4-core CPU / 8GB memory). In the embodiment, the training time < 1 min (656 data), and the prediction time for a single sample < 1 second. It supports cross-platform deployment (R / Python, etc.) and is applicable to multi-field scenarios such as pharmaceutical research and development (such as the analysis of substances in the Chinese Pharmacopoeia) and environmental detection (such as the screening of organic pollutants).
[0076] Prediction confidence evaluation and post - processing mechanism (such as automatically generating alternative solutions for samples with a confidence level < 85%), reducing the manual review requirement by more than 40% (in the embodiment, only 2 / 131 test samples require manual intervention), and at the same time providing interpretable decision - making support through the structural similarity analysis report, avoiding the limitations of traditional black - box models.
[0077] Example 2
[0078] This embodiment provides a storage medium containing computer - executable instructions. When the computer - executable instructions in the storage medium are executed by a computer processor, they are used to execute the structural apparent disease detection data standardization batch - processing method as described above. The storage medium provided in this embodiment is one of the key carriers for implementing the core functions of the present invention. By storing computer - executable instructions, it provides an efficient and automated solution for the standardization batch - processing of structural apparent disease detection data. Such a storage medium can be directly read and executed by a computer processor, thus transforming the complex liquid chromatography analysis aqueous phase type prediction process into a program that can be quickly deployed and run. It not only simplifies the cumbersome manual trial - and - error steps in the development process of traditional liquid chromatography analysis methods but also significantly improves the prediction accuracy and efficiency through the high - performance computing ability of machine - learning algorithms. In addition, the versatility of this storage medium enables it to be widely applied in multiple fields such as environmental detection, pharmaceutical research and development, and materials science, providing a convenient and reliable tool for researchers and engineers and promoting the wide application and rapid development of liquid chromatography analysis technology in different disciplines.
[0079] Example 3
[0080] This embodiment provides a computer program product, including a computer program which, when executed by a processor, implements the above-mentioned method for predicting the aqueous phase type in liquid chromatography analysis based on machine learning. The computer program product of this embodiment is a software implementation form of the technical solution of the present invention. Through the efficient execution of the computer program, it realizes the method for predicting the aqueous phase type in liquid chromatography analysis based on machine learning. This program product encapsulates steps such as complex quantum chemical calculations, feature extraction, data processing, model construction and verification into a complete software system. Users only need to input the chemical structure information of the organic compound to quickly obtain the prediction result of the aqueous phase type required for its liquid chromatography analysis. Such a program product is not only easy to operate and easy to get started, but also ensures high accuracy and reliability of the prediction by optimizing the algorithm and data processing process. It can automatically handle the problem of data imbalance, dynamically adjust model parameters, and provide alternative solutions and suggestions for manual review when the confidence level of the prediction result is low, greatly improving the efficiency and success rate of the development of liquid chromatography analysis methods. In addition, this program product has wide applicability and can meet the diverse needs of different users in fields such as environmental monitoring, pharmaceutical research and development, and materials science, providing strong technical support for scientific research and production in related fields.
[0081] The above description of the embodiments is to enable those of ordinary skill in the art to understand and use the invention. It is obvious that those skilled in the art can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art without departing from the scope of the present invention as disclosed should be within the protection scope of the present invention.
Claims
1. A method for predicting water phase type in liquid chromatography analysis based on machine learning, characterized in that: The following steps are involved: S1. Data collection and classification: Obtain the water phase type data including the liquid chromatography analysis method of organic matter, and divide the data into category Y that does not require buffer salt and category N that requires buffer salt according to whether buffer salt is added; S2. Structural feature extraction: obtaining the chemical structure information of the organic compound, and calculating the feature data set including 1D and 2D molecular descriptors and molecular fingerprints through quantum chemical software; S3, data processing and model construction: dividing the feature data set into a training set and a test set, preprocessing the training set to optimize feature quality, and constructing a water phase type prediction model through a machine learning algorithm; S4. Model validation and optimization: The model performance is double-validated by the Kappa value of the training set and the test set; S5. Prediction application: Extract the chemical structure characteristics of the organic matter to be tested and input them into the optimized model in S4 to predict the type of water phase required for its liquid chromatography analysis.
2. The method for predicting the water phase type of liquid chromatography analysis based on machine learning according to claim 1, characterized in that: In S2, the quantum chemistry software is PaDEL-Descriptor, and the output molecular descriptors include a combination of at least three of atom type, bond type, topological parameters, and electrical parameters.
3. The method for predicting water phase type of liquid chromatography analysis based on machine learning according to claim 1, characterized in that: In S3, the preprocessing includes at least one of the following operations: Eliminate variables with more than 95% identical data; Eliminate variables with null values; Variables with Pearson correlation coefficients > 0.9 were eliminated; In S3, the data division ratio is training set: test set = 8:2, and the division process adopts random sampling method.
4. The method for predicting the water phase type of liquid chromatography analysis based on machine learning according to claim 1, characterized in that: S3 also includes data balancing processing: When the number of class N samples in the training set exceeds that of class Y, the SMOTE or ADASYN oversampling algorithm is used to enhance class Y samples; The oversampling multiple is dynamically set according to the original data distribution, and the difference between the two types of samples after enhancement is controlled within ±10%; The oversampling multiple is calculated by the following formula: Enhancement multiple = N-type sample size × balance coefficient / Y-type sample size, The balance coefficient is between 0.8 and 1.2, and the final total sample size does not exceed 5 times the original data size.
5. The method for predicting water phase type of liquid chromatography analysis based on machine learning according to claim 1, characterized in that: In S3, the machine learning algorithm is selected from at least one of the following: XGBoost gradient boosting tree, random forest, support vector machine, deep neural network, and K nearest neighbor algorithm.
6. The method for predicting the water phase type of liquid chromatography analysis based on machine learning according to claim 5, characterized in that: When the XGBoost algorithm is selected, its hyperparameter combination meets the following conditions: Number of trees: 500-3000; Learning rate range: 0.01-0.1; Maximum tree depth range: 1-5 layers; The loss function uses the binary logistic regression function.
7. The method for predicting the water phase type of liquid chromatography analysis based on machine learning according to claim 1, characterized in that: In S4, model validation and optimization include the following mechanisms: When the Kappa value of the test set is ≤0.75, adjust the machine learning algorithm type; When the Kappa value of the training set is ≤ 0.75, adjust the hyperparameter combination of the selected algorithm.
8. The method for predicting water phase type of liquid chromatography analysis based on machine learning according to claim 1, characterized in that: S5 also includes the post-processing process of the prediction results: When the prediction confidence of the model output is less than 85%, the following actions are automatically performed: a) generating a structural similarity analysis report of the organic matter; b) provide at least two alternative aqueous phase solutions; c) Mark the output results as requiring manual review.
9. A storage medium containing computer executable instructions, characterized in that: The storage medium of the computer executable instructions is used to execute the standardized batch processing method for structural surface defect detection data according to any one of claims 1 to 8 when executed by a computer processor.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the method for predicting the water phase type of liquid chromatography analysis based on machine learning as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Self-supporting high-strength twisted-pair double-shielding computer cable processing and forming assembly line
CN114255932A