Method for predicting ionization efficiency of compound mass spectrum based on cosmo-rs and ann algorithm quantitative structure-activity relationship model

The quantitative structure-activity relationship model constructed by COSMO-RS and ANN algorithm solves the problem of low accuracy in the prediction of organic mass spectrometry ionization efficiency in the existing technology, realizes efficient and accurate prediction of mass spectrometry ionization efficiency, and is suitable for mass spectrometry detection of a variety of compounds.

CN115862762BActive Publication Date: 2025-10-10DALIAN POLYTECHNIC UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211393252.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-08
Publication Date
2025-10-10
Estimated Expiration
2042-11-08

AI Technical Summary

Technical Problem

In the prior art, the model for predicting the ionization efficiency of organic mass spectrometry is single and the prediction accuracy is not high.

Method used

A quantitative structure-activity relationship model was constructed using COSMO-RS and ANN algorithms. Feature descriptors were screened through stepwise linear regression. Combined with high-performance liquid spectrometry data collection, a nonlinear prediction model was established. Quantum chemical methods were used to optimize the molecular structure and construct a QSPR model.

Benefits of technology

It improves the accuracy and extensiveness of mass spectrometry ionization efficiency prediction, reduces experimental costs, is suitable for mass spectrometry detection without standards, and has good goodness of fit and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862762B_ABST
    Figure CN115862762B_ABST
Patent Text Reader

Abstract

The application discloses a method for predicting compound mass spectrum ionization efficiency based on a COSMO-RS and ANN algorithm, and belongs to the technical field of predicting compound mass spectrum ionization efficiency. On the basis of knowing the structure of a compound and an elution solvent environment, optimal structural property descriptors are obtained by using a COSMO-RS program to perform thermodynamic simulation calculation, the determined relative mass spectrum ionization efficiency is combined, a nonlinear regression algorithm of an artificial neural network is written, and a quantitative relationship between various structural property descriptors and the mass spectrum ionization efficiency is established. Then, the mass spectrum ionization efficiency of an organic compound can be quickly and effectively predicted. In the data set collection, gradient elution is considered, so that the application range is expanded and the data set is expanded. In addition, in the descriptor calculation, factors such as the compound and the elution solvent are considered, so that the efficiency and accuracy of the prediction of the mass spectrum ionization efficiency are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a method for predicting the ionization efficiency of a compound based on a COSMO-RS and ANN algorithm quantitative structure-property relationship model; and belongs to the technical field of predicting the ionization efficiency of a compound. BACKGROUND

[0002] Metabolomics and other non-targeted analysis of mass spectrometry can simultaneously detect thousands of compounds due to its high-throughput characteristics, which poses a great challenge to traditional mass spectrometry quantitative methods. Ionization efficiency (IE) refers to the efficiency of ionization of the analyte into a gaseous ion form inside the mass spectrometer and the ultimate detection efficiency, which is an important parameter for evaluating the response behavior of a compound in mass spectrometry and is also a key to mass spectrometry quantitative detection without standard samples. In order to get rid of the limitations of a limited number of standard samples which are generally expensive, it is particularly important to develop a simple and accurate theoretical prediction method for estimating the ionization efficiency of a compound.

[0003] Quantitative structure-property relationship (QSPR) refers to a computer modeling method that correlates the molecular structure of an organic compound with its physicochemical properties, environmental behavior and ionization efficiency parameters, which can reduce or replace related experiments, make up for the lack of experimental data and reduce experimental costs. At present, there are few methods for predicting the ionization efficiency of organic compounds, especially the study of nonlinear methods, and the substances studied in the existing models are relatively single in number, and the prediction accuracy of the model needs to be further improved. SUMMARY

[0004] [TECHNICAL PROBLEM]

[0005] The existing technology has a single model for predicting the ionization efficiency of organic compounds, and the prediction accuracy is not high.

[0006] [TECHNICAL SCHEME]

[0007] In view of the above technical problems, the present application considers that the environmental distribution behavior of various organic compounds is a complex process, and the distribution coefficient may involve some nonlinear relationship, therefore, it is necessary to construct a nonlinear prediction model for the ionization efficiency of mass spectrometry which covers a variety of compounds, has a clear algorithm, is easy to apply and popularize, and does not depend on experimental data, and the model is verified and characterized in accordance with the OECD guidelines.

[0008] Given the high dimensionality of feature descriptors, selecting the most useful subset of features from the original variables for modeling becomes increasingly important. To select more appropriate molecular descriptors for constructing QSPR models, a stepwise linear regression method was employed to reduce the dimensionality of the variables. To establish a reliable QSPR model, an artificial neural network (ANN) regression algorithm was employed as a nonlinear algorithm. This method is not only simple to implement but also exhibits robustness and excellent generalization capabilities.

[0009] The present invention provides a method for predicting the mass spectrometric ionization efficiency of organic compounds by constructing a quantitative structure-activity relationship model using quantum chemical methods. The method can quickly and effectively predict the IE value of the compound based on the structural property descriptors of the compound and the solvent in which it is located. This method is conducive to further exploration of the ESI source mass spectrometric ionization mechanism, facilitates mass spectrometry personnel to design experiments and analyze mass spectrometry data, and is also beneficial for providing basic data for auxiliary quantitative mass spectrometry detection without standard substances.

[0010] The object of the present invention is to provide a method for predicting the mass spectrometry ionization efficiency of a compound based on a quantitative structure-activity relationship model of COSMO-RS and ANN algorithm; the method comprises the following steps:

[0011] Step 1) Data Collection: Peak areas corresponding to a range of concentrations of several compounds under different isocratic and gradient elution conditions are collected using a high performance liquid mass spectrometer. A quantitative relationship model between the concentration and peak area of ​​each compound is constructed. Based on the slope of the constructed quantitative relationship model, the relative ionization efficiency (log IE) of the compound mass spectrometry data for the various elution conditions is calculated.

[0012] After sorting the log IE values ​​obtained above, we select one data point from every four log IE values ​​and put it into the validation set. The remaining data points are then put into the training set, i.e., the data points are divided into the training set and the validation set in a ratio of 4:1. The data points in the training set are used to build the model for internal validation, and the data points in the validation set are used for external validation of the model.

[0013] Step 2) Descriptor Calculation: The SMILES number of the organic compound was input into Turbomole software, and the initial molecular structure of the compound was optimized using the Becke-type three-parameter density functional model (B3LYP) algorithm in density functional theory (DFT) and the 6-311G** basis set. Quantum chemical calculations were performed on all atoms in the molecule at the DFT-BP 86TZVP level, using Turbomole's default convergence criteria. The optimized molecular structure of the compound was further calculated using quantum chemistry software (COSMO-RS) to obtain 45 corresponding molecular structure descriptors. In addition, 10 molecular structure descriptors involved in the pubchem public library were collected, for a total of 55 descriptors. After preprocessing, the final descriptors were screened using stepwise linear regression.

[0014] Step 3) Model Establishment: Based on the training set data described in step 1), the natural logarithm of the relative ionization efficiency (log IE) of the compound mass spectrometry data is used as the dependent variable, and the final descriptor selected in step 2) is used as the independent variable. An artificial neural network (ANN) nonlinear regression algorithm model is established using a Python program. The optimized parameters of the established QSPR model are selected using a 5-fold cross-validation algorithm to construct a QSPR model based on the optimal ANN algorithm.

[0015] Step 4) Model verification: Verify the model in two steps:

[0016] a) Evaluation of model goodness of fit and robustness;

[0017] b) Characterize the application domain and evaluate the performance of the model; after verification, proceed to step 5);

[0018] Step 5) Application domain characterization: Characterize the model application domain through Williams diagram;

[0019] Step 6) Model application: using the model to predict the mass spectrometry ionization efficiency of the target compound in a specified solvent environment.

[0020] In one embodiment, the several compounds in step 1) include glycine, alanine, leucine, isoleucine, valine, proline, phenylalanine, methionine, tryptophan, serine, glutamine, threonine, cysteine, asparagine, tyrosine, aspartic acid, glutamic acid, lysine, arginine, histidine, citrulline, taurine, hypotaurine, L-hydroxyproline, spermidine, N-Ω-acetylhistamine, L-pyroglutamic acid, tyramine, histamine, creatinine, L-dihydroorotic acid, L-carnosine, uric acid, betaine, sarcosine, cytosine, hypoxanthine, xanthine, one or more of purine, guanine, uracil, adenine, β-thymidine, cytidine, inosine, xanthine, guanosine, uridine, adenine, pyridoxine, pyridoxal, pyridoxamine, thiamine, biotin, D-pantothenic acid, niacin, nicotinamide, Harman, DMIP, PhIP, IQ[4,5-b], MeIQ, 7,8-DiMeIQx, 8-MeIQx, MeAαC, IQ, IQx, Phe-p-1, AαC, Norharman, and D3-hypoxanthine.

[0021] In one embodiment, the different isocratic elution and gradient elution conditions in step 1) are:

[0022] Isocratic elution: 5% B, 20% B, 35% B, 50% B, 65% B, 80% B, 95% B, the mobile phase ratio remains constant during the process, the elution flow rate is 0.3 mL / min; the elution time is 20 min, and the total ratio of phase A to phase B during the entire process is 1;

[0023] Gradient elution: 1% B, 0 min; 1% B, 0-1.5 min; 1-99% B, 1.5-13; 99% B, 13-16.5 min; 1% B, 16.6-20.0 min; 0.3 mL / min;

[0024] Phase A is pure water phase; phase B is pure methanol phase or phase A is 0.1% by volume formic acid-water solution; phase B is 0.1% by volume formic acid-methanol solution or both.

[0025] In one embodiment, the specific calculation method of the mass spectrometry relative ionization efficiency of the compound in step 1) is:

[0026] i. To eliminate the differences in measurement and model application between instruments and laboratories, the present invention defines the ratio of the analyte (M1) to the standard (M2) as the relative ionization efficiency (RIE). M1 ), which is calculated according to the following equation:

[0027]

[0028] Here, "slope" refers to the slope of the quantitative relationship between the mass spectrometric peak area of ​​the analyte M1 or standard M2 at different dilution concentrations under specified elution conditions. It is obtained by linear regression within the linear range of the signal concentration plot. To facilitate data presentation and analysis, the relative ionization efficiency (RIE) in this invention is presented on a logarithmic scale (log IE). The log IE values ​​of all analytes measured under specified elution conditions are subtracted from the log IE values ​​of M2 measured under the corresponding conditions.

[0029] ii. To facilitate the subsequent calculation of COSMO descriptors, it is necessary to determine the proportion of mobile phase B before entering the mass spectrometer during the IE determination of the analyte, because the ionization process of the analyte in the mass spectrometer is closely related to the solvent environment. Depending on the specific method used in the gradient elution process, since the peak time of the analyte IE determination is known, the proportion of mobile phase B before entering the mass spectrometer can be calculated according to the following equation:

[0030] B%=1 (t<=1.5min or t>=16.6min);

[0031] B%=1+(t-1.5)*(99-1) / (13-1.5) (1.5min <t<13min);

[0032] B%=99 (13min <t<16.5min);

[0033] Wherein t and B% represent the peak elution time of the test compound and the ratio of mobile phase B to the peak elution time, respectively.

[0034] In one embodiment, the preprocessing process in step 2) includes removing constant, near-constant, missing, and descriptors with correlation greater than 0.95.

[0035] In one embodiment, the final descriptors screened in step 2) include molecular size descriptors, polarity descriptors, basicity descriptors, and solubility descriptors.

[0036] In one embodiment, the molecular size descriptor relies solely on the intrinsic properties of the compound, including molecular volume (MV), logMV, molecular surface area (A), relative molecular mass (MW) * 、M+1 * , atomic weight, dipole moment (DM), hydrogen bond donor moment, hydrogen bond acceptor moment, σ2, σ3, σ4, σ5, σ6 * and dipole moment t value.

[0037] In one embodiment, the polarity descriptors rely solely on the intrinsic properties of the compound, including logD, molecular polar surface area (PSA) and WANS, COSMO total charge number, E_COSMO+dE* , E_COSMO-E_gas + dE * , E_gas, E_diel, and Averaging corrdE.

[0038] In one embodiment, the basicity descriptor is related to the compound's own properties and the solvent environment and temperature, including pK a (H2O), pK b (MeCN).

[0039] In one embodiment, the solubility descriptor is related to the compound's own properties and the solvent environment and temperature, including H-Bond interaction energy in mixture (H_HB), gas-liquid transition energy (H_glt), total average interaction energy (H_int), mismatch interaction energy (H_MF) * , chemical potential energy * , H_vdW, log(H) * , Gsolv * , log value of gas partial pressure (mbar) * , molecular free energy * , Henry's law coefficient * , solvent's Gibbs free energy * , ln(γ), log(pv), solvent density * , solvent molar mass and relative solubility.

[0040] In one embodiment, the relative solubility includes log10(x_RS) * , w_RS and log10(S_RS).

[0041] In one embodiment, the step 3) of constructing the QSPR model based on the optimal ANN algorithm specifically includes the following processes:

[0042] (1) Divide the entire data set into 5 sets, each of which will be rotated as a test set, and the remaining set will be used as a training set, and repeat the training and testing 5 times to ensure that each set will be verified once as a test set;

[0043] (2) Adjust different network parameters, optimization parameters, and regularization parameters, calculate and compare the average cross-validation accuracy of the 5 training before and after adjustment, and select the set of parameters with the highest cross-validation accuracy to apply to the QSPR model, that is, to construct the optimized model.

[0044] In one embodiment, the grid parameters of step (2) include the number of network layers and the activation function.

[0045] In one embodiment, the optimization parameters of step (2) include the learning rate and the number of iterations.

[0046] In one embodiment, the regularization parameter in step (2) includes a dropout ratio.

[0047] In one embodiment, when the model is verified in step 4), the goodness of fit and robustness evaluation indicators are: the coefficient of determination R 2 adj , training set root mean square error RMSE tra And the training set mean absolute error MAE tra .

[0048] In one embodiment, the Williams diagram is used to characterize the model application domain in step 5, specifically, the leverage value h is represented by the standard residual δ. i The Williams diagram characterizes the application domain of the model. When the absolute value of δ is greater than 3.0, the compound is an outlier. When the leverage value h i Greater than the warning value h * When , it indicates that the structure of this compound is significantly different from that of other compounds; i and h * Calculated by the following formula:

[0049] h i =x i T (X T X) -1 x i

[0050] h * =3(p+1) / n

[0051] where x i is the descriptor matrix of the ith compound; x i T is x i The transposed matrix of ; X is the descriptor matrix of all compounds; X T is the transposed matrix of X; (X T X) -1 is the matrix X T The inverse of X; p is the number of variables in the model, and n is the number of data points in the dataset.

[0052] [Beneficial Effects]

[0053] The present invention uses the COSMO-RS / Turbomole quantum chemistry software with a simple calculation process to obtain descriptors, more comprehensively considering the influence of the solvent environment on the mass spectrometry ionization efficiency, and at the same time constructing a QSPR prediction model with the help of an artificial neural network nonlinear regression algorithm. The model has good goodness of fit, robustness and predictive ability, and can quickly and efficiently predict the mass spectrometry ionization efficiency of other compounds in the application domain based on the molecular structure parameters of organic compounds. The establishment and verification of the mass spectrometry ionization efficiency prediction method of the present invention strictly follows the QSPR model development and use guidelines stipulated by the OECD. Therefore, the model constructed by the present invention has high reliability in predicting the ionization efficiency of the target compound, can accurately predict the mass spectrometry ionization efficiency of the target compound, and provides technical support for quantitative detection without standards. It has great promotion and application value and theoretical significance; it also has the following characteristics:

[0054] 1. Based on the OECD guidelines for the construction and use of QSPR models, a QSPR model with transparent algorithms was established, which has good goodness of fit, robustness and predictive ability;

[0055] 2. The model uses COSMO-RS / Turbomole quantum chemistry calculations to collect more characteristic descriptors. These descriptors can, to a certain extent, reflect the influence of the solvent environment on mass spectrometry ionization efficiency. Compared with previous models, this greatly improves prediction accuracy and reduces the workload of additional formula adjustments due to differences in solvent environments.

[0056] 3. The model is fitted using an artificial neural network nonlinear regression algorithm, which enhances the fitting effect and stability compared to the previous simple linear fitting model;

[0057] 4. Compared with existing models, this model has a wider application domain and higher utilization, making it suitable for application and promotion. This method is low-cost, simple and fast, and can save a lot of manpower, material and financial resources.

[0058] 5. Given that the existing model only uses isocratic elution to collect data to establish the log IE model, the model prediction accuracy is low. The present invention uses both isocratic elution and gradient elution to collect data sets for measurement. The purpose is to expand the scope of model application as much as possible and meet the needs of actual production experiments while increasing the sample capacity, while also improving the accuracy of model prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 This is a flow chart of the model prediction method of the present invention;

[0060] Figure 2 This is a fitting diagram of the experimental and predicted values ​​of the data set log IE in Example 1 of the present invention;

[0061] Figure 3 Williams plot was applied in the domain representation of the model of Example 1 of the present application. DETAILED DESCRIPTION

[0062] 1. The compounds involved in the present application include a series of nitrogen-containing compounds such as glycine, alanine, leucine, isoleucine, valine, proline, phenylalanine, methionine, tryptophan, serine, glutamine, threonine, cysteine, asparagine, tyrosine, aspartic acid, glutamic acid, lysine, arginine, histidine, citrulline, taurine, hypotaurine, L-hydroxyproline, spermidine, N-Ω-acetylhistamine, L-pyroglutamic acid, tyramine, histamine, creatinine, L-dihydroorotic acid, L-carnosine, uric acid, betaine, sarcosine, cytosine, hypoxanthine, xanthine, guanine, uracil, adenine, β-thymidine, cytidine, hypoxanthine nucleoside, xanthine nucleoside, guanine nucleoside, uracil nucleoside, adenine nucleoside, pyridoxol, pyridoxal, pyridoxamine, thiamine, biotin, D-pantothenic acid, nicotinic acid, nicotinamide, Harman, DMIP, PhIP, IQ[4,5-b], MeIQ, 7,8-DiMeIQx, 8-MeIQx, MeAαC, IQ, IQx, Phe-p-1, AαC, Norharman and D3-hypoxanthine, wherein D3-hypoxanthine is set as the standard M2 of the present application.

[0063] The specific dilution concentrations of different compounds involved in the ionization efficiency determination and data collection process are 10 -4 6 concentration gradients (including 0.1, 1, 10, 50, 200 and 500 ng / ml) in the range of ~1.0 mg / mL.

[0064] 2. The ionization efficiency determination instrument involved in the present application is a Q-Exactive-HF combined quadrupole-Orbitrap mass spectrometer, and the data acquisition uses Full MS scanning mode under ESI source positive ion mode. The MS and ESI parameters are only optimized by setting the target mass parameter, and the remaining parameters are set as follows: the pressure of the atomizer gas is 50 psi, the dry gas flow is 1 arb, the dry gas temperature is 370℃, and the capillary voltage between MS and atomizer is +3.6 kV.

[0065] 3. The following are the specific condition settings in different isocratic elution and gradient elution conditions in Example 1; data collection is carried out using two mobile phase schemes and two elution forms respectively.

[0066] The specific detection conditions of high performance liquid chromatography are as follows:

[0067] The injection volume is 1 μL;

[0068] Mobile phase: ① A phase is pure water phase, B phase is pure methanol phase; ② A phase is 0.1% formic acid-water solution by volume fraction, B phase is 0.1% formic acid-methanol solution by volume fraction;

[0069] Elution condition:

[0070] ① The proportion of B phase in the liquid phase elution process is adjusted and set to 5%, 20%, 35%, 50%, 65%, 80%, 95% and the like, which is called isocratic elution, and the proportion of the mobile phase is constant during the process, the elution flow rate is 0.3 mL / min, the elution time is 20 min, and the sum of the proportions of A phase and B phase during the whole process is 1; wherein A and B represent the mobile phase schemes ① and ②;

[0071] ② The proportion of B phase in the liquid phase elution process is adjusted according to the method described in Table 1, which is called gradient elution, and the proportion of A phase changes with the change of B phase during the process; wherein A and B represent the mobile phase schemes ① and ②.

[0072] Table 1, specific method of gradient elution

[0073] Time (min) Proportion of mobile phase B (%) Elution flow rate (mL / min) 0 1 0.3 0~1.5 1 0.3 1.5~13 1~99 0.3 13~16.5 99 0.3 16.6~20 1 0.3

[0074] 4. The specific calculation method of relative ionization efficiency is:

[0075] i. In order to eliminate the differences in measurement and model application of instruments and laboratories, the ratio of the analyte (M1) to the standard (M2) is defined as the relative ionization efficiency (RIE M1 ) in the present application, which is calculated according to the following equation:

[0076]

[0077] Wherein, slope refers to the slope value of the quantitative relationship between different dilution concentrations of the analyte M1 or the standard M2 and the mass spectrum peak area under the specified elution condition, which is obtained by linear regression in the linear range of the signal concentration graph. In order to make the data more easily presented and analyzed, the relative ionization efficiency in the present application is presented in logarithmic scale (log IE), and the log IE value measured under the specified elution condition of all analytes needs to be subtracted from the log IE value of M2 measured under the corresponding condition.

[0078] ii. In order to facilitate subsequent calculation of COSMO descriptor, it is necessary to specify the proportion of the mobile phase before entering the mass spectrometer during the determination of the analyte IE, because the ionization process of the analyte in the mass spectrometer is closely related to the solvent environment. According to the specific method adopted in the gradient elution process, since the peak time of the known analyte IE determination is known, the proportion of the mobile phase B before entering the mass spectrometer can be calculated according to the following equation:

[0079] B%=1 (t<=1.5min or t>=16.6min);

[0080] B%=1+(t-1.5)*(99-1) / (13-1.5) (1.5min <t<13min);

[0081] B%=99 (13min <t<16.5min);

[0082] Wherein t and B% represent the peak elution time of the test compound and the ratio of mobile phase B to the peak elution time, respectively.

[0083] Example 1

[0084] A method for predicting the mass spectrometry ionization efficiency of a compound using a quantitative structure-activity relationship model based on COSMO-RS and ANN algorithms comprises the following steps:

[0085] Step 1. Data Collection: A Q-Exactive-HF hybrid quadrupole-Orbitrap mass spectrometer was used to detect the mass spectrometric signals (peak areas) of 70 compounds at a range of concentrations using different isocratic and gradient elution conditions. A quantitative relationship model between concentration and peak area was constructed. Based on the slope of the quantitative relationship model constructed for each compound, 945 sets of compound mass spectrometric relative ionization efficiency data (natural logarithm (log IE) values) were calculated for the 70 compounds under different isocratic and gradient elution conditions. For the same substance, parallel data that significantly deviated from the overall value were eliminated and the average value was taken for model construction.

[0086] After sorting the log IE values ​​obtained above, we selected one data point from every four log IE values ​​to be included in the validation set, and the remaining data points to be included in the training set, that is, the data points were divided into training set and validation set in a ratio of 4:1. The training set data included 756 data points, and the validation set data included 189 data points. The data points in the training set were used to build the model for internal validation, and the data points in the validation set were used for external validation of the model.

[0087] Step 2. Descriptor calculation: The SMILES number of the organic compound was input into Turbomole software. The initial molecular structure of the compound was optimized using the Becke-type three-parameter density functional model (B3LYP) algorithm in density functional theory (DFT) and the 6-311G** basis set. Quantum chemical calculations were performed on all atoms in the molecule at the DFT-BP 86TZVP level, using Turbomole's default convergence criteria.

[0088] The optimized compound molecular structure is further calculated by quantum chemistry software (COSMO-RS) to obtain 45 corresponding molecular structure descriptors, and 10 molecular structure descriptors involved in the pubchem public library are collected, a total of 55, which are subjected to a series of pretreatment processes including removing constants, approaching constants, missing and correlation greater than 0.95.

[0089] On the basis of the training set described in step 1, the final descriptors (40) are screened out by stepwise linear regression as follows:

[0090]

[0091] Step 3, model establishment: based on the data of the training set described in step 1, the natural logarithm value log IE of the relative ionization efficiency data of the compound mass spectrum is the dependent variable, and the final descriptors screened out in step 2 are the independent variables. A kind of artificial neural network (ANN) nonlinear regression algorithm model is established by python program. The optimal parameters of the established QSPR model are selected by 5-fold cross-validation algorithm, and the QSPR model based on the optimal ANN algorithm is constructed;

[0092] The specific process of constructing the QSPR model based on the optimal ANN algorithm is as follows:

[0093] 1) The entire training set described in step 1 is divided into 5 sets, each of which will be rotated as a test set, and the remaining 4 sets will be used as training sets. This is repeated 5 times to ensure that each set is verified once as a test set.

[0094] 2) Adjust different network parameters (network layers, activation functions, etc.), optimize parameters (learning rate, iteration number, etc.), and regularize parameters (dropout ratio, etc.) hyperparameters. Calculate and compare the average cross-validation accuracy of 5 training before and after adjustment, and select the group of parameters with the highest cross-validation accuracy to apply to the QSPR model, that is, to construct the optimal model.

[0095] Step 4, model verification: the model is verified, which is divided into two steps:

[0096] a) Model fitting degree and robustness evaluation;

[0097] b) Application domain representation and performance evaluation of the model;

[0098] The fitting ability of the model is represented by the degree of freedom correction coefficient R 2 adj , the root mean square error RMSE tra of the training set, and the mean absolute error MAE tra of the training set.

[0099] The determination coefficient R of the freedom correction in the embodiment 2 adj = 0.850, the root mean square error RMSE of the training set tra = 0.343, the mean absolute error MAE of the training set tra = 0.264, the smaller the error value, the higher the fitting degree, indicating that the model has good fitting degree and robustness;

[0100] The fitting coefficient R between the prediction and the measurement is used for external verification 2 ext , Q 2 ext and the consistency correlation coefficient CCC represent the external prediction ability of the model.

[0101] The determination basis is: R 2 > 0.7, Q 2 > 0.6, R 2 - Q 2 < 0.3, and CCC > 0.85; in the embodiment, the final model is R 2 ext = 0.841, Q 2 ext = 0.800, and CCC = 0.869, indicating that the model has good external prediction ability, Figure 2 The fitting degree and the verification result of the model are given.

[0102] Step 5, application domain representation: the Williams graph of the lever value h i based on the standard residual δ is used to represent the application domain of the model, specifically: it is generally considered that when the absolute value of δ is greater than 3.0, the compound is an outlier, and when the lever value h i is greater than the warning value h * , it indicates that the structure of the compound is significantly different from other compounds; h i and h * are calculated by the following formula:

[0103] h i = x i T (X T X)-1x i (2)

[0104] h * = 3(p+1) / n (3)

[0105] Where x i is the descriptor matrix of the i th compound; x i T is x iThe transposed matrix of ; X is the descriptor matrix of all compounds; X T is the transposed matrix of X; (X T X) -1 is the matrix X T The inverse of X; p is the number of variables in the model, and n is the number of data points in the dataset.

[0106] like Figure 3 As shown, the model h * is 0.151, and the model is suitable for h i Prediction of compound log IE values ​​less than 0.151;

[0107] Step 6: Model application: Use the model to predict the mass spectrometry ionization efficiency of the target compound in a specified solvent environment.

[0108] Application Example 1

[0109] Given a compound glycine (CAS: 56-40-6), the elution environment is set as: 0.1% formic acid-water solution (phase A) / 0.1% formic acid-methanol solution (phase B) gradient elution, predict its log IE value.

[0110] First, the mobile phase ratio (1% B) of the compound's environment at the time of peak elution was calculated by the peak elution time (0.79 min), and then the SMILES number of glycine (OC(CN)=O) was input into the Turbomole software. The Becke-type 3-parameter density functional model B3LYP algorithm in the density functional theory (DFT) method and the 6-311G** basis set were used to optimize the initial molecular structure of the compound. Quantum chemical calculations were performed on all atoms in the molecule at the DFT-BP 86TZVP level, using the default convergence standard of Turbomole. The optimized molecular structure of the compound was further calculated using the quantum chemical software COSMO-RS to obtain 45 corresponding molecular structure descriptors. In addition, 10 molecular structure descriptors involved in the pubchem public library were collected, totaling 55. After preprocessing, 40 final descriptors were screened out by stepwise linear regression. The h of the substance was obtained according to the calculation formula. i The value is 0.072<0.151, so the compound is within the application domain of the model.

[0111] Substituting the values ​​of the above descriptors into the constructed model, the predicted value of log IE was -0.21, and the experimental value was -0.22, which were close to the experimental value.

[0112] Application Example 2

[0113] Given a compound uric acid (CAS: 69-93-2), the elution environment is set as: 0.1% formic acid-water solution (phase A) / 0.1% formic acid-methanol solution (phase B) gradient elution, and the log IE value thereof is predicted.

[0114] First, the proportion of the mobile phase (4.74% B) at the time of the peak of the compound is calculated by the peak time of the substance (1.88 min), and then the SMILES number of uric acid (O=C(N1)NC2=C1NC(NC2=O)=O) is input into the Turbomole software. The initial molecular structure of the compound is optimized by using the Becke type 3 parameter density functional model B3LYP algorithm in the density functional theory DFT method and the 6-311G** basis set; wherein the quantum chemical calculation of all atoms in the molecule is carried out at the DFT-BP 86TZVP level, and the default convergence standard of Turbomole is used; the optimized molecular structure of the compound is further calculated by using the quantum chemistry software COSMO-RS, 45 kinds of molecular structure descriptors are obtained, in addition, 10 kinds of molecular structure descriptors involved in the pubchem public library are collected, a total of 55 kinds, and 40 final descriptors are screened after pretreatment by stepwise linear regression. According to the calculation formula, the h i value of the substance is 0.067<0.151, so the compound is within the application domain of the model.

[0115] The value of the above descriptor is brought into the built model, and the predicted value of log IE is-0.77, and the experimental value is-0.91, which is close to the predicted value and the experimental value.

[0116] Application Example 3

[0117] Given a compound biotin (CAS: 58-85-5), the elution environment is set as: water (A phase) / methanol (B phase) gradient elution, and the log IE value thereof is predicted.

[0118] First, the proportion of the mobile phase (32.43% B) at the time of the peak of the compound is calculated by the peak time of the substance (5.13 min), and then the SMILES number of biotin (O=C1N[C@]2([H])[C@]([C@H](CCCCC(O)=O)SC2)([H])N1) is input into the Turbomole software, the initial molecular structure of the compound is optimized by using the Becke type 3-parameter density functional model B3LYP algorithm in the density functional theory DFT method and the 6-311G** basis set; wherein the quantum chemical calculation of all atoms in the molecule is carried out at the DFT-BP 86TZVP level, and the default convergence standard of Turbomole is used; the optimized compound molecular structure is further calculated by using quantum chemistry software COSMO-RS, 45 kinds of corresponding molecular structure descriptors are obtained, in addition, 10 kinds of molecular structure descriptors involved in the pubchem public library are collected, a total of 55 kinds, and 40 final descriptors are screened out after pretreatment by stepwise linear regression. According to the calculation formula, the h i value of the substance is 0.096<0.151, so the compound is within the application domain of the model.

[0119] The value of the above descriptor is brought into the established model, and the log IE prediction value is-0.53, and the experimental value is-1.01, which is close to the experimental value.

[0120] Application Example 4

[0121] Given a compound MeAαC (CAS: 68006-83-7), the elution environment is set as: water (phase A) / methanol (phase B) gradient elution, and the log IE value is predicted.

[0122] First, the proportion of the mobile phase (82.03% B) at the time of the peak of the compound is calculated by the peak time of the substance (10.95 min), and then the SMILES number of MeAαC (CC1=CC2=C(NC3=CC=C=C32)N=C1N) is input into the Turbomole software, the initial molecular structure of the compound is optimized by using the Becke type 3-parameter density functional model B3LYP algorithm in the density functional theory DFT method and the 6-311G** basis set; wherein the quantum chemical calculation of all atoms in the molecule is carried out at the DFT-BP 86TZVP level, and the default convergence standard of Turbomole is used; the optimized compound molecular structure is further calculated by using quantum chemistry software COSMO-RS, 45 kinds of corresponding molecular structure descriptors are obtained, in addition, 10 kinds of molecular structure descriptors involved in the pubchem public library are collected, a total of 55 kinds, and 40 final descriptors are screened out after pretreatment by stepwise linear regression. According to the calculation formula, the h iThe value is 0.029 < 0.151, so the compound is within the model application domain.

[0123] The value of the above descriptor is brought into the built model, and the log IE predicted value is 1.08, and the experimental value is 1.68, and the predicted value is close to the experimental value.

[0124] Comparative Example 1

[0125] The data set is only selected from the 812 log values collected in step 1 described by isocratic elution, and after sorting by size, every 4 log IE data is selected 1 data into the validation set, and the remaining data is into the training set, that is, it is divided into training set and validation set according to the ratio of 4:1; The training set data includes 650, and the validation set data includes 162. The data points in the training set are used to construct the model and perform internal validation. The data points in the validation set are used for external validation of the model.

[0126] Using quantum chemistry software (COSMO-RS), 45 molecular structure descriptors of organic compounds are obtained, and 10 molecular structure descriptors involved in the pubchem public library are collected, a total of 55. After a series of preprocessing processes including removing constants, approaching constants, missing and correlation greater than 0.95, on the basis of the above training set, the final descriptor (40) is screened out by stepwise linear regression; Then take the final descriptor as the independent variable, and the natural logarithm value log IE of the relative ionization efficiency data of the compound mass spectrum as the dependent variable, a standard artificial neural network (ANN) nonlinear regression algorithm program model is established by using python program.

[0127] In this comparative example, the degree of freedom corrected determination coefficient R 2 adj = 0.859, the root mean square error RMSE tra = 0.308 of the training set, the mean absolute error MAE tra = 0.244 of the training set; In the external validation of this embodiment, the final model is R 2 ext = 0.723, Q 2 ext = 0.723, CCC = 0.852.

[0128] In this embodiment, only the data set collected by isocratic elution is used to establish the model, and the effect is poor.

[0129] Comparative Example 2

[0130] A total of 45 molecular structure descriptors for organic compounds were obtained using quantum chemistry software (COSMO-RS). After a series of preprocessing steps, including removing constants, near-constants, missing descriptors, and descriptors with correlations greater than 0.95, a final set of 35 descriptors was selected using stepwise linear regression based on the training set described in step 1. A standard artificial neural network (ANN) nonlinear regression algorithm was then developed using a Python program, using the final descriptors as independent variables and the natural logarithm of the relative ionization efficiency (log IE) of the compound mass spectrometry data as the dependent variable.

[0131] The coefficient of determination R for the degree of freedom correction in this comparative example 2 adj =0.841, RMSE of the training set tra =0.336, mean absolute error MAE of the training set tra =0.261; In the external validation of this embodiment, the final model is R 2 ext =0.803, Q 2 ext =0.767, CCC=0.854.

[0132] In this embodiment, only the descriptors calculated by COSMO-RS software are used to build the model, which has a poor effect.

[0133] Comparative Example 3

[0134] Forty-five molecular structure descriptors of organic compounds were obtained using quantum chemistry software (COSMO-RS). An additional ten molecular structure descriptors were collected from the PubChem public library, for a total of 55. Using these 55 descriptors as independent variables and the natural logarithm of the relative ionization efficiency (log IE) of the compound mass spectrometry data as the dependent variable, a standard artificial neural network (ANN) nonlinear regression algorithm was developed using a Python program on the training set.

[0135] The coefficient of determination R for the degree of freedom correction in this comparative example 2 adj =0.850, RMSE of the training set tra =0.325, mean absolute error MAE of the training set tra =0.260; in the external validation of this embodiment, the final model is R 2 ext =0.819, Q 2 ext =0.791, CCC=0.871.

[0136] The effect of directly using the collected 55 descriptors to establish the model in this embodiment is not good.

[0137] Comparative Example 4

[0138] 45 kinds of molecular structure descriptors of organic compounds were obtained by quantum chemistry software (COSMO-RS), and 10 kinds of molecular structure descriptors related to the pubchem public library were collected, totaling 55 kinds. After a series of preprocessing processes including removing constants, near constants, missing, and descriptors with a correlation greater than 0.95.

[0139] On the basis of the training set described in Example 1, step 1, 40 descriptors were screened out by stepwise linear regression; on this basis, 35 descriptors were further screened out by stepwise linear regression; then taking the screened descriptors as independent variables and the natural logarithm value log IE of the relative ionization efficiency data of the compound mass spectrum as dependent variables, a standard artificial neural network (ANN) nonlinear regression algorithm program model was established by a python program.

[0140] In this comparative example, the determination coefficient R of the degree of freedom correction 2 adj = 0.848, the root mean square error RMSE of the training set tra = 0.327, the mean absolute error MAE of the training set tra = 0.257; in the external validation of this embodiment, the final model is R 2 ext = 0.805, Q 2 ext = 0.783, and CCC = 0.870.

[0141] The effect of finally screening 35 descriptors to establish the model in this embodiment is not good.

[0142] Comparative Example 5

[0143] 45 kinds of molecular structure descriptors of organic compounds were obtained by quantum chemistry software (COSMO-RS), and 10 kinds of molecular structure descriptors related to the pubchem public library were collected, totaling 55 kinds. After a series of preprocessing processes including removing constants, near constants, missing, and descriptors with a correlation greater than 0.95.

[0144] On the basis of the training set described in Example 1 Step 1, the final descriptors (40 kinds) were screened out by stepwise linear regression; then taking the final descriptors as the independent variable, the natural logarithm value log IE of the relative ionization efficiency data of the compound mass spectrum as the dependent variable, a standard artificial neural network (ANN) nonlinear regression algorithm program model was established for the training set by a python program.

[0145] In this comparative example, the determination coefficient R 2 adj = 0.853, the root mean square error RMSE tra = 0.334, the mean absolute error MAE tra = 0.266; in the external validation of this example, the final model is R 2 ext = 0.804, Q 2 ext = 0.783, CCC = 0.863.

[0146] In this example, 1421 new Mordred descriptors were added to establish the model, and the effect was poor, so this type of descriptor was not used in the final application example.

[0147] Comparative Example 6

[0148] 45 kinds of molecular structure descriptors of organic compounds were obtained by quantum chemistry software (COSMO-RS), and 10 kinds of molecular structure descriptors involved in the pubchem public library were collected, a total of 55 kinds. After a series of preprocessing processes including removing constants, near constants, missing and correlation greater than 0.95, on the basis of the training set described in Step 1, 40 kinds of descriptors were randomly selected; then taking the selected descriptors as the independent variable, the natural logarithm value log IE of the relative ionization efficiency data of the compound mass spectrum as the dependent variable, a standard artificial neural network (ANN) nonlinear regression algorithm program model was established for the training set by a python program.

[0149] In this comparative example, the determination coefficient R 2 adj = 0.853, the root mean square error RMSE tra = 0.322, the mean absolute error MAE tra = 0.257; in the external validation of this example, the final model is R 2 ext = 0.805, Q 2 ext = 0.777, CCC = 0.862.

[0150] In this embodiment, the effect of randomly screening descriptors to build a model is not good.

[0151] Comparative Example 7

[0152] Quantum chemistry software (COSMO-RS) was used to obtain 45 molecular structure descriptors of organic compounds. In addition, 10 molecular structure descriptors involved in the pubchem public library were collected, totaling 55. After a series of preprocessing processes, including removing constants, near constants, missing descriptors, and descriptors with correlation greater than 0.95, the descriptors were analyzed.

[0153] Based on the training set described in step 1 of Example 1, the final descriptors (40 types) were screened out by stepwise linear regression; then, a multiple linear regression algorithm program model was established using a Python program with the final descriptors as independent variables and the natural logarithm of the relative ionization efficiency data of the compound mass spectrum as the dependent variable.

[0154] The coefficient of determination R for the degree of freedom correction in this comparative example 2 adj =0.801, RMSE of the training set tra =0.376, mean absolute error MAE of the training set tra =0.304; in the external validation of this embodiment, the final model is R 2 ext =0.771, Q 2 ext =0.761, CCC=0.865. In this embodiment, the effect of using the multiple linear regression algorithm to establish the model is poor, and it can be optimized by replacing it with a nonlinear regression algorithm.

[0155] The present invention is not limited to the above-mentioned embodiments. On the basis of the technical solutions disclosed in the present invention, those skilled in the art can make some substitutions and modifications to some of the technical features therein according to the disclosed technical content without creative labor, and these substitutions and modifications are all within the protection scope of the present invention.

Claims

1. A method for predicting the mass spectrometry ionization efficiency of a compound based on a quantitative structure-activity relationship model of COSMO-RS and ANN algorithm, characterized in that: The method comprises the following steps: Step 1) Data Collection: Peak areas corresponding to a series of concentrations of several compounds under different isocratic and gradient elution conditions are collected using a high performance liquid mass spectrometer. A quantitative relationship model between the concentration and peak area of ​​each compound is constructed. Based on the slope of the constructed quantitative relationship model, the relative ionization efficiency (log IE) of the compound mass spectrometry under different elution conditions is calculated for each compound. After sorting the log IE values ​​obtained above, we select one data point from every four log IE values ​​and put it into the validation set. The remaining data points are then put into the training set, i.e., the data points are divided into the training set and the validation set in a ratio of 4:

1. The data points in the training set are used to build the model for internal validation, and the data points in the validation set are used for external validation of the model. Step 2) Descriptor calculation: The SMILES number of the organic compound was input into the Turbomole software, and the initial molecular structure of the compound was optimized using the Becke-type three-parameter density functional model B3LYP algorithm in the density functional theory (DFT) method and the 6-311G** basis set. Quantum chemical calculations were performed on all atoms in the molecule at the DFT-BP 86TZVP level, using the default convergence criteria of Turbomole. The optimized molecular structure of the compound was further calculated using the quantum chemistry software COSMO-RS to obtain 45 corresponding molecular structure descriptors. In addition, 10 molecular structure descriptors involved in the pubchem public library were collected, for a total of 55 descriptors. After preprocessing, the final descriptors were screened using stepwise linear regression. Step 3) Model Establishment: Based on the training set data described in step 1), the natural logarithm of the relative ionization efficiency (log IE) of the compound mass spectrometry data is used as the dependent variable, and the final descriptor selected in step 2) is used as the independent variable. An artificial neural network (ANN) nonlinear regression algorithm model is established using a Python program. The optimized parameters of the established QSPR model are selected using a 5-fold cross-validation algorithm to construct a QSPR model based on the optimal ANN algorithm. Step 4) Model verification: Verify the model in two steps: a) Evaluation of the model's goodness of fit and robustness; b) Characterize the application domain and evaluate the performance of the model; after verification, proceed to step 5); Step 5) Application domain characterization: Characterize the model application domain through Williams diagram; Step 6) Model application: using the model to predict the mass spectrometry ionization efficiency of the target compound in a specified solvent environment.

2. The method according to claim 1, characterized in that The compounds in step 1) include glycine, alanine, leucine, isoleucine, valine, proline, phenylalanine, methionine, tryptophan, serine, glutamine, threonine, cysteine, asparagine, tyrosine, aspartic acid, glutamic acid, lysine, arginine, histidine, citrulline, taurine, hypotaurine, L-hydroxyproline, spermidine, N-Ω-acetylhistamine, L-pyroglutamic acid, tyramine, histamine, creatinine, L-dihydroorotic acid, L-carnosine, uric acid, betaine, sarcosine, cytosine, hypoxanthine, xanthine, guanine one or more of: xanthine, uracil, adenine, β-thymidine, cytidine, inosine, xanthine, guanosine, uridine, adenine, pyridoxine, pyridoxal, pyridoxamine, thiamine, biotin, D-pantothenic acid, niacin, nicotinamide, Harman, DMIP, PhIP, IQ[4,5-b], MeIQ, 7,8-DiMeIQx, 8-MeIQx, MeAαC, IQ, IQx, Phe-p-1, AαC, Norharman, and D3-hypoxanthine.

3. The method according to claim 1, characterized in that Step 1) Different isocratic elution and gradient elution conditions: Isocratic elution: 5% B, 20% B, 35% B, 50% B, 65% B, 80% B, 95% B, the mobile phase ratio remains constant during the process, the elution flow rate is 0.3 mL / min; the elution time is 20 min, and the total ratio of phase A to phase B during the entire process is 1; Gradient elution: 1% B, 0 min; 1% B, 0-1.5 min; 1-99% B, 1.5-13; 99% B, 13-16.5 min; 1% B, 16.6-20.0 min; 0.3 mL / min; Phase A is pure water phase; phase B is pure methanol phase or phase A is 0.1% by volume formic acid-water solution; phase B is 0.1% by volume formic acid-methanol solution or both.

4. The method according to claim 1, wherein Step 2) The preprocessing process includes removing constant, near-constant, missing and descriptors with correlation greater than 0.

95.

5. The method according to claim 1, wherein The final descriptors screened out in step 2) include molecular size descriptors, polarity descriptors, basicity descriptors, and solubility descriptors.

6. The method according to claim 5, characterized in that The solubility descriptors are related to the properties of the compound itself, the solvent environment and the temperature, including H-Bond interaction energy, gas-liquid transition energy, total average interaction energy, mismatch interaction energy, chemical potential energy, etc. * 、H_vdW、log(H) * 、Gsolv * , log value of gas partial pressure*, molecular free energy * , Henry's law coefficient * , Gibbs free energy in solvent * , ln(γ), log(pv), solvent density * , solvent molar mass and relative solubility.

7. The method according to claim 1, characterized in that Step 3) The construction of the QSPR model based on the optimal ANN algorithm specifically includes the following process: (1) The entire data set is divided into 5 sets. Each set is used as a test set in turn, and the remaining sets are used as training sets. This is repeated 5 times to ensure that each set is verified once as a test set. (2) Debug different network parameters, optimization parameters, and regularization parameters, calculate and compare the average cross-validation accuracy of 5 training sessions before and after the adjustment, and select the set of parameters with the highest cross-validation accuracy to apply to the QSPR model to construct the optimal model.

8. The method according to claim 1, characterized in that When the model is verified in step 4), the goodness of fit and robustness evaluation indicators are: the coefficient of determination R after degree of freedom correction 2 adj , training set root mean square error RMSE tra And the training set mean absolute error MAE tra .

9. The method according to claim 1, characterized in that Step 5) The model application domain is characterized by the Williams diagram, specifically by using the leverage value h based on the standard residual δ i The Williams diagram characterizes the application domain of the model. When the absolute value of δ is greater than 3.0, the compound is an outlier. When the leverage value h i Greater than the warning value h * When , it indicates that the structure of this compound is significantly different from that of other compounds; i and h * Calculated by the following formula: h i =x i T (X T X) -1 x i h * =3(p+1) / n where x i is the descriptor matrix of the ith compound; x i T is x i The transposed matrix of ; X is the descriptor matrix of all compounds; X T is the transposed matrix of X; (X T X) -1 is the matrix X T The inverse of X; p is the number of variables in the model, and n is the number of data points in the dataset.

10. Use of the method according to any one of claims 1 to 9 in predicting mass spectrometry ionization efficiency of compounds.

Citation Information

Patent Citations

  • Chiral drug mass spectrometry quantitative analysis method based on chemical derivatization reaction and spectral deformation quantitative analysis theory

    CN108469467A

  • Method for predicting PA-water distribution coefficient of organic pollutant based on quantitative structural property relationship

    CN112086141A