A method for constructing a multi-class lipid retention time prediction universal model, a universal prediction model and a prediction system
By using multivariate polynomial nonlinear regression modeling, a multi-class lipid retention time prediction model was constructed, which solved the problems of high coverage and high accuracy in the prediction of lipid compound retention time in existing technologies, and realized efficient qualitative and quantitative analysis of lipid compounds.
Patent Information
- Application Number
- CN202510776242.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Existing technologies struggle to construct high-coverage, high-precision retention time prediction models for lipid compounds, resulting in insufficient utilization of mass spectrometry detection resources and an inability to meet the qualitative and quantitative analysis needs of various lipid compounds.
Multivariate polynomial nonlinear regression modeling was adopted, and characteristic structural parameters of lipid compounds, such as the total number of carbons and the number of carbon-carbon double bonds, were used to construct a general prediction model for the retention time of lipids in multiple categories. This model includes a general prediction model for five major lipid categories and a general prediction model for a single subclass, with the expression tRC=k0+K1c3+K2c2+K3c+K4.
It achieves efficient and high-coverage retention time prediction of lipid compounds, and can set a precise MRM data acquisition window for unknown lipid compounds, thereby improving the sensitivity and coverage of mass spectrometry detection.
Smart Images

Figure CN120913704B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to biological detection technology, more particularly to a method for constructing a universal model for predicting retention time of multiple categories of lipids, a universal prediction model and a prediction system. BACKGROUND
[0002] According to the lipid classification platform established by the Lipid Metabolites And Pathways Strategy (LIPID MAPS), lipids can be classified into eight categories, including free fatty acids and their modified compounds (such as acyl carnitines, hydroxy fatty acids and dicarboxylic acids, FAs), glycerolipids (GLs), glycerophospholipids (GPs), sphingolipids (SPs), sterol lipids (STs), prenol lipids (PRs), sugar lipids (SLs) and polyketides (PKs). According to the combination and function of different structural units inside lipids, lipids can be further classified into different subcategories of lipid compounds (theoretically containing more than 40,000 kinds of lipid compounds). Although lipidomics detection technology has become an important tool for studying the metabolic change rules of organisms under different physiological or pathological conditions, it is used to find various lipid biomarkers to diagnose or predict ischemic stroke, coronary sclerosis, lung cancer and other chronic cardiovascular diseases and other major diseases. However, the number of lipid compounds in organisms is large, and the corresponding standard product coverage and number are very small, so the current high-coverage qualitative and quantitative analysis of lipidomics is facing great challenges.
[0003] The existing qualitative analysis of the structure of lipid compounds is mainly based on non-targeted lipidomics analysis methods and targeted lipidomics analysis methods, mainly involving mass spectrometry data and retention time t R Data matching, etc. Due to the existence of many isomers and isobars in lipidomics, the qualitative false positive rate of mass spectrometry data, especially the secondary mass spectrometry spectrum (MS 2 ), is high. Since t R can effectively verify and improve the accuracy of mass spectrometry, the acquisition of t R data of endogenous or stable isotope-labeled lipid compound standards is one of the most important and ideal qualitative analysis strategies. In particular, in the process of high-sensitivity qualitative and quantitative analysis based on multiple reaction monitoring (MRM) scanning, known or high-precision predicted t R value is one of the important mass spectrometry parameters. For example, when the acquisition window is only 1 min, the deviation between the theoretical t R value of the lipid compound and the true value should be less than 0.5 min. However, the number of commercially available lipid standards is limited (about 60-70 kinds of commercial lipid standards), so there are still a large number of lipid compounds that lack reliable t RThe values are used for MRM qualitative and quantitative analysis. When the standard is absent, the t R values of the target lipid compounds can be constructed according to the quantitative structure-retention relationship (QSRR) based on the quantitative structure-retention relationship (QSRR) of the standard. R For example, when the total number of carbon (c) in the fatty acyl chain, ether chain, alkenyl ether chain or / and sphingosine skeleton is constant, the t R values of the lipid compounds of the same lipid subclass decrease with the increase of the number of double bonds (d) and meet the equation t R = k1d + k0 (k1 represents a coefficient and k0 is a constant). Similarly, when the number of double bonds is the same, the t R values of the lipid compounds of the same lipid subclass increase with the increase of the total number of carbon and meet the equation t 3 = k1c 2 + k2c R + k3c + k0 (k1-k3 represent coefficients and k0 is a constant).
[0004] However, the above multivariate polynomial regression algorithm needs to construct at least two QSRR prediction models for each lipid subclass, and the standard of the lipid compounds of many subclasses is less than 3 for constructing the QSRR prediction model, so it cannot meet the current high coverage and high precision t R prediction of lipidomics. Therefore, under the premise of limited standard of lipid compounds, how to construct a high-precision t R general prediction model for one or more categories of lipid subclasses is a difficult problem to be solved at present, so as to effectively predict the retention time of other lipid compounds not detected in each subclass for accurate setting of the MRM acquisition window, facilitate the full and reasonable use of mass spectrometry detection resources, and improve the sensitivity, stability and coverage. SUMMARY
[0005] In view of the above defects of the prior art, the present application provides a construction method and a prediction system of a general prediction model for the retention time of multiple categories of lipids. The present application takes the experimental retention time t R E as the dependent variable, takes the total number of carbon c, the number of carbon-carbon double bonds d in the fatty acyl chain, ether chain, alkenyl ether chain or / and sphingosine skeleton, and the screened 36 characteristic structural parameters as the independent variables, and quantitatively processes these characteristic structural parameters, to obtain a general prediction model for predicting five categories of lipids by using multivariate polynomial nonlinear regression modeling, and the expression is t R C = k0 + K1c 3 + K2c 2+K3c+K4, where k0 is a constant, and K1 to K4 are not constants, comprising different functions composed of various screened feature structure parameters. Simultaneously, this invention also employs regression modeling to obtain a general prediction model for the retention time of lipid compounds in a single subclass, the expression of which is t R C =k0+K1c 3 +K2c 2 +K3c+K4, where k0 is a constant, and K1~K4 are composed of d, d 2 With d 3 Different functions are composed of these components. Therefore, compared with traditional multivariate multinomial regression algorithms or machine learning algorithms, the new general prediction model constructed in this invention will effectively solve the problem of insufficient number of existing lipid compound standards or the number of identified lipid compounds, thereby satisfying the requirement of lipid compound t in each subclass. R High-precision prediction.
[0006] This invention provides a method for constructing a universal prediction model for the retention time of multiple classes of lipid compounds based on high-performance liquid chromatography-mass spectrometry (HPLC-MS), including using lipid compounds t in a training set. R E Using the characteristic structural parameters as the independent variables, the steps of regression modeling analysis are performed after quantifying the independent variables. The characteristic structural parameters include the total number of carbons (c) and the number of carbon-carbon double bonds (d) of the fatty acyl chain, ether chain, alkenyl ether chain, and / or sphingosine skeleton. The lipid compounds in the training set include one or more lipid classes such as free fatty acids and their modified compounds (e.g., acylcarnitine, hydroxy fatty acids, and dicarboxylic acids), glycerides, glycerophospholipids, sphingolipids, and sterol lipids. The characteristic structural parameters also include the types and numbers of skeletons contained in the lipid compounds, and the types and numbers of characteristic groups and residues in the lipid compounds. The characteristic groups are modifying groups or compound structures used to identify different lipid compounds. The modifying groups include fatty acylation modifications, peroxidation modifications, hydroxylation modifications, and / or glycosylation modifications. The parameters of the skeletons, characteristic groups, or residues are quantified according to the number of corresponding types of skeletons, characteristic groups, or residues contained in the lipid compounds. If none are present, the parameter is defined as 0. Regression modeling is used to obtain the t-values of multiple categories of lipid compounds. R The general prediction model is expressed as t R C =k0+K1c 3 +K2c 2 +K3c+K4, where k0 is a constant, and K1 to K4 include independent variables of various characteristic structural parameters other than the independent variable c. Simultaneously, this invention also employs regression modeling to obtain a general prediction model for the retention time of lipid compounds in a single subclass, the expression of which is t RC = k0 + K1c 3 + K2c 2 + K3c + K4, where k0 is a constant, K1 to K4 are functions of d, d 2 and d 3 consisting of different functions.
[0007] Preferably, the method, the lipid compounds in the training set include five major classes of lipids, free fatty acids and their modified compounds (such as acyl carnitines, hydroxylated fatty acids and dicarboxylic acids), glycerolipids, glycerophospholipids, sphingolipids and sterol lipids, the skeleton includes glycerol skeleton v, cholesterol skeleton L, sphingosine skeleton μ, dihydrosphingosine skeleton η, 22,23-ergostanol skeleton ε, brassicasterol skeleton θ, 24α-ethylcholesterol skeleton ξ, stigmasterol skeleton π, 3β-cholestanol skeleton ω and 24β-ethylcholesterol skeleton λ;
[0008] The characteristic group includes fatty acylated modification of hydroxylated fatty acid or hexose a, peroxidation modification group z, hydroxylated modification group h, phosphate p, glycosylation w, ether chain j, olefin ether chain f, glucuronic acid g, trimethylhomoserine t, dimethyl ethanolamine r, sialic acid y, ethanolamine e, choline q, inositol serine s, methylated ethanol ammonium x and ethanol u;
[0009] The residue includes carnitine residue m, carboxylic acid residue b, sulfonyl and sulfonylated galactose residue ψ.
[0010] Preferably, in the training set, the lipid compounds contain hydroxylated modification group h on the fatty acyl chain, the independent variable further includes the hydroxylated modification position n, that is, the carbon atom at the n position of the fatty acyl chain from the carboxyl end is connected with the hydroxylated modification group; if the lipid compound does not contain hydroxylated modification group h on the fatty acyl chain, h is defined as 0, n is defined as 0.
[0011] Preferably, the method, when the lipid compounds in the training set contain glycerol skeleton, the independent variable further includes α, β and γ corresponding to the fatty acyl chain, ether chain or olefin ether chain connection positions sn-1, sn-2 and sn-3 of the glycerol skeleton, respectively, the parameters α, β and γ are quantified according to the number of fatty acyl chain, ether chain or olefin ether chain connected at the sn-1, sn-2 and sn-3 positions of the glycerol skeleton.
[0012] Preferably, the method, when the lipid compounds in the training set contain sphingosine skeleton, the independent variable further includes sphingosine skeleton hydroxylation δ, the parameter δ is quantified according to the number of sphingosine skeleton hydroxylation in the lipid compound.
[0013] Preferably, the method, the t R K1 in the general prediction model is k1+k2d+k3d
[0014]
[0015] wherein k i is the equation coefficient corresponding to the QSRR model constructed by the multiple regression analysis.
[0016] For each subcategory of the five categories of lipid compounds, the above K1-K4 are simplified into new functions consisting of d, d 2 and d 3 .
[0017] In addition, the present application also provides a t R General prediction model, wherein K1 in the general prediction model is k1+k2d+k3d 2 +k4d 3 +k5h+k6m+k7b.
[0018] K2 is k8+k9d+k 10 d 2 +k 11 d 3 +k 12 h+k 13 m+k 14 n+k 15 b.
[0019] K3 is k 16 +k 17 d+k 18 d 2 +k 19 d 3 +k 20 h+k 21 m+k 22 n.
[0020] K4 is k 23 d 3 +k 24 h+k 25 m+k 26 n+k 27 b.
[0021] A general prediction model for the retention time of glycerolipid compounds constructed by the method of the present application is also provided, wherein K1 in the general prediction model is k1+k2d+k3d2 +k4d 3 +k5g+k6j+k7w+k8β+k9ψ;
[0022] K2 is k 10 +k 11 d+k 12 d 2 +k 13 d 3 +k 14 g+k 15 j+k 16 t+k 17 w+k 18 α+k 19 β+k 20 γ+k 21 ψ;
[0023] K3 is k 22 +k 23 d+k 24 d 2 +k 25 d 3 +k 26 g+k 27 h+k 28 j+k 29 n+k 30 t+k 31 w+k 32 z+k 33 α+k 34 β+k 35 γ+k 36 ψ;
[0024] K4 is k 37 d+k 38 d 2 +k 39 d 3 +k 40 g+k 41 h+k 42 j+k 43 n+k 44 t+k 45 w+k 46 z+k 47 β+k 48 γ+k 49 ψ.
[0025] Also provided is a glycerophospholipid-like lipid compound t R K1 is k
[0026] K2 is
[0027] K3 is k
[0028] K4 is k
[0029] Also provided is a sphingolipid lipid compound t R K1 is k1+k2d+k3d 2 +k4d 3 +k5w+k6μ+k7ψ;
[0030] K2 is k8+k9a+k 10 d+k 11 d 2 +k 12 d 3 +k 13 h+k 14 p+k 15 q+k 16 w+k 17 y+k 18 δ+k 19 μ+k 20 ψ;
[0031] K3 is k 21 +k 22 a+k 23 d+k 24 d 2 +k 25 d 3 +k 26 h+k 27 p+k 28 q+k 29 δ+k 30 η+k 31 μ+k 32 ψ;
[0032] K4 is k 33 a+k 34 d+k 35 d 2 +k 36 e+k 37 p+k 38 q+k 39 w+k 40 y+k 41 δ+k 42 μ+k 43 ψ.
[0033] Also provided is a solid sterol lipid compound t R The general prediction model has K1=k1+k2d+k3d 2 K2=k4+k5d+k6d 2 +k7d 3 +k8L; K3=k9+k 10 d+k 11 d 2 +k 12 d 3 +k 13 ε+k 14 L+k 15 a+k 16 ξ;
[0034] K4=k10+k
[0035] The present application also provides a multi-class lipid general compound retention time prediction system, which comprises a data acquisition module and a retention time prediction module.
[0036] The data acquisition module is used to quantitatively acquire the characteristic structural parameters of the lipid compounds to be identified according to the method of the present application, and submit them to the retention time prediction module.
[0037] The retention time prediction module is provided with a general prediction model of the retention time of the lipid compounds obtained according to the method of the present application or the general prediction model of the retention time according to the present application, which is used to calculate the theoretical retention time according to the general prediction model expression according to the quantitative data of the acquired characteristic structural parameters of the lipid compounds, and output the prediction result.
[0038] Overall, the above technical solutions conceived by the present application are superior to the prior art method. The construction method of the retention time prediction model of the lipid compounds provided by the present application can achieve the following beneficial effects:
[0039] (1) High efficiency and high coverage: compared with the prior art which can only establish a prediction model for each sub-class of lipids, the method provided by the present application can construct a general prediction model for five classes of t R The general prediction model has high efficiency and high coverage advantage.
[0040] (2) Accurate detection of unknown lipid compounds: the general prediction model of the retention time of the lipid compounds constructed by the method can establish a high-precision theoretical database of the t R Retention time of the lipid compounds, which can set a more accurate MRM data acquisition window for unknown lipid compounds, so that the qualitative and quantitative analysis of the lipid compounds without standard samples can be performed. BRIEF DESCRIPTION OF DRAWINGS
[0041] Figure 1 A is the prediction ability of QSRR model constructed when each of the training set contains 80% of the number of FFA (A), ACar (B), DAG (C), and TAG (D) lipid compounds, B is the prediction ability of QSRR model constructed when each of the training set contains 80% of the number of FFA (A), ACar (B), DAG (C), and TAG (D) lipid compounds, C is the prediction ability of QSRR model constructed when each of the training set contains 80% of the number of FFA (A), ACar (B), DAG (C), and TAG (D) lipid compounds, D is the prediction ability of QSRR model constructed when each of the training set contains 80% of the number of FFA (A), ACar (B), DAG (C), and TAG (D) lipid compounds, and E is the prediction ability of QSRR model constructed when each of the training set contains 80% of the number of FFA (A), ACar (B), DAG (C), and TAG (D) lipid compounds;
[0042] Figure 2 A is the prediction ability of QSRR model constructed when each of the training set contains 80% of the number of FFA (A), ACar (B), DAG (C), and TAG (D) lipid compounds, B is the prediction ability of QSRR model constructed when each of the training set contains 80% of the number of FFA (A), ACar (B), DAG (C), and TAG (D) lipid compounds, C is the prediction ability of QSRR model constructed when each of the training set contains 80% of the number of FFA (A), ACar (B), DAG (C), and TAG (D) lipid compounds, D is the prediction ability of QSRR model constructed when each of the training set contains 80% of the number of FFA (A), ACar (B), DAG (C), and TAG (D) lipid compounds, and E is the prediction ability of QSRR model constructed when each of the training set contains 80% of the number of FFA (A), ACar (B), DAG (C), and TAG (D) lipid compounds;
[0043] Figure 3 A is the prediction ability of QSRR model I constructed when the training set contains 29 FFA, 14 ACar, 10 2HOFA, 3 3HOFA, 5 DCA, and 1 FAHFA lipid compounds; and B is the prediction ability of QSRR model II constructed when the training set contains 33 FFA, and each contains 1 ACar, 1 HOFA, 1 HOFA, and 1 DCA lipid compound;
[0044] Figure 4 The number of different lipid subclasses detected in typical biological matrices under chromatographic conditions (M1) is shown, respectively;
[0045] Figure 5 The correlation and deviation between the predicted theoretical retention time and the experimentally determined t R
[0046] Figure 6 The correlation analysis results of the experimentally determined retention time and the corresponding QSRR calculated retention time of different lipid compounds in the five major categories of lipids under reversed-phase chromatographic conditions M2 are shown in the figure, wherein A to E represent free fatty acids and their modified compounds, glycerolipids, glycerophospholipids, sphingolipids, and sterol lipids, respectively. DETAILED DESCRIPTION
[0047] To further illustrate the method for efficiently, high coverage, high precision prediction of the retention time of different lipid compounds, the specific embodiments, structures, features and effects of the present application are described in detail below in combination with the drawings and specific implementation cases.
[0048] The existing lipid compound retention time prediction model is only applicable to the same subclass, and it is difficult to realize rapid and high coverage lipid quantitative detection. At the same time, when only one lipid compound (or one lipid compound standard) is contained in the subclass, the current method is not applicable to the construction of QSRR mathematical model. In order to overcome the shortcomings of the prior art, the present application aims to construct a general model for predicting the retention time of multiple categories of lipids.
[0049] The present application takes five categories (free fatty acids and their modified compounds, glycerolipids, glycerophospholipids, sphingolipids and sterol lipids) of lipids identified in biological matrix as the training set, and finally selects 38 characteristic structure parameters of compounds through experimental screening. The experimental retention time of the lipid compounds in the training set is used as the dependent variable, and the 38 characteristic structure parameters selected are used as the independent variable. After quantifying the independent variable, the general model for predicting the retention time of the five categories of lipid compounds can be obtained by regression modeling analysis, and its expression is t R C =k0+K1c 3 +K2c 2 +K3c+K4, wherein k0 is a constant, and K1-K4 include a function composed of various characteristic structure parameters other than the independent variable c. At the same time, the present application also obtains a general prediction model for the retention time of lipid compounds in a single subclass by regression modeling, and its expression is t R C =k0+K1c 3 +K2c 2 +K3c+K4, wherein k0 is a constant, and K1-K4 are different functions composed of d, d 2 and d 3 , and other independent variables are fixed values.
[0050] The characteristic structure parameters of the screened lipid compounds include the total number of carbon atoms (c), the number of double bonds (d), the number of the skeleton types and the number of the corresponding skeletons, the characteristic groups and the number of the corresponding residues in the lipid compounds. The skeletons include glycerol skeleton v, cholesterol skeleton L, sphingosine skeleton μ, dihydrosphingosine skeleton η, 22,23-ergostanol skeleton ε, brassicasterol skeleton θ, 24α-ethylcholesterol skeleton ξ, stigmasterol skeleton π, 3β-cholestanol skeleton ω and 24β-ethylcholesterol skeleton λ; the characteristic groups include fatty acylated modification of hydroxylated fatty acid or hexose a, peroxidation modification group z, hydroxylated modification group h and hydroxylated modification position n, phosphate p, glycosylation w, ether chain j, alkenyl ether chain f, glucuronic acid g, trimethylhomoserine t, dimethylethanolamine r, sialic acid y, ethanolamine e, choline q, inositol serine s, methylated ethanol ammonium x, and ethanol u; the residues include carnitine residue m, sphingosine skeleton hydroxylation δ, carboxylic acid residue b, sulfonyl and sulfonated galactose residue ψ; and the parameters α, β and γ corresponding to the positions sn-1, sn-2 and sn-3 of the fatty acyl chain, ether chain or alkenyl ether chain on the glycerol skeleton. The independent variables are quantitatively processed as follows:
[0051] The total number of carbon atoms c is quantified according to the total number of carbon atoms in the fatty acyl chain, sphingosine, ether chain and alkenyl ether chain in the lipid compound; the number of carbon-carbon double bonds d is quantified according to the total number of carbon-carbon double bonds in the fatty acyl chain, sphingosine, ether chain and alkenyl ether chain in the lipid compound; the skeleton, characteristic group or residue is quantified according to the number of the corresponding type of skeleton, characteristic group or residue in the lipid compound, and the parameter of the corresponding type of skeleton, characteristic group or residue is quantified by a numerical value, or the parameter is defined as 0 if the corresponding type of skeleton, characteristic group or residue is not contained.
[0052] For example, in the training set, if a lipid compound contains carnitine, m = 1 is defined; if it does not contain carnitine, m = 0 is defined. If the fatty acyl chain of a lipid compound contains a hydroxylated modification group, h = 1 is defined, and n is quantified according to the carbon position of the hydroxylated modification (e.g., if the second carbon at the carboxyl terminus of the fatty acid is hydroxylated, n = 2 is defined). If the fatty acyl chain is not hydroxylated, h = 0 and n = 0 are defined. If the fatty acyl chain of a lipid compound contains a carboxylic acid residue at its end, b = 1 is defined; if it does not contain a carboxylic acid residue, b = 0 is defined. If the fatty acyl chain or hexose contains... The hydroxyl group is esterified by another fatty acyl chain, and a = 1 is defined. If there is no esterification, a = 0 is defined. If the lipid compound contains a glycerol backbone (v), a 22,23-ergosterol backbone (ε), a campesterol backbone (θ), a 24α-ethylcholesterol backbone (ξ), a stigmasterol backbone (π), a 3β-cholesterol backbone ω, or a 24β-ethylcholesterol backbone λ, the corresponding independent variable parameter can be quantified according to the number of the corresponding backbone. If it does not contain the corresponding backbone, the independent variable parameter is defined as 0. For example, if it contains a glycerol backbone (v), v is defined as 1. If it does not contain a glycerol backbone, v is defined as 0.
[0053] Additionally, if the lipid compounds in the training set contain a glycerol backbone (v), the parameters α, β, and γ are quantified according to the total number of fatty acyl chains, ether chains, or alkenyl ether chains connected to the sn-1, sn-2, and sn-3 positions of the glycerol backbone. For example, if only a fatty acyl chain is connected to sn-1 on the glycerol backbone, α = 1, β = 0, and γ = 0 are defined; if only one fatty acyl chain is connected to sn-1 and sn-2 on the glycerol backbone, α = 1, β = 1, and γ = 0 are defined; when there are two glycerol backbones and one fatty acyl chain is connected to sn-1 and sn-2 on each glycerol backbone, α = 2, β = 2, and γ = 0 are defined.
[0054] If the lipid compounds in the training set contain a sphingosine skeleton, the independent variable also includes sphingosine skeleton hydroxylation, i.e., phytosphingosine δ. The parameter δ is numerically quantified according to the number of sphingosine skeleton hydroxylations in the lipid compound. For example, if a lipid compound contains one sphingosine skeleton hydroxylation, i.e., phytosphingosine, then the sphingosine skeleton μ = 1, and δ = 1. If the lipid compound contains alkenyl ether chain (f), ether chain (j), choline (q), dimethylethanolamine (r), ethanol (u), ethanolamine (e), glucuronic acid (g), glycosylation (w), serine (s), or inositol... Methylated ethanolamine (x), sialic acid (y), peroxidation modification (z), phosphoric acid (p), sulfonated galactose residues (ψ), sulfonyl group or trimethylhomoserine (t), quantified according to the number of corresponding structures, such as containing one enol ether chain (f) or ether chain (j), defining f or j = 1, and defining its argument parameter as 0 if not containing the corresponding structure. Other arguments are quantified in the same or similar manner. After the arguments are quantified, variable screening is completed by stepwise regression and full subset regression, and the optimal general prediction model is selected according to the adjusted goodness of fit (Goodness of fit, R 2 ) and prediction error by 10-fold cross-validation.
[0055] Based on this, the present application proposes a method for constructing a general prediction model of retention time of multi-class lipid compounds, which is based on high performance liquid chromatography-mass spectrometry technology, comprising the steps of taking the experimental retention time of lipid compounds in the training set as the dependent variable, and taking the characteristic structure parameters as the independent variables, quantifying the independent variables, and then using regression modeling analysis, wherein the characteristic structure parameters include the total number of carbons c, the number of carbon-carbon double bonds d;
[0056] The lipid compounds in the training set include one or more lipid superclasses in free fatty acids and their modified compounds (such as acyl carnitine, hydroxy fatty acid and dicarboxylic acid), glycerolipids, glycerophospholipids, sphingolipids and sterol lipids, and the independent variables further include the number of skeleton types and corresponding skeletons contained in the lipid compounds, characteristic groups and residue types in the lipid compounds, and the number of corresponding residues; the characteristic group is a modified group or compound structure used to identify different lipid compounds;
[0057] Wherein the skeleton, characteristic group or residue is quantified according to the number of corresponding skeleton, characteristic group or residue if the lipid compound contains the corresponding type of skeleton, characteristic group or residue, and the parameter is defined as 0 if it does not contain;
[0058] A general prediction model of retention time of single subclass lipid compounds is obtained by regression modeling, and its expression is t R C = k0+ K1c 3 + K2c 2 + K3c+ K4, wherein k0 is a constant, K1-K4 are different functions composed of d, d 2 and d 3 , and other arguments are constant values. In particular, a general prediction model of retention time of multi-class lipid compounds is obtained by regression modeling, and its expression is t R C = k0+ K1c 3 + K2c 2+K3c+K4, wherein k0 is a constant, K1-K4 are not constants, and K1-K4 are different functions respectively including a plurality of said characteristic structure parameters other than the independent variable c.
[0059] The general prediction model constructed by the method can establish high-coverage and high-precision lipid compound retention time t R The theoretical database, in combination with the retention time and secondary mass spectrum of the lipid compound, is used for identification and quantitative analysis of the lipid compound, and is especially used for identifying the lipid compound without commercial standard. Meanwhile, the general prediction model constructed based on the method is used for prediction and obtaining the predicted t R of the lipid compound, which is used for MRM acquisition parameter setting, is conducive to fully and reasonably using the mass spectrometry detection resource, and improves the sensitivity, stability and coverage.
[0060] In some embodiments, five categories of lipid compounds (free fatty acids and modified compounds thereof, glycerolipids, glycerophospholipids, sphingolipids and sterol lipids) in the identified biological matrix are used, based on reversed-phase chromatography, and the expression of the general model for predicting the retention time of the five categories of lipid compounds constructed by the above method is
[0061] K1 is
[0062] K2 is
[0063] K3 is
[0064] K4 is In the formula, k i is the equation coefficient corresponding to the QSRR model constructed by the multivariate polynomial regression analysis.
[0065] For each subcategory of the five categories of lipids, the above K1-K4 are simplified into new functions including d, d 2 and d 3 .
[0066] Among them, for different reversed-phase chromatography conditions (including different mobile phase compositions and pH values, column types and elution gradients), the covered lipid subcategories and the corresponding quantitative retention structure relationship will be different, so that different optimal QSRR mathematical models are constructed. For example, based on the reversed-phase chromatography conditions of acidic mobile phase A and mobile phase B with different pH values (wherein the pH of the mobile phase A is 2.5-3.0, and the pH of the mobile phase B is 2.5-5.0), the general prediction model expression for predicting the retention time of the five categories of different lipid compounds is
[0067] K1 is
[0068] K2 is
[0069] K3 is
[0070] K4 is
[0071] K1 is
[0072] K2 is
[0073] K3 is
[0074] K4 is
[0075] Special note: in the present application, the coefficients k i in the different general prediction model expressions for the retention time of the different classes of lipid compounds have the same subscript "i", which does not mean that the regression coefficient values are the same in the different model expressions. For example, the k1 in the prediction model expression for the retention time of free fatty acids and their modified compounds is the same as the k1 in the prediction model for glycerolipids or other classes of lipids, but the specific values can be the same or different, and the specific value of k1 is the coefficient value obtained by multivariate nonlinear analysis, and the values of other k i are the same.
[0076] Further, the above independent variables can be quantitatively processed according to the characteristic structures of the five classes of lipid compounds, and the corresponding identified lipid compounds are used to construct the structure-retention time general prediction model for free fatty acids and their modified compounds, glycerolipids, glycerophospholipids, sphingolipids and sterol lipids. Depending on the samples used for modeling, the specific equation coefficients obtained can differ, but all belong to the scope of the present application. The specific equations are as follows:
[0077] Based on the reversed-phase chromatographic conditions, the retention time universal prediction model of free fatty acids and their modified compounds (such as acylcarnitines, hydroxy fatty acids and dicarboxylic acids, FAs) constructed by the method, the expression of which is
[0078] K1 is k1+k2d+k3d 2 +k4d 3 +k5h+k6m+k7b;
[0079] K2 is k8+k9d+k 10 d 2 +k 11 d 3 +k 12 h+k 13 m+k 14 n+k 15 b;
[0080] K3 is k 16 +k 17 d+k 18 d 2 +k 19 d 3 +k 20 h+k 21 m+k 22 n;
[0081] K4 is k 23 d 3 +k 24 h+k 25 m+k 26 n+k 27 b;
[0082] As in some embodiments (such as M1 chromatographic conditions), the prediction model is constructed by including 33 free fatty acids (FFAs) and 37 free fatty acids and their modified compounds containing C16:0 acylcarnitines (ACar), 2-hydroxy fatty acids (2HOFA), 3-hydroxy fatty acids (3HOFA) and dicarboxylic acids (DCA) standards, respectively, the quantitative processing is as above, the variable screening is completed by stepwise regression and full subset regression, and the optimal model is selected according to the adjusted goodness of fit (Goodness of fit, R 2 ) and prediction error by 10-fold cross-validation, the expression of which is t R C =k0+(k1+k2d 2 +k3d 3 )c 3 +(k4d+k5d 2 +k6d 3 )c 2 +(k7+k8d 2 +k9d3 c+k 10 m+k 11 h+k 14 n+k 15 b.
[0083] According to the goodness of fit and prediction error, QSRR model is optimized by 10-fold cross-validation, and the optimal QSRR model expression of the retention time of free fatty acids and modified compounds of different lipid compounds is obtained as follows: k0=-4.12e 0 ; k1=-2.06e -4 ; k2=-5.83e -4 ; k3=-2.96e -5 ; k4=4.98e -3 ; k5=-3.67e -2 ; k6=1.43e -3 ; k7=7.38e -1 ; k8=-1.69e -2 ; k9=7.16e -1 ; k 10 =-4.71e 0 ; k 11 =-1.37e -1 ; k 12 =-2.73e 0 ; k 13 =-4.67e -1 ; k 14 =-3.90e -1 ; k 15 =-3.48e 0 .
[0084] The general prediction model of the retention time of glycerolipid (GLs) lipid compounds constructed by the method is as follows: K1 is k1+k2d+k3d 2 +k4d 3 +k5g+k6j+k7w+k8β+k9ψ;
[0085] K2 is k 10 +k 11 d+k 12 d 2 +k 13 d 3 +k 14 g+k 15 j+k 16 t+k 17 w+k 18 α+k 19 β+k 20 γ+k 21 ψ;
[0086] K3 is k 22 + k 23 d + k 24 d 2 + k 25 d 3 + k 26 g + k 27 h + k 28 j + k 29 n + k 30 t + k 31 w + k 32 z + k 33 a + k 34 b + k 35 g + k 36 y + k
[0087] K4 is k 37 d + k 38 d 2 + k 39 d 3 + k 40 g + k 41 h + k 42 j + k 43 n + k 44 t + k 45 w + k 46 z + k 47 b + k 48 g + k 49 y + k
[0088] A general retention time prediction model for glycerophospholipid (GPs) lipid compounds constructed by the method, the expression of which K1 is
[0089] K2 is
[0090] K3 is
[0091] K4 is
[0092] A general retention time prediction model for sphingolipid (SPs) lipid compounds constructed by the method, the expression (t R C = k0+ K1c 3 + K2c 2 + K3c + K4) in which
[0093] K1 is k1+k2d+k3d 2 +k4d3 + k5w + k6μ + k7ψ;
[0094] K2 = k8 + k9a + k 10 d + k 11 d 2 + k 12 d 3 + k 13 h + k 14 p + k 15 q + k 16 w + k 17 y + k 18 δ + k 19 μ + k 20 ψ;
[0095] K3 = k 21 + k 22 a + k 23 d + k 24 d 2 + k 25 d 3 + k 26 h + k 27 p + k 28 q + k 29 δ + k 30 η + k 31 μ + k 32 ψ;
[0096] K4 = k 33 a + k 34 d + k 35 d 2 + k 36 e + k 37 p + k 38 q + k 39 w + k 40 y + k 41 δ + k 42 μ + k 43 ψ.
[0097] The solid sterol lipids (STs) lipid compound retention time universal prediction model QSRR is constructed by the method, and the expression (t R C = k0+ K1c 3 + K2c 2 + K3c + K4) is
[0098] K1 = k1 + k2d + k3d 2 ;
[0099] K2 = k4 + k5d + k6d 2 + k7d 3 ;
[0100] K3 is k8+k9d+k 10 d 2 +k 11 ε+k 12 L;
[0101] K4 is
[0102] Furthermore, when the mobile phase used in the reversed-phase chromatography is acidic, the general retention time prediction model QSRR for free fatty acids and their modified compounds (such as acylcarnitine, hydroxy fatty acids, and dicarboxylic acids) constructed using this method is expressed as follows (t R C =k0+K1c 3 +K2c 2 +K3c+K4)
[0103] K1 is k1+k2d+k3m+k4b+k5d 2 +k6d 3 ;
[0104] K2 is k7d+k8m+k9h+k 10 n+k 11 b+k 12 d 2 +k 13 d 3 ;
[0105] K3 is k 14 +k 15 d+k 16 h+k 17 n+k 18 d 2 +k 19 d 3 ;
[0106] K4 is k 20 m+k 21 h+k 22 n+k 23 b;
[0107] Where k i (i = 1 to 25) are the equation coefficients related to the QSRR model, and their specific values are determined by multivariate polynomial regression analysis. The specific equation coefficients obtained may vary depending on the samples used for modeling, but all are within the scope of this invention, and the same applies below.
[0108] The general prediction model for the retention time of glycerol lipids (GLs) constructed using this method is expressed as (t R C= k0+ K1c 3 +K2c 2 +K3c+K4) in which
[0109] K1 is k1+k2d 2 +k3d 3 +k4β+k5w+k6j+k7ψ;
[0110] K2 is k8+k9d+k 10 d 2 +k 11 d 3 +k 12 β+k 13 γ+k 14 w+k 15 j+k 16 t+k 17 ψ;
[0111] K3 is k 18 +k 19 d+k 20 h+k 21 n+k 22 d 2 +k 23 d 3 +k 24 α+k 25 β+k 26 γ+k 27 z+k 28 w+k 29 t+k 30 ψ;
[0112] K4 is k 31 h+k 32 n+k 33 d+k 34 γ+k 35 w+k 36 g+k 37 ψ+k 38 d 2 +k 39 d 3 .
[0113] A general retention time prediction model for glycerophospholipid (GPs) lipid compounds constructed using the present method, expressed as (t R C = k0+ K1c 3 +K2c 2 +K3c+K4) in which
[0114] K1 is
[0115] K2 is
[0116] K3 is k
[0117] K4 is k
[0118] A general retention time prediction model for sphingolipid (SPs) lipid compounds constructed using the present method has the expression (t R C = k0+ K1c 3 + K2c 2 + K3c+ K4)
[0119] K1 is k1+k2d+k3d 2 +k4d 3 +k5w+k6ψ+k7μ;
[0120] K2 is k8+k9d+k 10 d 2 +k 11 d 3 +k 12 ψ+k 13 p+k 14 μ+k 15 δ;
[0121] K3 is k 16 +k 17 d+k 18 h+k 19 d 2 +k 20 d 3 +k 21 ψ+k 22 p+k 23 μ+k 24 η+k 25 δ;
[0122] K4 is k 26 d+k 27 w+k 28 ψ+k 29 d 2 +k 30 p+k 31 e+k 32 q+k 33 μ+k 34 δ+k 35 y.
[0123] A general retention time prediction model for sterol lipids (STs) lipid compounds constructed using the present method has the expression (t R C = k0+ K1c3 + K2c 2 + K3c + K4) in which K1 is k1 + k2d + k3m + k4d 2 ;
[0124] K2 is k4 + k5d + k6m + k7h + k8d 2 + k9h 3 + k10d
[0125] K3 is k11 + k12d + k13m + k14h + k15d 10 + k16h 11 + k17d 2 + k18m 12 + k19d 3 + k20m 13 + k21h 14 + k22d 15 + k23m 16 + k24h
[0126] K4 is k25 + k26m + k27h + k28d
[0127] In the case of the detection of the mobile phase of the reversed phase chromatography using the conventional mobile phase (in which the pH of the mobile phase A and the mobile phase B are the same, and the pH is neutral), the retention time of the free fatty acid and the modified compound (such as acyl carnitine, hydroxy fatty acid and dicarboxylic acid) constructed by the present method is used to construct the general prediction model (t R C = k0 + K1c 3 + K2c 2 + K3c + K4) in which K1 is k1 + k2d + k3m + k4d 2 + k5h
[0128] K2 is k6 + k7d + k8m + k9h + k10d 10 + k11h 2 + k12d 11 + k13m 3 ;
[0129] K3 is k14 + k15d + k16m + k17h + k18d 12 + k19h 13 + k20d 14 + k21m 15 + k22d 2 + k23m 16 + k24d 3 + k25m 17 + k26h
[0130] K4 is k27 + k28m + k29h + k30d 18 + k31h 19 + k32d 20 + k33m 3 ;
[0131] K1 is k1 + k2d + k3w + k4j + k5ψ + k6g in the general predictive model expression for the retention time of glycerolipid (GLs) lipid compounds constructed using the present method 2 + k3w + k4j + k5ψ + k6g;
[0132] K2 is k7 + k8d + k9d 2 + k 10 γ + k 11 w + k 12 j + k 13 t + k 14 ψ + k 15 α + k 16 g;
[0133] K3 is k 17 d + k 18 d 2 + k 19 d 3 + k 20 α + k 21 β + k 22 γ + k 23 w + k 24 t + k 25 ψ + k 26 j + k 27 g;
[0134] K4 is k 28 h + k 29 γ + k 30 w + k 31 g + k 32 d 2 + k 33 β + k 34 z + k 35 j + k 36 t;
[0135] K1 is k1 + k2d + k3w + k4j + k5ψ + k6g in the general predictive model expression for the retention time of glycerophospholipid (GPs) lipid compounds constructed using the present method
[0136] K2 is k7 + k8d + k9d
[0137] K3 is k
[0138] K4 is k
[0139] K1 is k1 + k2d + k3w + k4j + k5ψ + k6g in the general predictive model expression for the retention time of sphingolipid (SPs) lipid compounds constructed using the present method 2+k3d 3 +k4ψ+k5μ;
[0140] K2 is k6+k7h+k8d 2 +k9d 3 +k 10 w+k 11 ψ+k 12 p+k 13 q+k 14 μ+k 15 a+k 16 y;
[0141] K3 is k 17 +k 18 d+k 19 d 2 +k 20 d 3 +k 21 p+k 22 q+k 23 μ+k 24 a;
[0142] K4 is k 25 w+k 26 d 2 +k 27 μ+k 28 δ+k 29 a;
[0143] K1 is k1+k2d in the general prediction model expression of the retention time of the solid sterol lipids (STs) lipid compounds constructed by the method;
[0144] K2 is k3+k4d;
[0145] K3 is k5+k6d+k7L;
[0146] K4 is k8d+k9L+k 10 ξ+k 12 a+k 13 a.
[0147] In addition, the application also provides a multi-class lipid general compound retention time prediction system, which comprises a data acquisition module and a retention time prediction module;
[0148] The data acquisition module is used for acquiring characteristic structure parameters of a lipid compound to be identified, performing quantitative characterization according to the method, and submitting to the retention time prediction module;
[0149] The retention time prediction module is provided with a general prediction model of the retention time of lipid compounds constructed according to the method of the present application, which is used to predict the theoretical retention time of the lipid compounds according to the model expression based on the obtained characteristic structural quantitative data and output the prediction result.
[0150] Based on the elution of reverse phase chromatography, the data acquisition module is used to acquire the characteristic structural parameters of the lipid compounds to be identified, including the total carbon number c, the number of double bonds d, the number of hydroxylated modifications h and the position n on the fatty acyl chain, the number of carboxyl groups b at the end of the acyl chain, the acyl carnitine m and the number of fatty acyl modifications a on the modified hydroxyl position of the acyl chain (such as FAHFA) for free fatty acids and their modified compounds (such as acyl carnitine, hydroxy fatty acid and dicarboxylic acid).
[0151] For glycerolipid lipid compounds, the characteristic structural parameters of the lipid compounds to be identified include the total carbon number c, the number of double bonds d, the fatty acyl position α, β and γ on the glycerol skeleton and the ether chain j, the glycosylation w, the number of trimethylhomoserine t, the sulfonated galactosyl ψ and the glucuronic acid g.
[0152] For glycerophospholipid lipid compounds, the characteristic structural parameters of the lipid compounds to be identified include the total carbon number c, the number of double bonds d, ethanolamine e, alkenyl ether chain f, the number of hydroxylated modifications h and the position n on the fatty acyl chain, inositol ether chain j, phosphate p, choline q, glycerol v, peroxidation modification z, fatty acyl position α, β and γ on the glycerol skeleton, dimethyl ethanolamine r, ethanol u, methyl ethanolamine x and serine s.
[0153] For glycerophospholipid lipid compounds, the characteristic structural parameters of the lipid compounds to be identified include the total carbon number c, the number of double bonds d, ethanolamine e, alkenyl ether chain f, the number of hydroxylated modifications h and the position n on the fatty acyl chain, inositol
[0154] For glycerophospholipid lipid compounds, the characteristic structural parameters of the lipid compounds to be identified include the total carbon number c, the number of double bonds d, ethanolamine e, alkenyl ether chain f, the number of hydroxylated modifications h and the position n on the fatty acyl chain, inositol and 3β-cholestanol ω.
[0155] Application examples:
[0156] Typical biological matrices such as human body fluids (plasma, urine), animal tissues (mouse heart, liver, brain, feces, kidney and lung), bacteria (E. coli), fungi (yeast), lower plants (lichen male and female), higher plants (Arabidopsis thaliana) were used for non-targeted lipidomics analysis. The t R Data of the matched lipid compounds were obtained and a mathematical model training set was established, so as to further predict the t R of unidentified lipid compounds. The MRM acquisition window was optimized for better results. The biological matrix was plasma, and the Matyash extraction method (“Lipid extraction by methyl-tert-butyl ether for high-throughput lipidomics”) was used, and 80 μL of dichloromethane (DCM) / methanol (MeOH) mixed solution (1:1, v / v) was used for reconstitution, and finally recorded as plasma A. In order to quantitatively determine the acid, low abundance and / or strong hydrophilic lipids in human plasma, the optimization method proposed by Sarafian (“Objective set of criteria for optimization of sample preparation procedures for ultra-high throughput untargeted blood plasma lipid profiling by ultra performance liquid chromatography-mass spectrometry”) was used for extraction, and 36 μL of dichloromethane (DCM) / methanol (MeOH) mixed solution (1:1, v / v, 0.1% FA) was used for reconstitution, and finally recorded as plasma B. The lipids of other biological samples of the above-mentioned biological matrix except plasma were extracted by the extraction method proposed by Sarafian, and were diluted into two different injection samples. For example, the lipids extracted from the liver matrix were reconstituted with 150-300 μL of DCM (dichloromethane) / MeOH (methanol) mixed solution (1:1, v / v) as liver A; 50% of sample A by volume was taken, nitrogen was blown to dryness, and 40-60 μL of DCM / MeOH solution (1:1, v / v, 0.1% formic acid FA) was used for reconstitution, which was liver B. Each kind of biological matrix reconstitution solution was injected in turn and subjected to non-targeted lipidomics analysis according to sample A and sample B.
[0157] Data acquisition used SCIEX zenoTOF 7600 and X500R TOF (SCIEX, Chromos, Singapore) coupled with Shimadzu UPLC system (Kyoto, Japan) for non-targeted lipidomics analysis to obtain retention time (t R ) and mass spectrometry data of different standard lipid compounds and typical biological matrix.
[0158] The following are two reversed-phase chromatography elution conditions used in the examples. M1 is the optimized reversed-phase chromatography condition screened by our experiment, and M2 is the traditional conventional reversed-phase chromatography condition, as follows:
[0159] Table 1-1 M1 and M2 two reversed-phase chromatography condition settings
[0160]
[0161] Wherein the total volume of mobile phase A and mobile phase B is 100%, and the gradient elution G1 is as follows: 0-1 min: the volume ratio of mobile phase B is maintained at 5%; 1-1.5 min: the volume ratio of mobile phase B is increased from 5% to 25%; 1.5-6 min: the volume ratio of mobile phase B is increased from 25% to 90%; 6-8 min: the volume ratio of mobile phase B is increased from 90% to 92%; 8-9 min: the volume ratio of mobile phase B is increased from 92% to 99%; 9-11 min: the volume ratio of mobile phase B is maintained at 99%.
[0162] The gradient elution G2 is as follows: 0-0.5 min: the volume ratio of mobile phase B is maintained at 20%; 0.5-1.5 min: the volume ratio of mobile phase B is increased from 20% to 40%; 1.5-3 min: the volume ratio of mobile phase B is increased from 40% to 60%; 3-13 min: the volume ratio of mobile phase B is increased from 60% to 98%.
[0163] Example 1: Construction of new retention time prediction models for 9 lipid subclasses in five major categories of lipids
[0164] Based on the representative lipid subclasses in five major categories of lipids (including free fatty acids and their modified compounds, glycerolipids, glycerophospholipids, sphingolipids and sterol lipids), including FFA, ACar, diglyceride (DAG), triglyceride (TAG), phosphatidylcholine (PC), phosphatidylethanolamine (PE), ceramide (Cer), sphingomyelin (SM) and cholesterol ester (CE), the general prediction model of retention time was constructed, as follows:
[0165] Firstly, under the condition of reversed-phase chromatography M1, 46 FFA (including 37 standard compounds), 26 ACar (including 17 standard compounds), 64 DAG (including 8 standard compounds), 204 TAG (including 7 standard compounds), 120 PC (including 10 standard compounds), 95 PE (including 8 standard compounds), 29 Cer (including 1 standard compound), 61 SM (including 4 standard compounds) and 33 CE (including 8 standard compounds) were detected in typical biological matrix and corresponding lipid standard compounds, respectively, which were used as the data set for modeling in the present application (divided into training set and test set according to the ratio of 8:2).
[0166] Among them, we used 24 FFA (all standard), 19 ACar (including 12 standard compounds), 48 DAG (including 5 standard compounds), 148 TAG (including 5 standard compounds), 96 PC (including 9 standard compounds), 78 PE (including 7 standard compounds), 24 Cer (including 1 standard compound), 36 SM (including 2 standard compounds) and 23 CE (including 7 standard compounds) as the training set, respectively, to build QSRR new mathematical models of the above 9 sub-classes based on regression model. Then, the remaining 18 FFA, 7 ACar, 16 DAG, 56 TAG, 24 PC, 17 PE, 5 Cer, 25 SM and 10 CE in the data set were used as the test set.
[0167] Among them, the experimental retention time t R E The total number of carbon (c) and carbon-carbon double bond (d) contained in the fatty acyl chain, ether chain, alkenyl ether chain or / and sphingosine in the chemical structure were used as independent variables, and multivariate polynomial nonlinear regression modeling was used, and the general expression was
[0168] t R E = k0+ k1d+ k2d 2 +k3d 3 +(k4+k5d+k6d 2 +k7d 3 )c+(k8+k9d+k 10 d 2 +k 11 d 3 )c 2 +(k 12 +k 13 d+k 14 d 2 +k 15 d 3 )c 3, wherein k1-k 15 are coefficients, and k0 is a constant. In the new mathematical model of QSRR corresponding to each sub-class, the values of k0-k 15 are shown in Table 1-2.
[0169] Table 1-2 Equation coefficients (k i , i = 1-15) related to QSRR model of 9 lipid sub-classes under reversed phase chromatography condition M1
[0170]
[0171]
[0172] As shown in Figure 1 and Figure 2 , compared with the traditional QSRR mathematical model, the single QSRR mathematical model constructed for the above 9 sub-classes can well predict the retention time (t R C ) of other lipid compounds in the same sub-class, and has a good linear fitting relationship (R 2 ≥ 0.99) with the experimental retention time (t R E ). At the same time, the retention time prediction deviation (Δt R ) of the lipid compounds of FFA, ACar and PC in these verification sets is less than 0.3 min and 0.2 min, respectively, and the Δt R of the lipid compounds in DAG, TAG, PC, PE, Cer, SM and CE is less than 0.15 min. Therefore, the experimental results show that the single new general mathematical model constructed for the lipid compounds in the same sub-class can effectively predict the theoretical retention time of other lipid compounds.
[0173] Example 2 Construction of retention time general prediction model of free fatty acid and its modified compounds
[0174] In the process of lipidomic qualitative and quantitative analysis of biological matrix, the number of detected lipid compounds in different lipid subclasses is different. When the number of detected lipid compounds in some subclasses (such as sphingosine-1-phosphate, sphingosine) is less than 3, the traditional multivariate polynomial regression algorithm or the new general mathematical model described in embodiment 1 of the present application cannot be constructed. Therefore, in order to further solve this problem, we fully consider the similarities and differences in the chemical structures of lipid compounds in different subclasses, such as including FFA, ACar, 2HOFA, 3HOFA, DCA and FAHFA in free fatty acids and their modified compounds (such as acyl carnitine, hydroxy fatty acid and dicarboxylic acid). These lipid compounds mainly include the total number of carbon (c) and double bond (d) in different fatty acyl chains, whether acyl carnitine (m) is contained, the number (h) and modification position (n) of hydroxy modification on the fatty acyl chain, and the number (b) of carboxyl at the end of the acyl chain. Therefore, we quantitatively process the characteristic structures of these lipid compounds, and construct a general retention time prediction model for free fatty acids and their modified compounds by using multivariate regression nonlinear analysis.
[0175] We first selected 62 lipid compounds from the 95 lipid compounds identified in the typical biological matrix as the training set to construct the general retention time prediction model of free fatty acids and their modified compounds. R Construction of prediction model I (model I in table 2). Among them, the 62 lipid compounds include 29 free fatty acids (FFA) verified by standard, 14 ACar (including 7 standard verification), 10 HOFA (including 5 standard verification), 3 3HOFA (including 1 standard verification), 5 DCA (including 3 standard verification) and 1 FAHFA, which are detected by UPLC-QTOFMS, and the reverse phase chromatography condition uses M1.
[0176] Quantitative processing of different lipid compounds: we define the parameters m, h, n, b for carnitine, hydroxy modification and modification position, and dicarboxylic acid, respectively, define the total carbon number of fatty acyl chain as parameter c (such as carbon chain length 16, c = 16), and define the number of carbon-carbon double bonds as parameter d (such as containing three double bonds, d = 3; not containing double bond, d = 0); containing carnitine, define m = 1, not containing carnitine, define m = 0; free fatty acid and its modified compound is hydroxylated, define h = 1, and n is quantified according to the carbon position of hydroxylation (such as the 2nd carbon is hydroxylated, define n = 2, not hydroxylated, define h = 0, n = 0; free fatty acid and its modified compound contains dicarboxylic acid, define b = 1, does not contain dicarboxylic acid, define b = 0; the obtained experimental retention time t R EFor the dependent variable, the R language code was used to complete the variable combination, screening, modeling, and verification processes with carbon chain length (c), double bond number (d), carnitine content (m), hydroxylated modification (h) and position (n), and double carboxyl content (b) as independent variables to obtain the retention time of different lipid compounds in free fatty acid and its modified compounds. QSRR model I was established with a 60% data set, and its expression is:
[0177] t R C = k0+ (k1+k2m)c 3 +(k3md+k4h+k5n+k6b)c 2 +(k7+k8n+k9b)c+k 10 d+k 11 d 2 +
[0178] k 12 d 3 +k 13 m.
[0179] According to the goodness of fit and prediction error, the QSRR model was optimized by 10-fold cross-validation to obtain a better QSRR model for predicting the retention time of free fatty acid and its modified compounds. In its expression, k0=-4.16e 0 ; k1=-2.03e -4 ; k2=1.52e -4 ; k3=-4.02e -3 ; k4=-2.86e -3 ; k5=3.70e -3 ; k6=1.08e -2 ; k7=7.15e -1 ; k8=-7.67e -2 ; k9=-3.88e -1 ; k 10 =-8.90e -1 ; k 11 =8.98e -2 ; k 12 =-7.72e -3 ; k 13 =-2.10e 0 ; The adjusted R 2 -squared (adjusted R-squared) reached 0.9974, showing good fitting effect.
[0180] Validation of the predictive model: The retention times t of different lipid compounds in free fatty acids and their modified compounds were calculated according to the same quantification process as in the modeling, for compounds containing carnitine (m), hydroxylation modification (h) and position (n), and dicarboxyl groups (b). R C For example, the retention times (t) of 9 types of ACar, 12 types of FFA, 6 types of 2HOFA, 3 types of 3HOFA, and 3 types of DCA calculated using the above formula. R C ), calculate the obtained predicted retention time (t) R C ) and its experimental retention time value (t) R E The correlation analysis results between them are shown in Table 2 and Figure 3 As shown in A, Figure 3 The QSRR mathematical model constructed in section A is used to predict lipid compounds in ACar, 2HOFA, 3HOFA and DCA.
[0181] As shown in Table 2, in Model I, t R E and t R C The difference (△t) R All values were less than 0.30 min. Experimental results show that when the training set contains 60% of ACar, 2HOFA, 3HOFA and DCA, the QSRR model I has good predictive performance and can accurately predict the theoretical retention time of free fatty acids and their modified compounds without commercial standards, thus reducing the false positive rate of mass spectrometry identification of free fatty acids and their modified compounds.
[0182] In addition Figure 3 As shown in section A, the retention times (t) of lipid compounds in ACar, 2HOFA, 3HOFA, and DCA were calculated according to the QSRR model described above. R C ) and its experimental value (t) R E All showed good correlation (R). 2 The value is 0.998. This indicates that the retention time prediction model I constructed using this method for free fatty acids and their modified compounds has high accuracy and can meet the prediction requirements for the theoretical retention times of other lipid compounds in these lipid subclasses.
[0183] Table 2. Experimental values of retention times (t) for ACar, FFA, 2HOFA, 3HOFA, and DCA R E ) and the predicted model calculated value (t) R C The correlation between these lipid compounds (* indicates the retention time for standard validation). "-" represents the t-value of these lipid compounds.R Data in model I or II construction process as a training set, t R Data as a validation set.
[0184]
[0185]
[0186]
[0187]
[0188] In addition, in order to further evaluate whether the fitting ability and prediction ability of QSRR model for other lipid compounds are still strong when the number of lipid compounds in part of the subcategory in the training set is small, we further constructed QSRR model II. Among them, we selected 37 lipid compounds from 94 lipid compounds (except 1 FAHFA) as the training set, and the t R Construction of prediction model II (model II in Table 2). The training set includes 33 FFA verified by standard, and 4 standard containing C16:0 ACar, 2-HOFA, 3-HOFA and DCA respectively; among them, ACar, 2-HOFA, 3-HOFA and DCA containing C16:0 mean that the fatty acid chain of these lipid compounds is hexadecane saturated fatty acid, which is detected by UPLC-QTOFMS, and the reverse phase chromatographic condition uses M1. In addition, only 1 FAHFA is identified, and no other lipid compound can be used to verify the prediction ability of the model, so it is not further included in the training set.
[0189] The expression of QSRR model II is:
[0190] t R C = k0+ (k1+k2d 2 +k3d 3 )c 3 +(k4d+k5d 2 +k6d 3 )c 2 +(k7+k8d 2 +k9d 3 )c+k 10 m+k 11 h+k 14 n+k 15 b.
[0191] According to the goodness of fit and prediction error, the QSRR model is optimized by 10-fold cross-validation to obtain a better QSRR model for predicting the retention time of free fatty acid and its modified compounds, and the expression of which k0= -4.12e0 k1 = -2.06e -4 k2 = -5.83e -4 k3 = -2.96e -5 k4 = 4.98e -3 k5 = -3.67e -2 k6 = 1.43e -3 k7 = 7.38e -1 k8 = -1.69e -2 k9 = 7.16e -1 ;k 10 =-4.71e 0 ;k 11 =-1.37e -1 ;k 12 =-2.73e 0 ;k 13 = -4.67e -1 ;k 14 =-3.90e -1 ;k 15 = -3.48e 0 Adjusted R-value 2 The value reached 0.9991, indicating a good fit.
[0192] Validation of the predictive model: The retention times t of different lipid compounds in free fatty acids and their modified compounds were calculated according to the same quantification process as in the modeling, for compounds containing carnitine (m), hydroxylation modification (h) and position (n), and dicarboxyl groups (b). R C For example, the retention times (t) of 22 acylcarnitines, 8 FFAs, 15 2HOFAs, 5 3HOFAs, and 7 DCAs calculated using the above formula. R C ), calculate the obtained predicted retention time (t) R C ) and its experimental retention time value (t) R E The correlation analysis results between them are shown in Table 2 and Figure 3 As shown, Figure 3 B represents the predictive ability of the constructed QSRR mathematical model II for lipid compounds in ACar, 2HOFA, 3HOFA, and DCA.
[0193] As shown in Table 2, t R E and t R C The difference (△t) RAll were less than 0.50 min (except for ACar 26:0 and ACar 18:3), of which approximately 90% of the lipid compounds had a Δt value of less than 0.50 min. R Less than 0.4 min. Experimental results show that when the training set contains only one type of ACar, 2HOFA, 3HOFA, and DCA containing C16:0, the QSRR general model II constructed using this method still has good predictive performance, low false positive rate for lipid compound identification, and can accurately predict the theoretical retention time without commercial standards, thus reducing the false positive rate of mass spectrometry identification of free fatty acids and their modified compounds. Further identification through subsequent MRM quantitative / semi-quantitative analysis revealed that ACar 18:2 and 3HOFA 17:0 were detected at 3.25 min and 5.23 min, respectively, while the retention times t predicted by this QSRR general model for ACar 18:2 and 3HOFA 17:0 were... R C The times were 3.33 min and 5.39 min, respectively, Δt R A result less than 0.2 minutes indicates high accuracy in the prediction.
[0194] In addition Figure 3 As shown in section B, the retention times (t) of lipid compounds in ACar, 2HOFA, 3HOFA, and DCA were calculated according to the QSRR model described above. R C ) and its experimental value (t) R E All showed good correlation (R). 2 The value is 0.992. This indicates that the retention time prediction model for free fatty acids and their modified compounds constructed using this method has high accuracy and can meet the requirements of predicting the theoretical retention time of other lipid subclasses using only one lipid compound or a standard containing only one lipid compound.
[0195] Example 3: Construction of a universal lipid retention time prediction model for five major lipid classes
[0196] Given that Example 2 utilizes a general retention time prediction model for free fatty acids and their modified compounds, which can obtain highly accurate retention times for lipid compounds, this example constructs general retention time prediction models for multiple lipid classes (five major lipid classes: including free fatty acids and their modified compounds, glycerides, glycerophospholipids, sphingolipids, and sterol lipids) to further illustrate the universality of this method. Furthermore, the fitting ability of the general prediction models constructed under two different reversed-phase chromatographic conditions is compared, as detailed below:
[0197] In this embodiment, through experimental screening, 38 compound characteristic structures were ultimately selected as independent variables to construct a general prediction model for retention time of five major lipid compounds. Preferred QSRR models for each of the five major lipid compounds were also constructed, and the fitting ability of the prediction models under different reversed-phase chromatographic elution conditions was compared. The 38 compound characteristic structures include the total number of carbons (c), double bond number (d), number (h) and position (n) of hydroxyl modifications on the fatty acyl chain, ether chain, alkenyl ether chain, and / or sphingosine skeleton, of fatty acyl chains, ether chains, and / or sphingosine skeletons, of carboxyl groups at the end of the acyl chain (b), number of fatty acyl chains at the hydroxyl position (a), fatty acyl positions (α: sn1, β: sn2, and γ: sn3) on the glycerol skeleton, and the presence of acylcarnitine (m), alkenyl ether chain (f), ether chain (j), cholesterol (L), choline (q), dimethylethanolamine (r), ethanol (u), ethanolamine (e), glucuronic acid (g), glycerol (v), glycosylation (w), and inositol. The number of methylethanolamine (x), sialic acid (y), peroxidation modification (z), phosphoric acid (p), 22,23-ergosterol (ε), rapeseed sterol (θ), 24α-ethylcholesterol (ξ), stigmasterol (π), serine (s), dihydrosphingosine (η), sphingosine (μ) and its hydroxylation (δ), 3β-cholesterol (ω), sulfonated galactosyl (ψ), sulfonyl (φ) and trimethylhomoserine (t) and the number of 24β-ethylcholesterol skeletons λ.
[0198] First, based on high-performance liquid chromatography-mass spectrometry (HPLC-MS), five major lipid classes (including free fatty acids and their modified compounds, glycerol lipids, glycerophospholipids, sphingolipids, and sterol lipids) were detected under reversed-phase chromatographic elution conditions (M1 and M2), and the experimental retention times (t) of each lipid compound were obtained. R E Secondly, the experimental retention time (t) R E Using the 38 compound structural features as independent variables and taking them as dependent variables, new QSRR mathematical models for predicting the retention times of five major lipid classes under reversed-phase chromatography were constructed based on multivariate polynomial nonlinear equations after quantifying these structural feature parameters. The quantification of these 38 structural feature descriptors is detailed in Table 3.
[0199] Table 3. Quantifiable chemical descriptors for lipid compounds
[0200]
[0201]
[0202]
[0203] Finally, the lipid extracts of typical biological samples (human body fluids (plasma, urine), animal tissues (mouse heart, liver, brain, feces, kidney and lung), bacteria (E. coli), fungi (yeast), lower plants (lichen male and female), higher plants (Arabidopsis thaliana) were analyzed by non-targeted lipidomics using reversed-phase chromatography M1 and M2 conditions combined with high-resolution mass spectrometry technology. The data obtained were peak extracted, aligned and matched with the mass spectrometry data of lipid compounds using MS-DIAL software to identify the experimental retention time of the high-coverage lipid compounds including the identified lipid compounds. Among them, under M1 condition, 102 kinds of free fatty acids and their modified compounds (such as acyl carnitine, hydroxy fatty acid and dicarboxylic acid), 389 glycerolipid compounds, 807 glycerophospholipid compounds, 214 sphingolipid compounds and 33 sterol lipid compounds were identified Figure 4 ); under M2 condition, 69 kinds of free fatty acids and their modified compounds, 383 glycerolipid compounds, 714 glycerophospholipid compounds, 221 sphingolipid compounds and 38 sterol lipid compounds were identified. The lipid compounds identified by matching the mass spectrometry data of lipid compounds are prone to false positives, such as PC (16:0_20:5) at 8.19 min and 8.98 min.
[0204] To solve this problem, we further identified the t R E As the dependent variable, the structural characteristics of the above 38 known compounds were used as the independent variable to establish a high-coverage lipid compound t R The training set was used to complete the variable combination, screening, modeling and verification process using R language code, and the QSRR model for predicting the retention time of different lipid compounds in the five major categories of lipids under different reversed-phase chromatography conditions was obtained. According to the deviation (less than 0.3 min) between the theoretical calculation and experimental measurement of the retention time of the lipid compounds, the false positive lipid compounds identified by the above non-target mass spectrometry were further removed, and the mathematical model was established again according to the above modeling steps, i.e. a more optimal mathematical model, and finally the theoretical retention time of other undetected lipid compounds was obtained, as shown in Tables 4 and 5. It is worth noting that the traditional multivariate polynomial regression equation can only construct a specific mathematical model for each subcategory (i.e. at least hundreds of mathematical models), while the present application can construct a unified mathematical model for multiple subcategories in the same major category. At the same time, the present application can still accurately predict the t RC Values.
[0205] Table 4 Equation coefficients (k) of QSRR model related to five major categories of lipids under reversed phase chromatography condition (M1) i , i = 1-104)
[0206]
[0207]
[0208]
[0209]
[0210] Table 5 Equation coefficients (k) of QSRR model related to five major categories of lipids under reversed phase chromatography condition (M2) i , i = 1-106)
[0211]
[0212]
[0213]
[0214]
[0215] To further evaluate the QSRR fitting degree of the multivariate polynomial nonlinear equation to the lipid compounds of different lipid categories under different reversed phase chromatography conditions (equation correlation coefficients are shown in Tables 3 and 4), we evaluated the QSRR fitting ability of the above multivariate polynomial nonlinear equation to the lipid compounds under different reversed phase chromatography conditions, respectively using two reversed phase chromatography conditions (M1 and M2) for elution, to detect the experimental retention time (t R E ) of each lipid compound obtained and the retention time (t R C ) of the corresponding lipid compound calculated based on the multivariate polynomial nonlinear regression equation (see Tables 5 and 6) for correlation analysis, and the results are shown in Figure 5 and Figure 6 , in which the horizontal coordinate is the experimental retention time (t R E ), and the vertical coordinate is the retention time (t R C ) calculated according to the prediction model.
[0216] Figure 5The figure shows the correlation between the experimental retention times of different lipid compounds in five major lipid classes and the corresponding QSRR calculated retention times under the optimized reversed-phase chromatography conditions (M1) of this invention. In the figure, A to E represent the QSRR fitting ability evaluation results of free fatty acids and their modified compounds, glycerides, glycerophospholipids, sphingolipids and sterol lipids, respectively. Figure 6 For different lipid compounds in five major lipid classes under traditional reversed-phase chromatography conditions (M2), t R E t calculated with the corresponding QSRR R C Correlation: In the figure, A to E represent the QSRR fitting ability assessment results of free fatty acids and their modified compounds, glycerides, glycerophospholipids, sphingolipids, and sterol lipid compounds, respectively.
[0217] Depend on Figure 5 and Figure 6 The fitting results show that, under different reversed-phase chromatographic conditions, the QSRR fitting ability of different lipid compounds in the five major lipid classes constructed according to this method can all reach R. 2 >0.99, and 97% of the lipid compounds have a Δt R All values were less than 0.3 min, indicating that the general prediction model constructed according to the method provided in this invention has a high degree of fit for the vast majority of lipid compounds. Furthermore, a QSRR mathematical model for free fatty acids and their modified compounds was constructed based on the above M1 chromatographic conditions to predict the theoretical t-values of undetected lipid compounds. R Furthermore, the MRM acquisition mode was used to detect 8 types of 2HOFA, 3 types of 3HOFA, 8 types of ACar, 1 type of ACar-OH, 11 types of DCA, and 6 types of FFA (Δt). R All values were less than 0.3 min (see Table 6). These results indicate that the universal prediction model for the five major lipid classes provided in this invention has higher accuracy and is independent of chromatographic conditions, making it suitable for establishing high-coverage, high-precision prediction models for lipid compounds. R The theoretical database is used to set up a more precise MRM data acquisition window, providing high-coverage quantitative analysis for lipidomics, especially for other lipid compounds without commercially available standards.
[0218] Table 6 shows a novel mathematical model based on MRM scanning combined with QSRR for the effective detection of other free fatty acids and their modified compounds in biological matrices.
[0219]
[0220]
[0221] The above is only the preferred embodiment of the present application, and does not limit the present application in any form. Any person skilled in the art can make some changes or modifications to the above disclosed technical route without departing from the scope of the present application to obtain equivalent embodiments. Therefore, any simple modification, equivalent change and modification made to the above embodiments without departing from the technical route of the present application or according to the technical route of the present application still belongs to the scope of the present application.
Claims
1. A method for constructing a general prediction model for the retention time of multiple categories of lipid compounds, including using the experimental retention time t of lipid compounds in a training set. R E The dependent variable is a structural parameter, and the independent variables are their characteristic structural parameters. The independent variables are quantified and then subjected to regression modeling analysis. The characteristic structural parameters include the total number of carbons (c) and the number of carbon-carbon double bonds (d) of the fatty acyl chain, ether chain, alkenyl ether chain, and / or sphingosine skeleton. The lipid compounds in the training set include five major categories of lipids: free fatty acids and their modified compounds, glycerides, glycerophospholipids, sphingolipids, and sterol lipids. The characteristic structural parameters also include the types and numbers of backbones contained in the lipid compounds, the characteristic groups in the lipid compounds, and the types and numbers of residues. The characteristic groups are modification groups and / or compound structures used to identify different lipid compounds, and the modification groups include fatty acylation modification, peroxidation modification, hydroxylation modification and / or glycosylation modification groups; The backbone, characteristic groups or residues mentioned therein are numerically quantified according to the number of corresponding backbones, characteristic groups or residues contained in the lipid compound; if they are not contained, their parameters are defined as 0. A general QSRR prediction model for the retention time of multiple lipid compounds was obtained using regression modeling, and its expression is t. R C =k0+K1c 3 + K2c 2 + K3c+K4, where k0 is a constant, and K1~K4 are different functions composed of various feature structure parameters other than the independent variable c; wherein K1~K4 in the constructed single subclass general prediction model are composed of d, d 2 With d 3 Different functions composed of; If the lipid compound in the training set contains a hydroxyl modification group h on the fatty acyl chain, the independent variable also includes the hydroxyl modification position n, that is, the carbon atom at the nth position starting from the carboxyl end of the fatty acyl chain is connected to the hydroxyl modification group; if the lipid compound does not contain a hydroxyl modification group h on the fatty acyl chain, then h is defined as 0 and n is defined as 0. The skeleton includes a glycerol skeleton v, a cholesterol skeleton L, a sphingosine skeleton μ, a dihydrosphingosine skeleton η, a 22,23-ergosterol skeleton ε, a rapeseed sterol skeleton θ, a 24α-ethylcholesterol skeleton ξ, a stigmasterol skeleton π, a 3β-cholesterol skeleton ω, and a 24β-ethylcholesterol skeleton λ. The characteristic groups include fatty acid acylation modification a of hydroxylated fatty acids or hexoses, peroxidation modification group z, hydroxylation modification group h, phosphoric acid p, glycosylation w, ether chain j, alkenyl ether chain f, glucuronic acid g, trimethylhomoserine t, dimethylethanolamine r, sialic acid y, ethanolamine e, choline q, inositol ∀, serine s, methylated ammonium ethanol x, and ethanol u; The residues include carnitine residue m, carboxylic acid residue b, sulfonyl group φ, and sulfonated galactose residue ψ; When the lipid compounds in the training set contain a glycerol backbone, the independent variables also include parameters α, β, and γ corresponding to the fatty acyl chains, ether chains, or alkenyl ether chains connected at positions sn-1, sn-2, and sn-3 on the glycerol backbone, and the parameters α, β, and γ are numerically quantified according to the number of fatty acyl chains, ether chains, or alkenyl ether chains connected at positions sn-1, sn-2, and sn-3 on the glycerol backbone, respectively; when the lipid compounds in the training set contain a sphingosine backbone, the independent variables also include sphingosine backbone hydroxylation δ, i.e., phytosphingosine, and the parameter δ is numerically quantified according to the number of sphingosine backbone hydroxylations in the lipid compound; In the general prediction model for the retention time of the five major lipid compounds constructed using regression modeling, K1 is k1 + k2d + k3m + k4b + k5d 2 + k6d 3 + k7β + k8w + k9j + k 10 ψ + k 11 v + k 12 α + k 13 z + k 14 e + k 15 q+ k 16 ∀ + k 17 μ + k 18 h + k 19 g + k 20 p + k 21 f + k 22 r + k 23 s + k 24 x + k 25 u; K2 is k 26 + k 27 d + k 28 m + k 29 h + k 30 n + k 31 b + k 32 d 2 + k 33 d 3 + k 34 β + k 35 γ + k 36 w +k 37 j + k 38 t + k 39 ψ + k 40 v + k 41 p + k 42 α + k 43 z + k 44 f + k 45 e + k 46 q + k 47 ∀ + k 48 s +k 49 u + k 50 μ + k 51 δ + k 52 g + k 53 r + k 54 x + k 55 a + k 56 y+ k 57 L; K3 is k 58 + k 59 d + k 60 h + k 61 n + k 62 d 2 + k 63 d 3 + k 64 α + k 65 β + k 66 γ + k 67 z + k 68 w +k 69 t + k 70 ψ + k 71 v + k 72 p + k 73 j + k 74 f + k 75 r + k 76 q + k 77 ∀ + k 78 s + k 79 x + k 80 μ +k 81 η + k 82 δ + k 83 L + k 84 ε + k 85 m + k 86 g + k 87 e + k 88 u + k 89 a+ k 90 ξ; K4 is k 91 m + k 92 h + k 93 n + k 94 b + k 95 d + k 96 γ + k 97 w + k 98 g + k 99 ψ + k 100 d 2 +k 101 d 3 + k 102 v + k 103 p + k 104 α + k 105 β + k 106 z + k 107 f + k 108 e + k 109 q + k 110 ∀ + k 111 μ+ k 112 δ + k 113 y + k 114 L + k 115 ε + k 116 θ + k 117 ξ + k 118 π + k 119 φ + k 120 ω + k 121 j +k 122 t + k 123 r + k 124 s + k 125 x + k 126 u + k 127 a; In the formula k i The equation coefficients corresponding to the QSRR model constructed for multiple regression analysis; For lipid compounds of the same subclass within the five major lipid classes, K1~K4 are composed of d, d 2 With d 3 The functions are composed of different values, while the other independent variables are all fixed values.
2. A general prediction model for the retention time of free fatty acids and their modified compounds constructed using the method described in claim 1, characterized in that, In the general prediction model, K1 is k1 + k2d + k3d 2 + k4d 3 + k5h + k6m + k7b; K2 is k8 + k9d + k 10 d 2 + k 11 d 3 + k 12 h + k 13 m + k 14 n + k 15 b; K3 is k 16 + k 17 d + k 18 d 2 + k 19 d 3 + k 20 h + k 21 m + k 22 n; K4 is k 23 d 3 + k 24 h + k 25 m + k 26 n + k 27 b.
3. A general prediction model for the retention time of glycerol lipid compounds constructed using the method described in claim 1, characterized in that, In the general prediction model, K1 is k1 + k2d + k3d 2 + k4d 3 + k5g + k6j + k7w + k8β +k9ψ; K2 is k 10 + k 11 d + k 12 d 2 + k 13 d 3 + k 14 g + k 15 j + k 16 t + k 17 w + k 18 α + k 19 β + k 20 γ + k 21 ψ; K3 is k 22 + k 23 d + k 24 d 2 + k 25 d 3 + k 26 g + k 27 h + k 28 j + k 29 n + k 30 t + k 31 w + k 32 z +k 33 α + k 34 β + k 35 γ + k 36 ψ; K4 is k 37 d + k 38 d 2 + k 39 d 3 + k 40 g + k 41 h + k 42 j + k 43 n + k 44 t + k 45 w + k 46 z + k 47 β + k 48 γ + k 49 ψ。 4. A general prediction model for the retention time of glycerophospholipid compounds constructed using the method described in claim 1, characterized in that, In the general prediction model, K1 is k1 + k2d + k3d 2 + k4d 3 + k5e + k6f + k7∀ + k8j +k9p + k 10 q + k 11 r + k 12 s + k 13 u + k 14 v + k 15 x + k 16 z + k 17 α + k 18 β; K2 is k 19 + k 20 d + k 21 d 2 + k 22 d 3 + k 23 e + k 24 f + k 25 h + k 26 ∀ + k 27 j + k 28 p + k 29 q +k 30 r + k 31 s + k 32 u + k 33 v + k 34 x + k 35 z + k 36 α + k 37 β; K3 is k 38 + k 39 d + k 40 d 2 + k 41 d 3 + k 42 e + k 43 f + k 44 h + k 45 ∀ + k 46 j + k 47 p + k 48 q +k 49 r + k 50 s + k 51 u + k 52 v + k 53 x + k 54 z + k 55 α + k 56 β; K4 is k 57 d + k 58 d 3 + k 59 e + k 60 f + k 61 h + k 62 ∀ + k 63 j + k 64 n + k 65 p + k 66 q + k 67 r + k 68 s + k 69 u + k 70 v + k 71 x + k 72 z + k 73 α + k 74 β + k 75 γ。 5. A general prediction model for the retention time of sphingolipid compounds constructed using the method described in claim 1, characterized in that, In the general prediction model, K1 is k1 + k2d + k3d 2 + k4d 3 + k5w + k6μ + k7ψ; K2 is k8 + k9a + k 10 d + k 11 d 2 + k 12 d 3 + k 13 h + k 14 p + k 15 q + k 16 w + k 17 y + k 18 δ +k 19 μ + k 20 ψ; K3 is k 21 + k 22 a + k 23 d + k 24 d 2 + k 25 d 3 + k 26 h + k 27 p + k 28 q + k 29 δ + k 30 η + k 31 μ +k 32 ψ; K4 is k 33 a + k 34 d + k 35 d 2 + k 36 e + k 37 p + k 38 q + k 39 w + k 40 y + k 41 δ + k 42 μ + k 43 ψ。 6. A general prediction model for the retention time of sterol lipid compounds constructed using the method described in claim 1, characterized in that, In the general prediction model, K1 is k1 + k2d + k3d 2 K2 is k4 + k5d + k6d 2 + k7d 3 + k8L; K3 is k9 + k 10 d + k 11 d 2 + k 12 d 3 + k 13 ε + k 14 L+ k 15 a + k 16 ξ; K4 is k 17 d + k 18 w + k 19 d 2 + k 20 d 3 + k 21 L + k 22 ε + k 23 θ + k 24 ξ + k 25 π + k 26 φ + k 27 ω + k 28 a.
7. A universal compound retention time prediction system for multiple lipid categories, characterized in that, Includes a data acquisition module and a retention time prediction module; The data acquisition module is used to acquire the characteristic structural parameters of the lipid compound to be identified according to the method described in claim 1, perform numerical quantification, and submit them to the retention time prediction module. The retention time prediction module is equipped with a general retention time prediction model for lipid compounds obtained according to the method described in claim 1 or the general retention time prediction model described in claims 2 to 6. It is used to calculate and predict the theoretical retention time based on the quantitative data of the characteristic structural parameters of the lipid compounds obtained according to the general prediction model expression, and output the prediction result.
Citation Information
Patent Citations
Method for predicting organic matter PDMS film-air distribution coefficient through quantitative structure-activity relationship model
CN113722988A
Chromatographic retention index prediction method and device based on graph convolutional network
CN114121178A