A quantitative model of the relationship between lipid structure and its reversed-phase liquid chromatography retention and a retention time prediction system

By introducing characteristic structural parameters and regression analysis, a high-coverage and concise QSRR model was constructed, which solves the problems of model complexity and high cost in existing technologies and achieves efficient prediction of the retention time of unknown lipid compounds.

CN121662224BActive Publication Date: 2026-06-02FUDAN UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FUDAN UNIVERSITY
Filing Date
2026-02-06
Publication Date
2026-06-02

Smart Images

  • Figure CN121662224B_ABST
    Figure CN121662224B_ABST
Patent Text Reader

Abstract

The application discloses a quantitative relationship model between lipid structure and its reversed-phase liquid chromatography retention and a retention time prediction system. The model takes the experimental retention time of a specific lipid compound as a dependent variable, takes screened specific structure characteristic parameters as independent variables, adopts an LM() function, and obtains a model as follows: Si represents the number of carbon atoms and carbon-carbon double bonds on a sphingosine skeleton, the number of carbon atoms and carbon-carbon double bonds on a sterol skeleton, and other skeletons, characteristic groups or residues; the quantitative relationship model between lipid structure and its reversed-phase liquid chromatography retention is constructed by taking screened multiple structure characteristic parameters as independent variables, and the independent variables are simplified, so that the retention time of more than ten thousand lipid compounds can be predicted R C The value of the medical health index data analysis ability can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of healthcare informatics and specific computational models, and more specifically, relates to a quantitative relationship model of lipid structure and its retention by reversed-phase liquid chromatography and a retention time prediction system. Background Technology

[0002] Many lipid compounds are key biomarkers for assessing health data in community populations and for early prevention and treatment of major diseases. These include polyunsaturated fatty acids, acylcarnitines, triglycerides, alkenyl ether glycerophospholipids, ceramides, gangliosides, and cholesterol esters, covering five major lipid classes (free fatty acids and their modified compounds, glycerides, glycerophospholipids, sphingolipids, and sterols and their esters). Liquid chromatography-mass spectrometry (LC-MS / MS) is currently one of the important tools for lipidomics analysis, primarily relying on the mass spectrometric data and retention time (t) of lipid compounds. R The data underwent qualitative analysis of chemical structure. However, because the variety of commercially available lipid compound standards (commonly around 300-500) is far less than the number of lipid compounds in biological matrices (over 40,000), most current studies predict the theoretical retention time of unknown lipid compounds by constructing quantitative structure-retention time (QSRR) models. For example, cholesterol esters (ChE), a major representative, can be modeled using traditional QSRR models with only 18 carbon atoms in the fatty acyl chain. R C ~f(d) includes ChE18:0, ChE18:1, ChE18:2 and ChE18:3, but cannot predict other unknown ChE compounds.

[0003] Current methods for constructing general QSRR models based on machine learning and nonlinear regression algorithms require a large number of endogenous lipid compounds as training sets, with their experimental retention time as the dependent variable. For example, previous studies used open-source lipidomics data analysis software (such as MS-DIAL and LIPIDMAPS) and manual review to select experimental retention time data of over a thousand endogenous lipid compounds with high qualitative accuracy as dependent variables. These data were then combined with multiple structural feature parameters (used to describe the composition of fatty acyl chains, basic skeleton, and modifying feature groups, etc.) as independent variables to construct a high-coverage, high-precision general theoretical retention time prediction model for five major classes of lipid compounds (patent CN120913704A).

[0004] However, the general QSRR model construction method disclosed in patent CN120913704A heavily relies on the qualitative accuracy of over a thousand endogenous lipid compounds in the training set, requiring a high level of experience in manual qualitative review. Therefore, how to construct a low-cost, high-precision, and high-coverage retention time-specific calculation model for one or more major lipid compounds using hundreds of lipid compounds (prioritizing lipid standards) is a key technical problem that urgently needs to be solved. Summary of the Invention

[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a quantitative relationship model and retention time prediction system for lipid structure and its retention in reversed-phase liquid chromatography. The aim is to discover, based on the characteristic structural parameters of lipid compounds screened by patent CN120913704A, additional characteristic structural parameters including the number of carbons ε and the number of carbon-carbon double bonds π in the sphingosine skeleton, the number of carbons θ and the number of carbon-carbon double bonds ξ in the sterol skeleton (these four parameters are classified as skeleton independent variables), and N-acetylgalactosamine λ (classified as characteristic group independent variables). After quantifying these characteristic structural parameters, variable screening is completed through stepwise regression and full subset regression. It was discovered that only four structural features—the number of carbons in the fatty acyl chain and the number of carbon-carbon double bonds, the number of carbons in the sterol skeleton and the number of carbon-carbon double bonds—are needed to describe dozens of sterols. This allows for the construction of a high-coverage, concise, and efficient general QSRR model for sterols and their esters, which can be used to predict the theoretical retention times of over a thousand unknown sterols and their esters. The constructed QSRR models for five major classes of lipid compounds can predict the theoretical retention times of tens of thousands of unknown lipid compounds, thus solving the technical problem of overly complex formulas in existing general QSRR models based on endogenous lipid compounds.

[0006] To achieve the above objectives, according to a first aspect of the present invention, a quantitative relationship model between lipid structure and its retention by reversed-phase liquid chromatography is provided, which is constructed by the following method:

[0007] Using the structural features of lipid compounds as candidate independent variables, the optimal independent variables and their combinations are screened out through stepwise regression modeling as specific structural features. The screened specific structural features include one or more of the following: the number of carbons ε and the number of carbon-carbon double bonds π in the sphingosine skeleton, the number of carbons θ and the number of carbon-carbon double bonds ξ in the sterol skeleton, N-acetylgalactosamine λ, and / or one or more of the skeleton, characteristic groups and residues contained in the lipid compound.

[0008] t of specific lipid compound standards R E Using the selected specific structural features as independent variables and employing the LM() function, a quantitative relationship model was established between lipid structure and its retention in reversed-phase liquid chromatography. The constructed quantitative relationship model is as follows: , This indicates the specific structural features selected. This indicates the theoretical retention time of lipid compounds.

[0009] For example, based on a specific lipid compound standard, its structural characteristic parameters are quantified, and specific structural characteristic parameters are screened out through stepwise regression or full subset regression. The screened specific structural characteristic parameters include one or more of the following: the number of carbons ε and the number of carbon-carbon double bonds π in the sphingosine skeleton, the number of carbons θ and the number of carbon-carbon double bonds ξ in the sterol skeleton, as well as the skeleton, characteristic groups and / or residues contained in the lipid compound. The skeleton, characteristic groups or residues are numerically quantified according to the number of corresponding types of skeletons, characteristic groups or residues contained in the lipid compound. If they are not contained, their parameters are defined as 0.

[0010] Experimental retention time of specific lipid compounds Using the selected specific structural feature parameters as independent variables, a quantitative relationship model between lipid structure and its retention in reversed-phase liquid chromatography was constructed using the LM() function. The constructed model is as follows: S i This represents specific structural characteristic variables other than carbon atom c and carbon-carbon double bond d in fatty acid chains, ether chains, alkenyl ether chains, and sphingosine skeletons; This indicates the theoretical retention time of an unknown lipid compound.

[0011] Preferably, in the quantitative relationship model, the specific lipid compound standard includes sterols and specific cholesterol esters; the specific cholesterol ester has a carbon number c in the fatty acyl chain ranging from 0 to 22 and a carbon-carbon double bond number d in its fatty acyl chain ranging from 0 to 6; the constructed quantitative relationship model between the sterol lipid structure and its retention by reversed-phase liquid chromatography is t R C =k0+k1θ+k2ξ+k3c+k4dc+k5c 2 +k6c 3 In the formula, k0-k6 represent the weight coefficients of different terms in the regression equation.

[0012] Preferably, in the quantitative relationship model, the specific lipid compound standard includes a specific sphingolipid compound, wherein the total number of carbons (c) on the sphingosine and fatty acyl chains of the specific sphingolipid compound ranges from 20 to 70, and the number of carbon-carbon double bonds (d) on its fatty acyl chains ranges from 0 to 4; the constructed quantitative relationship model between the lipid structure containing two or more fatty acyl chains and a sphingosine backbone and its retention by reversed-phase liquid chromatography is t R C =k0+k1a+k2h+k3d+k4d 2 +k5p+k6w+k7y+k8δ+k9c+k 10 pc+k11 c 2 +k 12 c 3 In the formula, k0-k 12 These represent the weight coefficients of different terms in the regression equation.

[0013] Preferably, the quantitative relationship model is a constructed quantitative relationship model between lipid compounds containing a single fatty acyl chain, ether chain, alkenyl ether chain, or sphingosine skeleton and their retention by reversed-phase liquid chromatography, and its expression is t R C =k0+X1c 3 +X2c 2 +X3c+X4, where X1 is: k1+k2d+k3d 2 +k4d 3 +k5e+k6m+k7n+k8p+k9γ;

[0014] X2 is: k 10 +k 11 d+k 12 d 2 +k 13 d 3 +k 14 h+k 15 m+k 16 n+k 17 b+k 18 p+k 19 γ;

[0015] X3 is: k 20 +k 21 d+k 22 d 2 +k 23 d 3 +k 24 h+k 25 A+k 26 m+k 27 n+k 28 b+k 29 p+k 30 q+k 31 v+k 32 γ;

[0016] X4 is: k 33 a+k 34 d+k 35 d 2 +k 36 d 3 +k 37 e+k 38 f+k 39 h+k 40 A+k 41j+k 42 m+k 43 p+k 44 q+k 45 s+k 46 v+k 47 β+k 48 γ+k 49 δ+k 50 ε.

[0017] Preferably, in the quantitative relationship model, the total number of carbons c of the fatty acyl chain, ether chain, and / or alkenyl ether chain of the specific lipid compound standard ranges from 20 to 70, and the total number of carbon-carbon double bonds d on the fatty acyl chain ranges from 0 to 12. The quantitative relationship model constructed for lipid compounds containing two or more fatty acyl chains, ether chains, and / or alkenyl ether chains and their retention by reversed-phase liquid chromatography is expressed as t R C =k0+X1c 3 +X2c 2 +X3c+X4, where X1 is: k1+k2d 2 +k3e+k4p+k5γ;

[0018] X2 is: k6+k7d+k8d 2 +k9p+k 10 γ;

[0019] X3 is: k 11 +k 12 d+k 13 d 2 +k 14 p+k 15 q+k 16 v+k 17 γ+k 18 A;

[0020] X4 is: k 19 d 2 +k 20 f+k 21 j+k 22 p+k 23 q+k 24 s+k 25 v+k 26 w+k 27 γ+k 28 A.

[0021] Preferably, the quantitative relationship model is used to construct a quantitative relationship model between the structures of five major lipid classes and their retention in reversed-phase liquid chromatography. The specific lipid compound standards include specific free fatty acids, α-hydroxy fatty acids, β-hydroxy fatty acids, dicarboxylic acids, acylcarnitines and their hydroxylated modified compounds, monoglycerides, Sn1 and Sn2 type lysophospholipids and their alkenyl ether or / and ether forms, sphingosine and its modified compounds, specific sterols and cholesterol esters, specific diglycerides, triglycerides, glycerophosphatidylcholine, other glycerides and glycerophospholipids, specific dihydroceramides, ceramides and other sphingolipids; the specific... The number of carbon atoms (c) on the fatty acyl chain of the free fatty acid ranges from 9 to 30, and the number of carbon-carbon double bonds (d) on its fatty acyl chain ranges from 0 to 6; the number of carbon atoms (c) on the fatty acyl chain of the specific cholesterol ester ranges from 0 to 22, and the number of carbon-carbon double bonds (d) on its fatty acyl chain ranges from 0 to 6; the total number of carbon atoms (c) on the fatty acyl chain, ether chain, and / or alkenyl ether chain of the specific lipid compound ranges from 20 to 70, and the total number of carbon-carbon double bonds (d) on its fatty acyl chain ranges from 0 to 12; the total number of carbon atoms (c) on the sphingosine and fatty acyl chains of the specific sphingolipid compound ranges from 20 to 70, and the number of carbon-carbon double bonds (d) on its fatty acyl chain ranges from 0 to 4;

[0022] A model was constructed using multiple regression analysis; the five lipid classes include free fatty acids and their modified compounds, glycerides, glycerophospholipids, sphingolipids, and sterol lipids. Its expression is t... R C =k0+X1c 3 +X2c 2 +X3c+X4, k0 is a constant, and X1-X4 is composed as follows:

[0023] X1:k1+k2d+k3d 2 +k4d 3 +k5h+k6m+k7β+k8θ+k9μ;

[0024] X2:k 10 +k 11 b+k 12 d+k 13 d 2 +k 14 d 3 +k 15 e+k 16 h+k 17 m+k 18 n+k 19 p+k 20 q+k 21 β+k 22 θ+k 23 A;

[0025] X3:k24 +k 25 b+k 26 d+k 27 d 2 +k 28 d 3 +k 29 f+k 30 h+k 31 m+k 32 n+k 33 p+k 34 v+k 35 β+k 36 δ+k 37 θ+k 38 μ+k 39 A;

[0026] X4:k 40 a+k 41 b+k 42 d 2 +k 43 d 3 +k 44 f+k 45 h+k 46 j+k 47 m+k 48 s+k 49 w+k 50 y+k 51 α+k 52 δ+k 53 ε+k 54 θ+k 55 μ+k 56 ξ; where k0-k 56 These represent the weighting coefficients of different terms in the above equation.

[0027] Preferably, the quantitative relationship model, for a lipid subclass, uses the selected specific structural characteristic parameters as the total number of carbons c and the total number of carbon-carbon double bonds d of the fatty acyl chain to construct a general model for the quantitative relationship between the lipid subclass and its retention in reversed-phase liquid chromatography.

[0028] t R C =k0+k1d+k2d 2 +k3d 3 +(k4+k5d+k6d 2 +k7d 3 )c+(k8+k9d+k 10 d 2 +k 11 d 3 )c 2 +(k 12 +k 13d+k 14 d 2 +k 15 d 3 )c 3 The lipid subclasses include free fatty acids, cholesterol esters, triglycerides, and glycerophosphatidylcholine.

[0029] According to a second aspect of the present invention, a lipid retention time prediction system is also provided, which includes a data acquisition module and a retention time prediction module.

[0030] The data acquisition module is used to acquire specific structural feature parameters of lipid compounds, quantify them, and submit them as feature values ​​to the retention time prediction module. The specific structural feature parameters include one or more of the following: the number of carbons ε and the number of carbon-carbon double bonds π in the sphingosine skeleton; the number of carbons θ and the number of carbon-carbon double bonds ξ in the sterol skeleton; N-acetylgalactosamine λ; and / or one or more of the skeleton, characteristic groups, and residues contained in the lipid compound. The skeleton, characteristic groups, or residues are quantified according to the number of corresponding types of skeletons, characteristic groups, or residues contained in the lipid compound. If they are not contained, their parameters are defined as 0.

[0031] The retention time prediction module calculates the predicted theoretical retention time based on the received feature values ​​according to a built-in specific calculation model, and outputs the result according to the calculation result.

[0032] Preferably, the specific calculation model of the lipid retention time prediction system is the quantitative relationship model described in this invention.

[0033] Preferably, the lipid retention time prediction system is used to predict the theoretical retention time of five major lipid compounds, and the specific calculation model is as described in the quantitative relationship model of the present invention; the five major lipid compounds include free fatty acids and their modified compounds, glycerides, glycerophospholipids, sphingolipids and sterol lipids.

[0034] Preferably, the lipid retention time prediction system is used to predict the theoretical retention time of unknown sterol lipid compounds, and the specific calculation model is as described in the quantitative relationship model of the present invention.

[0035] The specific calculation model used to predict the theoretical retention time of unknown sphingolipid compounds containing two or more fatty acyl chains and a sphingosine skeleton is the quantitative relationship model described in this invention.

[0036] The specific calculation model used to predict the theoretical retention time of unknown lipid compounds containing a single fatty acyl chain, ether chain, alkenyl ether chain, or sphingosine skeleton is the quantitative relationship model described in this invention.

[0037] The specific calculation model used to predict the theoretical retention time of unknown lipid compounds containing two or more fatty acyl chains, ether chains, or / and alkenyl ether chains is the quantitative relationship model described in this invention.

[0038] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:

[0039] (1) Predictive accuracy: Compared with the patent CN120913704A, which requires the inclusion of dozens of variables to describe dozens of sterol structural features, the prediction model provided by this invention only requires four structural features to be constructed: the number of carbons c on the fatty acyl chain and the number of carbon-carbon double bonds d, the number of carbons θ on the sterol skeleton and the number of carbon-carbon double bonds ξ. The quantitative relationship model between the sterol and its ester lipid structure and its retention by reversed-phase liquid chromatography can be used to predict the theoretical retention time of more than a thousand unknown sterols and their esters. This model has the advantages of high coverage, simplicity and high efficiency.

[0040] (2) Low cost, high coverage and high precision: Compared with the existing technology, which requires the use of thousands of lipid compound standards or endogenous lipid compounds with high qualitative accuracy, the prediction model provided by this invention can be constructed based on only a small number of specific lipid standards. The model constructs a quantitative relationship model between the lipid structure and its retention by reversed-phase liquid chromatography, using the experimental retention time data of lipid standards as the dependent variable and multiple screened structural feature parameters as independent variables. This model can predict the theoretical retention time of tens of thousands of unknown lipid compounds. Attached Figure Description

[0041] Figure 1 The prediction performance of the QSRR model is based on lipid subclass standards of fatty acids, cholesterol esters, triglycerides, glycerol phosphatidylcholine, and ceramide lipids. Figure 1 In the table, A represents the prediction performance of the QSRR model built based on fatty acid standards, B represents the prediction performance of the QSRR model built based on cholesterol ester standards, C represents the prediction performance of the QSRR model built based on triglyceride standards, D represents the prediction performance of the QSRR model built based on glycerophosphatidylcholine standards, and E represents the prediction performance of the QSRR model built based on ceramide standards. Two-thirds of the standards (black circles) are used as the training set for model construction, and one-third of the standards (red circles) are used as the test set for model validation.

[0042] Figure 2 It is the experimental retention time (t) R E ) and retention time (t) calculated by the QSRR model R C The correlation between ) Figure 2In the first part, A is a lipid compound containing a fatty acyl chain, ether chain, alkenyl ether chain, or sphingosine skeleton. R E With t R C The correlation, B is the t of sterol esters. R E With t R C The correlation, C is a lipid compound containing two or more fatty acyl chains t R E With t R C The correlation, D is a lipid compound containing two or more fatty acyl chains and a sphingosine backbone. R E With t R C The correlation is given by E, representing other endogenous lipid compounds predicted in high-resolution mass spectrometry (HRMS). Two-thirds of the standards (black circles) were used as the training set to build the model; one-third of the standards (red circles) were used as the test set for model validation. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0044] Current methods for constructing QSRR calculation models for the five major lipid classes heavily rely on thousands of lipid compounds, resulting in high costs associated with manual qualitative review. Furthermore, only a few hundred commercially available lipid compound standards are available, with only free fatty acids, α- or β-hydroxy fatty acids (2HOFA or 3HOFA), acylcarnitine, cholesterol esters, glycerophosphatidylcholine, glycerophosphatidylethanolamine, and ceramide subclasses containing more than ten standards. Therefore, how to construct optimal QSRR models for the five major lipid classes using hundreds of lipid compounds (prioritizing lipid standards) as training sets, and how to predict the high-precision theoretical retention times of tens of thousands of unknown lipid compounds, is a critical technical challenge that most experimental platforms urgently need to address.

[0045] The QSRR general model construction method disclosed in patent CN120913704A uses only a binary classification approach to quantify seven different sterol skeletons with 0 or 1 values, i.e., seven structural feature variables, including cholesterol, 22,23-ergosterol, campesterol, 24α-ethylcholesterol, stigmasterol, 3β-cholesterol skeleton, and 24β-ethylcholesterol and its esters. However, this method of describing sterol structural features requires constructing dozens of variables to describe dozens of sterols, leading to excessive complexity in the QSRR general model formula. Therefore, how to select a few specific structural features to describe dozens of sterols and construct a high-coverage, concise, and efficient QSRR general model for sterols and their esters is another key technical problem that urgently needs to be solved.

[0046] To address the aforementioned deficiencies in existing technologies, this invention uses 267 specific lipid compounds (prioritizing lipid standards) as research objects. Based on the characteristic structural parameters of lipid compounds screened out by patent CN120913704A, new characteristic structural parameters are added, including the number of carbons ε and the number of carbon-carbon double bonds π in the sphingosine skeleton, the number of carbons θ and the number of carbon-carbon double bonds ξ in the sterol skeleton (these four parameters are classified as skeleton independent variables), and N-acetylgalactosamine λ (classified as characteristic group independent variables). After quantifying these characteristic structural parameters, variable selection is completed through stepwise regression and full subset regression, and 10-fold cross-validation is used to determine the adjusted goodness of fit (R0). 2 The optimal QSRR model is selected based on the prediction error.

[0047] The characteristic structural parameters screened in this invention include one or more of the following: the number of carbons ε and the number of carbon-carbon double bonds π in the sphingosine skeleton; the number of carbons θ and the number of carbon-carbon double bonds ξ in the sterol skeleton; and N-acetylgalactosamine λ. Furthermore, the characteristic structural parameters used to construct QSRR models for the five major lipid classes (free fatty acids and their modified compounds, glycerides, glycerophospholipids, sphingolipids, and sterol lipids) also include the total number of carbon atoms c in the fatty acyl chain, ether chain, alkenyl ether chain, and / or the sphingosine skeleton; the total number of carbon-carbon double bonds d; the types and numbers of skeletons contained in the lipid compound; and the characteristic groups, residue types, and residue numbers in the lipid compound. The skeletal structural features include the number of glycerol skeletons (v) and the number of fatty acyl chains, ether chains, and / or alkenyl ether chains (α, β, and γ) connected at the sn1, sn2, and sn3 positions in glycerol lipids and glycerophospholipids; the number of sphingosine skeletons (μ) and the number of carbons (ε) and carbon-carbon double bonds (π) in sphingolipids; and the number of carbons (θ) and carbon-carbon double bonds (ξ) on the sterol skeleton in sterol lipids. These two new sterol skeletal features can describe the structural characteristics of different sterol types, including cholesterol, coprosterol, 7-encholesterol, 7-dehydrocholesterol, sterol, campesterol, brassosterol, ergosterol, β-sitosterol, and β-stigmasterol.

[0048] The characteristic groups include α or β-hydroxylation modification (h and n) on the fatty acyl chain, peroxidation modification z on the fatty acyl chain, fatty acylation modification a of hydroxylated fatty acids or hexoses, ether bond j or olefinic ether bond f on the first carbon of the fatty alcohol chain, hydroxylation modification of sphingosine (i.e., phytosphingosine) δ, phosphate p, hexose w, glucuronic acid g, trimethylhomocysteine ​​t, dimethylethanolamine r, sialic acid y, N-acetylgalactosamine (commonly known as acetylgalactosamine) λ, sulfonated group φ, galactose sulfonic acid ψ, ethanolamine e, choline q, inositol A, serine s, methylethanolamine x, and ethanol u.

[0049] The residues include carnitine residue m, carboxylic acid modification b at the ω position of the fatty acyl chain, sulfonyl group, and sulfonated galactose residue ψ.

[0050] The experimental retention time t of a specific lipid compound R E Using the selected characteristic structural parameters as independent variables, a nonlinear regression mathematical model is constructed to establish a new quantitative relationship between chemical structure and retention time. Specifically, the LM() function is used to obtain a general QSRR prediction model for five major lipid compounds (free fatty acids and their modified compounds, glycerides, glycerophospholipids, sphingolipids, and sterol lipids). Its expression is t R C =k0+X1c 3 +X2c 2 +X3c+X4 (where k0 is a constant), where X1, X2, X3, and X4 are different functions composed of the aforementioned characteristic parameters. Its simplified expression is: , This indicates the specific structural features selected; This represents the theoretical retention time of the lipid compound. The backbone, characteristic groups, or residues are quantified numerically according to the number of corresponding types of backbones, characteristic groups, or residues present in the lipid compound; if none are present, their parameters are defined as 0.

[0051] Based on this, the present invention provides a quantitative relationship model between lipid structure and its retention in reversed-phase liquid chromatography, which is constructed based on a specific lipid compound according to the following method:

[0052] Using the structural features of lipid compounds as candidate independent variables, the optimal independent variables and their combinations are screened out through stepwise regression modeling as specific structural features. The screened specific structural features include one or more of the following: the number of carbons ε and the number of carbon-carbon double bonds π in the sphingosine skeleton, the number of carbons θ and the number of carbon-carbon double bonds ξ in the sterol skeleton, N-acetylgalactosamine λ, and / or one or more of the skeleton, characteristic groups and residues contained in the lipid compound.

[0053] t of specific lipid compound standardsR E Using the selected specific structural features as independent variables and employing the LM() function, a quantitative relationship model was established between lipid structure and its retention in reversed-phase liquid chromatography. The constructed quantitative relationship model is as follows: , This indicates the specific structural features selected. This indicates the theoretical retention time of lipid compounds.

[0054] For example, the structural characteristic parameters of specific lipid compounds can be quantified, and specific structural characteristic parameters can be screened out through stepwise regression and full subset regression. The screened specific structural characteristic parameters include one or more of the following: the number of carbons ε and the number of carbon-carbon double bonds π in the sphingosine skeleton; the number of carbons θ and the number of carbon-carbon double bonds ξ in the sterol skeleton; N-acetylgalactosamine λ; and / or one or more of the skeleton, characteristic groups, and residues contained in lipid standards. The retention time of the specific lipid compound is determined by reversed-phase liquid chromatography. Using the selected specific structural feature parameters as independent variables, a quantitative relationship model between lipid structure and its retention in reversed-phase liquid chromatography (RP-LC) was constructed using the LM() function. The obtained quantitative relationship model between lipid structure and its retention in RPL is as follows: ; This represents the theoretical retention time of lipid compounds and can be used to predict the theoretical retention time of unknown lipid compounds.

[0055] In some embodiments, the specific lipid compounds include specific free fatty acid standards, 1-10 α-hydroxy fatty acids, 1-10 β-hydroxy fatty acids, 1-10 dicarboxylic acids, 1-20 acylcarnitines and their hydroxylated modified compounds, 1-5 monoglycerides, 1-5 sn1 and sn2 type lysophosphatidic acids, 3-5 sn1 and sn2 type lysophosphatidylcholine and their alkenyl ether form or / and ether form, and 1-5 sn1 and sn2 type lysophosphatidylcholine compounds. The standards for n2 type lysophosphatidylethanolamine and its alkenyl ether form or / and ether form, 1 to 5 sn1 and sn2 type lysophosphatidylinositols, 1 to 5 sn1 and sn2 type lysophosphatidylserines, 1 to 5 sn1 and sn2 type lysophosphatidylglycerols, 1 to 5 sphingosines and their modified compounds; the specific free fatty acid has a carbon atom count c in the range of 9 to 30 and a carbon-carbon double bond count d in the range of 0 to 6.

[0056] A quantitative model was constructed to establish the relationship between lipid compounds containing a single fatty acyl chain, ether chain, alkenyl ether chain, or sphingosine skeleton and their retention by reversed-phase liquid chromatography. The expression is t. R C =k0+X1c 3 +X2c 2+X3c+X4 (where k0 is a constant), where X1, X2, X3, and X4 are different functions composed of the aforementioned characteristic parameters. For example:

[0057] X1:k1+k2d+k3d 2 +k4d 3 +k5e+k6m+k7n+k8p+k9γ;

[0058] X2:k 10 +k 11 d+k 12 d 2 +k 13 d 3 +k 14 h+k 15 m+k 16 n+k 17 b+k 18 p+k 19 γ;

[0059] X3:k 20 +k 21 d+k 22 d 2 +k 23 d 3 +k 24 h+k 25 A+k 26 m+k 27 n+k 28 b+k 29 p+k 30 q+k 31 v+k 32 γ;

[0060] X4:k 33 a+k 34 d+k 35 d 2 +k 36 d 3 +k 37 e+k 38 f+k 39 h+k 40 A+k 41 j+k 42 m+k 43 p+k 44 q+k 45 s+k 46 v+k 47 β+k 48 γ+k 49 δ+k 50 ε;

[0061] In the formula k iThe equation coefficients corresponding to the QSRR model constructed for multiple regression analysis.

[0062] For example, the specific lipid compounds mentioned are specific free fatty acid standards, 1-10 α-hydroxy fatty acids, 1-10 β-hydroxy fatty acids, 1-10 dicarboxylic acids, 1-20 acylcarnitines and their hydroxylated modified compounds, 1-5 monoglycerides, 1-5 sn1 and sn2 type lysophosphatidic acids, 3-5 sn1 and sn2 type lysophosphatidylcholine and their alkenyl ether form or / and ether form, 1-5 sn1 and sn2 type lysophosphatidylethanolamine and their alkenyl ether form or / and ether form, 1-5 sn1 and sn2 type lysophosphatidylinositol, 1-5 sn1 and sn2 type lysophosphatidylserine, and 1-5 sn1 and sn2 type lysophosphatidylglycerol standards. A quantitative relationship model was constructed between lipid compounds containing a single fatty acyl chain and their retention in reversed-phase liquid chromatography. The expression is t. R C =k0+X1c 3 +X2c 2 +X3c+X4 (where k0 is a constant), where X1, X2, X3, and X4 are different functions composed of the aforementioned characteristic parameters. For example:

[0063] X1:k1+k2d+k3d 2 +k4d 3 +k5e+k6m+k7n+k8p+k9γ;

[0064] X2:k 10 +k 11 d+k 12 d 2 +k 13 d 3 +k 14 h+k 15 m+k 16 n+k 17 b+k 18 p+k 19 γ;

[0065] X3:k 20 +k 21 d+k 22 d 2 +k 23 d 3 +k 24 h+k 25 A+k 26 m+k 27 n+k 28 b+k 29 p+k 30 q+k 31 v+k 32 γ;

[0066] X4:k 33 a+k 34 d+k 35 d 2 +k 36 d 3 +k 37 e+k 38 f+k 39 h+k 40 A+k 41 j+k 42 m+k 43 p+k 44 q+k 45 s+k 46 v+k 47 β+k 48 γ; where k i The equation coefficients corresponding to the QSRR model constructed for multiple regression analysis.

[0067] In some embodiments, the specific lipid compounds include specific cholesterol ester standards, coprosterol, 7-encholesterol, 7-dehydrocholesterol, sterols, campesterol, brassosterol, ergosterol, β-sitosterol, β-stigmasterol, other sterols, and their ester standards; the specific cholesterol esters have a fatty acyl chain with carbon number c ranging from 0 to 22 and a fatty acyl chain with carbon-carbon double bonds d ranging from 0 to 6. The quantitative relationship between the constructed sterol lipid compounds and their retention by reversed-phase liquid chromatography is expressed as t. R C =k0+k1θ+k2ξ+k3c+k4dc+k5c 2 +k6c 3 , where k0-k6 represent the weight coefficients of different terms in the above equation.

[0068] In some embodiments, the specific lipid compounds include 6-15 diglycerides, 1-10 galactosyldiglycerides, 1-10 digalactosyldiglycerides, 10-20 triglycerides, 1-5 phosphatidic acids, 15-30 phosphatidylcholine and its alkenyl ether form or / and ether form, 2-15 phosphatidylethanolamine and its alkenyl ether form or / and ether form, 1-5 phosphatidylinositol, 1-5 phosphatidylserine, and 1-5 phosphatidylglycerol; the total number of carbons c in the fatty acyl chain, ether chain, and / or alkenyl ether chain of the specific lipid compound ranges from 20 to 70, and the total number of carbon-carbon double bonds d on its fatty acyl chain ranges from 0 to 12; the quantitative relationship between the constructed lipid compound containing two or more fatty acyl chains, ether chains, and / or alkenyl ether chains and its retention by reversed-phase liquid chromatography is expressed as t. R C =k0+X1c 3 +X2c 2+X3c+X4 (where k0 is a constant), where X1, X2, X3, and X4 are different functions composed of the aforementioned characteristic parameters. For example:

[0069] X1:k1+k2d 2 +k3e+k4p+k5γ;

[0070] X2:k6+k7d+k8d 2 +k9p+k 10 γ;

[0071] X3:k 11 +k 12 d+k 13 d 2 +k 14 p+k 15 q+k 16 v+k 17 γ+k 18 A;

[0072] X4:k 19 d 2 +k 20 f+k 21 j+k 22 p+k 23 q+k 24 s+k 25 v+k 26 w+k 27 γ+k 28 A; where k0-k 28 These represent the weighting coefficients of different terms in the above equation.

[0073] In some embodiments, the specific lipid compounds include 10-30 specific ceramides and dihydroceramides, 1-5 hydroxylated fatty acyl phytosphingosine, 1-5 hydroxylated ceramides, 1-5 fatty acyl phytosphingosine, 1-5 fatty acyl-O-hydroxylated fatty acyl phytosphingosine, 1-5 fatty acyl-O-hydroxylated fatty acyl sphingosine, 1-5 fatty acyl sphingosine phosphate ethanolamine, 1-5 sphingomyelins, 1-10 hexosylceramides (including monohexosylceramides, disacylceramides, and polyhexosylceramides), and 1-10 gangliosides (including GM1, GM2, and GM3); the total number of carbons (c) on the sphingosine and fatty acyl chains of the specific sphingolipid compounds ranges from 20 to 70, and the number of carbon-carbon double bonds (d) on their fatty acyl chains ranges from 0 to 4; a quantitative relationship model is constructed between sphingolipid compounds containing two or more fatty acyl chains and a sphingosine skeleton and their retention by reversed-phase liquid chromatography. The expression is t. R C =k0+k1a+k2h+k3d+k4d 2+k5p+k6w+k7y+k8δ+k9c+k 10 pc+k 11 c 2 +k 12 c 3 ;where, k0-k 12 These represent the weighting coefficients of different terms in the above equation.

[0074] In some embodiments, the specific lipid compounds include 39 FFAs, 15 ACars, 6 ACar-OHs, 9 2HOFAs, 9 3HOFAs, 15 DCAs, 9 DAGs, 15 TAGs, 6 MGDGs, 9 DGDGs, 2 MAGs, 21 PCs, 4 LPCs, 1 LPC-P, 12 PEs, 3 PAs, 3 PGs, 3 PIs, 2 LPAs, 2 PSs, 14 CEs, and 10 sterol standards, used to construct a general prediction model for the retention time of five major lipid classes.

[0075] The experimental retention time t of lipid standards R E Using the selected characteristic structural parameters as independent variables, a general prediction model for the retention time of five major lipid compounds was constructed, with t as the dependent variable. R C =k0+X1c 3 +X2c 2 +X3c+X4 (where k0 is a constant), where X1, X2, X3, and X4 are derived from the aforementioned structural characteristic parameters ε, π, θ, ξ, d, d 2 d 3 Different functions composed of multiple characteristic parameters from a, b, c, d, e, f, g, h, j, m, n, p, q, r, s, t, u, v, w, x, y, α, β, γ, μ, A, φ, ψ, ε, π, δ, λ. For example:

[0076] X1:k1+k2d+k3d 2 +k4d 3 +k5h+k6m+k7β+k8θ+k9μ;

[0077] X2:k 10 +k 11 b+k 12 d+k 13 d 2 +k 14 d 3 +k 15 e+k 16 h+k 17 m+k 18 n+k 19 p+k 20 q+k21 β+k 22 θ+k 23 A;

[0078] X3:k 24 +k 25 b+k 26 d+k 27 d 2 +k 28 d 3 +k 29 f+k 30 h+k 31 m+k 32 n+k 33 p+k 34 v+k 35 β+k 36 δ+k 37 θ+k 38 μ+k 39 A;

[0079] X4:k 40 a+k 41 b+k 42 d 2 +k 43 d 3 +k 44 f+k 45 h+k 46 j+k 47 m+k 48 s+k 49 w+k 50 y+k 51 α+k 52 δ+k 53 ε+k 54 θ+k 55 μ+k 56 ξ; where k i The equation coefficients corresponding to the QSRR model constructed for multiple regression analysis.

[0080] The five major lipid categories include free fatty acids and their modified compounds (such as acylcarnitine, hydroxy fatty acids and dicarboxylic acids), glycerides, glycerophospholipids, sphingolipids and sterol lipids.

[0081] In some embodiments, for a lipid subclass, specific structural characteristic parameters are selected as the total number of carbons c and the total number of carbon-carbon double bonds d of the fatty acyl chain. A quantitative relationship model between the lipid subclass and its retention in reversed-phase liquid chromatography is constructed: t R C =k0+k1d+k2d 2 +k3d 3 +(k4+k5d+k6d 2+k7d 3 )c+(k8+k9d+k 10 d 2 +k 11 d 3 )c 2 +(k 12 +k 13 d+k 14 d 2 +k 15 d 3 )c 3 The lipid subclasses include free fatty acids, cholesterol esters, triglycerides, and glycerophosphatidylcholine.

[0082] The following are examples.

[0083] In this implementation case, 267 lipid compound standards were prepared into standard stock solutions at concentrations ranging from 1 to 100 μmol / L, and further prepared into lipid mixed standards at appropriate ratios for subsequent experiments on the retention time t of the 267 lipid standards. R E Acquisition. Simultaneously, typical biological matrices such as human bodily fluids (plasma, urine), animal tissues (mouse heart, liver, brain, feces, kidneys, and lungs), bacteria (Escherichia coli), fungi (yeast), lower plants (male and female lichens), and higher plants (Arabidopsis thaliana) were used to extract lipidomes using the Matyash method ("Lipid extraction by methyl-tert-butylether for high-throughput lipidomics") and the Sarafian method ("Objective set of criteria for optimization of sample preparation procedures for ultra-high-throughput untargeted blood plasma lipid profiling by ultra-performance liquid chromatography-mass spectrometry"). These lipidomes were then used for subsequent non-targeted lipidomics qualitative analysis of endogenous lipid compounds, obtaining verifiable data on thousands of lipid compounds.

[0084] The following implementation examples used a Shimadzu UPLC system (Kyoto, Japan) to obtain tRE data for the aforementioned lipid compound standards and typical biological matrices. An Agilent ZORBAX Eclipse Plus C18 (2.1 mm * 100, 1.8 μm) was used, with a column temperature of 50 °C and a flow rate of 0.35 mL / min. Mobile phases A and B were water / methanol / acetonitrile solution (1:1:1, V / V / V, 1 mM ammonium acetate, 0.1% formic acid) and isopropanol / acetonitrile solution (9:1, V / V / V, 10 mM ammonium acetate, 0.05% formic acid), respectively. The corresponding gradient elution was as follows: during the 0-1 min period, the volume percentage of mobile phase B remained at 5%; during the 1-1.5 min period, the volume percentage of mobile phase B increased from 5% to 25%; during the 1.5-6 min period, the volume percentage of mobile phase B increased from 25% to 90%; during the 6-8 min period, the volume percentage of mobile phase B increased from 90% to 92%; during the 8-9 min period, the volume percentage of mobile phase B increased from 92% to 99%; and during the 9-11 min period, the volume percentage of mobile phase B remained at 99%.

[0085] Mass spectrometry data were acquired using a SCIEX Zeno TOF 7600 and X500 RTOF system with the following parameters: DuoSpray ion source temperature 550℃; curtain gas (CUR) 35 psi; nebulizing gas (GS1) and drying gas (GS2) both 55 psi; cluster potential (DP) ±70 V; and ion spray voltage set to 5500 V or -4500 V. The mass ranges for TOF-MS and TOF-MS / MS were set to 100-2000 Da and 50-2000 Da, respectively, with corresponding collision energies (CE) of 10 V and 45 V (±15 V). Raw data acquisition and processing were performed using OS (v1.7, SCIEX, Chromos, Singapore), combined with MS-DIAL software for qualitative analysis of endogenous lipid compounds to obtain corresponding retention times and mass spectrometry data.

[0086] The process of constructing the quantitative relationship (QSRR) model between lipid chemical structure and reversed-phase chromatography retention time (R language code) is as follows: First, based on the standard t in different cases... R E The dataset is used as the dependent variable, and the structural features of the lipid compounds involved are used as candidate independent variables, including the number of carbons ε on the sphingosine skeleton, the number of carbon-carbon double bonds π on the sphingosine skeleton, the number of carbons θ on the sterol skeleton in sterol lipids, the number of carbon-carbon double bonds ξ on the sterol skeleton in sterol lipids, the total number of carbon-carbon double bonds d of fatty acyl chains, ether chains, alkenyl ether chains, and / or the sphingosine skeleton, and d. 2 d 31. Fatty acylation modification of hydroxylated fatty acids or hexoses: a. Carboxylic acid modification at the ω-position of fatty acyl chains; b. Total carbon atoms of fatty acyl chains, ether chains, alkenyl ether chains, and / or sphingosine skeletons; c. Ethanolamine; e. Alkenyl ether bond; f. Glucuronic acid; g. Hydroxylation modification at the α-position of fatty acyl chains; h. Ether bond at the 1st carbon of fatty alcohol chains; j. Carnitine; m. Hydroxylation modification at the β-position of fatty acyl chains; n. Phosphate; p. Choline; q. Dimethylethanolamine; r. Serine; s. Trimethylhomocysteine; t. Ethanol; u. Number of glycerol skeletons; v. Hexose; w. Methylethanolamine; x. Sialic acid; y. Glycerol skeleton; sn1. The number of fatty acyl chains, ether chains, and / or alkenyl ether chains connected at sn2 and sn3 (α, β, and γ), the number of sphingosine skeletons μ, inositol A, sulfonated groups φ, galactose sulfonic acid ψ, sphingosine hydroxylation modification i.e., phytosphingosine δ, and acetylgalactosamine λ; where, for each subclass of the QSRR model, only the total carbon atoms c of fatty acyl chains, ether chains, alkenyl ether chains, and / or sphingosine skeleton, and the total number of carbon-carbon double bonds d are considered; the primary model M0 is established using the LM() function in the R language package, and the significance of the independent variables and their combinations in different terms is checked using summary().

[0087] Furthermore, stepwise regression modeling is performed using the MASS function package in R. The stepAIC() function is used to optimize the initial model M0 to obtain the preferred independent variables and their combinations. The direction parameter of this function is set to "both", while other parameters are left at their default settings. Using standard product t... R E Using the dataset as the dependent variable and the above-mentioned preferred independent variables and their combinations as independent variables, the LM() function is used to establish a better model M1.

[0088] Then, the `crossval()` function in the R package is called to perform 10 cross-validations of the model (k=10) to obtain the optimal independent variables and their combinations. Using standard product t... R E Using the dataset as the dependent variable and the selected independent variables and their combinations as independent variables, the optimal model M2 is established using the LM() function; the simplified independent variables and their combinations are viewed using summary(), and the coefficient values ​​of the corresponding terms are obtained.

[0089] Finally, the structural characteristic parameters of other unknown lipid compounds were substituted into the above optimal model M2 to calculate the corresponding theoretical retention times.

[0090] Example 1: Quantitative Relationship Model between Single Lipid Subclasses and Their Retention in Reversed-Phase Liquid Chromatography

[0091] (1) Quantitative relationship model between free fatty acid (FFA) subclasses and their retention in reversed-phase liquid chromatography

[0092] Thirty-nine specific free fatty acid standards were selected from 267 lipid compound standards and divided into training and test sets at a 2:1 ratio. The specific free fatty acid (FFA) standards in the training set have a carbon number (c) range of 9–30 on the fatty acyl chain and a carbon-carbon double bond number (d) range of 0–6 (see Table 1-1), covering all FFA structural features involved in non-targeted or targeted lipidomics analysis methods.

[0093] For each subclass, the specific structural characteristic parameters are only S1 and S2, namely the number of carbons (c) and the number of carbon-carbon double bonds (d) of the fatty acyl chain. Based on 26 FFA standards in the training set, using their experimental retention time data as the dependent variable and the two structural characteristic parameters S1 and S2 (the total number of carbons (c) and the total number of carbon-carbon double bonds (d) of different fatty acyl chains and sphingosine skeletons) as independent variables, a general prediction model for free fatty acid retention time (FFA-QSRR) was constructed. The expression of the FFA-QSRR general prediction model was obtained. =k0+X1c 3 +X2c 2 +X3c+X4, specifically:

[0094] t R C =k0+k1d+k2d 2 +k3d 3 +(k4+k5d+k6d 2 +k7d 3 )c+(k8+k9d+k 10 d 2 +k 11 d 3 )c 2 +(k 12 +k 13 d+k 14 d 2 +k 15 d 3 )c 3 (k0~k) 15 (These are the equation coefficients).

[0095] Table 1-1 Training Set of Free Fatty Acid Standards

[0096]

[0097] The predictive ability of the above model was validated based on 13 FFA standards and biological samples in the test set. The results are shown in Tables 1-2 and 2-3. Figure 1 As shown in Figure A.

[0098] Table 1-2 Performance of the FFA-QSRR prediction model constructed based on test set and biological sample validation

[0099]

[0100] From Table 1-2 and Figure 1 As shown in A, the FFA-QSRR prediction model has good predictive ability, i.e., the correlation R between the experimental retention time and the theoretical retention time is good. 2 >0.99, and the corresponding retention time deviation Δt of the 13 FFA standards in the test set. R All were less than 0.13 min. Furthermore, based on this FFA-QSRR prediction model, this implementation case also identified 7 endogenous FFAs in the aforementioned 7 typical biological matrices, with corresponding retention time deviations Δt. R All values ​​were less than 0.04 min. Therefore, this implementation case demonstrates that the QSRR prediction model constructed based on 26 specific FFA standards can achieve low-cost, high-precision, and high-coverage calculation of the theoretical retention time of FFA, which can be used to meet the qualitative analysis needs of all FFAs in non-targeted or targeted lipidomics.

[0101] (2) Quantitative relationship model between cholesterol ester (ChE) subclasses and their retention in reversed-phase liquid chromatography

[0102] Cholesterol esters, as the main representatives of sterol lipids in lipidomics research, currently only contain 15 common commercially available standards, which is insufficient for constructing traditional t-models. R C ~f(c) and t R C The required number of standard samples for ~f(d) is determined. In this implementation case, 10 ChE standard samples are used as training research objects (see Table 2-1) to construct a prediction model similar to FFA-QSRR.

[0103] This implementation case selected 15 cholesterol ester standards from 267 lipid compound standards, dividing them into training and test sets at a 2:1 ratio. The training set contains 10 specific cholesterol ester (ChE) standards with a range of 0–22 carbons (c) and 0–6 carbon-carbon double bonds (d) on the fatty acyl chains (including free cholesterol, cholesterol esters containing short, medium, and long fatty acyl chains). This covers all ChE structural features involved in non-targeted and targeted lipidomics analysis methods.

[0104] For each subclass, the specific structural characteristic parameters are only S1 and S2, namely the number of carbons c in the fatty acyl chain and the number of carbon-carbon double bonds d. Based on 10 specific ChE standards in the training set, their experimental retention times t are used... R E Using data as the dependent variable and c and d as independent variables, a general prediction model for cholesterol ester retention time (ChE-QSRR) was constructed through nonlinear regression analysis. Its expression is as follows:

[0105] tR C =k0+k1d+k2d 2 +k3d 3 +(k4+k5d+k6d 2 +k7d 3 )c+(k8+k9d+k 10 d 2 +k 11 d 3 )c 2 +(k 12 +k 13 d+k 14 d 2 +k 15 d 3 )c 3 (k0~k) 15 (These are the equation coefficients).

[0106] Table 2-1 Performance of the Cholesterol Ester QSRR Prediction Model

[0107]

[0108] The predictive ability of the above model was validated based on five ChE standards and biological samples in the test set. The results are shown in Table 2-2 and... Figure 1 As shown in B.

[0109] Table 2-2 Performance of the Cholesterol Ester QSRR Prediction Model

[0110]

[0111] From Table 2-2 and Figure 1 As shown in B, the ChE-QSRR prediction model has good predictive ability, i.e., the correlation R between the experimental retention time and the theoretical retention time is good. 2 >0.99, and the corresponding retention time deviation Δt of the five ChE standards in the test set. R All were less than 0.11 min. Furthermore, based on this ChE-QSRR prediction model, this implementation case also identified the theoretical retention times of 10 unknown endogenous ChEs, and the corresponding retention time deviations Δt. R All values ​​were less than 0.15 min. Therefore, this case study demonstrates that the QSRR prediction model constructed based on 10 specific ChE standards can achieve low-cost, high-precision, and high-coverage calculation of theoretical ChE retention time, meeting the qualitative analysis needs of unknown ChEs in non-targeted or targeted lipidomics, especially in setting the acquisition time window (Δt) in multiple response monitoring (MRM) scanning mode. R All less than 0.5 min).

[0112] (3) Quantitative relationship model between triglyceride (TAG) subclasses and their retention in reversed-phase liquid chromatography

[0113] Triglycerides are among the most common glycerol lipids in lipidomics research. Their glycerol backbone can be linked with fatty acyl chains of varying carbon numbers, carbon-carbon double bond numbers, and positions at the sn1, sn2, and sn3 positions. Therefore, theoretically, there are approximately tens of thousands of triglyceride compounds in biological matrices. However, currently only a few dozen common commercially available standards are available, which is insufficient for constructing traditional model t-tests. R C ~f(c) and t R C The required number of standard samples for ~f(d) is determined. This implementation case uses 15 TAG standard samples as training research objects to construct a prediction model similar to FFA-QSRR.

[0114] This implementation case selected 15 triglyceride (TAG) standards from 267 lipid compound standards, dividing them into training and test sets at a 2:1 ratio. The training set included 10 specific TAG standards with a total number of carbons (c) on the total fatty acyl chain and a total number of carbon-carbon double bonds (d) ranging from 48 to 56 and 0 to 9 respectively (see Table 3-1), covering all TAG structural features involved in both non-targeted and targeted lipidomics analysis methods.

[0115] For each subclass, the specific structural characteristic parameters are only S1 and S2, namely the total number of carbons c and the total number of carbon-carbon double bonds d of different aliphatic acyl chains. Based on 10 specific TAG standards in the training set, their experimental retention times t are used... R E Using data as the dependent variable and c and d as independent variables, a general prediction model for triglyceride retention time (TAG-QSRR) was constructed through nonlinear regression analysis. Its expression is:

[0116] t R C =k0+k1d+k2d 2 +k3d 3 +(k4+k5d+k6d 2 +k7d 3 )c+(k8+k9d+k 10 d 2 +k 11 d 3 )c 2 +(k 12 +k 13 d+k 14 d 2 +k 15 d 3 )c 3 (k0~k) 15 (These are the equation coefficients).

[0117] Table 3-1 Performance of the QSRR Prediction Model for Triglycerides

[0118]

[0119] The predictive ability of the above model was verified based on five TAG standards in the test set and the validation set. The results are shown in Table 3-2 and... Figure 1 As shown in C.

[0120] Table 3-2 Performance of the QSRR Prediction Model for Triglycerides

[0121]

[0122] From Table 3-2 and Figure 1 As shown in C, the TAG-QSRR prediction model has good predictive ability, i.e., the correlation R between the experimental retention time and the theoretical retention time is good. 2 >0.99, and the corresponding retention time deviation Δt of the five TAG standards in the test set. R All were less than 0.02 min. Furthermore, based on this TAG-QSRR prediction model, this implementation case also identified the theoretical retention times of 122 unknown endogenous TAGs, and the corresponding retention time deviations Δt. R All were less than 0.18 min (the theoretical retention times of some unknown endogenous TAGs are shown in Table 3-2). Therefore, this case study demonstrates that the QSRR prediction model constructed based on 10 specific TAG standards can achieve low-cost, high-precision, and high-coverage calculation of TAG theoretical retention times, meeting the qualitative analysis needs of unknown TAGs in non-targeted or targeted lipidomics, especially in the acquisition time window setting (Δt) of multiple reaction monitoring scanning mode. R All less than 0.5 min).

[0123] (4) Quantitative relationship model between glycerol phosphatidylcholine (PC) subclasses and their retention in reversed-phase liquid chromatography

[0124] Phosphatidylcholine (PC) is one of the most common glycerophospholipids in lipidomics research. The sn1 and sn2 positions of its glycerol backbone can be linked to fatty acyl chains with varying numbers of carbons, carbon-carbon double bonds, and their positions, theoretically resulting in approximately a thousand PCs in biological matrices. However, currently only a few dozen common commercially available standards are available, which is insufficient for constructing traditional models. R C ~f(c) and t R C The required number of standard samples for ~f(d) is determined. In this implementation case, 21 PC standard samples are used as training research objects (see Table 4-1) to construct a prediction model similar to FFA-QSRR.

[0125] This implementation case selected 21 PC standards from 267 lipid compound standards and divided them into training and testing sets at a 2:1 ratio. The training set of 14 PC standards had a total number of carbons (c) and carbon-carbon double bonds (d) on the total fatty acyl chain ranging from 25 to 44 and 0 to 6, respectively, covering all PC structural features involved in non-targeted or targeted lipidomics analysis methods.

[0126] For each subclass, the specific structural characteristic parameters are only S1 and S2, namely the total number of carbons c and the total number of carbon-carbon double bonds d of different aliphatic acyl chains. Based on 14 specific PC standards in the training set, their experimental retention times t are used... R E With data as the dependent variable and c and d as independent variables, a general predictive model for phosphatidylcholine retention time (PC-QSRR) was constructed using nonlinear regression analysis. Its expression is:

[0127] t R C =k0+k1d+k2d 2 +k3d 3 +(k4+k5d+k6d 2 +k7d 3 )c+(k8+k9d+k 10 d 2 +k 11 d 3 )c 2 +(k 12 +k 13 d+k 14 d 2 +k 15 d 3 )c 3 (k0~k) 15 (These are the equation coefficients).

[0128] Table 4-1 Performance of the glycerophosphatidylcholine QSRR prediction model

[0129]

[0130] The predictive ability of the above model was verified based on seven TAG standards in the test set and the validation set. The results are shown in Table 4-2 and... Figure 1 As shown in D.

[0131] Table 4-2 Performance of the glycerophosphatidylcholine QSRR Prediction Model

[0132]

[0133] From Table 4-2 and Figure 1As can be seen from D, the PC-QSRR prediction model has good predictive ability, i.e., the correlation R between the experimental retention time and the theoretical retention time is good. 2 >0.99, and the corresponding retention time deviation Δt of the 7 PC standards in the test set. R All were less than 0.12 min. Furthermore, based on this PC-QSRR prediction model, this implementation case also identified the theoretical retention times of 77 unknown endogenous PCs, and the corresponding retention time deviations Δt. R All values ​​were less than 0.22 min (the theoretical retention times of some endogenous PCs are shown in Table 4-2). Therefore, this case study demonstrates that the QSRR prediction model constructed based on 14 specific PC standards can achieve low-cost, high-precision, and high-coverage calculation of the theoretical retention time of PCs, meeting the qualitative analysis needs of unknown PCs in non-targeted or targeted lipidomics, especially in the acquisition time window setting (Δt) of multiple reaction monitoring (MRM) scanning mode. R All less than 0.5 min).

[0134] (5) Quantitative relationship model between ceramide (Cer-NS) subclasses and their retention in reversed-phase liquid chromatography

[0135] Ceramides are among the most common sphingolipids in lipidomics research. Their sphingosine backbone can be linked to fatty acyl chains with varying numbers of carbons, carbon-carbon double bonds, and positions, theoretically resulting in approximately a thousand Cer-NS compounds in biological matrices. However, currently only a few dozen common commercially available standards are available, insufficient for constructing traditional t-models. R C ~f(c) and t R C The required number of standard samples for ~f(d) is discussed. This implementation case attempts to use 15 Cer-NS standard samples as training research objects to construct a prediction model similar to FFA-QSRR.

[0136] This implementation case selected 15 ceramide standards from 267 lipid compound standards and divided them into training and test sets at a 2:1 ratio. The training set contained 10 specific Cer-NS standards with a total number of carbons (c) and carbon-carbon double bonds (d) on the total fatty acyl chain ranging from 20 to 40 and 0 to 2, respectively (see Table 5-1), covering all FFA structural features involved in non-targeted or targeted lipidomics analysis methods.

[0137] For each subclass, the specific structural characteristic parameters are only S1 and S2, namely the total number of carbons c and the total number of carbon-carbon double bonds d in different fatty acyl chains and sphingosine skeletons. Based on 10 specific Cer-NS standards in the training set, their experimental retention times t are used. R EWith data as the dependent variable and c and d as independent variables, a general prediction model for ceramide retention time (Cer-QSRR) was constructed using nonlinear regression analysis. Its expression is t R C =k0+k1d+k2d 2 +k3d 3 +(k4+k5d+k6d 2 +k7d 3 )c+(k8+k9d+k 10 d 2 +k 11 d 3 )c 2 +(k 12 +k 13 d+k 14 d 2 +k 15 d 3 )c 3 (k0~k) 15 (These are the equation coefficients).

[0138] Table 5-1 Predictive effect of ceramide QSRR model

[0139]

[0140] The predictive ability of the above model was verified based on seven TAG standards in the test set and the validation set. The results are shown in Table 5-2 and... Figure 1 As shown in E.

[0141] Table 5-2 Cer-QSRR Prediction Model Performance

[0142]

[0143] From Table 5-2 and Figure 1 As can be seen from E, the Cer-NS-QSRR prediction model has good predictive ability, i.e., the correlation R between the experimental retention time and the theoretical retention time is good. 2 >0.99, and the corresponding Δt of the five Cer-NS standards in the test set R All were less than 0.10 min. Furthermore, based on this Cer-NS-QSRR prediction model, this implementation case also identified the theoretical retention times of 21 unknown endogenous Cer-NS, and the corresponding retention time deviations Δt. R All values ​​were less than 0.18 min. Therefore, this case study demonstrates that the QSRR prediction model constructed based on 10 specific Cer-NS standards can achieve low-cost, high-precision, and high-coverage calculation of the theoretical retention time of Cer-NS, meeting the qualitative analysis needs of unknown Cer-NS in non-targeted or targeted lipidomics, especially in the acquisition time window setting (Δt) of multiple reaction monitoring (MRM) scanning mode.R All less than 0.5 min).

[0144] Example 2: Quantitative Relationship Model of Five Major Lipid Classes and Their Retention by Reversed-Phase Liquid Chromatography

[0145] This implementation case prepared a mixed standard stock solution of 267 lipid compounds. The aforementioned ultra-high performance liquid chromatography-mass spectrometry (UHPLC-MS / MS) technique was used for experimental retention time and mass spectrometry data acquisition. The samples were divided into training and testing sets at a 2:1 ratio. Multiple structural characteristic parameters were combined to construct the optimal QSRR prediction model for five major classes of lipid compound standards. The 267 lipid compounds involved 49 lipid subclasses, including 39 FFAs, 15 ACars, 6 ACar-OHs, 9 2HOFAs, 9 3HOFAs, 15 DCAs, 9 DAGs, 15 TAGs, 6 MGDGs, 9 DGDGs, 2 MAGs, 21 PCs, 4 LPCs, 1 LPC-P, 12 PEs, 3 PAs, 3 PGs, 3 PIs, 2 LPAs, 2 PSs, 14 CEs, and 10 sterols (Tables 6-1 and 6-2).

[0146] Table 6-1 Distribution of 267 lipid compound standards in five major lipid classes

[0147]

[0148] Table 6-2 Distribution of 267 lipid compound standards in five major lipid classes (continued from Table 6-1)

[0149]

[0150] This embodiment, based on the aforementioned specific lipid standards, screened characteristic structural parameters including the number of carbons ε and the number of carbon-carbon double bonds π on the sphingosine skeleton, the total number of carbon atoms c and the total number of carbon-carbon double bonds d of the fatty acyl chain, ether chain, alkenyl ether chain, or / and the sphingosine skeleton. The skeleton structural features also include the number of glycerol skeletons v and the number of fatty acyl chains, ether chains, or / and alkenyl ether chains (α, β, and γ) connected at the sn1, sn2, and sn3 positions in glycerol lipids and glycerophospholipids; the number of sphingosine skeletons μ and the number of carbons ε and the number of carbon-carbon double bonds π in sphingolipids; and the number of carbons θ and the number of carbon-carbon double bonds ξ on the sterol skeleton in sterol lipids (these two new sterol skeletons). Features can describe the structural characteristics of different sterols, including cholesterol, coprosterol, 7-encholesterol, 7-dehydrocholesterol, sterol, campesterol, ergosterol, β-sitosterol, β-stigmasterol, hydroxylation modification at the α or β position of the fatty acyl chain (h and n), carboxylic acid modification at the ω position of the fatty acyl chain (b), fatty acylation modification of hydroxylated fatty acids or hexoses (a), ether bond (j) or enyl ether bond (f) on the first carbon of the fatty alcohol chain, hydroxylation modification of sphingosine (i.e., phytosphingosine) (δ), phosphate (p), carnitine (m), hexose (w), sialic acid (y), acetylgalactosamine (λ), ethanolamine (e), choline (q), inositol (A), and serine (s).

[0151] The selected specific structural characteristic parameters were quantified. The backbone, characteristic groups, or residues were quantified according to the number of corresponding backbones, characteristic groups, or residues present in the lipid compound; if not present, the parameter was defined as 0. Using the retention time of lipid compound standards as the dependent variable and the selected specific structural characteristic parameters as independent variables, a general QSRR model for lipids was constructed using nonlinear multiple regression analysis after quantification of the specific structural characteristic parameters. This yielded a general QSRR prediction model for the retention times of five major lipid compounds, expressed as follows: S i The above structural characteristic variables, excluding the total number of carbon atoms and the total number of carbon-carbon double bonds in fatty acyl chains, ether chains, alkenyl ether chains, and / or sphingosine skeletons, are represented.

[0152] Among them, the experimental retention time t using the above-mentioned 180+ specific lipid compound standards was... R E Using the aforementioned feature structure parameters as independent variables, and with the dependent variable of the training set, a general prediction model for the retention time of five major lipid compounds was constructed and optimized, namely t R C =k0+X1c 3 +X2c 2 +X3c+X4 (where k0 is a constant), where X1: k1+k2d+k3d 2 +k4d 3 +k5h+k6m+k7β+k8θ+k9μ;

[0153] X2:k 10 +k 11 b+k 12 d+k 13 d 2 +k 14 d 3 +k 15 e+k 16 h+k 17 m+k 18 n+k 19 p+k 20 q+k 21 β+k 22 θ+k 23 A;

[0154] X3:k 24 +k 25 b+k 26 d+k 27 d 2 +k 28 d 3 +k 29 f+k 30 h+k 31 m+k 32 n+k 33 p+k 34 v+k 35 β+k 36 δ+k 37 θ+k 38 μ+k 39 A;

[0155] X4:k 40 a+k 41 b+k 42 d 2 +k 43 d 3 +k 44 f+k 45 h+k 46 j+k 47 m+k 48 s+k 49 w+k 50 y+k 51 α+k 52 δ+k 53 ε+k 54 θ+k 55 μ+k 56 ξ; where k0-k 56 These represent the weighting coefficients of different terms in the above equation.

[0156] Furthermore, this implementation case classifies the aforementioned five categories of lipid compound standards based on their basic skeletal similarity and differences in the number of fatty acyl chains, alkenyl ether chains, and ether chains. These categories are: lipid compounds containing a single fatty acyl chain, ether chain, alkenyl ether chain, or sphingosine skeleton (MFLs); sterol lipid compounds (STs); lipid compounds containing two or more fatty acyl chains, ether chains, and / or alkenyl ether chains (PALs); and sphingolipid compounds containing two or more fatty acyl chains and a sphingosine skeleton (SPs). Thus, four general QSRR models are established. Their expressions are: t R C =k0+X1c 3 +X2c 2 +X3c+X4 (where k0 is a constant). Where X1: k1+k2d+k3d 2 +k4d 3 +k5e+k6m+k7n+k8p+k9γ;

[0157] X2:k 10 +k 11 d+k 12 d 2 +k 13 d 3 +k 14 h+k 15 m+k 16 n+k 17 b+k 18 p+k 19 γ;

[0158] X3:k 20 +k 21 d+k 22 d 2 +k 23 d 3 +k 24 h+k 25 A+k 26 m+k 27 n+k 28 b+k 29 p+k 30 q+k 31 v+k 32 γ;

[0159] X4:k 33 a+k 34 d+k 35 d 2 +k 36 d 3 +k 37 e+k 38 f+k 39 h+k 40 A+k 41j+k 42 m+k 43 p+k 44 q+k 45 s+k 46 v+k 47 w+k 48 y+k 49 β+k 50 γ+k 51 δ+k 52 ε+k 53 θ+k 54 ξ;

[0160] Table 7. Coefficients of different terms in the four general QSRR models.

[0161]

[0162]

[0163] (1) MFLs-QSRR model

[0164] Lipid compounds containing a single fatty acyl chain, ether chain, alkenyl ether chain, or sphingosine skeleton are one of the five major classes of lipid compounds in the biological matrix lipidome, and theoretically, there are thousands of endogenous compounds. However, apart from FFA, ACar, HOFA, and DCA, most lipid subclasses contain only a few compound standards, making traditional QSRR model construction methods unsuitable. Therefore, this case study uses only 76 lipid compound standards as the training set (see Table 8-1), including 26 FFAs, 10 ACars, 4 ACar-OHs, 6 2HOFAs, 6 3HOFAs, 10 DCAs, 1 MAG, 1 LPA, 3 LPCs, 1 LPC-P, 3 LPEs, 1 LPG, 1 LPI, 1 LPS, 1 S1P, and 1 SPH. The structural features of these lipid compound standards include the total number of carbon atoms (c, 9-30) and total number of carbon-carbon double bonds (d, 0-6) of the fatty acyl chain, ether chain, alkenyl ether chain, and / or sphingosine skeleton; the number of glycerol skeletons (v) and the number of fatty acyl chains, ether chains, and / or alkenyl ether chains (α, β) connected at the sn1 and sn2 positions; the number of sphingosine skeletons (μ) in the sphingolipid and its carbon number (ε) and carbon-carbon double bond number (π); hydroxylation modifications (h and n) at the α or β positions of the fatty acyl chain; carboxylic acid modification (b) at the ω position of the fatty acyl chain; ether bond (j) or alkenyl ether bond (f) on the first carbon of the fatty alcohol chain; hydroxylation modification (δ) of sphingosine; phosphoric acid (p); carnitine (m); hexose (w); ethanolamine (e); choline (q); inositol (A); and serine (s).

[0165] Table 8-1 Predictive efficacy of lipid compounds containing a single fatty acyl chain, ether chain, alkenyl ether chain, or sphingosine skeleton.

[0166]

[0167] The predictive ability of the above model was validated based on a test set of lipid compounds, and the results are shown in Table 8-2 and... Figure 2 A.

[0168] Table 8-2 Predictive efficacy of lipid compounds containing a single fatty acyl chain, ether chain, alkenyl ether chain, or sphingosine skeleton.

[0169]

[0170] Depend on Figure 2 As shown in Table A, the optimal QSRR prediction model constructed based on 76 MFLs (corresponding expressions are shown in Table 7) has good predictive ability, namely the correlation R between experimental retention time and theoretical retention time. 2 >0.99, and the corresponding retention time deviation Δt of the 38 MFLs standards in the test set. R All were less than 0.33 min (see Table 8-2). Furthermore, based on this MFLs-QSRR prediction model, this implementation case also identified the theoretical retention times of over 200 unknown endogenous MFLs, and the corresponding retention time deviations Δt. R All values ​​were less than 0.5 min. Therefore, this implementation case uses only 76 specific MFLs standards to construct an optimal QSRR prediction model, achieving low-cost, high-precision, and high-coverage calculation of the theoretical retention time of MFLs. This is used to meet the qualitative analysis needs of unknown MFLs in non-targeted or targeted lipidomics, especially for setting the acquisition time window (Δt) in multiple reaction monitoring scanning mode. R All less than 0.5 min).

[0171] (2) STs-QSRR model

[0172] For sterols and lipids, the published patent concerning the QSRR universal model construction method (patent CN120913704A) only uses a binary classification method to quantify seven different sterol skeletons with 0 or 1 values, i.e., seven structural feature variables, including cholesterol, 22,23-ergosterol, campesterol, 24α-ethylcholesterol, stigmasterol, 3β-cholesterol skeleton, and 24β-ethylcholesterol and its esters. However, this method of describing sterol structural features requires constructing dozens of variables to describe dozens of sterols, leading to excessive complexity in the QSRR universal model formula.

[0173] This implementation case selected 24 sterol lipid compound standards from 267 lipid compound standards, dividing them into training and testing sets at a 2:1 ratio. Based on 16 specific STs standards in the training set (Table 9-1), their experimental retention times t were used... R EThe data are the dependent variable, and the number of carbon atoms θ and carbon-carbon double bonds ξ in the sterol skeleton, and the fatty acylation modifications a, c, d, and w of hexose are the independent variables. A QSRR model is constructed through nonlinear regression analysis (the corresponding expressions are shown in Table 7). This implementation case only requires two new sterol skeleton features: the number of carbon atoms θ and carbon-carbon double bonds ξ in the sterol skeleton of sterol lipids, combined with the carbon atoms c and carbon-carbon double bonds d in the fatty acyl chain, to describe the structural features of more than 10 different types of sterols and their esters, including cholesterol, coprosterol, 7-encholesterol, 7-dehydrocholesterol, sterols, campesterol, ergosterol, β-sitosterol, β-stigmasterol and their esters, etc.

[0174] Table 9-1 Predictive Effects of Sterol Ester Compounds

[0175]

[0176] The predictive ability of the above model was verified based on eight sterol ester compounds in the test set and the validation set. The results are shown in Table 9-2 and... Figure 2 As shown in B.

[0177] Table 9-2 Predictive Results of Sterol Ester Compounds

[0178]

[0179] like Figure 2 As shown in Figure B, an optimal QSRR prediction model was constructed based on 16 sterol lipid compound standards (including 8 sterols and 8 cholesterol esters ChE4:0, ChE12:0, ChE16:0, ChE16:1, ChE17:0, ChE18:1, ChE20:3, and ChE22:6). This model demonstrated good predictive ability, specifically the correlation R between experimental retention time and theoretical retention time. 2 >0.99, and the corresponding retention time deviation Δt of the 8 MFLs standards in the test set. R All were less than 0.25 min (see Table 9-2). Furthermore, based on this STs-QSRR prediction model, this implementation case also identified the theoretical retention times of 23 unknown endogenous sterol ester compounds, and the corresponding retention time deviations Δt. R All values ​​were less than 0.4 min. Therefore, this implementation case uses only 16 specific STs standards to construct an optimal QSRR prediction model, achieving low-cost, high-precision, and high-coverage calculation of the theoretical retention time of STs. This is used to meet the qualitative analysis needs of unknown MFLs in non-targeted or targeted lipidomics, especially for the acquisition time window setting (Δt) in multiple response monitoring scanning mode. R All less than 0.5 min).

[0180] (3) PALs-QSRR model

[0181] Glycerides and glycerophospholipids in biological matrices are mainly lipid compounds containing two or more fatty acyl chains, alkenyl ether chains, or ether chains, and theoretically there are tens of thousands of such lipid compounds. However, apart from DAG, TAG, PC, and PE, most lipid subclasses only contain a few compound standards, making traditional QSRR model construction methods unsuitable.

[0182] This case study uses only 58 lipid compound standards as the training set (see Table 10-1), including 6 DAGs, 4 MGDGs, 6 DGDGs, 10 TAGs, 2 PAs, 14 PCs, 1 PC-O, 1 diPC-O, 8 PEs, 1 PE-P, 2 PGs, 2 PIs, and 1 PS. Based on these 58 specific PALs standards in the training set (see Table 10-1), their experimental retention times t... R E The data were used as the dependent variable, and multiple characteristic structural parameters were used as independent variables. A QSRR model was constructed through nonlinear regression analysis (the corresponding expressions are shown in Table 7). The specific structural features screened from these lipid compound standards included the total number of carbon atoms (c, 28–56) and the total number of carbon-carbon double bonds (d, 0–12) in the fatty acyl chain, ether chain, and alkenyl ether chain; the number of glycerol skeletons (v) and the number of fatty acyl chains, ether chains, and / or alkenyl ether chains (α, β, γ) connected at the sn1, sn2, and sn3 positions; the ether bond (j) or alkenyl ether bond (f) on the first carbon of the fatty alcohol chain; phosphate (p); hexose (w); ethanolamine (e); choline (q); inositol A; and serine (s).

[0183] Table 10-1 Predictive Effects of Lipid Compounds Containing Two or More Fatty Acyl Chains, Ether Chains, or Alkenyl Ether Chains

[0184]

[0185]

[0186] The predictive ability of the above model was verified based on the test set, and the results are shown in Table 10-2 and... Figure 2 As shown in C.

[0187] Table 10-2 Predictive Effects of Lipid Compounds Containing Two or More Fatty Acyl Chains, Ether Chains, or Alkenyl Ether Chains

[0188]

[0189] Depend on Figure 2 As shown in C, the optimal QSRR prediction model constructed based on 58 PALs standards has good predictive ability, i.e., the correlation R between experimental retention time and theoretical retention time. 2 >0.99, and the corresponding retention time deviation Δt of the 28 PALs standards in the test set.R All were less than 0.27 min (see Table 10-2). Furthermore, based on this PALs-QSRR prediction model, this implementation case also identified the theoretical retention times of 722 unknown endogenous PALs, including unknown PE-O, PC-P, PI-O, and TAG-O compounds, and the corresponding retention time deviations Δt. R All values ​​were less than 0.46 min. Therefore, this implementation case uses only 58 specific PALs standards to construct an optimal QSRR prediction model, achieving low-cost, high-precision, and high-coverage calculation of PALs theoretical retention times. This is used to meet the qualitative analysis needs of unknown PALs in non-targeted or targeted lipidomics, especially for setting the acquisition time window (Δt) in multiple response monitoring scanning mode. R All less than 0.5 min).

[0190] (4) SPs-QSRR model

[0191] Sphingolipids containing fatty acyl chains are an important component of the sphingolipidome in the biological matrix, and theoretically, thousands of such lipid compounds exist. However, apart from Cer-NDS and Cer-NS, most lipid subclasses only have a few compound standards, making traditional QSRR model construction methods unsuitable. This case study uses only 28 lipid compound standards as the training set (see Table 11-1), including 4 Cer-NDS, 10 Cer-NS, 1 Cer-ADP, 1 Cer-EODS, 1 Cer-EOP, 1 Cer-EOS, 1 Cer-HDS, 1 Cer-NDP, 1 Cer-NP, 1 CerPE, 1 GM1, 1 GM3, 1 Hex2Cer, 1 HexCer-NS, and 2 SM. Based on these 28 specific SPs standards in the training set (see Table 11-1), their experimental retention time t... R E The data were used as the dependent variable, and multiple characteristic structural parameters were used as independent variables. A QSRR model was constructed through nonlinear regression analysis (the corresponding expressions are shown in Table 7). The structural characteristics of these lipid compound standards include fatty acyl chains, the total number of carbon atoms (c) (20–68) and the total number of carbon-carbon double bonds (d) of the sphingosine skeleton, the number of sphingosine skeletons (μ) and their carbon number (ε) and carbon-carbon double bond number (π) in the sphingolipid, hydroxylation modifications (h and n) at the α or β positions on the fatty acyl chains, fatty acylation modifications (a) of hydroxylated fatty acids or hexoses, hydroxylation modifications of sphingosine (i.e., phytosphingosine δ), phosphate (p), hexose (w), sialic acid (y), acetylgalactosamine (λ), ethanolamine (e), and choline (q).

[0192] Table 11-1 Predictive Effects of Sphingolipid Compounds Containing Fatty Acyl Chains

[0193]

[0194] The predictive ability of the above model was verified based on the test set and validation set. The results are shown in Table 11-2 and... Figure 2 As shown in D.

[0195] Table 11-2 Predictive Effects of Sphingolipid Compounds Containing Fatty Acyl Chains

[0196]

[0197] Depend on Figure 2 As shown in D, the optimal QSRR prediction model constructed based on 28 SPs standards has good predictive ability, namely the correlation R between experimental retention time and theoretical retention time. 2 >0.99, and the corresponding retention time deviation Δt of the 15 SPs standards in the test set. R All were less than 0.23 min (see Table 11-2). Furthermore, based on this SPs-QSRR prediction model, this implementation case also identified the theoretical retention times of 189 unknown endogenous SPs, and the corresponding retention time deviations Δt. R All values ​​were less than 0.49 min. Therefore, this implementation case uses only 28 specific SPs standards to construct an optimal QSRR prediction model, achieving low-cost, high-precision, and high-coverage calculation of the theoretical retention time of SPs. This is used to meet the qualitative analysis needs of unknown SPs in non-targeted or targeted lipidomics, especially for the acquisition time window setting (Δt) in multiple reaction monitoring scanning mode. R All less than 0.5 min).

[0198] The above results indicate that t R C and t R E There is a linear correlation (R) 2 >0.996)( Figure 2 (from A to D), and their retention time deviation is small (Δt) R <0.33 min). Meanwhile, 2550 lipid compounds (covering 92 subclasses) were identified in seven typical biological samples, including human plasma, urine and cells, mouse liver tissue and feces, Arabidopsis thaliana leaves, and Escherichia coli, with 93% of these compounds having a t... R C and t R E Good consistency (R) 2 >0.996, t R <0.40min). The predictive ability of the model provided by this invention is significantly better than that of the reported models in terms of FFA (Δt). R <1.0min), 3HOFA (R2 ~0.984, Δt R <0.57min) and PE (R 2 ~0.980, Δt R The theoretical retention time is <0.40 min), and it is applicable to the calculation of the theoretical retention time of unknown compounds in lipid subclasses containing a few standards, such as lysophosphatidic acid, lysoglycerol phosphatidylcholine, dicarboxylic acids, and ceramides. Furthermore, the QSRR model and retention time prediction system provided by this invention offer optimal theoretical retention times for more than 20,000 lipid compounds (covering 190 subclasses) within a 1.5-minute multi-reaction monitoring scanning window.

[0199] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for constructing a model of the relationship between lipid structure and its retention quantification by reversed-phase liquid chromatography, characterized in that, Includes the following steps: Using the structural features of lipid compounds as candidate independent variables, the optimal independent variables and their combinations are screened out through stepwise regression modeling as specific structural features. The screened specific structural features include one or more of the following: the number of carbons ε and the number of carbon-carbon double bonds π in the sphingosine skeleton, the number of carbons θ and the number of carbon-carbon double bonds ξ in the sterol skeleton, N-acetylgalactosamine λ, and / or one or more of the skeleton, characteristic groups and residues contained in the lipid compound. t of specific lipid compound standards R E Using the selected specific structural features as independent variables and employing the LM() function, a quantitative relationship model was established between lipid structure and its retention in reversed-phase liquid chromatography. The constructed quantitative relationship model is as follows: , This indicates the specific structural features selected. This indicates the theoretical retention time of lipid compounds.

2. The construction method as described in claim 1, characterized in that, The specific lipid compound standards include sterols and specific cholesterol esters; the specific cholesterol esters have a carbon number (c) range of 0–22 on their fatty acyl chains and a carbon-carbon double bond number (d) range of 0–6 on their fatty acyl chains; the quantitative relationship model between the constructed sterol lipid structure and its retention by reversed-phase liquid chromatography is t. R C =k0+k1θ+k2ξ+k3c+k4dc+k5c 2 +k6c 3 In the formula, k0-k6 represent the weight coefficients of different terms in the regression equation.

3. The construction method as described in claim 1, characterized in that, The specific lipid compound standards include specific sphingolipid compounds, wherein the total number of carbons (c) on the sphingosine and fatty acyl chains of the specific sphingolipid compounds ranges from 20 to 70, and the number of carbon-carbon double bonds (d) on the fatty acyl chains ranges from 0 to 4; the quantitative relationship model between the sphingolipid compounds containing two or more fatty acyl chains and a sphingosine backbone and their retention by reversed-phase liquid chromatography is t R C =k0 + k1a + k2h + k3d + k4d 2 + k5p + k6w+ k7y + k8δ + k9c + k 10 pc + k 11 c 2 + k 12 c 3 In the formula, k0-k 12 δ represents the weight coefficients of different terms in the regression equation; a represents the fatty acylation modification of hydroxylated fatty acids or hexoses; h represents the hydroxylation modification at the α-position of the fatty acyl chain; d represents the total number of carbon-carbon double bonds in the fatty acyl chain, ether chain, alkenyl ether chain and / or sphingosine skeleton; p represents phosphoric acid; w represents hexoses; y represents sialic acid; δ represents sphingosine hydroxylation modification; and c represents the total number of carbon atoms in the fatty acyl chain, ether chain, alkenyl ether chain and / or sphingosine skeleton.

4. The construction method as described in claim 1, characterized in that, To construct a quantitative model for the retention of lipid compounds containing a single fatty acyl chain, ether chain, alkenyl ether chain, or sphingosine skeleton by reversed-phase liquid chromatography, the model expression is t. R C =k0+X1c 3 +X2c 2 +X3c+X4, where c represents the total number of carbon atoms in the fatty acyl chain, ether chain, alkenyl ether chain, and / or sphingosine skeleton, and X1 is: k1 + k2d + k3d 2 + k4d 3 + k5e + k6m + k7n + k8p + k9γ; X2 is: k 10 + k 11 d + k 12 d 2 + k 13 d 3 + k 14 h + k 15 m + k 16 n + k 17 b + k 18 p + k 19 γ; X3 is: k 20 + k 21 d + k 22 d 2 + k 23 d 3 + k 24 h + k 25 A + k 26 m + k 27 n + k 28 b + k 29 p + k 30 q + k 31 v + k 32 γ; X4 is: k 33 a + k 34 d + k 35 d 2 + k 36 d 3 + k 37 e + k 38 f + k 39 h + k 40 A + k 41 j + k 42 m + k 43 p + k 44 q + k 45 s + k 46 v + k 47 β + k 48 γ + k 49 δ + k 50 ε; In the formula, d represents the total number of carbon-carbon double bonds in the fatty acyl chain, ether chain, alkenyl ether chain, and / or sphingosine skeleton; e represents ethanolamine; m represents carnitine; n represents hydroxylation modification at the β-position of the fatty acyl chain; p represents phosphoric acid; γ represents the number of fatty acyl chains, ether chains, and / or alkenyl ether chains connected at the sn3 position in glycerol esters and glycerophospholipids; h represents hydroxylation modification at the α-position of the fatty acyl chain; b represents carboxylic acid modification at the ω-position of the fatty acyl chain; A represents inositol; q represents choline; v represents the number of glycerol skeletons in glycerol esters and glycerophospholipids; a represents fatty acylation modification of hydroxylated fatty acids or hexoses; f represents alkenyl ether bond; j represents ether bond on the first carbon of the fatty alcohol chain; s represents serine; v represents the number of glycerol skeletons in glycerol esters and glycerophospholipids; β represents the number of fatty acyl chains, ether chains, and / or alkenyl ether chains connected at the sn2 position in glycerol esters and glycerophospholipids; δ represents hydroxylation modification of sphingosine; and ε represents the number of carbons on the sphingosine skeleton in the sphingolipid.

5. The construction method as described in claim 1, characterized in that, To construct a quantitative model for the relationship between lipid compounds containing two or more fatty acyl chains, ether chains, and / or alkenyl ether chains and their retention in reversed-phase liquid chromatography, the model expression is t. R C =k0+X1c 3 +X2c 2 +X3c+X4, where c represents the total number of carbon atoms in the fatty acyl chain, ether chain, alkenyl ether chain, and / or sphingosine skeleton, and X1 is: k1 + k2d 2 + k3e + k4p + k5γ; X2 is: k6 + k7d + k8d 2 + k9p + k 10 γ; X3 is: k 11 + k 12 d + k 13 d 2 + k 14 p + k 15 q + k 16 v + k 17 γ + k 18 A; X4 is: k 19 d 2 + k 20 f + k 21 j + k 22 p + k 23 q + k 24 s + k 25 v + k 26 w + k 27 γ + k 28 A; In the formula, d represents the total number of carbon-carbon double bonds in the fatty acyl chain, ether chain, alkenyl ether chain and / or sphingosine skeleton, e represents ethanolamine, p represents phosphoric acid, γ represents the number of fatty acyl chains, ether chains and / or alkenyl ether chains linked at the sn3 position in glycerol esters and glycerophospholipids, q represents choline, v represents the number of glycerol skeletons in glycerol esters and glycerophospholipids, A represents inositol, f represents alkenyl ether bond, j represents the ether bond on the first carbon of the fatty alcohol chain, s represents serine, v represents the number of glycerol skeletons in glycerol esters and glycerophospholipids, and w represents hexose.

6. The construction method as described in claim 1, characterized in that, For the five major classes of lipid compounds, the general model for the quantitative relationship between the constructed lipid structure and its retention by reversed-phase liquid chromatography is t. R C =k0+ X1c 3 + X2c 2 + X3c+ X4, k0 is a constant, and c represents the total number of carbon atoms in the fatty acyl chain, ether chain, alkenyl ether chain and / or sphingosine skeleton; Where X1: k1 + k2d + k3d 2 + k4d 3 + k5h + k6m + k7β + k8θ + k9μ; X2: k 10 + k 11 b + k 12 d + k 13 d 2 + k 14 d 3 + k 15 and + k 16 h + k 17 m + k 18 n + k 19 p + k 20 q +k 21 β + k 22 θ + k 23 A; X3: k 24 + k 25 b + k 26 d + k 27 d 2 + k 28 d 3 + k 29 f + k 30 h + k 31 m + k 32 n + k 33 p + k 34 v +k 35 β + k 36 δ + k 37 θ + k 38 μ + k 39 A; X4: k 40 a + k 41 b + k 42 d 2 + k 43 d 3 + k 44 f + k 45 h + k 46 j + k 47 m + k 48 s + k 49 w + k 50 y +k 51 α + k 52 δ + k 53 ε + k 54 θ + k 55 μ + k 56 ξ; where d represents the total number of carbon-carbon double bonds in the fatty acyl chain, ether chain, alkenyl ether chain, and / or sphingosine skeleton; h represents the hydroxylation modification at the α-position of the fatty acyl chain; m represents carnitine; β represents the number of fatty acyl chains, ether chains, and / or alkenyl ether chains linked at the sn2 position in glycerol esters and glycerophospholipids; θ represents the number of carbons on the sterol skeleton in sterol esters; μ represents the number of sphingosine skeletons in sphingolipids; b represents the carboxylic acid modification at the ω-position of the fatty acyl chain; e represents ethanolamine; n represents the hydroxylation modification at the β-position of the fatty acyl chain; p represents phosphoric acid; and q represents... A represents choline, f represents inositol, v represents the number of glycerol backbones in glycerol lipids and glycerophospholipids, δ represents sphingosine hydroxylation modification, a represents fatty acylation modification of hydroxylated fatty acids or hexoses, j represents the ether bond on the first carbon of the fatty alcohol chain, s represents serine, w represents hexose, y represents sialic acid, α represents the number of fatty acyl chains, ether chains and / or enyl ether chains connected at the sn1 position in glycerol lipids and glycerophospholipids, ε represents the number of carbons on the sphingosine backbone in sphingolipids, and ξ represents the number of carbon-carbon double bonds on the sterol backbone in sterol lipids. The five major lipid classes include free fatty acids and their modified compounds, glycerides, glycerophospholipids, sphingolipids, and sterol lipids; For a lipid subclass, the general model for the quantitative relationship between the constructed lipid subclass and its retention in reversed-phase liquid chromatography is: t R C =k0+k1d+k2d 2 +k3d 3 +(k4+k5d+k6d 2 +k7d 3 )c+(k8+k9d+k 10 d 2 +k 11 d 3 )c 2 +(k 12 +k 13 d+k 14 d 2 +k 15 d 3 )c 3 ; In the formula, d represents the total number of carbon-carbon double bonds in the fatty acyl chain, ether chain, alkenyl ether chain and / or sphingosine skeleton, and c represents the total number of carbon atoms in the fatty acyl chain, ether chain, alkenyl ether chain and / or sphingosine skeleton. The lipid subclasses include free fatty acids, cholesterol esters, triglycerides, and glycerophosphatidylcholine.

7. A lipid retention time prediction system, characterized in that, Includes a data acquisition module and a retention time prediction module; The data acquisition module is used to acquire specific structural feature parameters of lipid compounds, quantify them as feature values, and submit them to the retention time prediction module. The specific structural feature parameters include one or more of the following: the number of carbons ε and the number of carbon-carbon double bonds π in the sphingosine skeleton, the number of carbons θ and the number of carbon-carbon double bonds ξ in the sterol skeleton, N-acetylgalactosamine λ, and / or one or more of the skeleton, feature groups, and residues contained in the lipid compound. The retention time prediction module calculates the predicted theoretical retention time based on the received feature values ​​according to a built-in specific calculation model, and outputs the result according to the calculation result. The specific calculation model is a quantitative relationship model as described in any one of claims 1 to 6.

8. The lipid retention time prediction system as described in claim 7, characterized in that, The specific computational model used to predict the theoretical retention time of unknown steroidal lipid compounds is the quantitative relationship model as described in claim 2; The specific calculation model used to predict the theoretical retention time of unknown sphingolipid compounds containing two or more fatty acyl chains and a sphingosine skeleton is the quantitative relationship model as described in claim 3. The specific calculation model used to predict the theoretical retention time of unknown lipid compounds containing a single fatty acyl chain, ether chain, alkenyl ether chain, or sphingosine skeleton is the quantitative relationship model as described in claim 4. The specific computational model used to predict the theoretical retention time of unknown lipid compounds containing two or more fatty acyl chains, ether chains, or / and alkenyl ether chains is the quantitative relationship model as described in claim 5.

9. The lipid retention time prediction system as described in claim 7, characterized in that, The specific computational model used to predict the theoretical retention time of compounds in five major lipid classes or lipid subclasses is the quantitative relationship model as described in claim 6. The five major lipid classes include free fatty acids and their modified compounds, glycerides, glycerophospholipids, sphingolipids, and sterol lipids; The lipid subclasses include free fatty acids, cholesterol esters, triglycerides, and glycerophosphatidylcholine.