Combined marker for grading diagnosis of non-alcoholic fatty liver disease and application of combined marker

Through metabolomics and lipidomics technology combined with machine learning algorithms, five plasma metabolic markers were identified and NAFLD hierarchical diagnostic models were constructed, which solved the invasiveness and low sensitivity of existing diagnostic methods, and achieved non-invasive, rapid and economical hierarchical diagnosis of the disease, improving the accuracy and accessibility of the diagnosis.

CN120334385APending Publication Date: 2025-07-18SHENYANG PHARMA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510284297.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-18

Smart Images

  • Figure HDA0005306677360000011
    Figure HDA0005306677360000011
  • Figure HDA0005306677360000012
    Figure HDA0005306677360000012
  • Figure HDA0005306677360000013
    Figure HDA0005306677360000013
Patent Text Reader

Abstract

The invention belongs to the technical field of medical detection, and particularly relates to a combined marker for grading diagnosis of a non-alcoholic fatty liver disease (NAFLD) and application of the combined marker. The invention discloses a combined marker for grading diagnosis of NAFLD (non-alcoholic flash disease). The combined marker consists of the following five markers: 5-aminolevulinic acid, mesaconic acid, shikimic acid, phosphatidylcholine O-35: 3 and phosphatidylinositol-36: 2. The grading diagnosis is divided into health, mild NAFLD, moderate NAFLD and severe NAFLD. According to the invention, by integrating metabonomics and lipidomics technologies, the metabolic change in the NAFLD progress process is comprehensively detected and analyzed; in combination with a machine learning algorithm, five biomarkers closely related to NAFLD progression in blood plasma are provided for the first time, and a diagnosis model for NAFLD disease grading is established. The established model is high in sensitivity and specificity, has the advantages of noninvasiveness, rapidness, economy and the like, and has a wide clinical application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical detection, and particularly relates to a combination biomarker for staging diagnosis of non-alcoholic fatty liver disease, and more precisely, a technique for identifying plasma metabolic biomarkers during the progression of non-alcoholic fatty liver disease and a method for constructing a diagnostic model. Background Art

[0002] Non-alcoholic fatty liver disease (NAFLD) is a liver disease closely related to metabolic disorders, usually manifested as fat accumulation in the liver without excessive alcohol consumption. The incidence of NAFLD has been increasing year by year and has become one of the common liver diseases worldwide, especially common in patients with metabolic syndromes such as obesity, diabetes, and hypertension.

[0003] Early identification and stratification of the severity of MAFLD are crucial for clinical decision-making and patient management. For patients with mild and moderate MAFLD, by changing lifestyle, such as adjusting diet structure, increasing physical exercise, and losing weight, liver steatosis can usually be effectively reversed; however, for patients with more severe conditions, additional drug interventions may be required to prevent the disease from further progressing. Therefore, through timely stratified diagnosis, it is possible to effectively prevent the aggravation of severe cases due to insufficient treatment and avoid unnecessary risks caused by over-treatment of mild cases. Currently, imaging techniques and liver biopsy are the main diagnostic methods for evaluating the severity of MAFLD. However, these methods have obvious limitations. Although imaging examinations are widely used, they are costly and have low sensitivity. Liver biopsy, as the current gold standard, although has high accuracy in evaluating liver lesions, due to its invasive operation, it is vulnerable to sampling errors and brings certain pain and risks to patients. These limitations further highlight the urgent need to develop non-invasive and reliable biomarkers. By accurately distinguishing different stages of MAFLD, new biomarkers can provide strong support for early diagnosis and important basis for formulating personalized treatment plans.

[0004] In recent years, metabolomics / lipidomics research has provided new ideas for disease diagnosis. By analyzing metabolites / lipids in patients' blood, biomarkers closely related to the disease can be identified. These biomarkers reflect the disorders of metabolism in the body and can be used as powerful tools for disease diagnosis. In addition, the application of machine learning technology makes the data analysis and pattern recognition based on metabolic biomarkers more accurate and efficient, and can construct a more accurate diagnostic model through big data analysis. Summary of the Invention

[0005] In view of the deficiencies of existing NAFLD diagnostic methods, the present invention proposes a combination of metabolic markers and a diagnostic model for the grading diagnosis of non-alcoholic fatty liver disease based on metabolomics, lipidomics and machine learning techniques, aiming to improve the diagnostic accuracy of NAFLD and the precision of disease staging. Compared with traditional liver biopsy, this model does not require invasive operations, thus avoiding related risks and discomfort.

[0006] The present invention adopts the following technical solutions:

[0007] In the first aspect of the present invention, a combined marker for the grading diagnosis of non-alcoholic fatty liver disease (NAFLD) is provided, which consists of the following 5 markers: 5-Aminolevulinic acid, Mesaconic acid, Shikinic acid, Phosphatidylcholine O-35:3 (PC O-35:3) and Phosphatidylinositol-36:2 (PI 36:2); the grading diagnosis is divided into healthy, mild NAFLD, moderate NAFLD and severe NAFLD.

[0008] In the second aspect of the present invention, the application of the aforementioned combined marker in the preparation of a diagnostic product for non-alcoholic fatty liver disease (NAFLD) is provided, and the NAFLD diagnosis is divided into healthy, mild NAFLD, moderate NAFLD and severe NAFLD.

[0009] In the above technical solution, further, the diagnostic product is a diagnostic kit.

[0010] In the above technical solution, further, the application of the reagent for detecting the content of the aforementioned combined marker in the preparation of a diagnostic kit for non-alcoholic fatty liver disease (NAFLD).

[0011] In the third aspect of the present invention, a grading diagnosis system for non-alcoholic fatty liver disease (NAFLD) is provided, including:

[0012] (1) A quantitative detection module: a module for detecting the normalized values of 5 plasma metabolites, namely 5-Aminolevulinic acid, Mesaconic acid, Shikinic acid, Phosphatidylcholine O-35:3 (PC O-35:3) and Phosphatidylinositol-36:2 (PI 36:2) in a sample;

[0013] (2) An F value calculation module:

[0014] F1 = 1.318X1 + 0.037X2 - 2.151X3 - 0.014X4 + 0.823X5 + 2.94;

[0015] F2 = -0.038X1 + 0.009X2 + 0.292X3 - 0.09X4 - 0.167X5 + 4.594;

[0016] F3 = -2.218X1 + 0.006X2 + 0.349X3 - 2.663X4 + 2.924X5 + 34.235;

[0017] In the formula, X1 to X5 respectively represent the normalized values of 5 - Aminolevulinic acid, Mesaconic acid, Shikinic acid, Phosphatidylcholine O - 35:3 (PC O - 35:3), and Phosphatidylinositol - 36:2 (PI 36:2) in plasma;

[0018] (3) Result output module: Using 0 as the cut - off value, output the results of comparing the values of F1, F2, and F3 with 0; if F1 > 0, the diagnosis result is positive, that is, NAFLD; if F2 < 0, it is mild NAFLD, if F2 > 0, it is moderate / severe NAFLD; if F3 < 0, it is moderate NAFLD, if F3 > 0, it is severe NAFLD.

[0019] The fourth aspect of the present invention provides a method for constructing a grading diagnosis model for non - alcoholic fatty liver disease (NAFLD):

[0020] S1. Collection of plasma samples from NAFLD patients with different disease severities: Use transient elastography (FibroTouch) to evaluate the liver stiffness index (CAP) of all subjects; grade the patients according to the CAP value, and collect the plasma samples of each group of patients;

[0021] S2. LC - MS / MS - based wide - targeted metabolomics analysis of plasma from NAFLD patients: Use the methanol precipitation protein method for plasma sample pretreatment; use liquid chromatography - tandem triple quadrupole mass spectrometry and monitor in the MRM acquisition mode; obtain the normalized values of each plasma metabolite;

[0022] S3. LC - MS / MS - based wide - targeted lipidomics analysis of plasma from NAFLD patients: Use the Bligh & Dyer method for plasma sample pretreatment; use liquid chromatography - tandem triple quadrupole mass spectrometry and monitor in the MRM acquisition mode; obtain the normalized values of each lipid;

[0023] S4. Identification of NAFLD grading diagnosis biomarkers:

[0024] Taking the normalized values of each substance as variables, the P and FC values of each variable among different groups were obtained; according to the criteria of P < 0.05, VIP > 1.0 and FC > 1.5 or FC < 0.667, differential metabolites between mild NAFLD vs. healthy group, moderate NAFLD vs. healthy group, and severe NAFLD vs. healthy group were screened out respectively;

[0025] Based on LASSO regression and machine learning algorithms, biomarkers that can reflect different degrees of NAFLD were identified;

[0026] S5. Construction of the NAFLD grading diagnosis model: Based on the identified biomarkers, a grading diagnosis model for NAFLD patients with different degrees of illness was established using the binary logistic regression-ROC curve method.

[0027] In the above technical solution, further, the specific operation steps of "grading patients according to the CAP value" in step S1 are as follows: A CAP value in the range of 230 - 248 dB / m is diagnosed as mild NAFLD; a CAP value in the range of 248 - 269 dB / m is diagnosed as moderate NAFLD; a CAP value ≥ 269 dB / m is diagnosed as severe NAFLD; at the same time, healthy people without NAFLD (CAP value lower than 230 dB / m) were selected as the control group.

[0028] In the above technical solution, further, 5 potential biomarkers reflecting different degrees of NAFLD were identified in step S3, which are 5-Aminolevulinic acid, Mesaconic acid, Shikinic acid, Phosphatidylcholine O-35:3 (PC O-35:3), and Phosphatidylinositol-36:2 (PI 36:2).

[0029] In the above technical solution, further, the specific operation steps of "pretreating plasma samples by the methanol precipitation protein method" in step S2 are as follows: Take 100 μL of plasma sample, add 10 μL of internal standard solution, add 300 μL of methanol, vortex for 3 min, centrifuge at 12000 rpm and 4 °C for 5 min; transfer the supernatant to an EP tube and dry it under a nitrogen stream; dissolve the residue in 50 μL of 50% methanol-water, vortex for 3 min, sonicate for 5 min, centrifuge at 12000 rpm and 4 °C for 5 min, and take the supernatant for subsequent injection analysis; the internal standard solution contains 50 μg / mL L-2-chlorophenylalanine and 100 μg / mL heptadecanoic acid;

[0030] The specific operation steps of "monitoring by liquid chromatography tandem triple quadrupole mass spectrometry in MRM acquisition mode" described in step S2 are as follows: Set the liquid chromatography conditions: mobile phase, flow rate, injection volume, elution program, etc.; Set the mass spectrometry conditions: ion source parameters, collision energy, parent ion, daughter ion, etc.; Equilibrate the system and inject the sample;

[0031] For the "liquid chromatography conditions" mentioned above, the chromatographic column is an octadecylsilyl-bonded silica gel column, and 0.1 vol% formic acid aqueous solution (A)-0.1 vol% formic acid acetonitrile (B) is used as the mobile phase, with a flow rate of 0.4 mL·min -1 , the column temperature is 30 °C, the injection volume is 2 μL, and gradient elution is carried out; among them, the elution program is: 0 - 2.1 min, 100% → 100% mobile phase A, 0% → 0% mobile phase B; 2.1 - 14.0 min, 100% → 5% mobile phase A, 0% → 95% mobile phase B; 14.0 - 16.0 min, 5% → 5% mobile phase A, 95% → 95% mobile phase B; 16.0 - 16.1 min, 5% → 100% mobile phase A, 95% → 0% mobile phase B;

[0032] For the "mass spectrometry conditions" mentioned above, an electrospray ionization source is used for detection in positive and negative ion modes; the source temperature is 500 °C; the ion spray voltages are 5500 V(+) and 4500 V(-) respectively; the vacuum degree (10e -5 ) is 3.1 Torr; nitrogen is used as the nebulizing gas, auxiliary heating gas and curtain gas, and the pressures are 50, 50 and 20 psi respectively; the metabolite ion pairs are introduced into the mass spectrometry for detection, and the CE voltage is optimized; the ion pairs with no response are deleted, and the stably appearing fragment ions are selected to achieve the best detection sensitivity. Finally, 269 ion pairs in the positive ion mode and 136 ion pairs in the negative ion mode are retained;

[0033] The specific operation steps of "pretreating plasma samples by the Bligh & Dyer method" described in step S3 are as follows: Take 200 μL of plasma sample in an EP tube, add 40 μL of methanol and 20 μL of internal standard, then vortex for 30 s, add 1.3 mL of methyl tert-butyl ether-methanol mixed solution (5:1.5, v / v), vortex mix for 3 min, ultrasonicate for 5 min, add 290 μL of water, centrifuge at 12000 r / min at 4 °C for 5 min, take the organic phase and place it in another EP tube to dry in an air stream, dissolve the residue in 200 μL of methanol, vortex mix for 3 min, ultrasonicate for 5 min, centrifuge at 12000 r / min for 5 min, and take the supernatant for injection analysis;

[0034] The specific operation steps of "monitoring in the MRM acquisition mode using liquid chromatography tandem triple quadrupole mass spectrometry" in step S3 are as follows: Set the liquid chromatography conditions: mobile phase, flow rate, injection volume, elution program, etc.; Set the mass spectrometry conditions: ion source parameters, collision energy, parent ion, daughter ion, etc.; Equilibrate the system and inject the sample;

[0035] For the "liquid chromatography conditions", the chromatographic column is an octadecylsilyl-bonded silica gel column, using acetonitrile-water (6:4, v / v) (A) - isopropanol-acetonitrile (9:1, v / v) (B) as the mobile phase, both containing 10 mmol / L ammonium acetate, and the flow rate is 0.4 mL·min -1 , the column temperature is 25 °C, the injection volume is 2 μL, and elution is carried out in a gradient elution manner; among them, from 0 to 5.0 min, 70% → 10% mobile phase A, 30% → 90% mobile phase B; from 5.0 to 13.0 min, 10% → 0% mobile phase A, 90% → 100% mobile phase B; from 13.0 to 13.1 min, 0% → 70% mobile phase A, 100% → 30% mobile phase B;

[0036] For the "mass spectrometry conditions", an electrospray ionization source is used, and detection is carried out in positive and negative ion modes; the source temperature is 500 °C; the ion spray voltages are 5500 V (+) and 4500 V (-) respectively; the vacuum degree is (10e -5 ) 3.1 Torr; nitrogen is used as the nebulizing gas, auxiliary heating gas and curtain gas, and the pressures are 50, 50 and 20 psi respectively; the obtained lipid reference substance is directly injected into the mass spectrometer through a syringe pump to optimize relevant mass spectrometry parameters; the "model - prediction" strategy is adopted, and the rules are summarized according to the chromatographic and mass spectrometric behaviors of existing lipid reference substances to predict the collision energy CE and declustering voltage DP of unknown lipids; the lipid ion pairs are introduced into the mass spectrometer for detection, the ion pairs with no response are deleted, and the stably appearing fragment ions are selected to achieve the best detection sensitivity. Finally, 1071 ion pairs in the positive ion mode and 121 ion pairs in the negative ion mode are retained.

[0037] In the above technical solution, further, in step S4, the least absolute shrinkage and selection operator (LASSO) is used to screen the characteristic variables of different disease degrees of NAFLD as potential biomarkers; through three machine learning algorithms: extreme gradient boosting (XGBoost), decision tree (DT), and logistic regression (LR), a machine model is constructed to verify the performance of potential biomarkers in the grading diagnosis of NAFLD.

[0038] In the above technical solution, further, the specific operation steps of step S5 are as follows: Using different disease stages as grouping variables, assigning 0 and 1, and the normalized data of the identified markers as covariates, perform binary logistic regression analysis to obtain the coefficients of each variable in the equation, and construct a NAFLD grading diagnosis expression; Using the Preg value as the test variable, the grouping values 0 and 1 as the state variables, and the state variable value set to 1, perform ROC curve analysis to evaluate the diagnostic ability of the established model, and construct the following NAFLD grading diagnosis model:

[0039] ① Healthy (-) vs. NAFLD (+):

[0040] F1 = 1.318X1 + 0.037X2 - 2.151X3 - 0.014X4 + 0.823X5 + 2.94;

[0041] ② Mild NAFLD (+) vs. Moderate / Severe NAFLD (-):

[0042] F2 = -0.038X1 + 0.009X2 + 0.292X3 - 0.09X4 - 0.167X5 + 4.594;

[0043] ③ Moderate NAFLD (+) vs. Severe group (-):

[0044] F3 = -2.218X1 + 0.006X2 + 0.349X3 - 2.663X4 + 2.924X5 + 34.235;

[0045] In the above formulas, X1 to X5 respectively represent the normalized values of 5-Aminolevulinic acid, Mesaconic acid, Shikinic acid, Phosphatidylcholine O-35:3 (PC O-35:3), and Phosphatidylinositol-36:2 (PI 36:2) in plasma.

[0046] For the NAFLD grading diagnosis model constructed by the present invention, in the discrimination between healthy and NAFLD patients, the area under the ROC curve is 1.00; in the discrimination between mild and moderate / severe NAFLD patients, the area under the ROC curve is 0.912; in the discrimination between moderate and severe NAFLD patients, the area under the ROC curve is 1.00. In addition, the overall accurate recognition rate of the model in the training set samples is 88.3%, and the accurate recognition rate in the test set samples is 91.7%, proving the good applicability and accuracy of the model.

[0047] Advantages of the present invention:

[0048] Through extensive targeted metabolomics and lipidomics technologies, combined with machine learning algorithms, the present invention for the first time provides a combination of five biomarker combinations closely related to the progression of NAFLD, which is applicable to the early diagnosis of NAFLD disease grading. The combined biomarkers of the present invention do not require invasive operations such as liver biopsy, and can achieve non-invasive, rapid, and economical diagnosis, providing a convenient detection tool for clinical practice, which can greatly improve the compliance of patients and the accessibility of diagnosis; the diagnostic model established by the present invention can accurately predict the pathological stage of NAFLD, improving the objectivity and accuracy of diagnosis; it can achieve the early diagnosis of NAFLD, help clinicians take timely intervention measures, slow down the disease progression, and improve the long-term prognosis of patients; it has a low cost, and compared with traditional imaging examinations and liver biopsy, it can be widely applied to primary medical institutions to help early detection and intervention of NAFLD. Description of the Drawings

[0049] Figure 1 It is the chromatogram of extensive targeted metabolomics of plasma samples of NAFLD patients; among them, Figure A is the chromatogram in positive ion mode; Figure B is the chromatogram in negative ion mode.

[0050] Figure 2 It is the chromatogram of extensive targeted lipidomics of plasma samples of NAFLD patients; among them, Figure A is the chromatogram in positive ion mode; Figure B is the chromatogram in negative ion mode.

[0051] Figure 3 It is the volcano plot of differential metabolites and differential lipids in plasma samples of NAFLD patients; among them, Figure A is the volcano plot of differential metabolites in mild NAFLD vs. healthy group, moderate NAFLD vs. healthy group, and severe NAFLD vs. healthy group; Figure B is the volcano plot of differential lipids in mild NAFLD vs. healthy group, moderate NAFLD vs. healthy group, and severe NAFLD vs. healthy group.

[0052] Figure 4 It is the result diagram of screening characteristic variables by LASSO regression method.

[0053] Figure 5 It is the result diagram of the diagnostic efficacy of different machine learning algorithms; among them, Figure A is the ROC curve of the training set, the ROC curve of the test set, and the confusion matrix diagram of the test set in the XGBoost algorithm; Figure B is the ROC curve of the training set, the ROC curve of the test set, and the confusion matrix diagram of the test set in the DT algorithm; Figure C is the ROC curve of the training set, the ROC curve of the test set, and the confusion matrix diagram of the test set in the LR algorithm.

[0054] Figure 6ROC curve diagram of the diagnostic model; among them, Figure A is the ROC curve for distinguishing healthy and NAFLD patients; Figure B is the ROC curve for distinguishing mild and moderate / severe NAFLD patients; Figure C is the ROC curve for distinguishing moderate and severe NAFLD patients.

[0055] Figure 7 Confusion matrix diagrams of the diagnostic model in the training set and the test set; among them, Figure A is the training set; Figure B is the test set. Specific implementation manner

[0056] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Many specific details are set forth in the following description in order to fully understand the present invention, but the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0057] Example 1: Screening of combined markers for non-alcoholic fatty liver disease grading diagnosis and construction of a diagnostic model (1) Collection of plasma samples from NAFLD patients with different degrees of illness

[0058] All subjects were evaluated for liver CAP value using transient elastography technology, and the patients were graded according to the CAP value: mild NAFLD was diagnosed when the CAP value was in the range of 230 - 248 dB / m; moderate NAFLD was diagnosed when the CAP value was in the range of 248 - 269 dB / m; severe NAFLD was diagnosed when the CAP value ≥ 269 dB / m; at the same time, healthy people without NAFLD (CAP value below 230 dB / m) were selected as the control group. After determining the different progression stages of NAFLD, plasma samples of each group of subjects were collected for subsequent metabolomics and lipidomics analysis. All plasma samples need to be separated and frozen to ensure that the samples are not contaminated or degraded before analysis.

[0059] Adopting the above scheme, the present invention collected plasma samples from 120 NAFLD patients and 40 healthy controls. The 120 NAFLD patients included 40 cases of mild, 40 cases of moderate, and 40 cases of severe.

[0060] (2) Wide-targeted metabolomics analysis of plasma of NAFLD patients

[0061] It includes the following steps:

[0062] a. Sample pretreatment: Take 100 μL of plasma sample, add 10 μL of internal standard solution (containing 50 μg / mL L-2-chlorophenylalanine and 100 μg / mL heptadecanoic acid), add 300 μL of methanol, vortex mix for 3 min, and centrifuge at 12000 rpm and 4 °C for 5 min. Transfer the supernatant to an EP tube and dry it under a nitrogen stream. Dissolve the residue in 50 μL of 50% methanol-water, vortex for 3 min, sonicate for 5 min, centrifuge at 12000 rpm and 4 °C for 5 min, and take the supernatant as the test solution;

[0063] b. Set chromatographic conditions: The chromatographic column is an octadecylsilane-bonded silica gel column. Use 0.1 vol% formic acid aqueous solution (A)-0.1 vol% formic acid acetonitrile (B) as the mobile phase, with a flow rate of 0.4 mL·min -1 , the column temperature is 30 °C, the injection volume is 2 μL, and elution is carried out in a gradient elution manner; among them, the elution program is: 0 - 2.1 min, 100% → 100% mobile phase A, 0% → 0% mobile phase B; 2.1 - 14.0 min, 100% → 5% mobile phase A, 0% → 95% mobile phase B; 14.0 - 16.0 min, 5% → 5% mobile phase A, 95% → 95% mobile phase B; 16.0 - 16.1 min, 5% → 100% mobile phase A, 95% → 0% mobile phase B;

[0064] c. Set mass spectrometry conditions: Use an electrospray ionization source, detect in positive and negative ion modes; the source temperature is 500 °C; the ion spray voltages are 5500 V (+) and 4500 V (-) respectively; the vacuum degree (10e -5 ) 3.1 Torr; nitrogen is used as the nebulizing gas, auxiliary heating gas and curtain gas, with pressures of 50, 50 and 20 psi respectively; Import the metabolite ion pairs into the mass spectrometry for detection, optimize the DP and CE voltages; Delete the ion pairs without response, select the fragment ions that appear stably to achieve the best detection sensitivity. Finally, 269 ion pairs are retained in the positive ion mode and 136 ion pairs are retained in the negative ion mode.

[0065] d. After equilibrating the system, take the processed sample for mass spectrometry detection.

[0066] Using the above scheme, a total of 405 metabolites in plasma were measured. The chromatograms in positive and negative ion modes are shown in Figure 1 . At the same time, the peak areas of each metabolite and internal standard in each sample were obtained, and the ratio (normalized value) of each metabolite to its corresponding peak area was calculated.

[0067] (3) Plasma wide-targeted lipidomics analysis of NAFLD patients

[0068] Including the following steps:

[0069] a. Sample pretreatment: Take 200 μL of plasma sample in an EP tube, add 40 μL of methanol and 20 μL of internal standard, then vortex for 30 s. Add 1.3 mL of methyl tert-butyl ether-methanol mixed solution (5:1.5, v / v), vortex for 3 min, sonicate for 5 min, add 290 μL of water, centrifuge at 4 °C (12,000 r / min) for 5 min. Take the organic phase and place it in another EP tube to dry in the air stream. Resuspend the residue with 200 μL of methanol, vortex for 3 min, sonicate for 5 min, centrifuge (12,000 r / min) for 5 min, and take the supernatant for injection analysis;

[0070] b. Set chromatographic conditions: The chromatographic column is an octadecylsilyl-bonded silica gel column. Acetonitrile-water (6:4, v / v) (A) - isopropanol-acetonitrile (9:1, v / v) (B) is used as the mobile phase (both containing 10 mmol / L ammonium acetate), and the flow rate is 0.4 mL·min -1 , the column temperature is 25 °C, the injection volume is 2 μL, and gradient elution is carried out. Among them, from 0 to 5.0 min, 70% → 10% mobile phase A, 30% → 90% mobile phase B; from 5.0 to 13.0 min, 10% → 0% mobile phase A, 90% → 100% mobile phase B; from 13.0 to 13.1 min, 0% → 70% mobile phase A, 100% → 30% mobile phase B;

[0071] c. Set mass spectrometry conditions: Use an electrospray ionization source and detect in positive and negative ion modes; the source temperature is 500 °C; the ion spray voltages are 5500 V (+) and 4500 V (-) respectively; the vacuum degree is (10e -5 ) 3.1 Torr; nitrogen is used as the nebulizing gas, auxiliary heating gas and curtain gas, and the pressures are 50, 50 and 20 psi respectively; Inject the obtained lipid reference substances directly into the mass spectrometer through a syringe pump to optimize the relevant mass spectrometry parameters; adopt the "model - prediction" strategy, summarize the rules based on the chromatographic and mass spectrometry behaviors of the existing lipid reference substances, and predict the collision energy CE and declustering voltage DP of unknown lipids (using the carbon chain length CCL and the number of double bonds DBN as independent variables, and the collision energy CE and declustering voltage DP as dependent variables, and establish a multiple linear regression model through SPSS). Import each lipid ion pair into the mass spectrometry detection, delete the ion pairs without response, and select the stably appearing fragment ions to achieve the best detection sensitivity. Finally, 1071 ion pairs in the positive ion mode and 121 ion pairs in the negative ion mode are retained.

[0072] d. After equilibrating the system, take the treated sample for mass spectrometry detection.

[0073] Using the above scheme, a total of 1192 lipids in plasma were measured. The chromatograms in positive and negative ion modes are shown in Figure 2Meanwhile, the peak areas of each lipid and internal standard in each sample were obtained, and the ratio of the peak area of each lipid to its corresponding internal standard was calculated as the normalized value of each lipid.

[0074] (4) Identification of biomarkers for NAFLD grading diagnosis

[0075] It includes the following steps:

[0076] a. Screening of differential metabolites: Using the normalized values of each substance (the ratio of the peak area to the peak area of its internal standard) as variables, multivariate statistical analysis was performed using SIMCA software to obtain the separation trend of samples at each stage and the VIP values of each variable between different groups; univariate statistical analysis was performed using SPSS software to obtain the P and FC values of each variable between different groups; according to the criteria of P < 0.05, VIP > 1.0, and FC > 1.5 or FC < 0.667, differential metabolites and differential lipids between mild NAFLD vs. healthy group, moderate NAFLD vs. healthy group, and severe NAFLD vs. healthy group were screened respectively.

[0077] Using the above-mentioned scheme, the present invention screened 91, 106, and 121 differential metabolites and 122, 295, and 341 differential lipids between mild NAFLD vs. healthy group, moderate NAFLD vs. healthy group, and severe NAFLD vs. healthy group respectively. The volcano plots are shown in Figure 3 .

[0078] b. Biomarker identification: LASSO regression was used to screen out characteristic variables that can reflect different degrees of NAFLD as potential biomarkers; 70% of the sample data was used as the training set, and machine learning model training was performed using Extreme Gradient Boosting (XGBoost), Decision Tree (DT), and Logistic Regression (LR) algorithms respectively; the remaining 30% of the samples were used as the test set, and the confusion matrix of the test set was output to verify the performance of potential biomarkers in NAFLD grading diagnosis.

[0079] Using the above-mentioned scheme, the present invention identified 5 potential biomarkers that can reflect different degrees of NAFLD, namely 5-Aminolevulinic acid, Mesaconic acid, Shikinic acid, Phosphatidylcholine O-35:3 (PC O-35:3), and Phosphatidylinositol-36:2 (PI 36:2), as shown in Figure 4 .

[0080] The diagnostic results of the 5 biomarkers in the three machine learning algorithms are shown in Figure 5。In the XGBoost algorithm, the AUC values of the control group, mild NAFLD, moderate NAFLD, and severe NAFLD in the training set were all 1.00, and the AUC values of the control group, mild NAFLD, moderate NAFLD, and severe NAFLD in the test set were 1.00, 0.95, 0.96, and 1.00 respectively; in the DT algorithm, the AUC values of the control group, mild NAFLD, moderate NAFLD, and severe NAFLD in the training set were all 1.00, and the AUC values of the control group, mild NAFLD, moderate NAFLD, and severe NAFLD in the test set were 1.00, 0.82, 0.90, and 1.00 respectively; in the LR algorithm, the AUC values of the control group, mild NAFLD, moderate NAFLD, and severe NAFLD in the training set were all 1.00, and the AUC values of the control group, mild NAFLD, moderate NAFLD, and severe NAFLD in the test set were 1.00, 0.97, 0.98, and 1.00 respectively.

[0081] (5) Construction of the NAFLD grading diagnosis model

[0082] Taking different disease stages as grouping variables, assigning 0 and 1, and using the normalized biomarker data as covariates, binary logistic regression analysis was performed to obtain the coefficients of each variable in the equation, and a NAFLD grading diagnosis model was constructed. Taking the Preg value as the test variable, the grouping values 0 and 1 as the status variables, and setting the status variable value to 1, ROC curve analysis was performed to evaluate the diagnostic ability of the established model.

[0083] Using the above scheme, the constructed NAFLD grading diagnosis model is as follows:

[0084] ① Healthy (-) vs. NAFLD (+): F1 = 1.318X1 + 0.037X2 - 2.151X3 - 0.014X4 + 0.823X5 + 2.94 (area under the ROC curve is 1.00).

[0085] ② Mild NAFLD (-) vs. Moderate / Severe NAFLD (+): F2 = -0.038X1 + 0.009X2 + 0.292X3 - 0.09X4 - 0.167X5 + 4.594 (area under the ROC curve is 0.912).

[0086] ③ Moderate NAFLD (-) vs. Severe group (+): F3 = -2.218X1 + 0.006X2 + 0.349X3 - 2.663X4 + 2.924X5 + 34.235 (area under the ROC curve is 1.00).

[0087] In the above formula, X1 to X5 respectively represent the normalized values of 5-aminolevulinic acid, mesaconic acid, shikinic acid, phosphatidylcholine O-35:3 (PC O-35:3), and phosphatidylinositol-36:2 (PI 36:2) in plasma. The ROC curves of the above models are shown in Figure 6 .

[0088] Substitute the normalized values of the 5 biomarkers in the plasma of the subjects into formula F1. If the result is negative, it is determined as healthy; if the result is positive, it is determined as an NAFLD patient. Substitute the normalized values of the 5 biomarkers in the plasma of NAFLD patients into formula F2. If the result is negative, it is determined as mild NAFLD; if the result is positive, it is determined as moderate / severe NAFLD. Substitute the normalized values of the 5 biomarkers in the plasma of moderate / severe NAFLD patients into formula F3. If the result is negative, it is determined as moderate NAFLD; if the result is positive, it is determined as severe NAFLD.

[0089] To verify the clinical practicability of the model, the accurate recognition rate of the training set samples for NAFLD patients at different stages is 88.3%, while that of the test set samples is 91.7%. These results further prove the good applicability and high accuracy of the model ( Figure 7 ).

Claims

1. A combined biomarker for the graded diagnosis of non-alcoholic fatty liver disease (NAFLD), characterized in that, It consists of the following 5 markers: 5-aminolevulinic acid, aconitic acid, shikimic acid, phosphatidylcholine O-35:3, and phosphatidylinositol-36:2; the hierarchical diagnosis is divided into healthy, mild NAFLD, moderate NAFLD, and severe NAFLD.

2. Use of the combined markers according to claim 1 in the preparation of a non-alcoholic fatty liver disease (NAFLD) diagnostic product, wherein the NAFLD diagnosis is divided into healthy, mild NAFLD, moderate NAFLD, and severe NAFLD.

3. The application according to claim 2, wherein The diagnostic product is a diagnostic kit.

4. A non-alcoholic fatty liver disease (NAFLD) grading and diagnosis system, characterized in that, It includes: (1) Quantitative detection module: a module for quantitatively detecting 5-aminolevulinic acid, aconitic acid, shikimic acid, phosphatidylcholine O-35:3, and phosphatidylinositol-36:2 in a sample and normalizing the values of 5 plasma metabolites; (2) F value calculation module: F1 = 1.318X1 + 0.037X2 - 2.151X3 - 0.014X4 + 0.823X5 + 2.94; F2 = -0.038X1 + 0.009X2 + 0.292X3 - 0.09X4 - 0.167X5 + 4.594; F3 = -2.218X1 + 0.006X2 + 0.349X3 - 2.663X4 + 2.924X5 + 34.235; In the formula, X1 to X5 respectively represent the normalized values of 5-aminolevulinic acid, aconitic acid, shikimic acid, phosphatidylcholine O-35:3, and phosphatidylinositol-36:2 in plasma; (3) Result output module: Using 0 as the cut-off value, output the results of comparing the values of F1, F2, and F3 with 0.

5. Method for constructing a non-alcoholic fatty liver disease (NAFLD) hierarchical diagnosis model: S1. Collection of plasma samples from NAFLD patients with different disease severities: Use the transient elastography technology FibroTouch to evaluate the liver stiffness index CAP of all subjects; classify the patients according to the CAP value and collect the plasma samples of each group of patients; S2. LC-MS / MS-based wide-targeted metabolomics analysis of plasma from NAFLD patients: Use the methanol precipitation protein method for plasma sample pretreatment; use liquid chromatography tandem triple quadrupole mass spectrometry and monitor in the MRM acquisition mode; obtain the normalized values of each plasma metabolite; S3. LC-MS / MS-based wide-targeted lipidomics analysis of plasma from NAFLD patients: Use the Bligh & Dyer method for plasma sample pretreatment; use liquid chromatography tandem triple quadrupole mass spectrometry and monitor in the MRM acquisition mode; obtain the normalized values of each lipid; S4. Identification of biomarkers for NAFLD hierarchical diagnosis: Taking the normalized values of each substance as variables, obtain the P and FC values of each variable between different groups; according to the criteria of P < 0.05, VIP > 1.0, and FC > 1.5 or FC < 0.667, screen out the differential metabolites between the mild NAFLD vs. healthy group, moderate NAFLD vs. healthy group, and severe NAFLD vs. healthy group respectively; Identifying biomarkers that can reflect different degrees of NAFLD based on LASSO regression and machine learning algorithms; S5. Construction of the NAFLD grading diagnosis model: Based on the identified biomarkers, a grading diagnosis model for NAFLD patients with different degrees of illness is established using the binary logistic regression-ROC curve method.

6. The construction method according to claim 5, characterized in that, The specific operation steps of "grading patients according to the CAP value" described in step S1 are as follows: When the CAP value is in the range of 230-248 dB / m, it is diagnosed as mild NAFLD; when the CAP value is in the range of 248-269 dB / m, it is diagnosed as moderate NAFLD; when the CAP value ≥ 269 dB / m, it is diagnosed as severe NAFLD; at the same time, healthy people without NAFLD with a CAP value below 230 dB / m are selected as the control group.

7. The construction method according to claim 5, wherein Step S3 identifies 5 potential biomarkers that can reflect different degrees of NAFLD, namely 5-aminolevulinic acid, mesaconic acid, shikimic acid, phosphatidylcholine O-35:3, and phosphatidylinositol-36:

2.

8. The construction method according to claim 5, characterized in that The specific operation steps of "identifying biomarkers that can reflect different degrees of NAFLD based on machine learning algorithms" described in step S4 are as follows: The least absolute shrinkage and selection operator LASSO is used to screen the characteristic variables of different degrees of NAFLD as potential biomarkers; through 3 machine learning algorithms: extreme gradient boosting XGBoost, decision tree DT, and logistic regression LR, a machine model is constructed to verify the performance of the potential biomarkers in the grading diagnosis of NAFLD.

9. The construction method according to claim 5, characterized in that, The specific operation steps of step S5 are as follows: Using different disease stages as the grouping variable, assigning 0 and 1, and the normalized data of the identified biomarkers as the covariate, binary logistic regression analysis is performed to obtain the coefficients of each variable in the equation, and a NAFLD grading diagnosis expression is constructed; using the Preg value as the test variable, the grouping values 0 and 1 as the status variables, and the status variable value set to 1, ROC curve analysis is performed to evaluate the diagnostic ability of the established model, and the NAFLD grading diagnosis model is constructed as follows: ① Healthy (-) vs. NAFLD (+): F1 = 1.318X1 + 0.037X2 - 2.151X3 - 0.014X4 + 0.823X5 + 2.94; ② Mild NAFLD (+) vs. Moderate / Severe NAFLD (-): F2 = -0.038X1 + 0.009X2 + 0.292X3 - 0.09X4 - 0.167X5 + 4.594; ③ Moderate NAFLD (+) vs. Severe group (-): F3 = -2.218X1 + 0.006X2 + 0.349X3 - 2.663X4 + 2.924X5 + 34.235; In the above formula, X1 to X5 respectively represent the normalized values of 5-aminolevulinic acid, mesaconic acid, shikinic acid, phosphatidylcholine O-35:3 (PC O-35:3), and phosphatidylinositol-36:2 (PI 36:2) in plasma.