Method, system and medium for constructing a liver fibrosis identification model for masld patients
Patent Information
- Application Number
- CN202410965549.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2044-07-18
AI Technical Summary
[0005]本发明提供一种MASLD患者肝纤维化识别模型的构建方法,用以解决现有非侵入性测试在识别较轻的、非晚期的肝纤维化MASLD患者中性能较差的问题,实现非侵入性识别较轻的、非晚期的肝纤维化MASLD患者
[0028]第五方面,本发明提供一种非暂态计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现上述MASLD患者肝纤维化的识别方法。
Smart Images

Figure CN118820953B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical informatics technology, and in particular to a method, system, and medium for constructing a liver fibrosis identification model for MASLD patients. Background Technology
[0002] Metabolic dysfunction-associated fatty liver disease (MASLD), also known as non-alcoholic fatty liver disease (NAFLD), affects approximately one-quarter of the adult population worldwide. Current clinical reference standards for assessing MASLD require biopsy, which is both expensive and invasive, and reportedly results in serious complications in 1% of cases. Therefore, there is an urgent need to develop minimally invasive tools for the effective diagnosis, staging, and monitoring of MASLD progression.
[0003] The phenotype of MASLD is defined as excessive accumulation of hepatic triglycerides, including a disease state transition from steatosis (non-alcoholic fatty liver, NAFL) to non-alcoholic steatohepatitis (NASH), characterized by ballooning hepatocytes and lobular inflammation as liver fibrosis progresses, eventually leading to cirrhosis and hepatocellular carcinoma. Therefore, in MASLD patients, liver fibrosis has been identified as a key histological feature associated with mortality and poor long-term outcomes of liver transplantation. Given the decisive role of liver fibrosis in the clinical outcomes of MASLD, non-invasive tests, such as the NAFLD Fibrosis Score (NFS), have been developed to identify advanced liver fibrosis (F≥3), but these tests have shown poor performance in identifying MASLD patients with milder, non-advanced liver fibrosis.
[0004] Therefore, how to identify patients with milder, non-late-stage liver fibrosis MASLD is a hot topic and a challenge for those skilled in the art. Summary of the Invention
[0005] This invention provides a method for constructing a liver fibrosis identification model for MASLD patients, which solves the problem that existing non-invasive tests have poor performance in identifying milder, non-late-stage MASLD patients with liver fibrosis, and achieves non-invasive identification of milder, non-late-stage MASLD patients with liver fibrosis.
[0006] In a first aspect, the present invention provides a method for constructing a liver fibrosis identification model for MASLD patients, comprising: Step 1: Obtain clinical indicator data and serum lipid molecule species content data of the first target population and the second target population. The first target population refers to non-liver fibrosis patients, and the second target population refers to liver fibrosis patients. Step 2: Calculate and analyze the content data of lipid molecular species that show significant differences in the first target population and the second target population, and select the lipid molecular species corresponding to the content data with p<0.05 and statistical significance as the first candidate identification factor; calculate and analyze the clinical indicator data that are significantly correlated with the content data corresponding to the first candidate identification factor, and select the clinical indicator data corresponding to the clinical indicator data with p<0.05 and statistical significance as the second candidate identification factor. Step 3: Construct a basic MASLD patient liver fibrosis identification model through Lasso regression analysis, and select the optimal MASLD patient liver fibrosis identification model from the basic MASLD patient liver fibrosis identification model; based on the optimal MASLD patient liver fibrosis identification model, select the optimal identification factor with a non-zero coefficient from the first candidate identification factor and the second candidate identification factor. Step 4: Establish a liver fibrosis identification model for MASLD patients based on the optimal identification factors, which can be used to identify MASLD patients with liver fibrosis.
[0007] In the above-described construction method, the liver fibrosis status of the target population can be confirmed in various ways, including invasive and non-invasive diagnostic methods, including but not limited to: pathological diagnosis, such as liver biopsy, fully quantitative detection using second harmonic / two-photon excitation fluorescence microscopy; imaging diagnosis; complete blood count and biochemical tests; liver fibrosis spectrum examination, etc. Imaging diagnosis includes but is not limited to: ultrasound, CT, MRI, transient elastography, etc. Complete blood count and biochemical tests include but are not limited to: serological markers (e.g., laminin, fibronectin, type IV collagen and type III procollagen N-terminal propeptide, APRI, capsidase protein, etc.). Preferably, the liver fibrosis status of the target population is confirmed by pathological diagnosis. In one specific embodiment, liver biopsy is used to confirm the liver fibrosis status of the target population. Further, the liver fibrosis status of the target population is classified according to the Kleiner classification. Furthermore, liver fibrosis can be classified into F0-F4 according to the Kleiner classification. In this invention, the Kleiner classification of the second target population can be F1-F4; more preferably, the Kleiner classification of more than 90% of the second target population can be F1-F2.
[0008] In the above construction method, the content data of lipid molecular species in the target population can be obtained through mass spectrometry-based lipidomics technology. The lipid molecular species can be any combination of two or more of the following major categories of lipid molecular species: fatty acyls, glycerides, glycerophospholipids, sterol esters, allyl esters, and sphingolipids. Further, the lipid molecular species can include any combination of one or more categories selected from glycerides, glycerophospholipids, sterol esters, and sphingolipids. Even further, sterol esters include oxidized sterols and sterols.
[0009] In the above construction method, the clinical indicator data of the target population may include any combination of two or more of the following indicators: liver fibrosis score, physiological indicators related to fatty liver, and blood biochemical indicators related to fatty liver.
[0010] Specifically, the liver fibrosis score can be selected from one or more of the following: NAFLD activity score, lobular inflammation NAS score, Knodell score, Scheuer score, Ishak score, METAVIR stage, FIB-4 index, etc. Physiological indicators related to fatty liver may include body mass index (BMI), height, weight, blood pressure (e.g., systolic blood pressure, diastolic blood pressure), etc. Blood biochemical indicators related to fatty liver may include alanine aminotransferase (ALT), aspartate aminotransferase (AST), gamma-glutamyl transferase (GGT), triglycerides, total cholesterol, high-density lipoprotein cholesterol, etc.
[0011] In the above construction method, in step 2, the content data of lipid molecular species of the first target population and the second target population are subjected to t-test, and the p value is corrected by multiple tests using the Benjamini Hochberg method. The lipid molecular species corresponding to the content data with p < 0.05 after correction are selected as the first candidate recognition factor. The correlation between the content data and clinical indicator data corresponding to the first candidate identification factor was tested to screen out the second candidate identification factor. Spearman correlation test was used for the two numerical variables, Kruskal-Wallis test was used for the categorical and numerical variables, and Fisher's Exact test was used for the two categorical variables.
[0012] In the above construction method, step 3, constructing a basic MASLD patient liver fibrosis identification model and selecting the optimal MASLD patient liver fibrosis identification model, includes the following steps: All data from the target population are divided into a training set and a validation set. The training set data is input into the basic MASLD patient liver fibrosis identification model. Hyperparameter (λ) tuning is performed, and a classification cutoff value is determined to obtain the optimal MASLD patient liver fibrosis identification model. Then, the performance of the optimal MASLD patient liver fibrosis identification model, including classification performance and prediction performance, is evaluated using the validation set data. In one specific implementation, the training set and validation set data are divided in a 7:3 ratio.
[0013] In the construction method of this invention, the optimal λ value for model performance can be selected through cross-validation (e.g., 5x cross-validation). The classification cutoff value can be determined by plotting an ROC curve and selecting the cutoff value that balances the true positive rate (TPR) and false positive rate (FPR), i.e., the point with the largest area under the curve (AUC). In a preferred embodiment, the "closest topleft method" is used to determine the classification cutoff value.
[0014] In the above construction method, the optimal identification factor obtained in step 3 may include any combination of two or more of the following clinical indicators: NAFLD activity score, lobular inflammation NAS score, body mass index, diastolic blood pressure, and aspartate aminotransferase level; and content data of any combination of two or more of the following lipid molecular species: diglycerides 34:0 (16:0-18:0), cerebroside sulfate d18:1 / 20:0, lysophosphatidylserine 18:1, glucosamine d18:1 / 22:0, bis(monoacyl) Glyceryl phosphate 38:6 (18:2-20:4), 7-keto-27-hydroxy-cholesterol, cerebroside sulfate d18:1 / 18:0h, triglycerides 52:4 (16:0), bis(monoacylglycerol) phosphate 38:5 (16:0-22:5), triglycerides 58:7 (20:4), cerebroside sulfate d18:1 / 18:1h, phosphatidic acid 32:2, bis(monoacylglycerol) phosphate 38:3 (18:0-20:3), 4β-hydroxycholesterol, glucosamine ceramide d18:1 / 20:0.
[0015] Specifically, diglyceride 34:0 (16:0-18:0) has the ID number HMDB0007156 in the HMDB database; cerebroside sulfate d18:1 / 20:0 has the ID number HMDB0012315 in the HMDB database; lysophosphatidylserine 18:1 has the ID number HMDB0240603 in the HMDB database; and glucosamine ceramide d18:1 / 22:0 has the ID number HMDB000 in the HMDB database. 4974; Bis(monoacylglycerol) phosphate 38:6 (18:2-20:4) refers to the total carbon chain length of the two fatty acid chains linked to the glycerol ester bond being 38, with a total unsaturation of 6. One fatty acid chain has a carbon chain length of 18 and an unsaturation of 2, while the other fatty acid chain has a carbon chain length of 20 and an unsaturation of 4. The CAS number for 7-keto-27-hydroxy-cholesterol is CAS148988-28-7. Cerebroside sulfate d18:1 / 18:0h in LIPID The ID in MAP is LMSP06020004; triglyceride 52:4 (16:0) refers to the total length of the three fatty acid chains linked to the glyceride bond being 52, with a total unsaturation of 4, one of which has a carbon chain length of 16 and an unsaturation of 0, while the length and unsaturation of the other two fatty acid chains are not further restricted; the carbon chain length and unsaturation information for bis(monoacylglycerol) phosphate 38:5 (16:0_22:5) and bis(monoacylglycerol) phosphate 38:3 (18:0_20:3) refer to bis(monoacylglycerol) phosphate 38:6 (18:2_20:4); the carbon chain length and unsaturation information for triglyceride 58:7 (20:4) refer to triglyceride 52:4 (16:0); the number of cerebroside sulfate d18:1 / 18:1h in the PubChem database is PubChem CID. 164449616; Phosphatidic acid 32:2 refers to the total carbon length of the two fatty acid chains linked to the glycerol ester bond being 32, and the total degree of unsaturation being 2; CAS number of 4β-hydroxycholesterol is 17320-10-4; ID number of glucose ceramide d18:1 / 20:0 in the HMDB database is HMDB0004973.
[0016] Furthermore, the optimal identification factors include NAFLD activity score, lobular inflammation NAS score, body mass index, diastolic blood pressure, aspartate aminotransferase level, diglyceride 34:0 (16:0-18:0) content data, cerebroside sulfate d18:1 / 20:0 content data, lysophosphatidylserine 18:1 content data, glucosamine ceramide d18:1 / 22:0 content data, bis(monoacylglycerol) phosphate 38:6 (18:2-20:4) content data, 7-keto-27-hydroxy-cholesterol content data, cerebroside sulfate d18:1 / 18:0h content data, triglyceride 52:4 (16:0) content data, bis(monoacylglycerol) phosphate 38:5 (16:0-22:5) content data, and triglyceride 58:7 (20:4) content data. Content data for: cerebroside sulfate d18:1 / 18:1h, phosphatidic acid 32:2, bis(monoacylglycerol) phosphate 38:3 (18:0-20:3), 4β-hydroxycholesterol, and glucosamine ceramide d18:1 / 20:0.
[0017] In the above construction method, step 4 also includes obtaining the probability value of liver fibrosis in MASLD patients through a logical function based on the best identification factors and coefficients selected in the previous steps.
[0018] Since the liver fibrosis identification model of this invention addresses a classification task (i.e., whether or not there is liver fibrosis) rather than a regression problem, linear regression may lead to decreased classification performance when solving some classification tasks (e.g., imbalanced class distribution, such as in the target population of this embodiment where the proportion of people with liver fibrosis is much greater than that of those without fibrosis). Therefore, it is preferable to further utilize a logistic function to transform linear regression into a 0-1 classification problem, thereby the output of the liver fibrosis identification model for MASLD patients is a probability value rather than a regression equation. Furthermore, the logistic function is a sigmoid function.
[0019] Secondly, the present invention provides a method for identifying liver fibrosis in patients with MASLD, comprising the following steps: Clinical data and serum lipid molecular species content data of patients with MASLD were obtained; the clinical data included NAFLD activity score, lobular inflammation NAS score, body mass index, diastolic blood pressure, and aspartate aminotransferase level; the lipid molecular species included diglycerides 34:0 (16:0-18:0), cerebroside sulfate d18:1 / 20:0, lysophosphatidylserine 18:1, glucosamine d18:1 / 22:0, and bis(monoacylglycerol) phosphate 38:6 (18: 2_20:4), 7-keto-27-hydroxy-cholesterol, cerebroside sulfate d18:1 / 18:0h, triglycerides 52:4 (16:0), bis(monoacylglycerol) phosphate 38:5 (16:0_22:5), triglycerides 58:7 (20:4), cerebroside sulfate d18:1 / 18:1h, phosphatidic acid 32:2, bis(monoacylglycerol) phosphate 38:3 (18:0_20:3), 4β-hydroxycholesterol, glucosamine ceramide d18:1 / 20:0; The clinical data and the content data of lipid molecular species in serum are input into any of the above-mentioned MASLD patient liver fibrosis identification models to obtain the probability value of the MASLD patient to be tested having liver fibrosis.
[0020] It should be noted that for the prediction results of classification prediction models, one approach is to directly use the predicted probability value as the prediction result, the magnitude of which represents the likelihood of a positive result. Another approach is to determine the category based on a probability threshold. The probability threshold can be set as a decision threshold to adjust the model's classification performance, such as using the default 0.5, or determined by certain parameters, such as the Youden index, the Clinical Decision Achievement (DCA) curve (incorporating the doctor's consideration of the risk-benefit ratio), etc.
[0021] In this invention, the above-mentioned identification method can be an identification method implemented by a computer.
[0022] In one specific embodiment, the method further includes a classification step: When the probability of a patient with MASLD having liver fibrosis is greater than or equal to a preset decision threshold, the patient is classified as a patient with liver fibrosis; when the probability of a patient with MASLD having liver fibrosis is less than the preset decision threshold, the patient is classified as a patient without liver fibrosis.
[0023] In one specific implementation, the "closest topleft method" is used to select the model critical value for disease classification as the preset decision threshold.
[0024] In one specific implementation, the preset decision threshold used in the above identification method is 0.6014.
[0025] Thirdly, the present invention provides a system for identifying liver fibrosis in patients with MASLD, comprising: The data acquisition module is used to acquire clinical data and serum lipid molecular species content data of patients with MASLD to be tested; the clinical data includes NAFLD activity score, lobular inflammation NAS score, body mass index, diastolic blood pressure, and aspartate aminotransferase level; the lipid molecular species include diglycerides 34:0 (16:0-18:0), cerebroside sulfate d18:1 / 20:0, lysophosphatidylserine 18:1, glucosamine d18:1 / 22:0, bis(monoacylglycerol) phosphate 38: 6 (18:2_20:4), 7-keto-27-hydroxy-cholesterol, cerebroside sulfate d18:1 / 18:0h, triglycerides 52:4 (16:0), bis(monoacylglycerol) phosphate 38:5 (16:0_22:5), triglycerides 58:7 (20:4), cerebroside sulfate d18:1 / 18:1h, phosphatidic acid 32:2, bis(monoacylglycerol) phosphate 38:3 (18:0_20:3), 4β-hydroxycholesterol, glucosamine ceramide d18:1 / 20:0; The data processing module is used to: input the clinical data and the content data of lipid molecular species in serum into the liver fibrosis identification model for MASLD patients constructed by any of the above construction methods, and obtain the probability value of liver fibrosis in the MASLD patient to be tested.
[0026] In one specific embodiment, the identification system further includes a classification module, used to: classify the MASLD patient under test as a liver fibrosis patient when the probability value of the MASLD patient under test having liver fibrosis is greater than or equal to a preset decision threshold; and classify the MASLD patient under test as a non-liver fibrosis patient when the probability value of the MASLD patient under test having liver fibrosis is less than the preset decision threshold.
[0027] Fourthly, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for identifying liver fibrosis in MASLD patients as described above.
[0028] Fifthly, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described method for identifying liver fibrosis in MASLD patients.
[0029] This invention analyzes and processes clinical characteristic data and serum lipid content data of MASLD patients to identify 5 clinical indicators and 15 lipids related to the degree of liver fibrosis in MASLD patients. Based on the 5 clinical indicators and 15 lipids, a liver fibrosis identification model for MASLD patients is constructed, which helps to identify mild liver fibrosis in MASLD patients, assists in clinical decision-making, and has the characteristics of being convenient, fast, and highly accurate. Attached Figure Description
[0030] Figure 1 This is a flowchart illustrating a method for constructing a liver fibrosis identification model for MASLD patients according to an embodiment of the present invention. Figure 2 The statistical results show the significant differences in lipid content between the fibrotic and non-fibrotic groups as displayed in the forest plot. Figure 3 Bar charts of lipids and clinical indicators with non-zero coefficients selected for Lasso regression model; Figure 4 The receiver operation characteristic curves for the three recognition models; Figure 5a The three recognition models demonstrate a balanced accuracy across the training and validation sets. Figure 5b The F1 scores of the three recognition models are shown on the training and validation sets. Figure 6 This is a flowchart illustrating a method for identifying liver fibrosis in MASLD patients according to an embodiment of the present invention. Figure 7 This invention provides a system for identifying liver fibrosis in MASLD patients according to an embodiment of the present invention. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, embodiments of this invention, and should not be construed as limiting the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. In the description of this invention, it should be understood that the terminology used is for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0032] Patients with histologically confirmed MASLD were included in this study. Exclusion criteria included: 1. heavy alcohol consumption (men >140 g / week, women >70 g / week); 2. use of drugs inducing steatosis; 3. diagnosis of viral hepatitis, autoimmune hepatitis, or other known chronic liver disease. A total of 519 MASLD patients were ultimately recruited for this study, which was approved by the local ethics committee of the First Affiliated Hospital of Wenzhou Medical University, China (No. 2016-0246). All enrolled participants (n=519) provided written informed consent.
[0033] Liver biopsies were performed on 519 patients with liver fibrosis (MASLD). Blood samples were collected after at least 8 hours of overnight fasting on the day of the biopsy. The degree of liver fibrosis in the 519 MASLD patients was scored according to the NAFLD Activity Scale (NAS). The fibrosis stages (F0-F4) were determined according to the Kleiner classification: F0 - no fibrosis; F1 - portal or periportal fibrosis only; F2 - perihepatic fibrosis combined with portal or periportal fibrosis; F3 - bridging fibrosis; F4 - cirrhosis. The 519 MASLD patients were divided into 195 non-fibrotic patients (F0 stage) and 324 fibrotic patients (of which 216 were F1 stage, 83 were F2 stage, 22 were F3 stage, and 3 were F4 stage). There were no statistically significant differences in mean age and sex distribution between the two groups.
[0034] Clinical data were collected from 519 patients with MASLD. These data included demographic data (age and sex), anthropometric data (height, weight, body mass index), NAFLD activity score (NAS), lobular inflammation NAS score (NAS_L), systolic blood pressure (mmHg), diastolic blood pressure (mmHg), waist circumference (cm), abdominal circumference, and white blood cell count (WBC *10). 9 / L), Red blood cell count (RBC) (*10 12 / L), hemoglobin (g / L), platelets (*10) 9 (l / L), ALT (U / L), AST (U / L), fasting blood glucose (mmol / L), total cholesterol (mmol / L), total triglycerides (mmol / L), high-density lipoprotein cholesterol (mmol / L), low-density lipoprotein cholesterol (mmol / L), glycated hemoglobin (%), fasting insulin (pmol / L), blood urea nitrogen (mmol / L), creatinine (μmol / L), glomerular filtration rate (EGFR) (ml / min / 1.73m 2 ), uric acid (μmol / L).
[0035] Serum samples were collected from 519 patients with metastatic atherosclerotic leukemia (MASLD), and the content of more than 1200 lipid molecular species in the serum was analyzed. The lipid content analysis method included: fasting serum samples from the 519 MASLD patients were incubated in an extraction solvent (a 1:2 mixture of chloroform and methanol, containing 1% butylated hydroxytoluene) at 4°C and 1500 rpm for 1 h. After incubation, 350 µL of MilliQ water at 0°C and 250 µL of chloroform at 0°C were added to induce phase separation. The samples were centrifuged at 16260 x g at 4°C for 5 min. The lower organic phase was transferred to a new tube. 450 µL of chloroform at 0°C was added to the remaining aqueous phase, and the extraction was repeated once. The organic phases were combined and dried in a SpeedVac vacuum concentrator. The dried sample was resuspended in a mixed solution containing internal standards (the mixed solution used chloroform and methanol in a volume ratio of 1:2, and the internal standards included d9-PC32:0(16:0 / 16:0)(Avanti 860352) and d 31 -PC(16:0 / 18:1)(Avanti860399), d9-PC36:1p(18:0p / 18:1)(Avanti 852475), d7-PE33:1(15:0 / 18:1)(Avanti791638), DMPE(Avanti 850745), d9-PE36:1p(18:0p / 18:1)(Avanti 852474), d 31-PS(16:0 / 18:1)(Forward 860403)、d7-PG33:1(15:0 / 18:1)(Forward 791640)、DMPG(Forward 840445)、d7-PI33:1(15:0 / 18:1)(Forward 86041) (8:0 / 8:0) (Forward 850181)、d7PA33:1(15:0 / 18:1)(Forward 330721)、PA 34:0(17:0 / 17:0)(Forward 830856)、C14-BMP(Forward 850181) d18:1 / 18:1(Forward 791649)、Cer d18:1 / d7-15:0(Forward 860681)、GluCer d18:1 / 8:0(Forward 860540)、d3-LacCer d18:1 / 16:0(Forward 860681) d18:1 / 18:0-d3(Matreya LLC 2052)、Gb3-d18:1 / 17:0(Cayman Chemicals 24876)、SL-d18:1 / 17:0(Avanti 860572)、d7-LPC 18:1 / 17:0(LPE 2053-LPE) 18:1(Forward 791644)、LPA-C17:0(Forward 857324)、LPI-C17:1(Forward 850103)、LPS-C17:1(Forward 858141)、LPG-C17:1(Forward 858127)、DAG(16:0 / 16:0)-d5(Avanti 110537)、DAG(18:1 / 18:1)-d5(Avanti800856)、S1P-d17:1(Avanti 860641)、Sph-d17:1(Avanti 110537) 860640)、TAG(14:0)3-d5(CDNIsotopes D-6958)、TAG(16:0)3-d5(CDN Isotopes D-5815)、TAG(18:0)3-d5(CDN IsotopesD-5816)、CdN Isotopes D-5816) D-5823)、d6-Cho(CDN Isotopes D-2139)、d3-16:0-carnitine(Cayman Chemicals 26569)、d 31-FFA-16:0 (Sigma 366897), d8-FFA-20:4 (Cayman Chemicals 390010) were analyzed by LC-MS to obtain various lipid content data. Phospholipid and sphingolipid analyses were performed on an Exion UPLC coupled with a Sciex 6500 Plus QTRAP, while neutral lipid analyses were performed on an Agilent 1260 HPLC coupled with a Sciex 5500 QTRAP.
[0036] The lipid extracts from the serum samples of the above 519 patients, after extraction and drying, were resuspended in 500 µL of ethanol solution containing 5 µg butylated hydroxytoluene. Internal standards were added, including d7-24-OHCho (Avanti 700018), d7-7β-OHCho (Avanti 700044), d6-25-OHCho (Avanti 700053), d7-27-OHCho (Avanti 700059), d7-7α,24-dihydroxycholestenone (Avanti 700120), d7-7K-Cho (Avanti 700046), d7-7α-hydroxycholestenone (Avanti 700112), d6-lanosterol (Avanti 700090), d7-lathosterol (Avanti 700056), and d6-TMAS (Avanti 700090). d7-β-sitosterol (Avanti 700074), d7-β-dehydrocholesterol (Avanti 700116), and d6-β-sitosterol (Avanti 700148) were extracted. The mixture was incubated at 4°C and 1200 rpm for 15 minutes. 250 µL of MilliQ water and 1 mL of n-hexane were added. The mixture was vortexed thoroughly and then centrifuged at 16260 x g for 5 minutes. The clear upper phase was transferred to a new tube. The remaining aqueous phase was extracted once more with 1 mL of n-hexane. The combined extracts were dried in a SpeedVac and derivatized to obtain oxidized sterols and pyridine carboxylates of sterols, followed by LC / MS analysis to obtain the contents of sterols and oxidized sterols.
[0037] Figure 1 This is a flowchart illustrating a method for constructing a liver fibrosis identification model for MASLD patients according to an embodiment of the present invention. Figure 1 As shown, the construction method includes the following steps: Step 1: Obtain clinical indicator data and serum lipid molecule species content data of the first target population and the second target population. The first target population refers to non-liver fibrosis patients, and the second target population refers to liver fibrosis patients. Step 2: Calculate and analyze the content data of lipid molecular species that show significant differences in the first target population and the second target population, and select the lipid molecular species corresponding to the content data with p<0.05 and statistical significance as the first candidate identification factor; calculate and analyze the clinical indicator data that are significantly correlated with the content data corresponding to the first candidate identification factor, and select the clinical indicator data corresponding to the clinical indicator data with p<0.05 and statistical significance as the second candidate identification factor. Step 3: Construct a basic MASLD patient liver fibrosis identification model through Lasso regression analysis, and select the optimal MASLD patient liver fibrosis identification model from the basic MASLD patient liver fibrosis identification model; based on the optimal MASLD patient liver fibrosis identification model, select the optimal identification factor with a non-zero coefficient from the first candidate identification factor and the second candidate identification factor. Step 4: Establish a liver fibrosis identification model for MASLD patients based on the optimal identification factors, which can be used to identify MASLD patients with liver fibrosis.
[0038] In one specific embodiment, in step 1, the present invention selects patients with imaging-confirmed hepatic steatosis and / or persistently elevated serum transaminase levels accompanied by metabolic risk factors for liver biopsy.
[0039] In step 2, the content data of lipid molecular species that show significant differences between the first target population and the second target population are calculated and analyzed. The lipid molecular species corresponding to the content data with p < 0.05 and statistical significance are selected as the first candidate identification factor. The clinical indicator data that are significantly correlated with the content data corresponding to the first candidate identification factor are calculated and analyzed. The clinical indicator data corresponding to the clinical indicator data with p < 0.05 and statistical significance are selected as the second candidate identification factor.
[0040] The content data of lipid molecular species in the liver fibrosis group and the non-liver fibrosis group were compared by t-test. The Benjamini Hochberg (BH) method was used to perform multiple test correction on the p-values of all lipids. The lipid molecular species corresponding to the content data with the corrected p-value (FDR) < 0.05 were selected as the first candidate recognition factor.
[0041] Based on the first candidate identification factor, the correlation between the content data and clinical indicator data corresponding to the first candidate identification factor was tested. Spearman correlation test was used for the two numerical variables, Kruskal-Wallis test was used for the categorical and numerical variables, and Fisher's Exact test was used for the two categorical variables. Multiple test correction was performed on the p-values of the nonparametric tests for the three data types. From the clinical indicator data, the clinical indicators corresponding to the clinical indicator data with p < 0.05 and statistical significance were selected as the second candidate identification factor.
[0042] Figure 2 The statistical results of lipid metabolite content data showing significant differences between the fibrotic and non-fibrotic groups as displayed in the forest plot are based on... Figure 2 Thirty lipids were selected as the first candidate recognition factors. Table 1 lists the correlation test results between clinical characteristic data and lipid content data. Based on Table 1, the following parameters were determined: weight, BMI, NAS score, NAS ballooning, NAS lobular inflammation, Fibrosis, diastolic blood pressure (mmHg), waist circumference (cm), abdominal circumference, and white blood cell count (WBC *10). 9 The following were identified as second candidate recognition factors: ALT (U / L), AST (U / L), and high-density lipoprotein cholesterol (mmol / L).
[0043] Table 1. Correlation test results between clinical characteristic data and lipid content data
[0044] Step 3: Construct a basic MASLD patient liver fibrosis identification model through Lasso regression analysis, and select the optimal MASLD patient liver fibrosis identification model from the basic MASLD patient liver fibrosis identification model; based on the optimal MASLD patient liver fibrosis identification model, select the best identification factor with a non-zero coefficient from the first candidate identification factor and the second candidate identification factor.
[0045] In one specific implementation, the entire queue (n=519) is randomly divided into a training set (70%) and a validation set (30%). The data from the training set is input into a Lasso regression model. Based on the hyperparameters (λ) and classification cutoff values obtained from the regression model, the optimal liver fibrosis identification model for MASLD patients is determined. Subsequently, the performance of the optimal liver fibrosis identification model for MASLD patients, including classification performance and prediction performance, is evaluated using validation set data.
[0046] Because a larger proportion of participants were classified as fibrotic (i.e., fibrotic group n=324, non-fibrotic group n=195), a "closest topleft method" strategy was used to select the model critical value for disease classification to achieve a balance between sensitivity and specificity (i.e., finding a critical value in selecting the classification model such that the point on the ROC curve is as close as possible to the top left corner (0,1). This point corresponds to the highest sensitivity (true positive rate) and the lowest 1-specificity (false positive rate)). Due to the inherent randomness in the training / validation dataset splitting and model λ tuning using 5x cross-validation, this embodiment repeated training and validation 100 times to obtain a more reliable model performance estimate. The average value of λ was calculated from the 100 training and validation iterations and used to train the Lasso regression model.
[0047] like Figure 3 As shown, the clinical parameters and lipids with non-zero coefficients selected for constructing a liver fibrosis identification model for MASLD patients include diglycerides 34:0 (16:0_18:0) [DAG34:0 (16:0_18:0)], cerebroside sulfate d18:1 / 20:0 [SL d18:1 / 20:0], phosphatidylserine 18:1 [LPS18:1], glucosamine ceramide d18:1 / 22:0 [GluCer d18:1 / 22:0], bis(monoacylglycerol) phosphate 38:6 (18:2_20:4) [BMP 38:6 (18:2_20:4)], 7-keto-27-hydroxy-cholesterol [7K-27OH-CHO], and cerebroside sulfate d18:1 / 18:0h [SL... d18:1 / 18:0h】、Triglycerides 52:4(16:0)
TAG 52:4(16:0)
BMP38:5(16:0_22:5)
TAG58:7(20:4)
SLd18:1 / 18:1h
PA 32:2
BMP 38:3(18:0_20:3)
[0048] The coefficients corresponding to the above clinical indicators and lipids are shown in Table 2.
[0049] Table 2. Coefficients corresponding to clinical indicators and lipid levels.
[0050] Step 4: Establish a liver fibrosis identification model for MASLD patients based on the optimal identification factors, which can be used to identify MASLD patients with liver fibrosis.
[0051] Based on the optimal identification factors and coefficients selected in the preceding steps, the probability value of liver fibrosis in MASLD patients is obtained through the sigmoid function, and a liver fibrosis identification model for MASLD patients is constructed. Simultaneously, a decision threshold is set based on the probability value, specifically 0.6014.
[0052] Furthermore, this invention also constructs a single clinical model and a lipid model, and evaluates the efficacy of the constructed clinical model, lipid model, and clinical + lipid model, such as... Figure 4 As shown, the area under the receiver operating curve (AUROC) for the clinical + lipid combination was the highest, at 0.775 (95% CI: 0.735–0.816). Further comparison of the performance of the two classification models using the DeLong test (comparing the performance of the two ROC curves) showed that the "clinical + lipid" combination significantly outperformed both the clinical model alone (P = 8.58E-07, DeLong test) and the lipid model alone (P = 3.26E-04, DeLong test).
[0053] like Figures 5a-5b As shown in Table 3, the balance accuracy and F1 score obtained by the "clinical + lipid" combined model are better than those of the clinical model and lipid model alone.
[0054] Table 3 Performance evaluation of the three identification models
[0055] Figure 6 This is a flowchart illustrating a method for identifying liver fibrosis in MASLD patients according to an embodiment of the present invention. Figure 6 As shown, it includes the following steps: Clinical data and serum lipid molecular species content data of patients with MASLD were obtained; the clinical data included NAFLD activity score, lobular inflammation NAS score, body mass index, diastolic blood pressure, and aspartate aminotransferase level; the lipid molecular species included diglycerides 34:0 (16:0-18:0), cerebroside sulfate d18:1 / 20:0, lysophosphatidylserine 18:1, glucosamine d18:1 / 22:0, and bis(monoacylglycerol) phosphate 38:6 (18: 2_20:4), 7-keto-27-hydroxy-cholesterol, cerebroside sulfate d18:1 / 18:0h, triglycerides 52:4 (16:0), bis(monoacylglycerol) phosphate 38:5 (16:0_22:5), triglycerides 58:7 (20:4), cerebroside sulfate d18:1 / 18:1h, phosphatidic acid 32:2, bis(monoacylglycerol) phosphate 38:3 (18:0_20:3), 4β-hydroxycholesterol, glucosamine ceramide d18:1 / 20:0; The clinical data and the content data of lipid molecular species in serum are input into any of the above-mentioned MASLD patient liver fibrosis identification models to obtain the probability value of the MASLD patient to be tested having liver fibrosis.
[0056] In one specific embodiment, the method further includes: classifying the MASLD patient as a liver fibrosis patient when the probability value of the MASLD patient to be tested is greater than or equal to a preset decision threshold; and classifying the MASLD patient as a non-liver fibrosis patient when the probability value of the MASLD patient to be tested is less than the preset decision threshold.
[0057] Figure 7 A system for identifying liver fibrosis in MASLD patients according to an embodiment of the present invention includes: The data acquisition module is used to acquire clinical data and serum lipid molecular species content data of patients with MASLD to be tested; the clinical data includes NAFLD activity score, lobular inflammation NAS score, body mass index, diastolic blood pressure, and aspartate aminotransferase level; the lipid molecular species include diglycerides 34:0 (16:0-18:0), cerebroside sulfate d18:1 / 20:0, lysophosphatidylserine 18:1, glucosamine d18:1 / 22:0, bis(monoacylglycerol) phosphate 38: 6 (18:2_20:4), 7-keto-27-hydroxy-cholesterol, cerebroside sulfate d18:1 / 18:0h, triglycerides 52:4 (16:0), bis(monoacylglycerol) phosphate 38:5 (16:0_22:5), triglycerides 58:7 (20:4), cerebroside sulfate d18:1 / 18:1h, phosphatidic acid 32:2, bis(monoacylglycerol) phosphate 38:3 (18:0_20:3), 4β-hydroxycholesterol, glucosamine ceramide d18:1 / 20:0; The data processing module is used to: input the clinical data and the content data of lipid molecular species in serum into the liver fibrosis identification model for MASLD patients constructed by any of the above construction methods, and obtain the probability value of liver fibrosis in the MASLD patient to be tested.
[0058] In one specific embodiment, the identification system further includes a classification module, used to: classify the MASLD patient under test as a liver fibrosis patient when the probability value of the MASLD patient under test having liver fibrosis is greater than or equal to a preset decision threshold; and classify the MASLD patient under test as a non-liver fibrosis patient when the probability value of the MASLD patient under test having liver fibrosis is less than the preset decision threshold.
[0059] This invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the program to implement the method for identifying liver fibrosis in MASLD patients as described above.
[0060] This invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, characterized in that the computer program, when executed by a processor, implements the method for identifying liver fibrosis in MASLD patients as described above.
[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a liver fibrosis identification model for MASLD patients, characterized in that, include: Step 1: Obtain clinical indicator data and serum lipid molecule species content data of the first target population and the second target population. The first target population refers to non-liver fibrosis patients, and the second target population refers to liver fibrosis patients. Step 2: Calculate and analyze the content data of lipid molecular species that show significant differences between the first and second target populations. Select the lipid molecular species corresponding to the content data with p < 0.05 and statistical significance as the first candidate identification factor. Calculate and analyze the clinical indicator data that are significantly correlated with the content data corresponding to the first candidate identification factor. Select the clinical indicator data corresponding to the clinical indicator data with p < 0.05 and statistical significance as the second candidate identification factor. Specifically, perform a t-test on the content data of lipid molecular species in the first and second target populations, and use the Benjamini Hochberg method to perform multiple test correction on the p-value. Select the lipid molecular species corresponding to the content data with p < 0.05 after correction as the first candidate identification factor. The correlation between the content data and clinical indicator data corresponding to the first candidate identification factor was tested to screen out the second candidate identification factor. Spearman correlation test was used for the two numerical variables, Kruskal-Wallis test was used for the categorical and numerical variables, and Fisher's Exact test was used for the two categorical variables. Step 3: Construct a basic MASLD patient liver fibrosis identification model through Lasso regression analysis, and select the optimal MASLD patient liver fibrosis identification model from the basic MASLD patient liver fibrosis identification model; based on the optimal MASLD patient liver fibrosis identification model, select the optimal identification factor with a non-zero coefficient from the first candidate identification factor and the second candidate identification factor. The optimal identification factors include NAFLD activity score, lobular inflammation NAS score, body mass index, diastolic blood pressure, aspartate aminotransferase (AST) level, diglyceride 34:0 (16:0-18:0) content data, cerebroside sulfate d18:1 / 20:0 content data, lysophosphatidylserine 18:1 content data, glucosamine ceramide d18:1 / 22:0 content data, bis(monoacylglycerol) phosphate 38:6 (18:2-20:4) content data, 7-keto-27-hydroxy-cholesterol content data, cerebroside sulfate d18:1 / 18:0h content data, triglyceride 52:4 (16:0) content data, bis(monoacylglycerol) phosphate 38:5 (16:0-22:5) content data, and triglyceride 58:7 (20:4) content data. Content data for: cerebroside sulfate d18:1 / 18:1h, phosphatidic acid 32:2, bis(monoacylglycerol) phosphate 38:3 (18:0-20:3), 4β-hydroxycholesterol, and glucosamine ceramide d18:1 / 20:0; Step 4: Establish a liver fibrosis identification model for MASLD patients based on the optimal identification factors, which can be used to identify MASLD patients with liver fibrosis.
2. The construction method according to claim 1, characterized in that, Step 4 also includes obtaining the probability value of liver fibrosis in MASLD patients based on the coefficients corresponding to the best identification factors selected and the best identification factors output by the best MASLD patient liver fibrosis identification model.
3. A method for identifying liver fibrosis in patients with MASLD, characterized in that, The steps include the following: Acquire clinical data and serum lipid molecular species content data of patients with MASLD to be tested; the clinical data include NAFLD activity score, lobular inflammation NAS score, body mass index, diastolic blood pressure, and aspartate aminotransferase level; the lipid molecular species include diglycerides 34:0 (16:0-18:0), cerebroside sulfate d18:1 / 20:0, lysophosphatidylserine 18:1, glucosamine d18:1 / 22:0, bis(monoacylglycerol) phosphate 38:6 (18: 2_20:4), 7-keto-27-hydroxy-cholesterol, cerebroside sulfate d18:1 / 18:0h, triglycerides 52:4 (16:0), bis(monoacylglycerol) phosphate 38:5 (16:0_22:5), triglycerides 58:7 (20:4), cerebroside sulfate d18:1 / 18:1h, phosphatidic acid 32:2, bis(monoacylglycerol) phosphate 38:3 (18:0_20:3), 4β-hydroxycholesterol, glucosamine ceramide d18:1 / 20:0; The clinical data and the content data of lipid molecular species in serum are input into the liver fibrosis identification model for MASLD patients according to any one of claims 1-2 to obtain the probability value of liver fibrosis in the MASLD patient to be tested.
4. The identification method according to claim 3, characterized in that, The method further includes: classifying the MASLD patient as a liver fibrosis patient when the probability value of the MASLD patient to be tested is greater than or equal to a preset decision threshold; and classifying the MASLD patient as a non-liver fibrosis patient when the probability value of the MASLD patient to be tested is less than the preset decision threshold.
5. A system for identifying liver fibrosis in patients with MASLD, characterized in that, include: The data acquisition module is used to acquire clinical data and serum lipid molecular species content data of patients with MASLD to be tested; the clinical data includes NAFLD activity score, lobular inflammation NAS score, body mass index, diastolic blood pressure, and aspartate aminotransferase level; the lipid molecular species include diglycerides 34:0 (16:0-18:0), cerebroside sulfate d18:1 / 20:0, lysophosphatidylserine 18:1, glucosamine d18:1 / 22:0, bis(monoacylglycerol) phosphate 38: 6 (18:2_20:4), 7-keto-27-hydroxy-cholesterol, cerebroside sulfate d18:1 / 18:0h, triglycerides 52:4 (16:0), bis(monoacylglycerol) phosphate 38:5 (16:0_22:5), triglycerides 58:7 (20:4), cerebroside sulfate d18:1 / 18:1h, phosphatidic acid 32:2, bis(monoacylglycerol) phosphate 38:3 (18:0_20:3), 4β-hydroxycholesterol, glucosamine ceramide d18:1 / 20:0; The data processing module is used to: input the clinical data and the content data of lipid molecular species in serum into the liver fibrosis identification model for MASLD patients constructed by the construction method of any one of claims 1-2, and obtain the probability value of liver fibrosis in the MASLD patient to be tested.
6. The identification system according to claim 5, characterized in that, The identification system further includes a classification module, used to: classify the MASLD patient as a liver fibrosis patient when the probability value of the MASLD patient to be tested is greater than or equal to a preset decision threshold; and classify the MASLD patient as a non-liver fibrosis patient when the probability value of the MASLD patient to be tested is less than the preset decision threshold.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for identifying liver fibrosis in MASLD patients as described in claim 3 or 4.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for identifying liver fibrosis in MASLD patients as described in claim 3 or 4.
Citation Information
Patent Citations
Construction method, prediction system and predication equipment of liver fibrosis prediction model based on machine learning method, and storage medium
CN112669960A
Model for evaluating hepatic fibrosis degree of hepatitis B patient
CN112837818A