Microbial marker for liver cirrhosis as well as screening method and application of microbial marker
By combining differential analysis of GWAS and 16S gut microbiota abundance data with Mendelian randomization, gut microbiota biomarkers associated with cirrhosis were screened out. Machine learning algorithms were then used to output early screening prediction risk values, solving the problem of early diagnosis and prevention of cirrhosis in existing technologies and achieving more accurate risk assessment and individualized intervention.
Patent Information
- Application Number
- CN202510999939.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Current research on biomarkers for cirrhosis is limited to a few biomarkers, the clinical cohorts are small, and the risk of future cirrhosis cannot be effectively predicted, making early diagnosis and prevention of cirrhosis difficult.
By combining genome-wide association study (GWAS) data and 16S gut microbiota abundance data, we used differential microbiota analysis, LDSC analysis, and two-sample Mendelian randomization to identify effective gut microbiota-related biomarkers and output early screening prediction risk values through machine learning algorithms.
It provides a more accurate method for early diagnosis and prevention of cirrhosis. By screening out gut microbiota biomarkers closely associated with cirrhosis, it improves the accuracy and sensitivity of the predictive model and guides individualized gut microbiota adjustment to reduce the risk of disease.
Smart Images

Figure CN120895110A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biomedical technology, in particular to a cirrhosis microbial marker and a screening method and application thereof. BACKGROUND
[0002] Cirrhosis is a pathological stage of various chronic liver diseases characterized by chronic inflammation of the liver, diffuse fibrosis, pseudolobule, regenerative nodule and proliferation of intrahepatic and extrahepatic vessels. The symptoms of cirrhosis are diverse, and early symptoms may be asymptomatic, or have non-specific symptoms such as fatigue, loss of appetite, diarrhea, etc. With the progression of the disease, patients may have symptoms such as jaundice, weight loss, fatigue, ascites, coma, etc. Cirrhosis is contagious, and the transmission routes include blood transmission, body fluid transmission and mother-to-child transmission, etc.
[0003] Cirrhosis ranks 11th in the world's death causes, and has become an important cause of morbidity and mortality of chronic liver disease patients worldwide. This is related to changes in people's lifestyle, and more than 25% of cirrhosis deaths worldwide are related to alcohol. With the increase in the number of people drinking and obesity, the disease burden brought by cirrhosis has become more and more serious, and has become a common and increasingly serious public health problem worldwide. In the past decade, the pathogenesis of cirrhosis has been studied in depth, but the treatment of cirrhosis is still largely passive, especially the prevention and treatment of cirrhosis complications still have great difficulties, which seriously affects the mental health and quality of life of cirrhosis patients.
[0004] Current research results on biomarkers for cirrhosis are mostly focused on P53 gene, growth factors, long non-coding RNA and miRNA, etc., which can help achieve early diagnosis of cirrhosis and prevent cirrhosis from evolving into hepatocellular carcinoma. In contrast, intestinal flora biomarkers have lower cost, simpler intervention means, and are increasingly valued. However, previous studies have been limited to a small number of biomarkers, and the clinical cohort size involved is small, and it is not clear whether these biomarkers can effectively predict future cirrhosis risk in diagnosis. SUMMARY
[0005] In order to solve these problems, the present application combines the use of whole genome association study data (GWAS) and 16S intestinal flora abundance data, adopts differential bacteria analysis, LDSC analysis and double-sample Mendelian randomization method to identify effective intestinal flora-related biomarkers, and then outputs early screening prediction risk value of cirrhosis through machine learning algorithm, to provide new methods and basis for early diagnosis and treatment strategy formulation.
[0006] In order to achieve the above object, the present application is realized by the following technical solutions: the screening method of the liver cirrhosis microbial marker provided by the present application uses whole genome association study data (GWAS) and 16S intestinal flora abundance data, combines difference bacteria analysis and double-sample Mendelian randomization method to identify effective intestinal flora related biomarkers, and specifically includes the following steps:
[0007] S1, 16S intestinal flora abundance data of liver cirrhosis and healthy samples is obtained, difference analysis of disease group and control group is carried out, and difference bacteria are screened out;
[0008] S2, GWAS summary statistics data of intestinal flora samples and liver cirrhosis samples is obtained;
[0009] S3, SNPs which are significantly related to intestinal flora characteristics in the whole genome and are independent are screened out, and are assigned to the corresponding flora characteristics;
[0010] S4, related data is extracted from the corresponding flora characteristics and GWAS data of fatty liver, and is sorted and combined, linkage disequilibrium score regression (LDSC) analysis is used to evaluate the genetic correlation between intestinal flora and liver cirrhosis, and LDSC genetically related bacteria are obtained;
[0011] S5, multi-sample Mendelian randomization (MR) analysis is used, combined with sensitivity analysis to evaluate the causal relationship between flora characteristics and liver cirrhosis, and finally the intestinal flora related to liver cirrhosis, i.e. MR related bacteria, is obtained;
[0012] S6, the intersection of difference bacteria, LDSC genetically related bacteria and MR related bacteria is taken as the microbial marker bacteria of liver cirrhosis.
[0013] Preferably, in step S1, the difference analysis of the disease group and the control group specifically uses the method of difference fold and Wilcoxon rank-sum test to calculate the logFC and P value of different flora of the disease group and the control group, and selects the bacteria with logFC>0.5 and p-value<0.05 as the difference bacteria.
[0014] Preferably, in step S3, the selection of instrumental variable SNPs needs to be closely related to intestinal flora and unrelated to liver cirrhosis, therefore, the clump algorithm in plink is used to screen out SNPs which are significantly related to each flora characteristic and p<1x10 -5 in the whole genome, and these SNPs are respectively assigned to the corresponding flora characteristics as their instrumental variables.
[0015] Preferably, in step S4, the related data mainly includes SNP site, allele, effect value, standard error and P value;
[0016] LDSC analysis was used to evaluate the genetic correlation between gut microbiota and cirrhosis, i.e., setting the linkage disequilibrium parameter r for the combined SNPs 2 The threshold was 0.01, the genetic distance was 10000 kb, the independent SNP genetic variation sites were screened out, and the microbiota with genetic correlation with cirrhosis was obtained.
[0017] Preferably, in step S5, the sensitivity analysis includes horizontal pleiotropy and heterogeneity,
[0018] Horizontal pleiotropy uses Mendelian randomization method to perform global test and outlier test on the screened independent SNPs, global test uses significance P value to evaluate the overall horizontal pleiotropy of all SNPs, and outlier test evaluates the existence of specific horizontal pleiotropy outliers by calculating the significance P value of the pleiotropy of each SNP;
[0019] When P≤0.05, it indicates that the overall SNPs have pleiotropy, then the outlier test is performed, and the abnormal SNP with the smallest pleiotropy P value is deleted; the global test is performed again on the remaining SNPs, and the process is repeated until the global test pleiotropy P>0.05 is no longer significant, then the SNPs without horizontal pleiotropy are retained as instrumental variables for subsequent MR analysis;
[0020] The heterogeneity between each SNP is evaluated by IVW analysis and Cochrane's Q test.
[0021] Preferably, it further includes performing two-sample MR analysis: using cirrhosis as the outcome variable of MR analysis, and using gut microbiota as the exposure factor, the gut microbiota used for analysis includes 36 characteristic microbiota at genus level; based on the pooled level data, the causal relationship between the microbiota characteristics and cirrhosis is obtained, specifically by determining the association between the SNP instrumental variables obtained after removing the horizontal pleiotropy and cirrhosis, only retaining microbiota characteristics with 3 or more SNP instrumental variables, and according to the final exposure factor and outcome variable files, the inverse variance weighting method is used to evaluate the relationship between each microbiota characteristic and cirrhosis;
[0022] Among them, the inverse variance weighting method is calculated by dividing the SNP-outcome variable estimate by the SNP-exposure factor estimate, and the determination standard of the influence of microbiota characteristics on cirrhosis is bilateral test P≤0.05.
[0023] A cirrhosis microbial marker is obtained according to the screening method of the cirrhosis microbial marker, and the biomarker bacteria of cirrhosis include the following species: Actinomyces, Veillonella, Prevotella 7, Rikenellaceae RC9gut group, Aggregatibacter.
[0024] The application of a liver cirrhosis microbial marker, an early screening prediction risk value of liver cirrhosis output by a machine learning algorithm, can be used for early screening prediction of liver cirrhosis, and can guide individualized intestinal flora adjustment to reduce the risk of disease, and specifically comprises the following steps:
[0025] (1) Obtain the abundance data corresponding to the liver cirrhosis microbial marker bacteria, divide the abundance data into a training set and a test set according to a set proportion, use the genetically related bacteria and the differential bacteria as input features, input the original machine learning model, cross-validation parameter tuning, determine the optimal model, then train the optimal machine learning model using the training set, test the optimal model using the test set and output the ROC curve, and then perform performance evaluation on the ROC curve of the test set output by the optimal machine learning model;
[0026] (2) Input the new sample data to be tested into the optimal machine learning model trained in step (1) for calculation and prediction, and output an early screening prediction risk value of liver cirrhosis.
[0027] Preferably, the original machine learning model includes Logistic regression, support vector machine, random forest, and Xgboost, wherein the parameters of the original machine learning model are optimized, trained, tested, and evaluated, and the specific steps include the following:
[0028] (11) Data set processing, the bacterial flora abundance data of the liver cirrhosis microbial marker bacteria is randomly divided into a 75% training set and a 25% test set according to a proportion;
[0029] (12) Constructing an original machine learning classifier, using genetically related bacteria and differential bacteria as input features, and sequentially inputting into the original machine learning models of logistic regression, SVM, random forest, and Xgboost;
[0030] (13) Selecting the optimal machine learning model, using cross-validation parameter tuning, selecting the parameter with the highest ROC-AUC score, determining the initial model of the optimal machine learning model, and then using the hyperparameters to fine-tune for 100 iterations to obtain the optimal machine learning model;
[0031] (14) Performance testing, inputting the training set into the optimal machine learning model of step (13) for training, then testing the optimal machine learning model using the test set, and outputting the ROC curve;
[0032] (15) Result evaluation, evaluating the ROC curve and AUC value output by the model in step (14), and evaluating the prediction performance of the optimal machine learning model.
[0033] Preferably, in step (2), specifically comprising: inputting the new sample data to be tested into the trained optimal machine learning model for calculation and prediction to obtain an early screening prediction risk value of liver cirrhosis,
[0034] The prediction risk value is between 0 and 1, wherein the risk value 0 represents that the sample corresponds to a healthy control individual, and the risk value 1 represents that the sample corresponds to a liver cirrhosis patient.
[0035] The determination result of the early screening prediction risk value of liver cirrhosis is as follows: (A) the risk value < 0.4 is determined as a low-risk state, and the intestinal flora does not need to be adjusted; (B) 0.4 <= the risk value <= 0.6 is determined as a medium-risk state, and the intestinal flora needs to be adjusted; (C) the risk value > 0.6 is determined as a high-risk population of liver cirrhosis, and clinical diagnosis is recommended.
[0036] The present application provides a liver cirrhosis microbial marker, a screening method and application thereof.
[0037] Beneficial effects:
[0038] (1) The screening method of the liver cirrhosis microbial marker uses whole genome association study data (GWAS) and 16S intestinal flora abundance data, and adopts differential analysis to screen out differential bacteria related to liver cirrhosis, linkage disequilibrium score regression (LDSC) analysis and double-sample Mendelian randomization method to screen out MR related bacteria and LDSC genetically related bacteria from the perspective of causal association and genetic correlation; the intersection of the differential bacteria, the MR related bacteria and the LDSC genetically related bacteria is taken as the marker, which not only considers the flora abundance difference, but also considers the genetic correlation, ensures the close correlation between the marker and liver cirrhosis, and can more accurately screen out the intestinal flora marker related to liver cirrhosis, thereby providing a new idea for early diagnosis and prevention of liver cirrhosis.
[0039] Meanwhile, the abundance data corresponding to the liver cirrhosis microbial marker bacteria is used as a data set, the genetically related bacteria and the differential bacteria are used as input features, and are input into the original machine learning model for optimization and parameter adjustment, and after the optimal model is determined, the training set and the test set are used to train and test the optimal machine learning model, which can significantly improve the accuracy and sensitivity of the liver cirrhosis prediction model. The optimal machine learning model of the present application can finally output an early screening prediction risk value of liver cirrhosis, realize early screening and risk stratification, not only can prompt the need for clinical diagnosis for the high-risk population, but also can provide guidance for intestinal flora adjustment for the medium-risk population, so as to reduce the risk of disease.
[0040] (2) The liver cirrhosis microbial marker provided by the application not only provides a new biomarker and method for early diagnosis and prevention of liver cirrhosis, but also lays a foundation for formulating individualized intestinal flora adjustment strategies and reducing the risk of liver cirrhosis by elucidating the potential correlation mechanism between intestinal flora and liver cirrhosis, which has important significance for reducing the disease burden caused by liver cirrhosis. BRIEF DESCRIPTION OF DRAWINGS
[0041] Figure 1 A flowchart of the screening method of the liver cirrhosis microbial marker in the application;
[0042] Figure 2 A prediction result data graph of the machine learning model in the application;
[0043] Figure 3 A risk prediction probability result data graph output by the machine learning model in the application. DETAILED DESCRIPTION
[0044] The following reference description of the drawings introduces a plurality of preferred embodiments of the application, so that the technical content of the application is clearer and easier to understand. The application can be embodied in many different forms of embodiments, and the protection scope of the application is not limited to the embodiments mentioned in the text.
[0045] Example 1
[0046] As shown in Figure 1 , the screening method of the liver cirrhosis microbial marker in the application combines the use of genome-wide association study data (GWAS) and 16S intestinal flora abundance data, and uses differential bacteria analysis and two-sample Mendelian randomization method to identify effective intestinal flora-related biomarkers, which specifically includes the following steps:
[0047] S1, obtain 16S intestinal flora abundance data containing liver cirrhosis and healthy samples, perform differential analysis of the disease group and the control group, and screen to obtain differential bacteria. Specifically, download the 16S sequencing data of intestinal flora of liver cirrhosis and healthy controls from the European Bioinformatics Institute (EBI).
[0048] Using the difference fold and Wilcoxon rank-sum test method, the logFC and P value of different flora of the disease group and the control group are calculated, and bacteria with logFC>0.5 and p-value<0.05 are screened out as differential bacteria.
[0049] S2, obtain GWAS summary statistics of gut microbiota samples and liver cirrhosis samples, specifically, the summary statistics of gut microbiota samples come from four datasets: multiple cohorts of MiBioGen consortium, Dutch Microbiome Project (DMP), German individuals, FINRISK 2002 (FR02) cohort.
[0050] Among them, the MiBioGen consortium cohort contains 21 sub-cohorts, with 18340 participants from different regions such as Europe, Asia, Africa and America, the cohort includes 211 taxonomic units, respectively 131 genera, 35 families, 20 orders, 16 classes and 9 phyla, but contains 15 unknown taxonomic units. The Dutch Microbiome Project includes 7738 participants from the Netherlands, and identifies 207 taxonomic units and 205 functional pathways through shotgun metagenomic sequencing of fecal samples, only using genus-level bacteria not included in the MiBioGen cohort. The German individual cohort of Ruhlemann et al. consists of 8956 participants, identified by 16S rRNA sequence, the cohort identifies 430 taxonomic units from phylum to genus, including abundance and prevalence formats, only using genus-level bacteria with abundance format and not included in the MiBioGen cohort. The FR02 cohort has 5959 Finnish participants, identified 473 taxonomic units from phylum to species through fecal shotgun sequencing, only using genus-level bacteria not included in the MiBioGen cohort and without unknown taxonomic units, and finally 144 genus-level bacteria are included.
[0051] The GWAS summary statistics of liver cirrhosis samples come from the FinnGen dataset, with participants of European ancestry and a sample size of 496506.
[0052] S3, screen out SNPs that are significantly associated with gut microbiota features in the whole genome and independent, and assign them to the corresponding microbiota features;
[0053] The selection of instrumental variables needs to be closely related to the exposure factor (gut microbiota) and unrelated to the outcome variable (liver cirrhosis), therefore, the clump algorithm in plink is used to screen out SNPs that are significantly associated with each microbiota feature in the whole genome and independent, and p < 1 x 10 -5 , and these SNPs are respectively assigned to the corresponding microbiota features as their instrumental variables.
[0054] S4, relevant data including SNP sites, alleles, effect values, standard errors and P values were extracted from corresponding gut flora characteristics and GWAS data of fatty liver, and were sorted and combined. Linkage disequilibrium score regression (LDSC) analysis was used to evaluate the genetic correlation between intestinal flora and cirrhosis, and to test the presence of sample overlap, with reference to LD scores from the European 1000 Genomes Project dataset. The combined SNP sites were set with linkage disequilibrium (LD) parameter r 2 threshold of 0.01 and genetic distance of 10000 kb to screen out independent SNP genetic variation sites, and obtain LDSC genetic correlation bacteria. In this embodiment, only genus-level bacteria were used for analysis, and the LDSC results showed that after controlling sample overlap, there was a whole genome genetic correlation between intestinal flora and cirrhosis.
[0055] S5, multi-sample Mendelian randomization (MR) analysis was used to evaluate the causal relationship between flora characteristics and cirrhosis, and finally the intestinal flora related to cirrhosis was obtained, i.e. MR correlation bacteria;
[0056] The sensitivity analysis includes horizontal pleiotropy and heterogeneity. The horizontal pleiotropy uses Mendelian randomization method to perform global test and outlier test on the screened independent SNPs. The global test uses significance P value to evaluate the overall horizontal pleiotropy of all SNPs, and the outlier test evaluates the presence of specific horizontal pleiotropy outliers by calculating the significance P value of each SNP pleiotropy. When P≤0.05, it indicates that the overall SNPs have pleiotropy, then the outlier test is performed, and the abnormal SNPs with the smallest pleiotropy P value are deleted. The remaining SNPs are tested again, and the process is repeated until the global test pleiotropy is no longer significant, at which time P>0.05, then the SNPs without horizontal pleiotropy are retained as instrumental variables for subsequent MR analysis. The heterogeneity between each SNP is evaluated by IVW analysis and Cochrane's Q test.
[0057] Two-sample MR analysis was performed, using cirrhosis as the outcome variable of Mendelian randomization analysis and intestinal flora as the exposure factor. The intestinal flora used for analysis included 36 characteristic flora at genus level.
[0058] Based on the aggregate level data, the causal relationship between the flora characteristics and liver cirrhosis is obtained, specifically including: determining the association of the SNP instrumental variable obtained after removing the multi-effect with liver cirrhosis, only retaining the flora characteristics of 3 SNP instrumental variables and above, and according to the final exposure factor and outcome variable file, using inverse variance weighting method to evaluate the relationship between each flora characteristic and liver cirrhosis; wherein, the inverse variance weighting method is calculated by dividing the SNP-outcome variable estimation value by the SNP-exposure factor estimation value, and the two-sample MR analysis is based on the aggregate level data, and there is no clear requirement such as normality for the analysis of exposure and outcome data. In this embodiment, the determination criterion of the influence of the flora characteristics on liver cirrhosis is two-sided test P≤0.05.
[0059] S6, taking the intersection of the difference bacteria and the LDSC genetically related bacteria and the MR related bacteria as the biomarker bacteria of liver cirrhosis, including 5 bacteria: Actinomyces, Veillonella, Prevotella 7, Rikenellaceae RC9 gut group and Aggregatibacter.
[0060] The liver cirrhosis microbial marker screened by the method of the application uses whole genome association study data (GWAS) and 16S intestinal flora abundance data, and adopts difference bacteria analysis to screen candidate microorganisms related to liver cirrhosis, linkage disequilibrium score regression (LDSC) analysis and two-sample Mendelian randomization method to screen MR related bacteria and LDSC genetically related bacteria from the perspective of causal correlation and genetic correlation; taking the intersection of the difference bacteria, the MR related bacteria and the LDSC genetically related bacteria as the marker, both the flora abundance difference and the genetic correlation are considered, the close correlation of the marker with liver cirrhosis is ensured, and the intestinal flora marker related to liver cirrhosis can be more accurately screened, thereby providing a new idea for early diagnosis and prevention of liver cirrhosis.
[0061] Embodiment 2
[0062] As shown in Figure 1 , the application further outputs an early screening prediction risk value of liver cirrhosis through a machine learning algorithm, and a supervised learning is a function generated by a corresponding relationship between a part of input data and output data, which maps the input to a suitable output. The sample data of the application has been clinically diagnosed and has a classified label, so the supervised machine learning classification model will be explored and selected. The abundance values corresponding to the biomarker bacteria of all samples are taken as input data, and the diagnosis results of the samples are taken as output classification labels.
[0063] The early screening prediction risk value of liver cirrhosis is output through a machine learning algorithm, specifically including the following steps:
[0064] (1) Obtain the abundance data corresponding to the microbial marker bacteria of liver cirrhosis, divide the abundance data into a training set and a test set according to a set proportion, input the genetic correlation bacteria and the differential bacteria as input features into an original machine learning model, cross-validation parameter tuning is performed to determine an optimal model, the training set is used to train the optimal machine learning model, the test set is used to test the optimal model and output an ROC curve, and performance evaluation is performed on the ROC curve of the test set output by the optimal machine learning model; the original machine learning model includes Logistic regression, support vector machine, random forest, and Xgboost, wherein, the parameters of the original machine learning model are optimized, trained, tested, and evaluated, and the specific steps include the following:
[0065] (11) Data set processing, the abundance data of the microbial marker bacteria of liver cirrhosis is randomly divided into a 75% training set and a 25% test set according to a proportion;
[0066] (12) Constructing an original machine learning classifier, the genetic correlation bacteria and the differential bacteria are combined as input features, and are sequentially input into the original machine learning models of logistic regression, SVM, random forest, and Xgboost;
[0067] (13) Selecting an optimal machine learning model, using cross-validation parameter tuning, selecting the parameters with the highest ROC-AUC score, determining the initial model of the optimal machine learning model, and using the hyperparameters to fine-tune for 100 iterations to obtain the optimal machine learning model;
[0068] (14) Performance testing, inputting the training set into the optimal machine learning model of step (13) for training, then testing the optimal machine learning model with the test set, and outputting an ROC curve;
[0069] (15) Result evaluation, evaluating the ROC curve and AUC value output by the model in step (14), evaluating the prediction performance of the optimal machine learning model, and storing the optimal machine learning model.
[0070] The prediction performance results of the optimal machine learning model evaluated by the present application are as follows:
[0071] Figure 2 The figure is the prediction result data graph of the machine learning model in the present application. In the figure, the horizontal axis of the ROC curve represents 1-Specificity, i.e. the false positive rate, ranging from 0 to 1, reflecting the proportion of the model incorrectly predicting negative examples as positive examples, and the vertical axis of the ROC curve represents Sensitivity, i.e. the true positive rate, measuring the ability of the model to correctly identify positive examples.
[0072] From Figure 2As can be seen, the blue area depicts the ROC curve of the model. In this invention, the ROC curve is close to the upper left corner, and there is a high true positive rate at each threshold. The area under the curve (AUC) is close to 1, indicating that the better the classification performance of the model, the better. The AUC value of this invention reaches 83.333%, indicating that the classifier performs well in distinguishing between positive and negative samples.
[0073] (2) Input the sample data to be tested into the optimal machine learning model trained in step (1) for calculation and prediction, and output the early screening prediction risk value of cirrhosis, specifically including:
[0074] The data of the sample to be tested is input into the trained optimal machine learning model for calculation and prediction to obtain the early screening prediction risk value of cirrhosis. The prediction risk value is between 0 and 1, where a risk value of 0 represents that the sample corresponds to a healthy control individual and a risk value of 1 represents that the sample corresponds to a patient with cirrhosis.
[0075] Figure 2 This is a data graph of the predicted probability results output by the machine learning algorithm in this embodiment. Figure 2 The probability distribution shows that when the predicted probability risk value is less than 0.4, it corresponds to the green healthy population distribution range, indicating a low to medium risk. When the predicted probability risk value is between 0.4 and 0.6, it corresponds to the gray transition area, indicating a certain risk of cirrhosis, and it is recommended to adjust the gut microbiota to reduce the risk of developing cirrhosis later. When the predicted probability risk value exceeds 0.6, it corresponds to the red disease population distribution range, indicating a high risk of cirrhosis. It is recommended to undergo examination and maintain healthy eating and lifestyle habits.
[0076] Therefore, the risk values for early screening of cirrhosis are determined as follows: (A) Risk value < 0.4 indicates a low-risk state, requiring no adjustment of the gut microbiota; (B) 0.4 ≤ risk value ≤ 0.6 indicates a moderate-risk state, requiring adjustment of the gut microbiota; (C) Risk value > 0.6 indicates a high-risk group for cirrhosis, and clinical diagnosis is recommended. This invention, through precise screening of suitable bacterial genera and the corresponding number of input variables, combined with an optimized machine learning model, ultimately achieves a prediction accuracy of 85.9%, laying the foundation for disease research.
[0077] In this embodiment, the abundance data corresponding to the microbial markers of liver cirrhosis are used as a dataset and input into the original machine learning model for optimization, parameter tuning, training, and testing. This significantly improves the accuracy and sensitivity of the liver cirrhosis prediction model. The model can ultimately output individual liver cirrhosis risk prediction labels, enabling early screening and risk stratification. It not only indicates clinical diagnostic needs for high-risk individuals but also provides guidance on gut microbiota adjustment for medium-risk individuals to reduce the risk of disease.
[0078] In summary, the liver cirrhosis microbial marker provided by the present application not only provides a new biomarker and method for early diagnosis and prevention of liver cirrhosis, but also clarifies the potential correlation mechanism between intestinal flora and liver cirrhosis, lays a foundation for formulating individualized intestinal flora adjustment strategies and reducing the risk of liver cirrhosis, and has important significance for reducing the disease burden caused by liver cirrhosis.
[0079] The preferred embodiments of the present application are described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and changes without creative labor based on the concept of the present application. Therefore, any technical solutions obtained by logical analysis, reasoning or limited experiments based on the prior art according to the concept of the present application shall be within the protection scope defined by the claims.
Claims
1. A method for screening microbial biomarkers for liver cirrhosis, characterized in that, Using genome-wide association study (GWAS) data and 16S gut microbiota abundance data, combined with differential microbiota analysis and two-sample Mendelian randomization, effective gut microbiota-related biomarkers were identified, specifically including the following steps: S1. Obtain 16S gut microbiota abundance data including liver cirrhosis and healthy samples, perform differential analysis between the disease group and the control group, and screen out differentially expressed bacteria; S2. Obtain the GWAS summary statistics for gut microbiota samples and liver cirrhosis samples; S3. Screen out SNPs that are significantly correlated with gut microbiota characteristics across the whole genome and are independent of them, and assign them to the corresponding microbiota characteristics; S4. Extract relevant data from the corresponding gut microbiota characteristics and GWAS data of fatty liver, organize and merge them, and use linkage disequilibrium fractional regression (LDSC) analysis to assess the genetic correlation between gut microbiota and liver cirrhosis, and obtain LDSC genetically related bacteria; S5. Using multi-sample Mendelian randomization (MR) analysis, combined with sensitivity analysis, we assessed the causal relationship between gut microbiota characteristics and cirrhosis, and finally obtained the gut microbiota associated with cirrhosis, namely MR-associated bacteria. S6. The intersection of differentially expressed bacteria with LDSC-related and MR-related bacteria is used as the microbial marker bacteria for liver cirrhosis.
2. The method for screening microbial biomarkers for liver cirrhosis according to claim 1, characterized in that, In step S1, the difference analysis between the disease group and the control group is specifically performed using the difference fold and Wilcoxon rank-sum test methods to calculate the logFC and p-value of different bacterial groups in the disease group and the control group, and bacteria with logFC > 0.5 and p-value < 0.05 are selected as differentially expressed bacteria.
3. The method for screening microbial biomarkers for liver cirrhosis according to claim 1, characterized in that, In step S3, the selection of instrumental variable SNPs needs to be closely related to the gut microbiota and unrelated to liver cirrhosis. Therefore, the clump algorithm in PLINK is used to screen for SNPs that are significantly associated with each microbiota characteristic and have p < 1 × 10⁻⁶ across the entire genome. -5 Independent SNPs were identified and assigned to the corresponding microbial community features according to their matching relationships, serving as their instrumental variables.
4. The method for screening microbial biomarkers for liver cirrhosis according to claim 1, characterized in that, In step S4, the relevant data mainly includes SNP loci, alleles, effect sizes, standard errors, and P-values. The LDSC analysis was used to assess the genetic association between gut microbiota and liver cirrhosis. Specifically, a linkage disequilibrium parameter r was set for the merged SNP loci. 2 With a threshold of 0.01 and a genetic distance of 10,000 kb, independent SNPs were screened to obtain bacterial communities that are genetically associated with cirrhosis.
5. The method for screening microbial markers for liver cirrhosis according to claim 1, characterized in that, In step S5, the sensitivity analysis includes level pleiotropy and heterogeneity. The level pleiotropy test uses Mendelian randomization to perform global and outlier tests on the selected independent SNPs. The global test uses the significance P-value to evaluate the overall level pleiotropy of all SNPs, while the outlier test evaluates the existence of outliers at a specific level of pleiotropy by calculating the significance P-value of each SNP's pleiotropy. When P ≤ 0.05, it indicates that all SNPs have pleiostatity as a whole. In this case, an outlier test is performed, and the outlier SNP with the smallest pleiostatity P value is removed. The remaining SNPs were then subjected to a global test again. This process was repeated until the global test showed that the pleiotropic effect P > 0.05 was no longer significant. At this point, the non-level pleiotropic SNPs were retained as instrumental variables for subsequent MR analysis. Heterogeneity among each SNP was assessed using IVW analysis and Cochrane's Q test.
6. The method for screening microbial biomarkers for liver cirrhosis according to claim 5, characterized in that, It also includes conducting a two-sample MR analysis: using liver cirrhosis as the outcome variable and gut microbiota as the exposure factor, the gut microbiota used for analysis includes 36 characteristic microbiota at the genus level; based on the pooled level data, the causal relationship between microbiota characteristics and liver cirrhosis is obtained, specifically by measuring the association between the SNP instrumental variables obtained after removing pleiotropic levels and liver cirrhosis, retaining only microbiota characteristics with 3 or more SNP instrumental variables, and using an inverse variance weighting method to assess the relationship between each microbiota characteristic and liver cirrhosis according to the final exposure factor and outcome variable file; The inverse variance weighting method is calculated by dividing the estimated value of the SNP-outcome variable by the estimated value of the SNP-exposure factor, and the criterion for determining the influence of microbial characteristics on cirrhosis is a two-tailed test with P ≤ 0.
05.
7. A microbial marker for liver cirrhosis, characterized in that, The method for screening microbial biomarkers for cirrhosis according to any one of claims 1-6, wherein the biomarker bacteria for cirrhosis include the following species: Actinomyces, Veillonella, Prevotella 7, Rikenellaceae RC9gut group, and Aggregatibacter.
8. The application of a liver cirrhosis microbial marker according to claim 7, characterized in that, Machine learning algorithms are used to output early screening risk values for cirrhosis, which can be used for early screening and prediction of cirrhosis, and guide individualized gut microbiota adjustment to reduce the risk of disease. The specific steps include: (1) Obtain the abundance data corresponding to the microbial marker bacteria of liver cirrhosis, divide the abundance data into training set and test set according to the set ratio, use genetically related bacteria and differential bacteria as input features, input into the original machine learning model, cross-validate and adjust parameters, determine the optimal model, train the optimal machine learning model with the training set, test the optimal model with the test set and output the ROC curve, and then evaluate the performance of the test set ROC curve output by the optimal machine learning model. (2) Input the new sample data to be tested into the optimal machine learning model trained in step (1) for calculation and prediction, and output the early screening prediction risk value of cirrhosis.
9. The application of a liver cirrhosis microbial marker according to claim 8, characterized in that, In step (1), the original machine learning model includes Logistic Regression, Support Vector Machine, Random Forest, and XGBoost. The specific steps for optimizing, training, testing, and evaluating the parameters of the original machine learning model are as follows: (11) Dataset processing: The abundance data of microbial marker bacteria for liver cirrhosis were randomly divided into a 75% training set and a 25% test set. (12) Construct the original machine learning classifier, using genetically associated bacteria and differentially related bacteria as input features, and input them sequentially into the original machine learning models of logistic regression, SVM, random forest and Xgboost; (13) Select the optimal machine learning model, use cross-validation to tune the parameters, select the parameters with the highest ROC-AUC scores, and determine the optimal machine learning model. After the initial model is determined, use hyperparameters to fine-tune it through 100 iterations to obtain the optimal machine learning model. (14) Performance testing: After training the training set into the optimal machine learning model in step (13), test the optimal machine learning model with the test set and output the ROC curve. (15) Result evaluation: Evaluate the ROC curve and AUC value output by the model in step (14) to evaluate the predictive performance of the optimal machine learning model.
10. The application of a liver cirrhosis microbial marker according to claim 9, characterized in that, Step (2) specifically includes inputting the new sample data to be tested into the trained optimal machine learning model for calculation and prediction to obtain the early screening prediction risk value for cirrhosis. The predicted risk value is between 0 and 1, where a risk value of 0 represents a healthy control individual and a risk value of 1 represents a patient with cirrhosis. The results of the early screening prediction risk values for cirrhosis are as follows: (A) Risk value < 0.4 indicates a low-risk state, and no adjustment of gut microbiota is required; (B) 0.4 ≤ risk value ≤ 0.6 indicates a moderate-risk state, and adjustment of gut microbiota is required; (C) Risk value > 0.6 indicates a high-risk group for cirrhosis, and clinical diagnosis is recommended.
Citation Information
Cited By
Chronic disease intestinal microbial marker screening and early screening model construction method
CN122157760A