Intestinal microbiota marker combination for diagnosis of liver cirrhosis and application thereof

Eight gut microbial biomarkers were screened through metagenomic sequencing and ensemble machine learning, and a high-precision non-invasive diagnostic model for liver cirrhosis was constructed. This solved the problem of insufficient sensitivity of existing diagnostic methods and enabled efficient and stable screening and monitoring of liver cirrhosis.

CN122104890APending Publication Date: 2026-05-29JINAN MICROECOLOGY & BIOMEDICINE PROVINCIAL LAB +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINAN MICROECOLOGY & BIOMEDICINE PROVINCIAL LAB
Filing Date
2026-02-14
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing non-invasive diagnostic methods for cirrhosis have insufficient sensitivity and limited specificity, and the existing biomarker combinations have poor stability, making it difficult to achieve large-scale population screening and dynamic disease monitoring, causing patients to miss the optimal intervention window for treatment.

Method used

Through a rigorous case-control study design, combined with metagenomic sequencing and integrated machine learning processes, a set of gut microbiota biomarkers with excellent diagnostic efficacy and strong stability was screened, including eight bacterial species such as Lactobacillus salivarius and Streptococcus constellations, to construct a high-precision non-invasive diagnostic model.

Benefits of technology

It achieves high-precision and stable non-invasive diagnosis of liver cirrhosis, with a feature simplification rate of up to 86.9% and a diagnostic performance retention rate of 102.2%. It has good clinical translation potential and provides a non-invasive screening tool suitable for hospitals and health check-up centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122104890A_ABST
    Figure CN122104890A_ABST
Patent Text Reader

Abstract

The application discloses an intestinal microbial marker combination for cirrhosis diagnosis and application thereof, which is composed of 8 specific intestinal bacterial species. The application accurately identifies the core combination from 181 differential characteristics by performing metagenomic sequencing on fecal samples of cirrhosis patients and healthy controls, and combining random forest, LASSO regression and recursive feature elimination algorithms. The random forest diagnosis model based on the 8 markers realizes an excellent performance of area under curve (AUC) 0.841, sensitivity 84.3% and specificity 83.7% on a completely independent test set, and the performance does not decrease compared with the full feature model in the case of reducing the number of characteristics by 86.9%. The marker combination provided by the application is highly simplified and has strong discrimination, and provides a clear and efficient microbiological basis and technical scheme for developing a cirrhosis non-invasive diagnosis kit, a gene chip and other products based on intestinal flora.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of biomedical detection and bioinformatics, specifically relating to a non-invasive disease diagnosis method based on gut microbiome characteristics. More specifically, it relates to the application of a group of gut microbial species, screened and validated through multi-step machine learning, as biomarkers in the preparation of non-invasive diagnostic kits for liver cirrhosis or the construction of related disease prediction models. Background Technology

[0002] Liver cirrhosis is a chronic, end-stage liver disease caused by various factors, including viruses, alcohol, and metabolic abnormalities. Its core pathological features are liver fibrosis, pseudolobule formation, and regenerative nodules. This disease is the most common precancerous lesion for hepatocellular carcinoma worldwide, with approximately 80%-90% of liver cancer cases secondary to cirrhosis, posing a significant public health burden. The clinical progression of cirrhosis is insidious; the compensated stage often lacks typical symptoms, but once it progresses to the decompensated stage, characterized by complications such as ascites and gastrointestinal bleeding, the patient's prognosis deteriorates rapidly, and the five-year survival rate decreases significantly.

[0003] While liver biopsy is considered the "gold standard" for the clinical diagnosis of cirrhosis, its invasive nature carries risks such as bleeding, infection, and sampling errors, and its low patient acceptance makes it unsuitable for large-scale population screening and dynamic disease monitoring. Existing non-invasive diagnostic methods, such as ultrasound elastography and serological markers (e.g., APRI, FIB-4), generally suffer from insufficient sensitivity, limited specificity, or strong operator dependence when identifying compensated cirrhosis and early liver fibrosis. This results in many patients missing the optimal intervention window at diagnosis, with treatment options passively limited to symptomatic support or costly, donor-scarce liver transplantation. Therefore, developing a novel biomarker system capable of non-invasive, accurate, and early identification of cirrhosis has become an urgent technological need to improve clinical diagnosis and treatment and patient prognosis.

[0004] In recent years, in-depth research into the "gut-liver axis" mechanism has revealed a close bidirectional link between the gut microbiota and liver health. Abundant evidence indicates a clear causal relationship between the development and progression of cirrhosis and the imbalance of gut microbial homeostasis. As the disease progresses from the compensated to the decompensated stage, the gut microbiota exhibits characteristic dysbiosis: the relative abundance of potentially pathogenic bacteria (such as Enterobacteriaceae and Enterococciceae) increases significantly, while the abundance of beneficial bacteria with mucosal barrier protection and anti-inflammatory effects (such as Trichophytonceae and Ruminococciceae) decreases accordingly. This dysbiosis is not only a biomarker reflecting the disease state, but it can also directly drive the occurrence and development of serious complications such as hepatic encephalopathy by promoting bacterial translocation and exacerbating systemic inflammatory responses. Therefore, the gut microbiome is considered a highly promising biomarker source for developing non-invasive diagnostic tools for cirrhosis.

[0005] Metagenomic sequencing technology enables high-throughput sequencing of all microbial DNA in intestinal samples such as feces, achieving precise quantification at the species and even strain level, providing crucial technical support for systematically analyzing the characteristics of gut microbiota associated with cirrhosis. Combined with supervised machine learning algorithms, key biological features can be efficiently mined from massive and complex metagenomic data, and high-precision predictive models can be constructed. However, most existing related studies are still limited to descriptive analyses of differences in microbiota composition. The reported biomarker combinations often have a large number of features and high redundancy, and generally lack systematic screening and efficacy confirmation through rigorous machine learning processes to prevent data leakage (such as using completely independent test sets for validation). This results in poor stability and insufficient clinical generalization ability of many reported potential biomarker combinations, making it difficult to directly translate them into stable and reliable in vitro diagnostic products. Summary of the Invention

[0006] The purpose of this invention is to overcome the aforementioned deficiencies of existing technologies. Through a rigorous case-control study design, metagenomic sequencing technology is applied to systematically analyze the gut microbiota characteristics of cirrhosis patients and healthy controls. An integrated advanced machine learning workflow is used for multiple rounds of biomarker screening, model construction, and rigorous multi-dimensional validation. The ultimate goal of this invention is to obtain a set of bacterial biomarkers with excellent diagnostic efficacy, high stability, and high simplicity. This not only provides a fully validated microbiological basis for the auxiliary diagnosis of cirrhosis but also provides a solid biomarker system and methodological foundation for the development of non-invasive detection technologies based on gut microbiota suitable for clinical application.

[0007] The technical solution adopted in this invention includes:

[0008] In a first aspect, the present invention provides the application of a combination of gut microbial biomarkers for the diagnosis of cirrhosis in the preparation of products for the auxiliary assessment of the risk of cirrhosis.

[0009] Preferably, the intestinal microbial marker combination is a bacterial species or a specific combination of nucleic acid sequences with diagnostic discrimination. The intestinal microbial marker combination for the diagnosis of cirrhosis consists of the following eight bacterial species: *Ligilactobacillus salivarius*, *Streptococcus constellatus*, *Lactobacillus sp. JM1*, *Enterobacter cloacae complex_sp.*, *Streptococcus anginosus*, *Veillonellarodentium*, *Lactobacillus crispatus*, and *Enterobacter chengduensis*.

[0010] Preferably, the product includes a kit, chip, or sequencing detection reagent, which contains reagents capable of specifically identifying the presence or relative abundance of various gut microbial markers.

[0011] Preferably, the reagent is a primer pair or probe capable of specifically amplifying or hybridizing to capture the specific nucleic acid sequences of each intestinal microbial marker.

[0012] Preferably, the product uses fecal samples as test samples.

[0013] Preferably, the chip includes a solid-phase carrier and probes or primers attached thereto that specifically recognize the nucleic acid sequences of each gut microbial biomarker in the gut microbial biomarker combination.

[0014] Secondly, the present invention provides a kit for assisting in the assessment of the risk of cirrhosis, which can be used for non-invasive auxiliary diagnosis of cirrhosis, the kit comprising reagents for detecting the combination of gut microbiota markers.

[0015] Preferably, the reagent is a primer pair or probe capable of specifically amplifying or hybridizing to capture the specific nucleic acid sequences of each intestinal microbial marker.

[0016] Preferably, the kit is suitable for detecting ex vivo fecal samples.

[0017] Thirdly, the present invention provides a gene chip comprising a solid-phase carrier and an oligonucleotide probe immobilized on the solid-phase carrier, wherein the probe is capable of specifically hybridizing to capture specific nucleic acid sequences of various intestinal microbial markers.

[0018] Fourthly, this invention provides an in vitro analytical method for assessing liver cirrhosis risk based on gut microbiota markers, comprising the following steps:

[0019] (a) Obtain ex vivo fecal samples from the individuals to be tested and extract total microbial DNA;

[0020] (b) Based on the extracted total microbial DNA, the presence or relative abundance of each intestinal microbial marker in the sample is detected to obtain abundance data; the intestinal microbial markers include: Ligilactobacillus salivarius, Streptococcus constellatus, Lactobacillus sp. JM1, Enterobacter cloacae complex_sp., Streptococcus anginosus, Veillonella rodentium, Lactobacillus crispatus, and Enterobacter chengduensis;

[0021] (c) Input the abundance data into a trained random forest classification model to obtain a risk score for auxiliary assessment;

[0022] The random forest classification model was trained based on the abundance differences of the eight bacterial species in the cirrhosis group and the healthy control group.

[0023] Preferably, in step (b), the relative abundance of the gut microbial markers is detected by quantitative polymerase chain reaction (qPCR) or high-throughput sequencing.

[0024] Preferably, the method for constructing the random forest classification model includes: acquiring a dataset and dividing the dataset into a training set, a validation set, and a test set; modeling on the training set; optimizing hyperparameters on the validation set; and evaluating model performance on the test set; wherein the division is completed before feature selection, and the test set does not participate in any feature selection or model training steps throughout the process.

[0025] Preferably, the area under the curve obtained on the test set is not less than 0.83.

[0026] As a preferred option, if the risk score is greater than 0.653, the individual is considered to have suspected or confirmed cirrhosis; otherwise, they are considered not to have cirrhosis.

[0027] Fifthly, this invention provides a method for screening a combination of gut microbial biomarkers for the diagnosis of cirrhosis. This method is achieved through a combination of fecal metagenomic sequencing, bioinformatics analysis, and machine learning algorithm screening, and specifically includes the following steps:

[0028] (1) Establishment of the research cohort and collection of stool samples

[0029] Fecal samples were collected from 60 patients with liver cirrhosis diagnosed by liver biopsy and 40 healthy volunteers. All participants had not used antibiotics, probiotics, or proton pump inhibitors within 4 weeks prior to sampling. Fecal samples were stored at -80°C for 2 hours after collection until DNA extraction.

[0030] (2) Fecal DNA extraction and metagenomic high-throughput sequencing

[0031] Total DNA was extracted from all fecal samples using the QIAGEN QIAamp PowerFecal Pro DNA Kit, and DNA concentration and integrity were quality controlled by a Qubit fluorometer and agarose gel electrophoresis. Samples that passed quality control were subjected to whole-genome shotgun sequencing on the Illumina NovaSeq 6000 platform, constructing paired-end libraries with 350 bp inserts. Each sample yielded an average of 10 Gb of raw sequencing data (150 bp paired-end reads).

[0032] (3) Bioinformatics processing and microbial feature extraction

[0033] The raw sequencing data were quality-assessed using Fastp software, and low-quality sequences and adapters were filtered out. The KneadData pipeline (based on the Bowtie2 alignment tool) was used to align the quality-controlled sequences with the human host reference genome, effectively removing host genome contamination and obtaining high-quality non-host microbial sequences. Furthermore, the microbial community composition was accurately identified using a species annotation pipeline combining MetaPhlAn4 and Kraken2. Based on this, differential analysis was performed using MaAsLin2: the original feature abundance table was filtered, retaining microbial features present in at least 50% of the samples; data preprocessing was performed using summative standardization (TSS) combined with logarithmic transformation (LOG); statistical tests were conducted using a linear model with experimental group as a fixed effect, and multiple test corrections were performed using the Benjamini-Hochberg method (significance threshold set at 0.01), while limiting the absolute value of the abundance difference (|coef|) to ≥2. Finally, 181 microbial features with significant abundance differences among different experimental groups were identified, forming the initial feature matrix for subsequent analyses.

[0034] (4) Feature integration screening and optimization based on multiple machine learning methods

[0035] In the screening and optimization stage of diagnostic biomarkers based on machine learning, this study adopted an integrated feature selection strategy to further refine 181 significantly different microbial features. The specific process is as follows: First, three machine learning methods were applied independently for screening: Random Forest selected high-contribution features based on the feature importance index (MeanDecreaseAccuracy > 2); LASSO regression selected 20 key features through a regularization path; and the recursive feature elimination method gradually eliminated redundant features, ultimately retaining 45 core microbial features. To obtain the most robust and discriminative feature set, this study took the intersection of the results obtained from the above three methods, and finally determined 61 key microbial features for subsequent modeling and validation analysis.

[0036] (5) Machine learning modeling and validation based on random forest algorithm

[0037] A random forest algorithm was used to construct a diagnostic prediction model for liver cirrhosis. Using 61 key microbial features identified in the initial screening as input, all 100 samples were randomly divided into a training set (60 cases), a validation set (20 cases), and a test set (20 cases) to strictly prevent data leakage. A 20-times repeated modeling strategy was employed, with hyperparameters optimized based on the validation set, and the model performance evaluated on the test set. Through importance ranking and further screening of the 61 features, a biomarker combination consisting of 8 core microbial features was finally determined, achieving a feature simplification rate of 86.9%. The combination includes: *Ligilactobacillus salivarius*, *Streptococcus constellatus*, *Lactobacillus sp. JM1*, *Enterobacter cloacae complex_sp.*, *Streptococcus anginosus*, *Veillonella rodentium*, *Lactobacillus crispatus*, and *Enterobacter chengduensis*. The random forest model constructed based on this simplified combination demonstrated excellent diagnostic efficacy on the test set: the mean area under the curve (AUC) was 0.841 ± 0.009, sensitivity was 84.3%, and specificity was 83.7%. It is worth emphasizing that, compared with the model using all 61 features, the model with 8 core biomarkers not only did not decline in the number of features, but actually slightly improved (performance retention rate of 102.2%), which fully demonstrates that the combination of biomarkers has significant redundancy removal capability and high efficiency while maintaining high diagnostic performance.

[0038] In a sixth aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned in vitro analysis method for assessing liver cirrhosis risk based on gut microbiota markers.

[0039] In a seventh aspect, the present invention also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, implement the aforementioned in vitro analysis method for assessing liver cirrhosis risk based on gut microbiota markers.

[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0041] 1. A highly efficient biomarker combination with rigorous multi-validation is provided: This invention discloses and validates for the first time a concise biomarker combination containing only 8 specific gut microbial species. This combination is cross-selected using three machine learning algorithms: random forest, LASSO regression, and recursive feature elimination, and its final efficacy is confirmed using a completely independent hold-out test set. Excellent diagnostic performance of 0.841, 84.3% sensitivity, and 83.7% specificity is achieved on the independent test set, demonstrating its high discriminative power and reliability.

[0042] 2. Achieved extreme simplification and efficient redundancy removal of biomarkers: Through the integrated screening strategy, 8 core biomarkers were ultimately extracted from 181 differential features, achieving a feature simplification rate of up to 86.9%. Crucially, the diagnostic model built based on this minimal combination did not experience a performance decrease compared to the model using all 61 key features, with a performance retention rate of 102.2%. This strongly demonstrates that the combination has extracted the most essential and non-redundant diagnostic information, ensuring high accuracy while simplifying detection complexity, resulting in significant advantages for clinical translation.

[0043] 3. Rigorous methodological design fundamentally eliminates overfitting: This invention adheres to a strict separation of "training-validation-testing" throughout the entire process, and all feature selection and model optimization steps are completed within the training and validation sets, while the test set remains a "black box" throughout. This design fundamentally avoids data leakage, ensuring the authenticity and objectivity of the reported performance metrics and the model's good generalization ability.

[0044] 4. A clear and feasible clinical commercialization path: The eight core biomarkers are all well-defined bacterial species, and based on their specific nucleic acid sequences, they can be conveniently and stably developed into standardized in vitro diagnostic products such as multiplex qPCR detection kits, gene chips, or targeted sequencing panels. This provides a clear and efficient technical solution for transforming research results into non-invasive screening tools for liver cirrhosis suitable for hospital laboratories and health check-up centers. Attached Figure Description

[0045] Figure 1 A graph showing the results of principal coordinate analysis of microbial communities based on Bray-Curtis distance, revealing significant differences between disease and healthy groups;

[0046] Figure 2 This is a graph showing the results of differential microbial species analysis related to health and disease status based on MaAsLin2 regression.

[0047] Figure 3The results of screening key microbial biomarkers for diseases based on multiple algorithms are shown in the following diagrams: (A) is the result of ranking the importance of random forest features, (B) is the result of statistical analysis of key features and coefficient distribution of LASSO regression, (C) is the result of feature selection frequency of recursive feature elimination, and (D) is the result of 61 core microbial features determined by the intersection of multiple algorithms.

[0048] Figure 4 Box plot comparing the relative abundance distribution of eight core biomarkers in the disease (cirrhosis) group and the healthy (control) group.

[0049] Figure 5 A comparison of receiver operating characteristic (ROC) curves of a random forest diagnostic model based on eight core biomarkers on the test and validation sets;

[0050] Figure 6 The chart shows the diagnostic performance comparison between the full-feature model and the simplified model on the validation and test sets. Detailed Implementation

[0051] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0052] First, it should be noted that the fecal samples used in this invention were obtained from a clinical research cohort approved by the Ethics Committee of the First Affiliated Hospital of Zhejiang University School of Medicine. The collection, processing, and use of all samples strictly adhered to relevant ethical guidelines, and written informed consent was obtained from all participants. The implementation of this invention does not involve any new invasive procedures on the human body.

[0053] The technical solution of the present invention will be described in detail below with reference to specific embodiments.

[0054] Example 1:

[0055] A method for screening a combination of gut microbial biomarkers for the diagnosis of cirrhosis, comprising:

[0056] Step 1: Establishing the research cohort and processing samples

[0057] Fecal total DNA data samples were collected from patients with liver cirrhosis diagnosed by liver biopsy and healthy volunteers; the samples were then divided into training, validation, and test sets according to a set ratio. In this embodiment:

[0058] The study employed a case-control design. The disease group (cirrhosis group) included 60 patients diagnosed with cirrhosis via liver biopsy, all meeting the following diagnostic criteria: (1) typical pathological features of cirrhosis; (2) a clear etiology of chronic liver disease (including but not limited to alcoholic liver disease, non-alcoholic fatty liver disease, viral hepatitis, or autoimmune liver disease). The control group included 40 age- and sex-matched healthy volunteers. Exclusion criteria for all participants included: (1) use of antibiotics, probiotics, or proton pump inhibitors within 4 weeks prior to sample collection; (2) gastrointestinal surgery or colonoscopy within 6 months prior to sampling; and (3) uncontrolled infection or systemic comorbidities. All stool samples were collected immediately after defecation and stored at -80°C for 2 hours until genomic DNA extraction. Total DNA was extracted from all stool samples using the QIAamp PowerFecal Pro DNA Kit from Qiagen. In this embodiment, all 100 samples are randomly divided into a training set (60 cases), a validation set (20 cases), and a test set (20 cases).

[0059] Step 2: Metagenomic Sequencing and Bioinformatics Analysis

[0060] The extracted DNA was quality controlled by a Qubit fluorometer and agarose gel electrophoresis. Qualified samples were sequenced using shotgun sequencing on the Illumina NovaSeq 6000 platform to construct paired-end libraries with 350bp inserts. Each sample yielded an average of 10Gb of raw data (150bp paired-end reads).

[0061] Step 3: Sequence Data Processing and Species Annotation

[0062] The raw sequencing data were quality-assessed using Fastp software, and low-quality sequences and adapter sequences were filtered out simultaneously. Subsequently, the KneadData analysis pipeline (based on the Bowtie2 alignment tool) was used to align the quality-controlled sequences with the human host reference genome, effectively removing host genome contamination and obtaining high-quality non-host microbial sequence data. Based on the obtained purified microbial sequences, an integrated analysis pipeline combining MetaPhlAn4 and Kraken2 was further used to identify the microbial species composition, ultimately generating a relative abundance spectrum for each sample, which served as the foundational data matrix for all subsequent analyses.

[0063] Step 4: Overall Difference Analysis of Microbial Community Structure

[0064] Based on the relative abundance spectra of species obtained above, the overall structural differences in the gut microbiota between the cirrhosis group and the healthy control group were first assessed. β-diversity analysis showed that, based on the principal coordinate analysis plot using Bray-Curtis distance, the two groups exhibited a clear separation trend. Further statistical tests (permutation multivariate ANOVA, number of permutations = 999) confirmed a significant difference in the overall gut microbiota structure between the two groups (R² = 6.68%, P < 0.001). Figure 1 This result establishes the objective fact of gut microbiota dysbiosis in patients with cirrhosis at a holistic level, providing an important premise and basis for subsequent screening of specific microbial biomarkers.

[0065] Step 5: Screening for Differential Microbial Characteristics

[0066] To screen for microbial biomarkers associated with cirrhosis, differential analysis was first performed using MaAsLin2. Abundance tables were filtered to retain features present in at least 50% of samples. Preprocessing was performed using summative method standardization (TSS) combined with logarithmic transformation (LOG). A linear model with group as a fixed effect was established, and multiple tests were performed using the Benjamini-Hochberg method (p < 0.01), while limiting the absolute value of the abundance difference (|coef|) to ≥ 2. Ultimately, 181 microbial biomarkers with significant abundance differences between the cirrhosis group and the healthy control group were identified, such as… Figure 2 As shown.

[0067] Step Six: Core biomarker selection and model building based on machine learning, including the following sub-steps:

[0068] (6.1) Integrated Feature Selection

[0069] To obtain a highly discriminative and stable biomarker combination, an ensemble feature selection strategy was employed for the aforementioned 181 differential features. Three algorithms—random forest (selection criterion: Mean Decrease Accuracy > 2), LASSO regression, and recursive feature elimination—were used for independent screening. To maximize the robustness of the biomarker combination, the intersection of the feature sets obtained by the three methods was taken, ultimately yielding 61 key microbial features for subsequent analysis. Figure 3 ).

[0070] (6.2) Diagnostic model construction and validation

[0071] It is worth noting that, in order to evaluate the model's generalization ability and strictly avoid data leakage, all 100 samples were randomly divided into a training set (60 samples), a validation set (20 samples), and a test set (20 samples). This division was completed before feature selection, and the test set did not participate in any feature selection or model training steps throughout the process.

[0072] A random forest algorithm was selected, and a diagnostic model was built on the training set using 61 key features as input. The model was trained by minimizing the error between the output and the ground truth as the loss function, resulting in a well-trained random forest classification model. The core hyperparameters were then optimized using a grid search on the validation set. To evaluate the stability of the model's performance, the above training and validation processes were repeated independently 20 times.

[0073] Based on the average feature importance of 20 modeling results, the eight core microbial biomarkers with the highest contribution were further identified from the 61 features, including: *Ligilactobacillus salivarius*, *Streptococcus constellatus*, *Lactobacillus sp. JM1*, *Enterobacter cloacae complex_sp.*, *Streptococcus anginosus*, *Veillonella rodentium*, *Lactobacillus crispatus*, and *Enterobacter chengduensis* (for comparison of their abundance differences, see [link to relevant documentation]). Figure 4 Subsequently, the random forest classification model was retrained using only these 8 markers to obtain the final trained model. The model outputs a risk score indicating whether the user has cirrhosis after inputting the features. If the risk score is greater than 0.653, the user is judged as suspected or diagnosed with cirrhosis; otherwise, they are judged as not having cirrhosis.

[0074] The final simplified random forest classification model demonstrated good discrimination ability on both the test and validation sets (its ROC curve is shown in [reference]). Figure 5 Finally, performance was evaluated on the hold-out test set, which did not participate in any modeling steps. The model demonstrated excellent diagnostic efficacy: the area under the receiver operating characteristic (AUC) was 0.844 (its ROC curve is shown in [reference needed]). Figure 5 The sensitivity was 84.3%, and the specificity was 83.7%. Notably, compared to a model built using all 61 features, this 8-marker model, with an 86.9% reduction in the number of features, achieved improved diagnostic performance on the test set (performance comparison see...). Figure 6This fully demonstrates the high efficiency and redundancy removal capability of the marker combination.

[0075] In summary, this embodiment demonstrates that the combination of eight gut microbial biomarkers obtained by the method described in this invention can be effectively used to construct a non-invasive diagnostic model for liver cirrhosis, exhibiting high accuracy, high stability, and promising clinical translation prospects.

[0076] Example 2: Products for assisting in the assessment of cirrhosis risk

[0077] This embodiment provides a kit for assisting in the assessment of liver cirrhosis risk. The kit contains reagents capable of specifically identifying the presence or relative abundance of various gut microbial markers, including: *Ligilactobacillus salivarius*, *Streptococcus constellatus*, *Lactobacillus sp. JM1*, *Enterobacter cloacae complex_sp.*, *Streptococcus anginosus*, *Veillonella rodentium*, *Lactobacillus crispatus*, and *Enterobacter chengduensis*.

[0078] The reagent is a primer pair or probe capable of specifically amplifying or hybridizing to capture specific nucleic acid sequences of various intestinal microbial markers, wherein the sequence of the primer pair is selected from at least one combination of the following:

[0079] (1) Primer pair for detecting *Ligilactobacillus salivarius*:

[0080] Forward primer: 5'-GCAATAAGCATTCCGCCTGG-3' (SEQ ID NO: 1)

[0081] Reverse primer: 5'-GCTGGCAACTGACAACAAGG-3' (SEQ ID NO: 2)

[0082] (2) Primer pair used for detecting Streptococcus constellatus:

[0083] Forward primer: 5'-GTGAGGTAACGGCTCACCAA-3' (SEQ ID NO: 3)

[0084] Reverse primer: 5'-GTGTCTCAGTCCCAGTGTGG-3' (SEQ ID NO: 4)

[0085] (3) Primer pair for detecting Enterobacter cloacae complex sp.:

[0086] Forward primer: 5'-TATCCTTTGTTGCCAGCGGT-3' (SEQ ID NO: 5)

[0087] Reverse primer: 5'-CGCTTCTCTTTGTATGCGCC-3' (SEQ ID NO: 6)

[0088] (4) Primer pair for detecting Streptococcus anginosus:

[0089] Forward primer: 5'-GAGTGCTAGGTGTTGGGTCC-3' (SEQ ID NO: 7)

[0090] Reverse primer: 5'-ACCACCTGTCACCGATGTTC-3' (SEQ ID NO: 8)

[0091] (5) Primer pair for detecting Veillonella rodentium:

[0092] Forward primer: 5'-CAAGCGGTGGAGTATGTGGT-3' (SEQ ID NO: 9)

[0093] Reverse primer: 5'-CGTGCACCACCTGTTTTCTG-3' (SEQ ID NO: 10)

[0094] (6) Primer pair for detecting Lactobacillus crispatus:

[0095] Forward primer: 5'-ACGCGAAGAACCTTACCAGG-3' (SEQ ID NO: 11)

[0096] Reverse primer: 5'-CCCAACATCTCACGACACGA-3' (SEQ ID NO: 12)

[0097] (7) Primer pair for detecting Enterobacter chengduensis:

[0098] Forward primer: 5'-TGCCTGATGGAGGGGGGATAA-3' (SEQ ID NO: 13)

[0099] Reverse primer: 5'-TTCCAGTGTGGCTGGTCATC-3' (SEQ ID NO: 14)

[0100] (8) Primer pairs for detecting Lactobacillus sp. JM1 were designed based on the whole genome sequence of the strain (accession number: NZ_CP044412.1), and their sequence is included within the scope of protection of this invention.

[0101] It should be noted that the above primer pairs are one example, and all primer pairs designed based on the whole genome sequences of the eight strains are included within the scope of protection of this invention.

[0102] 2.1 Kit Composition and Usage

[0103] The primer pairs are capable of specifically amplifying DNA fragments of corresponding microbial markers via polymerase chain reaction (PCR) or quantitative real-time PCR (qPCR). The kit also contains other necessary components for nucleic acid detection, including but not limited to DNA polymerase, dNTPs, reaction buffer, positive controls, and negative controls.

[0104] 2.2 Gene Chips

[0105] This embodiment also provides a gene chip, including a solid-phase carrier and oligonucleotide probes immobilized on the solid-phase carrier, wherein the probes are capable of specifically hybridizing to capture the specific nucleic acid sequences of the above-mentioned intestinal microbial markers.

[0106] 2.3 Applications

[0107] This invention provides the use of the kit or gene chip in the preparation of diagnostic products for assisting in the diagnosis of cirrhosis or assessing the risk of cirrhosis.

[0108] Example 3: An in vitro analytical method for assessing liver cirrhosis risk based on gut microbiota markers

[0109] This embodiment provides an in vitro analytical method for assessing liver cirrhosis risk based on gut microbiota markers, including the following steps:

[0110] (a) Obtain ex vivo fecal samples from the individuals to be tested and extract total microbial DNA;

[0111] (b) Based on the extracted total microbial DNA, the presence or relative abundance of each gut microbial biomarker in the sample is detected to obtain abundance data; wherein, the relative abundance of the gut microbial biomarkers is detected by quantitative polymerase chain reaction (qPCR) or high-throughput sequencing.

[0112] (c) Input the abundance data into a trained random forest classification model to obtain a risk score for auxiliary assessment;

[0113] The random forest classification model described above was trained using the methods described in the embodiments above.

[0114] Corresponding to the aforementioned embodiment of an in vitro analytical method for assessing the risk of liver cirrhosis based on gut microbiota markers, the present invention also provides an embodiment of an electronic device for implementing the aforementioned in vitro analytical method for assessing the risk of liver cirrhosis based on gut microbiota markers.

[0115] An electronic device provided in this invention includes one or more processors for implementing the aforementioned in vitro analysis method for assessing liver cirrhosis risk based on gut microbiota markers as described in the above embodiments.

[0116] The embodiments of the electronic device of the present invention can be applied to any device with data processing capabilities, such as a computer or other equipment or apparatus.

[0117] The device embodiment can be implemented through software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device in which it resides reading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, in addition to the processor, memory, network interface, and non-volatile memory, the data processing device in which the device in the embodiment resides may also include other hardware depending on the actual function of that data processing device, which will not be elaborated further.

[0118] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0119] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0120] This invention also provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements the aforementioned in vitro analysis method for assessing liver cirrhosis risk based on gut microbiota markers as described in the above embodiments.

[0121] The computer-readable storage medium can be an internal storage unit of any data processing device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units of any data processing device and external storage devices. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.

[0122] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. The application of a combination of gut microbiota biomarkers for the diagnosis of cirrhosis in the preparation of products to aid in the assessment of cirrhosis risk, characterized in that, The biomarker combination consists of the following eight bacteria: *Ligilactobacillus salivarius*, *Streptococcus constellatus*, *Lactobacillus sp. JM1*, *Enterobacter cloacae complex sp.*, *Streptococcus anginosus*, *Veillonella rodentium*, *Lactobacillus crispatus*, and *Enterobacter chengduensis*; the product contains reagents for the specific detection of the biomarkers.

2. The application according to claim 1, characterized in that, The product is a kit or gene chip, and the reagent is a primer pair or probe that specifically amplifies or hybridizes to capture the nucleic acid sequence of the marker.

3. A kit for assisting in the assessment of liver cirrhosis risk, characterized in that, Primer pairs or probes that specifically detect the eight intestinal microbial markers described in claim 1.

4. A gene chip, characterized in that, Oligonucleotide probes capable of specifically hybridizing and capturing the nucleic acid sequences of the eight intestinal microbial markers described in claim 1 are immobilized on a solid-phase support.

5. An in vitro analytical method for assessing liver cirrhosis risk based on gut microbiota markers, characterized in that, include: (a) Extracting total microbial DNA from ex vivo fecal samples of the individual being tested; (b) Detect the relative abundance of the eight gut microbial biomarkers described in claim 1 to obtain abundance data; (c) Input the abundance data into the trained random forest classification model and output a risk score; The random forest classification model is trained based on the abundance differences of the eight biomarkers in the cirrhosis group and the healthy control group, and the area under the curve (AUC) on the independent test set is not less than 0.

83.

6. The method according to claim 5, characterized in that, Step (b) involves detecting the relative abundance of the biomarker by quantitative polymerase chain reaction or high-throughput sequencing.

7. A method for screening a combination of gut microbiota biomarkers for the diagnosis of cirrhosis, characterized in that, include: (1) Obtain metagenomic sequencing data of fecal samples from patients with cirrhosis and healthy controls, and obtain a microbial characteristic abundance table after quality control, host removal, and species annotation; (2) MaAsLin2 was used for differential analysis. The threshold was |coef|≥2 and the corrected P value<0.

05. Differential abundance features were screened and an initial feature matrix was constructed. (3) The initial feature matrix is ​​independently filtered by three machine learning methods: random forest, LASSO regression and recursive feature elimination. The intersection of the three methods is used to obtain the second feature matrix. (4) Based on the training set, a random forest classification model is trained with the second feature matrix as input, and the hyperparameters are optimized through the validation set; (5) After 20 repeated modeling, based on the average feature importance ranking, the core marker combination consisting of the following 8 bacteria was finally screened: Ligilactobacillus salivarius, Streptococcus constellatus, Lactobacillus sp. JM1, Enterobacter cloacae complex sp., Streptococcus anginosus, Veillonella rodentium, Lactobacillus crispatus, and Enterobacter chengduensis.