Application of a chronic kidney disease microbial marker and risk prediction model and its construction method

By combining the mixed model of random forest and XGBoost algorithm, using specific microbial marker abundance data, a chronic kidney disease risk prediction model is constructed, which solves the accuracy of early diagnosis of chronic kidney disease, and realizes efficient and non-invasive early screening and auxiliary diagnosis, guides the adjustment of intestinal flora, and improves the therapeutic effect of chronic kidney disease.

CN116543899BActive Publication Date: 2025-08-22KANGMEIHUA GENE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310305184.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-24
Publication Date
2025-08-22
Estimated Expiration
2043-03-24

AI Technical Summary

Technical Problem

In the prior art, early diagnosis methods for chronic kidney disease lack sensitive indicators, resulting in a high missed diagnosis rate. The existing machine learning models have low accuracy in the prediction of intestinal flora and are limited in applicability, which cannot effectively assist in early screening and early warning of chronic kidney disease.

Method used

A hybrid machine learning model combined with random forest and XGBoost algorithm was used to construct a chronic kidney disease risk prediction model using abundance data of specific microbial markers such as Pratella, Fertilizer, Egertella and Broutella to help diagnose and predict chronic kidney disease risk by detecting the abundance of these markers in fecal samples.

Benefits of technology

It has achieved high sensitivity and high specificity of chronic kidney disease risk prediction, can early screening and assisted diagnosis, guide intestinal microbiota adjustment, improve treatment effect, and is non-invasive and economical in operation, and is suitable for non-invasive testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116543899B_ABST
    Figure CN116543899B_ABST
Patent Text Reader

Abstract

An application of a chronic kidney disease microbial marker and a risk prediction model and a construction method thereof, relating to the field of microbial technology; an application of a preparation for detecting microbial markers in the preparation of a product for detecting chronic kidney disease, wherein the microbial markers include Prevotella, Roseburia, Eggerthella tarda and Blautia. The method for constructing a chronic kidney disease risk prediction model comprises the following steps: S1, obtaining the abundance of the microbial markers in stool samples of healthy individuals and patients with chronic kidney disease, respectively, to construct a sample set; S2, inputting the sample set into a machine learning model, training and testing the model, and storing it to obtain a chronic kidney disease risk prediction model. The present invention can predict the positive probability of chronic kidney disease by detecting the abundance of microbial markers, and has high prediction accuracy and good sensitivity. It can be used as an auxiliary diagnostic method for chronic kidney disease and guide the direction of improving the intestinal flora environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of microbial technology, and specifically relates to an application of chronic kidney disease microbial markers and a risk prediction model and a construction method thereof. Background Art

[0002] The development of a new society has brought about changes in people's lifestyles. The fast-paced lifestyle has made it easier for people to neglect their health, reducing their vigilance for some "hidden" chronic diseases, such as chronic kidney disease (CKD). Currently, the incidence of CKD is rapidly increasing, and its mortality rate increases with declining renal function. It has become a major public health issue seriously affecting the health of the Chinese population, urgently requiring effective and accessible preventive and treatment methods and technologies. CKD is often difficult to detect early, and even missed, for three reasons: first, CKD can be completely asymptomatic or have subtle symptoms; second, people's awareness of preventive measures is low, and some doctors lack sufficient experience; third, current methods for assessing renal function have limitations and lack sensitive early markers, preventing the early diagnosis of chronic kidney disease. Therefore, new CKD diagnostic or screening methods are urgently needed to improve the early diagnosis rate and prognosis of this population.

[0003] In recent years, extensive research has revealed that patients with chronic renal failure (CKD) are prone to gastrointestinal dysfunction and intestinal microbiome disturbances. On the one hand, patients with CKD often experience decreased intestinal motility, which can lead to intestinal retention of nutrients such as protein and amino acids. On the other hand, the number, structure, and function of the intestinal microbiota in the colon of patients with CKD are significantly altered. This is manifested by a decrease in probiotics, which fail to produce mucosal barrier-protecting factors. For example, butyrate-producing bacteria (such as Faecalibacterium prausnitzii and Eubacterium) decrease, while pathogenic bacteria increase, producing a variety of mucosal barrier-damaging and inflammatory factors. When the CKD gut microbiota is disturbed, its metabolites (uremic toxins) are a key factor in disease progression. Academician Ren Fazheng of China Agricultural University and his team have discovered and validated the influence of the gut microbiota on toxin accumulation and renal disease phenotypes, revealing a relationship between gut microbial imbalance and metabolic disorders and clinical nephropathy. Given that the intestinal flora of CKD patients varies across different stages (stages 1-5), changes in intestinal flora abundance can serve as important indicators for determining disease status. These indicators may serve as microbial markers for chronic kidney disease, enabling the detection of microbial markers as an auxiliary diagnosis to improve early diagnosis and prognosis.

[0004] With the increase in research related to the intestinal flora, a massive amount of microbial sequencing data has been generated. As a major branch of machine learning, the random forest algorithm is often used to build disease classification models and screen core markers. For example, a study published in Advanced Science in 2020 used 5 OTUs obtained by random forest as non-invasive markers. The model distinguished patients from controls very well, but this study did not go into specific species. In fact, the sequences of 16S fragments at the species level are very similar, and there may be cases where it is difficult to distinguish them. Although the genus classification and above are more accurate, the inaccurate annotation makes most markers suitable for 16S sequencing and not very versatile in metagenomics or other methods of quantifying fungal abundance. At the same time, due to the independence and regionality between research cohorts within the field, some current disease studies have problems such as low accuracy of prediction methods, single model algorithms, and limited model applicability. In addition, CN109943636A discloses a microbial marker for colorectal cancer, and its results demonstrate the superiority of the machine learning Xgboost algorithm in disease prediction models based on microbial abundance; CN113736896A also applied the Xgboost algorithm to obtain a prediction model for hereditary angioedema, and its results also demonstrated the applicability and feasibility of this method in microbial abundance data. However, in chronic kidney disease, there is currently no Xgboost model based on species abundance information. In summary, it is of great significance to construct a hybrid machine learning model with good specificity and high sensitivity based on the combination of random forest and XGBoost algorithms, which can indicate the balance of intestinal bacterial content and guide the regulation of intestinal flora microbial markers for chronic kidney disease. Summary of the Invention

[0005] In order to overcome the shortcomings of the existing technology, one of the objectives of the present invention is to provide an application of a preparation for detecting microbial markers in the preparation of a product for detecting chronic kidney disease. By detecting the abundance of microbial markers, the positive probability of chronic kidney disease can be predicted. The prediction accuracy is high and the sensitivity is good. It can be used as an auxiliary diagnostic method for chronic kidney disease and guide the direction of improving the intestinal flora environment. It is suitable for non-invasive early screening and risk warning of chronic kidney disease.

[0006] A second object of the present invention is to provide a preparation for detecting microbial markers.

[0007] A third object of the present invention is to provide a kit for diagnosing chronic kidney disease.

[0008] The fourth object of the present invention is to provide a method for constructing a chronic kidney disease risk prediction model, which uses the abundance data of microbial markers to construct a chronic kidney disease risk prediction model, which helps to assist in the diagnosis or early warning of the probability of chronic kidney disease, and can be used for early screening, auxiliary diagnosis and prognosis of chronic kidney disease.

[0009] A fifth object of the present invention is to provide a chronic kidney disease risk prediction model.

[0010] One of the purposes of the present invention is achieved by the following technical solution:

[0011] A use of a preparation for detecting microbial markers in the preparation of a preparation for detecting chronic kidney disease, wherein the microbial markers include Prevotella copri, Roseburia faecis, Eggerthella lenta, and Blautia wexlerae.

[0012] Furthermore, the microbial markers also include any one or a combination of two or more of Eubacterium rectale, Eubacterium ventriosum, Lachnospira pectinoschiza, Ruminococcus bicirculans, Ruthenibacterium lactatiformans, Gordonibacter pamelaeae, Clostridium innocuum, Ruminococcus gnavus, Roseburia intestinalis, Megamonas funiformis and Ruminococcus torques.

[0013] Further, the microbial markers include Prevotella copri, Roseburia faecis, Eggerthella lenta, Blautiawexlerae, Eubacterium rectale, Eubacterium ventriosum, Lachnospira pectinoschiza, Ruminococcus bicirculans, Ruthenibacterium lactatiformans, Gordonibacter pamelaeae, Clostridium innocuum, Ruminococcus gnavus, Roseburia intestinalis, Megamonas funiformis and Ruminococcus torques.

[0014] The second object of the present invention is achieved by adopting the following technical solution:

[0015] A preparation for detecting microbial markers, wherein the preparation is used for detecting the abundance of the microbial markers in a stool sample.

[0016] The third object of the present invention is achieved by adopting the following technical solution:

[0017] A kit for diagnosing chronic kidney disease, comprising the preparation for detecting microbial markers.

[0018] The fourth object of the present invention is achieved by adopting the following technical solution:

[0019] A method for constructing a chronic kidney disease risk prediction model comprises the following steps:

[0020] S1, obtaining the abundance of the microbial markers described in the application in stool samples of healthy individuals and patients with chronic kidney disease, respectively, to construct a sample set;

[0021] S2, inputting the sample set into a machine learning model, training and testing the model, and storing it to obtain a chronic kidney disease risk prediction model.

[0022] Furthermore, in step S1, the abundance determination method includes any one or a combination of two or more of metagenomic sequencing, 16S sequencing, 18S sequencing, ITS sequencing and qPCR quantitative detection.

[0023] Furthermore, in step S2, the machine learning model is a random forest model and / or an XGBoost model.

[0024] Furthermore, in step S2, the specific operations are:

[0025] (1) Divide the sample set into a training set and a test set in a ratio of (6-8):(2-4);

[0026] (2) inputting the training set and the test set into a random forest model for training and prediction to obtain a first prediction result;

[0027] (3) inputting the training set and the test set into the XGBoost model for training and prediction to obtain a second prediction result;

[0028] (4) Inputting the first prediction result and the second prediction result into the combined model to obtain a classification prediction result, and constructing a chronic kidney disease risk prediction model; the prediction result calculation formula of the combined model is as follows:

[0029] AUC=(AUC1*probability1+AUC2*probability2) / 2

[0030] In the formula, AUC represents the AUC value of the combined model, AUC1 represents the internal test AUC value of the random forest model, probability1 represents the predicted probability value of the current sample in the random forest model, AUC2 represents the internal test AUC value of the XGBoost model, and probability2 represents the predicted probability value of the current sample in the XGBoost model.

[0031] The fifth object of the present invention is achieved by adopting the following technical solution:

[0032] A chronic kidney disease risk prediction model is constructed using the chronic kidney disease risk prediction model construction method.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] The present invention discloses a microbial marker preparation for use in the preparation of a product for detecting chronic kidney disease. The preparation, which utilizes a combination of Prevotella, Roseburia, Eggerthella tarda, and Blautiae as microbial markers, can effectively predict the probability of a positive diagnosis for chronic kidney disease with high accuracy and sensitivity, and can serve as an auxiliary diagnostic tool for chronic kidney disease. Furthermore, the microbial markers can indicate the status of the intestinal flora, guide adjustments to the intestinal microecology, and improve the therapeutic efficacy of chronic kidney disease.

[0035] The preparation for detecting microbial markers of the present invention can detect the abundance of microbial markers non-invasively.

[0036] The kit for diagnosing chronic kidney disease of the present invention is an economical, non-invasive, efficient and accurate product for early screening and diagnosis of chronic kidney disease.

[0037] The present invention provides a method for constructing a chronic kidney disease risk prediction model, which uses the abundance data of microbial markers to construct a chronic kidney disease risk prediction model, which helps to assist in the diagnosis or early warning of the probability of chronic kidney disease, and can be used for early screening, auxiliary diagnosis and prognosis of chronic kidney disease.

[0038] The chronic kidney disease risk prediction model of the present invention can predict the risk of chronic kidney disease with high specificity and sensitivity. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a schematic diagram of a research plan for a chronic kidney disease risk prediction model of the present invention.

[0040] Figure 2 This is a group rationality evaluation diagram in Example 1 of the present invention.

[0041] Figure 3 This is a diagram of the classifier model framework in Example 2 of the present invention.

[0042] Figure 4 This is an importance score graph of the 38 bacterial sample sets obtained in the XGBoost model in Example 2 of the present invention.

[0043] Figure 5 3 is a trend diagram of the ROC-AUC value of the test set in Example 2 of the present invention.

[0044] Figure 6 This is a diagram of the prediction results of the combined model in Example 2 of the present invention.

[0045] Figure 7 : is the probability distribution diagram of the test samples in Example 2 of the present invention; wherein N is the healthy group population, and Y is the disease group population.

[0046] Figure 8 This is a diagram of the prediction results of the combined model in Example 3 of the present invention.

[0047] Figure 9 3 is a probability distribution diagram of the test samples in Example 3 of the present invention, wherein N is the healthy group population and Y is the disease group population. DETAILED DESCRIPTION

[0048] The present invention will be further described below in conjunction with specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments.

[0049] Example 1

[0050] The specific method for screening microbial markers is as follows:

[0051] (1) A sample set was obtained from the NCBI (National Center for Biotechnology Information) database, which included metagenomic sequencing data of the intestinal flora of chronic kidney disease and healthy people, such as Figure 1 As shown, the sample set included 233 patients with chronic kidney disease and 69 healthy people.

[0052] (2) Intestinal fecal samples from healthy people and patients with kidney disease were collected and sequenced to obtain the DNA sequence of the intestinal flora. The obtained DNA sequence of the intestinal flora was quality controlled, and the abundance of the intestinal flora in the DNA sequence of the intestinal flora was obtained to form a sample set.

[0053] (3) The healthy group and the chronic kidney disease group were split at a ratio of 50% each, and then the split healthy groups were combined with the split chronic kidney disease groups to form two data set groups (denoted as Group A and Group B). Each group contained approximately 50% of healthy people and chronic kidney disease people. For example, Group A included 117 chronic kidney disease patients and 34 healthy people; Group B included 116 chronic kidney disease patients and 35 healthy people.

[0054] The rationality of the above groupings was evaluated: principal coordinate analysis (PcoA) was performed based on species abundance, and PERMANOVA test was performed using the dimensional coordinate distribution characteristics to obtain the rationality evaluation index of the grouping information.

[0055] Principal Coordinate Analysis (PcoA) is a visualization method that reduces the dimensionality of multidimensional data to study the similarity or difference of data, and can describe the relationship between grouped samples. Figure 2 As shown, the two-dimensional visualization of PCoA and the PERMANOVA test vividly show that there is no significant difference between the two groups of samples (p-value is greater than 0.05), which shows that the two groups of samples in this embodiment are reasonably distributed and there is no influence similar to batch effect.

[0056] (4) Analyze the groups obtained in step (3) using LEfSe to obtain the microbial species associated with the disease in each group.

[0057] LEfSe (Linear discriminant analysis Effect Size) identifies the features most likely to explain differences between classes by combining standard tests for statistical significance with tests of coded biological consistency and effect relevance, thereby finding species (i.e., biomarkers) that differ significantly in abundance between groups.Based on LEfSe analysis, a total of 38 potentially valuable bacteria were selected in the present invention, namely Bacteroides galacturonicus, Butyrivibrio crossotus, Coprococcus comes, Coprococcus eutactus, Dialister spCAG_357, Eubacterium eligens, Eubacterium ramulus, Eubacterium rectale, Eubacterium ventriosum, Lachnospira pectinoschiza, Megamonas funiformis, Parasutterella excrementihominis, Prevotella copri, Prevotella sp AM42_24, Prevotella sp CAG_279, Roseburia faecis, Roseburia intestinalis, Ruminococcus bicirculans, Ruminococcus callidus, Ruminococcus torques, Anaerotignum lactatifermentans, Blautia sp CAG_257, Blautia wexlerae, Clostridium innocuum, Clostridium spiroforme, Eggerthella lenta, Eisenbergiella massiliensis, Erysipelatoclostridium ramosum, Firmicutes bacterium CAG_145, Flavonifractorplautii, Fusobacterium mortiferum, Gordonibacter pamelaeae, Hungatellahathewayi, Monoglobus pectinilyticus, Ruminococcus gnavus, Ruthenibacteriumlactatiformans, Sellimonas intestinalis and Tyzzerella nexilis.

[0058] Example 2

[0059] Identification of microbial markers and construction of models

[0060] The present invention utilizes a combination of random forest model and XGBoost model to construct a machine learning combination model, and uses machine learning to select an adaptive prediction model. Supervised learning is to generate a function through the corresponding relationship between a part of the input data and the output data, and map the input to a suitable output, such as classification. The sample data of the present invention have been clinically confirmed and have classified labels, so they will be explored and selected in the supervised machine learning classification model. In this embodiment, the bacterial abundance values ​​of all samples are used as input data, and the diagnostic results of the samples are used as output classification labels. The algorithm is constructed specifically according to the following steps:

[0061] (1) Figure 1 As shown, the sample set in Example 1 is randomly divided into a training set accounting for 70% of the sample set population and a test set accounting for 30% of the sample set population;

[0062] (2) Figure 3 As shown, a machine learning classifier model was constructed using a random forest model and an XGBoost model; the abundance values ​​of all potentially valuable microbial species (38) in Example 1 were used as input data;

[0063] (3) The abundance data table is input into the random forest model and the XGBoost model with cross-validation processing, and then trained and tested to obtain the optimal result output;

[0064] (4) The XGBoost model in step (3) above can obtain the importance score graph of variable features, such as Figure 4 As shown, the number of bacterial variables is gradually increased according to the ranking of the scores;

[0065] (5) Using the abundance of the above-mentioned microbial species in the sample set as input data again, repeating steps (1) and (3) 200 times, obtaining the first prediction results obtained by multiple random forest models and the second prediction results obtained by the XGBoost model, constructing a receiver operating characteristic curve (ROC curve), inputting the first prediction result and the second prediction result into the combined model, calculating the area under the curve (AUC) of the ROC curve of the average test set, and obtaining the classification prediction result;

[0066] The area under the curve is calculated using a combination model, and the calculation formula is as follows:

[0067] AUC=(AUC1*probability1+AUC2*probabiriry2) / 2

[0068] Among them, AUC1 represents the internal test AUC value of the random forest model, probability1 represents the predicted probability value of the current sample in the random forest model, AUC2 represents the internal test AUC value of the XGBoost model, and probability2 represents the predicted probability value of the current sample in the XGBoost model.

[0069] (6) Specific bacteria selection; Based on the above step (5), the variables required for optimal ROC-AUC can be obtained, such as Figure 5 As shown, the relationship between the number of characteristic bacterial groups and the ROC-AUC value was obtained;

[0070] The results showed that the ROC-AUC value was at a higher level when the input feature variable was the bacterial abundance of more than 4 specific species. Figure 4 The importance score of the bacterial genus in the classification is determined. When it is determined that the microbial marker includes four or more species of the genera Prevotella copri, Roseburia faecis, Eggerthella lenta, and Blautia wexlerae, the ROC-AUC value is high and the classifier effect is good, indicating that the microbial marker has high sensitivity and specificity;

[0071] At the same time, from Figure 5 It can be seen that the ROC-AUC value is the largest when the input feature variable is the bacterial abundance of 15 specific species; and Figure 4 It can be seen that changes in input variables will produce different ROC-AUCs. The present invention optimizes the combination of the most suitable input variables and the model. That is, using the abundance of the 15 species described in the present invention as the input object can reduce the requirements for the microbial marker detection method under the condition of higher prediction accuracy.

[0072] Fifteen species were selected as microbial markers, including Prevotella copri, Roseburia faecis, Eggerthella lenta, Blautia wexlerae, Eubacterium rectale, Eubacterium ventriosum, Lachnospira pectinoschiza, Ruminococcus bicirculans, Ruthenibacterium lactatiformans, Gordonibacter pamelaeae, Clostridium innocuum, Ruminococcus gnavus, Roseburia intestinalis, and Megamonas monomorpha. funiformis and Ruminococcus torques, these 15 bacterial genera had the best sensitivity and specificity when used as microbial markers.

[0073] (7) Store the combined model; based on the 15 characteristic bacterial genera in step (6), select the combined model with the best ROC-AUC score. The ROC curve and probability distribution diagram of the combined model are as follows: Figure 6-7 shown.

[0074] Reference Figure 6 The AUC value of the selected combined model is as high as 1.000, and the authenticity of the detection is extremely high, indicating that the combined model can be used to predict the risk of chronic kidney disease in subsequent measurement data; Figure 7 The risk values ​​of the N (healthy group) population were all less than 0.4, and both the N (healthy group) and Y (disease group) populations were distributed in the interval of 0.4≤risk value<0.5. The risk values ​​of the Y (disease group) population were all greater than 0.5, and the changes were more obvious at the risk value of 0.7.

[0075] Therefore, the output value determination result of the combined model is as follows:

[0076] 1) Risk value <0.4, healthy people, no need to adjust the intestinal flora;

[0077] 2) 0.4≤Risk value<0.5, sub-healthy people, need to regulate intestinal flora;

[0078] 3) 0.5≤Risk value<0.7: People at low risk of kidney disease need long-term regulation of intestinal flora. Regular testing is recommended to see if the flora has improved.

[0079] 4) Risk value ≥ 0.7, high-risk group for kidney disease, clinical diagnosis is recommended.

[0080] Example 3

[0081] Clinically validated:

[0082] (1) Detection of relative abundance of intestinal microbial markers:

[0083] An independent external validation dataset was obtained, which included stool samples from 13 patients with chronic kidney disease and 24 healthy subjects. Metagenomic sequencing was performed on the samples to find the abundance of the microbial markers composed of the 15 bacterial genera in Example 2. The test data was input into the model to construct the ROC curve. The prediction effect was as follows: Figure 8 shown.

[0084] (2) Positive risk value prediction:

[0085] The test data obtained by sequencing analysis of the above external validation dataset was input into the combined model in Example 2 to obtain the probability between N (healthy group) and Y (disease group). The results are as follows: Figure 9 shown.

[0086] Ultimately, the Y (disease group) probability value was confirmed as the risk value. A value less than 0.4 was considered healthy, a value between 0.4 and 0.5 was considered sub-healthy, and recommended for certain intestinal flora adjustments. A value between 0.5 and 0.7 was considered low-risk for chronic kidney disease, and recommended for long-term and reasonable adjustment of the intestinal flora structure and regular flora testing to ensure that the flora structure has improved, thereby reducing the risk of subsequent colorectal cancer. A value exceeding 0.7 was considered high-risk for chronic kidney disease, and recommended for outpatient examination and diagnosis. For those without chronic kidney disease, intestinal flora adjustments were recommended. For those with chronic kidney disease, a gut flora profile was provided to clinicians for reference. The actual disease status and risk values ​​of the 37 subjects are shown in Table 1.

[0087] Table 1 Actual disease status and risk values ​​of the subjects

[0088]

[0089]

[0090]

[0091] As shown in Table 1, the chronic kidney disease microbial markers of the present invention can be effectively used to construct a chronic kidney disease risk prediction model with high prediction sensitivity and good specificity. Figure 8 and Figure 9 As shown, the chronic kidney disease risk prediction model of the present invention was verified using an external validation data set, with an AUC value of 0.962, that is, the accuracy rate reached 96.2%; and it can effectively distinguish positive results (healthy people) and negative results (chronic kidney disease patients) in multiple samples, providing effective data for early screening and mid-to-late stage treatment, laying the foundation for disease research.

[0092] In summary, the chronic kidney disease risk prediction model of the present invention uses the abundance of specific bacteria as an input indicator to construct a corresponding chronic kidney disease risk prediction model. In addition to auxiliary diagnosis, it can improve the intestinal flora through medical intervention to achieve the effect of auxiliary treatment. It can also be used for prevention and warning, guiding individuals to adjust their diet and other means to adjust the intestinal flora structure. It is easy to operate, economical and non-invasive, which is conducive to promotion and popularization, and is conducive to reducing the risk of chronic kidney disease and alleviating the possibility of chronic kidney disease.

[0093] The above embodiments are only preferred embodiments of the present invention and cannot be used to limit the scope of protection of the present invention. Any non-substantial changes and replacements made by technicians in this field on the basis of the present invention fall within the scope of protection required by the present invention.

Claims

1. Use of a preparation for detecting microbial markers in the preparation of a product for detecting chronic kidney disease, characterized in that: The microbial markers include: Prevotella copri, Roseburia faecis, Eggerthella lenta, Blautiawexlerae, Eubacterium rectale, Eubacterium ventriosum, Lachnospirapectinoschiza, Ruminococcus bicirculans, Ruthenibacterium lactatiformans, Gordonibacter pamelaeae, Clostridium innocuum, Ruminococcus gnavus, Roseburiaintestinalis, Megamonas funiformis and Ruminococcus torques; The applications include: Detecting the abundance of the microbial marker in the sample, and inputting the abundance data into a chronic kidney disease risk prediction model to predict the risk of chronic kidney disease and obtain a risk value; The determination results are drawn based on the risk value, including: (1) Risk value <0.4, indicating healthy population; (2) 0.4≤risk value<0.5, indicating sub-healthy population; (3) 0.5≤risk value<0.7, indicating a low-risk group for kidney disease; (4) Risk value ≥ 0.7, indicating a high-risk group for kidney disease.

2. A preparation for detecting microbial markers, characterized in that: The preparation for detecting microbial markers is used to detect the abundance of the microbial markers described in the application of claim 1 in a stool sample.

3. A kit for diagnosing chronic kidney disease, characterized in that: The invention also comprises the preparation for detecting microbial markers according to claim 2.

4. A method for constructing a chronic kidney disease risk prediction model, characterized in that: The following steps are involved: S1, obtaining the abundance of the microbial markers described in the application of claim 1 in stool samples of healthy individuals and patients with chronic kidney disease, respectively, to construct a sample set; S2, inputting the sample set into a machine learning model, training and testing the model, and storing it to obtain a chronic kidney disease risk prediction model.

5. The method for constructing a chronic kidney disease risk prediction model according to claim 4, wherein: In step S1, the abundance determination method includes any one or a combination of two or more of metagenomic sequencing, 16S sequencing, 18S sequencing, ITS sequencing and qPCR quantitative detection.

6. The method for constructing a chronic kidney disease risk prediction model according to claim 4, wherein: In step S2, the machine learning model is a random forest model and an XGBoost model.

7. A chronic kidney disease risk prediction system, characterized by: The method for constructing a chronic kidney disease risk prediction model according to any one of claims 4 to 6 is used.

Citation Information

Patent Citations

  • Marker for predicting hereditary angioedema attack and application thereof

    CN113736896A

  • Microbial marker of colorectal cancer and application of marker

    CN109943636A

  • Biomarkers for end-stage renal disease and application thereof

    CN110878349A

  • Power plant reheating flue gas baffle operation prediction method based on integrated hybrid model

    CN115526433A