Application of Candida albicans as a biomarker in the diagnosis of heart failure
By using Geotrichum candida and its combinations as biomarkers, combined with machine learning algorithms, the shortcomings of heart failure diagnosis have been addressed, enabling early diagnosis and effective intervention, and reducing the morbidity and mortality of heart failure.
Patent Information
- Application Number
- CN202510840276.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-01
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-07-01
AI Technical Summary
Current technologies have limited diagnostic methods for heart failure, lacking effective early diagnostic techniques, resulting in high morbidity and mortality rates. Furthermore, research on the association between oral flora and heart failure is insufficient.
Using Geotrichum candidum and its combination as biomarkers, a diagnostic model was constructed by detecting the abundance differences of specific strains in the oral microbiota. Metagenomic sequencing, 16S-rRNA sequencing and other technologies were combined with machine learning algorithms for the early diagnosis of heart failure.
It has enabled early diagnosis of heart failure, improved diagnostic efficiency, guided clinical intervention, and reduced readmission and mortality rates for heart failure.
Smart Images

Figure CN120350166B_ABST
Abstract
Description
[0001] The present application is a divisional application, and the parent application information is as follows: the title of the invention is the application of oral flora in the diagnosis of heart failure; the original application date is July 1, 2024, and the original application number is 2024108713528. TECHNICAL FIELD
[0002] The present application relates to the field of clinical medicine, in particular to the application of Geotrichum candidum as a biomarker in the diagnosis of heart failure. BACKGROUND
[0003] Heart failure (HF, hereinafter referred to as heart failure) is a group of clinical syndromes caused by various structural or functional abnormalities of the heart, which leads to impaired ventricular filling or ejection capacity. It is the end stage of various heart diseases. With the increasing aging of the population, the prevalence of chronic diseases of the cardiovascular system such as coronary heart disease and hypertension is gradually increasing, and the prevalence of heart failure is increasing year by year in China and even the world. According to statistics, the prevalence of heart failure in the adult population is about 1-2%, among which the prevalence of heart failure in people over 35 years old in China is about 1.3%, and the prevalence of heart failure in people over 70 years old is more than 10%. Chronic heart failure leads to decreased exercise tolerance, complications, and other serious effects on the quality of life of patients. In addition, heart failure has a high mortality rate. According to statistics, the mortality rate of patients with congestive heart failure within one year after discharge is about 16.5%, and the 5-year mortality rate of heart failure inpatients is more than 75%. Despite the continuous improvement of medical level, the readmission rate and mortality rate of heart failure patients remain high. The increasing prevalence of heart failure and high mortality rate have brought serious medical and economic burden to patients and society, and have become a global public health problem. However, the effective prevention and treatment measures for heart failure are still limited. Therefore, it is essential to study the pathogenesis of heart failure, achieve early diagnosis, and find protection strategies to reduce the harm of heart failure.
[0004] Recent studies have found that flora has a very close relationship with inflammation. Oral flora has been shown to be associated with many systemic inflammatory diseases. However, the evidence of the relationship between oral flora and heart failure risk is very limited, and is mostly based on the risk between periodontitis and heart failure. Therefore, the present application studies oral flora as a starting point to provide a new direction for the diagnosis and treatment of heart failure. SUMMARY
[0005] In view of the deficiencies of the prior art, the present application provides Geotrichum candidum as an independent biomarker for diagnosing heart failure, and a diagnostic marker combination comprising Geotrichum candidum.
[0006] To achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows:
[0007] The first aspect of the present application provides use of a reagent for detecting abundance of a biomarker in a sample in the manufacture of a product for diagnosing heart failure, the biomarker being Geotrichum candidum.
[0008] The second aspect of the present application provides use of a reagent for detecting abundance of a biomarker in a sample in the manufacture of a product for diagnosing heart failure, the biomarker being a combination of Geotrichum candidum, Neisseria perflava, Neisseria subflava, Veillonella parvula, Neisseria cinerea, Haemophilus parainfluenzae, Neisseria sicca, Streptococcus toyakuensis, Haemophilus influenzae, Veillonella massiliensis.
[0009] In the present application, the term "biomarker" refers to a specific strain in the oral microbiota. In some embodiments, the term also encompasses a measurable entity, such as the abundance of the biomarker, that has been determined to be indicative of a target output, e.g., one or more diagnostic, prognostic, predictive drug sensitivity, and / or therapeutic outputs.
[0010] The biomarker can be differentially present at any level, but is generally present at a level that is increased by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 110%, at least 120%, at least 130%, at least 140%, at least 150%, or more; or is generally present at a level that is decreased by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% (i.e., not present), preferably, the biomarker has a statistical difference (P<0.05).
[0011] In the present invention, the term "oral microbiota" refers to a collection of microorganisms colonized in the human oral cavity, the components of which include bacteria, fungi, viruses, etc., many of which are detected at low frequency and abundance in the oral cavity of healthy humans, and thus specific species can be used as biomarkers to distinguish patients from healthy people.
[0012] In the present invention, the term "abundance" refers to the proportion of the number of species of a biomarker in a sample to the total number of species of all microorganisms in the sample. This proportion is usually expressed in percentage.
[0013] Further, the method for detecting the abundance of the biomarker includes any one or more of metagenomic sequencing, 16S-rRNA sequencing, ITS sequencing, qRT-PCR method, Southern blotting method, in situ hybridization method.
[0014] Further, the abundance of the biomarker is determined by amplifying the ASV sequence of the biomarker in the sample of the subject and determining the proportion of the ASV sequence in the total sample.
[0015] Further, the reagent for detecting the abundance of the biomarker is a primer or probe capable of specifically amplifying the ASV sequence of the biomarker, and the ASV sequence is the sequence shown in SEQ ID NO: 1-10.
[0016] In the present invention, ASV is the abbreviation of Amplicon Seauence Varant, which refers to the determination and analysis of DNA sequences in microbial communities by high-throughput sequencing technology, so as to obtain the ASV number of each microorganism. The ASV number corresponds to the sequence of SEQ ID, which has sufficient nucleotide diversity to distinguish the genus and species of the biomarker.
[0017] In the context of the present invention, the term "sample" as used refers to a composition obtained or derived from a subject (e.g., an individual of interest) that contains cellular and / or other molecular entities to be characterized and / or identified with respect to, for example, physical, biochemical, chemical, and / or physiological characteristics. For example, a sample refers to any sample derived from a subject of interest that is expected or known to contain cellular and / or molecular entities to be characterized. Samples include, but are not limited to, tissue samples, primary or cultured cells or cell lines, cell cultures, cell supernatants, cell lysates, platelets, serum, plasma, vitreous humor, lymph, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, sputum, tears, sweat, mucus, saliva, dental plaque, carious lesion plaque, supragingival plaque, subgingival plaque, peri-implant submucosal plaque, intracanalicular plaque, dorsal tongue plaque, and other mucosal surface plaque samples, tissue culture fluid, tissue extract, homogenized tissue, cell extract, and combinations thereof.
[0018] In the present application, the subject refers to an individual, preferably a mammal, such as a mouse, a rat, a rabbit, a sheep, a monkey, a human, particularly preferably a human, who has or is suspected to have heart failure.
[0019] Further, the sample is a sample containing oral flora of the subject to be tested.
[0020] Further, the sample containing oral flora of the subject to be tested includes saliva, dental plaque, carious plaque, supra-gingival plaque, sub-gingival plaque, peri-implant mucosa plaque, intracanal plaque, dorsal tongue plaque, and other mucosal surface plaque samples.
[0021] Further, the sample containing oral flora of the subject to be tested is supra-gingival plaque.
[0022] Further, the product further comprises a sample processing reagent.
[0023] Further, the sample processing reagent includes a cryoprotectant, a buffer solution, and a storage medium.
[0024] The third aspect of the present application provides a method for constructing a heart failure diagnosis model, the steps of the method comprising: obtaining the abundance data of the biomarker of the second aspect of the present application in the sample and the clinical characteristics corresponding to the sample, and constructing a diagnosis model based on the abundance data and the clinical characteristics.
[0025] The clinical characteristics are heart failure patients and healthy controls.
[0026] The constructed diagnosis model is constructed by an algorithm.
[0027] As a skilled person knows, the step of associating the biomarker abundance data of the subject with a certain likelihood or risk can be implemented and realized in different ways. The abundance measurements of the biomarkers and one or more other markers are mathematically combined, and the combined value is associated with the underlying diagnostic problem. The measurements of the biomarker abundance data can be combined by any suitable prior art mathematical method, such as at least one of logistic regression, linear discriminant analysis, signature gene linear discriminant analysis, support vector machine, random forest, cross-validation, receiver operating characteristic curve, recursive partitioning tree, XGBoost, Shrunken Centroids, StepAIC, Kth-Nearest Neighbor, Boosting, neural network, Bayesian network, and hidden Markov model.
[0028] Further, the algorithm comprises at least one of logistic regression, linear discriminant analysis, linear discriminant analysis of features, support vector machine, random forest, cross-validation, receiver operating characteristic curve, recursive partitioning tree, XGBoost, ShrunkenCentroids, StepAIC, Kth-Nearest Neighbor, Boosting, neural network, Bayesian network, hidden Markov model.
[0029] The fourth aspect of the present application provides a computer-implemented heart failure diagnosis system, which comprises a classification unit, wherein the classification unit uses the diagnosis model obtained by the construction method of the third aspect of the present application to obtain a classification result of whether a subject has heart failure or is at risk of having heart failure, or obtain a classification result of whether a subject does not have heart failure.
[0030] Further, the diagnosis system further comprises an input unit and an output unit.
[0031] The input unit is used to input the abundance data of the biomarkers in claim 2.
[0032] The output unit is used to output the classification result of the classification unit.
[0033] The input unit and the output unit are connected to the classification unit in a certain way.
[0034] The present application provides a system programmed to implement the methods of the present application. The system can regulate various aspects of the oral microbiome analysis of the present application, such as, for example, matching data to known sequences. The system can be a user's electronic device or a computer system located remotely with respect to the electronic device. The electronic device can be a mobile electronic device.
[0035] Advantages and beneficial effects of the present application: The present application provides Geotrichum candidum and a biomarker combination comprising Geotrichum candidum, which has high heart failure diagnosis efficiency. The present application can realize early diagnosis of heart failure, guide clinicians to provide prevention or treatment programs for subjects, and thus better intervene in the disease progression of the occurrence and development of heart failure. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 Figure 1 is a random forest prediction result chart for the sample of the experimental set, wherein A is a random forest prediction model visualization image, and B is the importance ranking result of 10 species screened out.
[0037] Figure 2 Figure 2 is a classification result chart for the sample of the experimental set, wherein A is the abundance difference level of heart failure patients and healthy controls, and B is the ROC curve chart of the experimental set sample combination.
[0038] Figure 3 ROC curve of individual species of the experimental set sample.
[0039] Figure 4 Classification result graph of the verification set sample, wherein A is the abundance difference level of heart failure patients and healthy controls, and B is the ROC curve graph of the verification set sample combination.
[0040] Figure 5 ROC curve of individual species of the verification set sample. DETAILED DESCRIPTION
[0041] The various reagents involved in the technical solutions and experimental processes described in the present application are common reagents or commercial reagents that can be clearly known and easily obtained by those skilled in the art based on their professional knowledge and conventional practice. The description of reagents in the present application is intended to clearly illustrate the material basis involved in the technical solutions, and those skilled in the art can successfully obtain and correctly use these reagents based on their professional accomplishment and industry common sense to achieve the technical purpose of the present application.
[0042] The present application performs difference test on the abundance of all species in the oral flora of heart failure patients and healthy controls, and screens out the species with significant difference between the two groups through algorithm. The differentially expressed species preliminarily screened out are classified and modeled to select biomarkers that can be used for classification between groups. After analysis, 10 key strains are found, which are ASV10 (Streptococcus toyakuensis, Streptococcus), ASV24 (Neisseria perflava, Neisseria), ASV37 (Neisseria subflava, Neisseria subflava), ASV7 (Veillonella parvula, Veillonella), ASV215 (Neisseria cinerea, Neisseria cinerea), ASV1 (Haemophilus parainfluenzaes, Haemophilus parainfluenzaes), ASV127 (Neisseria sicca, Neisseria meningitidis), ASV443 (Geotrichum candidum, Geotrichum candidum), ASV25 (Haemophilus influenzae, Haemophilus influenzae), and ASV86 (Veillonella massiliensis, Veillonella massiliensis), and the ASV sequences are the sequences shown in SEQ ID NO: 1-10.
[0043] Further research found that the individual diagnostic AUC and combined diagnostic AUC of 10 species in the validation sample set were greater than 0.7, which had significant effectiveness in distinguishing heart failure patients and healthy controls. Among them, the AUC of C. albicans in the validation set was as high as 0.912, indicating that it could be used as an independent biomarker for heart failure.
[0044] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0045] Example 1 Screening and verification of biomarkers related to heart failure in oral flora
[0046] I. Experimental materials and experimental steps
[0047] 1. Experimental samples
[0048] 94 samples. The experimental group included 40 cases (17 HC, 23 HF), and the validation group included 54 cases (14 HC, 40 HF).
[0049] The samples were selected from 63 patients hospitalized in Peking Union Medical College Hospital from April 2022 to March 2023, who were diagnosed with HF by at least two experienced cardiologists according to the 2021 ESC heart failure guidelines. Detailed demographic, medical history, laboratory test results, and echocardiogram data were obtained through the electronic medical record system of PUMCH. Exercise tolerance was assessed using the New York Heart Association (NYHA) functional classification.
[0050] According to the clinical manifestations of patients, we divided them into: 1) Chronic HF (CHF): Patients who have received appropriate anti-HF treatment and have stable clinical manifestations for at least 1 month; 2) Acute HF (AHF): Patients with new-onset or acute decompensated heart failure; NYHA classification is II to IV at enrollment; have ≥1 symptoms (such as increased dyspnea, orthopnea) or ≥1 signs (such as rales, peripheral edema, serous effusion) at admission; and NT-proBNP (N-terminal of brain natriuretic peptide) >300 pg / ml or B-type natriuretic peptide (BNP) >100 pg / ml.
[0051] According to the left ventricular ejection fraction (LVEF) assessed by echocardiography, the subjects were divided into 1) heart failure with reduced ejection fraction (HFrEF): LVEF≤40%; 2) heart failure with mid-range ejection fraction (HFmrEF): LVEF 41%-49%; 3) heart failure with preserved ejection fraction (HFpEF): LVEF≥50%; 4) heart failure with improved ejection fraction (HFimpEF): LVEF≤40% at baseline, improved up to 40% within 1 month to 1 year after discharge, and increased by at least ≥10%.
[0052] A total of 31 subjects without symptoms or signs related to heart disease or HF were included in the control group. Subjects with one or more of the following conditions were excluded: 1) evidence of acute myocardial infarction; 2) received cardiac surgery within the previous 6 months; 3) systemic diseases such as systemic lupus erythematosus and malignancy; 4) end-stage renal disease (eGFR < 15 ml / kg / 1.73 m2, CKD-EPI); 5) use of antibiotics for more than 3 days within the previous 3 months.
[0053] The research protocol was approved by the Ethics Committee of Peking Union Medical College Hospital (K23C0687) and was performed in accordance with the principles of the Declaration of Helsinki. All subjects signed a written informed consent form to participate in the study.
[0054] 2. Sample collection
[0055] Sample type: Supragingival plaque.
[0056] Specific steps: Supragingival plaque was collected from all subjects using a sterile swab. All subjects rinsed their mouths with 10 milliliters of 0.9% sodium chloride solution for 15 seconds before sampling to remove residual food debris in the oral cavity, and none of them ate, drank water, brushed their teeth, or used any oral cleaning liquid within two hours before sampling. The collected samples were stored at -80°C.
[0057] 3. DNA extraction
[0058] The CTAB extraction method was used to extract the genomic DNA of the flora in the samples.
[0059] 4. Sequencing
[0060] The 16S rRNA and ITS rRNA genes were amplified using 16S V3-4 region and ITS 1-1F region, respectively. The sequencing library was generated using NEB Next® Ultra™ II FS DNA PCR-FREE Library Preparation Kit (New England Biolabs, USA, Catalog #: E7430L) according to the manufacturer's recommendations. The library was quantified using Qubit and real-time PCR, and size distribution was detected using a bioanalyzer. Based on the PE250 strategy, the quantified library was pooled and sequenced on an Illumina NovaSeq 6000. All the above procedures were completed at Novogene Bioinformatics Technology Co., Ltd. The ASV sequences of the biomarkers described in the present application are shown in Table 1.
[0061] Table 1. SEQ ID corresponds to ASV sequence table
[0062]
[0063]
[0064]
[0065]
[0066]
[0067]
[0068] 5. Species diversity analysis
[0069] Wilcoxon rank-sum test was used for species diversity analysis, and the specific steps were as follows: by ranking, the original data information was replaced by rank to perform the test. The data of two populations were mixed and sorted, and the average values of the two populations were calculated. If the average ranks of the two populations are not significantly different, it indicates that there is no significant difference between the two populations; otherwise, further statistical P value is calculated, and the species with P value less than 0.05 are screened out, i.e. the species with significant difference.
[0070] 6. Construction of classification model
[0071] The present application uses microbial species and clinical index variables to construct a random forest classifier. Five 10-fold cross-validation is used to select microbial indicators to construct the final classifier, and a receiver operating characteristic curve (AUC) is established to evaluate the performance of the classifier.
[0072] 1) Random forest analysis (Random Forest, RF)
[0073] Belongs to machine learning algorithm, is a classifier containing multiple decision trees, which can quickly and efficiently select the most important species categories for sample classification.
[0074] 2) Cross-validation analysis
[0075] Using 5 times 10-fold cross-validation (K-folder Cross-validation, K-folder CV), the combination of key species screened by random forest method is traversed to construct an efficient classifier with the lowest error rate using the optimal species combination.
[0076] 3) Receiver Operating Characteristic curve (ROC) analysis
[0077] Used to verify the classification performance of the classification model. According to the classification test results, the upper and lower limits, group distance and cutoff point of the measured value are determined, and the sensitivity and specificity of all cutoff points are calculated. The larger the area under the curve (Area Under the Curve, AUC), the better the performance of the classification model.
[0078] II. Experimental results
[0079] 1. Screening results
[0080] In the experimental group, 270 species with significant differences between the two groups were selected by differential test results, and the random forest classification model was trained from 270 species, as shown in Figure 1 A.
[0081] After feature selection based on 5-fold 10-fold cross-validation, there are 10 marker species retained with the best performance, which are ASV10 (Streptococcus toyakuensis, Streptococcus), ASV24 (Neisseria perflava, Neisseria), ASV37 (Neisseria subflava, Neisseria subflava), ASV7 (Veillonella parvula, Veillonella), ASV215 (Neisseria cinerea, Neisseria cinerea), ASV1 (Haemophilus parainfluenzaes, Haemophilus parainfluenzaes), ASV127 (Neisseria sicca, Neisseria meningitidis), ASV443 (Geotrichum candidum, Geotrichum candidum), ASV25 (Haemophilus influenzae, Haemophilus influenzae), and ASV86 (Veillonella massiliensis, Veillonella massiliensis). The 10 species are ranked according to the importance of the results as shown in Figure 1 B.
[0082] The ROC curve further verifies the classification performance of the classification model, as shown in Figure 2 , the AUC obtained by the combination of 10 species is as high as 0.96, indicating that the combination of 10 species has good classification performance in diagnosing heart failure.
[0083] 2. Verification results
[0084] The 10 species combination screened out from the experimental sample set is used to classify and verify the validation group, as shown in Figure 4 , according to the ROC curve, the AUC is as high as 0.96, indicating that the combination of 10 species has good classification performance in diagnosing heart failure and can be used as a biomarker for heart failure. The 10 species screened out are used to draw the ROC curve respectively, as shown in Figure 5 , it can be seen that among them, ASV443, i.e. Geotrichum candidum, has a diagnosis AUC as high as 0.912, indicating that Geotrichum candidum can be used as an independent biomarker for heart failure, and the prediction classification performance is excellent.
[0085] The application has been described in detail. For those skilled in the art, the application can be implemented in a wider range under the same parameters, concentrations and conditions without departing from the spirit and scope of the application and without unnecessary experiments. Although the application gives examples, it should be understood that further improvements can be made to the application. In summary, according to the principle of the application, the present application is intended to include any changes, uses or improvements of the application, including changes made by using conventional techniques known in the art, which depart from the scope disclosed in the present application.
Claims
1. The application of a reagent for detecting the abundance of biomarkers in a sample in the preparation of products for diagnosing heart failure, characterized in that, The biomarker is *Geotrichum candidum*. The abundance of the biomarkers was determined by amplifying the ASV sequences of the biomarkers in the subject samples and then calculating their proportion in the total sample. The reagent for detecting the abundance of the biomarker is a primer or probe capable of specifically amplifying the ASV sequence of the biomarker, wherein the ASV sequence is the sequence shown in SEQ ID NO:8; The sample is a sample containing the oral microbiota of the subject.
2. The application according to claim 1, characterized in that, The methods for detecting the abundance of the biomarkers include any one or more of sequencing, qRT-PCR, DNA blotting, and in situ hybridization.
3. The application according to claim 1, characterized in that, The samples containing the oral microbiota of the test subject include saliva, dental plaque, caries plaque, supragingival plaque, subgingival plaque, periimplant submucosal plaque, root canal plaque, dorsum of the tongue plaque, and other mucosal surface plaque samples.
4. The application according to claim 1, characterized in that, The sample containing the oral microbiota of the test subject is supragingival plaque.
5. The application according to claim 1, characterized in that, The product also includes reagents for sample processing.
6. The application according to claim 5, characterized in that, The reagents used for sample processing include cryoprotectants, buffer solutions, and storage culture media.
Citation Information
Patent Citations
Heart failure biomarker and application thereof
CN117683849A