Microbial marker combination for predicting breast cancer prognosis and application thereof

By constructing a risk scoring model based on a combination of microbial biomarkers, the challenge of predicting the prognosis of breast cancer has been solved, enabling accurate prediction of the prognosis of breast cancer patients and the development of personalized treatment plans.

CN121137147APending Publication Date: 2025-12-16SUZHOU JIANLIKANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511094208.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

There is a lack of effective methods in the current technology to predict the prognosis of breast cancer patients and to provide personalized treatment plans based on microbial composition and abundance.

Method used

A risk scoring model based on a combination of microbial biomarkers, including Dicipivirus, Tymovirus, Sclerodarnavirus, and Bafinivirus, was constructed. By measuring the abundance of microorganisms in tumor samples from breast cancer patients, a risk score was calculated using the Risk Score formula, and patients were divided into high-risk and low-risk groups to predict their prognosis.

Benefits of technology

This model can significantly predict the prognosis of breast cancer patients, help doctors develop personalized treatment plans, improve the effectiveness of treatment, and validate the model's good predictive performance through an external validation set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121137147A_ABST
    Figure CN121137147A_ABST
Patent Text Reader

Abstract

The invention relates to the field of medical detection, and provides a microbial marker combination for predicting breast cancer prognosis and application thereof. According to the present invention, the significant correlation between the microbial spectrum and the breast cancer prognosis is found, the comprehensive data analysis of the microbiome abundance, the clinical information and the transcriptome data is performed on the patient to determine the specific 22 microbial markers related to the breast cancer prognosis, and the risk scoring model for predicting the breast cancer prognosis is established; the microbial model is used for predicting the immunotherapy curative effect and response of breast cancer patients, and can help a doctor to more accurately judge which patients are more likely to benefit from immunotherapy before treatment, so that a more personalized and effective treatment scheme is formulated for the patients.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical detection, in particular to a microorganism marker combination for predicting the prognosis of breast cancer and application thereof. BACKGROUND

[0002] Breast cancer is the most common malignant tumor in women worldwide, and despite significant advances in early screening and treatment in recent years, it remains one of the leading causes of cancer death in women. Although tumor heterogeneity and the complexity of clinical manifestations have gradually been recognized, there are still many unknowns about the mechanism of breast cancer. There are trillions of microorganisms in the human body, and the microbiome not only plays an important role in digestion, immunity and metabolism, but also plays an important role in the emerging field of cancer immunotherapy. The mechanism by which the microbiome affects the occurrence and progression of breast cancer has been reported, and they can affect the immune environment, metabolic processes and treatment response of tumors through multiple pathways. The relationship between the microbiome and breast cancer, especially the potential role of the microbiome in the occurrence, progression and treatment of breast cancer, has become a research hotspot. Studies have shown that there are significant differences in the composition and abundance of the microbiome between breast cancer samples and normal samples. The TCGA database provides rich transcriptome data of breast cancer samples, which enables us to further analyze the correlation between breast cancer and the microbiome and further analyze the potential impact of the microbiome on the prognosis and treatment response of breast cancer patients. By identifying specific microbial populations associated with breast cancer, we may be able to develop microbiome-based interventions to improve patient outcomes and quality of life. SUMMARY

[0003] Objective: In order to overcome the deficiencies in the prior art, the present application provides a microorganism marker combination for predicting the prognosis of breast cancer and application thereof, which constructs a risk score model based on microorganism abundance data to predict the prognosis of breast cancer patients.

[0004] To solve the above technical problems, the technical scheme adopted by the present application is:

[0005] In a first aspect, the present invention provides a combination of biomarkers for predicting the prognosis of breast cancer, the combination of microbial biomarkers comprising: Dicipivirus, Tymovirus, Sclerodarnavirus, Bafinivirus, Roseiflexus, Lachnoclostridium, Aggregatibacter, Ornithobacterium, Psychroflexus, Pyrinomonas, Campylobacter, Eli zabethkingia, Desulfotalea, Candidatus_Nitrosopelagicus, Phaeobacter, Planktothricoides, Alt eromonas, Saccharomonospora, Rothia, Agreia, Vibrio, and Microvirga.

[0006] In some embodiments, the microbial biomarker combination comprises Dicipivirus, Tymovirus, Sclerodarnavirus, Bafinivirus, Roseiflexus, Lachnoclostridium, Aggregatibacter, Ornithobacterium, Psychroflexus, Pyrinomonas, Campylobacter, Elizabethkingia, Desulfotalea, Candidatus_Nitrosopelagicus, Phaeobacter, Planktothricoides, Alteromonas, Saccharomonospora, Rothia, Agreia, Vibrio, and Microvirga.

[0007] In a second aspect, the present invention provides the application of a reagent for detecting the abundance of each microorganism in a combination of microbial biomarkers as described in the first aspect in the preparation of a kit for predicting the prognosis of breast cancer.

[0008] In some embodiments, the microorganisms are derived from tumor samples from breast cancer patients.

[0009] Thirdly, the present invention provides a method for predicting the prognosis of breast cancer, comprising: collecting tumor samples from breast cancer patients, measuring the abundance of microorganisms in the tumor samples, calculating the risk score of breast cancer patients based on the abundance data and a risk scoring model, and predicting the prognosis of breast cancer.

[0010] the microorganism is Dicipivirus, Tymovirus, Sclerodarnavirus, Bafinivirus, Roseiflexus, Lachnoclostridium, Aggregatibacter, Ornithobacterium, Psychroflexus, Pyrinomonas, Campylobacter, Elizabethkingia, Desulfotalea, Candidatus_Nitrosopelagicus, Phaeobacter, Planktothricoides, Alteromonas, Saccharomonospora, Rothia, Agreia, Vibrio, Microvirga;

[0011] the model is Risk Score = g_Dicipivirus*0.0441 + g_Tymovirus*(-0.1151) + g_Sclerodarnavirus*(-0.0648) + g_Bafinivirus*0.1060 + g_Bafinivirus*0.0816 + g_Lachnoclostridium*0.1777 + g_Aggregatibacter*0.1354 + g_Ornithobacterium*0.0732 + g_Psychroflexus*(-0.1524) + g_Pyrinomonas*0.0059 + g_Campylobacter*(-0.0665) + g_Elizabethkingia*(-0.2616) + g_Desulfotalea*(-0.0712) + g_Candidatus_Nitrosopelagicus*0.0639 + g_Phaeobacter*(-0.0191) + g_Planktothricoides*0.0492 + g_Alteromonas*0.0857 + g_Saccharomonospora*(-0.0363) + g_Rothia*0.3957 + g_Agreia*(-0.0242) + g_Vibrio*(-0.1111) + g_Microvirga*(-0.4225);

[0012] wherein, Risk Score represents the risk score, g_microorganismLatinName represents the abundance of the microorganism of the genus in the patient sample.

[0013] In some embodiments, the median of the risk scores is taken as a division threshold, when the risk score of a breast cancer patient is greater than the median, the patient is determined to belong to a high-risk group, and when the risk score of a breast cancer patient is less than the median, the patient is determined to belong to a low-risk group, the overall survival of the low-risk group is significantly better than that of the high-risk group; in the low-risk group, the lower the risk score calculated based on the model, the better the prognosis of the breast cancer patient.

[0014] In a fourth aspect, the present application provides a system for predicting the prognosis of breast cancer, the system comprising:

[0015] a data collection module for collecting tumor samples of breast cancer patients and determining the abundance of each microorganism in the tumor samples, the microorganisms being Dicipivirus, Tymovirus, Sclerodarnavirus, Bafinivirus, Roseiflexus, Lachnoclost ridium, Aggregatibacter, Ornithobacterium, Psychroflexus, Pyrinomonas, Campylobacter, Elizabethkingia, Desulfotalea, Candidatus_Nitrosopelagicus, Phaeobacter, Planktothricoides, Alteromonas, Saccharomonospora, Rothia, Agreia, Vibrio, Microvirga;

[0016] a model calculation module for calculating a risk score based on the abundance data of the microorganisms according to the risk score model as described in the third aspect;

[0017] an output prediction module for predicting the prognosis of a breast cancer patient according to the risk score data of the patient, the lower the risk score of the patient, the better the prognosis.

[0018] In a fifth aspect, the present application provides a kit for predicting the prognosis of a breast cancer patient, comprising reagents for detecting the abundance of each microorganism in the microorganism marker combination as described in the first aspect.

[0019] Beneficial effects: The present application determines specific microorganism markers related to the prognosis of breast cancer and establishes a risk score model for predicting the prognosis of breast cancer. The study includes a total of 1071 breast cancer patients, comprehensive data analysis including microbiome abundance, clinical information and transcriptome data, and high-throughput sequence analysis and linear discriminant analysis are used to identify and visualize important microorganism markers.

[0020] Further, the present application uses single factor Cox and LASSOCox regression methods to verify these markers, and 22 microbial markers are obtained for finally constructing a risk score model. There is a significant correlation between the microbial spectrum and breast cancer progression, and specific microorganisms g_Alteromonas and g_Saccharomonospora show significant prognostic significance for breast cancer. The model can predict the efficacy and response of immunotherapy for breast cancer patients, help doctors more accurately judge which patients are more likely to benefit from immunotherapy before treatment, and thus develop more personalized and effective treatment plans for patients. BRIEF DESCRIPTION OF DRAWINGS

[0021] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, illustrate the embodiments of the present application and are used to explain the present application together with the embodiments of the present application to explain the present application, but do not constitute a limitation on the present application.

[0022] Figure 1 The proportions of the phylum, class, order, family, genus and species levels in the breast cancer samples and the paracancer samples in the embodiments of the present application are shown in the following table:

[0023] wherein, Figure 1 A is the phylum level proportion distribution graph; Figure 1 B is the class level proportion distribution graph, and the color blocks from right to left represent p_Caldiserica, p_Spirochaetes, p_Chlamydiae, p_Tenericutes, p_Crenarchaeota, p_Planctomycetes, p_Proteobacteria, p_Thaumarchaeota, p_Euryarchaeota, p_Synergistetes, p_Deinococcus, p_Bacteroidetes, p_Chlorobi, p_Actinobacteria, p_Firmicutes, p_Nitrospinae, p_Chrysiogenetes, p_Armatimonadetes, p_Thermotogae, p_Acidobacteria;

[0024] Figure 1C is the class level proportion distribution map, the color blocks from right to left represent: c_Caldisericia, c_Acidithiobacillia, c_Spirochaetia, c_Chlamydiia, c_Bacteroidia, c_Mollicutes, c_Phycisphaerae, c_Thermoprotei, c_Ktedonobacteria, c_Methanobacteria, c_Tissierellia, c_Erysipelotrichia, c_Halobacte ria, c_Gammaproteobacteria, c_Thermococci, c_Planctomycetia, c_Opitutae, c_Deltaproteobacteria, c_Alphaproteobacteria, c_Coriobacteriia;

[0025] Figure 1 D is the order level proportion distribution map, the color blocks from right to left represent: o_Mycoplasmatales, o_Streptomy cetales, o_Kiloniellales, o_Aeromonadales, o_Spirochaetales, o_Caldisericales, o_Vibrio nales, o_Acidithiobacillales, o_Desulfovibrionales, o_Pasteurellales, o_Chlamydiales, o_Geoderma tophilales, o_Thermogemmatisporales, o_Alteromonadales, o_Sulfolobales, o_Haloferacales, o_Brachyspirales, o_Bacteroidales, o_Xanthomonadales, o_Campylobacterales;

[0026] Figure 1E is a proportion distribution diagram of the genus level, and the color blocks from right to left represent f_Bacteroidaceae, f_Listeria ceae, f_Chlamydiaceae, f_Colwelliaceae, f_Rikenellaceae, f_Shewanellaceae, f_Desulfovibrio naceae, f_Waddliaceae, f_Sutterellaceae, f_Succinivibrionaceae, f_Gordoniaceae, f_Polydnaviridae, f_Nocardiaceae, f_Mycoplasmataceae, f_Pseudoalteromonadaceae, f_Nimaviridae, f_Retroviridae, f_Streptomycetaceae, f_Atopobiaceae, f_Kiloniellaceae;

[0027] Figure 1 F is a proportion distribution diagram of the genus level, and the color blocks from right to left represent g_Terrabacter, g_Bacteroides, g_Proteus, g_Listeria, g_Neisseria, g_Salmonella, g_Bordetella, g_Shigella, g_Aeromonas, g_Campylobacter, g_Streptomyces, g_Chlamydia, g_Nocardia, g_Sanguibacteroides, g_Vibrio, g_Colwellia, g_Enterococcus, g_Luteibacter, g_Desulfovibrio, and g_Leuconostoc.

[0028] Figure 2 Fig. 6 is a statistical diagram of LDA values of characteristic microorganisms based on LEfSe analysis in the embodiment of the present application.

[0029] Figure 3 Fig. 7 is an evolutionary branch diagram of characteristic microorganisms based on LEfSe analysis in the embodiment of the present application.

[0030] Figure 4 Fig. 8 is a forest plot of single-factor Cox regression results in the embodiment of the present application.

[0031] Figure 5 Fig. 9 is a graph of single-factor Cox regression results in the embodiment of the present application; wherein, Figure 5 A to F are KM curves of the top 6 microorganisms in the significance p value.

[0032] Figure 6 Figure 1 is a graph of the risk score model performance verification results of the high-risk group and the low-risk group based on the training set in the embodiments of the present application; wherein, Figure 6 A is the KM curve of the combination of the total survival data of the high-risk group and the low-risk group, Figure 6 B is the KM curve of the combination of the progression-free interval data of the high-risk group and the low-risk group, Figure 6 C is the ROC curve of the combination of the survival data of the high-risk group and the low-risk group.

[0033] Figure 7 Figure 2 is a graph of the risk score model performance verification results of the high-risk group and the low-risk group based on the verification set in the embodiments of the present application; wherein, Figure 7 A is the KM curve of the combination of the total survival data of the high-risk group and the low-risk group, Figure 7 B is the KM curve of the combination of the progression-free interval data of the high-risk group and the low-risk group, Figure 7 C is the ROC curve of the combination of the survival data of the high-risk group and the low-risk group.

[0034] Figure 8 Figure 3 is a nomogram and a ROC curve of the nomogram constructed in the embodiments of the present application; wherein, Figure 8 A is the nomogram of the training set, Figure 8 B is the nomogram of the verification set, Figure 8 C is the ROC curve of the nomogram of the training set, Figure 8 D is the ROC curve of the nomogram of the verification set, Figure 8 E is the C-Index of the nomogram and the risk score in the training set and the verification set.

[0035] Figure 9 Figure 4 is a box plot of the correlation between the risk score and the clinical characteristics in the embodiments of the present application.

[0036] Figure 10 Figure 5 is a heatmap of the correlation between the gene expression and the feature model containing the microbial species abundance in the embodiments of the present application.

[0037] Figure 11 Figure 6 is a graph of the enrichment analysis results of the significantly related gene functions in the embodiments of the present application; wherein, Figure 11 A is the pathway results graph of biological process enrichment, Figure 11 B is the pathway results graph of cell component enrichment, Figure 11 C is the pathway results graph of molecular function enrichment, Figure 11 D is the pathway results graph of KEGG enrichment.

[0038] Figure 12Figure 1 is a box plot of the difference in enrichment scores of immune infiltrating cells in the high-risk group and the low-risk group in the embodiments of the present application.

[0039] Figure 13 Figure 2 is a box plot of the difference in 13 immune function-related scores in the high-risk group and the low-risk group in the embodiments of the present application.

[0040] Figure 14 Figure 3 is a box plot of the difference in expression of immune checkpoints between the high-risk group and the low-risk group in the embodiments of the present application.

[0041] Figure 15 Figure 4 is a result plot of the risk score model evaluated by the external validation set PRJNA759366 in the embodiments of the present application. Figure 15 Figure 4A is a result plot of the risk score model evaluated by the external validation set PRJNA759366 in the embodiments of the present application. Figure 15 Figure 4B is a result plot of the risk score model evaluated by the external validation set PRJNA926328 in the embodiments of the present application. DETAILED DESCRIPTION

[0042] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The following description of at least one exemplary embodiment is actually only illustrative, but not as any limitation on the present application and its application or use.

[0043] Unless otherwise specifically indicated, numerical expressions and values set forth in these embodiments do not limit the scope of the present application. Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail, but should be considered part of the description of the application. In all examples shown and discussed herein, any specific value should be interpreted as merely illustrative, not as a limitation. Thus, other examples of the exemplary embodiments can also include different values. It should be noted that similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0044] The present application will be further described below in conjunction with the embodiments, and in the following embodiments, g represents the genus of microorganisms.

[0045] Example 1 Collection and data analysis of clinical information of TCGA breast cancer samples

[0046] 1071 breast cancer patient samples and their related data were collected through the TCGA database, and the collected data and analysis were summarized in 1-7 and Table 1 below, respectively:

[0047] 1. Microbiome abundance data:

[0048] https: / / ftp.microbio.me / pub / cancer_microbiome_analysis / TCGA / Kraken / Kraken-TCGA-Voom-SNM-Plate-Center-Filtering-Data.csv;

[0049] 2. Microbiome clinical data:

[0050] https: / / ftp.microbio.me / pub / cancer_microbiome_analysis / TCGA / Kraken / Kraken-TCGA-Raw-Data-17625-Samples.csv;

[0051] 3. Expression data:

[0052] https: / / portal.gdc.cancer.gov / projects / TCGA-BRCA;

[0053] 4. Survival data:

[0054] https: / / portal.gdc.cancer.gov / projects / TCGA-BRCA;

[0055] 5. Phenotype data:

[0056] https: / / portal.gdc.cancer.gov / projects / TCGA-BRCA;

[0057] 6. Ensembl database download human gtf file (Homo_sapiens.GRCh38.104.gtf.gz):

[0058] http: / / ftp.ensembl.org / pub / release-104 / gtf / homo_sapiens / Homo_sapiens.GRCh38.104.gtf.gz;

[0059] 7. Immune cell related gene set downloaded from Table S6. List of Pan-cancer Immune Metagenes of PMID: 28052254.

[0060] Table 1 Breast cancer sample clinical information from TCGA database

[0061]

[0062]

[0063] Example 2 Screening and identification of breast cancer prognosis-related microbial features

[0064] Based on the microbial abundance-related information obtained for each breast cancer patient sample in Example 1, the average abundance of microbial populations was then calculated based on the phylum, class, order, family, genus, and species levels. Due to the excessive number of species within the phylum, class, order, family, and genus, only the top 20 microorganisms in terms of abundance at each level were taken to plot the box plot.

[0065] As shown in FIG. 2, the distribution proportions of microorganisms at each level were relatively uniform in the breast cancer samples and the paracancer samples. Figure 1

[0066] Further, LEfSe analysis based on the breast cancer samples and the paracancer samples was performed using the R package microeco (v1.1.0), and the LDA values of the breast cancer-related feature microorganisms were obtained with an alpha < 0.05 significance test threshold. Then, based on the LDA values, the top 20 feature microorganisms with the highest LDA values were taken to plot the LDA value statistics chart, and the evolutionary branch diagram was plotted based on the top 150 microbial populations with the highest abundance.

[0067] As shown in FIG. 3, there were a total of 785 microorganism markers with significant differences between the breast cancer samples and the paracancer samples, of which 331 belonged to the microorganisms with higher LDA values in the breast cancer group, and the remaining 454 belonged to the microorganisms with higher LDA values in the control group (paracancer samples). Figure 2 As shown in FIG. 4, most of the bacteria in the k_Archaea and k_Viruses phyla played an important role in the breast cancer group, while most of the bacteria in the k_Bacteria phylum mainly played an important role in the control group (paracancer samples). Among the top 150 microbial populations with the highest abundance, there were a total of 692 differential microorganism markers, including the phylum, class, order, family, genus, and species levels. After integration, a total of 541 genus-level differential microorganism markers were obtained.

[0068] Figure 3

[0069] Example 3 Construction and verification of breast cancer microbial prognosis feature model

[0070] 1. Single factor Cox regression:

[0071] ​​​The median of the microbial abundance data was combined according to the TCGA patient ID number of the sample, and the significant feature microorganisms obtained by LEfSe analysis were used for subsequent analysis. The TCGA-BRCA data of breast cancer samples obtained in Example 1 were divided according to 2:1, of which 67% of the data was the training set for modeling, and 33% of the data was the validation set for verification. The training set had a total of 714 samples, and the validation set had a total of 357 samples. Further, based on the abundance of the 541 genus-level differential microbial markers in Example 2, further combined with the overall survival data (Overall survival, OS) of the patients, the R package survival (v3.2-7) and survminer (v0.4.8) were used to perform batch Cox single factor regression analysis on the feature microbial marker abundance values of the training set samples. The threshold of p<0.05 was used to screen the feature microbial markers significantly related to overall survival for subsequent analysis.

[0072] As shown in Figure 4 , based on the p<0.05 correlation threshold, 22 prognostically significant microbial markers were obtained, of which 9 microbial markers had HR>1, indicating that an increase in microbial abundance may cause poor prognosis, and the remaining 11 microbial markers had HR<1, indicating that a decrease in microbial abundance may cause poor prognosis.

[0073] Further, the top 6 microorganisms in terms of p value were plotted to generate KM curves, as shown in Figure 5 .

[0074] 2. LASSOCox regression analysis:

[0075] Further, lasso regression dimensionality reduction was performed on the prognostically related feature microbial markers, and the 22 microbial markers were used to construct a risk score model, which mainly relied on the R package glmnet (v4.0-2). In order to construct a more accurate regression model, first, lambda screening was performed using cross-validation, then the model corresponding to lamdba.min was selected, the abundance matrix of the related feature microbial markers in the model was further extracted, and the risk score of each sample was calculated based on the following formula:

[0076]

[0077] wherein abundance represents the abundance of the corresponding microbial marker, i represents the sample, j represents the microbial marker, β represents the regression coefficient (coef) of the corresponding microbial marker in the lasso regression result, and RScore represents the sum of the abundance of the significantly related microbial markers in each sample multiplied by the coef of the corresponding marker.

[0078] The final microbial marker prognostic risk score model is: Risk Score = g_Dicipivirus*0.0441 + g_Tymovirus*(-0.1151) + g_Sclerodarnavirus*(-0.0648) + g_Bafinivirus*0.1060 + g_Bafinivirus*0.0816 + g_Lachnoclostridium*0.1777 + g_Aggregatibacter*0.1354 + g_Ornithobacterium*0.0732 + g_Psychroflexus*(-0.1524) + g_Pyrinomonas*0.0059 + g_Campylobacter*(-0.0665) + g_Elizabethkingia*(-0.2616) + g_Desulfotalea*(-0.0712) + g_Candidatus_Nitrosopelagicus*0.0639 + g_Phaeobacter*(-0.0191) + g_Planktothricoides*0.0492 + g_Alteromonas*0.0857 + g_Saccharomonospora*(-0.0363) + g_Rothia*0.3957 + g_Agreia*(-0.0242) + g_Vibrio*(-0.1111) + g_Microvirga*(-0.4225).

[0079] Based on the above formula, the median of the risk score is taken as the critical value to divide the breast cancer patients into high-risk group (High-risk) and low-risk group (Low-risk). The survival analysis results show that in the training set and the validation set, the overall survival of the low-risk group is significantly better than that of the high-risk group, and the difference is statistically significant. The model can predict the survival rate of patients at different times, and the AUC curve also proves that the risk score can predict the survival rate of patients with high accuracy, further verifying the effectiveness of the model, and also emphasizing the importance of tumor microbial flora in patient prognosis evaluation.

[0080] 3. Training set validation of microbial marker prognostic risk score model:

[0081] Based on the risk score of the training set samples, the median is taken as the node to divide the high-risk group and the low-risk group, and the sample overall survival and progression free interval (PFI) data are combined to draw the KM curve, and p<0.05 is taken as the threshold to judge the significance of the difference between the high-risk group and the low-risk group. For example Figure 6As shown in FIG. 2A, the survival curves of the two groups were significantly different (p<0.05), as shown in FIG. 2B, the KM curves based on the progression-free interval data were also significantly different. Figure 6

[0082] Further, the sample risk score was used as the model prediction result, and the AUC value of the model was calculated in combination with the survival data, and then the ROC curve was drawn, as shown in FIG. 2C. Figure 6 As shown in FIG. 2C, the AUC values at 1 year, 3 years and 5 years were all greater than 0.6, indicating that the model had good performance.

[0083] 4. Verify the performance of the microbial marker prognostic risk score model in the verification set:

[0084] Further, the model performance was verified in the verification set. The sample risk score was calculated according to the formula, the sample risk score was used as the node to divide the high-risk group and the low-risk group based on the median, and the KM curve was drawn in combination with the overall survival data, as shown in FIG. 3A. Figure 7 As shown in FIG. 3A, the survival curves of the two groups were significantly different (p<0.05), as shown in FIG. 3B, the KM curves based on the progression-free interval data were also significantly different. Figure 7

[0085] Further, the sample risk score was used as the model prediction result, and the AUC value of the model was calculated in combination with the survival data, and then the ROC curve was drawn, as shown in FIG. 3C. Figure 7 As shown in FIG. 3C, the AUC values at 1 year, 3 years and 5 years were all greater than 0.6, indicating that the model had good performance.

[0086] 5. Construct a nomogram:

[0087] The Risk score, age, gender, stage data in the training set and the verification set data were used to draw a nomogram, and the ROC curve based on the nomogram was drawn. The nomogram can visually present the results of Cox regression. It formulates a scoring standard according to the size of all independent variable regression coefficients, gives each independent variable a score for each value level, and calculates a total score for each patient. The probability of outcome time occurrence of each patient is calculated through the conversion function between the score and the outcome occurrence probability. The R package rms (v6.1-0) and survival (v3.2-7) were used to draw the nomogram. First, the Cox proportional hazards regression model was constructed by combining the age, gender and stage data using the cph() function, then the survival probability of the patient was calculated using the survival() function, and finally the nomogram object was constructed using the nomogram() function, and the nomogram was drawn using the plot() function, as shown in FIG. 4A to FIG. 4D. Figure 8 As shown in FIG. 4A to FIG. 4D.

[0088] ​​In addition, the C-index of the training set and the validation set data was calculated, and a statistical chart was drawn. The C-index of the nomogram and the Risk score in the training set and the validation set was calculated and statistically analyzed, respectively, as shown in Table 2. Figure 8 As shown in Table 2, the C-index of both types of models was greater than 0.5, indicating positive predictive power.

[0089] Example 4. Clinical relevance study of the microbial marker prognostic risk score model

[0090] 1. Verification of the correlation between the prognostic risk score model and patient prognosis and clinical characteristics

[0091] The training set and the validation set data were combined, and the differences in the risk scores in the age, gender, pathologic_M, pathologic_N, pathologic_T, and stage groups were statistically analyzed in the total data using the wilcox.test function and the kruskal.test function, and a box plot was drawn, as shown in FIG. 6. Figure 9 .

[0092] 2. Correlation between characteristic microbial populations and gene expression

[0093] (1) Correlation between gene expression and microbial species abundance included in the characteristic model

[0094] The expression matrix of the sample protein-coding genes was obtained, and then the correlation between the protein-coding gene expression and the abundance of the characteristic microorganisms in the model was calculated. According to the threshold of |correlation|>0.3 and p<0.01, a correlation heat map was drawn. In this embodiment, a total of 19597 protein-coding genes were obtained, and based on the correlation threshold of p<0.01, 792 combinations with significant correlation were obtained, including 587 protein-coding genes. The correlation heat map is shown in FIG. 7. Figure 10 As shown in FIG. 7, the proportion of g_Campylobacter, g_Candidatus_Nitrosopelagicus, and g_Desulfotalea that were positively correlated with the related genes was higher, while the proportion of g_Sclerodama virus, g_Aggregatibacter, g_Saccharomonospora, and g_Pyrinomonas that were negatively correlated with the related genes was higher.

[0095] (2) Gene function enrichment analysis

[0096] Further, the R package clusterProfiler (v4.6.2) was used to perform functional enrichment analysis on the 587 protein-coding genes that were significantly correlated with the microbial markers in the model after filtering. As shown in FIG. 8, the top 20 enriched functions are shown in Table 3. Figure 11As shown, GO enrichment analysis is divided into 3 parts, namely biological process, cell component, molecular function, for genes, the biological process enrichment pathways are mainly meiotic nuclear division, meiotic cell cycle process, regulation of vasculature development, etc., the cell component enrichment pathways are mainly nucleosome, nuclear chromosome, protein-DNA complex, etc., the molecular function enrichment pathways are mainly structural constituent of chromatin, integrin binding, growth factor binding, etc., the KEGG enrichment pathways are mainly Malaria, Tyrosine metabolism, Fatty acid degradation, etc.

[0097] 3. Immune cells, immune functions, and immunotherapy prospects related to prognosis risk score models

[0098] (1) Differences in immune cell infiltration and immune function: use R package GSVA (v1.34.0) to calculate the enrichment scores of 28 immune infiltrating cells in cancer samples, and then use the scale() function to standardize the data after obtaining the results, and draw box plots of high and low risk groups. At the same time, calculate the immune function enrichment scores of the samples, and then use the scale() function to standardize the data after obtaining the results, and draw box plots. As shown in Figure 12 , there are differences in the proportions of 9 immune infiltrating cells in the high and low risk groups of microorganisms. The proportions of Activated CD4 T cell, Cd56 bright nk cell, and Memory B cell in the high risk group of microorganisms were significantly higher than those in the low risk group of microorganisms, and the remaining Activated B cell, Immature dendritic cell, MDSC, Natural killer cell, and Plasmacytoid dendritic cell were significantly higher in the low risk group of microorganisms than in the high risk group of microorganisms.

[0099] Further statistical analysis of the differences in 13 immune function-related scores between the high and low risk groups of microorganisms showed that Figure 13 , only 2 immune function-related scores had significant differences between the high and low risk groups of microorganisms, including Cytolytic activity and Parainflammation, both of which had higher enrichment scores in the low risk group of microorganisms.

[0100] (2) Immune checkpoint expression difference: obtain immune checkpoint information from literature (PMID: 32814346), count the expression difference of immune checkpoints between groups, and draw a box plot. As shown in Figure 14 Figure 5B, for 64 immune checkpoints, 35 have significant differences between the high and low risk groups of microorganisms, of which 15 are significantly highly expressed in the high risk group, and the remaining 20 are significantly highly expressed in the low risk group of microorganisms, further suggesting that the immune therapy of the low risk group may be better.

[0101] Example 5 External data verification of the prognosis risk score model

[0102] In order to further evaluate the prediction performance of the prognosis risk score model, this embodiment obtains two external verification patient data sets (PRJNA759366 and PRJNA926328) from the NCBI-SRA database, and performs ROC analysis using the selected microorganisms in the model according to the prognosis risk score model obtained in the foregoing embodiment. As shown in Figure 15 Figure 5C, the AUC of the external verification set PRJNA759366 is 0.76, and the AUC of the external verification set PRJNA926328 is 0.683, and the results of external verification show that the prognosis risk score model constructed by 22 microbial markers has good prediction performance.

[0103] Example 6

[0104] This embodiment provides a system for predicting the prognosis of breast cancer, the system comprising:

[0105] a data collection module for collecting tumor samples of breast cancer patients and determining the abundance of each microorganism in the tumor sample, the microorganism being Dicipivirus, Tymovirus, Sclerodarnavirus, Bafinivirus, Roseiflexus, Lachnoclostridium, Aggregatibacter, Ornithobacterium, Psychroflexus, Pyrinomonas, Campylobacter, Elizabethkingia, Desulfotalea, Candidatus_Nitrosopelagicus, Phaeobacter, Planktothricoides, Alteromonas, Saccharomonospora, Rothia, Agreia, Vibrio, Microvirga;

[0106] a model calculation module for calculating a risk score based on the abundance data of the microorganism according to the risk score model as described in the third aspect;

[0107] an output prediction module configured to predict a prognosis of the breast cancer patient according to the risk score data of the breast cancer patient, wherein the lower the risk score of the patient is, the better the prognosis is.

[0108] The above merely describes the preferred embodiments of the present application, and it should be noted that those skilled in the art can make several improvements and modifications without departing from the technical principles of the present application, and these improvements and modifications should also be considered as falling within the protection scope of the present application.

Claims

1. A combination of biomarkers for predicting breast cancer prognosis, characterized in that, The microbial marker combination includes: Dicipivirus, Tymovirus, Sclerodarnavirus, Bafinivirus, Roseiflexus, Lachnoclostridium, Aggregatibacter, Ornithobacterium, Psychroflexus, Pyrinomonas, Campy lobacter, Elizabethkingia, Desulfotalea, Candidatus_Nitrosopelagicus, Phaeobacter, Planktothricoides, Alteromonas, Saccharomonospora, Rothia, Agreia, Vibrio, Microvirga.

2. The biomarker combination for predicting breast cancer prognosis according to claim 1, characterized in that, The microbial marker combination consists of Dicipivirus, Tymovirus, Sclerodarnavirus, Bafinivirus, Roseiflexus, Lachnoclostridium, Aggregatibacter, Ornithobacterium, Psychroflexus, Pyrinomonas, Campylo bacter, Elizabethkingia, Desulfotalea, Candidatus_Nitrosopelagicus, Phaeobacter, Planktothricoides, Alteromonas, Saccharomonospora, Rothia, Agreia, Vibrio, Microvirga.

3. The use of reagents for detecting the abundance of each microorganism in the microbial biomarker combination as described in claim 1 or 2 in the preparation of a kit for predicting the prognosis of breast cancer.

4. The application according to claim 3, characterized in that, The microorganisms were derived from tumor samples from breast cancer patients.

5. A method for predicting the prognosis of breast cancer, characterized in that, include: Tumor samples were collected from breast cancer patients, the abundance of microorganisms in the tumor samples was measured, and the risk score of breast cancer patients was calculated based on the abundance data and risk scoring model to predict the prognosis of breast cancer. The microorganisms are Dicipivirus, Tymovirus, Sclerodarnavirus, Bafinivirus, Roseiflexus, Lachnoclostridium, Aggregatibacter, Ornithobacterium, Psychroflexus, Pyrinomonas, Campylob acter, Elizabethkingia, Desulfotalea, Candidatus_Nitrosopelagicus, Phaeobacter, Planktothricoides, Alteromonas, Saccharomonospora, Rothia, Agreia, Vibrio, Microvirga; The model is Risk Score = g_Dicipivirus*0.0441+g_Tymovirus*(-0.1151)+g_Sclerodarnavirus*(-0.0648)+g_Bafinivirus*0.1060+g_Bafinivirus*0.0816+g_Lachnoclostridium *0.1777+g_Aggregatibacter*0.1354+g_Ornithobacterium*0.0732+g_Psychroflexus*(-0.1524)+g_Pyrinomonas*0.0059+g_Campylobacter*(-0.0665)+g_ Elizabethkingia*(-0.2616)+g_Desulfotalea*(-0.0712)+g_Candidatus_Nitrosopelagicus*0.0639+g_Phaeobacter*(-0.0191)+g_Planktothricoides*0 .0492+g_Alteromonas*0.0857+g_Saccharomonospora*(-0.0363)+g_Rothia*0.3957+g_Agreia*(-0.0242)+g_Vibrio*(-0.1111)+g_Microvirga*(-0.4225); Here, Risk Score represents the risk score, and g_microbial Latin name represents the abundance of that genus of microorganism in the patient sample.

6. The risk scoring model for predicting breast cancer prognosis according to claim 5, characterized in that, The median risk score was used as the dividing line. When the risk score of a breast cancer patient was greater than the median, the patient was judged to be in the high-risk group. When the risk score of a breast cancer patient was less than the median, the patient was judged to be in the low-risk group. The overall survival of the low-risk group was significantly better than that of the high-risk group. In the low-risk group, the lower the risk score calculated based on the model, the better the prognosis for breast cancer patients.

7. A system for predicting the prognosis of breast cancer, characterized in that, The system includes: The data collection module is used to collect tumor samples from breast cancer patients and determine the abundance of various microorganisms in the tumor samples. The microorganisms are Dicipivirus, Tymovirus, Sclerodarnavirus, Bafinivirus, Roseiflexus, Lachnoclostridium, Aggregatibacter, Ornithobacterium, Psychroflexus, Pyrinomonas, Campylobacter, Elizabethkingia, Desulfotalea, Candidatus_Nitrosopelagicus, Phaeobacter, Planktothricoides, Alteromonas, Saccharomonospora, Rothia, Agreia, Vibrio, and Microvirga. The model calculation module is used to calculate a risk score based on the abundance data of microorganisms, according to the risk scoring model as described in claim 5. The output prediction module is used to predict the prognosis of breast cancer patients based on their risk score data. The lower the patient's risk score, the better the prognosis.

8. A kit for predicting the prognosis of breast cancer patients, characterized in that, Includes reagents for detecting the abundance of each microorganism in the microbial biomarker combination as described in claim 1 or 2.