A method for predicting intratumoral microbial-derived immunogenic peptides and applications thereof

By performing microbial diversity analysis and HLA-TCR motif correlation analysis on tumor RNA-seq data, immunogenic peptides of intratumoral microorganisms were predicted, which solved the problem that the immune infiltration mechanism in the tumor microenvironment was not fully explored, provided new tumor therapeutic targets, and improved treatment efficacy.

CN116246708BActive Publication Date: 2026-02-03EIGHTH AFFILIATED HOSPITAL SUN YAT SEN UNIV (SHENZHEN FUTIAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310162954.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-17
Publication Date
2026-02-03
Estimated Expiration
2043-02-17

AI Technical Summary

Technical Problem

In current tumor research, the mechanisms of immune infiltration in the tumor microenvironment have not been fully explored, resulting in a lack of effective targets and immunotherapy strategies for tumor treatment. Furthermore, the utilization rate of existing RNA-seq data is low, making it difficult to predict immunogenic peptides of intratumoral microorganisms.

Method used

By performing quality control and microbial diversity analysis on tumor RNA-seq data, differentially expressed microorganisms were screened. Combined with HLA typing and TCR CDR3 motif analysis, immunogenic peptides derived from intratumoral microorganisms were predicted. A "HLA-TCR motif-microorganism" triple relationship was constructed to screen out immunogenic peptides.

Benefits of technology

This approach enables in-depth analysis of tumor microbes and immune responses, providing new tumor-targeting and anti-cancer treatment strategies, improving the survival rate and prognosis of cancer patients, and solving the problem of low utilization of RNA-seq data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246708B_ABST
    Figure CN116246708B_ABST
Patent Text Reader

Abstract

The application discloses a method for predicting immunogenic peptide segments of tumor intratumoral microorganisms and application thereof, and mainly comprises the following steps: collecting tumor tissue RNA-seq data, screening microorganism markers related to tumor immunity, and predicting immunogenic peptide segments in microorganisms. The application first proposes a method for predicting immunogenic peptide segments in tumor microorganisms based on sufficient mining of tumor RNA-seq data, so that new targets for tumor immunotherapy are provided, and the possibility of improving the survival rate and prognosis effect of tumor patients is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and particularly relates to a method for predicting tumor intratumoral microbial-derived immunogenic peptide segments and application thereof. BACKGROUND

[0002] At present, malignant tumor has become one of the most serious diseases endangering human health. The global cancer statistics report released by the International Agency for Research on Cancer shows that there were about 1930 million new tumor cases and nearly 1000 million tumor patients died in the world in 2020. In 183 countries in the world, cancer is the first or second leading cause of death in people under the age of 70, and in other 23 countries, cancer ranks third or fourth, so each country is facing a heavy tumor burden. With the aggravation of population aging and natural environment extreme, it is estimated that the tumor burden in the world will increase by 50% in 2040 compared with that in 2020, among which the growth rate of countries with low Human Development Index (HDI) is 95%, which is the most significant group among the four HDI levels, and the growth rate of countries with very high HDI is 63%. Although China has entered the ranks of countries with high HDI, the situation of tumor prevention and control is still severe. Nowadays, the prevention and treatment of tumor has become a social problem that every family pays close attention to. In order to cope with the increasingly severe challenge of malignant tumor, the investment in tumor research has been increased worldwide to establish a standardized tumor treatment system and improve the efficacy of tumor treatment. Although the tumor research has developed rapidly, the tumor has not been conquered yet, and its pathogenesis, molecular mechanism and evolution process still need to be revealed.

[0003] At present, the main method for treating tumor in clinic mainly depends on early surgical resection and late radiotherapy and chemotherapy, but still faces a series of shortcomings: surgical resection of tumor is not complete and needs to face tumor metastasis and recurrence, and radiotherapy and chemotherapy not only kill tumor cells but also damage the body's immune system and cause gene mutation.

[0004] For a long time, the mechanism of immune infiltration in the tumor microenvironment is a key and core direction in tumor research, and is the key to improving the response rate and developing new immunotherapy strategies in tumor treatment. The key to tumor immunotherapy is to specifically clear tumor lesions, inhibit tumor growth, and break immune tolerance by activating immune cells in the body and enhancing anti-tumor immune responses. Tumor immunotherapy has become the fourth largest tumor treatment method. Tumor immunotherapy can complement conventional treatment strategies such as surgery, radiotherapy and chemotherapy, and is characterized by small side effects, significant treatment effects, and can effectively prevent tumor metastasis and recurrence. Due to the hypoxic, low-pH, high-osmotic-pressure microenvironment of tumors, some microorganisms can specifically target tumor necrosis zones and grow and reproduce, affecting the occurrence and development of tumors by recruiting immune cells and activating the immune system. For example, Salmonella typhi activates the NLRP3 inflammasome pathway through interaction with macrophages; Salmonella enterica infection can induce the release of large amounts of TNF-α, which helps the bacteria to colonize in the necrotic tumor tissue due to hypoxia, thereby helping to induce a strong immune response to inhibit tumor growth; Escherichia coli, as a facultative anaerobe, can induce the production of CD8+ T cells to inhibit the growth of tumor cells in colon cancer mice. Overall, tumor microorganisms are closely related to immune responses, making them potential targets for improving immunotherapy. Therefore, exploring the correlation between tumor microorganisms and immune responses and constructing the relationship chain of tumor-immune-microorganism will be beneficial to the development of new anticancer drugs or the development of new immunotherapy strategies and their application in clinical practice. SUMMARY

[0005] The present application aims to at least solve one of the above-mentioned technical problems in the prior art. To this end, the present application provides a method for predicting tumor intratumoral microbial-derived immunogenic peptides and its application. The prediction method in the present application can solve the problem of insufficient mining of transcriptome data in tumor sample RNA-seq data at the present stage, and further predict tumor intratumoral microbial immunogenic peptides, so as to explore the relationship between tumor intratumoral microorganisms and immune responses, predict the immunogenic peptides secreted by unexplored tumor intratumoral microorganisms, and realize the research on tumor targeting, carcinogenic and anticancer characteristics, thereby providing new ideas and approaches for tumor treatment.

[0006] In a first aspect of the present application, a method for predicting tumor intratumoral microbial-derived immunogenic peptides is provided, comprising the following steps:

[0007] (1) Collecting tumor sample RNA-seq data;

[0008] (2) Removing human host sequences in the RNA-seq data, performing microbial diversity analysis, and finding differential microorganisms;

[0009] (3) Differential expression analysis of human host genes in RNA-seq data, and enrichment analysis and functionalized aggregate analysis of differentially expressed genes;

[0010] (4) Analysis of the correlation between the expression of differentially expressed genes and immune infiltrating cells, and enrichment analysis of cell types in RNA-seq data;

[0011] (5) The enrichment modules obtained after enrichment analysis are sorted into a gene expression matrix, principal component analysis is performed on the gene expression matrix, the correlation between principal component 1 (PC1) and the relative abundance of microorganisms is calculated, and the correlation between the relative abundance of microorganisms and the immune cell infiltration score is calculated, and finally the specific microorganisms related to both the enrichment module and the immune are selected;

[0012] (6) Screening of transcripts in the sample, translation into proteins, and construction of a sample-specific microbial protein library; HLA typing of RNA-seq data, calculation of the occurrence frequency of TCR CDR3 motifs in samples with HLA alleles, Spearman correlation analysis of TCR CDR3 motifs and the abundance level of different microorganisms, and obtaining the "HLA-TCR motif-microorganism" triad relationship; screening of peptides with "HLA-TCR motif-microorganism" triad relationship for HLA binding affinity, and obtaining immunogenic peptides derived from intratumoral microorganisms of the tumor.

[0013] In some embodiments of the present application, the microorganisms include bacteria, viruses, fungi, protists, and algae.

[0014] In some embodiments of the present application, the microorganisms are bacteria, viruses, and fungi.

[0015] In some embodiments of the present application, the tumor sample RNA-seq data includes RNA-seq data from databases and clinical samples.

[0016] In some embodiments of the present application, the RNA-seq data from clinical samples is RNA-seq data obtained after sequencing of clinical tumor tissue samples.

[0017] In some embodiments of the present application, step (2) further includes, before removing human host sequences: quality control of RNA-seq data, and removal of artificial additives and low-quality sequences.

[0018] In some embodiments of the present application, the low-quality sequences refer to sequences with a Quality Phred score cutoff of ≤20.

[0019] In some embodiments of the present application, the artificial addition includes a primer, a linker or other artificial sequence added in the sequence.

[0020] In some embodiments of the present application, the filtering condition for the quality control is a default parameter.

[0021] In some embodiments of the present application, the microbial diversity analysis in step (2) is based on Kraken2 and Vegan in R package, and the obtained data is formatted and normalized.

[0022] In some embodiments of the present application, the normalization value is 1000000.

[0023] In some embodiments of the present application, step (2) further includes noise reduction processing (interference exclusion) of microbial contaminants.

[0024] In some embodiments of the present application, the microbial diversity analysis in step (2) adopts Alpha diversity analysis (including species richness and Shannon index) and Beta diversity analysis (including principal coordinate (PCoA) analysis based on Bray-Curtis distance).

[0025] In some embodiments of the present application, the criterion for judging the differentially expressed genes in step (3) is qvalue < 0.01 and |log2FC| > 1.5, where FC represents the ratio of expression levels between two groups.

[0026] In some embodiments of the present application, step (3) further includes, before the differential expression analysis, converting the human host genes (host transcriptome sequences) in the RNA-seq data into a gene expression profile matrix, and normalizing the expression profile by RPKM (Reads Per Kilo Bases per Million reads).

[0027] In some embodiments of the present application, the differential expression analysis in step (3) adopts Wilcoxon rank-sum test method.

[0028] In some embodiments of the present application, in step (3), the q value is obtained by Benjamini-Hochberg method for multiple test correction of p value.

[0029] In some embodiments of the present application, the enrichment analysis and functional aggregate analysis in step (3) include KEGG enrichment analysis, GO function cluster analysis and GSEA gene set enrichment analysis.

[0030] In some embodiments of the present application, step (4) further comprises commonality analysis of the assembled TCR sequences in the data, and alignment with T / B cell receptor database, and statistics of the length distribution of TCR beta chain CDR3 motif and the degree of TCR clonal expansion.

[0031] In some embodiments of the present application, the assembly of TCR sequences is achieved based on the de novo assembly of TCR CDR3 sequences V, J and C genes by TRUST4 software.

[0032] In some embodiments of the present application, the T / B cell receptor database comprises human immune database IMGT and T / B cell receptor database of GRCh38.

[0033] In some embodiments of the present application, the length distribution of TCR beta chain CDR3 motif and the degree of TCR clonal expansion are statistically analyzed by Immunarch in R package.

[0034] In some embodiments of the present application, the correlation of the expression of the differentially expressed genes with immune infiltrating cells in step (4) is analyzed by xCell algorithm.

[0035] In some embodiments of the present application, the ratio of the transcript to the corresponding microbial transcriptome data in step (6) is greater than 80%.

[0036] In some embodiments of the present application, the alignment of the transcript with the corresponding microbial transcriptome data in step (6) is based on Bowtie2.

[0037] In some embodiments of the present application, the translation of the transcript in step (6) is performed by EMBOSS tool getorf.

[0038] In some embodiments of the present application, the frequency of occurrence of the selected HLA allele is greater than or equal to 7.

[0039] In some embodiments of the present application, the HLA binding affinity screening standard is %RANK≤0.2.

[0040] The second aspect of the present application provides the use of the prediction method of the first aspect of the present application in the preparation of tumor cell intracellular microbial source antigens and / or antibodies.

[0041] In some embodiments of the present application, the tumor comprises but is not limited to breast cancer.

[0042] The present application has the following beneficial effects:

[0043] 1.The method for predicting immunogenic peptide segments in tumor microorganisms based on sufficient mining of tumor RNA-seq data is proposed for the first time, so as to provide new targets for tumor immunotherapy and improve the survival rate and prognosis of tumor patients.

[0044] 2.The method for predicting immunogenic peptide segments in tumor microorganisms based on sufficient mining of tumor RNA-seq data is proposed for the first time, so as to provide new targets for tumor immunotherapy and improve the survival rate and prognosis of tumor patients. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 The schematic diagram of the prediction method in the present application is shown.

[0046] Figure 2 The columnar chart of the correlation analysis between intratumoral bacteria and host genes in breast cancer (ERBC) is shown, wherein + in the heat map represents p<0.05, and * represents p<0.01.

[0047] Figure 3 The columnar chart of the correlation analysis between intratumoral bacteria and host genes in breast cancer (TNBC) is shown, wherein + in the heat map represents p<0.05, and * represents p<0.01.

[0048] Figure 4 The columnar chart of the correlation analysis between intratumoral bacteria and host genes in breast cancer (TNBC) is shown, wherein + in the heat map represents p<0.05, and * represents p<0.01. DETAILED DESCRIPTION

[0049] The content of the present application is further described in detail through specific examples. The raw materials, reagents or devices used in the examples and comparative examples can be obtained from conventional commercial channels, or can be obtained by the prior art method, unless otherwise specified. Unless otherwise specified, the test or test method is a conventional method in the art.

[0050] In the present application, the method for predicting immunogenic peptide segments in tumor microorganisms mainly includes the following steps: collecting tumor tissue RNA-seq data, screening tumor immune-related microbial markers, and predicting immunogenic peptide segments in microorganisms (the schematic diagram is shown in Figure 1

[0051] Specifically, the steps are as follows:

[0052] ​(1) Sample data collection:

[0053] Raw RNA-seq sequencing data and clinical information tables of tumor samples can be obtained from public high-throughput sequencing databases, or tumor tissue samples can be collected clinically and sent to sequencing companies to obtain raw RNA-seq sequencing data of tumor tissue samples.

[0054] (2) Intratumoral microbial analysis:

[0055] The obtained sequencing data were quality controlled using Trim Galore with default filtering parameters. This removed artificially added primers, adapters, and low-quality sequences from the sequencing and library construction processes, resulting in high-quality clean data. The clean data was then aligned to the human reference genome GRCh38 using Hisat2 software to remove host sequences for microbial classification.

[0056] The Kraken2 (κ-mer based) algorithm was used to annotate and classify microbial species in the data after removing host sequences, resulting in a species taxonomic count table. The microbial reference database used for classification was constructed by Kraken2 from RefSeq (https: / / www.ncbi.nlm.nih.gov / refseq / ). Then, Bracken was used to classify the species taxonomic count table based on Bayesian probability, and its more accurate taxonomic abundance at the phylum, class, order, family, genus, and species levels was re-estimated. To remove interference from potential microbial contaminants, the obtained microbial taxonomic table was compared with known common laboratory contaminants. For matching components, a literature search was conducted; if the microorganism had been reported in relevant tumor microbiology research literature, it was excluded from the "potential microbial contaminant" cohort.

[0057] Alpha diversity analysis (including species richness and Shannon index) and Beta diversity analysis (including principal coordinate analysis based on Bray-Curtis distance (PCoA)) were performed on the contaminant-excluded data using the Vegan package in R. Species taxonomic abundance tables were merged, along with sample names and grouping information, as input to the LEfSe software (Linux, version 1.0).

[0058] The obtained data was formatted, and considering the extremely low values ​​in the microbial classification data, the normalization value was set to 1,000,000. The formatted data was then analyzed using the default parameters, and calculations revealed microorganisms with significant biological and statistical differences. Finally, a bar chart was used to visualize the differentially expressed microorganisms.

[0059] (3) Host gene expression analysis:

[0060] The sequences that successfully aligned with GRCh38 (i.e., the host transcriptome sequences) were extracted. Based on the gene annotation file (Homo_sapiens.GRCh38.105), the number of reads aligned to each gene was counted using HTSeq to obtain the gene expression profile matrix. The expression profile was standardized using RPKM (Reads Per Kilo Bases per Million reads), and differential gene expression analysis was performed using the Wilcoxon rank-sum test. Genes meeting the criteria of qvalue < 0.01 and |log2FC| > 1.5 (FC stands for fold change, representing the ratio of expression levels between two samples (groups)) were selected as differentially expressed genes. To control the false discovery rate (FDR), the p-value was corrected using the Benjamini-Hochberg method after multiple testing to obtain the q-value. KEGG enrichment analysis, GO functional clustering analysis, and gene set enrichment analysis (GSEA) were performed on the selected differentially expressed genes using ClusterProfiler in the R package.

[0061] (4) Tumor immune response analysis:

[0062] CDR3 is the most important region of the TCR that defines antigen specificity. It contains the contact sites of antigenic peptide residues and is unique to each T cell clone. The unique combination of CDR3α and CDR3β sequences further defines the T cell clonal type. Through allelic exclusion mechanisms, most circulating lymphocytes express a single allelic copy of TCRβ, which is the fundamental principle behind using TCRβ as a molecular barcode to track antigens. Therefore, the CDR3β chain (CDR3) can represent a unique antigen-specific biological barcode for T cells in biological sample analysis. The V, J, and C genes of the TCR CDR3 sequence (from a TCR immune repertoire) were de novo assembled using TRUST4 software to report TCR sequence commonality. Contigs were aligned with the human immune database IMGT and the GRCh38 T / B cell receptor database to annotate gene information. Finally, Immunarch in the R package was used to statistically analyze the length distribution of the TCRβ chain (CDR3) and the degree of TCR clonal amplification in the immune repertoire. The xCell algorithm (http: / / xcell.ucsf.edu / ) was used to convert gene expression profiles into richness fractions for 64 immune and stromal cell types. Tumor immune infiltration data were extracted from the gene expression profiles, and cell type enrichment analysis was performed on the 64 immune and stromal cell types. The Wilcoxon rank-sum test was used to compare the proportion of immune cell types in tumor tissues and normal tissues.

[0063] (5) Correlation analysis between differentially expressed microorganisms and host genes:

[0064] After performing enrichment analyses (including KEGG, GO, and GESA), enrichment modules related to antimicrobial immune responses and tumorigenesis and development were selected based on the studied cancer type. First, a gene expression matrix was compiled based on the genes of each enrichment module (the previously selected KEGG, GO, and GESA enrichment modules), and "principal component 1" (PC1) was calculated to represent the gene expression profile of that module. The bacterial species classification tables for each sample output by Braken were summarized to obtain a bacterial abundance table for all samples, from which the relative abundance matrix of tumor-enriched bacteria was selected. The Pearson correlation coefficient (PCC) between PC1 and the relative bacterial abundance was calculated using prcomp in R.

[0065] The "principal components" are calculated using principal component analysis, which utilizes orthogonal transformations to linearly transform the observed values ​​of a series of potentially related variables, projecting them into a series of linearly unrelated variables called principal components. In this embodiment, principal component 1 (PC1) is calculated by moving the center of the coordinate axis to the center of the data, and then rotating the coordinate axis to maximize the variance of the data on the C1 axis, meaning that the projections of all n data points in this direction are most dispersed. This implies that more information is retained. C1 becomes the first principal component.

[0066] Then, the correlation between differentially enriched microorganisms and immune infiltration was analyzed: differentially infiltrating immune cells were selected, a relative abundance matrix of differentially enriched bacteria was prepared, and the PCC between the relative abundance of tumor-enriched bacteria and the immune cell infiltration fraction was calculated using prcomp on R.

[0067] (6) Predicting microbial immunogenic peptides:

[0068] We predict microbial immunogenic peptides by using the "HLA-TCR motif-microbe" triple relationship, and further screen them by the affinity between microbial peptides and HLA molecules.

[0069] Taking bacteria as an example. Since there are differences between microbial proteins in different samples and bacterial protein reference databases, it is first necessary to construct a sample-specific bacterial protein database. In the absence of a reference genome, transcripts can be assembled de novo using Trinity software. To ensure transcript quality, the obtained transcripts are compared with bacterial transcriptome data using Bowtie2. When the alignment rate is greater than 80%, it indicates that high-quality transcripts have been obtained and can be used for downstream analysis.

[0070] The high-quality transcripts obtained were translated into proteins using the EMBOSS tool getorf, thereby obtaining a sample-specific microbial protein library.

[0071] HLA typing of tumor RNA-seq data (BAM files) was performed using HLA-HD software: the frequency of HLA alleles was calculated, and high-frequency HLA alleles were selected. After removing conserved regions at both ends using LOGO analysis of the amino acid sequence of TCR CDR3, the frequency of the TCR CDR3 motif in the sample containing each HLA allele was calculated using Immunarch in the R package. Then, Spearman correlation analysis was performed on the abundance level of TCR CDR3 motif and differentially expressed bacteria to obtain the "HLA-TCR motif-bacteria" triple relationship. Finally, netCTLpan was used to calculate the differential expression.

[0072] The binding affinity of bacterial peptides to HLA molecules is determined by further screening the peptides obtained from the above screening to identify differentially expressed bacterial peptides with strong affinity for HLA molecules. The selected differentially expressed bacterial peptides are the predicted immunogenic peptides secreted by tumor-specific bacteria.

[0073] Example 1: Prediction of immunogenic peptides derived from intratumoral bacteria based on breast cancer RNA-seq data

[0074] To provide a more intuitive understanding of the method in this invention, the inventors used breast cancer RNA-seq data as an example to demonstrate the practical effects of the method. It should be understood that the implementation of this invention is not limited to this embodiment.

[0075] The specific steps are as follows:

[0076] (1) Sample data collection:

[0077] A breast cancer RNA-seq dataset (accession number GSE58135) was downloaded from NCBI's Sequence Read Archive (SRA). This dataset includes 42 triple-negative breast cancer tumor tissues (TNBC), 42 estrogen receptor-positive breast cancer tumor tissues (ERBC), 21 adjacent breast tissues (also known as peritumoral tissues) of TNBC primary tumors, and 30 adjacent breast tissues of ERBC primary tumors.

[0078] (2) Intratumoral microbial analysis:

[0079] Diversity analysis revealed differences in the microbial community composition between tumor tissue and adjacent normal tissue. After calculating differentially expressed bacteria, 10 bacterial genera were significantly enriched in ERBC tumor tissue, including *Pasteurella*, *Neisseria*, *Enterococcus*, *Pseudomonas*, *Escherichia*, *Paracoccus*, *Halomonas*, *Streptococcus*, *Rhodococcus*, and *Shimwellia*. Five types of bacteria enriched in TNBC tumors were found, three of which were also found in ERBC: *Neisseria*, *Halomonas*, and *Escherichia*. The remaining two were *Actinomyces* and *Staphylococcus*. Nine bacterial species were significantly enriched in ERBC tumor tissue, including Pasteurella multocida, Neisseria gonorrhoeae, Staphylococcus haemolyticus, Enterococcus faecalis, Escherichia coli, Paracoccus mutanolyticus, Pseudomonas putid, Halomonas sp. JS92-SW72, and Streptococcus pneumoniae. Four bacterial species were significantly enriched in TNBC tumor tissue, three of which were also enriched in ERBC tumor tissue: Neisseriagonorrhoeae, Escherichia coli, and Halomonas sp. JS92-SW72. The remaining bacterial species was Staphylococcus aureus.

[0080] (3) Host gene expression analysis:

[0081] Analysis of host gene expression revealed 3427 upregulated genes and 2444 downregulated genes in ERBC compared to adjacent normal tissues, while 4972 upregulated genes and 2795 downregulated genes were identified in TNBC tumor tissue. To further analyze the correlation between tumor microbiota and gene enrichment functional modules, these modules were screened, and the following modules were selected for analysis.

[0082] ① ERBC functional modules:

[0083] This module covers: fatty acid degradation, fat digestion and absorption, and adipokines signaling pathways; PPAR, PI3K-Akt and AMPK signaling pathways, cGMP-PKG signaling pathway, CAMP signaling pathway, neutrophil extracellular trap, neuroactive ligand-receptor interactions and transcriptional dysregulation in cancer, positive and negative regulation of cold-induced thermogenics, antimicrobial peptide-mediated antimicrobial humoral immune responses, antimicrobial humoral responses and cellular responses to bacterial-derived molecules, inflammatory cell apoptosis, acute inflammatory responses, positive regulation of inflammatory responses, protein kinase B signaling, endocrine hormone secretion, aging, mitotic spindle assembly checkpoint signaling and alcohol metabolism, G2M checkpoint and E2F target, IL2 STAT5 signaling, glycolysis, adipogenesis and fatty acid metabolism, late estrogen response, KRAS signaling, hypoxia and epithelial-mesenchymal transition.

[0084] ② TNBC's functional modules:

[0085] PI3k-Akt, Ras, PPAR, MAPK, and AMPK signaling pathways; regulation of transcriptional errors in cancer; neutrophil extracellular trap formation; cAMP and cGMP-PKG signaling pathways; endocrine resistance; neuroactive ligand-receptor interactions; regulation of lipolysis in adipocytes and ECM-receptor interactions; acute inflammatory responses; neuroinflammatory responses and positive regulation of inflammatory responses; antimicrobial peptide-mediated antimicrobial humoral immune responses; antimicrobial humoral responses and cellular responses to bacterial-derived molecules; G protein-coupled peptide receptor activity; mitotic spindle assembly checkpoint signaling; positive regulation of cold-induced thermogenesis and responses to alcohol; G2M checkpoints and E2F targets; early and late estrogen responses; inflammatory responses; cancer; increased Kras signaling; TNFA signaling via NFKB; angiogenesis; epithelial-mesenchymal transition; hypoxia-related factors.

[0086] (4) Tumor immune response analysis:

[0087] The relative abundance of tumor immune infiltrating cells was calculated using xCELL. The relative abundance of 60 immune and basal cell subsets in each sample was obtained. After filtering with p<0.05, the proportion of immune cell subsets was analyzed.

[0088] Analysis revealed that, compared to adjacent normal tissue, ERBC tumor tissue showed a significant increase in the abundance of NK T cells, naive CD8+ T cells, Th1 CD4+ T cells, Th2 CD4+ T cells, plasma dendritic cells, myeloid dendritic cells, macrophages, common lymphoid progenitor cells, plasma cells, and naive B cells. TNBC showed a significant increase in the abundance of NK T cells, CD8+ T cells, naive CD8+ T cells, effector CD8+ T cells, central memory CD8+ T cells, Th1 CD4+ T cells, Th2 CD4+ T cells, memory CD4+ T cells, central memory CD4+ T cells, plasma dendritic cells, macrophages, common lymphoid progenitor cells, B cells, plasma cells, naive B cells, and effector B cells.

[0089] (5) Correlation analysis between differentially expressed microorganisms and host genes:

[0090] Correlation analysis was used to screen for microorganisms associated with tumor progression and immune response. After selecting 36 and 33 host gene functional modules in ERBC and TNBC, respectively, using the above steps, correlation analysis was performed on the differentially enriched bacterial species in tumors and the host gene functional modules.

[0091] The results are as follows Figure 2 , Figure 3 and Figure 4 As shown.

[0092] In ERBCs, Halomonas sp. JS92-SW72 was found to be positively correlated with multiple functional modules related to antimicrobial immune responses, suggesting that this bacterial peptide may activate tumor immune responses by activating humoral immune responses. Furthermore, Halomonas sp. JS92-SW72 was also found to be positively correlated with functional modules related to enhanced Kras signaling, breast cancer risk factors, and breast cancer invasiveness. Therefore, Halomonas sp. JS92-SW72 may play an important role in constructing the tumor immune environment of ERBCs. Pasteurella multocida, Neisseria gonorrhoeae, Staphylococcus haemolyticus, Enterococcus faecalis, Escherichia coli, Paracoccus mutanolyticus, Pseudomonasputida, and Streptococcus pneumoniae are positively correlated with functional modules related to fatty acid metabolism, cold-induced thermogenesis, cGMP-PKG signaling pathway, and AMPK signaling pathway, and are significantly negatively correlated with functional modules related to ERBC progression, such as glycolysis, cancer transcriptional dysregulation, cAMP signaling pathway, hypoxia, and estrogen-induced dysregulation.

[0093] In TNBC, a positive correlation was observed between *Escherichia coli* and antimicrobial peptide-mediated humoral immune responses, indicating that this bacterial peptide can activate humoral immune responses. *Neisseria gonorrhoeae*, *Escherichia coli*, *Halomonas* sp. JS92-SW72, and *Staphyloccocus aureu* were all positively correlated with breast cancer proliferation, including mitotic spindle assembly checkpoint signaling, AMPK signaling, G2M checkpoint, lipolysis, adipocyte regulation, and hypoxia; these functional modules are crucial for breast cancer development. Therefore, these bacteria may be involved in regulating the occurrence and development of TNBC.

[0094] like Figure 4 As shown, in ERBC, Halomonas sp. JS92-SW72 was positively correlated with plasma cells, Th1 CD4+ T cells, naive CD8+ T cells, common lymphoid progenitor cells, and macrophages (p<0.05, statistically significant), as well as Th2 CD4+ T cells (p>0.05, no statistical significance). Therefore, it can be inferred that the peptide of Halomonas sp. JS92-SW72 can be recognized, processed, and presented to Th2 CD4+ T cells as an extracellular antigen by macrophages, further activating plasma cell-mediated humoral immune responses. Paracoccus mutanolyticus, Staphylococcus haemolyticus, and Enterococcus faecalis were all positively correlated with the activation of CTLs, central memory and naive CD8+ T cells, naive, central memory and Th1 CD4+ T cells, NK T cells, and myeloid dendritic cells, but negatively correlated with macrophages. Therefore, it can be inferred that these bacteria may invade cancer cells, present intracellular antigens after treatment, and activate naive CD8+ T cells to differentiate into central memory CD8+ T cells and CTLs, thereby recognizing and eliminating infected cells. Furthermore, these bacteria are positively correlated with the IL2-STAT5 signaling pathway, AMPK and cGMP-PKG signaling pathways, cold-induced thermogenicity, adipokines signaling pathways, and the negative regulation of fatty acid metabolism. Therefore, it can be further inferred that they not only activate inflammation to kill tumor cells but also participate in regulating important metabolic processes related to alcohol, fat, and thermogenicity, thereby influencing the occurrence and development of breast tumors.

[0095] In TNBC, Escherichia coli was significantly positively correlated with B cells, naive and memory B cells, Th1 cells, memory and central memory CD4+ T cells, central memory, effector memory and naive CD8+ T cells.

[0096] (6) Predicting microbial immunogenic peptides:

[0097] By analyzing the "HLA-TCR motif-bacteria" triplet relationship of all obtained tumor samples, and screening the binding affinity of bacterial-derived peptides to HLA types, it was found that some peptides enriched in the bacteria in the tumor samples have a strong affinity for the HLA molecules of the samples.

[0098] In ERBC, HLA-A*02:01, HLA-A*03:01, HLA-A*07:02, HLA-B*40:01, HLA-B*44:02, HLA-C*07:01, and HLA-C*07:02 were selected based on their frequency ranking. The following triadic relationships were identified: HLA-A*03:01-TDTQY (SEQ ID NO:1)-Halomonas sp.JS92-SW72-RSVAVLTCK (SEQ ID NO:2), HLA-B*07:02-TDTQY-Escherichia coli-SPLLFNIVL (SEQ ID NO:3), and HLA-B*07:02-NTEAF (SEQ ID NO:4)-Escherichia coli-NPIVSAQNL (SEQ ID NO:5) / SPLLFNIVL (represented as: corresponding HLA-corresponding TCR motif-originating bacterial name-peptide amino acid sequence). It was also found that although HLA-A*02:01 was the most frequent allele, it did not actually contain a strong binding peptide derived from bacteria enriched in tumors. The same situation also occurred in HLA-B*07:02, HLA-B*40:01, HLA-B*44:02, HLA-C*07:01, and HLA-C*07:02, where these alleles also did not have positive associated bacteria or strong binding peptides.

[0099] In TNBC, HLA-A*01:01, HLA-A*03:01, HLA-B*08:01, HLA-B*07:01, HLA-C*07:01, and HLA-C*07:02 were selected as screening alleles. Among them, HLA-A*07:02 was the most frequent HLA-A allele, present in 12 of the 42 tumor samples in TNBC. Simultaneously, the top-ranking TCR CDR3 motifs, including YNEQF (SEQ ID NO:6), TDTQY, SYEQY (SEQ ID NO:7), and TGELF (SEQ ID NO:8), were selected. These motifs were positively correlated with Escherichia coli or Neisseria gonorrhoeae. The obtained triads include: HLA-A*01:01-YNEQF / TDTQY / SYEQY / TGELF-Escherichia coli-LFADDMIVY (SEQ ID NO:9) and HLA-A*01:01-YNEQF / TGELF-Neisseria gonorrhoeae-QSLQEIWDY (SEQ ID NO:10) / VSSFSVLFF (SEQ ID NO:11) / FFSPSLWFY (SEQ ID NO:12). Furthermore, the triad HLA-A*03:01-TDTQY-Halomonassp.JS92-SW72-RSVAVLTCK also appeared in ERBCs. These results indicate that the "HLA-TCR motif-bacteria" triad can activate immune responses in different samples.

[0100] In addition, 13 immunogenic peptides were found to belong to HLA-B*08:01-YNEQF-Neisseriagonorrhoeae, which have similar amino acid compositions, such as ELKAKAREL (SEQ ID NO:13), ELKTKAPEL (SEQ ID NO:14), ELKTKAREL (SEQ ID NO:15) and ELREECRSL (SEQ ID NO:16), as well as YVKRPNLHL (SEQ ID NO:17) and YVKRPNLRL (SEQ ID NO:18). Furthermore, the inventors summarized three modes: HLA-B*08:01-SYEQY-Escherichia coli-FMLKTLNKL (SEQ ID NO:19) / MIKRHVKFF (SEQ ID NO:20) / NIRKSINVI (SEQ ID NO:21), HLA-C*07:01-YNEQF / TDTQY / SYEQY / TGELF-Escherichia coli-SPLLFNIVL, and HLA-C*07:02-NTEAF-Halomonas sp.JS92-SW72-LWWRSVAVL (SEQ ID NO:22). Specific results are shown in Table 1.

[0101] Table 1. Results of HLA-TCR-bacterial pattern and affinity analysis

[0102]

[0103]

[0104]

[0105]

[0106] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A method for predicting immunogenic peptides derived from intratumoral microorganisms in tumors, comprising the following steps: (1) Collect RNA-seq data from tumor samples; (2) Remove human host sequences from RNA-seq data, perform microbial diversity analysis, and search for differentially expressed microorganisms; (3) Perform differential expression analysis on human host genes in RNA-seq data, and perform enrichment analysis and functionalized aggregate analysis on differentially expressed genes; (4) Analyze the correlation between the expression of differentially expressed genes and immune infiltrating cells, and perform enrichment analysis on cell types in RNA-seq data; (5) The enrichment modules obtained after enrichment analysis are organized into a gene expression matrix. Principal component analysis is performed on the gene expression matrix to calculate the correlation between principal component 1 and the relative abundance of microorganisms, as well as the correlation between the relative abundance of microorganisms and the immune cell infiltration fraction. Finally, specific microorganisms that are related to both the enrichment module and immunity are selected. (6) Screen transcripts in the samples, translate them into proteins, and construct a sample-specific microbial protein library; perform HLA typing on RNA-seq data, calculate the frequency of occurrence of TCR CDR3 motifs in samples containing HLA alleles, perform Spearman correlation analysis on the abundance levels of TCR CDR3 motifs and differentially expressed microorganisms, and obtain the "HLA-TCR motif-microorganism" triple relationship; screen peptides with the "HLA-TCR motif-microorganism" triple relationship for HLA binding affinity to obtain immunogenic peptides derived from tumor intratumoral microorganisms.

2. The prediction method according to claim 1, characterized in that, The microorganisms include bacteria, viruses, fungi, and protozoa.

3. The prediction method according to claim 1, characterized in that, The microorganisms mentioned are bacteria, viruses, or fungi.

4. The prediction method according to claim 1, characterized in that, Step (2) before removing the human host sequence also includes: quality control of the RNA-seq data to remove artificial additives and low-quality sequences.

5. The prediction method according to claim 1, characterized in that, The microbial diversity analysis described in step (2) is implemented based on Kraken2 and Vegan in the R package, and the obtained data is formatted and normalized.

6. The prediction method according to claim 1, characterized in that, The criteria for judging differentially expressed genes in step (3) are: q value < 0.01 and |log2FC| > 1.5, where FC represents the ratio of expression levels between the two groups.

7. The prediction method according to claim 1, characterized in that, The enrichment analysis and functionalized cluster analysis mentioned in step (3) include KEGG enrichment analysis, GO functional cluster analysis and GSEA gene set enrichment analysis.

8. The prediction method according to claim 1, characterized in that, Step (4) also includes performing a commonality analysis on the assembled TCR sequences in the data, comparing them with the T / B cell receptor database, and statistically analyzing the length distribution of the TCR β chain CDR3 motif and the degree of TCR clonal amplification.

9. The prediction method according to claim 1, characterized in that, The correlation between the expression of differentially expressed genes and immune infiltrating cells in step (4) was analyzed using the xCell algorithm.

10. The prediction method according to claim 1, characterized in that, The transcripts described in step (6) have a correlation ratio with the corresponding microbial transcriptome data of more than 80%.

11. The application of the prediction method according to any one of claims 1 to 10 in the preparation of intracellular microbial antigens and / or antibodies of tumor cells; wherein the tumor includes breast cancer.