Tumor-specific microbial epitope recognition method

By identifying microbial transcriptional reprogramming regions in the tumor microenvironment using the MiTRI index, this method overcomes the limitations of specific difference identification methods that address the challenge of identifying microbial species similarity between tumors and normal tissues. It enables precise screening of tumor-specific microbial antigenic epitopes, reduces off-target effects and the risk of autoimmune toxicity, and is applicable to various solid tumor types.

CN121725883APending Publication Date: 2026-03-24HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-12
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies for identifying tumor-specific microbial antigens face limitations such as a lack of specificity due to the high similarity between the microbial species composition of tumors and normal tissues, a lack of quantitative assessment methods for microbial transcriptional activity, and limitations in detection methods, making it difficult to accurately screen for microbial antigenic epitopes with high immunogenicity.

Method used

By quantitatively analyzing the transcriptional adaptive changes of microorganisms in the tumor microenvironment, the MiTRI index was used to identify microbial transcriptional reprogramming regions, and tumor-specific microbial antigenic epitopes were screened out. Data processing and species classification were performed using whole transcriptome sequencing data. Combined with MiTRI index calculation and HLA binding affinity prediction, highly immunogenic tumor-specific candidate antigenic epitopes were identified and screened out.

Benefits of technology

This technology enables the precise identification of tumor-specific microbial epitopes in the context of similar microbial species composition between tumors and normal tissues. This reduces off-target effects and autoimmune toxicity risks associated with immunotherapy. The selected epitopes exhibit high immunogenicity and are applicable to various solid tumor types, while also lowering detection costs and technical barriers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121725883A_ABST
    Figure CN121725883A_ABST
Patent Text Reader

Abstract

The invention discloses a tumor specific microbial antigen epitope recognition method, and belongs to the technical field of bioinformatics and tumor immunology. The identification method comprises the following steps: based on whole transcriptome sequencing data of a sample, obtaining a high-quality non-host microorganism sequence, and carrying out species classification screening; the method comprises the following steps: defining a microorganism transcription programming region and a target sample specific microorganism transcription reprogramming region, calculating to obtain a microorganism transcription reprogramming index MiTRI, obtaining microorganisms which generate high transcription reprogramming activity in a target sample, and quantitatively analyzing the transcription adaptability change (namely transcription reprogramming) of the microorganisms in a tumor microenvironment, so as to obtain the high transcription reprogramming activity of the microorganisms. Therefore, the precise recognition of the tumor-specific microbial epitope with high immunogenicity is realized. And a new method is provided for screening and identifying the microbial epitope with tumor specificity and immunogenicity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the fields of bioinformatics and tumor immunology, specifically relating to a method for identifying tumor-specific microbial antigenic epitopes based on high-throughput sequencing data, and the application of the method in tumor-specific antigenic epitope screening and drug development. Background Technology

[0002] Tumor-specific antigens (TSAs) are ideal targets for tumor immunotherapy (such as tumor vaccines and TCR-T therapy). For a long time, research in this field has primarily focused on neoantigens derived from somatic mutations in the host genome. These antigens, containing mutated sequences, are considered "altered selves," capable of evading central immune tolerance and activating specific T cells. However, many malignancies with low mutation burdens (such as microsatellite stable colorectal cancer (MSS-CRC) and pancreatic cancer) often lack endogenous neoantigens, resulting in an "immunely cold" phenotype and poor response to existing immunotherapies. This unmet clinical need has prompted researchers to search for antigen classes beyond mutation sources.

[0003] In recent years, intratumoral microbial antigens have attracted attention as a highly promising complementary target. Unlike endogenous antigens, microbial antigens originate from bacteria and other microorganisms that colonize tumor tissues, and are essentially "pure non-self" components. The antigenic epitopes generated after these microbial proteins are processed by the host cell's antigen presentation mechanism theoretically possess stronger immunogenicity than autoantigens and are independent of the host's tumor mutational burden.

[0004] Despite the enormous potential of microbial antigens, current technologies face significant technical bottlenecks in translating them into safe and effective clinical therapeutic targets. These bottlenecks are mainly reflected in the following aspects: (1) Lack of specificity due to the high similarity of microbial species composition between tumor and normal tissues. Current technologies mainly rely on metagenomics or transcriptomics analysis of the "species classification" and "relative abundance" of microorganisms. However, a large amount of research data shows that the microbial species composition between tumor tissues and adjacent normal tissues (adjacent tissues) has a very high degree of similarity. This means that if targets are screened solely based on the presence or abundance differences of species, the selected microbial antigenic epitopes are very likely to also exist in normal tissues. This high degree of species similarity masks potential antigenic differences, directly leading to serious "off-target effects" and autoimmune toxicity risks in immunotherapy.

[0005] (2) Lack of quantitative assessment methods for microbial transcriptional activity. After colonizing the tumor microenvironment (such as hypoxia, acidity, and immunosuppressive environment), microorganisms undergo significant changes in their gene transcription profile (i.e., transcriptional reprogramming) in order to adapt to survival pressures. Current bioinformatics workflows lack indicators or algorithms that can quantitatively assess the transcriptional differences of microorganisms in different tissue environments, making it impossible to distinguish which genomic regions are truly tumor-specific transcriptional regions.

[0006] (3) Limitations of detection methods. Although traditional mass spectrometry (MS)-based immunopeptidomics methods can directly detect HLA-binding peptides, they require extremely high sample volumes and have insufficient sensitivity for detecting low-abundance microbial antigenic epitopes, making it difficult to meet the detection needs of routine clinical biopsy samples.

[0007] Therefore, there is an urgent need to develop a tumor-specific microbial antigen recognition method to systematically identify and quantify microbial transcriptional reprogramming from conventional transcriptome sequencing data, and thereby screen for microbial antigenic epitopes with tumor specificity and immunogenicity. Summary of the Invention

[0008] In view of this, the primary objective of this application is to provide a method for identifying tumor-specific microbial antigenic epitopes, which accurately identifies highly immunogenic tumor-specific microbial antigenic epitopes by quantitatively analyzing the transcriptional adaptive changes (i.e., transcriptional reprogramming) of microorganisms in the tumor microenvironment.

[0009] To achieve the above objectives, this application adopts the following technical solution: One aspect of this application discloses a method for recognizing tumor-specific microbial antigen epitopes, comprising the following steps: Whole transcriptome sequencing data of the samples were obtained, and high-quality non-host microbial sequences were obtained through data processing and species classification screening was performed; the samples included target samples and background control samples. Microbial genomic regions covered by transcriptome sequencing in at least one of the target sample or background control sample are identified as microbial transcriptional reprogramming regions. Microbial genomic regions covered only by transcriptome sequencing in the target sample but not in the background control sample are identified as target sample-specific microbial transcriptional reprogramming regions. The microbial transcriptional reprogramming index MiTRI is calculated according to formula (1) to obtain microorganisms with high transcriptional reprogramming activity in the target sample:

[0010] In formula (1), Bases in the target sample-specific microbial transcriptional reprogramming region Coverage depth; This represents the total number of bases in the transcriptional reprogramming region; This represents the total number of microbial reads detected in the sample. This represents the total number of reads in the sequencing library. The total length of the microbial transcription programming region; Based on microorganisms exhibiting high transcriptional reprogramming activity in target samples, tumor-specific candidate antigen epitopes are identified and screened.

[0011] Another aspect of this application discloses a system / device for recognizing tumor-specific microbial antigenic epitopes, which implements the recognition method described in this application during operation; The identification system / device includes: The data input module is configured to receive whole transcriptome sequencing data from the sample; The sequence processing and transcription analysis module is configured to perform data processing, obtain high-quality non-host microbial sequences, and perform species classification and screening. The region definition module is configured to define microbial transcriptional programming regions and target sample-specific transcriptional reprogramming regions based on differences in whole transcriptome sequencing coverage. The MiTRI calculation module is configured to calculate the microbial transcriptional reprogramming index according to the formula defined in this application; The epitope screening module is configured to translate and predict the affinity of sequences in tumor-specific transcriptional reprogramming regions, and output tumor-specific candidate antigen epitopes.

[0012] Another aspect of this application discloses tumor-specific microbial antigen peptides, which are identified by the identification method described in this application.

[0013] Another aspect of this application discloses a system / device for assisting in the evaluation of the efficacy of tumor treatment, said system / device comprising the following modules: The calculation module is configured to calculate the MiTRI value of the microbial transcriptional reprogramming index in the tumor tissue to be tested according to the formula (1) defined in this application; The data analysis module is configured to compare the MiTRI value with a preset reference value; The results analysis and output module is configured to output an evaluation result of the efficacy of tumor treatment based on the results of the data analysis module. If the MiTRI value is higher than the preset reference value, it indicates that the tumor tissue under test has a high response to the treatment. The treatment referred to herein is neoadjuvant therapy or immunotherapy.

[0014] The beneficial effects of this application are: This application discloses a method for identifying tumor-specific microbial antigenic epitopes and its application. This method can accurately identify truly tumor-specific microbial antigenic epitopes, providing support for the discovery and identification of tumor-specific microbial antigens, the screening of targets for tumor immunotherapy, the development of biomarkers for tumor microbiome characteristics, and the auxiliary formulation of personalized immunotherapy regimens. The specific advantages of this identification method are as follows: (1) An innovative quantitative assessment system for microbial transcriptional reprogramming has been established. The MiTRI index proposed in this application fills the gap in the quantitative assessment of microbial transcriptional activity and successfully achieves accurate identification and quantification of differences in functional transcription of microorganisms in tumor and normal tissues against the backdrop of highly similar microbial species composition.

[0015] (2) This method solves the problem of specificity and safety in microbial antigen screening. By precisely locating tumor-specific transcriptional reprogramming (MTR) regions, this method can effectively eliminate non-specific sequences expressed in normal tissues and identify truly tumor-specific microbial antigen epitopes, thereby significantly reducing the potential off-target effects and autoimmune toxicity risks of immunotherapy.

[0016] (3) The selected antigenic epitopes have been verified to have high immunogenicity. Taking colorectal cancer in the examples as an example, the microbial antigenic epitopes (such as the 21 peptides in the examples) screened by this method and verified by experiments have been shown to induce clonal expansion of patient T cells and specific immune responses, showing good potential for development into tumor vaccines or TCR-T therapies.

[0017] (4) Strong clinical applicability and low cost. This method only requires routine whole transcriptome sequencing (WTS) data to complete the analysis. It does not require a large amount of fresh tissue for mass spectrometry detection, nor does it require special experimental processing steps, which significantly reduces the technical threshold and cost of clinical testing.

[0018] (5) Significant clinical relevance and good universality. Clinical validation shows that the MiTRI value calculated based on this method is significantly positively correlated with the efficacy of neoadjuvant therapy for tumors, and has the potential to serve as a biomarker for predicting efficacy. Furthermore, pan-cancer analysis results show that the microbial transcriptional reprogramming pattern on which this application is based is prevalent in 11 types of solid tumors, demonstrating the broad prospects for the application of this method in various solid tumors. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the overall technical process of the identification method in this application.

[0020] Figure 2 This is a schematic diagram illustrating the calculation principle of the Microbial Transcription Reprogramming Index (MiTRI).

[0021] Figure 3 This image shows a comparison of the distribution of microbial species, protein-coding genes, and HLA-binding peptides in colorectal cancer tumor (T) tissues and adjacent normal (PT) tissues. Figure 3 Figure a shows a comparison of the distribution at the microbial species level, indicating a high degree of overlap in species composition between tumor tissue and adjacent normal tissue (overlap rate 85.6%). Figure 3 Figure b shows a comparison of the distribution of microbial protein-coding genes, indicating a significant decrease in the proportion of shared transcriptional genes between the two (overlap rate 40.9%). Figure 3 The image in Figure c shows a comparison of the distribution of HLA-binding peptides in microorganisms, revealing significant differences in antigenic epitopes between the two (overlap rate of only 19.0%), indicating that the same microbial species underwent transcriptional reprogramming in tumors and adjacent normal tissues.

[0022] Figure 4 The distribution of MiTRI values ​​for nine colorectal cancer-associated bacteria in colorectal cancer tumor (T) tissue and adjacent normal (PT) tissue.

[0023] Figure 5 This is a molecular structure model of the TCR-peptide-HLA ternary complex.

[0024] Figure 6 This is a comparison chart of spot counts from enzyme-linked immunospot (ELISpot) analysis in an in vitro activation experiment.

[0025] Figure 7 This is a comparative graph showing the distribution of the number of microbial species and the number of microbial HLA-binding peptides in the colorectal cancer response and non-response groups. Figure 7 Figure a shows a comparison of the distribution at the microbial species level, indicating that there is some overlap in species composition between the response group and the non-response group (overlap rate 50.9%). Figure 7 Figure b shows a comparison of the distribution of microbial HLA-binding peptides, revealing that the antigenic epitopes of the two groups of patients are almost completely different (overlap rate of only 0.8%), suggesting a significant correlation between specific antigens identified based on transcriptional reprogramming and treatment efficacy.

[0026] Figure 8 MiTRI values ​​of nine colorectal cancer-associated bacteria are shown in the colorectal cancer response group (R) and non-response group (NR). Detailed Implementation

[0027] The embodiments of this application will be clearly and completely described below. The technical solutions in the embodiments described below are exemplary and only possible technical implementations of this application, not all possible implementations. Those skilled in the art can combine the embodiments of this application to obtain other embodiments without creative effort, and these embodiments are also within the protection scope of this application.

[0028] To better understand this application, the following provides definitions and explanations of relevant terms.

[0029] As used in this application, the term "sample" includes any biological sample or source of data isolated from an individual, such as cells, organs, tumors, blood, or their sequencing data. The samples in this application are logically divided into target samples and background control samples.

[0030] As used in this application, the term "target sample" refers to the sample to be subjected to microbial antigen identification, primarily referring to the target tumor tissue or its biopsy sample. In some embodiments, the target sample may also include paired adjacent normal tissue (adjacent tissue) of the target tumor tissue, but paired tissue is not a necessary condition for implementing the method of this application. There are no particular limitations on the form of the sample; for example, it may be a fresh or frozen tissue sample, or a paraffin section of tissue, etc.

[0031] As used in this application, the term "background control sample" refers to a reference sample used to compare transcriptional activity with the target sample to determine the MiTRI value or to screen for specific regions. In this application, the background control sample may be one or more of the following: adjacent normal tissue (adjacent tissue), low microbial load tissue, and public control datasets.

[0032] As used in this application, the term "low microbial load tissue" refers to a tissue type that has a natural physiological barrier function and a microbial load that is significantly lower than that of other common body surface or mucosal tissues; these tissues have very few types and numbers of microorganisms and have strong mechanisms to resist microbial colonization; as a preferred example of this application, it includes, but is not limited to, at least one of healthy brain tissue, healthy testicular tissue, and normal brain tissue adjacent to a glioma.

[0033] As used in this application, the term "public control dataset" refers to a collection of whole transcriptome sequencing data from non-tumor or adjacent normal tissues derived from public databases. These public databases are biomedical online public databases that provide gene expression and transcriptomics data, and specific examples include, but are not limited to, TCGA, GTEx, and SRA.

[0034] As used in this application, the terms "adjacent" and "adjacent normal tissue" refer to the tissue surrounding the tumor relative to the tumor tissue, and are also referred to in the art as "adjacent tissue".

[0035] As used in this application, the term "tumor" primarily refers to a solid tumor. Specific examples of such solid tumors include, but are not limited to: lung cancer, gastric cancer, pancreatic cancer, glioblastoma, cervical cancer, ovarian cancer, liver cancer, bladder cancer, breast cancer, colon cancer, colorectal cancer, endometrial cancer or uterine cancer, esophageal cancer, kidney cancer, prostate cancer, thyroid cancer, neuroblastoma or head and neck cancer, etc. As an example, the tumor in this application is breast cancer or colorectal cancer.

[0036] In this application, the term "whole transcriptome sequencing data" refers to the collection of all transcripts (RNAs) of a target species, tissue, or cell under a specific physiological state, obtained through sequencing technology, including coding RNA and non-coding RNA. There are no particular limitations on obtaining this sequencing data; it can be achieved using high-throughput sequencing technologies well-known in the art, including next-generation sequencing (NGS), third-generation sequencing (NGS), and single-cell whole transcriptome sequencing. Preferably, the sequencing technologies are NGS and NGS. In this application, whole transcriptome sequencing data can be obtained by downloading from online databases or by collecting clinical samples. Furthermore, it is understood that to ensure the accuracy of the results, sequencing data generally requires rigorous data screening and quality control. In some examples of this application, the data comes from online databases, and the data screening criteria include: whole transcriptome sequencing technology using ribosomal desosomed rRNA followed by random primer library construction; and samples possessing clear clinicopathological information.

[0037] In this application, the term "host sequence" refers to the complete sequence of genetic material of the host organism itself, that is, the nucleic acid fragments in the host genome. In this application, the host sequence specifically refers to the human genome sequence.

[0038] In this application, the term "microbial sequence" or "microbial-derived sequence" refers to the genetic material sequence of a microorganism (including bacteria, viruses, fungi, archaea, etc.), including specific segments or complete genes in its genome.

[0039] In this application, the term "human genome" or "human reference genome" refers to a standard DNA sequence representing the human genome, intended to provide a universal framework for the human genome. In this application, the universal human reference genome refers to GRCh37 / hg19 or GRCh38 / hg38, preferably GRCh38 / hg38. However, such human reference genomes have some gaps; therefore, this application further compares telomere-to-telomere (T2T) human reference genomes.

[0040] In this application, the term "telomere-to-telomere (T2T)" refers to complete sequencing of chromosomes from one telomere to the other. "Telomere-to-telomere human reference genome" refers to a human reference genome assembled seamlessly from one telomere to the other using advanced sequencing technology, filling approximately 8% of the missing regions in the human reference genome. The most commonly used version is T2T-CHM13; as an example, the telomere-to-telomere human reference genome used in this application is CHM13v2.0.

[0041] As used in this application, "artificial vector sequence" refers to a DNA molecule artificially constructed to introduce exogenous DNA into a host cell, while "artificial vector sequence database" specifically contains detailed maps, full sequences, and functional elements of these modified vectors. UniVec is a non-redundant database maintained by the National Center for Biotechnology Information (NCBI) for the rapid identification and filtering of contaminating sequences such as vectors, adapters, linkers, and primers in nucleic acid sequences. In this application, this database is preferred as the artificial vector sequence database for data screening.

[0042] In this application, the term "sequencing coverage" or "transcriptome sequencing coverage" refers to the number of times a base at a specific location in the genome is covered by sequencing reads in transcriptome sequencing alignment results. In this application, this value is typically calculated using bioinformatics tools (such as bedtools or bamCoverage) to reflect the transcriptional expression level at that location.

[0043] In this application, the term "microbial transcriptional programming region (TP region)" is defined as: a set of genomic regions in the target microbial genome that have effective sequencing coverage (coverage depth ≥1) in at least one of the target sample or the background control sample. This region represents the potential transcribed genomic extent of the microorganism in all currently tested samples (i.e., the union of transcriptional potential).

[0044] In this application, the term "microbial transcriptional reprogramming region (MTR region)" is defined as: a genomic region in the target microbial genome that has effective sequencing coverage only in the target sample (e.g., target tumor tissue) and has no effective sequencing coverage in all introduced background control samples. This region represents a genomic segment whose transcription is specifically activated by the microorganism in response to the specific microenvironment of the target sample (e.g., the tumor microenvironment).

[0045] In this application, the term "Microbial Transcriptional Reprogramming Index (MiTRI)" is used to quantify the transcriptional adaptability of microorganisms in a target microenvironment, and is calculated using the following formula (1):

[0046] in, Bases in the target sample-specific microbial transcriptional reprogramming region (MTR region) Coverage depth; This represents the total number of bases in the transcriptional reprogramming region; This represents the total number of microbial reads detected in the sample. This represents the total number of reads in the sequencing library. This represents the total length of the microbial transcription programming region (TP region).

[0047] In this application, the term "antigen" refers to a molecule that, upon entering the body, can elicit an acquired immune response, which may involve antibody production, specific immunogenic cells, or both. Those skilled in the art will understand that any macromolecule, including virtually all proteins or peptides, can serve as an antigen. Furthermore, antigens can be derived from recombinant or genomic DNA or RNA.

[0048] In this application, the term "antigenic epitope" refers to a portion of an antigen that is recognized by the immune system (particularly by antibodies, B cells, or T cells) in a suitable context. An epitope can be a conformational epitope or a linear epitope. The epitope in this application is a linear epitope, defined by a linear, continuous amino acid sequence of a specific region of a protein. In this application, the "tumor microbial antigenic epitope" can bind to MHC molecules, resulting in the presentation of tumor microbial antigens, which are then recognized by T cells and induce T cell activation, thereby attacking tumor cells.

[0049] In this application, the terms "peptide" or "peptide segment" refer to a polymer comprising two or more amino acids covalently linked by peptide bonds. A "protein" may comprise one or more polypeptides or peptide segments, wherein the polypeptides interact with each other covalently or non-covalently. Unless otherwise stated, the terms "peptide," "peptide segment," and "protein" are used interchangeably in this application.

[0050] In this application, the terms "HLA binding affinity", "HLA affinity" or "MHC binding affinity" refer to the binding affinity between a specific antigen and a specific MHC allele.

[0051] In this application, the term "MHC" refers to the major histocompatibility complex, which is a gene complex involved in all vertebrates. MHC proteins or molecules play a function in signal transduction between lymphocytes and antigen-presenting cells during normal immune responses. Human MHC, also known as HLA (human leukocyte antigen), is located on chromosome 6 and mainly comprises MHC molecules. I and MHC II.

[0052] In this application, the term "MHC Class I" refers to major histocompatibility complex class I proteins or genes. In human MHC... I (HLA) I) Within the region, there is HLA A. HLA B. HLA C, HLA E, HLA F, HLA G subregion. MHC class I proteins are present on the surface of almost all cells, including most tumor cells. MHC Protein I carries antigens, typically derived from endogenous proteins or intracellular pathogens, which are then presented to cytotoxic T lymphocytes (CTLs). T cell receptors recognize and bind to MHC receptors. MHC class I molecules are peptides. Each cytotoxic T lymphocyte expresses a unique T cell receptor capable of binding to a specific MHC / peptide complex. MHC class I molecules primarily mediate the presentation of endogenous antigens.

[0053] In this application, the term "MHC class II" refers to major histocompatibility complex class II proteins or genes. MHC class II proteins are mainly expressed on antigen-presenting cells such as B cells, monocytes / macrophages, and dendritic cells. MHC class II molecules primarily mediate the presentation of exogenous antigens, presenting exogenous antigenic peptide molecules to Th cells (helper T cells).

[0054] In this application, the term "allele" generally refers to a pair of genes located at the same locus on a pair of homologous chromosomes that control contrasting traits; it can be a form of gene, a form of gene sequence, or a form of protein. The term "allele typing" refers to the accurate localization of alleles (or heterozygous sites) on the paternal or maternal chromosomes of a diploid (or even polyploid) genome, ensuring that all alleles from the same parent are aligned on the same chromosome.

[0055] HLA is classified into three main types based on different gene loci: type I, type II, and type III. Due to its high sequence variability, HLA has many different alleles. Antigenic peptides (including tumor-specific antigens or microbial-derived antigenic peptides) need to bind to HLA to assist T-cell receptor (TCR) recognition, thereby triggering an immune response. Therefore, predicting HLA typing is crucial for identifying tumor antigens. Currently, HLA genotyping mainly uses PCR technology, combined with allele-specific oligonucleotides (ASOs) or sequence-specific oligonucleotide probes (SSOs), to perform DNA typing of the HLA system. Next-generation sequencing data analysis for HLA typing can obtain information ranging from polymorphisms of individual SNV genotypes to haplotype information.

[0056] In this application, "BWA-MEM" is an efficient sequence alignment algorithm within the BWA software package, primarily used to align nucleotide sequencing reads or assembled contigs to a large reference genome (such as the human genome). While BWA-MEM is preferably used for sequence alignment in this application, it is understood that other sequence alignment algorithms in the art are also available, and those skilled in the art can make appropriate selections as needed.

[0057] In this application, the terms "NetMHCpan," "NetTCR," and "Panpep" refer to antigen prediction or affinity analysis tools commonly used in the field of bioinformatics. Those skilled in the art should understand that these tools are merely exemplary means of implementing this application, and other algorithms or software with similar functions (such as MHCflurry, MixMHCpred, etc.) also fall within the scope of protection of this application.

[0058] In this application, "computer-simulated translation" (also known as "virtual translation" or "in silicotranslation") refers to the computational process of using software tools to convert nucleic acid sequences (DNA or RNA) into their possible corresponding amino acid sequences according to the rules of the genetic code. This includes direct sequence conversion, as well as the process by which bioinformatics alignment tools (such as DIAMOND, BLASTX, etc.) dynamically translate nucleotide sequences into protein sequences during homology searches. It is not a real cellular translation process, but a predictive simulation based on sequence and codon tables.

[0059] In this application, "pVACtools" is an open-source computational toolkit specifically designed for cancer immunotherapy, typically used to identify, screen, prioritize, and visualize tumor-specific neoantigens. In this application, the toolkit is primarily used to process microbial protein fragment sequences, particularly utilizing its sliding window strategy to segment microbial proteins into candidate peptides of specific lengths to facilitate subsequent HLA binding affinity prediction.

[0060] The identification method disclosed in this application can be used to analyze microorganisms present in the tumor microenvironment. As an example, the microorganisms analyzed in this application include, but are not limited to, Fusobacterium nucleatum, Propionibacterium acnes, Helicobacter pylori, Porphyromonas gingivalis, Enterococcus faecalis, and Escherichia coli, but are not limited to these.

[0061] It should be understood that the core contribution of this application lies in revealing the phenomenon of transcriptional reprogramming (MTR) in the tumor microenvironment, and based on this, establishing an antigen discovery strategy based on identifying tumor-specific transcriptional regions (MTR regions) and using the MiTRI index for quantitative screening. This strategy provides a novel method to overcome the challenge of overlapping microbial taxonomy between tumors and normal tissues and to accurately identify tumor-specific microbial antigenic epitopes with high immunogenicity. Subsequent identification processes targeting MTR region sequences (such as antigen peptide prediction, HLA binding prediction, etc.) can be performed using methods known in the art, without particular limitations. Furthermore, after obtaining the corresponding peptides, peptides with TCR binding scores greater than a preset threshold can be further screened as high-confidence microbial antigenic epitopes using a TCR binding prediction algorithm. It should be noted that this is not a necessary step in the identification method of this application, but rather a further step set up for subsequent experimental validation or drug target screening after identifying the corresponding antigenic epitopes.

[0062] The present application will be further illustrated below with reference to specific embodiments. It should be noted that the specific embodiments below are for illustrative purposes only and do not limit the scope of the present application in any way.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.

[0064] In addition, unless otherwise specified, methods without detailed conditions or steps are conventional methods, and the reagents and materials used are commercially available.

[0065] All patient data used in this study strictly adhered to ethical guidelines, and the collection of all data was approved by the ethics committee. The research team ensured that all participants participated voluntarily with informed consent and that patient privacy was fully protected. The collection and use of all data strictly complied with relevant laws, regulations, and ethical requirements, ensuring the transparency and compliance of the research process.

[0066] Example 1: Tumor-specific microbial antigen epitope recognition method This embodiment uses colorectal cancer (CRC) as an example to demonstrate how to identify tumor-specific microbial antigen epitopes. The specific steps are as follows: 1.1 Data Sources and Sample Queues Paired sample data from 89 colorectal cancer patients were collected from the NCBI Sequence Read Archive (SRA) database, including whole transcriptome sequencing data of tumor tissue and adjacent normal tissue. These data came from 6 datasets: PRJNA596460, PRJNA737203, PRJNA778353, PRJNA815861, PRJNA846310, and PRJNA872190.

[0067] Data screening criteria included: whole transcriptome sequencing technology using ribosomal desosomed rRNA followed by random primer library construction to ensure effective capture of microbial transcripts.

[0068] 1.2 Data Preprocessing and Microbial Sequence Extraction To ensure the authenticity of the microbial signals, a rigorous decontamination process was implemented: (1) Quality control. The raw sequencing reads were processed using fastp (v0.23.2) to remove adapter sequences and low-quality bases, and high-quality reads were retained for subsequent analysis.

[0069] (2) Host sequence removal (two-step method). Step 1: Use BWA-MEM (v0.7.17) to align the quality-controlled reads to the human reference genome (GRCh38) and retain the unaligned reads; Step 2: Use BWA-MEM again to align the unaligned reads obtained in Step 1 to the telomere-to-telomere human reference genome (T2T-CHM13) to further remove host homologous sequences that are missing in GRCh38 but exist in the T2T genome.

[0070] (3) Vector contamination removal. The host-filtered reads were aligned to the UniVec database (NCBIVecScreen), and BWA-MEM was used to identify and remove contamination from artificial vector sequences such as plasmids and phages.

[0071] (4) Microbial species identification. GATK PathSeq (v4.2.6.1) was used to classify species by aligning to a tumor-associated microbial reference database. Reads with a cut alignment length ≥50 bp were retained, and only species with a detection rate ≥50% in each group were retained for subsequent calculation of species overlap analysis between tumor tissues and adjacent normal tissues.

[0072] 1.3 Quantification of Microbial Transcriptional Activity and Calculation of MiTRI Index To accurately quantify the transcriptional adaptation (i.e., transcriptional reprogramming) of microorganisms in the tumor microenvironment, this application establishes the Microbial Transcriptional Reprogramming Index (MiTRI) (the calculation principle is described in [link to article]). Figure 2 ).

[0073] (1) Reads extraction and re-comparison Input data: BAM files that have undergone two rounds of host removal and vector library decontamination, and have been filtered using the GATK PathSeq process (i.e., using PathSeqFilterSpark to remove duplicate and low-quality sequences).

[0074] Realignment: The reads in the BAM file above are extracted and converted to FASTQ format, and then realigned to the specific reference genome of the target bacterial species using BWA-MEM (v0.7.17).

[0075] (2) Comparison and Filtering To ensure the specificity of the alignment results, the following dual filtering criteria were applied to the aligned BAM files: Standard 1: Only primary alignments are retained. Secondary alignments and supplementary alignments are discarded to ensure that each read is counted only once.

[0076] Standard 2: Alignment consistency (Identity) ≥ 90%. Parameter definition: The formula for calculating alignment consistency is: Identity = (Matches - Deletions) / Read Length. Where Matches represents the total number of bases in the read sequence that completely match the reference genome (excluding mismatches and insertions / deletions), Deletions represents the number of missing bases relative to the reference genome, and Read Length is the full length of the read. This formula ensures high confidence in the alignment by additionally deducting the penalty for deletions.

[0077] The high-quality alignment results were retained, and single-base resolution transcriptional coverage maps were generated using deepTools bamCoverage for subsequent MiTRI calculations. For coverage calculations of each category, the samtools merge function was used to merge the data into a single bam file before calculating the coverage.

[0078] (3) Background subtraction and MiTRI calculation: A low-microbial-load biological background set was introduced as the first background (including data from healthy brain tissue, healthy testicular tissue, and normal brain tissue adjacent to gliomas). Regions that are transcribed only in the target sample but not covered in the background set were defined as MTR regions, and then MiTRI values ​​were calculated according to the formula defined above.

[0079] 1.4 Differential analysis of antigen libraries between tumors and adjacent normal tissues To demonstrate the differences between tumor tissue and adjacent normal tissue at the level of microbial antigens, thereby verifying the necessity of introducing the concept of "transcriptional reprogramming," this embodiment performs a panoramic prediction of the microbial antigen libraries of the two groups of samples and preliminarily presents the relevant results: (1) Identification of microbial protein-coding genes. Using the DIAMOND (v2.1.16) tool, the filtered high-quality microbial reads were aligned to the microbial protein sequence reference database. Strict filtering criteria: 100% alignment consistency; E-value < 1e-5; Species consistency verification: the taxonomy ID of the reads must be completely consistent with the source species ID of the protein sequence it is aligned to, in order to exclude cross-species erroneous alignments.

[0080] (2) Prediction of HLA-binding peptides. The identified microbial protein fragment sequences were cut using a sliding window using pVACtools to generate candidate peptides of 8-11 aa (for MHC class I) and 15 aa (for MHC class II). High-precision HLA allele typing was performed directly using the patient's whole transcriptome sequencing (WTS) data with the HLA-HD (v1.7.0) tool. The typing range covered the major loci of MHC class I (including classical loci HLA-A, HLA-B, HLA-C and non-classical loci HLA-E, HLA-F, HLA-G) and MHC class II (HLA-DR, HLA-DQ, HLA-DP) to ensure the personalization and accuracy of subsequent affinity prediction. Based on the patient's specific HLA typing information, NetMHCpan-4.2 (MHC-I) and NetMHCIIpan-4.3 (MHC-II) were used to predict peptide affinity. All peptides predicted as strongly bound (SB) or weakly bound (WB) according to the default parameters are retained as the "genome-wide background peptide library".

[0081] Analysis revealed that although CRC tumors and adjacent normal tissues shared a high degree of similarity in microbial species composition (85.6% species sharing), significant functional differences existed. Only 40.9% of microbial protein-coding genes and 19.0% of predicted HLA-binding peptides were found. Figure 3 The results were shared between the two groups. MiTRI analysis revealed that nine core-associated bacteria of colorectal cancer, when calculated to have MiTRI values, exhibited significant reprogramming of transcriptional activity in tumor tissue. Figure 4 These nine core bacteria associated with colorectal cancer include Fusobacterium animalis, Fusobacterium canifelinum, Fusobacterium hwasookii, Fusobacterium massiliense, Fusobacterium nucleatum, Fusobacterium periodonticum, Fusobacterium polymorphum, Fusobacterium pseudoperipheral, and Fusobacterium vincentii.

[0082] Example 2: Precise identification and peptide optimization of tumor-specific MTR region antigens This embodiment demonstrates how, under complex clinical sample conditions (such as unpaired adjacent normal tissue), the method of this application can be used to accurately locate antigens and, in combination with patient-specific single-cell TCR sequencing data, select high-confidence targets.

[0083] 2.1 MTR Region Determination Strategy for Unpaired Adjacent Normal Samples This embodiment introduces a case of a colorectal cancer neoadjuvant therapy response group patient (number P0006) collected from this research center, independent of the public dataset in Embodiment 1. Given the lack of paired adjacent normal tissue for this patient, a "public background set strategy" was used to replace paired adjacent normal tissue. The MTR region was determined and MiTRI was calculated according to the method in Embodiment 1, including the following steps: (1) Constructing a background atlas: Using the CRC adjacent normal tissue data collected in Example 1 from 89 public databases, a “background transcriptome” of the microbial transcriptome was constructed.

[0084] (2) Region locking: The transcriptional data of the P0006 tumor sample was compared with the "background map". Only the protein-coding regions that were transcribed in the P0006 tumor sample but were completely uncovered in the background map were retained and defined as tumor-specific MTR regions.

[0085] (3) Antigen screening: Only predicted HLA-binding peptides encoded by transcribed sequences falling within the MTR region were retained. Ultimately, among the 7101 microbial HLA-binding peptides screened against the genomes of 9 CRC-related bacteria, 446 (approximately 6.3%) were identified as originating from tumor-specific MTR regions.

[0086] 2.2 In vivo validation: Global association analysis of T cell clonal expansion and MTR antigen To verify whether the MTR antigens selected above elicited a real immune response in the patient, this embodiment introduced single-cell TCR sequencing data of peripheral blood mononuclear cells (PBMCs) of the patient (P0006) after treatment (based on the 10xGenomics platform).

[0087] Analysis method: The PanPep (v1.0.0) algorithm (which supports prediction of class I and class II) was used to predict the binding affinity between the TCR clones (clone size ≥ 2) expanded in the patient and all detected microbial HLA binding peptides (7101 in total, including MTR-derived and non-MTR-derived peptides).

[0088] Analysis Results: At the TCR clonal level, among the 93 amplified TCR clonal types detected, 74 (79.6%) were predicted (PanPep Score ≥ 0.7) to specifically recognize MTR-derived antigenic epitopes. At the TCR cell count level, the number of TCR cells specifically recognizing MTR antigens reached 494, accounting for 64.2% of all high-affinity amplified TCR cells (769). Calculations showed that the "binding density" of MTR-derived antigenic epitopes recruiting high-affinity amplified TCRs was more than 10 times that of background peptides across the entire genome. This enrichment phenomenon was consistent in both MHC-I and MHC-II restriction epitopes. The difference was statistically significant (Poisson rate ratio test, P < 0.0001).

[0089] Conclusion: The large-scale screening data based on PanPep shown above indicate that the MTR region is the main source of antigens inducing T cell clonal expansion in patients, validating the biological rationale for the identification method proposed in this application. Example 3: Optimization and Functional Verification of High-Confidence Antigenic Peptides Based on the findings of Example 2, this embodiment further establishes stringent criteria to select the core antigen peptide and verifies its potential as a vaccine / drug component through dry and wet experiments.

[0090] 3.1 Optimal Strategy for High-Confidence MTR Antigenic Peptides In order to screen the targets with the greatest clinical translation potential from the aforementioned 446 MTR candidate antigens, this embodiment established the following strict "high-confidence" selection criteria: (1) Dual-algorithm affinity verification. The candidate peptide is required to have a prediction binding score ≥ 0.7 in both PanPep (v1.0.0) and NetTCR-2.2, two independent TCR-antigen binding prediction algorithms. This "double insurance" strategy greatly reduces the false positives that may be caused by a single algorithm.

[0091] (2) Hyperclonal amplification association. The candidate peptide must be recognized by a highly amplified TCR clone in the patient's body. Specifically, the absolute clone size (Clone Size) of the TCR clone that recognizes the peptide in the single-cell sequencing data must be ≥10.

[0092] Based on the above criteria, 21 high-confidence MHC class I restriction MTR antigen peptides were finally selected from 446 candidates. The specific sequence information is shown in Table 1. Table 1. Screened high-confidence MTR antigen peptides

[0093] 3.2 Structural Verification: TCR-pMHC Molecular Dynamics Simulation To elucidate the mechanism by which T cell receptors (TCRs) recognize MTR antigens at the atomic level, this embodiment selected representative high-affinity ternary complexes from a preferred list for structural modeling and analysis: antigenic peptides EHFMHLVGI and HLA-B from the MTR region of Fusobacterium polymorphum. 14:02 (allele carried by patient P0006), and the TCR β chain CDR3 sequence CSVEYLGITGELFF (derived from the CDR3β region of the TCR clone that was frequently amplified in the patient's body).

[0094] (1) Structural modeling process. The initial crystal structure of the MHC molecule was obtained from the Protein Database (PDB), and amorphous water molecules were removed using PyMOL. The cocrystal peptide was mutated to the target peptide sequence using PyMOL's mutagenesis wizard. The peptide was re-aligned into the MHC binding groove using HDOCK (v1.1). From the top 10 pMHC complexes, the most reasonable model was selected based on docking score and binding geometry. The variable regions of the α and β chains of the TCR were tandemly submitted to AlphaFold3 for structure prediction, and the model with the highest pLDDT score was retained for docking.

[0095] (2) Molecular dynamics simulations. Molecular dynamics simulations were performed using GROMACS (v2024.3) and the AMBER99SB-ILDN force field. The ternary complex was solvated in a cubic box using the SPC / E water model with a minimum 1.5 nm filler and Na + or Cl - Ion neutralization. Energy minimization was performed using the steepest descent method until the maximum force was below 1000.0 kJ / mol / nm. Equilibrium was achieved in two steps: first, a 1 ns NVT ensemble was performed at 300 K using a V-rescale thermostat to impose positional constraints on the protein's heavy atoms; then, a 1 ns NPT ensemble was performed at 1.0 bar using a Parrinello-Rahman thermostat. Production was run for 100 ns to release all constraints.

[0096] (3) Simulation results and binding energy calculation. The simulation results show that ( Figure 5The TCR-pMHC complex maintained high structural stability during the 100 ns simulation, with the CDR3 ring of the TCR maintaining a tight interface with the antigen peptide-HLA complex. The binding free energy between TCR and pMHC was estimated using the MM / PBSA method via gmx_MMPBSA (v1.6.4), yielding ΔG = -17.71 kcal / mol. This suggests that microbial antigens derived from the MTR region can form a stable TCR-pMHC complex, supporting their potential as a T cell recognition target.

[0097] 3.3 In vitro validation: Antigen peptide immunogenicity assay (ELISpot) This embodiment verifies the ability of the selected antigenic peptides to activate immune cells through in vitro experiments. To fully verify the reliability of the prediction results, this embodiment validates the 21 high-confidence peptides (SEQ ID NO.1-SEQ ID NO.21) screened and synthesized in Example 3.1.

[0098] Experimental Methods: The 21 peptides were divided into two peptide pools for testing: Pool 1 (containing 10 peptides) and Pool 2 (containing 11 peptides). Peripheral blood mononuclear cells (PBMCs) were collected from patients in the response group (P0006) and cultured in vitro for peptide stimulation.

[0099] Experimental groups: The experimental groups were supplemented with MTR-derived antigen Pool 1 and Pool 2, respectively; the negative control group consisted of solvent controls corresponding to the peptide pools (Neg 1 and Neg 2) to eliminate background interference from the solvent (DMSO); the positive control group was supplemented with phytohemagglutinin (PHA). After 10 days of in vitro expansion culture, the level of interferon-γ (IFN-γ) secreted by T cells was detected by enzyme-linked immunospot assay (ELISpot). All experiments were performed in triplicate.

[0100] Experimental results: ELISpot results show ( Figure 6 In T cell wells stimulated by the MTR antigen peptide pool, the number of IFN-γ spots (SFUs) increased significantly. The number of IFN-γ spots in the Pool 1 stimulation well was 3.05 times higher than that in the negative control (Neg 1), and the number of IFN-γ spots in the Pool 2 stimulation well was 1.81 times higher than that in the negative control (Neg 2).

[0101] Statistical significance was assessed using the distribution-free resampling (DFR(2x)) method, with a 1.5-fold change threshold set and 100,000 bootstrap iterations performed. (P<0.001). This method confirmed that the differences between the two stimulation groups and their respective negative controls were statistically significant.

[0102] The above results demonstrate that the MTR-derived antigenic peptides (covering SEQ ID NO.1-SEQ ID NO.21) screened using the method described in this application can be specifically recognized by the patient's autologous T cells and induce strong IFN-γ secretion. This not only functionally validates the biological activity of the MTR region as an immunogenic hotspot, but also strongly supports the potential application value of all 21 high-confidence peptides (SEQ ID NO.1-SEQ ID NO.21) as components of tumor vaccines or drugs.

[0103] Example 4: Application of MiTRI index in predicting the efficacy of neoadjuvant cancer therapy This embodiment verifies the application value of the MiTRI index as a biomarker for predicting the efficacy of neoadjuvant therapy for colorectal cancer.

[0104] 4.1 Patient Enrollment and Clinical Treatment Protocol This study included 8 patients with pathologically confirmed primary colorectal adenocarcinoma, clinically staged as AJCC II-III (T2N+ or ≥T3, M0), and confirmed as microsatellite stable (MSS) by testing.

[0105] The patient received a standard neoadjuvant chemotherapy regimen: 85 mg / m² intravenously on day 1. 2 Platinum-based drugs and 400 mg / m 2 Leucovorin calcium, followed by a bolus injection of 400 mg / m² 2 Fluoropyrimidine drugs, administered continuously at 2400 mg / m² over 46-48 hours. 2 Fluoropyrimidine drugs. This regimen is repeated every 2-3 weeks for a total of 4-6 cycles.

[0106] 4.2 Efficacy assessment and grouping After treatment, efficacy was assessed according to the following criteria. Those who met at least two criteria were classified as Responders (R), otherwise as Non-Responders (NR): (1) Clinical assessment: No primary lesion was palpable during digital rectal examination; (2) Endoscopic assessment: The tumor has disappeared or only superficial scars are visible; (3) Imaging assessment: MRI showed tumor regression grade (mrTRG) of 1-2; (4) Pathological assessment: The postoperative pathological tumor regression grade (TRG) was 0-1.

[0107] Ultimately, 4 patients were assigned to the response group and 4 to the non-response group.

[0108] 4.3 Sample Collection and Sequencing Prior to treatment, tumor tissue was collected via endoscopic biopsy. Depending on the tumor size, 2 to 3 tissue samples were collected from each patient, each approximately 2-3 mm in diameter. Immediately after collection, the samples were rinsed with saline to remove blood and mucus, flash-frozen on dry ice for 2 hours, and stored at -80°C. Whole transcriptome sequencing (WTS) was performed using a ribosomal desosomed rRNA library preparation method, generating 150 bp paired-end sequencing data. Peripheral blood mononuclear cells (PBMCs) were separated using Ficoll-Paque plus density gradient centrifugation, and single-cell RNA and single-cell TCR sequencing were performed using a 10×Genomics platform.

[0109] 4.4 Association analysis between MiTRI and treatment response Using the identification method described in Example 1, MiTRI values ​​were calculated for nine core colorectal cancer-associated bacteria, including *Fusobacterium animalis*, *Fusobacterium canifelinum*, *Fusobacterium hwasookii*, *Fusobacterium massiliense*, *Fusobacterium nucleatum*, *Fusobacterium periodonticum*, *Fusobacterium polymorphum*, *Fusobacterium pseudoperipheral*, and *Fusobacterium vincentii*.

[0110] Statistical analysis showed that although the microbial species composition of the two groups was 50.9% similar (Jaccard index), the predicted overlap rate of HLA-binding peptides was only 0.8%. Figure 7 The MiTRI value of tumor-associated microorganisms in the response group was significantly higher than that in the non-response group (P<0.05, two-sided U test). Figure 8 ).

[0111] These results indicate that even though the two patient groups had high similarity in microbial species composition, their transcriptional activity status (MiTRI) differed significantly. This suggests that the MiTRI index can serve as a highly sensitive biomarker independent of species abundance, used to accurately predict the probability of response to neoadjuvant therapy in MSS-CRC patients before treatment.

[0112] It is worth noting that although this embodiment is limited by the sample size and does not define a fixed threshold, in actual clinical applications, the preset reference value can be determined through statistical methods. For example, based on the MiTRI distribution of the non-response group (NR) in a large sample cohort, the reference value can be set as the 95th percentile of the NR group; or, by constructing an ROC curve (Receiver Operating Characteristic curve), the MiTRI value corresponding to the maximum point of the Youden index can be calculated as the optimal cut-off value. When the MiTRI value of the sample to be tested is higher than this cut-off value, it indicates that it has a higher probability of treatment response.

[0113] Example 5: Pan-cancer analysis and validation 5.1 Multi-cancer microbial transcriptional reprogramming analysis To assess the generalizability of the MiTRI method in this application, this embodiment analyzed 1893 samples from 11 solid tumor types in the SRA database, including 1184 tumor samples and 709 samples of adjacent normal tissue. The cancer types covered included bladder cancer, breast cancer, cervical cancer, colorectal cancer, kidney cancer, liver cancer, lung cancer, oral squamous cell carcinoma, pancreatic cancer, gastric cancer, and thyroid cancer.

[0114] The analysis results showed that although tumors and adjacent normal tissues had a high degree of similarity in microbial composition (96.6% sharing rate at the genus level and 95.3% sharing rate at the species level), there were significant differences at the functional level (only 39.6% of the microbial protein-encoding genes were expressed in both tissue types, and the predicted HLA-binding peptide overlap rate was only 22.0%).

[0115] Spearman correlation analysis showed that tumor-specific MiTRI was significantly correlated with multiple immunological indicators: the correlation coefficient with the predicted number of HLA-binding peptides was ρ=0.65 ( P <2.2×10 -16 The correlation coefficient ρ between IEDB and immunogenicity prediction was 0.31. P <2.2×10 -16 The correlation coefficient ρ between BigMHC and BigMHC in predicting immunogenicity was 0.34. P <2.2×10 -16 The correlation coefficient ρ between DeepImmuno and DeepImmuno's prediction of immunogenicity was 0.38. P<2.2×10 -16 ).

[0116] The above results indicate that the microbial transcriptional reprogramming principle upon which this application is based is prevalent in various solid tumor types, demonstrating the broad application prospects of the MiTRI-based method in various solid tumors.

[0117] It should be noted that this application is not limited to the above-described embodiments. The above embodiments are merely examples, and any embodiments with the same structure and effect as the technical concept within the scope of this application are included in the technical scope of this application. Furthermore, various modifications that can be conceived by those skilled in the art to the embodiments, and other ways of constructing by combining some of the constituent elements of the embodiments, without departing from the spirit of this application, are also included in the scope of this application.

Claims

1. A method for recognizing tumor-specific microbial antigen epitopes, characterized in that, Includes the following steps: Whole transcriptome sequencing data of the samples were obtained, and high-quality non-host microbial sequences were obtained through data processing and species classification screening was performed; the samples included target samples and background control samples. Microbial genomic regions covered by transcriptome sequencing in at least one of the target sample or background control sample are identified as microbial transcriptional reprogramming regions. Microbial genomic regions covered only by transcriptome sequencing in the target sample but not in the background control sample are identified as target sample-specific microbial transcriptional reprogramming regions. The microbial transcriptional reprogramming index MiTRI is calculated according to formula (1) to obtain microorganisms with high transcriptional reprogramming activity in the target sample: In formula (1), Bases in the target sample-specific microbial transcriptional reprogramming region Coverage depth; This represents the total number of bases in the transcriptional reprogramming region; This represents the total number of microbial reads detected in the sample. This represents the total number of reads in the sequencing library. The total length of the microbial transcription programming region; Based on microorganisms exhibiting high transcriptional reprogramming activity in target samples, tumor-specific candidate antigen epitopes are identified and screened.

2. The tumor-specific microbial antigen epitope recognition method as described in claim 1, characterized in that, The target sample contains at least the target tumor tissue, and the background control sample is derived from at least one of the adjacent normal tissue, low microbial load tissue, or public control dataset of the target tumor tissue. The low microbial load tissue is selected from one or more of healthy brain tissue, healthy testicular tissue, and normal brain tissue adjacent to a glioma; and / or the data in the public control dataset comes from a biomedical online public database that provides gene expression and transcriptomics data.

3. The tumor-specific microbial antigen epitope recognition method as described in claim 1, characterized in that, The data processing includes sequentially performed steps of quality control, host sequence removal, and artificial vector sequence removal.

4. The tumor-specific microbial antigen epitope recognition method as described in claim 1 or 3, characterized in that, The data processing steps specifically include the following steps a) to d): a) Align the quality-controlled sequencing reads to the universal human reference genome; b) Align the unaligned reads from step a) to the telomere-to-telomere human reference genome; c) Align the unaligned reads from step b) to the artificial vector sequence database; d) Identify the unaligned reads from step c) as high-quality non-host microbial sequences.

5. The tumor-specific microbial antigen epitope recognition method as described in claim 1, characterized in that, Based on microorganisms exhibiting high transcriptional reprogramming activity in target samples, the steps for identifying and screening tumor-specific candidate antigenic epitopes include: antigenic peptide prediction; HLA binding prediction; and tumor-specific antigenic epitope screening.

6. The tumor-specific microbial antigen epitope recognition method as described in claim 5, characterized in that, The antigen peptide prediction is achieved by extracting microbial transcriptome sequencing reads that exhibit high transcriptional reprogramming activity in the target sample and using computer simulation translation to predict microbial protein fragment sequences. The HLA binding prediction involves using an antigen epitope prediction tool to segment the microbial protein fragment sequence into candidate peptides and predict the binding affinity of the candidate peptides to the patient's HLA molecules. The tumor-specific antigen epitope screening identifies HLA-binding microbial peptides from microbial transcriptional reprogramming regions as tumor-specific candidate microbial antigen epitopes within those regions.

7. The tumor-specific microbial antigen epitope recognition method as described in claim 6, characterized in that, The HLA binding prediction specifically includes the following steps: The microbial protein fragment sequence is cut using a sliding window to generate the candidate peptide, which consists of an MHC class I peptide of 8-11 amino acids and an MHC class II peptide of 12-18 amino acids. The percentile rank (%Rank) of the candidate peptides was calculated based on the HLA affinity prediction algorithm. Peptides predicted to be strongly or weakly bound are retained.

8. The tumor-specific microbial antigen epitope recognition method as described in claim 6 or 7, characterized in that, The method further includes a high-confidence screening step: after obtaining the HLA-binding microbial peptides, the method further combines a TCR binding prediction algorithm to screen peptides with TCR binding scores greater than a preset threshold as high-confidence microbial antigen epitopes for subsequent experimental verification or drug target screening.

9. A system / device for recognizing tumor-specific microbial antigenic epitopes, characterized in that, The identification system / device implements the identification method according to any one of claims 1-8 during operation; The identification system / device includes: The data input module is configured to receive whole transcriptome sequencing data from the sample; The sequence processing and transcription analysis module is configured to perform data processing, obtain high-quality non-host microbial sequences, and perform species classification and screening. The region definition module is configured to define microbial transcriptional programming regions and target sample-specific microbial transcriptional reprogramming regions based on differences in whole transcriptome sequencing coverage. The MiTRI calculation module is configured to calculate the microbial transcriptional reprogramming index according to formula (1) as defined in claim 1; The epitope screening module is configured to predict the affinity of microbial protein fragment sequences from specific transcriptional reprogramming regions of the target sample and output HLA-binding microbial peptides from the microbial transcriptional reprogramming regions.

10. A tumor-specific microbial antigen polypeptide, characterized in that, The polypeptide is obtained by identifying the target tumor tissue using the identification method described in any one of claims 1-8.

11. The tumor-specific microbial antigen polypeptide according to claim 10, characterized in that, The target tumor tissue is colorectal cancer, and the amino acid sequence of the screened polypeptide is shown in at least one of SEQ ID NO.1 to SEQ ID NO.

21.

12. A system / device for assisting in the evaluation of the efficacy of tumor treatment, characterized in that, The system / device includes the following modules: A calculation module configured to calculate the MiTRI value of the microbial transcriptional reprogramming index in the tumor tissue to be tested according to formula (1) as defined in claim 1; The data analysis module is configured to compare the MiTRI value with a preset reference value; The results analysis and output module is configured to output an evaluation result of the efficacy of tumor treatment based on the results of the data analysis module. If the MiTRI value is higher than the preset reference value, it indicates that the tumor tissue under test has a high response to the treatment. The treatment referred to herein is neoadjuvant therapy or immunotherapy.

Citation Information

Patent Citations

  • Tumor antigen prediction method based on whole transcriptome, and application of the same

    CN109801678A