Method for screening for tumor neoantigen and use thereof
By using NGS data analysis and bioinformatics methods to screen tumor neoantigens, the problem of failing to effectively screen RNA-level mutations and patient specificity in existing technologies has been solved, achieving highly accurate and efficient screening of tumor neoantigens and improving the efficacy of immunotherapy.
Patent Information
- Application Number
- PCT/CN2025/108060
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-11
- Filing Date
- 2025-07-11
- Publication Date
- 2026-01-15
AI Technical Summary
Existing technologies fail to effectively consider gene fusion mutations caused by RNA-level mutations when screening for tumor neoantigens, and the constructed peptide libraries do not distinguish them from peptides produced by the patient's own normal proteins. This results in some mutant peptides being identical to the patient's own highly expressed normal peptides, thus failing to achieve an effective immune response.
Based on next-generation sequencing (NGS) data, combined with immunology and bioinformatics, candidate SNV/INDEL and fusion mutations were screened by analyzing WES data from tumor and normal samples. Low-expression mutations were filtered out, and candidate mutant peptide sequences were calculated and extracted. Target peptide sequences were selected based on antigen binding, immune presentation, antigen similarity, and self-similarity. The sequences were then ranked by comprehensively considering clonalness and mutation abundance.
It improves the accuracy of tumor neoantigen screening, covers mutations with high probability of master clonal variation, encompasses the entire process from abnormal protein processing to T cell recognition, reduces false positives, and improves the prognostic accuracy of immunotherapy.
Smart Images

Figure CN2025108060_15012026_PF_FP_ABST
Abstract
Description
A method for screening tumor neoantigens and its application Technical Field
[0001] This invention relates to a method for screening tumor neoantigens and its application, belonging to the field of bioinformatics. Background Technology
[0002] Neoantigens are autoantigens produced by tumor cells due to genomic mutations. They are generated by gene mutations in tumor cells and can be presented by the major histocompatibility complex (MHC) of antigen-presenting cells. They are recognized by T cells, activate the body's immune system, and trigger a series of immune responses. Therefore, they are ideal targets for personalized cancer immunotherapy.
[0003] CN 115424740 A discloses a method and system for predicting the prognosis of tumor immunotherapy by constructing a peptide library with neoantigens based on missense mutation (SNV) and indel mutation information from NGS data. However, this method neglects gene fusion mutations generated at the RNA level, missing potentially existing and more distinct gene fusion neoantigens. Furthermore, the constructed peptide library does not differentiate between peptides generated from the patient's own normal proteins, leading to some mutant peptides potentially being identical to the patient's highly expressed normal peptides, thus failing to achieve an immune response.
[0004] CN 118028442 A discloses a method for precise screening of personalized neoantigen epitopes in tumors. This method predicts the affinity of peptides for MHC and selects the highest-ranking peptides as neoantigens based on affinity ranking. However, the immune process is highly complex. Affinity only involves the MHC binding step; prior to this, there is peptide generation, and subsequently, TCR recognition. This approach, relying solely on affinity ranking, is too simplistic and results in poor prognostic accuracy in immunotherapy.
[0005] Therefore, there is an urgent need to develop a more effective method for predicting and assessing tumor neoantigens. Summary of the Invention
[0006] To address the aforementioned technical problems, the present invention aims to provide a method for screening tumor neoantigens and its application. This screening method, based on next-generation sequencing (NGS) data, immunology, and bioinformatics, and considering the various biological processes involved in the immune efficacy of antigenic peptides, can more effectively predict and evaluate tumor neoantigens.
[0007] To achieve the above objectives, the present invention provides a method for screening tumor neoantigens, the method comprising:
[0008] We analyzed the WES data of tumor samples and normal samples from the individuals to be tested to obtain candidate SNV / INDEL data.
[0009] RNA-seq data from tumor samples were analyzed to obtain candidate fusion data;
[0010] The candidate SNV / INDEL data and candidate fusion data are filtered out, wherein low-expression mutations are filtered out to obtain filtered mutation sites; the filtered mutation sites are calculated and the sequences containing the mutation sites are extracted to obtain candidate mutant peptide sequences.
[0011] Candidate mutant polypeptide sequences are screened based on antigen binding, immune presentation, antigen similarity, and self-similarity to obtain target polypeptide sequences, from which neoantigens are selected.
[0012] In the above method, preferably, the process of screening candidate mutant polypeptide sequences includes:
[0013] From the candidate mutant polypeptide sequences, polypeptide sequences with MHC affinity lower than 2500 nM and binding half-life greater than 0.5 h were selected as general candidate sequences.
[0014] From the general candidate sequences, polypeptide sequences that meet at least one of the following categories are selected:
[0015] Type I: MHC affinity less than 100 nM, binding half-life greater than 1 hour
[0016] Category II: Presentation rank less than 2%
[0017] Category III: Antigen similarity higher than 80%
[0018] Category IV: Individual differences exceeding 80%
[0019] Selected polypeptide sequences that meet at least one of the above categories are included as shortlisted sequences. The shortlisted sequences are sorted as a whole according to their clonalness and mutation abundance, or sorted separately according to categories I-IV. Sequences with higher clonalness and mutation abundance are selected as target polypeptide sequences. Alternatively, sequences are sorted according to their comprehensive performance in antigen binding, immune presentation, antigen similarity, and self-similarity. Sequences with higher comprehensive performance are selected as target polypeptide sequences.
[0020] According to a specific embodiment of the present invention, when the shortlisted sequences include class III and / or class IV polypeptide sequences, the target polypeptide sequence preferably includes class III and / or class IV polypeptide sequences.
[0021] According to a specific embodiment of the present invention, when sorting sequences according to their clonalness and mutation abundance, both clonalness and mutation abundance range from 0 to 1. The sequences are sorted by multiplying their values and sorting from largest to smallest. Generally, in shortlisted sequences, expression status has higher priority than grouping criteria. When classifying general candidate sequences, the higher the rank number (I, II, III, IV), the higher its priority; that is, class IV has the highest priority, and so on. When a general candidate sequence simultaneously meets different classification criteria, it should be prioritized for inclusion in the higher priority group. This method can maximize the inclusion of all potential mutations.
[0022] Candidate mutant peptide sequences are screened based on antigen binding, immune presentation, antigen similarity, and self-similarity to obtain target peptide sequences. The classification and ranking methods for candidate mutant peptide sequences are not limited to the above combinations.
[0023] If the sample has a high mutation rate, other methods can also be used. Besides the methods corresponding to the principles for including samples in high-priority groups, a ranking accumulation method can also be used. That is, the original group represents the score; if a sample belongs to multiple groups, the scores from all groups are added together, and the ranking is based on this score.
[0024] According to a specific embodiment of the present invention, the process of ranking the shortlisted sequences based on a comprehensive performance of antigen binding, immune presentation, antigen similarity, and self-similarity includes: analyzing the epitopes of MHCI and MHCII in the shortlisted sequences for mutation sites, and scoring the shortlisted sequences based on the epitope information. The score for each polypeptide sequence is calculated according to the following formula:
[0025] The peptide sequence score = MHCII antigen binding score + MHCII immune presentation score + MHCII antigen binding score + MHCII immune presentation score + max{antigen similarity score, self-similarity score}
[0026] For each polypeptide, there are MHCII antigen binding score, MHCII immune presentation score, MHCII antigen binding score, MHCII immune presentation score, antigen similarity score, and self-similarity score, among which:
[0027] MHCI affinity below 100 nM and binding half-life greater than 1 h earn 1 point (0 points if this condition is not met, i.e., no points for this item); MHCII affinity below 100 nM and binding half-life greater than 1 h earn 1 point (0 points if this condition is not met, i.e., no points for this item); MHCI immunopresentation rank less than 2% earns 2 points (0 points if this condition is not met, i.e., no points for this item); MHCII immunopresentation rank less than 2% earns 2 points (0 points if this condition is not met, i.e., no points for this item); antigen similarity greater than 80% earns 3 points (0 points if this condition is not met, i.e., no points for this item); self-similarity greater than 80% earns 4 points (0 points if this condition is not met, i.e., no points for this item); max{antigen similarity score, self-similarity score} refers to taking the higher score between antigen similarity score and self-similarity score. For example: a polypeptide with an MHC I affinity that meets the criteria of "less than 100 nM and a binding half-life greater than 1 h" receives 1 point; an MHC II affinity that does not meet the criteria of "less than 100 nM and a binding half-life greater than 1 h" receives 0 points; an MHC I immunopresenting rank of less than 2% receives 2 points; an MHC II immunopresenting rank of less than 2% receives 2 points; and if antigen similarity is greater than 80% and self-similarity is greater than 80%, then the maximum score for both antigen similarity and self-similarity is 4 points. The scores for each of these categories can be scaled up or down proportionally.
[0028] The score of each polypeptide sequence obtained in the above manner represents the overall performance of that sequence in terms of antigen binding, immune presentation, antigen similarity, and self-similarity. Preferred sequences with higher scores are included in the target polypeptide sequences of this invention.
[0029] According to a specific embodiment of the present invention, when the number of neoantigens screened is not required, a neoantigen screening method with higher confidence can also be adopted. If the sample has a relatively high number of mutations, then after calculating the data of antigen binding, immune presentation, antigen similarity, and self-similarity of the candidate mutant polypeptide sequence, all dimensions of a single candidate mutant polypeptide sequence can be comprehensively evaluated, such as by functionalization methods. This method has stronger filtering capabilities and is suitable for finding a small number of high-confidence neoantigens.
[0030] According to some specific embodiments of the present invention, the process of screening candidate mutant polypeptide sequences may further include: using a functionalization method, calculating and normalizing the distances of the candidate mutant polypeptide sequences to antigen binding, immune presentation, antigen similarity and self-similarity standards respectively, with positive numbers for those better than the standards and negative numbers for those worse than the standards, and finally accumulating the scores to sort the candidate mutant polypeptide sequences, selecting target polypeptide sequences based on the sorting results, and selecting neoantigens from them.
[0031] According to some specific embodiments of the present invention, the process of screening candidate mutant polypeptide sequences may further include: using a functionalization method, taking some indicators from the four dimensions of antigen binding, immune presentation, antigen similarity and self-similarity as basic criteria, for example, excluding all candidate mutant polypeptide sequences whose self-difference is lower than a new threshold, ranking the candidate mutant polypeptide sequences according to the fitting function of the indicators of the other three dimensions, selecting the target polypeptide sequence according to the ranking result, and selecting the neoantigen from it.
[0032] According to a specific embodiment of the present invention, the method of selecting neoantigens from target polypeptide sequences is mainly by verifying the immunogenicity of the target polypeptide sequences. Specifically, the target polypeptide sequences can be tandemly linked or combined in other orders or set as multiple tandem sequences for verification, which does not affect the results. Alternatively, the obtained target polypeptide sequences can be verified individually.
[0033] In the above method, the purity of the tumor sample should be greater than or equal to 10%, preferably greater than or equal to 40%. If the purity of the tumor sample is less than 40%, a sample quality warning should be given; if the purity of the tumor sample is less than 10%, resampling and sequencing are required. The purity of the tumor sample can be calculated using the open-source Sequenza software.
[0034] In the above method, when filtering the candidate SNV / INDEL data and candidate fusion data, low-expression mutations are somatic mutations with an FPKM of less than 5 in transcript abundance data and / or a mutation frequency of less than 10% in DNA / RNA in mutation abundance data.
[0035] The calculation method for FPKM (Fragments Per Kilobase of exon model per Million mapped fragments) is as follows:
[0036] In the above method, preferably, when filtering the candidate SNV / INDEL data and candidate fusion data, the filtering also includes filtering in conjunction with a mutation priority database.
[0037] According to some specific embodiments of the present invention, by summarizing the performance of proto-oncogenes and tumor suppressor genes in the tumor immune process in existing literature and public databases, a mutation priority database can be obtained. Corresponding screening parameter combinations are adopted for mutations of different priorities. The higher the mutation priority, the more lenient the corresponding filtering criteria.
[0038] According to a specific embodiment of the present invention, preferably, the mutation priority is divided into 5 priorities, with the highest priority being level 1 and the lowest being level 5. Based on the standard that "low-expression mutations are somatic mutations with an FPKM of less than 5 in the transcript abundance data and / or a mutation frequency of less than 10% in DNA / RNA in the mutation abundance data", the actual filtering standard is the corresponding index multiplied by the priority / 5. That is, the low-quality mutations of the first priority are somatic mutations with an FPKM of less than 1 and / or a mutation frequency of less than 2% in DNA / RNA in the mutation abundance data.
[0039] In the above method, the mutation site can be truncated from the sequences containing the mutation site and flanking it in the candidate SNV / INDEL data and candidate fusion data.
[0040] In the above method, preferably, the peptide length of the candidate mutant polypeptide sequence is 17aa-29aa.
[0041] The present invention preferably extracts epitope peptides containing mutated MHCI and MHCII classes. MHCII epitopes are generally 15 aa longer than MHCI epitopes (9 aa). Therefore, in order to ensure the inclusion of mutation sites, the final peptide length of each epitope is between 17 aa and 29 aa.
[0042] According to a specific embodiment of the present invention, the method for screening tumor neoantigens further includes: analyzing WES data of tumor samples and normal samples to obtain clonal analysis data; analyzing RNA-seq data of tumor samples to obtain mutation abundance data and transcript abundance data; and then sorting the screened polypeptide sequences according to their clonal and mutation abundance.
[0043] According to a specific embodiment of the present invention, the method for screening tumor neoantigens of the present invention further includes:
[0044] HLA typing data of the individual to be tested is obtained, and then the candidate mutant peptide sequences are calculated based on the HLA typing data to obtain MHC affinity, binding half-life, and presentation rank data. The candidate mutant peptide sequences are compared with an immunogenic epitope database to obtain antigenic similarity data. The candidate mutant peptide sequences are also compared with an autoproteome database to obtain autodissimilarity data. Thus, peptide sequences that meet the criteria of class I, II, III, and IV are screened from the general candidate sequences.
[0045] According to a specific embodiment of the present invention, in the above method, a total of Q polypeptide sequences conforming to Class I, Class II, Class III, and Class IV are screened. When Q is less than the number of target polypeptide sequences, polypeptide sequences with high MHC affinity can be selected from general candidate sequences to supplement the shortlisted sequences, so that the number of shortlisted sequences is greater than or equal to the number of target polypeptide sequences.
[0046] This invention also provides the application of the above-described method for screening tumor neoantigens in the preparation of antitumor drugs or vaccines. The tumor neoantigens screened by the method according to this invention can be used to prepare antitumor drugs or vaccines.
[0047] Definitions of abbreviations and key terms:
[0048] Neoantigens: Self-antigens produced by tumor cells due to genomic mutations.
[0049] Machine learning: Automating the analysis and pattern recognition of data by building and training models.
[0050] Clonality: The prevalence of a specific mutation in a population of tumor cells, i.e., the proportion of tumor cells in which the mutation is present.
[0051] Dominant clonality: The dominant clone in a tumor or cell population; these mutations may be the primary drivers of tumorigenesis.
[0052] aa: Length of an amino acid unit
[0053] Fusion: Gene Fusion
[0054] The method of the present invention has the following advantages:
[0055] 1. Fusion was identified as a candidate neoantigen by RNA-SEQ, which showed high differences from its own protein.
[0056] 2. Based on a self-built mutation priority database, mutations with a high probability of master clonalization are further covered.
[0057] 3. It comprehensively considers four dimensions: antigen binding, immune presentation, self-similarity, and antigen similarity, covering the entire process of abnormal proteins from peptide processing to final recognition by T cells.
[0058] 4. Based on our self-built immunogenic epitope database and self-proteome database, we analyzed two new indicators: self-similarity and antigen-similarity, and included a large number of novel neoantigens that are easily overlooked by other indicators.
[0059] 5. In addition to the common self-peptides in the database, this invention can also compare wild-type peptides and mutant peptides in sequencing samples, removing sequences similar to wild-type peptides at other positions after mutation, thus reducing false positives.
[0060] 6. This invention can use machine learning models to predict antigen binding and immune presentation indicators. Furthermore, to overcome the inherent data bias of machine learning, antigen binding can simultaneously employ both MHC affinity and binding stability models. Immune presentation can utilize models that comprehensively predict both antigen processing and antigen presentation.
[0061] 7. Each mutation simultaneously predicts both MHC class I and MHC class II, increasing the probability of immune peptide hits.
[0062] 8. The proportion of four types of neoantigens in the target polypeptide sequence can be controlled. By appropriately distributing the proportions of the four types of neoantigens, the diversity of screened tumor neoantigens can be maximized. According to the strength required for the pathway to take effect, the screened tumor neoantigens can be combined to ensure that each feasible presentation process has sufficient tumor neoantigen coverage to cope with the highly individualized immune environment of patients, such as a certain pathway blockage or immune infiltration, so as to achieve better immune effects. Attached Figure Description
[0063] Figure 1 is a flowchart of the tumor neoantigen screening method of the present invention.
[0064] Figure 2 is a comparison experiment diagram of mutation detection in Example 1.
[0065] Figure 3 shows the CT26 immunogenicity prediction experiment in Example 2.
[0066] Figure 4 shows the experimental results for BALB_CT26 in BioNTech. Detailed Implementation
[0067] In order to provide a clearer understanding of the technical features, objectives and beneficial effects of the present invention, the technical solution of the present invention will now be described in detail below, but it should not be construed as limiting the scope of implementation of the present invention.
[0068] This invention provides a method for screening tumor neoantigens and its application, as shown in Figure 1, which is a flowchart of the tumor neoantigen screening method of this invention.
[0069] This invention provides a method for screening tumor neoantigens, the method comprising:
[0070] We analyzed the WES data of tumor samples and normal samples from the individuals to be tested to obtain candidate SNV / INDEL data.
[0071] RNA-seq data from tumor samples were analyzed to obtain candidate fusion data;
[0072] The candidate SNV / INDEL data and candidate fusion data are filtered out, wherein low-expression mutations are filtered out to obtain the filtered mutation sites; the mutation sites are calculated and the sequences containing the mutation sites are extracted to obtain candidate mutant peptide sequences.
[0073] Candidate mutant polypeptide sequences are screened based on antigen binding, immune presentation, antigen similarity, and self-similarity to obtain target polypeptide sequences, from which neoantigens are selected.
[0074] In the above method, preferably, the process of screening candidate mutant polypeptide sequences includes:
[0075] From the candidate mutant polypeptide sequences, polypeptide sequences with MHC affinity lower than 2500 nM and binding half-life greater than 0.5 h were selected as general candidate sequences.
[0076] From the general candidate sequences, polypeptide sequences that meet at least one of the following categories are selected:
[0077] Type I: MHC affinity less than 100 nM, binding half-life greater than 1 hour
[0078] Category II: Presentation rank less than 2%
[0079] Category III: Antigen similarity higher than 80%
[0080] Category IV: Individual differences exceeding 80%
[0081] Selected polypeptide sequences that meet at least one of the above categories are included as shortlisted sequences. The shortlisted sequences are sorted as a whole according to their clonalness and mutation abundance, or sorted separately according to categories I-IV. Target polypeptide sequences with higher clonalness and mutation abundance are selected.
[0082] In one specific embodiment of the invention, mutations can be annotated with master clonal characteristics based on the results of clonal analysis. Furthermore, mutations appearing in a mutation priority database are more likely to be master clonal and are present in most subclonal cells.
[0083] In one specific embodiment of the present invention, the MHC binding assay includes two evaluation perspectives: affinity and stability. The software NetMHCIIpan v4.0, used to predict MHC affinity, and the software Netmhcstabpan v1.0, used to predict peptide binding stability and calculate the binding half-life, were employed. Regarding immunogenicity, the immunopeptide processing was considered, and the immunopresenting performance of the mutant peptide compared to the native peptide was calculated using MHCflurry 2.0.0.
[0084] In one specific embodiment of the present invention, data from the public database IEDB and related literature are further filtered, classified, and summarized to obtain a self-built immunogenic epitope database and an autoproteome database. Based on these two databases, the similarity of candidate mutant peptide sequences to known antigens and their differences from their own antigens can be calculated, predicting the tendency of candidate mutant peptide sequences in the human immune environment.
[0085] In one specific embodiment of the present invention, candidate mutant peptide sequences are evaluated immunologically along four dimensions: antigen binding, immune presentation, antigen similarity, and intrinsic similarity. These dimensions cover the entire process of abnormal proteins from peptide processing to final recognition by T cells. Immune presentation starts from the antigen peptide cleavage process, increasing the weight of epitopes that are easily cleaved and presented. The antigen binding process is jointly evaluated from two perspectives: affinity prediction and peptide binding stability. Immunopresentation then compares the target antigen with other antigens, assessing and scoring the probability of the target antigen being presented. Antigen similarity is a new application of cross-recognition. The human body contains a large number of memory T cells containing pathogen-derived epitopes. MHC molecules have a certain degree of degeneracy; if the contact residues of a neoantigen are similar to the derived epitopes, it is easier to activate these memory T cells, and the secondary immune response induced by memory T cells is stronger. Intrinsic similarity assesses the impact of negative selection in the immune system. Mutations that are too similar to common non-epitopes may cause T cells with corresponding epitopes to be incapacitated due to negative selection; therefore, neoantigens need to exclude this influence to improve the actual response rate.
[0086] In one specific embodiment of the present invention, after calculating the four dimensions of candidate mutant polypeptide sequences, the candidate mutant polypeptide sequences are classified in conjunction with the epitope information of MHCI and II. Since patients actually have relatively few neoantigens, but numerous evaluation indicators, candidate mutant polypeptide sequences only need to be excellent in specific dimensions to trigger the corresponding immune presentation pathway. Candidate mutant polypeptide sequences with MHC affinity below 2500 nM and a binding half-life greater than 0.5 h are initially included in the general candidate neoantigen category. Then, based on whether they meet the criteria, they are divided into four categories of neoantigens. When multiple criteria are met, they are preferentially assigned to the higher-numbered group.
[0087] Generally, in classifying candidate sequences, the higher the rank number, the higher the priority; that is, class IV has the highest priority, and so on. When a general candidate sequence satisfies different classification criteria simultaneously, it should be prioritized for inclusion in the higher priority group.
[0088] In one specific embodiment of the present invention, if the total number of the four types of antigens is less than 20, the general candidate sequences should be sorted according to MHC affinity, and the general candidate sequences with high affinity should be selected first to make up the 20 neoantigens.
[0089] Example 1: Experimental Validation of Mutation Detection (Data Analysis Validation)
[0090] The accuracy of candidate SNV / INDEL data was verified by analyzing the WES data of tumor samples and normal samples of the test individuals using the gold standard dataset of the Sequencing Quality Control Phase 2 (SEQC2) Consortium.
[0091] WES data of tumor samples and normal samples from the individuals to be tested were preprocessed. Trimmomatic-0.39 was used to remove low-quality reads and sequencing adapters to obtain qualified reads after filtering.
[0092] The filtered qualified reads were aligned with the human reference genome hg38 using the alignment software bwa-0.7.17 to generate a bam file.
[0093] Samtools-1.15.1 was used to sort the BAM files, mark repetitive sequences, and recorrect base quality to generate high-quality BAM files for subsequent mutation detection.
[0094] The high-quality BAM file from the previous step was subjected to mutation detection using gatk-4.2.6.1 to generate the original VCF file; then, gatk filter was used to filter and annotate the original VCF file.
[0095] The accuracy was assessed by comparing the filtered and annotated VCF with the gold standard dataset in the reference (Kreiter, S., Vormehr, M., van de Roemer, N. et al. Mutant MHC class II epitopes drive therapeutic immune responses to cancer. Nature 520, 692–696 (2015)).
[0096] Figure 2 shows the F-score quantification results of three pairs of HCC1395 samples. In the SNV validation, two pairs of samples outperformed the literature results, and in the INDEL validation, one pair of samples outperformed the literature results, proving the authenticity and accuracy of the mutation process.
[0097] Example 2:
[0098] Immunogenicity assay validation (full-process experimental and data analysis validation)
[0099] Colorectal cancer tumor samples and normal whole blood samples were collected from one BALB_CT26 mouse. DNA and RNA were extracted using any existing method. The tumor tissue was then subjected to WES+RNA-SEQ sequencing using an Illumina sequencing platform or another next-generation paired-end sequencing platform. The normal whole blood samples were subjected to WES paired-end sequencing. This yielded WES data from the tumor samples, RNA-seq data, and WES data from the normal samples.
[0100] The sequencing data were preprocessed using Trimmomatic-0.39 to remove low-quality reads and sequencing adapters.
[0101] DNA data was aligned using bwa-0.7.17 alignment software, and RNA data was aligned using STAR 2.7.11b. The filtered qualified reads were then aligned with the mouse reference genome GRCm38 sequence to generate a bam file.
[0102] The bam file was sorted, repetitive sequences were marked, and base quality recorrection (BQSR) was performed using samtools-1.15.1 to generate a high-quality bam file for subsequent mutation detection.
[0103] The high-quality BAM file from the previous step was subjected to mutation detection using gatk-4.2.6.1 to generate the original VCF file; then, gatk filter was used to filter and annotate the original VCF file.
[0104] BALB_CT26 is an inbred mouse strain, and its MHC subtype is known and stable, so no further typing prediction is needed.
[0105] Based on WES data, tumor purity was calculated using Sequenza 2.2.0, and clonal analysis was performed.
[0106] Gene fusion detection, annotation, and filtering were performed using RnaFusion 1.0 based on RNA data.
[0107] Transcript abundance and mutation abundance were calculated using kallisto1.0 based on RNA data.
[0108] Mutations are summarized and low-expression mutations are filtered out, and peptides carrying candidate mutations are calculated.
[0109] Using NetMHCIIpan v4.0, Netmhcstabpan v1.0, MHCflurry 2.0.0, as well as immunogenic epitope databases and autoproteome databases, various indicators of immunogenic peptides were predicted.
[0110] Peptide sequences with MHC affinity below 2500 nM and binding half-life greater than 0.5 h were selected as general candidate sequences.
[0111] Peptide sequences that meet at least one of the following categories are selected from general candidate sequences:
[0112] Type I: MHC affinity less than 100 nM, binding half-life greater than 1 hour
[0113] Category II: Presentation rank less than 2%
[0114] Category III: Antigen similarity higher than 80%
[0115] Category IV: Individual differences exceeding 80%
[0116] The selected peptide sequences that meet at least one of the above categories are used as shortlisted sequences. These shortlisted sequences are then sorted according to their clonalness and mutation abundance, and sequences with higher clonalness and mutation abundance are selected as target peptide sequences. In this embodiment, the number of predetermined target peptide sequences is 20. Since there are fewer than 20 neoantigens in the four major classes of BALB_CT26, sequences from the general candidate sequences (Group denoted by N) are sorted according to MHC affinity, with priority given to sequences with high affinity. Four sequences with high MHC affinity are included for testing. The sequences and groups are shown in Table 1.
[0117] Table 1
[0118] According to the target polypeptide sequence number in Table 1, the target polypeptide sequences are sequentially tandem, with the universal linker GGSGGGGSGG (SEQ ID No. 21) as the separator, to construct the corresponding mRNA neoantigen vaccine.
[0119] Four mice were administered the drug via intramuscular injection on day 1 (D1) and a second injection on day 14. On day 21, spleen cells from the immunized mice were harvested for ELISpot assays. After isolating the mouse spleen cells, they were divided into 20 portions to verify the effects of 20 neoantigen epitopes. ELISpot plating and incubation were performed, followed by cell lysis, plate washing, and detection of antibody incubation effects and color development.
[0120] The ELISpot experimental results are shown in Figure 3. The results showed that among the 20 target peptide sequences tested, using the triggering of an immune response in each mouse as the standard for immunogenicity, 6 target peptide sequences exhibited immunogenicity based on the criterion of immune response in all four replicates, representing an immunogenicity rate of 30%. This rate is higher than the 21% measured by BioNTech on BALB_CT26, as shown in Figure 4. Among them, neoantigen M11 was selected using the new indicator of antigen similarity screening, and its antigen similarity ranked first among class III high-level neoantigens. The immunogenicity of this neoantigen was significantly higher than that of other neoantigens, demonstrating that this neoantigen likely activated previously existing memory T cells through the degeneracy of the TCR.
[0121] Example 3
[0122] The steps of sorting the shortlisted sequences in Example 2 according to antigen binding, immune presentation, antigen similarity and self-similarity include analyzing the epitopes of MHCI and MHCII at the mutation sites of the shortlisted sequences, scoring the shortlisted sequences according to the epitope information, obtaining the total score of each sequence, ranking the scores, and selecting the top-ranked polypeptide sequences.
[0123] The total score for each polypeptide sequence is calculated using the following formula: Polypeptide sequence score = MHCI antigen binding score + MHCI immune presentation score + MHCII antigen binding score + MHCII immune presentation score + max{antigen similarity score, self-similarity score}. Each polypeptide has an MHCI antigen binding score, an MHCI immune presentation score, an MHCII antigen binding score, an MHCII immune presentation score, an antigen similarity score, and a self-similarity score, where:
[0124] 1 point is awarded for MHC I affinity below 100 nM and binding half-life greater than 1 h; 1 point is awarded for MHC II affinity below 100 nM and binding half-life greater than 1 h; 2 points are awarded for MHC I immune presentation rank less than 2%; 2 points are awarded for MHC II immune presentation rank less than 2%; 3 points are awarded for antigen similarity greater than 80%; and 4 points are awarded for self-similarity greater than 80%.
[0125] Taking the target polypeptide sequence 11 of CT26 as an example, it has 1 point for binding to MHC class I antigens, 1 point for binding to MHC class II antigens, 2 points for MHC class I immunopresentation, 2 points for MHC class II immunopresentation, and 3 points for antigen similarity higher than 80%. Therefore, its score is 1+1+2+2+3=9.
Claims
1. A method for screening tumor neoantigens, the method comprising: We analyzed the WES data of tumor samples and normal samples from the individuals to be tested to obtain candidate SNV / INDEL data. RNA-seq data from tumor samples were analyzed to obtain candidate fusion data; The candidate SNV / INDEL data and candidate fusion data are filtered out, wherein low-expression mutations are filtered out to obtain filtered mutation sites; the filtered mutation sites are calculated and the sequences containing the mutation sites are extracted to obtain candidate mutant peptide sequences. Candidate mutant polypeptide sequences are screened based on antigen binding, immune presentation, antigen similarity, and self-similarity to obtain target polypeptide sequences, from which neoantigens are selected.
2. The method according to claim 1, wherein, The process of screening candidate mutant peptide sequences includes: From the candidate mutant polypeptide sequences, polypeptide sequences with MHC affinity lower than 2500 nM and binding half-life greater than 0.5 h were selected as general candidate sequences. From the general candidate sequences, polypeptide sequences that meet at least one of the following categories are selected: Type I: MHC affinity less than 100 nM, binding half-life greater than 1 hour Category II: Presentation rank less than 2% Category III: Antigen similarity higher than 80% Category IV: Individual differences exceeding 80% Selected polypeptide sequences that meet at least one of the above categories are included as shortlisted sequences. The shortlisted sequences are sorted as a whole according to their clonalness and mutation abundance, or sorted separately according to categories I-IV. Sequences with higher clonalness and mutation abundance are selected as target polypeptide sequences. Alternatively, sequences are sorted according to their comprehensive performance in antigen binding, immune presentation, antigen similarity, and self-similarity. Sequences with higher comprehensive performance are selected as target polypeptide sequences.
3. The method according to claim 1, wherein, The purity of the tumor sample should be greater than or equal to 10%, preferably greater than or equal to 40%.
4. The method according to claim 1, wherein, When filtering the candidate SNV / INDEL data and candidate fusion data, low-expression mutations are somatic mutations with an FPKM of less than 5 in transcript abundance data and / or a mutation frequency of less than 10% in DNA / RNA in mutation abundance data.
5. The method according to claim 1, wherein, The filtering of the candidate SNV / INDEL data and candidate fusion data also includes filtering in conjunction with a mutation priority database.
6. The method according to claim 1, wherein, The peptide length of the candidate mutant polypeptide sequence is 17aa-29aa.
7. The method according to claim 1, wherein, The method also includes: analyzing WES data of tumor samples and normal samples to obtain clonal analysis data; RNA-seq data from tumor samples were analyzed to obtain mutation abundance data and transcript abundance data; Therefore, the screened polypeptide sequences are sorted according to their clonalness and mutation abundance.
8. The method according to claim 1, wherein, The method also includes: HLA typing data of the individual to be tested is obtained, and then the candidate mutant peptide sequence is calculated in combination with the HLA typing data to obtain data on MHC affinity, binding half-life and presentation rank. The candidate mutant polypeptide sequences are compared with those in an immunogenic epitope database to obtain antigen similarity data of the candidate mutant polypeptide sequences. The candidate mutant polypeptide sequences are compared with their own proteome database to obtain the self-dissimilarity data of the candidate mutant polypeptide sequences. Thus, polypeptide sequences that conform to Class I, Class II, Class III, and Class IV are screened from general candidate sequences.
9. The method according to claim 1, wherein, Q peptide sequences that meet the criteria of Class I, II, III, and IV are selected. If Q is less than the number of target peptide sequences, peptide sequences with high MHC affinity are selected from general candidate sequences to supplement the shortlisted sequences, so that the number of shortlisted sequences is greater than or equal to the number of target peptide sequences.
10. The use of the method for screening tumor neoantigens according to any one of claims 1-9 in the preparation of antitumor drugs or vaccines.
Citation Information
Patent Citations
Method and device for recognizing tumor neoantigen on multi-omics level based on high-throughput sequencing technology and computer readable storage medium
CN116825188A
Precise screening method for personalized neoantigen epitopes of tumors
CN118028442A