A method and system for detecting and screening tumor neoantigens by combining molecular genomics and computational structure

By combining molecular omics and computational structure analysis, the three-dimensional structure of the TCR-pMHC protein complex was predicted, solving the problem of unpredictable binding affinity and stability of TCR and pMHC in existing technologies. This improved the accuracy and efficiency of tumor neoantigen detection and enhanced the therapeutic effect of personalized tumor vaccines.

CN114333999BActive Publication Date: 2026-02-06SHANGHAI PUDAI BIOTECHNOLOGY PARTNERSHIP (LLP)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111462256.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-04
Filing Date
2021-12-02
Publication Date
2026-02-06
Estimated Expiration
2041-12-02

AI Technical Summary

Technical Problem

Existing technologies cannot effectively predict the affinity and conformational stability of T-cell receptor (TCR) binding to tumor neoantigen-major histocompatibility complex (pMHC), resulting in insufficient accuracy and efficiency of tumor neoantigen detection methods.

Method used

Using a method combining molecular omics and computational structure, the three-dimensional structure of the TCR-pMHC protein complex was predicted using whole-genome and transcriptome sequencing data. This prediction was then validated by mass spectrometry, and tumor neoantigens with high expression, stability, and high affinity were screened out.

Benefits of technology

This enables accurate prediction of the binding affinity and conformational stability of TCR and pMHC, improving the accuracy and efficiency of tumor neoantigen detection and enhancing the therapeutic effect of personalized tumor vaccines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333999B_ABST
    Figure CN114333999B_ABST
Patent Text Reader

Abstract

The application discloses a tumor neoantigen detection screening method combined with molecular genomics and calculation structure. The tumor neoantigen detection method disclosed by the application can quickly and efficiently detect tumor neoantigens with high expression, high affinity and stable binding to HLA. The application also discloses a method for predicting the three-dimensional structure of TCR-pMHC protein, which evaluates the affinity and stability of TCR, tumor neoantigen and MHC from the aspect of protein structure, and integrates the affinity and stability prediction of TCR and pMHC, antigen peptide expression translation, hydrophobicity and mutation mode into the detection and analysis process of tumor neoantigen. The application further proposes to use mass spectrometry to verify whether the screened neoantigen really exists in cancer patients, and to evaluate the structural stability of the antigen peptide in the human body. The application further proposes a method and system for predicting the immune cell type and proportion of tumor microenvironment.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present application takes the invention patent application with the application date of December 4, 2020 and the application number of 202011398320.9 as the priority basis. TECHNICAL FIELD

[0002] The present application belongs to the field of tumor neoantigen detection and the field of protein three-dimensional structure prediction, and specifically relates to a tumor neoantigen detection method based on omics detection and a tumor neoantigen screening method for TCR-pMHC protein three-dimensional structure prediction. BACKGROUND

[0003] Tumor neoantigen, also known as tumor-specific antigen (TSA), refers to a polypeptide fragment specific to tumor cells, which can specifically bind to major histocompatibility complex (MHC) and T cell receptor (TCR). Research in the last century showed that tumor cells can specifically express some short peptides, which can be combined and presented by MHC, and these short peptides are tumor neoantigens. In the 1990s, Boon et al. found that tumor neoantigens can be recognized by CD8+ or CD4+ T cells, and TCR, tumor neoantigens and MHC play an immune effect by forming a ternary complex.

[0004] In recent years, immunotherapy represented by CAR-T and immune checkpoint inhibitors has achieved good efficacy in tumor treatment. As an immunotherapy, personalized tumor vaccines based on tumor neoantigens specific to cancer patients have also made great progress in recent years. In 2016, the Rosenberg research team published important research results in the New England Journal of Medicine. They designed tumor neoantigens for KRAS gene G12D mutation and returned lymphocytes expressing these tumor antigens to cancer patients, which relieved the symptoms of the patients. Catherine J. Wu research team published articles in Nature magazine in 2017 and 2019, respectively, reporting their research results on the treatment of melanoma and malignant glioma with tumor vaccines based on tumor neoantigens. Among them, there were cases of cancer no longer recurring in melanoma patients receiving tumor vaccine treatment. Although all malignant glioma patients receiving tumor vaccine treatment died of complications, the survival time of some patients was effectively prolonged, which showed that personalized tumor vaccines had good clinical effects on inhibiting tumor development.

[0005] In the last decade, the efficiency and accuracy of detecting tumor neoantigens have been greatly improved with the wide application of high-throughput sequencing and machine learning in biomedical field. So far, several computer programs for screening tumor neoantigens have been developed, including pVACseq, MuPeXI, Tlminer, OpenVax, NeoEpiScope, EpiSeq and CloudNeo. However, these computer programs can only predict the binding affinity of antigen peptide and MHC, and do not involve the binding of pMHC (complex of antigen peptide and MHC) and TCR. Since the TCR gene of each person is rearranged, there are thousands of TCRs in each human body, resulting in no fixed binding mode between TCR and pMHC. Therefore, the machine learning method based on protein sequence is not suitable for predicting the affinity of TCR and pMHC.

[0006] Currently, there is no method for predicting the binding affinity and conformational stability of TCR and pMHC for detecting tumor neoantigens in the market. Meanwhile, the characteristics of antigen peptides, such as hydrophobicity, mutation sites, length and amino acid distribution, are not fully utilized in the screening process. SUMMARY

[0007] To overcome the defects in the prior art, the present application provides a tumor neoantigen detection method based on second-generation genome sequencing, which can predict the binding affinity of antigen peptide and MHC, and can also predict the three-dimensional structure of T cell receptor-tumor neoantigen-major histocompatibility complex (TCR-pMHC) protein complex. The predicted TCR-pMHC protein structure is scored to evaluate the binding affinity and / or conformational stability of TCR and pMHC from the perspective of protein three-dimensional structure, which helps to detect tumor neoantigens from multiple angles.

[0008] The present application provides a complete set of tumor neoantigen detection and screening method, i.e. a tumor neoantigen detection and screening method based on molecular omics technology and TCR-pMHC three-dimensional structure prediction. The method can evaluate the binding affinity of TCR and pMHC by predicting the three-dimensional structure of TCR-pMHC, and for the first time, the prediction of TCR and pMHC binding affinity and / or conformational stability is included in the detection process of tumor neoantigens.

[0009] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:

[0010] The present application provides a tumor neoantigen detection and screening method combining molecular omics and computational structure, which is based on whole genome and / or whole exon and / or transcriptome sequencing data. The method includes the following steps:

[0011] (1) HLA molecule typing step. Before performing tumor neoantigen detection, the type of HLA (human leukocyte antigen, which is the expression product of MHC in the human body) needs to be determined first. The present application uses HLA molecule typing software to predict the most likely 6 HLA molecule types, covering the main subtypes of MHC I and MHC II, and selects 1 or 2 HLA with the highest frequency of occurrence in the local population (characteristic population to which the patient belongs) from the 6 HLA molecules as the final predicted HLA molecule. In an embodiment of the present application, the type of HLA is predicted based on whole genome and / or whole exon, and specifically HLA molecule typing software such as HLAminer and / or Polysolver is used.

[0012] (2) Tumor somatic gene mutation annotation step. Using whole genome and / or whole exon sequencing technology, tumor somatic gene mutations in the human body can be detected and annotated, including point mutations, insertion and deletion mutations, and fusion gene mutations, which need to be annotated on the chromosome, and the acquisition of the annotation results is assisted and verified by transcriptome gene expression detection to determine the expression. In an embodiment of the present application, the method of VEP (Variant Effect Prediction) is specifically used for annotation.

[0013] (3) Gene mutation peptide segment translation step. Translate the mutation nucleic acid sequence into an amino acid sequence, and truncate the amino acid sequence containing the gene mutation into a peptide segment of 8-17 amino acids in length, with 9 peptides as the default priority parameter and 11 peptides as the second priority parameter. In an implementation of the present application, the mutation antigen peptide is truncated to n (n = 8-13) amino acids in length. For an amino acid point mutation, a peptide segment of (2n-1) amino acids in length is truncated from the mutation amino acid sequence, a sliding window of n length is used, and a mutation antigen peptide of n amino acids in length is truncated from the peptide segment of (2n-1) amino acids in length; in an implementation of the present application, the length of the insertion mutation is m amino acids, and for the insertion mutation, a peptide segment of (m+2n-2) amino acids in length is truncated from the mutation amino acid sequence, a sliding window of n length is used, and a mutation antigen peptide of n amino acids in length is truncated from the peptide segment of (m+2n-2) amino acids in length; for a deletion mutation, a peptide segment of (2n-2) amino acids in length is truncated from the mutation amino acid sequence, a sliding window of n length is used, and a mutation antigen peptide of n amino acids in length is truncated from the peptide segment of (2n-2) amino acids in length; for a fusion gene mutation, a peptide segment of (2n-2) amino acids in length is truncated from the mutation amino acid sequence, a sliding window of n length is used, and a mutation antigen peptide of n amino acids in length is truncated from the peptide segment of (2n-2) amino acids in length. The amino acid sequence is processed with reference to the predicted HLA molecule typing matching, and at least one of MHC I and / or MHC II typing is ensured, the MHC I related typing is preferentially truncated to 9 peptides, 10 peptides or 11 peptides, and the MHC II related typing is preferentially truncated to 13 peptides, 14 peptides, 15 peptides or 16 peptides.

[0014] In the present application, the antigen peptide, the mutation antigen peptide, the mutation peptide, etc. refer to a series of immunogenic peptides generated by tumor-specific mutations, i.e. short peptides containing mutation amino acids generated by degradation of tumor-specific mutation genes by proteasomes, which, after binding to MHC class molecules, can be presented to the cell surface, and then recognized by corresponding T cell receptors, thereby activating the activity of cytotoxic T lymphocytes, so as to produce specific anti-tumor immune response in the body.

[0015] (4) Antigen peptide and HLA molecule affinity prediction step. The affinity of all antigen peptides and HLA is predicted using antigen peptide and HLA affinity prediction software. Software including netMHCpan and / or MetaMHCpan and / or PSSMHCpan can be used alone or in combination to predict the affinity of HLA and antigen peptides. In an implementation of the present application, the software netMHCpan is specifically used to predict the affinity of HLA and antigen peptides.

[0016] (5) Antigen peptide expression detection step. Antigen peptides must be expressed in cells to bind to MHC and TCR to form TCR-pMHC complexes and induce immune response reactions in the human body to kill tumor cells. Therefore, antigen peptide expression detection, i.e., antigen peptide gene expression detection, is particularly important. Here, based on transcriptome sequencing data, gene expression calculation software is used to calculate the expression of the antigen peptide gene, representing the expression of the antigen peptide. Specifically, software including HTSeq and / or Salmon can be used to calculate read counts and / or TPM and / or FPKM and / or RPKM of the antigen peptide gene as a measure of the expression level of the antigen peptide. In an implementation of the present application, the software HTSeq is specifically used to calculate the read counts of the antigen peptide gene as a measure of the expression level of the antigen peptide.

[0017] (6) Comprehensive quantitative antigen peptide screening step (antigen peptide screening step based on expression and affinity, etc.). The affinity threshold and expression threshold of the antigen peptide are used as screening criteria to further screen antigen peptides with certain expression and high affinity. In the screening step, antigen peptide hydrophobicity evaluation and amino acid mutation site paradigm are simultaneously included. In the case of similar antigen peptide affinity and expression value scores, antigen peptide hydrophobicity evaluation and amino acid mutation site paradigm are used for screening. In an implementation of the present application, in the screening step, antigen peptide hydrophobicity evaluation is simultaneously included. In the case of similar antigen peptide affinity and expression value scores, antigen peptides with weak overall hydrophobicity are specifically quantitatively measured by the proportion of hydrophobic amino acids, i.e., polypeptides with a lower proportion of hydrophobic amino acids are ranked first. Mutation amino acid sites are also considered. For antigen peptides with similar affinity, expression, and hydrophobicity, antigen peptides with amino acid mutation at position 2 are excluded, and antigen peptides with no mutation at position 2 and / or mutation at position 3 are preferred.

[0018] (7) Antigen peptide structure stability prediction step. The antigen peptides of 8-11 amino acids in length obtained through the above steps need to maintain stable protein structure to ensure that they are not degraded in cells and further combined with MHC and TCR. In this regard, in the antigen peptides remaining after steps (1)-(6) in the screening, the stability is estimated according to the length of the antigen peptide and the distribution of amino acids; at the same time, the antigen peptide structure stability prediction software is used to predict the protein structure stability of the antigen peptide to ensure that they are not degraded in cells and further combined with MHC and TCR. In an implementation mode of the present application, the software NetMHCstab is specifically used to predict the structure stability of the antigen peptide, and the antigen peptide with stronger stability is preferentially retained.

[0019] Among the remaining antigen peptides after the foregoing screening, the 9-mer and / or 10-mer antigen peptides are retained in the case of similar comprehensive scores for HLA-A, HLA-B, and HLA-C molecule binding; the 15-mer and / or 16-mer antigen peptides are retained in the case of similar comprehensive scores for HLA-DP, HLA-DQ, and HLA-DR molecule binding; the antigen peptide with lower relative content of the five amino acids M (Met), W (Trp), C (Cys), G (Gly), and T (Thr) is preferentially retained under the same conditions; in the embodiment of the present application, the MHC class I molecule is mainly used, so the 9-mer antigen peptide is preferentially retained under similar conditions, followed by the 10-mer antigen peptide.

[0020] Further, the method of the present application can further include:

[0021] (8) Immune cell type prediction of tumor immune microenvironment.

[0022] In the tumor immune microenvironment, the types and proportions of immune cells infiltrated therein affect the strength of the immune response induced by tumor neoantigens, and further affect the treatment effect of the personalized tumor vaccine, so it is necessary to predict the immune cell composition. In this regard, the types and proportions of immune cells and stromal cells in the tumor immune microenvironment are predicted based on transcriptome data using immune cell type prediction software. In an implementation mode of the present application, the software xCell or TIMER is specifically used to predict the types and proportions of immune cells and stromal cells, including various immune cells and stromal cells.

[0023] Further, the method of the present application can further include:

[0024] (9) Mass spectrometry verification step of antigen peptide.

[0025] Extracting samples from cancer patients, including blood, tissue, etc. (preferably extracting tissue), ex vivo samples, performing mass spectrometry analysis, using mass spectrometry database search tools, searching for polypeptides matching the peptides identified by mass spectrometry from the theoretical protein database containing mutations. The polypeptides screened with certain expression and high affinity and stability with HLA are compared with the polypeptides identified by mass spectrometry and / or the polypeptides identified by mass spectrometry from the public database, and the polypeptides coinciding with both are selected as the screened antigen peptides; or the polypeptides accumulated in the public database can be used as the reference for mass spectrometry verification, which is suitable for the case of no cancer patient for mass spectrometry detection data. Specifically, tools including but not limited to SEQUEST and / or TPP, etc. can be used for protein database search, and the reference protein database used is the corresponding human proteome version in UniprotKB, and the customization includes antigen peptide related mutations. In an implementation manner of the present application, the tool SEQUEST and / or MaxQuant is specifically used for protein database search, and the protein database used is UniprotKB, including the classic reference proteome and the theoretical spectrum library of antigen peptide related mutations.

[0026] Further, based on whole genome and / or whole exon and / or transcriptome sequencing data and / or TCR sequencing data, the method of the present application can further include:

[0027] (10) TCR-pMHC complex protein structure prediction step.

[0028] Due to the gene rearrangement of TCR genes, there are thousands of TCR protein sequences in the human body, therefore, the pMHC cannot have a clear binding mode with the TCR. By predicting the TCR-pMHC protein structure, the binding of a specific TCR with the pMHC can be observed from the perspective of protein structure, which helps to screen the TCR capable of binding with the pMHC from a series of TCRs. Here, known TCR-pMHC protein structures are searched from the protein structure database PDB (protein data bank) as a protein template library, the most similar protein structure to the input sequence is searched from the template library as a template according to the input TCR-pMHC protein sequence, and the protein structure of the TCR-pMHC is predicted based on the homology modeling method using the protein structure prediction software. In an implementation manner of the present application, the Modeller software is specifically used to predict the protein structure of the TCR-pMHC complex. Meanwhile, the software including but not limited to Rosetta, FlexPepDock, TCRFlexDock, etc. is comprehensively used to disturb and optimize the complex structure.

[0029] Further, the method of the present application can further include:

[0030] (11) New antigen specificity score TCR-pMHC protein structure scoring step.

[0031] For a TCR-pMHC protein sequence, the predicted TCR-pMHC complex protein structures are scored using the TCR-pMHC complex protein structure-based TCR-pMHC complex protein structure scoring function, the TCR-pMHC complex conformation stability and / or TCR sequence similarity matching indicators are fused, in the case of similar affinity scores, the antigen peptides that can improve the TCR-pMHC complex conformation stability and / or the similarity matching degree with the reported activatable TCR are preferentially screened, the predicted structures are sorted according to the scores, and the best protein structure is selected as the protein structure prediction result. In an implementation mode of the present application, the TCR-pMHC binding affinity scoring function published in the literature (hereinafter referred to as Literature 1) Tyler Borrman, Jennifer Cimons, Michael Cosiano, et al. “ATLAS: a database linking binding affinities with structures for wild-type and mutant TCR-pMHC complexes.” Proteins. 2017; 85 (5): 908-916. is used to score the predicted TCR-pMHC protein structure, and the TCR-pMHC complex structure stability score can be fused, and the new antigen TCR sequence similarity evaluation based on the TCR binding experimental data in IEDB and the existing experimental evidence support (using BLAST and other methods, based on BLOSUM62 or PAM250 scoring matrix) can be included. The TCR-pMHC protein structure scoring is optimized.

[0032] Among them, steps (4)-(11) can be reorganized according to the scene, including the combination order and type of steps and the measurement indicators generated by them, and the weight relationship between each other when combining; generally, steps (4)-(7) are used as the basic optional steps, and the remaining steps are used as optimization additions;

[0033] In the method of the present application, the combination and analysis method and idea including but not limited to steps (4)-(11) are also applicable to the screening of tumor neoantigens of point mutations, insertion and deletion mutations and fusion gene mutations.

[0034] The present application also provides a tumor neoantigen detection and screening system based on omics detection and TCR-pMHC protein three-dimensional structure prediction, which comprises the following modules:

[0035] 301) HLA molecule typing module. Using HLA molecule typing software, the scores of different types of HLA molecules are calculated, and the two best scored HLA-A, HLA-B, HLA-C molecules or the two best scored HLA-DP, HLA-DQ, HLA-DR molecules are selected, respectively, and one of the six HLA molecules with the highest frequency in the local population is selected as the final predicted HLA molecule according to the HLA type frequency distribution database.

[0036] 302) Tumor somatic gene variation annotation module. Using gene variation annotation software, the point mutations, insertion and deletion mutations and fusion gene variations detected by whole genome and / or whole exome sequencing are annotated.

[0037] 303) Gene variation peptide translation module. After translating the variation nucleic acid sequence into an amino acid sequence, the amino acid sequence containing the gene variation is truncated into a peptide segment of 8-17 amino acids in length, and reference is made to step (3) in the method described above.

[0038] 304) Antigen peptide and HLA affinity prediction module. Using antigen peptide and HLA affinity prediction software, the affinity of the HLA to all antigen peptides is predicted, and reference is made to step (4) in the method described above.

[0039] 305) Antigen peptide expression prediction module. Using gene expression prediction software, based on transcriptome sequencing data, the expression of the gene where the antigen peptide is located is predicted, and reference is made to step (5) in the method described above.

[0040] 306) Antigen peptide screening module. Using the affinity threshold and expression threshold of the antigen peptide as the screening basis, further screening is carried out to obtain antigen peptides with certain expression and high affinity. In the screening step, the hydrophobicity evaluation of the antigen peptide is also included, in the case of similar antigen peptide affinity and expression score, the antigen peptide with weaker overall hydrophobicity is preferred, here the hydrophobic amino acid proportion is quantitatively measured, i.e. the polypeptide with lower hydrophobic amino acid proportion is ranked first; and the mutation amino acid site will be considered, for antigen peptides with similar affinity, expression and hydrophobicity, the antigen peptide with mutation at position 2 is excluded, and the antigen peptide with no mutation at position 2 and / or mutation at position 3 is preferred.

[0041] 307) antigen peptide protein structure stability prediction screening module. For the high-affinity antigen peptides with certain expression obtained in the previous step, the 9-peptide and / or 10-peptide are retained in the case of similar comprehensive scores for antigen peptides binding to HLA-A, HLA-B, and HLA-C molecules; the 15-peptide and / or 16-peptide are retained in the case of similar comprehensive scores for antigen peptides binding to HLA-DP, HLA-DQ, and HLA-DR molecules; and the antigen peptides with relatively low content of the five amino acids M (Met), W (Trp), C (Cys), G (Gly), and T (Thr) are preferentially retained under the same conditions; then, the protein structure stability of the antigen peptides is predicted using antigen peptide protein structure stability prediction software such as NetMHCstab, and the antigen peptides with stable protein structures are preferentially selected for subsequent screening steps.

[0042] Further comprising, 308) tumor immune microenvironment immune cell type prediction module. The types and proportions of immune cells and stromal cells in the tumor immune microenvironment are predicted using immune cell type prediction software.

[0043] Further comprising, 309) antigen peptide mass spectrometry verification module. The mass spectrometry method is used to verify whether the detected antigen peptides actually exist in the human body, and the antigen peptides supported by mass spectrometry data are selected.

[0044] Further comprising, 310) TCR-pMHC protein structure prediction and scoring module. Taking the TCR-pMHC protein sequence as the input file, searching for the most similar structure as the template from the known TCR-pMHC protein structure, and using the protein structure prediction software to predict multiple TCR-pMHC protein structures based on the homology modeling method. Using the TCR-pMHC protein structure-based TCR-pMHC affinity scoring function, the predicted multiple TCR-pMHC complex protein structures are scored by fusing the TCR-pMHC complex conformational stability and the TCR sequence similarity matching degree (using BLAST and other methods based on BLOSUM62 or PAM250 scoring matrix), the predicted structures are sorted according to the scores, and the best protein structure is selected as the TCR-pMHC protein structure prediction result.

[0045] Currently, the function flow organization is process-oriented, and filtering and screening are performed through logical judgment and quantitative scoring functions. The system implementation of this screening process can also be implemented using other programming technology frameworks, such as antigen peptide affinity, expression value, hydrophobicity, mutation site; TCR-pMHC structure affinity and stability, TCR sequence similarity; antigen peptide mass spectrometry qualitative and quantitative; tumor immune microenvironment characteristics, which can be reorganized using machine learning frameworks, including decision trees, random forests, neural networks, and deep neural networks.

[0046] The application also provides a prediction method of a TCR-pMHC protein three-dimensional structure, which comprises the following steps based on TCR sequencing data and HLA molecule typing, obtaining a TCR-pMHC protein sequence, and based on a TCR-pMHC structure template library constructed according to the sequence:

[0047] Step (1) searching a template with the most similar sequence from the template library based on the HLA typing, antigen peptide and possible TCR sequence information determined through the previous steps;

[0048] Step (2) predicting a plurality of TCR-pMHC protein structures by using homology modeling;

[0049] Step (3) scoring the TCR-pMHC structure by using an affinity scoring function;

[0050] Step (4) evaluating the conformational stability of the TCR-pMHC complex by using a structure model quality;

[0051] Step (5) evaluating the binding quality of the TCR and the antigen peptide based on the TCR sequence similarity matching degree;

[0052] Step (6) selecting the TCR-pMHC protein structure with the highest score according to the comprehensive scores of steps (3)-(5).

[0053] The application also provides a prediction method of immune cell types and proportions of a tumor immune microenvironment based on transcriptome data.

[0054] The application provides a tumor neoantigen detection and screening method and system based on molecular omics detection technology and TCR-pMHC three-dimensional structure prediction, which has the following beneficial effects:

[0055] The tumor neoantigen detection and screening method and system based on omics detection can quickly and efficiently detect tumor neoantigens with high expression and high affinity binding with HLA, and simultaneously consider the antigen peptide binding stability, amino acid hydrophobicity proportion and mutation characteristics.

[0056] The application also evaluates the affinity and stability of TCR, tumor neoantigen and MHC from the aspect of protein structure, and integrates the affinity and stability prediction of TCR and pMHC, antigen peptide expression translation and hydrophobicity and mutation mode into the tumor neoantigen detection and analysis process.

[0057] The prediction method of the TCR-pMHC protein three-dimensional structure evaluates the affinity and conformational stability of TCR, tumor neoantigen and MHC from the aspect of protein structure and TCR sequence similarity, and integrates the affinity and conformational stability prediction of TCR and pMHC into the tumor neoantigen detection process for the first time.

[0058] The tumor neoantigen detection method of this invention screens out antigenic peptides with high expression levels and high affinity for HLA, compares them with peptides from cancer patients identified by mass spectrometry, and selects peptides that overlap between the two as candidate antigenic peptides.

[0059] This invention provides a tumor neoantigen detection method that considers the structural stability of antigenic peptides in the human body and assesses the protein structural stability of antigenic peptides. Furthermore, for antigenic peptides with similar affinity, binding stability, gene expression values, and protein profiles, amino acid hydrophobicity ratio, mutation characteristics, antigenic peptide length, and amino acid distribution patterns are incorporated as screening indicators.

[0060] This invention further uses mass spectrometry to verify whether the screened neoantigens actually exist in cancer patients, assess the structural stability of antigenic peptides in the human body, and predict the types and proportions of immune cells in the tumor microenvironment.

[0061] The tumor neoantigen detection method of this invention addresses the problem of immune cell infiltration and predicts the type and proportion of immune cells in the tumor immune microenvironment. Attached Figure Description

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0063] Figure 1 This is a flowchart of the necessary steps for detecting tumor neoantigens based on molecular omics and the Peptide-MHC complex structure prediction method in the embodiments of the present invention.

[0064] Figure 2 This is a flowchart of the additional steps for detecting tumor neoantigens based on molecular omics and TCR-pMHC complex structure prediction method in this embodiment of the invention.

[0065] Figure 3 This is a flowchart illustrating the specific process of predicting immune cell types in the tumor immune microenvironment according to an embodiment of the present invention.

[0066] Figure 4 This is a flowchart illustrating the specific process for mass spectrometry validation of antigenic peptides.

[0067] Figure 5 This is a flowchart illustrating the specific process for predicting the three-dimensional structure of the TCR-pMHC complex protein in this embodiment of the invention.

[0068] Figure 6 is a structural diagram of a computer system used in the tumor neoantigen detection and TCR-pMHC protein three-dimensional structure prediction method in the embodiments of the present application;

[0069] Figure 7 is a correlation analysis diagram between the predicted affinity value calculated by the affinity prediction function of pMHC and TCR and the experimental value in the TCR-pMHC protein three-dimensional structure prediction method of the present application. DETAILED DESCRIPTION

[0070] The present application will be further described in conjunction with the following specific examples and drawings. The process, conditions, experimental methods, etc. for implementing the present application are the general knowledge and common sense in the art, and the present application does not have special limitations.

[0071] As shown in Figure 1 , the tumor neoantigen detection and screening method based on omics detection technology and TCR-pMHC protein three-dimensional structure prediction of the present application comprises the following steps:

[0072] Step (1): HLA molecule typing.

[0073] Step (2): Tumor somatic gene mutation annotation.

[0074] Step (3): Gene mutation peptide translation.

[0075] Step (4): Antigen peptide and HLA affinity prediction.

[0076] Step (5): Gene expression detection.

[0077] Step (6): Antigen peptide screening according to expression, affinity, and antigen peptide hydrophobicity evaluation and amino acid mutation site paradigm, etc.

[0078] Step (7): Antigen peptide structure stability prediction screening.

[0079] As shown in Figure 2 ,

[0080] Further comprising, step (8): prediction of immune cell types in tumor immune microenvironment (see Appendix Figure 3 for details).

[0081] Further comprising, step (9): mass spectrometry verification of antigen peptide.

[0082] Further comprising, step (10): TCR-pMHC complex protein structure prediction.

[0083] As shown in Figure 4As shown, the mass spectrometry verification method for the antigen peptide of the present invention includes the following key steps:

[0084] (1) Determine whether individual proteomic data of the patient can be obtained.

[0085] (2) If individual proteomic data of patients are available, a customized proteome reference library is constructed, and the library is searched to verify the presence of antigenic peptides.

[0086] (3) If there is no individual patient proteomic data, human peptides accumulated in public databases are used as mass spectrometry verification references.

[0087] Specifically, tumor transcriptome data is subjected to data quality control and standardization to obtain a tumor gene expression profile matrix; combined with a reference dataset of immune cells and their characteristic genes, an immune cell characteristic gene expression profile matrix is ​​constructed; a machine learning model is built and deconvolution is performed; the optimal model is selected for deconvolution to obtain the immune cell type.

[0088] like Figure 5 As shown, the method for predicting the three-dimensional structure of the TCR-pMHC complex protein of the present invention includes the following steps:

[0089] (1) Based on the HLA typing, antigen peptide and possible TCR sequence information determined by the previous steps, search for the template with the most similar sequence from the template library;

[0090] (2) Homology modeling was used to predict the structures of multiple TCR-pMHC proteins;

[0091] (3) Use the affinity scoring function to score the TCR-pMHC structure;

[0092] (4) Use structural model quality comprehensive assessment to evaluate the conformational stability of TCR-pMHC complex.

[0093] (5) Assess the binding quality of TCR and antigen peptide based on TCR sequence similarity matching.

[0094] (6) Based on the comprehensive scoring in steps (3) to (5), select the TCR-pMHC protein structure with the highest score.

[0095] like Figure 6 As shown, this invention provides a tumor neoantigen detection and screening system based on omics detection and the three-dimensional structure of TCR-pMHC proteins. The system includes the following modules:

[0096] 301. HLA Molecular Typing Module.

[0097] 302. Tumor somatic cell gene variation annotation module.

[0098] 303. Gene variant peptide translation module.

[0099] 304. Antigen peptide and HLA affinity prediction module.

[0100] 305. Antigen peptide expression prediction module.

[0101] 306. Antigen peptide screening module according to expression and affinity.

[0102] 307. Antigen peptide protein structure stability prediction screening module.

[0103] 308. Tumor immune microenvironment immune cell type prediction module (optional).

[0104] 309. Antigen peptide mass spectrometry verification module (optional).

[0105] 310. TCR-pMHC protein structure prediction and scoring module (optional).

[0106] The following examples and drawings only further specifically illustrate the present application and cannot be understood as limiting the present application.

[0107] Example 1

[0108] This example uses a gastric cancer patient's tumor sample whole exome and transcriptome data, adopts a tumor neoantigen detection method based on second-generation sequencing, detects tumor neoantigen, and the specific operation is as follows:

[0109] (1) HLA molecule typing

[0110] In this example, the tumor neoantigen detection process CloudNeo is specifically used to complete this step. CloudNeo automatically calls HLA molecule type prediction software to predict the 2 most likely HLA-A, HLA-B, and HLA-C molecules through the tumor sample whole exome bam file. In this example, the 6 predicted HLA molecules are HLA-A*02:774, HLA-A*02:03P, HLA-B*38:02P, HLA-B*39:01P, HLA-C*03:02P, and HLA-C*07:02P. Referring to the HLA molecule type distribution frequency of the Chinese population in the database The Allele Frequency Net Database (website: http: / / allelefrequencies.net ), HLA-C*07:02P is selected as the predicted HLA type.

[0111] (2) Tumor somatic point mutation annotation

[0112] In this embodiment, the step is completed by using the computational pipeline CloudNeo. CloudNeo annotated 3461 point mutations from the vcf file of whole exome using the method of VEP (Variant Effect Prediction), each point mutation can get their chromosome position, gene number, transcript number, nucleic acid substitution, amino acid substitution and nucleic acid sequence information, etc.

[0113] (3) Mutant peptide translation

[0114] In this embodiment, the length of the antigen peptide is set to 9 amino acids, and the step is completed by using the computational pipeline CloudNeo. CloudNeo translates all genes containing mutations into protein sequences, and truncates the fragments containing mutations into antigen peptides of 9 amino acids in length, finally obtaining 44146 antigen peptides containing amino acid point mutations.

[0115] (4) Antigen peptide and HLA affinity prediction

[0116] In this embodiment, the step is completed by using the computational pipeline CloudNeo. For HLA-C*07:02P predicted in step (1), CloudNeo calls NetMHCpan to predict the affinity between HLA and 44146 mutant antigen peptides.

[0117] (5) Antigen peptide expression prediction

[0118] The read counts of the gene where the antigen peptide is located are predicted by HTSeq as the measurement value of the expression of the antigen peptide.

[0119] (6) Antigen peptide screening according to expression and affinity

[0120] The affinity threshold and expression threshold of the antigen peptide are used as the screening basis to further screen the antigen peptide. In this embodiment, for HLA-C*07:02P, if the affinity threshold of the antigen peptide and the HLA is set to <500nM, and the expression threshold of the antigen peptide is set to >0, 107 antigen peptides are screened, these antigen peptides involve multiple cancer-related genes, including tumor suppressor gene TP53, etc.

[0121] (7) Antigen peptide protein structure stability prediction screening

[0122] The antigen peptides obtained in step (6) are subjected to protein structure stability prediction using software NetMHCstab, and under the same conditions, the antigen peptides with relatively low relative content of the five amino acids M (Met), W (Trp), C (Cys), G (Gly), and T (Thr) are preferentially retained, and when the aforementioned characteristic ranking is similar, the priority of the length of the antigen peptides is arranged in descending order as follows: 9 peptides > 10 peptides > 11 peptides > 8 peptides, so as to ensure that the protein structure of these antigen peptides has good stability.

[0123] Further comprising, (8) immune cell type prediction of tumor immune microenvironment

[0124] The type and proportion of immune cells in the tumor immune microenvironment are predicted using immune cell type prediction software xCell, and the results show that immune cells such as memory B cells and CD4+ T cells exist in the tumor immune microenvironment.

[0125] The application also provides an optional step, (9) mass spectrometry verification of antigen peptides (optional). That is, whether the tumor neoantigens detected using the above steps actually exist in the human body is verified using mass spectrometry.

[0126] According to the tumor neoantigen detection method of the application, for a specific gastric cancer sample, the type of HLA is predicted to be HLA-C*07:02P, 3461 somatic point mutations are detected, 44146 antigen peptides with a length of 9 amino acids are generated according to the point mutations, and their expression amounts and their affinities with HLA are predicted. According to the affinity threshold value of 500nM and the expression threshold value of read counts of 0, 107 antigen peptides (affinity less than 500nM and read counts>0) are screened. Finally, the protein structure stability of these antigen peptides is predicted, and the type and proportion of immune cells in the tumor immune microenvironment are predicted. Optionally, whether these neoantigens actually exist in the body of the cancer patient can be verified using mass spectrometry.

[0127] Example 2

[0128] In addition to the embodiment 1 of the application, the application also provides a TCR-pMHC protein three-dimensional structure prediction method, which is helpful for screening TCRs that bind to pMHC. The TCR-pMHC protein structure prediction method of the application is based on the existing TCR-pMHC protein structure, and specifically uses software Modeller to predict the protein structure, uses the affinity prediction function ATLAS score mentioned in document 1 to predict the affinity of pMHC and TCR based on the TCR-pMHC protein structure. For example, Figure 7As shown, using the data set in document 1, the correlation analysis of the predicted value and the experimental value of the affinity of pMHC and TCR shows that the correlation reaches 0.42, which is obviously higher than the method of predicting the affinity of protein and protein based on protein structure commonly used, and document 1 mentions that the correlation of the common method is from-0.18 to 0.3. These results show that the TCR-pMHC protein structure prediction method of the present application has certain reliability, which helps to detect tumor neoantigens capable of simultaneously binding MHC and TCR.

[0129] The above two embodiments of the present application can be performed alone or in combination, and when the present application is used, one embodiment or a combination of the two embodiments can be selected as appropriate. The present application provides a tumor neoantigen detection and screening method based on omics detection and TCR-pMHC protein three-dimensional structure prediction, and provides a corresponding computer program, i.e., a system, which helps to quickly and efficiently detect tumor neoantigens in the human body and obtain suitable matching MHC and TCR information, and provides an innovative idea and a practical tool for the development of tumor personalized vaccines.

[0130] The protection scope of the present application is not limited to the above embodiments. Changes and advantages that can be thought of by those skilled in the art without departing from the spirit and scope of the present application are included in the present application, and are protected by the appended claims.

Claims

1. A method for screening tumor neoantigens by combining molecular omics and computational structure, characterized in that, The method is based on whole genome and / or whole exon and / or transcriptome sequencing data, comprising the following steps: Step (1): HLA molecule typing step: based on whole genome and / or whole exon data, the type of HLA molecule is predicted; Step (2): tumor somatic gene mutation annotation; Step (3): gene mutation peptide translation, obtaining variant antigen peptide; Step (4): affinity prediction of the antigen peptide and HLA molecule; Step (5): antigen peptide gene expression detection based on transcriptome sequencing data; Step (6): based on the affinity prediction results of step (4) and the antigen peptide gene expression detection results of step (5), and taking into account the antigen peptide hydrophobicity evaluation and amino acid mutation site paradigm, comprehensive quantitative antigen peptide screening is carried out; In step (6), when comprehensive quantitative antigen peptide screening is carried out, the affinity threshold and expression threshold of the antigen peptide are used as the screening basis; in the screening step, the antigen peptide hydrophobicity evaluation and amino acid mutation site paradigm are also taken into account, and in the case of similar antigen peptide affinity and expression value scores, screening is carried out according to the antigen peptide hydrophobicity evaluation and amino acid mutation site paradigm; In the antigen peptide screening, the affinity threshold and expression threshold of the antigen peptide are used as the screening basis, and antigen peptides with higher expression and higher affinity are further screened; and in the case of similar antigen peptide affinity and expression value scores, antigen peptides with weaker overall hydrophobicity and / or mutation sites meeting the fixed paradigm are further screened, where the hydrophobicity is evaluated by scoring or the proportion of hydrophobic residues; the mutation site paradigm focuses on amino acids at positions 2 and 3; Step (7): antigen peptide structure stability prediction screening based on the antigen peptide screening results of step (6); In step (7), when antigen peptide structure stability prediction screening is carried out, among the antigen peptides remaining after steps (1)-(6), the stability is estimated according to the length and amino acid distribution of the antigen peptide; at the same time, the protein structure stability of the antigen peptide is predicted using antigen peptide structure stability prediction software to ensure that it is not degraded in the cell and is further combined with MHC and TCR; the software NetMHCstab is used to predict and screen the structure stability of the antigen peptide, and the antigen peptides with strong stability are retained; Among the remaining antigen peptides after screening, for antigen peptides combined with HLA-A, HLA-B, and HLA-C molecules, 9-mer and / or 10-mer peptides are retained under the condition of similar comprehensive scores; for antigen peptides combined with HLA-DP, HLA-DQ, and HLA-DR molecules, 15-mer and / or 16-mer peptides are retained under the condition of similar comprehensive scores; further, under the same conditions, antigen peptides with relatively low content of the five amino acids M (Met), W (Trp), C (Cys), G (Gly), and T (Thr) are retained.

2. The method of claim 1, wherein, The method is suitable for screening of tumor neoantigens with point mutations, insertion-deletion mutations, and fusion gene mutations.

3. The method of claim 1, wherein, In the step (1), the method for typing HLA molecules comprises: predicting 6 HLA molecule types covering the main subtypes of MHC I and MHC II; selecting 1 or 2 HLA with the highest frequency in the patient's characteristic population from the 6 HLA molecules as the final predicted HLA molecules by referring to the HLA type frequency distribution database; and using HLA molecule typing software, including HLAminer and / or Polysolver to predict HLA molecule types.

4. The method of claim 1, wherein, In the step (2), the tumor somatic gene mutations detected by using whole genome and / or whole exon sequencing technology are annotated, including point mutations, insertion and deletion mutations and fusion gene mutations, and the gene mutations are annotated on the chromosome, and the annotation results are verified by transcriptome gene expression detection.

5. The method of claim 1, wherein, In the step (3), the variant nucleic acid sequence is translated into an amino acid sequence, the amino acid sequence containing the gene mutation is truncated into a peptide segment with a length of 8-17 amino acids, the variant antigen peptide with a length of n (n=8-17 amino acids) is obtained, for the amino acid point mutation, 2n-1 amino acid length peptide segments are obtained from the variant amino acid sequence by extending n-1 amino acids forward and backward respectively from the point mutation, and the 2n-1 amino acid length peptide segments are truncated into n amino acid length mutant antigen peptides by using a sliding window with a length of n; for the insertion mutation, the length of the insertion mutation is m amino acids, for the insertion mutation, m+2n-2 amino acid length peptide segments are obtained from the variant amino acid sequence by extending n-1 amino acids forward and backward from the insertion fragment, and the m+2n-2 amino acid length peptide segments are truncated into n amino acid length mutant antigen peptides by using a sliding window with a length of n; for the deletion mutation, 2n-2 amino acid length peptide segments are obtained from the variant amino acid sequence by extending n-1 amino acids forward and backward from the deletion site, and the 2n-2 amino acid length peptide segments are truncated into n amino acid length mutant antigen peptides by using a sliding window with a length of n; for the fusion gene mutation, 2n-2 amino acid length peptide segments are obtained from the variant amino acid sequence by extending n-1 amino acids forward and backward from the fusion site, and the 2n-2 amino acid length peptide segments are truncated into n amino acid length variant antigen peptides by using a sliding window with a length of n.

6. The method of claim 5, wherein, The amino acid sequence is processed by referring to the matching of the predicted HLA molecule typing; the MHC I and / or MHC II typing ensures at least 1; and / or, The MHC I related typing is truncated into 9 peptides, 10 peptides or 11 peptides, and / or the MHC II related typing is truncated into 13 peptides, 14 peptides, 15 peptides or 16 peptides.

7. The method of claim 1, wherein, In the step (4), the software including netMHCpan and / or MetaMHCpan and / or PSSMHCpan is used alone or in combination to predict the affinity of HLA and antigen peptides.

8. The method of claim 1, wherein, In the step (5), based on the transcriptome sequencing data, the expression amount of the gene where the antigen peptide is located is calculated by using a gene expression amount calculation software, and the expression amount of the antigen peptide is represented. The software includes HTSeq and / or Salmon, and the read counts and / or TPM and / or FPKM and / or RPKM of the gene where the antigen peptide is located are calculated as the measurement value for measuring the expression amount of the antigen peptide.

9. The method of claim 1, wherein, Further comprising a step (8) of predicting based on the immune cell types of the tumor immune microenvironment.

10. The method of claim 9, wherein, In the step (8), the types and proportions of immune cells and stromal cells in the tumor immune microenvironment are predicted based on the transcriptome data by using an immune cell type prediction software.

11. The method of claim 1, wherein, Further comprising a step (9) of mass spectrometry verification of the antigen peptide, and verifying whether the detected tumor neoantigen actually exists in the human body by using a mass spectrometry method.

12. The method of claim 11, wherein, In the step (9), the sample of the cancer patient is extracted, including tissue and / or blood, and the sample is taken out of the body for mass spectrometry analysis. A mass spectrometry database search tool is used to search for a polypeptide matching the identified peptide segment by mass spectrometry from a customized protein database containing mutations; or, a human-derived polypeptide accumulated in a public database is used as a reference for mass spectrometry verification, which is suitable for the case of no cancer patient for mass spectrometry detection data; the screened polypeptide with a certain expression amount and high affinity and stability with HLA is compared with the polypeptide identified by mass spectrometry analysis and / or the polypeptide identified by mass spectrometry from the public database, and the polypeptide coinciding with both is selected as the screened antigen peptide.

13. The method of claim 1, wherein, Further comprising a step (10) of predicting the protein structure of TCR-pMHC complex based on HLA molecular typing and TCR sequencing data.

14. The method of claim 13, wherein, In the step (10), known TCR-pMHC protein structures are searched from a protein structure database PDB as a template library, the most similar protein structure to the input sequence is searched from the template library as a template according to the input TCR-pMHC protein sequence, and the protein structure of TCR-pMHC is predicted by using a protein structure prediction software based on the method of homology modeling.

15. The method of claim 1, wherein, Further comprising a step (11) of scoring the TCR-pMHC protein structure.

16. The method of claim 15, wherein, In the step (11), for a TCR-pMHC protein sequence, a TCR-pMHC complex protein structure scoring function based on the TCR-pMHC complex protein structure is used to score a plurality of predicted TCR-pMHC complex protein structures, the predicted structures are sorted according to the scores, and the best protein structure with the highest score is selected as the protein structure prediction result. 17.A tumor neoantigen screening system based on omics detection, tumor neoantigen detection and TCR-pMHC protein three-dimensional structure prediction of tumor neoantigen screening system, characterized in that, The system implements the method of claims 1-16, and the system comprises the following modules: The HLA molecule typing module: using HLA molecule typing software, the scores of different types of HLA molecules are calculated, and the two best HLA-A, HLA-B and HLA-C molecules or the two best HLA-DP, HLA-DQ and HLA-DR molecules are selected according to the scores, and one or two HLA molecules with the highest frequency in the local population are selected from the six HLA molecules according to the HLA type frequency distribution database, which are the final predicted HLA molecules; The tumor somatic gene variation annotation module: using gene variation annotation software, the point mutations, insertion and deletion mutations and fusion gene variations detected by whole genome and / or whole exome sequencing are annotated; The gene variation peptide translation module: after translating the variation nucleic acid sequence into an amino acid sequence, the amino acid sequence containing the gene variation is truncated into a peptide segment with a length of 8-17 amino acids; The antigen peptide and HLA affinity prediction module: using the antigen peptide and HLA affinity prediction software to predict the affinity of the HLA and all antigen peptides; The antigen peptide expression prediction module: using gene expression prediction software, based on transcriptome sequencing data, the expression of the gene where the antigen peptide is located is predicted; The antigen peptide screening module; The antigen peptide protein structure stability prediction and screening module; The tumor immune microenvironment immune cell type prediction module: using immune cell type prediction software, the types and proportions of immune cells and stromal cells in the tumor immune microenvironment are predicted; The antigen peptide mass spectrometry verification module: using mass spectrometry method to verify whether the detected antigen peptide really exists in the human body, and selecting the antigen peptide supported by mass spectrometry data; The TCR-pMHC protein structure prediction and scoring module.

18. The system of claim 17, wherein, The process-oriented function flow organization is adopted, and filtering and screening are carried out through logical judgment and quantitative scoring functions; TCR-pMHC structure affinity and stability, TCR sequence similarity evaluation; antigen peptide mass spectrometry qualitative and quantitative; The tumor immune microenvironment feature prediction is reorganized using a machine learning framework, including decision tree, random forest, neural network and deep neural network.

19. The system of claim 17, wherein, In the antigen peptide screening module, the affinity threshold and expression threshold of the antigen peptide are used as the screening basis to further screen out antigen peptides with certain expression and high affinity. In the screening step, the hydrophobicity of the antigen peptide is also evaluated. In the case of similar antigen peptide affinity and expression, the antigen peptide with weak overall hydrophobicity is selected; wherein the proportion of hydrophobic amino acids is quantitatively measured, and the mutation amino acid site is considered. For antigen peptides with similar affinity threshold, expression and hydrophobicity, and / or, selecting antigen peptides with no mutation at position 2 and / or mutation at position 3.

20. The system of claim 17, wherein, The antigen peptide protein structure stability prediction module uses the antigen peptide protein structure stability prediction software NetMHCstab to predict the protein structure stability for the detected high-affinity antigen peptide with certain expression.

21. The system of claim 17, wherein, The TCR-pMHC protein structure prediction and scoring module uses the TCR-pMHC protein sequence obtained in the foregoing steps as an input file, searches for the most similar structure as a template from known TCR-pMHC protein structures, uses protein structure prediction software to predict multiple TCR-pMHC protein structures based on the homology modeling method, scores the predicted multiple TCR-pMHC complex protein structures using a TCR-pMHC affinity scoring function based on protein structure, fuses TCR-pMHC complex conformational stability and / or TCR sequence similarity matching degree indicators, sorts the predicted structures according to the scores, and selects the best protein structure as the TCR-pMHC protein structure prediction result.

22. A method for predicting the three-dimensional structure of a TCR-pMHC complex protein, characterized in that, The method comprises the following steps: Step (1) searching for the most similar template from the template library based on the determined HLA typing, antigen peptide, and TCR sequence information; In step (1), searching for known TCR-pMHC protein structures as a protein template library from the protein structure database PDB, and searching for the most similar protein structure as a template from the template library according to the input TCR-pMHC protein sequence; Step (2) using protein structure prediction software to predict multiple TCR-pMHC protein structures based on the homology modeling method; Step (3) for the TCR-pMHC protein sequence, using a TCR-pMHC complex protein structure-based TCR-pMHC affinity scoring function; Step (4) using a structure model quality to comprehensively evaluate the conformational stability of the TCR-pMHC complex; For the detected high-affinity antigen peptide with a certain expression amount, using antigen peptide protein structure stability prediction software NetMHCstab to predict the protein structure stability; Step (5) evaluating the binding quality of TCR and antigen peptide based on TCR sequence similarity matching degree; incorporating a new antigen-binding TCR sequence similarity evaluation indicator supported by IEDB TCR binding experimental data and existing experimental evidence to optimize the TCR-pMHC protein structure scoring; Step (6) selecting the highest-scored TCR-pMHC protein structure according to the comprehensive scoring of steps (3) to (5). Step (6) selecting the highest-scored TCR-pMHC protein structure according to the comprehensive scoring of steps (3) to (5).

Citation Information

Patent Citations

  • ECM1 polypeptide epitope capable of combining with human MHC-I molecule

    CN104163852A

  • Tumor neoantigen detection method and device based on next-generation sequencing, and storage medium

    CN108796055A

  • Tumor antigen prediction method based on whole transcriptome, and application of the same

    CN109801678A

  • A method to generate a cocktail of personalized cancer vaccines from tumor-derived genetic alterations for the treatment of cancer

    WO2019036043A2