A method and system for predicting new antigens in malignant solid tumors

The new antigen prediction method of multi-dimensional eigenvalue scoring and product scoring function solves the problem of unsatisfactory new antigen screening results in the prior art, and achieves higher prediction accuracy and reliability.

CN119446264BActive Publication Date: 2025-05-06TIANJIN CANCER HOSPITAL AIRPORT HOSPITAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510034286.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The prior art has poor screening results when predicting neoantigens of malignant solid tumors, mainly due to excessive reliance on single prediction software and hard screening thresholds, resulting in false positive and false negative problems.

Method used

A neoantigen prediction method and system is adopted to obtain multi-dimensional eigenvalues ​​through multiple biological information calculation methods based on the action mechanism of neoantigen, and a product-based neoantigen scoring function is designed to calculate the neoantigen score, and the neoantigen with the highest score is selected as the prediction result.

Benefits of technology

Improve the accuracy and reliability of neoantigens prediction, significantly improve the positive consistency rate with known valid neoantigens, and reduce false positives and false negatives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119446264B_ABST
    Figure CN119446264B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of tumor neoantigen detection, and specifically discloses a method and system for predicting neoantigens of malignant solid tumors, the method comprising: obtaining tumor samples and peripheral blood samples, and performing DNA and RNA sequencing; obtaining DNA sequencing data and RNA sequencing data through analysis and processing; analyzing and obtaining HLA typing results and sequences of mutant proteins; calculating neoantigen-related feature values; calculating neoantigen scores; intercepting the top N with the highest scores to obtain neoantigen prediction results of malignant solid tumors. Starting from the mechanism of action of neoantigens, the present invention selects multiple feature values ​​according to the key steps of the prediction process, does not use rigid filtering indicators, and uses a product method to calculate the neoantigen score according to the scores of each neoantigen-related feature value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tumor neoantigen detection, and in particular to a method and system for predicting neoantigens of malignant solid tumors. Background Art

[0002] Neoantigens are protein fragments produced by mutations in tumor cells that are different from normal cells. These fragments can be presented on the surface of tumor cells by human leukocyte antigens (HLA) molecules to form pMHC (peptide-major histocompatibility complex) complexes. Protein fragments that can be presented by HLA molecules and recognized by T cell receptors (TCRs) to stimulate immune responses are defined as neoantigens.

[0003] Neoantigens are key factors in activating and directing T cells to recognize and attack tumor cells. Because neoantigens are directly generated by mutations in tumor cells, they are highly tumor-specific, meaning they are unlikely to be expressed in normal cells, thus not being affected by pre-existing immune tolerance and unlikely to produce autoimmune reactions.

[0004] The effectiveness of cancer immunotherapy is based on a core premise: the presence of specific T cells inside the tumor that can recognize and respond to functional antigens. Tumor neoantigens are widely regarded as potential therapeutic targets due to their unique tumor specificity. Currently, the research and application of tumor neoantigens are driving the progress of personalized immunotherapy, including strategies such as personalized vaccines and cell therapy. The success of these therapies depends largely on the ability to efficiently, accurately and reproducibly identify new antigens that can stimulate tumor-specific immune responses from available samples.

[0005] There is significant heterogeneity in somatic mutations between cancer patients, and even for the same cancer, the mutation spectrum between different patients is not the same. In addition, the subtypes of human leukocyte antigens show high diversity among individuals. The difficulty in predicting tumor neoantigens using high-throughput sequencing data lies in the need to accurately identify tumor-specific mutations and predict whether these mutations can produce immunogenic neoantigens.

[0006] Based on the TESLA (Tumor Epitope SeLection Alliance) study, Wells, DK et al. Key Parameters of Tumor Epitope Immunogenicity Revealed Through a Consortium Approach Improve Neoantigen Prediction. Cell 183, 818-834 e813 (2020), only a small number of neoantigens have been reported to be immunogenic. In a list of 51 peptides provided by 28 research teams, only 6% (median) were proven to be immunogenic in subsequent validation. There are many aspects of current technology that need to be improved. Specifically, they include the following aspects: 1) For variant detection, only SNV (somatic single nucleotide variant) and Sindel (somatic small insertion and deletion variant) detection are considered, which narrows the screening scope of new antigens; 2) When considering the binding affinity between MHC (histocompatibility complex) and polypeptide, most of them use a single prediction software and a rigid screening threshold, resulting in a high false positive; 3) The expression of the mutant gene itself and whether it is expressed at the RNA / protein level are not considered, resulting in false negative predictions; 4) The single new antigen immunogenicity detection method is still immature and requires multi-dimensional evaluation indicators; 5) In the new antigen screening stage, the characteristics considered mostly use rigid thresholds, or the parameters considered are incomplete. Summary of the invention

[0007] The present invention aims to solve the problem of unsatisfactory screening results. To this end, the present invention provides a method and system for predicting neoantigens of malignant solid tumors, which selects multiple feature values ​​based on the mechanism of action of neoantigens according to the key steps of the prediction process, does not use rigid filtering indicators, and uses a multiplication method to calculate the neoantigen score based on the scores of the feature values ​​related to each neoantigen.

[0008] The present invention provides a method for predicting new antigens of malignant solid tumors, and the technical scheme adopted is as follows: comprising the following steps:

[0009] Step 1: Obtain tumor samples and peripheral blood samples and perform DNA and RNA sequencing;

[0010] Step 2: Perform data quality control, genome reference sequence comparison, UMI detection, and comparison data deduplication analysis on the DNA sequencing results to obtain DNA sequencing data.

[0011] The RNA sequencing results are sequentially subjected to data quality control and compared with the genome reference sequence to obtain RNA sequencing data;

[0012] Step 3: Analyze the DNA sequencing data and RNA sequencing data to obtain the HLA typing results and the sequence of the mutant protein;

[0013] Step 4: Calculate the new antigen-related feature value based on the HLA typing results and the sequence of the mutant protein;

[0014] Step 5: Calculate the neoantigen score based on the neoantigen-related feature values;

[0015] Step 6: Select the N neoantigens with the highest neoantigen scores as the neoantigen prediction results for malignant solid tumors.

[0016] Furthermore, in step 1, tumor sample DNA and tumor sample RNA are extracted from the tumor sample, and peripheral blood sample DNA is extracted from the peripheral blood sample;

[0017] According to the tumor sample DNA and the peripheral blood sample DNA, respectively, a sequencing library carrying a UMI tag is constructed and exon sequencing is performed to obtain a tumor sample DNA sequencing result and a peripheral blood sample DNA sequencing result;

[0018] RNA sequencing is performed on the tumor sample RNA to obtain the tumor sample RNA sequencing result.

[0019] Furthermore, in step 2,

[0020] The data quality control process is as follows: remove reads with N bases greater than or equal to 5 in the sequencing results; remove reads with base quality values ​​less than 15 and accounting for more than 40% of the total bases; cut the sequencing reads containing sequencing adapter sequences, and retain reads with a final length of more than 75;

[0021] The genome reference sequence was aligned using BWA software to align the sequence results with the human reference genome hg19;

[0022] After completing UMI detection and deduplication analysis of alignment data, GATK software was used to perform base correction on sequencing reads.

[0023] Furthermore, in step 3, the following steps are included:

[0024] Step 3.1: Based on the DNA sequencing data of peripheral blood samples, HLA analysis was performed using HLA-HD and Polysolver software, and the consistent HLA analysis results were used as HLA typing results;

[0025] Step 3.2: Based on the paired DNA sequencing data of tumor samples and peripheral blood samples, Mutect2 software was used to detect somatic mutations, and the VEP tool was used to annotate the somatic mutation detection results and screen for SNV and Sindel mutations that can affect protein coding;

[0026] Step 3.3: Based on the RNA sequencing data of the tumor sample, perform quantitative detection and analysis of gene expression to obtain gene expression, and perform annotation analysis on the detected fusion gene to obtain the protein sequence of the fusion gene;

[0027] Step 3.4: Obtain the sequence of the mutant protein based on the SNV and Sindel mutations, gene expression, and protein sequence of the fusion gene.

[0028] Furthermore, the sequence of the mutant protein includes a sliding window with a length of 8, 9, 10, and 11 amino acids around the mutation site of the target tumor mutation, and the sequence of the tumor mutant peptide derived from the mutation, as well as the sequence of the wild peptide corresponding to the mutation.

[0029] Furthermore, the new antigen-related feature values ​​include Pm, Cm, Sm, Fm, Dm, Dai and PRm.

[0030] Pm represents the percentage of mutant peptides with a predicted Rank value less than 2 in the eight softwares using BA as the model, and the sum of the percentage of mutant peptides with a predicted Rank value less than or equal to 2 in the two softwares using EL as the model; the eight softwares using BA as the model are SMM, SMMPMBEC, NetMHC-4, NetMHCcons, MHCnuggets, PickPoket, MHCflurry (BA) and NetMHCpan (BA); the two softwares using EL as the model are MHCflurryEL and NetMHCpanEL;

[0031] Cm represents the efficiency of mutant peptide cleavage by protease;

[0032] Sm represents the stability of mutant peptide binding to MHC class I;

[0033] Fm indicates the similarity between the mutant peptide and the IEDB antigen;

[0034] Dm indicates the degree of difference between the mutant peptide and the normal protein sequence that humans can use as an immunogen or recognize as an antigen;

[0035] Dai represents the difference index between the binding affinity of mutant peptides to MHC class I and the binding affinity of wild-type peptides to MHC class I;

[0036] PRm represents the ranked percentage of the probability score of the mutant peptide bound to MHC class I in complex to be recognized by TCR.

[0037] Furthermore, based on the HLA typing results and the sequence of the mutant protein, NetChop software was used to predict the probability score of each mutant peptide sequence being cleaved by the proteasome to obtain Cm;

[0038] According to the HLA typing results and the sequence of the mutant protein, 8 softwares based on BA model and 2 softwares based on EL model were used to calculate the binding affinity between the mutant peptide and MHC class I and the binding affinity between the wild peptide and MHC class I, respectively. The median value of the predicted binding affinity of the 10 softwares was taken as the final affinity value to calculate Dai;

[0039] According to the HLA typing results and the sequence of the mutant protein, 8 softwares based on BA model and 2 softwares based on EL model were used to calculate the binding affinity of the mutant peptide to HLA I respectively. Then, the percentage of binding affinity predicted by the 8 softwares based on BA model with a Rank value less than 2 and the sum of the percentage of binding affinity predicted by the 2 softwares based on EL model with a Rank value less than or equal to 2 were counted to obtain Pm;

[0040] According to the HLA typing results and the sequence of the mutant protein, Netmhcstabpan was used to calculate the binding stability between each mutant peptide and MHCI to obtain Sm;

[0041] According to the HLA typing results and the sequence of the mutant protein, the PRIME software was used to calculate the propensity score of each mutant peptide and each MHC class I molecule to be recognized by TCR to obtain PRm;

[0042] According to the HLA typing results and the sequence of the mutant protein, the antigen.garnish software was used to calculate the similarity between the mutant peptide and the known antigens in IEDB, as well as the degree of difference between the mutant peptide and the known normal protein sequence that can be used as an immunogen or recognize antigen in humans, to obtain Fm and Dm.

[0043] Furthermore,

[0044]

[0045] in, The final affinity value representing the binding affinity of the wild-type peptide to MHC class I, The final affinity value represents the binding affinity of the mutant peptide to MHC class I.

[0046] Furthermore, in step 5, the scoring function of the new antigen is:

[0047] Neo_Score=Pm×Cm×Sm×(1+Fm)×Dai×PRm / (1+Dm)

[0048] Among them, Neo_Score represents the neoantigen score.

[0049] The present invention also provides a new antigen prediction system for malignant solid tumors, which adopts the following technical solution: comprising: a sequencing module, a data processing module, a data analysis module, a feature value calculation module and a score calculation module connected in sequence,

[0050] The sequencing module is used to obtain tumor samples and peripheral blood samples and perform DNA and RNA sequencing;

[0051] The data processing module is used to sequentially perform data quality control, genome reference sequence comparison, UMI detection, and deduplication analysis on the DNA sequencing results to obtain DNA sequencing data, and sequentially perform data quality control and genome reference sequence comparison on the RNA sequencing results to obtain RNA sequencing data;

[0052] The data analysis module is used to analyze the HLA typing results and the sequence of the mutant protein according to the DNA sequencing data and the RNA sequencing data;

[0053] The characteristic value calculation module is used to calculate the characteristic value related to the new antigen according to the HLA typing result and the sequence of the mutant protein;

[0054] The score calculation module is used to calculate the neoantigen score according to the neoantigen-related feature values; and select N neoantigens with the highest neoantigen scores as the neoantigen prediction results of malignant solid tumors.

[0055] The above one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects:

[0056] 1. The present invention starts from the mechanism of action of neoantigens, from the possibility of neoantigens being cleaved by proteasomes to their affinity for binding to MHC, to the stability of the complex after binding, and finally the possibility of being recognized by TCR. For each key step, multiple bioinformatics calculation methods are selected to obtain multi-dimensional feature values ​​to simulate the mechanism of action of neoantigens in the human body as much as possible.

[0057] 2. In the entire prediction process, the present invention does not use rigid filtering indicators. It conducts fair scoring from multiple dimensions for all possible polypeptide fragments, and finally obtains the final ranking of neoantigens based on the scores of the characteristic values ​​related to each neoantigen. The present invention designs a new scoring function, which uses the product method so that neoantigens with high scores for key evaluation indicators in each process can be ranked first, and the individual characteristic values ​​do not have a decisive influence.

[0058] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0060] Figure 1 It is a flow chart of the method provided by the present invention.

[0061] Figure 2 It is a structural block diagram of the system provided by the present invention.

[0062] Reference numerals:

[0063] 1. Sequencing module; 2. Data processing module; 3. Data analysis module; 4. Eigenvalue calculation module; 5. Score calculation module. DETAILED DESCRIPTION

[0064] In order to make the purpose, technical scheme and advantages of the present invention clearer, the technical scheme of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.

[0065] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0066] Combine the following Figure 1 and Figure 2 The present invention is further described in detail, and a method and system for predicting new antigens of malignant solid tumors of the present invention are described as follows:

[0067] In this embodiment, Figure 1 As shown, a method for predicting new antigens of malignant solid tumors is provided, comprising the following steps:

[0068] Step 1: Obtain tumor samples and peripheral blood samples from the same patient and perform DNA and RNA sequencing.

[0069] The specific process of sequencing is:

[0070] Tumor sample DNA and tumor sample RNA are extracted from the tumor sample, and peripheral blood sample DNA is extracted from the peripheral blood sample;

[0071] According to the tumor sample DNA and the peripheral blood sample DNA, respectively, a sequencing library carrying a UMI tag is constructed and exon sequencing is performed to obtain a tumor sample DNA sequencing result and a peripheral blood sample DNA sequencing result;

[0072] RNA sequencing is performed on the RNA of the tumor sample to obtain the RNA sequencing result of the tumor sample. RNA sequencing is specifically whole transcriptome sequencing.

[0073] In this step, three sequencing results are obtained, namely, the DNA sequencing results of peripheral blood samples, the DNA sequencing results of tumor samples, and the RNA sequencing results of tumor samples. That is, the DNA sequencing results include the DNA sequencing results of peripheral blood samples and the DNA sequencing results of tumor samples; the RNA sequencing results are the RNA sequencing results of tumor samples.

[0074] Step 2: Perform data quality control, genome reference sequence comparison, UMI detection, and comparison data deduplication analysis on the DNA sequencing results to obtain DNA sequencing data.

[0075] The RNA sequencing results are sequentially subjected to data quality control and compared with the genome reference sequence to obtain RNA sequencing data;

[0076] This step performs basic data analysis on the three sequencing results obtained in step 1 and ensures the accuracy of subsequent mutation detection through strict quality control.

[0077] The data quality control process is as follows: remove reads containing N bases greater than or equal to 5 in the sequencing results; remove reads with base quality values ​​less than 15 that account for more than 40% of the total bases; cut sequencing reads containing sequencing adapter sequences, and retain reads with a final length of more than 75; Reads refer to the base sequence obtained by a single sequencing by the sequencer.

[0078] The genome reference sequence was compared with the human reference genome hg19 using BWA software.

[0079] After data quality control and comparison with the genomic reference sequence, the RNA sequencing data of the tumor sample is obtained.

[0080] After data quality control and alignment with the genomic reference sequence, the DNA sequencing results of peripheral blood samples and tumor samples also need to complete UMI detection and deduplication analysis of the alignment data, and use GATK software to perform base correction on the sequencing reads to obtain the DNA sequencing data of peripheral blood samples and tumor samples, respectively.

[0081] That is, the DNA sequencing data obtained in step 2 includes DNA sequencing data of peripheral blood samples and DNA sequencing data of tumor samples; and the RNA sequencing data is RNA sequencing data of tumor samples.

[0082] Step 3: Based on the DNA sequencing data and RNA sequencing data, analyze and obtain the HLA typing results and the sequence of the mutant protein.

[0083] The specific process includes the following steps:

[0084] Step 3.1: Based on the DNA sequencing data of the peripheral blood samples, HLA analysis was performed using HLA-HD and Polysolver software. When the HLA analysis results of the two software were consistent, they were used as the HLA typing results.

[0085] Step 3.2: Based on the DNA sequencing data of paired tumor samples and peripheral blood samples, Mutect2 software was used to detect somatic mutations. The Variant Effect Prediction (VEP) tool was used to annotate the results of somatic mutation detection and screen for SNV and Sindel mutations that could affect protein coding.

[0086] Step 3.3: Based on the RNA sequencing data of the tumor sample, perform quantitative detection and analysis of gene expression to obtain gene expression, and perform annotation analysis on the detected fusion gene to obtain the protein sequence of the fusion gene.

[0087] Step 3.4: Obtain the sequence of the mutant protein based on the SNV and Sindel mutations, gene expression, and protein sequence of the fusion gene.

[0088] The sequence of the mutant protein is a polypeptide sequence of 8-11 amino acids, including the mutation site surrounding the target tumor mutation, forming a sliding window of 8, 9, 10, and 11 amino acids in length, and the sequence of the tumor mutant peptide derived from the mutation, as well as the sequence of the wild peptide corresponding to the mutation. The sequence of the mutant protein contains all mutations (SNV, Sindel, and fusion genes) that cause changes in the protein sequence.

[0089] Step 4: Based on the HLA typing results and the sequence of the mutant protein, new antigens are predicted and the characteristic values ​​related to the new antigens are calculated.

[0090] In this example, seven new antigen-related feature values ​​were calculated, including Pm, Cm, Sm, Fm, Dm, Dai and PRm.

[0091] The calculation process of the 7 new antigen-related feature values ​​is as follows:

[0092] (1) Based on the HLA typing results and the sequence of the mutant protein, NetChop software was used to predict the probability score of each mutant peptide sequence being cleaved by the proteasome to obtain Cm, which represents the efficiency of the mutant peptide being cleaved by the protease.

[0093] (2) The binding affinity between a polypeptide sequence and HLA is a very important factor in determining whether a polypeptide sequence can produce immunogenicity. In this embodiment, 8 softwares based on BA as a model and 2 softwares based on EL as a model are used to predict and calculate Dai and Pm. In this embodiment, the 8 softwares based on BA as a model are SMM, SMMPMBEC, NetMHC-4, NetMHCcons, MHCnuggets, PickPoket, MHCflurry (BA) and NetMHCpan (BA); the 2 softwares based on EL as a model are MHCflurryEL and NetMHCpanEL.

[0094] According to the HLA typing results and the sequence of the mutant protein, 8 softwares based on BA model and 2 softwares based on EL model were used to calculate the binding affinity of mutant peptide to MHC class I and the binding affinity of wild-type peptide to MHC class I, respectively. The median value of the predicted binding affinity of the 10 softwares was taken as the final affinity value, and Dai was calculated. Dai represents the difference index of the binding affinity of mutant peptide to MHC class I and the binding affinity of wild-type peptide to MHC class I.

[0095] The calculation formula of Dai is:

[0096]

[0097] in, The final affinity value representing the binding affinity of the wild-type peptide to MHC class I, The final affinity value represents the binding affinity of the mutant peptide to MHC class I.

[0098] Because different software uses different training models, different algorithm focuses, different binding type advantages, especially the influence of different HLA subtype population proportions, different research depths involved, and different reference data set sizes. The more software that supports strong binding affinity between peptides and HLA I, the more credible the results will be. Finally, the sum of the support rates of the software predicted by two different models is used to characterize the possibility of the peptide binding to HLA.

[0099] According to the HLA typing results and the sequence of the mutant protein, 8 softwares based on BA model and 2 softwares based on EL model were used to calculate the binding affinity of the mutant peptide to HLA I, respectively. Then, the proportion of binding affinity predicted by Rank values ​​less than 2 in the 8 softwares based on BA model and the sum of the proportion of binding affinity predicted by Rank values ​​less than or equal to 2 in the 2 softwares based on EL model were counted to obtain Pm, which represents the sum of the proportion of mutant peptide binding affinity predicted by Rank values ​​less than 2 in the 8 softwares based on BA model and the sum of the proportion of binding affinity predicted by Rank values ​​less than or equal to 2 in the 2 softwares based on EL model.

[0100] (3) According to the HLA typing results and the sequence of the mutant protein, Netmhcstabpan was used to calculate the binding stability between each mutant peptide and MHC class I to obtain Sm, which represents the binding stability of the mutant peptide to MHC class I.

[0101] (4) Based on the HLA typing results and the sequence of the mutant protein, the PRIME software was used to calculate the propensity score of each mutant peptide and each MHC class I molecule to be recognized by TCR, and the PRm was obtained. PRm represents the ranking percentage of the probability score of the mutant peptide and MHC class I binding complex being recognized by TCR.

[0102] (5) Based on the HLA typing results and the sequence of the mutant protein, the antigen.garnish software used in the TESLA study was used to calculate the similarity between the mutant peptide and the known antigens in the IEDB, as well as the degree of difference between the mutant peptide and the known normal protein sequence that can be used as an immunogen or recognize an antigen in humans, to obtain Fm and Dm. Fm represents the similarity between the mutant peptide and the IEDB antigen, where the Immune Epitope Database (IEDB) records the experimental data of antibodies and T cells studied in humans, non-human primates and other animal species in the context of infectious diseases, allergies, autoimmunity and transplantation; Dm represents the degree of difference between the mutant peptide and the normal protein sequence that can be used as an immunogen or recognize an antigen in humans. The lower the score, the greater the possibility that the peptide sequence can induce immunogenicity.

[0103] Step 5: Calculate the neoantigen score based on the neoantigen-related feature values.

[0104] The new antigen scoring function is:

[0105] Neo_Score=Pm×Cm×Sm×(1+Fm)×Dai×PRm / (1+Dm)

[0106] Among them, Neo_Score represents the neoantigen score.

[0107] The 7 new antigen-related characteristic values ​​selected in this example have been mentioned separately in some studies as being related to the immunogenicity of new antigens. However, these characteristic values ​​do not work alone, and there is no unified standard or screening threshold. The immunogenicity of new antigens is determined by many factors. Currently, many methods only consider a single factor, which is a very important reason for false negative prediction of new antigens. Moreover, if a fixed threshold is set, false positives in new antigen screening will also occur.

[0108] This embodiment comprehensively considers all evaluation indicators of the key steps in the entire biological process. In terms of scoring, a new antigen scoring function using a product method is designed. This is to highlight that new antigens with high scores for key evaluation indicators in each process can be ranked at the top. The new antigen score itself is meaningless, and the final ranking will serve as the most important criterion for new antigen screening.

[0109] Step 6: Select the N neoantigens with the highest neoantigen scores as the neoantigen prediction results for malignant solid tumors.

[0110] The specific process is: sort the neoantigen scores from large to small. The higher the score, the stronger the immunogenicity of the peptide, and the more likely it is to be an effective neoantigen. The screening criteria (the N neoantigens with the highest scores) can be set according to different experimental requirements and actual application scenarios. Usually, neoantigens ranked in the top (10, 20, 30...100) are selected for subsequent research.

[0111] In order to verify the effectiveness of this prediction method, the following experiments were conducted:

[0112] In the TESLA study, the immunogenicity of the selected new peptides was determined by labeling the subject's matched TIL or peripheral blood mononuclear cells (PBMC) with HLA-I peptide polymers. This example collects 31 new antigen peptides that were validated in the TESLA study, and 3 new antigen peptides were not collected. The somatic mutation data and HLA typing results of 3 human lung cancer samples (TESLA2, TESLA12, TESLA16) and 2 human melanoma samples (TESLA1, TESLA3) from the TESLA study provided by Müller, M. et al. Machine learning methods and harmonized datasets improve immunogenic neoantigen prediction. Immunity 56, 2650-2663.e2656 (2023) were used.

[0113] Using the prediction method of this embodiment, the final neoantigen scores were sorted from large to small, and the same screening criteria in the TESLA study were referred to. The screening results are shown in Tables 1 and 2. Among the screened neoantigens, 23 neoantigens were validated by the TESLA study, with a positive consistency rate of 70.97% (22 / 31), which was significantly higher than the best positive predictive value of 58.82% (20 / 34) mentioned in the TESLA study.

[0114] Table 1 Statistics of positive prediction results

[0115]

[0116] p, l, t...f, g, b in Table 1 are the numbers of the prediction teams in the TESLA study.

[0117] Table 2 Screening results in the same way as TESLA study

[0118]

[0119] In the verification results of Table 2, negative means the verification result is negative; not_tested means there is no corresponding verification result; CD8 means the verification result is positive. The calculation results of Pm, Cm, Sm, PRm, Dai and Neo_Score are retained to three decimal places.

[0120] In this embodiment, Figure 2 As shown, a new antigen prediction system for malignant solid tumors is also provided, and the technical solution adopted is as follows: it includes: a sequencing module 1, a data processing module 2, a data analysis module 3, a feature value calculation module 4 and a score calculation module 5 connected in sequence.

[0121] The sequencing module 1 is used to obtain tumor samples and peripheral blood samples and perform DNA and RNA sequencing;

[0122] The data processing module 2 is used to sequentially perform data quality control, genome reference sequence comparison, UMI detection, and deduplication analysis on the DNA sequencing results to obtain DNA sequencing data, and sequentially perform data quality control and genome reference sequence comparison on the RNA sequencing results to obtain RNA sequencing data;

[0123] The data analysis module 3 is used to analyze the HLA typing results and the sequence of the mutant protein according to the DNA sequencing data and the RNA sequencing data;

[0124] The characteristic value calculation module 4 is used to calculate the characteristic value related to the new antigen according to the HLA typing result and the sequence of the mutant protein;

[0125] The score calculation module 5 is used to calculate the neoantigen score according to the neoantigen-related feature values; and select N neoantigens with the highest neoantigen scores as the neoantigen prediction results of malignant solid tumors.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting new antigens of malignant solid tumors, characterized in that: The following steps are involved: Step 1: Obtain tumor samples and peripheral blood samples and perform DNA and RNA sequencing; Step 2: Perform data quality control, genome reference sequence comparison, UMI detection, and comparison data deduplication analysis on the DNA sequencing results to obtain DNA sequencing data. The RNA sequencing results are sequentially subjected to data quality control and compared with the genome reference sequence to obtain RNA sequencing data; Step 3: Analyze the DNA sequencing data and RNA sequencing data to obtain the HLA typing results and the sequence of the mutant protein; Step 4: Calculate the new antigen-related feature value based on the HLA typing results and the sequence of the mutant protein; Step 5: Calculate the neoantigen score based on the neoantigen-related feature values; Neoantigen-related feature values ​​include Pm, Cm, Sm, Fm, Dm, Dai, and PRm. Pm represents the percentage of mutant peptides with a predicted Rank value less than 2 in the eight softwares using BA as the model, and the sum of the percentage of mutant peptides with a predicted Rank value less than or equal to 2 in the two softwares using EL as the model; the eight softwares using BA as the model are SMM, SMMPMBEC, NetMHC-4, NetMHCcons, MHCnuggets, PickPoket, MHCflurry (BA) and NetMHCpan (BA); the two softwares using EL as the model are MHCflurryEL and NetMHCpanEL; Cm represents the efficiency of mutant peptide cleavage by protease; Sm represents the stability of mutant peptide binding to MHC class I; Fm indicates the similarity between the mutant peptide and the IEDB antigen; Dm indicates the degree of difference between the mutant peptide and the normal protein sequence that humans can use as an immunogen or recognize as an antigen; Dai represents the difference index between the binding affinity of mutant peptides to MHC class I and the binding affinity of wild-type peptides to MHC class I; in, The final affinity value representing the binding affinity of the wild-type peptide to MHC class I, The final affinity value representing the binding affinity of the mutant peptide to MHC class I; PRm represents the ranked percentage of the probability score of the mutant peptide-MHC class I binding complex being recognized by TCR; In step 5, the scoring function of the new antigen is: Neo_Score=Pm×Cm×Sm×(1+Fm)×Dai×PRm / (1+Dm) Among them, Neo_Score represents the neoantigen score; Step 6: Select the N neoantigens with the highest neoantigen scores as the neoantigen prediction results for malignant solid tumors.

2. The method for predicting new antigens of malignant solid tumors according to claim 1, characterized in that: In step 1, tumor sample DNA and tumor sample RNA are extracted from the tumor sample, and peripheral blood sample DNA is extracted from the peripheral blood sample; According to the tumor sample DNA and the peripheral blood sample DNA, respectively, a sequencing library carrying a UMI tag is constructed and exon sequencing is performed to obtain a tumor sample DNA sequencing result and a peripheral blood sample DNA sequencing result; RNA sequencing is performed on the tumor sample RNA to obtain the tumor sample RNA sequencing result.

3. The method for predicting new antigens of malignant solid tumors according to claim 1, characterized in that: In step 2, The data quality control process is as follows: remove reads with N bases greater than or equal to 5 in the sequencing results; remove reads with base quality values ​​less than 15 and accounting for more than 40% of the total bases; cut the sequencing reads containing sequencing adapter sequences, and retain reads with a final length of more than 75; The genome reference sequence was aligned using BWA software to align the sequence results with the human reference genome hg19; After completing UMI detection and deduplication analysis of alignment data, GATK software was used to perform base correction on sequencing reads.

4. The method for predicting new antigens of malignant solid tumors according to claim 1, characterized in that: In step 3, the following steps are included: Step 3.1: Based on the DNA sequencing data of peripheral blood samples, HLA analysis was performed using HLA-HD and Polysolver software, and the consistent HLA analysis results were used as HLA typing results; Step 3.2: Based on the paired DNA sequencing data of tumor samples and peripheral blood samples, Mutect2 software was used to detect somatic mutations, and the VEP tool was used to annotate the somatic mutation detection results and screen for SNV and Sindel mutations that can affect protein coding; Step 3.3: Based on the RNA sequencing data of the tumor sample, perform quantitative detection and analysis of gene expression to obtain gene expression, and perform annotation analysis on the detected fusion gene to obtain the protein sequence of the fusion gene; Step 3.4: Obtain the sequence of the mutant protein based on the SNV and Sindel mutations, gene expression, and protein sequence of the fusion gene.

5. A method for predicting new antigens of malignant solid tumors according to claim 1 or 4, characterized in that: The sequence of the mutant protein includes a sliding window with a length of 8, 9, 10, or 11 amino acids around the mutation site of the target tumor, the sequence of the tumor mutant peptide derived from the mutation, and the sequence of the wild peptide corresponding to the mutation.

6. The method for predicting new antigens of malignant solid tumors according to claim 1, characterized in that: According to the HLA typing results and the sequence of the mutant protein, NetChop software was used to predict the probability score of each mutant peptide sequence being cleaved by the proteasome to obtain Cm; According to the HLA typing results and the sequence of the mutant protein, 8 softwares based on BA model and 2 softwares based on EL model were used to calculate the binding affinity between the mutant peptide and MHC class I and the binding affinity between the wild peptide and MHC class I, respectively. The median value of the predicted binding affinity of the 10 softwares was taken as the final affinity value to calculate Dai; According to the HLA typing results and the sequence of the mutant protein, 8 softwares based on BA model and 2 softwares based on EL model were used to calculate the binding affinity of the mutant peptide to HLA I respectively. Then, the percentage of binding affinity predicted by the 8 softwares based on BA model with a Rank value less than 2 and the sum of the percentage of binding affinity predicted by the 2 softwares based on EL model with a Rank value less than or equal to 2 were counted to obtain Pm; According to the HLA typing results and the sequence of the mutant protein, Netmhcstabpan was used to calculate the binding stability between each mutant peptide and MHC I to obtain Sm; According to the HLA typing results and the sequence of the mutant protein, the propensity score of each mutant peptide and each MHC class I molecule to be recognized by TCR was calculated using PRIME software to obtain PRm; According to the HLA typing results and the sequence of the mutant protein, the antigen.garnish software was used to calculate the similarity between the mutant peptide and the known antigens in IEDB, as well as the degree of difference between the mutant peptide and the known normal protein sequence that can be used as an immunogen or recognize antigen in humans, to obtain Fm and Dm.

7. A new antigen prediction system for malignant solid tumors, characterized in that: A method for predicting new antigens of a malignant solid tumor according to any one of claims 1 to 6, comprising: a sequencing module, a data processing module, a data analysis module, a feature value calculation module and a score calculation module connected in sequence, The sequencing module is used to obtain tumor samples and peripheral blood samples and perform DNA and RNA sequencing; The data processing module is used to sequentially perform data quality control, genome reference sequence comparison, UMI detection, and deduplication analysis on the DNA sequencing results to obtain DNA sequencing data, and sequentially perform data quality control and genome reference sequence comparison on the RNA sequencing results to obtain RNA sequencing data; The data analysis module is used to analyze the HLA typing results and the sequence of the mutant protein according to the DNA sequencing data and the RNA sequencing data; The characteristic value calculation module is used to calculate the characteristic value related to the new antigen according to the HLA typing result and the sequence of the mutant protein; The score calculation module is used to calculate the neoantigen score according to the neoantigen-related feature values; and select N neoantigens with the highest neoantigen scores as the neoantigen prediction results of malignant solid tumors.

Citation Information

Patent Citations

  • Prediction method of clinical individualized tumor neoantigen

    CN111415707A

  • Method and device for predicting tumor neoantigen

    CN116580771A

  • Tumor neoantigen recognition method and system based on next-generation sequencing data

    CN118351934A