Neoantigen load calculation method, equipment and program product

By calculating the load of neoantigens, the problem that existing biomarkers cannot accurately judge patients' response to immune checkpoint inhibitor treatment is solved, and a more accurate assessment of cancer diagnosis and treatment effect is achieved.

CN120048334AActive Publication Date: 2025-05-27SHANGHAI TENTH PEOPLES HOSPITAL
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510166358.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-27
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

Existing biomarkers, such as PD-L1 expression levels, TMB, MSI-H and dMMR, cannot accurately distinguish which patients can benefit from immune checkpoint inhibitor treatment, resulting in poor treatment results.

Method used

A method for calculating the neoantigens load is proposed. By obtaining the sequencing data of the person to be tested, the sequence and structural characteristics of the antigen peptide, the sequence and structural characteristics of the MHC binding tank are extracted, and the affinity of the MHC-antigens peptide and the probability of TCR recognition are calculated, and the neoantigens load of the person to be tested is finally calculated.

Benefits of technology

By computed neoantigen load as a biomarker, the clinical efficacy of immune checkpoint inhibitors can be more accurately judged, the accuracy of cancer diagnosis can be improved, and cancer treatment can be assisted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048334A_ABST
    Figure CN120048334A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent medical treatment, in particular to a new antigen load calculation method, equipment and a program product. Comprising the following steps: S1, obtaining sequencing data of a to-be-detected person, wherein the sequencing data is exon group sequencing data or panel sequencing data; s2, extracting the sequence and structural characteristics of the antigen peptide, the sequence and structural characteristics of an MHC binding slot and the sequence of a TCR variable region based on the sequencing data; s3, calculating MHC-antigen peptide affinity based on the sequence and structural characteristics of the MHC binding slot and the sequence and structural characteristics of the antigen peptide; s4, calculating the probability that the TCR recognizes the MHC-antigen peptide based on the sequence of the TCR variable region; s5, calculating the new antigen load of the to-be-detected person based on the probability that the TCR identifies the MHC-antigen peptide and the affinity of the MHC-antigen peptide, and the method can calculate the new antigen load and has good clinical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent medicine, and particularly relates to a method, device, program product and computer-readable storage medium for calculating the tumor neoantigen burden. Background Art

[0002] Anti-PD-L1 (programmed death ligand 1) therapy is an immune checkpoint inhibitor therapy aimed at enhancing the body's own immune system's ability to recognize and attack tumor cells. This treatment method relieves the inhibition of the immune system by tumor cells by blocking the interaction between PD-L1 and PD-1, thereby activating the T cell-mediated anti-tumor immune response. In the treatment of esophageal cancer with Anti-PD-L1, there is a lack of reliable biomarkers. In the clinical application of immune checkpoint inhibitors CTLA4 and PD-L1 monoclonal antibodies, it is very important to identify patients by detecting biomarkers. Currently, commonly used companion diagnostic markers include PD-L1 expression level, tumor mutation burden (TMB), microsatellite instability (MSI-H), and DNA mismatch repair defect (dMMR), etc. Although these prognostic markers play a guiding role in clinical applications, more and more clinical cohort studies have shown that these markers cannot accurately distinguish which patients can benefit from immune checkpoint inhibitor therapy. Therefore, the concept of tumor neoantigen burden (TNB) has been proposed. Tumor neoantigens are the actual number of mutations targeted by T cells and can better judge the clinical efficacy of immune checkpoint inhibitors. Some studies have shown that combining neoantigen quality with tumor CD8+ T cell infiltration can accurately identify patients with malignant gliomas with the longest survival, while other studies have proposed a neoantigen screening algorithm, which has been proven to be able to effectively distinguish pancreatic cancer patients with long and short survival, but models relying only on the number of antigens cannot achieve this. Summary of the Invention

[0003] In view of the above problems, the present invention provides a method for calculating the tumor neoantigen burden, specifically including: S1. Obtain the sequencing data of the subject to be tested, and the sequencing data is exome sequencing data or panel sequencing data; S2. Extract the sequence and structural features of antigen peptides, the sequence and structural features of the MHC binding groove, and the sequence of the TCR variable region based on the sequencing data; S3. Calculate the MHC-antigen peptide affinity based on the sequence and structural features of the MHC binding groove and the sequence and structural features of the antigen peptide; S4. Calculate the probability of TCR recognizing MHC-antigen peptide based on the sequence of the TCR variable region; S5. Calculate the neoantigen load of the subject based on the probability of the TCR recognizing the MHC-antigen peptide and the MHC-antigen peptide affinity.

[0004] The extraction obtains the sequence and structural features of the antigen peptide and the sequence and structural features of the MHC binding groove through a trained first deep network model; Optionally, the training process of the first deep network model is as follows: Obtain the antigen peptide sequence and the sequence of the MHC binding groove from the protein database; Extract the structural features from the sequences to obtain the structural features; Construct a feature training set based on the sequences and structural features; Input the feature training set into the first deep network model for training to obtain a trained first deep network model; Optionally, the structural features include one or more of the following: structural features, structural neighbor features, solvent accessible surface area; Optionally, the method of feature extraction adopts one or more of the following: physicochemical extraction, local structure entropy extraction, pairwise potential extraction, interaction tendency extraction; Optionally, the first deep network model includes one or more of the following: GBRT, XGBoost, LightGBM, CatBoost, Random Forest, AdaBoost.

[0005] Replace S2 with: Input the sequencing data into a trained second deep network model to obtain the MHC-antigen peptide affinity. The training process of the second deep network model is as follows: The first step: Obtain the MHC antigen peptide mass spectrometry dataset and IEDB data; The second step: Perform pre-training on the neural network model through the IEDB data to obtain the weights of the training model; The third step: Transfer the weights of the training model to the second deep network model and input the MHC antigen peptide mass spectrometry dataset into the second deep network model for training to obtain a trained second deep network model; Optionally, the first step further includes MHC typing, obtaining antigen peptide data and MHC molecules, typing the MHC molecules to obtain MHC class I data, and constructing an MHC class I antigen peptide mass spectrometry dataset of MHC class I data and antigen peptides based on mass spectrometry; Input the MHC class I antigen peptide mass spectrometry dataset into the training model for training to obtain the second deep network model; Optionally, the third step also includes antigen deep features, obtaining antigen gene expression level data, extracting features of the antigen gene expression level data to obtain antigen deep features, fusing the antigen deep features with the MHC-I class antigen peptide mass spectrum data and inputting them into the second deep network model for training to obtain a trained second deep network model.

[0006] Optionally, S2 is replaced by: inputting the sequencing data into a trained multi-source heterogeneous network model to obtain MHC-antigen peptide affinity, and the training process of the multi-source heterogeneous network model is: Obtaining the MHC molecule structure and the antigen peptide structure to generate the MHC network structure and the antigen peptide network structure; the MHC molecule includes the MHC-I subtype and the MHC-II subtype; Calculate the similarity network between MHC-I subtypes and MHC-II subtypes, and the similarity network of antigen peptides; Obtain the association between MHC molecules and antigen peptides; Based on the similarity network of MHC-I subtype and MHC-II subtype, the similarity network of antigen peptide, the association relationship between MHC molecules and antigen peptides is input into the third deep network model to generate a multi-source heterogeneous network; Optionally, the multi-source heterogeneous network model adopts one or more of the following: a two-way heterogeneous network, a heterogeneous information network, a graph neural network, and a hypergraph; Optionally, S2 further includes executing the second deep network model in parallel to obtain MHC-antigen peptide affinity, performing weight distribution calculation on the MHC-antigen peptide affinity obtained by the first deep network model and the second deep network model to obtain the second MHC-antigen peptide affinity, and calculating the new antigen load of the subject based on the probability of the TCR recognizing the MHC-antigen peptide and the second MHC-antigen peptide affinity; Optionally, S2 further includes executing the multi-source heterogeneous network model in parallel to obtain MHC-antigen peptide affinity, performing weight distribution calculation on the MHC-antigen peptide affinity obtained by the first deep network model and the multi-source heterogeneous network model to obtain a third MHC-antigen peptide affinity, and calculating the new antigen load of the subject based on the probability of the TCR recognizing the MHC-antigen peptide and the third MHC-antigen peptide affinity; Optionally, S2 also includes executing the second deep network model and the multi-source heterogeneous network model in parallel to obtain MHC-antigen peptide affinity, performing weight distribution calculation on the MHC-antigen peptide affinity obtained by the first deep network model, the multi-source heterogeneous network model, and the second deep network model to obtain a fourth MHC-antigen peptide affinity, and calculating the new antigen load of the subject based on the probability of TCR recognizing the MHC-antigen peptide and the fourth MHC-antigen peptide affinity.

[0007] The sequence extraction of the TCR variable region is obtained by extracting from the genomic data of the subject to be tested; Optionally, the probability calculation of the TCR recognizing the MHC-antigen peptide is calculated by a recognition model, and the construction process of the recognition model is as follows: Obtain antigen epitope-TCR interaction data, the sequences and structures of TCR and peptide chains; Extract TCR features based on the sequence and structure data of the TCR and peptide chains; Input the TCR features and antigen epitope-TCR interaction data into the fourth deep network model for training to obtain a recognition model; Optionally, the TCR features include one or more of the following: the charge, hydrophobicity, and two-dimensional structure features of CDR3; Optionally, the extraction of the TCR features is obtained by group sparse regularization regression analysis.

[0008] Optionally, the probability calculation of the TCR recognition further includes TCR-antigen epitope binding prediction, and the probability of TCR recognition is obtained through the recognition model and TCR-antigen epitope binding prediction; the TCR-antigen epitope binding prediction is obtained by predicting through a trained TCR-antigen epitope binding prediction model; Optionally, the construction process of the trained TCR-antigen epitope binding prediction model is as follows: Obtain the sequence of the CDR3 region, gene mutation data, and antigen peptide; Convert the sequence of the CDR3 region, gene mutation data, and antigen peptide into a feature matrix; Input the feature matrix into a neural network for training to obtain a TCR-antigen epitope binding prediction model; Optionally, the feature matrix of the antigen peptide includes the sequence features and structural features of the antigen peptide; Optionally, the TCR-antigen epitope binding prediction model further includes MHC-antigen peptide affinity calculation, and the affinity prediction result is obtained through the sequence features and structural features of the antigen peptide.

[0009] The probability calculation of the TCR recognition further includes the immunogenicity prediction of tumor antigens, and the probability of TCR recognition is obtained through the recognition model and the immunogenicity prediction of tumor antigens. The immunogenicity prediction of tumor antigens is obtained by predicting through a trained immunogenicity prediction model of tumor antigens; Optionally, the construction process of the immunogenicity prediction model of tumor antigens is as follows: Obtain RNA-seq data of different tumor-infiltrating lymphocytes, tumor cell genomes, and immunogenicity data; Performing TCR repertoire analysis on the RNA-seq data to obtain tumor gene expression profile data and obtaining CDR3 sequences from the RNA-seq data; Obtaining gene mutation and antigen peptide data based on the tumor cell genomic data; Performing feature extraction on the CDR3 sequences and antigen peptide data to obtain feature matrix data; Inputting the tumor gene expression profile data, gene mutations, and feature matrix data into a neural network for training to obtain an immunogenicity prediction model.

[0010] The calculation formula for the neoantigen load is as follows:

[0011] Where N is the neoantigen load, n represents the number of tumor clone lines analyzed from the exome sequencing data of this patient, b i represents the affinity, r i represents the probability of TCR recognition, c j represents the expression level of gene j.

[0012] The purpose of the present invention is to provide a method for predicting the prognosis of tumor treatment based on neoantigen load, including: Obtaining tumor slice data and serum data of the subject to be tested; Performing neoantigen screening based on the tumor slice data and serum data to obtain neoantigens; Calculating the load of the neoantigens based on the above-mentioned neoantigen load calculation method to obtain the neoantigen load; Performing prognosis prediction based on the neoantigen load to obtain a prediction result.

[0013] The purpose of the present invention is to provide a computer program product, which includes a computer program or instruction, and the computer program or instruction is executed by a processor to implement the above-mentioned neoantigen load calculation method or the method for predicting the prognosis of tumor treatment based on neoantigen load.

[0014] The purpose of the present invention is to provide a computer device, which includes a memory, a processor, and a computer program or instruction stored on the memory, and the computer program or instruction is executed by the processor to implement the above-mentioned neoantigen load calculation method or the method for predicting the prognosis of tumor treatment based on neoantigen load.

[0015] The purpose of the present invention is to provide a computer-readable storage medium, which stores a computer program or instruction, and the computer program or instruction is executed by a processor to implement the above-mentioned neoantigen load calculation method or the method for predicting the prognosis of tumor treatment based on neoantigen load.

[0016] Advantages of the present invention: 1. A new method for calculating neoantigen load is proposed, which is calculated through the affinity of MHC antigen peptides and the probability of TCR recognition. The calculated neoantigen load is used as a biomarker to assist cancer diagnosis, improve the diagnostic accuracy of cancer, and contribute to cancer treatment.

[0017] 2. For the calculation of the affinity of MHC antigen peptides, the affinity of MHC antigen peptides is predicted through the mass spectrometry data of MHC class I-binding antigen peptides, the sequence and structural characteristics of antigen peptides binding to the MHC binding groove, and the topological and ontological characteristics of the MHC antigen peptide isomer network. Multiple data and characteristics are integrated for the calculation of affinity, enhancing the accuracy and reliability of affinity, and further improving the accuracy and reliability of neoantigen load.

[0018] 3. For the calculation of the probability of TCR recognition, the relevant characteristics of antigen and TCR are used as inputs to predict the probability of TCR recognition and binding to tumor antigens; the gene mutations of different tumors are integrated to construct a gene mutation-TCR corresponding matrix, and the probability of tumor-specific antigens being recognized by TCR is analyzed from the perspective of tumor clonal evolution; combined with tumor-infiltrating lymphocyte (TIL) data and experimentally obtained immunogenicity data, the tumor gene expression profile and antigen peptide feature matrix are used as inputs to a deep neural network to predict the immunogenicity of tumor antigens, and then the probability of TCR recognition is predicted. Integrating multi-source data for TCR recognition can obtain an accurate probability of TCR recognition, which is beneficial to the calculation accuracy of neoantigen load. Brief Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0020] Figure 1 Schematic diagram of the calculation method flow of neoantigen load provided by the embodiment of the present invention; Figure 2 Schematic diagram of the calculation system of neoantigen load provided by the embodiment of the present invention; Figure 3 Schematic diagram of the calculation device of neoantigen load provided by the embodiment of the present invention; Figure 4 Multi-omics data-driven tumor neoantigen screening route provided by the embodiment of the present invention; Figure 5The present invention provides a method for training a high-precision MHC-antigen peptide affinity prediction model using transfer learning. (a) Train a convolutional neural network using data from IEDB to obtain a pan-cancer MHC-antigen peptide affinity prediction model; (b) Use transfer learning technology to perform a one-step convolutional neural network in combination with transcriptome and proteomic data to obtain a prediction model for a specific cancer type. Detailed implementation manners

[0021] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0022] In some processes described in the specification, claims and above-mentioned drawings of the present invention, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0023] Figure 1 The schematic diagram of the calculation method for the neoantigen load provided by the embodiments of the present invention specifically includes: S101: Obtain the sequencing data of the subject to be tested, and the sequencing data is exome sequencing data or panel sequencing data; In a specific embodiment, the present invention takes the tumor immune cycle as the research context, deeply analyzes the factors involved in tumor neoantigen screening from endogenous mechanisms, influencing variables and corresponding data sources, etc., extracts key problems and conducts research on them. By carefully analyzing the biological mechanism of these tumor antigens presenting to activate T cells, the present invention conducts research on two problems: the affinity between MHC-I and antigen peptides, and the recognition of antigen epitopes by T and TCR, because these are the keys to determining whether an antigen is presented and whether it can activate T cells after being presented to the cell surface.

[0024] S102: Extract the sequence and structural features of antigen peptides, the sequence and structural features of the MHC binding groove, and the sequence of the TCR variable region based on the sequencing data; In one embodiment, the extraction is to obtain the sequence and structural features of antigen peptides and the sequence and structural features of the MHC binding groove through a trained first deep network model.

[0025] In one embodiment, the training process of the first deep network model is as follows: Obtain the antigen peptide sequence and the sequence of the MHC binding groove from a protein database; Extract the structural features from the sequences to obtain structural features; Construct a feature training set based on the sequences and structural features; Input the feature training set into the first deep network model for training to obtain a trained first deep network model.

[0026] In one embodiment, the structural features include one or more of the following: structural features, structural neighbor features, solvent accessible surface area.

[0027] In one embodiment, the method for feature extraction adopts one or more of the following: physicochemical extraction, local structure entropy extraction, pairwise potential extraction, interaction tendency extraction.

[0028] In one embodiment, the first deep network model includes one or more of the following: GBRT, XGBoost, LightGBM, CatBoost, Random Forest, AdaBoost.

[0029] In one embodiment, S102 is replaced with: Input the sequencing data into a trained second deep network model to obtain the MHC-antigen peptide affinity. The training process of the second deep network model is as follows: First step: Obtain the MHC antigen peptide mass spectrometry dataset and IEDB data; Second step: Perform pre-training on a neural network model through the IEDB data to obtain the weights of the training model; Third step: Transfer the weights of the training model to the second deep network model and input the MHC antigen peptide mass spectrometry dataset into the second deep network model for training to obtain a trained second deep network model; Optionally, the first step further includes MHC typing, obtaining antigen peptide data and MHC molecules, typing the MHC molecules to obtain MHC class I data, and constructing an MHC class I antigen peptide mass spectrometry dataset of MHC class I data and antigen peptides based on mass spectrometry; Input the MHC class I antigen peptide mass spectrometry dataset into the training model for training to obtain the second deep network model.

[0030] In one embodiment, the third step further includes antigen deep features, obtaining antigen gene expression level data, extracting the features of the antigen gene expression level data to obtain antigen deep features, fusing the antigen deep features with the MHC class I antigen peptide mass spectrometry data, and inputting the fused data into the second deep network model for training to obtain a trained second deep network model.

[0031] In one embodiment, step S102 is replaced with: inputting the sequencing data into a trained multi-source heterogeneous network model to obtain the MHC-antigen peptide affinity. The training process of the multi-source heterogeneous network model is as follows: Obtain the MHC molecular structure and the antigen peptide structure to generate an MHC network structure and an antigen peptide network structure; the MHC molecules include MHC-I subtypes and MHC-II subtypes; Calculate the similarity network of MHC-I subtypes and MHC-II subtypes, and the similarity network of antigen peptides; Obtain the association relationship between MHC molecules and antigen peptides; Based on the similarity network of MHC-I subtypes and MHC-II subtypes, the antigen peptide similarity network, and the association relationship between MHC molecules and antigen peptides, input them into a third deep network model to generate a multi-source heterogeneous network.

[0032] In one embodiment, the multi-source heterogeneous network model adopts one or more of the following: two-pass heterogeneous network, heterogeneous information network, graph neural network, hypergraph.

[0033] In one embodiment, step S102 further includes parallelly executing a second deep network model to obtain the MHC-antigen peptide affinity, performing weighted distribution calculation on the MHC-antigen peptide affinities obtained by the first deep network model and the second deep network model to obtain a second MHC-antigen peptide affinity, and calculating the neoantigen load of the subject based on the probability of TCR recognizing MHC-antigen peptide and the second MHC-antigen peptide affinity.

[0034] In one embodiment, step S2 further includes parallelly executing a multi-source heterogeneous network model to obtain the MHC-antigen peptide affinity, performing weighted distribution calculation on the MHC-antigen peptide affinities obtained by the first deep network model and the multi-source heterogeneous network model to obtain a third MHC-antigen peptide affinity, and calculating the neoantigen load of the subject based on the probability of TCR recognizing MHC-antigen peptide and the third MHC-antigen peptide affinity.

[0035] In one embodiment, step S2 further includes parallelly executing a second deep network model and a multi-source heterogeneous network model to obtain the MHC-antigen peptide affinity, performing weighted distribution calculation on the MHC-antigen peptide affinities obtained by the first deep network model, the multi-source heterogeneous network model, and the second deep network model to obtain a fourth MHC-antigen peptide affinity, and calculating the neoantigen load of the subject based on the probability of TCR recognizing MHC-antigen peptide and the fourth MHC-antigen peptide affinity.

[0036] In one embodiment, the sequence of the TCR variable region is extracted from the genomic data of the subject.

[0037] S103: Calculate the MHC-antigen peptide affinity based on the sequence and structural features of the MHC binding groove and the sequence and structural features of the antigen peptide. In a specific embodiment, calculate the affinity between MHC-I molecules and antigen peptides: Integrate the large-scale mass spectrometry data of MHC-I binding antigen peptides, construct a mass spectrometry-based MHC-antigen peptide benchmark dataset and develop an online database; Analyze genomic or whole exome sequencing data for MHC typing, extract the sequence and structural features of antigen peptides and MHC binding grooves, and propose and develop a deep learning-based MHC-antigen peptide affinity prediction model; Construct an MHC-antigen peptide heterogeneous network that fuses multi-source data, combine network topology features and ontology features, and train a deep feedforward network to predict MHC-antigen peptide affinity.

[0038] S104: Calculate the probability that TCR recognizes MHC-antigen peptides based on the sequence of the TCR variable region. In one embodiment, the probability that TCR recognizes MHC-antigen peptides is calculated by a recognition model, and the construction process of the recognition model is as follows: Obtain antigen epitope-TCR interaction data, and the sequence and structure of TCR and peptide chains. Extract TCR features based on the sequence and structural data of TCR and peptide chains. Input the TCR features and antigen epitope-TCR interaction data into the fourth deep network model for training to obtain a recognition model.

[0039] In one embodiment, the TCR features include one or more of the following: the charge, hydrophobicity, and two-dimensional structural features of CDR3.

[0040] In one embodiment, the extraction of TCR features is obtained through group sparse regularization regression analysis.

[0041] In one embodiment, the calculation of the probability that TCR recognizes also includes TCR-antigen epitope binding prediction, and the probability that TCR recognizes is obtained through the recognition model and TCR-antigen epitope binding prediction; The TCR-antigen epitope binding prediction is obtained through a trained TCR-antigen epitope binding prediction model. Optionally, the construction process of the trained TCR-antigen epitope binding prediction model is as follows: Obtain the sequence of the CDR3 region, gene mutation data, and antigen peptides. Convert the sequence of the CDR3 region, gene mutation data, and antigen peptides into a feature matrix. Input the feature matrix into a neural network for training to obtain a TCR-antigen epitope binding prediction model.

[0042] In one embodiment, the characteristic matrix of the antigen peptide includes the characteristics of the antigen peptide sequence and the structure.

[0043] In one embodiment, the TCR-antigen epitope binding prediction model further includes MHC-antigen peptide affinity calculation, and the affinity prediction result is obtained through the characteristics of the antigen peptide sequence and the structure.

[0044] In one embodiment, the calculation of the probability of TCR recognition further includes the prediction of the immunogenicity of tumor antigens. The probability of TCR recognition is obtained through the recognition model and the prediction of the immunogenicity of tumor antigens. The prediction of the immunogenicity of tumor antigens is obtained through a trained immunogenicity prediction model of tumor antigens.

[0045] In one embodiment, the construction process of the immunogenicity prediction model of tumor antigens is as follows: Obtain RNA-seq data, tumor cell genomes, and immunogenicity data of different tumor-infiltrating lymphocytes; Perform TCR repertoire analysis on the RNA-seq data to obtain tumor gene expression profile data and obtain CDR3 sequences from the RNA-seq data; Obtain gene mutations and antigen peptide data based on the tumor cell genome data; Extract features from the CDR3 sequences and antigen peptide data to obtain feature matrix data; Input the tumor gene expression profile data, gene mutations, and feature matrix data into a neural network for training to obtain an immunogenicity prediction model.

[0046] S105: Based on the probability of TCR recognizing MHC-antigen peptides and MHC-antigen peptide affinity calculation, obtain the neoantigen load of the subject.

[0047] In one embodiment, the calculation formula of the neoantigen load is:

[0048] where N is the neoantigen load, n represents the number of tumor clone lines analyzed from the exome sequencing data of this patient, b i represents the affinity, r i represents the probability of TCR recognition, c j represents the expression level of gene j.

[0049] In a specific embodiment, tumor-specific antigens recognized by TCRs are screened: sequence and structural features of antigen peptides and TCR variable regions are extracted from genomic data, and combined with experimentally verified TCR-antigen peptide binding datasets for regression analysis and correlation analysis to discover key factors for TCR recognition of tumor antigens; a deep learning model is constructed with antigen- and TCR-related features as inputs to predict the probability of TCR recognition and binding to tumor antigens; gene mutations of different tumors are integrated to construct a gene mutation-TCR correspondence matrix to analyze the probability of tumor-specific antigens being recognized by TCRs from the perspective of tumor clonal evolution; combined with tumor-infiltrating lymphocyte (TIL) data and immunogenicity data obtained from experiments, with tumor gene expression profiles and antigen peptide feature matrices as inputs to a deep neural network, the immunogenicity of tumor antigens is predicted.

[0050] Define and calculate the neoantigen load index: The activation of T cell responses by tumor antigens is a multi-step complex process, in which antigen presentation and TCR recognition of the pMHC complex are the main steps affecting T cell recognition and killing of target cells, and failure of any step will lead to the failure of immunotherapy. For antigen i of tumor clone line j, assuming the affinity of the MHC-I molecule for this antigen peptide is bi and the probability of TCR recognition of this antigen is ri, the neoantigen load of this patient is defined as:

[0051]

[0052] where n represents the number of tumor clone lines analyzed from the exome sequencing data of this patient.

[0053] It should be noted that when developing the bioinformatics model of the present invention, multiple omics data will be used, including genomic or exome, transcriptomic, and mass spectrometry data. Once the prediction model is trained, only the patient's exome sequencing or large panel sequencing data is required during use, which helps to reduce the cost of clinical companion diagnosis and promote market application and promotion.

[0054] In one embodiment, the present invention aims to conduct tumor neoantigen screening driven by big data and deep learning. Based on the integration of multiple omics data such as genomics, transcriptomics, and proteomics, in-depth exploration is carried out on tumor neoantigens generated by different gene mutation types, the affinity between MHC molecules and antigen peptides, TCR recognition of antigen peptides, and clinical response biomarkers of PD-L1 blockers. High-dimensional heterogeneous features are extracted from the biological processes of tumor antigen generation, presentation, and immune response, and multiple intelligent algorithms including deep learning are developed and used to screen high-quality neoantigens with immunogenicity, such as Figure 4 shown.

[0055] In a specific embodiment, large-scale mass spectrometry data of MHC-I-bound antigenic peptides are integrated to construct a mass spectrometry-based MHC-antigenic peptide benchmark dataset and develop an online database; genomic or whole-exome sequencing data are analyzed for MHC typing, and the sequence and structural features of antigenic peptides and MHC binding grooves are extracted. A deep learning-based MHC-antigenic peptide affinity prediction model is proposed and developed; an MHC-antigenic peptide heterogeneous network integrating multi-source data is constructed, the network topology features and ontology features are combined, and a deep feedforward network is trained to predict the MHC-antigenic peptide affinity.

[0056] It is intended to improve the MHC-I and antigenic peptide affinity prediction performance from three aspects: 1) Integrate large-scale mass spectrometry data of MHC-I-bound antigenic peptides to construct a mass spectrometry-based MHC-antigenic peptide benchmark dataset. First, pre-train a deep neural network-based model using IEDB data, and then perform transfer learning on the pre-trained model with the mass spectrometry data to obtain a higher-performance prediction model, as Figure 5 shown; 2) Combine sequence and structural features to predict the MHC-I and antigenic peptide affinity: Extract the sequences of antigenic peptides and MHC-I binding grooves from the protein database, use one-hot encoding for amino acid sequences, and use various methods to extract sequence features including physicochemical properties, local structure entropy, pairing potential, interaction propensity, etc. At the same time, extract secondary structure features, structural neighbor features, solvent accessible surface area and other structural features. Combine these sequence- and structure-based features to construct a training set, learn a gradient boosting regression tree model, determine the important factors for the affinity between MHC-I molecules and antigenic peptides, and develop an online prediction service. 3) Calculate the MHC-I and MHC-II subtype similarity networks and antigenic peptide similarity networks, integrate the MHC I molecule-antigenic peptide associations in IEDB, thereby construct an MHC-antigenic peptide heterogeneous network integrating multi-source heterogeneous data, and use the random walk algorithm on the two-pass heterogeneous network to predict the probability of potential MHC molecules binding to antigenic peptides.

[0057] In a specific embodiment, TCR recognition of tumor-specific antigens and immunogenicity prediction: Extract the sequence and structural features of antigenic peptides and TCR variable regions from genomic data, combine with an experimentally verified TCR-antigenic peptide binding dataset, perform regression analysis and correlation analysis to discover the key factors for TCR recognition of tumor antigens; construct a deep learning model, use antigen- and TCR-related features as inputs to predict the probability of TCR recognition and binding to tumor antigens; integrate gene mutations of different tumors to construct a gene mutation-TCR correspondence matrix, and analyze the probability of tumor-specific antigens being recognized by TCR from the perspective of tumor clonal evolution; combine tumor-infiltrating lymphocyte (TIL) data and experimentally obtained immunogenicity data, use tumor gene expression profiles and antigenic peptide feature matrices as inputs to a deep neural network to predict the immunogenicity of tumor antigens.

[0058] CD8+ T lymphocytes recognize and kill tumor cells through TCR-dependent complementarity-determining region 3 (CDR3). The interaction between CDR3 and antigen peptides presented by MHC is the key to the adaptive immune response. TCR is determined by the V(D)J recombination process, and the human TCR repertoire can accommodate up to 10^15 different molecular types. Unless from the same clone, the TCRs of each T lymphocyte are different, but different TCRs can bind the same antigen, which poses great difficulties for computationally predicting immunogenic tumor antigens. We believe that the TCR regions of T cells recognizing the same pMHC complex contain conserved sequence features (motifs), and we plan to develop three methods to predict the probability of TCR recognizing tumor antigens and activating immune responses: 1) Identification of the main factors for TCR-epitope binding: Collect and integrate experimentally verified antigen epitope-TCR interaction data, extract the sequences and structures of TCRs and peptide chains, including the charge, hydrophobicity, two-dimensional structure features of CDR3 and the one-hot encoding of antigens, etc. Use group sparse regularization regression to analyze the influence of each feature on TCR recognition of antigens, construct a deep neural network model and train it with antigen-TCR interaction data to predict the probability of a specific TCR subtype recognizing a specific antigen. 2) TCR-epitope binding prediction based on CNN: Obtain the CDR3 region sequence from RNA-seq data, and use autocross covariation to convert CDR3 sequences of different lengths into feature matrices of the same dimension; obtain gene mutations and antigen peptides from genomic or exome sequencing data and convert them into feature matrices using one-hot encoding; The prediction of MHC-antigen peptide affinity and TCR-epitope binding share the features of antigen peptide sequences and structures, construct a multi-task deep learning model, and perform two tasks: predicting MHC-antigen peptide affinity and TCR-epitope binding simultaneously; 3) Profile the TCR repertoire of RNA-seq data from different tumor-infiltrating lymphocytes (TILs), combine genomic or exome sequencing data of tumor cells to obtain gene mutations, collect experimentally obtained immunogenic data sets, and use the tumor gene expression profile, antigen peptide feature matrix and CDR3 feature matrix as the input of a deep neural network to predict the immunogenicity of tumor antigens.

[0059] In a specific embodiment, the approval and marketing of immune checkpoint inhibitors such as anti-PD-L1 monoclonal antibody and anti-CTLA-4 monoclonal antibody have greatly promoted the research of tumor immunotherapy. However, the population that currently benefits from immunotherapy is very limited. For example, the objective response rate of anti-PD-L1 inhibitor monotherapy is only about 30%. The companion diagnostic indicators of immune checkpoint inhibitors, such as tumor mutational burden (TMB) and PD-L1 protein level, cannot accurately screen patients who can benefit from anti-PD-L1 monoclonal antibody treatment because cytotoxic T lymphocytes cannot recognize tumor-specific antigen epitopes. Therefore, screening individual-specific biomarkers by integrating multi-omics data as companion diagnostics for immune checkpoint inhibitors is the direction of the development of precision medicine. This project intends to collect and integrate clinical data such as the response of various tumors to the immune checkpoint inhibitor PD-L1, progression-free survival, and overall survival, construct a benchmark clinical dataset of PD-L1 inhibitors, and develop two methods for biomarker screening on this basis: 1) Analyze the association between neoantigens generated by tumor driver mutations and passenger mutations and the response and prognosis of PD-L1 inhibitors, and discover driver gene mutations with statistical significance; 2) Obtain gene mutations and MHC subtypes from the exome sequencing data of tumor cells and paired normal cells, obtain the gene expression and TCR repertoire profiles from RNA-seq data, integrate relevant features such as antigen peptide sequences, tumor gene expression profiles, MHC typing, and TCR repertoire, combine large-scale clinical immune response data, extract features from high-dimensional heterogeneous data, and train a deep neural network model to achieve high-precision prediction of tumor-specific immune checkpoint response and prognosis, providing more reference indicators for the companion diagnosis of immune checkpoint clinical treatment.

[0060] The disclosed embodiments of the present invention also provide a computer program product or system, including a computer program, which implements the steps of the above method for calculating neoantigen load when executed by a processor.

[0061] Figure 2 The schematic diagram of the system for calculating neoantigen load provided by the embodiments of the present invention specifically includes: Acquisition unit: Acquire the sequencing data of the subject to be tested, and the sequencing data is exome sequencing data or panel sequencing data; Extraction unit: Extract the sequence and structural features of antigen peptides, the sequence and structural features of MHC binding grooves, and the sequences of TCR variable regions based on the sequencing data; Affinity unit: Calculate the MHC-antigen peptide affinity based on the sequence and structural features of MHC binding grooves and the sequence and structural features of antigen peptides; Recognition unit: Calculate the probability of TCR recognizing MHC-antigen peptides based on the sequences of TCR variable regions; Calculation unit: Calculate the neoantigen load of the subject based on the probability of MHC-antigen peptide recognition by the TCR and the MHC-antigen peptide affinity.

[0062] Figure 3 Schematic diagram of the calculation device for neoantigen load provided by the embodiments of the present invention, specifically including: A memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, any of the above-mentioned neoantigen load calculation methods.

[0063] The disclosed embodiments of the present invention also provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, any of the above-mentioned neoantigen load calculation methods.

[0064] The embodiments of the present invention provide a method for predicting the prognosis of tumor treatment based on neoantigen load, including: Obtain the tumor slice data and serum data of the subject; Perform neoantigen screening based on the tumor slice data and serum data to obtain neoantigens; Calculate the load of the neoantigens based on the above-mentioned neoantigen load calculation method to obtain the neoantigen load; Perform prognosis prediction based on the neoantigen load to obtain a prediction result.

[0065] In a specific embodiment, sequencing of tumor and serum clinical samples and collection of clinical data of esophageal cancer patients: Among the patients first diagnosed with esophageal cancer in this hospital, cases meeting the following conditions were selected as research objects: aged between 18 and 70 years old; clinical stage IIIA-IV, unable to undergo surgical treatment; no history of malignant tumors; no history of chemotherapy; before chemotherapy, the patient's blood routine, heart, liver and kidney functions were all normal. All research objects adopted a combined chemotherapy regimen mainly based on cisplatin or carboplatin. Informed consent forms. All patients received 4-6 cycles of chemotherapy, and PD-L1 inhibitors were used during chemotherapy. According to their clinical chemotherapy data, the chemotherapy efficacy of the patients was evaluated according to the RECIST standard: complete remission (CR) - all target lesions disappeared; partial remission (PR) - the sum of the longest diameters of the target lesions decreased by at least 30%; stable disease (SD) - taking the minimum value of the sum of the longest diameters at the start of treatment as a reference, not meeting the PR standard and not meeting the PD standard; progression (PD) - the sum of the longest diameters of the target lesions increased by at least 20%, taking the minimum value of the sum of the longest diameters at the start of treatment or when one or more new lesions appeared as a reference.

[0066] The chemotherapy-related grade 3 and 4 toxic and side reactions of all patients were evaluated according to the evaluation criteria of the National Cancer Institute (NCI 3.0). The adverse reactions caused by chemotherapy included leukopenia, neutropenia, thrombocytopenia, anemia, nausea, vomiting, diarrhea, etc. The adverse reactions were routinely divided into three groups for subsequent analysis: (i) all grade 3 or 4 toxicity reactions; (ii) all grade 3 or 4 hematological toxicity reactions; (iii) all grade 3 or 4 gastrointestinal toxicity reactions.

[0067] Referring to the strategies for the construction of the NIH specimen bank and the tumor cohort study in the United States, an esophageal cancer specimen bank was established that collected complete clinical treatment and follow-up data, experimental research data, and tumor tissue and body fluid samples at each treatment and follow-up stage according to the unified standards of this project. A third-party sequencing agency was commissioned to perform whole-exome sequencing on the tumor and serum samples, obtain the original fastq files and perform mutation analysis, extract the tumor antigen sequences, and use the developed neoantigen screening model to calculate the neoantigen load index for each patient. At the same time, prognostic indicators such as PD-L1 expression level, TMB, and MSI were collected as the basis for control analysis.

[0068] Clinical verification of the prognostic role of neoantigen load in the anti-PD-L1 treatment of esophageal cancer patients: 100 patients were enrolled in a randomized controlled study. The treatment group was treated with concurrent chemoradiotherapy combined with the PD-L1 antibody Imfinzi (durvalumab), and the control group was treated with concurrent chemoradiotherapy combined with a placebo. The treatment process was carried out strictly in accordance with the NCCN clinical practice guidelines, and the clinical data of the patients, such as biochemistry, immunity, and imaging, were recorded. At the same time, for the related toxic and side effects, the hematological toxicity, gastrointestinal reactions, heart damage, lung damage, skin damage, etc. were mainly evaluated. According to the genomic sequencing data of the TCGA public database, the critical value of neoantigen load for patient stratification in esophageal cancer patients was calculated. The patients were divided into two groups, and the prognostic indicators such as OS, objective response rate (ORR), overall response rate (OR), progression-free survival (PFS), etc. of the two groups of patients were statistically analyzed to determine the statistical significance of neoantigen load in each prognostic indicator. In particular, it was compared with indicators such as PD-L1+, TMB, and MSI to verify the prognostic role of neoantigen load.

[0069] In a specific embodiment, the approved marketing of immune checkpoint inhibitors such as anti-PD-L1 monoclonal antibody and anti-CTLA-4 monoclonal antibody has greatly promoted the research of tumor immunotherapy. However, currently, the population that benefits from immunotherapy is very limited. For example, the objective response rate of anti-PD-L1 inhibitor monotherapy is only about 30%. The companion diagnostic indicators TMB and PD-L1 protein level of immune checkpoint inhibitors cannot accurately screen patients who can benefit from anti-PD-L1 monoclonal antibody treatment because cytotoxic T lymphocytes cannot recognize tumor-specific antigen epitopes. Therefore, screening individual-specific biomarkers by integrating multi-omics data as companion diagnostics for immune checkpoint inhibitors is the development direction of precision medicine. The present invention intends to collect and integrate clinical data such as the responses of various tumors to the immune checkpoint inhibitor PD-L1, progression-free survival, and overall survival, construct a benchmark PD-L1 inhibitor clinical data set, and develop two methods for biomarker screening on this basis: 1) Analyze the correlation between neoantigens generated by tumor driver mutations and passenger mutations and the response and prognosis of PD-L1 inhibitors respectively, and identify driver gene mutations with statistical significance; 2) Obtain gene mutations and MHC subtypes from the exome sequencing data of tumor cells and paired normal cells, obtain the spectra of gene expression and TCR repertoire from RNA-seq data, integrate relevant features such as antigen peptide sequences, tumor gene expression profiles, MHC typing, and TCR repertoire, combine large-scale clinical immune response data, extract features from high-dimensional heterogeneous data, and train a deep neural network model to achieve high-precision prediction of tumor-specific immune checkpoint response and prognosis, providing more reference indicators for the companion diagnosis of immune checkpoint clinical treatment.

[0070] In a specific embodiment, for the anti-PD-L1 clinical treatment of esophageal cancer patients, the marketed PD-L1 antibody Imfinzi (durvalumab) is used to treat the enrolled patients. The treatment process is strictly carried out in accordance with the NCCN clinical diagnosis and treatment plan, and the clinical data of the patients, such as biochemistry, immunity, and imaging, are recorded. At the same time, for the related side effects, the side effects such as hematological toxicity, gastrointestinal reactions, heart damage, lung damage, and skin damage are mainly evaluated. According to the calculated neoantigen load, the esophageal cancer patients are stratified and grouped, and statistical analysis is performed on the prognostic indicators such as OS, objective response rate (ORR), overall response rate (OR), and progression-free survival (PFS) of the patients to determine the statistical significance of the neoantigen load in each prognostic indicator. In particular, it is compared with indicators such as PD-L1+, TMB, and MSI to verify the prognostic effect of the neoantigen load.

[0071] The verification results of this verification embodiment show that assigning fixed weights to indications can improve the performance of this method compared to the default settings. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk, or optical disc, etc.

[0072] Those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned medium storage can be read-only memory, magnetic disk, or optical disc, etc.

[0073] The above has introduced in detail a computer device provided by the present invention. For those of ordinary skill in the art, according to the idea of the embodiments of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for calculating a new antigen load, characterized in that: include: S1. Obtaining sequencing data of a subject, wherein the sequencing data is exome sequencing data or panel sequencing data; S2, extracting the sequence and structural features of antigenic peptides, the sequence and structural features of MHC binding grooves, and the sequence of TCR variable regions based on sequencing data; S3, calculating the MHC-antigen peptide affinity based on the sequence and structural characteristics of the MHC binding groove and the sequence and structural characteristics of the antigen peptide; S4, calculating the probability of TCR recognizing MHC-antigen peptide based on the sequence of TCR variable region; S5. The new antigen load of the subject is calculated based on the probability of the TCR recognizing the MHC-antigen peptide and the MHC-antigen peptide affinity.

2. The method for calculating the neoantigen load according to claim 1, characterized in that: The extraction is to obtain the sequence and structural features of the antigen peptide and the sequence and structural features of the MHC binding groove through the trained first deep network model; Optionally, the training process of the first deep network model is: Obtain the antigen peptide sequence and the sequence of the MHC binding groove from the protein database; Extracting structural features from the sequence to obtain structural features; Constructing a feature training set based on the sequence and structural features; Inputting the feature training set into the first deep network model for training to obtain a trained first deep network model; Optionally, the structural features include one or more of the following: structural features, structural neighbor features, volume accessible surface area; Optionally, the feature extraction method adopts one or more of the following: physical and chemical extraction, local structure entropy extraction, pairing potential energy extraction, interaction tendency extraction; Optionally, the first deep network model includes one or more of the following: GBRT, XGBoost, LightGBM, CatBoost, Random Forest, AdaBoost.

3. The method for calculating the neoantigen load according to claim 1, characterized in that: The S2 is replaced by: inputting the sequencing data into a trained second deep network model to obtain MHC-antigen peptide affinity, and the training process of the second deep network model is: The first step is to obtain the MHC antigen peptide mass spectrometry data set and IEDB data; The second step is to pre-train the neural network model using the IEDB data to obtain the weight of the training model; Step 3: Migrate the weights of the training model to the second deep network model and input the MHC antigen peptide mass spectrometry dataset into the second deep network model for training to obtain a trained second deep network model; Optionally, the first step further includes MHC typing, obtaining antigen peptide data and MHC molecules, typing the MHC molecules to obtain MHC-I class data, and constructing an MHC-I class antigen peptide mass spectrum data set of the MHC-I class data and antigen peptides based on mass spectrometry; inputting the MHC-I class antigen peptide mass spectrum data set into the training model for training to obtain a second deep network model; Optionally, the third step further comprises antigen depth features, obtaining antigen gene expression level data, extracting features of the antigen gene expression level data to obtain antigen depth features, fusing the antigen depth features with the MHC-I class antigen peptide mass spectrum data and then inputting them into the second deep network model for training to obtain a trained second deep network model; Optionally, S2 is replaced by: inputting the sequencing data into a trained multi-source heterogeneous network model to obtain MHC-antigen peptide affinity, and the training process of the multi-source heterogeneous network model is: Obtaining the MHC molecule structure and the antigen peptide structure to generate the MHC network structure and the antigen peptide network structure; the MHC molecule includes the MHC-I subtype and the MHC-II subtype; Calculate the similarity network between MHC-I subtypes and MHC-II subtypes, and the similarity network of antigen peptides; Obtain the association between MHC molecules and antigen peptides; Based on the similarity network of MHC-I subtype and MHC-II subtype, the similarity network of antigen peptide, the association relationship between MHC molecules and antigen peptides is input into the third deep network model to generate a multi-source heterogeneous network; Optionally, the multi-source heterogeneous network model adopts one or more of the following: a two-way heterogeneous network, a heterogeneous information network, a graph neural network, and a hypergraph; Optionally, S2 further includes executing the second deep network model in parallel to obtain MHC-antigen peptide affinity, performing weight distribution calculation on the MHC-antigen peptide affinity obtained by the first deep network model and the second deep network model to obtain the second MHC-antigen peptide affinity, and calculating the new antigen load of the subject based on the probability of the TCR recognizing the MHC-antigen peptide and the second MHC-antigen peptide affinity; Optionally, S2 further includes executing the multi-source heterogeneous network model in parallel to obtain MHC-antigen peptide affinity, performing weight distribution calculation on the MHC-antigen peptide affinity obtained by the first deep network model and the multi-source heterogeneous network model to obtain a third MHC-antigen peptide affinity, and calculating the new antigen load of the subject based on the probability of the TCR recognizing the MHC-antigen peptide and the third MHC-antigen peptide affinity; Optionally, S2 also includes executing the second deep network model and the multi-source heterogeneous network model in parallel to obtain MHC-antigen peptide affinity, performing weight distribution calculation on the MHC-antigen peptide affinity obtained by the first deep network model, the multi-source heterogeneous network model, and the second deep network model to obtain a fourth MHC-antigen peptide affinity, and calculating the new antigen load of the subject based on the probability of TCR recognizing the MHC-antigen peptide and the fourth MHC-antigen peptide affinity.

4. The method for calculating the neoantigen load according to claim 1, characterized in that: The sequence of the TCR variable region is extracted from the genome data of the subject; Optionally, the probability calculation of the TCR recognizing the MHC-antigen peptide is obtained by calculating the recognition model, and the construction process of the recognition model is: Obtain epitope-TCR interaction data, TCR and peptide chain sequences and structures; Extracting TCR characteristics based on the sequence and structure data of the TCR and peptide chain; Inputting the TCR characteristics and antigen epitope-TCR interaction data into a fourth deep network model for training to obtain a recognition model; Optionally, the TCR characteristics include one or more of the following: charge, hydrophobicity, and two-dimensional structural characteristics of CDR3; Optionally, the TCR feature is extracted by group sparse regularized regression analysis; Optionally, the probability calculation of TCR recognition also includes TCR-epitope binding prediction, and the probability of TCR recognition is obtained by the recognition model and TCR-epitope binding prediction; the TCR-epitope binding prediction is obtained by predicting the trained TCR-epitope binding prediction model; Optionally, the trained TCR-antigen epitope binding prediction model is constructed by: Obtain the sequence, gene mutation data and antigenic peptides of the CDR3 region; Converting the sequence of the CDR3 region, gene mutation data, and antigenic peptides into a feature matrix; Inputting the feature matrix into a neural network for training to obtain a TCR-antigen epitope binding prediction model; Optionally, the characteristic matrix of the antigenic peptides includes the characteristics of the sequence and structure of the antigenic peptides; Optionally, the TCR-antigen epitope binding prediction model also includes MHC-antigen peptide affinity calculation, and the affinity prediction result is obtained by combining the sequence characteristics and structural characteristics of the antigen peptide.

5. The method for calculating the neoantigen load according to claim 4, characterized in that: The probability calculation of TCR recognition also includes the prediction of the immunogenicity of the tumor antigen, and the probability of TCR recognition is obtained by the recognition model and the prediction of the immunogenicity of the tumor antigen, and the prediction of the immunogenicity of the tumor antigen is obtained by predicting the trained tumor antigen immunogenicity prediction model; Optionally, the process of constructing the immunogenicity prediction model of the tumor antigen is: Obtain RNA-seq data of different tumor-infiltrating lymphocytes, tumor cell genomes, and immunogenicity data; Performing TCR library spectrum analysis on the RNA-seq data to obtain tumor gene expression spectrum data and obtaining CDR3 sequences from the RNA-seq data; Acquiring gene mutation and antigen peptide data based on the tumor cell genome data; Feature extraction is performed on CDR3 sequence and antigen peptide data to obtain feature matrix data; The tumor gene expression profile data, gene mutation, and feature matrix data are input into a neural network for training to obtain an immunogenicity prediction model.

6. The method for calculating the neoantigen load according to claim 1, characterized in that: The calculation formula of the new antigen load is: Where N is the neoantigen load, n represents the number of tumor clones analyzed from the patient's exome sequencing data, b i Express affinity, r i represents the probability of TCR recognition, c j represents the expression level of gene j.

7. A method for predicting tumor treatment prognosis based on neoantigen load, characterized in that: include: Obtaining tumor slice data and serum data of the subject to be tested; Performing new antigen screening based on the tumor slice data and serum data to obtain new antigens; Calculating the load of the neoantigen based on the neoantigen load calculation method described in claims 1-6 to obtain the neoantigen load; Prognosis prediction is performed based on the neoantigen load to obtain a prediction result.

8. A computer program product comprising a computer program or instructions, characterized in that: The computer program or instructions are executed by the processor to implement the method for calculating the neoantigen load described in any one of claims 1 to 6 or to implement the method for predicting tumor treatment prognosis based on the neoantigen load described in claim 7.

9. A computer device comprising a memory, a processor and a computer program or instruction stored in the memory, characterized in that: The computer program or instructions are executed by the processor to implement the method for calculating the neoantigen load described in any one of claims 1 to 6, or to implement the method for predicting tumor treatment prognosis based on the neoantigen load described in claim 7.

10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: The computer program or instructions are executed by the processor to implement the method for calculating the neoantigen load described in any one of claims 1 to 6, or to implement the method for predicting tumor treatment prognosis based on the neoantigen load described in claim 7.

Citation Information

Patent Citations

  • Rapid screening method for effective new antigenic peptide of tumor individualized vaccine

    CN110257478A

  • Method and system for calculating tumor neoantigen load

    CN112309502A

  • Tumor neoantigen detection and screening method and system combining molecular omics and computational structure

    CN114333999A

  • Molecular conjugation epitope prediction method and device, storage medium and computer equipment

    CN116959621A

  • Screening method of tumor neoantigen

    CN119049554A