Method, apparatus and program product for screening of neoantigens

By processing tumor slices and serum data, and using a deep network model to screen the affinity between HLA/MHC pseudosequences and antigen peptides and the TCR recognition probability, the problem of inaccurate new antigen screening in existing technologies was solved, and the predictive effect of immune checkpoint inhibitor treatment in patients with esophageal cancer was improved.

CN120048333BActive Publication Date: 2025-10-213201 HOSPITAL +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510166355.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-10-21
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately screen new antigens that can activate T cells, resulting in poor therapeutic effects of immune checkpoint inhibitors and a lack of reliable biomarkers. In particular, it is difficult to determine which patients with esophageal cancer can benefit from immune checkpoint inhibitor treatment.

Method used

By obtaining the patient's tumor slice data and serum data, data processing is performed to obtain HLA/MHC pseudo sequences, antigen peptide data and antigen gene expression levels. New antigens are screened using a deep network model, combining the affinity of HLA/MHC pseudo sequences and antigen peptides and TCR recognition probability for screening.

Benefits of technology

It improves the accuracy and reliability of neoantigen screening and can better judge the efficacy of immune checkpoint inhibitors, especially has a better predictive effect in Anti-PD-L1 treatment of patients with esophageal cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048333B_ABST
    Figure CN120048333B_ABST
Patent Text Reader

Abstract

The application relates to the field of intelligent medical treatment, in particular to a screening method, device and program product of a new antigen. The method comprises the following steps: S1, acquiring tumor section data and serum data of a patient; S2, performing data processing on the tumor section data and the serum data to obtain HLA / MHC pseudo-sequence data, antigen peptide data and antigen gene expression level data; the data processing comprises data sequencing, HLA / MHC typing, somatic mutation calculation, fusion gene calculation and gene expression spectrum calculation; and S3, screening a new antigen based on the HLA / MHC pseudo-sequence data, the antigen peptide data and the antigen gene expression level data to obtain the new antigen. The application can screen the new antigen and has good clinical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent medicine, and specifically to a method, device, program product, and computer-readable storage medium for screening new antigens. Background Art

[0002] Esophageal squamous cell carcinoma (ESCC) accounts for over 90% of esophageal cancer cases. For patients with unresectable or ineligible ESCC, traditional chemotherapy, radiotherapy, and targeted therapies have been ineffective. In recent years, substantial progress has been made in tumor immunotherapy, particularly immune checkpoint inhibition. In 2017, the FDA approved pembrolizumab as a second-line treatment for PD-L1-positive recurrent locally advanced or metastatic ESCC. This has sparked a surge in research on esophageal cancer immunotherapy and its companion diagnostic markers both domestically and internationally. However, reliable biomarkers are lacking for anti-PD-L1 therapy in esophageal cancer. Therefore, biomarker detection is crucial for patient identification in the clinical application of immune checkpoint inhibitors like CTLA4 and PD-L1 monoclonal antibodies. Currently, commonly used companion diagnostic markers include PD-L1 expression levels, tumor mutation burden (TMB), microsatellite instability (MSI-H), and DNA mismatch repair deficiency (dMMR). Although these prognostic markers play a guiding role in clinical applications, a growing number of clinical cohort studies have shown that these markers cannot accurately distinguish which patients will benefit from immune checkpoint inhibitor treatment. Therefore, the concept of tumor neoantigen burden (TNB) has been proposed. Tumor neoantigens are the number of mutations actually targeted by T cells and can better judge the clinical efficacy of immune checkpoint inhibitors. However, the prediction of binding affinity for screening neoantigens is imprecise, and immunogenicity assessment is difficult (not all peptides that can bind to MHC can effectively activate T cells, and existing methods cannot accurately predict which mutations will lead to truly immunogenic neoantigens). Summary of the Invention

[0003] In response to the above problems, the present invention provides a method for screening new antigens, which specifically includes:

[0004] S1. Obtain the patient's tumor slice data and serum data;

[0005] S2. Processing the tumor slice data and serum data to obtain HLA / MHC pseudo sequences, antigen peptide data, and antigen gene expression levels; the data processing includes data sequencing, HLA / MHC typing, somatic mutation calculation, fusion gene calculation, and gene expression profile calculation;

[0006] S3. Screening new antigens based on the HLA / MHC pseudo sequence, antigen peptide data, and antigen gene expression level data to obtain new antigens.

[0007] The data processing process is as follows:

[0008] S21, performing exome sequencing based on the tumor section data to obtain exome sequencing data;

[0009] S22, performing transcriptome sequencing based on the serum data to obtain transcriptome sequencing data;

[0010] S23. Performing HLA / MHC typing based on the exome sequencing data to obtain HLA / MHC pseudo sequences and performing somatic mutation calculation to obtain mutation data;

[0011] S24, calculating gene expression profile data and fusion gene data based on the transcriptome sequencing data;

[0012] S25. Calculating antigen gene expression level data based on the gene expression profile data;

[0013] S26. The fusion gene and the mutation data are fused to obtain antigen peptide data.

[0014] The screening is performed by inputting HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression level data into a deep network model to obtain the affinity of the HLA / MHC pseudo-sequences and antigen peptide data, combining the HLA / MHC pseudo-sequences with the antigen peptides to obtain HLA / MHC-antigen peptides, and then screening to obtain new antigens based on the probability of TCR recognizing the HLA / MHC-antigen peptides;

[0015] Optionally, the training process of the trained deep network model is:

[0016] The first step is to obtain HLA / MHC pseudo sequences, antigen peptide data, antigen gene expression level data, and IEDB data;

[0017] The second step is to pre-train the initial deep network model using the IEDB data to obtain the model weights;

[0018] Step 3: Migrating the weights of the model to the deep network model to obtain a migrated deep network model;

[0019] Step 4: Vectorize the HLA / MHC pseudo sequence and antigen peptide data to obtain HLA / MHC pseudo sequence vector data and antigen peptide vector data;

[0020] Step 5: Extract features from the antigen gene expression level data to obtain the antigen depth feature vector;

[0021] Step 6: Fuse the HLA / MHC pseudo sequence vector data, antigen peptide vector data, and antigen deep feature vector to obtain fused data, and input it into the migration deep network model for training to obtain a trained deep network model;

[0022] Optionally, the training process of the affinity prediction model further includes HLA / MHC antigen peptide feature data, obtaining sequence feature data and structural feature data of the HLA / MHC antigen peptide, and inputting the sequence feature data and structural feature data of the HLA / MHC antigen peptide and the HLA / MHC antigen peptide mass spectrometry data set into the training model for training to obtain the affinity prediction model.

[0023] The training process of the deep network model also includes sequence feature data, obtaining sequence data of antigen peptide sequences and HLA / MHC binding grooves, performing feature extraction on the sequence data of the antigen peptide sequences and HLA / MHC binding grooves to obtain sequence features, and inputting the HLA / MHC pseudo sequence, antigen peptide data, antigen gene expression level data, and sequence features into the deep network model to obtain the affinity of the HLA / MHC pseudo sequence and antigen peptide data;

[0024] The sequence features are extracted through the feature model, and the feature model construction process is as follows:

[0025] Obtain the antigen peptide sequence and the sequence of the HLA / MHC binding groove from the protein database;

[0026] performing feature extraction on the sequence to obtain sequence features;

[0027] performing structural feature extraction on the sequence to obtain structural features;

[0028] Constructing a feature training set based on the sequence features and structural features;

[0029] Inputting the feature training set into a second deep network model for training to obtain a feature model;

[0030] Optionally, the structural features include one or more of the following: structural features, structural neighbor features, volume accessible surface area;

[0031] Optionally, the feature extraction method adopts one or more of the following: physical and chemical extraction, local structure entropy extraction, pairing potential energy extraction, interaction tendency extraction;

[0032] Optionally, the second deep network model includes one or more of the following: GBRT, XGBoost, LightGBM, CatBoost, Random Forest, AdaBoost.

[0033] The training process of the deep network model also includes similar feature data, obtaining the structures of HLA / MHC molecules and antigen peptides to generate HLA / MHC network structures and antigen peptide network structures, extracting features from the HLA / MHC network structures and antigen peptide network structures to obtain similar features, and inputting the HLA / MHC pseudo sequence, antigen peptide data, antigen gene expression level data, and similar features into the deep network model to obtain the affinity of the HLA / MHC pseudo sequence and antigen peptide data;

[0034] Optionally, the similar features are extracted by a similar feature model, and the similar feature model is constructed by: obtaining the structures of HLA / MHC molecules and antigen peptides to generate an HLA / MHC network structure and an antigen peptide network structure; the HLA / MHC includes HLA / MHC-I subtype and HLA / MHC-II subtype;

[0035] Calculate the similarity network between HLA / MHC-I subtypes and HLA / MHC-II subtypes, and the antigen peptide similarity network;

[0036] Obtain the association between HLA / MHC molecules and antigen peptides;

[0037] Based on the HLA / MHC-I subtype and HLA / MHC-II subtype similarity network and the antigen peptide similarity network, the association relationship between HLA / MHC molecules and antigen peptides is input into the similarity feature model in the third deep network model;

[0038] Optionally, the third deep network model adopts one or more of the following: a two-way heterogeneous network, a heterogeneous information network, a graph neural network, and a hypergraph;

[0039] Optionally, the training process of the deep network model also includes sequence feature data and similar features, obtaining sequence data of antigen peptide sequences and HLA / MHC binding grooves, and performing feature extraction on the sequence data of the antigen peptide sequences and HLA / MHC binding grooves to obtain sequence features; obtaining the structures of HLA / MHC molecules and antigen peptides to generate HLA / MHC network structures and antigen peptide network structures, and performing feature extraction on the HLA / MHC network structures and antigen peptide network structures to obtain similar features; and inputting the HLA / MHC pseudo-sequences, antigen peptide data, antigen gene expression level data, sequence features, and similar features into the deep network model to obtain the affinity of the HLA / MHC pseudo-sequences and antigen peptide data.

[0040] The TCR recognition probability is calculated by a trained recognition model, and the construction process of the recognition model is as follows:

[0041] Obtain epitope-TCR interaction data, TCR and peptide chain sequences and structures;

[0042] Extracting TCR characteristics based on the sequence and structural data of the TCR and peptide chain;

[0043] Inputting the TCR characteristics and antigen epitope-TCR interaction data into a fourth deep network model for training to obtain a recognition model;

[0044] Optionally, the TCR characteristics include one or more of the following: charge, hydrophobicity, and two-dimensional structural characteristics of CDR3;

[0045] Optionally, the extracted TCR features are obtained by group sparse regularized regression analysis;

[0046] Optionally, the probability calculation of TCR recognition further includes TCR-epitope binding prediction, and the probability of TCR recognition is obtained by the recognition model and TCR-epitope binding prediction; the TCR-epitope binding prediction is obtained by predicting using a trained TCR-epitope binding prediction model;

[0047] Optionally, the process of constructing the trained TCR-antigen epitope binding prediction model is:

[0048] Obtain the sequence, gene mutation data, and antigenic peptides of the CDR3 region;

[0049] Converting the sequence of the CDR3 region, gene mutation data, and antigenic peptides into a feature matrix;

[0050] Inputting the feature matrix into a neural network for training to obtain a TCR-antigen epitope binding prediction model;

[0051] Optionally, the characteristic matrix of the antigenic peptides includes sequence characteristics and structural characteristics of the antigenic peptides;

[0052] Optionally, the TCR-antigen epitope binding prediction model also includes MHC-antigen peptide affinity calculation, and the affinity prediction results are obtained based on the sequence characteristics and structural characteristics of the antigen peptide;

[0053] Optionally, the probability calculation of TCR recognition also includes prediction of the immunogenicity of the tumor antigen, and the probability of TCR recognition is obtained by the recognition model and the immunogenicity prediction of the tumor antigen, and the immunogenicity prediction of the tumor antigen is obtained by predicting the immunogenicity prediction model of the trained tumor antigen;

[0054] Optionally, the process of constructing the tumor antigen immunogenicity prediction model is:

[0055] Obtain RNA-seq data of different tumor-infiltrating lymphocytes, tumor cell genomes, and immunogenicity data;

[0056] Performing TCR library analysis on the RNA-seq data to obtain tumor gene expression profile data and obtaining CDR3 sequences from the RNA-seq data;

[0057] Acquiring gene mutation and antigen peptide data based on the tumor cell genome data;

[0058] Feature extraction is performed on CDR3 sequence and antigen peptide data to obtain feature matrix data;

[0059] The tumor gene expression profile data, gene mutation, and feature matrix data are input into a neural network for training to obtain an immunogenicity prediction model.

[0060] The method further includes calculating the neoantigen load by calculating the affinity of the HLA / MHC-antigen peptide of the neoantigen and the probability of TCR recognizing the HLA / MHC-antigen peptide;

[0061] Optionally, the formula for calculating the neoantigen load is:

[0062]

[0063] Where N is the neoantigen load, n represents the number of tumor clones analyzed from the patient's exome sequencing data, b i Expresses affinity, r i represents the probability of TCR recognition, c j represents the expression level of gene j.

[0064] The object of the present invention is to provide a computer program product comprising a computer program or instructions, which are executed by a processor to implement the above-mentioned method for screening new antigens.

[0065] An object of the present invention is to provide a computer device comprising a memory, a processor, and a computer program or instructions stored on the memory, wherein the computer program or instructions are executed by the processor to implement the above-mentioned method for screening new antigens.

[0066] An object of the present invention is to provide a computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions are executed by a processor to implement the above-mentioned method for screening new antigens.

[0067] Advantages of the present invention:

[0068] 1. Propose a new antigen screening method, using a two-step screening process to obtain new antigens. The first step is to screen the affinity of HLA / MHC-antigen peptides, and the second step is to screen the probability of TCR recognition of HLA / MHC-antigen peptides. The MHC-antigen peptide affinity screening is carried out by constructing a mass spectrometry-based MHC-antigen peptide benchmark dataset, analyzing genome or whole-exome sequencing data for MHC typing, extracting the sequence and structural features of the antigen peptide and the MHC binding groove, proposing and developing a deep learning-based MHC-antigen peptide affinity prediction model, and constructing an MHC-antigen peptide heterogeneous network that integrates multi-source data. The network topology features and ontology features are combined to train a deep feedforward network to predict MHC-antigen peptide affinity. Affinity calculation is performed from multiple perspectives and multiple features to improve the accuracy and reliability of affinity calculation.

[0069] 2. Research on computational prediction of TCR binding to tumor antigens is just beginning, benefiting from the accumulation of experimentally validated TCR-antigen data in recent years. A few researchers have attempted to develop machine learning models to predict TCR-antigen affinity. However, currently available training sets are far insufficient compared to the vast human TCR repertoire. Therefore, this paper extracts sequence and structural features of antigen peptides and TCR variable regions from genomic data. Combined with experimentally validated TCR-antigen peptide binding datasets, regression and correlation analyses are performed to identify key factors in TCR recognition of tumor antigens. A deep learning model is constructed that uses antigen- and TCR-related features as input to predict the probability of TCR recognition and binding to tumor antigens. Furthermore, a mutation-TCR correspondence matrix is ​​constructed by integrating genetic mutations from different tumors to analyze the probability of TCR recognition of tumor-specific antigens from the perspective of tumor clonal evolution. Similarly, the probability of TCR recognition is calculated from multiple perspectives to improve the reliability of TCR recognition probabilities, thereby improving the accuracy of calculated neoantigen loads and contributing to the accuracy of predicting outcomes after tumor treatment.

[0070] 3. A method for calculating neoantigen load is proposed. It quantifies neoantigen load by calculating the affinity of MHC-antigen peptides and the probability of TCR recognition. Neoantigen load is then used to predict the prognosis of tumors after treatment, especially for esophageal cancer after anti-PD-L1 treatment. Compared with existing TMB, MSI-H, and dMMR biomarkers, this method has better predictive effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0072] Figure 1 A schematic diagram of a process for screening a new antigen according to an embodiment of the present invention;

[0073] Figure 2 Schematic diagram of a screening system for new antigens provided in an embodiment of the present invention;

[0074] Figure 3 Schematic diagram of a screening device for neoantigens provided in an embodiment of the present invention;

[0075] Figure 4 The present invention provides a high-precision MHC-antigen peptide affinity prediction model trained through transfer learning. (a) A convolutional neural network is trained using data from the IEDB to obtain a pan-cancer MHC-antigen peptide affinity prediction model. (b) A one-step convolutional neural network is trained using transfer learning techniques, combining transcriptome and protein profile data, to obtain a cancer-specific prediction model.

[0076] Figure 5 The clinical trial provided in the embodiments of the present invention verifies the diagnostic and prognostic role of neoantigen load in immune checkpoint inhibitor treatment of esophageal cancer: (a) Tumor and serum samples of esophageal cancer patients were sequenced and data analyzed, and neoantigens were screened and neoantigen load was calculated by combining tumor antigens, HLA typing, and antigen host gene expression levels; (b) Enrolled patients were stratified, and clinical immune checkpoint inhibitor treatment was performed, and clinically relevant indicators and toxic side effects were recorded; (c) Combined with neoantigen load, statistical analysis of immune response and prognosis was performed to verify the role of neoantigen load in diagnosis and prognosis. DETAILED DESCRIPTION

[0077] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0078] In some of the processes described in the specification and claims of the present invention and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to be different types.

[0079] Figure 1 A schematic diagram of a method for screening new antigens provided in an embodiment of the present invention specifically includes:

[0080] S101: Obtaining the patient's tumor slice data and serum data;

[0081] In one embodiment, the tumor includes one or more of the following: brain tumor, glioma, breast cancer, prostate cancer, esophageal cancer, liver cancer, lymphoma, and melanoma.

[0082] In one example, esophageal cancer is a heterogeneous tumor. Screening for neoantigens based on gene mutations within tumor clones and defining and calculating tumor neoantigen burden is a challenging task. This project proposes to use maximum likelihood analysis to integrate multiple variables, including clones, MHC-antigen affinity, and TCR-antigen epitopes, to define a tumor neoantigen burden metric. A critical threshold for neoantigen burden in specific tumors will be determined based on clinical trial data and public data.

[0083] In one specific example, a randomized controlled study enrolled 100 patients. The treatment group received concurrent chemoradiotherapy combined with the PD-L1 antibody Imfinzi (durvalumab), while the control group received concurrent chemoradiotherapy combined with a placebo. Neoantigen load was calculated to stratify esophageal cancer patients. Combined with follow-up data after clinical treatment, the role of neoantigen load as a companion diagnostic and prognostic indicator in esophageal cancer immunotherapy was validated.

[0084] In a specific embodiment, sequencing of tumor and serum clinical samples and collection of clinical data from esophageal cancer patients:

[0085] The present invention team selected patients diagnosed with esophageal cancer for the first time at our hospital and met the following criteria for study: age between 18 and 70 years; clinical stage IIIA-IV, inoperable; no history of malignancy; no history of chemotherapy; and normal blood count, heart, liver, and kidney function before chemotherapy. All subjects received a combination chemotherapy regimen primarily based on cisplatin or carboplatin. Informed consent was obtained. All patients underwent 4-6 cycles of chemotherapy, with PD-L1 inhibitors used during chemotherapy. Based on their clinical chemotherapy data, chemotherapy efficacy was evaluated according to RECIST criteria: complete response (CR) - disappearance of all target lesions; partial response (PR) - at least a 30% reduction in the sum of the longest diameters of target lesions; stable disease (SD) - based on the minimum sum of the longest diameters at the start of treatment, failing to meet either PR or PD criteria; progressive disease (PD) - at least a 20% increase in the sum of the longest diameters of target lesions, based on the minimum sum of the longest diameters at the start of treatment or the appearance of one or more new lesions.

[0086] All patients were evaluated for grade 3 or 4 chemotherapy toxicities according to the National Cancer Institute (NCI 3.0) criteria. Chemotherapy-induced adverse reactions included leukopenia, neutropenia, thrombocytopenia, anemia, nausea, vomiting, and diarrhea. Adverse reactions were routinely divided into three groups for subsequent analysis: (i) all grade 3 or 4 toxicities; (ii) all grade 3 or 4 hematologic toxicities; and (iii) all grade 3 or 4 gastrointestinal toxicities.

[0087] Drawing on the strategies of the US NIH specimen library and tumor cohort studies, a esophageal cancer specimen library was established using the unified standards of this invention. This library contains complete clinical treatment and follow-up data, experimental research data, and tumor tissue and body fluid samples from various treatment and follow-up stages. A third-party sequencing agency was commissioned to perform whole-exome sequencing on tumor and serum samples, obtain raw FastQ files, perform mutation analysis, extract tumor antigen sequences, and use the developed neoantigen screening model to calculate the neoantigen load index for each patient. Furthermore, prognostic indicators such as PD-L1 expression levels, TMB, and MSI were collected as a basis for control analysis.

[0088] S102: Processing the tumor slice data and serum data to obtain HLA / MHC pseudo sequences, antigen peptide data, and antigen gene expression levels; the data processing includes data sequencing, HLA / MHC typing, somatic mutation calculation, fusion gene calculation, and gene expression profile calculation;

[0089] In one embodiment, the data processing process is as follows:

[0090] S21, performing exome sequencing based on the tumor section data to obtain exome sequencing data;

[0091] S22, performing transcriptome sequencing based on the serum data to obtain transcriptome sequencing data;

[0092] S23. Performing HLA / MHC typing based on the exome sequencing data to obtain HLA / MHC pseudo sequences and performing somatic mutation calculation to obtain mutation data;

[0093] S24, calculating gene expression profile data and fusion gene data based on the transcriptome sequencing data;

[0094] S25. Calculating antigen gene expression level data based on the gene expression profile data;

[0095] S26. The fusion gene and the mutation data are fused to obtain antigen peptide data.

[0096] S103: Screening new antigens based on the HLA / MHC pseudo sequence, antigen peptide data, and antigen gene expression level data to obtain new antigens;

[0097] In one embodiment, the screening is performed by inputting HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression level data into a deep network model to obtain the affinity of the HLA / MHC pseudo-sequences and antigen peptide data. The HLA / MHC pseudo-sequences are combined with the antigen peptides to obtain HLA / MHC-antigen peptides, and then the probability of TCR recognition of the HLA / MHC-antigen peptides is used to screen and obtain new antigens.

[0098] In one embodiment, the training process of the trained deep network model is:

[0099] The first step is to obtain HLA / MHC pseudo sequences, antigen peptide data, antigen gene expression level data, and IEDB data;

[0100] The second step is to pre-train the initial deep network model using the IEDB data to obtain the model weights;

[0101] Step 3: Migrating the weights of the model to the deep network model to obtain a migrated deep network model;

[0102] Step 4: Vectorize the HLA / MHC pseudo sequence and antigen peptide data to obtain HLA / MHC pseudo sequence vector data and antigen peptide vector data;

[0103] Step 5: Extract features from the antigen gene expression level data to obtain the antigen depth feature vector;

[0104] Step 6: Fuse the HLA / MHC pseudo-sequence vector data, antigen peptide vector data, and antigen deep feature vector to obtain fused data, and input it into the migration deep network model for training to obtain a trained deep network model.

[0105] In one embodiment, the training process of the affinity prediction model further includes obtaining HLA / MHC antigen peptide feature data, obtaining sequence feature data and structural feature data of the HLA / MHC antigen peptide, and inputting the sequence feature data and structural feature data of the HLA / MHC antigen peptide and the HLA / MHC antigen peptide mass spectrometry data set into the training model for training to obtain the affinity prediction model.

[0106] In one embodiment, the training process of the deep network model also includes sequence feature data, obtaining sequence data of antigen peptide sequences and HLA / MHC binding groove sequence data, performing feature extraction on the antigen peptide sequences and HLA / MHC binding groove sequence data to obtain sequence features, and inputting the HLA / MHC pseudo sequence, antigen peptide data, antigen gene expression level data, and sequence features into the deep network model to obtain the affinity of the HLA / MHC pseudo sequence and antigen peptide data.

[0107] In one embodiment, the sequence features are extracted through a feature model, and the feature model construction process is as follows:

[0108] Obtain the antigen peptide sequence and the sequence of the HLA / MHC binding groove from the protein database;

[0109] performing feature extraction on the sequence to obtain sequence features;

[0110] performing structural feature extraction on the sequence to obtain structural features;

[0111] Constructing a feature training set based on the sequence features and structural features;

[0112] The feature training set is input into the second deep network model for training to obtain a feature model.

[0113] In one embodiment, the structural features include one or more of the following: structural features, structural neighbor features, and volume accessible surface area.

[0114] In one embodiment, the feature extraction method adopts one or more of the following: physical and chemical extraction, local structure entropy extraction, pairing potential energy extraction, and interaction tendency extraction.

[0115] In one embodiment, the second deep network model includes one or more of the following: GBRT, XGBoost, LightGBM, CatBoost, Random Forest, and AdaBoost.

[0116] In one embodiment, the training process of the deep network model further includes similar feature data, obtaining the structures of HLA / MHC molecules and antigen peptides to generate HLA / MHC network structures and antigen peptide network structures, performing feature extraction on the HLA / MHC network structures and antigen peptide network structures to obtain similar features, and inputting the HLA / MHC pseudo sequence, antigen peptide data, antigen gene expression level data, and similar features into the deep network model to obtain the affinity of the HLA / MHC pseudo sequence and antigen peptide data.

[0117] In one embodiment, the similar features are extracted by a similar feature model, and the similar feature model is constructed by: obtaining the structures of HLA / MHC molecules and antigen peptides to generate an HLA / MHC network structure and an antigen peptide network structure; the HLA / MHC includes HLA / MHC-I subtype and HLA / MHC-II subtype;

[0118] Calculate the similarity network between HLA / MHC-I subtypes and HLA / MHC-II subtypes, and the antigen peptide similarity network;

[0119] Obtain the association between HLA / MHC molecules and antigen peptides;

[0120] Based on the HLA / MHC-I subtype and HLA / MHC-II subtype similarity network and the antigen peptide similarity network, the association relationship between HLA / MHC molecules and antigen peptides is input into the similarity feature model in the third deep network model;

[0121] In one embodiment, the third deep network model adopts one or more of the following: a two-way heterogeneous network, a heterogeneous information network, a graph neural network, and a hypergraph.

[0122] In one embodiment, the training process of the deep network model further includes sequence feature data and similarity features, obtaining sequence data of antigen peptide sequences and HLA / MHC binding grooves, and performing feature extraction on the sequence data of the antigen peptide sequences and HLA / MHC binding grooves to obtain sequence features; obtaining the structures of HLA / MHC molecules and antigen peptides to generate HLA / MHC network structures and antigen peptide network structures, and performing feature extraction on the HLA / MHC network structures and antigen peptide network structures to obtain similarity features; and inputting the HLA / MHC pseudo-sequences, antigen peptide data, antigen gene expression level data, sequence features, and similarity features into the deep network model to obtain the affinity of the HLA / MHC pseudo-sequences and antigen peptide data.

[0123] In one embodiment, the TCR recognition probability is calculated using a trained recognition model, and the recognition model is constructed as follows:

[0124] Obtain epitope-TCR interaction data, TCR and peptide chain sequences and structures;

[0125] Extracting TCR characteristics based on the sequence and structural data of the TCR and peptide chain;

[0126] The TCR characteristics and antigen epitope-TCR interaction data are input into the fourth deep network model for training to obtain a recognition model.

[0127] In one embodiment, the TCR characteristics include one or more of the following: charge, hydrophobicity, and two-dimensional structural characteristics of CDR3.

[0128] In one embodiment, the extracted TCR features are obtained through group sparse regularized regression analysis.

[0129] In one embodiment, the probability calculation of TCR recognition also includes TCR-epitope binding prediction, and the probability of TCR recognition is obtained through the recognition model and TCR-epitope binding prediction; the TCR-epitope binding prediction is obtained by predicting through a trained TCR-epitope binding prediction model.

[0130] In one embodiment, the process of constructing the trained TCR-epitope binding prediction model is as follows:

[0131] Obtain the sequence, gene mutation data, and antigenic peptides of the CDR3 region;

[0132] Converting the sequence of the CDR3 region, gene mutation data, and antigenic peptides into a feature matrix;

[0133] The feature matrix is ​​input into a neural network for training to obtain a TCR-antigen epitope binding prediction model.

[0134] In one embodiment, the feature matrix of the antigenic peptides includes sequence features and structural features of the antigenic peptides.

[0135] In one embodiment, the TCR-antigen epitope binding prediction model further includes MHC-antigen peptide affinity calculation, and the affinity prediction results are obtained based on the sequence characteristics and structural characteristics of the antigen peptide.

[0136] In one embodiment, the probability calculation of TCR recognition also includes the prediction of the immunogenicity of the tumor antigen. The probability of TCR recognition is obtained by the recognition model and the immunogenicity prediction of the tumor antigen. The immunogenicity prediction of the tumor antigen is obtained by predicting the immunogenicity prediction model of the trained tumor antigen.

[0137] In one embodiment, the process of constructing the tumor antigen immunogenicity prediction model is as follows:

[0138] Obtain RNA-seq data of different tumor-infiltrating lymphocytes, tumor cell genomes, and immunogenicity data;

[0139] Performing TCR library analysis on the RNA-seq data to obtain tumor gene expression profile data and obtaining CDR3 sequences from the RNA-seq data;

[0140] Acquiring gene mutation and antigen peptide data based on the tumor cell genome data;

[0141] Feature extraction is performed on CDR3 sequence and antigen peptide data to obtain feature matrix data;

[0142] The tumor gene expression profile data, gene mutation, and feature matrix data are input into a neural network for training to obtain an immunogenicity prediction model.

[0143] In one embodiment, the neoantigen load is calculated by calculating the affinity of the neoantigen's antigenic peptide to the MHC molecule and the probability of TCR recognition.

[0144] In one embodiment, the formula for calculating the neoantigen load is:

[0145]

[0146] Where N is the neoantigen load, n represents the number of tumor clones analyzed from the patient's exome sequencing data, b i Expresses affinity, r i represents the probability of TCR recognition, c j represents the expression level of gene j.

[0147] In a specific embodiment, the present invention conducts big data and deep learning-driven tumor neoantigen screening. Based on the integration of multi-omics data such as genomic, transcriptomic, and proteomic data, in-depth exploration is conducted on the generation of tumor neoantigens by different gene mutation types, the affinity between MHC molecules and antigenic peptides, TCR recognition of antigenic peptides, and clinical response biomarkers of PD-L1 blockers. High-dimensional heterogeneous features are extracted from the biological processes of tumor antigen generation, presentation, and immune response, and a variety of intelligent algorithms, including deep learning, are developed and used to screen for high-quality immunogenic neoantigens.

[0148] In a specific embodiment, the affinity between MHC-I molecules and antigen peptides is predicted by integrating large-scale mass spectrometry data of MHC-I class binding antigen peptides, constructing a mass spectrometry-based MHC-antigen peptide benchmark dataset and developing an online database; analyzing genome or whole exome sequencing data for MHC typing, extracting sequence and structural features of antigen peptides and MHC binding grooves, proposing and developing an MHC-antigen peptide affinity prediction model based on deep learning; constructing an MHC-antigen peptide heterogeneous network that integrates multi-source data, merging network topology features and ontology features, and training a deep feedforward network to predict MHC-antigen peptide affinity.

[0149] The prediction performance of MHC-I and antigen peptide affinity is improved from three aspects: 1) Integrate large-scale mass spectrometry data of MHC-I class binding antigen peptides to build a mass spectrometry-based MHC-antigen peptide benchmark dataset, first use IEDB data to pre-train a deep neural network-based model, and then combine the mass spectrometry data to perform transfer learning on the pre-trained model to obtain a higher performance prediction model, such as Figure 4As shown; 2) Combining sequence and structural features to predict the affinity of MHC-I and antigen peptides: Extract the sequences of antigen peptides and MHC-I binding grooves from the protein database, use one-hot encoding of amino acid sequences, and use a variety of methods to extract sequence features including physical chemistry, local structural entropy, pairing potential, interaction tendency, etc., and simultaneously extract structural features such as secondary structure features, structural neighbor features, and solvent accessible surface area. Combine these sequence- and structure-based features to construct a training set, learn a gradient boosting regression tree model, determine the important factors of the affinity between MHC-I molecules and antigen peptides, and develop an online prediction service. 3) Calculate the MHC-I and MHC-II subtype similarity network and the antigen peptide similarity network, integrate the MHC I molecule-antigen peptide association in the IEDB, and thus construct an MHC-antigen peptide heterogeneous network that integrates multi-source heterogeneous data. Use a random walk algorithm on a two-way heterogeneous network to predict the probability of potential MHC molecules binding to antigen peptides.

[0150] In a specific embodiment, TCR recognition of tumor-specific antigens and immunogenicity prediction: extract the sequence and structural characteristics of antigen peptides and TCR variable regions from genomic data, combine with experimentally verified TCR-antigen peptide binding data sets, perform regression analysis and correlation analysis, and discover the key factors for TCR recognition of tumor antigens; construct a deep learning model, using antigen and TCR-related characteristics as input to predict the probability of TCR recognizing and binding to tumor antigens; integrate gene mutations of different tumors, construct a gene mutation-TCR correspondence matrix, and analyze the probability of tumor-specific antigens being recognized by TCR from the perspective of tumor clonal evolution; combine tumor infiltrating lymphocyte (TIL) data with experimentally obtained immunogenicity data, use tumor gene expression profiles and antigen peptide feature matrices as inputs to the deep neural network, and predict the immunogenicity of tumor antigens.

[0151] CD8+ T lymphocytes rely on the complementarity-determining region (CDR3) of the TCR to recognize and kill tumor cells. The binding of CDR3 to antigenic peptides presented by MHC is crucial for the adaptive immune response. TCRs are determined by the V(D)J recombination process, and the human TCR repertoire can accommodate up to 1015 different molecular types. Unless derived from the same clone, each T lymphocyte has a unique TCR, yet different TCRs can bind to the same antigen. This poses significant challenges in the computational prediction of immunogenic tumor antigens. We believe that the TCR region of T cells that recognize the same pMHC complex contains conserved sequence features (motifs). We plan to develop three methods to predict the probability of TCR recognizing tumor antigens and activating an immune response: 1) Identification of the main factors of TCR binding to antigen epitopes: Collect and integrate experimentally verified antigen epitope-TCR interaction data, extract the sequence and structure of TCR and peptide chains, including the charge, hydrophobicity, two-dimensional structural characteristics of CDR3 and one-hot encoding of antigens, use group sparse regularized regression to analyze the impact of each feature on TCR recognition of antigens, construct a deep neural network model and use antigen-TCR interaction data for training, and predict the probability of a specific TCR subtype recognizing a specific antigen. 2) CNN-based TCR-antigen epitope binding prediction: The sequence of the CDR3 region is obtained from RNA-seq data, and autocross covariation is used to convert CDR3 sequences of different lengths into a feature matrix of the same dimension; gene mutations and antigen peptides are obtained from genome or exome sequencing data and converted into a feature matrix using one-hot encoding; MHC-antigen peptide affinity prediction and TCR-antigen epitope binding prediction share the characteristics of antigen peptide sequence and structure, and a multi-task deep learning model is constructed to simultaneously perform two tasks: MHC-antigen peptide affinity and TCR-antigen epitope binding prediction; 3) TCR library profiling is performed on RNA-seq data of different tumor-infiltrating lymphocytes (TILs), and gene mutations are obtained by combining tumor cell genome or exome sequencing data. Experimentally obtained immunogenicity datasets are collected, and the tumor gene expression profile, antigen peptide feature matrix, and CDR3 feature matrix are used as inputs to the deep neural network to predict the immunogenicity of tumor antigens.

[0152] In one specific embodiment, the neoantigen load indicator is defined and calculated: Tumor antigen activation of T cell responses is a complex multi-step process, in which antigen presentation and TCR recognition of pMHC complexes are the main steps affecting T cell recognition and killing of target cells. Failure in any of these steps will lead to immunotherapy failure. For antigen i of tumor clone j, assuming the affinity of the MHC-I molecule for the antigen peptide is bi and the probability of TCR recognition of the antigen is ri, the neoantigen load of the patient is defined as:

[0153]

[0154] Where n represents the number of tumor clones analyzed from the patient's exome sequencing data.

[0155] It should be noted that the present invention uses multi-omics data, including genome or exome, transcriptome and mass spectrometry data, when developing bioinformatics models. Once the prediction model training is completed, the patient's exome sequencing or large panel sequencing data is required during use, which is conducive to reducing the cost of clinical companion diagnosis and promoting market application and promotion.

[0156] In a specific embodiment, an esophageal cancer specimen library is established, which is collected according to the unified standards of the present invention and contains complete clinical treatment and follow-up data, experimental research data, and tumor tissue and body fluid samples at various treatment and follow-up stages. Whole-exome sequencing is performed on tumor and serum samples to obtain raw exome sequencing data for mutation analysis, obtaining non-synonymous mutations (NSVs) and insertion / deletion mutations (InDels). HLA typing analysis can also be performed; fusion genes and genomic expression profiles are analyzed from transcriptome sequencing; tumor antigen sequences are extracted based on mutations, and antigen peptides, HLA typing, and antigen host gene expression levels are integrated as input to the new antigen screening model to screen high-quality new antigens, and calculate the new antigen load index for each patient, such as Figure 5 At the same time, prognostic indicators such as PD-L1 expression level, TMB and MSI were collected as the basis for control analysis.

[0157] In a specific embodiment, the approval and marketing of immune checkpoint inhibitors such as PD-L1 and CTLA-4 monoclonal antibodies has greatly promoted the research of tumor immunotherapy. However, the population that currently benefits from immunotherapy is very limited. For example, the objective response rate of PD-L1 inhibitors alone is only about 30%. The companion diagnostic indicators of immune checkpoint inhibitors, TMB and PD-L1 protein levels, cannot accurately screen patients who can benefit from PD-L1 monoclonal antibody treatment because cytotoxic T lymphocytes cannot recognize tumor-specific antigen epitopes. Therefore, combining multi-omics data to screen individual-specific biomarkers as companion diagnostics for immune checkpoint inhibitors is the direction of development in precision medicine. The present invention intends to collect and integrate clinical data such as the response, progression-free survival, and overall survival of various tumors to the immune checkpoint inhibitor PD-L1, to construct a benchmark PD-L1 inhibitor clinical dataset. On this basis, two methods are developed for biomarker screening: 1) Neoantigens generated by tumor driver mutations and passenger mutations are analyzed for their correlation with response to PD-L1 inhibitors and prognosis, and statistically significant driver gene mutations are discovered; 2) Gene mutations and MHC subtypes are obtained from exome sequencing data of tumor cells and paired normal cells, and gene expression and TCR library spectra are obtained from RNA-seq data. Relevant features such as antigen peptide sequences, tumor gene expression profiles, MHC typing, and TCR library are integrated, combined with large-scale clinical immune response data, to extract features from high-dimensional heterogeneous data and train deep neural network models, to achieve high-precision prediction of tumor-specific immune checkpoint responses and prognosis, and provide more reference indicators for companion diagnosis of immune checkpoint clinical treatment.

[0158] In a specific embodiment, in the clinical treatment of esophageal cancer patients with anti-PD-L1, the enrolled patients were treated with the marketed PD-L1 antibody Imfinzi (durvalumab). The treatment process was strictly carried out in accordance with the NCCN clinical diagnosis and treatment plan, and the patient's biochemical, immune, imaging and other clinical data were recorded. At the same time, the relevant side effects were mainly evaluated, including hematological toxicity, gastrointestinal reactions, heart damage, lung damage, skin damage and other side effects. According to the calculated neoantigen load, the patients with esophageal cancer were stratified and divided into groups. The prognostic indicators such as OS, objective response rate (ORR), overall response rate (OR), progression-free survival (PFS) of the patients were statistically analyzed to determine the statistical significance of the neoantigen load in various prognostic indicators, such as Figure 5 In particular, the prognostic role of neoantigen load was verified by comparison with indicators such as PD-L1+, TMB, and MSI.

[0159] The disclosed embodiments of the present invention also provide a computer program product or system, including a computer program, which implements the above-mentioned steps of the screening method for new antigens when executed by a processor.

[0160] Figure 2 A schematic diagram of a screening system for neoantigens provided in an embodiment of the present invention specifically includes:

[0161] Acquisition unit: acquires the patient's tumor slice data and serum data;

[0162] Processing unit: Processing the tumor slice data and serum data to obtain HLA / MHC pseudo sequence, antigen peptide data, and antigen gene expression level; the data processing includes data sequencing, HLA / MHC typing, somatic mutation calculation, fusion gene calculation, and gene expression profile calculation;

[0163] Screening unit: Screening new antigens based on the HLA / MHC pseudo sequence, antigen peptide data, and antigen gene expression level data to obtain new antigens.

[0164] Figure 3 A schematic diagram of a screening device for neoantigens provided in an embodiment of the present invention specifically includes:

[0165] A memory and a processor; the memory is used to store program instructions; the processor is used to call program instructions, and when the program instructions are executed, any one of the above-mentioned new antigen screening methods is performed.

[0166] The disclosed embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs any of the above-mentioned methods for screening new antigens.

[0167] The validation results of this validation example demonstrate that assigning inherent weights to indications can improve the performance of the present method compared to the default settings. Those skilled in the art will readily appreciate that, for ease of description and brevity, the specific operating processes of the systems, devices, and units described above can be referenced to the corresponding processes in the aforementioned method embodiments and will not be further elaborated upon here. It should be understood that the disclosed systems, devices, and methods can be implemented in other ways within the several embodiments provided herein. For example, the device embodiments described above are merely illustrative. For example, the division of units described is merely a logical functional division. In actual implementation, other divisions may be employed, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, the coupling, direct coupling, or communication connection shown or discussed may be through interfaces, indirect coupling, or communication connection between devices or units, and may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the units may be selected to achieve the objectives of the present embodiment as needed. In addition, the functional units in the various embodiments of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units. Those skilled in the art will understand that all or part of the steps in the various methods of the above-mentioned embodiments may be completed by instructing the relevant hardware through a program, and the program may be stored in a computer-readable storage medium, which may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0168] Those skilled in the art will understand that all or part of the steps in the above-mentioned embodiment method can be implemented by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned medium storage can be a read-only memory, a disk or an optical disk, etc.

[0169] The above is a detailed introduction to a computer device provided by the present invention. For those skilled in the art, according to the concept of the embodiments of the present invention, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A method for screening new antigens, characterized in that: include: S1. Obtain the patient's tumor slice data and serum data; S2. Processing the tumor slice data and serum data to obtain HLA / MHC pseudo sequences, antigen peptide data, and antigen gene expression levels; the data processing includes data sequencing, HLA / MHC typing, somatic mutation calculation, fusion gene calculation, and gene expression profile calculation; S3. Screening new antigens based on the HLA / MHC pseudo sequence, antigen peptide data, and antigen gene expression level data to obtain new antigens; The screening is performed by inputting HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression level data into a deep network model to obtain the affinity of the HLA / MHC pseudo-sequences and antigen peptide data, combining the HLA / MHC pseudo-sequences with the antigen peptides to obtain HLA / MHC-antigen peptides, and then screening to obtain new antigens based on the probability of TCR recognizing the HLA / MHC-antigen peptides; The method further includes calculating the neoantigen load by calculating the affinity of the HLA / MHC-antigen peptide of the neoantigen and the probability of TCR recognizing the HLA / MHC-antigen peptide; The formula for calculating the neoantigen load is: Where N is the neoantigen load, n represents the number of tumor clones analyzed from the patient's exome sequencing data, b i Expresses affinity, r i represents the probability of TCR recognition, c j Indicates the expression level of gene j, clone j represents the tumor clone line of gene j.

2. The method for screening new antigens according to claim 1, characterized in that: The data processing process is as follows: S21, performing exome sequencing based on the tumor section data to obtain exome sequencing data; S22, performing transcriptome sequencing based on the serum data to obtain transcriptome sequencing data; S23. Performing HLA / MHC typing based on the exome sequencing data to obtain HLA / MHC pseudo sequences and performing somatic mutation calculation to obtain mutation data; S24, calculating gene expression profile data and fusion gene data based on the transcriptome sequencing data; S25. Calculating antigen gene expression level data based on the gene expression profile data; S26. The fusion gene and the mutation data are fused to obtain antigen peptide data.

3. The method for screening new antigens according to claim 1, characterized in that: The training process of the deep network model is: The first step is to obtain HLA / MHC pseudo sequences, antigen peptide data, antigen gene expression level data, and IEDB data; The second step is to pre-train the initial deep network model using the IEDB data to obtain the model weights; Step 3: Migrating the weights of the model to the deep network model to obtain a migrated deep network model; Step 4: Vectorize the HLA / MHC pseudo sequence and antigen peptide data to obtain HLA / MHC pseudo sequence vector data and antigen peptide vector data; Step 5: Extract features from the antigen gene expression level data to obtain the antigen depth feature vector; Step 6: Fuse the HLA / MHC pseudo-sequence vector data, antigen peptide vector data, and antigen deep feature vector to obtain fused data, and input it into the migration deep network model for training to obtain a trained deep network model.

4. The method for screening new antigens according to claim 3, characterized in that: The training process of the deep network model also includes sequence feature data, obtaining sequence data of antigen peptide sequences and HLA / MHC binding grooves, performing feature extraction on the sequence data of the antigen peptide sequences and HLA / MHC binding grooves to obtain sequence features, and inputting the HLA / MHC pseudo sequence, antigen peptide data, antigen gene expression level data, and sequence features into the deep network model to obtain the affinity of the HLA / MHC pseudo sequence and antigen peptide data; The sequence features are extracted through the feature model, and the feature model construction process is as follows: Obtain the antigen peptide sequence and the sequence of the HLA / MHC binding groove from the protein database; performing feature extraction on the sequence to obtain sequence features; performing structural feature extraction on the sequence to obtain structural features; Constructing a feature training set based on the sequence features and structural features; The feature training set is input into the second deep network model for training to obtain a feature model.

5. The method for screening new antigens according to claim 4, characterized in that: The structural features include one or more of the following: secondary structure features, structural neighbor features, and solvent accessible surface area.

6. The method for screening new antigens according to claim 4, characterized in that: The feature extraction method adopts one or more of the following: physical and chemical extraction, local structure entropy extraction, pairing potential energy extraction, and interaction tendency extraction.

7. The method for screening new antigens according to claim 4, characterized in that: The second deep network model includes one or more of the following: GBRT, XGBoost, LightGBM, CatBoost, Random Forest, and AdaBoost.

8. The method for screening new antigens according to claim 3, characterized in that: The training process of the deep network model also includes similar feature data, obtaining the structures of HLA / MHC molecules and antigen peptides to generate HLA / MHC network structures and antigen peptide network structures, extracting features from the HLA / MHC network structures and antigen peptide network structures to obtain similar features, and inputting HLA / MHC pseudo sequences, antigen peptide data, antigen gene expression level data, and similar features into the deep network model to obtain the affinity of the HLA / MHC pseudo sequences and antigen peptide data.

9. The method for screening new antigens according to claim 8, characterized in that: The similar features are extracted through a similar feature model. The similar feature model is constructed by: obtaining the structures of HLA / MHC molecules and antigen peptides to generate HLA / MHC network structures and antigen peptide network structures; the HLA / MHC includes HLA / MHC-I subtype and HLA / MHC-II subtype; Calculate the similarity network between HLA / MHC-I subtypes and HLA / MHC-II subtypes, and the antigen peptide similarity network; Obtain the association between HLA / MHC molecules and antigen peptides; Based on the similarity network of HLA / MHC-I subtype and HLA / MHC-II subtype, and the antigen peptide similarity network, the association relationship between HLA / MHC molecules and antigen peptides is input into the similarity feature model in the third deep network model.

10. The method for screening new antigens according to claim 9, characterized in that: The third deep network model adopts one or more of the following: heterogeneous information network, graph neural network, and hypergraph.

11. The method for screening new antigens according to claim 3, characterized in that: The training process of the deep network model also includes sequence feature data and similarity features, obtaining sequence data of antigen peptide sequences and HLA / MHC binding grooves, and performing feature extraction on the sequence data of the antigen peptide sequences and HLA / MHC binding grooves to obtain sequence features; Acquiring the structures of HLA / MHC molecules and antigen peptides to generate an HLA / MHC network structure and an antigen peptide network structure, and extracting features from the HLA / MHC network structure and the antigen peptide network structure to obtain similar features; The HLA / MHC pseudo sequence, antigen peptide data, antigen gene expression level data, sequence features, and similarity features are input into the deep network model to obtain the affinity of the HLA / MHC pseudo sequence and antigen peptide data.

12. The method for screening new antigens according to claim 1, characterized in that: The probability of TCR recognition is calculated by a trained recognition model, and the construction process of the recognition model is as follows: Obtain epitope-TCR interaction data, TCR and peptide chain sequences and structures; Extracting TCR features based on the sequence and structural data of the TCR and peptide chain; The TCR characteristics and antigen epitope-TCR interaction data are input into the fourth deep network model for training to obtain a recognition model.

13. The method for screening new antigens according to claim 12, characterized in that: The TCR characteristics include one or more of the following: charge, hydrophobicity, and two-dimensional structural characteristics of CDR3.

14. The method for screening new antigens according to claim 12, characterized in that: The extracted TCR features are obtained through group sparse regularized regression analysis.

15. The method for screening new antigens according to claim 12, characterized in that: The probability calculation of TCR recognition also includes TCR-epitope binding prediction, and the probability of TCR recognition is obtained through the recognition model and TCR-epitope binding prediction; the TCR-epitope binding prediction is obtained by predicting through a trained TCR-epitope binding prediction model.

16. The method for screening new antigens according to claim 15, characterized in that: The process of constructing the trained TCR-antigen epitope binding prediction model is as follows: Obtain the sequence, gene mutation data, and antigenic peptides of the CDR3 region; Converting the sequence of the CDR3 region, gene mutation data, and antigenic peptides into a feature matrix; The feature matrix is ​​input into a neural network for training to obtain a TCR-antigen epitope binding prediction model.

17. The method for screening new antigens according to claim 16, characterized in that: The characteristic matrix of the antigen peptides includes the sequence characteristics and structural characteristics of the antigen peptides.

18. The method for screening new antigens according to claim 15, characterized in that: The TCR-antigen epitope binding prediction model also includes MHC-antigen peptide affinity calculation, and the affinity prediction results are obtained based on the sequence characteristics and structural characteristics of the antigen peptide.

19. The method for screening new antigens according to claim 12, characterized in that: The probability calculation of TCR recognition also includes the prediction of the immunogenicity of tumor antigens. The probability of TCR recognition is obtained through the recognition model and the immunogenicity prediction of tumor antigens. The immunogenicity prediction of tumor antigens is obtained by predicting the immunogenicity prediction model of trained tumor antigens.

20. The method for screening new antigens according to claim 19, characterized in that: The construction process of the tumor antigen immunogenicity prediction model is as follows: Obtain RNA-seq data of different tumor-infiltrating lymphocytes, tumor cell genomes, and immunogenicity data; Performing TCR library analysis on the RNA-seq data to obtain tumor gene expression profile data and obtaining CDR3 sequences from the RNA-seq data; Acquiring gene mutation and antigen peptide data based on the tumor cell genome data; Feature extraction is performed on CDR3 sequence and antigen peptide data to obtain feature matrix data; The tumor gene expression profile data, gene mutation, and feature matrix data are input into a neural network for training to obtain an immunogenicity prediction model.

21. A computer program product comprising a computer program or instructions, characterized in that: The computer program or instructions are executed by a processor to implement the method for screening new antigens according to any one of claims 1 to 20.

22. A computer device comprising a memory, a processor, and a computer program or instruction stored in the memory, wherein: The computer program or instructions are executed by a processor to implement the method for screening new antigens according to any one of claims 1 to 20.

23. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: The computer program or instructions are executed by a processor to implement the method for screening new antigens according to any one of claims 1 to 20.

Citation Information

Patent Citations

  • Deep learning model for predicting tumor-specific neoantigen MHC class I or class II immunogenicity

    CN117136410A