New antigen screening method, equipment and program product

By processing tumor sections and serum data of esophageal cancer patients and screening deep learning models, the problem of inaccurate affinity prediction in neoantigen screening is solved, more accurate neoantigen screening and immunogenicity evaluation are achieved, and the efficacy prediction of immune checkpoint inhibitor treatment is improved.

CN120048333AActive Publication Date: 2025-05-273201 HOSPITAL +1

Patent Information

Application Number
CN202510166355.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-27
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

In the prior art, when screening esophageal carcinoma neoantigens, the affinity prediction is inaccurate, and it is difficult to evaluate immunogenicity, and it is difficult to accurately predict which mutations will lead to truly immunogenic neoantigens.

Method used

By obtaining the patient's tumor section data and serum data, data processing was performed to obtain the HLA/MHC pseudo-sequence, antigen peptide data and antigen gene expression levels. Neoantigens were screened using a deep network model, and affinity prediction and TCR recognition probability calculation were performed in combination with HLA/MHC pseudo-sequence, antigen peptide data and antigen gene expression level data.

Benefits of technology

It improves the accuracy and reliability of neoantigens screening, can predict the immunogenicity of neoantigens more accurately, and helps judge the efficacy of immune checkpoint inhibitor treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048333A_ABST
    Figure CN120048333A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent medical treatment, in particular to a new antigen screening method, equipment and a program product. Comprising the following steps: S1, acquiring tumor slice data and serum data of a patient; s2, performing data processing on the tumor slice data and the serum data to obtain an HLA / MHC pseudo sequence, antigen peptide data and an antigen gene expression level; the data processing comprises data sequencing, HLA / MHC typing, somatic mutation calculation, fusion gene calculation and gene expression profile calculation; s3, screening new antigens based on the HLA / MHC pseudo sequence, the antigen peptide data and the antigen gene expression level data to obtain the new antigens, the application can perform new antigen screening, and has good clinical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent medicine, and specifically relates to a screening method, device, program product and computer-readable storage medium for neoantigens. Background Art

[0002] Among esophageal cancer cases, esophageal squamous cell carcinoma (ESCC) accounts for more than 90%. For patients with inoperable and unsuitable esophageal cancer, the curative effects of traditional chemotherapy, radiotherapy and targeted therapy are not optimistic. In recent years, tumor immunotherapy, especially immune checkpoint inhibition, has made substantial progress. In 2017, the FDA approved Pembrolizumab for the second-line clinical treatment of recurrent locally advanced or metastatic esophageal squamous cell carcinoma with PD-L1 positivity. The research on immunotherapy for esophageal cancer and its companion diagnostic markers has set off a wave of enthusiasm at home and abroad. However, there is a lack of reliable biomarkers for Anti-PD-L1 treatment of esophageal cancer. In the clinical application of immune checkpoint inhibitors CTLA4 and PD-L1 monoclonal antibodies, it is very important to identify patients by detecting biomarkers. Currently, commonly used companion diagnostic markers include PD-L1 expression level, tumor mutation burden (TMB), microsatellite instability (MSI-H), and DNA mismatch repair defect (dMMR), etc. Although these prognostic markers play a guiding role in clinical applications, more and more clinical cohort studies have shown that these markers cannot accurately distinguish which patients can benefit from immune checkpoint inhibitor treatment. Therefore, the concept of Tumor Neoantigen Burden (TNB) has been proposed. Tumor neoantigens are the actual number of mutations targeted by T cells, which can better judge the clinical efficacy of immune checkpoint inhibitors. However, the screening of neoantigens combined with affinity prediction is inaccurate, and immunogenicity evaluation is difficult (not all peptide segments that can bind to MHC can effectively activate T cells, and existing methods are difficult to accurately predict which mutations will lead to truly immunogenic neoantigens). Summary of the Invention

[0003] In view of the above problems, the present invention provides a screening method for neoantigens, which specifically includes: S1. Obtain tumor slice data and serum data of a patient; S2. Process the tumor slice data and serum data to obtain HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression levels; the data processing includes data sequencing, HLA / MHC typing, somatic mutation calculation, fusion gene calculation, and gene expression profile calculation; S3. Screen neoantigens based on the HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression level data to obtain neoantigens.

[0004] The process of the data processing is as follows: S21. Perform whole-exome sequencing on the tumor section data to obtain whole-exome sequencing data; S22. Perform transcriptome sequencing on the serum data to obtain transcriptome sequencing data; S23. Perform HLA / MHC typing on the whole-exome sequencing data to obtain HLA / MHC pseudo-sequences and calculate somatic mutations to obtain mutation data; S24. Calculate gene expression profile data and fusion gene data based on the transcriptome sequencing data; S25. Calculate antigen gene expression level data based on the gene expression profile data; S26. Fuse the fusion gene and the mutation data to obtain antigen peptide data.

[0005] The screening is performed by inputting the HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression level data into a deep network model to obtain the affinity of the HLA / MHC pseudo-sequences and antigen peptide data. The HLA / MHC pseudo-sequences bind to the antigen peptides to obtain HLA / MHC-antigen peptides, and then the probability of TCR recognizing the HLA / MHC-antigen peptides is used for screening to obtain neoantigens; Optionally, the training process of the trained deep network model is as follows: The first step: Obtain HLA / MHC pseudo-sequences, antigen peptide data, antigen gene expression level data, and IEDB data; The second step: Perform initial pre-training on the deep network model through the IEDB data to obtain the weights of the model; The third step: Transfer the weights of the model to the deep network model to obtain a transferred deep network model; The fourth step: Perform vectorization processing on the HLA / MHC pseudo-sequences and antigen peptide data to obtain HLA / MHC pseudo-sequence vector data and antigen peptide vector data; The fifth step: Extract features from the antigen gene expression level data to obtain antigen deep feature vectors; The sixth step: Fuse the HLA / MHC pseudo-sequence vector data, antigen peptide vector data, and antigen deep feature vectors to obtain fusion data, and input it into the transferred deep network model for training to obtain a trained deep network model; Optionally, the training process of the affinity prediction model further includes HLA / MHC antigen peptide feature data, obtaining sequence feature data and structural feature data of the HLA / MHC antigen peptide, and inputting the sequence feature data, structural feature data of the HLA / MHC antigen peptide and the HLA / MHC antigen peptide mass spectrometry dataset into the training model for training to obtain an affinity prediction model.

[0006] The training process of the deep network model further includes sequence feature data, obtaining the antigen peptide sequence and the sequence data of the HLA / MHC binding groove, performing feature extraction on the antigen peptide sequence and the sequence data of the HLA / MHC binding groove to obtain sequence features, and inputting the HLA / MHC pseudo-sequence, antigen peptide data, antigen gene expression level data, and sequence features into the deep network model to obtain the affinity of the HLA / MHC pseudo-sequence and antigen peptide data; The sequence features are extracted by a feature model, and the construction process of the feature model is as follows: Obtain the antigen peptide sequence and the sequence of the HLA / MHC binding groove from a protein database; Perform feature extraction on the sequence to obtain sequence features; Perform structural feature extraction on the sequence to obtain structural features; Construct a feature training set based on the sequence features and structural features; Input the feature training set into a second deep network model for training to obtain a feature model; Optionally, the structural features include one or more of the following: structural features, structural neighbor features, solvent accessible surface area; Optionally, the method of feature extraction adopts one or more of the following: physicochemical extraction, local structure entropy extraction, pairwise potential extraction, interaction tendency extraction; Optionally, the second deep network model includes one or more of the following: GBRT, XGBoost, LightGBM, CatBoost, Random Forest, AdaBoost.

[0007] The training process of the deep network model further includes similar feature data, obtaining the structures of the HLA / MHC molecule and the antigen peptide to generate the HLA / MHC network structure and the antigen peptide network structure, performing feature extraction on the HLA / MHC network structure and the antigen peptide network structure to obtain similar features, and inputting the HLA / MHC pseudo-sequence, antigen peptide data, antigen gene expression level data, and similar features into the deep network model to obtain the affinity of the HLA / MHC pseudo-sequence and antigen peptide data; Optionally, the similar features are extracted by a similar feature model. The construction process of the similar feature model is as follows: Obtain the structures of HLA / MHC molecules and antigen peptides to generate an HLA / MHC network structure and an antigen peptide network structure; the HLA / MHC includes HLA / MHC-I subtypes and HLA / MHC-II subtypes; Calculate the similarity networks of HLA / MHC-I subtypes and HLA / MHC-II subtypes, and the antigen peptide similarity network; Obtain the association relationship between HLA / MHC molecules and antigen peptides; Based on the similarity networks of HLA / MHC-I subtypes and HLA / MHC-II subtypes, the antigen peptide similarity network, and the association relationship between HLA / MHC molecules and antigen peptides, input them into the similar feature model in the third deep network model; Optionally, the third deep network model adopts one or more of the following: two-pass heterogeneous network, heterogeneous information network, graph neural network, hypergraph; Optionally, the training process of the deep network model further includes sequence feature data and similar features. Obtain the sequence data of antigen peptide sequences and HLA / MHC binding grooves, and perform feature extraction on the sequence data of the antigen peptide sequences and HLA / MHC binding grooves to obtain sequence features; obtain the structures of HLA / MHC molecules and antigen peptides to generate an HLA / MHC network structure and an antigen peptide network structure, and perform feature extraction on the HLA / MHC network structure and the antigen peptide network structure to obtain similar features; input HLA / MHC pseudo-sequences, antigen peptide data, antigen gene expression level data, sequence features, and similar features into the deep network model to obtain the affinity of HLA / MHC pseudo-sequences and antigen peptide data.

[0008] The TCR recognition probability is calculated by a trained recognition model. The construction process of the recognition model is as follows: Obtain antigen epitope-TCR interaction data, and the sequences and structures of TCRs and peptide chains; Extract TCR features based on the sequence and structure data of TCRs and peptide chains; Input the TCR features and antigen epitope-TCR interaction data into the fourth deep network model for training to obtain the recognition model; Optionally, the TCR features include one or more of the following: the charge, hydrophobicity, and two-dimensional structure features of CDR3; Optionally, the extraction of TCR features is obtained by group sparse regularization regression analysis; Optionally, the calculation of the probability of TCR recognition further includes TCR-epitope binding prediction, and the probability of TCR recognition is obtained through the recognition model and TCR-epitope binding prediction; the TCR-epitope binding prediction is obtained through a trained TCR-epitope binding prediction model; Optionally, the construction process of the trained TCR-epitope binding prediction model is as follows: Obtain the sequence of the CDR3 region, gene mutation data, and antigen peptides; Convert the sequence of the CDR3 region, gene mutation data, and antigen peptides into a feature matrix; Input the feature matrix into a neural network for training to obtain a TCR-epitope binding prediction model; Optionally, the feature matrix of the antigen peptide includes the sequence feature and structural feature of the antigen peptide; Optionally, the TCR-epitope binding prediction model further includes MHC-antigen peptide affinity calculation, and the affinity prediction result is obtained through the sequence feature and structural feature of the antigen peptide; Optionally, the calculation of the probability of TCR recognition further includes immunogenicity prediction of tumor antigens, and the probability of TCR recognition is obtained through the recognition model and immunogenicity prediction of tumor antigens, and the immunogenicity prediction of tumor antigens is obtained through a trained immunogenicity prediction model of tumor antigens; Optionally, the construction process of the immunogenicity prediction model of tumor antigens is as follows: Obtain RNA-seq data of different tumor-infiltrating lymphocytes, tumor cell genomes, and immunogenicity data; Perform TCR repertoire analysis on the RNA-seq data to obtain tumor gene expression profile data and obtain the CDR3 sequence from the RNA-seq data; Obtain gene mutation and antigen peptide data based on the tumor cell genome data; Extract features from the CDR3 sequence and antigen peptide data to obtain feature matrix data; Input the tumor gene expression profile data, gene mutations, and feature matrix data into a neural network for training to obtain an immunogenicity prediction model.

[0009] The method further includes neoantigen load calculation, which is obtained by calculating the affinity of the neoantigen HLA / MHC-antigen peptide and the probability of TCR recognizing HLA / MHC-antigen peptide; Optionally, the formula for the neoantigen load calculation is:

[0010] Wherein, N is the neoantigen load, and n represents the number of tumor clone lines analyzed from the patient's exome sequencing data. b i represents the affinity, r i represents the probability of TCR recognition, c j represents the expression level of gene j.

[0011] An object of the present invention is to provide a computer program product, which includes a computer program or instruction, and the computer program or instruction is executed by a processor to implement the above-mentioned method for screening neoantigens.

[0012] An object of the present invention is to provide a computer device, which includes a memory, a processor, and a computer program or instruction stored on the memory, and the computer program or instruction is executed by the processor to implement the above-mentioned method for screening neoantigens.

[0013] An object of the present invention is to provide a computer-readable storage medium, which stores a computer program or instruction, and the computer program or instruction is executed by a processor to implement the above-mentioned method for screening neoantigens.

[0014] Advantages of the present invention: 1. The screening of neoantigens is proposed. Neoantigens are obtained through two-step screening. The first step of screening is the affinity screening of HLA / MHC-antigen peptides, and the second step of screening is the screening of the probability of TCR recognition of HLA / MHC-antigen peptides. Among them, the affinity screening of MHC-antigen peptides is carried out by constructing a mass spectrometry-based MHC-antigen peptide benchmark dataset, analyzing genomic or whole exome sequencing data for MHC typing, extracting the sequence and structural features of antigen peptides and MHC binding grooves, proposing and developing a deep learning-based MHC-antigen peptide affinity prediction model, and constructing an MHC-antigen peptide heterogeneous network that integrates multi-source data, combining network topology features and ontology features, and training a deep feedforward network to predict the affinity of MHC-antigen peptides. Affinity calculations are performed from multiple angles and multiple features to improve the accuracy and reliability of affinity calculations.

[0015] 2. Research on calculating and predicting the binding of TCR to tumor antigens has just started, thanks to the accumulation of experimentally verified TCR-antigen data in recent years. A few researchers have tried to develop machine learning models to predict the affinity between TCR and antigens. However, the currently available training sets are far from sufficient compared to the vast human TCR library. Therefore, this invention extracts the sequence and structural features of antigen peptides and TCR variable regions from genomic data, combines the experimentally verified TCR-antigen peptide binding data set, conducts regression analysis and correlation analysis to discover the key factors for TCR to recognize tumor antigens; constructs a deep learning model, uses the relevant features of antigens and TCR as inputs to predict the probability of TCR recognizing and binding to tumor antigens; integrates the gene mutations of different tumors, constructs a gene mutation-TCR correspondence matrix, and analyzes the probability of tumor-specific antigens being recognized by TCR from the perspective of tumor clonal evolution. Similarly, the probability of TCR recognition is calculated from multiple perspectives to improve the credibility of the TCR recognition probability, and further improve the calculation accuracy of neoantigen load, which helps to improve the accuracy of the prediction results after tumor treatment.

[0016] 3. A method for calculating neoantigen load is proposed, which is calculated through the affinity of MHC-antigen peptides and the probability of TCR recognition, quantifies the neoantigen load, and conducts prognostic prediction after tumor treatment through the neoantigen load, especially the prognostic prediction of esophageal cancer after Anti-PD-L1 treatment. Compared with the existing TMB, MSI-H, and dMMR biomarkers, it has a better prediction effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 Schematic flow chart of the neoantigen screening method provided by the embodiment of the present invention; Figure 2 Schematic diagram of the neoantigen screening system provided by the embodiment of the present invention; Figure 3 Schematic diagram of the neoantigen screening device provided by the embodiment of the present invention; Figure 4 Training a high-precision MHC-antigen peptide affinity prediction model by transfer learning. (a) Training a convolutional neural network using the data in IEDB to obtain a pan-cancer MHC-antigen peptide affinity prediction model; (b) Adopting transfer learning technology, combining transcriptome and proteomic data for a one-step convolutional neural network to obtain a prediction model for a specific cancer type; Figure 5 To provide a clinical trial to verify the diagnostic and prognostic roles of neoantigen burden in the treatment of esophageal cancer with immune checkpoint inhibitors in the embodiments of the present invention; (a) Sequencing and data analysis of tumor and serum samples of esophageal cancer patients, screening for neoantigens and calculating neoantigen burden by combining tumor antigens, HLA typing, and antigen host gene expression levels; (b) Stratifying the enrolled patients and performing clinical immune checkpoint inhibitor treatment, recording clinical relevant indicators and side effects; (c) Combining the neoantigen burden, performing statistical analysis on immune response and prognosis, and verifying the roles of neoantigen burden in diagnosis and prognosis. Detailed implementation manners

[0019] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0020] In some processes described in the specification, claims, and above-mentioned drawings of the present invention, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0021] Figure 1 Schematic diagram of the neoantigen screening method provided by the embodiments of the present invention, specifically including: S101: Obtain the tumor section data and serum data of the patient; In one embodiment, the tumor includes one or more of the following: brain tumor, glioma, breast cancer, prostate cancer, esophageal cancer, liver cancer, lymphoma, melanoma.

[0022] In one embodiment, esophageal cancer is a heterogeneous tumor. Screening for neoantigens based on gene mutations within the tumor clone line and defining and calculating the tumor neoantigen burden is a highly challenging task. This project intends to use the maximum likelihood method to integrate multiple variables such as clone line, MHC-antigen affinity, and TCR-epitope, define the tumor neoantigen burden index, and determine the neoantigen burden critical value of a specific tumor based on clinical trial data and public data.

[0023] In a specific embodiment, 100 patients were enrolled in a randomized controlled study. The treatment group was treated with concurrent chemoradiotherapy combined with the PD-L1 antibody Imfinzi (durvalumab), and the control group was treated with concurrent chemoradiotherapy combined with a placebo. The neoantigen load was calculated to stratify esophageal cancer patients, and combined with the follow-up data after clinical treatment, the companion diagnostic and prognostic roles of neoantigen load in esophageal cancer immunotherapy were verified.

[0024] In a specific embodiment, sequencing of tumor and serum clinical samples and collection of clinical data of esophageal cancer patients: The research team of the present invention will select cases that meet the following conditions as research subjects among the patients first diagnosed with esophageal cancer in this hospital: aged between 18 and 70 years old; clinical stage IIIA-IV, unable to undergo surgical treatment; no history of malignant tumors; no history of chemotherapy; before chemotherapy, the patient's blood routine, heart, liver and kidney functions are normal. All research subjects were treated with a combined chemotherapy regimen mainly based on cisplatin or carboplatin. Informed consent form. All patients received 4-6 cycles of chemotherapy, and PD-L1 inhibitors were used during chemotherapy. According to their clinical chemotherapy data, the chemotherapy efficacy of the patients was evaluated according to the RECIST standard: complete remission (CR) - all target lesions disappeared; partial remission (PR) - the sum of the longest diameters of the target lesions decreased by at least 30%; stable disease (SD) - taking the minimum value of the sum of the longest diameters at the start of treatment as a reference, not meeting the PR standard and not meeting the PD standard; progression (PD) - the sum of the longest diameters of the target lesions increased by at least 20%, taking the minimum value of the sum of the longest diameters at the start of treatment or when one or more new lesions appeared as a reference.

[0025] According to the evaluation criteria of the National Cancer Institute of the United States (NCI 3.0), the grade 3 and 4 toxic and side reactions of chemotherapy for all patients were evaluated. The adverse reactions caused by chemotherapy include leukopenia, neutropenia, thrombocytopenia, anemia, nausea, vomiting, diarrhea, etc. The adverse reactions were routinely divided into three groups for subsequent analysis: (i) all grade 3 or 4 toxicity reactions; (ii) all grade 3 or 4 hematological toxicity reactions; (iii) all grade 3 or 4 gastrointestinal toxicity reactions.

[0026] Referring to the strategies for the construction of the NIH specimen bank in the United States and tumor cohort studies, an esophageal cancer specimen bank was established that collected complete clinical treatment and follow-up data, experimental research data, and tumor tissue and body fluid samples at each treatment and follow-up stage according to the unified standards of the present invention. A third-party sequencing agency was commissioned to perform whole-exome sequencing on tumor and serum samples, obtain the original fastq files and perform mutation analysis, extract tumor antigen sequences, and use the developed neoantigen screening model to calculate the neoantigen load index for each patient. At the same time, the PD-L1 expression level, prognostic indicators such as TMB and MSI were collected as the basis for control analysis.

[0027] S102: Perform data processing on the tumor section data and serum data to obtain HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression levels; the data processing includes data sequencing, HLA / MHC typing, somatic mutation calculation, fusion gene calculation, and gene expression profile calculation; In one embodiment, the process of the data processing is as follows: S21: Perform exome sequencing on the tumor section data to obtain exome sequencing data; S22: Perform transcriptome sequencing on the serum data to obtain transcriptome sequencing data; S23: Perform HLA / MHC typing on the exome sequencing data to obtain HLA / MHC pseudo-sequences and perform somatic mutation calculation to obtain mutation data; S24: Calculate gene expression profile data and fusion gene data based on the transcriptome sequencing data; S25: Calculate antigen gene expression level data based on the gene expression profile data; S26: Fuse the fusion gene and the mutation data to obtain antigen peptide data.

[0028] S103: Screen for neoantigens based on the HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression level data to obtain neoantigens; In one embodiment, the screening is performed by inputting the HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression level data into a deep network model to obtain the affinity of the HLA / MHC pseudo-sequences and antigen peptide data, combining the HLA / MHC pseudo-sequences with the antigen peptides to obtain HLA / MHC-antigen peptides, and then screening for neoantigens through the probability of TCR recognizing HLA / MHC-antigen peptides.

[0029] In one embodiment, the training process of the trained deep network model is as follows: First step: Obtain HLA / MHC pseudo-sequences, antigen peptide data, antigen gene expression level data, and IEDB data; Second step: Perform pre-training on the initial deep network model through the IEDB data to obtain the weights of the model; Third step: Transfer the weights of the model to the deep network model to obtain a transferred deep network model; Fourth step: Perform vectorization processing on the HLA / MHC pseudo-sequences and antigen peptide data to obtain HLA / MHC pseudo-sequence vector data and antigen peptide vector data; Fifth step: Perform feature extraction on the antigen gene expression level data to obtain an antigen deep feature vector; Step 6: Fuse the HLA / MHC pseudo-sequence vector data, antigen peptide vector data, and antigen depth feature vectors to obtain fused data, and input it into the transfer deep network model for training to obtain a trained deep network model.

[0030] In one embodiment, the training process of the affinity prediction model further includes HLA / MHC antigen peptide feature data, obtaining sequence feature data and structural feature data of the HLA / MHC antigen peptide, and inputting the sequence feature data and structural feature data of the HLA / MHC antigen peptide and the HLA / MHC antigen peptide mass spectrometry dataset into the training model for training to obtain an affinity prediction model.

[0031] In one embodiment, the training process of the deep network model further includes sequence feature data, obtaining the antigen peptide sequence and the sequence data of the HLA / MHC binding groove, extracting features from the antigen peptide sequence and the sequence data of the HLA / MHC binding groove to obtain sequence features, and inputting the HLA / MHC pseudo-sequence, antigen peptide data, antigen gene expression level data, and sequence features into the deep network model to obtain the affinity of the HLA / MHC pseudo-sequence and antigen peptide data.

[0032] In one embodiment, the sequence features are extracted by a feature model, and the construction process of the feature model is as follows: Obtain the antigen peptide sequence and the sequence of the HLA / MHC binding groove from a protein database; Extract sequence features from the sequence; Extract structural features from the sequence; Construct a feature training set based on the sequence features and structural features; Input the feature training set into a second deep network model for training to obtain a feature model.

[0033] In one embodiment, the structural features include one or more of the following: structural features, structural neighbor features, and solvent accessible surface area.

[0034] In one embodiment, the method of feature extraction adopts one or more of the following: physicochemical extraction, local structure entropy extraction, pairwise potential extraction, and interaction tendency extraction.

[0035] In one embodiment, the second deep network model includes one or more of the following: GBRT, XGBoost, LightGBM, CatBoost, Random Forest, AdaBoost.

[0036] In one embodiment, the training process of the deep network model further includes similar feature data, obtaining the structures of HLA / MHC molecules and antigen peptides to generate HLA / MHC network structures and antigen peptide network structures, extracting features from the HLA / MHC network structures and antigen peptide network structures to obtain similar features, and inputting HLA / MHC pseudo-sequences, antigen peptide data, antigen gene expression level data, and similar features into the deep network model to obtain the affinity of HLA / MHC pseudo-sequences and antigen peptide data.

[0037] In one embodiment, the similar features are extracted by a similar feature model, and the construction process of the similar feature model is as follows: obtaining the structures of HLA / MHC molecules and antigen peptides to generate HLA / MHC network structures and antigen peptide network structures; the HLA / MHC includes HLA / MHC-I subtypes and HLA / MHC-II subtypes; Calculating the similarity network of HLA / MHC-I subtypes and HLA / MHC-II subtypes and the antigen peptide similarity network; Obtaining the association relationship between HLA / MHC molecules and antigen peptides; Based on the similarity network of HLA / MHC-I subtypes and HLA / MHC-II subtypes and the antigen peptide similarity network, and the association relationship between HLA / MHC molecules and antigen peptides, inputting them into the third deep network model to obtain the similar feature model; In one embodiment, the third deep network model adopts one or more of the following: two-pass heterogeneous network, heterogeneous information network, graph neural network, and hypergraph.

[0038] In one embodiment, the training process of the deep network model further includes sequence feature data and similar features, obtaining antigen peptide sequences and sequence data of HLA / MHC binding grooves, extracting features from the antigen peptide sequences and sequence data of HLA / MHC binding grooves to obtain sequence features; obtaining the structures of HLA / MHC molecules and antigen peptides to generate HLA / MHC network structures and antigen peptide network structures, extracting features from the HLA / MHC network structures and antigen peptide network structures to obtain similar features; inputting HLA / MHC pseudo-sequences, antigen peptide data, antigen gene expression level data, sequence features, and similar features into the deep network model to obtain the affinity of HLA / MHC pseudo-sequences and antigen peptide data.

[0039] In one embodiment, the TCR recognition probability is calculated by a trained recognition model, and the construction process of the recognition model is as follows: Obtaining antigen epitope-TCR interaction data, and the sequences and structures of TCRs and peptide chains; Extracting TCR features based on the sequence and structure data of TCRs and peptide chains; Input the TCR features and antigen epitope-TCR interaction data into a fourth deep neural network model for training to obtain an identification model.

[0040] In one embodiment, the TCR features include one or more of the following: the charge, hydrophobicity, and two-dimensional structural features of CDR3.

[0041] In one embodiment, the extraction of TCR features is obtained through group sparse regularization regression analysis.

[0042] In one embodiment, the calculation of the probability of TCR recognition further includes TCR-antigen epitope binding prediction. The probability of TCR recognition is obtained through the identification model and TCR-antigen epitope binding prediction; the TCR-antigen epitope binding prediction is obtained through a trained TCR-antigen epitope binding prediction model.

[0043] In one embodiment, the construction process of the trained TCR-antigen epitope binding prediction model is as follows: Obtain the sequence of the CDR3 region, gene mutation data, and antigen peptides. Convert the sequence of the CDR3 region, gene mutation data, and antigen peptides into a feature matrix. Input the feature matrix into a neural network for training to obtain a TCR-antigen epitope binding prediction model.

[0044] In one embodiment, the feature matrix of the antigen peptides includes the sequence features and structural features of the antigen peptides.

[0045] In one embodiment, the TCR-antigen epitope binding prediction model further includes MHC-antigen peptide affinity calculation, and the affinity prediction result is obtained through the sequence features and structural features of the antigen peptides.

[0046] In one embodiment, the calculation of the probability of TCR recognition further includes immunogenicity prediction of tumor antigens. The probability of TCR recognition is obtained through the identification model and immunogenicity prediction of tumor antigens, and the immunogenicity prediction of tumor antigens is obtained through a trained immunogenicity prediction model of tumor antigens.

[0047] In one embodiment, the construction process of the immunogenicity prediction model of tumor antigens is as follows: Obtain RNA-seq data of different tumor-infiltrating lymphocytes, tumor cell genomes, and immunogenicity data. Perform TCR repertoire analysis on the RNA-seq data to obtain tumor gene expression profile data and obtain the CDR3 sequence from the RNA-seq data. Obtain gene mutation and antigen peptide data based on the tumor cell genome data. Feature extraction is performed on the CDR3 sequence and antigen peptide data to obtain feature matrix data; The tumor gene expression profile data, gene mutations, and feature matrix data are input into a neural network for training to obtain an immunogenicity prediction model.

[0048] In one embodiment, the neoantigen load calculation is obtained by calculating the affinity between the antigen peptide of the neoantigen and the MHC molecule and the probability of TCR recognition.

[0049] In one embodiment, the formula for calculating the neoantigen load is:

[0050] where N is the neoantigen load, n represents the number of tumor clone lines analyzed from the patient's exome sequencing data, b i represents the affinity, r i represents the probability of TCR recognition, c j represents the expression level of gene j.

[0051] In a specific embodiment, the present invention conducts tumor neoantigen screening driven by big data and deep learning. Based on the integration of multi-omics data such as genomics, transcriptomics, and proteomics, it deeply explores tumor neoantigens generated by different gene mutation types, the affinity between MHC molecules and antigen peptides, the recognition of antigen peptides by TCR, and the clinical response biomarkers of PD-L1 blockers. High-dimensional heterogeneous features are extracted from the biological processes of tumor antigen generation, presentation, and immune response, and a variety of intelligent algorithms including deep learning are developed and used to screen high-quality neoantigens with immunogenicity.

[0052] In a specific embodiment, prediction of the affinity between MHC-I molecules and antigen peptides: Integrate large-scale mass spectrometry data of MHC-I-bound antigen peptides, construct a mass spectrometry-based MHC-antigen peptide benchmark dataset and develop an online database; analyze genomic or whole exome sequencing data for MHC typing, extract the sequence and structural features of antigen peptides and MHC binding grooves, and propose and develop a deep learning-based MHC-antigen peptide affinity prediction model; construct an MHC-antigen peptide heterogeneous network that fuses multi-source data, combine network topology features and ontology features, and train a deep feedforward network to predict the affinity between MHC-antigen peptides.

[0053] It is proposed to improve the prediction performance of MHC-I and antigen peptide affinity from three aspects: 1) Integrate large-scale mass spectrometry data of MHC-I binding antigen peptides, construct a mass spectrometry-based MHC-antigen peptide benchmark dataset, first pre-train a deep neural network-based model using IEDB data, and then perform transfer learning on the pre-trained model with mass spectrometry data to obtain a higher-performance prediction model, as Figure 4 shown; 2) Combine sequence and structural features to predict MHC-I and antigen peptide affinity: Extract the sequences of antigen peptides and MHC-I binding grooves from the protein database, use one-hot encoding for amino acid sequences, and use various methods to extract sequence features including physicochemical properties, local structure entropy, pairing potential, interaction tendency, etc. At the same time, extract secondary structure features, structural neighbor features, solvent accessible surface area and other structural features, combine these sequence- and structure-based features to construct a training set, learn a gradient boosting regression tree model, determine the important factors for the affinity between MHC-I molecules and antigen peptides, and develop an online prediction service. 3) Calculate the similarity networks of MHC-I and MHC-II subtypes and the antigen peptide similarity network, integrate the MHC I molecule-antigen peptide associations in IEDB, thereby construct an MHC-antigen peptide heterogeneous network integrating multi-source heterogeneous data, and use the random walk algorithm on the two-pass heterogeneous network to predict the probability of potential MHC molecules binding antigen peptides.

[0054] In a specific embodiment, TCR recognizes tumor-specific antigens and immunogenicity prediction: Extract the sequence and structural features of antigen peptides and TCR variable regions from genomic data, combine the experimentally verified TCR-antigen peptide binding dataset, perform regression analysis and correlation analysis to discover the key factors for TCR to recognize tumor antigens; construct a deep learning model, use antigen- and TCR-related features as inputs to predict the probability of TCR recognizing and binding tumor antigens; integrate gene mutations of different tumors to construct a gene mutation-TCR correspondence matrix, and analyze the probability of tumor-specific antigens being recognized by TCR from the perspective of tumor clonal evolution; combine tumor-infiltrating lymphocyte (TIL) data and experimentally obtained immunogenicity data, use tumor gene expression profiles and antigen peptide feature matrices as inputs to a deep neural network to predict the immunogenicity of tumor antigens.

[0055] CD8+ T lymphocytes recognize and kill tumor cells through the complementary determining region (CDR3) of the T cell receptor (TCR). The interaction between CDR3 and antigen peptides presented by MHC is the key to the adaptive immune response. TCR is determined by the V(D)J recombination process, and the human TCR repertoire can accommodate up to 10^15 different molecular types. Unless from the same clone, the TCRs of each T lymphocyte are different, but different TCRs can bind to the same antigen, which poses great difficulties for computationally predicting immunogenic tumor antigens. We believe that the TCR regions of T cells recognizing the same pMHC complex contain conserved sequence features (motifs), and we plan to develop three methods to predict the probability of TCR recognizing tumor antigens and activating immune responses: 1) Identification of the main factors for TCR-antigen epitope binding: Collect and integrate experimentally verified antigen epitope-TCR interaction data, extract the sequences and structures of TCRs and peptide chains, including the charge, hydrophobicity, two-dimensional structural features of CDR3 and the one-hot encoding of antigens, etc. Use group sparse regularization regression to analyze the influence of each feature on TCR recognition of antigens, construct a deep neural network model and train it with antigen-TCR interaction data to predict the probability of a specific TCR subtype recognizing a specific antigen. 2) Prediction of TCR-antigen epitope binding based on CNN: Obtain the sequence of the CDR3 region from RNA-seq data, and use autocross covariation to convert CDR3 sequences of different lengths into feature matrices of the same dimension; obtain gene mutations and antigen peptides from genomic or exome sequencing data and convert them into feature matrices using one-hot encoding; The prediction of MHC-antigen peptide affinity and TCR-antigen epitope binding share the sequence and structural features of antigen peptides. Construct a multi-task deep learning model to perform two tasks simultaneously: prediction of MHC-antigen peptide affinity and TCR-antigen epitope binding; 3) Profile the TCR repertoire of RNA-seq data from different tumor-infiltrating lymphocytes (TILs), combine genomic or exome sequencing data of tumor cells to obtain gene mutations, collect experimentally obtained immunogenic data sets, and use the tumor gene expression profile, antigen peptide feature matrix and CDR3 feature matrix as the input of a deep neural network to predict the immunogenicity of tumor antigens.

[0056] In a specific embodiment, a neoantigen load index is defined and calculated: The activation of T cell responses by tumor antigens is a multi-step complex process, in which antigen presentation and TCR recognition of the pMHC complex are the main steps affecting T cell recognition and killing of target cells. Failure of any step will lead to the failure of immunotherapy. For antigen i of tumor clone line j, assuming the affinity between MHC-I molecules and this antigen peptide is bi, and the probability of TCR recognizing this antigen is ri, the neoantigen load of this patient is defined as:

[0057]

[0058] Among them, n represents the number of tumor clone lines analyzed from the exome sequencing data of this patient.

[0059] It should be noted that when developing the bioinformatics model of the present invention, multi-omics data will be used, including genomic or exome, transcriptome, and mass spectrometry data. Once the prediction model is trained, only the exome sequencing or large panel sequencing data of the patient is required during use, which is beneficial to reducing the cost of clinical companion diagnosis and promoting market application and promotion.

[0060] In a specific embodiment, an esophageal cancer specimen bank is established with complete clinical treatment and follow-up data, experimental research data, and tumor tissue and body fluid samples at each treatment and follow-up stage collected according to the unified standard of the present invention. Whole exome sequencing is performed on the tumor and serum samples to obtain the original exome sequencing data for mutation analysis, and non-synonymous mutations (NSVs), insertion and deletion mutations (InDels) are obtained. HLA typing analysis can also be performed; fusion genes and genomic expression profiles are analyzed from transcriptome sequencing; tumor antigen sequences are extracted according to the mutations, and antigen peptides, HLA typing, and antigen host gene expression levels are integrated as the input of the neoantigen screening model to screen high-quality neoantigens and calculate the neoantigen load index for each patient, such as Figure 5 shown in (a) of. At the same time, prognostic indicators such as PD-L1 expression level, TMB, and MSI are collected as the basis for control analysis.

[0061] In a specific embodiment, the approved marketing of immune checkpoint inhibitors such as anti-PD-L1 monoclonal antibody and anti-CTLA-4 monoclonal antibody has greatly promoted the research of tumor immunotherapy. However, currently, the population that benefits from immunotherapy is very limited. For example, the objective response rate of anti-PD-L1 inhibitor monotherapy is only about 30%. The companion diagnostic indicators of immune checkpoint inhibitors, TMB and PD-L1 protein level, cannot accurately screen patients who can benefit from anti-PD-L1 monoclonal antibody treatment because cytotoxic T lymphocytes cannot recognize tumor-specific antigen epitopes. Therefore, combining multi-omics data to screen individual-specific biomarkers as companion diagnostics for immune checkpoint inhibitors is the development direction of precision medicine. The present invention intends to collect and integrate clinical data such as the responses of various tumors to the immune checkpoint inhibitor PD-L1, progression-free survival, and overall survival, construct a benchmark PD-L1 inhibitor clinical data set, and on this basis, develop two methods for biomarker screening: 1) Analyze the correlation between neoantigens generated by tumor driver mutations and passenger mutations and the response and prognosis of PD-L1 inhibitors respectively to find driver gene mutations with statistical significance; 2) Obtain gene mutations and MHC subtypes from the exome sequencing data of tumor cells and paired normal cells, obtain the spectra of gene expression and TCR repertoire from RNA-seq data, integrate relevant features such as antigen peptide sequences, tumor gene expression profiles, MHC typing, and TCR repertoire, combine large-scale clinical immune response data, extract features from high-dimensional heterogeneous data, and train a deep neural network model to achieve high-precision prediction of tumor-specific immune checkpoint response and prognosis, providing more reference indicators for the companion diagnosis of clinical treatment of immune checkpoints.

[0062] In a specific embodiment, for the anti-PD-L1 clinical treatment of esophageal cancer patients, the enrolled patients are treated with the marketed PD-L1 antibody Imfinzi (durvalumab). The treatment process is carried out strictly in accordance with the NCCN clinical diagnosis and treatment plan, and the patients' clinical data such as biochemistry, immunity, and imaging are recorded. At the same time, for the related side effects, the side effects such as hematological toxicity, gastrointestinal reactions, heart damage, lung damage, and skin damage are mainly evaluated. According to the calculated neoantigen load, the esophageal cancer patients are stratified and grouped, and statistical analysis is performed on the prognostic indicators such as the overall survival (OS), objective response rate (ORR), overall response rate (OR), and progression-free survival (PFS) of the patients to determine the statistical significance of the neoantigen load in each prognostic indicator, as Figure 5 shown. In particular, compared with indicators such as PD-L1+, TMB, and MSI, the prognostic role of the neoantigen load is verified.

[0063] The disclosed embodiment of the present invention also provides a computer program product or system, including a computer program, which when executed by a processor, implements the steps of the above-mentioned neoantigen screening method.

[0064] Figure 2 Schematic diagram of a neoantigen screening system provided by an embodiment of the present invention, specifically including: Acquisition unit: acquiring tumor slice data and serum data of a patient; Processing unit: performing data processing on the tumor slice data and serum data to obtain HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression levels; the data processing includes data sequencing, HLA / MHC typing, somatic mutation calculation, fusion gene calculation, and gene expression profile calculation; Screening unit: screening neoantigens based on the HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression level data to obtain neoantigens.

[0065] Figure 3 Schematic diagram of a neoantigen screening device provided by an embodiment of the present invention, specifically including: A memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions to perform any one of the above-mentioned neoantigen screening methods.

[0066] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it performs any one of the above-mentioned neoantigen screening methods.

[0067] The verification results of this verification embodiment show that assigning fixed weights to indications can improve the performance of this method compared to the default settings. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk, or optical disc, etc.

[0068] Those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned medium storage can be read-only memory, magnetic disk, or optical disc, etc.

[0069] The above has introduced in detail a computer device provided by the present invention. For those of ordinary skill in the art, according to the idea of the embodiments of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for screening new antigens, characterized in that: include: S1. Obtain the patient's tumor slice data and serum data; S2. Processing the tumor slice data and serum data to obtain HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression levels; the data processing includes data sequencing, HLA / MHC typing, somatic mutation calculation, fusion gene calculation, and gene expression profile calculation; S3. Screening new antigens based on the HLA / MHC pseudo sequence, antigen peptide data, and antigen gene expression level data to obtain new antigens.

2. The method for screening new antigens according to claim 1, characterized in that: The data processing process is as follows: S21, performing exome sequencing based on the tumor section data to obtain exome sequencing data; S22, performing transcriptome sequencing based on the serum data to obtain transcriptome sequencing data; S23, performing HLA / MHC typing based on the exome sequencing data to obtain HLA / MHC pseudo sequences and performing somatic mutation calculation to obtain mutation data; S24, calculating gene expression profile data and fusion gene data based on the transcriptome sequencing data; S25, calculating and obtaining antigen gene expression level data based on the gene expression profile data; S26. The fusion gene and the mutation data are fused to obtain antigen peptide data.

3. The method for screening new antigens according to claim 1, characterized in that: The screening is performed by inputting HLA / MHC pseudo-sequences, antigen peptide data, and antigen gene expression level data into a deep network model to obtain the affinity of HLA / MHC pseudo-sequences and antigen peptide data, combining HLA / MHC pseudo-sequences with antigen peptides to obtain HLA / MHC-antigen peptides, and then screening to obtain new antigens through the probability of TCR recognizing HLA / MHC-antigen peptides; Optionally, the training process of the deep network model is: The first step is to obtain HLA / MHC pseudo sequences, antigen peptide data, antigen gene expression level data, and IEDB data; The second step is to pre-train the initial deep network model using the IEDB data to obtain the model weights; Step 3: Migrating the weights of the model to the deep network model to obtain a migrated deep network model; Step 4: vectorize the HLA / MHC pseudo sequence and antigen peptide data to obtain HLA / MHC pseudo sequence vector data and antigen peptide vector data; Step 5: Extract features from antigen gene expression level data to obtain antigen depth feature vectors; Step 6: Fuse the HLA / MHC pseudo-sequence vector data, antigen peptide vector data, and antigen deep feature vector to obtain fused data, and input them into the migration deep network model for training to obtain a trained deep network model; Optionally, the training process of the affinity prediction model also includes HLA / MHC antigen peptide feature data, obtaining sequence feature data and structural feature data of the HLA / MHC antigen peptide, and inputting the sequence feature data and structural feature data of the HLA / MHC antigen peptide and the HLA / MHC antigen peptide mass spectrum data set into the training model for training to obtain the affinity prediction model.

4. The method for screening new antigens according to claim 3, characterized in that: The training process of the deep network model also includes sequence feature data, obtaining sequence data of antigen peptide sequence and HLA / MHC binding groove, extracting features of the sequence data of the antigen peptide sequence and HLA / MHC binding groove to obtain sequence features, and inputting HLA / MHC pseudo sequence, antigen peptide data, antigen gene expression level data, and sequence features into the deep network model to obtain the affinity of HLA / MHC pseudo sequence and antigen peptide data; The sequence features are extracted through the feature model, and the feature model construction process is as follows: Obtain antigen peptide sequences and HLA / MHC binding groove sequences from the protein database; Performing feature extraction on the sequence to obtain sequence features; Extracting structural features from the sequence to obtain structural features; Constructing a feature training set based on the sequence features and structural features; Inputting the feature training set into a second deep network model for training to obtain a feature model; Optionally, the structural features include one or more of the following: structural features, structural neighbor features, volume accessible surface area; Optionally, the feature extraction method adopts one or more of the following: physical and chemical extraction, local structure entropy extraction, pairing potential energy extraction, interaction tendency extraction; Optionally, the second deep network model includes one or more of the following: GBRT, XGBoost, LightGBM, CatBoost, Random Forest, AdaBoost.

5. The method for screening new antigens according to claim 3, characterized in that: The training process of the deep network model also includes similar feature data, obtaining the structures of HLA / MHC molecules and antigen peptides to generate HLA / MHC network structures and antigen peptide network structures, extracting features from the HLA / MHC network structures and antigen peptide network structures to obtain similar features, and inputting HLA / MHC pseudo sequences, antigen peptide data, antigen gene expression level data, and similar features into the deep network model to obtain the affinity of the HLA / MHC pseudo sequences and antigen peptide data; Optionally, the similar features are extracted by a similar feature model, and the construction process of the similar feature model is: obtaining the structures of HLA / MHC molecules and antigen peptides to generate HLA / MHC network structures and antigen peptide network structures; the HLA / MHC includes HLA / MHC-I subtype and HLA / MHC-II subtype; Calculate the similarity network of HLA / MHC-I subtypes and HLA / MHC-II subtypes, and the antigen peptide similarity network; Obtain the association between HLA / MHC molecules and antigen peptides; Based on the similarity network of HLA / MHC-I subtype and HLA / MHC-II subtype and the antigen peptide similarity network, the association relationship between HLA / MHC molecules and antigen peptides is input into the similarity feature model in the third deep network model; Optionally, the third deep network model adopts one or more of the following: a two-way heterogeneous network, a heterogeneous information network, a graph neural network, and a hypergraph; Optionally, the training process of the deep network model further includes sequence feature data and similarity features, obtaining sequence data of antigen peptide sequences and HLA / MHC binding grooves, and performing feature extraction on the sequence data of the antigen peptide sequences and HLA / MHC binding grooves to obtain sequence features; Acquiring the structures of HLA / MHC molecules and antigen peptides to generate HLA / MHC network structures and antigen peptide network structures, and extracting features from the HLA / MHC network structures and antigen peptide network structures to obtain similar features; The HLA / MHC pseudo sequence, antigen peptide data, antigen gene expression level data, sequence features, and similar features are input into the deep network model to obtain the affinity of the HLA / MHC pseudo sequence and antigen peptide data.

6. The method for screening new antigens according to claim 3, characterized in that: The probability of TCR recognition is calculated by a trained recognition model, and the construction process of the recognition model is as follows: Obtain epitope-TCR interaction data, TCR and peptide chain sequences and structures; Extracting TCR characteristics based on the sequence and structure data of the TCR and peptide chain; Inputting the TCR characteristics and antigen epitope-TCR interaction data into a fourth deep network model for training to obtain a recognition model; Optionally, the TCR characteristics include one or more of the following: charge, hydrophobicity, and two-dimensional structural characteristics of CDR3; Optionally, the extracted TCR features are obtained by group sparse regularized regression analysis; Optionally, the probability calculation of TCR recognition also includes TCR-epitope binding prediction, and the probability of TCR recognition is obtained by the recognition model and TCR-epitope binding prediction; the TCR-epitope binding prediction is obtained by predicting the trained TCR-epitope binding prediction model; Optionally, the trained TCR-antigen epitope binding prediction model is constructed by: Obtain the sequence, gene mutation data and antigenic peptides of the CDR3 region; Converting the sequence of the CDR3 region, gene mutation data, and antigenic peptides into a feature matrix; Inputting the feature matrix into a neural network for training to obtain a TCR-antigen epitope binding prediction model; Optionally, the characteristic matrix of the antigenic peptides includes the characteristics of the sequence and structure of the antigenic peptides; Optionally, the TCR-antigen epitope binding prediction model also includes MHC-antigen peptide affinity calculation, and the affinity prediction results are obtained by combining the sequence characteristics and structural characteristics of the antigen peptide; Optionally, the probability calculation of TCR recognition also includes the prediction of the immunogenicity of the tumor antigen, and the probability of TCR recognition is obtained by the recognition model and the prediction of the immunogenicity of the tumor antigen, and the prediction of the immunogenicity of the tumor antigen is obtained by predicting the trained immunogenicity prediction model of the tumor antigen; Optionally, the process of constructing the immunogenicity prediction model of the tumor antigen is: Obtain RNA-seq data of different tumor-infiltrating lymphocytes, tumor cell genomes, and immunogenicity data; Performing TCR library spectrum analysis on the RNA-seq data to obtain tumor gene expression spectrum data and obtaining CDR3 sequences from the RNA-seq data; Acquiring gene mutation and antigen peptide data based on the tumor cell genome data; Feature extraction is performed on CDR3 sequence and antigen peptide data to obtain feature matrix data; The tumor gene expression profile data, gene mutation, and feature matrix data are input into a neural network for training to obtain an immunogenicity prediction model.

7. The method for screening new antigens according to claim 1, characterized in that: The method further comprises calculating the neoantigen load by calculating the affinity of the HLA / MHC-antigen peptide of the neoantigen and the probability of TCR recognizing the HLA / MHC-antigen peptide; Optionally, the formula for calculating the neoantigen load is: Where N is the neoantigen load, n represents the number of tumor clones analyzed from the patient's exome sequencing data, b i Express affinity, r i represents the probability of TCR recognition, c j represents the expression level of gene j.

8. A computer program product comprising a computer program or instructions, characterized in that: The computer program or instructions are executed by a processor to implement the method for screening new antigens according to any one of claims 1 to 7.

9. A computer device comprising a memory, a processor and a computer program or instruction stored in the memory, characterized in that: The computer program or instructions are executed by a processor to implement the method for screening new antigens according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: The computer program or instructions are executed by a processor to implement the method for screening new antigens according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for screening tumor neoantigen based on HLA typing and structure

    CN110675913A

  • Neoantigen prediction method and device based on next-generation sequencing and storage medium

    CN110752041A

  • Tumor immunotherapy effect prediction system based on NGS and deep learning

    CN115424740A

  • Deep learning model for predicting tumor-specific neoantigen MHC class I or class II immunogenicity

    CN117136410A

  • MHC-I type molecular neoantigen recognition method based on multi-instance learning

    CN118553308A

Cited By

  • Tumor neoantigen screening method, device, equipment, storage medium and product

    CN120708721A

  • Screening method and screening device of neoantigen

    CN120954512A