Neopeptide sequence generation method and system based on deep learning

By generating novel antigen sequences through deep learning, the problem of insufficient immune capacity in the design of novel antigen vaccines has been solved, achieving better binding with HLA and T cell recognition, and supporting personalized vaccine development.

CN116403639BActive Publication Date: 2026-02-10NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310331805.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2026-02-10
Estimated Expiration
2043-03-30

AI Technical Summary

Technical Problem

Current technologies lack research on the generation of neoantigen sequences in unknown protein spaces, resulting in insufficient immunogenicity in neoantigen vaccine design.

Method used

Using a deep learning-based approach, novel antigens with better affinity for HLA are generated by acquiring, preprocessing, mutating, and screening new antigen sequences. Combined with biological validation, novel antigen sequences with immunogenicity are screened out.

Benefits of technology

This significantly increases the exploration of protein space, generates new antigens with better affinity for HLA, improves T cell recognition capabilities, and provides guidance for personalized vaccine development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116403639B_ABST
    Figure CN116403639B_ABST
Patent Text Reader

Abstract

The application discloses a kind of new antigen sequence generation method and system based on deep learning, the method includes: obtaining original new antigen data;Original new antigen data is preprocessed, and original new antigen sequence and its corresponding HLA sequence are obtained;According to original new antigen sequence and its corresponding HLA sequence, original new antigen sequence is mutated, and monte carlo search is carried out in sequence space to original new antigen sequence, and the new antigen sequence after mutation that meets design target is determined;Through molecular docking, presentation ability prediction and immune ability prediction, the screening of new antigen sequence after mutation is carried out, and the final new antigen sequence is obtained.The application generates new new antigen based on the method of deep learning, and the generated new antigen sequence is screened, so that the new antigen finally obtained has more optimal affinity and immune capacity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of neoantigen sequence generation technology, and in particular to a method and system for generating neoantigen sequences based on deep learning. Background Technology

[0002] Neoantigens, also known as neoantigens, are new proteins formed on tumor cells (including malignant tumor cells, i.e., cancer cells) when certain mutations occur in their DNA. Neoantigens are selective and can elicit a T-cell response against the tumor, thereby eliminating it. This makes neoantigens a key element in the design of cancer vaccines. Neoantigens play a crucial role in helping the body mount an immune response against cancer cells, and are being researched for use in vaccines and other types of immunotherapies to treat many types of cancer.

[0003] Existing technologies typically identify and characterize immunogenic tumor neoantigens, using algorithmic design based on existing neoantigens. However, these neoantigens currently only occupy a small portion of the protein space, and the vast unknown protein space has not yet been explored. In other words, current research on neoantigens focuses on neoantigen identification and neoantigen vaccine design, lacking research on generating neoantigen sequences. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a method and system for generating novel antigen sequences based on deep learning. The deep learning method generates entirely new neoantigens, which exhibit better affinity for human leukocyte antigens (HLA) compared to existing neoantigens. By screening the generated neoantigen sequences, new neoantigens that can also be recognized and remembered by T cells are obtained. With subsequent biological validation, these neoantigens can achieve better immune responses compared to existing neoantigens, providing guidance for the development of personalized vaccines.

[0005] Firstly, this disclosure provides a method for generating novel antigen sequences based on deep learning.

[0006] A method for generating novel antigen sequences based on deep learning includes:

[0007] Obtain the original neoantigen data;

[0008] The original neoantigen data was preprocessed to obtain the original neoantigen sequence and its corresponding HLA sequence;

[0009] Based on the original neoantigen sequence and its corresponding HLA sequence, the original neoantigen sequence is mutated, and a Monte Carlo search is performed on the original neoantigen sequence in the sequence space to determine the mutated neoantigen sequence that meets the design objectives.

[0010] The mutated neoantigen sequences were screened by molecular docking, presentation capability prediction, and immune capability prediction to obtain the final neoantigen sequences.

[0011] A further technical solution, the preprocessing includes:

[0012] The original neoantigen data was screened to identify human neoantigen data, including both neoantigen data and potential neoantigen data.

[0013] Filter the potential neoantigen data to select those with high confidence scores;

[0014] For each screened human neoantigen data, HLA-I type data were selected, and the original neoantigen sequence and its corresponding HLA sequence for each data point were obtained.

[0015] A further technical solution, wherein the mutation of the original neoantigen sequence based on the original neoantigen sequence and its corresponding HLA sequence includes:

[0016] Based on the original neoantigen sequence and its corresponding HLA sequence, the attention weight of each amino acid in the original neoantigen sequence to the HLA sequence is calculated. Based on the attention weight, amino acid positions are selected for mutation.

[0017] A further technical solution, wherein performing a Monte Carlo search in the sequence space on the original neoantigen sequence to determine the mutated neoantigen sequence that meets the design target, includes:

[0018] The mutated new antigen sequence is input into the protein structure prediction model, and the predicted protein structure is output.

[0019] The mutated neoantigen sequence and its corresponding HLA sequence are input into the peptide-HLA binding affinity prediction model, and the predicted affinity is output.

[0020] Using structural confidence and affinity score as loss functions, a Monte Carlo search is performed on the original neoantigen sequence in the sequence space to determine the final protein structure. Then, a protein sequence design model is used to output the corresponding amino acid sequence.

[0021] A further technical solution, wherein the screening of mutated neoantigen sequences through molecular docking, presentation capability prediction, and immune capability prediction to obtain the final neoantigen sequence includes:

[0022] The binding affinity of the generated neoantigen sequence to its corresponding HLA sequence is determined by molecular docking. Based on the binding affinity, a new antigen sequence with higher affinity is selected.

[0023] Further technical solutions also include:

[0024] By using a deep learning-based presentation capability model, we can capture the binding information of HLA-binding peptides, predict whether HLA-binding peptides can be presented on the cell surface, and screen for new antigen sequences that predict HLA-binding peptides can be presented on the cell surface.

[0025] Further technical solutions also include:

[0026] Using neoantigen characteristics as a filtering criterion, the neoantigen characteristics are incorporated into the immunogenicity score to calculate the immunogenicity score; the neoantigen characteristics include the dissociation constant and binding stability of the neoantigen-HLA complex and the expression of mutated genes in the tumor.

[0027] The neoantigen with the highest score is used as the output to obtain the final neoantigen sequence.

[0028] Secondly, this disclosure provides a new antigen sequence generation system based on deep learning.

[0029] A novel antigen sequence generation system based on deep learning includes:

[0030] The data acquisition module is used to acquire raw neoantigen data;

[0031] The data preprocessing module is used to preprocess the raw neoantigen data to obtain the raw neoantigen sequence and its corresponding HLA sequence;

[0032] The neoantigen sequence design module is used to mutate the original neoantigen sequence based on the original neoantigen sequence and its corresponding HLA sequence, and to perform a Monte Carlo search in the sequence space on the original neoantigen sequence to determine the mutated neoantigen sequence that meets the design target.

[0033] The neoantigen sequence screening module is used to screen mutated neoantigen sequences through molecular docking, presentation capability prediction, and immunogenicity prediction to obtain the final neoantigen sequence.

[0034] Thirdly, this disclosure also provides an electronic device, including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the steps of the method described in the first aspect.

[0035] Fourthly, this disclosure also provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the steps of the method described in the first aspect.

[0036] The above one or more technical solutions have the following beneficial effects:

[0037] 1. This invention provides a method and system for generating novel antigen sequences based on deep learning. The method generates entirely new antigens based on deep learning, which have better affinity for human leukocyte antigens (HLA) compared to the original antigens. By screening the generated novel antigen sequences, new antigens that can also be recognized and remembered by T cells are obtained, that is, new antigens with immune capabilities are obtained.

[0038] 2. This invention utilizes deep learning to design novel antigens, which greatly increases the exploration of the protein space compared to traditional approaches that only target existing antigens.

[0039] 3. With subsequent biological validation, this invention can achieve better immune response compared to existing neoantigens, providing guidance for the development of personalized vaccines. Attached Figure Description

[0040] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0041] Figure 1 This is an overall flowchart of the deep learning-based new antigen sequence generation method described in Embodiment 1 of the present invention;

[0042] Figure 2 This is a flowchart of the neoantigen design algorithm in Embodiment 1 of the present invention. Detailed Implementation

[0043] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0044] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0045] Example 1

[0046] This embodiment provides a method for generating neoantigen sequences based on deep learning. First, the data from the collected original neoantigen database is preprocessed. Then, deep learning methods are used to mutate the input neoantigen sequence, resulting in a reasonably mutated neoantigen sequence with superior affinity for its corresponding HLA. Finally, the generated neoantigen is screened multiple times using a series of methods, combined with biological verification, to obtain a neoantigen with stronger immunogenicity as the final result. The method described in this embodiment is as follows... Figure 1 As shown, it includes the following steps:

[0047] Step S1: Obtain the original neoantigen data;

[0048] Step S2: Preprocess the original neoantigen data to obtain the original neoantigen sequence and its corresponding HLA sequence;

[0049] Step S3: Based on the original neoantigen sequence and its corresponding HLA sequence, mutate the original neoantigen sequence and perform a Monte Carlo search in the sequence space to determine the mutated neoantigen sequence that meets the design target.

[0050] Step S4: Screen the mutated neoantigen sequence through molecular docking, presentation capability prediction and immune capability prediction to obtain the final neoantigen sequence.

[0051] First, in step S1, raw neoantigen data is acquired to construct a raw neoantigen dataset. Given the high correlation between accurate and complete datasets and the efficiency of deep learning, collecting high-quality datasets is a crucial starting point for generating neoantigen sequences. In this embodiment, data is downloaded from databases such as NeoPeptide, TSNAdb, and TransLnc, which contain approximately 170,000 experimental neoantigens, 1.3 million, and 400,000 potential neoantigens, respectively, covering a range of cancers including melanoma, breast cancer, lung cancer, and squamous cell carcinoma.

[0052] Secondly, in step S2, the original neoantigen data is preprocessed to obtain the original neoantigen sequence and its corresponding HLA sequence. Before analyzing the input original neoantigen data, this data needs to be preprocessed. First, because the database contains data from multiple species, such as humans and mice, and this embodiment focuses on human neoantigens, neoantigens from other species are filtered out to select human neoantigens. This filtering method includes: for each data entry in the database, determining whether the HLA in each data entry is human HLA (e.g., HLA-A, B, C, etc.), and sequentially determining whether the data entry originates from humans. Then, because the data contains not only experimental neoantigens but also potential neoantigens, the potential neoantigen portion needs further filtering to select potential neoantigens with high confidence scores. This selection method includes: calculating the confidence score using a neoantigen identification tool (such as netMHCpan). The result obtained using netMHCpan is %rank. If %rank < 0.5, then the neoantigen is considered to have a high affinity for the corresponding HLA, and the potential neoantigen is selected as input. Finally, HLA-I type data were selected, and the necessary portion of each data point was extracted so that each data point contained only the neoantigen sequence and its corresponding HLA sequence. Here, the HLA sequence refers to the amino acid sequence corresponding to HLA; the following descriptions will all use HLA sequences.

[0053] Then, in step S3, a new antigen sequence is designed. This involves mutating the original new antigen sequence based on its corresponding HLA sequence, and performing a Monte Carlo search in the sequence space to determine the mutated new antigen sequence that meets the design objective. The algorithm is as follows: Figure 2 As shown, the method described in this embodiment is based on deep learning and explores the protein space of the neoantigen extensively by mutating the input original neoantigen sequence. This method has no restrictions on the structure of the original sequence.

[0054] Step S3.1: Based on the original neoantigen sequence and its corresponding HLA sequence, calculate the attention weight of each amino acid in the original neoantigen sequence to the HLA sequence, sort according to the attention weight, and select amino acid positions for mutation.

[0055] Attention is calculated for the original neoantigen sequence and the corresponding HLA sequence, using the following formula:

[0056]

[0057] Here, K and V represent the neoantigen sequence, and Q represents the corresponding HLA sequence. Each character in both the neoantigen and HLA sequences represents an amino acid; therefore, the parameters mentioned above actually represent amino acid sequences. After the softmax operation in the formula, each amino acid in the neoantigen sequence receives its own attention weight.

[0058] Next, the intermediate calculation results are taken. Specifically, for HLA, the attention weight of each amino acid element in the neoantigen sequence is averaged, and the resulting value is used as the importance score of the corresponding amino acid. The higher the score, the more important the corresponding amino acid may play in binding to HLA. The scores are sorted from low to high, and mutations are only performed on the amino acids at the top of the list.

[0059] Step S3.2: Input the mutated neoantigen sequence into the protein structure prediction model, outputting the predicted protein structure. Input the mutated neoantigen sequence and its corresponding HLA sequence into the peptide-HLA binding affinity prediction model, outputting the predicted affinity. Using structure confidence and affinity score as loss functions, perform a Monte Carlo search in the sequence space on the original neoantigen sequence to determine the final protein structure. Then, design a protein sequence model to output the corresponding amino acid sequence. The aforementioned Monte Carlo search refers to Markov chain Monte Carlo optimization; the aforementioned structure confidence refers to the model output values ​​plddt and ptm, which represent the difference between the model's predicted output structure and the actual structure. plddt represents the confidence of each residue in the predicted structure by calculating the intrinsic distance of the predicted structure in the absence of the native structure; ptm represents the template modeling score of the model prediction, which evaluates the structure at the overall level.

[0060] A Monte Carlo search is performed on the original neoantigen sequence in sequence space, combining a protein structure prediction model and a peptide-HLA binding affinity prediction model. The protein structure prediction model takes an amino acid sequence as input and outputs the predicted protein structure. It extracts features from the sequence, encodes and decodes these features into the 3D location of the protein structure to obtain the final output. The peptide-HLA binding affinity prediction model takes a peptide and its corresponding HLA as input and outputs the binding affinity between the peptide and the corresponding HLA. The training data scale is increased by using binding affinity data and elution ligand data, and the peptide-HLA binding affinity prediction model uses the NNAlign_MA architecture, enabling the model to handle both types of data. In this embodiment, during the training of the protein structure prediction model and the peptide-HLA binding affinity prediction model, the training datasets for the protein structure prediction model are the protein structure dataset PDB and the protein sequence dataset Uniclust30, while the training datasets for the peptide-HLA binding affinity prediction model are the binding affinity dataset and the elution ligand dataset.

[0061] The outputs of the two models described above are used as the loss function for Monte Carlo search. By combining Monte Carlo and the above models, sequences are continuously searched until the obtained structure meets the design goal. Finally, a protein sequence design model is used to redesign the sequence. This model, unlike the protein structure prediction model, takes the protein structure as input and outputs the corresponding amino acid sequence. Through a message-passing neural network, an encoder-decoder architecture, and a flexible decoding order, the model achieves higher performance. The training dataset for the protein sequence design model is the Protein Structure Dataset (PDB). The sequences obtained through the search need to have their structures predicted using the protein structure prediction model. For each predicted structure, it is then input into the protein sequence design model to generate the final corresponding sequence, thereby reducing overfitting during the search process.

[0062] Although the neoantigens obtained through the above design methods are expected to bind better to HLA, only a small number of HLA-binding peptides can actually be presented on the cell surface. Therefore, it is not possible to obtain neoantigens that can better stimulate immunity based solely on affinity. In this embodiment, the following method is also used for subsequent screening.

[0063] That is, in step S4, the mutated neoantigen sequence is screened by molecular docking, presentation capability prediction and immune capability prediction to obtain the final neoantigen sequence.

[0064] After obtaining the amino acid sequence of the neoantigen through step S3 above, further screening of the obtained neoantigen sequence is still required. During the subsequent screening process, if structural data is needed, a structure prediction model can be used to obtain the corresponding structure for screening, ultimately determining the required neoantigen sequence after screening.

[0065] Molecular docking is widely used in the study of protein-ligand interactions and in drug discovery and development. Typically, given a target with a known structure, molecular docking is used to predict the binding conformation and binding free energy of small molecules to the target, determining how they bind. In this embodiment, molecular docking is used to determine the binding affinity between a designed neoantigen sequence and its corresponding HLA sequence. Based on the binding affinity, a screening process is performed to obtain neoantigen sequences with higher affinity.

[0066] Neoantigens not only need to bind to HLA, but also need to be presented on the cell surface to be recognized by T cells and trigger an immune response. To address this, this embodiment utilizes a deep learning-based presentation capability model to predict whether HLA-binding peptides can be presented on the cell surface. This model makes predictions by capturing the binding information and circulation patterns of HLA-binding peptides. The presentation capability model is trained using an eluted ligand dataset containing both monoalletic and multialletic HLA antigens. This deep learning model filters out neoantigens that cannot be presented on the cell surface, leading to more convincing results.

[0067] To further screen for neoantigens with immunogenicity, in this embodiment, neoantigen characteristics are used as filtering criteria, including the dissociation constant and binding stability of the neoantigen-HLA complex and the expression of mutated genes in the tumor. These characteristics are incorporated into the immunogenicity score to calculate a score. The immunogenicity score is calculated for the neoantigen sequence, and the neoantigen with the highest score is taken as the output result to obtain the final neoantigen sequence.

[0068] The above-described scheme in this embodiment utilizes deep learning to design novel neoantigens, which greatly increases the exploration of the protein space compared to traditional schemes that only target existing neoantigens. Furthermore, after using deep learning to design neoantigens, the design results are further screened to obtain neoantigens with immunogenicity, making them more practically significant and providing guidance for the development of subsequent personalized vaccines.

[0069] Example 2

[0070] This embodiment provides a new antigen sequence generation system based on deep learning, including:

[0071] The data acquisition module is used to acquire raw neoantigen data;

[0072] The data preprocessing module is used to preprocess the raw neoantigen data to obtain the raw neoantigen sequence and its corresponding HLA sequence;

[0073] The neoantigen sequence design module is used to mutate the original neoantigen sequence based on the original neoantigen sequence and its corresponding HLA sequence, and to perform a Monte Carlo search in the sequence space on the original neoantigen sequence to determine the mutated neoantigen sequence that meets the design target.

[0074] The neoantigen sequence screening module is used to screen mutated neoantigen sequences through molecular docking, presentation capability prediction, and immunogenicity prediction to obtain the final neoantigen sequence.

[0075] Example 3

[0076] This embodiment provides an electronic device, including a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When the processor executes the computer instructions, it completes the steps in the deep learning-based new antigen sequence generation method described above.

[0077] Example 4

[0078] This embodiment also provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the steps in the deep learning-based new antigen sequence generation method described above.

[0079] The steps and methods involved in Embodiments 2 to 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0080] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0081] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0082] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for generating novel antigen sequences based on deep learning, characterized in that, include: Obtain the original neoantigen data; The original neoantigen data was preprocessed to obtain the original neoantigen sequence and its corresponding HLA sequence; Based on the original neoantigen sequence and its corresponding HLA sequence, the original neoantigen sequence is mutated, and a Monte Carlo search is performed in the sequence space to determine the mutated neoantigen sequence that meets the design objectives, including: The mutated new antigen sequence is input into the protein structure prediction model, and the predicted protein structure is output. The mutated neoantigen sequence and its corresponding HLA sequence are input into the peptide-HLA binding affinity prediction model, and the predicted affinity is output. Using structural confidence and affinity score as loss functions, a Monte Carlo search is performed on the original neoantigen sequence in the sequence space to determine the final protein structure. Then, a protein sequence design model is used to output the corresponding amino acid sequence. The mutated neoantigen sequences were screened using molecular docking, presentation capability prediction, and immunogenicity prediction to obtain the final neoantigen sequences, including: The binding affinity between the generated neoantigen sequence and its corresponding HLA sequence is determined by molecular docking. Based on the binding affinity, the neoantigen sequence with higher affinity is selected. By predicting presentation capability, a deep learning-based presentation capability model is used to capture the binding information of HLA-binding peptides, predict whether HLA-binding peptides can be presented on the cell surface, and screen for new antigen sequences that predict HLA-binding peptides can be presented on the cell surface. By predicting immune capacity and using neoantigen characteristics as a filtering criterion, neoantigen characteristics are incorporated into the immunogenicity score to calculate the immunogenicity score; the neoantigen characteristics include the dissociation constant and binding stability of the neoantigen-HLA complex and the expression of mutated genes in the tumor. The neoantigen with the highest score is used as the output to obtain the final neoantigen sequence.

2. The method for generating novel antigen sequences based on deep learning as described in claim 1, characterized in that, The preprocessing includes: The original neoantigen data was screened to identify human neoantigen data, including both neoantigen data and potential neoantigen data. Filter the potential neoantigen data to select those with high confidence scores; For each screened human neoantigen data, HLA-I type data were selected, and the original neoantigen sequence and its corresponding HLA sequence for each data point were obtained.

3. The method for generating novel antigen sequences based on deep learning as described in claim 1, characterized in that, The mutation of the original neoantigen sequence based on the original neoantigen sequence and its corresponding HLA sequence includes: Based on the original neoantigen sequence and its corresponding HLA sequence, the attention weight of each amino acid in the original neoantigen sequence to the HLA sequence is calculated. Based on the attention weight, amino acid positions are selected for mutation.

4. A novel antigen sequence generation system based on deep learning, characterized in that, include: The data acquisition module is used to acquire raw neoantigen data; The data preprocessing module is used to preprocess the raw neoantigen data to obtain the raw neoantigen sequence and its corresponding HLA sequence; The neoantigen sequence design module is used to mutate the original neoantigen sequence based on the original neoantigen sequence and its corresponding HLA sequence, and to perform a Monte Carlo search in the sequence space to determine the mutated neoantigen sequence that meets the design objectives, including: The mutated new antigen sequence is input into the protein structure prediction model, and the predicted protein structure is output. The mutated neoantigen sequence and its corresponding HLA sequence are input into the peptide-HLA binding affinity prediction model, and the predicted affinity is output. Using structural confidence and affinity score as loss functions, a Monte Carlo search is performed on the original neoantigen sequence in the sequence space to determine the final protein structure. Then, a protein sequence design model is used to output the corresponding amino acid sequence. The neoantigen sequence screening module is used to screen mutated neoantigen sequences through molecular docking, presentation capability prediction, and immunogenicity prediction to obtain the final neoantigen sequence, including: The binding affinity between the generated neoantigen sequence and its corresponding HLA sequence is determined by molecular docking. Based on the binding affinity, the neoantigen sequence with higher affinity is selected. By predicting presentation capability, a deep learning-based presentation capability model is used to capture the binding information of HLA-binding peptides, predict whether HLA-binding peptides can be presented on the cell surface, and screen for new antigen sequences that predict HLA-binding peptides can be presented on the cell surface. By predicting immune capacity and using neoantigen characteristics as a filtering criterion, neoantigen characteristics are incorporated into the immunogenicity score to calculate the immunogenicity score; the neoantigen characteristics include the dissociation constant and binding stability of the neoantigen-HLA complex and the expression of mutated genes in the tumor. The neoantigen with the highest score is used as the output to obtain the final neoantigen sequence.

5. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, complete the steps of a deep learning-based method for generating novel antigen sequences as described in any one of claims 1-3.

6. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, complete the steps of a deep learning-based method for generating novel antigen sequences as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Method of producing tumor-reactive t cell composition using modulatory agents

    US20230032934A1

  • Method and system for screening neoantigens, and uses thereof

    US20230047716A1