Method for predicting T cell activity against peptide-MHC and analysis device

Through gene data analysis and neural network model prediction of T cell activity, the problem of difficult to predict T cells on antigenic tumor cell-specific peptide-MHC complex in the prior art was solved, and the effect of efficient screening of high T cell activity antigens was achieved, and the effect of cancer immunotherapy was improved.

JP7672638B2Active Publication Date: 2025-05-08KOREA ADVANCED INST OF SCI & TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023560163
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-30
Filing Date
2021-12-16
Publication Date
2025-05-08
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict the activity of T cells on antigenic tumor cell-specific peptide-MHC complexes, affecting the effect of cancer immunotherapy.

Method used

By entering the patient's genetic data, the amino acid sequences of MHC and tumor cell antigens are identified, a matrix representing the interaction between the two sequences is generated, and the matrix is ​​input into the trained neural network model to predict whether T cells secrete cytokines, such as interferon-γ.

Benefits of technology

A method of rapidly screening out high T cell active antigens is achieved. Using interferon-γ secretion as a reference, the activity of T cells on antigenic tumor cells is accurately predicted, and the effect of cancer immunotherapy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672638000004
    Figure 0007672638000004
  • Figure 0007672638000005
    Figure 0007672638000005
  • Figure 0007672638000006
    Figure 0007672638000006
Patent Text Reader

Abstract

The method for predicting T cell activity against an MHC-peptide includes the steps of: receiving genetic data of a patient into an analysis device; identifying a first amino acid sequence of an MHC (major histocompatibility complex) and a second amino acid sequence of an antigen produced by a tumor cell based on the genetic data; generating a matrix showing the interrelationship between the first amino acid sequence and the second amino acid sequence in units of one amino acid by the analysis device; and inputting the matrix into a trained neural network model by the analysis device to determine whether or not T cells secrete cytokines above a threshold level due to the binding of the MHC and the antigen.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The technology described below relates to a technique for predicting T cell activity against an antigen peptide-MHC. [Background technology]

[0002] Neoantigens are tumor cell-specific proteins. Neoantigens are expressed by tumor cell-specific mutations. Epitopes of neoantigens are expressed on major histocompatibility complexes (MHC) located on the surface of tumor cells, and T cells recognize MHC-epitopes to elicit immune responses.

[0003] Cancer immunotherapy is a treatment that activates the body's immune system to destroy tumor cells. Research is currently underway to discover effective neoantigens in the field of cancer immunotherapy. Summary of the Invention [Problem to be solved by the invention]

[0004] The technology described below aims to provide an in silico technique for discovering neoantigens that are highly reactive to T cells. [Means for solving the problem]

[0005] The method for predicting T cell activity against a peptide-MHC includes the steps of: receiving genetic data of a patient into an analysis device; identifying a first amino acid sequence of a major histocompatibility complex (MHC) and a second amino acid sequence of an antigen produced by a tumor cell based on the genetic data; generating a matrix showing the interrelationship between the first amino acid sequence and the second amino acid sequence in units of amino acids by the analysis device; and inputting the matrix into a trained neural network model by the analysis device to determine whether or not T cells secrete cytokines above a threshold level due to the binding of the MHC and the antigen.

[0006] The analysis device for predicting T cell activity against peptide-MHC includes an input device for inputting genetic data of a patient, a storage device for storing a neural network model for predicting the amount of cytokine secretion by T cells based on a matrix showing the interrelationship between the amino acid sequence of MHC (major histocompatibility complex) and the amino acid sequence of an antigen produced by tumor cells, and a calculation device for identifying a first amino acid sequence of MHC and a second amino acid sequence of an antigen produced by tumor cells from the genetic data, generating a matrix showing the interrelationship between the first amino acid sequence and the second amino acid sequence in units of amino acids, inputting the generated matrix into the neural network model, and determining whether the MHC-antigen of the patient induces the secretion of interferon-gamma by T cells. Effect of the Invention

[0007] The technology described below uses a dip-running model to rapidly select neoantigens with high T cell activation from patient candidate peptides. The technology described below uses the amount of interferon gamma secreted as a criterion to accurately predict T cell activation against antigen peptide-MHC. [Brief description of the drawings]

[0008] [Figure 1] This is an example of a system for predicting T cell activity against peptide-MHC. [Diagram 2] This is an example of the process of developing a custom-made anti-cancer vaccine. [Diagram 3] 1 is an example of the process of learning a neural network model. [Figure 4] 1 is an example of a process for generating a matrix showing peptide-MHC interactions. [Diagram 5] 1 is an example of a process for predicting T cell activity against peptide-MHC. [Figure 6] 1 is an example of an analytical device for predicting T cell activity against peptide-MHC. [Figure 7] 13 is an example of an experimental result verifying a neural network model. [Figure 8] This is another example of experimental results verifying the neural network model. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] The technology described below can be modified in various ways and can have various embodiments, and a specific embodiment will be illustrated in the drawings and described in detail. However, this is not intended to limit the technology described below to a specific embodiment, and it should be understood that the technology described below includes all modifications, equivalents, and alternatives included in the spirit and technical scope of the technology described below.

[0010] Terms such as first, second, A, B, etc. may be used to describe various components, but the components are not limited by the terms and are used merely to distinguish one component from another. For example, a first component may be named a second component, and similarly, a second component may be named a first component, without departing from the scope of the technology described below. The term "and / or" includes a combination of multiple related listed items or any item of multiple related listed items.

[0011] In the terms used in this specification, the singular term "a" or "an" should be understood to include the plural term unless clearly interpreted otherwise, and terms such as "comprise" or "comprise" should be understood to mean the presence of a stated feature, number, step, operation, component, part, or combination thereof, and not to exclude the presence or additional possibility of one or more other features, numbers, step operations, components, parts, or combinations thereof.

[0012] Before proceeding to a detailed description of the drawings, it is to be clear that the components in this specification are merely classified according to the main function of each component. That is, two or more components described below may be combined into one component, or one component may be divided into two or more components according to more specific functions. Furthermore, each component described below may perform a part or all of the functions of other components in addition to its own main function, and of course, some of the main functions of each component may be performed by other components.

[0013] Furthermore, in carrying out a method or method of operation, the steps making up the method may be performed out of the order specified unless a particular order is expressly recited herein, i.e., the steps may be performed in the same order as specified, substantially simultaneously, or in the reverse order.

[0014] The terms used in the following description will be explained below.

[0015] An antigen is a substance that induces an immune response.

[0016] Neoantigens are tumor cell specific antigens resulting from tumor cell mutations or post-translational modifications. Neoantigens may include polypeptide sequences or nucleotide sequences. Here, mutations may include any genomic or expression modifications that cause frameshifts, insertions, deletions, substitutions, splice site changes, genomic rearrangements, gene fusions or new open reading frames (ORFs). Mutations may also include splice variants. Tumor cell specific post-translational modifications may include aberrant phosphorylation. Tumor cell specific post-translational modifications may also include spliced ​​antigens produced by the proteasome.

[0017] Epitope can refer to the specific part of an antigen that an antibody or T-cell receptor normally binds.

[0018] MHC is a peptide structure that functions as a mediator that recognizes the target substance of the immune response as an antigen. Human MHC is also called HLA (human leukocyte antigen). Hereinafter, MHC is used to include human HLA.

[0019] Peptide means a polymer of amino acids. The technology described below corresponds to a technique for discovering neoantigens. In the following description, peptide means an amino acid polymer or an amino acid sequence expressed in tumor cells. Therefore, in the following, peptide means a tumor-specific amino acid polymer or an amino acid sequence expressed on the surface of tumor cells.

[0020] Peptide-MHC (pMHC) or peptide-MHC complexes are peptide and MHC structures expressed on the surface of tumor cells. T cells recognize peptide-MHC complexes and initiate immune responses.

[0021] The degree of binding refers to the degree of binding between the MHC and a peptide. The binding preference or binding affinity refers to the degree of binding affinity between the MHC molecule and a peptide.

[0022] By sample is meant a single cell or multiple cells, cell fragments, body fluids, etc., in an individual to be analyzed.

[0023] A subject includes a cell, tissue, or organism. Typically, a subject is obtained from a patient with a particular tumor. A subject is primarily, but not limited to, a human subject.

[0024] An exome is a subset of the genome that encodes proteins. An exome can refer to the collection of exons present in a cell, a group of cells, or an individual.

[0025] Genetic data refers to genetic information calculated by analyzing a sample. For example, genetic data may include base sequences obtained from deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or protein from cells, tissues, etc., gene expression data, genetic variations compared to standard genetic data, DNA methylation, etc. Genetic data can be obtained using traditional sequencing methods, next-generation sequencing (NGS), etc. Genetic data is generally digital data and can be calculated in a specific file format (e.g., FATSQ).

[0026] Machine learning is a field of artificial intelligence that develops algorithms to enable computers to learn. Learning models include decision trees, random forests, KNN (K-nearest neighbor), Naive Bayes, SVM (support vector machine), and artificial neural networks. The techniques described below can utilize artificial neural networks. The following explanation focuses on artificial neural networks and neural network models.

[0027] An artificial neural network is a statistical learning algorithm that mimics biological neural networks. Various neural network models have been researched. Recently, deep learning networks (DNNs) have attracted attention. DNNs are artificial neural network models that consist of multiple hidden layers between an input layer and an output layer. Like general artificial neural networks, DNNs can model complex non-linear relationships. Various types of DNN models have been researched. Examples include CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), GAN (Generative Adversarial Network), and RL (Relation Networks).

[0028] The analyzer is a device that extracts neoantigens of specific tumors from patient samples. The analyzer predicts the activity of T cells against specific peptide-MHC. The analyzer can process and analyze data using installed programs or codes.

[0029] Fig. 1 is an example of a system 100 for predicting T cell activity against peptide-MHC. In Fig. 1, analysis devices 130, 140, and 150 predict T cell activity. In Fig. 1, the analysis devices are shown in the form of a server 130 and computer terminals 130 and 140. It should be noted that the analysis devices 130, 140, and 150 may be embodied in various forms.

[0030] The genetic analyzer 110 analyzes a patient sample and generates genetic data. For example, the genetic analyzer 110 may be an NGS analyzer. Since peptides expressed by tumor cells are the subject of analysis, the genetic analyzer 110 may perform sequencing targeting exomes. A detailed description of whole exome sequencing will be omitted. The genetic analyzer 110 may also store the generated genetic data in a separate DB 120.

[0031] The server 130 receives the genetic data from the genetic analyzer 110 or the DB 120. The server 130 analyzes the genetic data and provides a service for predicting T cell activity.

[0032] The computer terminal 140 receives the genetic data from the genetic analyzer 110 or the DB 120. The computer terminal 140 analyzes the genetic data and predicts T cell activity.

[0033] The computer terminal 150 receives the genetic data via a medium (e.g., a USB, an SD card, etc.) in which the genetic data generated by the genetic analyzer 110 is stored. The computer terminal 150 analyzes the genetic data and predicts T cell activity.

[0034] The analyzers 130, 140, 150 predict the degree of T cell activity for the peptide-MHC currently being analyzed, and if the degree of T cell activity is equal to or greater than a threshold, the current peptide can be extracted as a neoantigen candidate. The analyzers 130, 140, 150 input information about the peptide-MHC into a pre-constructed neural network model to predict the degree of T cell activity. The process by which the analyzers 130, 140, 150 predict T cell activity for a specific peptide-MHC using the neural network model will be described later.

[0035] The users 10, 20, 30 may be researchers or medical staff who develop neoantigens and vaccines. The users 10, 20, 30 can confirm the degree of T cell activity against a specific peptide-MHC of a sample. The users 10, 20, 30 can also grasp the neoantigens effective for the sample. The user 10 can connect to the server 130 via a user terminal (PC, smartphone, etc.) and confirm the analysis results performed by the server 130. The user 20 can confirm the M analysis results via the computer terminal 140 that the user uses. The user 30 can confirm the analysis results via the computer terminal 150 that the user uses.

[0036] Figure 2 shows an example of a process (200) for developing a custom-made anti-cancer vaccine. Figure 2 includes a process in which an analyzer predicts T cell activity using a neural network model and discovers neoantigens based on the prediction results.

[0037] The analysis device learns the neural network model using the learning data and constructs the neural network model (210). The neural network model learning process will be described later. The neural network model learning process can be performed by a separate computer device instead of the analysis device. That is, the entity that learns the neural network model and the entity that performs analysis using the neural network model can be different entities.

[0038] The analytical device receives input of gene sequence analysis results (genetic data) for a patient sample (e.g., tumor tissue). The analytical device can identify a tumor-specific mutated sequence from the genetic data. The analytical device can identify a mutated sequence from the tumor tissue sequence based on a normal tissue sequence or a reference sequence. The analytical device can identify the tumor-specific mutated sequence as a tumor-specific antigen (220). That is, the analytical device can select the tumor-specific mutated sequence as a neoantigen candidate. The analytical device can select multiple neoantigen candidates. The following description will be given assuming that multiple neoantigen candidates have been identified.

[0039] The analyzer selects a specific candidate from among the neoantigen candidates and predicts T cell activity. The analyzer can identify a gene sequence for the specific candidate using the genetic data. The analyzer can determine an amino acid sequence of the candidate antigen based on the gene sequence for the specific candidate. The analyzer can also identify an MHC sequence for the patient using the genetic data. The analyzer can determine an amino acid sequence of the MHC based on the MHC sequence.

[0040] The analysis device predicts the T cell activity level using a neural network model previously constructed for the candidate antigen (230). The analysis device inputs information about the candidate antigen into the neural network model and predicts the T cell activity level. The analysis device can input the amino acid sequence of the candidate antigen and the amino acid sequence of the MHC into the neural network model and predict the T cell activity level. The analysis device can generate a matrix showing the interaction and affinity between the candidate antigen and the MHC based on the amino acid sequence of the candidate antigen and the amino acid sequence of the MHC. The analysis device can input the generated matrix into the neural network model and predict the T cell activity level for the candidate antigen. The operation of the neural network model will be described later.

[0041] The neural network model outputs information regarding the presence or absence of T cell activation for the input candidate antigen. For example, the neural network model can output whether T cells are activated or inactivated for the candidate antigen being analyzed. The presence or absence of T cell activation can be determined by whether the activation level is above a certain threshold.

[0042] The analyzer determines whether the activity of the T cells is above a critical value (240). The presence or absence of T cell activity can be determined based on the amount of cytokines secreted by the T cells. If the activity of the T cells against the current candidate antigen is above a critical value (YES in 240), the analyzer adds the current candidate antigen (peptide) to a target candidate group (250). The target candidate group is composed of tumor-specific neoantigen candidates that can be targets for immune anti-cancer therapy.

[0043] The analyzer can check whether prediction of T cell activity for the specific antigen of the sample is complete (260). If prediction of T cell activity for the candidate antigen of the sample is not complete (NO in 260), the analyzer selects the next specific antigen from the candidate antigens for which T cell activity is not predicted (270) and repeats the process of determining T cell activity for that antigen. The analyzer performs the process up to extracting a target candidate group.

[0044] Once the prediction of T cell activity against all candidate antigens is complete, researchers can perform further validation experiments on the current group of candidate targets (280). Furthermore, researchers can design vaccines that target neoantigens that currently elicit high T cell activity in patients (290). Anticancer vaccines produced through this process will be customized for each patient and will target only tumor cells.

[0045] The process of constructing the above-mentioned neural network model will be explained below. The explanation will be based on the process of a researcher actually constructing a neural network model.

[0046] The researchers collected information on peptide-MHC (pMHC) from public databases. The public databases may include the IEDB, etc. The researchers collected pMHC data for humans and mice from the IEDB, IMMA2, MHCBN, and other sources. The researchers also collected information on T cell activity in association with pMHC. The activity of T cells was determined based on the amount of cytokines secreted by the T cells. More specifically, the researchers evaluated the activity of T cells for a specific pMHC based on the amount of interferon gamma (IFNγ) secreted by the T cells. To this end, the researchers selected data from the pMHC data in the public database that contained data on the amount of IFNγ secreted by T cells for the pMHC.

[0047] The data the researchers collected included HLA type and peptide length, which is a 9-mer for MHC class I and a 15-mer for MHC class II, as well as immunogenicity labeling values ​​for each pMHC.

[0048] Unlike MHCI, MHCII is a heterodimer of HLA-DP and HLA-DQ, so the experimental data is given in the HLA-DQA / HLA-DQB pair. As described later, the molecular distance between the antigen and MHC was used as the learning data, so the researchers used only the β chain that acts directly on the antigen. The learning data was adjusted to have a good balance between antigenic and non-antigenic data. Ultimately, the researchers prepared 13,128 MHCI data and 6,650 MHCII data. The data collected by the researchers is shown in Table 1 below. In Table 1, "individual research" means data obtained through individual research. Immunogenic peptides means tumor-specific neoantigens.

[0049] [Table 1]

[0050] FIG. 3 is an example of a process (300) for training a neural network model.

[0051] Peptide DB-A can store information related to peptide-MHC. FIG. 3 shows only a public DB such as IEBD. However, the learning data may include not only public DBs but also data obtained by researchers (developers) in separate experiments. Peptide DB-A may be a device present on a network. Alternatively, peptide DB-A may be a device connected to or built into computer device B.

[0052] The neural network model can be trained by computer device B using the training data. Computer device B may be an analysis device that predicts T cell activity, or may be a device dedicated to training.

[0053] Computer device B extracts learning data from peptide DB-A (310). The learning data may include MHC class, amino acid sequence of antigen candidate, and amount of IFNγ secreted by T cells in response to the peptide-MHC, as shown in Fig. 3. The amount of IFNγ secreted can be classified into cases where it is equal to or higher than a critical value (high, H) as a criterion for judging T cell activity, and cases where it is lower than the critical value (low, L).

[0054] The computer device B generates a matrix showing the correlation between peptides and MHC for the peptide-MHC pairs for the specific antigen being analyzed (320). The correlation means the affinity between the peptide and the MHC. The matrix generation process will be described later.

[0055] Computer device B inputs the generated matrix into the neural network model to predict the presence or absence of T cell activity for the current peptide-MHC. The neural network model outputs information such as T cell activity (high IFNγ secretion) or inactivity (low IFNγ secretion) for the peptide-MHC. Computer device B updates the weights of the neural network model based on the label value (IFNγ secretion amount) for the currently input peptide-MHC (330). The neural network model can learn through a backpropagation process.

[0056] Researchers used the CNN model to predict T cell activity in response to peptide-MHC. Of course, other neural network models can be used to predict T cell activity instead of CNN. Here we will briefly explain the neural network model, focusing on CNN.

[0057] The CNN may include a convolution layer (Conv), a pooling layer, and a fully connected layer (FC). The convolution layer and the pooling layer may be arranged repeatedly.

[0058] The CNN model 400 predicts the peptide-MHC binding degree based on the input data (interaction map). The CNN model 400 includes multiple convolutional layers 410, 420, a fully connected layer 430, and an output layer 440. As shown in FIG. 5, the convolutional layer may be composed of two layers.

[0059] The convolution layer performs a convolution operation on the input data and outputs a value obtained by applying a rectified linear unit (ReLU) function to the convolution value. The convolution operation is an operation in which the input value is multiplied by a weight matrix. The weights can be updated through the learning process. The convolution layer extracts interaction features for peptide-MHC. The input data may include a parameter indicating the degree of interaction between amino acid pairs.

[0060] The fully connected layer integrates the input information. The fully connected layer receives the output values ​​from the convolution layer. The fully connected layer can perform ReLU operations.

[0061] The output layer uses a sigmoid function to output information regarding the degree of T cell activity or the presence or absence of activity for a given peptide-MHC.

[0062] The final output value of the neural network model can be a value between 0 and 1. The analysis device can compare the value output by the neural network model with a critical value to determine activation or inactivation of the T cells.

[0063] An additional explanation will be given based on the model shown in FIG.

[0064] A convolution layer uses a certain number of kernels or weight matrices to perform convolution. The convolution may be a one-dimensional or two-dimensional operation, etc. All convolution results are transformed by ReLU, which transforms negative values ​​to 0. Figure 3 shows two convolution layers.

[0065] The first convolution layer detects a connection pattern from the input data. The first convolution layer can use a window with a moving distance of 1. The operation of the convolution layer is as shown in Equation 1 below. The second convolution layer may have the same structure as the first convolution layer. Alternatively, the second convolution layer may have a different window size or stride width from the first convolution layer.

[0066]

number

[0067] X is the input data, i is the index indicating the position of the output, and k is the kernel index. Each convolution kernel W k corresponds to a weight matrix of size M×N, where M is the window size and N is the number of input channels.

[0068] Pooling layers may not be used. Pooling is a process of reducing the dimensionality of data. Amino acids that are relatively far from each other can also affect the interaction of peptide-MHC complexes with T cell receptors. Therefore, CNN can extract features without utilizing pooling layers while maintaining the size of the input data.

[0069] The fully connected layer FC takes all the outputs from the second convolutional layer as input. The fully connected layer combines the inputs from the previous layers. The fully connected layer performs the ReLU(WX) function, where X is the input value and W is the weight matrix for the fully connected layer.

[0070] The output layer can output a value between 0 and 1 using a sigmoid function. The value output by the output layer is the activation H or inactivation L of the T cell. The output layer performs a sigmoid function Sigmoid(WX), where X is the input value and W is the weight matrix for the sigmoid output layer. Note that the output layer can also use an activation function such as softmax or ReLU instead of sigmoid.

[0071] The CNN model is trained in the direction of minimizing the objective function. The training process corresponds to the process of optimizing the weights used in the CNN model. For example, the gradient descent method can be used to optimize the weights.

[0072] The objective function is defined as the sum of negative log likelihood (NLL) and a regularization term. The objective function for the CNN model is expressed as Equation 2 below.

[0073]

number

[0074] s is the index of the training data. t is the index of the interaction feature. Y t s is the label value (0 or 1) for the T cell activity for the training data s. t (X s ) is the neural network model that calculates the input data X s These are the predicted results for T cell activity.

[0075] In addition, MHCI and MHCII have different functional characteristics and protein lengths. Therefore, it is preferable to set up separate neural network models for MHCI and MHCII. Researchers have also constructed separate neural network models for MHCI and MHCII using separate training data.

[0076] The neural network model receives as input a matrix for peptide-MHC. The computer generates a matrix for peptide-MHC during the learning process. The analysis device generates matrices for each peptide-MHC during the analysis process. FIG. 4 is an example of the process (400) for generating a matrix showing peptide-MHC interactions. For ease of explanation, it is assumed in FIG. 4 that computer device B generates the matrix. Note that the computer device may be a PC, a server, etc.

[0077] The computer device B receives the amino acid sequence for the peptide-MHC (410). The computer device may receive the amino acid sequence via an input device, a storage medium, or communication. The amino acid sequence is the amino acid sequence of the MHC and the amino acid sequence of the antigen. Alternatively, the computer device may store a specific MHC amino acid sequence in advance, and only the amino acid sequence of the antigen may be input.

[0078] Computer device B generates a matrix for pairs of amino acid sequences of MHC and antigen. In the amino acid sequence of MHC, individual amino acids can be identified by order from 1 to n. In the amino acid sequence of antigen, individual amino acids can be identified by order from a to z.

[0079] Computer device B determines an interaction value for each amino acid pair between the amino acid sequence of the MHC (designated as the first amino acid sequence) and the amino acid sequence of the antigen (designated as the second amino acid sequence). For example, the computer device determines an interaction value between amino acid 1 of the first amino acid sequence and amino acid a of the second amino acid sequence. In this manner, the computer device determines interaction values ​​for all amino acid pairs that can be formed between the first amino acid sequence and the second amino acid sequence.

[0080] Computer device B can determine an interaction value on a specific amino acid by referring to previously known protein structures. Protein structure DB-A stores information on previously known protein structures. Protein structure DB-A can hold information on amino acids constituting a protein structure and distances between amino acids. Protein structure DB-A can hold information on a number of protein structures.

[0081] The computer device B can determine the distance (proximity) between a specific first amino acid of the first amino acid sequence and a specific second amino acid of the second amino acid sequence by referring to the protein structure DB-A. The protein structure DB-A can hold distance information for a large number of identical amino acid pairs. The computer device B can determine the interaction value of the first amino acid-second amino acid pair using various criteria. For example, (i) the computer device B can determine the average distance of the first amino acid-second amino acid pair in the protein structure DB-A as the interaction value of the first amino acid-second amino acid pair. (ii) the computer device B can determine the interaction value of the first amino acid-second amino acid pair based on the proximity frequency of the first amino acid-second amino acid in the protein structure DB-A. The computer device B can determine that the amino acid pair is proximal when the first amino acid-second amino acid is located within a certain reference distance in the second or third space in the protein structure DB-A. Next, computer device B can determine an interaction value of the first amino acid-second amino acid pair based on the frequency at which the first amino acid and the second amino acid are adjacent in protein structure DB-A. Computer device B can determine the number of times at which the first amino acid and the second amino acid are adjacent in protein structure DB-A as the interaction value. Alternatively, computer device B can determine the interaction value by processing the frequency at which the first amino acid and the second amino acid are adjacent in protein structure DB-A.

[0082] The interaction value between amino acid pairs can be determined in units of regions that divide the protein structure into certain regions, and can be determined based on the distance between Cα (alpha carbon) atoms present in the protein structure.

[0083] Computer device B refers to protein structure DB-A and extracts proximity information (such as distance or proximity frequency) of specific amino acid pairs (420). Computer device B determines an interaction value for each amino acid pair that constitutes the first amino acid sequence and the second amino acid sequence, and generates a matrix (430). Since the matrix indicates the interaction of the amino acid sequences, it can also be called an interaction matrix. The matrix for the first amino acid sequence and the second amino acid sequence is composed of information indicating the degree of interaction (affinity or proximity) for each amino acid pair.

[0084] An example of an interaction map is shown in the lower part of Fig. 4. The interaction map is a two-dimensional matrix having a horizontal axis and a vertical axis. The horizontal axis corresponds to the amino acid sequence of MHC labeled with 1 to n, and the vertical axis corresponds to the amino acid sequence of antigen labeled with a to z.

[0085] The matrix includes an interaction value for each amino acid pair. The interaction value may be a numerical value. Furthermore, the matrix may be in the form of a map in which the degree of interaction is indicated by a certain color.

[0086] In addition, the length of the amino acid sequence of the antigen may differ depending on the source data or MHC class. Therefore, the computer device can pad the matrix based on the largest input data.

[0087] FIG. 5 is an example of a process (500) for predicting T cell activity against a peptide-MHC.

[0088] The analytical device receives (510) genetic data of a sample. The sample may be tissue from a particular tumor patient. The genetic data may include information on multiple antigens. For ease of explanation, the description will be based on one peptide-MHC.

[0089] In addition, the analysis device can select a neural network model that has been previously constructed according to the MHC class. As described above, different neural network models can be prepared according to the MHC class. Therefore, the analysis device can select a neural network model that matches the MHC class of the current analysis target, and then proceed with the analysis process.

[0090] The analyzer extracts MHC amino acid sequences and antigen amino acid sequences from the genetic data. The analyzer can use a program or model to predict the MHC structure. For example, the analyzer can predict the HLA (Human Leukocyte Antigen) structure using HLAminer. The analyzer can also identify the antigen amino acid sequence from the genetic data using a certain program. For example, the analyzer can search for nonsynonymous mutations flanking amino acid sequences from the genetic data using the idfetch program to detect the antigen amino acid sequence.

[0091] As described with reference to FIG. 4, the analyzer can generate a matrix for the amino acid sequences of the MHC and the amino acid sequences of the antigen (520).

[0092] The analysis device inputs the generated matrix into the neural network model and performs analysis (530). The analysis device can determine the presence or absence of T cell activity for the peptide-MHC currently being analyzed based on the information (T cell activity or inactivity) that the neural network model outputs for the input matrix (540).

[0093] The analyzer can compare the value output by the neural network model with a critical value to determine whether or not there is T cell activity. The researchers constructed separate neural network models for MHCI and MHCII. In the case of the neural network models constructed using the training data described in Table 1, the MHCI neural network model determined that a value greater than 0.5 indicated T cell activity, and the MHCII neural network model determined that a value greater than 0.7 indicated T cell activity.

[0094] Furthermore, the analytical device can now determine that an antigen is a candidate target for an anti-cancer vaccine when T cells are activated against the peptide-MHC.

[0095] 6 shows an example of an analysis device 600 for predicting T cell activity against peptide-MHC. The analysis device 600 corresponds to the analysis device 130, 140 or 150 in FIG.

[0096] The analysis device 600 can predict the peptide-MHC binding degree using the above-mentioned neural network model. The analysis device 600 may be physically embodied in various forms. For example, the analysis device 600 may have the form of a PC, a smart device, a server on a network, a chip set dedicated to data processing, etc.

[0097] The analysis device 600 may include a storage device 610 , a memory 620 , a computing device 630 , an interface device 640 , a communication device 650 and an output device 660 .

[0098] The storage device 610 stores a neural network model that predicts the degree of activity of T cells. The neural network model is as described above. The neural network model must be trained in advance. The neural network model can output the amount of cytokine secretion corresponding to a measure of T cell activity. For example, the neural network model can output the amount of IFNγ secreted by the T cells. The neural network model can output information indicating the activity (secretion of a large amount of IFNγ) or inactivity (secretion of a small amount or no IFNγ) of the T cells.

[0099] Furthermore, the storage device 610 can store programs and source codes required for data processing.

[0100] The storage device 610 can store input genetic data, the amino acid sequence of an antigen to be analyzed, and the MHC amino acid sequence to be analyzed.

[0101] Storage device 610 may also store programs that identify MHC and / or antigen sequences from genetic data.

[0102] The storage device 610 can store the activity level of T cells against a specific peptide-MHC, which is the analysis result. The storage device 610 can store the above-mentioned neoantigen candidates.

[0103] The memory 620 can store data and information generated in the process of the analysis device 600 analyzing the activity of T cells.

[0104] The interface device 640 is a device to which certain commands and data are input from the outside. The interface device 640 can input genetic data of a patient from a physically connected input device or an external storage device.

[0105] Alternatively, the interface device 640 can receive the MHC amino acid sequence and / or the amino acid sequence of the antigen to be analyzed.

[0106] The interface unit 640 can receive a learning model for data analysis. The interface unit 640 can also receive learning data, information and parameter values ​​for training the learning model.

[0107] The interface unit 640 can receive the distance or the adjacent frequency of a specific amino acid pair in a protein structure from the protein structure DB.

[0108] The communication device 650 refers to a configuration that receives and transmits certain information via a wired or wireless network. The communication device 650 can receive genetic data from an external object. The communication device 650 can also receive data for model learning. The communication device 650 can receive an MHC amino acid sequence and / or an antigen amino acid sequence to be analyzed.

[0109] The communication device 650 can transmit the analysis result for the input sample to an external object. The analysis result may be the T cell activity for a specific peptide-MHC. Alternatively, the analysis result may be whether or not the peptide is a neoantigen candidate in the specific peptide-MHC.

[0110] The communication device 650 can receive the distance or the frequency of proximity of a specific amino acid pair in a protein structure from the protein structure DB.

[0111] The communication device 650 and the interface device 640 are devices to which certain data and commands are transmitted from the outside. The communication device 650 and the interface device 640 can be named input devices.

[0112] The output device 660 is a device that outputs certain information. The output device 660 can output an interface required for the data processing process, an analysis result, and the like.

[0113] The computing device 630 can identify a first amino acid sequence of the MHC and a second amino acid sequence of an antigen produced by a tumor cell from the genetic data. The computing device 630 can identify the first amino acid sequence and / or the second amino acid sequence from the genetic data using a specific program.

[0114] The arithmetic device 630 can generate a matrix for the first amino acid sequence and the second amino acid sequence by referring to the known protein structure information as described above. The arithmetic device 630 can calculate the distance and the proximity frequency of the specific amino acid pair to be evaluated by referring to the already known protein structure from the protein structure DB. The arithmetic device 630 can determine the interaction value based on the distance and the proximity frequency of the specific amino acid pair.

[0115] The computing device 630 can input the interaction matrix into the neural network model and predict the presence or absence of T cell activity against a specific peptide-MHC. The computing device 630 can predict the amount of IFNγ secreted by T cells against a specific peptide-MHC or the presence or absence of secretion. Furthermore, when T cell activity against a specific peptide-MHC is high, the computing device 630 can determine that peptide is a neoantigen candidate.

[0116] The computing device 630 may be a device such as a processor, AP, or a chip with an embedded program that processes data and performs certain operations.

[0117] The results of experiments verifying the effectiveness of the above-mentioned T cell activation method will now be described.

[0118] To validate the neural network model, the researchers performed ELISPOT analysis for EMT6. They selected mutations in the sample genes with a variant allele frequency (VAF) higher than 0.3. They selected the 25 peptides with the highest scores (candidate neoantigens) and the 5 peptides with the lowest scores (controls) based on the neural network prediction score.

[0119] The researchers performed ELISPOT analysis for each of the 25 and 5 peptides. The researchers performed ELISPOT analysis for H2-Dd / H2-Ld (class 1 alleles) and H2-IAd and H2-IEd (class 2 alleles). The researchers calculated the ELISPOT analysis results (ELISPOT.count) for all 30 peptides. In addition, an in silico model that measures the degree of peptide-MHC binding was used as a reference model. The reference model used was NetMHCIIpan.

[0120] Figure 7 shows an example of experimental results verifying the neural network model. Figure 7 shows the results of ELISPOT analysis for the 30 peptides mentioned above. In Figure 7, the peptides are arranged in order of increasing ELISPOT.count value. In other words, the peptides with higher T cell activation are shown as they move to the right in the graph of Figure 7. The lower part of Figure 7 shows the prediction results of the neural network model (shown as Target) and the reference model (shown as Reference). The white blocks indicate nonimmunogenic peptides, and the shaded blocks indicate antigenic peptides. The results of Figure 7 show that the reference model used in many conventional studies does not accurately predict T cell activation. In contrast, the neural network model mentioned above showed high accuracy overall, except for two peptides in the control group (nonimmunogenic). Therefore, it can be seen that the neural network model developed by the researchers showed significantly better performance than the widely used in silico model.

[0121] The data previously collected by the researchers is described in Table 1. The researchers selected a portion of the collected data as training data to train models for MHCI and MHCII separately. The researchers also used a portion of the data as validation data. The researchers selected 13,128 for MHCI and 6,650 for MHCII. The researchers divided the selected data into training data and validation data in a ratio of 7:3.

[0122] The accuracy of the neural network model was verified by comparing the results calculated by neural networks trained on human or mouse peptide-MHC pairs with experimentally known values.

[0123] Figure 8 shows another example of experimental results verifying the neural network model. The neural network model was trained with separate models for MHCI and MHCII. Figure 8(A) shows the experimental results for MHCI. As a result of the verification, the AUC (area under the curve) of the MHCI neural network model was 0.7787. Figure 8(B) shows the experimental results for MHCII. The AUC of the MHCII neural network model was 0.8083. Therefore, it can be said that the predictive accuracy of the developed neural network model is very high.

[0124] In addition, the above-mentioned method for predicting T cell activity or the method for discovering neoantigens may be embodied in a program (or application) including an executable algorithm that can be executed by a computer. The program may be provided by being stored in a non-transitory or non-transitory computer readable medium.

[0125] The non-transitory readable medium refers to a medium that stores data semi-permanently and can be read by a device, rather than a medium that stores data for a short moment such as a register, cache, memory, etc. Specifically, the various applications or programs described above may be stored and provided in a non-transitory readable medium such as a CD, DVD, hard disk, Blu-ray disk, USB, memory card, ROM (read-only memory), PROM (programmable read only memory), EPROM (Erasable PROM, EPROM), EEPROM (Electrically EPROM), flash memory, etc.

[0126] The temporary readable medium refers to various types of RAM such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synclink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).

[0127] The present embodiment and the drawings attached to this specification merely clearly show a portion of the technical ideas contained in the above-mentioned technology, and it is self-evident that any modified example and specific embodiment that can be easily inferred by a person skilled in the art within the scope of the technical ideas contained in the specification and drawings of the above-mentioned technology is included in the scope of the above-mentioned technology.

Claims

1. receiving genetic data of a patient into an analytical device; The analyzing device distinguishes a first amino acid sequence of a major histocompatibility complex (MHC) and a second amino acid sequence of an antigen produced by a tumor cell based on the genetic data; generating a matrix showing the interrelationship between the first amino acid sequence and the second amino acid sequence on an amino acid basis by the analytical device; inputting the matrix into a trained neural network model by the analysis device and determining whether or not T cells secrete a cytokine above a threshold level in response to the binding of the MHC and the antigen; Including, The neural network model is trained in advance using training data; the learning data includes input values ​​of MHC-neoantigen amino acid sequence pairs and label values ​​of amounts of cytokine secretion by T cells for each of the pairs; The neural network model outputs a cytokine secretion level of T cells for the input pair of the MHC and the antigen, The cytokine is interferon-γ. Method for predicting T cell activity against peptide-MHC.

2. 2. The method of claim 1, wherein the matrix includes proximity information of the amino acid pairs in an actual protein structure based on previously known protein structure information for each amino acid pair between the first amino acid sequence and the second amino acid sequence.

3. The method of claim 1, wherein the analysis device determines the antigen as a target candidate for an anti-cancer vaccine when the result output by the neural network model is cytokine secretion equal to or greater than the critical value.

4. The method for predicting T cell activity against peptide-MHC according to claim 1 , wherein the neural network model is a convolutional neural network (CNN).

5. an input device for inputting genetic data of a patient; a storage device for storing a neural network model for predicting the amount of cytokine secretion of T cells based on a matrix showing the correlation between the amino acid sequence of MHC (major histocompatibility complex) and the amino acid sequence of an antigen produced by a tumor cell; a computing device for identifying a first amino acid sequence of MHC and a second amino acid sequence of an antigen produced by a tumor cell from the genetic data, generating a matrix showing a correlation between the first amino acid sequence and the second amino acid sequence in units of one amino acid, inputting the generated matrix into the neural network model, and determining whether the MHC-antigen of the patient induces the secretion of interferon-gamma by T cells, The neural network model is trained in advance using training data; The learning data includes input values ​​of MHC-neoantigen amino acid sequence pairs and label values ​​of amounts of interferon gamma secreted by T cells for each of the pairs, The neural network model outputs the degree of interferon gamma secretion by T cells in response to the input pair of MHC and antigen. An assay device for predicting T cell activity against peptide-MHC.

6. 6. The analytical device for predicting T cell activity against peptide-MHC according to claim 5, wherein the matrix includes proximity information of the amino acid pairs in an actual protein structure based on previously known protein structural information for each amino acid pair between the first amino acid sequence and the second amino acid sequence.

7. The analytical device for predicting T cell activity against peptide-MHC according to claim 5, wherein the computing device determines the antigen as a target candidate for an anti-cancer vaccine when the result output by the neural network model is interferon gamma secretion above a critical value.

8. The analysis device for predicting T cell activity against peptide-MHC according to claim 5 , wherein the neural network model is a Convolutional Neural Network (CNN).

Citation Information

Patent Citations

  • Predicting immunogenicity of t cell epitopes

    KR1020160030101A

  • Reagents and methods for identifying, enriching, and / or expanding antigen-specific t cells

    KR1020170078619A

  • Method and Apparatus for Predicting a Binding Affinity between MHC and Peptide

    KR1020180052959A

  • Prediction method for binding preference between MHC and peptide on cancer cell and analysis apparatus

    KR102184720B1

  • Method and systems for prediction of HLA class ii-specific epitopes and characterization of CD4+ t cells

    US20200279616A1