TCR engineering using deep reinforcement learning to enhance efficacy and safety of TCR-T immunotherapy
TCRPPO optimizes TCRs by leveraging reinforcement learning to enhance binding to target peptides and minimize interactions with self-peptides, addressing inefficiencies in existing methods and improving the efficacy and safety of TCR-T therapy.
Patent Information
- Application Number
- JP2024536521
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-02-27
- Filing Date
- 2023-02-28
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2043-02-28
AI Technical Summary
Existing computational methods for optimizing T cell receptors (TCRs) for TCR-T therapy are time-consuming and inefficient, as they do not adequately consider the validity and specificity of generated sequences, leading to suboptimal recognition of target peptides and potential safety issues.
A novel reinforcement learning framework, TCRPPO, is employed to optimize TCRs by maximizing binding to target peptides while minimizing interaction with self-peptides, using a reward function that combines reconstruction-based and density estimation-based scores, and a buffering mechanism to handle hard-to-optimize sequences.
TCRPPO effectively generates TCRs with high validity scores and recognition probabilities, significantly outperforming baseline methods, enhancing the efficacy and safety of TCR-T immunotherapy.
Smart Images

Figure 0007681196000071 
Figure 0007681196000072 
Figure 0007681196000073
Abstract
Description
[Technical field]
[0001] The present invention relates to T cell receptors, and more particularly to T cell receptor (TCR) engineering using deep reinforcement learning to enhance the efficacy and safety of TCR-T immunotherapy. [Background technology]
[0002] 2. Description of Related Art T cells monitor cellular health by identifying foreign peptides displayed on their surface. The T cell receptor (TCR) is a protein complex on the surface of T cells that is capable of binding to these peptides. This process is called TCR recognition and constitutes a key step in the immune response. Optimizing TCR sequences for TCR recognition is a fundamental step towards the development of personalized therapeutics that trigger an immune response that kills cancer cells or virus-infected cells. Summary of the Invention
[0003] We present a method to implement deep reinforcement learning with a T cell receptor (TCR) mutation policy to generate binding TCRs that recognize target peptides for immunotherapy. The method includes: extracting peptides to identify viruses or tumor cells; collecting a library of TCRs from a target patient; predicting interaction scores between the extracted peptides and the TCRs from the target patient by a deep neural network; developing a deep reinforcement learning (DRL) framework with a TCR mutation policy to generate TCRs with maximum binding scores; defining a reward function based on a reconstruction-based score and a density estimation-based score; randomly sampling a batch of TCRs and mutating the TCRs according to a policy network; outputting the mutated TCRs; ranking the output TCRs and utilizing top-ranked TCR candidates targeting the virus or the tumor cells for immunotherapy; for each top-ranked TCR candidate, iteratively identifying a set of self-peptides that the top-ranked TCR candidate binds; and greedily optimizing the top-ranked TCR candidates by maximizing a sum of interaction scores with a set of given peptide antigens while minimizing a sum of interaction scores with a set of self-peptides until one or more stopping criteria of efficacy and safety are met.
[0004] A non-transitory computer readable storage medium is presented that includes a computer readable program for performing deep reinforcement learning with a T cell receptor (TCR) mutation policy to generate binding TCRs that recognize target peptides for immunotherapy. The computer readable program, when executed on a computer, causes the computer to perform the steps of: extracting peptides to identify viruses or tumor cells; collecting a library of TCRs from a target patient; predicting interaction scores between the extracted peptides and the TCRs from the target patient by a deep neural network; developing a deep reinforcement learning (DRL) framework with a TCR mutation policy to generate TCRs with the highest binding score; defining a reward function based on a reconstruction-based score and a density estimation-based score; randomly sampling a batch of TCRs and applying the policy neural network to the binding score; and generating a binding score of the TCRs with the highest binding score. the method includes the steps of: mutating the TCR according to a gene network; outputting the mutated TCR; ranking the output TCRs and utilizing the top-ranked TCR candidates that target the virus or the tumor cells for immunotherapy; and for each top-ranked TCR candidate, iteratively identifying a set of self-peptides bound by the top-ranked TCR candidate, and further greedily optimizing the top-ranked TCR candidates by maximizing a sum of interaction scores with a set of given peptide antigens while minimizing a sum of interaction scores with a set of self-peptides until one or more stopping criteria of efficacy and safety are met.
[0005] A system is presented that performs deep reinforcement learning with a T cell receptor (TCR) mutation policy to generate a binding TCR that recognizes a target peptide for immunotherapy, the system includes a memory and one or more processors in communication with the memory, the processor extracts peptides to identify viruses or tumor cells, collects a library of TCRs from a target patient, predicts an interaction score between the extracted peptides and the TCRs from the target patient by a deep neural network, develops a deep reinforcement learning (DRL) framework with a TCR mutation policy to generate a TCR with a maximum binding score, defines a reward function based on a reconstruction-based score and a density estimation-based score, randomly samples a batch of TCRs, and generates a policy network. The method is configured to mutate the TCR according to a set of hypotheses, output the mutated TCRs, rank the output TCRs, utilize top-ranked TCR candidates that target the virus or the tumor cell for immunotherapy, for each top-ranked TCR candidate, iteratively identify a set of self-peptides bound by the top-ranked TCR candidate, and further greedily optimize the top-ranked TCR candidates by maximizing a sum of interaction scores with a set of given peptide antigens while minimizing a sum of interaction scores with the set of self-peptides until one or more stopping criteria of efficacy and safety are met.
[0006] These and other features and advantages will become apparent from the following detailed description of illustrative embodiments thereof, which is to be read in connection with the accompanying drawings. [Brief description of the drawings]
[0007] The present disclosure provides details in the following description of preferred embodiments with reference to the following figures.
[0008] [Figure 1] FIG. 1 is a block / flow diagram of an exemplary model architecture for T Cell Receptor Proximal Policy Optimization (TCRPPO), in accordance with an embodiment of the present invention.
[0009] [Diagram 2] FIG. 1 is a block / flow diagram of an exemplary data flow for TCR PPO and T cell receptor autoencoder (TCR-AE) training, according to an embodiment of the present invention.
[0010] [Diagram 3] FIG. 1 is a block / flow diagram of an exemplary practical application for T cell receptor optimization using reinforcement learning and mutation policies for precision immunotherapy, in accordance with an embodiment of the present invention.
[0011] [Figure 4] FIG. 1 illustrates an exemplary processing system for optimization of T cell receptors using reinforcement learning and mutation policies for precision immunotherapy, in accordance with an embodiment of the present invention.
[0012] [Diagram 5] FIG. 1 is a block / flow diagram of an exemplary method for optimization of T cell receptors using reinforcement learning and mutation policies for precision immunotherapy, in accordance with an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] Immunotherapy is a fundamental treatment for human diseases that harnesses the human immune system to fight diseases. In the immune system, immune responses are triggered by cytotoxic T cells, which are activated by the binding of their T cell receptors (TCRs) to immunogenic peptides presented by major histocompatibility complex (MHC) proteins on the surface of infected or cancer cells. Recognition of these foreign peptides is determined by the interaction of the peptides with the TCR on the surface of T cells. This process is called TCR recognition and constitutes a key step in the immune response. Adoptive T cell immunotherapy (ACT), a promising cancer treatment, involves genetically modifying autologous T cells taken from a patient in the laboratory and then injecting the modified T cells into the patient's body to fight cancer.
[0014] TCR-T cell (TCR-T) therapy, which is one type of ACT therapy, can directly modify the TCR of T cells to increase binding affinity and effectively recognize and kill tumor cells. TCR is a heterodimeric protein with an α chain and a β chain. Each chain has three loops, CDR1, CDR2, and CDR3, as complementary determining regions (CDRs). CDR1 and CDR2 are mainly responsible for interacting with MHC, and CDR3 interacts with peptides. CDR3 of the β chain has many variations and is therefore thought to be mainly involved in recognizing foreign peptides. In an exemplary embodiment, the optimization of the CDR3 sequence of the β chain of the TCR is focused on to increase binding affinity to peptide antigens, and the optimization is performed by reinforcement learning. The success of the exemplary approach may guide the design of TCR-T therapies. For simplicity, when the exemplary method refers to TCR, it means the CDR3 of the β chain of the TCR.
[0015] Despite the great promise of TCR-T therapy, optimizing TCRs for therapeutic purposes remains a time-consuming process, typically requiring exhaustive screening of high-affinity TCRs either in vitro or in silico. To accelerate this process, computational methods have recently been developed that leverage experimental peptide-TCR binding data and TCR sequences to predict peptide-TCR interactions. However, these peptide-TCR binding prediction tools cannot immediately direct the rational design of new high-affinity TCRs. Existing computational methods for biological sequence design include search-based methods, generative methods, optimization-based methods, and reinforcement learning (RL)-based methods. However, all of these methods generate sequences without considering additional conditions such as peptides, and therefore cannot optimize TCRs tuned to recognize different peptides. In addition, these methods do not consider the validity of the generated sequences. This is important for TCR optimization, as an effective TCR should follow certain characteristics.
[0016] Exemplary embodiments present a novel reinforcement learning (RL) framework based on proximal policy optimization (PPO), called TCRPPO, to computationally optimize TCRs through mutation policies. In particular, TCRPPO learns a joint policy to optimize a customized TCR for a given peptide. In TCRPPO, a novel reward function is presented that measures both the likelihood that a mutant sequence is a valid TCR and the probability that the TCR recognizes the peptide.
[0017] In the reward design, besides maximizing the sum of the interaction scores between the mutated TCR under consideration and the set of given peptide antigens, the exemplary embodiment also minimizes the sum of the interaction scores between the mutated TCR under consideration and the set of self-peptides derived from normal human tissues that are most similar to the set of given peptide antigens. After TCR optimization, the exemplary embodiment identifies a set of self-peptides that the optimized TCR potentially binds. To ensure the safety of TCR-T therapy, the exemplary embodiment further mutates the TCR to maximize the sum of the interaction scores between the mutated TCR under consideration and the set of given peptide antigens while minimizing the sum of the interaction scores between the TCR and the set of identified self-peptides. These steps are repeated until convergence or a specified immunotherapy safety management criterion is met.
[0018] To measure the validity of TCRs, a TCR autoencoder called TCR-AE is developed, and the reconstruction error is utilized from the TCR-AE and the latent space distribution quantified by a Gaussian mixture model (GMM) to calculate a novel validity score. To measure peptide recognition, the exemplary method leveraged the state-of-the-art peptide-TCR binding prediction tool ERGO to predict peptide-TCR binding. TCRPPO is a flexible framework, and ERGO can be replaced by other binding predictors. In addition, a novel buffering mechanism called Buf-Opt is presented to correct difficult-to-optimize TCRs. Extensive experiments were performed with 7 million TCRs from TCRdb200 (Figure 2), 10 peptides from McPAS, and 15 peptides from VDJDB. Experimental results showed that TCRPPO generates qualified TCRs with high validity scores and high recognition probabilities, with an improvement of 58.2% and 26.8% for McPAS and VDJDB peptides, respectively, significantly outperforming the best baselines.
[0019] The recognition ability of a TCR sequence for a given peptide is given by s r The probability that a sequence is a valid TCR is measured by the recognition probability, denoted as s v A qualifying TCR is measured by a score expressed as s v >σ r And s v >σ c is defined as an array of σ r and σ c is a predefined threshold. The goal of TCRPPO is to mutate existing TCR sequences with low recognition probability for a given peptide into eligible ones. A peptide p or a TCR sequence c is represented as its sequence of amino acids. <o 1 ,o 2 ,...,o i ,...,o l >,o i is one of the 20 natural amino acids at position i in the sequence, and l is the length of the sequence. The TCR mutation process was formulated as a Markov decision process (MDP) M = {S, A, P, R} with the following components:
[0020] S: State space. Each state s∈S is a tuple of potential TCR sequences c and peptides p, i.e. s=(c,p). The subscript t (t=0,...,T) denotes the step of s, i.e. s t =(c t ,p). t Note that state s may not be a valid TCR. t is a qualified c t If t contains t or t reaches the maximum step limit T, then it is a terminal state, and s T Also, p is written as s 0 Note also that t is sampled at t and does not change over time t.
[0021] A: Action space, where:
number
number
[0022] P: State transition probability, where
number
number
number
[0023] R: Reward function in a state. In TCRPPO, state s t The intermediate rewards at (t=0,...,T-1) are all 0. T Only the final reward at is used as the optimization guide.
[0024] For a mutation policy network, TCRPPO mutates one amino acid in sequence c in a step to modify c into a qualified TCR. Specifically, in the first step t = 0, a peptide p is sampled as a target and a valid TCR c is generated. 0 is sampled and 0 =(c 0 ,p) is initialized. t =(c t ,p)(t>0), the mutation policy network of TCRPRO is t By mutating one amino acid in c, it is possible to obtain a TCR that binds to p. t+i The action of changing
number
[0025] For amino acid encoding, each amino acid o is represented by concatenating three vectors.
number
number
number
number
number
[0026] Regarding state embedding, s t =(c t ,p) is the associated sequence c t and p are embedded by embedding c t Each amino acid in i,t For example, the method is i,t and c t The context information in the hidden vector is then extracted using a one-layer bidirectional LSTM (Long Short Term Memory) as follows:
number
number
[0027] Where:
number
number
[0028]
number
number
[0029]
number
number
[0030]
number
number
number
[0031] The peptide sequence is then used to generate the hidden vector using another bidirectional LSTM in the same way.
number
[0032] Regarding motion prediction, at time t,
number
number
number
number
[0033] Where:
number
number
number
number
[0034] Given a predicted position i, TCRPPO is t In i It is necessary to predict the new amino acid that will replace the . TCRPPO calculates the probability that each amino acid type will be a new substitute as follows:
number
[0035] Here, U j (j=1,2,3) is a learnable matrix, and softmax(·) converts the 20-dimensional vector into a probability for the 20 types of amino acids. Next, the original amino acid type o i,t The type of converted amino acid is determined by sampling from the distribution excluding
[0036] Regarding the efficacy measurement of potential TCRs, a quantitative measurement is made of the likelihood that a given sequence c is a valid TCR (e.g., s v A novel scoring function (which calculates ) is presented, which is part of the reward of TCRPPO. Specifically, the exemplary method trained a novel autoencoder model, called TCR-AE, from only valid TCRs. The reconstruction accuracy of sequences in the TCR-AE was used to measure the validity of the TCR. Intuitively, since the TCR-AE is trained only from valid TCRs, its encoding-decoding process will follow the "rules" of true TCR sequences, and therefore it will not be able to successfully reproduce non-TCR sequences from the TCR-AE. However, if the TCR-AE learns common patterns common to TCRs and non-TCRs and cannot detect irregularities or the model complexity of the TCR-AE is high, it is possible that non-TCR sequences can obtain high reconstruction accuracy from the TCR-AE. To mitigate this, the exemplary method additionally evaluates the latent space in the TCR-AE using a Gaussian mixture model (GMM), hypothesizing that non-TCRs deviate from the dense regions of TCRs in the latent space.
[0037] TCR-AE150 presents an autoencoder TCR-AE, as shown in TCRPPO100 in Figure 1. TCR-AE150 uses bidirectional LSTM to convert an input sequence c into
number
number
number
number
[0038] This is done by decoding the sequence via the decoder 140.
number
number
number
[0039] Where:
number
number
number
number
number
number
[0040] Note that TCR-AE150 is trained end-to-end from TCR, independent of TCRPPO100. Supervisory constraints are applied during training to ensure that the decoded sequences have the same length as the input sequences, and therefore cross-entropy loss is applied to optimize TCR-AE150. As a standalone module, TCR-AE150 achieves a score s v The input sequence c to TCR-AE150 is encoded using only the BLOSUM matrix, as empirically BLOSUM encoding has been shown to yield good reconstruction performance and fast convergence compared to other encoding combinations.
[0041] Using the fully trained TCR-AE150, the TCR efficacy score based on rearrangement of sequence c was calculated as follows:
number
[0042] where TCR-AE(c) represents the sequence of c reconstructed from TCR-AE, lev(c, TCR-AE(c)) is the Levenshtein distance, which is an index based on the edit distance between c and TCR-AE(c), and l c is the length of c. rA higher c indicates a higher probability that c is a valid TCR. Note that when TCR-AE150 is used in testing, the length of the reconstructed sequence may not be the same as the input c, because TCR-AE150 cannot accurately predict the end of the sequence and the reconstructed sequence may be too short or too long. Therefore, the Levenshtein distance is the length of the input sequence l. c If the distance is greater than the length of the sequence, then r r Note that (c) can be negative. Negative values have no effect on the use of the score (e.g., a negative r r (c) shows that TCR-AE(c) and c are very different).
[0043] To better distinguish between valid and invalid TCRs, the TCRPPO100 uses GMM145
number
[0044] For a given sequence c, TCR PPO100 computes the likelihood score of c falling within the Gaussian mixture region of the training TCRs as follows:
number
[0045] Where:
number
number
[0046] Combining rearrangement-based scoring and density estimation-based scoring, we developed a new scoring method to measure TCR validity as follows.
number
[0047] This method is used to evaluate whether a sequence is likely to be a valid TCR, which is used in the reward function.
[0048] With respect to TCRPPO learning and with respect to the final reward, an exemplary method is r Score and s v Based on the scores, the final rewards for TCRPPO100 are defined as follows:
number
[0049] Here, s r (c T ,p) is the predicted recognition probability by ERGO160, and σ c is T is very likely to be a valid TCR (σ c = 1.2577), and α is s r ands v is a hyperparameter used to control the tradeoff between
[0050] With respect to policy learning, an exemplary method employs proximal policy optimization (PPO) to optimize the policy network.
[0051] The objective function of PPO is defined as follows:
number
number
[0052] Where:
number
number
number
number
number
[0053]
number
number
[0054] where γ∈(0,1) is the discount factor that determines the importance of future rewards, and δ t =r t +γV(s t+1 )-V(s t) is V(s t ) is the time difference error, which is the value function, and λ∈(0,1) is V(s t ) is a parameter to balance the bias and variance of the peptide embedding using a multilayer perceptron (MLP).
number
number
[0055] The objective function for V(·) is:
number
[0056] Where:
number
number
number
number
[0057] The final objective function of TCRPPO100 is defined as follows:
number
[0058] Here, α 1 and α 2 are two hyperparameters that control the tradeoff between the PPO objective, the value function, and the entropy regularization term.
[0059] TCRPPO100 implements a novel buffering and reoptimization mechanism called Buf-Opt to deal with hard-to-optimize TCRs and generalize the optimization ability to a more diverse range of TCRs. This mechanism has a buffer to store TCRs that cannot be optimized to be eligible. These hard sequences are sampled from the buffer again according to the following probability distribution and further optimized by TCRPPO100:
number
[0060] In Equation 10, S is the final reward R(c T Based on the ξ(c,p), we measure how hard it is to optimize c for p, ξ is a hyperparameter (e.g., ξ=5), and Σ converts S(c,p) as a probability. By sampling and reoptimizing, TCRPPO100 is trained to learn from hard sequences, and it is hoped that hard sequences will have a chance to be better optimized by TCRPPO100. If a hard sequence is still not optimized to be eligible, it is returned to the buffer with a 50% probability. When the buffer becomes full (size 2,000 in our experiments), the sequence that was allocated to the buffer the earliest is removed. We call TCRPPO+b TCRPPO100 plus Buf-Opt.
[0061] In conclusion, an exemplary embodiment of the present invention formulates the search for optimized TCRs as a RL problem and presents a framework, TCRPPO, with a mutation policy using proximal policy optimization (PPO). TCRPPO mutates TCRs into valid ones that can recognize a given peptide. TCRPPO utilizes a reward function that combines the likelihood that a mutated sequence is a valid TCR, measured by a novel scoring function based on deep autoencoders, and the probability that a mutated sequence recognizes the peptide, obtained from a peptide-TCR interaction predictor. Comparison of TCRPPO with multiple baseline methods shows that TCRPPO significantly outperforms all baseline methods in generating positive binding and valid TCRs. These results demonstrate the potential of TCRPPO in both precision immunotherapy and peptide discovery that recognizes TCR motifs.
[0062] The exemplary method further presents a deep reinforcement learning system with a TCR mutation policy to generate binding TCRs that recognize target peptides. The library of predefined peptides can be obtained from the genome of a virus such as SARS-CoV-2 or from sequencing patient tumor samples. Thus, the presented exemplary system can be used for immunotherapy to target specific types of viruses or tumors using TCR engineering.
[0063] Given a viral genome or tumor cells, the exemplary method performs sequencing followed by a commercially available peptide processing pipeline to extract peptides that can uniquely identify the virus or tumor cells. The exemplary method also collects a library of TCRs from the target patient. Using this library of peptides from the virus or tumor and targeting the given TCR, the system can generate optimized or mutated TCRs that can trigger an immune response to kill the virus or tumor cells.
[0064] The exemplary method first trains a deep neural network on the publicly available IEDB, VDJdb, and McPAS-TCR datasets or downloads a pre-trained model such as ERGO to predict binding interactions between peptides and TCRs. Based on this pre-trained peptide-TCR interaction score prediction model, the exemplary method evolves a DRL system with a TCR mutation policy to generate TCRs with high binding scores that are the same as or differ at most by d amino acids from the library of provided TCRs. Specifically, using a pre-trained predictive deep model to define a reward function and starting from a random or existing TCR, the exemplary method then pre-trains a DRL system to learn a good TCR mutation policy that transforms the given random TCR into a peptide-recognizing TCR with high binding interaction score. Based on this trained DRL system with a pre-trained TCR mutation policy, the exemplary method randomly samples a batch of TCRs from the provided library and mutates the TCRs according to the policy network. During the mutation process, if the mutated TCR already differs from the starting TCR by d amino acids, the process ends and the TCR is output as the final TCR. The final mutant TCRs that recognize the given peptide are output, and the collection of edited mutant TCRs is ranked, with the top ranked ones being used as promising artificial TCRs for immunotherapy to target specific viruses or tumor cells.
[0065] In the reward design, besides maximizing the sum of the interaction scores between the mutated TCR under consideration and the set of given peptide antigens, the exemplary embodiment also minimizes the sum of the interaction scores between the mutated TCR under consideration and the set of self-peptides derived from normal human tissues that are most similar to the set of given peptide antigens. After TCR optimization, the exemplary embodiment identifies a set of self-peptides that the optimized TCR potentially binds. To ensure the safety of TCR-T therapy, the exemplary embodiment further mutates the TCR to maximize the sum of the interaction scores between the mutated TCR under consideration and the set of given peptide antigens while minimizing the sum of the interaction scores between the TCR and the set of identified self-peptides. These steps are repeated until convergence or a specified immunotherapy safety management criterion is met.
[0066] FIG. 3 is an exemplary practical application for optimization of T cell receptors using reinforcement learning and mutation policies for precision immunotherapy, in accordance with an embodiment of the present invention.
[0067] In one example 300, peptides are processed by the TCRPPO 100 in the peptide mutation environment 110 by the mutation policy network 120 to generate new qualified peptides 310 that are displayed on a screen 312 and analyzed by a user 314. For every peptide selected from the same database (e.g., 10 peptides from McPAS and 15 peptides from VDJDB), the exemplary method trains one TCRPPO agent, which optimizes the training sequences (e.g., the 7,281,105 TCRs in FIG. 2 ) to be qualified against one of the selected peptides. The ERGO model trained on the corresponding database determines the recognition probability s of the TCRPPO agent. r It should be noted that one ERGO model is trained on all peptides in each database (e.g., one ERGO predicts TCR-peptide binding for multiple peptides). Thus, the ERGO model is trained on multiple peptides in the exemplary setting.r It is also noted that in the exemplary method, we trained one TCRPRO corresponding to each database, because the peptides and TCRs in these two databases are very different, and ERGO training the two databases together proved to perform poorly.
[0068] TCRPPO mutates each sequence up to 8 steps (T = 8), which is large enough since the most common length of TCR is 15. In TCRPPO training (Figure 2), the initial TCR sequence (e.g., s 0 c 0 ) is S trn are randomly sampled from and mutated in the following way: 0 are randomly sampled at and remain the same under the following conditions (e.g., s t =(c t ,p)). TCRPPO100 is S trn After being thoroughly trained by tst will be tested.
[0069] Experimental results comparing with generation-based and mutation-based methods for TCR optimization show that TCRPPO100 significantly outperforms baseline methods. Analysis of the TCRs generated by TCRPPO100 shows that TCRPPO100 can successfully learn the conserved patterns of TCRs. Experimental comparison of the generated TCRs with existing TCRs shows that TCRPPO100 can generate TCRs similar to existing human TCRs, which can be used for further medical evaluation and research. The comparative results of TCR detection are shown in the s v The score indicates that the non-TCR sequences can be detected very effectively. v Analysis of the distribution of scores indicates that TCRPPO100 mutates sequences along a trajectory not far from the effective TCR.
[0070] FIG. 4 is an exemplary processing system for optimization of T cell receptors using reinforcement learning and mutation policies for precision immunotherapy, in accordance with an embodiment of the present invention.
[0071] The processing system includes at least one processor (CPU) 404 operatively coupled to other components via a system bus 402. A graphical processing unit (GPU) 405, a cache 406, a read-only memory (ROM) 408, a random access memory (RAM) 410, an input / output (I / O) adapter 420, a network adapter 430, a user interface adapter 440, and a display adapter 450 are operatively coupled to the system bus 402. Additionally, the TCRPPO 100 is employed within the peptide mutation environment 110 by using the mutation policy network 120.
[0072] The storage devices 422 are operably coupled to the system bus 402 by the I / O adapter 420. The storage devices 422 may be any type of disk storage device (e.g., magnetic or optical disk storage device), solid-state magnetic device, or the like.
[0073] The transceiver 432 is operably coupled to the system bus 402 by the network adapter 430 .
[0074] User input device(s) 442 are operatively coupled to system bus 402 by user interface adapter 440. User input device 442 may be any of a keyboard, mouse, keypad, image capture device, motion sensing device, microphone, a device incorporating the functionality of at least two of the aforementioned devices, and the like. Of course, other types of input devices may be used while maintaining the spirit of the invention. User input device 442 may be the same type of user input device or different types of user input devices. User input device 442 is used to input and output information to and from the processing system.
[0075] A display device 452 is operatively coupled to the system bus 402 by a display adapter 450 .
[0076] Of course, the processing system may include other elements (not shown) or omit certain elements, as would be readily contemplated by one of ordinary skill in the art. For example, various other input and / or output devices may be included in the system, depending on the particular implementation, as would be readily understood by one of ordinary skill in the art. For example, various types of wireless and / or wired input and / or output devices may be used. Furthermore, various configurations of additional processors, controllers, memory, etc. may also be utilized, as would be readily understood by one of ordinary skill in the art. These and other variations of the processing system are readily contemplated by one of ordinary skill in the art given the teachings of the invention provided herein.
[0077] FIG. 5 is a block / flow diagram of an exemplary method for optimizing T cell receptors using reinforcement learning and mutation policies for precision immunotherapy, in accordance with an embodiment of the present invention.
[0078] In block 501, peptides are extracted to identify viruses or tumor cells.
[0079] In block 503, a library of TCRs is collected from a target patient.
[0080] In block 505, a deep neural network predicts interaction scores between the extracted peptides and TCRs from the target patient.
[0081] In block 507, a deep reinforcement learning (DRL) framework with a TCR mutation policy is developed to generate a TCR with a maximum combined score.
[0082] At block 509, a reward function is defined based on the reconstruction-based score and the density estimation-based score.
[0083] In block 511, a batch of TCRs is randomly sampled and the TCRs are mutated according to the policy network.
[0084] At block 513, the mutated TCR is output.
[0085] In block 515, the output TCRs are ranked and the top ranked TCR candidates are utilized to target viruses or tumor cells for immunotherapy.
[0086] In block 517, for each top-ranked TCR candidate, the set of self-peptides to which the top-ranked TCR candidate binds is iteratively identified, and the top-ranked TCR candidate is further optimized by maximizing the sum of its interaction scores with a given set of peptide antigens while minimizing the sum of its interaction scores with the set of self-peptides until one or more stopping criteria of efficacy and safety are met.
[0087] In conclusion, the exemplary method proposes a DRL system with a TCR mutation policy to create a binding TCR that recognizes a given peptide antigen. The presented system can be used for generating TCRs for immunotherapy targeting a specific type of virus or tumor. The reward design is based on the TCR intradistribution score and the binding interaction score. The exemplary method optimizes the DRL model using PPO, outputs the final mutated TCR, and ranks the compiled set of mutated TCRs. The top ranked ones are used for immunotherapy as promising candidates to target a specific virus or tumor. For each top ranked TCR, a set of self-peptides that the optimized TCR may bind is identified. To ensure the safety of TCR-T therapy, the TCR is mutated to maximize the sum of the interaction scores between the mutated TCR under consideration and the set of given peptide antigens while minimizing the sum of the interaction scores between the TCR and the set of identified self-peptides (such steps are repeated until convergence or some stopping criterion is met, outputting the final set of optimized TCRs).
[0088] As used herein, "data," "content," "information," and similar terms may be used interchangeably to refer to data that may be captured, transmitted, received, displayed, and / or stored in accordance with various exemplary embodiments. Thus, use of such terms should not be construed as limiting the spirit and scope of the present disclosure. Additionally, where a computing device is described herein as receiving data from another computing device, the data may be received directly from the other computing device or may be received indirectly via one or more intermediary computing devices, such as, for example, one or more servers, relays, routers, network access points, base stations, and / or the like.
[0089] As will be appreciated by those skilled in the art, aspects of the present invention may be embodied as a system, method, or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be referred to generally herein as a "circuit," "module," "computer," "apparatus," or "system." Additionally, aspects of the present invention may take the form of a computer program product embodied in one or more computer readable medium(s) having computer readable program code embodied thereon.
[0090] Any combination of one or more computer readable media may be utilized. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer readable storage media include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable CD-ROM, an optical data storage device, a magnetic data storage device, or any suitable combination of the foregoing. In this specification, a computer readable storage medium is any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0091] A computer-readable signal medium may include a propagated data signal having computer-readable program code embodied therein, for example as part of a baseband or carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electromagnetic waves, optical, or any suitable combination thereof. A computer-readable signal medium is not a computer-readable storage medium, but may be any computer-readable medium that can communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0092] The program code embodied in the computer readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, etc., or any suitable combination of the foregoing.
[0093] Computer program code for carrying out operations for aspects of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language. The program code may run entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a LAN or WAN, or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0094] Aspects of the present invention are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, executed via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram blocks or modules.
[0095] These computer program instructions may also be stored on a computer-readable medium that can instruct a computer, other programmable data processing apparatus, or other device to function in a particular manner to produce an article of manufacture, where the instructions stored on the computer-readable medium include instructions that implement the functions / operations specified in the blocks or modules of the flowcharts and / or block diagrams.
[0096] Computer program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause the computer, other programmable apparatus, or other device to perform a series of operational steps, creating a computer-implemented process, such that the instructions executing on the computer or other programmable apparatus provide a process for implementing the functions / operations specified in the flowchart and / or block diagram blocks or modules.
[0097] It should be understood that the term "processor" as used herein is intended to include any processing device, such as, for example, one that includes a CPU and / or other processing circuitry, and that the term "processor" may refer to more than one processing device, and that various elements associated with a processing device may be shared by other processing devices.
[0098] The term "memory" as used herein is intended to include memory associated with a processor or CPU, such as, for example, RAM, ROM, fixed memory devices (e.g., hard drives), removable memory devices (e.g., diskettes), flash memory, etc. Such memory is considered a computer-readable storage medium.
[0099] Additionally, the phrase "input / output device" or "I / O device" as used herein is intended to include, for example, one or more input devices (e.g., a keyboard, mouse, scanner, etc.) for inputting data to a processing unit, and / or one or more output devices (e.g., speakers, displays, printers, etc.) for presenting results related to a processing unit.
[0100] The foregoing is understood in all respects to be illustrative and illustrative, but not restrictive, and the scope of the invention disclosed herein is to be determined not from the detailed description, but from the claims which are interpreted in accordance with the full breadth permitted by the patent laws. It will be understood that the embodiments shown and described herein are merely illustrative of the principles of the invention, and that various modifications can be made by those skilled in the art without departing from the scope and spirit of the invention. Various other feature combinations can be implemented by those skilled in the art without departing from the scope and spirit of the invention. Having thus described aspects of the invention with the detail and particularity required by the patent laws, what is desired to be claimed and protected by Letters Patent is set forth in the appended claims.
Claims
1. 1. A method of performing deep reinforcement learning with a T cell receptor (TCR) mutation policy to generate a binding TCR that recognizes a target peptide for immunotherapy, comprising: Extracting peptides to identify viruses or tumor cells; collecting a library of TCRs from a target patient; predicting an interaction score between the extracted peptides and the TCR from the target patient by a deep neural network; Developing a deep reinforcement learning (DRL) framework with a TCR mutation policy to generate a TCR that maximizes the combined score; defining a reward function based on the reconstruction-based score and the density estimation-based score; randomly sampling a batch of TCRs and mutating said TCRs according to a policy network; outputting a mutated TCR; ranking the output TCRs and utilizing the top ranked TCR candidates to target the virus or the tumor cells for immunotherapy; and iteratively identifying, for each top-ranked TCR candidate, a set of self-peptides bound by said top-ranked TCR candidate, and further greedily optimizing said top-ranked TCR candidates by maximizing a sum of interaction scores with a given set of peptide antigens while minimizing a sum of interaction scores with the set of self-peptides until one or more stopping criteria of efficacy and safety are met.
2. 2. The method of claim 1, wherein the reward function measures both the likelihood that a mutant sequence is a valid TCR and the probability that the TCR recognizes the peptide.
3. 3. The method of claim 2, wherein measuring the likelihood that the mutant sequence is a valid TCR is made possible by a TCR autoencoder (TCR-AE) trained solely on TCRs.
4. The method of claim 3 , wherein density estimates over the latent space within the TCR-AE are estimated using a Gaussian mixture model (GMM).
5. 4. The method of claim 3, wherein the TCR-AE uses a bidirectional long short-term memory (LSTM) to encode an input sequence into a hidden vector by concatenating the last hidden vectors from two LSTM directions.
6. The method of claim 1 , wherein a buffering and re-optimization framework including a buffer is employed to handle difficult-to-optimize TCRs and generalize the optimization capability to more diverse TCRs.
7. The method of claim 1 , wherein the TCRs and the extracted peptides are encoded into a distributed embedding space by a TCR-AE, and a mapping between the embedding space and the TCR mutation policy is learned.
8. 1. A non-transitory computer readable storage medium comprising a computer readable program for performing deep reinforcement learning with a T cell receptor (TCR) mutation policy to generate a binding TCR that recognizes a target peptide for immunotherapy, the computer readable program, when executed on a computer, causes the computer to: Extracting peptides to identify viruses or tumor cells; collecting a library of TCRs from a target patient; predicting an interaction score between the extracted peptides and the TCR from the target patient by a deep neural network; Developing a deep reinforcement learning (DRL) framework with a TCR mutation policy to generate a TCR that maximizes the combined score; defining a reward function based on the reconstruction-based score and the density estimation-based score; randomly sampling a batch of TCRs and mutating the TCRs according to a policy network; outputting the mutated TCR; ranking the output TCRs and utilizing the top ranked TCR candidates to target the virus or the tumor cells for immunotherapy; and for each top-ranked TCR candidate, iteratively identifying a set of self-peptides to which the top-ranked TCR candidate binds, and further greedily optimizing the top-ranked TCR candidate by maximizing a sum of interaction scores with a given set of peptide antigens while minimizing a sum of interaction scores with the set of self-peptides until one or more stopping criteria of efficacy and safety are met.
9. 9. The non-transitory computer-readable storage medium of claim 8, wherein the reward function measures both the likelihood that a variant sequence is a valid TCR and the probability that the TCR recognizes a peptide.
10. 10. The non-transitory computer readable storage medium of claim 9, wherein measuring the likelihood that the variant sequence is a valid TCR is enabled by a TCR autoencoder (TCR-AE) trained solely by TCRs.
11. 11. The non-transitory computer-readable storage medium of claim 10, wherein density estimation over a latent space in the TCR-AE is evaluated using a Gaussian mixture model (GMM).
12. 11. The non-transitory computer-readable storage medium of claim 10, wherein the TCR-AE uses a bidirectional long short-term memory (LSTM) to encode an input sequence into a hidden vector by concatenating the last hidden vectors from two LSTM directions.
13. 9. The non-transitory computer-readable storage medium of claim 8, wherein a buffering and re-optimization framework including a buffer is employed to handle difficult to optimize TCRs and generalize the optimization capabilities to more diverse TCRs.
14. 9. The non-transitory computer-readable storage medium of claim 8, wherein the TCRs and the extracted peptides are encoded into a distributed embedding space by a TCR-AE, and a mapping between the embedding space and the TCR mutation policy is learned.
15. 1. A system for performing deep reinforcement learning with a T cell receptor (TCR) mutation policy to generate a binding TCR that recognizes a target peptide for immunotherapy, comprising: Memory, and one or more processors in communication with the memory, the processors comprising: Extract peptides to identify viruses or tumor cells collecting a library of TCRs from a target patient; predicting an interaction score between the extracted peptides and the TCR from the target patient by a deep neural network; Develop a deep reinforcement learning (DRL) framework with a TCR mutation policy to generate a TCR that maximizes the combined score; Define a reward function based on the reconstruction-based score and the density estimation-based score; Randomly sampling a batch of TCRs and mutating said TCRs according to a policy network; Outputting mutated TCR, ranking the output TCRs and utilizing the top ranked TCR candidates to target the virus or the tumor cells for immunotherapy; The system is configured to iteratively identify, for each top-ranked TCR candidate, a set of self-peptides to which the top-ranked TCR candidate binds, and further greedily optimize the top-ranked TCR candidate by maximizing a sum of interaction scores with a given set of peptide antigens while minimizing a sum of interaction scores with the set of self-peptides until one or more stopping criteria of efficacy and safety are met.
16. 16. The system of claim 15, wherein the reward function measures both the likelihood that a variant sequence is a valid TCR and the probability that the TCR recognizes the peptide.
17. The system of claim 16, wherein measuring the likelihood that the variant sequence is a valid TCR is enabled by a TCR autoencoder (TCR-AE) trained solely on TCRs.
18. The system of claim 17, wherein density estimates over the latent space within the TCR-AE are evaluated using a Gaussian mixture model (GMM).
19. 18. The system of claim 17, wherein the TCR-AE uses a bidirectional long short-term memory (LSTM) to encode an input sequence into a hidden vector by concatenating the last hidden vectors from two LSTM directions.
20. 16. The system of claim 15, wherein a buffering and re-optimization framework including a buffer is employed to handle difficult to optimize TCRs and generalize the optimization capabilities to more diverse TCRs.
Citation Information
Patent Citations
Information processing method and learning model
JP2020077206A
Machine Learning Algorithm for Identifying Peptides that Contain Features Positively Associated with Natural Endogenous or Exogenous Cellular Processing, Transportation and Histocompatibility Complex (MHC) Presentation
US20190311781A1
Peptide-based vaccine generation system
US20210319847A1
Methods for identifying and using disease-associated antigens
WO2020243469A1