Optimizing T-cell receptors using reinforcement learning and mutation policies for precision immunotherapy

TCRPPO optimizes TCRs using a reinforcement learning framework with a novel reward function and autoencoder to generate TCRs with high binding affinity and validity, addressing inefficiencies in existing methods and enhancing immunotherapy efficacy.

JP7785170B2Active Publication Date: 2025-12-12NEC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024523718
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-01-09
Filing Date
2023-01-11
Publication Date
2025-12-12
Estimated Expiration
2043-01-11

AI Technical Summary

Technical Problem

Existing computational methods for optimizing T cell receptors (TCRs) are time-consuming and inefficient, as they do not effectively tailor TCRs to recognize specific peptides and lack consideration for sequence validity, hindering the development of personalized immunotherapies.

Method used

A reinforcement learning framework, TCRPPO, is developed to optimize TCRs using a novel reward function that combines reconstruction-based and density estimation-based scores, along with a TCR autoencoder and a buffering mechanism, to generate TCRs with high binding affinity and validity.

Benefits of technology

TCRPPO significantly improves the generation of qualified TCRs with high recognition probabilities, outperforming baseline methods by 58.2% and 26.8% for McPAS and VDJDB peptides, respectively, demonstrating its effectiveness in precision immunotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007785170000071
    Figure 0007785170000071
  • Figure 0007785170000072
    Figure 0007785170000072
  • Figure 0007785170000073
    Figure 0007785170000073
Patent Text Reader

Abstract

A method is presented to implement deep reinforcement learning with a T cell receptor (TCR) mutation policy to generate binding TCRs that recognize target peptides for immunotherapy. The method includes extracting peptides to identify viruses or tumor cells, collecting a library of TCRs from target patients, predicting interaction scores between the extracted peptides and TCRs from the target patients by a deep neural network, developing a deep reinforcement learning (DRL) framework with the TCR mutation policy to generate TCRs with the highest binding score, defining a reward function based on a reconstruction-based score and a density estimation-based score, randomly sampling a batch of TCRs, mutating the TCRs according to the policy network, outputting the mutated TCRs, and ranking the output TCRs to utilize the top-ranked TCR candidates that target viruses or tumor cells for immunotherapy.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to T cell receptors, and more particularly to optimizing T cell receptors using reinforcement learning and mutation policies for precision immunotherapy. [Background technology]

[0002] 2. Description of Related Art T cells monitor cellular health by recognizing foreign peptides displayed on their surface. The T cell receptor (TCR) is a protein complex on the surface of T cells that can bind to these peptides. This process, called TCR recognition, constitutes a critical step in the immune response. Optimizing TCR sequences for TCR recognition is a fundamental step toward developing personalized therapies that trigger immune responses that kill cancer cells or virus-infected cells. Summary of the Invention

[0003] A method for generating binding TCRs that recognize target peptides for immunotherapy using deep reinforcement learning with a TCR mutation policy is presented. The method includes extracting peptides to identify viruses or tumor cells, collecting a library of TCRs from target patients, predicting interaction scores between the extracted peptides and the TCRs from the target patients using a deep neural network, developing a deep reinforcement learning (DRL) framework with the TCR mutation policy to generate TCRs with the highest binding scores, defining a reward function based on a reconstruction-based score and a density estimation-based score, randomly sampling a batch of TCRs and mutating the TCRs according to a policy network, outputting the mutated TCRs, and ranking the output TCRs to select the top-ranked TCR candidates that target the virus or tumor cells for immunotherapy.

[0004] A non-transitory computer-readable storage medium is provided, comprising a computer-readable program for performing deep reinforcement learning using a T cell receptor (TCR) mutation policy to generate binding TCRs that recognize target peptides for immunotherapy. When executed on a computer, the computer-readable program causes the computer to perform the following steps: extracting peptides to identify viruses or tumor cells, collecting a library of TCRs from target patients, predicting interaction scores between the extracted peptides and the TCRs from the target patients using a deep neural network, developing a deep reinforcement learning (DRL) framework using the TCR mutation policy to generate TCRs that maximize binding scores, defining a reward function based on a reconstruction-based score and a density estimation-based score, randomly sampling a batch of TCRs and mutating the TCRs according to a policy network, outputting the mutated TCRs, and ranking the output TCRs and utilizing the top-ranked TCR candidates that target the virus or the tumor cells for immunotherapy.

[0005] A system for performing deep reinforcement learning using a T cell receptor (TCR) mutation policy to generate binding TCRs that recognize target peptides for immunotherapy is presented. The system includes a memory and one or more processors in communication with the memory, the processors being configured to: extract peptides to identify viruses or tumor cells, collect a library of TCRs from a target patient, predict interaction scores between the extracted peptides and the TCRs from the target patient using a deep neural network, develop a deep reinforcement learning (DRL) framework using the TCR mutation policy to generate TCRs with the highest binding score, define a reward function based on a reconstruction-based score and a density estimation-based score, randomly sample a batch of TCRs, mutate the TCRs according to a policy network, output the mutated TCRs, rank the output TCRs, and utilize the top-ranked TCR candidates that target the virus or the tumor cells for immunotherapy.

[0006] These and other features and advantages will become apparent from the following detailed description of illustrative embodiments thereof, which is to be read in connection with the accompanying drawings. [Brief explanation of the drawings]

[0007] The present disclosure provides details in the following description of preferred embodiments with reference to the following figures.

[0008] [Figure 1] FIG. 1 is a block / flow diagram of an exemplary model architecture for T-cell receptor proximal policy optimization (TCRPPO), according to an embodiment of the present invention.

[0009] [Figure 2] FIG. 1 is a block / flow diagram of an exemplary data flow for TCRPPO and T-cell receptor autoencoder (TCR-AE) training, according to an embodiment of the present invention.

[0010] [Figure 3]FIG. 1 is a block / flow diagram of an exemplary practical application for T cell receptor optimization using reinforcement learning and mutation policies for precision immunotherapy, according to an embodiment of the present invention.

[0011] [Figure 4] FIG. 1 illustrates an exemplary processing system for optimization of T cell receptors using reinforcement learning and mutation policies for precision immunotherapy, according to an embodiment of the present invention.

[0012] [Figure 5] FIG. 1 is a block / flow diagram of an exemplary method for T cell receptor optimization using reinforcement learning and mutation policies for precision immunotherapy, according to an embodiment of the present invention.

[0013] [Figure 6] FIG. 1 is a block / flow diagram of an exemplary method for T cell receptor optimization using reinforcement learning and mutation policies for precision immunotherapy, according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0014] Immunotherapy is a fundamental treatment for human diseases that harnesses the human immune system to fight disease. In the immune system, immune responses are triggered by cytotoxic T cells, which are activated by the binding of T cell receptors (TCRs) to immunogenic peptides presented by major histocompatibility complex (MHC) proteins on the surface of infected or cancer cells. Recognition of these foreign peptides is determined by the interaction of the peptides with the TCR on the surface of T cells. This process, called TCR recognition, constitutes a critical step in the immune response. Adoptive T cell immunotherapy (ACT), a promising cancer treatment, involves genetically modifying autologous T cells collected from patients in the laboratory and then injecting the modified T cells into the patient's body to fight cancer.

[0015] TCR-T cell (TCR-T) therapy, a type of ACT therapy, directly modifies the TCR of T cells to increase binding affinity and effectively recognize and kill tumor cells. TCRs are heterodimeric proteins with an α chain and a β chain. Each chain contains three loops as complementarity-determining regions (CDRs): CDR1, CDR2, and CDR3. CDR1 and CDR2 are primarily responsible for interactions with MHC, while CDR3 interacts with peptides. The CDR3 of the β chain exhibits high variation and is therefore thought to be primarily responsible for recognizing foreign peptides. In an exemplary embodiment, we focus on optimizing the CDR3 sequence of the TCR β chain to increase binding affinity to peptide antigens, and this optimization is performed using reinforcement learning. The success of the exemplary approach may serve as a guide for the design of TCR-T therapies. For simplicity, when the exemplary method refers to a TCR, it refers to the CDR3 of the TCR β chain.

[0016] Despite the great promise of TCR-T therapy, optimizing TCRs for therapeutic purposes remains a time-consuming process, typically requiring exhaustive screening of high-affinity TCRs either in vitro or in silico. To accelerate this process, computational methods have recently been developed that utilize experimental peptide-TCR binding data and TCR sequences to predict peptide-TCR interactions. However, these peptide-TCR binding prediction tools cannot immediately direct the rational design of novel high-affinity TCRs. Existing computational methods for biological sequence design include search-based, generative, optimization-based, and reinforcement learning (RL)-based methods. However, all of these methods generate sequences without considering additional conditions such as the peptide, making them unable to optimize TCRs tailored to recognize different peptides. Additionally, these methods do not consider the validity of the generated sequences, which is important for TCR optimization because effective TCRs should conform to certain characteristics.

[0017] The exemplary embodiment presents a novel reinforcement learning (RL) framework, called TCRPPO, based on proximal policy optimization (PPO), for computationally optimizing TCRs through mutation policies. Specifically, TCRPPO learns a joint policy for optimizing a customized TCR for a given peptide. In TCRPPO, a novel reward function is presented that measures both the likelihood that a mutated sequence is a valid TCR and the probability that the TCR recognizes the peptide. To measure the validity of a TCR, a TCR autoencoder called TCR-AE is developed. A novel validity score is calculated using the reconstruction error from the TCR-AE and the latent space distribution quantified by a Gaussian mixture model (GMM). To measure peptide recognition, the exemplary method leverages the state-of-the-art peptide-TCR binding prediction tool ERGO to predict peptide-TCR binding. TCRPPO is a flexible framework, and ERGO can be replaced with other binding predictors. Furthermore, a novel buffering mechanism, called Buf-Opt, is presented to refine TCRs that are difficult to optimize. We performed extensive experiments using 7 million TCRs from TCRdb200 (Figure 2), 10 peptides from McPAS, and 15 peptides from VDJDB. The experimental results showed that TCRPPO generated qualified TCRs with high efficacy scores and high recognition probabilities for McPAS and VDJDB peptides, with improvements of 58.2% and 26.8%, respectively, significantly outperforming the best baselines.

[0018] The recognition ability of a TCR sequence for a given peptide is r The probability that a sequence is a valid TCR is measured by the recognition probability, denoted as s v A qualified TCR is measured by a score expressed as s v >σ r Kats v >σ c is defined as an array of σ r and σ cis a predefined threshold. The goal of TCRPPO is to mutate existing TCR sequences with low recognition probability for a given peptide into eligible ones. A peptide p or TCR sequence c is represented as its amino acid sequence. <o1,o2,...,o i ,...,o l >, o i is one of the 20 natural amino acids at position i in the sequence, and l is the length of the sequence. The TCR mutation process was formulated as a Markov decision process (MDP) M = {S, A, P, R} with the following components:

[0019] S: State space. Each state s∈S is a tuple of potential TCR sequence c and peptide p, i.e., s=(c,p). The subscript t (t=0,...,T) denotes the step of s, i.e., s t =(c t ,p) is used to indicate t Note that state s may not be a valid TCR. t is a qualified c t If t contains , or t reaches the maximum step limit T, then it is a terminal state, and s T Also note that p is sampled at s0 and does not change over time t.

[0020] A: Action space, where

number

number

[0021] P: State transition probability, where

number

number

number

[0022] R: Reward function in a state. In TCRPPO, state s t The intermediate rewards at (t=0,...,T-1) are all 0. T Only the final reward at is used as the optimization guide.

[0023] Regarding the mutation policy network, TCRPPO mutates one amino acid in sequence c in a step, modifying c into a qualified TCR. Specifically, at the first step t = 0, peptide p is sampled as a target, and a valid TCR c0 is sampled to initialize s0 = (c0, p). The state s t =(c t ,p)(t>0), the mutation policy network of TCRPRO is c t By mutating one amino acid in c, the final TCR is likely to bind to p. t+i The action of changing

number

[0024] Regarding amino acid encoding, each amino acid o is represented by concatenating three vectors.

number

number

number

number

number

[0025] Regarding the state embedding, s t =(c t ,p) is the associated array c t and p are embedded by embedding c t Each amino acid in i,t For example, i,t and c t The context information in the hidden vector is then extracted using a single-layer bidirectional LSTM (Long Short Term Memory) as follows:

number

number

[0026] where:

number

number

[0027]

number

number

[0028]

number

number

[0029]

number

number

number

[0030] The peptide sequence is then used in the same way to generate the hidden vectors

number

[0031] Regarding the motion prediction, the motion at time t

number

number

number

number

[0032] where:

number

number

number

number

[0033] Given a predicted position i, TCRPPO calculates c t In i It is necessary to predict the new amino acid that will replace the . TCRPPO calculates the probability that each amino acid type will become a new substitution as follows:

number

[0034] where U j (j=1,2,3) is a learnable matrix, and softmax(·) converts the 20-dimensional vector into probabilities for the 20 amino acids. Next, the original amino acid type o i,t The type of converted amino acid is determined by sampling from the distribution excluding

[0035] Regarding the efficacy measurement of potential TCRs, quantitatively measure the likelihood that a given sequence c is a valid TCR (e.g., s v A novel scoring function (calculating ) is presented, which is part of the reward for TCRPPO. Specifically, the exemplary method trained a novel autoencoder model called the TCR-AE from only valid TCRs. The sequence reconstruction accuracy in the TCR-AE was used to measure the validity of that TCR. Intuitively, because the TCR-AE is trained only from valid TCRs, its encoding-decoding process follows the "rules" of true TCR sequences, and therefore it cannot successfully reproduce non-TCR sequences from the TCR-AE. However, if the TCR-AE learns general patterns common to both TCRs and non-TCRs and is unable to detect irregularities, or if the TCR-AE model is highly complex, it is possible that non-TCR sequences can be reconstructed with high accuracy from the TCR-AE. To mitigate this, the exemplary method additionally evaluated the latent space within the TCR-AE using a Gaussian mixture model (GMM), hypothesizing that non-TCRs deviate from the densely populated regions of TCRs in the latent space.

[0036] TCR-AE150 presents an autoencoder TCR-AE, as shown in TCRPPO100 in Figure 1. TCR-AE150 uses bidirectional LSTM to encode an input sequence c by concatenating the last hidden vectors from the two LSTM directions.

number

number

number

number

[0037] This is done via the decoder 140

number

number

number

[0038] where:

number

number

number

number

number

number

[0039] Note that TCR-AE150 is trained end-to-end from TCR, independently of TCRPPO100. Supervised constraints are applied during training to ensure that the decoded sequences have the same length as the input sequences, and therefore cross-entropy loss is applied to optimize TCR-AE150. As a standalone module, TCR-AE150 achieves a score s v The input sequence c to TCR-AE150 is encoded using only the BLOSUM matrix, as empirical evidence shows that BLOSUM encoding provides better reconstruction performance and faster convergence than other encoding combinations.

[0040] Using the fully trained TCR-AE150, the TCR efficacy score based on rearrangement of sequence c was calculated as follows:

number

[0041] where TCR-AE(c) represents the sequence of c reconstructed from TCR-AE, lev(c, TCR-AE(c)) is the Levenshtein distance, which is an index based on the edit distance between c and TCR-AE(c), and lc is the length of c. r r A higher c indicates a higher probability that c is a valid TCR. Note that when TCR-AE150 is used in testing, the length of the reconstructed sequence may not be the same as the input c, because TCR-AE150 cannot accurately predict the end of the sequence and the reconstructed sequence may be too short or too long. Therefore, the Levenshtein distance is the length of the input sequence l. c If the distance is greater than the length of the sequence, then r r Note that (c) can be negative. Negative values ​​do not affect the use of the score (e.g., a negative r r (c) shows that TCR-AE(c) and c are very different).

[0042] To better distinguish between valid and invalid TCRs, TCRPPO100 was analyzed using GMM145.

number

[0043] For a given sequence c, TCR P O 100 calculates the likelihood score of c falling within the Gaussian mixture region of the training TCRs as follows:

number

[0044] where:

number

number

[0045] Combining rearrangement-based scoring and density estimation-based scoring, we developed a new scoring method to measure TCR validity as follows.

number

[0046] This method is used to assess whether a sequence is likely to be a valid TCR, which is used in the reward function.

[0047] In terms of TCRPPO learning and final reward, an exemplary method is r Score and s v Based on the scores, the final rewards for TCRPPO100 are defined as follows:

number

[0048] where s r (c T ,p) is the predicted recognition probability by ERGO160, and σ c is c T is very likely to be a valid TCR (σ c =1.2577), and α is s r and s v is a hyperparameter used to control the tradeoff between

[0049] With respect to policy learning, an exemplary method employs proximal policy optimization (PPO) to optimize the policy network.

[0050] The objective function of PPO is defined as follows:

number

number

[0051] where:

number

number

number

number

number

[0052]

number

number

[0053] where γ∈(0,1) is the discount factor that determines the importance of future rewards, and δ t =r t +γV(s t+1 )-V(s t) is V(s t ) is the time difference error, which is the value function, and λ∈(0,1) is V(s t ) is a parameter to balance the bias and variance of the peptide embedding using a multilayer perceptron (MLP).

number

number

[0054] The objective function for V(·) is:

number

[0055] where:

number

number

number

number

[0056] The final objective function of TCRPPO100 is defined as follows:

number

[0057] where α1 and α2 are two hyperparameters that control the tradeoff between the PPO objective, the value function, and the entropy regularization term.

[0058] To address hard-to-optimize TCRs and generalize its optimization capabilities to a wider variety of TCRs, TCRPPO100 implements a novel buffering and re-optimization mechanism called Buf-Opt. This mechanism has a buffer that stores TCRs that cannot be optimized to qualify. These hard sequences are sampled from the buffer again according to the following probability distribution and further optimized by TCRPPO100:

number

[0059] In Equation 10, S is the final reward R(c T Based on the (c, p) parameter, ξ measures how difficult it is to optimize c for p, where ξ is a hyperparameter (e.g., ξ = 5), and Σ converts S(c, p) as a probability. By sampling and reoptimizing, TCRPPO100 is trained to learn from hard sequences and is expected to provide opportunities for hard sequences to be better optimized by TCRPPO100. If a hard sequence is still not optimized to qualify, it is returned to the buffer with a 50% probability. When the buffer becomes full (size 2,000 in our experiments), the sequence that was allocated earliest to the buffer is removed. TCRPPO100 plus Buf-Opt is called TCRPPO+b.

[0060] In conclusion, an exemplary embodiment of the present invention formulates the search for optimized TCRs as a RL problem and presents a framework, TCRPPO, with a mutation policy using proximal policy optimization (PPO). TCRPPO mutates TCRs into effective ones that can recognize a given peptide. TCRPPO utilizes a reward function that combines the likelihood that a mutated sequence is an effective TCR, as measured by a novel scoring function based on a deep autoencoder, with the probability that the mutated sequence recognizes the peptide, obtained from a peptide-TCR interaction predictor. Comparison of TCRPPO with multiple baseline methods showed that TCRPPO significantly outperformed all baseline methods in generating positive binding and effective TCRs. These results demonstrate the potential of TCRPPO for both precision immunotherapy and peptide discovery that recognizes TCR motifs.

[0061] The exemplary method further presents a deep reinforcement learning system with a TCR mutation policy to generate binding TCRs that recognize target peptides. The library of predefined peptides can be obtained from the genome of a virus such as SARS-CoV-2 or from sequencing patient tumor samples. Therefore, the presented exemplary system can be used for immunotherapy that targets specific types of viruses or tumors using TCR engineering.

[0062] Given a viral genome or tumor cells, an exemplary method performs sequencing, followed by a commercially available peptide processing pipeline to extract peptides that can uniquely identify the virus or tumor cells. The exemplary method also collects a library of TCRs from the target patient. By targeting this library of peptides from the virus or tumor with the given TCR, the system can generate optimized or mutated TCRs that can trigger an immune response to kill the virus or tumor cells.

[0063] An exemplary method first trains a deep neural network on the publicly available IEDB, VDJdb, and McPAS-TCR datasets, or downloads a pre-trained model such as ERGO, to predict binding interactions between peptides and TCRs. Based on this pre-trained peptide-TCR interaction score prediction model, the exemplary method develops a DRL system with a TCR mutation policy to generate TCRs with high binding scores that are identical to or differ from a library of provided TCRs by at most d amino acids. Specifically, using a pre-trained predictive deep model to define a reward function and starting from a random or existing TCR, the exemplary method then pre-trains the DRL system to learn a good TCR mutation policy that converts the given random TCR into a peptide-recognizing TCR with a high binding interaction score. Based on this trained DRL system with the pre-trained TCR mutation policy, the exemplary method randomly samples a batch of TCRs from the provided library and mutates the TCRs according to the policy network. During the mutation process, if the mutated TCR already differs from the starting TCR by d amino acids, the process terminates and the TCR is output as the final TCR. The final mutant TCR that recognizes the given peptide is output, and the set of edited mutant TCRs is ranked, with the top ranked ones being used as promising artificial TCRs to target specific viruses or tumor cells for immunotherapy.

[0064] FIG. 3 is an exemplary practical application for optimization of T cell receptors using reinforcement learning and mutation policies for precision immunotherapy, in accordance with an embodiment of the present invention.

[0065] In one example 300, peptides are processed by the TCRPPO 100 in the peptide mutation environment 110 by the mutation policy network 120 to generate new qualified peptides 310 that are displayed on a screen 312 and analyzed by a user 314. For all peptides selected from the same database (e.g., 10 peptides from McPAS and 15 peptides from VDJDB), the exemplary method trains one TCRPPO agent, which optimizes the training sequences (e.g., 7,281,105 TCRs in FIG. 2) to be qualified against one of the selected peptides. The ERGO model trained on the corresponding database determines the TCRPPO agent's recognition probability s r It should be noted that one ERGO model is trained on all peptides in each database (e.g., one ERGO predicts TCR-peptide binding for multiple peptides). Therefore, the ERGO model is trained on multiple peptides in the exemplary setting. r It should also be noted that in the exemplary method, one TCRPRO corresponding to each database was trained because the peptides and TCRs in these two databases are very different, and training the two databases together proved to be inferior in performance.

[0066] TCRPPO mutates each sequence up to 8 steps (T = 8), which is large enough since the most common length of TCR is 15. In TCRPPO training (Figure 2), the initial TCR sequence (e.g., c0 of s0) is mutated to S trn A peptide p is randomly sampled from s0 and remains the same in the following states (e.g., s t =(c t ,p)). TCRPPO100 is S trn After being fully trained by S tst It will be tested.

[0067] Experimental results comparing TCR optimization with generation-based and mutation-based methods show that TCRPPO100 significantly outperforms baseline methods. Analysis of TCRs generated by TCRPPO100 indicates that TCRPPO100 can successfully learn the conserved patterns of TCRs. Experimental comparison of the generated TCRs with existing TCRs shows that TCRPPO100 can generate TCRs similar to existing human TCRs, which can be used for further medical evaluation and research. The comparative results of TCR detection are presented in the example framework. v The scores indicate that the non-TCR sequences can be detected very effectively. v Analysis of the distribution of scores indicates that TCRPPO100 mutates sequences along a trajectory not far from the effective TCR.

[0068] FIG. 4 is an exemplary processing system for optimization of T cell receptors using reinforcement learning and mutation policies for precision immunotherapy, according to an embodiment of the present invention.

[0069] The processing system includes at least one processor (CPU) 404 operatively coupled to other components via a system bus 402. A graphical processing unit (GPU) 405, a cache 406, a read-only memory (ROM) 408, a random access memory (RAM) 410, an input / output (I / O) adapter 420, a network adapter 430, a user interface adapter 440, and a display adapter 450 are operatively coupled to the system bus 402. Additionally, the TCRPPO 100 is employed within the peptide mutation environment 110 by using a mutation policy network 120.

[0070] Storage devices 422 are operably coupled to the system bus 402 by I / O adapter 420. Storage devices 422 may be disk storage devices (e.g., magnetic or optical disk storage devices), solid-state magnetic devices, or the like.

[0071] The transceiver 432 is operably coupled to the system bus 402 by the network adapter 430 .

[0072] User input device(s) 442 are operatively coupled to system bus 402 by user interface adapter 440. User input device 442 may be any of a keyboard, mouse, keypad, image capture device, motion sensing device, microphone, a device incorporating the functionality of at least two of the foregoing devices, etc. Of course, other types of input devices may be used while maintaining the spirit of the present invention. User input device(s) 442 may be the same type of user input device or different types of user input devices. User input device 442 is used to input and output information to and from the processing system.

[0073] A display device 452 is operatively coupled to the system bus 402 by a display adapter 450 .

[0074] Of course, the processing system may include other elements (not shown) or omit certain elements, as would be readily contemplated by one skilled in the art. For example, various other input and / or output devices may be included in the system, depending on the particular implementation, as would be readily understood by one skilled in the art. For example, various types of wireless and / or wired input and / or output devices may be used. Furthermore, various configurations of additional processors, controllers, memory, etc. may also be utilized, as would be readily understood by one skilled in the art. These and other variations of the processing system will be readily contemplated by one skilled in the art given the teachings of the present invention provided herein.

[0075] FIG. 5 is a block / flow diagram of an exemplary method for optimizing T cell receptors using reinforcement learning and mutation policies for precision immunotherapy, according to an embodiment of the present invention.

[0076] In block 501, a target peptide and a library of patient TCRs are extracted.

[0077] In block 503, a deep neural network is trained or a pre-trained model such as ERGO is downloaded to predict interaction scores between peptide antigens and TCRs.

[0078] In block 505, a reward function is defined using a pre-trained interaction prediction deep model, and starting from existing TCRs, a deep reinforcement learning (DRL) system is pre-trained to learn a good TCR mutation policy that transforms a given TCR into an optimized TCR with a high interaction score.

[0079] In block 507, based on this trained DRL system with a pre-trained TCR mutation policy, randomly sample a batch of TCRs from the provided library and mutate the TCRs according to the policy network.

[0080] In block 509, the final mutant TCRs that target the given peptide antigen are output and the compiled set of mutant TCRs is ranked.

[0081] In block 511, the top ranked ones are used as promising candidates for precision immunotherapy by TCR engineering to target specified viruses or tumor cells.

[0082] FIG. 6 is a block / flow diagram of an exemplary method for T cell receptor optimization using reinforcement learning and mutation policies for precision immunotherapy, according to an embodiment of the present invention.

[0083] In block 601, peptides are extracted to identify viruses or tumor cells.

[0084] In block 603, a library of TCRs is collected from a target patient.

[0085] In block 605, a deep neural network predicts interaction scores between the extracted peptides and TCRs from the target patient.

[0086] In block 607, a deep reinforcement learning (DRL) framework using a TCR mutation policy is developed to generate TCRs with the highest combined score.

[0087] At block 609, a reward function is defined based on the reconstruction-based score and the density estimation-based score.

[0088] In block 611, a batch of TCRs is randomly sampled and the TCRs are mutated according to the policy network.

[0089] At block 613, the mutated TCR is output.

[0090] In block 615, the output TCRs are ranked and the top ranked TCR candidates are utilized to target viruses or tumor cells for immunotherapy.

[0091] In conclusion, the exemplary method proposes a DRL system using a TCR mutation policy to generate binding TCRs that recognize a given peptide antigen. The presented system can be used to generate TCRs for immunotherapy targeting specific types of viruses or tumors. The reward design is based on the TCR distribution score and binding interaction score. The exemplary method optimizes the DRL model using PPO, outputs the final mutant TCRs, and ranks the compiled set of mutant TCRs. The top-ranked TCRs are used as promising candidates for immunotherapy targeting specific viruses or tumors.

[0092] As used herein, "data," "content," "information," and similar terms may be used interchangeably to refer to data that may be captured, transmitted, received, displayed, and / or stored in accordance with various exemplary embodiments. Accordingly, the use of such terms should not be construed to limit the spirit and scope of the present disclosure. Furthermore, when a computing device is described herein as receiving data from another computing device, the data may be received directly from the other computing device or indirectly through one or more intermediary computing devices, such as, for example, one or more servers, relays, routers, network access points, base stations, and / or the like.

[0093] As will be appreciated by those skilled in the art, aspects of the present invention may be embodied as a system, method, or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be referred to generally herein as a "circuit," "module," "computer," "device," or "system." Furthermore, aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable medium(s) having computer-readable program code embodied thereon.

[0094] Any combination of one or more computer-readable media may be utilized. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include an electrical connection having one or more wires, a portable computer diskette, a hard disk, RAM, ROM, erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable CD-ROM, an optical data storage device, a magnetic data storage device, or any suitable combination of the foregoing. As used herein, a computer-readable storage medium is any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0095] A computer-readable signal medium may include a propagated data signal having computer-readable program code embodied therein, for example, as part of a baseband or carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electromagnetic waves, optical waves, or any suitable combination thereof. A computer-readable signal medium is not a computer-readable storage medium, but may be any computer-readable medium that can communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0096] The program code embodied in the computer readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, etc., or any suitable combination of the foregoing.

[0097] Computer program code for carrying out operations for aspects of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" programming language. The program code may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a LAN or WAN, or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).

[0098] Aspects of the present invention are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in the flowchart and / or block diagram blocks or modules.

[0099] These computer program instructions may also be stored on a computer-readable medium that can direct a computer, other programmable data processing apparatus, or other device to function in a particular way to produce an article of manufacture, where the instructions stored on the computer-readable medium include instructions that implement the functions / acts specified in the blocks or modules of the flowcharts and / or block diagrams.

[0100] Computer program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause the computer, other programmable apparatus, or other device to perform a series of operational steps, creating a computer-implemented process such that the instructions executing on the computer or other programmable apparatus provide a process for performing the functions / operations specified in the flowchart and / or block diagram blocks or modules.

[0101] It should be understood that the term "processor," as used herein, is intended to include any processing device, such as one that includes a CPU and / or other processing circuitry. It should also be understood that the term "processor" may refer to more than one processing device, and that various elements associated with a processing device may be shared by other processing devices.

[0102] As used herein, the term "memory" is intended to include memory associated with a processor or CPU, such as, for example, RAM, ROM, fixed memory devices (e.g., hard drives), removable memory devices (e.g., diskettes), flash memory, etc. Such memory is considered a computer-readable storage medium.

[0103] Furthermore, the phrase "input / output device" or "I / O device" as used herein is intended to include, for example, one or more input devices (e.g., keyboard, mouse, scanner, etc.) for inputting data into a processing unit, and / or one or more output devices (e.g., speakers, displays, printers, etc.) for presenting results related to a processing unit.

[0104] The foregoing is understood in all respects to be illustrative and exemplary, but not restrictive, the scope of the invention disclosed herein being determined not from the detailed description but from the claims which are interpreted in accordance with the full breadth permitted by the patent laws. It will be understood that the embodiments shown and described herein are merely illustrative of the principles of the invention, and that various modifications can be made by those skilled in the art without departing from the scope and spirit of the invention. Various other feature combinations can be implemented by those skilled in the art without departing from the scope and spirit of the invention. Having thus described aspects of the invention with the detail and particularity required by the patent laws, what is desired to be claimed and protected by Letters Patent is set forth in the appended claims.

Claims

1. 1. A method of performing deep reinforcement learning using a T cell receptor (TCR) mutation policy to generate CDR3 sequences of a beta chain of a binding TCR that recognizes a target peptide for immunotherapy, comprising: Extracting peptides to identify viruses or tumor cells; collecting a library of TCR β chain CDR3 sequences from target patients; predicting an interaction score between the extracted peptide and the CDR3 sequence of the β chain of the TCR from the target patient using a deep neural network trained with teacher data that receives as input the extracted peptide and a library of CDR3 sequences of the β chain of the TCR and outputs an interaction score between the peptide and the CDR3 sequence of the β chain of the TCR; Using the deep neural network, a deep reinforcement learning (DRL) framework is developed using a TCR mutation policy to generate a CDR3 sequence of the β chain of the TCR that maximizes the interaction score; defining a reward function based on the reconstruction-based score and the density estimation-based score; Randomly sampling a batch of TCR β chain CDR3 sequences generated by the DRL, and mutating the TCR β chain CDR3 sequences according to a policy network; Outputting the mutated TCR beta chain CDR3 sequence; ranking the output TCR β chain CDR3 sequences and using the top-ranked TCR β chain CDR3 sequence candidates to target the virus or the tumor cell for immunotherapy; A method in which the reward function measures both the likelihood that a variant sequence is a valid TCR β chain CDR3 sequence and the probability that the TCR β chain CDR3 sequence recognizes a peptide.

2. The method of claim 1, wherein the likelihood that the mutant sequence is a valid TCR β chain CDR3 sequence is determined by a TCR autoencoder (TCR-AE) trained only with TCR β chain CDR3 sequences.

3. The method of claim 2, wherein density estimation on the latent space within the TCR-AE is evaluated using a Gaussian mixture model (GMM).

4. 3. The method of claim 2, wherein the TCR-AE uses a bidirectional long short-term memory (LSTM) to encode an input sequence into a hidden vector by concatenating the last hidden vectors from two LSTM directions.

5. The method of claim 1, wherein a buffering and re-optimization framework including a buffer is employed to handle difficult-to-optimize TCR β chain CDR3 sequences and generalize the optimization capabilities to more diverse TCR β chain CDR3 sequences.

6. 2. The method of claim 1, wherein the CDR3 sequence of the TCR beta chain and the extracted peptides are encoded into a distributed embedding space by a TCR-AE, and a mapping between the embedding space and the TCR mutation policy is learned.

7. 1. A non-transitory computer-readable storage medium comprising a computer-readable program for performing deep reinforcement learning using a T cell receptor (TCR) mutation policy to generate CDR3 sequences of a beta chain of a binding TCR that recognizes a target peptide for immunotherapy, the computer-readable program, when executed on a computer, causing the computer to: Extracting peptides to identify viruses or tumor cells; collecting a library of TCR β chain CDR3 sequences from target patients; predicting an interaction score between the extracted peptide and the CDR3 sequence of the β chain of the TCR from the target patient using a deep neural network trained with teacher data that receives as input the extracted peptide and a library of CDR3 sequences of the β chain of the TCR and outputs an interaction score between the peptide and the CDR3 sequence of the β chain of the TCR; Using the deep neural network to develop a deep reinforcement learning (DRL) framework with a TCR mutation policy to generate a TCR β chain CDR3 sequence that maximizes the interaction score; defining a reward function based on the reconstruction-based score and the density estimation-based score; Randomly sampling a batch of TCR β chain CDR3 sequences generated by the DRL, and mutating the TCR β chain CDR3 sequences according to a policy network; outputting the mutated TCR β chain CDR3 sequence; ranking the output TCR β chain CDR3 sequences and using the top-ranked TCR β chain CDR3 sequence candidates to target the virus or the tumor cell for immunotherapy; A non-transitory computer-readable storage medium, wherein the reward function measures both the likelihood that a variant sequence is a valid TCR beta chain CDR3 sequence and the probability that the TCR beta chain CDR3 sequence recognizes a peptide.

8. The non-transitory computer-readable storage medium of claim 7, wherein the likelihood that the mutant sequence is a valid TCR β chain CDR3 sequence is determined by a TCR autoencoder (TCR-AE) trained only with TCR β chain CDR3 sequences.

9. 9. The non-transitory computer-readable storage medium of claim 8, wherein density estimation on the latent space in the TCR-AE is evaluated using a Gaussian mixture model (GMM).

10. 9. The non-transitory computer-readable storage medium of claim 8, wherein the TCR-AE uses a bidirectional long short-term memory (LSTM) to encode an input sequence into a hidden vector by concatenating the last hidden vectors from two LSTM directions.

11. The non-transitory computer-readable storage medium of claim 7, wherein a buffering and re-optimization framework including a buffer is employed to handle TCR β chain CDR3 sequences that are difficult to optimize and generalize the optimization capabilities to more diverse TCR β chain CDR3 sequences.

12. 8. The non-transitory computer-readable storage medium of claim 7, wherein the CDR3 sequence of the beta chain of the TCR and the extracted peptides are encoded into a distributed embedding space by a TCR-AE, and a mapping between the embedding space and the TCR mutation policy is learned.

13. 1. A system for performing deep reinforcement learning with a T cell receptor (TCR) mutation policy to generate CDR3 sequences of a beta chain of a binding TCR that recognizes a target peptide for immunotherapy, comprising: Memory and one or more processors in communication with the memory, the processors comprising: Extract peptides to identify viruses or tumor cells collecting a library of TCR β chain CDR3 sequences from target patients; predicting an interaction score between the extracted peptide and the CDR3 sequence of the β chain of the TCR from the target patient using a deep neural network trained with teacher data that receives the extracted peptide and a library of CDR3 sequences of the β chain of the TCR as input and outputs an interaction score between the peptide and the CDR3 sequence of the β chain of the TCR; using the deep neural network to develop a deep reinforcement learning (DRL) framework with a TCR mutation policy to generate a TCR β chain CDR3 sequence that maximizes the interaction score; defining a reward function based on the reconstruction-based score and the density estimation-based score; Randomly sampling a batch of CDR3 sequences of the TCR β chain generated by the DRL, and mutating the CDR3 sequences of the TCR β chain according to a policy network; Outputting the CDR3 sequence of the mutated TCR β chain; ranking the output TCR β chain CDR3 sequences and using the top-ranked TCR β chain CDR3 sequence candidates to target the virus or the tumor cell for immunotherapy; A system in which the reward function measures both the likelihood that a variant sequence is a valid TCR β chain CDR3 sequence and the probability that the TCR β chain CDR3 sequence recognizes a peptide.

14. The system of claim 13, wherein the likelihood that the mutant sequence is a valid TCR β chain CDR3 sequence is determined by a TCR autoencoder (TCR-AE) trained only with TCR β chain CDR3 sequences.

15. The system of claim 14 , wherein density estimation on the latent space within the TCR-AE is evaluated using a Gaussian mixture model (GMM).

16. 15. The system of claim 14, wherein the TCR-AE uses a bidirectional long short-term memory (LSTM) to encode an input sequence into a hidden vector by concatenating the last hidden vectors from two LSTM directions.

17. The system of claim 13, wherein a buffering and re-optimization framework including a buffer is employed to handle TCR β chain CDR3 sequences that are difficult to optimize and generalize the optimization capabilities to more diverse TCR β chain CDR3 sequences.