TCR Optimization via Deep Reinforcement Learning for Immunotherapy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computational methods for optimizing T-cell receptors (TCRs) for immunotherapy are time-consuming and fail to generate high-affinity TCRs tailored to recognize specific peptides, as they do not consider the validity of generated sequences and do not optimize TCRs for different peptides effectively.

Innovation Solution

A deep reinforcement learning framework, TCRPPO, is developed to optimize TCRs using a proximal policy optimization (PPO) algorithm, which learns a joint policy to maximize interaction scores with peptide antigens while minimizing interactions with self-peptides, ensuring TCR validity through a novel reward function and auto-encoder-based scoring.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If current computational methods are used to optimize TCRs, then the optimization process can be performed, but it is time-consuming and fails to generate high-affinity TCRs tailored to specific peptides

Engineering Contradiction:
Improveoptimization speedVSAvoidTCR affinity precision
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent replaces traditional mechanical/computational optimization methods with a deep reinforcement learning system (TCRPPO) that uses neural networks to learn optimal TCR sequences. The system substitutes conventional iterative computational approaches with a trained AI model that can predict and generate high-affinity TCRs directly, significantly reducing optimization time while improving precision through the learned policy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent performs preliminary training of the reinforcement learning agent on a large dataset of TCR-peptide interactions before actual optimization. This preliminary action allows the system to pre-learn patterns and relationships, enabling it to quickly generate high-affinity TCRs for specific peptides without time-consuming iterative optimization during actual use.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If traditional methods optimize TCRs, then general optimization can be performed, but they do not consider TCR validity and fail to optimize for different peptides effectively

Engineering Contradiction:
Improvepeptide-specific optimization capabilityVSAvoidTCR sequence validity
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where the reinforcement learning agent receives rewards based on both binding affinity predictions and TCR validity assessments. The system uses autoencoders to encode TCR sequences and provides feedback on whether generated sequences maintain structural validity, allowing the agent to learn to generate both high-affinity and valid TCRs simultaneously for different peptides.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent creates a universal reinforcement learning framework (TCRPPO) that can optimize TCRs for any given peptide target. The system uses a multi-functionality approach where the same trained model handles diverse peptide-specific optimization tasks, adapting to different targets while maintaining TCR validity through the learned policy and validity assessment mechanisms.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If TCR optimization focuses on binding affinity, then recognition probability increases, but interaction with self-peptides may increase causing safety issues

Engineering Contradiction:
Improvepeptide recognition accuracyVSAvoidself-peptide interaction
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent applies preliminary anti-action by training the reinforcement learning system to explicitly minimize interactions with self-peptides while maximizing binding to target peptides. The reward function incorporates penalties for self-peptide binding, allowing the system to pre-learn avoidance strategies and generate TCRs that are both highly specific to target peptides and safe regarding self-peptide recognition.

Inventive Principle:
Principle #9Preliminary anti-action

Solution Approach 2:

The patent uses local quality by differentiating the optimization criteria for different peptide types. The system applies different weightings and constraints when evaluating binding to target peptides versus self-peptides, allowing high recognition accuracy for targets while maintaining low interaction with self-peptides through localized optimization adjustments in the reward function.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20230304189A1TCR engineering with deep reinforcement learning for increasing efficacy and safety of TCR-t immunotherapy
Publication Date: 2023.09.28 NEC LABORATORIES AMERICA INC
  • US20230304189A1 patent drawing
  • US20230304189A1 patent drawing
  • US20230304189A1 patent drawing

AI summary

A method for implementing deep reinforcement learning with T-cell receptor (TCR) mutation policies to generate binding TCRs for immunotherapy includes extracting peptides to identify a virus or tumor cells, collecting a library of TCRs from patients, predicting interaction scores between the extracted peptides and the TCRs from the patients, developing a deep reinforcement learning framework with TCR mutation policies to generate TCRs with maximum binding scores, defining reward functions, outputting mutated TCRs, ranking the outputted TCRs to utilize top-ranked TCR candidates to target the virus or the tumor cells, and for each top-ranked TCR candidate, repeatedly identifying a set of self-peptides that the top-ranked TCR candidate binds to and further optimizing it greedily by maximizing a sum of its interaction scores with a given set of peptide antigens while minimizing a sum of its interaction scores with the set of self-peptides until stopping criteria of efficacy and safety are met.