Predicting T cell receptor repertoire selection by physical model-augmented pseudolabeling
By employing physical modeling and data-augmented pseudo-labeling, the method enhances the accuracy and efficiency of TCR-peptide interaction predictions, overcoming data scarcity challenges in existing models for personalized medicine and targeted vaccines.
Patent Information
- Application Number
- JP2024521283
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-10-20
- Filing Date
- 2022-10-21
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-10-21
AI Technical Summary
Existing deep learning models for predicting T cell receptor (TCR)-peptide interactions are inefficient and inaccurate due to the lack of diverse TCRs and peptides in their datasets, leading to suboptimal performance in personalized medicine and targeted vaccines.
The method involves using physical modeling to generate an extended training dataset through docking energy scores, applying pseudo-labeling to classify TCR-peptide pairs, and iteratively retraining a deep learning model until convergence, utilizing a conditional variational autoencoder (cVAE) for improved prediction.
This approach significantly enhances the accuracy and efficiency of TCR-peptide interaction predictions, addressing data scarcity issues and improving the reliability of personalized medicine and targeted vaccine development.
Smart Images

Figure 0007763337000022 
Figure 0007763337000023 
Figure 0007763337000024
Abstract
Description
[Technical Field]
[0001] The present invention relates to the prediction of T cell receptor (TCR)-peptide interactions, and more particularly to the prediction of TCR-peptide interactions by training and utilizing a conditional variational autoencoder (cVAE) for TCR generation and classification using physical modeling and data-augmented pseudo-labeling. [Background technology]
[0002] 2. Description of Related Art Predicting interactions between T cell receptors (TCRs) and peptides is crucial for personalized medicine and targeted vaccines in immunotherapy. Previous systems, methods, and current datasets for training deep learning models for such predictions are inaccurate and inefficient, limited at least in part by the lack of diverse TCRs and peptides in the datasets. Recently, several deep learning approaches have been utilized to predict interactions between TCRs and peptides, including those utilizing long short-term memory (LSTM) and autoencoders. These include the complementarity-determining region 3 (CDR3) beta chain (e.g., ERGO autoencoder), the CDR3 alpha chain, V and J genes, MHC type, T cell type (e.g., ERGO II autoencoder), Gaussian processing, stacked convolutional networks for TCR-peptide prediction (e.g., NetTCR 1.0), and CDR3 alpha and beta chain pairs (e.g., NetTCR 2.0). However, these systems and methods are limited, at least in part, by the lack of diverse TCRs and peptides in their datasets, resulting in inefficient and inaccurate prediction of interactions between TCRs and peptides, which suffer from the problems mentioned above. Summary of the Invention
[0003] According to one aspect of the present invention, there is provided a method for predicting T cell receptor (TCR)-peptide interactions, comprising: determining a multiple sequence alignment (MSA) of TCR-peptide pair sequences from a dataset of TCR-peptide pair sequences using a sequence analyzer; constructing TCR and peptide structures using the MSA and corresponding structures from the Protein Data Bank (PDB) using MODELLER; and generating an extended TCR-peptide training dataset based on docking energy scores determined by docking peptides to the TCR using physical modeling based on the TCR and peptide structures constructed using MODELLER. Using pseudo-labels based on the docking energy scores, TCR-peptide pairs are classified and labeled as positive or negative pairs, and the deep learning model is iteratively retrained based on the extended TCR-peptide training dataset and the pseudo-labels until convergence is achieved.
[0004] According to another aspect of the present invention, there is provided a system for predicting T cell receptor (TCR)-peptide interactions, comprising a processor operatively coupled to a non-transitory computer-readable storage medium, for training a deep learning model for predicting the TCR-peptide interactions by using a sequence analyzer to determine a multiple sequence alignment (MSA) of TCR-peptide pair sequences from a dataset of TCR-peptide pair sequences, constructing TCR and peptide structures using the MSA and corresponding structures from the Protein Data Bank (PDB) using MODELLER, and generating an extended TCR-peptide training dataset based on docking energy scores determined by docking peptides to the TCR using physical modeling based on the TCR and peptide structures constructed using MODELLER. Pseudo-labels based on the docking energy scores are used to classify and label TCR-peptide pairs as positive or negative pairs. The deep learning model is iteratively retrained based on the extended TCR-peptide training dataset and the pseudo-labels until convergence is achieved.
[0005] According to another aspect of the present invention, there is provided a non-transitory computer-readable storage medium comprising content configured to cause a computer to execute a method for training a deep learning model for predicting TCR-peptide interactions by: determining a multiple sequence alignment (MSA) of sequences of multiple TCR-peptide pairs from a dataset of TCR-peptide pair sequences using a sequence analyzer; constructing TCR and peptide structures using the MSA and corresponding structures from the Protein Data Bank (PDB) using MODELLER; and generating an extended TCR-peptide training dataset based on docking energy scores determined by docking peptides to TCRs using physical modeling based on the TCR and peptide structures constructed using MODELLER. Pseudo-labels based on the docking energy scores are used to classify and label TCR-peptide pairs as positive or negative pairs. The deep learning model is iteratively retrained based on the extended TCR-peptide training dataset and the pseudo-labels until convergence is achieved.
[0006] These and other features and advantages will become apparent from the following detailed description of illustrative embodiments thereof, which is to be read in connection with the accompanying drawings. [Brief explanation of the drawings]
[0007] The present disclosure provides details in the following description of preferred embodiments with reference to the following figures.
[0008] [Figure 1] 1 is a block diagram illustrating an exemplary processing system to which the present invention may be applied, in accordance with an embodiment of the present invention.
[0009] [Figure 2] FIG. 1 illustrates an exemplary high-level diagram of a method for conjugating a T cell receptor (TCR) to a peptide, according to an embodiment of the present invention.
[0010] [Figure 3] FIG. 1 illustrates an exemplary high-level method for training a deep learning model for predicting T cell receptor (TCR)-peptide interactions, according to an embodiment of the present invention.
[0011] [Figure 4] FIG. 1 is a block / flow diagram illustrating an exemplary system and method for training a deep learning model for predicting T cell receptor (TCR)-peptide interactions with multiple encoders, according to an embodiment of the present invention.
[0012] [Figure 5] FIG. 1 is a block / flow diagram illustrating an exemplary system and method for predicting T cell receptor (TCR)-peptide interactions by training a deep learning model and calculating docking energies, according to embodiments of the present invention.
[0013] [Figure 6] FIG. 1 is a block / flow diagram illustrating an exemplary system and method for predicting T cell receptor (TCR)-peptide interactions by training a conditional variational autoencoder (cVAE) for TCR generation and classification, according to an embodiment of the present invention.
[0014] [Figure 7] FIG. 1 is a block / flow diagram illustrating an exemplary system and method for predicting T cell receptor (TCR)-peptide interactions by augmenting a training dataset using pseudo-labeling, according to an embodiment of the present invention.
[0015] [Figure 8] FIG. 1 is a block / flow diagram illustrating an exemplary system and method for predicting T cell receptor (TCR)-peptide interactions by training a conditional variational autoencoder (cVAE) for TCR generation and classification and augmenting the training dataset with pseudo-labeling, in accordance with embodiments of the present invention.
[0016] [Figure 9] FIG. 1 is a block / flow diagram illustrating an exemplary method for predicting and classifying T cell receptor (TCR)-peptide interactions by training and utilizing a neural network, according to an embodiment of the present invention.
[0017] [Figure 10] FIG. 1 illustrates an exemplary system for predicting and classifying T cell receptor (TCR)-peptide interactions by training and utilizing a neural network, according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0018] According to embodiments of the present invention, systems and methods are provided for predicting T cell receptor (TCR)-peptide interactions by using physical modeling and data-augmented pseudo-labeling to train and utilize a conditional variational autoencoder (cVAE) for TCR generation and classification.
[0019] In various embodiments, to combat the data scarcity problem, the present invention extends the training dataset through physical modeling of TCR-peptide pairs, thereby increasing the efficiency and accuracy of TCR-peptide interaction predictions (e.g., for use in personalized medicine in immunotherapy and targeted vaccines). In some embodiments, the docking energies between supplemental unknown TCR-peptide pairs can be utilized as additional example-label pairs to train a neural network learning model in a supervised manner. The area under the curve (AUC) scores of the model's predictions can be further fine-tuned and improved by pseudo-labeling such unknown TCR-peptide pairs and retraining the model using these pseudo-labeled TCR-peptide pairs. Experimental results demonstrate that training a deep neural network with physical modeling and data-augmented pseudo-labeling, in accordance with aspects of the present invention, significantly improves the accuracy and efficiency of TCR-peptide interaction predictions over baseline and conventional systems and methods.
[0020] In various embodiments, the present invention can be utilized to train deep learning models for predicting TCR-peptide interactions from three losses: supervised cross-entropy loss from given known TCR-peptide pairs, supervised cross-entropy loss based on the docking energy of unknown TCR-peptide pairs, and Kullback-Leibler (KL)-divergence loss from pseudo-labeled unknown TCR-peptide pairs, as described in further detail below in accordance with aspects of the present invention.
[0021] Predicting T cell receptor (TCR) and peptide-major histocompatibility complex (pMHC) interactions is essential for developing repertoire-based biomarkers (e.g., predicting whether a host is exposed to a target) and can be used in personalized medicine in immunotherapy and targeted vaccines, according to embodiments of the present invention. However, due to a lack of experimental data covering both a large number of peptides and a large number of TCRs, such predictions have traditionally been computationally inefficient, and traditional systems and methods can produce inaccurate results.
[0022] The embodiments described herein may be entirely hardware, entirely software, or contain both hardware and software elements. In a preferred embodiment, the invention is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.
[0023] Embodiments may include a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. A computer-usable or computer-readable medium may include any apparatus that stores, communicates, propagates, or transports a program for use by or in connection with an instruction execution system, apparatus, or device. The medium may be a magnetic, optical, electronic, electromagnetic, infrared, or semiconductor system (or apparatus or device) or propagation medium. The medium may include computer-readable storage media such as semiconductor or solid-state memory, magnetic tape, removable computer diskettes, random access memory (RAM), read-only memory (ROM), rigid magnetic disks, and optical disks.
[0024] Each computer program can be tangibly stored on a machine-readable storage medium or device (e.g., program memory or magnetic disk) readable by a general-purpose or special-purpose programmable computer to configure and control the operation of the computer when the storage medium or device is read by the computer to perform the procedures described herein. The system of the present invention can also be considered to be embodied in a computer-readable storage medium configured with a computer program, where the configured storage medium causes the computer to operate in a particular, predetermined manner to perform the functions described herein.
[0025] A data processing system suitable for storing and / or executing program code may include at least one processor coupled directly or indirectly to memory elements via a system bus. The memory elements may include local memory employed during the actual execution of the program code, bulk storage, and cache memory that provides temporary storage of at least some program code to reduce the number of times the code is retrieved from bulk storage during execution. Input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) may be coupled to the system either directly or through intervening I / O controllers.
[0026] Network adapters may also be coupled to the system to enable the data processing system to become coupled to other data processing systems or remote printers or storage devices through intervening private or public networks. Modems, cable modems, and Ethernet cards are just a few of the types of network adapters currently available.
[0027] Aspects of the present invention are described below with reference to flowchart illustrations and / or block diagrams of methods, systems, and computer program products according to embodiments of the invention. It should be noted that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions.
[0028] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in the flowcharts or block diagrams may represent a module, segment, or portion of code, which may consist of one or more executable instructions for implementing a specified logical function. In some alternative implementations of the present invention, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may actually be executed substantially simultaneously, or may be executed in reverse order, or may be executed in another order depending on the functionality of the particular embodiment.
[0029] It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special-purpose hardware systems that perform particular functions / operations, or by combinations of special-purpose hardware and computer instructions in accordance with the present principles.
[0030] As employed herein, the terms “hardware processor subsystem” or “processor” or “hardware processor” can refer to a processor, memory, software, or a combination thereof that cooperate to perform one or more specific tasks. In useful embodiments, the hardware processor subsystem can include one or more data processing elements (e.g., logic circuits, processing circuits, instruction execution devices, etc.). The one or more data processing elements can be included in a central processing unit, an image processing unit, and / or a separate processor or computing element-based controller (e.g., logic gates, etc.). The hardware processor subsystem can include one or more on-board memories (e.g., cache, dedicated memory array, read-only memory, etc.). In some embodiments, the hardware processor subsystem can include one or more memories that can be on-board or off-board or dedicated for use by the hardware processor subsystem (e.g., ROM, RAM, basic input / output system (BIOS), etc.).
[0031] In some embodiments, the hardware processor subsystem may include and execute one or more software elements, which may include an operating system and / or one or more applications and / or specific code to achieve a particular result.
[0032] In other embodiments, the hardware processor subsystem may include dedicated, dedicated circuitry to perform one or more electronic processing functions to achieve a specified result. Such circuitry may include one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and / or programmable logic arrays (PLAs).
[0033] These and other variations of the hardware processor subsystem are also contemplated in accordance with embodiments of the present invention.
[0034] Referring now to the drawings, in which like numerals represent the same or similar elements, and initially to FIG. 1 , an exemplary processing system 100 to which the present principles may be applied is illustratively depicted in accordance with an embodiment of the present principles.
[0035] In some embodiments, processing system 100 may include at least one processor (CPU) 104 operably coupled to other components via a system bus 102. Cache 106, read-only memory (ROM) 108, random access memory (RAM) 110, input / output (I / O) adapter 120, audio adapter 130, network adapter 140, user interface adapter 150, and display adapter 160 are operably coupled to system bus 102.
[0036] A first storage device 122 and a second storage device 124 are operably coupled to the system bus 102 by an I / O adapter 120. The storage devices 122 and 124 may be disk storage devices (e.g., magnetic or optical disk storage devices), solid-state magnetic devices, etc. The storage devices 122 and 124 may be the same type of storage device or different types of storage devices.
[0037] Speakers 132 are operably coupled to system bus 102 by audio adapter 130. Transceiver 142 is operably coupled to system bus 102 by network adapter 140. Display device 162 is operably coupled to system bus 102 by display adapter 160. One or more neural network training devices 164 may be further coupled to system bus 102 by any suitable connection system or method (e.g., Wi-Fi, wired, network adapter, etc.) in accordance with aspects of the present invention.
[0038] First user input device 152, second user input device 154, and third user input device 156 are operably coupled to system bus 102 by user interface adapter 150. User input devices 152, 154, and 156 may be any of a keyboard, mouse, keypad, image capture device, motion sensing device, microphone, a device incorporating the functionality of at least two of the foregoing devices, and the like. Of course, other types of input devices may be used while maintaining the spirit of the principles of the present invention. User input devices 152, 154, and 156 may be the same type of user input device or different types of user input devices. User input devices 152, 154, and 156 are used to input and output information to and from system 100.
[0039] Of course, processing system 100 may include other elements (not shown) or omit certain elements, as would be readily contemplated by one skilled in the art. For example, various other input and / or output devices may be included in processing system 100, depending on the particular implementation, as would be readily understood by one skilled in the art. For example, various types of wireless and / or wired input and / or output devices may be used. Furthermore, various configurations of additional processors, controllers, memory, etc. may also be utilized, as would be readily understood by one skilled in the art. These and other variations of processing system 100 will be readily contemplated by one skilled in the art given the teachings of the present principles provided herein.
[0040] Furthermore, it should be understood that systems 400, 500, 600, 700, 800, and 1000, described below with respect to Figures 4, 5, 6, 7, 8, and 10, respectively, are systems for practicing respective embodiments of the present invention. Part or all of processing system 100 may be implemented in one or more of the elements of systems 400, 500, 600, 700, 800, and 1000, in accordance with aspects of the present invention.
[0041] Furthermore, it should be understood that processing system 100 is capable of performing at least some of the methods described herein, including, for example, at least some of methods 200, 300, 400, 500, 600, 700, 800, and 900 described below with respect to Figures 2, 3, 4, 5, 6, 7, 8, and 9, respectively. Similarly, some or all of systems 400, 500, 600, 700, 800, and 1000 can be used to perform at least some of methods 200, 300, 400, 500, 600, 700, 800, and 900 of Figures 2, 3, 4, 5, 6, 7, 8, and 9, respectively, in accordance with embodiments of the present invention.
[0042] Referring now to FIG. 2, a diagram illustrating an exemplary high-level view of a method 200 for conjugating a T cell receptor (TCR) to a peptide, according to an embodiment of the present invention, is illustratively shown.
[0043] First, it should be noted that recognition of peptide / major histocompatibility complex (MHC) 206 by TCR 208 is a key interaction performed by a person's adaptive immune system. TCR 208 is a protein complex found on the surface of T cells 204 (or T lymphocytes) and is responsible for recognizing fragments of antigens (e.g., contained on tumor or virus-infected cells 202) as peptides 210 bound to MHC molecules 206. In some embodiments, tumor cells 202 can be eliminated by determining and utilizing tumor antigen-specific T cells 204 by targeting TCR 208-peptide 210 / MHC 206 interactions on the surface of tumor cells 202, in accordance with aspects of the present invention.
[0044] Although large databases of TCR208 and peptides210 are available, in practice, information about TCR-peptide binding specificity is limited and insufficient for determining most TCR-peptide binding. In some embodiments, deep learning models can be trained to determine information about binding specificity and predict TCR-peptide interactions from three losses in accordance with aspects of the present invention: supervised cross-entropy loss from given known TCR-peptide pairs, supervised cross-entropy loss based on the docking energy of unknown TCR-peptide pairs, and Kullback-Leibler (KL)-divergence loss from pseudo-labeled unknown TCR-peptide pairs.
[0045] In various embodiments, D TCRdb The introduction of puts the learning problem in a semi-supervised setting. Besides pseudo-labeling by physical modeling (described in more detail herein below), established semi-supervised techniques can be leveraged to further improve results. Pseudo-labeling has proven to be a successful technique in semi-supervised learning. According to various embodiments, an algorithm can be utilized that first labels unlabeled examples using a model that was initially trained on a labeled dataset, and then the model can be retrained on a labeled training dataset using the augmented pseudo-labeled examples according to aspects of the present invention.
[0046] In an exemplary embodiment,
number
number
number
number
[0047] Referring now to FIG. 3 , there is illustratively shown a diagram illustrating a high-level view of a method 300 for training a deep learning model for predicting T cell receptor (TCR)-peptide interactions in accordance with an embodiment of the present invention.
[0048] In some embodiments, the TCR-peptide model 312 (e.g., ERGO) is based on a labeled dataset 304 (e.g., D TCRdb ) example (which can include the TCR from block 310). label ), pseudo-labeled examples (e.g., the training dataset (D train ) and / or peptides 308 from the training model as inputs. pseudo ), and the physical properties between the TCR 306, 310 and the peptide 308 (L physical It is trained using three (3) losses, including a cross-entropy loss based on total ) can be determined.
[0049] For ease of explanation, let t be a TCR sequence, p be a peptide sequence, and x=(t,p) be a TCR-peptide pair.i ,y i )}(i=1,2,...,n), where n is the size of the dataset D and x i represents a TCR-peptide pair, and y i is either 1 (indicating a positive pair) or 0 (indicating a negative pair). train 302, test dataset D test The goal is to train a model 312 that performs well on the ensemble (not shown). train 302 and D test is a partition of the dataset D. train The above-mentioned data shortage problem in 302 is test Therefore, to further improve performance, the present invention uses the TCR dataset 304 (D TCRdb :{t j}), where j=1,2,...,N, and N is D TCRdb represents the number of TCRs 310 in 304. N>>n, and D TCRdb TCR310 in 304 is D train It can be assumed that the peptide has no known interactions with the
[0050] The TCR-peptide model 312 (ERGO-I) discussed herein can be used as a base model for illustrative purposes and experimental results, but it should be understood that other types of TCR-peptide models and / or modelers can be used in accordance with various embodiments of the present invention. ERGO-II is an improvement over ERGO-I by further considering auxiliary information (e.g., the α chain of CDR3, V and J genes, MHC type, and T cell type). However, ERGO-I is used herein to illustrate that the present invention performs physical modeling between TCRs and peptides, thereby enabling fine-tuning and improving any machine learning model for predicting bimolecular interactions. ERGO-I is a general framework applicable to any protein-protein interaction prediction, whereas ERGO-II is applicable only to TCR-peptide interaction prediction.
[0051] The present invention is not limited to TCR-peptide interactions, and therefore, models such as ERGO-I (or the like) can be utilized as a base model in accordance with embodiments of the present invention. While the system and method 300 are described herein below as utilizing the ERGO-I model as a base model, it should be understood that, as noted above, the principles of the present invention can be applied in accordance with embodiments of the present invention using any type of model as a base model.
[0052] Referring now to FIG. 4, an exemplary system and method 400 for training a deep learning model for predicting T cell receptor (TCR)-peptide interactions having multiple encoders is shown in accordance with an embodiment of the present invention.
[0053] According to various embodiments, and for ease of explanation, a base model for fair experimental comparison (e.g., a TCR-peptide model, applicable to any protein-protein interaction prediction) is hereinafter referred to as ERGO. In some embodiments, ERGO is
number
number
number
number
number
number
number
number
[0054] Referring now to FIG. 5, a system and method 500 for predicting T cell receptor (TCR)-peptide interactions by training a deep learning model and calculating docking energies is illustratively depicted in accordance with an embodiment of the present invention.
[0055] In various embodiments, the supervised training dataset D train To remedy the lack of diverse TCR502 and peptide 512 pairs in the training dataset D, the present invention provides, in accordance with an embodiment of the present invention, train The physical properties between the auxiliary TCR and the peptide can be exploited to expand the range of binding.
[0056] In some embodiments, for a given sequence of TCR 502 and peptide 512, a sequence analyzer 504, 514 (e.g., BLASTp) can be utilized to determine the MSA for the TCR in block 506 and / or the MSA for the peptide in block 516, in accordance with aspects of the present invention. In various embodiments, a large TCR database D with diverse TCRs 502 can be used. TCRdb However, these TCR502 are D train The interaction between the peptide 512 in the TCR 502 is unknown. The docking energy 524 between the TCR 502 and the peptide 512 can be selected as an indicator of the interaction, and the docking energy 524 reflects the binding affinity between the molecules by treating the molecules as rigid bodies. Docking of the peptide 512 with the TCR 502 can determine the arrangement of the two rigid bodies with minimal energy by moving the peptide 512 around the surface of the TCR 502, and a relatively small docking energy can indicate a positive pairing between the given TCR 502 and the peptide 512 according to an embodiment of the present invention.
[0057] Note that docking 522 (e.g., using HDock, etc.) is physics-based modeling that can first utilize known structures of TCR 502 and peptide 512. In some embodiments, D TCRdb TCR sequence t' sampled from, and D trainGiven a peptide sequence p' from
[0000] , the structure of TCR 510 t' and the structure of peptide 520 p can be constructed by using a sequence analyzer 504, 514 (e.g., BLASTp) to find homologous sequences with known structures. Docking can be considered a computational technique developed to predict the structure of protein complexes (e.g., dimers of two molecules). Docking can explore the configuration of the complex by minimizing an energy scoring function 524, and the determined final docking energy 524 between the TCR and the peptide can be used as a surrogate binding label for this TCR-peptide pair according to an embodiment of the present invention.
[0058] For ease of explanation, HDock will be described as the docking algorithm utilized; however, any other docking algorithm or method can be utilized in accordance with embodiments of the present invention. For unstructured TCR / peptide sequences, HDock can first use a fast protein sequence search algorithm to search for a multiple sequence alignment (MSA) of the target sequence and its corresponding structure in the Protein Data Bank (PDB). HDock can then perform docking using the structure constructed from the MSA and the structure of a known homologous sequence. The learning algorithm can utilize the final docking score as a surrogate label for the TCR-peptide pair and can utilize a threshold to divide the TCR-peptide pair into negative pairs, positive pairs, and other categories. This is described in more detail below.
[0059] In blocks 508 and 518, MODELLER 508, 518 can be utilized to build structures of TCR 510 and peptide 520 in accordance with aspects of the present invention. In some embodiments, the MSA and corresponding structures from the Protein Data Bank (PDB) are utilized by MODELLER 508, 518 to build the TCR / peptide structure. Finally, with the given structures of TCR 510 and peptide 520, docking (e.g., using HDock) of the TCR with the peptide can be performed in block 522 to calculate the docking energy in block 524.
[0060] In some embodiments, once the structures of the TCR 510 and peptide 520 have been determined, docking (e.g., using HDock, etc.) can be performed to dock the TCR and peptide in block 522 in accordance with aspects of the present invention. In this manner, for example, 80K TCR-peptide pairs with docking energy scores 524 can be generated. Pairs can be pseudo-labeled with the bottom 25% energy scores indicating positive pairs and the top 25% energy scores indicating negative pairs. A dataset is thus generated, and the docking energy D dock It can contain pseudo labels by x',y'∈D dock , y' is the pseudo-label by docking, and the learning objective can be expressed as:
number
[0061] Referring now to FIG. 6, a system and method 600 for predicting T cell receptor (TCR)-peptide interactions by training a conditional variational autoencoder (cVAE) for TCR generation and classification is illustratively depicted in accordance with an embodiment of the present invention.
[0062] In various embodiments, the cVAE can be trained on TCR generation conditions for peptides using various datasets (e.g., MCPAS, VDJdb, etc.) using peptide sequence 602, p, and TCR 608 as inputs. A peptide encoder 604 can be utilized to generate latent peptides in block 606, and a TCR encoder 610 can be utilized to generate latent TCRs in block 612. The latent peptides 606 and latent TCRs 612 are generated using a function
number
[0063] Referring now to FIG. 7, there is illustratively depicted a system and method 700 for predicting T cell receptor (TCR)-peptide interactions by augmenting a training dataset using pseudo-labeling, in accordance with an embodiment of the present invention.
[0064] In one embodiment, the classifier 708 (e.g., an ERGO classifier) uses the limited labeled data (e.g., from McPAS-TCR) D from block 702 for the TCR 704 and peptides 706. trainThe initial training model 708 may be pre-trained using the dc parameter set 706 to generate actual labels 710. The parameters may be copied in block 709 for use in a subsequent classifier (e.g., ERGO) model in block 718. ... initial training model 708 may be pre-trained using the dc parameter set 706 to generate actual labels 710. The initial training model 708 may be pre-trained using the dc parameter set 706 to generate actual labels 710. The initial training model 708 may be pre-trained using the dc parameter set 706 to generate actual labels 710. The initial training model 708 may be pre-trained using the dc parameter set 706 to generate actual labels 710. The initial training model 708 may be pre-trained using the dc parameter set 706 to generate actual labels 710. The initial training model 708 may be pre-trained using the dc parameter set 706 to generate TCRdb The data from can be used as a training model to generate pseudo scores and / or pseudo labeled TCR-peptide pairs (e.g., TCR' 714, peptide 716) by a classifier 718 (e.g., ERGO) in block 720.
[0065] In some embodiments, parameters can be copied in block 719 for use by a next classifier (e.g., ERGO) model in block 728. The model 728 uses the original dataset D from block 702. train sample data 722 from block 712, and the augmented pseudo-labeled dataset D TCRdb The classifier 728 can be retrained using data from block 724. Inputs to the classifier 728 can include the TCR / TCR' from block 724 and the peptides from block 726. A prediction (pred'|pred) can be generated in block 730, and the real labels 710 and pseudo labels 720 can be utilized along with the prediction 730 to generate a final docking score as output y'|y in block 732 based on the combined sampled dataset input from block 722, including the TCR' / TCR from block 724 and the peptides from block 726, in accordance with an embodiment of the present invention.
[0066] Referring now to FIG. 8 , a system and method 800 for predicting T cell receptor (TCR)-peptide interactions by training a conditional variational autoencoder (cVAE) for TCR generation and classification and augmenting the training dataset using pseudo-labeling is illustratively depicted in accordance with an embodiment of the present invention.
[0067] In one embodiment, TCRs 802 from multiple databases (e.g., McPAS-TCR, BM_data_CDR3s-TCR, etc.) may be received by a TCR encoder 804 and generate latent TCRs in block 806. In accordance with aspects of the present invention, the latent TCRs 806 may be modeled using a mean and variance function (e.g., as a Gaussian function), which may be utilized for concatenation / reparameterization in block 824, resulting in a latent variable Z t It can be combined with 828 and used as input for KL divergence 826.
[0068] In some embodiments, a TCR decoder 808 can be utilized to generate a TCR, t', in block 810, which can be classified using a TCR classifier 820, which can receive an additional input of peptides 822 from one of the databases discussed above (e.g., McPAS-TCR). The TCR classifier 820 can generate a new training set 830 for additional training to further fine-tune and improve the performance of the classifier 820 for TCR-peptide binding prediction based on the generative model and pseudo-labeling, in accordance with aspects of the present invention.
[0069] According to various embodiments, learning from physical modeling can effectively extend the training dataset, but it should be noted that the success of learning also depends on the quality of the physical modeling. A model can be trained such that the auxiliary learning from physical modeling is optimized for the primary learning objective, e.g., by meta-learning to minimize validation loss. This meta-learning algorithm can introduce time-consuming gradient-on-gradient learning. However, according to aspects of the present invention, to improve processing speed and accuracy, instead of minimizing validation loss, it can be approximated by minimizing the training loss of the current batch (e.g., optimizing learning from physical modeling such that learning from this auxiliary objective reduces the training loss of the current batch).
[0070] For example, for each training iteration of batch (x,y), the loss
number
number
number
number
[0071] Referring now to FIG. 9, there is illustratively depicted a method 900 for predicting and classifying T cell receptor (TCR)-peptide interactions by training and utilizing a neural network, in accordance with an embodiment of the present invention.
[0072] In some embodiments, at block 902, peptide embedding vectors and TCR embedding vectors can be concatenated from a dataset of positive and negative binding TCR-peptide pairs. At block 904, a deep neural network (DNN) classifier can be trained to predict a binary binding score, and at block 906, an autoencoder (e.g., Wasserstein) can be trained by combining unlabeled TCRs from a large TCR database (e.g., TCRdb) with labeled TCRs from the training data. At block 908, peptides can be docked to TCRs using physical modeling. At block 910, the docking energy can be utilized to generate additional positive and negative TCR-peptide pairs as training data and fine-tune an existing TCR-peptide interaction prediction classifier using the training data.
[0073] In some embodiments, new TCRs can be generated based on the autoencoder trained in block 912, and the new TCRs paired with the selected peptides can be labeled using a physical model and pseudo-labeling to generate new training data, which can be used to further fine-tune the existing TCR-peptide interaction prediction classifier. According to various embodiments, pseudo-labeling (e.g., self-training) can correspond to training a first (e.g., supervised) model on a labeled dataset and using the trained first (e.g., supervised) model to pseudo-label an unlabeled dataset. According to aspects of the present invention, a new model can be trained from a joint dataset of the original labeled dataset and the augmented pseudo-labeled dataset.
[0074] In some embodiments, learning using the trained DNN from block 904 may include using unlabeled examples by matching model predictions with weakly augmented and heavily augmented examples, and learning pseudo-labels by gradient-based metalearning (e.g., the pseudo-labels may be optimized to minimize validation loss for the target task). The present invention, in accordance with aspects of the present invention, can be viewed as a semi-supervised problem by using a large database of TCR sequences (e.g., TCRdb) and assigning pseudo-scores to unknown pairs (e.g., by a supervised model) and / or assigning pseudo-labels from determined properties of physical modeling of TCR-peptide pairs. In some embodiments, the steps of blocks 908, 910, 912, and / or 914 may be repeated until convergence at block 916, in accordance with aspects of the present invention.
[0075] Referring now to FIG. 10, there is illustratively depicted an exemplary system 1000 for predicting and classifying T cell receptor (TCR)-peptide interactions by training and utilizing neural networks in accordance with an embodiment of the present invention.
[0076] In some embodiments, one or more database servers 1002 may contain large amounts of unlabeled and / or labeled TCRs and / or peptides (or other data) for use as input in accordance with aspects of the present invention. According to aspects of the present invention, a peptide encoder 1004 may be utilized to generate latent peptides, and a TCR encoder / decoder 1006 may be utilized to generate latent TCRs (encoders) and new TCRs, l' (decoders). A neural network 1008 may be utilized, which may include a neural network training / learning unit 1010, which may include one or more processor units 1024, for performing training of one or more models (e.g., ERGO), and a sequence analyzer 1012 (e.g., BLASTp), which may be utilized to determine TCR MSAs and / or peptide MSAs for one or more sequences of TCRs and / or peptides.
[0077] In various embodiments, an autoencoder 1014 (e.g., Wasserstein) can be trained by combining unlabeled TCRs from a large TCR database (e.g., TCRdb) with labeled TCRs from training data and utilized to generate new TCRs using a TCR generator / classifier 1016. In accordance with aspects of the invention, the TCR generator / classifier 1016 can classify one or more new TCRs, / ′, generated using potential TCRs from the TCR encoder / decoder 1006 and can use gradients to constrain the generated TCRs to positively bind to the conditional peptide. A MODELLER 1018 can be utilized to construct structures of the TCRs and peptides; in some embodiments, corresponding structures from the MSA and the Protein Data Bank (PDB) can be utilized by the MODELLER 1018 to construct TCR / peptide structures in accordance with aspects of the invention.
[0078] In some embodiments, a TCR-peptide docker 1020 (e.g., HDock) may be utilized to perform docking of TCRs with peptides using the TCR and peptide structures constructed by the MODELLER 1018 to calculate docking energies in accordance with aspects of the present invention. In one embodiment, the label generator 1022 uses limited labeled data (e.g., from McPAS-TCR) D for TCRs generated by, for example, the TCR generator / classifier 1016 (e.g., the ERGO classifier). train In one embodiment, in accordance with aspects of the present invention, the label generator 1022 may generate actual labels for the TCRs and peptides by pre-training the classifier 1016 using, for example, the initial trained model as a teacher model to generate pseudo-scores, and may generate actual labels for the TCRs and peptides for which the labels are generated using an auxiliary dataset (e.g., from CRD3s-TCRs) D. TCRdb Pseudo-labels for TCRs generated by the TCR generator / classifier 1016 (e.g., ERGO classifier) can be generated by pseudo-labeling TCR-peptide pairs using data from.
[0079] In the embodiment shown in FIG. 10, the elements are interconnected by a bus 1001. However, in other embodiments, other types of connections may be used. Furthermore, in one embodiment, at least one of the elements of system 1000 is processor-based and / or logic circuitry and may include one or more processor units 1024. Furthermore, while one or more elements may be shown as separate elements, in other embodiments, these elements may be combined into a single element. The reverse is also applicable; while one or more elements may be part of another element, in other embodiments, one or more elements may be implemented as independent elements. These and other variations of the elements of system 1000 are readily determined by one of ordinary skill in the art given the teachings of the present principles provided herein, while maintaining the spirit of the present principles.
[0080] References in the specification to "one embodiment" or "one embodiment" of the present invention, as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with an embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the phrase "in one embodiment" or "in one embodiment," as well as any other variations thereof, in various places throughout this specification do not necessarily all refer to the same embodiment. However, it should be understood that features of one or more embodiments may be combined given the teachings of the present invention provided herein.
[0081] For example, in the case of "A / B," the use of any of the following " / ," "and / or," "at least one," such as "A and / or B" or "at least one of A and B" will be understood to be intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), the selection of only the first and third listed alternatives (A and C), the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). This can be expanded as many times as there are listed items.
[0082] The foregoing is understood in all respects to be illustrative and exemplary, but not restrictive, and the scope of the invention disclosed herein is to be determined not from the detailed description, but from the claims which are interpreted in accordance with the full breadth permitted by the patent laws. It will be understood that the embodiments shown and described herein are merely exemplary of the invention, and that those skilled in the art could make various modifications without departing from the scope and spirit of the invention. Various other feature combinations could be implemented by those skilled in the art without departing from the scope and spirit of the invention. Having thus described aspects of the invention with the detail and particularity required by the patent laws, what is desired to be claimed and protected by Letters Patent is set forth in the appended claims.
Claims
1. 1. A method for predicting T cell receptor (TCR)-peptide interactions, comprising: training a deep learning model to predict the TCR-peptide interaction, determining a multiple sequence alignment (MSA) of a plurality of TCR-peptide pair sequences from the dataset of TCR-peptide pair sequences using a sequence analyzer; constructing TCR and peptide structures using the MSA and corresponding structures from the Protein Data Bank (PDB) using MODELLER; generating an expanded TCR-peptide training dataset based on docking energy scores determined by docking peptides to the TCR using physical modeling based on the TCR structure and peptide structures constructed using the MODELLER; training the deep learning model, which includes classifying and labeling TCR-peptide pairs as positive pairs or negative pairs using pseudo-labels based on the docking energy scores; and iteratively retraining the deep learning model based on the expanded TCR-peptide training dataset and the pseudo-labels until convergence.
2. 2. The method of claim 1, wherein the TCR-peptide pair dataset comprises positive and negative binding TCR-peptide pairs.
3. 2. The method of claim 1, wherein classifying and labeling the TCR-peptide pairs further comprises pseudo-labeling the TCR-peptide pairs with the top x percent of the docking energy scores as negative pairs and the bottom y percent of the docking energy scores as positive pairs.
4. 2. The method of claim 1, further comprising concatenating peptide-embedded vectors and TCR-embedded vectors from the dataset of sequences of TCR-peptide pairs.
5. 10. The method of claim 1, further comprising training an autoencoder by combining unlabeled TCRs from a TCR database with labeled TCRs from the expanded TCR-peptide training dataset.
6. 2. The method of claim 1, wherein the deep learning model is trained based on a standard cross-entropy loss from sequences of the plurality of TCR-peptide pairs, a divergence loss from the pseudo-labeled TCR-peptide pairs, and a cross-entropy loss based on physical properties between TCRs and peptides using physical modeling.
7. Final total loss (L total )but, [Equation 1] where L labeled represents the standard cross-entropy loss from the sequences of the plurality of TCR-peptide pairs, and L pseudo-labeled represents the divergence loss from the pseudo-labeled TCR-peptide pair, and L physical The method of claim 6 , wherein: ĥ represents the cross-entropy loss based on physical properties.
8. 1. A system for predicting T cell receptor (TCR)-peptide interactions, comprising: A processor operably coupled to a non-transitory computer-readable storage medium, the processor comprising: training a deep learning model to predict the TCR-peptide interaction, determining a multiple sequence alignment (MSA) of a plurality of TCR-peptide pair sequences from the dataset of TCR-peptide pair sequences using a sequence analyzer; constructing TCR and peptide structures using the MSA and corresponding structures from the Protein Data Bank (PDB) using MODELLER; generating an expanded TCR-peptide training dataset based on docking energy scores determined by docking peptides to the TCR using physical modeling based on the TCR structure and peptide structures constructed using the MODELLER; training the deep learning model, which includes classifying and labeling TCR-peptide pairs as positive pairs or negative pairs using pseudo-labels based on the docking energy scores; iteratively retraining the deep learning model based on the expanded TCR-peptide training dataset and the pseudo-labels until convergence.
9. 9. The system of claim 8, wherein the TCR-peptide pair dataset comprises positive and negative binding TCR-peptide pairs.
10. 9. The system of claim 8, wherein classifying and labeling the TCR-peptide pairs further comprises pseudo-labeling the TCR-peptide pairs with the top x percent of the docking energy scores as negative pairs and the bottom y percent of the docking energy scores as positive pairs.
11. 9. The system of claim 8, wherein the processor is further configured to concatenate peptide embedding vectors and TCR embedding vectors from the dataset of sequences of TCR-peptide pairs.
12. 10. The system of claim 8, wherein the processor is further configured to train an autoencoder by combining unlabeled TCRs from a TCR database with labeled TCRs from the expanded TCR-peptide training dataset.
13. 9. The system of claim 8, wherein the deep learning model is trained based on a standard cross-entropy loss from sequences of the plurality of TCR-peptide pairs, a divergence loss from the pseudo-labeled TCR-peptide pairs, and a cross-entropy loss based on physical properties between TCRs and peptides using physical modeling.
14. Final total loss (L total )but, [Equation 2] where L labeled represents the standard cross-entropy loss from the sequences of the plurality of TCR-peptide pairs, and L pseudo-labeled represents the divergence loss from the pseudo-labeled TCR-peptide pair, and L physical The system of claim 13 , wherein:
15. 1. A non-transitory computer readable storage medium comprising a computer readable program operatively coupled to a processor device for predicting T cell receptor (TCR)-peptide interactions, the computer readable program, when executed on a computer, causing the computer to: training a deep learning model to predict the TCR-peptide interactions, determining a multiple sequence alignment (MSA) of a plurality of TCR-peptide pair sequences from the dataset of TCR-peptide pair sequences using a sequence analyzer; constructing TCR and peptide structures using the MSA and corresponding structures from the Protein Data Bank (PDB) using MODELLER; generating an expanded TCR-peptide training dataset based on docking energy scores determined by docking peptides to the TCR using physical modeling based on the TCR structure and peptide structures constructed using the MODELLER; training the deep learning model, which includes classifying and labeling TCR-peptide pairs as positive pairs or negative pairs using pseudo-labels based on the docking energy scores; and iteratively retraining the deep learning model based on the expanded TCR-peptide training dataset and the pseudo-labels until convergence.
16. 16. The non-transitory computer-readable storage medium of claim 15, wherein the TCR-peptide pair dataset comprises positive and negative binding TCR-peptide pairs.
17. 16. The non-transitory computer-readable storage medium of claim 15, wherein classifying and labeling the TCR-peptide pairs further comprises pseudo-labeling the TCR-peptide pairs with the top x percent of the docking energy scores as negative pairs and the bottom y percent of the docking energy scores as positive pairs.
18. 16. The non-transitory computer-readable storage medium of claim 15, further comprising concatenating peptide-embedded vectors and TCR-embedded vectors from the dataset of sequences of TCR-peptide pairs.
19. 16. The non-transitory computer-readable storage medium of claim 15, wherein the deep learning model is trained based on a standard cross-entropy loss from sequences of the plurality of TCR-peptide pairs, a divergence loss from the pseudo-labeled TCR-peptide pairs, and a cross-entropy loss based on physical properties between TCRs and peptides using physical modeling.
20. Final total loss (L total )but, [Equation 3] where L labeled represents the standard cross-entropy loss from the sequences of the plurality of TCR-peptide pairs, and L pseudo-labeled represents the divergence loss from the pseudo-labeled TCR-peptide pair, and L physical 20. The non-transitory computer-readable storage medium of claim 19, wherein: represents the cross-entropy loss based on physical properties.
Citation Information
Patent Citations
Prediction method and prediction system for G-protein-coupled receptor-ligand interaction relationship
CN108647487A
Method for predicting binding free energy of protein and ligand based on progressive neural network
CN110910951A
Method for predicting protein and ligand molecule binding free energy based on convolutional neural network
CN112185458A
Image enhancement method, image enhancement device, electronic equipment and storage medium
CN113393391A
Machine learning and molecular simulation-based methods for enhancing binding and activity predictions
JP2021515233A