Generation of Peptide Bond Motifs
A reinforcement learning model trains a neural network to optimize peptide binding to MHC proteins, addressing the challenge of identifying binding peptides by generating robust motifs for immune response induction.
Patent Information
- Application Number
- JP2024566479
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-18
- Filing Date
- 2023-05-19
- Publication Date
- 2025-06-03
AI Technical Summary
Identifying peptides that bind to specific major histocompatibility complex (MHC) proteins is challenging due to the large search space of possible peptides.
A reinforcement learning model is used to train a neural network for a peptide mutation policy, optimizing peptides to bind to specific MHC proteins by stepwise changing amino acids, and calculating binding motifs to screen a library of peptides.
The method effectively generates robust binding motifs with high correlation to experimentally obtained motifs, enabling the identification of peptides that can induce an immune response against specific pathogens or tumors.
Smart Images

Figure 2025517173000001_ABST
Abstract
Description
Technical Field
[0001] Related Application Information This application claims the priority of U.S. Patent Application No. 63 / 344,081, filed on May 20, 2022, the entire disclosure of which is hereby incorporated by reference in its entirety.
Background Art
[0002] The present invention relates to the identification of binding peptides, and more specifically, to a reinforcement learning model for generating binding peptides. Description of Related Art
[0003] Immunotherapy aims to enhance the patient's immune system against pathogens and tumor cells. The immune response is triggered when immune cells recognize foreign peptides presented by major histocompatibility complex (MHC) proteins on the cell surface. For recognition to occur, the foreign peptide needs to bind to MHC class I proteins. The resulting peptide-MHC complex interacts with the T cell receptor. These interactions can be utilized to generate peptide-based vaccines for disease prevention.
[0004] However, identifying peptides that bind to specific MHC proteins is a major challenge because the search space for all possible peptides is prohibitively large.
Summary of the Invention
[0005] The method for peptide generation includes training a neural network of a peptide mutation policy using reinforcement learning that includes a peptide presentation score as a reward. A new peptide is generated using the peptide mutation policy. Using the new peptide, the binding motif of the major histocompatibility complex (MHC) is calculated. According to the binding motif, a library of peptides is screened.
[0006] The peptide generation system includes a hardware processor and a memory storing a computer program. When executed by the hardware processor, the computer program causes the hardware processor to train a neural network of a peptide mutation policy using reinforcement learning that includes a peptide presentation score as a reward. Using the peptide mutation policy, a plurality of new peptides are generated. Using the plurality of new peptides, MHC binding motifs are calculated. According to the binding motifs, a plurality of library peptides are screened.
[0007] These and other features and advantages will become apparent from the following detailed description of its exemplary embodiments, read in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0008] The present disclosure provides details in the following description of preferred embodiments with reference to the following figures.
[0009]
Figure 1
[0010]
Figure 2
[0011]
Figure 3
[0012]
Figure 4
[0013]
Figure 5
[0014]
Figure 6
Mode for Carrying Out the Invention
[0015] Foreign peptides that bind to a specific major histocompatibility complex (MHC) may be identified using a reinforcement learning model. This model learns a mutation policy that optimizes peptides by stepwise changing amino acids, increasing the likelihood that the mutated peptides will be presented by a specific MHC protein. The generated motifs are robust, with random initial peptides leading to the same motif after stepwise mutations and having a high correlation with experimentally obtained motifs.
[0016] Referring to FIG. 1, a diagram of peptide-MHC protein binding is shown. Peptide 102 is shown binding to MHC protein 104, and the complementary two-dimensional interface in the figure suggests the complementary shapes of these three-dimensional structures. MHC protein 104 may be attached to cell surface 106.
[0017] MHC is a region on the DNA strand that encodes cell surface proteins used by the immune system. MHC molecules are used by the immune system and contribute to the interaction between white blood cells and other cells. For example, MHC proteins affect organ compatibility during transplantation and are also important for vaccine development.
[0018] On the one hand, a peptide may be part of a protein. When a pathogen presents a peptide recognized by an MHC protein, the immune system triggers a reaction to destroy the pathogen. Therefore, by finding the peptide structure that binds to the MHC protein, it may be possible to intentionally induce an immune response without taking in the pathogen itself into the body. In particular, if an existing peptide that binds well to MHC protein 104 is provided, a new peptide 102 may be automatically identified according to the desired characteristics and attributes.
[0019] The interaction between peptides and MHC plays a role in cellular immunity, regulation of immune responses, and transplant rejection reactions. Predicting the binding of peptides to proteins is useful for the search and design of peptides that can be used in vaccines and other pharmaceuticals. Given a library of known peptides, new peptide sequences can be generated using a mutagenesis policy. The resulting mutant peptides may be within a threshold number of amino acid differences from the peptide library. If the peptide library is derived from a specific pathogen such as a virus or tumor sample, the mutant peptides can be used to target a specific pathogen or tumor. This makes it possible, for example, to identify and target a specific cancer in an individual.
[0020] Therefore, given a specific genome (e.g., one sequenced from tumor cells), peptide sequences can be extracted to generate a library of peptides that uniquely identify the pathogen. By targeting this library, peptides that bind to MHC present on the cell surface can be screened / selected, and an immune response can be induced to kill the pathogen or tumor cells.
[0021] For this purpose, a deep neural network can be trained using a training dataset to predict a peptide presentation score when an MHC allele sequence and a peptide sequence are given. The peptide presentation score may be, for example, a combination of peptide-MHC binding affinity and antigen processing score.
[0022] Based on the trained peptide presentation model, deep reinforcement learning can be used to generate binding peptide motifs. The pre-trained presentation score prediction model can be used to define a reward function starting from random peptides. The deep reinforcement learning system may be trained to learn an appropriate peptide mutation policy by converting a given random peptide into a peptide with a high presentation score.
[0023] Applying a reinforcement learning system to this process, the "state" is interpreted as a specific MHC allele sequence and peptide sequence, and the "action" is interpreted as the editing of the peptide sequence. Such editing may replace the current amino acid at a specific position in the peptide sequence with a new amino acid.
[0024] The amino acid sequence is embedded using a one-dimensional convolutional layer on top of the concatenated amino acid embedding and the fully connected layer of the neural network model to generate MHC allele expression. The bidirectional long short-term memory (LSTM) layer further processes the amino acid embedding to obtain peptide representation. The deep policy network can learn the conditional probabilities of various actions when a state is given. At each time step, a positive reward value may be assigned if the peptide presentation score of the mutant peptide based on the action increases beyond a threshold, and a negative reward value may be assigned otherwise.
[0025] The peptide scoring model can be trained to accept a peptide p and an MHC protein m as inputs and generate an output score r(p,m) representing the binding affinity between the peptide p and the protein m, particularly the probability that the peptide p is presented on the cell surface by the protein m. In some cases, the presentation score may be a composite score of antigen processing prediction and binding affinity prediction, where the former predicts the probability that the peptide is delivered to the endoplasmic reticulum by a transporter associated with the antigen processing protein complex and where the peptide can bind to the MHC protein.
[0026] The mutant policy network may also be trained. The mutant policy network guides how the peptide sequence is changed. As will be described in detail below, this policy network guides a reinforcement learning system, receives a peptide and an MHC protein as inputs, and outputs a change or "mutation" of the peptide. The policy network selects a mutation for the purpose of improving the presentation score of the mutated peptide to the MHC protein. A library of peptides can be sampled, and this sampling can be performed randomly. The sampled peptides may then be mutated according to the mutant policy.
[0027] In this framework, a peptide can be represented as an amino acid sequence p = <o 1 ,o 2 ,...,o l >. Here, o is one of the set of natural amino acids, and l is the length of the sequence, for example, in the range from 8 to 15. The reinforcement learning agent explores the peptide mutation environment for high presentation peptide generation. Thus, when given an input pair (p, m), the reinforcement learning agent explores and utilizes the peptide mutation environment by repeatedly mutating the peptide and observing the resulting presentation score. The agent thereby learns a mutation policy π(·) that repeatedly mutates the amino acids of any peptide to generate a high presentation score. In this way, the peptide mutation environment and the mutant policy network are determined.
[0028] With the peptide mutation environment, the reinforcement learning agent can gradually improve the mutant policy by performing peptide mutations through trial and error and adjusting the parameters of the mutant policy network. During learning, the reinforcement learning agent continues to mutate the peptide and continues to determine its presentation score as the reward signal. The reward helps to reinforce the agent's mutation behavior, and mutation behaviors that produce high presentation scores are encouraged.
[0029] The mutant environment includes a state space, an action space, and a reward function. The state includes the current mutant peptide and the MHC protein. The actions and rewards represent the possible mutant activities to be executed by the reinforcement learning agent, and respectively generate new presentation scores for the mutated peptides.
[0030] The state of the environment can be defined as s at time t of the pair (p, m). The MHC protein can be represented, for example, as a pseudo-sequence consisting of 34 amino acids, and each amino acid may contact the bound peptide, for example, within a distance of 4.0 Å. In the case of a peptide of length l and an MHC protein, the state s t can be a tuple s t =(E t 、E p 、E m ), where E p and E m are the encoding matrices of the peptide and the MHC protein, respectively. The state s 0 may be initialized by sampling a peptide sequence from a library and using an MHC class I protein. During training, any appropriate peptide sequence and MHC protein can be used. The terminal state s T can be defined as a state with a maximum time step T or a state with a presentation score exceeding a predetermined threshold σ. When the terminal state s T is reached, the mutation of the peptide stops.
[0031] To optimize the peptide by replacing one amino acid with another, multiple discrete action spaces can be defined. At time t, given a peptide p t , the action of the reinforcement learning agent may be to determine the position of the amino acid O i to be replaced and predict the type of the new amino acid at that position. The reward function guides the optimization of the reinforcement learning agent, and only the terminal state can receive a reward from the peptide mutant environment. The final reward is determined as r(p T , m), and the peptide pT is in the final state s T and is present in
[0032] In one exemplary reward function, the score may be a composite score of an antigen processing prediction and a binding affinity prediction. The former predicts the probability that a peptide is delivered to the endoplasmic reticulum by a transporter associated with an antigen processing protein complex, where the peptide can bind to an MHC protein. The latter predicts the binding strength between the peptide and the MHC protein. The higher the presentation score, the higher the antigen processing and binding affinity scores, indicating a higher probability that the peptide is presented on the cell surface by a particular MHC protein.
[0033] Referring now to FIG. 2, a method for generating a binding peptide is shown. Block 202 determines a scoring function, and the output of the scoring function characterizes the quality of the binding between a peptide sequence and an MHC allele sequence. This score may be implemented as a presentation score that provides a combination of peptide-MHC binding affinity and an antigen processing score. The scoring function can be implemented, for example, as a deep neural network trained on a publicly available peptide dataset or can reflect a pre-trained scoring model.
[0034] The peptide mutation policy is trained (204) based on the scoring function, for example, using a deep reinforcement learning system. The peptide mutation policy receives a peptide sequence as input and generates an output peptide that includes one or more changes (referred to herein as mutations). A reward function is defined using the scoring function, and starting from the peptide sequence, the deep reinforcement learning system is trained to learn an appropriate peptide mutation policy that converts a particular input peptide into a peptide with a high presentation score.
[0035] Block 206 generates a binding peptide based on the input peptide using a scoring function and a trained peptide mutation policy. The input peptide may be randomly sampled from any suitable dataset at block 210. Block 212 uses the sampled peptide as input and applies the trained peptide mutation policy to generate a new peptide sequence.
[0036] Block 214 calculates the binding motifs for all MHCs of interest, including rare MHCs for which there is little experimental data. The binding motif may include a position weight matrix that includes the probability of an amino acid at each motif position. Given the binding motif for a particular MHC, peptides within the array library can be screened at block 216.
[0037] In the first example of peptide screening, for each position within the binding motif, a weighted block substitution matrix (BLOSUM) representation of the amino acid can be calculated, and for example, using the amino acid probabilities within the position weight matrix for each position, the BLOSUM representation of the amino acid can be weighted. Subsequently, a weighted sum is calculated as the final representation for each position. Next, a pairwise Euclidean distance can be used between the calculated motif BLOSUM representation and the BLOSUM representation of the screening peptide. In the second example of peptide screening, the log-likelihood of the peptide can be calculated based on the position weight matrix of the motif.
[0038] To learn the peptide mutation policy of block 204, the reinforcement agent learns to mutate one amino acid at a time in each step in the input peptide sequence with the aim of maximizing the presentation score of the mutated peptide. Both the peptide and the MHC protein can be encoded into a distributed embedding space, and then the mapping between the embedding space and the mutation policy can be learned by optimization of gradient descent.
[0039] Multiple encoding methods may be used to represent amino acids within a peptide sequence and MHC protein. Each amino acid is represented by a concatenated encoding vector from BLOSUM
Number
Number
Number
Number
Number
Number
Number
[0040] Each amino acid O i in the peptide sequence p is, for example,
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0041] To embed the MHC protein into a continuous latent vector, the encoding matrix E m is a vector
Number
Number
Number
Number
[0042] At each time step t, the peptide sequence p t is the potential embedding
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0043] The objective function for learning the mutation policy is defined as follows.
Number
Number
Number
Number
Number
Number
Number
[0044] The value function V(s t ) uses a multi - layer perceptron to predict the future return value of the current state s
Number
Number
Number
Number
Number
Number
[0045] To stabilize training and improve performance, an expert policy π ept can be derived from existing data. For each MHC protein m with sufficient binding peptide data, the amino acid distribution <p 1 (o|m), p 2 (o|m),..., p l (o|m)> of peptides of length l can be determined. When a peptide p is given, the position I is selected as follows.
Number
Number
Number
[0046] After determining the position, an amino acid can be sampled from the distribution
Number
[0047] Using the expert policy, the policy network can be pre-trained. The objective function for pre-training can minimize the following cross-entropy loss.
Number
Number
Number
[0048] To increase the diversity of the generated peptides, a non-deterministic policy can be used to generate diverse activities. Such a policy can increase the exploration in a large state space, thereby finding diverse appropriate activities.
[0049] To promote exploration, entropy regularization can be included in the objective function. To explicitly enforce the learning of diverse activities of the policy, a diversity-promoting experience buffer can be used to save the trajectories that may generate qualified peptides. In each iteration, the visited state and activity pairs of the mutant trajectories of the qualified peptides can be added to the buffer. The state and activity pairs are maintained with low-frequency activities, and those with high-frequency activities are deleted so that the buffer is not occupied by high-frequency activities. A batch of state and activity pairs containing low-frequency activities can be sampled from the buffer.
[0050] The cross-entropy loss L defined for a batch of pairs of states and actions including low-frequency activities B By including it in the final objective function, the policy network can be encouraged to reproduce low-frequency activities that may bring high rewards. [Number] Here, H is the entropy of the policy network, and α 1 , α 2 , α 3 are pre-determined coefficients.
[0051] Based on this trained DRL system with a pre-trained peptide mutation policy, binding peptides are generated from peptides randomly sampled at block 212. Block 214 calculates the binding motifs (position weight matrices, probabilities of amino acids at each motif position) for all MHCs, including rare MHCs with little experimental data. Block 216 can rapidly screen all peptides in the array library using the generated motifs of a given MHC to identify new antigens. This screening is robust as it is based on binding motifs and takes into account various variations / mutations of peptides in the peptide library, resulting in better results than a single interaction score predicted by a classifier.
[0052] The actual motifs may be characterized from experimental data, and there is an exemplary database with 149 human MHC proteins and 309,963 peptides included in the experimental dataset. In the case of calculated motifs, a predetermined number (e.g., 1,000) of peptides may be generated for each human MHC protein. Generated peptides with a presentation score below a predetermined threshold (e.g., 0.75) may be excluded as they have low binding affinity.
[0053] Referring now to FIG. 3, a method for treating a disease is shown. Block 206 generates a set of binding peptides of MHC proteins as described above. In the case of a specific disease such as a viral infection, block 302 generates a set of peptide vaccine candidates, such as by identifying peptides that may be presented by the infectious agent. When the infectious agent is present in the human body, MHC may use these peptides to recognize the pathogen and trigger an immune response.
[0054] Block 304 uses the binding motif of block 306 to determine the matching score of the vaccine candidate. These matching scores represent the binding affinity between the vaccine candidate and the MHC protein and reflect the ability of the peptide to generate an immune response targeting the pathogen. Block 306 generates a vaccine, such as by generating a new antigen incorporating the selected peptide vaccine candidate, based on the matching score. Then, block 308 administers the vaccine to prevent the disease.
[0055] Next, referring to FIGS. 4 and 5, exemplary neural network architectures are shown, which can be used to implement a part of this model. A neural network is a generalized system, and its function and accuracy are improved by being exposed to additional empirical data. A neural network is learned by being exposed to empirical data. During training, the neural network stores and adjusts a plurality of weights applied to the input empirical data. By applying the adjusted weights to the data, it is possible to identify that the data belongs to a specific class predefined from a set of classes, or to output the probability that the input data belongs to each class.
[0056] Empirical data (also called training data) obtained from a series of examples is formatted as a string of values and supplied as input to a neural network. Each example is associated with a known result or output. Each column is represented as a pair (x, y), where x represents the input data and y represents the known output. The input data can have various data types and may contain multiple different values. The network can have one input node for each value that makes up the input data of an example, and separate weights can be applied to each input value. The input data can be formatted as a vector, an array, or a string, depending on the architecture of the neural network being constructed and trained.
[0057] The neural network "learns" by comparing the neural network output generated from the input data with the known values of the examples and adjusting the stored weights to minimize the difference between the output value and the known value. The adjustment can be done to the stored weights through backpropagation, and the influence of the weights on the output value is determined by calculating the mathematical gradient and adjusting the weights in a way that shifts the output to the minimum difference. This optimization, called gradient descent, is a non-limiting example of how training is performed. A subset of examples with known values that were not used in training can be used to test and validate the accuracy of the neural network.
[0058] During operation, the trained neural network can be used for new data that was not previously used for training or validation through generalization. The adjusted weights of the neural network can be applied to the new data, and the weights estimate a function developed from the training examples. The parameters of the estimated function captured by the weights are based on statistical inference.
[0059] In a layered neural network, nodes are arranged in layers. An exemplary simple neural network has an input layer 420 of source nodes 422 and a single computational layer 430 having one or more computational nodes 432 that also function as output nodes, with a single computational node 432 for each possible category into which input examples can be classified. The input layer 420 can have a number of source nodes 422 equal to the number of data values 412 of the input data 410. The data values 412 of the input data 410 can be represented as a column vector. Each computational node 432 of the computational layer 430 generates a linear combination of weighted values from the input data 410 supplied to the input nodes 420 and applies a non - linear activation function that is differentiable with respect to the sum. An exemplary simple neural network can perform classification for linearly separable examples (e.g., patterns).
[0060] Deep neural networks, such as multi - layer perceptrons, can have an input layer 420 of source nodes 422, one or more computational layers 430 having one or more computational nodes 432, and an output layer 440 having one output node 442 for each category into which input examples may be classified. The input layer 420 can have a number of source nodes 422 equal to the number of data values 412 of the input data 410. The computational nodes 432 of the computational layer 430 are between the source nodes 422 and the output nodes 442 and are not directly observable, so they are also called hidden layers. Each node 432, 442 of the computational layer generates a linear combination of weighted values from the values output from the nodes of the previous layer and applies a non - linear activation function that is differentiable over the range of the linear combination. The weights applied to the values from each previous node are, for example, w 1 , w 2 ,... w n-l , w nIt can be represented by. The output layer provides the overall response of the network to the input data. A deep neural network may be a fully connected case where each node in the computational layer is connected to all nodes in the previous layer, or the connection between layers may be in other configurations. If there are missing links between nodes, the network is called partially connected.
[0061] The training of a deep neural network has two phases: a forward phase where the weights of each node are fixed and the input is propagated through the network, and a backward phase where the error value is backpropagated through the network to update the weight values.
[0062] The computational nodes 432 of one or more computational (hidden) layers 430 perform a non-linear transformation on the input data 412 that generates a feature space. Classes or categories may be more easily separable in the feature space than in the original data space.
[0063] Referring to FIG. 6, an exemplary arithmetic unit 600 according to an embodiment of the present invention is shown. The arithmetic unit 600 is configured to perform enhancement of a classifier.
[0064] The arithmetic unit 600 can be realized as any type of arithmetic unit or computer device that can execute the functions described herein. This includes, but is not limited to, computers, servers, rack-based servers, blade servers, workstations, desktop computers, laptop computers, notebook computers, tablet computers, mobile arithmetic units, wearable arithmetic units, network devices, web devices, distributed arithmetic systems, processor-based systems, and / or consumer electronic devices. Further, or alternatively, the arithmetic unit 600 may be embodied as one or more computing threads, memory threads, or other racks, threads, arithmetic chassis, or other components of physically distributed arithmetic devices.
[0065] As shown in FIG. 6, the computing device 600 illustratively includes a processor 610, an input / output subsystem 620, a memory 630, a data storage device 640, and a communication subsystem 650, and / or other components and devices commonly found in a server or similar computing device. The computing device 600 may include other or additional components (e.g., various input / output devices) as commonly found in a server computer in other embodiments. Further, in some embodiments, one or more of the exemplary components may be incorporated into, or alternatively form a part of, another component. For example, the memory 630, or a portion thereof, may be incorporated into the processor 610 in some embodiments.
[0066] The processor 610 may be embodied as any type of processor capable of executing the functions described herein. The processor 610 may be embodied as a single processor, a multiprocessor, a central processing unit (CPU), a graphics processing unit (GPU), a single or multi-core processor, a digital signal processor, a microcontroller, or other processor or processing circuitry.
[0067] Memory 630 can be embodied as any type of volatile or non-volatile memory or data storage capable of performing the functions described herein. During operation, memory 630 can store various data and software used during the operation of arithmetic unit 600, such as an operating system, applications, programs, libraries, and drivers. Memory 630 is communicatively coupled to processor 610 via I / O subsystem 620 and can be embodied as circuitry and / or components for facilitating input / output operations between processor 610, memory 630, and other components of arithmetic unit 600. For example, I / O subsystem 620 can be embodied as a memory controller hub, an input / output control hub, a platform controller hub, an integrated control circuit, a firmware device, a communication link (e.g., a point-to-point link, a bus link, a wire, a cable, a light guide, a printed circuit board trace, etc.) and / or other components and subsystems for facilitating input / output operations, or alternatively, may include these. In some embodiments, I / O subsystem 620 forms part of a system-on-chip (SOC) and may be incorporated into a single integrated circuit chip together with processor 610, memory 630, and other components of arithmetic unit 600.
[0068] The data storage device 640 can be embodied as any type of device or apparatus configured for short-term or long-term storage of data, such as, for example, a memory device and circuits, a memory card, a hard disk drive, a solid state drive, or other data storage devices. The data storage device 640 can store program code 640A for executing the training of the mutation policy network, program code 640B for generating peptides using the mutation policy, and / or program code 640C for screening the generated peptides. The communication subsystem 650 of the computing device 600 can be embodied as any network interface controller or other communication circuit, device, or collection thereof that can enable communication between the computing device 600 and other remote devices via a network. The communication subsystem 650 can be configured to effectuate such communication using any one or more communication technologies (e.g., wired or wireless communication) and associated protocols (e.g., Ethernet, InfiniBand®, Bluetooth®, Wi-Fi®, WiMAX® etc.).
[0069] As shown, the computing device 600 can also include one or more peripheral devices 660. The peripheral devices 660 may include any number of additional input / output devices, interface devices, and / or other peripheral devices. For example, in some embodiments, the peripheral devices 660 can include a display, a touch screen, a graphics circuit, a keyboard, a mouse, a speaker system, a microphone, a network interface, and / or other input / output devices, interface devices, and / or peripheral devices.
[0070] Of course, the computing device 600 can also include other elements (not shown), and specific elements can also be omitted, as would be readily apparent to those skilled in the art. For example, various other sensors, input devices, and / or output devices can be included in the computing device 600 depending on the specific implementation of the same, as would be readily understood by those skilled in the art. For example, various types of wireless and / or wired input and / or output devices can be used. Further, processors, controllers, memories, etc. can be added and utilized in various configurations. These and other variations of the processing system 600 are readily contemplated by those skilled in the art in view of the teachings of the present invention provided herein.
[0071] The embodiments described herein may be entirely in hardware, entirely in software, or may include both hardware elements and software elements. In a preferred embodiment, the present invention is implemented in software, including but not limited to firmware, resident software, microcode, etc.
[0072] Embodiments can include a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. A computer-usable medium or computer-readable medium can include any device that stores, communicates, propagates, or transports a program for use by or in connection with an instruction execution system, apparatus, or device. The medium can be magnetic, optical, electronic, electromagnetic, infrared, or semiconductor system (or apparatus or device), or a propagation medium. The medium can include computer-readable storage media such as semiconductor or solid state memory, magnetic tape, removable computer diskette, random access memory (RAM), read-only memory (ROM), rigid magnetic disk, and optical disk.
[0073] Each computer program can be tangibly stored on a machine-readable storage medium or device (such as a program memory or magnetic disk) that is readable by a general-purpose or special-purpose programmable computer to configure and control the operation of the computer when the storage medium or device is read by the computer for executing the procedures described herein. The system of the present invention can also be considered to be implemented on a computer-readable storage medium constituted by a computer program, in which case the configured storage medium causes the computer to operate in a specific predetermined manner to execute the functions described herein.
[0074] A data processing system suitable for storing and / or executing program code may include at least one processor directly or indirectly coupled to memory elements via a system bus. The memory elements can include local memory employed during actual execution of the program code, bulk storage, and cache memory that provides at least some temporary storage of at least some program code to reduce the number of times code is retrieved from bulk storage during execution. Input / output or I / O devices (including, but not limited to, keyboards, displays, pointing devices, etc.) can be coupled to the system directly or via intervening I / O controllers.
[0075] A network adapter can also be coupled to the system to enable the data processing system to become coupled to other data processing systems or remote printers or storage devices via intervening private or public networks. Modems, cable modems, Ethernet cards are but a small part of the types of network adapters currently available.
[0076] As used herein, the term "hardware processor subsystem" or "hardware processor" can refer to a processor, memory, software, or a combination thereof that cooperate to perform one or more specific tasks. In useful embodiments, the hardware processor subsystem can include one or more data processing elements (e.g., logic circuits, processing circuits, instruction execution devices, etc.). The one or more data processing elements can include a central processing unit, an image processing unit, and / or a controller based on a separate processor or computing element (e.g., logic gates, etc.). The hardware processor subsystem can include one or more on-board memories (e.g., cache, dedicated memory arrays, read-only memory, etc.). In some embodiments, the hardware processor subsystem can include one or more memories that can be on-board or off-board, or dedicated for use by the hardware processor subsystem (e.g., ROM, RAM, basic input / output system (BIOS), etc.).
[0077] In certain embodiments, the hardware processor subsystem can include and execute one or more software elements. The one or more software elements can include an operating system and / or one or more applications and / or specific code to achieve a particular result.
[0078] In other embodiments, the hardware processor subsystem can include dedicated circuits that execute one or more electronic processing functions to achieve a specified result. Such circuits can include one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and / or programmable logic arrays (PLAs).
[0079] These and other variations of the hardware processor subsystem are also contemplated in accordance with embodiments of the present invention.
[0080] In the specification, references to "an embodiment" or "an embodiment" of the present invention, and other variations, mean that the particular features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment of the present invention. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places in this specification, and any other variations, do not necessarily all refer to the same embodiment. However, it should be understood that, in view of the teachings of the present invention provided herein, the features of one or more embodiments can be combined.
[0081] For example, in the case of "A / B", the use of any of the following, " / ", "and / or", "at least one of A and B", is intended to encompass the selection of only the first listed option (A), or only the second listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to encompass the selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or the selection of only the first and second listed options (A and B), the selection of only the first and third listed options (A and C), the selection of only the second and third listed options (B and C), or the selection of all three options (A and B and C). This can be extended for as many items as are described.
[0082] The foregoing is illustrative and exemplary in every respect and not restrictive, and the scope of the invention disclosed herein is determined from the claims construed in accordance with the full breadth permitted by the Patent Act, rather than from the detailed description. The embodiments shown and described herein are merely illustrative of the invention, and it should be understood that those skilled in the art can make various modifications without departing from the scope and spirit of the invention. Those skilled in the art can implement various other combinations of features without departing from the scope and spirit of the invention. Thus, while the aspects of the invention have been described with the detail and particularity required by the Patent Act, what is desired to be claimed and protected by Letters Patent is as set forth in the appended claims.
Claims
1. A computer-implemented method for peptide generation, comprising: training a neural network of a peptide mutation policy using reinforcement learning that includes a peptide presentation score as a reward (204); generating a plurality of new peptides using the peptide mutation policy (212); calculating a binding motif of a major histocompatibility complex (MHC) using the plurality of new peptides (214); screening a plurality of library peptides according to the binding motif (216).
2. The method according to claim 1, wherein calculating the binding motif of the MHC includes generating a plurality of binding motifs for each of the plurality of MHCs, and screening includes screening according to the plurality of binding motifs.
3. The method according to claim 1, wherein training the neural network of the peptide mutation policy 【Number 1】 maximizes an objective function such as, where θ represents the parameters of the neural network of the peptide mutation policy, 【Number 2】 is the expected value with respect to the time step t, and r t (θ) is the probability ratio between the activities under the current policy and the activities under the previous policy, and 【Number 3】 is the average value at time step t, clip(·) is a clipping function, 【Number 4】 is the size of the clipping interval.
4. The method according to claim 1, wherein training the neural network of the peptide mutation policy includes pre-training using an expert policy.
5. The method according to claim 1, wherein screening the plurality of library peptides includes determining a pairwise Euclidean distance between a block substitution matrix (BLOSUM) representation of the plurality of library peptides and a BLOSUM representation of the binding motif.
6. The method according to claim 1, wherein screening the plurality of library peptides includes determining the log-likelihood of the plurality of library peptides under the weighted positions of the binding motif.
7. The method according to claim 1, wherein generating the plurality of new peptides includes sampling a random starting peptide and applying changes to the random starting peptide according to the peptide mutation policy.
8. The method according to claim 1, wherein Training the neural network of the peptide mutation policy is a method that includes changing the input peptide sequence as an activity and determining a reward for the activity based on the peptide presentation score of the changed input peptide sequence.
9. The method according to claim 1, further comprising determining a method for comparing the plurality of screened library peptides with a candidate vaccine peptide to determine whether the candidate vaccine peptide binds to MHC.
10. The method according to claim 9, further comprising generating a vaccine based on the candidate vaccine peptide and administering the vaccine to prevent disease.
11. A system for peptide generation, a hardware processor (610), when executed by the hardware processor, causes the hardware processor to, using reinforcement learning including a peptide presentation score as a reward, train a neural network of a peptide mutation policy (204), generate a plurality of new peptides using the peptide mutation policy (212), calculate a binding motif of a major histocompatibility complex (MHC) using the plurality of new peptides (214), and a memory (630) storing a computer program for causing the hardware processor to execute a procedure (216) for screening a plurality of library peptides according to the binding motif.
12. The system according to claim 11, wherein the computer program further causes the processor to execute a procedure for generating a plurality of additional binding motifs for a plurality of respective additional MHCs, and the screening includes screening according to the plurality of additional binding motifs.
13. The system according to claim 11, wherein the computer program causes the processor to, 【Number 5】 further execute a procedure for training the neural network of the peptide mutation policy by maximizing an objective function such as, where θ represents the parameters of the neural network of the peptide mutation policy, 【Number 6】 is the expected value with respect to time step t, rt(θ) is the probability ratio between the activity with the current policy and the activity with the previous policy, 【Number 7】 is the average value at time step t, and clip(·) is a clipping function. 【Number 8】 A system that is the size of the clipping interval.
14. The system according to claim 11, wherein the computer program further causes the processor to perform a procedure of pre-training the neural network of the peptide mutation policy using an expert policy.
15. The system according to claim 11, wherein the computer program further causes the processor to perform a procedure of determining a pairwise Euclidean distance between the block substitution matrix (BLOSUM) representation of the plurality of library peptides and the BLOSUM representation of the binding motif.
16. The system according to claim 11, wherein the computer program further causes the processor to perform a procedure of determining the log-likelihood of the plurality of library peptides under the weighted position of the binding motif.
17. The system according to claim 11, wherein the computer program further causes the processor to perform a procedure of sampling a random starting peptide and applying changes to the random starting peptide according to the peptide mutation policy.
18. The system according to claim 11, wherein the computer program further causes the processor to perform a procedure of changing the input peptide sequence as an activity and a procedure of determining a reward for the activity based on the peptide presentation score of the changed input peptide sequence.
19. The system according to claim 11, wherein the computer program further causes the processor to perform a procedure of comparing the plurality of screened library peptides with a candidate vaccine peptide to determine a method by which the candidate vaccine peptide binds to MHC.
20. The system according to claim 19, wherein the computer program further causes the processor to perform a procedure of generating a vaccine based on the candidate vaccine peptide.
Citation Information
Patent Citations
Biological sequence retrieval method and device, electronic equipment and storage medium
CN114005493A
Protein sequence design method, protein structure design method, device and electronic equipment
CN114155912A
Information processing method and learning model
JP2020077206A
Novel immunogenic peptides
JP2020079271A
Method and system for binding affinity prediction and method of generating a candidate protein-binding peptide
WO2020234188A1