Deep learning for the maturation (correction) and improvement of the properties of novel antibody affinity

Machine learning and Computational Directed Evolution enhance antibody maturation by predicting affinity and expression, addressing inefficiencies in directed evolution and reducing production costs.

JP7846083B2Active Publication Date: 2026-04-14FLAGSHIP PIONEERING INNOVATIONS VI LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
FLAGSHIP PIONEERING INNOVATIONS VI LLC
Filing Date
2021-07-28
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Current methods for antibody maturation, such as directed evolution, are inefficient in improving antibody affinity and specificity, particularly when targeting multiple properties simultaneously.

Method used

Employing machine learning techniques to computationally mature antibody sequences through a process similar to directed evolution, using a pre-trained model to predict affinity and expression, and applying Computational Directed Evolution (CDE) for constrained search in the antibody sequence space.

Benefits of technology

This approach significantly enhances antibody affinity, reducing the number of structures needed for testing and production costs, while achieving higher specificity and lower doses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007846083000030
    Figure 0007846083000030
  • Figure 0007846083000031
    Figure 0007846083000031
  • Figure 0007846083000032
    Figure 0007846083000032
Patent Text Reader

Abstract

Controlling antibody affinity and antibody expression is key for clinical application. High-affinity antibodies correlate with higher specificity and therefore can be used at lower doses. Currently, antibody maturation is approached by directed evolution, where an initial library of mutant binders is seeded into the process and affinity is improved through multiple rounds of mutation and selection. However, the present disclosure employs machine learning techniques to computationally mature antibody sequences using a process similar to directed evolution. These antibody sequences can be engineered into physical antibodies after their computation and validation. Additionally, this method has the potential to surpass directed evolution in targeting specific affinities and is applicable to general protein-protein interactions.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 057,376, filed on 28 July 2020. All teachings of the above application are incorporated herein by reference. [Background technology]

[0002] Antibody maturation is the process by which a given antibody improves its affinity for its antigen. In immunology, an antigen is a toxin or foreign substance that triggers an immune response. One example of such an immune response is the production of antibodies that bind to an antigen in order to neutralize it. Antibodies can also be designed in a wide variety of ways. [Overview of the project] [Means for solving the problem]

[0003] Controlling antibody affinity and expression is key to clinical applications. High-affinity antibodies correlate with higher specificity and can therefore be used at lower doses. Currently, antibody maturation is addressed by the directed evolution method, in which an initial library of mutagenesis conjugates is implanted into the process, and affinity is improved through multiple mutations and selections. However, this disclosure employs machine learning techniques to computationally mature antibody sequences using a process similar to directed evolution. These antibody sequences can then be manufactured into physical antibodies after their computation and validation. In addition, this method has the potential to surpass directed evolution when targeting specific affinity and is applicable to general protein-protein interactions.

[0004] In one embodiment, computationally directed evolution (CDE) can be used in multi-objective and multi-model situations (for example, when selecting two or more properties for improvement nearly simultaneously). In this case, one or more models may be used. Each of these models may have one or more properties to be optimized. One or more models may be combined into a single objective function, and CDE may optimize that objective. The objective function is optimized by an optimization procedure (in this case, computationally directed evolution). Once optimized, the objective function produces an antibody sequence that improves the one or more antibody properties in question.

[0005] In one embodiment, a method for determining antibody sequences having improved properties includes training a machine learning model on a first group of antibody sequences. Each antibody sequence in the database is marked by one or more corresponding properties. The method further generates a refined machine learning model by training a machine learning model on a second group of antibody sequences, each marked by corresponding properties. The second group of antibody sequences are associated with the antigen in question. The method further includes generating antibody sequences based on the refined machine learning model.

[0006] Those skilled in the art recognize that a finely tuned machine learning model (also called a finely tuned machine learning model) is a machine learning model that is first trained on a general dataset and then finely tuned on a larger, specific dataset. Fine-tuning can also be described as the process of taking a machine learning model trained on a given task and then training this machine learning model to perform a second task.

[0007] In one embodiment, a method for determining an antibody sequence having improved properties includes generating scores for each of a plurality of machine learning models. Each machine learning model is trained on each plurality of antibody sequences. Each antibody sequence in each plurality of antibody sequences is denoted by the properties corresponding to the plurality of antibody sequences and the values ​​of the properties corresponding to each antibody sequence. Generating scores for each machine learning model produces scores for each of the plurality of machine learning models, and the scores indicate the contribution of each machine learning model to predicting the properties corresponding to the respective machine learning model. The method further includes generating antibody sequences by using the plurality of machine learning models by weighting the outputs of each machine learning model according to each generated score and then combining the weighted outputs into a weighted sum.

[0008] In some embodiments, the contributions include one or more empirical derivations by modulating the importance of properties to production, protein expression, patient immunogenicity, expression, developability, interaction with other models, orthogonality with other models, and the production process.

[0009] In one embodiment, generating an antibody sequence further includes selecting an antibody sequence from a proposal distribution based on a finely tuned machine learning model. Generating an antibody sequence further includes determining whether the selected antibody sequence has an acceptance probability above a certain threshold, and if so, analyzing the antibody, and otherwise selecting the next antibody sequence from the proposal distribution.

[0010] In one embodiment, generating an antibody sequence further includes comparing a first characteristic (determined by a finely tuned machine learning model) of an antibody selected from the proposed distribution with a second characteristic (determined by a finely tuned machine learning model) of the antibody having the best characteristic in the current search. If the first characteristic is greater than the second characteristic, the method exchanges the antibody with the best characteristic for the antibody selected from the proposed distribution.

[0011] In one embodiment, training a machine learning model further includes providing a set of amino acid sequences marked by at least one characteristic. Training the machine learning model further includes masking a portion of the set of amino acid sequences to provide a masked set of amino acid sequences. The remainder of the set of amino acid sequences is an unmasked set of amino acid sequences. Training the machine learning model further includes training the machine learning model to estimate each of the masked sets of amino acid sequences based on (1) at least one characteristic marking each masked amino acid sequence and (2) an unmasked set of amino acid sequences and the marked characteristic of each unmasked amino acid sequence.

[0012] In one embodiment, generating a finely tuned machine learning model further includes weighting each sequence characteristic of a second set of antibody sequences. Generating a finely tuned machine learning model further includes determining the optimal model parameters for generating the second set of antibody sequences using the machine learning model. Generating a finely tuned machine learning model further includes applying the optimal model parameters to the machine learning model. The resulting model with the applied optimal model parameters is the finely tuned machine learning model.

[0013] In one embodiment, the corresponding property is affinity (e.g., binding affinity) or expression. In some other embodiments, examples of properties (e.g., function values) can be one or more of the following: binding affinity, binding specificity, catalytic (e.g., enzyme) activity, fluorescence, solubility, thermal stability, conformation, immunogenicity, protein aggregation, proteolytic stability, expression, off-target effects, and any other functional property of a biopolymer sequence. This process is applicable to any protein for which we have an example of a (small) starting set having the property. Next, from here, we can use the process described herein to modify the property to be more suitable for our needs. This can lead to an increase in the value (e.g., an increase in the catalytic reaction rate) to reach a target value or a value within a certain range (e.g., corresponding to a specific binding affinity) or a decrease in the property value (e.g., a decrease in immunogenicity).

[0014] In one embodiment, the method includes selecting an antibody sequence candidate that is within a defined acceptance criterion from a proposed distribution based on a fine-tuned machine learning model. The method further includes replacing the best-known antibody sequence with the antibody sequence candidate if the property of the antibody sequence candidate is better than the best-known antibody sequence or otherwise ignoring the antibody sequence candidate.

[0015] In one embodiment, the method also includes generating an antibody having the generated antibody sequence. In one embodiment, the method can also include providing a manufactured antibody having the generated antibody sequence and analyzing the antibody with respect to the property.

[0016] In one embodiment, a system for determining an antibody having improved characteristics includes a processor and a memory having computer code instructions stored thereon. The processor and the memory are configured by the computer code instructions to train a machine learning model based on a first plurality of antibody sequences for the system. Each antibody sequence in the database is labeled with a corresponding characteristic. The processor is further configured to generate a fine-tuned machine learning model by training the machine learning model based on a second plurality of antibody sequences and corresponding characteristics. The second plurality of antibody sequences are related to the antigen in question. The processor is further configured to generate an antibody sequence based on the fine-tuned machine learning model.

[0017] In one embodiment, a method of antibody maturation may include providing a first antibody sequence to the system or method described above and obtaining the generated antibody sequence from the system.

[0018] In one embodiment, an isolated antibody may be generated by the method described above. In one embodiment, the isolated antibody is generated by recombinant technology. In one embodiment, the isolated antibody is chemically synthesized.

[0019] In one embodiment, the method includes determining or generating an antibody sequence having improved characteristics, and the determining or generating is performed by a fine-tuned machine learning model. The fine-tuned machine learning model can be generated by (1) training a machine learning model based on a first plurality of antibody sequences each labeled with a corresponding characteristic, and (2) training the machine learning model based on a second plurality of antibody sequences each labeled with a corresponding characteristic and related to the antigen in question, thereby generating a fine-tuned machine learning model. Optionally, training the machine learning model, generating the fine-tuned machine learning model, or both can be performed by a third party separate from generating the antibody sequence.

[0020] As used herein, an antibody sequence refers to a regular sequence of amino acids that can be stored digitally or in another format. An antibody refers to the physical expression of an antibody. Those skilled in the art will recognize that antibody sequences produced by the systems and methods of this disclosure can be manufactured or otherwise produced as antibodies.

[0021] In one embodiment, a method for determining an antibody sequence having improved properties includes training multiple machine learning models. Each machine learning model is trained on a corresponding initial set of antibody sequences. Each antibody sequence in the initial set of antibody sequences is marked by a corresponding property. Those skilled in the art will understand that each machine learning model may be trained on a different set of antibody sequences. Next, the method generates a set of finely tuned machine learning models. Each finely tuned machine learning model is generated by training each machine learning model on a second set of antibody sequences. Each antibody sequence in the second set of antibody sequences is marked by a corresponding property, and the second set of antibody sequences is associated with the antigen in question. Next, the method generates antibody sequences based on an objective function by using the set of finely tuned machine learning models weighted by corresponding hyperparameters.

[0022] In one embodiment, one or more antibody sequences are associated with the antigen.

[0023] In one embodiment, a method for determining an antibody sequence having improved properties includes providing scores for each of a plurality of machine learning models, each trained on each of a plurality of antibody sequences. Each antibody sequence in the plurality of antibody sequences is denoted by a property corresponding to the plurality of antibody sequences and a value for the property corresponding to each antibody sequence. Each score for each of the plurality of machine learning models indicates the contribution to predicting the property corresponding to each machine learning model. The method further includes generating antibody sequences by using the plurality of machine learning models by weighting the outputs of each machine learning model according to each provided score and then combining the weighted outputs into a weighted sum.

[0024] In one embodiment, a method for determining an antibody sequence having improved properties includes generating antibody sequences by using multiple machine learning models. The generation of antibody sequences is performed by weighting the output of each machine learning model according to the respective scores of the multiple machine learning models. Each machine learning model is trained on each of the multiple antibody sequences. Each antibody sequence in each of the multiple antibody sequences is denoted by the properties corresponding to the multiple antibody sequences and the values ​​of the properties corresponding to each antibody sequence. Each score of each machine learning model in the multiple machine learning models indicates its contribution to predicting the properties corresponding to each machine learning model. The method further includes combining the weighted outputs into a weighted sum.

[0025] This patent or application file includes at least one figure drawn in color. A copy of this patent or patent application publication, including the color drawing, will be provided by the Patent and Trademark Office upon request and payment of the required fees.

[0026] The foregoing will become clear from the following more specific description of the exemplary embodiments shown in the accompanying drawings. In the accompanying drawings, similar reference letters refer to the same parts throughout the various figures. The accompanying drawings are not necessarily scaled, and emphasis is rather placed on illustrating the embodiments. [Brief explanation of the drawing]

[0027] [Figure 1A] A diagram showing exemplary embodiments of the method of this disclosure. [Figure 1B] This flowchart illustrates exemplary embodiments of the processes adopted in this disclosure. [Figure 2] This graph shows a random walk sequence within the antibody sequence space for discovering antibody sequences using computationally oriented evolution. [Figure 3] This graph shows the improvement in the generated antibody sequences compared to existing datasets. [Figure 4] This document describes a computer network or similar digital processing environment in which several embodiments of the present invention may be implemented. [Figure 5] Figure 4 is a diagram illustrating the exemplary internal structure of a computer (e.g., a client processor / device or server computer) within a computer system. [Modes for carrying out the invention]

[0028] The following is a description of exemplary embodiments.

[0029] Figure 1A is a block diagram 100 illustrating an exemplary embodiment of the method of the present disclosure. The present disclosure employs a novel method for closed-loop antibody affinity maturation. The iterative process begins with unsupervised pre-training of a deep machine learning model on a large and important set of protein sequence data and their properties (102). In some embodiments, this includes unsupervised pre-training of a language model on a dataset containing n antibody sequences. This conditions the model to learn the underlying statistical relationships between amino acids in the antibody sequences.

[0030] Next, this process fine-tunes a pre-trained model on a smaller supervised learning task by using data specific to the desired antibody-antigen pair (104). This fine-tuned machine learning model is trained on both affinity and expression as data specific to the desired antibody-antigen pair, although in practice any number or type of additional characteristics may be used in the fine-tuning process. Once the fine-tuned model is trained (104), it is employed downstream to perform affinity maturation by solving an optimization problem, which is further described below. After fine-tuning, computationally directed evolution performs a constrained search of the antibody sequence space for high-affinity antibody sequences for the selected target. The solution to this optimization problem is one or more affinity-mature antibody sequences (e.g., sequences predicted by the machine learning model to have higher affinity than those observed in the supervised training task) (106). Next, the method constructs these affinity-mature antibody sequences (i.e., improved antibody sequences). The method further experimentally analyzes the constructed antibody sequences for affinity and expression. New data from the analyzed sequences are incorporated into the supervised learning dataset, and process 100 is repeated until a suitable candidate is found (102). In one embodiment, a high affinity candidate is qualified by an equilibrium dissociation constant (K) less than 10 picomolars (pM). D (This can be an equilibrium dissociation constant) (for example, in the case of an antibody against fluorescein).

[0031] As used herein, "antibody" refers to an immunoglobulin molecule such as a carbohydrate, polynucleotide, lipid, polypeptide, etc. that can specifically bind to a target (via at least one antigen recognition portion) disposed within the variable region of the immunoglobulin molecule. As used herein, the term "antibody" refers not only to intact (i.e., full-length) monoclonal antibodies, but also to antigen-binding fragments (such as Fab, Fab', F(ab')2, Fv, etc.), single-chain variable fragments (scFv: single chain variable fragment), mutants thereof, fusion proteins containing antibody portions, humanized antibodies, chimeric antibodies, bispecific antibodies, linear antibodies, single-chain antibodies, single-domain antibodies (such as camelid or llama VHH antibodies), multispecific antibodies (such as bispecific antibodies), and any other modified configuration of an immunoglobulin molecule containing an antigen recognition portion of the required specificity (including glycosylation variants of the antibody, amino acid sequence variants of the antibody, and covalently modified antibodies).

[0032] As used herein, the term "K", also referred to as "binding constant", "equilibrium dissociation constant", or "affinity constant", D is a measure of the degree of reversible association between two molecular species (such as an antibody and a target protein), and includes both the actual binding affinity and the apparent binding affinity. The binding affinity can be determined by using methods known in the art (including, for example, by measurement of surface plasmon resonance (such as using a BIAcore system and analysis)).

[0033] In some embodiments, the antibody has a K of 10 -4 M, 10 -5 M, 10 -6 M, 10 -7 M, 10 -8 M, 10 -9 M, 10 -10 M, 10 -11 M, or less than 10 -12 M for the K DIt binds to the target protein at concentrations of less than 1000 nM, or alternatively less than 900 nM, alternatively less than 800 nM, alternatively less than 700 nM, alternatively less than 600 nM, alternatively less than 500 nM, alternatively less than 400 nM, alternatively less than 300 nM, alternatively less than 200 nM, alternatively less than 100 nM, alternatively less than 90 nM, alternatively less than 80 nM, alternatively less than 70 nM, alternatively less than 60 nM, alternatively less than 50 nM, or alternatively less than 40 K less than nM, alternatively less than 30 nM, or alternatively less than 20 nM, or alternatively less than 15 nM, alternatively less than 10 nM, or alternatively less than 9 nM, or alternatively less than 8 nM, or alternatively less than 7 nM, or alternatively less than 6 nM, alternatively less than 5 nM, or alternatively less than 4 nM, or alternatively less than 3 nM, alternatively less than 2 nM, or alternatively less than 1 nM, or alternatively less than 1000 pM, or alternatively less than 10 pM, or alternatively less than 1 pM D They can be joined together.

[0034] In some embodiments, antibody sequences identified by the methods disclosed herein correspond to antibodies that bind to the target protein with an affinity of at least 70%, alternatively at least 75%, or alternatively at least 80%, or alternatively at least 85%, or alternatively at least 90%, or alternatively at least 95%, compared to the affinity of a reference antibody. In some embodiments, the antibody binds to the target protein with a higher affinity than the reference antibody. The reference antibody may be, for example, the antibody with the highest reported (e.g., published) similarity with respect to a given target.

[0035] In fact, process 100, described in Figure 1A, rapidly improves antibody affinity after a single step by generating a fluorescein antibody sequence with affinity greater than 10 pM. This process can reduce the cost and time required to produce better therapeutic antibodies by reducing the number of structures that need to be made and tested. Details of this process are described below.

[0036] In the initial experiments, the process employed two datasets: a large unsupervised set of antibody sequences, and a small set of sequence affinity pairs for the antibodies in question to computationally bootstrap their maturation. Specifically, pre-training used a complete antibody repertoire taken from a single individual. This dataset is described in more detail in Bryan, et al., “Commonality despite exceptional diversity in the baseline human antibody repertoire,” Nature 566.7744(2019):393 (hereafter “Bryan”), which is incorporated herein by reference in its entirety. Unsupervised dataset

number

number

number

number

[0037] The machine learning model is trained to act as a trait oracle for both affinity and expression traits. The model used in this process (called “Omniprot”) is a modified version of the BERT masked language model, which is described in more detail in Devlin, Jacob, et al. “Bert: Pre-training of deep bidirectional transformers for language understanding.” arXiv preprint arXiv:1810.04805 (2018) (hereafter “Devlin”), which is incorporated herein by reference in its entirety. However, in principle, any model that can be pre-trained in an unsupervised manner can be used. Omniprot is a mask of the protein sequence s masked It is a deep transformer model trained by learning to reconstruct it, with a random 15% of its masked amino acids. The training task is defined as follows:

number

number

number

[0038] A set of model parameters

number

number

number

number

number

[0039] Once the model is trained, it can be used to find antibody sequences with improved affinity. This process is called sequence maturation. In one exemplary embodiment, a finely tuned Omniprot model can predict the affinity and expression of candidate sequences and explore the entire effective domain of possible antibody sequences. The optimization problem can be expressed as follows:

number

number

[0040] To implement the Metropolis-Hastings method, the following set of objective functions are used as inputs: a) Initial starter antibody sequence

number

number

number

number

number

[0041] Given the above, the CDE method proceeds as follows (note that this is an implementation of Metropolis-Hastings): a) Start with the initial antibody sequence:

number

number

number

number

number

[0042] Designers may wish to optimize multiple properties of an antibody sequence. To do so, models may be employed to estimate each property of the antibody sequence, and the objective function synthesizes the weighted results of each model using hyperparameters. Hyperparameters and their use are described below.

[0043] We consider a scenario in which two objective functions are optimized. Those skilled in the art will recognize that any two properties can be analyzed. However, in this example, binding affinity and antibody solubility are used for illustrative purposes. For example, the first model determines the binding affinity of a sequence to our target (m1(s)). The second model determines the antibody solubility (m2(s)). Both models take the sequence (s) and return a scalar quantity representing the measured property. The combined objective function (e.g., the objective function) can synthesize two scalar outputs as follows: objective(s):=(1-α)m1(s)+αm2(s) Here, α lies on a closed interval between 0 and 1. Determining this α hyperparameter is part of the process of finding a good objective function and may require testing. Generally, the objective function can be expressed as a weighted sum of a set of models. Optimizing multiple models involves determining how to weight each model when synthesizing them to obtain the desired antibody sequence. Such optimization may be done manually, with the antibody designer prioritizing some properties over others. Hyperparameters may represent one or more of the following: the importance of a property to manufacturing (e.g., manufacturability), protein expression, patient immunogenicity, developability, interaction with other models, orthogonality with other models, and empirical derivations by taming the production process. Manufacturability is a factor based on the ease or difficulty of producing a protein sequence (e.g., a drug) using standard biochemical techniques. Manufacturability factors include how easily the protein is expressed, how likely the protein is to aggregate, how stable the protein is, etc. All these concerns relate to the cost and feasibility of production. Developability is a factor based on attributes related to the clinical success of the protein sequence (e.g., a drug). Factors influencing development feasibility include how easily the protein is expressed, the likelihood of protein aggregation, the stability of the protein, its specificity to the target, and so on.

[0044] Generally, for n models, the multi-model type objective function can be expressed as follows:

number

[0045] Figure 1B is a flowchart 150 illustrating an exemplary embodiment of the process employed in this disclosure. The process begins with selecting an initial antibody s0 (152). The initial antibody s0 is an antibody that is improved by a machine learning process. To initialize, the method selects the best known antibody s b Set this to s0 of the first pass. Next, this method is defined in the proposal distribution above.

number

number

[0046] The primary distribution induced by multi-sequence alignment is the empirical distribution of latent amino acids found at specific locations within the training set. The zero-order distribution is the empirical distribution of all amino acids, regardless of location. The zero-order distribution is similar to the distribution obtained by taking all remaining amino acids from each protein in the training set, placing them in a bag, and then sampling from this bag without exchange. The zero-order distribution does not preserve location, while the primary distribution considers the distribution at each location within the protein. The process described above induces both the distribution of latent sequences and a conditional distribution given the start sequence.

[0047] Probability g(s,s c If ) is above the configurable threshold, then s is s c It is set to (158). In one embodiment, the configurable threshold can be set by the designing user.

[0048] In another embodiment, the configurable threshold may be set automatically by at least one factor. Whether automatic or manual, the configurable threshold may be set by considering the following factors: the acceptance rate of the proposal and the mixing rate of the MCMC procedure. If the threshold is set too low, the acceptance rate will be high, but the mixing rate will be low, and convergence will be slow. If the threshold is set too high, the mixing rate will be high, but the acceptance rate will be low, and therefore convergence will be slow again. Therefore, to maximize the performance of the algorithm, an intermediate value that balances the proposal acceptance rate and the mixing rate is ideal. As described above, g(s,s c ) represents a mapping from a pair of sequences (e.g., an antibody sequence and a candidate next sequence) to a unit interval that indicates the probability of accepting the candidate sequence as the next step in a random walk.

[0049] Next, the method

number

number

[0050] Figure 2 is a graph showing a random walk sequence in the antibody sequence space for discovering antibody sequences using computationally directed evolution. Antibody affinity is optimized by using CDE with a finely tuned model. Seed antibody sequence s0202 is one or more initial antibody sequences. The best sequence of these sequences s b Sequence 204 is selected for initialization, where the best sequence is the one with the highest affinity and expression characteristics. Next, a new candidate sequence 206 is drawn from the proposed distribution. Then, a pair of sequences s is s c It will be set to s. b If the characteristic value result of the finely tuned model is smaller than that of s, then s b The value is set to s. This process is then repeated until a suitable antibody is found.

[0051] The results of the methods described herein are improved antibody sequences. bThis provides a random walk, which is repeated to generate the same number of sequences required for the test. In some embodiments, this random walk may be a randomized, supervised, or hybrid method. These sequences are then analyzed for similarity to the antigen and expression. This data may then be fed back into the fine-tuning process and may be repeated as desired until a clinically significant antibody sequence is generated. By using the method described herein, antibody sequences with improved fluorescein antibody affinity, exceeding the highest affinity antibody found in the dataset by an order of magnitude, can be generated. In this case, the pre-training and fine-tuning used the dataset described above, but other datasets may be used.

[0052] Figure 3 is Graph 300, showing the improvement of the generated antibody sequence 302 compared to an existing dataset. This graph shows the fluorescein antibody sequences generated by using a novel antibody affinity maturation process. By using a novel antibody affinity maturation process, fluorescein antibody sequences with affinity of 100 pM or less are generated. In Graph 300, antibody sequences are plotted against affinity (X axis) versus expression (Y axis). The upper right quadrant contains the generated antibody sequences 302 having the highest values ​​for both properties. Those skilled in the art will recognize that these generated antibody sequences 302 (shown in red) possess the desired properties through this process.

[0053] Figure 4 shows a computer network or similar digital processing environment in which several embodiments of the present invention may be implemented.

[0054] The client computer / device 50 and server computer 60 provide input / output devices for processing, storing, and executing application programs, etc. The client computer / device 50 may also be linked to other computing devices (including other client devices / processes 50 and server computers 60) via a communication network 70. The communication network 70 may be part of a remote access network, a global network (e.g., the Internet), a worldwide collection of computers, a local area network or wide area network, and a gateway that currently uses its respective protocol (such as TCP / IP, Bluetooth®) to communicate with each other. Other electronic device / computer network architectures are also preferred.

[0055] Figure 5 is a diagram illustrating the exemplary internal structure of a computer (e.g., client processor / device 50 or server computer 60) within the computer system of Figure 4. Each computer 50, 60 includes a system bus 79, which is a set of hardware lines used for data transfer between components of a computer or processing system. The system bus 79 is essentially a shared conduit connecting various elements of a computer system (e.g., processor, disk storage, memory, input / output ports, network ports, etc.) that enable the transfer of information between elements. Attached to the system bus 79 is an I / O device interface 82 for connecting various input and output devices (e.g., keyboard, mouse, display, printer, speaker, etc.) to the computers 50, 60. The network interface 86 allows the computer to connect to various other devices attached to a network (e.g., network 70 in Figure 5). Memory 90 provides volatile storage for computer software directives 92 and data 94 used to implement one embodiment of the present invention (e.g., the machine learning model module detailed above and the fine-tuned machine learning model module code). The disk storage 95 provides non-volatile storage for computer software directives 92 and data 94 used to implement one embodiment of the present invention. A central processor unit 84 is also mounted on the system bus 79 and executes computer directives.

[0056] In one embodiment, the processor routine 92 and data 94 are a computer program product (generally referred to by reference numeral 92) (including a non-temporary computer-readable medium (e.g., one or more removable storage media such as DVD-ROMs, CD-ROMs, diskettes, tapes)) that provides at least a portion of the software instructions for the system of the present invention. The computer program product 92 may be installed by any preferred software installation procedure as is well known in the art. In another embodiment, at least a portion of the software instructions may also be downloaded over cable communications and / or wireless connections. In yet another embodiment, the program of the present invention is a computer program propagated signal product that is embodied on a propagated signal (e.g., radio waves, infrared waves, laser waves, sound waves, or radio waves propagated over a global network such as the Internet or other networks). Such a carrier medium or signal may be employed to provide at least a portion of the software instructions for the routine / program 92 of the present invention.

[0057] All patents, published applications, and references cited herein are incorporated in their entirety by reference.

[0058] While exemplary embodiments have been specifically shown and described, it will be understood by those skilled in the art that various modifications of form and detail can be made without departing from the spirit and scope of the embodiments contained in the appended claims. The present invention provides, for example, the following items: (Item 1) A method for determining an antibody sequence having improved properties, wherein the method is: The process involves generating scores for each of multiple machine learning models, each machine learning model being trained on each of multiple antibody sequences, each antibody sequence being represented by a characteristic corresponding to the multiple antibody sequences and a value of the characteristic corresponding to each antibody sequence, and generating scores for each of the multiple machine learning models indicating their contribution to predicting the characteristic corresponding to each machine learning model, and A method for generating an antibody sequence by using a plurality of machine learning models, weighting the output of each machine learning model according to the generated score of each model and combining the weighted outputs into a weighted sum. (Item 2) The generation of the aforementioned antibody sequence further involves, Selecting antibody sequences from proposal distributions based on the aforementioned multiple machine learning models; and The method according to item 1, comprising determining whether the selected antibody sequence has an acceptable probability exceeding a specific threshold, analyzing the antibody sequence if so, and otherwise selecting the next antibody sequence from the proposed distribution. (Item 3) The generation of the aforementioned antibody sequence further involves, Comparing a first characteristic value of an antibody sequence selected from a proposed distribution, determined by a function of the multiple machine learning models, with a second characteristic value of an antibody sequence having the best characteristic value in the current search, determined by the multiple machine learning models; and The method according to item 1 or 2, comprising, if the first characteristic value is greater than the second characteristic value, replacing the antibody sequence having the best characteristic value with the antibody sequence selected from the proposed distribution. (Item 4) The generation of the aforementioned fine-tuned machine learning model is further, Weighting the sequence characteristics of each of the aforementioned second antibody sequences; Using the aforementioned machine learning model, determine the optimal model parameters for generating the second set of antibody sequences; and The method according to any one of items 1 to 3, comprising applying the optimal model parameters to the machine learning model, wherein the resulting model having the applied optimal model parameters is the finely tuned machine learning model. (Item 5) The method according to any one of items 1 to 4, wherein the corresponding characteristic is at least one of affinity, expression, protein aggregation, proteolytic stability, expression, and off-target effect. (Item 6) Selecting antibody sequence candidates that fall within the specified acceptance criteria from the proposal distribution based on the aforementioned multiple machine learning models; and The method according to any one of items 1 to 5, further comprising: exchanging the antibody sequence candidate with the antibody sequence candidate if the characteristics of the antibody sequence candidate are better than those of the antibody sequence known to be the best; or otherwise ignoring the antibody sequence candidate. (Item 7) The method according to any one of items 1 to 6, further comprising generating an antibody having the generated antibody sequence. (Item 8) A method according to any one of items 1 to 7, further comprising providing a manufactured antibody having the generated antibody sequence, and analyzing the antibody with respect to the properties. (Item 9) The method according to any one of items 1 to 8, further comprising training one or more machine learning models, each machine learning model being trained on a plurality of antibody sequences, each of the plurality of antibody sequences being denoted by a characteristic corresponding to the plurality of antibody sequences and a value of the characteristic corresponding to the respective antibody sequence. (Item 10) Training one or more of the aforementioned machine learning models can be further: To provide a set of amino acid sequences marked by at least one characteristic; Masking a portion of a set of amino acid sequences in order to provide a masked set of amino acid sequences, wherein the remainder of the set of amino acid sequences is an unmasked set of amino acid sequences: and The method of item 9, further comprising training one or more machine learning models to estimate each of the masked amino acid sequences based on (1) at least one characteristic that identifies each masked amino acid sequence and (2) a set of unmasked amino acid sequences and the identified characteristic of each unmasked amino acid sequence. (Item 11) The method according to any one of items 1 to 10, wherein the antibody sequence is generated by employing MCMC sampling. (Item 12) The method according to any one of items 1 to 11, wherein the plurality of antibody sequences are related to the antigen. (Item 13) The aforementioned contributions will improve the following characteristics: The importance of the aforementioned characteristics for the production of the aforementioned antibody sequence, The immunogenicity of the aforementioned antibodies in the patient, The antibody expression level, Development feasibility, Interaction with other models, Orthogonality with other models, and empirical derivation by adjusting the generation process, The method described in any one of items 1 through 12, which is for predicting at least one of the following. (Item 14) A system for determining an antibody sequence having improved characteristics, comprising a processor and a memory storing computer code instructions, The processor and the memory, according to the computer code directive, perform the following: The process involves generating scores for each of multiple machine learning models, each machine learning model being trained on each of multiple antibody sequences, each antibody sequence being represented by a characteristic corresponding to the multiple antibody sequences and a value of the characteristic corresponding to each antibody sequence, and generating scores for each of the multiple machine learning models indicating their contribution to predicting the characteristic corresponding to each machine learning model, and The antibody sequence is generated by using the multiple machine learning models, weighting the output of each machine learning model according to the generated score, and then combining the weighted outputs into a weighted sum. A system configured to cause the aforementioned system to perform the following actions. (Item 15) The generation of the aforementioned antibody sequence further involves, Selecting antibody sequences from a proposal distribution based on the aforementioned fine-tuned machine learning model; and To determine whether the selected antibody sequence has an acceptable probability exceeding a specific threshold. The system according to item 14, comprising, if so, analyzing the antibody sequence, and otherwise selecting the next antibody sequence from the proposed distribution. (Item 16) The generation of the aforementioned antibody sequence further involves, Comparing a first characteristic value of an antibody sequence selected from the proposed distribution, determined by the function of the finely tuned machine learning model, with a second characteristic value of an antibody sequence having the best characteristic value in the current search, determined by the finely tuned machine learning model; and The system according to any one of items 14 to 15, comprising exchanging the antibody sequence having the best characteristic value with the antibody sequence selected from the proposed distribution if the first characteristic value is greater than the second characteristic value. (Item 17) The generation of the aforementioned fine-tuned machine learning model is further, Weighting the sequence characteristics of each of the aforementioned second antibody sequences; Using the aforementioned machine learning model, determine the optimal model parameters for generating the second set of antibody sequences; and The system according to any one of items 14 to 16, comprising applying the optimal model parameters to the machine learning model, wherein the resulting model having the applied optimal model parameters is the finely tuned machine learning model. (Item 18) The system according to any one of items 14 to 17, wherein the corresponding characteristic is at least one of affinity and expression. (Item 19) The aforementioned processor further: Based on the aforementioned fine-tuned machine learning model, antibody sequence candidates that fall within the specified acceptance criteria are selected from the proposal distribution; and The system according to any one of items 14 to 18, configured to replace the best known antibody sequence with the antibody sequence candidate if the characteristics of the antibody sequence candidate are better than those of the best known antibody sequence, or otherwise to ignore the antibody sequence candidate. (Item 20) The system according to any one of items 14 to 19, wherein the processor is further configured to train one or more machine learning models, each machine learning model being trained on each of the plurality of antibody sequences, and each of the plurality of antibody sequences is denoted by a characteristic corresponding to the plurality of antibody sequences and a value of the characteristic corresponding to each of the antibody sequences. (Item 21) Training the aforementioned machine learning model further involves, To provide a set of amino acid sequences marked by at least one characteristic; Masking a portion of a set of amino acid sequences in order to provide a masked set of amino acid sequences, wherein the remainder of the set of amino acid sequences is an unmasked set of amino acid sequences; and The system according to item 20, comprising training the machine learning model to estimate each of the masked sets of amino acid sequences based on (1) at least one characteristic that identifies each masked amino acid sequence and (2) the unmasked sets of amino acid sequences and the identified characteristics of each unmasked amino acid sequence. (Item 22) The system described in any one of items 14 to 21, wherein the generation of the antibody sequence is performed by employing MCMC sampling. (Item 23) The system according to any one of items 14 to 22, wherein the plurality of antibody sequences are associated with the target antigen. (Item 24) The aforementioned contributions will improve the following characteristics: The importance of the aforementioned characteristics for the production of the aforementioned antibody sequence, Expression of the aforementioned antibody sequence, The immunogenicity of the aforementioned antibodies in the patient, The antibody expression level, Development feasibility, Interaction with other models, Orthogonality with other models, and empirical derivation by adjusting the generation process, A system described in any one of items 14 through 23, for predicting at least one of the following. (Item 25) A method for maturing an antibody, comprising: providing a first antibody sequence to a system described in any one of items 9 to 16; and obtaining the generated antibody sequence from the system. (Item 26) Isolation antibodies produced by the method described in item 25. (Item 27) The aforementioned isolation antibody is the isolation antibody described in item 26, produced by recombinant technology. (Item 28) The isolation antibody is a chemically synthesized isolation antibody as described in any one of items 26 to 27. (Item 29) To provide a first antibody sequence for use in any one of items 1 to 9; and A method for maturing an antibody, comprising obtaining the generated antibody sequence from the system. (Item 30) Isolation antibodies produced by the method described in item 29. (Item 31) The aforementioned isolation antibody is the isolation antibody described in item 30, produced by recombinant technology. (Item 32) The isolation antibody is chemically synthesized, as described in any one of items 30 to 31. (Item 33) A method for determining an antibody sequence having improved properties, wherein the method is: To generate scores for each of a plurality of fine-tuned machine learning models trained on a plurality of corresponding initial antibody sequences, wherein each antibody sequence in the plurality of initial antibody sequences is marked by a corresponding characteristic, and each fine-tuned machine learning model is further generated by training each machine learning model on a plurality of second antibody sequences, wherein each antibody sequence in the plurality of second antibody sequences related to the target antigen is marked by a corresponding characteristic; and A method comprising generating antibody sequences based on an objective function by using a plurality of finely tuned machine learning models weighted by corresponding hyperparameters. (Item 34) A method for determining an antibody sequence having improved properties, wherein the method is: Training a machine learning model based on a first set of antibody sequences, each marked by corresponding characteristics; To generate a finely tuned machine learning model by training the machine learning model based on a second set of antibody sequences, wherein each antibody sequence of the second set of antibody sequences related to the target antigen is marked by a corresponding characteristic; and A method comprising generating an antibody sequence based on the aforementioned finely tuned machine learning model. (Item 35) A method for determining an antibody sequence having improved properties, wherein the method is: Each provides a score for each of several machine learning models trained on each of several antibody sequences, wherein each antibody sequence of each of the several antibody sequences is denoted by a characteristic corresponding to the several antibody sequences and a value for the characteristic corresponding to each of the several antibody sequences, and each score for each of the several machine learning models indicates and provides the contribution to predicting the characteristic corresponding to each of the machine learning models; and A method for generating an antibody sequence by using a plurality of machine learning models, weighting the output of each machine learning model according to each provided score and then combining the weighted outputs into a weighted sum. (Item 36) A method for determining an antibody sequence having improved properties, Weighting the output of each machine learning model according to the respective scores of the aforementioned multiple machine learning models, wherein each machine learning model is trained on respective multiple antibody sequences, each antibody sequence is denoted by a characteristic corresponding to the respective antibody sequence and a value of the characteristic corresponding to the respective antibody sequence, and the respective scores of each machine learning model indicate the contribution to predicting the characteristic corresponding to the respective machine learning model; and A method comprising generating an antibody sequence by using multiple machine learning models, which involves combining the weighted outputs into a weighted sum.

Claims

1. A method for determining an antibody sequence having improved properties, wherein the method is: To generate scores for each of multiple machine learning models, each machine learning model being trained on each of multiple antibody sequences, each antibody sequence being represented by a characteristic corresponding to each of the multiple antibody sequences and the value of the characteristic corresponding to the antibody sequence, and to generate scores for each of the multiple machine learning models indicating the contribution of the machine learning model to predicting the characteristic corresponding to each of the multiple antibody sequences on which the machine learning model is trained, and A method comprising generating antibody sequences by using a plurality of machine learning models, weighting the output of each machine learning model according to the generated score of each model and combining the weighted outputs into a weighted sum, wherein each antibody sequence is the amino acid sequence of an antibody.

2. The generation of the aforementioned antibody sequence further involves, Selecting antibody sequences from the proposed distribution based on the aforementioned multiple machine learning models; and The method according to claim 1, comprising determining whether the selected antibody sequence has an acceptable probability exceeding a specific threshold, analyzing the antibody sequence if so, and otherwise selecting the next antibody sequence from the proposed distribution.

3. The generation of the aforementioned antibody sequence further involves, Comparing a first characteristic value of an antibody sequence selected from a proposed distribution, determined by a function of the multiple machine learning models, with a second characteristic value of an antibody sequence having the best characteristic value in the current search, determined by the multiple machine learning models; and The method according to claim 1, further comprising replacing the antibody sequence having the best characteristic value with the antibody sequence selected from the proposed distribution if the first characteristic value is greater than the second characteristic value.

4. Training a machine learning model based on a first plurality of antibody sequences; and A finely tuned machine learning model is generated by training the machine learning model based on a second set of antibody sequences. Further including, generating the aforementioned fine-tuned machine learning model, Weighting the sequence characteristics of each of the second plurality of antibody sequences; Determining optimal model parameters for generating the second set of antibody sequences by using a machine learning model; and The method according to any one of claims 1 to 3, comprising applying the optimal model parameters to the machine learning model in order to generate an resulting model, wherein the resulting model is the finely tuned machine learning model.

5. The method according to any one of claims 1 to 4, wherein the characteristic corresponding to each of the plurality of antibody sequences is at least one of affinity, expression, protein aggregation, protein degradation stability, and off-target effect.

6. Based on the aforementioned multiple machine learning models, select antibody sequence candidates that fall within the specified acceptance criteria from the proposed distribution; and The method according to claim 1, further comprising: if the characteristics of the antibody sequence candidate are better than those of the best known antibody sequence, exchanging the best known antibody sequence with the antibody sequence candidate; otherwise, ignoring the antibody sequence candidate.

7. The method according to any one of claims 1 to 6, further comprising generating an antibody having the generated antibody sequence.

8. The method according to any one of claims 1 to 7, further comprising: providing a manufactured antibody having the generated antibody sequence; and analyzing the manufactured antibody with respect to a given property.

9. Training one or more machine learning models, each machine learning model being trained on the respective plurality of antibody sequences, each antibody sequence being denoted by the characteristic corresponding to the respective plurality of antibody sequences and the value of the characteristic corresponding to the antibody sequence, further comprising training one or more machine learning models: To provide a set of amino acid sequences marked by at least one characteristic; Masking a portion of a set of amino acid sequences in order to provide a masked set of amino acid sequences, wherein the portion of the set of amino acid sequences not included is an unmasked set of amino acid sequences: and The method according to any one of claims 1 to 8, comprising training one or more machine learning models to estimate each of the masked sets of amino acid sequences based on (1) at least one characteristic that identifies each masked amino acid sequence and (2) the unmasked sets of amino acid sequences and at least one characteristic that identifies each unmasked amino acid sequence.

10. The method according to any one of claims 1 to 9, wherein the generation of the antibody sequence is performed by employing MCMC sampling.

11. The method according to any one of claims 1 to 10, wherein each of the plurality of antibody sequences is related to the target antigen.

12. The aforementioned contributions will improve the following characteristics: The importance of at least one characteristic for the production of the antibody sequence, Expression of the aforementioned antibody sequence, Immunogenicity of antibodies in patients, Antibody expression level, Development feasibility, Interaction with other models, Orthogonality with other models, and Empirical derivation by adjusting the generation process, The method according to any one of claims 1 to 11, wherein the method is to predict at least one of the following.

13. A system for determining an antibody sequence having improved characteristics, comprising a processor and a memory storing computer code instructions, The processor and the memory, according to the computer code directive, perform the following: To generate scores for each of multiple machine learning models, each machine learning model being trained on each of multiple antibody sequences, each antibody sequence being represented by a characteristic corresponding to each of the multiple antibody sequences and the value of the characteristic corresponding to the antibody sequence, and to generate scores for each of the multiple machine learning models indicating the contribution of the machine learning model to predicting the characteristic corresponding to each of the multiple antibody sequences on which it was trained, and The antibody sequence is generated by using the multiple machine learning models, weighting the output of each machine learning model according to the generated score, and then combining the weighted outputs into a weighted sum. A system configured to allow the aforementioned system to function, wherein each antibody sequence is an amino acid sequence of an antibody.

14. In order to generate the antibody sequence, the processor and the memory, according to the computer code command, Selecting antibody sequences from the proposed distribution based on the aforementioned multiple machine learning models; and To determine whether the selected antibody sequence has an acceptable probability exceeding a specific threshold, and if so, to analyze the antibody sequence; otherwise, to select the next antibody sequence from the proposed distribution. The system according to claim 13, configured to cause the system to perform the above action.

15. In order to generate the antibody sequence, the processor and the memory, according to the computer code command, Comparing a first characteristic value of an antibody sequence selected from a proposed distribution, determined by a function of the multiple machine learning models, with a second characteristic value of an antibody sequence having the best characteristic value in the current search, determined by the multiple machine learning models; and If the first characteristic value is greater than the second characteristic value, the antibody sequence having the best characteristic value is exchanged with the antibody sequence selected from the proposed distribution. The system according to claim 13, configured to cause the system to perform the above action.

16. The processor and the memory, according to the computer code command, Training a machine learning model based on a first set of antibody sequences; and A finely tuned machine learning model is generated by training the machine learning model based on a second set of antibody sequences. The system is further configured to allow the fine-tuned machine learning model to be generated, Weighting the sequence characteristics of each of the second plurality of antibody sequences; Determining optimal model parameters for generating the second set of antibody sequences by using a machine learning model; and The system according to any one of claims 13 to 15, comprising applying the optimal model parameters to the machine learning model in order to generate an resulting model, wherein the resulting model is the finely tuned machine learning model.

17. The system according to any one of claims 13 to 16, wherein the characteristic corresponding to each of the plurality of antibody sequences is at least one of affinity, expression, protein aggregation, protein degradation stability, and off-target effect.

18. The processor and the memory, according to the computer code directive, perform the following: Based on the aforementioned multiple machine learning models, select antibody sequence candidates that fall within the specified acceptance criteria from the proposal distribution; and If the characteristics of the antibody sequence candidate are better than those of the best known antibody sequence, the best known antibody sequence and the antibody sequence candidate are exchanged, or otherwise the antibody sequence candidate is ignored. The system according to claim 13, further configured to cause the system to perform the above action.

19. The processor and the memory are configured to cause the system to further train one or more machine learning models by computer code commands, each machine learning model being trained based on the respective plurality of antibody sequences, each antibody sequence being denoted by the characteristic corresponding to the respective plurality of antibody sequences and the value of the characteristic corresponding to the antibody sequence, and the training of one or more machine learning models is further, To provide a set of amino acid sequences marked by at least one characteristic; Masking a portion of a set of amino acid sequences in order to provide a masked set of amino acid sequences, wherein the portion of the set of amino acid sequences not included in the masked portion is an unmasked set of amino acid sequences; and The system according to any one of claims 13 to 18, comprising training one or more machine learning models to estimate each of the masked sets of amino acid sequences based on (1) the at least one characteristic that identifies each masked amino acid sequence and (2) the at least one characteristic that identifies the unmasked sets of amino acid sequences.

20. The system according to any one of claims 13 to 19, wherein the generation of the antibody sequence is performed by employing MCMC sampling.

21. The system according to any one of claims 13 to 20, wherein each of the plurality of antibody sequences is related to a target antigen.

22. The aforementioned contributions will improve the following characteristics: The importance of at least one characteristic for the production of the antibody sequence, Expression of antibody sequences, Immunogenicity of antibodies in patients, Antibody expression level, Development feasibility, Interaction with other models, Orthogonality with other models, and Empirical derivation by adjusting the generation process, The system according to any one of claims 13 to 21, wherein the system predicts at least one of the following.

23. A method for determining an antibody sequence having improved properties, wherein the method is: To generate scores for each of a plurality of fine-tuned machine learning models trained on a plurality of corresponding initial antibody sequences, wherein each antibody sequence in the plurality of corresponding initial antibody sequences is marked by a corresponding characteristic, and each fine-tuned machine learning model is further generated by training each machine learning model on a plurality of second antibody sequences, wherein each antibody sequence in the plurality of second antibody sequences related to the target antigen is marked by a corresponding characteristic; and A method comprising generating antibody sequences based on an objective function by using a plurality of finely tuned machine learning models weighted by corresponding hyperparameters, wherein each antibody sequence is an amino acid sequence of an antibody.

24. A method for determining an antibody sequence having improved properties, wherein the method is: Training a machine learning model based on a first set of antibody sequences, wherein each antibody sequence in the first set of antibody sequences is denoted by a corresponding characteristic; To generate a finely tuned machine learning model by training the machine learning model trained on the first plurality of antibody sequences on a second plurality of antibody sequences, wherein each antibody sequence of the second plurality of antibody sequences related to the target antigen is marked by a corresponding characteristic; and A method comprising generating antibody sequences based on the aforementioned fine-tuned machine learning model, wherein each antibody sequence is an amino acid sequence of an antibody.

25. A method for determining an antibody sequence having improved properties, wherein the method is: To provide scores for each of a plurality of machine learning models, each machine learning model being trained on each plurality of antibody sequences, each antibody sequence being denoted by a characteristic corresponding to each plurality of antibody sequences and a value of the characteristic corresponding to the antibody sequence, and the scores provided to each of the plurality of machine learning models indicating and providing the contribution of the machine learning model to predicting the characteristic corresponding to each plurality of antibody sequences on which it is trained; and A method comprising generating antibody sequences by using a plurality of machine learning models, weighting the outputs of each machine learning model according to each provided score and synthesizing the weighted outputs into a weighted sum, wherein each antibody sequence is the amino acid sequence of an antibody.

26. A method for determining an antibody sequence having improved properties, Weighting the output of each machine learning model according to the respective scores of multiple machine learning models, wherein each machine learning model is trained on each of several antibody sequences, each antibody sequence is denoted by a characteristic corresponding to each of the several antibody sequences and the value of the characteristic corresponding to the antibody sequence, and the score of each machine learning model indicates the contribution of the machine learning model to predicting the characteristic corresponding to each of the several antibody sequences on which it is trained; and The weighted outputs are combined into a weighted sum. A method comprising generating antibody sequences by using multiple machine learning models, wherein each antibody sequence is an amino acid sequence of an antibody.

Citation Information

Patent Citations

  • Synthesizing vaccines, immunogens, and antibodies

    US20180282376A1

  • System and methods for machine learning for drug design and discovery

    US20190304568A1