Antibody competition model using hidden variable affinity
Patent Information
- Application Number
- JP2024500487
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-07-08
- Filing Date
- 2022-07-08
- Publication Date
- 2025-07-16
AI Technical Summary
Traditional methods for monoclonal antibody discovery, particularly epitope binning, provide limited insight into how antibodies compete when binding antigens, leading to resource-intensive and time-consuming processes.
A system and method for deriving hidden variables based on antibody competition data using an optimization engine to generate training data, which processes pairwise competition data to determine hidden variable affinity scores, allowing for more sophisticated prediction of antibody competition patterns.
This approach reduces the need for extensive experimental runs by predicting competition between antibodies without direct observation, improving resource efficiency and providing higher fidelity in understanding binding patterns, enabling more robust monoclonal antibody discovery and manufacturing.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT This invention was made with United States Government support under D18AC00002 awarded by the Defense Advanced Research Projects Agency. The United States Government has certain rights in this invention.
[0002] Field Embodiments of the present disclosure generally relate to deriving hidden variables based on antibody competition data to discover binding patterns. [Background technology]
[0003] background Monoclonal antibody ("mAB") discovery is a complex, time-consuming, and resource-intensive technical challenge. One component of mAB discovery involves understanding how antibodies compete when binding to antigens. Epitope binning is a valuable technique that can further understand this. However, traditional approaches can only achieve models with limited insight into how these antibodies compete. Summary of the Invention [Means for solving the problem]
[0004] Abstract An embodiment of the present disclosure is directed to a system and method for deriving hidden variables based on antibody competition data to discover binding patterns. Antibody competition data for a plurality of antibodies and antigens can be received, the antibody competition data including data values indicative of pair-wise competition between the antibodies. The antibody competition data can be processed to generate training data. Using the training data and an optimization engine, a plurality of hidden variables and affinity scores for the hidden variables can be derived, the affinity scores for the hidden variables are derived for each antibody, and the hidden variables represent competitive factors for the antigens that cause competition between the antibodies.
[0005] Features and advantages of the embodiments will be set forth in the description that follows, or will be obvious from the description, or may be learned by practice of the disclosure.
[0006] Further embodiments, details, advantages and modifications will become apparent from the following detailed description of the preferred embodiments, taken in conjunction with the accompanying drawings. [Brief description of the drawings]
[0007] [Figure 1] FIG. 1 illustrates a system for deriving hidden variables based on antibody competition data for discovering binding patterns according to an exemplary embodiment.
[0008] [Diagram 2] FIG. 2 illustrates a simplified diagram of a computing system in accordance with an exemplary embodiment.
[0009] [Diagram 3] FIG. 3 shows a conventional heat map showing competition data for monoclonal antibodies.
[0010] [Figure 4] FIG. 4 shows a previous network approach for binning monoclonal antibodies based on competition data.
[0011] [Diagram 5] FIG. 5 shows the competitive dynamics for monoclonal antibodies.
[0012] [Figure 6] FIG. 6 shows a flow diagram for deriving hidden variables based on antibody competition data for discovering binding patterns according to an exemplary embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] Detailed Description The embodiments derive hidden variable information indicative of competition patterns between monoclonal antibodies based on pairwise antibody competition data. For example, a predictive mathematical model of antibody-antigen binding can be discovered by an optimization engine. In some embodiments, the optimization engine can derive a set of hidden variables that form the basis for generating predictions regarding whether pairs of antibodies will compete with each other. These hidden variables can be loosely thought of as the epitope binding resources or "antigen real estate" used by antibodies during binding.
[0014] In some embodiments, the variables are "hidden" because the model is agnostic as to where these resources actually reside on the antigen surface. For example, each hidden variable may be a placeholder for some epitope resource on the antigen that the antibody uses to bind. In some implementations, some hidden variables may represent some other competing factors (e.g., other than epitope / positional competition).
[0015] In some embodiments, the optimization engine can generate logit values for the hidden variables and compare these logit values to observed competition data values (e.g., pairwise antibody competition) present in the training data for the antibodies. In some embodiments, the loss function can be optimized by implementing a gradient that adjusts the logit values of the antibody hidden variables until the loss function is optimized and / or a metric is achieved (e.g., convergence is achieved). For example, optimization of the logit values of the antibody hidden variables can achieve an affinity score for the hidden variables that indicates / predicts the level of competition of the antibody with respect to the competition factor represented by the hidden variable (e.g., with respect to the epitope on the antigen represented by the hidden variable). In some embodiments, pairwise competition prediction scores between antibodies can be generated using the logit values for the antibodies.
[0016] Embodiments can also implement ensemble learning techniques by combining predictions (e.g., competition scores) from multiple hidden variable models trained on different antibody competition data. For example, each hidden variable model can be trained using competition data for a different set of antibodies. A prediction as to whether two antibodies will compete can be generated by combining the competition scores from several trained hidden variable models.
[0017] In some embodiments, the landmark antibody correlation model can use competition measurements for a given set of landmark antibodies to predict pairwise competition (e.g., for a particular antigen) between unmeasured antibodies. For example, given a pair of antibodies for which competition prediction is desired, the correlation between the competition measurements of each antibody can be calculated using the landmark antibodies. Based on the correlation values, the competition likelihood can be predicted.
[0018] Traditional epitope binning involves testing antibodies (e.g., using a device that performs "experimental runs") in a combinatorial manner (e.g., pairwise) to derive competition data that is analyzed such that antibodies competing for the same binding region (e.g., epitope) are grouped together in bins. Epitope binning experiments generate large amounts of data. For example, in some binning experiments, a data point (e.g., a numerical value) is generated for each pair of antibodies involved for each experimental run. Some runs may include up to 384 antibodies per experiment, meaning that there are up to 384*384=147,456 observations of pairwise competition between different antibodies. Furthermore, sometimes it is advantageous to perform epitope binning across an even larger set of antibodies than current devices support in a single experimental run, or to extend a previous epitope binning run with newly discovered antibodies without performing competition experiments with all pairs of these antibodies.
[0019] The embodiments achieve an improved model for analyzing and understanding the results of single or multiple epitope binning runs. Additionally, the improved model can attribute experimental results to the properties of individual antibodies such that they can be grouped together in a more informative way than simply assigning each antibody to a single bin. The embodiments support techniques for combining results from multiple epitope binning experiments, which are currently limited to 384 antibodies at a time by device limitations. The embodiments also enable the extension of current epitope binning runs with new antibodies without repeating the entire experiment. Additionally, for antibodies involved in different epitope binning runs (e.g., for which there is no direct experimental information on whether they compete), the embodiments of the model support predictions on whether these antibodies compete.
[0020] Embodiments optimize techniques for collecting and organizing pairwise antibody competition measurements for a particular antigen by using models that can predict these pairwise antibody competition measurements before (or without) performing experimental runs that actually measure them. Thus, embodiments can significantly reduce the number of experimental runs required to generate the desired antibody competition data (and significantly improve resource efficiency) when compared to traditional epitope binning approaches.
[0021] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. In the following detailed description, several specific details are set forth to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the embodiments. Wherever possible, like reference numbers will be used for like elements.
[0022] 1 illustrates a system for deriving hidden variables based on antibody competition data for discovering binding patterns according to an exemplary embodiment. The system 100 includes antibody competition data 102, a processing module 104, an optimization engine 106, and an analysis module 108. For example, the antibody competition data 102 may include data generated from a surface plasmon resonance ("SPR") experimental technique that generates numerical results characterizing antibodies and their interactions with antigens (e.g., pairwise competition). In some implementations, the antibody competition data is generated using a Carterra® LSA™ instrument.
[0023] In some embodiments, antibody competition data 102 may include data from several experimental runs. For example, an experimental run may generate a numerical value indicative of pairwise competition between two antibodies for a given antigen, and may include data regarding interactions between several (e.g., tens, hundreds, or thousands) of antibodies in total antibody competition data 102. In some embodiments, competition data 102 includes binary pairwise competition data that indicates whether two antibodies compete using a binary value (e.g., 1 or 0, true or false, etc.).
[0024] The processing module 104 can process the antibody competition data 102 such that training data is generated for the optimization engine 106. For example, the antibody competition data 102 may include data from multiple experimental runs, and the processing module 104 can combine this data in a manner suitable for processing by the optimization engine 106. An embodiment of the processing module 104 can also transform numerical values derived from the competition data 102 using a function (e.g., a function that assigns binary values) or perform other suitable data transformations.
[0025] The optimization engine 106 can derive the hidden variables and affinity scores for the hidden variables for the involved antibodies based on the training data generated by the processing module 104. For example, the optimization engine 106 can generate logit values for the hidden variables (e.g., logit values representing affinity scores for the antibody hidden variables) and compare these logit values to observed competition data values (e.g., pairwise antibody competition) present in the training data for the antibodies. In some embodiments, the loss function can be optimized by implementing a gradient that adjusts the logit values of the antibody hidden variables until the loss function is optimized and / or a metric is achieved (e.g., convergence is achieved).
[0026] For example, optimization of the logit value of the latent variable for an antibody can arrive at an affinity score for the latent variable that indicates / predicts the level of competition of the antibody for the competitor represented by the latent variable (e.g., the epitope on the antigen represented by the latent variable). In some implementations, the latent variable may correlate with the competition factor for antigen binding across the epitope location (e.g., interference / competitor across competition for the same binding location).
[0027] The analysis module 108 can generate competitive information for antibodies based on the output from the optimization engine 106. For example, the optimization engine 106 can output a model for predicting / discovering competition between multiple antibodies. In particular, the model generated by the optimization engine 106 can discover antibodies that compete for different competitors (e.g., different epitopes or other competitors). Thus, the analysis module 108 can be used to generate a panel of antibodies with different hidden variable affinity values (e.g., antibodies that compete for antigens in different ways). Such a panel can provide diverse pathways to positive treatment outcomes, thus representing an improvement to the manufacturing / discovery of monoclonal antibodies that result in positive health outcomes.
[0028] FIG. 2 is a schematic diagram of a computing system 200 according to an embodiment. As shown in FIG. 2, the system 200 may include a bus 210, as well as other elements, configured to communicate information between a processor 212, data 214, memory 216, and / or other components of the system 200. The processor 212 may include one or more general or special purpose processors configured to execute commands, perform calculations, and / or control functions of the system 210. The processor 212 may include a single integrated circuit, such as a microprocessing device, or may include multiple integrated circuit devices and / or circuit boards operating in combination. The processor 212 may execute software, such as an operating system 218, an optimization engine 220, and / or other applications stored in the memory 216.
[0029] The communications component 222 may enable connections between components of the system 200 and other devices, such as by processing (e.g., encoding) data transmitted from one or more components of the system 200 to another device over a network (not shown) and processing (e.g., decoding) data received from another system over the network for one or more components of the system 200. For example, the communications component 222 may include a network interface card configured to provide wireless network communications. Any suitable wireless communication protocol or technology may be implemented by the communications component 222, such as Wi-Fi, Bluetooth, Zigbee, radio, infrared, and / or mobile communications technologies and protocols. In some embodiments, the communications component 222 may provide wired network connections, technologies, and protocols, such as Ethernet.
[0030] The system 200 includes a memory 216 that can store information and instructions for the processor 212. An embodiment of the memory 216 contains components for retrieving, reading, writing, modifying, and storing data. The memory 216 may store software that performs functions when executed by the processor 212. For example, the operating system 218 (and the processor 212) can provide operating system functions for the system 200. The optimization engine 220 (and the processor 212) can generate models for predicting / discovering antibody competition according to embodiments. An embodiment of the optimization engine 220 can be implemented as an in-memory configuration. The software modules of the memory 216 may include the operating system 218, the optimization engine 220, as well as other application modules (not depicted).
[0031] Memory 216 includes non-transitory computer-readable media accessible by the components of system 200. For example, memory 216 may include any combination of random access memory ("RAM"), dynamic RAM ("DRAM"), static RAM ("SRAM"), read-only memory ("ROM"), flash memory, cache memory, and / or any other type of non-transitory computer-readable media. Database 214 is communicatively coupled (such as via bus 210) to other components of system 200 to provide storage for the components of system 200. An embodiment of database 214 may store data in an integrated collection of logically related records or files.
[0032] The database 214 may be a data warehouse, a distributed database, a cloud-based database, a secure database, an analytical database, a production database, a non-production database, an end-user database, a remote database, an in-memory database, a real-time database, a relational database, an object-oriented database, a hierarchical database, a multidimensional database, a Hadoop Distributed File System ("HFDS"), a NoSQL database, or any other database known in the art. The components of the system 200 are further coupled (e.g., via bus 210) to a display 224, an I / O device 226 such as a keyboard, and an I / O device 228 such as a computer mouse or any other suitable I / O device that enables the processor 212 to display information, data, and any other suitable display matter to a user.
[0033] In some embodiments, system 200 may be an element of a system architecture, a distributed system, or other suitable system. For example, system 200 may include one or more additional functional modules, which may include various modules of a Carterra® LSA™ instrument, any other suitable device for generating antibody competition data, or any other suitable module.
[0034] An embodiment of system 200 may provide related functionality remotely for separate devices. In some embodiments, one or more components of system 200 may not be implemented. For example, system 200 may be a tablet, smartphone, or other wireless device that includes a display, one or more processors, and memory, but does not include one or more other components of system 200 shown in FIG. 2. In some embodiments, an implementation of system 200 may include additional components not shown in FIG. 2. Although FIG. 2 depicts system 200 as a single system, the functionality of system 200 may be implemented in different locations, as a distributed system, in a cloud infrastructure, or in any other suitable manner. In some embodiments, memory 216, processor 212, and / or database 214 are distributed (across multiple devices or computers that illustrate system 200). In one embodiment, system 200 may be part of a computing device (e.g., a smartphone, a tablet, a computer, etc.).
[0035] Monoclonal antibody ("mAB") discovery is a complex, time-consuming, and resource-intensive technical challenge. One component of mAB discovery involves understanding how antibodies compete when binding to an antigen. Epitope binning is a useful approach to further understand this. In particular, traditional epitope binning involves testing antibodies (e.g., using a device that performs an "experimental run") in a combinatorial manner (e.g., pairwise) to derive competition data that is analyzed such that antibodies competing for the same binding region (e.g., epitope) are grouped together in bins. Exemplary competition data generated by an experimental run is depicted in heatmap 300 of FIG. 3. Heatmap 300 includes different antibodies across rows and columns, with the numerical values at the intersections of two antibodies indicating pairwise competition between them.
[0036] Some previous approaches to epitope binning are based on graph clustering algorithms. In these approaches, antibodies are assigned to a single cluster based on the proximity or number of connecting edges in a competition graph. Figure 4 shows a conventional network approach for binning monoclonal antibodies based on competition data. The network graph 400 includes bins 402 that use previous graph clustering approaches. As depicted in Figure 4, each antibody is assigned to a single bin 402, or cluster, based on the antibody's competition profile.
[0037] Rather than assigning antibodies to a single bin, embodiments assign affinity score numbers to each antibody based on a set of hidden variables. These hidden variables can be used, for example, to predict competition between unobserved pairs of antibodies, and they are also inherently useful in understanding the binding patterns that antibodies use to bind antigens.
[0038] Several advantages are achieved by the hidden variable approach implemented by the embodiments. For example, the embodiments can model observed experimental data with higher fidelity than cluster-based models that assign each antibody to a single cluster. In particular, when the experimental data shows a non-transitive pattern of antibody competition, this cannot be well represented using a model that assigns antibodies to a single cluster. Figure 5 shows the competitive dynamics for a monoclonal antibody that exhibits this shortcoming in previous approaches. In particular, diagram 500 shows: Group A competes heavily with Group B Group B will compete heavily with Group C Group A does not compete significantly with Group C Describe the following.
[0039] A cluster-based model cannot determine a single cluster to which these groups should be assigned. However, a "hidden variable" model can explain this pattern of competition by assigning the different affinity antibodies of each group to two different hidden variables: [Table 1]
[0040] Thus, while previous approaches had limited insight, higher fidelity analyses can be derived using embodiments of the hidden variable approach. It can be useful to think of hidden variables as allowing a single antibody to belong to more than one cluster at the same time, and with affinity numerical values rather than binary decisions about membership to a particular group. Together, these properties allow the model to make predictions about antibody competition.
[0041] Another limitation of the cluster-based approach is that the resulting model cannot make robust predictions about whether antibodies will compete, such as by combining competition data to generate a larger competition matrix across different runs of an experimental setup. Rather than simply assigning each antibody to a single cluster, embodiments generate models that assign affinity numerical values for each antibody to different hidden variables. This approach supports numerical predictions about whether antibodies from different epitope binning runs will compete with each other.
[0042] The advantage of combining data from multiple epitope binning runs together can be considered as a novel approach to the commonly known matrix competition problem. For example, multiple epitope binning runs often do not completely intersect, so a matrix describing all pairs of antibody competition is incomplete (e.g., if antibodies across multiple runs are listed in a table format on rows and columns, data for some intersections will be missing). When a matrix represents pairwise antibody competition, the hidden variable affinity approach taken by the embodiment can "complete" the incomplete matrix by means of optimization based on available competition data.
[0043] In some embodiments, the numerical predictions derived from the model can be interpreted as confidence scores that allow the model to include noisy and / or conflicting experimental evidence and thus make predictions with greater or lesser confidence (e.g., depending on the strength of the evidence). An additional benefit of the hidden variable affinity scores is that the model can aid in dimensionality reduction techniques such as t-SNE or Uniform Manifold Approximation and Projection (UMAP) to generate two-dimensional clustering plots showing the relationships between groups of antibodies across multiple epitope binning runs. For example, a pairwise distance matrix can be calculated between the affinities of the hidden variables for each antibody using any suitable distance metric such as Euclidean, Manhattan, etc., and the distance matrix can be run through a dimensionality reduction system such as UMAP. Some techniques can also impute a full competitive matrix for a set of epitope binning runs, use the distances between its columns and rows to calculate a pairwise distance matrix for the antibodies, and send the pairwise distance matrix through a dimensionality reduction system.
[0044] The affinity scores of the hidden variables in the embodiment can be stored as a table of "hidden logits" that represent the affinity of each antibody with each hidden variable. For example, the hidden logits can be any finite number. In some embodiments, positive values represent higher affinity and negative values represent lower affinity. Below is an example table showing the hidden logit values of each antibody for an antibody and a set of three hidden variables: [Table 2]
[0045] Before using the hidden logits in the model, in some embodiments, the values can be passed through a sigmoid function. For example, the sigmoid function can transform them into the range of (0...1), where a hidden logit value greater than 0 results in a hidden variable greater than 0.5. Below is an example of a hidden logit: [Table 3]
[0046] The label "hidden logit" is used because the value of "hidden logit" represents the normalized affinity score between each antibody and each hidden variable. In some embodiments, model fitting is used to saturate the affinities of the hidden variables as close to 0 or 1 as possible so that they can be interpreted as a binary decision as to whether an antibody requires a particular hidden variable to bind, although this is not always possible due to conflicting evidence and other factors.
[0047] In some embodiments, the prediction of whether two antibodies will compete is based on a criterion of how much overlapping hidden variable resources these two antibodies require. One approach to achieve this criterion is a dot product operation, which multiplies together the values in the corresponding columns in the row in question. Finally, the predicted competition score is, in some embodiments, transmitted by a sigmoid operation, whose value is in the range of (0...1).
[0048] Below is an example algorithm showing how the model predicts whether two antibodies compete, according to some embodiments:
number
[0049] In the above algorithm, HV denotes a lookup in a table of hidden variables (e.g., hidden logits after being transformed to a range (0...1) using a sigmoid function). The value α may represent a temperature parameter on the outer sigmoid, which may take any suitable value (e.g., 5, or any other suitable value). Note that this embodiment of the algorithm implements two applications of the sigmoid function: 1) the table of hidden variables when creating the HV, and 2) the outermost operation when calculating the competition function.
[0050] The embodiments can also implement ensemble learning techniques by combining predictions (e.g., competition scores) derived from multiple hidden variable models trained on different antibody competition data. For example, each hidden variable model can be trained using competition data for a different set of antibodies (e.g., randomized training sets). A prediction on whether two antibodies will compete can be generated by combining the competition scores (e.g., calculated by a dot product operation, as disclosed above) derived from several trained hidden variable models. The combined score can be an average, a weighted average, or a combination calculated by any other suitable mathematical operation.
[0051] In some embodiments, multiple versions of the hidden variable model are trained using different subsets of antibody competition training data. For example, most pairwise competition measurements for groups of antibodies in a given subset of training data are completely removed. In other words, rather than simply removing random pairwise competition measurements to generate a subset of training data, the embodiment selectively removes most of the pairwise competition data for groups of antibodies. This selective removal of competition data for groups of antibodies in different subsets of training data achieves uncorrelated trained hidden variable models. Uncorrelated models achieve better results when they are combined in an ensemble approach.
[0052] In some embodiments, most of the pairwise competition data for the group of antibodies is removed to generate a subset of training data, but some competition data for the group can be maintained. For example, a given set of antibodies from the full set of training data can be designated as persistent antibodies, and the competition data for these persistent antibodies can be maintained across the subset of training data. In these embodiments, when selectively removing pairwise competition data for the group of antibodies to generate a subset of training data, the pairwise competition data between the group of antibodies and the persistent antibodies is maintained in the subset of training data. These embodiments can train uncorrelated models that each benefit from training aided by competition data for persistent antibodies.
[0053] One advantage of this ensemble technique is that the variance in predictions (e.g., calculated variance metrics) from individual models in the ensemble provides a measure of the reliability or certainty of the ensemble model in its overall predictions. Furthermore, these variances also provide an indication of how much the predictions may change if the underlying data distribution changes. In some embodiments, the ensemble technique can combine pairwise antibody competition predictions from one or more models of hidden variables and any other suitable models (e.g., landmark correlation models).
[0054] Some embodiments utilize a dot product operation to generate the prediction scores. Although some simple dot product models exist for constructing and optimizing certain problems in the field of machine learning, there are differences between the optimization model embodiments and some of the existing models: Certain natural language processing models, such as word2vec, model probability distributions over neighboring words, whereas embodiments model antibody competition as a binary event. For example, embodiments aim to saturate the outer sigmoid when possible, whereas this is less desirable in word2vec, and due in part to this difference, a shift forward of the outer sigmoid is implemented; Some current models can use more hidden variables than embodiments, such as large word embeddings. For example, when hidden variable models according to some embodiments are compared to natural language processing models, some differences are that the training data in hidden variable model embodiments is smaller; and hidden variable model embodiments are used to explore hidden variables to uncover patterns.
[0055] In some embodiments, the number of hidden variables and the degree of shift after the dot product are tunable hyperparameters of the model. For example, empirically, 5-10 hidden variables are sufficient to represent many competing patterns in the data for some datasets, although any other suitable number of hidden variables can be implemented. For ease of interpretation, the values of the hidden variables are constrained in some embodiments to be in the range (0...1).
[0056] In some embodiments, model training is used to calculate values for the hidden variables. For example, values for the hidden variables are derived based on experimental data from an epitope binning run. Examples of experimental data generated by a run are as follows: [Table 4]
[0057] Note that the tabular experimental data aids in the concatenation, or stacking, of experimental results from different epitope binning runs. This concatenation is used together with a numerical optimization procedure implemented by some embodiments to achieve a concatenated optimization of the values of hidden variables using data from multiple different epitope binning runs. Note that a sufficient set of cross-run antibodies that participated in both epitope binning runs is kept so that the same hidden variables have the same meaning for the antibodies that were present in the different runs. To optimally select cross-run antibodies, one or more predictive models can identify a small set of antibodies from the first run that are least correlated with each other in their competitive behavior, and these can be selected as the cross-run antibodies to be used in subsequent runs.
[0058] The embodiments derive the hidden variables and values of the hidden variables (e.g., affinity scores of the hidden variables) using numerical optimization techniques, such as a form of gradient descent, to optimize the values so that the model predictions match the actual experimental data according to a loss function. Several potential algorithms or techniques can be used to accomplish this task. Below is an example of an optimization procedure according to some embodiments. · 0) We start with a single hidden variable H0 initialized with a hidden logit value of 0.0 for all antibodies, implying an arbitrary hidden variable value of 0.5. Note that this is the point on the sigmoid curve where the slope is highest, i.e. the sigmoid is least saturated. 1) Optimize the cross-entropy loss of the models' competing predictions on the training set using an optimization procedure such as the Broyden-Fletcher-Goldfarb-Shanno algorithm ("BFGS") or L-BFGS. With only a single hidden variable, this optimization problem can be convex. The solution tends to represent the single largest competing group in the dataset and use this single hidden variable to make predictions. Optionally add a small L2 norm term to the cross entropy. 2) After the optimization has converged, add a new hidden variable column for each antibody that is reinitialized to a median logit value of 0.0, which means a hidden variable value of 0.5 for each antibody. 3) Jointly optimize all hidden variables for all antibodies to predict pairwise competition events using the same loss function and optimization procedure derived from step 1. This optimization problem is no longer convex, but because the first hidden variable has likely been moved away from 0.5 for some antibodies, the symmetry is broken and the optimizer can move forward. At this point, the optimizer will often pick the second most common group of competitors in the dataset. New hidden variables can be added incrementally by repeating steps 2 and 3 until the predictive performance of the model on the provided validation set converges. Note that this optimization procedure is entirely deterministic.
[0059] This technique of incrementally adding hidden variables is reminiscent of the idea of boosting from machine learning, where when an iterative series of weak learners is trained, each one corrects the errors made by the model in previous iterations. The optimization procedure described above differs from boosting in that old and new hidden variables are jointly optimized, allowing hidden variables from earlier iterations to be refined each time a new one is added.
[0060] The approach of incrementally adding hidden variables is also vaguely similar to the representation of matrices using lower-rank approximations, as is done in principal component analysis ("PCA"). However, the optimization procedure described above differs from PCA in that the matrix reconstruction function contains nonlinearities and does not track the structure of the standard truncated singular value decomposition, allowing the reconstruction of incomplete matrices with missing values, and the loss function is one that allows multiple epitope binning runs to be concatenated together.
[0061] Hidden variable model embodiments also differ from natural language processing models such as word2vec, for example, due at least to differences in training. Some further differences are as follows: Hidden variable model embodiments optimize using the full training set rather than a batch. For example, the training set in embodiments is smaller in scale than many natural language processing applications. Thus, the full training set can be used for each round of gradient calculation, and quadratic optimization methods such as LBFGS can be used to help speed convergence and avoid hyperparameter tuning. The quadratic optimizer may perform better on embodiments of the hidden variable model topology because of the two-layer sigmoid. If any sigmoid saturates, the gradient signal resulting from it may be very weak, and the embodiment pushes the sigmoid towards saturation. The implementation of the hidden variable model does not require a large number of hidden variables or random initialization for them. · The implementation of the hidden variable model can incrementally add hidden variables, rather than starting with a fixed size embedding.
[0062] In some embodiments, other optimization techniques can be implemented. For example, another option for optimization may be to start from a fixed number of hidden variables. In these embodiments, a random initialization of the hidden variables can be implemented to break the symmetry of the problem and allow the optimizer to progress; however, this implementation avoids the need to incrementally add hidden variables. Any other suitable optimization technique can be implemented.
[0063] In some embodiments, once the model is trained, the affinity scores of the hidden variables can be used to understand the tendency of competition between antibodies. For example, a novel property of the embodiments is that, compared to previously implemented clustering techniques, the embodiments associate the antibodies with multiple groups. This can be achieved by thresholding the affinity scores of each antibody with respect to the hidden variables at some cutoff value, such as 0.5. In some embodiments, an antibody may have an affinity higher than 0.5 for multiple hidden variables and thus belong to more than one group. This novel group membership allows further questioning of how these groups of antibodies intersect.
[0064] Some embodiments can be used to detect a pattern of competition referred to herein as a "competitive sandwich." A competitive sandwich comprises three groups of antibodies with the competitive profile shown in FIG. 5, namely: Group A competes heavily with Group B Group B will compete heavily with Group C Group A does not compete significantly with Group C.
[0065] In this case, groups A and C are the "slices of bread" and group B is the "fillings" of the sandwich. Figure 5 depicts a graph-based overview of the competitive sandwiches.
[0066] In some circumstances, a competitive sandwich may indicate partial adjacency or ordering of epitopes, one "sandwiched" between two others. A common way to think of this idea is that it identifies groups in a connectivity graph where transitivity does not hold. This effect can be interesting in many contexts. As an example, consider how a competitive sandwich can indicate associations between groups of people: Almost everyone in group A knows each other, and almost everyone in group B knows each other. Almost everyone in group B knows each other, and almost everyone in group C knows each other. · Nearly everyone in group C knows each other, but hardly anyone in group A knows each other.
[0067] This may indicate that there are two different / separate underlying reasons that these populations of people know about each other, and for some people (e.g., group B), both reasons apply. Similarly, competitive sandwiches may indicate different / separate competition between antibodies in some situations.
[0068] In some embodiments, the derived latent variables and affinity scores for the latent variables can provide a model for predicting competition between antibodies without requiring an explicit experimental run to observe the competition. In other words, the derived model can be used to predict competition for antibody pairs that have not been experimentally tested and observed. Embodiments of the derived model can be considered as a predictive or simulation tool for antibody competition. Thus, embodiments improve antibody competition testing by improving resource and time efficiency.
[0069] In some embodiments, the derived model can also predict high fidelity competitive dynamics between antibodies. For example, antibodies exhibiting different latent variable affinity scores may exhibit different binding mechanisms, while antibodies exhibiting similar latent variable affinity scores may exhibit similar binding mechanisms. The latent variable affinity score descriptors provide a more sophisticated view of competitive dynamics when compared to previous approaches that simply associate antibodies with individual bins. Embodiments of the derived model allow for the selection of antibodies exhibiting diverse latent variable affinities to achieve a more robust monoclonal antibody discovery and manufacturing process.
[0070] In some implementations, antibody competition can be predicted using a landmark antibody correlation model. For example, the landmark antibody correlation model can characterize each antibody in terms of its competition with a set of given landmark antibodies. One way to consider these competition measurements with a given landmark antibody is as a substitute for a hidden variable in the hidden variable model disclosed herein, however, the competition measurements with a given landmark antibody are not hidden.
[0071] Below is an example showing a landmark antibody competition model that references a set of landmark antibody competition measurements. [Table 5]
[0072] In this example, antibodies A, B, and C are landmark antibodies, and measurements have been taken that indicate competition with each of four other antibodies, D, E, F, and G. For a set of antibodies for which pairwise predictions are desired (e.g., antibodies D, E, F, and G), an embodiment of the landmark antibody correlation model utilizes previously taken competition measurements against the same set of given landmark antibodies (e.g., A, B, C). However, pairwise competition measurements between these other non-landmark antibodies have not been taken, and therefore these values are unknown (e.g., it is unknown whether antibodies D and E compete with each other).
[0073] To predict the likelihood of competition between antibodies D and E, the correlation between the columns of these two antibodies is calculated. According to the above example, the correlation between columns D and F is 1.0, predicting a high likelihood that antibodies D and F will compete with each other since they have exactly the same competition profile against the landmark antibodies. On the other hand, the correlation between antibodies D and E is close to 0, indicating that they are unlikely to compete with each other. Note that the standard deviation of antibody G's column is 0, and the correlation of antibody G with all other antibodies is undefined since the standard deviation of each column appears in the denominator of the correlation coefficient.
[0074] Figure 6 shows a flow diagram for deriving hidden variables based on antibody competition data for discovering binding patterns according to an exemplary embodiment. In one embodiment, the functions of Figure 6 are implemented by software stored in memory or other computer-readable or tangible medium and executed by a processor. In other embodiments, the respective functions may be performed by hardware (e.g., by using application specific integrated circuits ("ASICs"), programmable gate arrays ("PGAs"), field programmable gate arrays ("FPGAs"), etc.), or by any combination of hardware and software.
[0075] At 602, antibody competition data can be received for a plurality of antibodies and antigens, the antibody competition data including data values indicative of pair-wise competition between antibodies. For example, an experimental run (e.g., data generated from a surface plasmon resonance ("SPR") experimental technique) can generate combinatorial (e.g., pair-wise) competition data for a set of antibodies.
[0076] In some embodiments, the received antibody competition data includes data from multiple experimental runs, each experimental run generating data values indicative of pair-wise competition between a set of antibodies, and the multiple experimental runs generating antibody competition data for different sets of antibodies.
[0077] At 604, the antibody competition data can be processed to generate training data. For example, processing the competition data can include data transformation, such as a mathematical conversion to a binary representation of the competition. In some embodiments, processing the antibody competition data includes combining antibody competition data from multiple experimental runs.
[0078] At 606, the training data and the optimization engine can be used to derive a plurality of hidden variables and affinity scores for the hidden variables, where an affinity score for the hidden variables is derived for each antibody, where the hidden variables represent a competitive factor for the antigen that causes competition between the antibodies. For example, a first hidden variable may represent a first competitive factor for the antigen, and a derived affinity score for the first hidden variable associated with a given antibody indicates a degree of competition of the given antibody with respect to the first competitive factor. In some embodiments, the first competitive factor corresponds to an epitope of the antigen that causes competition between the antibodies. In some embodiments, deriving the plurality of hidden variables and affinity scores for the hidden variables includes deriving affinity scores for the antibodies from different sets of antibodies (e.g., different sets of antibodies associated with different experimental runs).
[0079] In some embodiments, the hidden variables are derived by optimizing hidden logit values for the antibodies using pairwise competition data values from the training data, hidden logit values that represent the affinity scores of the antibodies with respect to the hidden variables. For example, the hidden logit values of the antibodies may be optimized using a loss function, pairwise competition data values from the training data, and a gradient technique that adjusts the hidden logit values to optimize the loss function.
[0080] In some embodiments, the hidden variables and affinity scores for the hidden variables are derived by first optimizing a hidden logit value of the antibody for a first hidden variable, and successively adding additional hidden variables after the initial optimization of the first hidden variable, and jointly optimizing the hidden logit value of the antibody for the first hidden variable and each successively added additional hidden variable.
[0081] In some embodiments, pairwise competition score predictions for two antibodies can be generated using the hidden logit values optimized for the two antibodies. For example, the received antibody competition data (e.g., processed to generate training) may not include pairwise competition data for the two antibodies. In some embodiments, pairwise competition score predictions are generated in part by performing a dot product operation on the hidden logit values for the two antibodies.
[0082] In some embodiments, the derived latent variables and affinity scores for the latent variables can provide a model for predicting competition between antibodies without requiring an explicit experimental run to observe the competition. In other words, the derived model can be used to predict competition for antibody pairs that have not been experimentally tested and observed. Embodiments of the derived model can be considered as a predictive or simulation tool for antibody competition. Thus, embodiments improve antibody competition testing by improving resource and time efficiency.
[0083] In some embodiments, the derived model can also predict high fidelity competitive dynamics between antibodies. For example, antibodies exhibiting different latent variable affinity scores may exhibit different binding mechanisms, while antibodies exhibiting similar latent variable affinity scores may exhibit similar binding mechanisms. The latent variable affinity score descriptors provide a more sophisticated view of competitive dynamics when compared to previous approaches that simply associate antibodies with individual bins. Embodiments of the derived model allow for the selection of antibodies exhibiting diverse latent variable affinities to achieve a more robust monoclonal antibody discovery and manufacturing process.
[0084] The features, structures, or characteristics of the present disclosure described throughout this specification can be combined in any suitable manner in one or more embodiments. For example, the use of "one embodiment," "some embodiments," "certain embodiments," "certain embodiments," or other similar terms throughout this specification refers to the fact that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the present disclosure. Thus, the appearance of the phrases "one embodiment," "some embodiments," "certain embodiments," "certain embodiments," or other similar terms throughout this specification do not necessarily all refer to the same group of embodiments, but rather the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0085] Those skilled in the art will readily appreciate that the embodiments discussed above can be implemented with steps in a different order and / or elements in a different configuration than those disclosed. Thus, while the present disclosure contemplates the outlined embodiments, it will be apparent to those skilled in the art that certain modifications, variations, and alternative constructions will become apparent while remaining within the spirit and scope of the present disclosure. Accordingly, reference should be made to the appended claims to determine the metes and bounds of the present disclosure.
Claims
**Claim 1** A method for deriving hidden variables based on antibody competition data for discovering binding patterns, comprising: Receiving antibody competition data regarding a plurality of antibodies and antigens, including data values indicating pairwise competition between antibodies; Processing the antibody competition data to generate training data; Using the training data and an optimization engine to derive a plurality of hidden variables and affinity scores regarding the hidden variables, wherein the affinity scores regarding the hidden variables are derived for each antibody, and the hidden variables represent competitive factors regarding the antigens that cause competition between the antibodies; and Generating a pairwise competition score prediction regarding the two antibodies using the affinity scores regarding the two antibodies A method comprising the above steps. **Claim 2** The method according to claim 1, wherein a first hidden variable represents a first competitive factor regarding the antigen, and the derived affinity score regarding the first hidden variable associated with a given antibody indicates the degree of competition of the given antibody regarding the first competitive factor. **Claim 3** The method according to claim 2, wherein the first competitive factor corresponds to an epitope of the antigen that causes competition between the antibodies. **Claim 4** The method according to claim 2, wherein the received antibody competition data includes data derived from a plurality of experimental runs, each experimental run generating data values indicating pairwise competition between sets of antibodies, and the plurality of experimental runs generating antibody competition data regarding different sets of antibodies. **Claim 5** The method according to claim 4, wherein processing the antibody competition data includes combining the antibody competition data derived from the plurality of experimental runs. **Claim 6** The method according to claim 5, wherein deriving the plurality of hidden variables and the affinity scores regarding the hidden variables includes deriving affinity scores regarding the antibodies from different sets of the antibodies. **Claim 7** The method according to claim 1, wherein the hidden variables are derived by optimizing the hidden logit values regarding the antibodies using the pairwise competition data values derived from the training data and the hidden logit values representing the affinity scores of the antibodies regarding the hidden variables. **Claim 8** The method according to claim 7, wherein the hidden logit value of the antibody is optimized using a loss function, the pairwise competition data value derived from the training data, and a gradient technique for adjusting the hidden logit value to optimize the loss function.
9. The hidden variable and the affinity score for the hidden variable are first optimizing the hidden logit value of the antibody for a first hidden variable; and subsequently adding additional hidden variables successively after the first optimization of the first hidden variable, and jointly optimizing the hidden logit value of the antibody for the first hidden variable and each successively added additional hidden variable The method according to claim 8, derived by
10. The method according to claim 1, wherein the received antibody competition data does not include pairwise competition data for the two antibodies.
11. A system for deriving hidden variables based on antibody competition data for discovering binding patterns, a processor; and a memory storing instructions for execution by the processor, the instructions causing the processor to receive antibody competition data for a plurality of antibodies and antigens, including data values indicating pairwise competition between antibodies; process the antibody competition data to generate training data; derive a plurality of hidden variables and an affinity score for the hidden variables using the training data and an optimization engine, the affinity score for the hidden variables being derived for each antibody, the hidden variables representing competitive factors for the antigen causing competition between the antibodies; and generate a pairwise competition score prediction for the two antibodies using the affinity scores for the two antibodies. A system including a memory.
12. The system according to claim 11, wherein a first hidden variable represents a first competitive factor for the antigen, and the derived affinity score for the first hidden variable associated with a given antibody indicates the degree of competition of the given antibody for the first competitive factor.
13. The system according to claim 12, wherein the first competitive factor corresponds to an epitope of the antigen causing competition between the antibodies.
14. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to derive hidden variables based on antibody competition data for discovering binding patterns, and that, when executed, cause the instructions to cause the processor to receive antibody competition data regarding a plurality of antibodies and antigens, including data values indicating pairwise competition between antibodies; process the antibody competition data to generate training data; use the training data and an optimization engine to derive a plurality of hidden variables and affinity scores regarding the hidden variables, wherein the affinity scores regarding the hidden variables are derived for respective antibodies, and the hidden variables represent competing factors regarding the antigens that cause competition between the antibodies; and generate a pairwise competition score prediction regarding the two antibodies using the affinity scores regarding the two antibodies, a non-transitory computer-readable medium.