Designing biomolecular sequence variants with pre-specified attributes
Patent Information
- Application Number
- JP2024541017
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-14
- Filing Date
- 2023-01-05
- Publication Date
- 2026-01-14
AI Technical Summary
Traditional antibody screening methods can only explore small sequence spaces, resulting in insufficient antibody binding affinity, development potential and immunogenicity-profiles, and deep mutation methods have problems such as inefficiency and difficulty in data processing.
Combining high-throughput experimental biology and machine learning, by training machine learning models, predicting the binding affinity and nature of antibody sequence variants, multiple attributes are optimized to improve the development potential and safety of antibodies.
It has achieved efficient exploration of large-scale antibody sequence space, improved the success rate and development efficiency of antibody screening, and reduced experimental costs and time.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Background technology]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 398,222, filed August 15, 2022. This application claims the benefit of U.S. Provisional Patent Application No. 63 / 339,450, filed May 7, 2022. This application claims the benefit of U.S. Provisional Patent Application No. 63 / 338,398, filed May 4, 2022. This application claims the benefit of U.S. Provisional Patent Application No. 63 / 338,433, filed May 4, 2022. This application claims the benefit of U.S. Provisional Patent Application No. 63 / 320,067, filed March 15, 2022. This application claims the benefit of U.S. Provisional Patent Application No. 63 / 297,679, filed January 7, 2022. This application claims the benefit of U.S. Patent Application No. 18 / 046,849, filed October 14, 2022. Each of the priority applications is incorporated herein by reference in its entirety.
[0002] Incorporation by Reference of Electronically Submitted Material The Sequence Listing is part of this disclosure and is submitted herewith as an XML file. The XML containing the Sequence Listing is named "457548_Seqlisting.xml", was created on January 4, 2023, and is 68,540 bytes in size. The subject matter of the Sequence Listing is incorporated herein by reference in its entirety.
[0003] Monoclonal antibodies can exhibit suboptimal binding affinity to their target antigens. Affinity can be improved by "maturation" of the antibody sequence, often through combinatorial mutagenesis of CDRs. However, the combinatorial mutation space is so large that it would require a significant amount of time and resources to thoroughly explore it experimentally. Therefore, wet lab screening solutions (enhanced antibody affinity) can be inefficient and time-consuming.
[0004] Biological drug discovery is a complex combinatorial challenge. The number of possible monoclonal antibody CDR variants exceeds the number of atoms in the universe. Traditional antibody screening approaches explore a small sequence space (e.g., hundreds to thousands of variants) and often result in drug candidates with poor binding affinity, developability concerns, and poor immunogenicity profiles. Biological drug discovery too often fails. Specifically, despite billions of dollars invested annually, only an estimated 4% of drug leads are successful in their journey from development to market. Even worse, only 18% of drug leads that pass preclinical trials ultimately pass Phase I and II clinical trials, suggesting that the majority of drug candidates are risky or ineffective. While much of this failure rate is due to an incomplete understanding of the underlying biology and pathology, inadequate drug lead optimization contributes to numerous failures.
[0005] Traditional antibody screening approaches can only explore a small sequence space, which can limit results to sequences that confer suboptimal properties, such as poor binding affinity, limited developability, and poor immunogenicity profiles. In contrast, deep mutagenesis coupled with screening or selection allows for the exploration of a larger antibody sequence space, potentially yielding more and better drug leads. However, deep mutagenesis comes with its own challenges. For example, most mutations degrade the binding affinity of a given antibody rather than improving it, thereby reducing screening efficiency. Furthermore, the combinatorial space of antibody sequence variants grows exponentially with the mutational burden (i.e., the number of mutations simultaneously introduced into each sequence variant), rapidly exceeding the capabilities of experimental assays by orders of magnitude.
[0006] Furthermore, the development of candidate biomolecules (e.g., antibodies) into therapeutic drugs is a complex process that involves a high degree of risk. This risk often stems from numerous challenges in production, formulation, efficacy, and adverse reactions. Modeling these risks has been a tremendous challenge for the industry due to the difficulty in obtaining the relevant data, especially at large scale.
[0007] Furthermore, most antibody screening approaches are limited to screening only one characteristic at a time, limiting the simultaneous optimization of drug efficacy and developability. Simultaneous, rather than sequential, antibody optimization is a more advantageous therapeutic strategy because improving a single characteristic can negatively impact other characteristics. A strategy that simultaneously co-optimizes all characteristics can yield better therapeutic products. Deep neural networks are an emerging tool for overcoming the limitations of experimental screening capabilities. A common approach involves training a model on a small amount of experimental data and applying it to predict which sequences are most likely to improve the measured trait. Several promising approaches have been proposed, but only a limited number of in silico predictions from these models have been validated in the laboratory. While sufficient as proof-of-principle, such demonstrations are limited for practical design by the drawbacks of the screening platforms used to generate the training data: binary readouts (rather than sequential) with limited throughput. Overall, this limits the quantitative accuracy of models and their ability to extrapolate to higher mutational burdens.
[0008] An additional concern with antibody screening approaches is that improved binding affinity may come at the expense of reduced developability and immunogenic properties, an issue that will remain unaddressed by trained machine learning models regardless of other properties.
[0009] Thus, there is a need for technologies that couple high-throughput experimental biology with machine learning to improve upon traditional biomolecular screening approaches, thereby enabling greater exploration of the correct sequences. Summary of the Invention
[0010] In one aspect, a computing system for identifying biomolecular sequence variants of interest is provided, the computing system comprising: (a) one or more processors; and (b) one or more non-transitory computer-readable media having stored thereon: (i) a machine learning model trained using training data, wherein the training data includes one or more training biomolecular sequence variants, each having a respective measured binding property representing the respective ability of the variant to bind to a corresponding respective binding partner, and wherein the machine learning model is configured to output predicted biomolecular binding properties of the input biomolecular sequence variants. and (ii) instructions that, when executed by one or more processors, cause a computing system to: (1) process one or more biomolecular sequence variants with the machine learning model to generate one or more predicted binding properties, each corresponding to a respective one of the one or more biomolecular sequence variants; (2) analyze the one or more predicted binding properties to identify one or more biomolecular sequence variants of interest from among the sequence variants, each of the one or more biomolecular sequence variants of interest having one or more desired properties; and (3) provide the one or more biomolecular sequence variants of interest as output.
[0011] In another aspect, a computer-implemented method for training a machine learning model to identify biomolecular sequence variants of interest includes: (1) generating one or more biomolecular sequence variants by programmatically mutating a reference biomolecule; (2) receiving screening data that includes ranking the biomolecular sequence variants according to one or more training binding properties; and (3) training a machine learning model using the received screening data to predict one or more desired binding properties of the input biomolecular sequence variants.
[0012] In another aspect, a computing system for predicting naturalness of biomolecular sequence variants, the computing system including one or more processors; and one or more non-transitory computer-readable media having stored thereon a machine learning model trained using training data, the training data including one or more training biomolecular sequence variants, the machine learning model configured to output predicted naturalness characteristics for each of the one or more biomolecular sequence variants; and instructions that, when executed by the one or more processors, cause the computing system to: (i) process one or more input biomolecular sequence variants with the machine learning model to generate a respective predicted naturalness characteristic for each of the one or more input biomolecular sequence variants; and (ii) provide at least one of the predicted naturalness characteristics as an output.
[0013] In yet another aspect, a computing system for predicting naturalness of biomolecular sequence variants, the computing system including one or more processors; and one or more non-transitory computer-readable media having stored thereon a machine learning model trained using training data, the training data including one or more training biomolecular sequence variants, the machine learning model configured to output predicted naturalness characteristics for each of the one or more biomolecular sequence variants; and instructions, when executed by the one or more processors, that cause the computing system to: (1) process one or more input biomolecular sequence variants with the machine learning model to generate a respective predicted naturalness characteristic for each of the one or more input biomolecular sequence variants; and (2) provide at least one of the one or more predicted naturalness characteristics as an output.
[0014] In some embodiments, the present disclosure provides a method for generating training data. In one embodiment, a method for generating training data for a machine learning model is provided, comprising: a) expressing a biomolecule variant library in host cells; b) measuring: (i) the expression levels and (ii) affinity values for a binding partner of interest of two or more biomolecule variants expressed in (b); c) sorting the host cells into a distribution of cell subpopulations based on the measured expression levels and the measured affinity values, thereby collecting cells across the affinity distribution; d) sequencing the expressed biomolecule variants from the collected cells of (c); e) calculating an enrichment score for each sequenced biomolecule variant, the enrichment score and the biomolecule variant sequence being capable of training a machine learning model capable of performing sequence-based affinity prediction.
[0015] In another embodiment, a method as described above is provided, wherein the library of biomolecule variants is generated by randomly mutating a nucleic acid encoding a reference biomolecule. In another embodiment, a method as described above is provided, wherein the library of biomolecule variants is generated by random mutagenesis, error-prone PCR mutagenesis, oligonucleotide-directed mutagenesis, cassette mutagenesis, shuffling, saturation mutagenesis, homology-directed mutagenesis, activation-induced cytidine deaminase (AID)-mediated mutagenesis, or transposon mutagenesis. In yet another embodiment, a method as described above is provided, wherein the library of biomolecule variants comprises at least 10 to 10 unique biomolecule variant sequences. In yet another embodiment, a method as described above is provided, wherein the library of biomolecule variants is displayed on the host cell surface. In another embodiment, a method as described above is provided, wherein the library of biomolecule variants is expressed and maintained in the host cytoplasm.
[0016] In another embodiment, the aforementioned method is provided wherein the host cell is an E. coli cell. In yet another embodiment, the aforementioned method is provided wherein the E. coli cell is an E. coli 521 cell. In another embodiment, the aforementioned method is provided wherein the E. coli cell comprises one or more or all of the following: a) modified gene function of at least one gene encoding a transporter protein for an inducer of the at least one inducible promoter; b) a reduced level of gene function of at least one gene encoding a protein that metabolizes the inducer of the at least one inducible promoter; c) a reduced level of gene function of at least one gene encoding a protein involved in the biosynthesis of the inducer of the at least one inducible promoter; d) modified gene function of a gene that affects the reducing / oxidizing environment of the cytoplasm of the host cell; e) a reduced level of gene function of a gene encoding a reductase; f) at least one expression construct encoding at least one disulfide bond isomerase protein; g) at least one polynucleotide encoding a form of DsbC lacking a signal peptide; and / or h) at least one polynucleotide encoding Ervlp.
[0017] In yet another embodiment, the aforementioned method is provided, wherein step (c) optionally measures one or more of the binding specificity, biological activity, stability, and / or solubility of the expressed biomolecule variant.
[0018] In yet another embodiment, the aforementioned method is provided, wherein affinity is quantified by measuring the binding dissociation constant (KD) of the biomolecule variant for a binding partner of interest. In one embodiment, the binding partner of interest is a fluorescently labeled antigen.
[0019] In yet another embodiment, the aforementioned method is provided, wherein the expression level of the biomolecule variant is quantified by measuring anti-IgG binding capacity. In another embodiment, the aforementioned method is provided, wherein the expression level of the biomolecule variant is quantified using an anti-IgG antibody conjugated to a fluorophore. In yet another embodiment, the aforementioned method is provided, wherein the expression level of the biomolecule variant is quantified by measuring non-antigen binding capacity.
[0020] The present disclosure also provides, in some embodiments, the aforementioned method, wherein the measuring in step (c) and the sorting in step (d) comprise a fluorescence-activated cell sorting (FACS) assay. In another embodiment, the aforementioned method is provided, optionally further comprising measuring the binding affinity of the sequenced biomolecule variants prior to calculating the enrichment score. In one embodiment, the binding affinity is measured using an assay selected from the group consisting of a surface plasmon resonance (SPR)-based binding assay, biolayer interferometry, and / or a flow cytometry-derived binding curve.
[0021] In another embodiment, the aforementioned method is provided, wherein the sequencing in step (e) is obtained by a method selected from the group consisting of deep sequencing, next generation sequencing, long read nanopore sequencing, and single molecule real time long read sequencing (pacbio). In another embodiment, the sequencing in step (e) is obtained by a method selected from the group consisting of deep sequencing, next generation sequencing, long read nanopore sequencing, and single molecule real time long read sequencing (pacbio). In another embodiment, the aforementioned method is provided, wherein the nucleic acid encoding the biomolecule variant is modified prior to sequencing to include a barcode sequence comprising a unique molecular identifier (UMI).
[0022] The present disclosure also provides, in one embodiment, the aforementioned method, wherein the biomolecule variant is selected from the group consisting of a monoclonal antibody, a bispecific antibody, a multispecific antibody, a humanized antibody, a chimeric antibody, a camelized antibody, a single domain antibody, a single-chain Fv (ScFv), a single-chain antibody, a Fab fragment, a F(ab') fragment, a disulfide-linked Fv (sdFv), or an anti-idiotypic (anti-Id) antibody. In another embodiment, the aforementioned method is provided, wherein the biomolecule variant is selected from the group consisting of a monoclonal antibody, a bispecific antibody, a multispecific antibody, a humanized antibody, a chimeric antibody, a camelized antibody, a single domain antibody, a single-chain Fv (ScFv), a single-chain antibody, a Fab fragment, a F(ab') fragment, a disulfide-linked Fv (sdFv), or an anti-idiotypic (anti-Id) antibody. In yet another embodiment, the aforementioned method is provided, wherein the biomolecule variant is selected from the group consisting of peptides, polypeptides, proteases, oxidoreductases, transferases, hydrolases, lyases, isomerases, ligases, enzymes, antibodies, cytokines, chemokines, nucleic acids, metabolites, small molecules (<1 kDa), and synthetic molecules.
[0023] In yet another embodiment, a method for generating training data for a machine learning model is provided, comprising: a) expressing a biomolecular variant library in host cells; b) measuring: (i) the expression levels and (ii) affinity values for a binding partner of interest of two or more biomolecular variants expressed in (b); c) sorting the host cells into a distribution of cell subpopulations based on the measured expression levels and the measured affinity values, thereby collecting cells across the affinity distribution; d) isolating nucleic acids encoding the biomolecular variants from the collected host cells of (c), amplifying the nucleic acids using selective rolling circle amplification (sRCA), and sequencing the nucleic acids encoding the biomolecular variants; and e) calculating an enrichment score for each sequenced biomolecular variant, wherein the enrichment scores and the biomolecular variant sequences can be used to train a machine learning model capable of performing sequence-based affinity prediction. [Brief explanation of the drawings]
[0024] [Figure 1] FIG. 1 illustrates an exemplary computing environment for implementing the present technology, according to some aspects.
[0025] [Figure 2] FIG. 2 illustrates an exemplary computer-implemented method for training one or more machine learning models to identify one or more biomolecular sequence variants of interest, according to some embodiments.
[0026] [Figure 3A] FIG. 3A depicts a computer-implemented method of operating a trained machine learning model to identify one or more biomolecular sequence variants of interest, according to some embodiments.
[0027] [Figure 3B] FIG. 3B depicts a computer-implemented method for training a machine learning model to identify biomolecular sequence variants of interest, according to some embodiments.
[0028] [Figure 4A] FIG. 4A shows an exemplary data flow diagram depicting training and prediction of biomolecular sequence variants of interest, which may correspond to FIG. 2, according to some embodiments.
[0029] [Figure 4B] 4B depicts an exemplary block flow diagram for performing the assay of FIG. 4A, according to some embodiments. Strains expressing unique antibody sequence variants may be added to fixed and permeabilized cells, and a probe may be added.
[0030] [Figure 4C] FIG. 4C depicts an example of an affinity prediction chart, according to some embodiments.
[0031] [Figure 4D] FIG. 4D depicts an exemplary affinity prediction validation chart, according to some embodiments.
[0032] [Figure 4E-1] FIG. 4E depicts an exemplary denoised data chart, according to some embodiments. [Figure 4E-2] Same as above.
[0033] [Figure 4F] FIG. 4F depicts an exemplary conceptual diagram illustrating naturalness training and prediction, according to some embodiments.
[0034] [Figure 4G] FIG. 4G depicts an exemplary naturalness score validation chart, according to some embodiments.
[0035] [Figure 4H] FIG. 4H depicts an exemplary naturalness / developability correlation chart, according to some embodiments.
[0036] [Figure 4I] FIG. 4I depicts an exemplary naturalness and immunogenicity correlation chart, according to some embodiments.
[0037] [Figure 4J] FIG. 4J depicts an exemplary naturalness and mutation load correlation chart, according to some embodiments.
[0038] [Figure 4K] FIG. 4K depicts an exemplary chart of affinity prediction improvement when enriching with naturalness data, according to some embodiments.
[0039] [Figure 4L] FIG. 4L depicts an exemplary conceptual diagram of in silico sequence variant generation and optimization, according to some embodiments.
[0040] [Figure 4M]FIG. 4M depicts an exemplary chart of affinity predictions from trastuzumab, according to some embodiments.
[0041] [Figure 4N-1] FIG. 4N depicts an exemplary affinity prediction chart from different parent antibodies, according to some embodiments. [Figure 4N-2] Same as above. [Figure 4N-3] Same as above. [Figure 4N-4] Same as above. [Figure 4N-5] Same as above.
[0042] [Figure 4O-1] FIG. 4O depicts an exemplary visualization of optimization for affinity and naturalness, according to some embodiments. [Figure 4O-2] Same as above.
[0043] [Figure 4P] FIG. 4P shows another exemplary affinity prediction chart including a comparison of binding affinity measurements, according to some embodiments.
[0044] [Figure 4Q] FIG. 4Q depicts an exemplary affinity prediction chart including a comparison of binding affinity measurements generated by SPR and model prediction trained on SPR data, while retaining all points with higher affinity than wild-type trastuzumab, according to some embodiments.
[0045] [Figure 5A] FIG. 5 depicts an example of an AI-enhanced antibody optimization diagram, according to some embodiments. [Figure 5B] Same as above.
[0046] [Figure 6A] FIG. 6A depicts an example of an exemplary workflow proof-of-concept diagram, according to some embodiments.
[0047] [Figure 6B]FIG. 6B depicts the predictive performance of a model trained on qaACE scores of variants from 90% of trust-1, evaluated on a 10% holdout dataset, according to some embodiments.
[0048] [Figure 6C] FIG. 6C depicts a comparative analysis of replicated qaACE measurements and qaACE scores predicted from models trained on individual qaACE replicates, according to some embodiments.
[0049] [Figure 6D] FIG. 6D depicts a comparison of ACE scores measured by two replicate FACS sorts, according to some embodiments.
[0050] [Figure 6E-1] FIG. 6E depicts an all-to-all comparison of ACE scores measured by one of two replicate FACS sorts against ACE scores predicted by a model trained on data from only one of the two replicates, according to some embodiments. [Figure 6E-2] Same as above.
[0051] [Figure 6F] FIG. 6F depicts the correlation between qaACE affinity scores and log-transformed SPR KD measurements, according to some embodiments.
[0052] [Figure 6G] FIG. 6G depicts prediction performance for a holdout set that is uniformly distributed with respect to binding affinity, according to some embodiments.
[0053] [Figure 7A] FIG. 7A depicts predictions from a model trained on SPR-measured log10 KD values, according to some embodiments.
[0054]
[0055] [Figure 7B]FIG. 7B depicts a comparative analysis of replicate-log10 KD measurements and-log10 KD predicted from models trained on individual SPR replicates, according to some embodiments.
[0056] [Figure 7C] FIG. 7C depicts predictions from a model trained on log10 k on values, according to some embodiments.
[0057] [Figure 7D] FIG. 7D depicts predictions from a model trained with a −log10 koff value, according to some embodiments.
[0058] [Figure 7E] FIG. 7E depicts a comparison of −log 10 KD values measured by two SPR experiments, according to some embodiments.
[0059] [Figure 7F-1] FIG. 7F depicts an all-to-all comparison of -log KD values measured by one of two replicate SPR experiments to -log KD values predicted by a model trained on data from only one of the two replicates, according to some embodiments. [Figure 7F-2] Same as above.
[0060] [Figure 7G] FIG. 7G depicts prediction of -log10KD using SPR training data alone or supplemented with ACE measurements, according to some embodiments.
[0061] [Figure 7H] FIG. 7H depicts a comparison of log 10 k on values measured by two SPR experiments, according to some embodiments.
[0062] [Figure 7I-1]FIG. 7I depicts a comparative analysis of replicate log10kon measurements and log10kon predicted from models trained on individual SPR replicates, according to some embodiments. [Figure 7I-2] Same as above.
[0063] [Figure 7J] Figure 7J depicts an all-to-all comparison of log10kon values measured from one of two replicate SPR experiments to log10kon values predicted by a model trained on data from only one of the two replicates, according to some embodiments.
[0064] [Figure 7K] FIG. 7K depicts a comparison of −log 10 koff values measured by two SPR experiments, according to some embodiments.
[0065] [Figure 7L-1] FIG. 7L depicts a comparative analysis of replicate -log10 koff measurements and -log10 koff predicted from models trained on individual SPR replicates, according to some embodiments. [Figure 7L-2] Same as above.
[0066] [Figure 7M] FIG. 7M depicts an all-to-all comparison of −log koff values measured from one of two replicate SPR experiments against −log koff values predicted by a model trained on data from only one of the two replicates, according to some embodiments.
[0067] [Figure 7N] FIG. 7N depicts a 90:10 training:holdout split of ACE scores from the trust-1 dataset, according to some embodiments.
[0068] [Figure 7O] FIG. 7O depicts a 10-fold cross-validation using −log10 KD values from the trast-2 dataset, according to some embodiments.
[0069] [Figure 7P] FIG. 7P depicts a scatter plot of random model embedding versus binding affinity, without pre-training and without fine-tuning, according to some embodiments.
[0070] [Figure 7Q] FIG. 7Q depicts a scatter plot of model embedding versus binding affinity without pre-training and fine-tuning with binding affinity data using the trast-2 dataset, according to some embodiments.
[0071] [Figure 7R] FIG. 7R depicts a scatter plot of model embedding versus binding affinity without pre-training and fine-tuning using OAS-derived sequences, according to some embodiments.
[0072] [Figure 7S] FIG. 7S depicts a scatter plot of model embedding versus binding affinity, using pre-training with OAS-derived sequences and fine-tuning with binding affinity data using the trast-2 dataset, according to some embodiments.
[0073] [Figure 8A] FIG. 8A depicts a density plot of predicted (design) and measured (validation) binding affinities of 50 sequences designed to span approximately two orders of magnitude of KD (Set A), according to some embodiments.
[0074] [Figure 8B] FIG. 8B depicts a density plot of predicted (design) and measured (validation) binding affinities of 50 sequences designed to bind HER2 more tightly than parent trastuzumab (Set B), according to some embodiments.
[0075] [Figure 8C]Figure 8C depicts the empirical distribution function (ECDF) of the measured (validation) binding affinities of 50 sequences from Design Set B, with the line indicating the measured -log10 KD (or deviation by -0.1 or -0.5 log) of trastuzumab, according to some embodiments.
[0076] [Figure 8D] Figure 8D depicts a density plot of binding affinities from Set B as in Figure 8B, predicted by a model trained on the full trast-2 dataset (design, original prediction) or re-predicted by a model trained on a version of the trast-2 dataset in which any variant binding was depleted more strongly than parent trastuzumab (training, KD cap) (design, prediction with KD cap training), according to some embodiments.
[0077] [Figure 8E] FIG. 8E depicts a scatter plot of predicted (design) and measured (validation)-log10 KD values, data referring to Design Set A of FIG. 8A, according to some embodiments.
[0078] [Figure 8F] FIG. 8F depicts a scatter plot of measured (validation)-log10 KD values in individual SPR replicates, data referring to Design Set A of FIG. 8A, according to some embodiments.
[0079] [Figure 8G] FIG. 8G depicts a scatter plot of predicted (design) and measured (validation)-log10 KD values, data referring to Design Set B in FIGS. 8B-8D, according to some embodiments.
[0080] [Figure 8H] FIG. 8H depicts a scatter plot of measured (validation)-log10 KD values in individual SPR replicates, data referring to Design Set B of FIGS. 8B-8D, according to some embodiments.
[0081] [Figure 8I]FIG. 8I depicts a chart of model predictions for variants with desirable binding properties for naive library screening, according to some embodiments.
[0082] [Figure 9A] FIG. 9A depicts an illustration of a combinatorial mutagenesis strategy for the trast-3 dataset, up to a triple mutant at position 20 of trastuzumab (10 for CDRH2 and 10 for CDRH3) screened using ACE, according to some embodiments.
[0083] [Figure 9B] FIG. 9B depicts the predictive performance of a model trained on the trust-3 dataset using 20% of the data in the holdout set, according to some embodiments.
[0084] [Figure 9C] FIG. 9C depicts validation of models trained on up to triple mutants against a holdout set of up to triple mutants, as well as against a holdout set of quadruple and quintuple mutants, extrapolating predictions to higher mutation burdens than seen in the training set, according to some embodiments.
[0085] [Figure 9D] FIG. 9D depicts a line plot showing the accuracy of models on a common holdout validation set across different training set sizes, according to some embodiments, in which: (i) the shaded area indicates the standard deviation across folds; (ii) for each training subset size, the performance of each of the OAS pre-trained model and the randomly initialized model is shown, each trained using a subset of the high-fidelity trust-3 dataset or a low-fidelity version of the dataset; and (iii) included under each subset size is an indication of the proportion of training data used, the size of the training dataset, and the percent of the complete mutation space covered by the training subset.
[0086] [Figure 9E] FIG. 9E depicts the performance of modeling with randomized ACE scores (trast-3 dataset), according to some embodiments.
[0087] [Figure 9F] FIG. 9F depicts an extrapolation of predictions (trast-3 dataset) to higher mutation burden for quadruple mutants, according to some embodiments.
[0088] [Figure 9G] FIG. 9G depicts an extrapolation of predictions (trast-3 dataset) to higher mutation burden for quintuple mutants, according to some embodiments.
[0089] [Figure 9H] FIG. 9H depicts a plot illustrating that the effect of individual mutations, according to some embodiments, can vary strongly in the presence of other mutations for a range (minimum to maximum) of incremental effects on predicted binding affinity from a model trained on the trast-3 dataset at each individual substitution across all possible single variants of trastuzumab.
[0090] [Figure 9I] FIG. 9I depicts a plot illustrating that the effect of individual mutations, according to some embodiments, can vary strongly in the presence of other mutations for a range (minimum to maximum) of incremental effects on predicted binding affinity from a model trained on the trast-3 dataset at each individual substitution across all possible double mutants of trastuzumab.
[0091] [Figure 9J-1] FIG. 9J depicts a sequence logo plot illustrating the composition of high affinity variants of trastuzumab, according to some embodiments. [Figure 9J-2] Same as above. [Figure 9J-3] Same as above. [Figure 9J-4] Same as above.
[0092] [Figure 9K] FIG. 9K depicts a heatmap illustrating the epistatic effect across all possible pairs of substitutions, according to some embodiments.
[0093] [Figure 10A] FIG. 10A shows predicted binding affinities for single mutants from a model trained on the trast-3 dataset, according to some embodiments, where (i) positions carrying mutations included CDRH2 (10 positions starting at R55) and CDRH3 (10 positions starting at W107); (ii) the reference trastuzumab sequence is highlighted with a cross; and (iii) mutations at each position include all possible substitutions with natural amino acids except cysteine, sorted alphabetically (i.e., X ∈ [A, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y ]).
[0094] [Figure 10B] FIG. 10B shows predicted binding affinities for double mutants from a model trained on the trast-3 dataset, according to some embodiments, in which (i) positions carrying mutations included CDRH2 (10 positions starting at R55) and CDRH3 (10 positions starting at W107); (ii) the reference trastuzumab sequence is highlighted with a cross; and (iii) mutations at each position include all possible substitutions with natural amino acids except cysteine, sorted alphabetically (i.e., X ∈ [A, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y ]).
[0095] [Figure 10C-1] FIG. 10C depicts the regression performance of a model trained on 10% of the CR9114 dataset, according to some embodiments. [Figure 10C-2] Same as above. [Figure 10C-3] Same as above. [Figure 10C-4] Same as above.
[0096] [Figure 10D-1] FIG. 10D depicts the regression performance of a model trained on 1% of the CR9114 dataset, according to some embodiments. [Figure 10D-2] Same as above. [Figure 10D-3] Same as above. [Figure 10D-4] Same as above.
[0097] [Figure 10E-1] FIG. 10E depicts the regression performance of a model trained on 0.1% of the CR9114 dataset, according to some embodiments. [Figure 10E-2] Same as above. [Figure 10E-3] Same as above. [Figure 10E-4] Same as above.
[0098] [Figure 10F-1] FIG. 10F depicts the regression performance of a mixture model trained with 10% of CR9114, according to some embodiments. [Figure 10F-2] Same as above.
[0099] [Figure 10G-1] FIG. 10G depicts the regression performance of a mixture model trained on 1% of CR9114, according to some embodiments. [Figure 10G-2] Same as above.
[0100] [Figure 10H-1] FIG. 10H depicts the regression performance of a mixture model trained with 0.1% of CR9114, according to some embodiments. [Figure 10H-2] Same as above.
[0101] [Figure 11A]FIG. 11A depicts a language model pre-trained with an antibody repertoire that can be utilized to calculate the naturalness of an antibody sequence conditioned on a given species, where the naturalness score was examined for association with four antibody properties, according to some embodiments.
[0102] [Figure 11B] FIG. 11B depicts immunogenicity using anti-drug antibody (ADA) responses to a humanized clinical-stage antibody reported by Marks et al.
[28] (n=97), according to some embodiments.
[0103] [Figure 11C] FIG. 11C depicts the failure of developability predicted by Therapeutic Antibody Profiler (TAP) for round 3 enriched phage display hits from the Gifford library
[29] (n=882), according to some embodiments.
[0104] [Figure 11D] FIG. 11D depicts expression levels in HEK-293 cells (mg / L) of clinical-stage humanized antibodies from Jain et al.
[30] (n=67), according to some embodiments.
[0105] [Figure 11E] FIG. 11E depicts a naturality density plot of 6,710,401 trastuzumab variants divided by mutational burden, with the dashed line corresponding to the naturality of the parent trastuzumab sequence, according to some embodiments.
[0106] [Figure 11F] FIG. 11F depicts the correlation between naturality and antibody immunogenicity, according to some embodiments.
[0107] [Figure 11G] FIG. 11G depicts the naturalness scores of the clinical-stage humanized antibodies used in the immunogenicity analysis of FIG. 11B, according to some embodiments.
[0108] [Figure 11H] FIG. 11H depicts the naturalness scores of round 3 enrichment stage display hits from the Gifford library (n=882) used in developability analysis with Therapeutic Antibody Profiler (TAP), as in FIG. 11C, according to some embodiments.
[0109] [Figure 11I] FIG. 11I depicts the naturalness scores of trastuzumab triple mutants from the combinatorial space sampled from the trast-3 dataset (n=6710401) used in developability analysis with TAP, as in FIG. 11H, according to some embodiments.
[0110] [Figure 11J] FIG. 11J depicts the naturalness scores of clinical-stage humanized antibodies used in the analysis of HEK-293 expression titers, as in FIG. 11D, according to some embodiments.
[0111] [Figure 11K] FIG. 11K depicts the naturalness scores corresponding to FIG. 11I and the TAP-predicted developability failure for the trastuzumab triple mutant (n=6710401) from the combinatorial space sampled from the trast-3 dataset, according to some embodiments, where P-values may be calculated using the Jonckheere-Terpstra trend test for binary data.
[0112] [Figure 11L] FIG. 11L depicts a density map of the complete Trap-3 search space, according to some embodiments.
[0113] [Figure 11M] FIG. 11M depicts a density map of variants with higher predicted ACE scores than trastuzumab, according to some embodiments.
[0114] [Figure 12A]FIG. 12A depicts a diagram in which each line tracks the average predicted qaACE score of the best 100 sequences observed across an evolutionary trajectory, with shaded areas indicating the standard deviation, according to some embodiments.
[0115] [Figure 12B] FIG. 12B depicts a plot of the average naturalness of the best 100 sequences observed across evolutionary trajectories, in which the shaded region indicates the standard deviation, according to some embodiments.
[0116] [Figure 12C] Figure 12C depicts a diagram of the qaACE scores and naturalness scores of the best 100 sequences determined through three search strategies: genetic algorithm, exhaustive search, and random search, according to some embodiments; in which the dashed line indicates the predicted score for trastuzumab; and the purple dashed line indicates the maximum score predicted across the entire combination space.
[0117] [Figure 12D] FIG. 12D depicts a histogram showing the first generation in which each of the top 100 sequences observed along the evolutionary trajectory was identified, according to some embodiments.
[0118] [Figure 13A] FIG. 13A depicts representative parent gating for all ACE sorts, according to some embodiments.
[0119] [Figure 13B] FIG. 13B depicts the specific expression and collection gating for each ACE library sort, according to some embodiments.
[0120] [Figure 14] FIG. 14 depicts a flowchart 1400 with the number of sequences filtered out and retained after each preprocessing step, according to some embodiments.
[0121] [Figure 15A]FIG. 15A depicts the output of a grid search over hyperparameter values performed on a pilot dataset, according to some aspects.
[0122] [Figure 15B] FIG. 15B depicts the output of a grid search over hyperparameter values performed on a subset of the pilot dataset of FIG. 15A, which includes 500 randomly selected sequences.
[0123] [Figure 16] FIG. 16 depicts a graph depicting the performance of models trained on ACE+SPR data with different ACE:SPR loss ratios, according to some embodiments.
[0124] [Figure 17-1] FIG. 17 depicts a graph of hyperparameter optimization for the XGBoost baseline on the pilot dataset, according to some embodiments. [Figure 17-2] Same as above. [Figure 17-3] Same as above. [Figure 17-4] Same as above.
[0125] [Figure 18A] FIG. 18A depicts a density plot of naturality distribution for different sequences, according to some embodiments.
[0126] [Figure 18B] FIG. 18B depicts a diagram of the relationships between sequence spaces, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0127] Detailed Description The present disclosure addresses the need for artificial intelligence (AI) and machine learning (ML) models that are trained using mappings between antibody sequence variants and experimental measurements (e.g., binding affinity, pH, and other data types). As described herein, once trained, the models are capable of predicting binding affinities of unseen sequence variants. This technology, combined with high-throughput and low-throughput binding affinity data, provides several orders of magnitude (e.g., four orders of magnitude) of K D The technology includes deep contextual language models that can predict the binding affinity of a wide range of unseen antibody sequence variants. This technology allows for measuring the "naturalness" of biomolecule (e.g., antibody) sequence variants, a broadly applicable metric shown herein to be relevant to downstream issues related to drug developability and immunogenicity. This technology may accelerate and improve biomolecule (e.g., antibody) engineering and increase the success rate of practical applications (e.g., antibody drug candidate development).
[0128] A major challenge to building accurate machine learning models is the lack of suitable large-scale training datasets. Directed evolution platforms are well suited for this because they rely on concatenating biological sequence data (DNA, RNA, protein) to phenotypic output. Indeed, it has been proposed to use ML models trained on data generated by mutagenesis libraries as a means to guide protein engineering. In recent years, access to deep sequencing and parallel computing has made it possible to build deep learning models capable of predicting molecular phenotypes from sequence data. Deep learning incorporates multiple hidden layers to decipher relationships embedded in large, high-dimensional datasets, such as millions of reads collected from a single deep sequencing experiment. Well-trained models can then be used to make predictions regarding completely unseen and novel variants. This application of model extrapolation is perfectly suited to protein engineering because it provides a way to explore a much larger sequence space than is physically possible. Here, we address this problem by combining deep mutational scanning and a bacterial display system to generate training datasets for ML models to learn sequence-function relationships.
[0129] An activity-specific cell enrichment (ACE) assay, which identifies host cells expressing an active gene product (e.g., as used herein, a biomolecule) of interest rather than inactive material, is described in WO 2021 / 146626, which is incorporated herein in relevant part. Active gene products can be distinguished from inactive material, for example, by the ability of the active gene product to specifically bind to a binding partner molecule or by the gene product's ability to participate in a chemical or enzymatic reaction. The presence of properly formed disulfide bonds in a polypeptide gene product indicates that it is correctly folded and presumably active. In this cell enrichment method, the active gene product of interest is detected, for example, by using an appropriate labeled conjugate that specifically binds to the active gene product of interest, such as a labeled antigen, if the gene product of interest is an antibody or Fab; by using a labeled ligand that specifically binds to the active conformation of the receptor, if the gene product of interest is a receptor or receptor fragment; or by using a labeled substrate or labeled substrate analog, if the gene product of interest is an enzyme. For any gene product of interest, if there is an antibody or antibody fragment available that specifically binds to the active gene product but not the inactive gene product, that antibody or antibody fragment, when attached to a detectable moiety, can be used to label the active gene product of interest.
[0130] A key strength of ACE is its ability to screen tens of thousands of "units of variation" in a single run. However, ongoing AI efforts applied to drug discovery add additional requirements to wet-lab-only screening, which imposes additional optimization of ACE to generate a suitable dataset for AI. Wet-lab-only screening aimed at selecting top-performing variants does not require strict quantitative accuracy from the assay. Indeed, due to the iterative nature of such screening, hits from step n-1 are rescreened in step n, effectively eliminating false positives from step n-1. Furthermore, wet-lab screening is often tuned to select only the desired population of interest (e.g., higher affinity variants), and as such, the assay may not be quantitative over a large dynamic range of the parameter of interest (e.g., antibody affinity). However, AI models for predicting quantitative predictions benefit from quantitative sequence variant training data. As such, quantitative sequence variant training data needs to be accurate for the model to produce meaningful predictions below the line. The present disclosure addresses these needs and shortcomings.
[0131] In various embodiments, the present disclosure provides an enhanced qaACE assay—quantitative affinity qaACE (qaACE)—as a method for sampling the affinity of antibody variants in a high-throughput manner using flow cytometry and next-generation sequencing to generate a qaACE score that correlates with KD. The primary goal of this method is to generate highly quantitative, high-throughput training data for AI models to perform sequence-based affinity prediction. This method can be applied to any antibody format, including mab, fab, scFv, scFAB, VHH, nanobody, and conceivably, other conjugated drug formats.
[0132] In one embodiment, the first step in the qaACE process is to generate a mutationally diverse antibody library, thereby evenly sampling sequence space around the starting antibody molecule, which contains variants that span a moderate range of mutational distances from the original sequence.
[0133] In some embodiments, including the examples herein, the method provides flow cytometry readouts from antibodies expressed in SoluPro E. coli that bind to fluorescently labeled antigen probes. In the qaACE assay, the set expression of antibody molecules is normalized so that variations in fluorescent signal in cells are attributable to the different affinities of antibody variants expressed in the cells that bind to the fluorescent antigen probe. This normalization is achieved through a common target molecule probe that binds to all variants and whose signal is in an orthogonal fluorescent channel to the antigen probe. In this setting, we show that the fluorescent signal of a variant is proportional to the measured KD of the antibody variant within a range. Assuming this proportionality, FACS can be used to sort cells containing antibody variants that span a range (e.g., distribution) of affinities.
[0134] After sorting across a range of affinity values using gating across the library population distribution, the cellular material is sequenced and quantified for the abundance of observed variants across affinity gates (bins, tubes). Using the quantification, an enrichment score is calculated for each variant. The enrichment scores generated via qaACE are an ideal data type for AI modeling purposes due to their accuracy and throughput.
[0135] In one exemplary workflow, the present disclosure provides a qaACE assay that includes some or all of the following general steps:
[0136] 1) Generation of antibody or other drug molecule libraries for screening via qaACE expressed in host cells such as SoluPro E. coli.
[0137] 2) Identification of fluorescently labeled antigen or binding partner probes for use in, for example, FACS through the early cytometry development process.
[0138] 3) Use of generic probes to target molecule variants that allow for detection of intracellular expression levels, which are used to gate homogeneous expression populations to resolve ambiguities in expression signals related to affinity and epitope binding signals.
[0139] 4) Sorting of cells across affinity distributions.
[0140] 5) Sequencing of cells sorted across the affinity distribution.
[0141] 6) During sequencing, DNA barcodes or UMIs can be added via PCR amplification of the region of interest. These UMIs allow absolute quantification of variants recovered from the gate.
[0142] 7) Generation of affinity correlated enrichment scores for each observed variant.
[0143] 8) AI model training using enrichment scores and antibody variant sequences.
[0144] As described herein, the present disclosure provides methods for generating highly quantitative, high-throughput training data for ML models to perform, for example, sequence-based affinity prediction. In some embodiments of the present disclosure, sequences of a highly diverse library of biomolecular variants, expressed in or on the surface of host cells, serve as input to an experiment (e.g., an assay to determine expression and / or affinity, among other readouts). In some embodiments, variants are sorted into multiple bins based on high-throughput measurements of binding affinity values (KD) normalized for variant expression levels, and variant sequences in each bin are obtained and tabulated by deep DNA sequencing. In some embodiments, the method then outputs multiple enrichment scores correlating the sequence information of all biomolecular variants in each bin with KDs across the complete experimental affinity distribution (i.e., from non-binders, low-binders, and high-binders). The enrichment scores generated via the qaACE assay of the present disclosure are an ideal data type for AI modeling purposes due to their accuracy and throughput. The combined method of obtaining affinity and sequence data for biomolecule variants is therefore referred to herein as the quantitative affinity activity specific cell enrichment (qaACE) assay.
[0145] As used herein, the term "quantitative affinity activity specific cell enrichment assay or qaACE assay" refers to a high-throughput assay for obtaining affinity and sequence data for biomolecular variants (U.S. Provisional Patent Application No. 63 / 371,474, filed August 15, 2022, which is incorporated by reference in its entirety).
[0146] As used herein, the term "affinity distribution" refers to the distribution of KD values for antigen binding to all possible sequence variants in a randomized library of biomolecule variants. Comparison with the KD value of a reference biomolecule provides an indication of whether the variant binds with higher or lower affinity.
[0147] This technology demonstrates the ability to improve the binding affinity of antibodies to their target antigens using deep contextual language models and quantitative, high-throughput experimental binding affinity data. We show that the model can quantitatively predict the binding affinity of unseen antibody variants with high accuracy, provide the ability to perform in silico drug screening, and ultimately enhance accessible sequence space by orders of magnitude. In this sense, the trained learner fulfills the role of a general surrogate for the black-box problem of assigning functional annotation from sequence alone. Novel variants with defined properties can be consistently designed by using the model as an oracle for various frameworks trained on protein fitness landscapes. We confirm the predictions and resulting designs in the lab with a success rate much higher than that achieved by traditional screening.
[0148] The present deep contextual language model includes a large language model (e.g., for antibody engineering using high-quality binding affinity measurements of trastuzumab sequence variants) that can predict binding affinities of unseen sequence variants spanning one or more orders of magnitude (e.g., four orders of magnitude) with high accuracy, resulting in the ability to perform drug screening entirely in silico. Here, by introducing natural antibody sequences into our language model, the present technology can characterize the "naturalness" of any given sequence for a host species. Empirical testing has shown that high naturalness scores are associated with improved metrics of immunogenicity and developability, thereby highlighting the importance of simultaneously optimizing multiple antibody properties during drug lead screening. To address this challenge, we present a genetic algorithm for the extremely efficient identification of sequences with both strong binding affinity and high naturalness.
[0149] These models may be used to identify variants with improved binding affinity, and many of these variants (e.g., 76%) may be confirmed to have higher binding affinity than wild-type trastuzumab based on rigorous wet-lab screening. These developments in combining deep contextual language models with in silico screening may enable significantly accelerated antibody lead optimization, improving the therapeutic potential of antibody candidates.
[0150] As will be appreciated by those skilled in the art, a much larger mutation space can be explored because in silico modeling is much faster and cheaper than wet-lab experiments. As provided herein, in some embodiments, the models can provide quantitative predictions of binding affinity (K), as opposed to the recently published state-of-the-art (Mason et al., Nat Biomedical Engineering, 2021, 5, 600-612), which can only provide qualitative predictions (binders vs. non-binders, i.e., in Mason, the models are rudimentary classifiers). D (i.e., the model is a regressor). In silico screening allows for the exploration of a much larger sequence space compared to using wet-lab methods. The technology can include wet-lab aspects (e.g., for model training), which greatly accelerates the generation of highly accurate training data. The AI-assisted workflow described herein thus provides higher yields and better results with less effort.
[0151] As will be appreciated by those skilled in the art, while the present disclosure provides various embodiments relating to antibody-antigen binding properties, the models and methods described herein can be applied to any biomolecule, including, for example, proteins, nucleic acids, receptors, ligands, and the like, in which the biomolecule is capable of binding to or otherwise interacting with a binding partner (which in some embodiments is the same or a different type of biomolecule). A "biomolecular sequence variant of interest" thus refers, in some embodiments, to a variation (e.g., mutation) in the sequence of a biomolecule (such as, for example, an antibody or antibody fragment), as described herein. The present techniques involve generating variant sequences in silico and then synthesizing those sequences in a laboratory setting.
[0152] Embodiments of the present disclosure provide compositions and methods for using the model to identify antibody sequences that confer higher binding affinity (e.g., to their antigen-binding partners). The model may be an artificial neural network. In one embodiment, the neural network architecture is a Transformer-Encoder model, as described in RoBERTa (Liu et al. 2019 arXiv) and "Attention is all you need" (Vaswani et al. 2017 arXiv). The RoBERTa architecture belongs to the "Transformer" family of neural networks and is primarily used in natural language processing. Those skilled in the art will understand that RoBERTa refers to the entire training setup, rather than solely to the model architecture (the architecture is a Transformer Encoder). The fine-grained architecture (e.g., number of hidden layers, embedding size, etc.) may be parameterized according to numerous alternative RoBERTa configurations, and experimental results show comparable performance. As such, the fine-grained architecture is likely not a major factor in performance. Alternative Transformer and non-Transformer deep neural network architectures may yield similar performance. Alternative non-deep neural network machine learning algorithms may yield similar or slightly inferior performance.
[0153] In some aspects, the core architecture of the model is the RoBERTa model (e.g., its PyTorch implementation within the Hugging Face framework (Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. Huggingface's transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019.) The trunk of the model may include a number of hidden layers (e.g., 16), with a number of attention heads per layer (e.g., 12), and a size of hidden layers (e.g., 768). Regression tasks may include one or more hidden layers of a given size (e.g., 768), followed by a projection layer with the required number of outputs. For example, the total size of the model may be 114 million parameters.
[0154] As discussed below, the technology may involve first pre-training a binding affinity model (e.g., on immunoglobulin sequences) in a self-supervised regime (e.g., using an Observational Antibody Space Database (OAS)). Immunoglobulin chains may be represented by tokens that code for the species from which the biomolecule (e.g., antibody) originates, followed by a concatenation of the complementarity determining regions (CDRs), e.g., IMGT (Lefranc, M.-P., Pommie, C., Ruiz, M., Giudicelli, V., Foulquier, E., Truong, L., Thouvenin-Contet, V., and Lefranc, G. Imgt unique numbering for immunoglobulin and t cell receptor variable domains and ig superfamily v-like domains. Developmental & Comparative Immunology, 27(1):55-77, 2003.) and Martin (Abhinandan, K. and Martin, ACA analysis and improvements to kabat and structurally correct numbering of antibody variable domains. Molecular Immunology, 45(14):3832-3839, 2008.) labeling schemes (e.g., with separator tokens between them). For example, as discussed below, the technology can include incoming data (e.g., the OAS database) that includes unpaired immunoglobulin chains and excluded tests (e.g., tests whose samples are also part of another test present in the database, tests derived from immature B cells, B cell-related cancers, etc.). In some aspects, the technology can include further filtering sequences that fail multiple quality checks and / or extracting representations of the desired chains (e.g., extended CDRs or near-complete) and / or de-duplicating the resulting sequences across the entire database.The technique may further include filtering sequences that have been observed only once in a single test, delineating the dataset and training configuration for each pre-training task, as shown in the table below: [Table 1]
[0155] Antibody variants, including sequence variants (e.g., within a region or regions of a reference antibody), are described in detail herein. A sequence variant is a sequence that deviates from a reference / wild-type sequence by one or more mutations. Mutations are often introduced in the CDRs (as specified according to any common definition, e.g., IMGT, Martin, Kabat, Chothia, etc.), but in principle can also be introduced in the framework. This technique involves intentionally "mutating" a sequence to generate that sequence variant.
[0156] The data and predictions provided by this disclosure have been generated using Fabs; however, alternative scaffolds, such as mAbs, scFvs, VHHs, etc., as well as heavy versus light chains, are possible with very similar modeling approaches.
[0157] As described herein, in one embodiment, the model is first pre-trained using natural antibody sequencing data from multiple species (e.g., any sequence data related to any suitable species currently known or later developed), including human, mouse, and camelid, among others. Pre-training, in one embodiment, is performed using a masked language modeling objective: some positions in the antibody sequence are randomly masked, and the model is tasked with predicting which amino acid was present at the masked position (a classification task). By doing so, the model gains an understanding of the "grammar" governing the antibody sequence (i.e., it gains an understanding of "naturalness" and / or "humanness"), thereby making subsequent fine-tuning of the model using affinity data more efficient. Pre-training does not require labeling data: only antibody sequences are needed, without knowledge of their antigen specificity or other properties. The only requirement is that these sequences be natural sequences. These sequences were sourced from the Observed Antibody Space (OAS), a database published by the Oxford Protein Informatics Group (OPIG) at the University of Oxford. OAS does not contribute novel sequences; it is an aggregation of data sourced from multiple publications. However, it re-annotates raw data from such heterogeneous sources in a unified pipeline (Kovaltsuk, et al., J Immunol, 2018, 201(8)2502-2509). OAS re-annotation is useful, but is likely not a major factor in modeling performance. Similarly, aggregation of data from multiple studies is useful, but a single study is unlikely to be essential for modeling performance.
[0158] The present disclosure provides, in some embodiments, that pre-training improves affinity prediction when proprietary affinity data is limited. As the size of the proprietary affinity dataset increases, the benefits of pre-training diminish. As such, while pre-training is characterized as optional in one embodiment, we still emphasize the benefits of pre-training in "low-N" settings (i.e., small affinity datasets). This concept is well established in protein engineering (Biswas, S., Khimulya, G., Alley, EC et al. Low-N protein engineering with data-efficient deep learning. Nat Methods 18, 389-396 (2021). https: / / doi.org / 10.1038 / s41592-21-01100-y), but no antibody-specific demonstration has been made to date. The distinction between general protein engineering and antibody engineering is important because not all machine learning methods developed for proteins work with antibodies. For example, AlphaFold2 can predict protein complexes, but cannot predict antibody-antigen interactions. (Evans, R., et al., bioRxiv 2021.10.4.463034) Additionally, in some embodiments, pre-training is beneficial for reasons unrelated to low N settings, such as the ability to make predictions about variants with high mutational burden using training data from variants with only (among other things) low mutational burden.
[0159] After pre-training (or random initialization, in embodiments where pre-training is not performed), the model may be trained (i.e., fine-tuned) using affinity data generated using a workflow encompassing a primary screen, for example, using a high-throughput, quantitative, activity-based method (WO / 2021 / 146626) and targeted rescreening with Carterra LSA SPR (high accuracy) (carterra-bio.com / lsa / ). In other embodiments, other display technologies (e.g., yeast display, mRNA display, phage display, ribosome display, etc.) are used in a deep mutational scanning (DMS) setting (Kyrin, R., et al., Trends in Pharmacological Sciences, 2021, ISSN 0165-6147). In still other embodiments, Carterra LSA SPR can be replaced with low-throughput / conventional SPR, BLI, or similar technologies. While data collection is described herein in two steps to maximize throughput (primary screening) and accuracy (secondary rescreening), in some embodiments, the present disclosure provides either of two steps for generating training data. Because the affinity data is antigen-specific, the model is also antigen-specific. For new antibody / antigen pairs, pre-training is not repeated, but affinity data collection and model fine-tuning are.
[0160] During model pre-training and fine-tuning, the model can perform affinity predictions for unseen sequence variants. For low combinatorial space (e.g., triple mutants in two CDRs), the number of sequence combinations is small enough to be addressed by exhaustively predicting the affinities of all possible sequence variants.
[0161] For highly combinatorial spaces, the number of sequence combinations is too large to be exhaustively computationally predicted. In that case, two solutions are possible, as provided by this disclosure: (1) converting the predictive model into a generative model using the Plug and Play language model (Dathathri et al., arXiv:1912.02164 [cs.CL]); however, several other generative strategies have been published (e.g., Zachary Wu, Kadina E. Johnston, Frances H. Arnold, Kevin K. Yang, Protein sequence design with deep generative models, Current Opinion in Chemical Biology, Volume 65, 2021, Pages 18-27, ISSN 0222-2893). 1367-5931, https: / / doi.org / 10.1016 / j.cbpa.2021.4.004. (See https: / / www.sciencedirect.com / science / article / pii / S136759312100051X)), and (2) combining model predictions with more traditional optimization techniques, such as genetic algorithms and the like (e.g., simulated annealing). Still further methods are available to those skilled in the art (See, e.g., “Controllable Neural Text Generation,” https: / / lilianweng.github.io / lil-log / 2021 / 01 / 02 / controllable-neural-text-generation.html (Retrieved 1 / 7 / 2022)).
[0162] The methods and models provided by the present disclosure can be used for a number of purposes. In one embodiment, the methods and models provided by the present disclosure are used for affinity maturation of weak-binding antibodies, including commercially available antibodies. Such antibodies may be weak binders due to humanization of animal-derived antibodies, novel hits from library screening, poor target immunogenicity (e.g., mammalian antigens), or other reasons.
[0163] In one embodiment, the methods and models provided by the present disclosure are used for simultaneous affinity maturation of two or more antigens. Such antigens may be homologous proteins belonging to different species, often one human and the other a non-human species, such as cynomolgus monkeys. Engineering antibodies that bind to the same antigen from different species allows for in vivo testing during development. Alternatively, the antigens may be variants of the same protein. One example is in infectious diseases, where certain variants may escape antibody binding, thereby abrogating therapeutic efficacy. Restoring affinity for escape variants without compromising affinity for non-escape variants is valuable for providing antibodies with broad neutralizing activity. Alternatively, the antigens may be distinct members of the same family. This is beneficial when multiple members of the same family must be associated or neutralized, either because it increases therapeutic potential or because there is functional redundancy across the family, making it ineffective to associate / block a single member.
[0164] In another embodiment, the methods and models provided by the present disclosure are used for affinity maturation of the same antigen under different conditions, for example, such conditions may include varying pH, which may change in different microenvironments, thereby affecting binding.
[0165] While the emphasis in affinity maturation is often to increase binding affinity as much as possible, the models and methods described herein make and enable quantitative predictions, for example, (i) decreasing affinity rather than increasing it; (ii) manipulating affinity to be within defined lower and upper limits, e.g., to facilitate in vivo antibody clearance and / or limit target antigen association or blocking if side effects are present; and (iii) when performing multi-antigen affinity maturation, the goal will not necessarily be enhancement for all antigens—it may be desirable to increase affinity for one or more antigens while decreasing / neutralizing affinity for one or more other antigens (e.g., this may be advantageous if associating / blocking an antigen provides a therapeutic benefit while associating / blocking a related antigen leads to toxicity. Similarly, one may want to enhance affinity for a particular target while decreasing / neutralizing nonspecific binding to related, but undesirable, antigens).
[0166] While the emphasis is on affinity maturation of antibodies to antigens, in some embodiments, the same strategy can be applied to any protein-protein interaction. Depending on the nature of the two interactors, it may be beneficial to perform model pre-training on a protein database, such as Uniref90, rather than, or in addition to, OAS. For example, cytokine sequences can be engineered to increase / decrease binding affinity for receptors. Similarly, next-generation antibody scaffolds as well as antibody mimetic scaffolds (such as DARPins) can also be engineered. In another example, the Fc region of an antibody may be engineered to increase / decrease binding to a specific Fc receptor using the same strategy described herein for variable region affinity for antigen. In other aspects, models may be pre-trained on multiple databases, such as OAS and Uniref90, in a combined or sequential manner.
[0167] In one embodiment, the model requires training data specific to the pairwise interaction being optimized (i.e., affinity data of antibody sequence variants to a single antigen). A new pairwise interaction of interest requires a new training dataset specific to that interaction. However, the model architecture, model pre-training, and workflow (from data generation to model prediction) remain the same. While the model, in one embodiment, is used with training data specific to the optimized pairwise interaction (i.e., affinity data of antibody sequence variants to a single antigen) as described herein, the following embodiments are also provided herein: (a) the antigen (and not the antibody) is mutated while the antibody (and not the antigen) is fixed. This is useful, for example, for predicting which antigen variants are likely to escape from the antibody; (b) library-on-library screening and training, where both the antibody and the antigen (and not just the antibody) are mutagenized simultaneously. This is useful, for example, when engineering antibodies against antigen variants (e.g., escape variants for infectious diseases) without having to screen for a predetermined / fixed set of variants. This is also useful when insight into paratope / epitope residues is required; (c) Libraries of antibody sequences that are no longer variants of the reference sequence (i.e., separated by only one or a few mutations), but rather randomized CDR sequence libraries on a selected antibody scaffold. This may allow the model to learn on multiple unrelated binders to the same antigen. Different binders may target different antigenic sites / epitopes; (d) Libraries of antigens that are no longer variants of the reference sequence, but rather randomized peptides, either linear or conformational / cyclic.This allows the model to learn which motifs (linear or 3D) bind to antibody sequences; (e) a combination of the previous two points (library on library), which may allow the model to learn near-universal relationships between any antibody sequence and any structural motif (captured by a peptide library). This may ultimately allow de novo in silico antibody design, where one identifies select epitopes in terms of an alphabet of structural motifs and then asks the model to generate sequences with affinity for those structural motifs.
[0168] In one embodiment of the present disclosure, affinity is K D (or K D It is given numerically as a single-valued alternative / correlation of K D are the association and dissociation constants (K a and K. d ) and the same K D is K a and K. d Since the model can be derived from different combinations of K D Rather, K a and K. d The task can be to predict the affinity of a given molecule, which is useful when specific association / dissociation rates are desired as opposed to overall affinity.
[0169] In one embodiment, the focus is on antigen affinity, while the models provided herein map sequences to numerical features. As such, the same modeling strategy (including pre-training with natural antibody sequences) can be used to predict any quantitative outcome for which there is sufficient training data. The decision to deploy a model to predict a different numerical property of an antibody is not dependent on the modeling strategy, which is invariant, but on the assay throughput, which should be sufficient to generate sufficient data for the fine-tuning process. It is important to note that, unlike affinity, most other properties of antibodies are not specific to a given antigen but depend exclusively on the antibody's sequence. This is the case, for example, with antibody biophysical / developability properties, such as solubility and viscosity. On the other hand, this means that training data obtained for one project may be combined with training data obtained for a different project, even if these two projects involve different antibodies / antigens. On the other hand, generalizable predictions require training data that span greater sequence variation than simply local variation around a few reference / wild-type sequences.
[0170] As used herein, the term "developability" refers to the feasibility of a molecule to successfully progress from discovery to development through evaluation of its physicochemical properties. The term "developability" can also include concepts related to the ability of a molecule (e.g., an antibody) to bind to a desired target molecule, as well as other considerations (e.g., feasibility of manufacturing, stability in storage, and absence of off-target stickiness). As used herein, "binding partner" refers to a molecule with which another molecule forms a physical interaction. For example, the binding partner of an antibody is its antigen. As used herein, the term "binding properties" includes, but is not limited to, the equilibrium dissociation constant (K), which is a metric for measuring binding affinity. D ), dissociation constant (K d ), and the association constant (K a ) is included. Dmay be defined as the concentration of ligand such that half of the ligand binding sites on the protein are occupied at system equilibrium. It is the same as the rate constant (K off ) is calculated by dividing the rate at which the forward reaction occurs during the formation of the protein-ligand complex (K on )
[0171] The biophysical / developability properties of the present technology represent a significant and advantageous improvement over conventional techniques. In particular, careful data preprocessing and preparation in the present technology, particularly in embodiments of the present technology using model pre-training and / or model fine-tuning, significantly improves upon conventional methods based on random model initialization. Specifically, sequences generated by the present modeling technology inherently have better developability properties because the model is informed by the natural immune repertoire and trained with antibodies that actually exist in humans. In this way, the modeling results are better and more developable than any results based on randomized approaches.
[0172] Those skilled in the art will understand that it is possible, even if unlikely, to find high-affinity biomolecules that involve sequences not found in humans or other animals and thus are not developable. Sequence identification is a multivariate problem, and affinity binding is only one variable. Finding high-affinity antibodies that suffer from developability issues (poor solubility, excessive viscosity, human immunogenicity, etc.) is a major research and development obstacle overcome by this technology. These obstacles are particularly prevalent when using phage display technology.
[0173] Furthermore, while the modeling techniques may be biased toward "human" (i.e., sequences that are more similar to those found in humans and thus more likely to be exploitable), this does not mean that the techniques cannot output sequences that appear unnatural. Indeed, in some cases, the strength of affinity of a non-natural biomolecule may override the penalty for unnaturalness that results from pre-training data.
[0174] This technique is highly sensitive, and therefore, in practice, experimental error / noise generated during binding assays may affect modeling output in some ranges. For example, given the range of predictions introduced during the assay (e.g., between 0.1 pmol and 0.2 pmol), successive experimental runs of the trained model may produce results with different orderings (e.g., the top two affinity variants may be transposed). This may, in some cases, prevent relative ranking of variants by affinity. However, as discussed herein, this technique still represents a significant and advantageous improvement over conventional techniques limited to binary classification, while the technique is quantitative and generally allows mathematical manipulations (e.g., ranking, sorting, averaging, limiting, etc.) over several orders of magnitude.
[0175] The following references describe various aspects of protein engineering and modeling: protein engineering using machine learning (Biswas et al.; and Evans et al.); antibody affinity engineering using machine learning (Mason et al., and Hanning et al., Trends Pharm Sciences, 2021, doi.org / 10.1016 / j.tips.2021.11.10); generative NLP models (Dathathri et al.); antibody sequencing data (Kovaltsuk et al.); model architecture (Liu et al. 2019 arXiv, arXiv:1907.11692); and machine learning guided polypeptide design (WO2021026037A1 and WO2020167667A1).
[0176] The in silico screening aspect of the present invention offers many important advantages over wet-lab assays, including dramatically higher throughput and multi-objective optimization. D We demonstrate the ability of deep learning models to accurately predict binding affinities across a range of sequences, which may include an implicit understanding of sequence naturalness and provide powerful measures of various developability metrics. This technology represents an advantageous and practical step toward enhanced in silico antibody design for therapeutic applications.
[0177] Exemplary Computer-Implemented Machine Learning Training and Operation 1 depicts an exemplary computing environment 100 for training and / or operating one or more machine learning (ML) models, according to some aspects. The environment 100 includes a client computing device 102, a molecular modeling server 104, an assay device 106, and an electronic network 108. Some embodiments may include multiple client devices 102, multiple molecular modeling servers 104, and / or multiple assay devices 106. Generally, the one or more molecular modeling servers 104 operate to perform the training and operation of fully or partially in silico molecular models as described herein.
[0178] The client computing device 102 may be an individual server, a group (e.g., a cluster) of servers, or another suitable type of computing device or system (e.g., a collection of computing resources). For example, the client computing device 102 may be any suitable computing device (e.g., a server, a mobile computing device, a smartphone, a tablet, a laptop, a wearable device, etc.). In some embodiments, one or more components of the client device 102 may be embodied in one or more virtual appearances (e.g., a cloud-based virtualization service) and / or may be included in a respective remote data center (e.g., a cloud computing environment, a public cloud, a private cloud, a hybrid cloud, etc.). The client computing device 102 includes a processor and a network interface controller (NIC). The processor may include any suitable number of processors and / or processor types, such as a CPU and one or more graphics processing units (GPUs). Generally, the processor is configured to execute software instructions stored in memory. The memory may include one or more persistent memories (e.g., hard drives / solid-state memory) and may store one or more sets of computer-executable instructions / modules. For example, the executable instructions may receive and / or present results generated by the server 104 .
[0179] The client computing device 102 may include a respective input device and a respective output device. Each input device may include any suitable device or devices for receiving input, such as one or more microphones, one or more cameras, a hardware keyboard, a hardware mouse, a capacitive touchscreen, etc. Each output device may include any suitable device for communicating output, such as hardware speakers, a computer monitor, a touchscreen, etc. In some cases, the input device and the output device may be combined into a single device, such as a touchscreen device that receives user input and presents output. The NIC of the client computing device may include any suitable network interface controller, such as a wired / wireless controller (e.g., an Ethernet controller), etc., and may facilitate bidirectional / multiplexed networking across a network between the client computing device 102 and other components of the environment 100.
[0180] The molecular modeling server 104 includes a processor 150, a network interface controller (NIC) 152, and memory 154. The molecular modeling server 104 may further include a data repository 180. The data repository 180 may be a Structured Query Language (SQL) database (e.g., a MySQL database, an Oracle database, etc.) or another type of database (e.g., not just an SQL (NoSQL) database). In some embodiments, the data repository 180 may include a file system (e.g., an EXT file system, an Apple File System (APFS), a Network File System (NFS), a local file system, etc.), an object store (e.g., Amazon Web Services S3), a data lake, etc. The data repository 180 may include multiple data types, such as pre-training data and fine-tuning data sourced from public data sources (e.g., OAS data). The fine-tuning data may be proprietary affinity data sourced from the quantitative assay ACE, Carterra, or any other suitable source.
[0181] Server 104 may include a library of client bindings for accessing data repository 180. In some embodiments, data repository 180 is located remotely from molecular modeling server 104. For example, data repository 180 may, in some aspects, be implemented using a RESTdb.IO database, Amazon Relational Database Service (RDS), or the like. In some aspects, molecular modeling server 104 may include a client-server platform technology, such as Python, PHP, ASP.NET, Java J2EE, Ruby on Rails, Node.js, web services, or an online API, that is responsive to receive and respond to electronic requests. Additionally, molecular modeling server 104 may include a set of instructions for performing machine learning operations, which may be integrated with client-server platform technology, as discussed below.
[0182] The assay device 106 may be a surface plasmon resonance (SPR) machine, such as a Carterra SPR machine. The device 106 may be physically connected to either the molecular modeling server 104 or the data repository 180, as depicted. The device 106 may be located in a laboratory and may be accessible from one or more computers in the laboratory (not depicted) and / or from the molecular modeling server 104. The device 106 may generate data and upload the data to the data repository 180 directly and / or via a laboratory computer. The assay device 106 may include instructions for receiving one or more sequences (e.g., mutant sequences) and for synthesizing those sequences. Synthesis may sometimes be performed via another technique (e.g., via a different device or via a human). In some aspects, the device 106 may not be configured as a device, but as an alternative assay capable of measuring protein-protein interactions as listed in other sections of this application. For example, device 106 may alternatively be configured as a complete device / workflow, including plate and liquid handling. In general, device 106 may be replaced with appropriate hardware and / or software, optionally including a human operator, to generate affinity data.
[0183] Network 108 may be a single communications network or may include multiple communications networks of one or more types (e.g., one or more wired and / or wireless local area networks (LANs) and / or one or more wired and / or wireless wide area networks (WANs), such as the Internet). Network 108 may, for example, enable bidirectional communication between client computing device 102 and molecular modeling server 104.
[0184] Processor 150 may include any suitable number of processors and / or processor types, such as a CPU and one or more graphics processing units (GPUs). Generally, processor 150 is configured to execute software instructions stored in memory 154. Memory 154 may include one or more persistent memories (e.g., hard drives / solid-state memory) and stores one or more sets of computer-executable instructions / modules 160, including an input / output (I / O) module 162, a variant module 164, an assay module 166, a sequencing module 168, a machine learning training module 170, a machine learning operation module 172; and a variant identification module 174.
[0185] Each of the modules 160 performs particular functions related to the present technology, as described further below. The modules 160 may store machine-readable instructions, including one or more applications, one or more software components, and / or one or more APIs, which may be implemented to facilitate or perform features, functions, or other disclosures described herein, such as any methods, processes, elements, or limitations illustrated, depicted, or described in the various flowcharts, illustrations, figures, drawings, and / or other disclosures herein. In some embodiments, multiple modules 160 may perform particular technologies cooperatively. For example, machine learning operation module 172 may load information from one or more other models before, during, and / or after initiating an inference operation. As such, the modules 160 may exchange data via any suitable technology, such as inter-process communication (IPC), representational state transfer (REST) APIs, etc., within a single computing device, such as the molecular modeling server 104. In some embodiments, one or more modules 160 may execute on multiple computing devices (e.g., multiple servers 104). Modules 160 may exchange data between multiple computing devices over a network, such as network 108. Modules 160 of FIG. 1 will now be described in more detail.
[0186] Generally, I / O module 162 includes instructions that enable a user (e.g., an employee of a company) to access and operate molecular modeling server 104 (e.g., via client computing device 102). For example, the employee may be a software developer using ML training module 170 to train one or more ML models in preparation for using the trained ML model(s) to generate outputs used in an antibody modeling project. Once the ML model(s) are trained, the same user (or another user) may access molecular modeling server 104 via the I / O module to initiate the molecular modeling process. I / O module 162 may include instructions for generating one or more graphical user interfaces (GUIs) (not depicted) that collect and store parameters related to biomolecular modeling, such as user selection of specific reference proteins, biomolecules, binding partners, etc., from lists stored in data repository 180.
[0187] Variant module 164 may include computer-executable instructions for generating one or more mutation sequence variants based on one or more reference biomolecules. For example, a user may use I / O module 162 to parameterize variant module 164 to selectively vary the manner in which reference biomolecule mutations are performed, which may involve performing repeated mutations, each of which variant module 164 may store in data repository 180 using a series of mutation storage instructions. In this manner, a user may retrieve previously performed parameterized mutation sequence variants or load the results of that mutation via I / O module 162.
[0188] Assay module 166 may include computer-executable instructions for retrieving / receiving one or more synthesized mutational variants (e.g., if stored via memory 154 and / or via data repository 180) and for controlling assay machine 106. For example, assay module 166 may include instructions for causing assay machine 106 to analyze the synthesized mutational variants. The assay module may include instructions for performing next-generation sequencing to determine binding kinetics to determine measured binding affinities, as shown in FIG. 2. Assay module 166 may store the determined measured binding affinities in association with one or more mutational variants in data repository 180, allowing another module / process (e.g., sequencing module 168) to retrieve the variants along with their measured binding affinities and other related data.
[0189] The sequencing module 168, in some embodiments, may include computer-executable instructions for manipulating gene sequences and for converting data generated by the assay module 166 and for the operation of the assay machine 106. The sequencing module 168 may store the converted assay data, for example, in a separate database table in the electronic data repository 180. The sequencing module 168 may also, in some cases, include software libraries for accessing third-party data sources, such as OAS.
[0190] Exemplary Computer-Implemented Machine Learning Model Training and Model Operation Generally, a computer program or computer-based product, application, or code (e.g., a model, such as a machine learning model, or other computing instructions described herein) may be stored on a computer-usable storage medium or tangible, non-transitory computer-readable medium (e.g., standard random access memory (RAM), an optical disk, a universal serial bus (USB) drive, or the like) having such computer-readable program code or computer instructions embodied therein, and the computer-readable program code or computer instructions may be installed on or otherwise adapted to be executed by processor 150 (e.g., operating in conjunction with a respective operating system in memory 154) to facilitate, implement, or perform the machine-readable instructions, methods, processes, elements, or limitations illustrated, depicted, or described in the various flowcharts, illustrations, figures, drawings, and / or other disclosure herein. In this regard, the program code may be implemented in any desired programming language and may be implemented as machine code, assembly code, bytecode, interpretable source code, or the like (e.g., via Golang, Python, C, C++, C#, Objective-C, Java, Scala, ActionScript, JavaScript, HTML, CSS, XML, etc.).
[0191] In some embodiments, computing module 160 may include an ML model training module 170 that includes a set of computer-executable instructions that implement machine learning training, configuration, parameterization, and / or storage functions. ML model training module 170 may initialize, train, and / or store one or more ML models as discussed herein. Trained ML models and their weights / parameters may be stored in a data repository 180 that is accessible to or otherwise communicatively associated with molecular modeling server 104.
[0192] For example, the ML training module 170 may train one or more ML models (e.g., artificial neural networks (ANNs)). One or more training datasets may be used for model training in the present technology, as discussed herein. The input data may have a particular shape, which may affect the ANN network architecture. The elements of the training dataset may include tensors scaled to small values (e.g., in the range (-1.0, 1.0)). In some embodiments, a preprocessing layer may be included in the training (and operation) that applies principal component analysis (PCA) or another technique to the input data. PCA or another dimensionality reduction technique may be applied during training to reduce the dimensionality from a high number to a relatively smaller number. Reducing the dimensionality may result in a substantial reduction in the computational resources (e.g., memory and CPU cycles) required to train and / or analyze the input data.
[0193] Generally, training an ANN may include establishing a network architecture, or topology, adding layers, including an activation function for each layer (e.g., "leaky" rectified linear unit (ReLU), softmax, hyperbolic tangent (tanh), etc.), a loss function, and an optimizer. In one aspect, the ANN may use different activation functions at each layer or between the hidden layer and the output layer. Suitable optimizers may include Adam and Nadam optimizers. In one aspect, different neural network types (e.g., recurrent neural networks, deep learning neural networks, etc.) may be chosen. The training data may be split into training, validation, and test data. For example, 20% of the training dataset may be retained for later validation and / or testing. In that example, 80% of the training dataset may be used for training. In that example, the training dataset data may be shuffled before being so split. Splitting the dataset may also be performed in a cross-validation setting, for example, when the dataset is small. Data input to an artificial neural network may be coded in an N-dimensional tensor, array, matrix, and / or other suitable data structure. In some embodiments, training may be performed by continuously evaluating (e.g., looping) the network using labeled training samples. The process of training an ANN may change the weights, or parameters, of the ANN. The weights may be initialized to random values. The weights may be adjusted as the network is continuously trained to reduce the loss and converge the values output by the network to expected, or "learned," values by using one or more gradient descent algorithms. In one embodiment, regression without an activation function may be used, in which the input data may be normalized by mean centering, and a mean squared error loss function may be used in addition to mean absolute error to determine a suitable loss and quantify the accuracy of the output.
[0194] In some embodiments, the ML training module 170 may include computer-executable instructions for performing ML model pre-training, ML model fine-tuning, and / or ML model self-supervised training. Model pre-training is known as transfer learning, and, for example, it may enable the training of a base model that is universal in the sense that it can be used as a common grammar for all antibody sequences. The term "pre-training" may be used to describe scenarios in which a second training may occur (i.e., when the model may be "fine-tuned"). Transfer learning refers to the model's ability to utilize the results (weights) of a first pre-training to better initialize a second training, which may otherwise require random initialization. The second training, i.e., fine-tuning, may be performed using proprietary affinity data, as discussed herein. Techniques that combine pre-training and fine-tuning advantageously enhance performance in that the results of training on affinity data perform better after pre-training (e.g., using natural antibody sequences from the described OAS) than if pre-training were not performed. Model fine-tuning may, in some embodiments, be performed for a given antibody-antigen pair. ML model self-supervised learning can be performed to give the model an understanding of antibody grammar during pre-training.
[0195] Generally, ML models can be trained as described herein using supervised, semi-supervised, or unsupervised machine learning programs or algorithms. The machine learning programs or algorithms may use neural networks, which may be convolutional neural networks, deep learning neural networks, transformers, autoencoders, and / or combined learning modules or programs that learn on two or more features or feature datasets (e.g., structured data, unstructured data, etc.) in a particular domain of interest. The machine learning programs or algorithms may also include natural language processing, semantic analysis, automated inference, regression analysis, support vector machine (SVM) analysis, decision tree analysis, random forest analysis, K-nearest neighbor analysis, naive Bayes analysis, clustering, reinforcement learning, and / or other machine learning algorithms and / or techniques (e.g., generative algorithms, genetic algorithms, etc.).
[0196] In some embodiments, an ML algorithm or technique may be selected for a particular input based on the problem set size of the input. In some embodiments, the artificial intelligence and / or machine learning-based algorithm may be based on or otherwise incorporate aspects of one or more machine learning algorithms included as libraries or packages executed on server 104. For example, the libraries may include TensorFlow-based libraries, Pytorch libraries (e.g., PyTorch Lightning), Keras libraries, Jax libraries, the HuggingFace ecosystem (e.g., the transformer, dataset, and / or tokenizer libraries therein), and / or the scikit-learn Python library. However, these popular open-source libraries are attractive and not required. The techniques may be implemented using other frameworks / languages.
[0197] Machine learning can involve identifying and recognizing patterns in existing data (e.g., binding affinity) to facilitate prediction, classification, and / or identification of subsequent data (e.g., using a trained model to predict variants with high binding affinity, etc.). Machine learning models can be created and trained based on example data (e.g., training data) inputs or data (which may be referred to as "features" and "labels") to make valid and reliable predictions for new inputs. In supervised machine learning, a machine learning program running on a server, computing device, or other processor is provided with example inputs (e.g., "features") and their associated or observed outputs (e.g., "labels"), and the machine learning program or algorithm determines or discovers rules, relationships, patterns, or otherwise machine learning "models" by which such inputs (e.g., "features") are mapped to outputs (e.g., labels), for example, by determining and / or assigning weights or other metrics to the model across its various trait categories. Such rules, relationships, or otherwise models are then provided with subsequent inputs and the model running on the server, computing device, or other processor predicts an expected output based on the discovered rules, relationships, or models.
[0198] For example, the ML training module 170 may analyze labeled data at an input layer of a model having a network layer architecture (e.g., an artificial neural network, a convolutional neural network, a deep neural network, etc.) to generate an ML model. The training data may be, for example, sequence variants labeled according to affinity. During training, the labeled data may be propagated through one or more connected deep layers of the ML model to establish weights for one or more nodes or neurons in each layer. Initially, the weights may be initialized to random values, and one or more appropriate activation functions may be selected for the training process, as will be understood by those skilled in the art. The ML training module 170 may include training the output layer of each of the one or more machine learning models. The output layer may be trained to output predictions. For example, an ML model trained herein may predict the binding affinity of an unseen sequence variant by analyzing labeled examples provided during training. In some embodiments, binding affinity may be expressed as a real number (e.g., in regression analysis). In some embodiments, binding affinity may be expressed as a Boolean value (e.g., in classification). In some aspects, multiple ANNs may be trained and / or operated separately, for example, individual models may be fine-tuned (i.e., trained) based on a pre-trained model using transfer learning for multiple different antibody-antigen pairs.
[0199] In unsupervised or semi-supervised machine learning, a server, computing device, or other processor may be required to discover its own structure in unlabeled example inputs, where multiple training iterations are performed by the server, computing device, or other processor to train multiple generations of models, for example, until a satisfactory model is generated. In the present technology, semi-supervised learning may be used, among other things, for natural language processing purposes and to learn the grammar of antibody sequences using objectives, such as masked language model objectives. Supervised learning and / or unsupervised machine learning may also include retraining, retraining, or otherwise updating a model with new or different information that may be received, incorporated, generated, or otherwise used over time. In various aspects, training an ML model herein may include generating an ensemble model including multiple models or sub-models, including models trained with the same and / or different AI algorithms and configured to operate together, as described herein.
[0200] Once model training module 170 initializes one or more ML models, which may be ANNs or recurrent networks, for example, model training module 170 trains the ML models by inputting labeled data into the models (e.g., affinity-tagged antibody variants). The trained ML models can be expected to provide accurate affinity predictions given antibody variant inputs not previously seen by the models (i.e., not used during training).
[0201] The model training module 170 may split the labeled data into respective training and test data sets. The model training module 170 may train an ANN using the labeled data. The model training module 170 may calculate an accuracy / error metric (e.g., cross-entropy) using the test data and the test-corresponding label set. The model training module 170 may serialize the trained model and store the trained model in a database (e.g., data repository 180). Of course, those skilled in the art will understand that the model training module 170 may train and store more than one model. For example, the model training module 170 may train a separate model for each antibody-antigen pair. It should be understood that the structure of the described network may vary depending on the embodiment.
[0202] In some aspects, computing module 160 may include a machine learning operations module 172 that includes a set of computer-executable instructions implementing machine learning loading, configuration, initialization, and / or operation functions. ML operations module 172 may include instructions for storing the trained model (e.g., in electronic data repository 180, as pickled binaries, etc.). Once trained, the trained ML model may be operated in an inference mode, where when the model is presented with new inputs not previously presented, the model may output one or more predictions, classifications, etc., as described herein. In unsupervised learning aspects, a loss minimization function may be used, for example, to teach the ML model to generate outputs that resemble known outputs (i.e., ground truth examples).
[0203] Once a model has been trained by model training module 170, model manipulation module 172 may load one or more trained models (e.g., from data repository 180). Model manipulation module 172 typically applies new data not previously analyzed by the trained model to the trained model. For example, model manipulation module 172 may load a serialized model, deserialize the model, and load the model into memory 154. Model manipulation module 172 may load new molecular variant data that was not used to train the trained model. For example, the new molecular data may include antibody sequence data, antigen sequence data, etc., as described herein, encoded as input tensors. Model manipulation module 172 may apply one or more input tensors to the trained ML model. Model manipulation module 172 may receive output (e.g., tensors, feature maps, etc.) from the trained ML model. The output of the ML model may be a prediction of affinity associated with the input sequence. In this way, the present technology advantageously provides a means for quantitatively estimating molecular affinities that is much more accurate and data-rich than conventional industry practices. The advantage is that measuring these molecular affinities is time-consuming and expensive because it must be performed in a laboratory. By using ML, the present technology only requires laboratory measurements to generate a training set, and then, rather than requiring the continuous use of a wet lab, can predict unmeasured sequence variants / KD pairs in a relatively inexpensive and rapid manner due to its in silico capabilities.
[0204] Model manipulation module 172 may be accessed by another element (e.g., a web service) of molecular modeling server 104. ML manipulation module 172 may pass its output to variant identification module 174 for further processing / analysis. Alternatively, variant identification module 174 may receive results stored by ML manipulation module 172 in electronic data repository 180. For example, variant identification module 174 may evaluate the output of ML manipulation module 172 using a set of rules to identify one or more variants of interest (e.g., those having highest binding, lowest binding, or other properties as discussed herein). Variant identification module 174 may include further instructions for providing one or more sequence variants of interest as input (e.g., via email, as a visualization such as a chart / graph, as an element of a GUI in a computing device such as client computing device 102, etc.). In some embodiments, users may interact with ML models during training and / or operation using command line tools, application programming interfaces (APIs), software development kits (SDKs), Jupyter notebooks, etc.
[0205] With respect to modules 160, those skilled in the art will understand that in some aspects, the software instructions comprising modules 160 may be organized differently and may include more or fewer modules. For example, one or more of modules 160 may be omitted or combined. In some aspects, additional modules may be added (e.g., a localization module). In some embodiments, software libraries implementing one or more modules (e.g., Python code) may be combined, such that, for example, ML training module 170 and ML operation module 172 are a single set of executable instructions used to train and generate predictions. In still further examples, modules 160 may not include assay module 166 and / or sequencing module 168. For example, laboratory computers and / or assay devices 106 may implement those modules and / or others of modules 160. In that case, assays and sequencing may be performed in the laboratory to generate training data that is stored in data repository 180 and accessed by server 104.
[0206] Exemplary Computer-Implemented Method Model Training FIG. 2 depicts an exemplary data flow block diagram of a computer-implemented method 200 for training a machine learning model to predict binding of previously unseen sequence variants, according to some embodiments of the present technology.
[0207] As discussed above, the present techniques may include a one-step or two-step training procedure. Either of these techniques may involve model training using a limited number of data points generated in a wet lab to build and train one or more models capable of predicting variants with desired properties (e.g., high affinity). As described above, in some embodiments, the training process includes both pre-training using human antibody sequences and fine-tuning using affinity data. At training time, training data (e.g., KD measurements for specific sequence variants) are obtained from wet lab experiments / assays performed on synthesized variant sequences. Method 200 may perform pre-training and / or fine-tuning. Using both pre-training and fine-tuning has been empirically shown to provide the best performance. Using fine-tuning alone provides second-best performance, while using pre-training alone provides third-best performance. Specifically, method 200 may include receiving screening data (block 204) that includes a ranking of biomolecular sequence variants according to one or more training binding properties. The training binding properties may include ranking according to affinity and may be performed by one or more "wet lab" binding assays of the synthesized biomolecule sequence variants. For example, the assays may include activity-based screening techniques, SPR techniques, and / or others, as discussed herein.
[0208] In some embodiments, method 200 may include receiving rescreening data corresponding to biomolecule sequence variants to amplify / improve the training binding properties, and using the rescreening data to further train the machine learning model to improve the accuracy of the model (block 206). For example, the rescreening data may be used to amplify / improve the training binding properties of the biomolecule sequence variants. DThis may increase precision in range. Note that method 200 may include generating a graph of the measured binding properties (e.g., measured binding affinities) to provide a visual demonstration of the relative affinity of each sequence, as shown in block 206. Information determined by assays, such as binding kinetics and next-generation sequencing, may be received in block 206.
[0209] Method 200 may include creating an AI / ML model and training the machine learning model using the received screening training data to predict one or more desired binding properties of the input biomolecular sequence variants (block 208). The training data may include the assayed synthetic biomolecular sequence variants from blocks 204 and / or 206, each with a respective measured binding property (e.g., affinity) representing each variant's ability to bind to a corresponding respective binding partner biomolecule (e.g., an antigen if the biomolecular sequence is associated with an antibody, and an antibody if the biomolecular sequence is associated with an antigen). As discussed, the training in block 208 may apply transfer learning, in which a generalized / universal (i.e., pre-trained) model endowed with knowledge of antibody grammar is used in conjunction with a fine-tuning process that includes specific antibody-antigen training based on affinity data. Method 200 may also include cross-validating the machine learning model, as discussed below with respect to Example 1, and using measured K D and the predicted KD (block 210).
[0210] Through the training process of method 200, method 200 may adjust the weights of the machine learning model so that the model learns the rules underpinning antibody / antigen interactions. In this way, as discussed in the next section, the trained model can reliably predict biomolecular binding properties of input biomolecular sequence variants that the model has not previously seen (block 212). The previously unseen (i.e., novel or simulated) biomolecular sequence variants input into the trained model may be generated via the process of block 202 or may come from another source entirely.
[0211] In some embodiments, model training (e.g., model training of method 200) may enable a trained model to predict one or more antibody variant properties, e.g., affinity, using only a limited amount of the entire possible variant / mutation space to train the model. Variant space, as discussed herein, may refer to the combinatorial search space of biomolecule (e.g., antibody / antigen, etc.) variants. For example, in some embodiments, 10% of the variant space may be used to train a model. Empirical testing has shown that using a lower percentage of the entire possible variant space still provides accurate predictions. Thus, in some embodiments, less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, or less than 1% of the entire possible variant space may be used to train a model. A percentage of less than 0.5%, less than 0.4%, less than 0.3%, less than 0.2%, or less than 0.1% of the total possible variant space may be used and still achieve a Pearson R value above 0.6.
[0212] Those skilled in the art will appreciate that using a relatively small percentage of the mutation space for the training data (e.g., 0.3%) while achieving high correlation values represents a significant improvement and advantageously enables the present technique to be used even in scenarios where limited computing power is available (e.g., via mobile devices, wearable devices, etc.). This also advantageously enables the present technique to perform training and retraining rapidly, thereby reducing latency and improving performance of applications that train and operate the disclosed models. Finally, the relatively low training data requirements advantageously reduce the amount of training data that must be stored.
[0213] Model inference The present technology may involve using the trained ML model described above to make predictions in silico to obtain a silico predicted KD. At a high level, this is similar to "simulating" an experiment in silico, since the only way to measure KD in a laboratory is to actually perform the assay, sequencing, etc. However, in silico, the present technology advantageously benefits from not having to simulate or determine every single step to obtain a KD. Rather, the AI can learn the relationship between sequence variants and KD and make KD predictions for unseen variants. These predictions are equivalent to laboratory-based experiments, even though the ML model does not explicitly simulate the experiment, assay, or sequencing run, but rather can output a predicted KD given novel (i.e., previously unseen) sequence variants.
[0214] FIG. 3A depicts a computer-implemented method 300 for operating a trained machine learning model to identify one or more biomolecular sequence variants of interest, according to some embodiments. Method 300 may include receiving one or more simulated / unseen biomolecular sequence variants. In some aspects, method 300 may generate the simulated biomolecular sequence variants via mutation, as discussed herein. For example, training data may be generated in a wet lab, and one or more ML models may be trained, as discussed above. Method 300 may then include using the trained ML model to predict the KD for any sequences that are varied in silico. For example, method 300 may include changing amino acids in silico and inputting those sequences into the trained ML model to obtain the KD after these changes. In some aspects, this process may be optimized, in which the task of predicting the KD using a model (i.e., inference / prediction), given a sequence variant, is repeated (i.e., optimization) until a sequence variant with a desired predicted KD is found.
[0215] Inputs to the trained model may be generated using a suitable generative technique (e.g., a generative adversarial technique). Both generation and optimization have the same goal: to provide a KD to the trained model and obtain a sequence from it. Generation attempts to do so directly by running an ML algorithm in the opposite direction to guessing. Optimization instead utilizes classical guessing, which continues to run guesses until a sequence with the desired KD is found. Optimization is more efficient than an exhaustive search of all possible sequence variants, which would be inefficient and in some cases impossible. An example of a generative technique is plug-and-play, as discussed above, while an example of optimization is a genetic algorithm.
[0216] After the machine learning model has been trained, method 300 may include receiving a previously unseen biomolecule sequence variant, as discussed herein. Of course, method 300 may also receive previously seen variants (e.g., those used during training). Method 300 may also include processing one or more previously unseen (e.g., simulated) antibody sequence variants with the machine learning model to generate one or more predicted binding affinities, each corresponding to a respective one of the one or more previously unseen (e.g., simulated) antibody sequence variants (block 302). Specifically, as discussed herein, the machine learning model may have been pre-trained using a masked language model objective with grammar-dominated antibody sequence understanding. Alternatively, or in addition, the machine learning model may have been fine-tuned using affinity data, with model weights updated with affinity measurements. Previously unseen antibody sequence variants may be generated, for example, by mutagenesis of a reference biomolecule (e.g., antibody, antigen, etc.) using mutational techniques as discussed herein.
[0217] For example, given an antibody sequence, the machine learning model may generate a list of predicted binding affinities, each one associated with a respective antigen. Method 300 may further include analyzing the one or more predicted binding properties to identify one or more biomolecular sequence variants of interest from among the simulated sequence variants, each of the one or more biomolecular sequence variants of interest having a respective one or more desired properties. The desired properties may include aspects such as upper / lower limits of the predicted binding affinity, as well as many others, as discussed herein.
[0218] As discussed herein, the variant identification module 174 may include computer-executable instructions for analyzing the output of the machine learning model to identify one or more biomolecular sequence variants of interest (block 304). For example, as discussed herein, the properties of interest in one or more variants of interest may include one or more of the following: (i) an increase in at least one predicted binding affinity of the variant of interest; (ii) a decrease in at least one predicted binding affinity of the variant of interest; (iii) an upper limit in at least one predicted binding affinity of the variant of interest; (iv) a lower limit in at least one predicted binding affinity of the variant of interest; (v) an increase in the affinity of a first predicted binding affinity of the variant of interest for a first antigen and a decrease in the affinity of a second predicted binding affinity of the variant of interest for a second antigen; (vi) the ability of the cytokine sequence of the variant of interest to increase or decrease binding affinity to a receptor; (vii) the suitability of the variant of interest for use as a next-generation antibody scaffold and / or antibody mimetic scaffold; (viii) the ability of the variant of interest in the Fc region of an antibody to bind to an Fc receptor; or (ix) the developability of the variant of interest as indicated by tolerability upon administration.
[0219] Method 300 may include providing one or more biomolecular sequence variants of interest as an output (block 306). For example, method 200 may cause the variants of interest to be stored in an electronic database (e.g., the electronic database of FIG. 1), displayed on a display screen (e.g., the display of computing device 102 of FIG. 1), or otherwise transmitted to a user (e.g., via email).
[0220] In some aspects, training of the ML model may include a fixed antibody and a mutated antigen. In that case, inference may include an antigen search space (i.e., input of a previously unseen antigen). In the case of a fixed antibody / mutated antigen, the antibody sequence is constant and the antigen sequence may vary. However, even when the antigen sequence is varied, the antigen may be the same or may simply vary at some residues. The model may then output binding affinities. Thus, the search universe in this setting includes binding affinities for all antigen variants in the sequence space of all possible / desired antigen variants. In the case of a fixed antigen / mutated antibody, the antigen sequence is constant and the antibody sequence may vary. However, even when the antibody sequence is varied, the antibody may be the same or may simply vary at some residues. The model may then output binding affinities. Thus, the search universe in this setting includes binding affinities for all antibody variants in the sequence space of all possible / desired antibody variants.
[0221] In other aspects, training of the ML model may include fixed antigens and mutated antibodies, in which case inference may include antibody search space (i.e., input of previously unseen antibodies). In still further aspects, training of the ML model may include mutations of each of the antigen and antibody, in which case inference may include antigen and antibody search space (i.e., input of previously unseen antigens and antibodies).
[0222] 3B depicts a computer-implemented method 350 for training a machine learning model to identify biomolecular sequence variants of interest, according to some embodiments. The method may, in some embodiments, be performed by a computer, such as the molecular modeling server 102 of FIG. 1.
[0223] Method 350 may include generating one or more biomolecule sequence variants by programmatically mutating a reference biomolecule (block 352). As used herein, the terms "programmatic" or "programmatically" mean according to a method or system (e.g., via a computer-implemented method, via a computer program, via a computing system, etc.). Mutation may be performed according to principles discussed herein. For example, method 350 may include generating an antibody library that evenly samples sequence space around a starting antibody molecule. Method 350 may perform in silico mutations by a processor executing computer-executable instructions (e.g., by CPU 150 of FIG. 1 using variant module 164 of molecular modeling server 102).
[0224] Method 350 may include receiving screening data including a ranking of biomolecular sequence variants according to one or more training binding properties (block 354). As discussed, screening in the present technology may include wet-lab screening (e.g., using qaACE) and / or in silico screening. Method 350 may include training a machine learning model using the screening data to predict one or more desired binding properties of the input biomolecular sequence variants (block 356).
[0225] In some embodiments, method 350 may include receiving re-screening data corresponding to the biomolecular sequence variants to amplify one or more training binding properties; and further training a machine learning model using the re-screening data to improve the accuracy of the machine learning model. In some embodiments, the training binding property of method 350 may include binding affinity (KD). In some embodiments, the screening data of method 350 may be received from one or both of (i) a human experimenter and (ii) an assay device. In some embodiments, the one or more biomolecular sequence variants of method 350 include an antibody or an antigen.
[0226] Exemplary Data Flow Diagram 4A depicts an exemplary data flow diagram depicting training and prediction of biomolecular sequence variants of interest, according to some embodiments. In some embodiments, the data flow diagram of FIG. 4A may correspond to method 200 of FIG. 2.
[0227] Generally, Figure 4A depicts the goal 402 of identifying higher affinity monoclonal antibodies relative to wild-type monoclonal antibodies. Figure 4A also depicts a workflow overview 404 showing laboratory assay measurements being input into an AI model (e.g., the trained model or models discussed in Figure 2) to generate one or more predictions. Figure 4A includes an additional level of detail, indicating that the assayed biomolecule may be, for example, a library of SoluPro® strains and / or one or more sequence variants (e.g., trastuzumab Fab CDRH3 variants). The method includes a proprietary primary screen that ranks input variants by affinity (e.g., using the ACE Assay™) to identify a given K of interest. D This may include rescreening the input to increase accuracy in coverage and training an AI model to screen for unseen variants in silico.
[0228] Figure 4B depicts an exemplary block flow diagram for performing the assay of Figure 4A, according to some embodiments. Strains expressing unique antibody sequence variants may be added to fixed and permeabilized cells, and a probe may be added (Blocks 1 and 2). Such SoluPro® strains may contain labeled antigen reports for affinity and labeled scaffold-binding protein reports for potency. Strains may be screened and sorted by flow cytometry (Block 3). Next-generation sequencing may be performed, and an ACE affinity score may be generated (Blocks 4 and 5).
[0229] 4C depicts an example of an affinity prediction chart, according to some embodiments. In FIG. 4C, the model predicted K for trastuzumab CDRH3 sequence variants that bind to Her2. D and SPR measurement K D The observed correlation between is shown to span nearly four orders of magnitude, with an R Pearson correlation coefficient of 0.85.
[0230] Figure 4D depicts an exemplary affinity prediction validation chart, according to some embodiments. Figure 4D shows 20 sequence variants (15 antibodies predicted to bind Her2 more strongly than trastuzumab and 4 antibodies predicted to be slightly weaker binders than trastuzumab) that have already been validated by mid-throughput SPR and have undergone additional secondary validation by low-throughput BLI. The correlation coefficient in this case is R=0.94.
[0231] Example Noise Removal There is abundant literature on using deep learning for noise reduction in the field of single-cell RNA-seq, but noise reduction has not been used in the field of antibody engineering.
[0232] The model used to perform denoising can be the same as that used in Example 1 in terms of architecture / algorithm, except that instead of training using Carterra SPR sequences, the model can be trained using ACE data. Pre-training using natural antibody sequences is the same. Model inputs can be ACE-derived affinity scores with different degrees of accuracy. The higher the coverage (number of cells divided by the number of unique variants), the more accurate the score. ACE scores and Carterra K DA negative correlation of approximately 0.7 between ACE scores and ACE variants was observed, but such correlation decreases as coverage decreases. Thus, ACE data is fed to the model as training data, including sequence variants and their respective ACE scores. The model can then predict ACE scores for the same sequence variants used for training. The model predictions are not identical to the training set because the model attempts to generalize. We refer to these predictions as "denoised" scores, and empirical testing has shown that such scores are significantly higher in SPR K than the original (hard-measured) ACE scores. D It has been shown that the AF ratio correlates well with the AF ratio. Thus, noise removal is non-traditionally used in the present technology to maintain or improve the accuracy of predictions (e.g., affinity) while allowing for increased throughput. Measurements taken in an assay (e.g., an ACE assay) can be correlated with affinity, as shown in Figure 4E.
[0233] FIG. 4E depicts an exemplary denoising data chart, according to some embodiments. FIG. 4E depicts the original measurements (row 410a) and model-based denoising scores (row 410b), along with saturated (column 412a) and unsaturated (column 412b) libraries. ACE libraries are saturated when the sorting and sequencing capacity significantly exceeds the library size (i.e., has high coverage). Saturated libraries generally exhibit a significant correlation between ACE scores and SPR-derived K D This results in the highest accuracy in terms of correlation with K. When coverage is low, the library becomes unsaturated, resulting in poor accuracy. The present technology can include model training using sequence variants and ACE scores from an unsaturated library, and the model predicting ACE scores for the same sequence variants used in training. Such model-derived ACE scores (i.e., denoised ACE scores) of training sequence variants are more accurate than hard measurements derived from SPR. D In contrast, no enhanced model rendering accuracy was observed when the library was fully saturated.
[0234] A compromise always exists between assay throughput and assay accuracy. Adding more sequence variants and identifying more interesting potential sequences comes at the cost of a decrease in accuracy. This is because smaller libraries generally have more redundant measurements, which means the same sequence can be resampled multiple times, eliminating errors.
[0235] The model for performing denoising may be trained in the same way as other models described herein, for example, using ACE training data. However, prediction may no longer be unseen sequence variants (i.e., sequences that were not in the training data set). Rather, the predicted sequence variants are the same sequence variants used in the training data set. Rather than predicting unseen variants, in the context of denoising, prediction is that of the training data.
[0236] Denoising provides a significant benefit, restoring the correlation coefficient in the unsaturated library chart at row 410b, column 412b to its corresponding saturated chart at row 410b, column 412b in the depicted example of FIG. 4E. Essentially, this result means that accuracy can be maintained while significantly increasing throughput. The model denoises imprecise measurements, making them more accurate in the sense that they correlate better with ground truth (i.e., SPR) measurements. Another way to think of improved denoising techniques is as enabling the collection and use of less accurate data that would otherwise go unused.
[0237] The naturalness of exemplary machine learning The technology can include providing natural antibody sequences to teach naturalness to one or more machine learning models. The source of antibody sequences can be those used for pre-training (e.g., the OAS database), optionally supplemented with proprietary sequences, such as those from Totient.
[0238] FIG. 4F depicts an exemplary conceptual diagram illustrating naturalness training and prediction, according to some embodiments. Generally, “naturalness” is a measure of whether an input resembles training data. This similarity may be expressed simply in terms of shape, as in FIG. 4F. That is, during the training phase 460a, a model may be trained (e.g., by the ML model training module 170 of FIG. 1) using examples of polygons. The model may learn that polygons have certain characteristics, such as straight lines, closed geometric shapes, etc. The model may learn that other characteristics are not determinative of whether a given input is a polygon (e.g., color, line thickness, rotation, etc.). The model may be trained to generate a score for each input representing the probability that the input is a polygon. During the prediction phase 460b, the trained model may be used by inputting a set of individual shapes (e.g., circles, triangles, lines, etc.) to obtain a naturalness prediction for each. As shown, the model infers that a triangle is the most polygonal of those inputs.
[0239] The sequence variants tested in the training set can be generated by combinatorially enumerating all possible sequences when defining the mutational load and positions to be mutated, synthesizing all or a subsample (e.g., randomly selected) of them in the laboratory, and then testing them. The affinity measurements (ACE score and / or SPR KD) can then be fed into the model.
[0240] Of course, "naturalness" may be expressed differently depending on which inputs are used to train the model, and the characteristics that determine such similarity may not be very intuitive. For example, a similarity process can be applied to a biomolecule of interest (e.g., antibody sequences). Such a model may be trained using training data containing many (e.g., millions or more) examples of antibody sequences. These sequences may be from one or multiple species. Once trained, such a model can score the naturalness of previously unseen sequences, such as variants of a parent antibody. In particular, such trained models may be used to determine the naturalness of new antibodies generated purely in silico. This is useful for many reasons; companies seeking to design therapeutic agents may find it highly beneficial to weed out antibodies that cannot be used in humans.
[0241] FIG. 4G depicts an exemplary naturalness score validation chart, according to some embodiments. As discussed, a model may be trained to determine the naturalness score of a biomolecule, such as an antibody. Scores derived from such a model may behave as expected in technical validation. For example, as shown in FIG. 4G, different sequences may be visually presented according to their respective naturalness. In this example, numerous antibodies from the OAS database that pass the quality control (QC) filter ("positives") are shown in the distribution. QC failures (e.g., those whose annotation indicates missing beginning and / or ending CDR residues ("negatives")) are shown in the naturalness score distribution near zero. Scores for low-abundance sequences were also calculated using the modeling techniques described herein, which are expected to encompass pure antibodies as well as sequencing errors, albeit rare. Sequencing errors should be randomly distributed, thereby having little impact on conserved antibody residues, which would most likely penalize naturalness. Consistently, the naturalness distribution of low-abundance antibodies shows only a slight downward shift compared to QC-passing antibodies.
[0242] For sequences that fail the QC filter, "low abundance" means an abundance of 1 count across the dataset. "Missing start / end CDR residues" means that the antibody sequence annotation (typically done using a tool called ANARCI) is missing the start or end residue of one CDR. The OAS filter is described as follows: From the full unpaired dataset, studies are excluded if they have overlapping samples (e.g., 'Bonsignori et al., 2016', 'Halliley et al., 2015', 'Thornqvist et al., 2018'). Diseases can also be excluded: "light chain amyloidosis", "CLL". Type B can be excluded: "immature B cell", "pre-B cell". Regarding the sequences themselves, we can exclude those that: have a stop codon; are marked as non-productive or out-of-frame; have a non-conserved cysteine site; have j_identity<50; have no AA in FWR2 or FWR3; have more than 37 AA in CDR3; or are missing the first two or last two positions on any CDR according to IMGT. Then, for model training, we can exclude sequences with a cumulative redundancy of 1.
[0243] Figure 4H depicts an exemplary naturalness / developability correlation chart, according to some embodiments. To analyze the relationship between antibody naturalness and developability, the present technology can be used to score the naturalness of the number of hits (e.g., 5,000) by enrichment in the final panning round of a publicly available phage display library (Liu et al., Bioinformatics 36:2126 (2020)) (Chart 480a). The same sequences can also be analyzed using a therapeutic antibody profiler (Raybould et al., PNAS 116:4025 (2019)), recording the percentage that received at least one amber or red developability flag, or that could not be modeled at all (developability failures, Chart 480b). The top and bottom 10% of sequences by naturalness are then compared to observe the depletion and enrichment of developability failures, respectively, demonstrating the association between naturalness and developability. It will be understood by those skilled in the art that other characteristics may be compared and potentially correlated (eg, cohesion, viscosity, thermal stability, oxidation, etc.).
[0244] Figure 4I depicts an exemplary naturalness and immunogenicity correlation chart, according to some embodiments. In some embodiments, the relationship between antibody naturalness and immunogenicity may be explored using the present modeling techniques. For example, the model may use a CDR-only model to score the naturalness of therapeutic antibodies administered to humans (Phase I, Phase II, Phase III, or clinically approved), binned by origin. ((Marks et al., Bioinformatics 37:4041 (2021)). As shown, fully human antibodies can yield higher naturalness scores than antibodies of other classes (Chart 490a). A naturalness threshold may be defined below which no humanized, chimeric, or hybrid (humanized + chimeric) antibody sequences were found. The reported portion of patients who developed anti-drug antibody (ADA) responses to fully human antibodies (Marks et al., Bioinformatics 37:4041 (2021)) may be divided and analyzed according to a previously defined naturalness threshold, and lower immunogenicity to human antibodies was observed when naturalness was above the threshold (Chart 490b), suggesting that naturalness is inversely associated with immunogenicity. This is an important result because the ability to identify naturalness above a threshold, despite potential confounding factors, may aid in identifying drug candidates that are less likely to be rejected by the human immune system.
[0245] Figure 4J depicts an exemplary naturalness and mutational load correlation chart according to some embodiments. In some embodiments, the present technology may include scoring the naturalness of trastuzumab variants as a function of CDRH3 mutational load. As mutational load increases, the median naturalness may decrease. This observation suggests that a larger and larger proportion of random samples of combinatorial sequence space may fail downstream development as more mutations are introduced, given the previously discussed link between developability and immunogenicity. As a result, model-guided optimization of naturalness may be a superior strategy compared to screening random samples of antibody variants.
[0246] 4K depicts an exemplary chart of affinity prediction improvement when enriching with naturalness data, according to some embodiments. Feeding the model with example antibody sequences not only enabled the calculation of naturalness scores; it also enhanced the accuracy of affinity prediction via transfer learning. This allows for the prediction of K using a model trained on both unlabeled native antibody sequences and affinity measurements of trastuzumab variants (chart 492a, also depicted in FIG. 4C) or the latter alone (chart 492b). D and SPR measurement K D This allows for the observed correlation between
[0247] Figure 4L depicts an exemplary conceptual diagram of in silico sequence variant generation and optimization, according to some embodiments. Because the affinities of unseen variants (sequences not present in the training set, Figure 4C) were predicted by a model trained on affinity measurements of trastuzumab variants, screening experiments may be simulated in silico. However, naive simulations can be inefficient with large sequence spaces, assuming a defined mutation load and estimating the K of all possible sequence variants. D Alternatively, K D can be optimized at the cost of only a fraction of the computation using generation techniques.
[0248] To find sequence variants that are very good according to two criteria (e.g., naturalness and affinity), genetic algorithms may be used in conjunction with the deep learning models described herein to generatively find the best sequence variants without the need to combinatorially generate sequences. D To maximize K D or to minimize a specific K D It should be understood that the above may be used to find a value.
[0249] 4M depicts an exemplary chart of affinity predictions from trastuzumab, according to some embodiments. FIG. 4M shows K driven by a model trained on affinity measurements of trastuzumab CDRH3 sequence variants from the start of the trastuzumab sequence. D Depicts the incremental optimization (minimization, maximization, or adjustment to a particular value) of
[0250] FIG. 4N depicts an exemplary affinity prediction chart from different parent antibodies, according to some embodiments. D K driven by a model trained on affinity measurements of trastuzumab CDRH3 sequence variants from multiple starting points for trastuzumab variants spanning three orders of magnitude. D Depict the incremental optimization (minimization or maximization) of
[0251] FIG. 4O depicts an exemplary visualization of optimization for affinity and naturalness, according to some embodiments. D Co-optimization of (minimize, left, or maximize, right) and naturalness (maximize) is depicted. The starting point (black dot) was the trastuzumab sequence.
[0252] As shown above, the present technology involves deep learning models that predict binding affinity. In some embodiments, the technology involves training a number (e.g., three) of affinity prediction models based on HT and SPR data: one using only SPR, one using only HT data, and one using both in a multitask setting. In such embodiments, model performance can be evaluated, directed toward addressing two questions: first, self-prediction, or how well does the model reproduce the data used to supervise its training (cross-validation)? And second, K D Prediction, or how well does the model predict the actual KD (measured by SPR)? Empirical means that the models considered here compare self-prediction and true KD. DThe results show that the HT and SPR measurements are strongly predictive across a significant number (e.g., 4) of orders of magnitude for both predictions, as depicted in the table below and in Figure 4P, which shows performance statistics for trastuzumab affinity prediction (e.g., data points with SPR measurements also have HT measurements, resulting in a total combined training set size of 5689): [Table 2]
[0253] K D The HT model for prediction can be similar to the predictive power of laboratory measured data. In some embodiments, of the three models, the combined multitask model has the best K D The prediction performance was demonstrated.
[0254] As discussed with respect to conventional techniques, developing candidate biomolecules (e.g., antibodies) into therapeutic drugs is a complex process that involves a high degree of risk, particularly with regard to modeling these risks. The present technique makes it possible to model productive patterns and mitigate these issues using examples of antibodies in natural systems. For example, the present technique can include using the pre-training techniques described above (e.g., based on natural OAS sequences) to evaluate new sequences for "naturalness." This measure of naturalness can then be used as an additional measure for in silico optimization.
[0255] In some embodiments, to determine the usefulness of a naturalness score, the technology may evaluate an independent measure of therapeutic outcome (e.g., Therapeutic Antibody Profiler (TAP) (Raybould et al., 2019), which reports five criteria for developability for antibodies). In some embodiments, the technology has demonstrated a strong association between naturalness and TAP for sequences from phage display libraries (Liu, G., Zeng, H., Mueller, J., Carter, B., Wang, Z., Schilz, J., Horny, G., Birnbaum, M.E., Ewert, S., and Gifford, D.K. Antibody complementarity determining region design using high-capacity machine learning. Bioinformatics, 36(7):2126-2133, November 2019a. doi:10.1093 / bioinformatics / btz895.URL https: / / doi.org / 10.1093 / bioinformatics / btz895.), fewer than half of the sequences failed to meet one or more developability criteria in the high naturalness domain (10th percentile) versus the low naturalness domain (90th percentile) (7.6% versus 17.8%).In another example, the second assessment may use sequences from a study of production titers in HEK-293 cell lines for clinical-stage antibodies (Jain, T., Sun, T., Durand, S., Hall, A., Houston, N.R., Nett, J.H., Sharkey, B., Bobrowicz, B., Caffry, I., Yu, Y., Cao, Y., Lynaugh, H., Brown, M., Baruah, H., Gray, L.T., Krauland, E.M., Xu, Y., Vasquez, M., and Wittrup, K.D. Biophysical properties of the clinical-stage antibody landscape. Proceedings of the National Academy of Sciences, 114(5):944-949, January 2017. doi:10.1073 / pnas.1616408114. URL https: / / doi.org / 10.1073 / pnas.1616408114). 10.1073 / pnas.1616408114.).
[0256] In that case, empirical evidence shows that sequences in high naturalness domains exhibited more than a three-fold increased production titer compared to low naturalness domains (180 mg / L vs. 130 mg / L). In yet another example, naturalness can be evaluated against reported immunogenicity measurements for 217 therapeutic antibodies compiled by Marks et al. (Marks, C., Hummer, A. M., Chin, M., and Deane, C. M. Humanization of antibodies using a machine learning approach on large-scale repertoire data. Bioinformatics, 37(22):4041-4047, June 2021. doi:10.1093 / bioinformatics / btab434. URL https: / / doi.org / 10.1093 / bioinformatics / btab434). Empirical evidence shows that the top quartile of native sequences was half as likely to be immunogenic compared with the bottom quartile (median ADA immunogenicity 2.6% vs. 5.4%).
[0257] Furthermore, this technology has demonstrated improved antibodies both through screening and through model-guided design. For example, an initial screen consisting of high-throughput enrichment followed by SPR identified several sequences with binding affinities greater than wild-type (n=87). This allows the SPR model to predict novel sequences in this range. This technology allows the identification of strong binders with K values exceeding those seen in the laboratory assays used to train them. D It may even be possible for values.
[0258] The present technology can include outcome testing by setting aside all data with measured affinity values higher than wild-type trastuzumab in a holdout set, then using the remaining data to train a model and predict the affinity of both the training set and the holdout set. In some embodiments, some models may be able to predict the exact K of the holdout data points. D In some cases, predictions cannot be made (because they were outside the distribution seen by the model). However, some models may place these points near the upper end of the predicted range, as shown in Figure 4Q, allowing for expansion of the predicted range using virtual screening of sequence space and additional laboratory experiments. In support of the practical value of our SPR model, we empirically tested a large number of sequences predicted to have greater than wild-type binding affinity. SPR screening confirmed that 76% of these sequences were greater than wild-type, and 94% of predictions were within 0.5-fold of their measured values.
[0259] Illustrative Results FIG. 5A depicts an exemplary AI-enhanced antibody optimization diagram 500, according to some embodiments. FIG. 5A shows that a deep learning model fed with qaACE and / or SPR measurements can quantitatively predict the affinity of novel sequence variants, thereby enabling the in silico design of antibodies with desired binding properties. A deep language model can predict the binding affinity of sequence variants. This technology postulates that artificial intelligence (AI) can learn the mapping between variants of biological sequences (e.g., antibodies) and quantitative readouts (e.g., binding affinity) from experimental data. With this capability, AI models can be used to simulate experiments in silico on novel sequence variants, thereby potentially accessing a larger sequence space and identifying more and better variants with desirable properties at a fraction of the time and cost, as depicted in FIG. 5A.
[0260] Training deep learning models generally requires large, high-quality datasets. To generate high-throughput measurements of antibody binding affinity, the present technology developed and incorporated the Activity-Specific Cell Enrichment (ACE) assay (or qaACE assay), a fluorescence-activated cell sorting (FACS) and next-generation sequencing (NGS) method for binning antibody variants based on affinity, as shown below in Figure 5B. In some aspects, the assay is an improved version of previous work [s20]. qaACE utilizes intracellular soluble overexpression of folded antibodies in the SoluPro™ E. coli B strain. Cells expressing antibody variants are fixed, permeabilized, and stained with fluorescently labeled antigen and scaffold probes, which allow for simultaneous differentiation of cells based on variant affinity and titer. The variant library is sorted and binned based on these signals. The collected DNA sequences are then amplified via PCR and sequenced.
[0261] The ACE score is calculated from sequencing read counts (Methods, see below). The ACE affinity score is proportional to binding affinity and is calculated using surface plasmon resonance (SPR) K D The results are highly correlated with measurements. To assess whether sequence affinity relationships can be modeled and predicted, the technique involves generating variants of the HER2-binding antibody trastuzumab in a fragment antigen-binding (Fab) format. Mutagenesis of CDRH2 and CDRH3 was prioritized because these regions house the highest density of paratope residues, both in general and for trastuzumab. Over the course of this study, up to five simultaneous amino acid substitutions were randomly introduced in up to two CDRs in the parent antibody, allowing for all naturally occurring amino acids except cysteine (which was removed to avoid potential disulfide bond association tendencies).
[0262] The following table summarizes the dataset used to train the model in this example: [Table 3]
[0263] The dataset table above depicts the characteristics of the trastuzumab variant dataset; specifically, the dataset used to train and evaluate the model. The positions hosting substitutions (IMGT numbering), the number of simultaneous substitutions (mutational burden), and the number of tolerated amino acids (all except cysteine) determine the combinatorial complexity of the sequence space. A subset of sequences was sampled from the combinatorial sequence space according to the indicated design strategy to construct libraries for screening by qaACE or SPR. The number of QC-passed amino acid sequence variants during screening and analysis is shown, resolved by mutational burden. *Random sampling of combinatorial space. **Uniform sampling by affinity from the trast-1 dataset. ***Random sampling of combinatorial space per mutational burden bin, with a defined abundance ratio of the mutational burden bin. Quadruple and quintuple mutants were only used to evaluate the performance of predictions from models trained with up to triple mutants.
[0264] In addition to high-throughput (HT) qaACE data, this technique also provides low-throughput, but highly accurate, SPR K data. D The readout was utilized to assess binding affinity. SPR was used (i) for targeted rescreening of sequence variants from the primary screen with ACE; and (ii) to validate model predictions.
[0265] As a proof-of-concept for this workflow, the technology generated a library containing all sequence variants with up to two mutations across eight positions of the trastuzumab CDRH3. Figure 6A depicts a diagram of this library. Figure 6A illustrates up to double mutants at eight positions of the trastuzumab CDRH3 that were screened using the combinatorial mutagenesis strategy:ACE on the trast-1 dataset.
[0266] Figure 6B depicts the predictive performance of a model trained on the qaACE scores of variants from 90% of the trust-1 dataset, evaluated for the remaining 10% of the sequences. Using the qaACE assay, the technology measured the binding affinity of 8,932 variants (97% of the combinatorial space) to create the trust-1 dataset in the table above. The technology trained a deep language model using 90% of the trust-1 dataset and evaluated model predictions using the remaining 10% of the holdout data. The measured and predicted qaACE scores for the holdout dataset were highly correlated, indicating that the language model was able to predict binding affinity with high accuracy, as shown in Figure 6B.
[0267] Figure 6C depicts a comparative analysis of replicate qaACE measurements and qaACE scores predicted from a model trained on individual qaACE replicates, according to some embodiments. Any deviation of the regression metrics from the theoretical optimum (1 for correlation and 0 for RMSE) is contributed by both inaccuracies in prediction and inaccuracies in measurement. To disentangle these two effects, the present technique may consider agreement between measurement replicates, e.g., as shown in Figure 6D, using the same metrics previously used to evaluate the predictive performance of the models disclosed herein. In particular, evaluating model performance relative to agreement between measurement replicates indicated that the majority of the error between prediction and measurement can be attributed to experimental noise, as shown in Figures 6C and 6E.
[0268] The holdout set evaluated in Figures 6A-B was randomly drawn from the trast-1 dataset. Thus, the training and holdout sets had similar distributions of qaACE scores, with the presence of low-affinity binders due to the deleterious effects of most mutations. This design of the training and holdout sets addressed the question of whether the model could simulate in silico experiments. However, a more challenging test would require evaluating predictions using a holdout set that was uniformly distributed with respect to binding affinity. This holdout set would be enriched for stronger binders compared to the training set. To reduce the presence of weak binders in our new holdout set, we sampled >200 sequences from the trast-1 dataset. The sampled sequences were rescreened by SPR to generate the trast-2 dataset shown in the dataset table above.
[0269] FIG. 6F shows a graph of qaACE affinity scores versus log-transformed SPR K D The correlation between the measurements is depicted. In particular, Figure 6F includes a plot showing the qaACE scores from trast-1 for sequence variants crossed with trast-2. Empirical testing revealed that the qaACE scores of trast-2 sequences correlated well with the SPR-derived log-1 as shown in Figure 6F. 10 K D Strong agreement between the values was observed, confirming the near-uniform distribution of this dataset.
[0270] Figure 6G depicts predictive performance for a uniformly distributed holdout set for binding affinity (ACE scores from trast-1 for the sequences shown in panel Figure 6F), according to some embodiments. As shown in Figure 6G, the present technology can include using trast-2 sequences as a holdout set for models trained on trast-1 qaACE scores, which confirmed strong predictive performance.
[0271] In general, the present technology demonstrates that a deep language model trained on the SPR-generated trast-2 dataset can quantitatively predict antibody binding affinity, as shown in Figures 7A-7D, where performance is assessed by pooled 10-fold cross-validation. In particular, Figure 7A shows the SPR measurements -log 10 K D Figure 7B depicts the predicted log-log replicates predicted from a model trained on individual SPR replicates; Figure 7B depicts the predicted log-log replicates predicted from a model trained on individual SPR replicates. 10 K D Measurements and -log 10 K D Figure 7C depicts a comparative analysis of log 10 k on and Figure 7D depicts predictions from a model trained on -log 10 k off Render predictions from a model trained on values.
[0272] Since SPR measurements are collected in some embodiments, such as using the trust-2 dataset depicted in Figure 6F, this technique can be used to investigate whether this dataset alone is sufficient to train a deep language model and directly predict the binding coefficient. Due to the relatively small size of the dataset (n = 215), all models can be trained using 10-fold cross-validation, and model performance was evaluated using pooled out-of-fold prediction. For example, this technique can be used to calculate the -log 10 K D The trast-2 data set was used to initially train a model that predicts the binding affinities of trast-2 and found that the correlation between measured and predicted values was slightly lower than that observed in the high-throughput trast-1 dataset, as shown in Figure 7A. However, 87% of the predicted binding affinities deviated from their respective measured values by less than half a log. Similar to trast-1, the techniques of the present invention can include evaluating the results of trast-2 relative to the best possible performance, defined as the degree of agreement between measurement replicates (e.g., as shown in Figure 7B and Figures 7E and 7F).
[0273] In addition to the equilibrium binding constant, SPR also determines the association coefficient (k on ) and dissociation coefficient (k off ) provides the measured log(s) between replicates in the trust-2 SPR dataset. Models trained to predict these coefficients also performed successfully, as shown in Figures 7C and 7D; and Figures 7H-7J and 7K-7M, opening up the potential for this AI-based technique to aid in the specific manipulation of association and dissociation properties in addition to overall binding affinity. Figures 7H-7J show the measured log(s) between replicates in the trust-2 SPR dataset. 10 k on Value and predicted log 10 k on Figures 7K-7M depict the comparison of log-log measurements between replicates in the trust-2 SPR dataset. 10 k off Values and predictions - log 10 k off Depicts a comparison of values. on Note that the lower correlation observed for k is due to the small range of variation observed. Consistently, the agreement of measurement replicates also off About k on is lower for , which further highlights the need to consider measurement noise when assessing prediction performance.
[0274] Finally, the present technique allows one to determine whether a model trained simultaneously on two affinity data types can improve the performance of a model fed only a single data type. To this end, for example, the present technique can be used to train a model using both qaACE (trust-1) data and SPR (trust-2) data in a multi-task setting, resulting in a -log 10 K DThe predicted values were used. Empirical testing found that this model performed slightly better than a model trained on trast-2 SPR data alone, as shown in Figure 7G. Figure 7G depicts a model trained using SPR data from trast-2, or a model co-trained using both ACE (trast-1) and SPR (trast-2) data; the models were evaluated using 10-fold cross-validation, comparing the -log of sequences in each out-of-fold validation set. 10 K D and / or combine predictions across folds to obtain SPR measurements -log 10 K D compared with the value.
[0275] In some embodiments, all models trained on the trust-1 and trust-2 datasets are deep language models pre-trained on immunoglobulin sequences from the OAS database (Methods, see below). In embodiments, these models can be compared against baselines using either a 90:10 train:holdout split from the trust-1 dataset or pooled 10-fold cross-validation from the trust-2 dataset. For example, the technique can first be used to train a deep language model with the same architecture but without pre-training (i.e., randomly initialized weights) to assess the impact of transfer learning. Second, an XGBoost model is trained to determine whether the deep language model enhanced predictive accuracy compared to "shallow" machine learning. In some embodiments, the pre-trained model outperformed both baselines147 for both the trust-1 and trust-2 datasets, as shown in Figures 7N and 7O, with a stronger gain seen for the smaller trust-2 dataset, consistent with previous observations
[23] .
[0276] For example, Figures 7N and 7O depict a comparison of pre-trained language model performance against baselines, comparing the predictive performance of two OAS pre-trained deep language models against two baselines: (1) a deep language model of the same architecture but with randomly initialized weights, and (2) an XGBoost model. In some embodiments, the models were trained against a 90:10 train:holdout split of ACE scores from the trust-1 dataset or a -log 2 split of ACE scores from the trust-2 dataset. 10 K D The training and evaluation were performed using 10-fold cross-validation with the values.
[0277] To understand why pretraining improves model performance, we can use this technique to examine model embeddings from all combinations of pretraining vs. no pretraining and fine-tuning vs. no fine-tuning. Even without fine-tuning, the embeddings from OAS pretraining appear to have a structure with distinct patches enriched for high (or low) binding affinity. This organization simplifies subsequent fine-tuning using binding data, making it easier to update the model weights and provide enhanced binding affinity predictions, as shown in Figures 7P-7S, which generally depict the structure of model embeddings for binding affinity. The embeddings were calculated using a forward pass using sequences from the trust-2 dataset, reduced to two dimensions using UMAP, and graphically encoded by measuring binding affinity.
[0278] Improved antibody variants discovered through model-guided design As discussed above, this technology uses holdout set and cross-validation to demonstrate AI prediction performance.This technology further demonstrates that model is used to design a series of sequences with desired binding properties, and then uses dedicated SPR experiments to validate.For example, in one embodiment, this technology can use the design sequence that spans two orders of magnitude of equilibrium dissociation constant (referred to herein as design set A) to train the model that is trained on the trast-2 data set.
[0279] Generally, model-enabled design involves exhaustive prediction in combinatorial sequence space, followed by sampling sequences with predicted binding affinities that match the requirements. Figures 8A-8D depict a deep language model trained on an SPR-generated trast-2 dataset that can design unseen sequence variants for validation in independent SPR experiments, according to some embodiments. Figures 8E-8H depict a deep language model trained with an SPR-generated trast-2 dataset that can design unseen sequence variants, such as those shown in Figures 8A-8D, for validation in independent SPR experiments. Empirically, this technique found excellent agreement between prediction and validation for design set A (as shown in Figures 8A, 8E, and 8F).
[0280] This technique can be further used to validate more challenging designs, seeking variants with stronger binding than trastuzumab (referred to herein as Design Set B). For previous designs, for example, 50 sequences were validated by SPR using this technique, and 74% of the variants were found to be tighter binders than the parent antibody (as shown in Figures 8B-C and in the table below), and 100% met the design specifications within less than 0.5 log tolerance (as shown in Figures 8C and 8G). [Table 4-1] [Table 4-2]
[0281] The Design B Sequence Variants Table above lists the 50 designed and validated trastuzumab sequence variants from Design Set B. For reference, see Parent Trastuzumab Prediction and Validation - log 10 K D were 8.3M and 8.25M, respectively.
[0282] This performance is competitive when compared to replicate measurements; if we selected tighter binders based on experimental measurements from the first replicate, a similar fraction would have lower binding affinity in the second replicate (Figure 8H).
[0283] This design's small -log 10 K D Due to the range, the correlation between prediction and measurement was low, as shown in Figure 8G. However, the above k on As previously observed in modeling, and as depicted in Figures 7H-7J and 8H, even measurement replicates correlate poorly with each other whenever the affinity range is narrow. In contrast to correlation, other experimental metrics, such as RMSE and the percentage of predictions that deviate from the measurements by less than 0.5 logs, remained consistent with previously observed performance, as shown in Figure 8G. In some embodiments, these metrics are generally more informative when working with a set of sequences packed in a narrow affinity range.
[0284] The validation rate for Design Set B compared very well to a naive approach to library screening, in which the proportion of tighter binders than trastuzumab was minimal, as shown in Figure 8I. Specifically, Figure 8I depicts a chart showing that model predictions strongly enrich for variants with desirable binding properties compared to naive library screening. In Figure 8I, sequences of interest in Design Set B can be defined as antibody variants that bind more tightly to HER2 than parent trastuzumab (i.e., top binders). Figure 8C depicts the presence of top binders in the combinatorial space (i.e., laboratory-only screening) estimated by the validation rate of top binders in Design Set B (i.e., AI-assisted screening) versus the proportion of top binders in model predictions for the entire combinatorial space, adjusted for the validation rate of Design Set B.
[0285] The strong enrichment provided by model predictions for variants of interest is a key finding that enables in silico experiments and AI-assisted antibody optimization, as shown in Figure 5.
[0286] As noted, the model used to design sequence set B may be trained on the trast-2 dataset, which includes several binders more potent than trastuzumab (see Figure 7A). Thus, in one example, the technique may involve determining whether a model that has never been fed any sequences as extreme (affinity-wise) as those it is tasked with designing can still prioritize top binders. This question has practical value, since it is conceivable that some applications may ultimately be faced with a large sequence space and a low abundance of positives, which would likely result in a training set lacking positive examples.
[0287] To test the performance of the models described herein in an out-of-distribution affinity prediction setting, some examples may include removing any binders tighter than trastuzumab from the trast-2 training set and then training a model using the remaining data to predict the affinity of Design Set B. Experimentally, as expected, models trained in this manner predicted accurate K for Design B. D The model was no longer able to make predictions. Notably, however, the model was still able to place the binding affinity of Design B variants at the top of their known distribution, as shown in Figure 8D. This result demonstrates that some AI-based aspects of the technology may still enable prioritization of top binders by sampling top-ranked predictions, even when the laboratory experiments generating the training data did not observe the full affinity range.
[0288] AI prediction performance is maintained when scaling to a larger sequence space Figures 9A-9D depict how high-throughput binding scores from the ACE-generated trast-3 dataset can expand predictive capabilities to a larger mutation space, according to some embodiments. In another example, to test model performance in a larger sequence space, the technology can be used to perform combinatorial mutagenesis of up to three mutations across more than 10 amino acids each in CDRH2 and CDRH3. This example involved constructing a library by sampling less than 1% of this sequence space and measuring the binding affinity of the sampled sequence variants using the qaACE assay, as shown in the trast-3 dataset table above and in Figure 9A. This example may further include training a model using 80% of the trast-3 data and evaluating its performance on the remaining 20% of holdout sequences. Experimentally, model performance was comparable to a double-mutant library (Figure 9B).
[0289] As a negative control, a model trained on a dataset with randomly shuffled qaACE scores was confirmed to have no predictive power, as shown in Figure 9E. Specifically, a model trained in this manner cannot predict ACE scores from randomly shuffled data. Because the trust-3 sequence space was very large, all models were trained and evaluated only on qaACE data. This is because the correlation between qaACE scores and -log 10 K D This is because a high correlation between the values had already been demonstrated (see Figure 6F).
[0290] Given the accuracy of the trast-3 model's predictions for variants with up to three mutations from the original trastuzumab sequence, experiments were also performed to test whether the model could accurately predict the qaACE scores of variants with four or five mutations from trastuzumab (Figure 9C).
[0291] For example, Figure 9F depicts predictive performance for a set of quadruple mutants, according to some embodiments. Figure 9G depicts predictive performance for a set of quintuple mutants, according to some embodiments. In Figures 9F and 9G, models may be trained on the trast-3 dataset of up to triple mutants of the parent trastuzumab sequence and evaluated on holdout sets of, for example, only quadruple or quintuple mutants, respectively.
[0292] As shown, in some embodiments, the model can predict qaACE scores for quadruple mutants with slightly less accuracy than for triple mutants, as shown in Figure 9F. Although the prediction accuracy for quintuple mutants was much lower, as shown, for example, in Figure 9G, the model could still be used to distinguish between high and low binders in a classification setting. These results indicate that the triple mutant model can be extrapolated to quantitatively predict binding scores for up to four simultaneous mutations from the original sequence and qualitatively predict binding scores for five mutations. As the mutation distance increased, the accuracy of the model deteriorated, especially for variants with low affinity, as depicted in Figures 9F and 9G.
[0293] Deep language models are highly sample-efficient In general, the predictive ability of any deep learning model is highly dependent on the quality and quantity of its training data. The trast-3 dataset contains binding affinities for approximately 50,000 unique antibody sequences, covering 0.7% of the complete combinatorial mutation space for this design, as shown in the dataset table above. To determine the relationship between model performance and the quality and quantity of the training dataset, some examples may involve training a cohort of models to predict affinities from a range of dataset sizes sampled from datasets of varying fidelity, as shown in Figure 9D. In some examples, the original trast-3 dataset may be treated as a high-fidelity dataset, and low-fidelity datasets may be generated by isolating a single DNA variant for each sequence from a single FACS sort (Methods, see below). In this example, the size of the training subset may range from 44,165 sequences (total 233 training datasets) to 350 sequences (1 / 128 of the total training dataset), and the models may be evaluated on a common holdout validation dataset containing 10% of all sequences in the high-fidelity dataset. At each training subset size, the performance of four models may be compared: (1) an OAS pre-trained model trained on a subset from the high-fidelity dataset; (2) an OAS pre-trained model trained on a subset from the low-fidelity dataset; (3) a randomly initialized model trained on a subset from the high-fidelity dataset; and (4) a randomly initialized model trained on a subset from the low-fidelity dataset.
[0294] As the size of the training dataset decreased, model performance deteriorated: models trained on low-fidelity data consistently performed worse than their counterparts trained on high-fidelity data, highlighting the importance of high-quality experimental assays.
[0295] Pre-training models on the OAS dataset typically improved model performance; however, the performance gains from pre-training decreased when the models were trained using either a smaller, high-fidelity dataset or a larger, low-fidelity dataset.
[0296] Given that the model requires at least 2,760 sequences to maintain Pearson R correlation above 0.8, it is not practical to model this mutation space using only SPR training data; a higher throughput assay, such as qaACE, is required.Since Pearson R correlation remains above 0.8 for all high-fidelity training subsets, covering at least 0.4% of potential search space, the model learns to predict approximately 2,500 sequences for all sequences in training dataset.Therefore, the deep language model of this technology can expand the search space of experimental datasets by at least an order of magnitude.
[0297] Deep language models provide interpretable analysis of antibody binding landscapes Once trained, the AI model can be used as an oracle to predict binding affinity scores for all sequences in the combinatorial space that match the design of the training set. Fast and accurate predictions can inform how antibodies can be affected by different engineering strategies and help guide experimental efforts.
[0298] To gain insight into the mutational landscape of trastuzumab, some examples may include using this technology to thoroughly evaluate the effect of all single, double, and triple mutations in CDRH2 and CDRH3. Trastuzumab has a high binding affinity for its target antigen, HER2 (-log β of 8.25 M in Fab format), as shown in Figure 7A. 10 K DThus, most mutations were predicted to have a detrimental effect on binding affinity, as shown in Figure 5. When multiple mutations were considered, empirical tests suggested that most combinations were predicted to have a detrimental effect on binding affinity, as shown in Figures 9H and 9I.
[0299] In particular, Figures 9H and 9I show that positions 55, 107, 111, 112, and 113 were predicted to have deleterious effects when mutated and tended to interact epistatically with other mutations, as discussed later. This is consistent with previous alanine scanning and structural studies, which pointed to strong contributions to binding affinity from these residues
[22] . For reference, the effect of each individual mutation on trastuzumab is shown with identical dots in both Figures 9H and 9I. Mutations at each position include all possible substitutions with natural amino acids, excluding cysteine, sorted alphabetically (i.e., X ∈ [A, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y]).
[0300] Analyzing the incremental effects of mutations across variants showed that positions 59, 62, and 110 were relatively tolerant to mutations, as depicted in Figures 9H and 9I. This suggested that they make a relatively small contribution to binding affinity and may be ideal candidates for optimization for other antibody properties. Some single mutations in CDRH2, such as Y57D / E, N62E, or T65D / E, were predicted to increase binding affinity (see Figure 5). In addition to single mutations, combining multiple mutations may also provide improved high-affinity variants. Indeed, as mutation burden increased, the number of predicted high-affinity sequences increased, but their proportion decreased. For example, two single mutants (0.56%), 192 double mutants (0.31%), and 7,063 triple mutants (0.11%) predicted qaACE scores higher than trastuzumab in the trast-3 dataset.
[0301] This AI-based technique allows for the identification of diverse clusters of high-affinity variants of trastuzumab. In one example, this technique was used to perform clustering analysis of model-derived embeddings of high-affinity sequences (predicted qaACE score > 8.0). For example, Figure 9J depicts a sequence logo plot illustrating the composition of high-affinity clusters of embeddings (predicted ACE score > 8.0). Clusters can be generated by reducing the dimensionality of the embeddings, followed by HDBSCAN clustering and filtering by the average predicted ACE score. A minimum number of sequences per cluster can be required (e.g., 40). The logo plot in Figure 9J shows the relative frequency of each specific substitution in the sequences within each cluster. In some embodiments, binding affinity predictions were generated from a model trained on the trast-3 dataset.
[0302] While the triple mutant space provided many potential high-affinity candidate sequences, these tended to form compact clusters containing specific substitutions at a few positions, as shown in Figure 9J. Notably, the mutation Y57D / E was observed in several clusters. Also, while the majority of high-affinity triple mutants had two or three mutations in CDRH2 (particularly at positions 57 and 62 or adjacent positions), fewer solutions contained one mutation in CDRH2 and two mutations in CDRH3. This finding also highlights the important role of the CDRH2 region in antigen binding by trastuzumab, as described by others [22, 24].
[0303] Empirical tests have also demonstrated that the effect of a given mutation on binding affinity varies widely with the presence of other mutations in the sequence, a phenomenon known as contingency
[25] .
[0304] Figures 9H and 9I depict that a given mutation can have a greater, lesser, or even opposite effect compared to the effect it would have on the parent trastuzumab sequence depending on the presence of only another single mutation. Furthermore, in the presence of two mutations, the range of possible effects for an additional (third) mutation becomes wider.
[0305] Similarly, epistasis refers to the deviation from additivity of the effects of two coexisting mutations compared to their individual effects.
[26] The epistatic interactions between mutations for all double mutants of trastuzumab are depicted in Figure 9K.
[0306] Specifically, Figure 9K depicts antagonistic epistasis commonly found between key paratope residues in trastuzumab. The heat map in Figure 9K depicts the epistatic effect across all possible substitution pairs. Epistasis refers to deviations from additivity in the effects of two mutations when both are present. Antagonistic epistasis refers to a smaller-than-expected change in binding affinity when two mutations co-occur. Mutations at each position include all possible substitutions with natural amino acids, except for cysteine, sorted alphabetically (i.e., X ∈ [A, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y]), according to some embodiments.
[0307] Given the negative effect that many mutations have on binding affinity, antagonistic, positive epistasis is often observed (i.e., double mutants exhibit higher binding affinity than would be expected based on their constituent single mutations). This is particularly evident in pairs of mutations involving positions 55, 107, 111, 112, and 113, which are critical to the binding affinity of trastuzumab
[22] . Epistatic interactions are also highly contingent on the presence of other mutations in the sequence. The complex interactions between mutations directly affect the biochemical properties of antibodies.
[0308] Taken together, the diversity of high-affinity sequences and their rarefaction as a function of mutational load highlights the value of thoroughly evaluating the space of possible variants. Such large-scale evaluations are only feasible with the aid of computational models. The models discussed herein, and their results, are in excellent agreement with previous functional and structural studies and can provide unique insights into how mutations interact to shape antibody binding affinity. The pervasiveness of epistasis effects also highlights the need for AI models to accurately predict and guide antibody optimization.
[0309] 10A and 10B depict comprehensive sequence affinity mapping of trastuzumab variants according to embodiments of the present technology.
[0310] AI demonstrates robust predictive performance in a second case study involving simultaneous binding predictions for three antigen variants The modeling approach discussed above, established for trastuzumab, can be easily extended to other antibodies. In one example, public binding data for variants of the broadly neutralizing (bn) Ab CR9114 can be used (see Supplementary Information, below)
[27] . Because the CR9114 dataset provides binding data for three different influenza subtypes of the target antigen hemagglutinin (HA), this technique can be used to extend the model to support multitask affinity prediction for multiple targets simultaneously. Furthermore, this technique can be used to explore the model's ability to combine classification and regression in a single mixture model, since many of the CR9114 variants lose binding to one or more HA subtypes. Furthermore, this technique can be used to assess the impact of training set size on model performance.
[0311] [Table 5]
[0312] The performance table above depicts the joint model affinity prediction performance of CR9114 against multiple influenza strains of the hemagglutinin (HA) antigen. For each training set size (10%, 1%, and 0.1% of 65,091), four models were trained (Reg: regression-only model; Mix: mixture classification / regression model; PT: initialized with pre-trained OAS model weights; NPT: initialized with random weights). Results are shown for these models using pooled CV. The complete CR9114 dataset contains 63,419 (97%) H1, 7,174 (11%) H3, and 198 (0.3%) FluB positive binders.
[0313] The above table and Figures 10C-10H together demonstrate that a single model can be trained using this technique to jointly predict the affinity of a given antibody sequence for multiple different antigen targets.
[0314] 10C depicts the regression performance of models trained on 10% of the CR9114 dataset, according to some embodiments. Figure 10C includes results for a regression-only model 1032, results for a mixture classification / regression model 1034, results for a model initialized with pre-trained OAS model weights 1036, and results initialized with random weights 1038.
[0315] 10D depicts the regression performance of a model trained on 1% of the CR9114 dataset, according to some embodiments. Figure 10D includes results for a regression-only model 1042, results for a mixture classification / regression model 1044, results for a model initialized with pre-trained OAS model weights 1046, and results initialized with random weights 1048.
[0316] 10E depicts the regression performance of models trained on 0.1% of the CR9114 dataset, according to some embodiments. Figure 10E includes results for a regression-only model 1052, results for a mixture classification / regression model 1054, results for a model initialized with pre-trained OAS model weights 1056, and results initialized with random weights 1058.
[0317] Figures 10C-10E depict the results for a model using pooled CV for positive binders only (-log 10 K D >B c , in the formula, B c are the lower boundaries for each target determined in the original publication; 7 for H1, and 6 for H3 and FluB. The complete CR 9114 dataset contains 63,419 (97%) H1; 7,174 (11%) H3; and 198 (0.3%) FluB positive binders.
[0318] 10F depicts the regression performance of a mixture model trained with 10% of CR9114, according to some embodiments. FIG. 10F includes results 1062 for a model initialized with pre-trained OAS model weights and results 1064 for a model initialized with random weights.
[0319] 10G depicts the regression performance of a mixture model trained with 1% of CR9114, according to some embodiments. FIG. 10G includes results 1072 for a model initialized with pre-trained OAS model weights and results 1074 for a model initialized with random weights.
[0320] 10H depicts the regression performance of a mixture model trained with 0.1% of CR9114, according to some embodiments. FIG. 10H includes results 1082 for a model initialized with pre-trained OAS model weights and results 1084 for a model initialized with random weights.
[0321] Figures 10F-10H depict the results for these models using pooled CV. For each model and target, Figures 10F-10H each depict the respective precision-recall curve plot and calibration curve (true probability vs. predicted probability at different scoring bins).
[0322] As expected, the predictive power of such models was lower for the FluB target compared to H1 and H3 because the complete dataset contained only 193 positive FluB binders. This left only 19 positive examples when using the 10% training set and one to two positive examples in the 1% and 0.1% training sets (a minimum of one positive and one negative example for each target was required when selecting cross-validation folds; see Supplementary Information, below). Nevertheless, even with only 19 training examples, 88% of the model's predictions for FluB were within 0.5 logarithms of their measurements when using initial weights pretrained on the OAS dataset, compared to only 73% when using random initial weights. Using pretrained weights improved performance in all cases where the number of training examples was below 1,000.
[0323] The mixture model performed well on classification tasks without a significant loss of performance on regression tasks compared to regression-only models. The balanced accuracy of the model's predictions was above 0.84 in all cases where the training set contained at least 7 positive examples and 7 negative examples, but even a training set of only 65 variants (7 positive variants and 58 negative variants on average) achieved a balanced accuracy score of 0.91 on the H3 binding task.
[0324] Naturalness In general, the development of candidate antibodies into therapeutic drugs is a complex process involving high preclinical and clinical risks. This risk often stems from numerous challenges related to production, formulation, efficacy, and adverse reactions. Modeling these risks has been a significant challenge for the industry due to the difficulty in obtaining information and relevant data, especially at large scale.
[0325] Figure 11 depicts the relationship between antibody naturalness, immunogenicity, developability, and other properties. In some instances, it has been hypothesized that using this technology to learn sequence patterns across natural antibodies from different species may be useful for identifying and prioritizing "human-like" antibody variants, disregarding unnatural sequences, and ultimately mitigating drug development risks, as depicted in Figure 11. To this end, in one example, antibody sequences may be evaluated for their naturalness score using a language model pre-trained with OAS (see Materials and Methods, below). In this context, naturalness is a score calculated by the pre-trained language model that measures how likely a given antibody sequence is to originate from the organism of interest. In this way, naturalness can be used as a metric for antibody design and engineering.
[0326] In Figures 11B-11D, the four bins (low, low-medium, medium-high, and high) can correspond to dividing the naturalness range into four equally sized sections (Figures 11G-11J). P values can be calculated using the Jonckheere-Terpstra test for trend. While the datasets in Figures 11B and 11D were scored for both chains, the datasets in panels 11C and 11E included only heavy chain variants and were consistently scored by the heavy chain model alone. To determine the usefulness of the naturalness scores, their association with four antibody properties was evaluated. For example, immunogenicity properties are reflected in Figure 11B, which were aggregated across multiple studies of clinical-stage antibodies by Marks et al.
[28] . A potential confounding factor in naturalness-immunogenicity association analyses is that some antibodies are fully human in origin, while others are humanized, chimeric, or murine.
[0327] Scoring antibodies of different origins by naturalness would amount to binning them primarily by species, which is trivial and uninformative. In contrast, scoring antibodies belonging to the same class would amount to ranking them purely from most natural to least natural. Only two antibody classes in Marks et al. are large enough to support statistical analysis: human antibodies and humanized antibodies. In one example, the latter were investigated using this technology because of their greater immunogenic potential, thereby providing an ideal case study.
[0328] A scatterplot between the proportion of patients positive for anti-drug antibodies (ADA) and naturalness reveals a weak, non-significant correlation, as shown in Figure 11F. However, closer inspection of the ADA responses indicates that the majority of data points fall within the 0-10% range, with a few outliers above 20%. Empirical testing and reasoning suggest that such outliers, if present, may obscure the relationship between naturalness and immunogenicity. To mitigate the influence of outliers, we used this technique to bin the naturalness scores and calculate the median ADA response per naturalness bin, as shown, for example, in Figure 11G. As in Figures 11B-11D, the ranges in Figures 11G-11J (low, low-medium, medium-high, high) may correspond to dividing the naturalness range into four equally sized sections.
[0329] This analysis revealed that the most natural antibodies elicited lower median ADA responses than the least natural antibodies, as depicted in FIG. 11B.
[0330] The second antibody characteristic considered using this technology is developability, which can be estimated with the Therapeutic Antibody Profiler (TAP)
[31] . Using this technology, we calculated naturalness scores (using the modeling techniques discussed herein) and developability scores (using TAP) for heavy chain sequences from a high-diversity phage display dataset
[29] ("Gifford Library," Figure 11H) and a low-diversity trastuzumab variant (Figure 11I). For both, empirical review found a strong correlation between naturalness and TAP-derived developability, as shown in Figures 11C and 11K. This result is particularly striking because naturalness scores were obtained during training using exclusively example natural antibody sequences, while TAP was calibrated using distributions of five metrics collected from therapeutic antibodies
[31] . Thus, the correlation between naturalness and TAP-assessed developability suggests that developable antibodies are enriched for human-like antibodies. FIG. 11K depicts the association between naturalness and failure of developability of trastuzumab variants relative to FIG. 11I.
[0331] The third property investigated using this technique was expression level in mammalian (HEK-293) cells, which has been reported for clinical-stage antibodies by Jain et al.
[30] . Regarding immunogenicity, the dataset included several classes of antibodies, and using this technique, we refocused on humanized antibodies. Empirical results suggested that highly natural antibodies performed better than antibodies scored less favorably by the modeling technique used herein, as depicted in Figures 7D and 11J.
[0332] Finally, the fourth property considered using this technology is mutational burden, which measures the number of amino acid substitutions between the parent antibody sequence and the variant.
[0333] For example, this technology was used to calculate naturalness scores for 6,710,400 single-, double-, and triple-mutated trastuzumab variants, as depicted in Figure 11E. Empirical testing of the results found that naturalness was negatively associated with mutation load. This finding is consistent with the general concept that most mutations have adverse effects and emphasizes the need to actively optimize naturalness along affinity, because introducing mutations into a parent antibody without considering naturalness is likely to degrade naturalness.
[0334] Generation of sequence variants with desired properties Generally, antibody optimization can be performed to a limited extent for individual properties using a number of established laboratory approaches. For example, deep mutational scanning has been used to improve the binding affinity of antibody candidates [5]. However, the large mutation space cannot be exhaustively screened by these methods, limiting the scope of potential improvements. Library screening methods, such as phage display, can overcome this obstacle; however, selecting for a single property (e.g., binding) at a time can result in the unintended degradation of other properties of interest. For example, using this technique, we showed that increasing the mutation load resulted in a lower median naturalness, as depicted in Figure 11D.
[0335] To further demonstrate this, we used this technology to exhaustively predict the qaACE scores and naturalness of all variants with up to three mutations from trastuzumab. Of the 6.7 million variants, only 46,931 (0.7%) had predicted qaACE scores higher than trastuzumab, as shown in Figure 11L. Even if we could experimentally screen for all variants with improved affinity, only 4,003 (8.5%) of these variants had naturalness scores equivalent to or higher than trastuzumab, as shown in Figure 11M. The dashed lines in Figures 11L and 11M indicate the naturalness and / or predicted ACE scores of trastuzumab in a density map of the fitness landscape for the exhaustive trast-3 search space.
[0336] Random screening of this space using the approximately 50,000 member trust-3 library yielded only 60 variants with higher qaACE scores and naturalness.
[0337] In silico screening provides a way to address this problem by optimizing for multiple properties simultaneously with a design objective function. This technique, which can involve a genetic algorithm (GA) on top of the affinity and naturalness model oracle, has been shown to significantly improve the throughput of the in silico screening process.
[0338] Figure 12 generally depicts how genetic algorithms can efficiently maximize, minimize, or target a specific qaACE score while maximizing naturalness. As an example, the present technology may be used to minimize, maximize, or target a specific qaACE score in a search space of over 6.7 million 419 sequence variants, as depicted in Figure 12A, while simultaneously maximizing naturalness, as shown in Figure 12B.
[0339] In this example, after 20 generations, the GA performed almost as well as an exhaustive search of mutation space, as shown in Figure 12C: 85 of the top 100 variants identified by the GA were among the top 100 variants overall. Additionally, all of the top variants identified by the GA were within 5% of the maximum achievable qaACE score (9, resulting from nine sorting gates) and had a naturalness score 424 higher than trastuzumab. At baseline, this technology can be used to perform a random search by querying the same number of sequences as the GA. In this example, the search was able to find only two sequences with a qaACE score and naturalness higher than trastuzumab, as depicted in Figure 12C.
[0340] Unlike an exhaustive search of mutation space, GA-driven optimization is highly efficient. For example, at each generation, the GA samples 200 new variants, resulting in only 4,000 total sequences sampled across all 20 generations. Moreover, more than half of the top 100 individuals were selected by the GA in the first 12 generations, as depicted in Figure 12D. Collectively, these results demonstrate that a genetic algorithm built on top of predictive models for binding affinity and naturalness can rapidly and efficiently identify a set of top candidates for downstream development. The value of optimization techniques coupled with AI oracles will increase as in silico design is applied to increasingly large combinatorial sequence spaces.
[0341] Consideration In general, deep learning methods have demonstrated rapid progress in modeling proteins, including sequence, structure, and function. Similarly, protein interactions are receiving increasing attention for therapeutic design purposes. A key limitation for many of these efforts is the ability to synthesize large libraries of proteins and evaluate their quantitative attributes. Here, we demonstrate that the qaACE assay discussed herein is a powerful complement to deep learning models, providing the throughput and fidelity to accurately model antibody binding affinities with up to four or five mutations in two CDRs (combinatorial space: 108–1010) from a single experiment. The qaACE assay offers advantages over existing methods for large-scale antibody variant matching, such as Tite-Seq
[32] , SORTCERY
[33] , and phage display
[34] . First, qaACE utilizes the SoluPro™ E. coli B strain to lytically express antibodies intracellularly, avoiding binding artifacts associated with surface display formats. Additionally, qaACE leverages the genetic tools available for E. coli, allowing for faster library generation cycles and increased transformation efficiency compared to other model organisms. Finally, the qaACE assay is a true screening method in which all variants are measured regardless of affinity strength, in contrast to selection, such as phage display, in which only high-affinity binders are preferentially isolated.
[0342] The predictive capabilities of the deep learning models demonstrated here are made possible by the quantitative capabilities of the improved qaACE assay, which offers two distinct advantages from a modeling perspective. First, the expanded ability of models trained with quantitative data for overall increased performance and quantitative prediction, which is particularly useful when the goal is to adjust binding affinity rather than simply maximizing it. Second, quantitative training data also enables intelligent selection of sequences for downstream quantification using low-throughput gold-standard SPR assays. Random sequence space is vast and heavily biased toward deleterious mutations. A common approach to this problem is to bias mutation libraries toward specific positions or critical mutations, but the strength of the epistatic effects identified by these models suggests that these approaches provide insufficient coverage. Our pre-quantification step using the improved qaACE assay provided us with the opportunity to measure a more uniform distribution without bias, thereby increasing the generalization ability of the model.
[0343] The diversity of available high-affinity sequences and their dilution with mutational burden highlights the value of thoroughly and accurately evaluating the space of possible variants. Such large-scale evaluation is only feasible with the aid of computational modeling. Modeling results are in excellent agreement with previous functional and structural studies and can provide unique insights into how different mutations interact to shape an antibody's binding affinity landscape. The pervasiveness of epistasis effects also highlights the need for highly flexible AI models to accurately predict and guide antibody optimization.
[0344] Only a small percentage of sequences within the vast combinatorial antibody sequence space are found in nature—over 10120 possible unique CDR sequences for the longest reported human sequence, compared to 108 in the OAS database. The naturalness model presented herein helps determine whether a novel sequence belongs to this category, and the technology may be used to roughly estimate the size of this natural space as 1060, as shown in Figures 18A and 18B.
[0345] Figures 18A and 18B depict plots and diagrams related to estimating sequence space size for heavy chain human CDRs. For example, in Figure 18A, one million random sequences may be generated using random amino acids that match the length distribution of the OAS. Based on the lower tail of OA naturality, a threshold of 0.15 may be chosen to estimate the size of the natural sequence space. In Figure 18B, the circles are not to scale. The size of the total possible sequence space may be estimated from 20 amino acid possibilities across 61 positions (the longest human sequence in the filtered OAS dataset). The natural CDR space may be roughly approximated by fitting a skew-normal distribution to the random sequences and calculating the fraction exceeding the naturalness threshold.
[0346] While one can account for statistical uncertainties in this calculation, it is clear that the natural space is much larger than one could possibly screen in the laboratory or in silico, and at the same time, these natural sequences are rare enough to be ignored in screening random sequence variants.
[0347] The solution presented here is to apply a model trained on both naturality and affinity data targeted to a particular antibody, the intersection of which effectively allows for evaluation of a larger sequence space than can be physically evaluated, while also focusing screening on the most relevant "natural" sequences.
[0348] This co-optimization of two antibody properties can be extended to the co-optimization of n antibody properties. Training models on multiple affinity datasets unlocks binding predictions for multiple antigens or antigen variants, as shown here for CR9114. In principle, multi-antigen prediction could facilitate engineering breadth (co-optimization for antigen escape variants), specificity (co-optimization to increase binding to desired members of the same family while decreasing binding to undesired members of a protein family), and species cross-reactivity (co-optimization for human and cynomolgus orthologs), to name just a few.
[0349] We have demonstrated above that pre-training with natural sequences improves the predictive performance of the model. Similarly, the model can continue to improve with the addition of new data for new antibodies and with the addition of new performance or developability attributes. This technology can be used to demonstrate that naturalness is a useful proxy for common metrics, and may actually surpass these associations if intuitive relationships between naturalness and desirable properties hold. Additional properties can be added along with data from their respective assays, such as conditional pH binding, effector function, melting temperature, autoaggregation, viscosity, and more. For most of these, a single model trained with a robust dataset may be useful for a variety of antibodies of interest, and the power of the binding affinity model can be further improved through multi-task training.
[0350] Importantly, the framework presented herein facilitates tuning antibody traits toward desired values and is not necessarily limited to selecting for variants at the extremes of a given range. Furthermore, while the model presented herein focuses on antibody optimization for target affinity and naturalness traits, the approach can in principle be applied to tuning the interaction of any protein with its target.
[0351] While AI-assisted optimization of biological sequences is of practical use in reducing therapeutic development time, it does not, by itself, provide fully in silico substitutions. To this end, fully generative modeling approaches are needed. However, their training and validation face even greater data challenges because the complete novel combinatorial space considered without parent sequence anchors is dramatically larger, and potent selective binders are an infinitesimally small slice of that space.
[0352] Structure-based approaches have shown increasing capabilities and may be useful for bridging this gap. The language models presented here offer the potential to serve as in silico oracles within their extrapolation range, which can provide an effective training ground for generative models. The integration of optimization and full generation may be the next major step in data-driven treatment design.
[0353] Materials and Methods Library Cloning Antibody variants were cloned and expressed in Fab format. To generate qaACE and SPR datasets for model training and evaluation (Table 1), DNA variants were synthesized spanning CDRH2 and CDRH3 in a single oligonucleotide using ssDNA oligo pools (Twist Bioscience). Codons were randomly selected from the two most common in E. coli B strains
[35] for each variant. Two synonymous DNA sequences were synthesized for each amino acid variant (5 or 10 for the parent trastuzumab and positive / negative controls). Amplification of the Twist Bioscience ssDNA oligo pools was performed by PCR according to Twist Bioscience's recommendations, with the exception that Platinum SuperFi II DNA polymerase (ThermoFisher) was used instead of KAPA polymerase. Briefly, 20 μL reactions consisted of 1× Platinum SuperFi II Mastermix, 0.3 μM each of forward and reverse primers, and 10 ng of the oligo pool. Reactions were first denatured at 95°C for 3 min, followed by 13 cycles: 95°C for 20 s; 66°C for 20 s; 72°C for 15 s; and a final extension at 72°C for 1 min. DNA amplification was confirmed by agarose gel electrophoresis, and the amplified DNA was subsequently purified (DNA Clean and Concentrate Kit, Zymo Research).
[0354] To construct a library for SPR validation of the model design in independent experiments, oligonucleotides (59 nt) spanning CDRH3 and the immediately upstream / downstream adjacent nucleotides were synthesized by Integrated DNA Technologies (IDT). Codon usage was identical for all variants except for the mutation position. Oligonucleotides were pooled so that each oligonucleotide was represented in an equimolar manner within the pool. This single-stranded oligonucleotide pool was used directly in cloning reactions (see below) without prior amplification.
[0355] To generate the linearized vector, a two-step PCR was performed to split the Absci plasmid vector carrying trastuzumab in fab format into two fragments in a manner that provided approximately 30 nucleotides (nt) of cloning overlap at the 5' and 3' ends of the amplified Twist Bioscience library or 18 nt of cloning overlap at the 5' and 3' ends of the IDT oligonucleotide. The vector linearization reaction was digested with DPN1 (New England Biolabs) and purified from a 0.8% agarose gel (Gel DNA Recovery Kit, Zymo Research) to eliminate parent vector contamination. Cloning reactions consisted of 50 fmol of each purified vector fragment, 100 fmol of either the purified library (Twist Bioscience) or 10 pmol (IDT) insert, and NEBuilder HiFi DNA Assembly (New England Biolabs) at a final concentration of 1x. Reactions were incubated at 50°C for either two hours (Twist Bioscience libraries) or 25 minutes (IDT libraries) before purification (DNA Clean and Concentrate Kit, Zymo Research). Transformax Epi300 (Lucigen) E. coli was transformed by electroporation (BioRad MicroPulser) with the purified assembly reactions and grown overnight at 30°C on LB agar plates containing 50 μg / ml kanamycin. The following morning, colonies were scraped from the LB plates, and plasmids were extracted (Plasmid Midi Kit, Zymo Research) and submitted for QC sequencing.
[0356] QC The antibody variant library was amplified by PCR across the CDRH2 and CDRH3 regions and sequenced with 2 × 150 nt reads using an Illumina NextSeq 1000 P2 platform with 20% PhiX. PCR reactions used 10 nM primer concentration, Q5 2× Master Mix (NEB), and 1 ng of input DNA diluted in MGH20. Reactions were initially denatured at 98°C for 3 minutes, followed by 30 cycles of 98°C for 10 seconds, 59°C for 30 seconds, and 72°C for 15 seconds; with a final extension at 72°C for 2 minutes.
[0357] Sequencing results were analyzed for mutation distribution, variant representation, library complexity, and expected sequence recovery. Metrics included the coefficient of variation of sequence representation, the read share of the top 1% most common sequences, and the percentage of designed library sequences observed within the library. Activity-specific cell enrichment (ACE / qaACE) assay Antibody expression in SoluPro™ E. coli B strain SoluPro™ E. coli B strain was transformed by electroporation (Bio-Rad MicroPulser). Cells were allowed to recover in 1 ml of SOC medium for 90 minutes at 30°C with 250 rpm shaking. The recovered growth was centrifuged at 8,000 g for 5 minutes, and the supernatant was removed. The resulting cell pellet was resuspended in 1 ml of induction medium (IBM) (4.5 g / L potassium phosphate 547 monobasic, 13.8 g / L ammonium sulfate, 20.5 g / L yeast extract, 20.5 g / L glycerol, 1.95 g / L citric acid) containing inducer and supplements (260 μM arabinose, 50 μg / mL kanamycin, 8 mM magnesium sulfate, 1 mM propionic acid, 1X Korz trace metals) and then added to 100 ml of IBM containing inducer and supplements in a 1 L baffled flask. Antibody Fab induction was allowed to proceed at 30°C with shaking at 250 rpm for 24 hours. At the end of 24 hours, 1 ml aliquots of the induced culture were adjusted to 25% v / v glycerol and stored at -80°C.
[0358] Cell preparation High-throughput quantitative selection of antigen-specific Fab-expressing cells was adapted from the approach described by Liu et al.
[20] . For staining, thawed glycerol stocks from induced cultures at OD600 = 2 were transferred to 0.7 ml matrix tubes and centrifuged at 3300 g for 3 minutes. The resulting pelleted cells were washed three times with PBS + 1 mM EDTA. The washed cells were thoroughly resuspended in 250 μL of 33 mM phosphate buffer (NaHPO4) by pipetting and then fixed by adding an additional 250 μL of 32 mM phosphate buffer with 0.5% paraformaldehyde and 0.4% glutaraldehyde. After 40 minutes of incubation on ice, the cells were washed three times with PBS, resuspended in lysozyme buffer (20 mM Tris, 50 mM glucose, 10 mM EDTA, 5 μg / ml lysozyme), and incubated on ice for 8 minutes. Fixed and lysozyme-treated cells were equilibrated in staining buffer by washing three times in 0.1% saponin buffer (1×PBS, 1 mM EDTA, 0.1% saponin, 1% heat-inactivated FBS).
[0359] staining Prior to library staining, the Her2 probe was titrated against a reference strain to determine the 75% effective concentration (EC75). After lysozyme treatment and equilibration, the trast-1 library was resuspended in 250 μL of saponin buffer and transferred to a new matrix tube. The trast-3 library was incubated in AlphaLISA immunoassay buffer (Perkin Elmer; 25 mM HEPES, 0.1% casein, 1 mg / ml dextran-500, 0.5% Triton X-100, and 0.5% kathon) for 20 minutes for additional permeabilization before equilibration and resuspension in saponin buffer. A 2x concentration of staining reagent—100 nM human HER2: AF647 (Acro Biosystems) and 60 nM anti-kappa light chain: AF488 (BioLegend)—was prepared in saponin buffer, and then 250 μL of probe solution was transferred to the prepared cells, bringing the total staining volume to 500 μL with 50 nM Her2 and 30 nM anti-kappa LC. The library was incubated with the probe overnight (16 h) with end-to-end rotation at 4°C, protected from light. After incubation, the cells were pelleted, washed 3x with PBS, and then resuspended in 500 μL of PBS by vigorous pipetting.
[0360] Screening Library The sorted library was sorted on a FACSymphony S6 (BD Biosciences) instrument. Immediately before sorting, 50 μL of prepared sample was transferred to a flow tube containing 1 mL of PBS + 3 μL of propidium iodide. Aggregates, debris, and impermeable cells were removed using singlet, size, and PI+ parent gating. To reduce expression bias, an additional parent gate was set at the middle 65% of peak expression-positive cells. Collection gates were drawn to evenly sample the log range of binding signal. The rightmost gate was set to collect the brightest 10,000 events over the allotted sorting time, estimated by including the five brightest events for every 65,000 in the expression parent gate. Seven additional gates were then set to fractionate the positive binding signal, and one gate collected the negative binding population, as shown in Figure S13A and Figure S13B.
[0361] Specifically, Figure 13A depicts representative parent gating for all ACE sorts, according to some embodiments. Two singlet gates were drawn to exclude SoluPro™ aggregated regions previously identified by dual fluorescence of the GFP and mCherry reporter strains, and propidium iodide was used to exclude non-permeabilized cells. Figure 13B depicts specific expression and collection gating for each ACE library sort, according to some embodiments. A parent gate encompassing approximately 65% of expression-positive cells and centered over the peak expression signal was drawn before setting a collection gate on the probe-specific binding signal.
[0362] Libraries were sorted simultaneously on two instruments with the photointensifier adjusted to normalize fluorescence intensity, and collected events were processed independently as technical replicates.
[0363] Next-generation sequencing Cell material from various gates was collected in a diluted PBS mixture (VWR) in 1.5 mL tubes (Eppendorf). After sorting, samples were spun down at 3,800 g, and the tube volume was normalized to 20 μl. Amplicons for sequencing were generated using one of two methods. The first method directly used the collected cell material as a template to amplify the CDRH2 and CDRH3 regions via two-step PCR. During the initial PCR step, unique molecular identifiers (UMIs) and partial Illumina adapters were added to the CDRH2 and CDRH3 amplicons via four PCR cycles. The second-step PCR added the remaining Illumina sequencing adapters, as well as Illumina i5 and i7 sample indices. The initial PCR reaction used a 1 nM UMI primer concentration, Q5 2x Master Mix (NEB), and 20 μl of sorted cell material input suspended in diluted PBS (VWR). Reactions were first denatured at 98°C for 3 minutes, followed by four cycles of 98°C for 10 seconds; 59°C for 30 seconds; and 72°C for 30 seconds, with a final extension at 72°C for 2 minutes. After the first PCR, 0.5 μM secondary sample index primer was added to each reaction tube. Reactions were then denatured at 98°C for 3 minutes, followed by 29 cycles of 98°C for 10 seconds; 62°C for 30 seconds; and 72°C for 15 seconds, with a final extension at 72°C for 2 minutes. The second method amplifies the CDRH2 and CDRH3 regions without the addition of UMIs. This single-step PCR used 10 nM primer concentrations, Q5 2x Master Mix (NEB), and 20 μl of sorted cell material input suspended in diluted PBS (VWR). Reactions were first denatured at 98°C for 3 minutes, followed by 30 cycles of 98°C for 10 seconds; 59°C for 30 seconds; 72°C for 15 seconds, with a final extension at 72°C for 2 minutes. After amplification by either method, samples were run on a 2% agarose gel at 75V for 60 minutes, and bands of the appropriate length were excised and purified using a Zymoclean Gel DNA Recovery Kit (Zymo Research).The resulting DNA samples were quantified using a Qubit fluorometer (Invitrogen), normalized, and pooled. Pool sizes were verified via Tapestation 1000 HS and sequenced on an Illumina NextSeq 1000 P2 (2 × 150 nt) using 20% PhiX.
[0364] ACE assay analysis To generate quantitative binding scores from the reads, the following processing and quality control steps were performed:
[0365] 1. Paired-end reads were merged using FLASH2
[36] with the maximum allowed overlap set according to amplicon size and sequencing read length (150 bases for all libraries described herein).
[0366] 2. If UMIs were added during amplification, the downstream UMI tag (the last 8 bases) was moved to the start of the read and any PCR duplicates were removed using the UMI Collapse tool
[37] in FASTQ mode. Only perfectly identical sequences were considered duplicates, and no error correction was performed at this stage.
[0367] 3. Primers were removed from both ends of merged reads using the cutadapt tool
[38] and reads were discarded where no primers were detected.
[0368] 4. Reads were aggregated across all FACS sorting gates and aligned to a reference sequence (parental version of the amplicon) in amino acid space. Alignment was performed using the Needleman-Wunsch algorithm implemented in Biopython
[39] with the following parameters: PairwiseAligner, mode=global, match_score=5, mismatch_score=-4, open_gap_score=-20, extend_gap_score=-1. Parameters were chosen by manual inspection across multiple processed libraries.
[0369] 5. Reads were then discarded if (1) the average base quality fell below 20 or (2) the sequence (in DNA space) was seen in fewer than 10 reads across all gates (or in fewer than 10 unique molecules after UMI deduplication, if available).
[0370] 6. The procedure also flagged: (1) sequences aligned to the reference with a low score (defined as less than 0.6 of the score obtained by aligning the reference to itself); (2) sequences containing stop codons outside the region of interest; and (3) sequences containing frameshift insertions or deletions. Flagged sequences were not included in any mutation-related statistics but were used to normalize counts for binding score calculations. Sequencing quality control metrics were generated using FastQC
[40] and MultiQC
[41] .
[0371] 7. For each gate, the abundance of each sequence (read / UMI count relative to the total number of reads / UMIs from all sequences in that gate) was normalized to 672 counts in a million.
[0372] 8. An binding score (ACE score) was assigned to each unique DNA sequence by taking the weighted average of the normalized counts across the sorting gates. For all experiments, weights were assigned linearly using an integer scale: the gate capturing the lowest fluorescent signal was assigned a weight of 1, the next lowest a weight of 2, etc.
[0373] 9. Any detected sequences that were not present in the originally designed and 679 synthesized library were removed.
[0374] 10. For each unique amino acid variant, the qaACE scores from the synonymous DNA sequences were averaged.
[0375] 11. qaACE scores were averaged across independent FACS sorts, and sequences with a standard deviation of replicates greater than 1.25 were removed. Amino acid variants were retained only if at least three independent QC pass observations were collected between the synonymous DNA variant and the replicate FACS sort.
[0376] Surface Plasmon Resonance (SPR) Antibody expression in SoluPro™ E. coli B strain Individual SoluPro™ E. coli B strain colonies expressing antibody Fab variants were inoculated into LB medium in a 96-well deep block (Labcon) and grown at 30°C for 24 hours to generate seed cultures for induction of expression. The seed cultures were then inoculated into IBM containing inducer and supplements in a 96-well deep block and grown for an additional 24 hours at 30°C. Post-induction, samples were transferred to a 96-well plate (Greiner Bio-One), pelleted, and lysed in 50 μL lysis buffer (1X BugBuster Protein Extraction Reagent containing 0.01K Benzonase Nuclease and 1X Protease Inhibitor Cocktail). Plates were incubated at 30°C for 15-20 minutes and then centrifuged to remove insoluble debris. After lysis, samples were adjusted to a final volume of 260 μL with 200 μL of SPR running buffer (10 mM HEPES, 150 mM NaCl, 3 mM EDTA, 0.01% w / v Tween-20, 0.5 mg / mL BSA) and filtered into a 96-well plate. The lysed samples were then transferred from the 96-well plate to a 384-well plate for high-throughput SPR using a Hamilton STAR automated liquid handler. Colonies were prepared in two sets of independent replicates before lysis, and each replicate was measured in two separate experimental runs. In some instances, a single replicate was used as indicated.
[0377] SPR experiments High-throughput SPR experiments were performed on a microfluidic Carterra LSA SPR instrument using SPR running buffer (10 mM HEPES, 150 mM NaCl, 3 mM EDTA, 0.01% w / v Tween-20, 0.5 mg / mL BSA) and SPR wash buffer (10 mM HEPES, 150 mM NaCl, 3 mM EDTA, 0.01% w / v Tween-20). Carterra LSA SAD200M chips were pre-functionalized with 20 μg / mL biotinylated antibody capture reagent for 600 seconds before experiments. Lysate samples in a 384-well block were immobilized on the 25 / 74 chip surface for 600 seconds, followed by a 60-second wash step for baseline stabilization. Antigen binding was performed using a non-regenerative kinetic method with a 300-second association phase followed by a 900-second dissociation phase. For analyte injection, six leading blanks were introduced to create a consistent baseline before monitoring antigen binding kinetics. After the leading blank, five concentrations of HER2 extracellular domain antigen (ACRO Biosystems, prepared in three-fold serial dilutions from a starting concentration of 500 nM) were injected into the instrument, and the time-series response was recorded. In most experiments, measurements of individual DNA variants were repeated four times. Typically, each experimental run consisted of two complete measurement cycles (ligand immobilization, leading blank injection, analyte injection, chip regeneration), thereby providing two duplicate measurement runs per clone per run. In most experiments, technical replicates measured in separate runs further doubled the number of measurement runs per clone to four.
[0378] Sensorgram baseline subtraction Sensorgrams were generated from the raw data using the Carterra Kinetics GUI software application provided with the Carterra LSA instrument. Sensorgram response values versus time for 384 regions of interest (ROIs) on the Carterra chip were corrected using a double referencing and alignment technique implemented by the Carterra manufacturer. This technique incorporates both the time-synchronous response of interspot reference regions adjacent to the ROI, as well as the asynchronous response from a preceding blank buffer injection run over the same ROI during a previous experimental run cycle, to estimate and subtract background responses. Corrected sensorgrams were exported from the Kinetics software package for offline analysis.
[0379] Kinetic binding parameters Kinetic binding parameters are defined as t c Estimates were made via nonlinear regression using a standard 1:1 binding model modified by incorporation of a vector of parameters. For a single analyte concentration, the association phase model is:
number
[0380] An additional concentration-dependent time offset parameter t c was required for the unique measurement system used by Carterra, in which successive association phase measurements at each new analyte concentration are taken before the analyte from the previous phase has completely dissociated, leading to response curves that do not start from a zero response at t=0. The time offset parameter represents the predicted time intercept of each association response curve, i.e., the amount of time before the onset of the association phase, at which measurements would have to begin in order to reach the actual observed response at t=0. The dissociation phase was modeled as a standard decay.
number
[0381] Regressions were performed using R language
[42] scripts. Minpack.lm
[43] is an R-ported copy of MINPACK-1
[44]
[45] , a FORTRAN-based software package that implements the Levenberg-Marquardt
[46]
[47] nonlinear least-squares parameter search algorithm, and was used to perform the parameter search.
[0382] Next-generation sequencing To identify the DNA sequences of individual antibody variants evaluated in SPR, next-generation sequencing (NGS) was performed on the measured variants. Individual colonies were picked from LB agar plates containing 50 μg / mL kanamycin (Teknova) into 96 deep-well plates containing 1 mL of LB medium (Teknova). The culture plates were grown overnight in a shaking incubator at 30°C. 200 μl of the overnight culture was transferred to a new 96-well plate (Labcon) and spun down at 3,500 g. A portion of the pelleted material was transferred via a pinner (Fisher Scientific) into a 96-well PCR (Thermo-Fisher) plate, which contained reagents for adding Illumina adapters and performing the initial PCR of a two-stage PCR for sequencing. The reaction volume used was 25 μl. During the initial PCR stage, partial Illumina adapters were added to the CDRH2 and CDRH3 amplicons via four PCR cycles. The second-stage PCR added the remaining Illumina sequencing adapters and Illumina i5 and i7 sample indexes. The initial PCR reaction used a 0.45 μM UMI primer concentration and 12.5 μl of Q5 2x Master Mix (NEB). Reactions were first denatured at 98°C for 3 minutes, followed by four cycles of 98°C for 10 seconds, 59°C for 30 seconds, and 72°C for 30 seconds, with a final extension at 72°C for 2 minutes. After the initial PCR, 0.5 μM secondary sample index primer was added to each reaction tube. Reactions were then denatured at 98°C for 3 minutes, followed by 29 cycles of 98°C for 10 seconds, 62°C for 30 seconds, and 72°C for 15 seconds, with a final extension at 72°C for 2 minutes. Reactions were then pooled into 1.5 mL tubes (Eppendorf). The pooled samples were size-selected using a 1x AMPure XP (Beckman Coulter) bead procedure, and the resulting DNA samples were quantified using a Qubit fluorometer.Pool sizes were verified via Tapestation 1000 HS and sequenced on an Illumina MiSeq Micro (2 × 150 nt) using 20% PhiX.
[0383] After sequencing, amplicon reads were merged according to their sample index. Merging was performed using a custom Python script. The script merged R1 and R2 reads based on overlapping sequences. The occurrence of unique amplicon sequences within each sample was counted and tallied. A custom R script was then applied to calculate the sequence frequency ratio and Levenshtein distance between the dominant and secondary sequences observed within the sample. These calculations were used for downstream quality filtering to ensure clonality of SPR measurements. The dominant sequences within each sample were then combined with the companion Carterra SPR measurements.
[0384] QC SPR matches were excluded if any of the following criteria were met: Fewer than three analyte concentrations providing a usable fit; handling errors described by the operator; · Non-physical fits (e.g., upward-sloping dissociation phase signals even after sensorgram baseline subtraction); Non-convergent fit -log coupled with an estimated signal-to-noise ratio of less than 10 for the highest analyte concentration ca included in the fit (typically 500 nM) 10 K D Values ≤ 8.5; -log conjugated with an estimated signal-to-noise ratio that is less than 70 for the highest analyte concentration included in the fit 10 K D values >8.5; ·t c < -300 s or t c t for the highest analyte concentration included in the fit so that s > 0 c value; ·Failed NGS; non-clonal sequences (dominant sequences less than 100 times more abundant than secondary sequences if the Levenshtein distance between them is greater than 2); and / or The sequence does not match any design variant in the synthetic oligo pool (within a sequence identity tolerance to accommodate sequencing error).
[0385] K D and k off -log 10 While the transformation on Log 10 The distribution of kinetic parameters was visually inspected for the absence of significant batch effects. Multiple measurements of the same antibody variant (usually (a) duplicate consecutive measurements of the same clone in the same SPR run; (b) technical replicates of the same clone from duplicate 384-well plates measured in separate runs; (c) two DNA variants with identical translations, if available; and (d) independent clones of a variant) were averaged in logarithmic space - log 10 K D Variants whose measurements showed a coefficient of variation greater than 5% upon aggregation were removed.
[0386] Observed Antibody Space (OAS) database processing The OAS database of unpaired immunoglobulin chains
[48] was downloaded on February 1, 2022. From the entire database, the following exclusions were applied to the raw OAS data: first, studies whose samples come from another study in the database (author column, Bonsignori et al., 2016; Halliley et al., 2015; Thornqvist et al., 2018); second, studies derived from immature B cells (B column, immature B cells and pre-B cells) and B cell-related cancers (disease column, light chain amyloidosis, CLL); and finally, sequences were excluded if any of the following criteria were met: The sequence includes a stop codon. · Arrays are unproductive. · The V and J segments are outside the frame. Framework region 2 is missing. Framework region 3 is missing. · CDR3 is longer than 37 amino acids. J segment sequence identity with the closest germline is less than 50% The sequence lacks amino acids at the beginning or end of any CDR. · It lacks a conserved cysteine residue. · The locus does not match the strand type.
[0387] From the resulting sequences, and for each of the two (heavy / light) chains, two types of subsequences were extracted: "CDR" and "near-full length (NF)." In the CDR dataset, we extracted only the CDR1, CDR2, and CDR3 segments defined by the combination of the IMGT
[49] and Martin
[50] labeling schemes. In the NF dataset, we included positions 21 to 128 of IMGT (for the light chain and 127 for the heavy chains from rabbit and camel).
[0388] For all four datasets, duplicate sequences were removed while redundancy information (i.e., the number of times a particular sequence was observed in each test) was tallied. Sequences with a redundancy of 1 (i.e., observed only once in a single test) were removed due to insufficient evidence of a genuine biological sequence, as opposed to sequencing error. FIG. 14 depicts a flow diagram illustrating a computer-implemented method 1400, according to some embodiments, with the number of sequences filtered and retained after each preprocessing step. Method 1400 may include constructing four databases by processing the OAS dataset. Method 1400 may include constructing the databases by extracting two subsets of sequences (CDR-only and near-full-length antibody sequences) for each of the two chains, heavy and light, as discussed above. Method 1400 may include using models trained on the CDR dataset for binding affinity and naturalness prediction, with the exception of the CR9144 case study, where a model trained on the near-full-length dataset may be used due to the location of the mutation positions. The numbers labeled "H" and "L" in Figure 14 depict the unique heavy and light chain sequences, respectively, that were filtered or retained at each step. Method 1400 can be performed by a computing device, such as computing device 102 of Figure 1.
[0389] Model Architecture Protein language models have shown great promise across a variety of protein manipulation tasks [17, 51-55]. In some embodiments, the present architecture is based on the RoBERTa model
[56] and its PyTorch implementation within the Hugging Face framework
[57] . In some embodiments, the model contains 16 hidden layers with 12 attention heads per layer. In some embodiments, the hidden layer size is 768 and the hidden layer size is 3072. In total, the model may contain 114 million parameters. In pilot studies, larger and smaller models were tested and their respective losses were compared in both masked language modeling and regression tasks. It was observed that smaller models performed poorly, while larger models did not provide significant performance enhancement, confirming the suitability of the selected model size. However, it will be understood by those skilled in the art that differently constructed and parameterized models may be used in the present technology to achieve its objectives.
[0390] Model Training Pre-training with OAS antibody sequences In some embodiments, the models for predicting binding affinity presented herein can be derived from a RoBERTa architecture pre-trained with immunoglobulin sequences from four datasets resulting from OAS database processing (see Observed Antibody Space (OAS) Database Processing, above). Thus, four or more models can be trained on heavy or light chain, CDR, or NF sequences. In some embodiments, the training sequences include species tokens (e.g., h for human, m for mouse, etc.) for conditioning the language model.
[58] Additionally, the input sequences to the CDR model can include CDR-unrestricted tokens, allowing originally non-contiguous CDR segments to be concatenated into a single input sequence.
[0391] In some aspects, the CDR model may be used for all binding affinity and naturality predictions, except for the CR9114 case study, where the NF model may be used due to some framework substitutions present in the dataset.
[0392] In some embodiments, model training may be performed in a self-supervised manner, following a dynamic masking procedure as described in Wolf et al.
[56] , whereby 15% of the tokens in the sequence are randomly masked with a special [MASK] token. For masking, the DataCollatorForLanguageModeling class from the Hugging face framework may be used, whereby, unlike Wolf et al.
[56] , all randomly selected tokens are simply masked. Training may be performed over a period of time, e.g., 10 -6 This may be implemented using the LAMB optimizer
[59] with an ε of 0.003, a weight decay of 0.003, and a clamp value of 10. In some embodiments, the maximum learning rate used is 10 ~3 with linear decay and 1000-step warming, a dropout probability of 0.2, a weight decay of 0.01, and a batch size of 416. The model may be trained for up to 10 epochs in some embodiments.
[0393] Fine-tuning using affinity data Transfer learning can be used in some embodiments to leverage the OAS pre-trained model by adding a dense hidden layer with a large number of nodes (e.g., 768), followed by a projection layer with the required number of outputs. All layers can be left unfrozen to update all model parameters during training. Training can be performed using the AdamW optimizer
[60] , e.g., 10 -5 with a learning rate of 0.01, a weight decay of 0.01, a dropout probability of 0.2, a linear learning rate decay with a warming step of 100, a batch size of 64, and mean squared error (MSE) as the loss function.
[0394] In some embodiments, the model may be trained for 25,000 steps. The number of steps, batch size, and learning rate for all runs were determined through hyperparameter search using a pilot dataset. For example, a grid search was performed using three learning rates (10 -4 , 10 -5 , 10 -6 ), three batch sizes (64, 128, 256), and two number of runs (25,000, 50,000). Each hyperparameter set may be used to fine-tune the OAS pre-trained RoBERTa model, for example, using a 90:10 training:holdout split from the pilot dataset as shown in FIG. 15A and a subset of 500 randomly selected sequences from the pilot dataset as shown in FIG. 15B. The hyperparameter sweep on the pilot dataset in FIGS. 15A and 15B may include three metrics of predictive accuracy on the test set for each model, along with the time required to train the model. To minimize model training time while maintaining model performance, the final hyperparameters may be adjusted, for example, using 10 for learning rate and 10 for holdout. -5 , a batch size of 64, and 25,000 training steps.
[0395] Co-training with qaACE and SPR data In some embodiments, the model may be utilized to predict both qaACE and SPR values from sequences, using a weighted sum of mean squared errors for each regression task as a loss function. For example, a sweep through weights showed that a weighting of 1:50 for SPR provided the best combined accuracy, as shown in Figure 16. In some embodiments, the model may be evaluated using pooled out-of-fold predictions in a 10-fold cross-validation setting, and using data from both ACE and SPR experiments simultaneously, combining the losses from each dataset using a weighted sum. The out-of-fold predictions were pooled in a 10-fold cross-validation to obtain a -log 2 metric as measured by SPR. 10 K D It may be compared to a value.
[0396] Model characterization Baseline To evaluate the effectiveness of fine-tuning a pre-trained model, two baselines may be evaluated. First, an XGBoost
[61] model may be implemented using one-hot encoding of amino acids. For example, as depicted in Figure 17, the following seven XGBoost hyperparameters may be selected using a grid search on the pilot dataset: eta = 0.5, gamma = 0, n_estimators = 1000, subsample = 0.6, max_depth = 9, min_child_weight = 1, col_sample_by_tree = 1. Figure 17 depicts the test set accuracy of 10-fold cross-validation across each of the seven optimized XGBoost hyperparameters. An exhaustive grid search may be performed across a large number (e.g., 3) of values for each hyperparameter using both the full pilot dataset and a subset of the pilot dataset containing a large number (e.g., 500) of selected sequences.
[0397] In some embodiments, default values may be used for other hyperparameters. Second, a RoBERTa model with the same architecture as the pre-trained model may be trained with affinity data starting from randomly initialized weights without OAS pre-training.
[0398] Out-of-distribution prediction of binding affinity To assess the predictive power for binding affinities outside the distribution seen in the training set, the technique was used to predict binding affinities higher than that of the parent trastuzumab. 10 K D This may include fine-tuning the model by excluding from the training set any variants with: Further, the model may be tasked with predicting the affinity of a set of sequences highly enriched in binders stronger than trastuzumab, as validated by SPR.
[0399] Assessing the size and fidelity of the training data In some embodiments, models may be trained using subsets of different sizes from datasets of varying fidelity. The trast-3 dataset was processed as a high-fidelity dataset. Low-fidelity datasets were generated by isolating a single DNA variant for each sequence from a single FACS sort using the same preprocessing workflow. Each training dataset may be divided (e.g., evenly divided into 1, 2, 4, 8, 16, 32, 64, and 128 subsets, respectively). Each training subset is used both to directly train a model with randomly initialized weights and to fine-tune an OAS pre-trained model. In some embodiments, a common holdout dataset containing 10% of the data from the original trast-3 dataset may be used to evaluate all models, regardless of data source or training set size. These sequences may be removed from both datasets before constructing the training subsets.
[0400] embedded The embedding can be generated by taking the average pool of activations from the last hidden layer of the model, excluding the head. The resulting size of the embedding for each sequence may be, for example, 768. The size of the embedding can be reduced with, for example, the Uniform Manifold Approximation and Projection (UMAP) algorithm implemented in the RAPIDS library
[62] .
[0401] In the first study, embeddings were compared from four different models, resulting from the presence or absence of OAS pre-training and the presence or absence of fine-tuning of binding affinities using the trast-2 dataset. In the second study, embeddings were used to cluster close variants in the internal representation space. To this end, the reduced-dimension embeddings were filtered to retain only strong binders based on predicted qaACE scores, and the 3D embeddings were clustered using HDBSCAN
[63] with a minimum cluster size of 40 sequences. Sequence logo plots for each cluster can be generated using Logomaker
[64] .
[0402] Epistasis Epistatic interactions between mutations can be assessed by considering the predicted affinity scores for the double mutant, the constituent single mutants, and the parent antibody sequence. Specifically, the epistatic effect between two mutations, m1 and m2, can be calculated as follows: Epistasis (m1, m2) = (y 1,2 -y wt )―(y1-y wt )―(y2-y wt ) In the formula, y i is a mutant with mutation i, or y wt The predicted ACE score for the parent sequence is displayed.
[0403] In the formula, y i is a mutant with mutation i, or y wt The predicted qaACE scores for the parent sequences are displayed.
[0404] Naturality of antibodies Naturalness of arrays n s can be defined as the inverse of its pseudo-perplexity, following the definition by Salazar et al.
[65] for a masked language model (MLM). Recall that for a sequence S with N tokens, the pseudo-likelihood that an MLM with parameters Θ assigns to this sequence is given by:
number
[0405] The pseudo-perplexity is obtained by first normalizing the pseudo-likelihood by the sequence length and then applying a negative exponential function:
number
[0406] Thus, the naturalness of arrays is:
number
[0407] A naturalness score may be calculated using the two pre-trained models described above (see pre-training with OAS antibody sequences). Several antibody properties (immunogenicity, developability, expression level, and mutational load) may be analyzed to examine potential relationships with sequence naturalness. For datasets where members show variation in both chains (immunogenicity and expression level), the reported naturalness score may be the average of the individual heavy and light chain scores. For datasets where members show variation only in the heavy chain (developability and mutational load), only the heavy chain naturalness score may be calculated. In some embodiments, a naturalness score is reported in all cases from a model trained on a CDR dataset (pre-training with OAS antibody sequences, see above).
[0408] Naturalness-related plots To assess the relationship between the naturalness of interest and antibody characteristics, the present technology may utilize a naturalness-related plot. To construct such a plot, a two-column table may be considered, in which each row contains a pair (antibody naturalness, antibody characteristic value). The following procedure may then be followed: Binning the naturalness values into four equal intervals (low, low-medium, medium-high, high) resulting from dividing the naturalness range by 4. For each naturalness interval, take the characteristic value of the antibodies whose naturalness falls within that interval and summarize the values by median (continuous characteristic) or percentage of positives (binary characteristic). The summary (Y axis) by naturalness interval (X axis) is then plotted as a box plot (continuous characteristic) or bar plot (binary). · Calculate the Jonckheere-Terpstra test for increasing or decreasing trend and record the p-value. Add dashed lines connecting medians (for box plots) or heights (for bar plots).
[0409] For box plots, the whisker parameter may be set to 1.5 and outliers may not be plotted. In general, the rationale for aggregating continuous variables using median and binning is to reduce the influence of outliers and noisy data points.
[0410] immunogenicity This technique may include obtaining immunogenic responses reported as the percentage of patients with anti-drug antibody (ADA) responses from Marks et al.
[28] . Of all classes of antibodies (human, humanized, chimeric, hybrid, murine), only humanized antibodies may be included in some analytical aspects because (1) inter-class comparisons are trivial, thus weighing on simple species discrimination, while intra-class comparisons are difficult and practically relevant; (2) humanized antibodies represent the largest class in this dataset (n=97), thereby providing the greatest statistical power; and (3) compared to the second largest class in the dataset, it is human antibodies, which, in principle, have more immunogenic potential due to the animal origin of their CDRs, thereby providing a practically relevant case study for evaluating the degree of human naturalness achieved during engineering / humanization.
[0411] Development Potential Sequence developability can be defined as a binary variable indicating whether an antibody sequence fails at least one of the developability flags calculated by the Therapeutic Antibody Profiler (TAP) tool.
[31] For detailed definitions of these flags, see the TAP analysis subsection, above.
[0412] The technique may involve scoring hit (positively enriched in round 3 compared to round 2) heavy chain sequences (n=882) from the phage display library described in Liu et al.
[29] , which we refer to as the Gifford Library. The technique may also involve analyzing trastuzumab variants with up to three simultaneous amino acid substitutions at 10 positions in CDRH2 and 10 positions in CDRH3 (following the same mutagenesis strategy of the trast-CDRH3 dataset).
[0413] Processing the Gifford Library Dataset This technique may involve obtaining phage display data from the Gifford library described in Liu et al.
[29] . Specifically, raw FASTQ files for rounds 2 (E1_R2) and 3 (E1_R3) of enrichment may be downloaded from the NIH Sequence Read Archive (SRA) under accession number SRP158510. The guidelines for processing data according to Liu et al.
[29] may then be followed. First, the CDRH3 sequence may be derived using the flanking DNA sequences TATTATTGCGCG (SEQ ID NO: 26) and TGGGGTCAA. Next, sequences containing N or that cannot be translated (divisible by 3) may be filtered out. Next, the DNA sequence may be translated into protein sequence and deleted if it contains a premature stop codon. Next, CDRH3 sequences shorter than 8 amino acids or longer than 20 amino acids may be filtered out. Finally, the number of occurrences of each unique sequence may be determined, and sequences occurring less than 6 times may, in some embodiments, be considered noise and deleted / excluded.
[0414] Construction of scaffold sequences in the Gifford library dataset The full-length antibody sequence may be required for analysis using the NF language model. However, the raw sequencing data from Liu et al.
[29] only includes CDRH3, and the original study provides scaffold gene names rather than sequences. Therefore, in some embodiments of the present technology, the antibody sequence can be reconstructed from the provided gene names. For the heavy chain, the present technology may use the IGHV3-23 germline and perform the following modifications: if CDRH3 ends in DY, add the IGHJ4 sequence to IGHV3-23; however, if CDRH3 ends in DV, use IGHJ6 instead, following Liu et al.
[29] . The heavy chain gene may also be cross-validated using IgBlast.
[0415] The resulting scaffold sequences are reported below, where the __CDR3__ placeholder was replaced with the selected NGS-derived CDRH3 sequence.
[0416] >HC_DY(SEQ ID NO:1)
[0417] EVQLVESGGGLVQPGGSLRLSCAASGFTFSSYAMSWVRQAPGKGLEWVSAISGSGGSTYYADSVKGRFTISRDNS-KNTLYLQMNSLRAEDTAVYYC__CDR3__DYWGQGTLVTVSS
[0418] >HC_DV 1016 (SEQ ID NO: 27)
[0419] EVQLVESGGGLVQPGGSLRLSCAASGFTFSSYAMSWVRQAPGKGLEWVSAISGSGGSTYYADSVKGRFTISRDNS-KNTLYLQMNSLRAEDTAVYYC__CDR3__DVWGKGTTVTVSS
[0420] TAP analysis This technology may calculate a developability score using the Therapeutic Antibody Profiler (TAP) described in Raybouldet et al.
[66] . A commercially licensed virtual machine image of the tool may be used (e.g., last updated on February 7, 2022). TAP calculates five developability metrics: total CDR length, surface patch, hydrophobicity (PSH), positively charged patch (PPC), negatively charged patch (PNC), and structural Fv charge symmetry parameter (SFvCSP). It also generates a flag indicating whether the metrics are acceptable relative to a reference set of therapeutic antibodies. Metrics that fall outside the reference distribution may be flagged as "red," whereas metrics that fall within the most extreme 5% of the distribution may be "amber," and metrics that fall within the body of the distribution above the 5% threshold may be "green" and acceptable.
[0421] TAP may be used to analyze sequence hits from the Gifford Library dataset as well as trastuzumab variants. TAP flags may be used to determine whether an antibody has an acceptable developability score; an antibody variant was considered a failure if at least one of the TAP flags was not green.
[0422] Antibody expression in HEK-293 cells In some embodiments, clinical-stage antibody expression levels in HEK-293 cells can be collected from Jain et al.
[30] . The dataset can be heterogeneous with respect to antibody type (e.g., human, humanized, chimeric, etc.). For the same reasons as illustrated for immunogenicity, the technique can focus on humanized antibodies (n=67).
[0423] In addition to HEK-293 titers, Jain et al. reported additional biophysical measurements.
[0424] In our testing, we did not specifically identify associations between naturalness scores and biophysical parameters other than potency. However, those skilled in the art will understand that clinical-stage antibody datasets are inevitably already biased toward antibodies endowed with favorable properties, meaning that the distribution of biophysical parameters strongly depletes poorly performing antibodies. The availability of positive, but not negative, examples severely limits our ability to detect associations between biophysical parameters and other metrics, such as naturalness.
[0425] Mutation load Mutational burden may be defined as the number of amino acid substitutions in an antibody variant compared to its parent sequence. The technique may analyze the distribution of naturalness scores across 6,710,401 trastuzumab variants with a mutational burden between 1 and 3 (10 positions in CDRH2 and 10 positions in CDRH3, allowing all natural amino acids except cysteine). The technique may include assessing the statistical significance of differences in naturalness score distribution by mutational burden using a Jonckheere-Terpstra test for trend.
[0426] Genetic Algorithm To generate sequence variants with desired properties (e.g., high / low / target qaACE scores and high naturalness), the technique may involve a genetic algorithm (GA) using, for example, a tailored version of the DEAP library in Python
[67] . In this GA, each individual sequence variant may be reduced to the CDR representation (a fusion of the IMGT and Martin definitions) described above. Each GA run may be initialized from a single trastuzumab sequence. The predicted qaACE score and naturalness score of each sequence may be evaluated using the model described above. A cyclic selection-regeneration-mutation-thinning process was applied to a common starting sequence pool in a μ + λ GA
[68] .
[0427] Each progeny pool may, for example, contain the original 100 parents as well as 200 new, unique 1062 individuals. Of the progeny, 30% were generated from single point mutations (excluding cysteines) in the parents, and 70% were generated from two-point crosses between the two parents.
[0428] For example, because the GA is initialized from a single sequence, the first offspring pool contained 299 individuals, all of which were generated using a single point mutation from trastuzumab. In some embodiments, all sequences can be constrained to stay within the trast-3 library computational space (up to triple mutants at 10 positions in CDRH2 and CDRH3, respectively). If unique offspring cannot be produced within these constraints, individuals randomly generated within the constraints can be added to the offspring pool. Tournament selection without replacement (tournament size = 3) can be performed to thin the population (size = 300) and select individuals for the next generation (size = 100).
[0429] This process represented one "generation" of the GA. The GA can be run for multiple generations (e.g., 20). To strike an appropriate balance between the qaACE score and the naturalness goal, the fitness objective may be defined as follows:
number
[0430] To test the generative capabilities of the present technical model, the GA can be run in the following configuration: Target qaACE score = 9 (maximize qaACE score), while maximizing naturalness Target qaACE score = 1 (minimize qaACE score), while maximizing naturalness Target qaACE score = 6, while maximizing naturalness For example, the GA queries 300 individuals in the first generation and 200 individuals in each subsequent generation, so that the GA queries 4,100 (non-unique) sequences over 20 generations.
[0431] As a baseline, the technique may involve randomly selecting 4,100 sequences from the full mutation search space and selecting the top 100 individuals with the highest fitness as described above. The fitness function may also be used to identify the top 100 individuals from an exhaustive search of the mutation space and from the trust-3 dataset.
[0432] References 1.SMPaul, DSMytelka, CTDunwiddie, CCPersinger, BHMunos, SRLindborg, and ALSchacht, “How to improve R&D productivity: the pharmaceutical industry's grand challenge,” Nature Reviews Drug Discovery, vol. 9, pp. 203-214, Feb 2010.
[0433] 2. S. Yamaguchi, M. Kaneko, and M. Narukawa, “Approval success rates of drug candidates based on target, action, modality, application, and their combinations,” Clinical and Translational Science, vol. 14, pp. 1113-1122, April 2021.
[0434] 3. JPHughes, S. Rees, SB Kalindjian, and KL Philpott, “Principles of early drug discovery,” British Journal of Pharmacology, vol. 162, pp. 1239-49, Mar 2011.
[0435] 4.J.Ministro,A.Manuel,and J.Goncalves,“Therapeutic antibody engineering and selection strategies,” Advances in biochemical engineering / biotechnology,vol.171,pp.55-86,Nov 2019.
[0436] 5.K.R.Hanning,M.Minot,A.K.Warrender,W.Kelton,and S.T.Reddy,“Deep mutational scanning for therapeutic antibody engineering,” Trends in Pharmacological Sciences,vol.43,no.2,pp.123-135,2022.
[0437] 6.T.Kuramochi,T.Igawa,H.Tsunoda,and K.Hattori,“Humanization and simultaneous optimization of monoclonal antibody,” Methods in Molecular Biology,vol.1060,pp.123-137,2014.
[0438] 7.C.Schneider,A.Buchanan,B.Taddese,and C.M.Deane,“DLAB-Deep learning methods for structure-based virtual screening of antibodies,” Bioinformatics,vol.38,pp.377-383,Sep 2021.
[0439] 8.A.Khan,A.I.Cowen-Rivers,D.-G.-X.Deik,A.Grosnit,K.Dreczkowski,P.A.Robert,V.Greiff,R.Tutunov,D.Bou-Ammar,J.Wang,and H.Bou-Ammar,“AntBO:Towards real-world automated antibody design with combinatorial bayesian optimisation,” arXiv:2201.12570 [q-bio.BM],2022.
[0440] 9.W.Jin,J.Wohlwend,R.Barzilay,and T.S.Jaakkola,“Iterative refinement graph neural network for antibody sequence-structure co-design,” arXiv:2110.04624 [q-bio.BM],2022.
[0441] 10.W.Jin,D.Barzilay,and T.Jaakkola,“Antibody-antigen docking and design via hierarchical structure refinement,” in Proceedings of the 39th International Conference on Machine Learning(K.Chaudhuri,S.Jegelka,L.Song,C.Szepesvari,G.Niu,and S.Sabato,eds.),vol.162 of Proceedings of Machine Learning Research,pp.10217-10227,PMLR,17-23 Jul 2022.
[0442] 11.S.Luo,Y.Su,X.Peng,S.Wang,J.Peng,and J.Ma,“Antigen-specific antibody design and optimization with diffusion-based generative models,” bioRxiv doi:10.1101 / 2022.7.10.499510,2022.
[0443] 12.S.P.Mahajan,J.A.Ruffolo,R.Frick,and J.Gray,“Hallucinating structure-conditioned antibody libraries for target-specific binders,” BioRxiv doi:10.1101 / 2022.6.6.494991,2022.
[0444] 13.J.J.Jeffrey A.Ruffolo,Jeremias Sulam,“Antibody structure prediction using interpretable deep learning,” Patterns,vol.3,p.100406,Feb 2022.
[0445] 14.R.W.Shuai,J.A.Ruffolo,and J.J.Gray,“Generative language modeling for antibody design,” bioRxiv doi:10.1101 / 2021.12.13.472419,2021.
[0446] 15.D.M.Mason,S.Friedensohn,C.R.Weber,C.Jordi,B.Wagner,S.M.Meng,R.A.Ehling,L.Bonati,J.Dahinden,P.Gainza,B.E.Correia,and S.T.Reddy,“Optimization of therapeutic antibodies by predicting antigen specificity from antibody sequence via deep learning,” Nature Biomedical Engineering,pp.600-612,Apr 2021.
[0447] 16.K.Saka,T.Kakuzaki,S.Metsugi,D.Kashiwagi,K.Yoshida,M.Wada,H.Tsunoda,and R.Teramoto,“Antibody design using LSTM based deep generative model from phage display library for affinity maturation,” Scientific Reports,vol.11,p.5852,Mar 2021.
[0448] 17.E.C.Alley,G.Khimulya,S.Biswas,M.AlQuraishi,and G.M.Church,“Unified rational protein engineering with sequence-only deep representation learning,” Nature Methods,vol.12,pp.1315-1322,Mar 2019.
[0449] 18.Z.Ren,J.Li,F.Ding,Y.Zhou,J.Ma,and J.Peng,“Proximal exploration for model-guided protein sequence design,” bioRxiv doi 10.1101 / 2022.4.12.487986,2022.
[0450] 19.L.A.Rabia,A.A.Desai,H.S.Jhajj,and P.M.Tessier,“Understanding and overcoming trade-offs between antibody affinity,specificity,stability and solubility,” Biochemical engineering journal,vol.137,pp.365-374,Sep 2018.
[0451] 20.J.Liu,“Activity-specific cell enrichment,” Patent Publication No.WO 2021 / 146626,22.7.2021.
[0452] 21.R.Akbar,P.A.Robert,M.Pavlovio,J.R.Jeliazkov,I.Snapkov,A.Slabodkin,C.R.Weber,L.Scheffer,E.Miho,I.H.Haff,D.T.T.Haug,F.Lund-Johansen,Y.Safonova,G.K.Sandve,and V.Greiff,“A compact vocabulary of paratope-epitope interactions enables predictability of antibody-antigen binding,” Cell Reports,vol.34,p.108856,Mar 2021.
[0453] 22.J.Bostrom,S.-F.Yu,D.Kan,B.A.Appleton,C.V.Lee,K.Billeci,W.Man,F.Peale,S.Ross,C.Wiesmann,and G.Fuh,“Variants of the antibody herceptin that interact with HER2 and VEGF at the antigen binding site,” Science,vol.323,pp.1610-1614,Mar 2009.
[0454] 23.S.Biswas,G.Khimulya,E.C.Alley,K.M.Esvelt,and G.M.Church,“Low-n protein engineering with data-efficient deep learning,” Nature Methods,vol.18,pp.389-396,Apr 2021.
[0455] 24.A.Burkovitz,O.Leiderman,I.Sela-Culang,G.Byk,and Y.Ofran,“Computational identification of antigen-binding antibody fragments,” The Journal of Immunology,vol.190,pp.2327-2334,Jan 2013.
[0456] 25.V.C.Xie,J.Pu,B.P.Metzger,J.W.Thornton,and B.C.Dickinson,“Contingency and chance erase necessity in the experimental evolution of ancestral proteins,” eLife,vol.10,Jun 2021.
[0457] 26.P.C.Phillips,“Epistasis-the essential role of gene interactions in the structure and evolution of genetic systems,” Nature Reviews Genetics,vol.9,pp.855-867,Nov 2008.
[0458] 27.A.M.Phillips,K.R.Lawrence,A.Moulana,T.Dupic,J.Chang,M.S.Johnson,I.Cvijovic,T.Mora,A.M.Walczak,and M.M.Desai,“Binding affinity landscapes constrain the evolution of broadly neutralizing anti-influenza antibodies,” eLife,vol.10,p.e71393,Sep 2021.
[0459] 28.C.Marks,A.M.Hummer,M.Chin,and C.M.Deane,“Humanization of antibodies using a machine learning approach on large-scale repertoire data,” Bioinformatics,vol.37,pp.4041-4047,Jun 2021.
[0460] 29.G.Liu,H.Zeng,J.Mueller,B.Carter,Z.Wang,J.Schilz,G.Horny,M.E.Birnbaum,S.Ewert,and D.K.Gifford,“Antibody complementarity determining region design using high-capacity machine learning,” Bioinformatics,vol.36,pp.2126-2133,Nov 2019.
[0461] 30.T.Jain,T.Sun,S.Durand,A.Hall,N.R.Houston,J.H.Nett,B.Sharkey,B.Bobrowicz,I.Caffry,Y.Yu,Y.Cao,H.Lynaugh,M.Brown,H.Baruah,L.T.Gray,E.M.Krauland,Y.Xu,M.Vasquez,and K.D.Wittrup,“Biophysical properties of the clinical-stage antibody landscape,” Proceedings of the National Academy of Sciences,vol.114,pp.944-949,Jan 2017.
[0462] 31.M.I.J.Raybould,C.Marks,K.Krawczyk,B.Taddese,J.Nowak,A.P.Lewis,A.Bujotzek,J.Shi,and C.M.Deane,“Five computational developability guidelines for therapeutic antibody profiling,” Proceedings of the National Academy of Sciences,vol.116,pp.4025-4030,Feb 2019.
[0463] 32.R.M.Adams,T.Mora,A.M.Walczak,and J.B.Kinney,“Measuring the sequence-affinity landscape of antibodies with massively parallel titration curves,” eLife,vol.5,p.e23156,Dec 2016.
[0464] 33.L.L.Reich,S.Dutta,and A.E.Keating,“SORTCERY-a high-throughput method to affinity rank peptide ligands,” Journal of Molecular Biology,vol.427,pp.2135-50,Jun 2015.
[0465] 34.C.E.Z.Chan,A.P.C.Lim,P.A.MacAry,and B.J.Hanson,“The role of phage display in therapeutic antibody discovery,” International Immunology,vol.26,pp.649-657,Aug 2014.
[0466] 35.I.T.Nakamura Y,Gojobori T,“Codon usage tabulated from international DNA sequence databases:status for the year 2000,” Nucleic Acids Research,vol.28,p.292,Jan 2000.
[0467] 36.T.Magoc and S.L.Salzberg,“FLASH:fast length adjustment of short reads to improve genome assemblies,” Bioinformatics,vol.27,pp.2957-2963,Sep 2011.
[0468] 37.L.D.,“Algorithms for efficiently collapsing reads with unique molecular identifiers,” PeerJ,vol.7,p.e8275,Dec 2019.
[0469] 38.M.Martin,“Cutadapt removes adapter sequences from high-throughput sequencing reads,” EMBnet.journal,vol.17,May 2011.
[0470] 39.P.J.Cock,T.Antao,J.T.Chang,B.A.Chapman,C.J.Cox,A.Dalke,I.Friedberg,T.Hamelryck,F.Kauff,B.Wilczynski,and M.J.L.de Hoon,“Biopython:freely available python tools for computational molecular biology and bioinformatics,” Bioinformatics,vol.25,pp.1422-1423,Jun 2009.
[0471] 40.S.Andrews,“FastQC.A quality control tool for high throughput sequence data.” Babraham Bioinformatics,Babraham Institute,Cambridge,United Kingdom,https: / / www.bibsonomy.org / bibtex / 2b6052877491828ab53d3449be9b293b3 / ozborn,2010.
[0472] 41.P.Ewels,M.Magnusson,S.Lundin,and M.Kaller,“MultiQC:summarize analysis results for multiple tools and samples in a single report,” Bioinformatics,vol.32,pp.3047-3048,Jun 2016.
[0473] 42.R Core Team,“R:A language and environment for statistical computing.” R Foundation for Statistical Computing,Vienna,Austria,https: / / www.R-project.org,2021.
[0474] 43.T.V.Elzhov,K.M.Mullen,A.-N.Spiess,and B.Bolker,minpack.lm:R Interface to the Levenberg-Marquardt Nonlinear Least-Squares Algorithm Found in MINPACK,Plus Support for Bounds.https: / / cran.r-project.org / web / packages / minpack.lm / minpack.lm.pdf,2022.
[0475] 44.J.J.More,“The Levenberg-Marquardt algorithm:Implementation and theory,” in Lecture Notes in Mathematics,pp.105-116,Springer Berlin Heidelberg,1978.
[0476] 45.J.J.More,B.S.Garbow,and K.E.Hillstrom,Implementation Guide for MINPACK-1.https: / / www.osti.gov / biblio / 5171554,1980.
[0477] 46.K.Levenberg,“A method for the solution of certain non-linear problems in least squares,” Quarterly of applied mathematics,vol.2,pp.164-168,Jul 1944.
[0478] 47.D.W.Marquardt,“An algorithm for least-squares estimation of nonlinear parameters,” Journal of the society for Industrial and Applied Mathematics,vol.11,no.2,pp.431-441,1963.
[0479] 48.T.H.Olsen,F.Boyles,and C.M.Deane,“Observed Antibody Space:A diverse database of cleaned,annotated,and translated unpaired and paired antibody sequences,” Protein Science,vol.31,pp.141-146,Jan 2022.
[0480] 49.M.-P.Lefranc,C.Pommie,Q.Kaas,E.Duprat,N.Bosc,D.Guiraudou,C.Jean,M.Ruiz,I.Da Piedade,M.Rouard,E.Foulquier,V.Thouvenin,and G.Lefranc,“IMGT unique numbering for immunoglobulin and T cell receptor constant domains and Ig superfamily C-like domains,” Developmental & Comparative Immunology,vol.29,pp.185-203,Mar 2005.
[0481] 50.K.Abhinandan and A.C.Martin,“Analysis and improvements to Kabat and structurally correct numbering of antibody variable domains,” Molecular Immunology,vol.45,pp.3832-3839,Aug 2008.
[0482] 51.R.Rao,N.Bhattacharya,N.Thomas,Y.Duan,X.Chen,J.Canny,P.Abbeel,and Y.S.Song,“Evaluating Protein Transfer Learning with TAPE,” in Neural Information Processing Systems,vol.32,pp.9689--9701,Cold Spring Harbor Laboratory,Jun 2019.
[0483] 52.A.Elnaggar,M.Heinzinger,C.Dallago,G.Rihawi,Y.Wang,L.Jones,T.Gibbs,T.Feher,C.Angerer,M.Steinegger,D.Bhowmik,and B.Rost,“ProtTrans:Towards Cracking the Language of Life’s Code Through Self-Supervised Deep Learning and High Performance Computing,” bioRxiv doi:10.48550 / arXiv.2007.06225,2020.
[0484] 53.A.Rives,J.Meier,T.Sercu,S.Goyal,Z.Lin,J.Liu,D.Guo,M.Ott,C.L.Zitnick,J.Ma,and R.Fergus,“Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences,” Proceedings of the National Academy of Sciences,vol.118,no.15,p.e2016239118,2021.
[0485] 54.J.Meier,R.Rao,R.Verkuil,J.Liu,T.Sercu,and A.Rives,“Language models enable zero-shot prediction of the effects of mutations on protein function,” in Advances in Neural Information Processing Systems(A.Beygelzimer,Y.Dauphin,P.Liang,and J.W.Vaughan,eds.),vol.34,pp.29287-29303,2021.
[0486] 55.R.M.Rao,J.Liu,R.Verkuil,J.Meier,J.Canny,P.Abbeel,T.Sercu,and A.Rives,“MSA transformer,” in Proceedings of the 38th International Conference on Machine Learning(M.Meila and T.Zhang,eds.),vol.139 of Proceedings of Machine Learning Research,pp.8844-8856,PMLR,Jul 2021.
[0487] 56.Y.Liu,M.Ott,N.Goyal,J.Du,M.Joshi,D.Chen,O.Levy,M.Lewis,L.Zettlemoyer,and V.Stoyanov,“RoBERTa:A robustly optimized BERT pretraining approach,” arXiv:1907.11692 [cs.CL],2019.
[0488] 57.T.Wolf,L.Debut,V.Sanh,J.Chaumond,C.Delangue,A.Moi,P.Cistac,T.Rault,R.Louf,M.Funtowicz,J.Davison,S.Shleifer,P.von Platen,C.Ma,Y.Jernite,J.Plu,C.Xu,T.L.Scao,S.Gugger,M.Drame,Q.Lhoest,and A.M.Rush,“Huggingface’s transformers:State-of-the-art natural language processing,” arXiv:1910.03771 [cs.CL],2019.
[0489] 58.N.S.Keskar,B.McCann,L.R.Varshney,C.Xiong,and R.Socher,“CTRL:A conditional transformer language model for controllable generation,” arXiv:1909.05858 [cs.CL],2019.
[0490] 59.Y.You,J.Li,S.Reddi,J.Hseu,S.Kumar,S.Bhojanapalli,X.Song,J.Demmel,K.Keutzer,and C.-J.Hsieh,“Large batch optimization for deep learning:Training bert in 76 minutes,” arXiv:1904.00962 [cs.LG],2019.
[0491] 60.I.Loshchilov and F.Hutter,“Fixing weight decay regularization in Adam,” https: / / openreview.net / forum?id=rk6qdGgCZ,2018.
[0492] 61.T.Chen and C.Guestrin,“XGBoost:A scalable tree boosting system,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,KDD ’16,(New York,NY,USA),pp.785-794,ACM,2016.
[0493] 62.R.D.Team,RAPIDS:Collection of Libraries for End to End GPU Data Science,2018.
[0494] 63.R.J.G.B.Campello,D.Moulavi,and J.Sander,“Density-based clustering based on hierarchical density estimates,” in Advances in Knowledge Discovery and Data Mining,pp.160-172,Springer Berlin Heidelberg,2013.
[0495] 64.A.Tareen and J.B.Kinney,“Logomaker:beautiful sequence logos in Python,” Bioinformatics,vol.36,pp.2272-2274,Dec 2019.
[0496] 65. J. Salazar, D. Liang, T. Q. Nguyen, and K. Kirchhoff, “Masked language model scoring,” arXiv:1910.14659 [cs.CL], 2019.
[0497] 66. M. I. J. Raybould, C. Marks, K. Krawczyk, B. Taddese, J. Nowak, A. P. Lewis, A. Bujotzek, J. Shi, and C. M. Deane, “Five computational developability guidelines for therapeutic antibody profiling,” Proceedings of the National Academy of Sciences, vol. 116, pp. 4025 - 4030, Feb 2019.
[0498] 67. F.-A. Fortin, F.-M. De Rainville, M.-A. Gardner, M. Parizeau, and C. Gagne, “DEAP: Evolutionary algorithms made easy,” Journal of Machine Learning Research, vol. 13, pp. 2171 - 2175, Jul 2012.
[0499] 68. H. Beyer and H. Schwefel, “Evolution strategies - a comprehensive introduction,” Natural Computing, vol. 1, pp. 3 - 52, Mar 2002.
[0500] Supplementary information Tite-seq CR9114 dataset Dataset processing The Tite-seq CR9114 dataset
[27] contains affinity data for 65,091 variants of the CR9114 bnAb heavy chain against three different influenza hemagglutinin (HA) antigen subtypes (H1, H3, and FluB). Each variant contains binary mutations at up to 16 positions based on differences between the CR9114 germline and somatic sequences.
[0501] Of the 65,091 variants, 63,419 (97%) were H1-binding, 7,174 (11%) were H3-binding, and 198 (0.3%) were FluB-binding. This technique may involve downloading a dataset from https: / / cdn.elifesciences.org / articles / 71393 / elife-71393-fig1-data1-v3.csv. The amino acid sequence of each variant was inferred from the binary mutation information using a custom Python script using germline sequences. JPEG2025504384000014.jpg35163
[0502] Sequence (16 somatic mutations highlighted in red, and trimmed sequence shown in strikethrough). The first 19 amino acids may be trimmed for compatibility with the NF heavy model (starting at the 21st amino acid in the IMGT numbering scheme).
[0503] Model Architecture An NF heavy model may be used, initialized with weights trained on the OAS dataset (PT), as well as a model initialized with random weights (NPT). To support predictions for all three antigen targets, the sum of mean squared errors may be used for each regression task in some embodiments as a loss function for regression-only models (Reg). Models may also be trained using a mixture model that combines classification and regression tasks in a joint model (mixture). For example, the loss function for the mixture model may be defined as follows:
[0504]
number
[0505] where C is the number of targets (e.g., 3), N is the number of training examples in the batch (e.g., 256),
[0506] y n,c is the measured affinity of sample n to target c,
number
number
[0507] Model Training In some embodiments, the present technology may involve training four types of models (Reg-PT, Reg-BMS, Mix-PT, and Mix-BMS) using three training set sizes (10%, 1%, and 0.1% of 65,091), each using 10 cross-validation folds. For the 1% and 0.1% experiments, the 10 folds may be randomly selected, requiring each fold to contain at least one positive and one negative example for each target in the training set. To support early stopping and classifier calibration, 10% of each test set may be allocated as a separate validation set. In some embodiments, transfer learning may be used to leverage OAS pre-trained models by adding a dense hidden layer with a large number of nodes (e.g., 768), followed by a projection layer with the required number of outputs. All layers may remain unfrozen to update all model parameters during training. Training is performed using the AdamW optimizer, with (e.g.,) 10 -5 The training was performed with a learning rate of 0.01, a weight decay of 0.01, a dropout probability of 0.2, a linear learning rate decay with a warming step of 100, and a batch size of 256. All models were trained until the loss on the validation set stopped improving for a number of epochs (e.g., 50, 250, 2500) for each training size (e.g., 10%, 1%, and 0.1%), respectively.
[0508] Model evaluation Unlike typical cross-validation experiments, the training set may be smaller than the test set, and therefore each variant may be present in multiple test sets. For each variant, predictions may be randomly selected from a single model instead of using the average predicted value to avoid introducing ensemble effects. During inference on the test set, the predicted regression value may be calculated as follows:
number
[0509] The raw classification logits may be converted to probabilities and calibrated using the calibrated classifier CV class of scikit-learn using cv="prefit" and method="isotonic", e.g., classification metrics may in some embodiments be calculated using the scikit-learn functions balanced_accuracy_score, f1_score, precision_score, recall_score, and average_precision_score. Specifically, balanced accuracy is defined as:
number
[0510] The average precision is defined as follows:
number
[0511] where P n and R n are the precision and recall at the n-th threshold on the precision-recall curve.
[0512] antibody As used herein, the term "antibody" refers to any antibody that interacts with an epitope on a target antigen (e.g., by binding, steric hindrance, stabilization / destabilization, spatial distribution). Naturally occurring "antibodies" are glycoproteins comprising at least two heavy (H) chains and two light (L) chains interconnected by disulfide bonds. Each heavy chain is composed of a heavy chain variable region (abbreviated herein as VH) and a heavy chain constant region. The heavy chain constant region is composed of three domains: CH1, CH2, and CH3. Each light chain is composed of a light chain variable region (abbreviated herein as VL) and a light chain constant region. The light chain constant region is composed of one domain: CL. The VH and VL regions can be further subdivided into regions of hypervariability called complementarity-determining regions (CDRs), interspersed with more conserved regions called framework regions (FRs). Each VH and VL is composed of three CDRs and four FRs, arranged from the amino terminus to the carboxy terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4. The variable regions of the heavy and light chains contain a binding domain that interacts with an antigen. The constant region of the antibody can mediate the binding of the immunoglobulin to host tissues or factors, including various cells of the immune system (e.g., effector cells) and the first component (Clq) of the classical complement system. The term "antibody" includes, for example, monoclonal antibodies, human antibodies, humanized antibodies, camelized antibodies, chimeric antibodies, single-chain Fvs (scFvs), disulfide-linked Fvs (sdFvs), Fab fragments, F(ab') fragments, and anti-idiotypic (anti-Id) antibodies (e.g., including anti-Id antibodies against an antibody of the present invention), as well as epitope-binding fragments of any of the above. The antibody can be of any isotype (e.g., IgG, IgE, IgM, IgD, IgA, and IgY), class (e.g., IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2), or subclass. An antibody or epitope-binding fragment can be, or be a component of, a multispecific molecule.
[0513] Both the light and heavy chains are divided into regions of structural and functional homology. The terms "constant" and "variable" are used functionally. In this regard, it will be understood that the variable domains of both the light (VL) and heavy (VH) chain portions determine antigen recognition and specificity. Conversely, the constant domains of the light (CL) and heavy chains (CH1, CH2, or CH3) confer important biological properties such as secretion, transplacental transfer, Fc receptor binding, and complement fixation. By convention, the numbering of constant region domains increases as they become more distal from the antigen-binding site or amino-terminus of the antibody. The N-terminus is the variable region, and the constant region is at the C-terminus. The CH3 and CL domains actually comprise the carboxy-terminus of the heavy and light chains, respectively.
[0514] The term "antibody fragment," as used herein, refers to one or more portions of an antibody that retain the ability to specifically interact with a target epitope (e.g., by binding, steric hindrance, stabilization / destabilization, spatial distribution). Examples of binding fragments include, but are not limited to, a Fab fragment, which is a monovalent fragment consisting of VL, VH, CL, and CH1; a F(ab)2 fragment, which is a bivalent fragment comprising two Fab fragments linked by a disulfide bridge at the hinge region; an Fd fragment consisting of the VH and CH1 domains; an Fv fragment consisting of the VL and VH domains of a single antibody arm; a dAb fragment consisting of the VH domain (Ward et al., (1989) Nature 341:544-546); and isolated complementarity-determining regions (CDRs). Furthermore, although the two domains of an Fv fragment, VL and VH, are encoded by separate genes, they can be joined using recombinant methods by a synthetic linker that allows them to be produced as a single protein chain in which the VL and VH regions pair to form a monovalent molecule (known as a single-chain Fv (scFv); see, e.g., Bird et al., (1988) Science 242:423-426; and Huston et al., (1988) Proc. Natl. Acad. Sci. 85:5879-5883). Such single-chain antibodies are also intended to be encompassed within the term "antibody fragment." These antibody fragments are obtained using conventional techniques known to those skilled in the art, and the fragments are screened for utility in the same manner as intact antibodies.
[0515] As described herein, antibodies may include biologically active derivatives, variants, or fragments. As used herein, a "biologically active derivative" or "biologically active variant" includes any derivative or variant of an antibody that has substantially the same functional and / or biological properties, such as binding properties, as the antibody (e.g., a wild-type antibody), and / or the same structural basis, such as a peptide backbone or basic polymer unit including a framework region.
[0516] As will be understood by those skilled in the art, an "analog," such as a "variant" or "derivative," is an antibody that is substantially similar in structure to a native antibody or a wild-type antibody or another reference antibody, and that has the same biological activity, albeit, in certain cases, to a different extent. For example, an antibody variant refers to an antibody that shares a substantially similar structure with a reference antibody and has the same biological activity. Variants or analogs differ in their amino acid sequence composition compared to the reference antibody from which the analog is derived, based on one or more mutations, including (i) deletion of one or more amino acid residues at one or more termini of the antibody and / or at one or more internal regions (e.g., fragments) of the antibody, (ii) insertion or addition of one or more amino acids at one or more termini of the antibody (typically "additions" or "fusions") and / or at one or more internal regions (typically "insertions") of the antibody sequence, or (iii) substitution of one or more amino acids for other amino acids in the antibody sequence. By way of example, a "derivative" is a type of analog, and refers to an antibody that shares the same or substantially similar structure as a reference antibody that has been chemically modified, for example.
[0517] In some embodiments, the variant or sequence variant is a mutant in which one, two, three, four, five, six, or more amino acids in one or more CDRs are mutated compared to the reference antibody. In some embodiments, the CDRs on the light chain, the heavy chain, or both the heavy and light chains are mutated. In some embodiments, one or more framework amino acid residues are mutated compared to the reference antibody.
[0518] In substitutional variants, one or more amino acid residues of an antibody, for example, in the CDR region, are removed and replaced with alternative residues. In one aspect, these substitutions are conservative in nature, and conservative substitutions of this type are well known in the art. Alternatively, the present disclosure encompasses substitutions that are also non-conservative. Exemplary conservative substitutions are described in Lehninger, Biochemistry, 2nd Edition; Worth Publishers, Inc., New York (1975), pp. 71-77.
[0519] Antibodies contemplated herein include full-length antibodies, biologically active subunits or fragments of full-length antibodies, and biologically active derivatives and variants of any of these forms of therapeutic proteins. Thus, antibodies include those having (1) an amino acid sequence with greater than about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% or more amino acid sequence identity to a reference antibody (e.g., encoded by a reference nucleic acid or amino acid sequence described herein) over a region of at least about 25, about 50, about 100, about 200, about 300, about 400, or more amino acids. According to the present disclosure, the term "recombinant protein" or "recombinant antibody" includes any protein obtained via recombinant DNA technology. In certain embodiments, the term encompasses the antibodies described herein.
[0520] In some embodiments, the antibodies or antibody variants described herein are expressed from one or more expression constructs and / or in the cells or lines described herein.
[0521] Exemplary wild-type or reference antibodies include commercially available antibodies or other known antibodies, including therapeutic monoclonal antibodies. Reference antibodies according to the present disclosure may include any antibody currently known or later developed, including those that are not clinically and / or commercially available.
[0522] Cells and expression constructs cell Antibodies of the present disclosure include wild-type (WT) antibodies and variant antibodies, and in some embodiments are produced in cells. Cells comprising one or more of the expression constructs described herein are contemplated in various embodiments of the present disclosure.
[0523] Prokaryotic Host Cells: In some embodiments of the present disclosure, expression constructs designed for expression of gene products, including the fusion proteins described herein, are provided in host cells, such as prokaryotic host cells. Prokaryotic host cells include those derived from archaea (e.g., Haloferax volcanii, Sulfolobus solfataricus), gram-positive bacteria (e.g., Bacillus subtilis, Bacillus licheniformis, Brevibacillus choshinensis, Lactobacillus brevis, Lactobacillus buchneri, Lactococcus lactis, and Streptomyces lividans), or alphaproteobacteria (e.g., Agrobacterium tumefaciens). tumefaciens, Caulobacter crescentus, Rhodobacter sphaeroides, and Sinorhizobium meliloti), Betaproteobacteria (Alcaligenes eutrophus), and Gammaproteobacteria (Acinetobacter calcoaceticus, Azotobacter vinelandii, Escherichia coli, Pseudomonas aeruginosa, and Pseudomonas putida).Preferred host cells include gammaproteobacteria of the Enterobacteriaceae family, such as Enterobacter, Erwinia, Escherichia coli (including E. coli), Klebsiella, Proteus, Salmonella (including Salmonella typhimurium), Serratia (including Serratia marcescans), and Shigella.
[0524] Eukaryotic Host Cells. Many additional types of host cells can be used in the expression systems of the present disclosure, including eukaryotic cells such as yeast (Candida shehatae, Kluyveromyces lactis, Kluyveromyces fragilis, other Kluyveromyces species, Pichia pastoris, Saccharomyces cerevisiae, Saccharomyces pastorianus, also known as Saccharomyces carlsbergensis, Schizosaccharomyces pombe, and the like). pombe, Dekkera / Brettanomyces, and Yarrowia lipolytica; other fungi (Aspergillus nidulans, Aspergillus niger, Neurospora crassa, Penicillium, Tolypocladium, Trichoderma lesia, etc.); reesia); insect cell lines (Drosophila melanogaster Schneider 2 cells and Spodoptera frugiperda Sf9 cells); and mammalian cell lines, immortalized cell lines (Chinese hamster ovary (CHO) cells, HeLa cells, baby hamster kidney (BHK) cells, monkey kidney (COS) cells, human embryonic kidney (HEK, 293, or HEK-293) cells, and human hepatocellular carcinoma (Hep G2) cells). The above host cells are available from the American Type Culture Collection.
[0525] As described in WO / 2017 / 106583, the entirety of which is incorporated herein by reference, the production of gene products, such as therapeutic proteins, on a commercial scale and in soluble form is addressed by providing suitable host cells that can grow to high cell densities in fermentation cultures and produce soluble gene products in the cytoplasm of oxidizing host cells through highly controlled, inducible gene expression. Host cells of the present disclosure with these qualities are produced by combining some or all of the following features: (1) The host cell is genetically modified to have an oxidizing cytoplasm through increasing the expression or function of oxidized polypeptides in the cytoplasm and / or decreasing the expression or function of reduced polypeptides in the cytoplasm. Specific examples of such genetic modifications are provided herein. Optionally, the host cell can also be genetically modified to express chaperones and / or cofactors that assist in the production of the desired gene product and / or to glycosylate the polypeptide gene product. (2) The host cell comprises one or more expression constructs designed for expression of one or more gene products of interest, and in certain embodiments, at least one expression construct comprises an inducible promoter and a polynucleotide encoding a gene product expressed from the inducible promoter. (3) The host cell comprises additional genetic modifications designed to improve certain aspects of gene product expression from the expression constructs.In certain embodiments, the host cell has (A) a modification of gene function of at least one gene encoding a transporter protein for an inducer of at least one inducible promoter, where, as another example, the gene encoding the transporter protein is selected from the group consisting of araE, araE, araG, araH, rhaT, xylF, xylG, and xylH, or in particular araE, or more specifically, the modification of gene function is expression of araE from a constitutive promoter, and / or (B) a modification of gene function of at least one gene encoding a protein that metabolizes an inducer of at least one of the inducible promoters. and / or (C) a reduced level of gene function of at least one gene encoding a protein involved in the biosynthesis of an inducer of at least one inducible promoter, which in further embodiments is selected from the group consisting of scpA / sbm, argK / ygfD, scpB / ygfG, scpC / ygfH, rmlA, rmlB, rmlC, and rmlD.
[0526] A host cell having an oxidized cytoplasm. The expression system of the present disclosure is designed to express a gene product, and in certain embodiments of the present disclosure, the gene product is expressed in the host cell. Examples of host cells are provided that allow for efficient and cost-effective expression of gene products, including components of multimeric products. Host cells can include isolated cells in culture as well as cells that are part of a multicellular organism or cells grown in different organisms or biological systems. In certain embodiments of the present disclosure, the host cell is a microbial cell, such as a yeast (e.g., Saccharomyces, Schizosaccharomyces), or a bacterial cell, or a Gram-positive or Gram-negative bacterium, or E. coli, or an E. coli B strain, or an E. coli (B strain) EB0001 cell (also referred to as an E. coli ASE (DGH) cell), or an E. coli (B strain) EB0002 cell. Growth experiments with E. coli host cells with an oxidizing cytoplasm, specifically the E. coli B strains SHuffle® Express (NEB Catalog No. C3028H) and SHuffle® T7 Express (NEB Catalog No. C3029H), and the E. coli K strain SHuffle® T7 (NEB Catalog No. C3026H), have shown that these E. coli B strains with an oxidizing cytoplasm can grow to much higher cell densities than their closest corresponding E. coli K strain (WO / 2017 / 106583).
[0527] Modifications to Host Cell Gene Function. Certain modifications can be made to gene function in host cells containing an inducible expression construct to facilitate efficient and uniform induction of a host cell population by an inducer. Preferably, the combination of expression construct, host cell genotype, and induction conditions results in at least 75% (more preferably at least 85%, and most preferably at least 95%) of the cells in culture expressing the gene product from each induced promoter, as measured by the method of Khlebnikov et al., described in Example 9 of WO / 2017 / 106583. For host cells other than E. coli, these modifications can involve the function of a gene structurally similar to an E. coli gene or a gene that performs a function in the host cell similar to the function of an E. coli gene. Modifications to host cell gene function include eliminating or reducing gene function by deleting the gene protein-coding sequence in its entirety or by deleting a sufficiently large portion of the gene, inserting a sequence into the gene, or otherwise modifying the gene sequence so that a reduced level of functional gene product is produced from that gene. Modifications to host cell gene function also include increasing gene function, for example, by modifying a native promoter to create a stronger promoter that drives higher levels of transcription of the gene, or by introducing missense mutations in the protein coding sequence that result in a more highly active gene product. Modifications to host cell gene function include altering gene function in any way, including, for example, modifying a native inducible promoter to create a constitutively activated promoter. In addition to modifying gene function for inducer transport and metabolism, as described herein with respect to modifying inducible promoters and / or expression of chaperone proteins, it is also possible to modify the reduced-oxidative environment of the host cell.
[0528] Host cell reducing and oxidizing environment. In bacterial cells, such as Escherichia coli, proteins requiring disulfide bonds are typically transported into the periplasm, where disulfide bond formation and isomerization are catalyzed by the Dsb system, including DsbABCD and DsbG. Increased expression of the cysteine oxidase DsbA, the disulfide isomerase DsbC, or a combination of Dsb proteins, all of which are normally transported into the periplasm, has been utilized in the expression of heterologous proteins requiring disulfide bonds (Makino et al., Microb Cell Fact 2011 May 14;10:32). It is also possible to express cytoplasmic forms of these Dsb proteins, such as cytoplasmic versions of DsbA and / or DsbC ('cDsbA' or 'cDsbC'), which lack signal peptides and therefore are not transported into the periplasm. Cytoplasmic Dsb proteins, such as cDsbA and / or cDsbC, are useful for making the host cell cytoplasm more oxidative, thus contributing to disulfide bond formation in heterologous proteins produced in the cytoplasm. The host cell cytoplasm can also be made less reducing and thus more oxidative by directly modifying the thioredoxin and glutaredoxin / glutathione enzyme systems: mutants defective in glutathione reductase (gor) or glutathione synthetase (gshB), together with thioredoxin reductase (trxB), confer cytoplasmic oxidation. These strains cannot reduce ribonucleotides and therefore cannot grow in the absence of exogenous reducing agents such as dithiothreitol (DTT).Suppressor mutations in the gene ahpC encoding the peroxiredoxin AhpC (e.g., ahpC* and ahpCA; Lobstein et al., Microb Cell Fact 2012 May 8;11:56; doi:10.1186 / 1475-2859-11-56) convert it into a disulfide reductase that generates reduced glutathione, allowing electron channeling onto the enzyme ribonucleotide reductase, allowing cells deficient in gor and trxB, or cells deficient in gshB and trxB, to grow in the absence of DTT. Different classes of mutant AhpC can allow strains deficient in gamma-glutamylcysteine synthetase (gshA) activity and deficient in trxB to grow in the absence of DTT. These include AhpC V164G, AhpC S71F, AhpC E173 / S71F, AhpC E171Ter, and AhpC dupl62-169 (Faulkner et al., Proc Natl Acad Sci USA 2008 May 6;105(18):6735-6740, Epub 2008 May 2). In such strains with an oxidative cytoplasm, exposed protein cysteines are readily oxidized, resulting in disulfide bond formation, in a process catalyzed by thioredoxin, reversing their physiological function. Other proteins that may help reduce the effects of oxidative stress in oxidative cytoplasmic host cells are HPI (hydroperoxidase I) catalase peroxidase, encoded by E. coli katG, and HPII (hydroperoxidase II) catalase peroxidase, encoded by E. coli katE, which dismutate peroxides to water and O2 (Farr and Kogoma, Microbiol Rev. 1991 Dec;55(4):561-585; Review). Increasing the levels of KatG and / or KatE proteins in host cells through induced co-expression or through increased levels of constitutive expression is an aspect of some embodiments of the present disclosure.
[0529] Another modification that can be made to the host cell is to express the sulfhydryl oxidase Ervlp from the inner membrane space of yeast mitochondria in the cytoplasm of the host cell, which has been shown to increase production of a variety of complex disulfide-bonded proteins of eukaryotic origin in the cytoplasm of E. coli, even in the absence of mutations in gor or trxB (Nguyen et al, Microb Cell Fact 2011 Jan 7;10:1).
[0530] Host cells containing the expression construct preferably also express cDsbA and / or cDsbC and / or Ervlp, are deficient in trxB gene function, are deficient in any of gor, gshB, or gshA gene function, optionally have increased levels of katG and / or katE gene function, and express an appropriate mutant form of AhpC such that the host cells can grow in the absence of DTT.
[0531] Chaperones. In some embodiments, a desired gene product is co-expressed with other gene products, e.g., chaperones, that are beneficial to the production of the desired gene product. Chaperones are proteins that assist in the folding or unfolding and / or assembly or disassembly of other gene products, but do not occur in the resulting monomeric or multimeric gene product structures when the structures are performing their normal biological function (completing the folding and / or assembly process). Chaperones can be expressed from an inducible or constitutive promoter in an expression construct, or can be expressed from the host cell chromosome; preferably, expression of the chaperone protein in the host cell is at a level high enough to produce the co-expressed gene products that are properly folded and / or assembled into the desired product. Examples of chaperones present in E. coli host cells include the folding factors DnaK / DnaJ / GrpE, DsbC / DsbG, GroEL / GroES, IbpA / IbpB, Skp, Tig (trigger factor), and FkpA, which have been used to prevent protein aggregation of cytoplasmic or periplasmic proteins. Because DnaK / DnaJ / GrpE, GroEL / GroES, and ClpB can function synergistically in supporting protein folding, expression of these chaperones in combination has been shown to be beneficial for protein expression (Makino et al., Microb Cell Fact 2011 May 14;10:32). When expressing eukaryotic proteins in prokaryotic host cells, eukaryotic chaperone proteins, such as protein disulfide isomerase (PDI) from the same or related eukaryotic species, are, in certain embodiments, co-expressed or inducibly co-expressed with the desired gene product.
[0532] One chaperone that can be expressed in host cells is protein disulfide isomerase from the soil nematode (soft-rot fungus) Humicola insolens. The amino acid sequence of Humicola insolens PDI is shown as SEQ ID NO:1 in WO / 2017 / 106583. It lacks the native protein signal peptide, allowing it to remain in the host cytoplasm. The nucleotide sequence encoding PDI was optimized for expression in E. coli, and the PDI expression construct is shown as SEQ ID NO:2 in WO / 2017 / 106583. SEQ ID NO:2 contains a GCTAGC Nhel restriction site at its 5' end, an AGGAGG ribosome binding site at nucleotides 7-12, the PDI coding sequence at nucleotides 21-1478, and a GTCGAC Sail restriction site at its 3' end. The nucleotide sequence of SEQ ID NO:2 was designed to be inserted immediately downstream of a promoter, such as an inducible promoter. The Nhel and Sail restriction sites in SEQ ID NO:2 can be used to insert it into a vector multiple cloning site, such as that of the pSOL expression vector (SEQ ID NO:3 of WO / 2017 / 106583) described in published U.S. patent application US2015353940A1. Other PDI polypeptides can also be expressed in host cells, including PDI polypeptides from various species (Saccharomyces cerevisiae (UniProtKB PI 7967), Homo sapiens (UniProtKB P07237), Mus musculus (UniProtKB P09103), Caenorhabditis elegans (UniProtKB Q17770 and Q17967), Arabdopsis thaliana (UniProtKB 048773, Q9XI01, Q9S G3, Q9LJU2, Q9MAU6, Q94F09, and Q9T042), Aspergillus niger (UniProtKB Q12730), and modified forms of such PDI polypeptides.In certain embodiments of the present disclosure, the PDI polypeptide expressed in a host cell of the present disclosure shares at least 70%, or 80%, or 90%, or 95% amino acid sequence identity over at least 50% (or at least 60%, or at least 70%, or at least 80%, or at least 90%) of the length of SEQ ID NO: 1 of WO / 2017 / 106583, where the amino acid sequence identity is det...
Claims
1. 1. A computing system for identifying biomolecular sequence variants of interest, the computing system comprising: one or more processors; and One or more non-transitory computer-readable media having stored thereon: A machine learning language model trained using training data, comprising: the training data includes one or more training biomolecule sequence variants, each having a respective measured binding property representing each of their abilities to bind to a corresponding respective binding partner; and the machine learning language model is configured to output predicted biomolecular binding properties of input biomolecular sequence variants; and instructions that, when executed by the one or more processors, cause the computing system to: processing one or more biomolecular sequence variants with the machine learning language model to generate one or more predicted binding properties, each corresponding to a respective one of the one or more biomolecular sequence variants; analyzing the one or more predicted binding properties to identify one or more biomolecular sequence variants of interest from among the one or more biomolecular sequence variants, each of the one or more biomolecular sequence variants of interest having one or more desired properties; and and instructions to cause the one or more biomolecular sequence variants of interest to be provided as output.
2. the one or more training biomolecule sequence variants: Less than 10% of the total possible variant space, Less than 9% of the total possible variant space, Less than 8% of the total possible variant space, Less than 7% of the total possible variant space, Less than 6% of the total possible variant space, Less than 5% of the total possible variant space, Less than 4% of the total possible variant space, Less than 3% of the total possible variant space, Less than 2% of the total possible variant space, Less than 1% of the total possible variant space, Less than 0.5% of the total possible variant space, Less than 0.4% of the total possible variant space, Less than 0.3% of the total possible variant space, less than 0.2% of the total possible variant space, or The computing system of claim 1 , comprising less than 0.1% of the total possible variant space.
3. The one or more non-transitory computer-readable media, when executed by the one or more processors, provide the computing system with: at least one of the one or more training biomolecule sequence variants, (i) Observed Antibody Space (OAS) database; (ii) Uniref90 protein database; (iii) any Uniref-derived dataset; (iv) BFD dataset; (v) Mgnify dataset; (vi) any metagenomic dataset derived from a JGI or EBI compendium; (vii) any corpus of assembled protein sequences; or (viii) any dataset of natural antibody sequences obtainable by BCR sequencing or other means.
4. The one or more non-transitory computer-readable media, when executed by the one or more processors, provide the computing system with: retraining the machine learning model using data output by at least one different binding assay corresponding to a different antibody-antigen pair; 10. The computing system of claim 1, further having stored thereon further instructions to cause retraining, wherein the retraining includes generating different sets of antibody-antigen specific weights corresponding to the different antibody-antigen pairs.
5. The one or more non-transitory computer-readable media, when executed by the one or more processors, provide the computing system with:
10. The computing system of claim 1, further having stored thereon further instructions for determining at least one respective measured binding characteristic based on an environmental condition.
6. the non-transitory computer-readable medium comprising: (i) Artificial Neural Networks; (ii) Transformer Neural Network; (iii) Convolutional Neural Networks; (iv) recurrent neural networks; (v) deep learning networks; (vi) autoencoder; (vii) regression model; (viii) plug-and-play language model; (iv) a generative model; or The computing system of claim 1 , further comprising: (x) a genetic algorithm.
7. 1. A computer-implemented method for training a machine learning model to identify biomolecular sequence variants of interest, the method comprising: generating one or more biomolecule sequence variants by programmatically mutating a reference biomolecule; receiving screening data including a ranking of the biomolecular sequence variants according to one or more training binding properties; and and training the machine learning model using the screening data to predict one or more desired binding properties of input biomolecular sequence variants.
8. 1. A computing system for improving accuracy and throughput through predictive denoising, the computing system comprising: one or more processors; and One or more non-transitory computer-readable media having stored thereon: A machine learning model trained using training data, the training data includes one or more training biomolecule sequence variants, each having a respective measured binding property representing each of their abilities to bind to a corresponding respective binding partner; and the machine learning model is configured to output predicted, denoised biomolecular binding properties of one or more training biomolecular sequence variants; and instructions that, when executed by the one or more processors, cause the computing system to: processing the one or more training biomolecular sequence variants with the machine learning model to generate one or more denoised predicted binding properties, each corresponding to a respective one of the one or more training biomolecular sequence variants; and and instructions to cause the one or more training biomolecular sequence variants and their respective denoised predicted binding properties to be provided as output.
9. The computing system of claim 8 , wherein the training biomolecule sequence variants include one or more unsaturated sequence variants.
10. 1. A computing system for predicting the naturalness of a biomolecular sequence variant, the computing system comprising: one or more processors; and One or more non-transitory computer-readable media having stored thereon: A machine learning model trained using training data, the training data includes one or more training biomolecule sequence variants; and the machine learning model is configured to output a predicted naturalness characteristic for each of one or more biomolecular sequence variants; and instructions that, when executed by the one or more processors, cause the computing system to: processing one or more input biomolecular sequence variants with the machine learning model to generate a respective predicted naturalness characteristic for each of the one or more input biomolecular sequence variants; and and instructions to cause at least one of the predicted naturalness characteristics to be provided as an output.
11. The non-transitory computer-readable medium, when executed by the one or more processors, provides the computing system with:
11. The computing system of claim 10, further having stored thereon: instructions for comparing the respective predicted naturalness characteristics for each of the one or more input biomolecular sequence variants with one or both of (i) published phage data and (ii) a therapeutic antibody profiler to determine one or more correlations between at least one respective naturalness characteristic and developability characteristic.
12. The non-transitory computer-readable medium, when executed by the one or more processors, provides the computing system with: generating origin binned data by comparing the respective predicted naturalness characteristics for each of the one or more input biomolecule sequence variants to published naturalness characteristics of therapeutic antibodies administered to humans in Phase I, Phase II, Phase III, or clinical phases using a CDR-only model; 11. The computing system of claim 10, having further instructions stored thereon to determine an immunogenicity score by dividing the origin-binned data according to whether a patient developed an anti-drug antibody response to fully human antibodies.
13. The non-transitory computer-readable medium, when executed by the one or more processors, provides the computing system with:
11. The computing system of claim 10, further having stored thereon further instructions for scoring at least one of (i) the naturalness of a sequence variant as a function of CDRH3 mutational burden, or (ii) the naturalness of trastuzumab.
14. The non-transitory computer-readable medium, when executed by the one or more processors, provides the computing system with:
11. The computing system of claim 10, further storing thereon further instructions to process the one or more biomolecular sequence variants with the machine learning model to generate the one or more predicted binding properties, analyze the one or more predicted binding properties, and identify one or more biomolecular sequence variants of interest from among the one or more biomolecular sequence variants using generation techniques, thereby avoiding exhaustively predicting the affinity of all possible sequence variants in sequence space.
15. The non-transitory computer-readable medium, when executed by the one or more processors, provides the computing system with:
15. The computing system of claim 14, further having stored thereon further instructions for processing the one or more biomolecular sequence variants with the machine learning model to generate the one or more predicted binding properties based on a predicted naturalness of each of the one or more biomolecular sequence variants.