Machine learning-guided generation of cross-reactive neutralizing antigen binding molecules against viral proteins
A machine-learning-guided method optimizes amino acid sequences for antibodies to target conserved viral domains, enhancing cross-reactivity and neutralization efficacy against diverse viral strains by iteratively refining models with experimental data.
Patent Information
- Application Number
- PCT/US2025/032314
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-06
- Filing Date
- 2025-06-04
- Publication Date
- 2025-12-11
AI Technical Summary
Existing methods for generating antibodies against viral proteins often target immunodominant epitopes that are poorly conserved across viral strains, leading to rare and poorly potent neutralization, and lack scalability and predictability in identifying cross-reactive neutralizing antibodies.
A machine-learning-guided approach using sequence-to-function models to generate and optimize candidate amino acid sequences for antibodies that bind to and neutralize both a first target and related targets, leveraging graph neural networks and iterative design cycles to enhance cross-reactivity.
The method effectively identifies antibodies with robust neutralization capabilities against multiple viral strains, including variants like Omicron, by focusing on conserved antigen domains and iteratively improving cross-reactive potency through machine learning and experimental validation.
Smart Images

Figure US2025032314_11122025_PF_FP_ABST
Abstract
Description
MACHINE LEARNING-GUIDED GENERATION OF CROSS-REACTIVE NEUTRALIZING ANTIGEN BINDING MOLECULES AGAINST VIRAL PROTEINSRELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 656,934, filed on June 6, 2024. The entire teachings of the above application are incorporated herein by reference.INCORPORATION BY REFERENCE OF MATERIAL IN XML
[0002] This application incorporates by reference the Sequence Listing contained in the following extensible Markup Language (XML) file being submitted concurrently herewith: File name: 5708.1078001_Sequence_Listing.xml; created June 4, 2025, 11,906 Bytes in size.BACKGROUND
[0003] Antibody response against selected surface viral proteins is important for providing protection from infection. Antibodies frequently target immunodominant epitopes that are often poorly conserved across viral strains. Antibodies targeting conserved antigen domains may be highly valuable in the context of active (e.g., vaccines) or passive (e.g., monoclonal antibodies) immunoprophylaxis; however, antibodies binding to such conserved antigen domains can be rare and / or frequently have poor neutralization potency.SUMMARY
[0004] In some embodiments, provided herein are methods of, and systems for, generating, in silico with a machine-learning model, multiple candidate amino acid sequences based on a reference amino acid sequence.
[0005] In some embodiments, method(s) include generating, in silico with a machinelearning model, multiple candidate amino acid sequences based on a reference amino acid sequence and sequence to function model(s). In some embodiments, a sequence to function model(s) predict binding to and neutralization of a first target. In some embodiments, method(s) include forming a sequence to function model for each in vitro evaluation of function of candidate amino acid sequences. In some embodiments, function(s) include, for example, binding to and neutralization of related targets. In some embodiments, method(s) include selecting a set of candidate amino acid sequences based on function(s) of eachrespective candidate amino acid sequence. In some embodiments, each candidate amino acid sequence in a selected set binds to a first target and related target(s), neutralizes the first target and related target(s), or a combination thereof. In some embodiments, method(s) include retraining a machine-learning model with a selected set of candidate amino acid sequences and at least one respective sequence to function model formed.
[0006] In some embodiments, method(s) may include generating multiple candidate amino acid sequences by determining amino acid substitutions, additions, and / or deletions to a reference amino acid sequence based on a sequence to function model(s).
[0007] In some embodiments, method(s) may include deriving a reference amino acid sequence landscape based on a structure in complex with (e.g., bound to) an antigen using a machine-learning model. In some embodiments, a machine-learning model may be a graph neural network for non-limiting example.
[0008] In some embodiments, retraining a machine-learning model may include updating a sequence to function model(s) with in vitro evaluation measurements associated with a selected set of candidate amino acid sequences.
[0009] In some embodiments, a sequence to function model for a candidate amino acid sequence(s) may measure binding to and / or neutralization of a related target(s).
[0010] In some embodiments, method(s) may include iterating any actions described herein, after retraining, for non-limiting example.
[0011] In some embodiments, selecting a set of amino acid sequences may include presenting a user interface of amino acid sequences and respective scores from a sequence to function model(s) and / or receiving user selections of amino acid sequences.
[0012] In some embodiments, selecting a set of amino acid sequences may include automatically analyzing, in silico, respective scores from a sequence to function model(s) and automatically selecting a set of amino acid sequences.
[0013] In some embodiments, selecting a set of amino acid sequences may include automatically analyzing, in silico, respective scores from a sequence to function model(s), presenting, in a user interface, an indication of recommended amino acid sequences, and / or receiving user input selecting a set of amino acid sequences.
[0014] In some embodiments, reference and candidate amino acid sequences are sequences of antibodies and / or antigen-binding fragments of antibodies.
[0015] In some embodiments, a reference amino acid sequence is in complex with (e.g., bound to) a first target amino acid sequence for non-limiting example.
[0016] In some embodiments, related targets may be targets with a threshold percent sequence identity of a first target or a threshold number of additions, substitutions, and / or deletions from a first target. In some embodiments, a threshold percent sequence identity can be 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%. In some embodiments, a threshold percent sequence identity can be at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, and / or at least 99%. In some embodiments, a threshold percent sequence identity can be at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and / or at least about 99%. In some embodiments, a threshold percent sequence identity can be about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, and / or about 99%. In some embodiments, a number of additions, substitutions, and / or deletions from a first target can be between 1 and 20 additions, substitutions, and / or deletions. In some embodiments, a number of additions, substitutions, or deletions from a first target can be between 1 and 5 additions, substitutions, and / or deletions. In some embodiments, a number of additions, substitutions, and / or deletions from a first target can be between 6 and 10 additions, substitutions, and / or deletions. In some embodiments, a number of additions, substitutions, and / or deletions from a first target can be between 11 and 15 additions, substitutions, and / or deletions. In some embodiments, a number of additions, substitutions, and / or deletions from a first target can be between 16 and 20 additions, substitutions, and / or deletions. In some embodiments, a number of additions, substitutions, and / or deletions from a first target can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and / or 20 additions, substitutions, and / or deletions.
[0017] In some embodiments, a system includes a memory and a processor operatively connected to the memory. In some embodiments, a memory stores instructions thereon that, when loaded and executed, cause a processor to generate, in silico with a machine-learning model, candidate amino acid sequences based on a reference amino acid sequence and sequence to function model(s). In some embodiments, a sequence to function model(s) may measure binding to and / or neutralization of a first target. In some embodiments, instructions may cause a processor to form a sequence to function model for each of a set of candidate amino acid sequences based on a respective in vitro evaluation of function(s) of eachrespective candidate amino acid sequence. In some embodiments, a function includes binding to or neutralization of related targets, or both. In some embodiments, instructions may cause a processor to select a set of candidate amino acid sequences based on a function of each respective candidate amino acid sequence. In some embodiments, each candidate amino acid sequence in a selected set binds to a first target and related target(s), neutralizes a first target and related target(s), or a combination thereof. In some embodiments, instructions may cause a processor to retrain a machine-learning model with a selected set of candidate amino acid sequences and a respective sequence to function model formed.
[0018] In some embodiments, generating candidate amino acid sequences includes determining amino acid substitutions, additions, and / or deletions to a reference amino acid sequence based on a sequence to function model(s).
[0019] In some embodiments, a processor may be configured to generate candidate amino acid sequences using a machine-learning model. In some embodiments, a machine-learning model is a graph neural network for non-limiting example.
[0020] In some embodiments, retraining a machine-learning model may include updating a sequence to function model(s) with a formed sequence to function model(s) associated with a selected set of candidate amino acid sequences.
[0021] In some embodiments, a sequence to function model for each candidate amino acid sequence measure(s) binding to or neutralization of related targets. In some embodiments, a sequence to function model for each evaluated function(s) predicts binding to related targets. In some embodiments, a sequence to function model for each evaluated function(s) predicts a neutralization of related targets. In some embodiments, a sequence to function model for each evaluated function(s) predicts binding to and neutralization of related targets.
[0022] In some embodiments, a processor may be configured to iterate actions, disclosed herein, after retraining for non-limiting example.
[0023] In some embodiments, selecting a set of amino acid sequences may include presenting a user interface of amino acid sequences and respective scores from sequence to function models and / or receiving user selections of a set of amino acid sequences.
[0024] In some embodiments, selecting a set of amino acid sequences may include automatically analyzing, in silico, respective scores from sequence to function models and / or automatically selecting a set of amino acid sequences.
[0025] In some embodiments, selecting a set of amino acid sequences may include automatically analyzing, in silico, respective scores from a sequence to function model(s), presenting, in a user interface, an indication of recommended amino acid sequences, and receiving user input selecting a set of amino acid sequences.
[0026] In some embodiments, reference and candidate amino acid sequences are sequences of antibodies and / or antigen-binding fragments of antibodies for non-limiting examples.
[0027] In some embodiments, a reference amino acid sequence is in complex with a first target amino acid sequence for non-limiting example.
[0028] In some embodiments, related targets are targets with a threshold percent sequence identity of a first target or a threshold number of additions, substitutions, and / or deletions from the first target. In some embodiments, a threshold percent sequence identity can be 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, and / or 99%. In some embodiments, a threshold percent sequence identity can be at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, and / or at least 99%. In some embodiments, a threshold percent sequence identity can be at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and / or at least about 99%. In some embodiments, a threshold percent sequence identity can be about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, and / or about 99%. In some embodiments, a number of additions, substitutions, and / or deletions from a first target can be between 1 and 20 additions, substitutions, and / or deletions. In some embodiments, a number of additions, substitutions, and / or deletions from a first target can be between 1 and 5 additions, substitutions, and / or deletions. In some embodiments, a number of additions, substitutions, and / or deletions from a first target can be between 6 and 10 additions, substitutions, and / or deletions. In some embodiments, a number of additions, substitutions, and / or deletions from a first target can be between 11 and 15 additions, substitutions, and / or deletions. In some embodiments, a number of additions, substitutions, and / or deletions from a first target can be between 16 and 20 additions, substitutions, and / or deletions. In some embodiments, a number of additions,substitutions, and / or deletions from a first target can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and / or 20 additions, substitutions, and / or deletions.BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The foregoing will be apparent from the following more particular description of example embodiments, as illustrated in the accompanying drawings in which like reference characters refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating embodiments.
[0030] Fig. l is a block diagram illustrating example embodiments of an iterative workflow carried out over multiple cycles of protein design to generate novel anti-RBD binders.
[0031] Fig. 2 is a diagram illustrating example embodiments.
[0032] Fig. 3 A is a flow diagram illustrating example embodiments.
[0033] Fig. 3B is a flow diagram illustrating example embodiments.
[0034] Figs. 4A and 4B are diagrams illustrating example embodiments of a protein design span of derived final cross-reactive SARS-CoV-2 neutralizers carried out over five serial cycles of protein design in a span of 10 months.
[0035] Fig. 5 is a graph illustrating improved cross-reactive potency against multiple SARS-CoV-2 variants that were generated across Cycles 1-5 as described in relation to Figs. 4 A and 4B.
[0036] Fig. 6 illustrates a computer network or similar digital processing environment in which various embodiments described herein may be implemented.
[0037] Fig. 7 is a diagram of an example internal structure of a computer (e.g., client processor / device or server computers) in the computer system(s) of Fig. 6.DETAILED DESCRIPTION
[0038] A description of example embodiments follows.
[0039] The term cross-reactive, as employed herein, may be defined as pertaining to the capability of antigen-binding molecules e.g., monoclonal antibodies or antigen-binding fragments thereof) to recognize and bind to not only a specific target antigen but also to one or more related antigens. These related targets may have a threshold percent sequence identity or a number of structural variations (such as additions, substitutions, and / or deletions) from the primary target. Cross-reactive molecules exhibit the ability to neutralizethe activity of these target antigens, offering broad-spectrum efficacy against multiple strains or types of pathogens, particularly in the context of antigenically heterogeneous viral proteins. This cross-reactivity is useful for developing therapeutic or preventive treatments that remain effective against various mutations or variants of the target pathogen.
[0040] Conserved antigen domains, as employed herein, may be defined as antigen domains with sequence and / or structural similarities.
[0041] An approach to focus an antibody response on conserved antigen domains shared by viral strains includes immunizing with different, yet related, antigens so that an immune response is directed against shared epitopes. Cross-reactivity between antigens occurs when an antibody directed against one specific antigen is successful in binding with another, different antigen. Traditionally, immunization of experimental animals with receptor binding domain (RBD) antigens representative of different Sarbecoviruses has reportedly led to the identification of a monoclonal antibody that is cross-reactive within the Sarbecovirus subgenus. (Burnett, et al., "Immunizations with diverse sarbecovirus receptor-binding domains elicit SARS-CoV-2 neutralizing antibodies against a conserved site of vulnerability;" Cell Reports, Immunity 54, December 14, 2021, 2908-2921). By way of another example, an antibody capable of neutralizing multiple SARS-CoV-2 variants was reportedly isolated from a SARS-CoV-1 convalescent individual subsequently vaccinated with SARS-CoV-2 vaccine. (Cao, et al., December 20, 2022, Cell Reports 41, 111845) However, these existing approaches lack predictability and scalability. In contrast, some embodiment disclosed herein may focus on areas that are conserved, while with an immunization approach the development of the antibody response against specific domains is not controlled.
[0042] Herein are provided methods and corresponding systems that leverage cooptimization capabilities to acquire antibody sequences that bind and neutralize across multiple viruses, selected for cross-reactivity. To provide one non-limiting example, a previously identified SARS-CoV-2 antibody targeting a conserved epitope within the Sarbecovirus subgenus did not neutralize SARS-CoV-2 Omicron variants. By acquiring binding and neutralization data for SARS-CoV-2 and non-SARS-CoV-2 Sarbecoviruses across multiple optimization cycles, the present inventors used methods as described herein to identify antibodies with robust neutralization metrics of Omicron variants for which binding and neutralization to non-SARS-CoV-2 Sarbecoviruses has also been preserved. Overall, systems and methods as described herein provide agile, modular, and scalableapproaches for the identification of cross-reactive neutralizing antigen binding molecules (e.g., monoclonal antibodies and / or antigen binding fragments thereof) against antigenically heterogeneous viral proteins.
[0043] Herein are provided a combined computational and experimental screening methods for generating of antigen binding molecules (e.g., monoclonal antibodies and / or antigen binding fragments thereof) targeting conserved antigen domains and capable of neutralizing multiple viral strains. Traditionally, cross-reactive antibodies have been identified by, for example, screening B-cells isolated from immunized experimental animals or human donors for binding to desired antigens, potentially followed by affinity maturations steps to improve binding. Such approaches have shortcomings, including but not limited to: a) Inability to predict the output of the immune response even in experimental animals (e.g., epitope specificity, affinity level, neutralization potencies, etc.). b) Overreliance on binding for antibody selection and optimization, which may not necessarily correlate with neutralization.
[0044] In contrast, the present inventors have addressed these issues by: a) Selecting a target domain on a viral protein based on desired biological properties (e.g., sequence conservation across viral strains, structural conformation, relevance to viral infection, for non-limiting examples). b) Exploring sequence space and acquisition of both binding and neutralization data for generated antibodies (e.g., no down selection based solely on binding since binding may not always predict neutralization). c) Optimizing based on seq2func model predictions to enable selection of cross- reactive antibodies.
[0045] In some embodiments, methods and systems as described herein are employed to optimize reference antibodies targeting a specific epitope. To give but one non-limiting example, systems and methods as described herein can be used to optimize a reference antibody targeting the class 4 receptor binding domain (RBD) epitope that is conserved among several Sarbecoviruses but heavily mutated in SARS-CoV-2 Omicron sublineages. In some embodiments, a reference antibody fails to neutralize several Omicron variants, and optimization of the reference antibody would be advantageous. In some embodiments, a reference antibody is described in Tortorici et al., “Broad Sarbecovirus Neutralization by a Human Monoclonal Antibody,” Nature, Vol. 597, Page 103 (2021), which is hereby incorporated by reference in its entirety. In some embodiments, more information about areference antibody is available at, see, e.g., PDB: 7M7W H, PDB: 7M7W L, and Starr et al., SARS-CoV-2 RBD antibodies that maximize breadth and resistance to escape, Nature 597(7874):97-102 (2021).
[0046] Fig. 1 is a block diagram 100 illustrating example embodiments of an iterative workflow carried out over multiple cycles of protein design to generate novel anti-RBD binders. A local landscape structure 102 representing an initial reference antibody in complex with a pathogen is used by a generative model 104 to generate a predicted sequence landscape. The local landscape structure 102 is a co-crystal structure of the reference antibody in complex with a pathogen, such as SARS-CoV-2 RBD. In one non-limiting example use case, the co-crystal structure can be the reference antibody in complex with SARS-CoV-2 RBD (7M7W); however, a person having ordinary skill in the art can recognize that any co-crystal structure can be used for exploration. A predicted sequence landscape can be represented as a second-order Potts model of interaction energies modeled on a given spatial conformance, with a single energy value for any given residue in each site within a sequence, and pairwise energy values between each residue pair combination.
[0047] As shown by Fig. 1, in some embodiments, methods as described herein can employ an iterative design workflow using multiple cycles. In some embodiments, in a first cycle, protein design can be considered explorative and aims to generate a diverse set of designs around a predicted sequence landscape based on a local landscape structure 102 of an initial reference antibody. In some embodiments, subsequent design cycles can be considered exploitative and search a predicted sequence landscape around best designs from a previous round. In some embodiments, all cycles employ a graph-based neural network to predict a sequence landscape compatible with a binding model derived from an input co-crystal structure. In some embodiments, all cycles can use a same predicted landscape generated from an initial reference structure. In some embodiments, cycles subsequent to an initial cycle can use a different available or predicted structure if it is available.
[0048] In some embodiments, in conjunction with a predicted sequence landscape provided by a generative model 104, machine learning models 106 are used in sequence optimization to generate a batch of sequences 108a-d that are based on sequence-to-function models relating to a first target / antigen / pathogen in complex with a reference antibody sequence. A person of ordinary skill in the art can understand that a batch of novel sequences, e.g., 108a-d, can be more or less than four sequences. In some embodiments, machinelearning models 106 generate between 50 and 2,000 (and in some embodiments, between 91and 1092) novel sequences in a batch 108a-d for non-limiting example. In some embodiments, machine learning models (e.g., seq2func models) 106 are utilized in an ensemble of Markov chain Monte Carlo-based sequence sampling strategies that co-optimize a predicted sequence landscape alongside surrogate models predicting additional properties- of-interest (see, e.g., seq2func models below). A person of ordinary skill in the art can recognize that Monte Carlo methods include computational methods that use random sampling to achieve numerical results. In the context of protein design, a Monte Carlo-based approach can be utilized to explore a sequence space and propose new designs that may possess desirable properties for non-limiting example. In some embodiments, desirable properties are predicted from machine-learning models being used.
[0049] In some embodiments, once a batch of sequences 108a-d is generated, each sequence may be evaluated for one or more levels of function for a related target or multiple related targets. In some embodiments, a related target is a target (e.g., a protein, such as the Spike protein of SARS-CoV-2, or a portion thereof (e.g., receptor-binding domain (RBD) of Spike)), that is different from but structurally similar to a reference target (e.g, an initial target in complex with a reference antibody). For example, Spike protein variants of different SARS-CoV-2 strains, such as Delta, BA.l, BA.2, BA.2.12, and BA.4 / 5, are non-limiting examples of related targets. In some embodiments, to assess each sequence of the batch 108a- d, antibody variants are generated as IgGs and evaluated 110 for their relative levels of function (e.g, SARS-CoV-2 pseudovirus neutralization against multiple strains), affinity (e.g., binding estimated with DELFIA against Spike proteins representative of SARS-CoV-2 strains and non-SARS-CoV-2 Sarbecoviruses), and / or developability (e.g., AC-SINS selfaffinity, HPLC-SEC monomericity, and polyspecific reactivity). These evaluations form sequence-to-function (seq2func) models 112 for each respective sequence of the batch 108a- d.
[0050] In some embodiments, in addition to forming seq2func models 112, evaluations 110 are used to perform a selection 114 of sequences as one or more seeds for a next iteration of a design process. In some embodiments a selection 114 can be performed manually, for example, by presenting a design user with evaluations individually or a composite fitness score. In some embodiments, a selection 114 can be performed in silico, for example, by selecting a composite fitness score over a given threshold automatically, or by other criteria for non-limiting examples.
[0051] In some embodiments, a next cycle of protein design begins by updating ML models 106. In some embodiments, updating ML models 106 is performed to improve cross- reactive neutralization potency. In some embodiments, ML models 106 update relative levels as a function of a current cycle by retraining using a collected functional data (e.g., binding to or neutralization against related targets) from a previous round of designs. In some embodiments, each current cycle updates ML models 106 by using experimental data gathered in a previous round(s) to train supervised regression models predicting cross- reactive potency and / or affinity from sequence. In this manner, in some embodiments, ML models 106 learn to generate antibodies against not only a first antigen / target, but also against related antigens / targets. In some embodiments, sequence-to-protein function (seq2func) models are included in a co-optimization objective to locally search a landscape 104 around top cross-reactive binders from a previous round (which may be referred to here as “seeds”). In some embodiments, for each round, each batch of a batch of sequences is evaluated. In some embodiments, a seq2func model is trained for each function measured. In some embodiments, all seq2func models are used jointly together in sequence optimization (e.g., Monte Carlo-based sampling) to generate novel designs. In some embodiments, some, but not all, seq2func models are used jointly together in sequence optimization (e.g., Monte Carlobased sampling) to generate novel designs. In some embodiments, one or more seeds are selected, as described above, and are used as a starting point for a next cycle of designs. In some embodiments, seeds can be selected by rank ordering designs having composite scores that prioritize cross-reactivity followed by filtering of designs with unsatisfactory developability measurements.
[0052] Fig. 2 is a diagram 200 illustrating example embodiments. A sequence landscape 202 representing a complex of a reference antibody in complex with a pathogen and one or more sequence to function model(s) 204 are input to a machine-learning model 206. In some embodiments, a reference sequence is a prediction of a sequence of an antibody given an antibody’s structure, generated by a model. In some embodiments, a reference sequence can be represented as a second-order model that is based on pairwise interactions of a crystal structure. In some embodiments, a machine-learning model 206 is trained on co-crystal structure and generates second order models. In some embodiments, second-order models are employed in a subsequent antibody generation process. In some embodiments, a second-order model(s) represents whether a particular sequence matches a particular structure. In responseto a proposed sequence and for non-limiting example, a second-order model can be used to inform whether a predicted 3D structure of a sequence is favorable.
[0053] In some embodiments, a machine-learning model 206 generates a batch of sequences within the sequence landscape 202 that have desired functions relating to a first antigen / pathogen / target in complex with a reference sequence using the function model(s) 204. In some embodiments, a batch of sequences is referred to as candidate amino acid sequences 208. In some embodiments, sequences 208 are evaluated in vitro 210 in a lab setting for various functions relating to a related target relative to a pathogen in complex with a reference antibody. In some embodiments, those evaluations are used to generate new seq2func models 212 for each respective candidate amino acid sequence 208, resulting in sequence and seq2func model pairs 214. In some embodiments, based on evaluations of each sequence, one or more sequences are selected 216 as seeds for a next round of design optimization.
[0054] In some embodiments, a machine-learning model 206 is updated with seq2func models 220 and / or a selected set of sequences 218. In some embodiments, a machine learning model 206 is updated to better select sequences using the seq2function models 220. In some embodiments, a machine learning model 206 can update a sequence landscape 202 based on selected sequences using a co-optimization procedure.
[0055] In some embodiments, a machine learning model 206 can be retrained based on seq2func models 220 and selected seeds. In some embodiments, Fig. 2 represents selected seeds as a selected set of sequences 218. In some embodiments, each retraining fine tunes a machine learning model so that it generates sequences with more cross-reactivity than a previous iteration. In some embodiments, retraining updates a predicted structure landscape.
[0056] Fig. 3A is a flow diagram 300 illustrating example embodiments. In some embodiments, methods generate, with a (first) machine-learning model, candidate amino acid sequences based on a reference amino acid sequence and seq2func models that relate to a first target / antigen / pathogen (302). In some embodiments, generating candidate amino acid sequences includes generating, using a (second) model, a predicted structure landscape based on a reference amino acid sequence in complex with an antigen, and generating candidate amino acid sequences that are structurally stable in the predicted structure landscape. In some embodiments, a first machine-learning model generates amino acid sequences that have high scores on seq2func models.
[0057] In some embodiments, after generation (302), each sequence is evaluated for function (304). In some embodiments, evaluation occurs in vitro after physically producing an amino acid for each amino acid sequence. In some embodiments, sequences are evaluated for functions, such as binding to or neutralization against a related target / antigen / pathogen, or both.
[0058] In some embodiments, evaluated functions are used to train seq2func model(s) for each evaluated function (306). In some embodiments, a set of candidate amino acid sequences are selected based on functions of each sequence (308). For example, in some embodiments, a composite function score of each sequence can be generated to aid in selection of amino acid sequences.
[0059] In some embodiments, disclosed methods iterate by generating additional candidate amino acid sequences (302).
[0060] Fig. 3B is a flow diagram 350 illustrating example embodiments. In some embodiments, the disclosed methods generate, with a (second) machine learning model, a predicted structure landscape based on a reference amino acid sequence in complex with an antigen (354). Then, in some embodiments, a (first) machine-learning model generates candidate amino acid sequences based on a predicted structure landscape and seq2func models that relate to a first target / antigen / pathogen (352). In some embodiments, a first machine-learning model generates amino acid sequences that have high scores on seq2func models that are stable in a predicted structure landscape.
[0061] In some embodiments, each generated amino acid sequence is evaluated for function (314). In some embodiments, evaluation occurs in vitro after physically producing an amino acid for each amino acid sequence. In some embodiments, sequences are evaluated for functions, such as binding to or neutralization against a related target / antigen / pathogen, or both. In some embodiments, evaluated functions are used to form seq2func models for each evaluated function (316). In some embodiments, a set of candidate amino acid sequences are selected based on functions of each sequence (318). For example, in some embodiments, a composite function score of each sequence can be generated to aid in a selection of amino acid sequences.
[0062] In some embodiments, in subsequent cycles, a seq2func model(s) can be retrained using newly measured functional evaluations. In some embodiments, retraining fine tunes a model so that it generates sequences with more cross-reactivity than a previous iteration. In some embodiments, retraining updates a predicted structure landscape. In someembodiments, disclosed methods iterate by generating additional candidate amino acid sequences (352).
[0063] Figs. 4A and 4B are diagrams (400A, 400B) illustrating example embodiments of a protein design span of derived final cross-reactive SARS-CoV-2 neutralizers carried out over five serial cycles of protein design in a span of 10 months for non-limiting example. A person of ordinary skill in the art can understand that the described methods and corresponding systems are not limited to SARS-CoV-2 neutralizers. During a period of time, several variants of concern (VOC) emerged and rapidly spread in the United States see, e.g., covi d . cdc . gov / covi d-data-tracker / #vari ant- summary) .
[0064] Cycles 1-5 404, 406, 408, 410, 412 were performed using the systems and methods described herein and generated a plurality of seed antibody sequences.
[0065] In some embodiments, a model begins with a first cycle 404, starting with a cocrystal structure of reference antibody in complex with SARS-CoV-2 RBD 402. In some embodiments, Cycle 1 404 was performed using systems and methods described herein and generated a plurality of seed antibody sequences. In some embodiments, Cycle 1 404 evaluated 414 a function of sequences against SARS-CoV Delta, binding against SARS- CoV-1, SARS-CoV-2 Delta, and SARS-CoV-2 Omicron BA. l, and developability metrics, such as High-performance liquid chromatography - size exclusion chromatograph (HPLC- SEC), Affinity-capture self-interaction nanoparticle spectroscopy (AC-SINS), and Polyspecificity for non-limiting examples.
[0066] In some embodiments, Cycle 2 406 began, after an emergence of Omicron (BA. l). Cycle 2 406 received seq2func model(s) 424 of Cycle 1 404 as training for its models and generates additional seeds. In some embodiments, these seq2func model(s) 424 are generated based on seeds being evaluated 414 from Cycle 1 404. In some embodiments, Cycle 2 406 evaluates 416 its additionally generated seeds for function against SARS-CoV-2 Delta and SARS-CoV-2 Omicron BA.l, binding against SARS-CoV- 1, SARS-CoV-2 Delta, and SARS-CoV-2 Omicron BA. l, and developability metrics, such as HPLC-SEC, AC-SINS, and Polyspecificity for non-limiting examples.
[0067] In some embodiments, Cycle 3 408 began after the emergence of the BA.2 variant. In some embodiments, Cycle 3 408 received seq2func model(s) 426 of Cycle 2 406 as training for its models and generates additional seeds. In some embodiments, these seq2func model(s) 426 are generated based on seeds being evaluated 416 from Cycle 2 406. In some embodiments, Cycle 3 408 evaluates 418 its additionally generated seeds forfunction against SARS-CoV-2 Delta and SARS-CoV-2 Omicron BA.2, binding against SARS-CoV-1, SARS-CoV-2 Delta, SARS-CoV-2 Omicron BA.l, and SARS-CoV-2 Omicron BA.2, and developability metrics, such as HPLC-SEC, AC-SINS, and Polyspecificity.
[0068] In some embodiments, Cycle 4 410 began, after an emergence of the BA.2 variant. In some embodiments, Cycle 4 410 received seq2func model(s) 428 of Cycle 3 408 as training for its models and generates additional seeds. In some embodiments, these seq2func model(s) 428 are generated based on seeds being evaluated 418 from Cycle 3 408. In some embodiments, Cycle 4 410 evaluates 420 its additionally generated seeds for function against SARS-CoV-2 Delta, SARS-CoV-2 Omicron BA.2, SARS-CoV-2 Omicron BA.2.12, and SARS-CoV-2 Omicron BA.4 / 5, binding against SARS-CoV-1, SARS-CoV-2 Delta, SARS-CoV-2 Omicron BA.2.12.1, and SARS-CoV-2 Omicron BA.4 / 5, and developability metrics, such as HPLC-SEC, AC-SINS, and polyspecificity.
[0069] In some embodiments, Cycle 5 412 began after an emergence of the BA.2.12.1 variant and the BA.4 / 5 variants. In some embodiments, Cycle 5 412 received seq2func model(s) 430 of Cycle 4 410 as training for its models and generates additional seeds. In some embodiments, seq2func model(s) 430 are generated based on seeds being evaluated 420 from Cycle 4 410. In some embodiments, Cycle 5 evaluates 422 its additionally generated seeds for function against SARS-CoV-2 Delta, SARS-CoV-2 Omicron BA.2.12, SARS-CoV- 2 Omicron BA.4 / 5, and WIV1, binding against SARS-CoV-1, SARS-CoV-2 Delta, SARS- CoV-2 Omicron BA.2.12.1, and SARS-CoV-2 Omicron BA.4 / 5, and developability metrics, such as HPLC-SEC, AC-SINS, and Polyspecificity for non-limiting examples.
[0070] In response to a rapid evolution of SARS-CoV-2, in some embodiments, a model was adjusted to include binding affinity and neutralization measurements against novel variants of concern (VOCs) soon after they emerged. By doing so, in some embodiments, a model identifies which current designs harbor the most cross-reactive potential and uses these as starting points for a next generation of designs. In some embodiments, a model can train seq2func models which learn patterns of amino acid substitutions associated with cross- reactive potency and use these models to influence sequence optimization to generate designs with improved potency.
[0071] Fig. 5 is a graph 500 illustrating improved cross-reactive potency against multiple SARS-CoV-2 variants that were generated across Cycles 1-5 (404, 406, 408, 410, 412) as described in relation to Figs. 4A and 4B. Ultimately, in some embodiments and for non-limiting example, method(s) generated five final designs with potent neutralization of BA.4 / 5 pseudovirus (mean IC50 = 46 ng / ml) compared to an initial reference antibody that lacks significant neutralization of recent Omicron variants. From the graph 500, it can be seen that subsequent iterations improve on neutralization of pseudovirus compared to previous iterations and an initial seed sequence, showing a model’s ability to determine cross-reactive antibodies. For example, in some embodiments, by Cycle 5, a generated antibody not only neutralizes the Delta variant, but also neutralizes newer B A.2.12.1 and BA 4 / 5 variants, better relative to a reference antibody or any other previously generated antibody from other cycles. In other words, in some embodiments, an ability to neutralize BA.2.12.1 and BA.4 / 5 improved, and performance against an earlier Delta strain also improved, as opposed to becoming less effective against it.
[0072] Fig. 6 illustrates a computer network or similar digital processing environment in which various embodiments described herein may be implemented.
[0073] In some embodiments, client computer(s) / devices 50 and server computer(s) 60 provide processing, storage, and input / output devices executing application programs and the like. In some embodiments, client computer(s) / devices 50 can be linked through communications network 70 to other computing devices, including other client devices / processes 50 and server computer(s) 60. In some embodiments, communications network 70 can be part of a remote access network, a global network (e.g., the Internet), a worldwide collection of computers, local area or wide area networks, and gateways that currently use respective protocols (TCP / IP, Bluetooth®, etc.) to communicate with one another. In some embodiments, other electronic device / computer network architectures are suitable.
[0074] Fig. 7 is a diagram of an example internal structure of a computer (e.g., client processor / device 50 or server computers 60) in the computer system(s) of Fig. 6. In some embodiments, each computer 50, 60 contains a system bus 79, where a bus is a set of hardware lines used for data transfer among components of a computer or processing system. In some embodiments, a system bus 79 is essentially a shared conduit that connects different elements of a computer system (e.g., processor, disk storage, memory, input / output ports, network ports, etc.) that enables transfer of information between elements. In some embodiments, an I / O device interface 82 is attached to a system bus 79 for connecting various input and output devices (e.g., keyboard, mouse, displays, printers, speakers, etc., for non-limiting examples) to a computer 50, 60. In some embodiments, a network interface 86allows a computer 50, 60 to connect to various other devices attached to a network (e.g., network 70 of Fig. 6 for non-limiting example). In some embodiments, memory 90 provides volatile storage for computer software instructions 92A and data 94A used to implement some embodiments described herein (e.g., sequence landscape model, sequence cooptimization model(s), seq2func model, and composite function score code detailed above). In some embodiments, disk storage 95 provides non-volatile storage for computer software instructions 92B and data 94B used to implement some embodiments described herein. In some embodiments, a central processor unit 84 is attached to a system bus 79 and provides for execution of computer instructions.
[0075] In some embodiments, processor routines 92A-B and data 94 are a computer program product (generally referenced 92), including a non-transitory computer-readable medium (e.g., a removable storage medium, such as one or more DVD-ROM’s, CD-ROM’s, diskettes, tapes, etc., for non-limiting example) that provides at least a portion of software instructions for a system(s) described herein. In some embodiments, a computer program product 92 can be installed by any suitable software installation procedure, as is well known in the art. In some embodiments, at least a portion of software instructions may be downloaded over a cable communication and / or wireless connection. In some embodiments, programs are a computer program propagated signal product embodied on a propagated signal on a propagation medium (e.g., a radio wave, an infrared wave, a laser wave, a sound wave, or an electrical wave propagated over a global network such as the Internet, or other network(s)). Such carrier medium or signals may be employed to provide at least a portion of the software instructions for routines / program 92.
[0076] In some embodiments, information regarding an example reference antibody follows:
[0077] Heavy Chain Amino Acid Sequence: QVQLVQSGAEVKKPGSSVKVSCKASGGIFNTYTISWVRQAPGQGLEWMGRIILMSG MANYAQKIQGRVTITADKSTSTAYMELTSLRSDDTAVYYCARGFNGNYYGWGDDD AFDIWGQGTLVTVYSASTKGPSVFPLAPSSKSTSGGTAALGCLVKDYFPEPVTVSWN SGALTSGVHTFP AVLQS SGLYSLS S VVTVPS S SLGTQTYICNVNHKPSNTKVDKRVEP KSCDKTHTCPPCPAPELLGGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSHEDPEVKF NWYVDGVEVHNAI<TI<PREEQYNSTYRVVSVLTVLHQDWLNGI<EYI<CI<VSNI<ALP APIEKTISKAKGQPREPQVYTLPPSREEMTKNQVSLTCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFFLYSKLTVDKSRWQQGNVFSCSVLHEALHSHYTQKSLSLS PGK (SEQ ID NO:1)
[0078] Light Chain Amino Acid Sequence:QTVLTQPPSVSGAPGQRVTISCTGSNSNIGAGYDVHWYQQLPGTAPKLLICGNSNRPS GVPDRFSGSKSGTSASLAITGLQAEDEADYYCQSYDSSLSGPNWVFGGGTKLTVLGQ PKAAPSVTLFPPSSEELQANKATLVCLISDFYPGAVTVAWKADSSPVKAGVETTTPSK QSNNKYAASSYLSLTPEQWKSHRSYSCQVTHEGSTVEKTVAPTECS (SEQ ID N0:2)
[0079] Heavy Chain Variable Region (VH):QVQLVQSGAEVKKPGSSVKVSCKASGGIFNTYTISWVRQAPGQGLEWMGRIILMSG MANYAQKIQGRVTITADKSTSTAYMELTSLRSDDTAVYYCARGFNGNYYGWGDDD AFDIWGQGTLVTVYS (SEQ ID NO 3)
[0080] Light Chain Variable Region (VL):QTVLTQPPSVSGAPGQRVTISCTGSNSNIGAGYDVHWYQQLPGTAPKLLICGNSNRPS GVPDRFSGSKSGTSASLAITGLQAEDEADYYCQSYDSSLSGPNWVFGGGTKLTVL (SEQ ID NO: 4)
[0081] Heavy chain complementarity determining region 1 (HCDR1): GGIFNTYT (SEQ. ID NO. 5)
[0082] Heavy chain complementarity determining region 2 (HCDR2): IILMSGMA (SEQ. ID NO. 6)
[0083] Heavy chain complementarity determining region 3 (HCDR3): ARGFNGNYYGWGDDDAFDI (SEQ. ID NO. 7)
[0084] Light chain complementarity determining region 1 (LCDR1): NSNIGAGYD (SEQ ID NO. 8)
[0085] Light chain complementarity determining region 2 (LCDR2):GNS
[0086] Light chain complementarity determining region 3 (LCDR3): QSYDSSLSGPNWV (SEQ ID NO. 10)
[0087] The teachings of all patents, published applications and references cited herein are incorporated by reference in their entirety.
[0088] While example embodiments have been particularly shown and described, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the embodiments encompassed by the appended claims.
Claims
CLAIMSWhat is claimed is:
1. A method comprising: generating, in silico with a machine-learning model, a plurality of candidate amino acid sequences based on a reference amino acid sequence and one or more sequence to function models, wherein the one or more sequence to function models measure at least one of binding to and neutralization of a first target; forming a sequence to function model based on an in vitro evaluation of at least one function of each respective candidate amino acid sequence, the at least one function being one or more of binding to and neutralization of related targets; selecting a set of the plurality of candidate amino acid sequences based on the at least one function of each respective candidate amino acid sequence, wherein each candidate amino acid sequence in the selected set binds to the first target and one or more related targets, neutralizes the first target and one or more related targets, or a combination thereof; and retraining the machine-learning model with the selected set of the plurality of candidate amino acid sequences and at least one respective sequence to function model formed.
2. The method of Claim 1, wherein generating the plurality of candidate amino acid sequences includes determining one or more amino acid substitutions, additions, deletions, or a combination thereof, to the reference amino acid sequence based on the one or more sequence to function models.
3. The method of Claim 1 or 2, further comprising deriving the reference amino acid sequence based on a structure in complex with an antigen using a machine-learning model, the machine-learning model being a graph neural network.
4. The method of any of the above claims, wherein retraining the machine-learning model includes updating the one or more sequence to function models with the formed sequence to function models associated with the selected set of the plurality of candidate amino acid sequences.
5. The method of any of the above claims, wherein the sequence to function model for each of the plurality of candidate amino acid sequences measures the binding to or neutralization of the plurality of related targets.
6. The method of any of the above claims, further comprising iterating the steps of Claim 1 after retraining.
7. The method of any of the above claims, wherein selecting the set of the plurality of amino acid sequences further includes presenting a user interface of the plurality of amino acid sequences and respective scores from the sequence to function models and receiving user selections of the set of the plurality of amino acid sequences.
8. The method of any of the above claims, wherein selecting the set of the plurality of amino acid sequences further includes automatically analyzing, in silico, respective scores from the sequence to function models and automatically selecting the set of the plurality of amino acid sequences.
9. The method of any of the above claims, wherein selecting the set of the plurality of amino acid sequences further includes automatically analyzing, in silico, respective scores from the sequence to function models, presenting, in a user interface, an indication of recommended amino acid sequences, and receiving user input selecting the set of the plurality of amino acid sequences.
10. The method of Claim 1, wherein the reference and candidate amino acid sequences are sequences of antibodies or antigen-binding fragments of antibodies.
11. The method of Claim 1, wherein the reference amino acid sequence is in complex with a first target amino acid sequence.
12. The method of Claim 1, wherein the related targets are targets having at least one of a threshold percent sequence identity of the first target and a threshold number of additions, substitutions, deletions, or a combination thereof, from the first target.
13. A system comprising: a memory; and a processor operatively connected to the memory, the memory storing instructions thereon that, when executed, cause the processor to: generate, in silico with a machine-learning model, a plurality of candidate amino acid sequences based on a reference amino acid sequence and one or more sequence to function models, wherein the one or more sequence to function models measure at least one of binding to and neutralization of a first target; form a sequence to function model for each of the plurality of candidate amino acid sequences based on a respective in vitro evaluation of at least one function of each respective candidate amino acid sequence, the at least one function being one or more of binding to and neutralization of related targets; select a set of the plurality of candidate amino acid sequences based on the at least one function of each respective candidate amino acid sequence, wherein each candidate amino acid sequence in the selected set binds to the first target and one or more related targets, neutralizes the first target and one or more related targets, or a combination thereof; and retrain the machine-learning model with the selected set of the plurality of candidate amino acid sequences and at least one respective sequence to function model formed.
14. The system of Claim 13, wherein generating the plurality of candidate amino acid sequences includes determining one or more amino acid substitutions, additions, deletions, or a combination thereof, to the reference amino acid sequence based on the one or more sequence to function models.
15. The system of any of Claims 13-14, wherein the processor is further configured to generate the reference amino acid sequence using a machine-learning model, the machine-learning model being a graph neural network.
16. The system of any of Claims 13-15, wherein retraining the machine-learning model includes updating the one or more sequence to function models with additionalevaluated functional measurements associated with the selected set of the plurality of candidate amino acid sequences.
17. The system of any of Claims 13-16, wherein the sequence to function model for each of the plurality of candidate amino acid sequences predicts the binding to or neutralization of the plurality of related targets.
18. The system of any of Claims 13-17, wherein the processor is configured to iterate the steps of Claim 13 after retraining.
19. The system of any of Claims 13-18, wherein selecting the set of the plurality of amino acid sequences further includes presenting a user interface of the plurality of amino acid sequences and respective scores from the sequence to function models and receiving user selections of the set of the plurality of amino acid sequences.
20. The system of any of Claims 13-19 wherein selecting the set of the plurality of amino acid sequences further includes automatically analyzing, in silico, respective scores from the sequence to function models and automatically selecting the set of the plurality of amino acid sequences.
21. The system of any of Claims 13-20, wherein selecting the set of the plurality of amino acid sequences further includes automatically analyzing, in silico, respective scores from the sequence to function models, presenting, in a user interface, an indication of recommended amino acid sequences, and receiving user input selecting the set of the plurality of amino acid sequences.
22. The system of any of Claims 13-21, wherein the reference and candidate amino acid sequences are sequences of antibodies or antigen-binding fragments of antibodies.
23. The system of any of Claims 13-22, wherein the reference amino acid sequence is in complex with a first target amino acid sequence.
4. The system of any of Claims 13-23, wherein the related targets are at least one of a threshold percent sequence identity of the first target and a threshold number of additions, substitutions, deletions, or a combination thereof, from the first target.
Citation Information
Patent Citations
Designing biomolecule sequence variants with pre-specified attributes
US20230268026A1
Machine learning for designing antibodies and nanobodies in-silico
WO2023049466A2
Systems and methods for determination of protein interactions
WO2024025963A1