A computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens

A computational method using protein language models predicts bnAb precursors and designs immunogens to stimulate broad neutralization, addressing the challenge of inefficient HIV-1 vaccine immunogen development by iteratively introducing mutations to achieve broad neutralization.

WO2025221679A1PCT designated stage Publication Date: 2025-10-23DUKE UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/024564
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-14
Filing Date
2025-04-14
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Current methods are difficult and time-consuming in designing immunogens that can stimulate broadly neutralizing antibodies (bnAbs) for HIV-1 vaccines, as they require targeting rare B cell lineage traits and specific mutations.

Method used

A computational approach using machine learning algorithms, specifically protein language models like ProteinBERT and ESM-2, to predict neutralization capacity of antibody sequences, identify bnAb precursors, and design immunogens that induce these antibodies by iteratively introducing mutations to achieve broad neutralization breadth.

Benefits of technology

This method efficiently identifies potential bnAbs and their precursors, facilitating the development of HIV-1 vaccines by selecting for appropriate mutations, thereby enhancing the likelihood of broad neutralization in an outbred population.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025024564_23102025_PF_FP_ABST
    Figure US2025024564_23102025_PF_FP_ABST
Patent Text Reader

Abstract

A method is described herein comprising training a protein language model using a first dataset to predict neutralization values of individual antibody sequences against corresponding Env sequences, using the trained model on a second dataset to identify neutralizing values of individual antibody sequences in the second dataset against Env sequences in a virus panel, summing the identified neutralization values of antibodies in the second dataset to identify a first set of broadly neutralizing antibody (bnAb) precursors, wherein an antibody sequence qualifies as a precursor when its summed neutralization values exceed a threshold, structurally comparing the first set of bnAb precursors to a known bnAb, using information of the structural comparison to identify a second set of bnAb precursors, maturing the second set of precursors to breadth, and identifying immunogens that select for the matured precursors.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A COMPUTATIONAL PROCESS FOR IDENTIFYING HIV-1 BROADLY NEUTRALIZING ANTIBODIES, THEIR PRECURSORS, AND INTERACTIVE IMMUNOGENS

[0002] RELATED APPLICATIONS

[0003] This International Patent Application claims the benefit of and priority to U.S. Application No. 63 / 634,200, filed April 15, 2024, entitled, “A COMPUTATIONAL PROCESS FOR IDENTIFYING HIV-1 BROADLY NEUTRALIZING ANTIBODIES, THEIR PRECURSORS, AND INTERACTIVE IMMUNOGENS”, and U.S. Provisional Application No. 63 / 788,344, filed on April 14, 2025, entitled, “A COMPUTATIONAL PROCESS FOR IDENTIFYING HIV- 1 BROADLY NEUTRALIZING ANTIBODIES, THEIR PRECURSORS, AND INTERACTIVE IMMUNOGENS” the entirety of which is incorporated by reference as if set forth herein.

[0004] STATEMENT OF GOVERNMENTAL INTEREST

[0005] This invention was made with government support under grant UM1-AH44371 awarded by the NIH, NIAID, Division of AIDS and HHS. The government has certain rights in the invention.

[0006] TECHNICAL FIELD

[0007] The disclosure herein involves the use of machine learning approaches for the analysis of antibody sequences as well as incorporating multiple biological and structural data sets pursuant to understanding the rules for designing immunogens for induction of pathogen (e.g., HIV) or broadly neutralizing protective antibodies.

[0008] BACKGROUND

[0009] HIV-1 broadly neutralizing antibodies (bnAbs) were first described in individuals infected with human immunodeficiency virus 1 (HIV-1). These antibodies neutralize multiple HIV-1 strains and target conserved viral epitopes, in contrast to antibodies that are not bnAbs, that bind to single HIV-1 strains and target unique epitopes. Immunogens for effective HIV-1 vaccines should specifically target and select bnAbs so that the vaccine can stimulate antibodies that neutralize multiple HIV-1 strains. Currently, it is difficult and time-consuming to readily design immunogens that can stimulate bnAb B cell lineage development.

[0010] INCORPORATION BY REFERENCE

[0011] Each patent, patent application, and / or publication mentioned in this specification is herein incorporated by reference in its entirety to the same extent as if each individual patent, patent application, and / or publication was specifically and individually indicated to be incorporated by reference.

[0012] BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a flow chart depicting the training and use of a model for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens, under an embodiment.

[0014] Figure 2A is a protein language model, under an embodiment.

[0015] Figure 2B is a protein language model, under an embodiment.

[0016] Figure 3 illustrates training of the protein language model, under an embodiment.

[0017] Figure 4 illustrates prediction capability of the protein language model, under an embodiment.

[0018] Figure 5 are sample entries of the Observed Antibody Space Database, under an embodiment.

[0019] Figure 6 show results from a previous experimental screen of donor NIAID45 and corresponding model prediction outcomes, under an embodiment.

[0020] Figure 7 illustrates use of the model to identify bnAb-like precursors, under an embodiment.

[0021] Figure 8 show structure / sequence based curation, under an embodiment.

[0022] Figure 9 illustrates that breadth predictions do not necessarily correlate with sequence similarity, under an embodiment. Figure 10 presents a DH270 like panel featuring the precursors identified by the structure / sequence based curation, under an embodiment.

[0023] Figure 11 illustrates steps of in silico maturation, under an embodiment.

[0024] Figure 12 illustrates steps of immunogen design, under an embodiment.

[0025] Figures 13A-13C illustrates data on antibody binding to a panel of HIV-1 Envelope SOSZPs. (13A) is information on various antibodies. (13B) and (13C) show example data on binding of antibodies to a panel of SOSIPs. In (13B), binding data for one antibody is excluded (see cross-hatch) to show binding of the other antibodies (note log(AUC) scale on y-axis 0 to over 1.0). In (13C) binding data for the antibody excluded in (13B) is shown (note log(AUC) scale on y-axis 0 to over 10).

[0026] Figures 14A-14B: (14A) depicts a plot of the t-distributed stochastic neighbor embedding (tSNE) reduction of AntiBERTa2 embeddings for CH235 clonal lineage members (paired heavylight). Color bar shows hamming distance from CH235.12. (14B) shows K-means clustering of tSNE space for CH235 (15 clusters).

[0027] Figure 15 depicts a table of heavy chain positions with high mutual information in CH235 clones. Selection of sequences is biased towards sampling of unique pairs of residues at these positions. Mutual information defined as measurement of how much knowing the identity of position X reduces the uncertainty of the identity at position Y, and vice-versa.

[0028] Figure 16 illustrates a tSNE plot where initially selected CH235 clones were made by taking a set amount from each cluster. Half the sequences in each cluster selected through greedy search algorithm to maximize diversity (red x's) for a total of 379 selections. Greedy selection based on maximizing Hamming distance and minimizing cosine similarity (averaged for each residue) while maximizing the selection of unique pairs at heavy chain positions with high mutual information.

[0029] Figure 17 depicts downselection of thousands of CH235 mutants identified. Top graph shows each mutant is compared to their wildtype sequence and scored on their impact on the AntiBERTaZ embedding: Score = (Average per-residue Euclidean Distance) / ( # Mutations * Average per-residue Cosine Similarity). Bottom graph shows 621 mutants are selected from this remaining set by pooling sequences by the number of mutations present: Final set of selected mutants distributed so that 35% had 3 mutations, 20% had 2, 20% had 4, 12.5% had 1, and the final 12.5% had 5

[0030] For each pool, 75% of the selected mutants came from the highest scoring mutants while the final 25% were randomly selected.

[0031] Figure 18 illustrates visualization of the downselected CH235 mutants through tSNE dimensionality reduction of their AntiBERTa2 embeddings, colored by the distance they have shifted from the wildtype embedding.

[0032] Figure 19 depicts a heatmap of the selected CH235 mutant binding ECso values to known HIV immunogens.

[0033] Figure 20 depicts a plot of the phylogenic lines for VRC01 antibody VH and VL chains, showing how VH and VL pairing was performed consistent with identified clades.

[0034] Figures 21A-21B: (21A) depicts a tSNE plot of AntiBERTa2 embeddings for VRC01 (paired heavy -light). Color bar shows hamming distance from VRC01 (UCA in red). (2 IB) shows K-means clustering of tSNE space for VRC01 (16clusters).

[0035] Figure 22 depicts a table of heavy chain positions with high mutual information in VRC01 clones. Selection of sequences is biased towards sampling of unique pairs of residues at these positions.

[0036] Figure 23 illustrates a tSNE reduction where initially selected VRC01 clones were made by taking a set amount from each cluster. 16 sequences from each cluster selected through greedy search algorithm to maximize diversity (red x's) for a total of 256 selections. Greedy selection based on maximizing Hamming distance and minimizing cosine similarity (averaged for each residue) while maximizing the selection of unique pairs at heavy chain positions with high mutual information.

[0037] Figure 24 depicts downselection of thousands of VRC01 mutants identified. Top graph shows each mutant is compared to their wildtype sequence and scored on their impact on the pLM embedding: Score = (Average per-residue Euclidean Distance) / ( # Mutations * Average per- residue Cosine Similarity). Bottom graph shows 229 mutants are selected from this remaining set by pooling sequences by the number of mutations present: Final set of selected mutants distributed so that 35% had 3 mutations, 20% had 2, 20% had 4, 12.5% had 1, and the final 12.5% had 5.

[0038] For each pool, 75% of the selected mutants came from the highest scoring mutants while the final 25% were randomly selected.

[0039] Figure 25 illustrates visualization of the downselected VRC01 mutants through tSNE dimensionality reduction of their AntiBERTa2 embeddings, colored by the distance they have shifted from the wildtype embedding. tSNE space of clonal and selected mutant sequences with x's denoting mutants, colored by their tSNE distance from their wildtype (VRC01 in red, UCA in orange).

[0040] Figure 26 depicts a heatmap of the selected VRC01 mutant binding ECso values to known HIV immunogens.

[0041] Figures 27A-27B: (27A) depicts a tSNE reduction of AntiBERTa2 embeddings for CH103 (paired heavy-light). Color bar shows hamming distance from CH103. (27B) shows K-means clustering of tSNE space for CH 103 (4 clusters).

[0042] Figure 28 depicts a table of heavy and light chain positions with high mutual information in CHI 03 clones. Selection of sequences is biased towards sampling of unique pairs of residues at these positions.

[0043] Figure 29 illustrates visualization of the CH103 mutants through tSNE dimensionality reduction of their AntiBERTa2 embeddings, colored by the distance they have shifted from the wildtype embedding. Mutants create stronger clustering from the initial small pool of clonal sequences.

[0044] Figure 30 depicts a heatmap of the selected CHI 03 mutant binding ECso values to known HIV immunogens.

[0045] SUMMARY OF THE INVENTION

[0046] In embodiment, a method is described herein comprising training a protein language model using a first dataset to predict neutralization values of individual antibody sequences against corresponding Env sequences, using the trained model on a second dataset to identify neutralizing values of individual antibody sequences in the second dataset against Env sequences in a virus panel, summing the identified neutralization values of antibodies in the second dataset to identify a first set of broadly neutralizing antibody (bnAb) precursors, wherein an antibody sequence qualifies as a precursor when its summed neutralization values exceed a threshold, structurally comparing the first set of bnAb precursors to a known bnAb, using information of the structural comparison to identify a second set of bnAb precursors, maturing the second set of precursors to breadth, and identifying immunogens that select for the matured precursors.

[0047] In embodiments, the model comprises the ESM-2 model.

[0048] In embodiments, the first dataset is curated from the CATNAP (Compile, Analyze and Tally NAb Panels) database.

[0049] In embodiments, each entry of the first dataset comprises an antibody’s heavy chain sequences, the antibody’s light chain sequences, an Env gpl60 sequence, and a neutralization value.

[0050] The method further comprises under an embodiment preparing the first dataset for training the model, wherein the preparing comprises individually pairing each antibody sequence of the first dataset to its corresponding Env sequence.

[0051] In embodiments, the trained model receives an antibody sequence and Env sequence pairing.

[0052] In embodiments, the trained model outputs a neutralization value wherein the neutralization value indicates effectiveness of the antibody sequence in neutralizing the corresponding Env sequence.

[0053] In embodiments, the known bnAb comprises DH70.6.

[0054] In embodiments, the second dataset is curated from the OAS (Observed Antibody Space) database.

[0055] In embodiments, the second dataset omits antibodies with known neutralization properties.

[0056] The method further comprises under an embodiment projecting information of the first set of bnAb precursors and the known bnAb onto a subspace.

[0057] The method further comprises under an embodiment using information of the subspace to structurally compare the first set of bnAb precursors to the known bnAb. In embodiments, the maturing the second set of precursors comprises in silico maturation.

[0058] In embodiments, the in silico maturation comprises receiving as input the second set of precursors.

[0059] In embodiments, the in silico maturation comprises mutating the second set of precursors.

[0060] In embodiments, the in silico maturation comprises using the trained model to screen the mutated set of precursors for breath to identify an updated second set of precursors.

[0061] In embodiments, the in silico maturation comprises iteratively repeating the receiving, the mutating, and the screening of the second set of precursors to identify the matured precursors.

[0062] In embodiments, the identifying immunogens comprises receiving as input a first set of Env sequences and the matured precursors.

[0063] In embodiments, the identifying immunogens comprises mutating the first set of Env sequences.

[0064] In embodiments, the identifying immunogens comprises screening the first set of mutated Env sequences for breadth to identify an updated first set of Env sequences by using the model to predict a number of the matured precursors that neutralize Env sequences of the first set.

[0065] In embodiments, the identifying immunogens comprises iteratively repeating the receiving, the maturing, and the screening the first set of Env sequences to identify immunogens that select for breadth based on the matured precursors.

[0066] In embodiments, the maturing the second set of precursors comprises experimental testing.

[0067] In embodiments, the experimental testing comprises testing the matured precursors against a panel of Envs based on model predictions for neutralization.

[0068] In embodiments, the experimental testing comprises the introduction of known breadth conferring mutations, and then again testing the mutated precursors against the panel.

[0069] In embodiments, the experimental testing comprises the introduction of known breadth conferring mutations, swapping in bnAb HCDR3, and then again testing the mutated precursors against the panel.

[0070] DETAILED DESCRIPTION

[0071] HIV vaccine induction of broadly neutralizing antibodies (bnAbs) has been difficult because bnAbs are disfavored by their rare traits required for binding to HIV-1 Env neutralizing epitopes (Haynes BF et al. Nature Reviews in Immunol., 23: 142-158, 2023). From studies of antibody-virus co-evolution has come the strategy of the need to target the naive B cell receptor or unmutated common ancestors (UCAs), also called germlines, of bnAb B cell lineages and to select lineage members that have acquired key, improbable mutations by optimally designed sequential Env immunizations (Nature Biotechnology 30(5):423-33, 2012). The field of HIV-1 vaccine design has made progress in design of UCA-targeting immunogens (Nature Reviews in Immunol., 23: 142-158, 2023). In humans, broadly neutralizing antibody CD4 binding site precursors have been induced (Science doi: 10.1126 / science.add6502), and proof of concept that HIV-1 heterologous neutralizing antibodies can be induced in humans in the HVTN 133 clinical trial that induced gp41 MPER bnAb precursor and mature heterologous neutralizing antibodies (MedRxiv, doi: 10.1101 / 2024.03.15.243043052023). Thus, the next task en route to developing the first successful prototype HIV-1 vaccine, is to design optimal boosts that select for improbable mutations required for potent bnAb development (Cell Host and Microbe 23(6):759- 765, 2018). In some embodiments, an HIV-1 vaccine plan is to have four or more bnAb types induced by a prototypic vaccine to avoid transmitted / founder (TF) virus escape during transmission. This can include the development of a set of sequential Envs for each of 4 or more bnAb B cell lineages to select for bnAb B cell lineage mutations that will keep the lineages on track, and also not take any of the B cell lineages off track.

[0072] What is needed is a machine learning algorithm that can evaluate the needs of multiple boosts of multiple B cell lineages to multiple sites on the HIV envelope, and to be able to choose vaccine candidate envelop immunogens that will stimulate each bnAb B cell lineage, keep each on track to affinity maturation and robust bnAb neutralization strength and breadth.

[0073] A computational approach for identifying potential HIV-1 broadly neutralizing antibodies (bnAbs) and their precursors is disclosed herein. This approach, which aids in vaccine immunogen design, involves an algorithm that pairs an antibody heavy and light chain variable region (Fv) amino acid sequence with a series of HIV-1 Envelope (Env) viral variant sequences. The algorithm then returns a prediction for the pair’s neutralization capacity. This is achieved through fine-tuning of protein language models, such as ProteinBERT or ESM-2, using known antibody and virus sequences and their respective neutralization potency values (ICso) from the CATNAP database. An outlier removal algorithm was used to determine an average ICso value for entries containing more than two experimentally determined neutralization values. A series of models using differing data subsets and hyperparameters were used to train an ensemble of models. This ensemble is used to predict neutralization potency, or lack of neutralization, based on a model vote in which the value selected is determined by majority selection. This algorithm allows the prediction of antibody HIV-1 neutralization breadth against viral panels, enabling the screening of large libraries of paired heavy and light chain variable region sequences for antibodies with HIV-1 neutralization breadth potential. The algorithm, by design, specifically identifies which HIV-1 isolates are likely neutralized, facilitating the careful selection of variants for experimental validation and further development of candidate HIV-1 bnAbs. By clustering on fine-tuned model weights and / or layer outputs, we can identify antibodies similar to known HIV- 1 bnAbs for experimental development.

[0074] Antibody precursors are identified based on characteristics such as heavy chain complementarity determining region 3 (HCDR3) length and heavy chain VDJ and light chain VJ gene usage. However, these characteristics do not guarantee that these antibodies can be matured to neutralization breadth by acquiring specific mutations. Our virus neutralization algorithm, which can identify variant neutralization across diverse panels of HIV-1 variants, offers a solution. By introducing mutations into potential precursors with successive rounds to panel neutralization prediction, we can identify mutations that may confer breadth. This significantly advance our understanding of HIV-1 vaccine development and may potentially lead to the discovery of new, more effective antibodies. Here, we used a genetic algorithm to mature precursors computationally for downstream production. The algorithm takes the Fv amino acid sequence as input, generates a set of mutated antibody sequences, screens for breadth using the fine- tuned protein language model, selects sequences with enhanced breadth for further mutation, and cycles through the mutation / prediction / sel ection process multiple times until breadth is achieved. In this way, a potential precursor can be confirmed as an actual precursor. This information is then used in immunogen design to target these antibodies' selection and maturation specifically. Immunogen design can maximize the likelihood of inducing broad neutralization by vaccination in an outbred population by identifying large numbers of precursors and the mutations needed to confer breadth through this process. Finally, immunogens that target each antibody and their respective breadth conferring mutations can be designed using the same genetic algorithm by taking as input one or several variant Env sequences and a panel of targeted antibodies with mutations made in the Env with each successive iteration. Together, these tools provide a high-throughput means by which to screen large antibody sequence panels for bnAbs and their precursors, mature precursors, and to develop vaccine immunogens that select for breadth based on those precursors.

[0075] A machine learning approach for the analysis of antibody sequences is described herein to (1) find precursors, (2) mature them to breadth, and (3) identify immunogens that select for them. Figure 1 provides a flow chart representation of this approach. The method involves training a model on the CATNAP dataset. CATNAP (or rather Compile, Analyze and Tally Neutralizing Antibody Panels) CATNAP (Compile, Analyze and Tally NAb Panels) is a database and web server created to respond to the newest advances in HIV neutralizing antibody research. It is a comprehensive platform focusing on neutralizing antibody potencies in conjunction with viral sequences. CATNAP integrates neutralization and sequence data from published studies, and allows users to analyze that data for each HIV Envelope protein sequence position and each antibody. As seen in Figure 1, the model is first trained on data curated from the CATNAP dataset. The model then predicts neutralization values for a series of Ab sequences versus an Env sequence panel and thereby predicts Abs of interest. Based on the model’s results, the method then (i) prepares monoclonal antibodies for testing and (ii) matures monoclonal antibodies in silico thereby generating “Artificial Lineages”. The method then identifies Envs that bind and mature Ab panels.

[0076] The approach uses the following CATNAP dataset information: Heavy chain sequences, light chain sequences, aligned Env gp!60 sequences, and neutralization values. Curating the dataset produces 433 paired antibody variable region sequences. Each antibody sequence is paired individually to each Env sequence (gp!60) yielding a dataset of ~43K Ab / Env / Neut. Entries. (Note that for the ESM-2 embodiment, the Env gp!60 sequence is unaligned and with hypervariable regions and the V3 sequence truncated to fit within the required 1024 sequence token limit).

[0077] Machine learning models are selected according to their capabilities for handling the complexities of HIV- 1 neutralization. Protein language models are particularly adept at this task. Under one embodiment, the machine learning approach incorporates the ProteinBERT which is a deep language model specifically designed for proteins. The pretraining scheme combines language modeling with a novel task of Gene Ontology (GO) annotation prediction. The model introduces novel architectural elements that make the model highly efficient and flexible to long sequences. The architecture of ProteinBERT consists of both local and global representations. Figure 2B shows a schematic representation of ProteinBERT.

[0078] Under another embodiment, the machine learning approach described herein adopts the ESM-2 which is a state-of-the-art protein model trained on a masked language modelling objective. It is suitable for fine-tuning on a wide range of tasks that take protein sequences as input. The method described below uses the ESM-2 model. Figure 2B shows a schematic representation of the ESM-2 model.

[0079] Figure 3 illustrates training of the protein language model with the CATNAP paired heavy+light+virus sequences as input and CATNAP neutralization values as known output. The trained model demonstrates approximately 90 percent test set accuracy. Using this machine learning approach, we can examine which sequence positions the model is paying attention to when making predictions, use this information to find interesting virus and / or antibody features, and examine how different residues are coupled in their effect on neutralization.

[0080] Figure 4 shows use of the model to predict a neutralization tier from finetuning the protein language model to classify neutralization for one paired heavy+light+virus input sequence. Neutralization is classified as either potent (ICso < 0.1), weak (0.1 < ICso < 1), or nonneutralizing (IC50 > 1). Count O henceforth refers to the total number of virus sequences an antibody is predicted to neutralize in the potent class, Count_l for the weak class, and Count_2 for the non-neutralizing class. This one-sequence-in, one-neutralization-value-out approach allows us to test virus panels to predict overall breadth, identify viruses likely neutralized, and compare features of neutralized vs. non-neutralized. We set a particular series of cut-offs for neutralization IC50 values to create neutralization categories, such as 0 to 0.1 for potent, or Count O, 0.1 to 1 for weak, or Count l, and greater than 1 for not-neutralized, or Count_2.

[0081] Once in possession of a fine-tuned model, we turn to other datasets for screening, e.g. the Observed Antibody Space (OAS) Database which provides immune repertoires for use in large- scale analysis. The Observed Antibody Space (OAS) Database contains approximately 1.8M Human Paired Heavy and Light chain sequences. Figure 5 provides a subset of OAS Database entries. Of particular interest in testing the efficacy of the model, the OAS Database includes results from a previous experimental screen of donor NIAID45. The model proves effective in identifying bnAbs in this dataset. As seen in Figure 6, the model identifies bnAbs in this subset of data. Note entry 788403 indicating 100 percent heavy chain identity corresponding to a high Count O and Count l sums. The sequence identity indicates how similar the antibody for which predictions are made is to an antibody with known neutralization potential. The counts indicate the number of viruses neutralized which tends to correlate with sequence similarity to antibodies in the training dataset.

[0082] Note that under an embodiment, the model may be fine tuned with certain bnAbs left out of the training set. Even when the model is trained with certain bnAbs excluded, the model still recognizes those particular bnAbs. Accordingly, the model can identify breadth but is conservative in assigning virus neutralization.

[0083] The question then becomes whether any antibodies in non-HIV-1 datasets may be predicted to neutralize viruses. One does not necessarily expect to find true breadth. However, if antibodies in the dataset are predicted to neutralize multiple viruses, they may contain bnAb precursor features. In other words, the model may identify sequences in the OAS Database as bnAbs when that sequence is not known to have neutralizing properties.

[0084] As indicated above, the OAS Database contains approximately 1.8M Human Paired Heavy and Light chain sequences. However, approximately 7,000 of these sequences have known neutralizing properties against HIV-1 Env sequence panels. Those particular sequences are removed from consideration. The model is then applied to the remaining entries and identifies approximately 6,000 sequences predicted to neutralize more than 10 HIV-1 viruses in a 213 virus panel. Figure 7 provides partial results of this analysis. In particular, the results feature model results, i.e. Count_0, Count_l, and Count_2. Sequences with a Count_0 over ten are identified as precursors.

[0085] These antibodies may not have true breadth but could have bnAb precursor features. In order to examine their potential, attention is turned to DH270-like antibodies. Of course, DH270 is a known bnAb. The approximately 6,000 OAS database sequences identified by the model as potential precursors are then visually depicted along with the DH270 sequence using UMAP dimensionality reduction to project high-dimensional antibody representations from the final hidden output layer of the protein language model down to five dimensions, of which the first two are used for plotting. Euclidean distances in this five-dimensional space are used to identify antibodies in the OAS database that the model predicts to be structurally and / or functionally similar to DH270.6. Structure based screening is used to down select for production. Using the visual data representation of Figure 8, the approach chooses antibodies clustering near DH270.6.

[0086] Figure 9 illustrates that breadth predictions do not necessarily correlate with sequence similarity. Despite the presence of sequences with greater than 70% sequence identity to DH270.6 and the DH270 UCA, the model predicts sequences with less than 50% identity to be the most structurally and / or functionally similar to DH270.6, indicating an ability to interpret antibody sequence beyond sequence homology.

[0087] Figure 10 presents a DH270 like panel featuring the precursors identified by the structure / sequence based curation. Figure 10 also identifies viruses predicted to be neutralized by the DH270 panel where viruses highlighted in red are common to all precursors in the panel.

[0088] The precursor candidates are then experimentally tested and also subjected to in silico maturation. The testing procedures are described below with respect to the 10 DH270.6 like putative precursors. However, the same testing may be applied to BF520.1 (V3-glycan) with respect to precursors identified by the approach described above.

[0089] The 10 DH270.6 like precursors are experimentally tested against a panel of Envs (based on model predictions for neutralization), are then tested again after introduction of known breadth conferring mutations, and are then tested again after introduction of known breadth conferring mutations and swapping in bnAb HCDR3.

[0090] Under an embodiment, the identified precursors are also matured in silico. As indicated above, a virus neutralization algorithm is used to identify variant neutralization across diverse panels of HIV-1 variants. By introducing mutations into potential precursors with successive rounds to panel neutralization prediction, we can identify mutations that may confer breadth. This significantly advances our understanding of HIV- 1 vaccine development and may potentially lead to the discovery of new, more effective antibodies. Here, we used a genetic algorithm to mature precursors computationally for downstream production. The algorithm takes the Fv amino acid sequence as input, generates a set of mutated antibody sequences 1102, screens for breadth using the fine tuned protein language model 1104, and selects sequences with enhanced breadth for further mutation 1106. The algorithm then again introduces amino acid mutations in the antibody sequence at differing sites either randomly or based on known amino acid substitution probabilities 1108. Accordingly, the algorithm cycles through the mutation / prediction / sel ection process multiple times until breadth is achieved. In this way, a potential precursor can be confirmed as an actual precursor.

[0091] This information can be used in immunogen design to target these antibodies' selection and maturation specifically. Immunogen design can maximize the likelihood of inducing broad neutralization by vaccination in an outbred population by identifying large numbers of precursors and the mutations needed to confer breadth through this process. Finally, immunogens that target each antibody and their respective breadth conferring mutations can be designed using the same genetic algorithm by taking as input one or several variant Env sequences and a panel of targeted antibodies with mutations made in the Env with each successive iteration. As seen in Figure 12, the algorithm introduced mutations into Env sequences 1202, screens for breadth using the fine tuned protein language model 1204, i.e. predicts the number of matured precursors that neutralize the mutated Env, and selects mutated Env sequences that improve antibody breadth for further mutation 1206. The algorithm then again introduces amino acid mutations in the immunogen sequence at differing sites either randomly or based on known amino acid substitution probabilities 1208. Accordingly, the algorithm cycles through the mutation / prediction / selection process multiple times to improve antibody breadth for as many viruses as possible. Together, these tools provide a high-throughput means by which to screen large antibody sequence panels for bnAbs and their precursors, mature precursors, and to develop vaccine immunogens that select for breadth based on those precursors.

[0092] The ProteinBERT model is described in Brandes, Navad et al. (2022). ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics, 38(8), 2102- 2110. (academic. oup.com / bioinformatics / article / 38 / 8 / 2102 / 6502274). The publication is incorporated herein by reference in its entirety.

[0093] The ESM-2 model is described in Lin, Zeming et al. (2023). Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637), 1123- 1130. (www.science.org / doi / 10.1126 / science.ade2574). The publication is incorporated herein by reference in its entirety.

[0094] Computer networks suitable for use with the embodiments described herein include local area networks (LAN), wide area networks (WAN), Internet, or other connection services and network variations such as the world wide web, the public internet, a private internet, a private computer network, a public network, a mobile network, a cellular network, a value-added network, and the like. Computing devices coupled or connected to the network may be any microprocessor controlled device that permits access to the network, including terminal devices, such as personal computers, workstations, servers, mini computers, main-frame computers, laptop computers, mobile computers, palm top computers, hand held computers, mobile phones, TV set-top boxes, or combinations thereof. The computer network may include one of more LANs, WANs, Internets, and computers. The computers may serve as servers, clients, or a combination thereof.

[0095] The computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens can be a component of a single system, multiple systems, and / or geographically separate systems. The computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens can also be a subcomponent or subsystem of a single system, multiple systems, and / or geographically separate systems. The components of computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens can be coupled to one or more other components (not shown) of a host system or a system coupled to the host system.

[0096] One or more components of the computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens and / or a corresponding interface, system or application to which the computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens is coupled or connected includes and / or runs under and / or in association with a processing system. The processing system includes any collection of processor-based devices or computing devices operating together, or components of processing systems or devices, as is known in the art. For example, the processing system can include one or more of a portable computer, portable communication device operating in a communication network, and / or a network server. The portable computer can be any of a number and / or combination of devices selected from among personal computers, personal digital assistants, portable computing devices, and portable communication devices, but is not so limited. The processing system can include components within a larger computer system.

[0097] The processing system of an embodiment includes at least one processor and at least one memory device or subsystem. The processing system can also include or be coupled to at least one database. The term “processor” as generally used herein refers to any logic processing unit, such as one or more central processing units (CPUs), digital signal processors (DSPs), application-specific integrated circuits (ASIC), etc. The processor and memory can be monolithically integrated onto a single chip, distributed among a number of chips or components, and / or provided by some combination of algorithms. The methods described herein can be implemented in one or more of software algorithm(s), programs, firmware, hardware, components, circuitry, in any combination.

[0098] The components of any system that include the computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens can be located together or in separate locations. Communication paths couple the components and include any medium for communicating or transferring files among the components. The communication paths include wireless connections, wired connections, and hybrid wireless / wired connections. The communication paths also include couplings or connections to networks including local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), proprietary networks, interoffice or backend networks, and the Internet. Furthermore, the communication paths include removable fixed mediums like floppy disks, hard disk drives, and CD-ROM disks, as well as flash RAM, Universal Serial Bus (USB) connections, RS-232 connections, telephone lines, buses, and electronic mail messages.

[0099] Aspects of the computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens and corresponding systems and methods described herein may be implemented as functionality programmed into any of a variety of circuitry, including programmable logic devices (PLDs), such as field programmable gate arrays (FPGAs), programmable array logic (PAL) devices, electrically programmable logic and memory devices and standard cell-based devices, as well as application specific integrated circuits (ASICs). Some other possibilities for implementing aspects of the computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens and corresponding systems and methods include: microcontrollers with memory (such as electronically erasable programmable read only memory (EEPROM)), embedded microprocessors, firmware, software, etc. Furthermore, aspects of the computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens and corresponding systems and methods may be embodied in microprocessors having softwarebased circuit emulation, discrete logic (sequential and combinatorial), custom devices, fuzzy (neural) logic, quantum devices, and hybrids of any of the above device types. Of course the underlying device technologies may be provided in a variety of component types, e.g., metal - oxide semiconductor field-effect transistor (MOSFET) technologies like complementary metal- oxide semiconductor (CMOS), bipolar technologies like emitter-coupled logic (ECL), polymer technologies (e.g., silicon-conjugated polymer and metal-conjugated polymer-metal structures), mixed analog and digital, etc.

[0100] It should be noted that any system, method, and / or other components disclosed herein may be described using computer aided design tools and expressed (or represented), as data and / or instructions embodied in various computer-readable media, in terms of their behavioral, register transfer, logic component, transistor, layout geometries, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, non-volatile storage media in various forms (e.g., optical, magnetic or semiconductor storage media) and carrier waves that may be used to transfer such formatted data and / or instructions through wireless, optical, or wired signaling media or any combination thereof. Examples of transfers of such formatted data and / or instructions by carrier waves include, but are not limited to, transfers (uploads, downloads, e-mail, etc.) over the Internet and / or other computer networks via one or more data transfer protocols (e.g., HTTP, FTP, SMTP, etc.). When received within a computer system via one or more computer-readable media, such data and / or instruction-based expressions of the above described components may be processed by a processing entity (e.g., one or more processors) within the computer system in conjunction with execution of one or more other computer programs.

[0101] Unless the context clearly requires otherwise, throughout the description and the claims, the words “comprise,” “comprising,” and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in a sense of “including, but not limited to.” Words using the singular or plural number also include the plural or singular number respectively. Additionally, the words “herein,” “hereunder,” “above,” “below,” and words of similar import, when used in this application, refer to this application as a whole and not to any particular portions of this application. When the word “or” is used in reference to a list of two or more items, that word covers all of the following interpretations of the word: any of the items in the list, all of the items in the list and any combination of the items in the list. The above description of embodiments of the computational process for identifying HIV- 1 broadly neutralizing antibodies, their precursors, and interactive immunogens is not intended to be exhaustive or to limit the systems and methods to the precise forms disclosed. While specific embodiments of, and examples for, the computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens and corresponding systems and methods are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the systems and methods, as those skilled in the relevant art will recognize. The teachings of the computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens and corresponding systems and methods provided herein can be applied to other systems and methods, not only for the systems and methods described above.

[0102] The elements and acts of the various embodiments described above can be combined to provide further embodiments. These and other changes can be made to the computational process for identifying HIV-1 broadly neutralizing antibodies, their precursors, and interactive immunogens and corresponding systems and methods in light of the above detailed description.

[0103] One embodiment would be to incorporate the ARMADiLLO program into the machine learning algorithm to incorporate into the algorithm the ability to plan immunogen design targeted at key functional, but rare bnAb mutations.

[0104] An embodiment would be to train this algorithm to incorporate similar data sets for other viral pathogens that require germline targeting and sequential immunizations to convert subdominant B cell responses to dominant B cell responses for making antibody responses to sub dominant protective epitopes.

[0105] EXAMPLE 1

[0106] Under an embodiment, the model predicts bnAb similar antibodies from the curated OAS dataset. The model predicts neutralization values based on one paired heavy+light+virus sequence input. Using UMAP dimensionality reduction to project high-dimensional antibody representations from the final hidden output layer of the protein language model down to five dimensions, of which the first two are used for plotting. Euclidean distances in this fivedimensional space are used to identify antibodies in the OAS database that the model predicts to be structurally and / or functionally similar to bnAbs (in particular to DH270.6). Under an embodiment, the model identified 51 possible DH270 precursors from the curated OAS dataset. Structural model -based screening identified 10 promising candidates. As further described below, the approach then adds DH270 breadth conferring mutations with and / or without a DH270.6 HCDR3 swap thereby conferring binding breadth in several antibodies.

[0107] The V3-glycan targeting DH270 lineage bnAb is a major target for mutation guided design:

[0108] • The mutations needed in this lineage are known.

[0109] • Rapid selection of HIV envelopes that bind to neutralizing antibody B cell lineage members with functional improbable mutations, Cell Reports, 2021

[0110] • A priming immunogen is known.

[0111] • Targeted selection of HIV-specific antibody mutations by engineering B cell maturation, Science, 2019

[0112] • Mutation guided design works for this lineage are known.

[0113] • Targeted selection of HIV-specific antibody mutations by engineering B cell maturation, Science, 2019

[0114] • Mutation-Guided Vaccine Design: A Strategy for Developing Boosting Immunogens for HIV Broadly Neutralizing Antibody Induction, bioRxiv, 2022

[0115] • Engineering immunogens that select for specific mutations in HIV broadly neutralizing antibodies, bioRxiv, 2023

[0116] The V3-glycan precursors are rare.

[0117] • A generalized HIV vaccine design strategy for priming of broadly neutralizing antibody responses, Science, 2019

[0118] The model finds antibodies with diverse heavy and light chain genes that are DH270-like precursor bnAbs to expand the priming target pool.

[0119] • For mutation-guided design to work, a precursor can have a known set of mutations that give it breadth.

[0120] The model described herein was used to screen approximately 1 .8 million antibody sequences. In one embodiment, the model narrowed the pool of candidates to 52. The 52 candidates were examined manually to select antibodies to prepare. Antibodies were prepared and tested as below. A study was performed to test predictions made by the modeling. Example DH270 precursor (DHp) antibodies were made that have the sequences indicated in Figure 13 A. Two sets of these antibodies were made. First, DHp containing known 12 breadth-conferring mutations were made (DHp-min). Second, DHp containing known 12 breadth-conferring mutations and containing the DH270.6 HCDR3 were made (DHp-min-HCDR3). Molecules encoding these antibodies were expressed via transient transfection of HEK203 cells. The antibodies were purified and tested for binding against a panel of HIV- 1 Envelope SOSIPs using ELISA.

[0121] Example data from this study are shown in Figure 13B and 13C. These plots of antibody binding to a panel of HIV-1 Envelope SOSIPs. The results indicate several of the antibodies interact with one or more of the Envelopes. These data indicate the model described herein can identify candidate precursor antibodies.

[0122] EXAMPLE 2

[0123] A protein language model was used (AntiBERTa2, paper: www.biorxiv.org / content / 10.1101 / 2023.12.12.569610vl.full.pdf) to generate embeddings from clonal antibody sequences and clustered these sequences through tSNE dimensionality reduction and K-means clustering (Figures 14A-B, 21A-B, 27A-B). For CH235 and VRC01, initially clones were selected by taking a set amount from each cluster (Figures 16, 23) - since there were a limited amount of CHI 03 clones, all of them were used. The goal of our sequence selection was to maximize the diversity (evaluated through hamming distance of amino acid sequences and cosine similarity of embeddings) while also biasing selection towards sampling unique pairs at positions in the sequence that showed higher degrees of evolutionary coupling, evaluated through mutual information (Figures 15, 22, 28).

[0124] From these selected sequences, more diversity was added by introducing new mutations from other clusters. In this way, model training can be enhanced by altering phylogenetic signatures and creating a stronger training incentive to learn how epistatic interactions between bnAb mutations impact antigen binding and thermal stability. For CH235, sites to swap mutations between clusters were determined by finding the least similar positions using cosine similarity on the per-residue embeddings (Figure 18). For CH103 (Figure 29) and VRC01 (Figure 25), we performed mutations at positions with high mutual information. Figures 17 and 24 describe how we downselected mutants to produce, which we visualize through tSNE dimensionality reduction of their AntiBERTa2 embeddings, colored by the distance they have shifted from the wildtype embedding (Figures 18, 25). The ECso heatmaps (Figures 19, 26, 30) show the ELISA binding results of the antibodies selected and produced via this process.

Claims

CLAIMS1. A method comprising, training a protein language model using a first dataset to predict neutralization values of individual antibody sequences against corresponding Env sequences; using the trained model on a second dataset to identify neutralizing values of individual antibody sequences in the second dataset against Env sequences in a virus panel; summing the identified neutralization values of antibodies in the second dataset to identify a first set of broadly neutralizing antibody (bnAb) precursors, wherein an antibody sequence qualifies as a precursor when its summed neutralization values exceed a threshold; structurally comparing the first set of bnAb precursors to a known bnAb; using information of the structural comparison to identify a second set of bnAb precursors; maturing the second set of precursors to breadth; identifying immunogens that select for the matured precursors.

2. The method of claim 1, wherein the model comprises the ESM-2 model.

3. The method of claim 2, wherein the first dataset is curated from the CATNAP (Compile, Analyze and Tally NAb Panels) database.

4. The method of claim 3, wherein each entry of the first dataset comprises an antibody’s heavy chain sequences, the antibody’s light chain sequences, an Env gp 160 sequence, and a neutralization value.

5. The method of claim 4, further comprising preparing the first dataset for training the model, wherein the preparing comprises individually pairing each antibody sequence of the first dataset to its corresponding Env sequence.

6. The method of claim 5, wherein the trained model receives an antibody sequence and Env sequence pairing.

7. The method of claim 6, wherein the trained model outputs a neutralization value wherein the neutralization value indicates effectiveness of the antibody sequence in neutralizing the corresponding Env sequence.

8. The method of claim 1, wherein the known bnAb comprises DH70.6.

9. The method of claim 1, wherein the second dataset is curated from the OAS (Observed Antibody Space) database.

10. The method of claim 9, wherein the second dataset omits antibodies with known neutralization properties.

11. The method of claim 10, further comprising projecting information of the first set of bnAb precursors and the known bnAb onto a subspace.

12. The method of claim 11, further comprising using information of the subspace to structurally compare the first set of bnAb precursors to the known bnAb.

13. The method of claim 1, wherein the maturing the second set of precursors comprises in silica maturation.

14. The method of claim 13, wherein the in silica maturation comprises receiving as input the second set of precursors.

15. The method of claim 14, wherein the in silica maturation comprises mutating the second set of precursors.

16. The method of claim 15, wherein the in silica maturation comprises using the trained model to screen the mutated set of precursors for breath to identify an updated second set of precursors.

17. The method of claim 16, wherein the in silica maturation comprises iteratively repeating the receiving, the mutating, and the screening of the second set of precursors to identify the matured precursors.

18. The method of claim 17, wherein the identifying immunogens comprises receiving as input a first set of Env sequences and the matured precursors.

19. The method of claim 18, wherein the identifying immunogens comprises mutating the first set of Env sequences.

20. The method of claim 19, wherein the identifying immunogens comprises screening the first set of mutated Env sequences for breadth to identify an updated first set of Env sequences by using the model to predict a number of the matured precursors that neutralize Env sequences of the first set.

21. The method of claim 20, wherein the identifying immunogens comprises iteratively repeating the receiving, the maturing, and the screening the first set of Env sequences to identify immunogens that select for breadth based on the matured precursors.

22. The method of claim 1, wherein the maturing the second set of precursors comprises experimental testing.

23. The method of claim 22, wherein the experimental testing comprises testing the matured precursors against a panel of Envs based on model predictions for neutralization.

24. The method of claim 23, wherein the experimental testing comprises the introduction of known breadth conferring mutations, and then again testing the mutated precursors against the panel.

25. The method of claim 24, wherein the experimental testing comprises the introduction of known breadth conferring mutations, swapping in bnAb HCDR3, and then again testing the mutated precursors against the panel.

Citation Information

Patent Citations

  • Methods to identify immunogens by targeting improbable mutations

    US20220185871A1

  • Reinforcement learning (RL) for protein design

    WO2023246834A1