Antibody generation

A fine-tuned protein large language model addresses inefficiencies in antibody discovery by generating diverse and functional human antibodies against various targets, overcoming limitations of conventional methods and AI-based techniques that rely on structure-based training data.

WO2025245380A1PCT designated stage Publication Date: 2025-11-27VANDERBILT UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/030650
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-22
Filing Date
2025-05-22
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Conventional antibody discovery methods are inefficient, costly, and limited by the need for large amounts of practical trial and error, while existing AI-based techniques struggle with complex biological datasets and require antibody-antigen structures that are less available than sequences, limiting their training data.

Method used

A computer-implemented method using a fine-tuned protein large language model (PLM) for autoregressive training on antigen-antibody sequences, enabling the generation of human antibodies without requiring antibody-antigen structures, and incorporating filtering and ranking criteria for humanness, germline identity, and mutational load.

Benefits of technology

The method generates diverse and functional human antibodies with demonstrated binding specificity against targets like SARS-CoV-2, H5N1, and RSV-A, showcasing efficient and accurate antibody design capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000048_0000
    Figure 00000048_0000
  • Figure 00000049_0000
    Figure 00000049_0000
  • Figure 00000050_0000
    Figure 00000050_0000
Patent Text Reader

Abstract

An example method of training a large language model for antibody generation includes: receiving a training corpus comprising a plurality of antigen amino acid sequences and a corresponding plurality of antibody sequences, generating a plurality of training strings, where each training string comprises an antigen sequence and a corresponding antibody sequence, fine-tuning a pretrained protein language model by autoregressively training on the plurality of training strings, and outputting a fine-tuned protein language model configured to output human antibodies in response to receiving an input string comprising an antigen.
Need to check novelty before this filing date? Find Prior Art

Description

ANTIBODY GENERATIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. provisional patent applicationNo. 63 / 650,551, filed on May 22, 2024, and titled “DE NOVO GENERATION OF PAIRED HEAVY-LIGHT CHAIN ANTIBODY SEQUENCES AGAINST SARS-COV-2 USING LARGE LANGUAGE MODELS / ’ the disclosure of which is expressly incorporated herein by reference in its entirety.STATEMENT REGARDING FEDERALLY FUNDED RESEARCH

[0002] This invention was made with Government Support under Grant Nos. R01AH75245 and R01 AH52693 awarded by the National Institutes of Health. The Government has certain rights in the invention.REFERENCE TO SEQUENCE LISTING

[0003] The sequence listing submitted on May 22, 2025, as an .XML file entitled“10644-180W01 Sequence Listing” created on May 22, 2025 and having a file size of 43,956 bytes is hereby incorporated by reference pursuant to 37 C.F.R. § 1.52(e)(5).BACKGROUND

[0004] Antigens are molecules that trigger an immune response. Antibodies are proteins produced by the immune system. Antibodies function by binding to antigens. Antibodies can be defined by the sequence of proteins that makes up the antibody. Antibodies can be used as treatments for various diseases and conditions. Improvements to antibodies, and the design of antibodies, enable new and improved medical treatments.SUMMARY

[0005] In some aspects, implementations of the present disclosure include a computer-implemented method of training a protein large language model for antibody generation, the computer-implemented method including: receiving a training corpus including a plurality of antigen amino acid sequences and a corresponding plurality of antibody sequences; generating a plurality of training strings, wherein each training string includes an antigen sequence and a corresponding antibody sequence; fine-tuning apretrained protein large language model by autoregressively training on the plurality of training strings; and outputting a fine-tuned protein large language model configured to output human antibodies in response to receiving an input string including an antigen.

[0006] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein the plurality of antibody sequences include a plurality of light-chain sequences and a plurality of heavy-chain sequences.

[0007] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein the corresponding antibody sequence includes a heavy chain sequence and light chain sequence.

[0008] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein the plurality of training strings further include a plurality of special tokens to separate the heavy chain sequence and light chain sequence.

[0009] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein the pretrained protein language model includes a transformer model.

[0010] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein the training corpus includes a dataset of antigenspecific antibody sequences.

[0011] In some aspects, implementations of the present disclosure include a computer-implemented method of generating an antibody, the method including: inputting an antigen sequence to a fine-tuned pretrained protein language model, wherein the fine-tuned pretrained protein language model is autoregressively trained on a plurality' of antigen ammino acid sequences and corresponding plurality of antibody sequences; outputting, by fine-tuned pretrained protein language model, a plurality of antibody sequences; filtering the plurality of antibody sequences; ranking the plurality of antibody sequences; selecting, based on the ranked antibody sequences, an antibody sequence to neutralize the antigen sequence.

[0012] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein the antibody sequence includes a monoclonal antibody.

[0013] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein the corresponding antibody sequence includes a heavy chain sequence and light chain sequence.

[0014] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein filtering the plurality’ of antibody sequences includes selecting antibody sequences based on humanness.

[0015] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein filtering the plurality of antibody sequences includes selecting antibody sequences based on a germline identity of the antibody sequences.

[0016] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein filtering the plurality of antibody sequences includes selecting antibody sequences based on a length of the antibody sequences.

[0017] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein filtering the plurality of antibody sequences includes selecting antibody sequences based on a mutational load measurement.

[0018] In some aspects, implementations of the present disclosure include a computer-implemented method, wherein ranking the plurality of antibody sequences include applying a diversity-aware heuristic.

[0019] In some aspects, implementations of the present disclosure include a non- transitory computer readable medium having instructions stored thereon, which, when executed by a processor, cause the processor to perform the methods described herein.

[0020] In some aspects, implementations of the present disclosure include an isolated recombinant antibody or antigen-binding fragment that specifically binds to avian- influenza H5 hemagglutinin, wherein the antibody includes: a heavy-chain variable region with a CDR1 including a sequence with at least 70% identity to SEQ ID NO: 21; a CDR2 including a sequence with at least 70% identity to SEQ ID NO: 23; and a CDR3 including a sequence with at least 70% identity to SEQ ID NO: 25; and a light chain variable region with a CDL1 including a sequence with at least 70% identity to SEQ ID NO: 22; a CDL2 including a sequence with at least 70% identity to SEQ ID NO: 24; and a CDL3 including a sequence with at least 70% identity to SEQ ID NO: 26.

[0021] In some aspects, implementations of the present disclosure include an recombinant antibody or antigen-binding fragment where the heavy-chain variable region comprises a variable heavy chain region comprising a sequence with at least 70 % identity to SEQ ID NO: 25, and wherein the light-chain variable region comprises a variable light chain region comprising a sequence with at least 70 % identity to SEQ ID NO: 26.

[0022] In some aspects, implementations of the present disclosure include an recombinant antibody or antigen-binding fragment where the heavy-chain variable region comprises a variable heavy chain region comprising a sequence with at least 70 % identity to SEQ ID NO: 27, and wherein the light-chain variable region comprises a variable light chain region comprising a sequence with at least 70 % identity to SEQ ID NO: 28.

[0023] In some aspects, implementations of the present disclosure include an recombinant antibody or antigen-binding fragment where the heavy-chain vanable region comprises a variable heavy chain region comprising a sequence with at least 70 % identity to SEQ ID NO: 29, and wherein the light-chain variable region comprises a variable light chain region comprising a sequence with at least 70 % identity to SEQ ID NO: 30.

[0024] In some aspects, implementations of the present disclosure include an recombinant antibody or antigen-binding fragment where the antibody is a monoclonal antibody.

[0025] In some aspects, implementations of the present disclosure include an recombinant antibody or antigen-binding fragment where the monoclonal antibody is configured as a pharmaceutical composition.

[0026] It should be understood that the above-described subject matter may also be implemented as a computer-controlled apparatus, a computer process, a computing system, or an article of manufacture, such as a computer-readable storage medium.

[0027] Other systems, methods, features and / or advantages will be or may become apparent to one with skill in the art upon examination of the following drawings and detailed description. It is intended that all such additional systems, methods, features and / or advantages be included within this description and be protected by the accompanying claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The components in the drawings are not necessarily to scale relative to each other. Like reference numerals designate corresponding parts throughout the several views.

[0029] FIG. 1 illustrates an example system for training machine learning models to design antibodies, according to implementations of the present disclosure.

[0030] FIG. 2 A illustrates an example method of training machine learning models to design antibodies, according to implementations of the present disclosure.

[0031] FIG. 2B illustrates an example method of designing antibodies using a trained machine learning model.

[0032] FIG. 3 illustrates an example computing device.

[0033] FIG. 4A illustrates an example system for fine-tuning a protein LLM, according to implementations of the present disclosure.

[0034] FIG. 4B illustrates antibody -antigen pairs used in a training database,

[0035] FIG. 4C illustrates training counts of non CoV antigen groups, according to a study of an example implementation of the present disclosure.

[0036] FIG. 4D illustrates results showing a percentage of 1000 antibodies generated against RBD for each combination of light and heavy genes, according to a study of an example implementation of the present disclosure.

[0037] FIG. 4E illustrates generated variable heavy and variable light sequences aligned to the training data, according to a study of an example implementation of the present disclosure.

[0038] FIG. 4F illustrates mean a mean across all RBD-generated sequences, according to a study of an example implementation of the present disclosure.

[0039] FIG. 5A illustrates a schematic of antibody selection after generation, according to a study of an example implementation of the present disclosure.

[0040] FIG. 5B illustrates Levenshtein distance for the VH or VL. according to a study of an example implementation of the present disclosure.

[0041] FIG. 5C illustrates ELISAarea-under-the-curve, according to a study of an example implementation of the present disclosure.

[0042] FIG. 5D illustrates an experimental relationship between minimum VH andVL, according to a study of an example implementation of the present disclosure.

[0043] FIG. 5E illustrates BLI sensorgrams for binding of high-affinity IgG antibodies, according to a study of an example implementation of the present disclosure.

[0044] FIG. 6A illustrates publicness of binding antibodies, according to a study of an example implementation of the present disclosure.

[0045] FIG. 6B illustrates a sum of edits w ithin each VH region for binding antibodies, compared to a closest sequence match in training data, according to a study of an example implementation of the present disclosure.

[0046] FIG. 6C illustrates a table of sequence characteristics of RBD binding antibodies, according to a study of an example implementation of the present disclosure.

[0047] FIG. 6D illustrates ELISA AUC for binding curve dilutions, according to a study of an example implementation of the present disclosure.

[0048] FIG. 6E illustrates IC50 values for pseudovirus neutralization, according to a study of an example implementation of the present disclosure.

[0049] FIG. 6F illustrates full pseudovirus neutralization curves, according to a study of an example implementation of the present disclosure.

[0050] FIG. 7A illustrates log fold changes showing an increase in same antigenspecificity clones for RSV-A and H5 / TX / 24 prompts compared to WT RBD prompt, according to a study of an example implementation of the present disclosure.

[0051] FIG. 7B illustrates Heatmap showing percent of 1000 generated antibody encoding different variable genes for each antigen prompt. For 1000 generated antibodies against each prompt according to a study of an example implementation of the present disclosure.

[0052] FIG. 7C illustrates minimum VH Levenshtein distance to any training antibody, according to a study of an example implementation of the present disclosure.

[0053] FIG. 7D illustrates minimum VL Levenshtein distance to any training antibody, according to a study of an example implementation of the present disclosure.

[0054] FIG. 7E illustrates percent identity to VH germline, according to a study of an example implementation of the present disclosure.

[0055] FIG. 7F illustrates percent identity to VL germline, according to a study of an example implementation of the present disclosure.

[0056] FIG. 8A illustrates ELISA dilution curves for designed antibodies againstH5 / TX / 24 hemagglutinin, according to a study of an example implementation of the present disclosure.

[0057] FIG. 8B illustrates percent somatic hypermutation in heavy and light chain for binding antibodies, calculated across VH and VL genes, according to a study of an example implementation of the present disclosure.

[0058] FIG. 8C illustrates minimum distance to training antibody sequences.Distance represents number of residues different when compared to the heavy and light chain sequences from the training match with the lowest total distance (VH + VL), according to a study of an example implementation of the present disclosure.

[0059] FIG. 8D illustrates edit distance by VH region to closest training sequence match, according to a study of an example implementation of the present disclosure.

[0060] FIG. 8E illustrates neutralization dilution curves against H5 / TX / 24 hemagglutinin, according to a study of an example implementation of the present disclosure.

[0061] FIG. 8F illustrates IC50 values calculated from curves shown in FIG. 8E.

[0062] FIG. 9 A illustrates Full ELISA dilution curves for designed antibodies against RSV-A pre-fusion, according to a study of an example implementation of the present disclosure.

[0063] FIG. 9B illustrates minimum distance to training antibody sequences where distance represents number of residues different when compared to the heavy and light chain sequences from the training match with the lowest total distance (VH + VL), according to a study of an example implementation of the present disclosure.

[0064] FIG. 9C illustrates percent somatic hypermutation in heavy and light chain for binding antibodies, calculated across VH and VL genes, according to a study of an example implementation of the present disclosure.

[0065] FIG. 9D illustrates distance by VH region to closest training sequence match, according to a study of an example implementation of the present disclosure.

[0066] FIG. 9E illustrates antibody neutralization dilution curves against RSV-A, according to a study of an example implementation of the present disclosure.

[0067] FIG. 10A illustrates an overview of 3.4 A resolution cryo-EM structure ofRSV F bound to fragments of antigen binding, according to a study of an example implementation of the present disclosure.

[0068] FIG. 10B illustrates and RSV-2245 heavy chain, according to a study of an example implementation of the present disclosure.

[0069] FIG. 10C illustrates an RSV-2245 light chain, according to a study of an example implementation of the present disclosure.

[0070] FIG. 10D illustrates an RSV-3301 heavy chain, according to a study of an example implementation of the present disclosure.

[0071] FIG. 10E illustrates an RSV-3301 light chain, according to a study of an example implementation of the present disclosure.

[0072] FIG. 11 A illustrates Training and evaluation loss for 5 epochs (iterations over training data) of finetuning, according to a study of an example implementation of the present disclosure.

[0073] FIG. 1 IB illustrates counts of training data by source, according to a study of an example implementation of the present disclosure.

[0074] FIG. 11C illustrates counts of binding (LSS >2, UMI > 30) cells screened by LIBRA-seq across 4 donors, according to a study of an example implementation of the present disclosure.

[0075] FIG. 12A illustrates a scatterplot showing relationship between VH identity and the OASis humanness score 963 for antibodies generated against RBD.

[0076] FIG. 12B illustrates a scatterplot showing relationship between VH identity and the OASis humanness score 963 for antibodies generated against RBD.

[0077] FIG. 12C illustrates a scatterplot showing relationship between VH identity and the OASis humanness score 963 for antibodies generated against RBD.

[0078] FIG. 12D illustrates a scatterplot showing relationship between VH identity and the OASis humanness score 963 for antibodies generated against RBD.

[0079] FIG. 12E illustrates a scatterplot showing relationship between VH identity and the OASis humanness score 963 for antibodies generated against RBD.

[0080] FIG. 13A illustrates HCDR3 identity selection group, according to a study of an example implementation of the present disclosure.

[0081] FIG. 13B illustrates VH identity' selection group, according to a study of an example implementation of the present disclosure.

[0082] FIG. 13C illustrates unbiased selection group, according to a study of an example implementation of the present disclosure.

[0083] FIG. 13D illustrates ELISA AUCs from curve fit to dilution series, according to a study of an example implementation of the present disclosure.

[0084] FIG. 14A illustrates BLI sensorgrams for the association and dissociation kinetics of high-affinity antibodies binding to immobilized SARS-CoV-2 RBD-SD1. Data were fit to a 1 :2 bivalent analyte model, according to a study of an example implementation of the present disclosure.

[0085] FIG. 14B illustrates BLI sensorgrams for binding of low-affinity antibodies to immobilized SARS-CoV-2 RBD-SD1, according to a study of an example implementation of the present disclosure.

[0086] FIG. 14C illustrates BLI sensorgrams showing antibody specificity forSARS-CoV-2 RBD where binding was measured for antibodies to immobilized SARS-CoV- 2 RBD-SD1 or prefusion-stabilized RSV F (DS-Cavl), according to a study of an example implementation of the present disclosure.

[0087] FIG. 15 illustrates BLI sensorgrams showing antibody specificity’ forSARS-CoV-2 RBD, where binding was measured for antibodies to immobilized SARS-CoV- 2 RBD-SD1 or prefusion-stabilized RSV F (DS-Cavl), according to a study of an example implementation of the present disclosure.

[0088] FIG. 16 illustrates neutralization dilution curves for RBD-binders against full panel of spike variants, according to a study of an example implementation of the present disclosure.

[0089] FIG. 17A illustrates an initial ELISA screening for MAGE-generated antibodies against H5 / TX / 24 1009 hemagglutinin was performed at a concentration of1 Opg / mL, according to a study of an example implementation of the present disclosure.

[0090] FIG. I7B illustrates ELISA area-under-the-curve for H5 perfusion binding antibodies, according to a study of an example implementation of the present disclosure.

[0091] FIG. 17C illustrates initial ELISA screening for MAGE-generated antibodies against RSV-A prefusion was 1014 performed at a concentration of 1 Opg / mL, according to a study of an example implementation of the present disclosure.

[0092] FIG. 17D illustrates ELISA area-under-the-curve (AUC) for RSV-A prefusion binding antibodies, according to a study of an example implementation of the present disclosure.

[0093] FIG. 17E illustrates neutralization against H5 / VN / 04. according to a study of an example implementation of the present disclosure.

[0094] FIG. 17F illustrates neutralization against H5 / MI / 15, according to a study of an example implementation of the present disclosure.

[0095] FIG. 18A illustrates RSV-2245 and RSV-3301 pre-fusion, according to a study of an example implementation of the present disclosure.

[0096] FIG. 18B illustrates area under the curve (AUC) values for germline- reverted ELISA dilution curves, according to a study of an example implementation of the present disclosure.

[0097] FIG. 18C illustrates BLI sensorgrams for binding of RSV-2245 Fab to immobilized RSV-A, according to a study of an example implementation of the present disclosure.

[0098] FIG. 18D illustrates BLI sensorgrams for binding of RSV-3301 Fab to immobilized RSV-A, according to a study of an example implementation of the present disclosure.DETAILED DESCRIPTION

[0099] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Methods and materials similar or equivalent to those described herein can be used in the practice ortesting of the present disclosure. As used in the specification, and in the appended claims, the singular forms “a,” "an." "the" include plural referents unless the context clearly dictates otherwise. The term "comprising” and variations thereof as used herein is used synonymously with the term “including” and variations thereof and are open, non-limiting terms. The terms “optional” or “optionally” used herein mean that the subsequently described feature, event or circumstance may or may not occur, and that the description includes instances where said feature, event or circumstance occurs and instances where it does not. Ranges may be expressed herein as from "about" one particular value, and / or to "about" another particular value. When such a range is expressed, an aspect includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value forms another aspect. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint. While implementations will be described for the design of certain antibodytreatments, it will become evident to those skilled in the art that the implementations are not limited thereto, but are applicable for any kind of antibody design.

[0100] As used herein, the term “antibody” refers to an antigen-binding protein. As used herein, the term “antigen” refers to proteins, peptides, polysaccharides, lipids, and / or nucleic acids that bind to antibodies.

[0101] The design of antibodies is an important problem of modem medicine.Antibodies are promising treatments for many different diseases, and the rapid development of antibodies can allow for rapid response to novel diseases (e.g., new viruses). Conventional antibody discovery can require large amounts of practical trial and error, which can be extremely high cost and / or inefficient.

[0102] Machine learning models can overcome some problems with conventional antibody design. But conventional machine learning models are limited in ways that prevent them from adequately solving the of antibody design.

[0103] One type of machine learning model is a large language model (LLMs).LLMs are a type of machine learning trained on sequences of text data. Large language models use deep learning architectures (e.g., transformers) to learn statistical patterns and contextual relationships between words (which are often represented as “tokens” of text). LLMs are commonly applied to natural language problems like summarization, sentiment analysis, and search. However, conventional LLMs can struggle with specific scientific tasks,data analysis, and design. Conventional LLMs can be limited in fields like biology with complex datasets.

[0104] One approach to improving LLMs for biology includes protein LLMs.Because proteins can be encoded as sequences of text, protein LLMs can be configured to learn statistical patterns and relationships in protein data using similar computational techniques to LLMs for natural language. However, the training and fine tuning of protein LLMs require specific methodologies to create useful outputs.

[0105] Implementations of the present disclosure include systems and methods for improving antibody design by fine-tuning protein LLMs to configure them for the purpose of antibody design. Implementations of the present disclosure provide improvements to protein LLMs that enable their use for accurate antibody design, overcoming limitations of conventional LLMs and protein LLMs, as well as conventional antibody discovery methods. Implementations of the present disclosure further include antibodies configured to treat SARS-CoV-2, H5N1, and RSV-A.

[0106] Some existing Al-based techniques for antibody design are based on pairs of antibody-antigen structures. Antibody-antigen structures are less available than sequences, and thus far more antibody-antigen sequences are known than antibody-antigen structures. This limits Al-based techniques that rely on antibody-antigen structures by limiting the training data that can be used for structure-based techniques. Implementations of the present disclosure can be purely sequence based and therefore do not require antibody-antigen structures for training or inference. Accordingly, implementations of the present disclosure can take advantage of far more training data than existing structure-based techniques.Additionally, implementations of the present disclosure that are purely sequence based can be more efficient to train and use in inference mode than conventional structure-based techniques.

[0107] With reference to FIG. 1, an example system is shown according to implementations of the present disclosure. The system of FIG. 1 can be implemented using any number of computing devices (e.g., as described with reference to FIG. 3). The system can include an input module 102, a training module 110, and an output module 120.

[0108] The input module 102 is configured to receive both curated data 104 and a protein large language model 106. The curated data 104 can include any representations of amino acid sequences and corresponding antibodies. For example, the curated data can integrate paired antibody sequences and antigen specificity labels from any number ofdatabases or sources. Additional examples of curated data 104 that can be used are described in the Example, below.

[0109] The input module 102 can further include a protein large language model106. The protein large language model 106 can be any LLM that has been trained using protein sequences. For example, the protein large language model 106 can be trained to predict amino acid sequences by next-token prediction. An example protein LLM is Progen2, which is described in additional detail in the Example, below.

[0110] In some embodiments, the protein large language model 106 can be replaced with, or supplemented by, an encoder-decoder transformer, a retrieval-augmented LLM, or any ensemble thereof. Thus, it should be understood that the fine-tuning techniques described herein are not limited to the particular protein large language models 106 described herein.

[0111] The training module 110 is configured perform fine-tuning on the protein large language model 106. The training module 110 can include an antigen-specific antibody database 112 obtained from the curated data and an autoregressive fine tuning module 114. The antigen-specific antibody database 112 can store strings of antigens and corresponding antibody sequences. Optionally, the antibody sequences are stored as light chains and heavy chains separated by delimiter characters.

[0112] The autoregressive fine tuning module 114 adapts the protein large language model 106 to the data stored in the antigen-specific antibody database 112. Autoregressive fine tuning can include exposing the protein large language model 106 to each string in the antigen-specific antibody database 112 and updating the weighs of the protein large language model 106 in response. The use of autoregressive fine tuning with the antigen-specific antibody database 112 allows for the protein large language model 106 to be adapted to antibody generation, without requiring as much training data as training a new protein large language model 106. The output of the autoregressive fine tuning module 114 can be a fine-tuned protein LLM 124.

[0113] Optionally the training module 110 can be configured to fine tune multiple models. For example a first model can be configured to output light chains and a second model can be configured to output heavy chains. Alternatively or additionally, the training module 110 can include one or more models configured to assess the viability (e.g., physical possibility, human compatibility, etc.) of the antibody chains.

[0114] Still with reference to FIG. 1, the system can include an inference module120. The inference module 120 can be configured to receive an input antigen 122 (e.g., a target for antibody development). The input antigen 122 can be input to the fine-tuned protein LLM 124. The input antigen 122 can be a full-length amino acid sequence or an epitope. The input antigen 122 can be represented in any format.

[0115] The fine-tuned protein LLM 124 can output any number of antibody sequences in response to the input antigen 122. A filtering module 126 can be used to select output antibody sequences from the fine-tuned protein LLM 124 that are most likely to be successful candidate antibodies to treat the input antigen 122. The selected antibody sequences can further be ranked (e.g., based on the filtering criteria). Filtering criteria can include any measure of manufacturability or developability.

[0116] The system can then output the antibody sequences 128. Optionally, implementations of the present disclosure further include outputting the antibody sequences 128 and / or any information about the antibody sequences 128 for display. Examples of outputting the antibody sequences for display are shown in the Example, below.

[0117] With reference to FIG. 2A, implementations of the present disclosure include computer-implemented methods of training a protein large language model for antibody generation. The method of FIG. 2A can optionally be implemented using the systems described with reference to FIG. 1.

[0118] At step 202, the method includes receiving a training corpus comprising a plurality of antigen amino acid sequences and a corresponding plurality of antibody sequences. For example, the training corpus can be a dataset of antigen-specific antibody sequences. The plurality of antibody sequences can optionally include a plurality of lightchain sequences and a plurality of heavy-chain sequences.

[0119] At step 204, the method includes generating a plurality of training strings, wherein each training string comprises an antigen sequence and a corresponding antibody sequence. The corresponding antibody sequence can include both a heavy chain sequence and a light chain sequence. Optionally, special tokens can be used to separate the antigen sequence and corresponding antibody sequence. Alternatively or additionally, special tokens can be used to separate the heavy chain sequence and light chain sequence. The special tokens can include any character or sequence of characters.

[0120] At step 206, the method includes fine-tuning a pretrained protein large language model by autoregressively training on the plurality of training strings. The pretrained protein large language model can optionally be a transformer model. In someimplementations of the present disclosure, the loss function used is configured to mask the antigen sequence and calculate the mean negative loglikelihood loss across the antibody sequence only during the training step. Additional details of fine tuning are provided in the example, hereto.

[0121] At step 208, the method includes outputting a fine-tuned protein large language model configured to output human antibodies in response to receiving an input string comprising an antigen.

[0122] With reference to FIG. 2B, implementations of the present disclosure include computer-implemented methods of generating antibodies configured to neutralize antigen sequences.

[0123] At step 252, the method includes inputting an antigen sequence to a finetuned pretrained protein large language model. The fine-tuned pretrained protein large language model can be any protein large language model trained or fine tuned according to the methods described herein. For example, the protein large language model can be autoregressively trained on a plurality of antigen ammino acid sequences and corresponding plurality of antibody sequences.

[0124] At step 254, the method includes outputting, by the fine-tuned pretrained protein language model, a set of antibody sequences.

[0125] At step 256, the method includes filtering the set of antibody sequences. As described in greater detail in the example hereto, filtering can include measuring the humanness, germline identity, length, and / or mutational load of each antibody sequence in the set of antibody sequences.

[0126] At step 258, the method includes ranking the set of antibody sequences.Optionally, ranking can include applying a diversity-aware heuristic (i.e.. biased selections) to select antibodies from the set of antibody sequences that provide coverage of the set of antibody sequences.

[0127] At step 260, the method includes selecting, based on the ranked antibodysequences, an antibody sequence from the set of antibody sequences to neutralize the antigen sequence. Optionally, the method can be configured to output a monoclonal antibodysequence.

[0128] Example Implementation

[0129] An example implementation was tested and experimentally validated.Human monoclonal antibodies are a diverse class of therapeutics that can theoretically target any protein with exquisite specificity, making them promising candidates for treating a widevariety of diseases. Until recently, antibody development has been primarily driven by discovery-based experimental methods, typically through screening human or animal samples with prior exposure to an antigen target of interest. Even with recent developments that have drastically improved the throughput of antibody discovery methods, this process is laborious, slow, and cost-ineffective. The continued growth of the therapeutic market and range of applications for monoclonal antibodies presents an increased demand for in silico tools that accelerate and expand the capabilities of antibody discovery.

[0130] Recent breakthroughs in artificial intelligence (Al), most notably the unmatched performance of transformer-based Large Language Models (LLMs) and diffusion models on various tasks, have enabled a surge in computational approaches for antibody - related design tasks. Such methods include affinity maturation( 1,2 ), antibody redesign(3-5), and generation of single-domain antibodies(6,7). However, such methods lack the ability7to design template-free, antigen-specific antibodies. Existing approaches are limited to antibody redesign, with a focus on generation of complementarity determining regions (CDRs), requiring an initial antibody template to provide variable genes and framework regions for the antibody. Additionally, such models are primarily structure-based and require antibodyantigen complexes for training, which is significantly limiting due to insufficient data, especially in the context of paired, human antibodies.

[0131] The traditional process of antibody discovery is limited by inefficiency, high costs, and low success rates. Recent approaches employing artificial intelligence (Al) have been developed to optimize existing antibodies and generate antibody sequences in a target-agnostic manner. The study described herein included an example implementation of the present disclosure titled “MAGE” (Monoclonal Antibody GEnerator). MAGE includes a sequence-based Protein Language Model (PLM) fine-tuned for the task of generating paired human variable heavy and light chain antibody sequences against targets of interest. The study herein shows that MAGE can generate novel and diverse antibody sequences with experimentally validated binding specificity against SARS-CoV-2, an emerging avian influenza H5N1, and respiratory syncytial virus A (RSV-A). MAGE represents a first-inclass model capable of designing human antibodies against multiple targets with no starting template.

[0132] MAGE is a protein language model capable of generating paired heavy and light chain antibody variable sequences with binding specificity against input antigen sequences. MAGE was developed by finetuning Progen2. an auto-regressive decoder LLM that was pretrained on general protein sequences (8). Progen2 learns from observed aminoacid sequences by next-token prediction, using self-attention to capture complex dependencies within input sequences. The example implementation leverages this pretrained model's learned representation of amino acid sequences as a starting point for learning human antibody sequence features associated with binding specificity to diverse antigen targets. The study shows that the example implementation is capable of generating antibodies that exhibit diverse sequence features, including heavy and light chain variable gene usage, levels of somatic hypermutation (SHM), and novel CDRs not observed in the training data. When prompted with SARS-CoV-2 wildtype receptor binding domain (RBD), binding specificity was successfully confirmed for 9 / 20 of experimentally validated MAGE-generated antibodies, including one antibody with better than 10 ng / mL potency of SARS-CoV-2 neutralization. Binding antibodies were also designed and validated against RSV-A prefusion F (7 / 23 antibodies), which was significantly less represented in the training data. The study determined a cryo-EM structure of two MAGE-designed antibodies in complex with RSV F, demonstrating that MAGE generates antibodies with diverse binding modes and can incorporate impactful residues at key binding interfaces. Finally, MAGE-designed antibodies were validated against H5 / TX / 24 hemagglutinin (HA) (5 / 18 antibodies), demonstrating zeroshot learning capabilities against an influenza virus strain that was not present in the training data. MAGE therefore represents a first-in-class model capable of designing novel human antibodies with demonstrated functionali ty against antigen targets of interest, without having to provide any part of the antibody sequence as a starting template.

[0133] Fine-tuning a PLMfor antigen-specific antibody generation

[0134] The example implementation included a protein LLM called MAGE, finetuned for generating paired heavy and light chain antibody variable sequences that bind to a prompted antigen sequence. Toward this goal, the pretrained model Progen2(8), was finetuned on a training database of 18,507 antibody -antigen sequence pairs curated from literature and existing databases (FIG. 4A). The largest group of sequences were sourced from the Coronavirus Antibody Database (CoV-AbDab) (9), from which 10,043 human antibodies with published heavy and light chain sequences were selected as shown in FIG.11 A, then mapped back to their cognate antigen sequences based on reported binding specificities. In addition, sequences for antibody-antigen pairs across a diverse range of antigens were pulled from the Structural Antibody Database (SAbDab) (10) (n = 2,113) and the Patent and Literature Antibody Database (PLAbDab)(ll)(n = 987). Finally, antibodysequences were manually curated from literature containing high-throughput quantitative binding data for paired antibody sequences (n = 2,030)(12 — 24).

[0135] In addition to published data, the study collected an original dataset of antigen-specific antibody sequences against diverse viral antigens using LIBRA-seq (Linking B-cell Receptors to Antigen-speci ficity through Sequencing), a high-throughput method for identification of antigen-specific B cell receptors (BCRs) against an antigen panel(20). A panel of 18 diverse antigens was used to screen peripheral blood mononuclear cells (PBMCs) from 20 donors distributed across four groups (HIV infected, influenza vaccinated, COVID- 19 convalescent, and healthy). After filtering based on LIBRA-seq Scores (LSS), this dataset yielded 1,924 BCR sequences with LIBRA-seq signal for at least one antigen in the panel as shown in FIG. 11B.

[0136] In total, 67%(12, 480 / 18, 507) of the training data consisted of antibodies against CoV-related antigens, with 3,000(16.2%) antibodies included against the exact wildtype RBD sequence used for prompting (FIG. 4B). Of the remaining training examples, the most abundant target groups were viral antigens including influenza, RSV, and HIV-1 (FIG. 4C). There were, however, 987 other training antibodies with specificities against 535 different antigens sourced from the SAbDab, many of which were not viral proteins. Using this diverse training dataset, the study aimed to present a model capable of generating functional, target-specific antibodies against input antigen sequences.

[0137] Even with the inclusion of antibody -antigen pairs from these various data sources, the training dataset was far too small to train an LLM from scratch. General protein models have been shown to have superior performance on antibody-specific tasks due to the complex nature of understanding the input antigens, antibodies, and the interactions between them (5,8). The study finetuned the general protein model Progen2-base (8). which was pretrained on over a billion protein sequences across diverse domains, for the task of antigenspecific antibody generation. This was achieved by providing antibody-antigen pairs as concatenated sequences, separated by tokens between the heavy and light chains ( [LC] ) and between the antibody and antigen sequences ([SEP]). Progen2-base has a context window of 2,048 tokens, well beyond the length of larger antigens including full SARS-2 spike (1,261 residues), enabling the model to process these concatenated sequences during training.

[0138] Following finetuning, the trained model could be prompted w ith an antigen sequence of interest to generate an output containing an antibody variable heavy and light chain sequence. To evaluate the ability of MAGE to generate antigen-specific antibodysequences, the study selected three targets that spanned the range of training data representation. The study first tested generation against SARS-CoV-2 RBD, which had disproportionately higher representation in the training data. Then, to validate that the model can successfully work for antigen targets with less training data, the study tested against two additional antigens. To assess the quality' and diversity' of sequences generated by MAGE, 1,000 antibody sequences were generated against RBD and aligned to a human germline reference using IMGT numbering(25) and then filtered using the following criteria (described in detail in Methods): 1) Removal of sequences without a recognizable heavy or light chain, 2) removal of sequences with any missing CDRs or frameyvork regions (FWRs), and 3) removal of variable heavy or light sequences less than 100 amino acids in length. Almost all (991 / 1,000) of the generated sequences passed these filters. Additionally, sequences were scored for ’humanness1using the open-source platform BioPhi OASis(26). Based on suggested thresholds, sequences yvith an OASis percentile score less than 70% yvere removed, with only 2.2%(22 / 991) sequences falling below this humanness threshold as shown in FIG. 12A. While these sequences could represent viable, particularly novel sequences, this model was intended to generate human antibodies for further characterization and these loyv-scoring antibodies by OASis yvere removed accordingly. In total, 969 of 1,000 generated sequences were retained for further analysis and down-selection for in vitro characterization.

[0139] The RBD-prompted sequences displayed diverse sequence features, using37 unique variable heavy chain genes and 30 unique variable light chain genes, not accounting for different alleles. In total, 322 different pairs of heavy and light variable genes were represented in the generated sequences, with the most frequently used pair (IGHV3- 53 / 66: IGK.V1-33) representing only 13.9% ( 135 / 969 ) of sequences (FIG. 4D). Generated sequences also showed diverse CDRs, yvith heavy chain CDR3 (CDRH3) lengths ranging from 5 to 28 amino acids (mean = 16 ), and light chain CDR3s (CDRL3) lengths ranging from 7 to 12 amino acids (mean = 10 ) as shown in FIG. 12B. The light chains were more biased towards germline, yvith 50.1% (486 / 969) of containing no mutations, compared to 18.1% (175 / 969) for the heavy chains as sho vn in FIG. 12C. These results suggest that rather than simply using a single dominant heavy -light chain combination, MAGE is capable of generating diverse populations of antibody sequences.

[0140] The study determined the novelty of generated antibodies at an individual level. In an attempt to quantify this novelty, each generated sequence was compared to allsequences seen during fine-tuning to identify the most similar training example based on the minimum Levenshtein distance between each pair of sequences. This distance, which can be intuitively interpreted as the number of amino acid differences, was first calculated separately for the heavy and light chains (FIG. 4E). The study observed that the generated heavy chains contained more differences from training data sequences on average (mean = 11.7 differences), compared to light chains which exhibited substantially lower levels of differences (mean = 1.4 differences). When separated based on antibody sequence region, the distances to the nearest training sequence were highest in the CDR3s (FIG. 4F), as could be expected due to the high diversify in this region. Distances were higher for the heavy chains than light chains across all regions aside from framework region 4 (FWR4). The study also compared the similarity of generated heavy chain CDRH3s specifically to the training RBD sequences by finding the maximum sequence identify based on Levenshtein distance. The generated CDRH3s were largely novel, covering a range of identities centered at a mean of 72.4% sequence identify, with 7.4% of the generated antibodies containing CDRH3s identical to an antibody seen in training illustrated by FIG. 12E. The distribution of similarity to training data based on both heavy and light chain distance and CDR3 identify was broad as illustrated in FIG. 12D, suggesting that the generated antibody sequences cover a range of uniqueness with respect to sequences seen in training. Together, these results indicated that MAGE generated sequences with differences in all regions of the antibody, rather than only designing CDRs.

[0141] Following basic filtering of the 1,000 generated antibody sequences, the study used a simple pipeline to select antibodies for experimental validation of binding (FIG. 5A). From the 969 antibody sequences that remained after filtering. 10 antibodies were first chosen in an unbiased manner, without comparison against RBD-specific antibodies: to test a diverse unbiased set, the 20thand 80thpercentile of VH germline identify antibodies from sequences using the top 5 most frequently generated VH genes were selected. Another set of 10 antibodies was selected based on similarity to known RBD-specific antibodies: for this biased selection, the top 5 antibodies with the highest CDRH3 identify to any CoV-AbDab antibody, and top 5 antibodies with the highest VH identify to any CoV-AbDab antibody were chosen. In total, a set of 20 antibodies w as selected for in vitro validation, containing a range of sequence characteristics and novelty which aimed to represent the distribution of generated sequences. When compared to the most similar training antibody, the selected antibodies ranged from a minimum VH distance of 3 residues (RBD-238) to 24 differentresidues (RBD-153) (FIG. 5B). The respective light chains were more similar to those seen in training, with a minimum distance ranging from 0 residues (RBD-135) to 9 residues (RBD- 727).

[0142] The 20 antibodies selected for experimental validation were tested for binding to RBD from the SARS-CoV-2 index strain by ELISA (enzyme-linked immunosorbent assay) (FIG. 5C, FIGS. 13A-13D). From these results, 9 / 20(45%) of the tested antibodies were identified as binding hits for further characterization based on a minimum of 2-fold signal over background at the highest antibody concentration tested (10 pg / mL ). In the biased selection group, 2 / 5 of the CDRH3 matches (RBD-159, RBD-951) and 4 / 5 of the VH matches (RBD-238, RBD-409, RBD-413, RBD-446) displayed binding by ELISA. In the unbiased selection group. 3 / 10 antibodies (RBD-61, RBD-404, RBD-839) displayed binding, with RBD-839 displaying the strong binding signal, on par with the positive control antibody S309(27). All of these binding antibodies showed no detectable ELISA signal to BG505, an HIV-1 envelope trimer. While the binding antibodies generally exhibited lower minimum distances from both VH and VL training sequences compared to the antibodies that showed no binding, the binding antibodies nevertheless exhibited substantial novelty, with a range of 5-25 (mean: 13.6) total distance to closest training antibody (FIG. 5D). In particular, RBD-839 showed a higher minimum distance from the nearest training antibody (total distance =18 residues) than 67% of the non-binding antibodies (FIG. 5D). The study observed a wider range of VH distances to training data compared to VL, in alignment with the lower diversity of light chains the study previously observed in the pool of generated antibodies and training data.

[0143] Binding was further validated using biolayer interferometry (BLI) to measure association and dissociation kinetics for IgG binding to immobilized, monomeric SARS-CoV-2 RBD (RBD-SD1) (FIG. 5E, FIGS. 14A-14C). Apparent K_D (K_D1 ) values were determined by fitting the resulting binding curves to a 1 :2 bivalent analyte model(28), which accounts for the slower observed dissociation rate due to avidity. Of the hits identified by ELISA, 8 of 9 demonstrated measurable binding to RBD-SD1, with no binding observed for RBD-404 at the highest concentration tested (l,024nM). Four antibodies from the biased- selection groups (RBD159, RBD-238, RBD-409, and RBD-951) and one from the unbiased- selection group (RBD-839) demonstrated apparent high-affinity binding, with K_D1 values in the nanomolar to sub-nanomolar range. RBD-61, RBD-413, and RBD-446 also bound to RBD-SDL albeit with reduced apparent affinity as illustrated in FIG. 15. Although a small amount of non-specific binding was detected for one antibody (RBD-951, FIGS. 14A-14C),for the other 7 / 8 antibodies no binding was observed by BLI to a prefusion-stabilized RSV F trimer (DSCavl(29)), which is consistent with the specificity observed by ELISA.

[0144] Although FIG. 5D shows that the antibody sequences are distinct from the training data, exhibiting a range of novelty, the study sought to assess how similar the generated binding antibodies are at the population level. Public antibody clones, commonly defined by matching variable genes and CDR3 identity >70%, represent a set of criteria for grouping similar antibodies found in different individuals that are likely to share the same binding specificity(30-32). When comparing the generated binding antibodies to antibodies from the CoV-AbDab using this definition, the study found a range of 'publicness', from zero public clones for RBD-61 and RBD-404 to the highly public RBD-238 with over 100 public clones (FIG. 6A). The study further included analyzing a sum of edits within each VH region for binding antibodies, compared to a closest sequence match in training data, as shown in FIG. 6B.

[0145] The study observed that all generated binding antibodies had >70%CDRH3 identity with at least one CoV AbDab antibody, which is not surprising given the vast diversity and size of the database. Some of these antibodies, e.g. RBD-951, shared sequence features with many training antibodies at a population level, while others, e.g. RBD-61, appeared much less public as shown in FIG. 15. When comparing each generated binding antibody to its closest training match based on VH distance, the study observ ed that the majority of differences were in the CDRH3. but almost all (8 / 9) of the binding antibodies contained at least one difference outside of the CDRH3 (FIG. 6D), demonstrating the ability of MAGE to generate distinct full variable sequences rather than only designing CDRs. In addition to vary ing levels of publicness and locations of mutations, the study demonstrates that RBD-specific antibodies generated by MAGE have diverse sequence features including CDR lengths, variable gene usage, and humanness scores (FIG. 6C).

[0146] Experimental validation

[0147] Generated RBD antibodies bind full-length spike and neutralize SARS-CoV-2.

[0148] The 9 binding antibodies to RBD were tested for binding to full-lengthSARS-CoV2 spike (index), along w ith a highly mutated variant (XBB. 1) and SARS-CoV spike (SARS-1). Although MAGE was prompted using RBD only, the study wanted to interrogate whether the generated antibodies would be compatible with and bind full-length spike. Of the RBD binding antibodies. 3 / 9 showed low to no binding to full-length spike in ELISA, suggesting that these antibodies may bind epitopes that are occluded or bind inconformations that may be sterically hindered on the spike trimers (FIG. 6D). All of the 6 / 9 antibodies which did bind to full-length spike also displayed binding to XBB. 1, and two also displayed weak signal to SARS-CoV spike. These results further emphasize that MAGE was able to generate antibodies with diverse characteristics and binding properties, exhibiting cross-reactivity to different coronavirus spike variants.

[0149] Following validation of binding by ELISA, the study aimed to determine whether the generated antibodies also exhibited virus neutralization in a pseudovirus assay(32). Four of the RBD binding antibodies displayed neutralization against SARS-CoV -2 index pseudovirus, with RBD-409 displaying highly potent neutralization ( IC_50=6.7ng / mL ) (FIG. 6E, FIG. 16). Out of the 6 antibodies that bound full spike in ELISA, all but one showed neutralization potency of <lp" " g / mL against at least one spike variant. None of these antibodies were able to neutralize XBB. 1.5, although this was unseen in training as the newest RBD variant included in training was Omicron BA.5. Nevertheless, RBD-409 displayed high neutralization potency against the SARS-CoV-2 spike Gamma ( IC_50=17ng / mL ) and Delta ( IC_50=4. Ing / mL ) variants and was able to retain neutralization against several Omicron variants including BA.2, BA.2.75, and BJ. l (FIG. 6F).

[0150] Generating functional antibodies against diverse targets with lower representation in the training datasets. While the training dataset used to fine-tune MAGE was highly biased towards coronavirus antibody-antigen pairs, the dataset did contain other diverse antigen specificities to enable generation against different prompts. To that end, antibodies were designed and tested for binding against a newly emerging highly pathogenic avian influenza virus (33) (H5) and the RSV-A glycoprotein prefusion F (RSV-A). For RSV- A, there were 292 training antibodies against the exact RSV F sequence used for prompting, along with 753 antibodies against related RSV antigens including RSV-B and post-fusion RSV F. Hence, the number of exact prompt training antibodies for RSV-A represented approximately 1 / 10 of the size of the training antibodies for SARS-CoV-2 RBD. In addition to validating antibody designs against a target with limited training data, the study also sought to test the capability of MAGE to generate antibodies against a target not seen in training (zero-shot). Toward this goal, the studyprompted MAGE using hemagglutinin from the avian influenza (H5N1 clade 2.3.4.4b virus), an emerging public health threat with multiple reported interspecies transmissions, including human infections (33,34). While this exact antigen sequence was not seen in training and was not even reported at the time of training MAGE, a total of 472 H5N1 -specific antibodies were included in training. These training antibodies were primarily specific to the hemagglutinin variantA / Indonesia / 05 / 2005(17), which has 91.5% sequence identity to the more recent A / Texas / 37 / 2024 used for prompting. This target therefore represents a realistic use-case where MAGE can generate antibodies against an emerging threat without pre-existing knowledge of binding antibodies for that specific target antigen sequence.

[0151] To explore the behavior of MAGE when prompted with different antigens,1,000 sequences were generated against A / cattle / TX / 2024 H5 and RSV-A F. Notably, there was a significant enrichment of RSV-A and H5-specific clones generated when prompting with the respective antigens as opposed to prompting with SARS-CoV-2 RBD (FIG. 7A), suggesting that MAGE can enrich for prompt-specific antibody sequences. Further, each of the three prompts yielded antibody sequences with unique distributions of VH gene usage, with antibodies generated with the SARS-CoV-2 RBD prompt most frequently using IGHV3- 53 / 66, in alignment with reported gene usage biases in SARS-CoV-2-specific repertoires(35), while RSV-A sequences heavily biased towards IGHV1-18 and H5 sequences towards IGHV4-34 (FIG. 7B). The antibody sequences generated against each prompt were then compared to the training data to find the minimum Levenshtein distance for each heavy and light chain, indicating that H5 and RSV-A prompted antibodies were more novel, on average, than the RBD prompted antibodies (FIGS. 7C-7D). Additionally, the study found that the H5 and RSV-A prompted sequences exhibited higher levels of somatic hypermutation (SHM) than the RBD-prompted antibodies (FIGS. 7E-7F).

[0152] The study next sought to experimentally validate the binding specificity for a subset of these generated sequences against the H5 and RSV-A prompts. For H5, the study compared the generated sequences to H5 training antibodies and selected a validation set of 18 designed sequences for experimental validation, aiming to capture a range of novelty compared to the training examples seen (see methods section Antibody selection for experimental validation of H5N1 antibodies. The study confirmed strong binding by ELISA for 5 / 18 (28%) of these designs (FIG. 8A), along with another seven weak binding antibodies ( > 2-fold signal over background and > 0.5 absorbance) at lOjU g / mL ELISA as shown in FIG. 17 A). FIG. 17B illustrates the ELISA AUC for H5 prefusion binding antibodies calculated from the curve shown in FIG. 8A. The minimum distance to training antibodies for the binding antibodies ranged from 4-16 residues for the heavy chain, and 1 — 8 residues for the light chain (FIG. 8B). The levels of SHM ranged from 6 — 11% for the heavy chain, and 7 — 8% for the light chain (FIG. 8C). Similarly to the antibodies designed against RBD, novel residues in these H5-prompted antibodies were found across the entireVH region (FIG. 8D) and were not limited to the CDRs. Notably, all five of the strong H5 binding antibodies were neutralizing against influenza strains A / Texas / 37 / 2024, A / Vietnam / 1203 / 2004, and A / Michigan / 45 / 2015 (FIG. 8E, FIG. 8F, FIG. 17E and FIG. 17F). Antibodies H5-242 and H5-384 were the most potent (IC IC50< lOOng / mL, FIG. 8F), with IC50S comparable to the positive control CR9114, a potently and broadly neutralizing stembinding antibody (36). FIG. 17D illustrates ELISA AUC for RSV-A prefusion binding antibodies, calculated from the curve shown in FIG. 9A.

[0153] For the RSV-A prompt, the study generated a larger pool of 10,000 antibodies, followed by selection for validation of biased and unbiased selections using a similar stratification method as used for RBD (see methods section - Antibody selection for experimental validation of RSV antibodies), yielding a set of 23 antibodies for experimental validation of binding. Following initial screening illustrated in FIG. 17C, the study confirmed binding by ELISA for 7 / 23 (30%) of these designs, including three antibodies that were selected without biasing towards known RSV-specific antibodies (FIG. 9A). While the seven binding antibodies had at least one heavy chain clone in the training data (>70% CDRH3 identity, same VH gene), they nevertheless included many variations throughout the heavy and light chains ranging from a minimum heavy chain distance of six residues for RSV-6479 to 21 residues for RSV-2954 (FIG. 9B). In the light chain, the distances compared to training sequences range from 4 for RSV- 4314 to 12 for RSV3301. The MAGE-designed RSV binding antibodies showed SHM levels ranging from 3-21% for the heavy chain, and 2-12% for the light chain variable region (FIG. 9C), suggesting that MAGE does not simply leam germline-level antibody sequences. Compared to the closest training antibodies, the study showed that the RSV binding antibodies included a range of differences across the VH regions, including differences in at least 4 / 7 regions for all seven binding antibodies (FIG. 9D). There was also a range of novelty for the light chains in this set of antibodies, with the minimum VL distance to training antibodies ranging from 412 residues. The binding antibodies were further characterized by pseudovirus neutralization assays (FIG. 9E).Notably, 3 / 7 of the binding antibodies were able to neutralize RSV-A (FIG. 9E); while ICso values were not determined, RSV-2245 and RSV-4314 appeared to be strongly neutralizing with neutralization >50% at 0. 1 |ig / mL. Notably, RSV-2245 was from the unbiased selection group, with a VH distance of 17 amino acids to the closest training antibody and a SHM level of 10%, representing a highly mutated antibody with a notably distinct sequence.

[0154] To investigate the epitopes targeted by MAGE-generated antibodies from the unbiased selection group with both high levels of SHM and high distances from training,the study determined a cryo-EM structure of RSV prefusion F (PR-DM(37)) bound to fragments of antigen binding (Fabs) for RSV-2245 and RSV3301 (FIG. 10A), with details shown in FIGS. 10B-10E. For this complex, 140,634 particles were extracted from 1,323 micrographs to generate a 3.4 A resolution reconstruction with 3 copies of each Fab bound to the RSV F trimer. The structure revealed that RSV-2245 binds to an epitope primarily within prefusion-specific antigenic site V, burying 850" A"A2 on a single F protomer. Antibodies that target Site V are common in the human repertoire and tend to be potently neutralizing(12), consistent with the neutralization efficacy the study observed for RSV-2245. RSV preF is contacted by all three CDRs of the RSV-2245 heavy chain and CDRs 1 and 2 of the light chain. The interface is centered on the strands of the P3-P4 hairpin, with a large network of hydrophobic contacts mediated by CDRH3 and Tyr53 of CDRH2. The sidechain of Tyr53cDRH2 additionally contacts a single residue within P2, forming a hydrogen bond with the sidechain of Tyr53r. The RSV-2245 light chain contributes additional interactions within P4 and with residues that flank the P3-P4 hairpin. Of note, Asp30cDRL, w hich is mutated from asparagine in the germline sequence, forms a salt bridge with the Lysl92Fsidechain. This mutation was only observed in 1 / 292(0.34%) of the training RSV-specific antibodies, with the corresponding training antibody showing low' similarity ( 73% LCDR1 identity and only 50% CDRH3 identity), demonstrating the ability of the model to leam sequence features from individual training sequences and integrate them into novel antibodies. The RSV-2245 epitope further extends to include residues within antigenic Site II, mediated by polar contacts between CDRH2 and the helix-tum-helix formed by the a6 and a7 helices.

[0155] RSV-3301 represents the most highly mutated antibody of the validatedRSV-specific set. The structure revealed that RSV-3301 buries approximately 715A2within the membrane-proximal lobe of one F protomer, targeting an epitope that lies almost entirely within antigenic Site I. This site is typically considered to be postfusion F-specific but is largely conserved in both pre- and postfusion conformations (38-40). The interaction is dominated by CDRH3, which extends into the cavity formed between the a8 helix and the curved ?-sheet formed in part by ? 10, 39, ?7. and ?2. Backbone atoms within Asp99CDRH3and ArglOOCDRH3form hydrogen bonds with the sidechains of Asn380Fand Asp344F, respectively, bridging ct8 and / ?9. Notably, Argl00CDRH3was observed in training antibodies but was not found in the most similar training antibody (VH distance = 12 ) despite having a highly similar CDRH3 ( 94.4% identity). CDRH1 and CDRH2 make polar and hydrophobiccontacts within and proximal to the a8 helix, including two salt bridges formed between Arg32 CDRHI and Glu378Fand Asp58CDRH2 and Lys390F. The RSV-3301 light chain also buries surface area on F between «8 and / ?2, primarily through hydrophobic contacts mediated by Tyr32CDRL1and Tyr92CDRL3• Additionally, CDRL1 and LFR3 contact residues within ?22, extending the RSV-3301 epitope into antigenic Site IV.

[0156] Together, the structural characterization of these two antibodies demonstrates that MAGE generates antibodies with diverse binding properties. Not only do RSV-2245 and RSV-3301 target different binding sites of the RSV-A F protein, but these two antibodies display different binding properties. RSV-2245 contains binding residues distributed across both the heavy and light chains, whereas RSV-3301 binding is dominated by interactions within CDRH3. Although both antibodies contained MAGE-generated mutations in key binding residues, there were many mutations introduced into framework regions that did not interface with the antigen surface. To test the impact of these nongermline mutations, the study reverted the VH genes to germline and tested for binding by ELISA, with the germline-reverted RSV-3301 showing substantial reduction in binding by ELISA, while germline-reverted RSV-2245 retained comparable binding to its fully mutated form as shown in FIGS. 18A and 18B. Further, BLI was used to characterize the binding of RSV2245 Fab and RSV-3301 Fab to immobilized RSV-A F as shown in FIGS. 18C and 18D. For RSV-2245, a 1 : 1 binding model was used to determine binding affinity ( KD=1.5 x 10-7M ). Due to suspected heterogeneity in the epitope targeted by RSV-3301, these data were fit to a heterogeneous ligand model to determine two KDvalues (KD1=6.7 x 10-6M and KD2= 4.5 x 10-9M). Together, these results show that MAGE can generate antibodies with a variety of SHM changes in different regions of the antibody sequence and with differing impacts on antigen recognition and binding affinity.

[0157] Database curation

[0158] Paired antibody sequences with antigen specificity labels were curated from public databases and literature, primarily the CoV-AbDab, PlAbDab. and SabDab. (10- 12)For the PlAbDab, due to vague antigen labels (e.g. "flu"), sequences were manually curated from referenced literature. (13, 17-20, 50, 51)Sequences were also sources from previously published LIBRA-seq datasets (21. 22, 24, 25) and other recent literature.(16, 23) For all training sequences, heavy and light chains less than 100 amino acid residues in length w ere removed. When not provided by sources, antigens sequences were obtained from Uniprot(52). For all antigen sequences, signal peptides were removed using SignalP 6.0.(53).

[0159] Fine-tuning

[0160] The pretrained general protein model Progen2-base was fine-tuned for the task of generating paired heavy and light chain antibody sequences in response to an antigen prompt. This was achieved by autoregressive finetuning of Progen2-base's 764 million parameters on sequences consisting of paired heavy and light chain antibodies and their cognate antigens. For each training example, these sequences were input to the model as a single concatenated vector with separation tokens between each sequence:<| bos > [ Input antigen sequence] [SEP] [heavy chain sequence] [LC] [light chain sequence] < |eos| >

[0161] with two new special tokens, '[SEP]' and '[LC]' added to the Progen2 tokenizer. The training loss dropped rapidly in less than a single epoch as shown in FIG. 11C, demonstrating the ability of the pretrained model to quickly adapt to the new task and prompt format. During training, 10% of the data was held-out for evaluation during training to monitor overfitting and generalization. Fine-tuning was performed using HuggingFace(54) Trainer for causal language modeling, with the Adam optimizer. The loss function was modified to mask the antigen sequence so that the mean negative loglikelihood loss was calculated across the antibody sequence only during training (excluding the antigen sequence and [SEP] ). Loss was calculated after each step in training, while evaluation loss and accuracy were calculated every' 500 steps using the evaluation data, for a total of 5 epochs (10,415 training steps). A learning rate of 1 x 10“5was used with the default linear learning rate scheduler in HuggingFace. A training batch size of 8 was used, distributed across 4 Tesla VI 00 GPUs.

[0162] Antibody generation and basic filtering

[0163] Based on the evaluation loss minimum, antibodies were generated using the model checkpoint saved after epoch 4. For generation of antibodies against an antigen target, the fine-tuned model was prompted with the entire antigen amino acid sequence using a probability threshold of 0.9 and a temperature of 1 . The maximum sequence length w'as limited to the length of the input antigen sequence plus 250, allow ing for a total length of 250 for the combined heavy and light chain. This length limit was not necessary’ for the model to generate quality antibody sequences, as multiple pad tokens were always generated following the end of the light chain sequence. In alignment with the format of the training data, the generated sequences were in the format:<|bos> [Input antigen sequence] [SEP] [heavy chain sequence] [LC] [light chain sequence] < I pad |>„ ... <| eos |>

[0164] In order to separate out the generated antibody chains without introducing bias, the chains were simply selecting by splitting the string at the '[SEP]' and '[LC]' tokens, then truncating the light chain at the appearance of the first pad token (<|pad|>). Since the maximum length allowed for generation was 250 residues, most sequences contained extra pad tokens at the end of the sequences following the light chain, which were removed.

[0165] The methods described herein can include human-performed steps for filtering and selecting antibody sequences output by the computer-implemented methods. Thus, the antibody sequences described herein are a result of human selection by applying human judgement to a large number of outputs from the Al systems. For example, following selection of the generated heavy and light variable regions, the antibody sequences were annotated using ANARCI (55) with IMGT (26) numbering. If a sequence had either a heavy or light chain that was not recognized as a human variable chain by ANARCI, it was discarded, along with any sequences missing framework or CDR regions following alignment. Heavy chains or light chains shorter than 100 amino acids were also discarded, although few sequences under these lengths survived the previous filtering steps. Heavy chains which were identical to a training example were discarded, although there was only- one occurrence of this in the sequences generated against SARS-CoV-2 RBD. Finally, antibodies were assessed for mutational load based on identity to germline variable genes, and humanness using the BioPhi OASis software.

[0166] Perplexity7was calculated by performing a backward pass through the trained model, with the model in evaluation mode, to calculate the exponential of the average negative log-likelihood for each generated sequence.

[0167] Antibody selection for experimental validation of RBD antibodies

[0168] In addition to the filtering outlined above, additional criteria were applied to select candidates for experimental validation. Optionally, the additional criterial and selection can be performed by a human, as described herein. For the heavy variable gene, only sequences with > 85% identity- were retained. For humanness, sequences below 70thpercentile were discarded based on the recommended threshold for OASis (27). In order to test novel sequences generated by the model, antibodies with a CDRH3 identical to any- training example ( n = 72 ). or a VH germline identity of 100% ( n = 175 ) were removed, leaving a selection pool of 732 antibodies. In addition, any sequences with both > 90%VHidentity and > 90% CDRH3 identity to CoV-AbDab RBD binding antibody sequences were removed. Following these filtering steps, the following automated selected pipeline was applied:1. From the top 5 most frequently generated VH genes, select sequences with 20thand 80thpercentile VH identities.2. Select top 5 sequences by rank of maximum CDRH3 identity to known binding antibodies.3. Select top 5 sequences by rank of maximum VH identity’ to know binding antibodies.

[0169] Together, these three selection steps yield 20 antibodies per antigen target.Antibodies from Step 1 represent selection independent of known binding antibodies to avoid any bias, while antibodies from Steps 2 and 3 yield testing antibodies with high similarity to known binding antibodies.

[0170] Antibody selection for experimental validation of RSV antibodies

[0171] MAGE was prompted using the RSV-A Fusion glycoprotein F0(UniProt(52) entry P03420) amino acid sequence to generate 10,000 sequences for down selection and validation. Basic filtering was applied as described above, along with filtering based on perplexity (PPL < 1.5), heavy chain germline identity (percent VH identity < 98%), and sequence identity to training antibodies (maximum CDRH3 identity to any training antibody < 95% ). Due to the overrepresentation of CoV-specific antibodies seen in training, the remaining generated sequences were compared to CoV-AbDab antibodies to remove sequences with CDRH3s similar to CoV-specific antibodies (CDRH3 percent identity > 70%).

[0172] Following filtering, antibodies were selected for validation based on three separate criteria groups. First, an unbiased group was selected by clustering generated CDRH3s using hierarchical clustering based on a Levenshtein identity matrix with a maximum identity distance of 20% within each cluster. One sequence was then randomly sampled from each of the top 10 largest clusters, for a total of 10 unbiased sequences selected for validation. For the unbiased group, the study selected generated sequences with CDRH3 identity > 75% and equal CDRH3 length compared RSV-A-specific training antibodies. From these matches, 10 antibodies w ere randomly selected from unique CDRH3 clusters. Finally, three generated antibodies with CDRH3 identity > 70% to RSV-A-specific trainingantibodies and CDRH3 identity > 60% to MPV-A-specific training antibodies were selected. In total, 23 antibodies were selected for experimental validation.

[0173] Antibody selection for experimental validation of H5N1 antibodies

[0174] MAGE was prompted using the highly pathogenic avian influenza virusH5 / TX / 24 hemagglutinin sequence (Strain A / Texas / 37 / 2024, GenBank accession number PP577943. 1(34)) amino acid sequence as a prompt to generate 1,000 sequences for down selection and validation. Basic filtering was applied as described above, along with filtering based on heavy chain germline identity (percent VH identity < 100%). Antibodies were then selected for validation based on CDRH3 Levenshtein identity to H5 -specific training antibodies. Generated sequences were randomly selected from four CDRH3 identity bins: [80%-85%) ( n = 3 ), [85%-90%) ( n = 6). [ 90% - 95% ) ( n = 6 ), and [ 95% - 99% ) ( n = 3 ) for a total of 18 antibodies. Since this exact flu strain sequence was not seen in training, no unbiased group was selected for testing.

[0175] Discussion

[0176] The study validated the example implementation of the present disclosure which included a purely sequence-based model capable of generating paired heavy-light chain antibody sequences with prompt-specific binding. Once trained, the MAGE model presented here requires no template antibody or protein structural information. When prompted with an antigen amino acid sequence, MAGE produces full human variable heavy and light chains, including novel designs with changes from germline sequence introduced throughout the entire variable sequences. The results confirm that generative LLM models like MAGE are capable of the complex task of generating full paired heavy and light chain antibody sequences, demonstrating validated binding against RBD, H5 hemagglutinin, and RSV-A prefusion F. MAGE-generated antibodies showed diverse sequence characteristics and binding properties, including potent neutralization for a subset of the binding antibodies designed against each antigen. While MAGE is not conditioned on neutralization, this demonstrates the functionality of these antibodies and validates the ability of MAGE to produce useful, clinically relevant antibodies in the context of therapeutic discovery. For RBD and RSV-A, a subset of validated, target-specific designs were selected with no bias towards known antibodies, demonstrating design of potently neutralizing antibodies without the need for a starting template antibody sequence or structure. The design of neutralizing antibodies against H5 / TX / 24 hemagglutinin demonstrates zero-shot learning capabilities, where MAGE was able to generate antibodies against an unseen influenza strain by trainingon previously characterized antibodies with specificity against a related, but divergent H5N 1 strain. This demonstrates a realistic use-case for this approach, where MAGE could be used to generate antibodies against an emerging health threat more rapidly than traditional antibody discover}' methods that would rely on access to specialized biological materials (e.g. blood samples or antigen protein).

[0177] The antibodies designed and characterized here display a range of sequence characteristics, including differential gene usage, CDR properties, and levels of SHM. While a subset of the validated binding antibodies have CDRH3s that are similar to those seen in training, it is well-established that individual amino acid substitutions can disrupt antibodyantigen binding(41, 42), even within non-interfacing framework regions(43, 44). As such, the ability of the model presented here to generate binding - and in some cases potently neutralizing - antibody sequences highlights the utility of generative algorithms in creating solutions that differ from those seen in the training data while retaining antibody-antigen recognition properties. In addition to designs with low numbers of edits introduced to training antibodies, the study also validated binding for more novel antibodies with >20 total amino acid differences to the most similar training examples (RSV-2245 and RSV-3301). Structural characterization of these antibodies showed that they target different sites on RSV F with different modes of binding which utilize residues not found in the closest training antibody matches. Additionally, the Site I epitope targeted by RSV-3301 is not well characterized an this is the first structure showing a human antibody targeting this epitope in prefusion F(45).

[0178] It should be understood that the example implementation is not restricted to redesigning existing antibodies, rather it is able to sample the distribution of known binding sequences to leam the complex sequence features associated with antigen-binding specificity and then generate a pool of diverse antibodies that is highly enriched for binding antibodies, providing candidates for further characterization, down-selection, and development. The study only sampled a fraction of this sequence space for validation but envision that this candidate pool could be further mined to find antibodies with desired properties that have not yet been explored.

[0179] In the study, the example implementation validated against viral antigen targets as a proof-of-concept. However, data generation methods are constantly improving, and large-scale efforts using high-throughput methods such as LIBRA-seq could soon yield datasets of sufficient scale for training such models to efficiently generate antibodies against diverse antigen targets beyond what is included in the training datasets. The development of these datasets, along w ith the subsequent experimental validation of generated antibodieswhich can be incorporated into training data, will enable iterative improvement of MAGE. Since applications of LLMs in other fields have shown evidence of generalization (46-48). Provided enough data, implementations of the present disclosure can be capable of learning the more general rules of residue-level interactions that govern antibody-antigen binding with the capability to generate antibodies against completely unseen targets. Such approaches can revolutionize the field of antibody discovery. The models described herein therefore include an LLM capable of antigen specific paired heavy-hght chain antibody sequence generation enable Al-accelerated antibody discovery.

[0180] As used herein, the terms "about" or "approximately" when referring to a measurable value such as an amount, a percentage, and the like, is meant to encompass variations of ±20%, ±10%, ±5%, or ±1% from the measurable value.

[0181] “Administration” of “administering” to a subject includes any route of introducing or delivering to a subject an agent. Administration can be carried out by any suitable means for delivering the agent. Administration includes self-administration and the administration by another.

[0182] The term “subject” is defined herein to include animals such as mammals, including, but not limited to, primates (e.g., humans), cows, sheep, goats, horses, dogs, cats, rabbits, rats, mice and the like. In some embodiments, the subject is a human.

[0183] The following examples are put forth so as to provide those of ordinary skill in the art with a complete disclosure and description of how the compounds, compositions, articles, devices and / or methods claimed herein are made and evaluated, and are intended to be purely exemplary and are not intended to limit the disclosure. Efforts have been made to ensure accuracy with respect to numbers (e.g., amounts, temperature, etc.), but some errors and deviations should be accounted for. Unless indicated otherwise, parts are parts by weight, temperature is in °C or is at ambient temperature, and pressure is at or near atmospheric.

[0184] It should be appreciated that the logical operations described herein with respect to the various figures may be implemented (1) as a sequence of computer implemented acts or program modules (i.e., software) running on a computing device (e.g., the computing device described in FIG. 3), (2) as interconnected machine logic circuits or circuit modules (i.e., hardware) within the computing device and / or (3) a combination of software and hardware of the computing device. Thus, the logical operations discussed herein are not limited to any specific combination of hardware and software. The implementation isa mater of choice dependent on the performance and other requirements of the computing device. Accordingly, the logical operations described herein are referred to variously as operations, structural devices, acts, or modules. These operations, structural devices, acts and modules may be implemented in software, in firmware, in special purpose digital logic, and any combination thereof. It should also be appreciated that more or fewer operations may be performed than shown in the figures and described herein. These operations may also be performed in a different order than those described herein.

[0185] Referring to FIG. 3, an example computing device 300 upon which the methods described herein may be implemented is illustrated. It should be understood that the example computing device 300 is only one example of a suitable computing environment upon which the methods described herein may be implemented. Optionally, the computing device 300 can be a well-known computing system including, but not limited to, personal computers, servers, handheld or laptop devices, multiprocessor systems, microprocessorbased systems, network personal computers (PCs), minicomputers, mainframe computers, embedded systems, and / or distributed computing environments including a plurality of any of the above systems or devices. Distributed computing environments enable remote computing devices, which are connected to a communication network or other data transmission medium, to perform various tasks. In the distributed computing environment, the program modules, applications, and other data may be stored on local and / or remote computer storage media.

[0186] In its most basic configuration, computing device 300 typically includes at least one processing unit 306 and system memory 304. Depending on the exact configuration and type of computing device, system memory 304 may be volatile (such as random access memory (RAM)), non-volatile (such as read-only memory (ROM), flash memory’, etc.), or some combination of the two. This most basic configuration is illustrated in FIG. 3 by box 302. The processing unit 306 may be a standard programmable processor that performs arithmetic and logic operations necessary’ for operation of the computing device 300. The computing device 300 may also include a bus or other communication mechanism for communicating information among various components of the computing device 300.

[0187] Computing device 300 may have additional features / functionality. For example, computing device 300 may include additional storage such as removable storage 308 and non-removable storage 310 including, but not limited to, magnetic or optical disks or tapes. Computing device 300 may also contain network connect! on(s) 316 that allow the device to communicate with other devices. Computing device 300 may also have inputdevice(s) 314 such as a keyboard, mouse, touch screen, etc. Output device(s) 312 such as a display, speakers, printer, etc. may also be included. The additional devices may be connected to the bus in order to facilitate communication of data among the components of the computing device 300. All these devices are well known in the art and need not be discussed at length here.

[0188] The processing unit 306 may be configured to execute program code encoded in tangible, computer-readable media. Tangible, computer-readable media refers to any media that is capable of providing data that causes the computing device 300 (i.e., a machine) to operate in a particular fashion. Various computer-readable media may be utilized to provide instructions to the processing unit 306 for execution. Example tangible, computer- readable media may include, but is not limited to, volatile media, non-volatile media, removable media and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. System memory 304, removable storage 308, and non-removable storage 310 are all examples of tangible, computer storage media. Example tangible, computer-readable recording media include, but are not limited to, an integrated circuit (e.g., field-programmable gate array or application-specific IC), a hard disk, an optical disk, a magneto-optical disk, a floppy disk, a magnetic tape, a holographic storage medium, a solid- state device, RAM, ROM, electrically erasable program read-only memory (EEPROM), flash memory or other memory technology. CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices.

[0189] In an example implementation, the processing unit 306 may execute program code stored in the system memory 304. For example, the bus may carry data to the system memory 304, from which the processing unit 306 receives and executes instructions. The data received by the system memory 304 may optionally be stored on the removable storage 308 or the non-removable storage 310 before or after execution by the processing unit 306.

[0190] It should be understood that the various techniques described herein may be implemented in connection with hardware or software or, where appropriate, with a combination thereof Thus, the methods and apparatuses of the presently disclosed subject matter, or certain aspects or portions thereof, may take the form of program code (i.e., instructions) embodied in tangible media, such as floppy diskettes, CD-ROMs, hard drives, or any other machine-readable storage medium wherein, when the program code is loaded intoand executed by a machine, such as a computing device, the machine becomes an apparatus for practicing the presently disclosed subject matter. In the case of program code execution on programmable computers, the computing device generally includes a processor, a storage medium readable by the processor (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. One or more programs may implement or utilize the processes described in connection with the presently disclosed subject matter, e.g., through the use of an application programming interface (API), reusable controls, or the like. Such programs may be implemented in a high level procedural or object-oriented programming language to communicate with a computer system. However, the program(s) can be implemented in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language and it may be combined with hardware implementations.

[0191] Machine Learning. “Artificial intelligence” as used in this specification, refers to computational techniques (e.g., computer systems and computer-implemented methods) that can approximate cognitive functions including perception, pattern recognition, prediction, and / or planning. Non-limiting examples include rule-based expert systems, statistical and symbolic machine-learning methods, representation-learning approaches, and multilayer or “deep” neural networks. Machine learning (ML) is a subset of Al where statistical models iteratively improve their performance at a task by optimizing objective functions derived from data.

[0192] Example Machine learning techniques include, but are not limited to, logistic-regression classifiers, support vector machines (SVMs), decision-tree ensembles (e.g., random forest, gradient-boosted trees), probabilistic models such as Naive Bayes, and artificial neural networks. “Representation Learning” denotes ML techniques that automatically derive intermediate features paces from raw observations, which can obviate manual feature engineering for downstream prediction or classification tasks. Representation learning techniques include, but are not limited to, autoencoders and embeddings. The term “deep learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, classification, etc., using layers of processing. Deep learning techniques include, without limitation, convolutional, recurrent, transformer-based, and multilayer-perceptron neural-network architectures.

[0193] Machine learning paradigms include supervised, semi-supervised, and unsupervised learning. In a supervised learning, a model is trained on labeled examples (x, y)to approximate a target function f : X -> Y that generalizes to unseen inputs. In unsupervised learning, the algorithm infers latent structures such as clusters, manifolds, or density estimates from unlabeled data. In a semi-supervised model, the model leams a function that maps an input (also known as a feature or features) to an output (also known as a target) during training with both labeled and unlabeled data.

[0194] Neural Networks. An artificial neural network (ANN) model comprises a set of parameterized units ("‘neurons7’) arranged in layers. Connections between layers can be dense (fully connected) or sparse (e.g., convolutional kernels, attention heads), depending on the architecture. This disclosure contemplates that the nodes can be implemented using a computing device (e.g., a processing unit and memory' as described herein). The nodes can be arranged in a plurality of layers such as an input layer, an output layer, and optionally one or more hidden layers with different activation functions. An ANN having hidden layers can be referred to as a deep neural network or multilayer perceptron (MLP). Each node is connected to one or more other nodes in the ANN. For example, each layer is made of a plurality of nodes, where each node is connected to all nodes in the previous layer. The nodes in a given layer are not interconnected with one another, i.e.. the nodes in a given layer function independently of one another. As used herein, nodes in the input layer receive data from outside of the ANN, nodes in the hidden layer(s) modify the data between the input and output layers, and nodes in the output layer provide the results. Each node is configured to receive an input, implement an activation function (e.g., binary step, linear, sigmoid, tanh, or rectified linear unit (ReLU)), and provide an output in accordance with the activation function. Additionally, each node is associated with a respective weight. ANNs are trained with a dataset to maximize or minimize an objective function. In some implementations, the objective function is a cost function, which is a measure of the ANN’S performance (e.g., error such as LI or L2 loss) during training, and the training algorithm tunes the node weights and / or bias to minimize the cost function. This disclosure contemplates that any algorithm that finds the maximum or minimum of the objective function can be used for training the ANN. ANN parameters are typically optimized by stochastic-gradient-descent variants that employ back-propagation of error derivatives (e.g., Adam, RMSProp, L-BFGS, etc.). It should be understood that an ANN is provided only as an example machine learning model. This disclosure contemplates that the machine learning model can be any supervised learning model, semi-supervised learning model, or unsupervised learning model. Optionally, the machine learning model is a deep learning model. Machine learning models are known in the art and are therefore not described in further detail herein.

[0195] A convolutional neural network (CNN) is a ty pe of deep ANN whose layers apply learned convolutional filters across one or more spatial dimension (e.g., 2-D images. 1- D time series, and / or 3-D volumetric data). Feature maps produced by the filters are typically channel-wise stacked, providing width-by-height-by-channel tensors. CNNs can include different types of layers, e.g., convolutional, pooling, and fully-connected (also referred to herein as “dense”) layers. A convolutional layer includes a set of filters and performs the bulk of the computations. A pooling layer is optionally inserted between convolutional layers to reduce the computational power and / or control overfitting (e.g., by downsampling). A fully- connected layer includes neurons, where each neuron is connected to all of the neurons in the previous layer. The layers are stacked similar to traditional neural networks. GCNNs (graph convolutional neural networks) are CNNs that have been adapted to work on structured datasets such as graphs. Graph-convolutional networks (GCNs) generalize the convolution operation to non-Euclidean domains such as graphs, enabling message passing along edges to learn node or graph-level embeddings.

[0196] Other Supervised Learning Models. A logistic-regression (LR) classifier is a linear model that maps features to the log-odds of a binary target via the sigmoid (logistic) function and can be trained by minimizing cross-entropy loss. LR classifiers are trained with a data set (also referred to herein as a “dataset”) to maximize or minimize an objective function, for example, a measure of the LR classifier's performance (e.g., error such as LI or L2 loss), during training. This disclosure contemplates that any algorithm that finds the minimum of the cost function can be used. LR classifiers are known in the art and are therefore not described in further detail herein.

[0197] A Naive Bay es’ (NB) classifier applies Bayes’ rule under a conditionalindependence assumption where each feature can be modeled as statistically independent given the class label. NB classifiers are trained with a data set by computing the conditional probability7distribution of each feature given a label and applying Bayes’ Theorem to compute the conditional probability7distribution of a label given an observation. NB classifiers are known in the art and are therefore not described in further detail herein.

[0198] A k-nearest-neighbor (k-NN) classifier assigns a query sample the majority label its k closest training samples as measured by a distance metric (e.g.. Euclidean or cosine). The k-NN classifiers are trained with a data set (also referred to herein as a “dataset”) to maximize or minimize a measure of the k-NN classifier’s performance during training. This disclosure contemplates any algorithm that finds the maximum or minimum.The k-NN classifiers are known in the art and are therefore not described in further detail herein.

[0199] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

[0200] References1. B. L. Hie et al., Efficient evolution of human antibodies from general protein language models. Nat Biotechnol 42, 275-283 (2024).2. T. A. Desautels et al., Computationally restoring the potency of a clinical antibody against Omicron. Nature, (2024).3. A. Shanehsazzadeh et al., In vitrovalidated antibody design against multiple therapeutic antigens using generative inverse folding. bioRxiv, 2023.2012.2008.570889 (2023).4. M. Haraldson Hoie et al., AntiFold: Improved antibody structure-based design using inverse folding. 2024 (10.48550 / arXiv.2405.03370).5. B. L. Hie et al., Efficient evolution of human antibodies from general protein language models. Nature Biotechnology 42, 275-283 (2024).6. N. R. Bennett et al., Atomically accurate de novo design of single-domain antibodies. bioRxiv, 2024.2003.2014.585103 (2024).7. R. W. Shuai, J. A. Ruffolo, J. J. Gray, IgLM: Infilling language modeling for antibody sequence design. Cell Syst 14, 979-989. e974 (2023).8. E. Nijkamp, J. A. Ruffolo, E. N. Weinstein, N. Naik, A. Madani, ProGen2: Exploring the boundaries of protein language models. Cell Systems 14, 968-978.e963 (2023).9. M. I. J. Raybould, A. Kovaltsuk. C. Marks, C. M. Deane, CoV-AbDab: the coronavirus antibody database. Bioinformatics 37, 734-735 (2020).10. J. Dunbar et al., SAbDab: the structural antibody database. Nucleic Acids Research 42, D1140D1146 (2013).11. B. Abanades et al.. The Patent and Literature Antibody Database (PLAbDab): an evolving reference set of functionally diverse, literature-annotated antibody sequences and structures. Nucleic Acids Research 52, D545-D551 (2023).12. M. S. Gilman et al.. Rapid profiling of RSV antibody repertoires from the memory B cells of naturally infected adult donors. Sci Immunol 1, (2016).13. Y. Zurbuchen et al., Human memory B cells show plasticity and adopt multiple fates upon recall response to SARS-CoV-2. Nature Immunology 24, 955-965 (2023).14. K. J. Kramer et al., Single-cell profiling of the antigen-specific response to BNT162b2 SARS-CoV-2 RNA vaccine. Nature Communications 13. 3466 (2022).15. A. Shanehsazzadeh et al., Unlocking de novo antibody design with generative artificial intelligence. bioRxiv, 2023.2001.2008.523187 (2024).S. F. Andrews et al., Immune history profoundly affects broadly protective B cell responses to influenza. Sci Transl Med 7, 316ral92 (2015). M. G. Joyce et al., Vaccine-Induced Antibodies that Neutralize Group 1 and Group 2 Influenza A Viruses. Cell 166, 609-623 (2016). T. Weber et al., Analysis of antibodies from HCV elite neutralizers identifies genetic determinants of broad neutralization. Immunity 55, 341-354. e347 (2022). Z. A. Bomholdt et al., Isolation of potent neutralizing antibodies from a survivor of the 2014 Ebola virus outbreak. Science 351. 1078-1083 (2016). I. Setliff et al., High-Throughput Mapping of B Cell Receptor Sequences to Antigen Specificity. Cell 179, 1636-1646.el615 (2019). L. M. Walker et al., High-Throughput B Cell Epitope Determination by Next- Generation Sequencing. Front Immunol 13, 855772 (2022). E. C. Chen et al., Systematic analysis of human antibody response to ebolavirus glycoprotein shows high prevalence of neutralizing public cl ono t pes. Cell Rep 42, 112370 (2023). A. R. Shiakolas et al., Efficient discovery of SARS-CoV -2 -neutralizing antibodies via B cell receptor sequencing and ligand blocking. Nature Biotechnology 40, 1270-1275 (2022). A. R. Shiakolas et al., Cross-reactive coronavirus antibodies with diverse epitope specificities and Fc effector functions. Cell Reports Medicine 2. 100313 (2021). M. P. Lefranc et al., IMGT, the international ImMunoGeneTics information system. Nucleic Acids Res 37, D1006-1012 (2009). D. Prihoda et al., BioPhi: A platform for antibody design, humanization, and humanness evaluation based on natural antibody repertoires and deep learning. MAbs 14, 2020203 (2022). D. Pinto et al., Cross-neutralization of SARS-CoV-2 by a human monoclonal SARS- CoV antibody. Nature 583, 290-295 (2020). D. Apiyo, "Biomolecular Binding Kinetics Assays on the Octet® BLI Platform," (2022). J. S. McLellan et al.. Structure-based design of a fusion glycoprotein vaccine for respiratory syncytial virus. Science 342, 592-598 (2013). E. C. Chen et al., Convergent antibody responses to the SARS-CoV-2 spike protein in convalescent and vaccinated individuals. Cell Rep 36, 109604 (2021). I. Setliff et al., Multi-Donor Longitudinal Antibody Repertoire Sequencing Reveals the Existence of Public Antibody Clonotypes in HIV-1 Infection. Cell Host Microbe 23, 845-854.e846 (2018). S. C. Wall et al., SARS-CoV-2 antibodies from children exhibit broad neutralization and belong to adult public clonotypes. Cell Reports Medicine 4. 101267 (2023). T. M. Uyeki et al., Highly Pathogenic Avian Influenza A(H5N1) Virus Infection in a Dairy7Farm Worker. New England Journal of Medicine 390, 2028-2029 (2024).Y. Medina-Armenteros, D. Cajado-Carvalho, R. das Neves Oliveira, M. Apetito Akamatsu, P. Lee Ho, Recent Occurrence, Diversity, and Candidate Vaccine Virus Selection for Pandemic H5N1 : Alert Is in the Air. Vaccines 12, 1044 (2024). M. Yuan et al., Structural basis of a shared antibody response to SARS-CoV-2. Science 369, 1119-1123 (2020). C. Dreyfus et al., Highly conserved protective epitopes on influenza B viruses. Science 337, 1343-1348 (2012). A. Krarup et al., A highly stable prefusion RSV F vaccine derived from structural analysis of the fusion mechanism. Nature communications 6, 8143 (2015). J. A. Lopez et al., Antigenic structure of human respiratory7syncytial virus fusion glycoprotein. Journal of virology 72, 6922-6928 (1998). L. Anderson, J. C. Hierholzer, Y. Stone, C. Tsou, B. Femie, Identification of epitopes on respiratory7syncytial virus proteins by competitive binding immunoassay. Journal of clinical microbiology 23, 475-480 (1986). I. Rossey, J. S. McLellan, X. Saelens, B. Schepens, Clinical potential of prefusion RSV F-specific antibodies. Trends in microbiology 26, 209-219 (2018). D. A. Dougan, R. L. Malby, L. C. Gruen, A. A. Kortt. P. J. Hudson, Effects of substitutions in the binding surface of an antibody on antigen affinity7. Protein Eng 11, 65-74 (1998). K. Winkler et al.. Changing the antigen binding specificity by single point mutations of an antip24 (HIV-1) antibody. The Journal of Immunology7165, 4505-4514 (2000). J. Foote, G. Winter, Antibody framework residues affecting the conformation of the hypervariable loops. Journal of molecular biology 224, 487-499 (1992). F. Klein et al ., Somatic mutations of the immunoglobulin framework are generally required for broad and potent HIV-1 neutralization. Cell 153, 126-138 (2013). 1. Rossey et al.. A vulnerable, membrane-proximal site in human respiratory syncytial virus F revealed by7a prefusion-specific single-domain antibody. Journal of virology 95, 10. 1128 / jvi. 02279-02220 (2021). S. Bubeck et al., Sparks of artificial general intelligence: Early experiments with gpt- 4. arXiv preprint arXiv:2303. 12712, (2023). H. Naveed et al., A comprehensive overview of large language models. arXiv preprint arXiv:2307.06435, (2023). A. Power, Y. Burda, H. Edwards, I. Babuschkin, V. Misra, Grokking: Generalization beyond overfitting on small algorithmic datasets. arXiv preprint arXiv:2201.02177, (2022). S. A. Ehrhardt et al., Polyclonal and convergent antibody response to Ebola virus vaccine rVSVZEBOV. Nature Medicine 25, 1589-1600 (2019). Y. Liu et al., Cross-lineage protection by human antibodies binding the influenza B hemagglutinin. Nature Communications 10, 324 (2019). T. U. Consortium, UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Research 51. D523-D531 (2022).F. Teufel et al., SignalP 6.0 predicts all five types of signal peptides using protein language models. Nature Biotechnology 40, 1023-1025 (2022). T. Wolf et al., Huggingface's transformers: State-of-the-art natural language processing. arXiv preprint arXiv: 1910.03771, (2019). J. Dunbar, C. M. Deane, ANARCI: antigen receptor numbering and receptor classification. Bioinformatics 32, 298-300 (2015). D. J. Sheward et al., Omicron sublineage BA.2.75.2 exhibits extensive escape from neutralising antibodies. Lancet Infect Dis 22, 1538-1540 (2022). A. Creanga et al., A comprehensive influenza reporter virus panel for high-throughput deep profiling of neutralizing antibodies. Nature Communications 12, 1722 (2021). 1. S. Georgiev et al., Single-Chain Soluble BG505.SOSIP gpl40 Trimers as Structural and Antigenic Mimics of Mature Closed HIV-1 Env. J Virol 89, 5318-5329 (2015). A. A. Abu-Shmais et al., Antibody sequence determinants of viral antigen specificity. mBio 0, e01560-01524. S. A. Rush et al.. Characterization of prefusion-F-specific antibodies elicited by natural infection with human metapneumovirus. Cell Rep 40, 111399 (2022). J. S. McLellan et al., Structure of RSV fusion glycoprotein trimer bound to a prefusion-specific neutralizing antibody. Q. Zhu et al., A highly potent extended half-life antibody as a potential RSV vaccine surrogate for all infants. Sci Transl Med 9. (2017). M. M. Leuthold, A. D. Koromyslova, B. K. Singh, G. S. Hansman, Production of Human Norovirus Protruding Domains in E. coli for X-ray Crystallography. JoVE, e53845 (2016). X. BrocheL M. P. Lefranc. V. Giudicelli, IMGT / V-QUEST: the highly customized and integrated system for IG and TR standardized V-J and V-D-J sequence analysis. Nucleic Acids Res 36, W503-508 (2008). D. N. Mastronarde. Automated electron microscope tomography using robust prediction of specimen movements. Journal of structural biology 152, 36-51 (2005). A. Punjani, J. L. Rubinstein, D. J. Fleet, M. A. Brubaker, cryoSPARC: algorithms for rapid unsupervised cryo-EM structure determination. Nature methods 14, 290-296 (2017). J. L. Rubinstein, M. A. Brubaker, Alignment of cryo-EM movies of individual particles by optimization of image translations. Journal of structural biology 192, 188- 195 (2015). R. Sanchez-Garcia et al., DeepEMhancer: a deep learning solution for cryo-EM volume postprocessing. Communications biology 4, 874 (2021). J. Abramson et al., Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 1-3 (2024). E. F. Pettersen et al., UCSF ChimeraX: Structure visualization for researchers, educators, and developers. Protein science 30, 70-82 (2021).P. D. Adams et al., PHENIX: building new software for automated crystallographic structure determination. Acta Crystallographica Section D: Biological Crystallography 58, 1948-1954 (2002). P. Emsley, K. Cowtan, Coot: model-building tools for molecular graphics. Acta crystallographica section D: biological crystallography 60, 2126-2132 (2004). T. I. Croll, ISOLDE: a physically realistic environment for model building into low- resolution electron-density maps. Acta Crystallographica Section D: Structural Biology 74, 519-530 (2018).

Claims

WHAT IS CLAIMED:

1. A computer-implemented method of training a protein large language model for antibody generation, the computer-implemented method comprising: receiving a training corpus comprising a plurality of antigen amino acid sequences and a corresponding plurality7of antibody sequences; generating a plurality7of training strings, wherein each training string comprises an antigen sequence and a corresponding antibody sequence; fine-tuning a pretrained protein large language model by autoregressively training on the plurality7of training strings; and outputting a fine-tuned protein large language model configured to output human antibodies in response to receiving an input string comprising an antigen.

2. The computer-implemented method of claim 1, wherein the plurality of antibody sequences comprise a plurality of light-chain sequences and a plurality of heavychain sequences.

3. The computer-implemented method of claim 1 or claim 2, wherein the corresponding antibody sequence comprises a heavy chain sequence and light chain sequence.

4. The computer-implemented method of claim 3, wherein the plurality of training stnngs further comprise a plurality of special tokens to separate the heavy chain sequence and light chain sequence.

5. The computer-implemented method of any one of claims 1-4, wherein the pretrained protein language model comprises a transformer model.

6. The computer-implemented method of any one of claims 1-5, wherein the training corpus comprises a dataset of antigen-specific antibody sequences.

7. A computer-implemented method of generating an antibody, the method comprising: inputting an antigen sequence to a fine-tuned pretrained protein language model, wherein the fine-tuned pretrained protein language model is autoregressively trained on a plurality of antigen ammino acid sequences and corresponding plurality of antibody sequences; outputting, by fine-tuned pretrained protein language model, a plurality of antibody sequences; filtering the plurality of antibody sequences; ranking the plurality of antibody sequences; selecting, based on the ranked antibody sequences, an antibody sequence to neutralize the antigen sequence.

8. The computer-implemented method of claim 7, wherein the antibody sequence comprises a monoclonal antibody.

9. The computer-implemented method of claim 7 or claim 8. wherein the corresponding antibody sequence comprises a heavy chain sequence and light chain sequence.

10. The computer-implemented method of any one of claims 7-9, wherein filtering the plurality' of antibody sequences comprises selecting antibody sequences based on humanness.

11. The computer-implemented method of any one of claims 7-10, wherein filtering the plurality of antibody sequences comprises selecting antibody sequences based on a germline identity of the antibody sequences.

12. The computer-implemented method of any one of claims 7-11, wherein filtering the plurality of antibody sequences comprises selecting antibody sequences based on a length of the antibody sequences.

13. The computer-implemented method of any one of claims 7-12, wherein filtering the plurality of antibody sequences comprises selecting antibody sequences based on a mutational load measurement.

14. The computer-implemented method of any one of claims 7-13, wherein ranking the plurality of antibody sequences comprise applying a diversity-aware heuristic.

15. An isolated recombinant antibody or antigen-binding fragment that specifically binds to avian-influenza H5 hemagglutinin, wherein the antibody compnses: a heavy-chain variable region with a CDR1 comprising a sequence with at least 70% identity to SEQ ID NO: 21; a CDR2 comprising a sequence with at least 70% identity toSEQ ID NO: 23; anda CDR3 comprising a sequence with at least 70% identity' toSEQ ID NO: 25; and a light-chain variable region with a CDL1 comprising a sequence with at least 70% identity to SEQ ID NO: 22; a CDL2 comprising a sequence with at least 70% identity to SEQ ID NO: 24; and a CDL3 comprising a sequence with at least 70% identity to SEQ ID NO: 26.1 . The antibody of claim 15, wherein the heavy-chain variable region comprises a variable heavy chain region comprising a sequence with at least 70 % identity to SEQ ID NO: 25, and wherein the light-chain variable region comprises a variable light chain region comprising a sequence with at least 70 % identity to SEQ ID NO: 26.

17. The antibody of claim 15, wherein the heavy-chain variable region comprises a variable heavy chain region comprising a sequence with at least 70 % identity to SEQ ID NO: 27, and wherein the light-chain variable region comprises a variable light chain region comprising a sequence with at least 70 % identity to SEQ ID NO: 28.

18. The antibody of claim 15, wherein the heavy-chain variable region comprises a variable heavy chain region comprising a sequence with at least 70 % identity to SEQ ID NO: 29, and wherein the light-chain variable region comprises a variablelight chain region comprising a sequence with at least 70 % identity to SEQ ID NO: 30.

19. The antibody of claim 15, wherein the antibody is a monoclonal antibody .

20. The antibody of claim 19, wherein the monoclonal antibody is configured as a pharmaceutical composition.

Citation Information

Patent Citations

  • Antibodies against h5n1 strains of influenza a virus

    US20100278834A1