Methods enabling in vitro sampling of biological sequence models and synthesis of biological molecule libraries at petascales
The use of a manufacturing-aware generative model with variational synthesis enables efficient, large-scale production of high-quality biological sequence libraries, addressing the limitations of conventional methods by integrating chemical synthesis processes and reducing costs.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2026-03-19
AI Technical Summary
Conventional methods for synthesizing biological sequences using generative models are limited by high costs and resource constraints, resulting in small libraries of uniformly random sequences that are not efficiently scalable, and downstream testing is expensive and time-consuming, hindering the discovery of sequences with desired properties.
A method and system utilizing a manufacturing-aware generative model with variational synthesis, integrating knowledge of chemical synthesis processes, allows for the concurrent in silico and in vitro generation of biological sequences, enabling the production of large, high-quality libraries at petascale.
This approach significantly reduces synthesis costs and time while producing diverse, realistic biological sequence libraries, overcoming the limitations of conventional methods by efficiently generating and validating sequences at a scale comparable to internet-scale datasets.
Smart Images

Figure US2025046210_19032026_PF_FP_ABST
Abstract
Description
[0001]Attorney Docket No. 063640-514001WO METHODS AND SYSTEMS ENABLING IN VITRO SAMPLING OF BIOLOGICAL SEQUENCE MODELS AND SYNTHESIS OF BIOLOGICAL MOLECULE LIBRARIES AT PETASCALES CROSS REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of US Provisional Application No. 63 / 693,950, filed on September 12, 2024, the entire contents of which are hereby incorporated by reference. SEQUENCE LISTING The instant application contains a Sequence Listing which has been filed electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on September 12, 2025, is named 063640-514001WO_Seq-Listing_ST26.xml and is 9,800 bytes in size. TECHNICAL FIELD The present disclosure is directed to generating biological molecule libraries designed by, for example, generative models. BACKGROUND Generative models used to design novel and functional biological sequences are constrained by the cost of building the functional biological sequences that are generated by the generative models. Downstream testing of the designs for biological sequences generated by the generative models are expensive and require resources and time. This limits the ability to discover sequences with desired properties, and limits the ability to improve and refine the machine learning models. The conventional approach to synthesizing generated designs is to draw samples computationally and then synthesize each of these designs individually. While generative models can produce astronomical numbers of novel sequences, and high-throughput experimental assays can evaluate billions of candidates or more, exact synthesis of individual sequences is limited by cost, and as a result most libraries do not exceed 105candidates in practice. Degenerate codon methods can produce larger libraries but they are uniformly random with no connection to a given generative model. Attorney Docket No. 063640-514001WO SUMMARY The present disclosure describes methods and systems for generative, a novel procedure for building generative biological sequence models and synthesizing samples from those models. The present disclosure demonstrates, inter alia, successful manufacturing of about of 10-100 quadrillion (1016-1017) samples from generative models in the real world across several important applications in protein design using a generative model. Disclosed are improved procedures for building generative biological sequence models and synthesizing samples from those models. In some aspects, provided herein is a method of designing a library of nucleic acids or peptides in silico and synthesizing the library in vitro, optionally at petascale, the method including: (a) assembling a training dataset including nucleotide or amino acid sequences; (b) training a generative model on the training dataset, wherein an architecture of the generative model is a manufacturing-aware architecture suited to an in vitro synthesis platform configured to execute stochastic chemical reactions of synthesizing nucleic acids or peptides; (c) generating sequences of the nucleic acids or peptides in silico; (d) mapping in silico parameters of the generative model onto parameters of the in vitro synthesis platform and generating in vitro synthesis protocols; and (e) designing and synthesizing the library of the nucleic acids or peptides in silico and in vitro, optionally at petascale, based at least in part on learned parameters of the generative model and the in vitro synthesis protocols by controlling the stochastic chemical reactions with the in silico parameters. In some instances, the in vitro synthesis and the in silico design occur concurrently. In some aspects, provided herein is a method of designing one or more nucleic acid, peptide, or polypeptide sequence(s) in silico for in vitro production, the method including: (a) assembling a training dataset including nucleotide or amino acid sequences; (b) training a generative model on the training dataset, wherein an architecture of the generative model is a manufacturing-aware architecture configured to execute stochastic chemical reactions of synthesizing the nucleic acids or peptides; optionally wherein the manufacturing-aware architecture is suited to an in vitro synthesis platform; and (c) generating the sequence(s) of the one or more nucleic acids, peptides, or polypeptides in silico; optionally wherein the sequences of the nucleic acids, the peptides, or the polypeptides form a library, which optionally is at petascale. In some instances, the in vitro synthesis and the in silico design occur concurrently. Such a method may further comprise mapping in silico parameters of the generative model onto the parameters of the in vitro synthesis platform and generating in vitro synthesis protocols. In some instances, the one or more nucleic acids, peptides, or polypeptides are Attorney Docket No. 063640-514001WO synthesized based at least in part on learned parameters of the generative model and the in vitro synthesis protocols by controlling the stochastic chemical reactions with the in silico parameters. In some embodiments, the training dataset used in any of the methods disclosed herein may include a validation dataset for tuning the generative model parameters to prevent overfitting. In some embodiments,, the training step of any of the methods disclosed herein may further include optimizing the in vitro parameters to match a distribution of the nucleotide or amino acid sequences in the training dataset. In any of the methods disclosed herein, the sequences of the in silico and in vitro synthesized nucleic acids, peptides, or optionally polypeptides are fed back into the generative model for iterative model improvement. In some embodiments, the method disclosed herein may further include validating the library. In some examples, the validating step may comprise sequencing a sample of the in vitro synthesized nucleic acids or peptides. In some instances, a subset of the nucleotide or amino acid sequences used in the method disclosed herein is held out as a heldout dataset. In some examples, the generative model's predictive performance is evaluated on the heldout dataset. In some examples, such a method may further comprise assessing distributional match between the samples of the in vitro or in silico synthesized nucleic acids or peptides and the heldout dataset. In specific examples, the distributional match can be assessed with summary statistic evaluations and nonparametric evaluations. In some instances, the nonparametric evaluations are Bayesian Embedded AutoRegressive (BEAR) and Maximum Mean Discrepancy (MMD) tests. In some embodiments, the generative model used in the method disclosed herein generates a large, high quality dataset for training a machine learning model that maximizes distributional overlaps with target data and minimizes generalization errors. In some embodiments, the training dataset used in the method disclosed herein includes nucleotide sequences. In other embodiments, the training dataset used in the method disclosed herein includes amino acid sequences. In some embodiments, the nucleotide or amino acid sequences in the method disclosed herein can be generated by another generative machine learning model, a supervised machine learning model, a virtual screen or predictor, or an experimental screen or assay. In some embodiments, the nucleic acid or amino acid sequences of the training dataset used in the method disclosed herein may include non-natural sequences. In some Attorney Docket No. 063640-514001WO examples, the non-natural sequences represent non-canonical nucleic acids or peptides. Alternatively, or in addition, the nucleic acid library or the peptide library synthesized in vitro by any of the methods disclosed herein may include non-natural nucleic acids or non-natural peptides. In some examples, the non-natural nucleic acids or non-natural peptides can be non- canonical nucleic acids or peptides. In some examples, the in silico designed and / or in vitro synthesized library produced by the method disclosed herein includes nucleic acids. In some examples, the library includes DNAs. In other examples, the library includes RNAs. In some instances, the nucleic acids in the library synthesized in vitro encoding proteins or functional fragments thereof. Examples include, but are not limited to, antibodies or heavy and / or light chain complementarity determining regions thereof (e.g., heavy chain CDR3 regions), enzymes or functional domains thereof, neoantigens, therapeutic proteins, and antigenic epitopes such as epitopes capable of being presented by a specific HLA molecule. In other instances, the in vitro synthesized library comprises peptides, for examples, peptides less than 50-amino acids in length (e.g., <30-amino acids in length, <20-amino acids in length, or peptides with 5-20 amino acids in length). In further aspects, the present disclosure provides a system for synthesizing a library of nucleic acids or peptides in vitro, optionally at petascale, the system including: (a) a generative model with a manufacturing-aware architecture trained on a training dataset including nucleotide or amino acid sequences; and (b) an in vitro synthesis platform configured to execute stochastic chemical reactions of synthesizing the library of nucleic acids or peptides in vitro optionally at petascale based at least in part on learned parameters of the generative model and in vitro synthesis protocols; wherein in silico parameters of the generative model map onto the parameters of the in vitro synthesis platform. In other aspects, the present disclosure provides a system for designing one or more nucleic acid, peptide, or polypeptide sequences in silico for production of the one or more nucleic acids, peptides, or polypeptides, the system including: a generative model with a manufacturing-aware architecture trained on a training dataset including biological sequences; optionally wherein the system is for designing a library of nucleic acids, peptides, or polypeptides, optionally at petascale. In some examples, the system may further comprise an in vitro synthesis platform configured to execute stochastic chemical reactions of synthesizing the library of nucleic acids or peptides in vitro optionally at petascale based at least in part on learned parameters of the generative model and in vitro synthesis protocols. In some instances, in silico parameters of the generative model of the system map onto the parameters of the in vitro synthesis platform. Attorney Docket No. 063640-514001WO In any of the systems disclosed herein, the training dataset may include a validation dataset for tuning the generative model parameters to prevent overfitting. Alternatively, or in addition, the training of the generative model with the training dataset may include or further include optimizing the parameters of the in vitro synthesis platform to match a distribution of the biological sequences in the training dataset. In any of the system disclosed herein, the sequences of the in silico and in vitro synthesized nucleic acids and peptides may be fed back into the generative model for iterative model improvement. In some embodiments, the system disclosed herein may further include a validation system. In some examples, the validation system includes sequencing a sample of the in vitro synthesized nucleic acids or peptides. In some instances, a subset of the nucleotide or amino acid sequences used in the system disclosed herein can be held out as a heldout dataset. In some instances, the generative model's predictive performance is evaluated on the heldout dataset. In some examples, the validation system may include a system for assessing distributional match between the samples of the in vitro or in silico synthesized nucleic acids or peptides and the heldout dataset or the training dataset. For example, the distributional match can be assessed with summary statistic evaluations and nonparametric evaluations. In some instances, the nonparametric evaluations are Bayesian Embedded AutoRegressive (BEAR) and Maximum Mean Discrepancy (MMD) tests. In some embodiments, the trained generative model used in the system disclosed herein may generate a large, high-quality dataset for training a machine learning model that maximizes distributional overlaps with a target data and minimizes generalization errors. In some embodiments, the training dataset in the system disclosed herein may comprise nucleotide sequences. Alternatively, the training dataset in the system may comprise peptide sequences. In some examples, the nucleotide or amino acid sequences can be generated by another generative machine learning model, a supervised machine learning model, a virtual screen or predictor, or an experimental screen or assay. In some instances, the nucleotide or amino acid sequences of the training dataset may include non-natural nucleotide or amino acid sequences, for example, non-natural sequences representing non-canonical nucleic acids or peptides. In some embodiments, the library of the system disclosed herein is a library of nucleic acids. In some examples, the library is a DNA library. In some specific examples, the DNA library may comprise nucleic acids encoding antibodies, antigenic epitopes, enzymes, Attorney Docket No. 063640-514001WO neoantigens, or therapeutic proteins. Moreover, the present disclosure features a method comprising: obtaining a training dataset including biological sequences; training a generative model with manufacturing-aware architecture to output one or more biological sequences and one or more instructions for synthesis of the one or more biological sequences; and generating a library of biological sequences based on the training dataset by synthesizing the output biological sequences output by the generative model by executing the output synthesis instructions. Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material. BRIEF DESCRIPTION OF THE DRAWINGS The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which: FIGs.1A-1C include diagrams illustrating an exemplary process flow and system schematics of exemplary processes as disclosed herein. FIG.1A: Schematic representation of a generative model sampling process for generating libraries of biological sequences using a manufacturing-aware architecture and synthesis of DNA samples, followed by protein expression from the synthesized DNA samples. FIG.1B: a flow chart representing a process for utilizing a manufacturing-aware, generative, variational synthesis model, in accordance with some embodiments of the present disclosure. FIG.1C: a flow chart representing a method Attorney Docket No. 063640-514001WO of generating a library of biological sequences using a manufacturing-aware, generative, variational synthesis model, in accordance with some embodiments of the present disclosure. FIG.2 is a functional block diagram of a machine in the example form of computer system, within which a set of instructions for causing the machine to perform any one or more of the methodologies, processes or functions discussed herein may be executed in accordance with some embodiments of the present disclosure. FIGs.3A–3I include diagrams showing an exemplary process flow and data comparing sequences generated in silico or manufactured in vitro by the variational synthesis (VS) models of the disclosure as compared with the held out data. FIG.3A: a flow chart representing the experimental design for generating antibody libraries. FIG.3B: a plot showing a low-dimensional representation of sequences from the held-out data. FIG.3C: a plot showing a low-dimensional representation of sequences from the held-out data, together with sequences generated in silico by the variational synthesis model. FIG.3D: a plot showing a low-dimensional representation of sequences from the held-out data, together with sequences manufactured in vitro by variational synthesis model. FIG.3E: a plot showing a low- dimensional representation of sequences from the held-out data, together with sequences manufactured in vitro using degenerate codon synthesis (NNK codons). FIG.3F: a plot showing per-residue discrepancy between generated sequences (in silico or manufactured). FIG.3G: a bar graph showing per-sequence Bayes factor of the BEAR two-sample test, comparing sequences generated in silico (s) or in vitro (m). FIG.3H: a bar graph showing estimated maximum mean discrepancy (MMD) between sequences generated in silico (s) or in vitro (m). FIG.3I: a plot showing the fraction of generated sequences that are considered clearly realistic based on a high MMD witness function score as compared with the held-out data. FIGs.4A–4J include diagrams showing an exemplary process flow and data comparing peptides generated in silico or in vitro by the variational synthesis (VS) models of the disclosure as compared with the held out data. FIG.4A: a flow chart representing the generation of peptides. FIG.4B: a plot showing a low-dimensional representation of sequences from the held-out data. FIG.4C: a plot showing a low-dimensional representation of sequences from the held-out data, together with sequences generated in silico by the variational synthesis model. FIG.4D: a plot showing a low-dimensional representation of sequences from the held-out data, together with sequences manufactured in vitro by variational synthesis model. FIG.4E: a plot showing a low-dimensional representation of sequences from the held- out data set, together with sequences generated in silico with NNK codons. FIG.4F: a plot Attorney Docket No. 063640-514001WO showing the per-residue discrepancy between generated sequences (in silico or manufactured). FIG.4G: a bar graph showing the per-sequence Bayes factor of the BEAR two-sample test, comparing generated sequences (in silico or in vitro). i.s.: in silico; i.v. in vitro. FIG.4H: a bar graph showing the maximum mean discrepancy (MMD) between generated sequences (in silico or in vitro). i.s.: in silico; i.v. in vitro. FIG.4I: a plot showing the fraction of generated sequences (in vitro or in silico) that are considered clearly realistic based on a high MMD witness function score as compared with the held-out data. FIG.4J: a plot showing the fraction of generated sequences (in silico or in vitro) classified as non-, weak, or strong binders by NetMHCpan-4.1 as compared with the held-out data. FIGs.5A–5I include diagrams showing an exemplary process flow and data comparing final segments for generating CAR antigens by the variational synthesis (VS) models of the disclosure as compared with the held out data. FIG.5A: a flow chart representing a workflow for generating CAR antigens. FIG.5B: a plot showing a low- dimensional representation of sequences from the held-out data. FIG.5C: a plot showing a low-dimensional representation of sequences from the held-out data, together with sequences generated in silico by the variational synthesis model. FIG.5D: a plot showing a low- dimensional representation of sequences from the held-out data, together with sequences manufactured in vitro by variational synthesis. FIG.5E: a plot showing a low-dimensional representation of sequences from the held-out data set, together with sequences generated in silico with NNK codons. FIG.5F: a plot showing the per-residue discrepancy between generated sequences (in silico or manufactured) and the held-out data. FIG.5G: a bar graph showing the per-sequence Bayes factor of the BEAR two-sample test, comparing generated sequences (in silico or in vitro). i.s.: in silico; i.v. in vitro. and held-out data. FIG.5H: a bar graph showing the maximum mean discrepancy (MMD) between generated sequences (in silico or in vitro). i.s.: in silico; i.v. in vitro and the held-out data. FIG.5I: a plot showing the fraction of generated sequences that are considered clearly realistic based on a high MMD witness function score as compared with the held-out data. FIGs.6A–6D include diagrams showing an exemplary process flow, schematic of an exemplary assembled second generation CAR, and data comparing final segments for generating CAR antigens by the variational synthesis (VS) models of the disclosure as compared with the held out data. FIG.6A: a graphical illustration of assembled second generation CAR construct with CDRH3 region (VS) generated from the output of a generative model utilizing variational synthesis and the resulting library as disclosed herein. FIG.6B: a schematic overview of the multiplexed screening assay wherein the CAR library is screened Attorney Docket No. 063640-514001WO against barcoded dextramers. Heatmap shows the raw data from single-cell sequencing of the resultant hits. FIG.6C: a plot showing the fraction of clearly realistic sequences based on the witness function score with specific hits from the screen included (hits). FIG.6D: a plot showing the distribution of witness function score for held-out post-selection human antibody data versus specific hits from the screen (hits). FIG.7 is a plot showing the length distribution in the antibody OAS dataset in accordance with some embodiments of the present disclosure. FIGs.8A -8C are plots showing comparisons between the amino-acid usage and length distributions in the in silico (designed) and in vitro (manufactured) variational synthesis libraries of the disclosure. FIG.8A: Antibody dataset. FIG.8B: Epitope dataset. FIG.8C: DNA polymerase dataset. FIGs.9A-9F are plots showing comparisons between the length distributions in the held-out data compared with the in vitro variational synthesis libraries or in silico the NNK libraries as disclosed herein. FIG.9A: Comparisons between length distributions for held-out data and in vitro variational synthesis data for antibodies. FIG.9B: Comparisons between length distributions for held-out data and in vitro variational synthesis data for epitopes. FIG. 9C: Comparisons between length distributions for held-out data and in vitro variational synthesis data for DNA Polymerases dataset. FIG.9D: Comparisons between length distributions for held-out data and in vitro NKK library data for antibodies. FIG.9E: Comparisons between length distributions for held-out data and in silico NKK library data for epitopes. FIG.9F: Comparisons between length distributions for held-out data and in silico NKK library data for DNA Polymerases. FIGs.10A-10E are plots showing data from witness function analysis of antibodies. FIG.10A: Plot showing shows the percentage of generated sequences (in silico or in vitro) with witness function values above the real-data percentile threshold given on the x-axis. VS: variational synthesis. FIG.10B: Plot showing comparisons between distribution of witness function scores ˆ f(x) for in silico samples from IgLM as compared with held-out data. FIG.10C: Plot showing comparisons between the distribution of witness function scores for variational synthesis in silico samples as compared with held-out data. FIG.10D: Plot showing comparisons between witness function scores for variational synthesis in vitro manufactured samples as compared with held-out data. FIG.10E: Plot showing comparisons between distribution of witness function scores for NNK in vitro manufactured sequences as compared with held-out data. FIGs.11A-11B are plots showing witness function analysis of epitopes and DNA Attorney Docket No. 063640-514001WO polymerases. FIG.11A: Epitopes. FIG.11B: DNA Polymerases. DETAILED DESCRIPTION Disclosed are systems and methods for generating improved libraries of biological sequences. In some implementations, the improved library may be based at least in part on outputs produced by a generative model trained for the synthesis of biological sequences. Biological sequences can include, without limitation, antibodies, human antibodies, peptides with specific binding patterns, tumor antigen binding elements, vaccine candidates, DNA polymerases, and the like. The disclosed systems and methods can produce an improved library of biological sequences having improved quality, including complexity, variability, high accuracy, and minimal errors. As disclosed herein, the sequences generated by a generative model with manufacturing-aware architectures produces improved sequences that can be synthesized more efficiently, by requiring less resources and in shorter time periods. Traditional protein-engineering methods for generating medicines and enzymes are often labor-intensive, time-consuming, and / or limited by understanding of protein function. Conventional generative biological sequence models provide mechanisms for designing novel proteins with desired properties. However, conventional generative biological sequence models remain limited in that they may provide protein designs that cannot be feasibly built in a laboratory setting, requiring expensive and laborious techniques. In some implementations, the generative models described herein can include variational synthesis models which provide many improvements to conventional generative biological sequence models. In some implementations, the variable synthesis models can be constructed with a manufacturing-aware architecture. Variable synthesis models may design and output biological sequences including protein sequences that have comparable or better quality than the outputs of conventional generative models in terms of factors such as realism and diversity. In some implementations, the variable synthesis models may be manufacturing- aware. For example, manufacturing-aware may refer to the variable synthesis model having awareness of the constraints and possibilities of chemical manufacturing processes. In some implementations, a manufacturing-aware variable synthesis model may output not only a target protein sequence but also a series of steps or experimental parameters for manufacturing the target protein sequence. In some implementations, the variable synthesis model may design a DNA sequence and introduce noise via randomness in order to generate novel protein sequences. In some implementations, generative models utilizing variational synthesis can be Attorney Docket No. 063640-514001WO used to generate sequences (e.g., nucleic acid sequences or amino acid sequences) that provide structural information for corresponding biological molecules (e.g., nucleic acid / peptides) in silico. A corresponding in vitro synthesis platform in communication with the outputs of the generative model can be configured to produce biological molecules (e.g., nucleic acids / peptides) based on the output of the in silico process. Resulting protein sequences can be large in scope and size. For example, some implementations of the current subject matter reduce the cost of synthesizing proteins and other biological sequences designed by a generative model by 10-billion-fold. The disclosed systems and methods train a generative model with manufacturing-aware architecture on sequences data. Generative models demonstrated herein by training and synthesizing samples from generative models of e.g., antibodies, T cell antigens, and DNA polymerases. For example, a manufacturing-aware generative model (i.e., generative model) can be trained on 325 million observed human antibodies and ∼1017generated designs from the model can be synthesized, achieving a sample quality comparable to a state-of-the-art protein language model. If using currently available methods, synthesis of a library of the same accuracy and size would cost roughly a quadrillion dollars, by contrast implementations of the current subject matter provide cost, time, and resource savings. The embodiments of the present disclosure are described in greater details below. The following descriptions and examples illustrate embodiments of the present disclosure in detail. Although the present disclosure has been described in some details by way of illustration and example for purposes of clarity and understanding, it will be apparent that certain changes and modifications can be practiced within the scope of the appended claims. The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. Although various features of the disclosure can be described in the context of a single embodiment, the features can also be provided separately or in any suitable combination. Conversely, although the present disclosure can be described herein in the context of separate embodiments for clarity, the present disclosure can also be implemented in a single embodiment. It is to be understood that the present disclosure is not limited to the particular embodiments described herein and as such can vary. Those of skill in the art will recognize that there are variations and modifications of the present disclosure, which are encompassed within its scope. It is intended that every maximum numerical limitation given throughout this Attorney Docket No. 063640-514001WO specification includes every lower numerical limitation, as if such lower numerical limitations were expressly written herein. Every minimum numerical limitation given throughout this specification will include every higher numerical limitation, as if such higher numerical limitations were expressly written herein. Every numerical range given throughout this specification will include every narrower numerical range that falls within such broader numerical range, as if such narrower numerical ranges were all expressly written herein. All patent filings, websites, other publications, accession numbers and the like cited above or below are incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be so incorporated by reference. If different versions of a sequence are associated with an accession number at different times, the version associated with the accession number at the effective filing date of this application is meant. The effective filing date means the earlier of the actual filing date or filing date of a priority application referring to the accession number if applicable. Likewise, if different versions of a publication, website or the like are published at different times, the version most recently published at the effective filing date of the application is meant unless otherwise indicated. Any feature, step, element, embodiment, or aspect of the disclosure can be used in combination with any other unless specifically indicated otherwise. GENERAL DEFINITIONS All terms are intended to be understood as they would be understood by a person skilled in the art. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the disclosure pertains. The following definitions supplement those in the art and are directed to the current application and are not to be imputed to any related or unrelated cases, e.g., to any commonly owned patent or application. Although any methods and materials similar or equivalent to those described herein can be used in the practice for testing of the present disclosure, the preferred materials and methods are described herein. Accordingly, the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting. In this application, the use of the singular includes the plural unless specifically stated otherwise. It must be noted that, as used in the specification, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. Attorney Docket No. 063640-514001WO In this application, the use of “or” means “and / or” unless stated otherwise. The terms “and / or” and “any combination thereof” and their grammatical equivalents as used herein, can be used interchangeably. These terms can convey that any combination is specifically contemplated. Solely for illustrative purposes, the following phrases “A, B, and / or C” or “A, B, C, or any combination thereof” can mean “A individually; B individually; C individually; A and B; B and C; A and C; and A, B, and C”. The term “or” can be used conjunctively or disjunctively, unless the context specifically refers to a disjunctive use. Furthermore, the use of the term “including” as well as other forms, such as “include”, “includes” and “included”, is not limiting. Reference in the specification to “some embodiments”, “an embodiment”, “one embodiment” or “other embodiments” means that a particular feature, structure, or characteristic described in connection with the embodiments is included in at least some embodiments, but not necessarily all embodiments, of the present disclosures. As used in this specification and claim(s), the words “comprising” (and any form of comprising, such as “comprise” and “comprises”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”) or “containing” (and any form of containing, such as “contains” and “contain”) are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. It is contemplated that any embodiment discussed in this specification can be implemented with respect to any method or composition of the disclosure, and vice versa. Furthermore, compositions of the present disclosure can be used to achieve methods of the present disclosure. The term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. In another example, the amount “about 10” includes 10 and any amounts from 9 to 11. In yet another example, the term “about” in relation to a reference numerical value can also include a range of values plus or minus 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% from that value. Alternatively, particularly with respect to biological systems or processes, the term “about” can mean within an order of magnitude, preferably within 5-fold, and more preferably within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range Attorney Docket No. 063640-514001WO for the particular value should be assumed. It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, can also be provided in combination in a single embodiment. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, can also be provided separately or in any suitable sub-combination. All combinations of the embodiments pertaining to the disclosure are specifically embraced by the present disclosure and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub-combinations of the various embodiments and elements thereof are also specifically embraced by the present disclosure and are disclosed herein just as if each and every such sub combination was individually and explicitly disclosed herein. GENERATIVE MODEL WITH MANUFACTURING-AWARE ARCHITECTURE Provided herein is an approach to building generative biological sequence models, which allows their output to be manufactured physically at large scales. In some implementations, knowledge of DNA synthesis is integrated into the model architecture. Generative models built with a manufacturing-aware architecture where knowledge of DNA synthesis is integrated into the model architecture can produce designs that can be just as diverse and realistic as those of other modern generative models. However, their designs can also be synthesized in vitro in extreme parallel. Such models are referred to as generative models herein. FIGs.1A and 1B provides a schematic representation of a process for generating libraries of biological sequences using a generative model with manufacturing-aware architecture. As demonstrated in FIGs.1A and 1B each step of the sampling process can be implemented chemically through a stochastic reaction, adding new nucleotides to growing DNA molecules. The experimental parameters of each reaction, such as the concentration of different nucleotides, can be set based on the learned parameters of the synthesis model. Second, the process can be repeated across a large array of separate reaction compartments, each corresponding to a distinct latent variable zi. The total number of synthesized samples can be equal to the total number of synthesized DNA molecules, ∼1016-1017. Lastly, synthesized DNA can be assembled into relevant backbones, and expressed as proteins for downstream assays. The upper bound for the number of samples can be given by the number of synthesized DNA molecules. A generative model utilizing variational synthesis describe a distribution pθ*(x) over Attorney Docket No. 063640-514001WO amino acid sequences x that can be determined by parameters θ. Each model parameter in θ can correspond to a controllable experimental parameter in a DNA synthesis protocol. The distribution pθ(x) can be the distribution of amino acid sequences produced in the laboratory by reactions that use the experimental parameters θ. As with other generative models, generative models utilizing variational synthesis can be trained on data (e.g., biological sequences) by adjusting the parameters θ to find an optimal θ*, such that pθ*(x) describes the data distribution. With a generative model utilizing variational synthesis, however, samples from the trained model pθ*(x) can be manufactured in the physical world in parallel by running synthesis reactions with the learned parameters θ*. In practice, pθ⋆(^^) takes the form of a mixture model, with hundreds or thousands of mixture components, corresponding to different reaction compartments. This makes pθ⋆(^^) highly expressive, capable of capturing complex correlation structure in the data distribution. The overall workflow of a generative model utilizing variational synthesis can start with a generative model having a manufacturing-aware architecture that is suited to the experimental synthesis platform available. This model can be trained on a large dataset of biological sequence examples. Then, designs from this generative model can be drawn not only in silico but also in vitro, by running DNA synthesis reactions according to the learned parameters of the model. The result can be a very large number of manufactured samples from the generative model, in the real world. In this manner, large libraries of biological sequences having high quality can be generated. While conventional generative models are limited in that the only available way to synthesize designs is to sample designs on the computer, and then synthesize each individually, a generative model with variational synthesis and manufacturing-aware architectures is able to more deeply integrate computational sampling and chemical synthesis. This is due at least in part to the fact that the parameters of a generative model with variational synthesis are not simply un-interpretable weights and biases. Rather, each in silico parameter of the disclosed model maps directly to a specific experimental parameter of the DNA synthesis procedure, such as the concentration of certain reagent, or the timing of a particular reaction. A generative model thus can provide a highly specific experimental protocol along with the generative model, giving step-by-step instructions for large-scale laboratory synthesis of the model’s designs. Variational synthesis models enable synthesis of extremely large numbers of model samples via the controlled use of stochastic chemical reactions. All generative models depend on controlled randomness to produce diverse designs. Conventionally, this randomness is Attorney Docket No. 063640-514001WO introduced by a random number generator on the computer, while chemical synthesis proceeds deterministically. Variational synthesis models can instead exploit the physical randomness inherent in stochastic chemical reactions. As the synthesis process proceeds, each molecule of DNA randomly encounters a different series of reactants. As a result, every DNA molecule corresponds to a new design, an independent sample from the model. So, if a series of synthesis reactions yield a picomole of DNA, six hundred billion samples will be synthesized from the generative model. Despite the randomness, however, the population of DNA molecules as a whole remains highly controlled, accordingly, a generative model utilizing variational synthesis can have the experimental parameters of each stochastic reaction set by the learned parameters of the generative model. Additionally, generative models utilizing variational synthesis can integrate data from many sources beyond just sequences. For example, this can include functional properties of sequences measured by experimental testing of samples from an earlier generative model, or domain-specific proxies of biological function such as 3D protein structure. Further, to address the risk of compounding errors by training one generative model on data produced by another generative model, new ways of training conditioning generative models on auxiliary information or design goals can be used. The disclosed systems and methods can produce large, relevant datasets, or libraries that can be an input for downstream assays and the creation of new datasets. In some implementations, these libraries can be in the petascale size. By training a generative model utilizing variational synthesis on real biological sequences, the resulting datasets can contain realistic, in-distribution examples. By conditioning the generative model utilizing variational synthesis on auxiliary information, the datasets can further be focused on relevant examples rather than having to learn from sparse data or randomly observed examples. By manufacturing sequences efficiently, the output datasets can be very large and fully represent the desired space of sequences to be explored. In short, generative models utilizing variational synthesis opens the door to the creation of vast new datasets about the functional properties of proteins, RNA and DNA. Indeed, with petascale libraries, these biological datasets can begin to rival in size the internet-scale datasets available for images and text. As illustrated in FIGs.1A and 1B, in some implementations a manufacturing-aware, generative, variational synthesis model can be initialized by specifying model hyperparameters and initial weights. The hyperparameters and initial weights may correspond to chemical synthesis parameters used in the synthesis of biological sequences. The initialized model may be trained on sequence data collected from humans or other organisms, or from output from Attorney Docket No. 063640-514001WO another machine learning based model. In some implementations, pθ⋆(^^) is the variational synthesis model and pθ⋆(^^) is the data distribution that the model is learning. In some implementations, the model may output a series of instructions that are capable of being executed by the synthesis platform in order to generate or synthesize a biological sequence output by the model. The model may output sequence designs which can be synthesized or generated by a synthesis platform. Each parameter of the trained variational synthesis model θ★i determines an experimental parameter in a stochastic chemical reaction. To synthesize complete sequence designs, a series of such controlled reactions can be run using the experimental parameters. The process can be repeated across different reaction compartments, with each compartment corresponding to a different latent variable zi in the generative model. The end result, using many reaction compartments, is synthesized samples from the generative model. Each molecule of DNA is an independent sample X ∼ pθ⋆ (^^). The total number of molecules / samples reaches ∼1016-1017. Additionally, sample quality and diversity of the resulting synthesized sequences can be evaluated with summary statistics, non-parametric two- sample tests and the like. In some embodiments, the resulting evaluations can be used to retrain, modify, or update the model. As illustrated in FIG.1C, in some implementations a process for utilizing a manufacturing-aware, generative, variational synthesis model can include inputting a target sequence data for training the manufacturing-aware generative model, executing library- specific step-by-step instructions by following the generative model parameters, validating the library by sequencing and performing downstream experimental tests. FIG.1C provides a flow chart for a method of generating a library of biological sequences using a manufacturing-aware, generative, variational synthesis model. As illustrated in FIG.1C a method can include the steps of: obtaining sequence data 101, training a generative model including variational synthesis on the obtained sequence data and parameters corresponding to synthesis constraints 103, generating one or more biological sequences by applying the trained variational synthesis model 105, and synthesizing the generated one or more biological sequences 107. In some implementations the manufacturing-aware, generative, variational synthesis model utilizes stochastic synthesis methods and procedures for optimizing stochastic synthesis methods to produce biological samples. As opposed to synthesizing large numbers of samples from generative sequence models, a variational synthesis model may utilize stochastic DNA synthesis and optimize the parameters of the laboratory synthesis protocol to produce samples Attorney Docket No. 063640-514001WO from a distribution close the distribution of the target generative model. In this manner, ultra- large scale libraries can be built. For example, ultra-large scale libraries including those at a petascale. In some implementations the manufacturing-aware, generative, variational synthesis model can include a transformer structure with hyperparameters indicative of the parameters that map directly to the actual DNA synthesis process. These parameters are then optimized to maximize the likelihood of producing a desired sequence distribution (i.e., the training data). The parameters provide a set of instructions to construct a pooled library of sequences in the lab that reproduces the target distribution. In some implementations, the disclosed manufacturing-aware, generative model utilizing variational synthesis can be used to synthesize a library of nucleic acids or peptides in silico and in vitro. The resulting library can be utilized in petascale computing. To generate the library, a method can include the steps of: assembling a training dataset comprising nucleotide or amino acid sequences; training a generative model on the training dataset, where the architecture of the generative model is a manufacturing-aware architecture suited to an in vitro synthesis platform configured to execute stochastic chemical reactions of synthesizing nucleic acids or peptides; generating sequences of the nucleic acids or peptides in silico; mapping in silico parameters of the generative model onto parameters of the in vitro synthesis platform and generating in vitro synthesis protocols; and synthesizing the library of the nucleic acids or peptides in silico and in vitro optionally at petascale based at least in part on learned parameters of the generative model and the in vitro synthesis protocols by controlling the stochastic chemical reactions with the in silico parameters. In some implementations, a method for designing one or more nucleic acid, peptide, or polypeptide sequence(s) in silico for in vitro production, can include the steps of: assembling a training dataset comprising nucleotide or amino acid sequences; training a generative model on the training dataset, wherein the architecture of the generative model is a manufacturing-aware architecture configured to execute stochastic chemical reactions of synthesizing the nucleic acids or peptides; optionally wherein the manufacturing-aware architecture is suited to an in vitro synthesis platform; and generating the sequence(s) of the one or more nucleic acids, peptides, or polypeptides in silico; optionally wherein the sequences of the nucleic acids, the peptides, or the polypeptides form a library, which optionally is at petascale. Methods for designing the one or more nucleic acids, peptides, or polypeptide sequence(s) can also include a step of mapping in silico parameters of the generative model Attorney Docket No. 063640-514001WO onto the parameters of the in vitro synthesis platform and generating in vitro synthesis protocols. The resulting one or more nucleic acids, peptides, or polypeptides can be synthesized based at least in part on learned parameters of the generative model and the in vitro synthesis protocols by controlling the stochastic chemical reactions with the in silico parameters. In some implementations, the generative model utilizing variational synthesis can be trained using a training dataset. Optionally, the training dataset can include a validation dataset for tuning the generative model parameters to prevent overfitting. Additionally, training the generative model utilizing variational synthesis can include optimizing the in vitro parameters to match a distribution of the biological sequences in the training dataset. Further, the sequences of the in silico and in vitro synthesized nucleic acids, peptides, or polypeptides can be fed back into the generative model for iterative model improvement, or retraining. In some embodiments, the resulting library can be validated. For example, validating results output by the generative model can include sequencing a sample of the in vitro synthesized nucleic acids or peptides. Validation may be performed on data similar to the training data, such as a subset of the biological sequences in the training data set that can be held out as a heldout dataset. In other words, the generative model’s predictive performance can be evaluated on the heldout dataset. In some embodiments, the generative model can be trained to optimize a distributional match between samples of the in vitro or in silico synthesized nucleic acids or peptides and the heldout dataset. The distributional match can be evaluated or assessed with summary statistic evaluations and nonparametric evaluations. Examples of nonparametric evaluations include but are not limited to Bayesian Embedded AutoRegressive (BEAR) and Maximum Mean Discrepancy (MMD) tests. In some embodiments, the training data may be generated by a separate generative model. For example, a separate generative model may maximize the distributional overlaps within the target data and minimize generalization errors. For example, the biological sequences are generated by another generative machine learning model, a supervised machine learning model, a virtual screen or predictor, or an experimental screen or assay. In some embodiments, the biological sequences of the training dataset can include non-natural biological sequences. Examples include non-canonical nucleic acids or peptides. In some embodiments, the biological sequences of the training dataset can include nucleotide or peptide sequences. As a result, the in vitro synthesized nucleic acids or peptides can include a repertoire of nucleic acids or peptides. In some embodiments, the in silico synthesized nucleic acids or peptides includes a designed library of nucleic acids or peptides Attorney Docket No. 063640-514001WO compiled from the output of the generative models. In some embodiments the in vitro synthesis and the in silico synthesis occur concurrently. The repertoire of nucleic acids or peptides can include non-natural sequences. Examples include non-canonical nucleic acids, peptides, or polypeptides. In some embodiments, the outputs of the generative model utilizing variational synthesis can include in vitro or in silico synthesized nucleic acids. For example, the synthesized nucleic acids can be DNA sequences or RNA sequences. Optionally, the in vitro or in silico synthesized nucleic acids can be a library of hybridization probes or a library of nucleic acid barcodes. The library of hybridization probes and / or the library of nucleic acid barcodes can be used for sequencing or tracking sequence diversities. In some embodiments, the outputs of the generative model utilizing variational synthesis can include a repertoire of a peptides or polypeptides. Examples include an antibody repertoire, an epitope repertoire, an enzyme repertoire, a neoantigen repertoire, a therapeutic protein repertoire, and the like. In some embodiments of the present disclosure, a system can be configured to synthesize a library of nucleic acids or peptides in vitro optionally at petascale. For example, the system can include a generative model coupled with a platform for synthesis. A system can include a generative model with a manufacturing-aware architecture trained on a training dataset comprising nucleotide or amino acid sequences; and an in vitro synthesis platform configured to execute stochastic chemical reactions of synthesizing the library of nucleic acids or peptides in vitro optionally at petascale based at least in part on learned parameters of the generative model and in vitro synthesis protocols. In some implementations, the in silico parameters of the generative model map onto the parameters of the in vitro synthesis platform. The system can be configured to perform the steps of any one of the methods described herein. In some implementations, a generative model utilizing variational synthesis can be trained on a training dataset that includes nucleotide sequences. The generative model may output designs for nucleotide sequences in silico including, for example, coding sequences for proteins or functional fragments thereof. An in vitro synthesis of nucleic acid libraries based on the in silico design (DNA or RNA libraries) can be performed to generate corresponding nucleic acid libraries. In some implementations, a generative model utilizing variational synthesis can be trained on a training dataset that includes amino acid sequences of proteins or functional fragments thereof. The generative model may output designs for nucleotide sequences that encode the proteins or functional fragments in silico. An in vitro synthesis of nucleic acid Attorney Docket No. 063640-514001WO libraries comprising the encoded nucleic acids generated by the model can be performed. In some implementations, a generative model utilizing variational synthesis can be trained on a training dataset including amino acid sequences of peptides. The generative model may generate designs for peptide sequences in silico. A synthesis process may generate a library including peptides based on the in silico design including peptides (e.g., less than 50 amino acids). In some instances, the peptides may have a length ranging from about 5-40 amino acids. In other instances, the peptides may have a length ranging from about 5-30 amino acids. In yet another example, the peptides may have a length ranging from about 5-20 amino acids. FIG.2 illustrates a functional block diagram of a machine in the example form of computer system 1200, within which a set of instructions for causing the machine to perform any one or more of the methodologies, processes or functions discussed herein may be executed. In some examples, the machine may be connected (e.g., networked) to other machines as described above. The machine may operate in the capacity of a server or a client machine in a client-server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine may be any special-purpose machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine for performing the functions described herein. Further, while only a single machine is illustrated, the term “machine” shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein. In some examples, the computing system or processor executing the process of FIG.1C may be implemented by the example machine shown in FIG.2 (or a combination of two or more of such machines). Example computer system 1200 may include processing device 1203, memory 1207, data storage device 1209 and communication interface 1215, which may communicate with each other via data and control bus 1201. In some examples, computer system 1200 may also include display device 1213 and / or user interface 1211. In some embodiments, the user interface 1211 may include a graphical user interface. Processing device 1203 may include, without being limited to, a microprocessor, a central processing unit, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP) and / or a network processor. Processing device 1203 may be configured to execute processing logic 1205 for performing the operations described herein. In general, processing device 1203 may include any suitable special-purpose processing device specially programmed with processing logic 1205 to perform the operations Attorney Docket No. 063640-514001WO described herein. Memory 1207 may include, for example, without being limited to, at least one of a read-only memory (ROM), a random access memory (RAM), a flash memory, a dynamic RAM (DRAM) and a static RAM (SRAM), storing computer-readable instructions 1217 executable by processing device. In general, memory 1207 may include any suitable non- transitory computer readable storage medium storing computer-readable instructions 1217 executable by processing device 1203 for performing the operations described herein. Although one memory device 1207 is illustrated in FIG.2, in some examples, computer system 1200 may include two or more memory devices (e.g., dynamic memory and static memory). Computer system 1200 may include communication interface device 1215, for direct communication with other computers (including wired and / or wireless communication), and / or for communication with a network. In some examples, computer system 1200 may include display device 1213 (e.g., a liquid crystal display (LCD), a touch sensitive display, etc.). In some examples, computer system 1200 may include user interface 1211 (e.g., an alphanumeric input device, a cursor control device, etc.). In some examples, computer system 1200 may include data storage device 1209 storing instructions (e.g., software) for performing any one or more of the functions described herein. Data storage device 1209 may include any suitable non-transitory computer- readable storage medium, including, without being limited to, solid-state memories, optical media and magnetic media. EXAMPLES These Examples are provided for illustrative purposes only and not to limit the scope of the claims provided herein. In some of the following examples auxiliary information was integrated by using an auxiliary model (i.e., NetMHCpan and ProGen). Samples were conditioned and generated from the auxiliary model, then a generative model utilizing variational synthesis was trained to describe those samples. Generative modeling offers a powerful paradigm for designing novel functional DNA, RNA and protein sequences. The disclosed systems and methods efficiently synthesize designs output by generative models in the real world. Embodiments of the present disclosure include an integrated machine learning and wet lab procedure, which implements generative sampling algorithms physically through controlled stochastic chemical reactions. In one example below, ∼1016designs are synthesized from a generative model of human antibodies, Attorney Docket No. 063640-514001WO at a level of realism and diversity comparable to state-of-the-art protein language models. Sequencing verifies the quality of the manufactured designs. The library yields therapeutic scFv CAR candidates against HLA-presented intracellular tumor antigens. The method is effective across diverse sequence families. Generative models have achieved dramatic successes in designing novel and functional biological sequences. Trained on natural data, these models learn the underlying constraints imposed by evolution and capture the range of outstanding biological possibilities. Sampling from the model produces a diverse set of realistic designs. However, downstream testing of these designs is severely constrained by the cost of building them. This bottleneck limits the ability to discover sequences with desired properties, and limits the ability to improve and refine machine learning models. Prior approaches to synthesizing generated designs draw samples computationally and then synthesize each of these designs individually. While generative models can produce astronomical numbers of novel sequences, and high-throughput experimental assays can evaluate billions of candidates or more, exact synthesis of individual sequences is limited by cost, and as a result most libraries do not exceed 105candidates in practice. Degenerate codon methods can produce larger libraries, but they are uniformly random, with no connection to a given generative model. Accordingly, prior approaches are limited by the vast amount of resources required (e.g., cost, memory, processing, speed) and quality of outputs. Disclosed is a model utilizing variational synthesis, a procedure for building generative biological sequence models and synthesizing samples from those models. The below Examples illustrate implementation of the variational synthesis in practice, demonstrating successful manufacturing of about 10-100 quadrillion (1016-1017) samples from generative models in the real world across several important applications in protein design. As will be provided in the examples below, generative models utilizing variational synthesis models achieve large scale synthesis via a manufacturing-aware model architecture, which accounts for the possibilities and constraints of stochastic DNA synthesis technology. It is envisioned that by incorporating DNA assembly or other synthesis technologies, longer sequences (e.g., greater than 60 amino acids) could be designed and synthesized. Indeed, variational synthesis models that control assembly have been demonstrated in silico. Developing and implementing such models in vitro will be important for many application areas going forward. It may also be fruitful to develop novel DNA synthesis technologies that are specifically suited to variational synthesis applications. Variational synthesis models are capable of integrating information beyond just Attorney Docket No. 063640-514001WO biological sequences. For many design applications, generative models that condition on auxiliary information, such as experimental measurements of function, or a proxy of function such as 3D structure can be used. Below, one example integrates such auxiliary information by using auxiliary models (i.e. NetMHCpan and ProGen): first the model is conditioned and generates samples with the auxiliary model, then a variational synthesis model is trained on those samples. Additional methods for conditioning variational synthesis models on auxiliary information and design goals are envisioned. Additionally, for the safe usage of variational synthesis, integration with appropriate biosecurity measures is envisioned. Existing methods and regulations are focused on securing deterministic synthesis orders, by checking if the list of ordered sequences includes a sequence of concern. Variational synthesis is capable of generating accurate sequences, but it uses stochastic synthesis protocols; it produces synthesis orders that correspond to the parameters of a generative model, not individual sequences. Going forward, to secure variational synthesis, methods for protecting against harmful outputs from language models and other generative models can be integrated. The below examples demonstrate the design of a library of nucleic acids or peptides in silico and synthesizing the library in vitro. EXAMPLE 1. Building Manufacturing-Aware Generative Biological Sequence Models One or more manufacturing-aware generative biological sequence models were designed based on different stochastic reaction technologies. Customized synthesis protocols and customized generative models describing the distribution of DNA and amino acid sequences produced by these specific protocols were developed herein. The variation synthesis models developed accounted for complex manufacturing constraints, and the model-designed reactions have been carried out repeatedly at scale with automation. With appropriate calibration and quality control, samples from a specified generative model with 97% accuracy per base could reliably produced (Table 1). Yields on the order of >10 nanomoles of full- length DNA, corresponding to >1016model samples (petascale synthesis) were achieved. Table 1. Manufacturing accuracy of generative models on different libraries Attorney Docket No. 063640-514001WO Optimization procedures were developed which allowed the generative models disclosed herein to be trained on datasets with hundreds of millions of sequences. As with other generative models, GPU-accelerated stochastic optimization with minibatches of data were utilized. However, many experimental parameters in synthesis reactions were discrete and hence did not admit backpropagation. Therefore, conventional stochastic gradient descent methods with stochastic discrete optimization methods were combined, based on adaptive expectation maximization. EXAMPLE 2. Synthesizing a Nucleic Acid Library Encoding a Pan-Human Antibody Repertoire Antibodies are central to the human adaptive immune system and a key source of new medicines, including monoclonal antibody therapies, antibody-drug conjugates, and cell therapies. Large scale antibody synthesis allows for detailed experimental study of antibodies’ diverse properties, enabling characterization of human immunity at the population level, and discovery of novel therapeutic candidates. Post-selection antibodies collected from humans are especially valuable therapeutically, since their safety in humans has already been demonstrated in an individual. Indeed, the post-selection antibodies that comprise convalescent serum are routinely transferred between patients; and the immunotherapy with the fastest time from discovery to FDA approval time was an anti-spike monoclonal antibody found in a convalescent patient during the SARS-Cov2 pandemic (bebtelovimab). Previous synthetic “fully human” antibody libraries have used human Ig genes but random CDR3s, e.g. from a degenerate codon library. Given that the CDR3 plays a central role in determining antibody binding, such libraries are unlikely to fully inherit the safety, low immunogenicity, and functionality of real post-selection human antibodies. This example illustrates the design and build of a large scale library of antibody CDR3 sequences that closely matched those observed in human post-selection antibody repertoires. FIG.3A provides an overview of the experimental design. First, a variational synthesis model was trained on a large and diverse dataset of post-selection human antibodies. Then the model was used to generate CDRH3 samples of diverse lengths, in silico and in vitro. The model’s in silico performance was validated by checking how well the distribution of samples matched held-out data. The in vitro performance was validated by sequencing a subsample of manufactured designs and similarly evaluating them against held-out data. Results of this experiment indicate that the generative model performs well in silico, accurately predicting held-out human sequences and generating realistic novel sequences. Attorney Docket No. 063640-514001WO Additionally, the generative model performs well in vitro, manufacturing a petascale library that closely resembles real, unseen human antibody CDRH3s – as much as do in silico designs from an in silico-only protein language model. FIGs.3A-3I illustrate an exemplary process for synthesizing a pan-human repertoire. FIG.3A provides an overview of the process. Data including antibody CDRH3 sequences collected from eleven thousand human repertoires was obtained. A generative, manufacturing-aware model utilizing variational synthesis was trained on this obtained data to generate samples from the model in silico and in vitro. Sample quality was also evaluated. Antibody sequences spanning the full range of human post-selection antibody diversity were designed and built. To accomplish this, a generative model using variational synthesis was trained on a large and diverse dataset of post-selection human antibodies. The variational synthesis model was trained on the human CDRH3 sequences from the Observed Antibody Space (OAS) database. All human repertoires available in the database on 2024-06-07 were collected: 13,265 repertoires total 756 donors. Then, all repertoires labeled with diseases that may lead to production of potentially harmful antibodies were excluded: the full list in the OAS nomenclature was ‘Healthy / celiac-disease’, ‘Allergy / NoSIT’, ‘Allergy / SIT’, ‘MS’, ‘MuSK- MG’, ‘AChR-MG’, ‘SLE’, ‘Light-Chain-Amyloidosis’, ‘Asthma’, ‘Allergic-Rhinitis-Out-Of- Season’, ‘Allergic-Rhinitis-In-Season’.11,271 repertoires from individuals aged 3 to 92 (42 ± 19) that contained 327,341,295 unique productive CDR3 amino acid sequences (CDR3 was determined as in OAS, i.e., positions 105 to 117 in the IMGT31 numbering scheme) were used. As illustrated in FIG.7, the distribution of lengths of these sequences were evaluated and extreme cases such as those sequences with length below 4 or above 40 were excluded, which resulted in 325,596,608 unique sequences. This set of sequences were uniformly split into 90% training and 10% heldout data. FIG.3B illustrates a low-dimensional representation of sequences from the held-out data. In connection with FIGs.3A and 3B, a training dataset of 325 million unique human heavy-chain antibody sequences was compiled from a diverse set of 11,271 repertoires in the OAS database. Patients with autoimmune disorders and related conditions were excluded to ensure the training data consists of functional and non-autoreactive sequences. Trained on heavy-chain CDR3 sequences, the variational synthesis model estimates the distribution underlying all healthy human antibody CDRH3 sequences, covering both seen and unseen sequences. The model achieves an average per-residue perplexity on a held-out test set of 10.6. Attorney Docket No. 063640-514001WO Training set perplexity is similar to the measured test set perplexity of 10.6, indicating strong generalization from seen to unseen antibodies. FIG.3C illustrates low-dimensional representation of sequences from the held-out data, together with sequences generated in silico by the variational synthesis model. The in silico sample quality of the variational synthesis model, i.e. its performance in the absence of any manufacturing errors was evaluated by comparing summary statistics of the generated samples to held-out data. As shown in FIG.3F the variational synthesis model accurately captures the average amino acid usage across positions. Additionally, as shown in FIGs.3B and 3C, low-dimensional representations suggest qualitatively that samples from the variational synthesis model closely follow the distribution of the held-out data. FIG.3D illustrates low-dimensional representation of sequences from the held-out data, together with sequences manufactured in vitro by variational synthesis. FIG.3E illustrates low-dimensional representation of sequences from the held-out data, together with sequences manufactured in vitro using degenerate codon synthesis (NNK codons). FIG.3F illustrates per-residue discrepancy between generated sequences (in silico or in vitro) and the held-out data. The discrepancy is based on total variation distance. To more stringently evaluate sample quality, a nonparametric two-sample test evaluations based on the Bayesian embedded autoregressive (BEAR) model and the maximum mean discrepancy (MMD) with a biological sequence kernel was used. These methods rigorously checked the entire distribution of model samples for mismatch with the held-out data distribution, rather than just checking a finite number of metrics. They thus help avoid the possibility that the generative model performs well according to hand-picked scores, but poorly along other dimensions. The two evaluations are complementary, with BEAR more focused on short-range correlations (it relies on variable-length k-mer statistics) and MMD focused on more global structure (it relies on padded Hamming distances between sequences). Performance of the variational synthesis model to an antibody language model, IgLM. IgLM was compared. The antibody language model was also trained on a very similar dataset of human antibodies, also based on OAS. While the antibody language model produces realis- tic human antibody sequences in silico, it lacks a manufacturing-aware architecture, so in vitro samples must be synthesized individually. FIG.3G illustrates per-sequence Bayes factor of the BEAR two-sample test, comparing generated sequences (in silico: i.s. or in vitro: i.v.) and the held-out data. The dashed line indicates a Bayes factor of one; and values lower than one favor the hypothesis that the model and data distributions are identical. FIG.3H illustrates estimated maximum mean discrepancy (MMD) between generated sequences and the held-out data. An Attorney Docket No. 063640-514001WO MMD of zero indicates a perfect distribution match. As shown in FIGs.3G and 3H, in silico samples from the variational synthesis model achieve a good match to held-out data, with sample quality similar to IgLM. In detail, the BEAR test provides a Bayes factor describing the odds in favor of the model and data distributions being different. The per-sequence geometric average Bayes factor is reported, which illustrates the evidence contributed by each sequence (in units of odds). Values above one indicate the model samples are distinguishable from the data. Applied to the data, the BEAR test says that both generative models produce an average Bayes factor near one: 1.31 for the variational synthesis model and 1.02 for IgLM (see FIG. 3G). The MMD quantifies the mismatch between model samples and held-out data in terms of the worst-case difference in their average phenotype. The MMD confirms that variational synthesis model samples are realistic, and indeed says the variational synthesis model outperforms IgLM (See FIG.3H). Overall, these evaluations suggest that variational synthesis models are powerful generative models that can achieve comparable in silico performance to deep generative models. FIG.3I illustrates the fraction of generated sequences that are considered clearly realistic based on a high MMD witness function score. Dashed line indicates the percentile threshold set to define clearly realistic based on the held-out data. The fraction of in silico model samples that are highly plausible human antibodies was quantified using the witness function of the MMD, which scores how much a sequence looks like real data as opposed to a sample from the model. Low scores suggest a sequence is an outlier and might not have the functional properties of a real human antibody. To be conservative, a model-generated sequence was considered to be clearly realistic only if its score is in the top 67% of witness function scores for real, held-out human antibody sequences. By this metric, 28% of samples from the variational synthesis model are clearly realistic, versus 49% of samples from IgLM. As discussed above, nucleic acid molecules encoding CDRH3 samples of length 4-40 amino acids were manufactured from the trained variational synthesis model. A yield of roughly 150 nanomoles, 9 × 1016molecules was achieved. A random subset of these samples was sequenced: six million after paired-end read assembly and filtering. The manufactured variational synthesis library has only a slight degradation in performance compared to in silico predictions, with the manufactured samples maintaining similar average amino acid usage (as shown in FIG.3F) and similar low-dimensional representations to real held-out data (as shown in FIG.3D). As an additional comparison, as shown in FIG.3E and F, a degenerate codon library was constructed based on NNK codons, where N is randomly A, C, G or T and K is randomly Attorney Docket No. 063640-514001WO G or T. NNK libraries can achieve similar sequence diversity to variational synthesis libraries, as they also rely on stochastic synthesis reactions, but they are uniformly random with no relationship to a trained generative model. The results indicate that sequences in the manufactured NNK library show dramatically different amino acid usage and low-dimensional representations compared to samples from the variational synthesis model (see FIGs.3E and 3F). The BEAR and MMD tests are designed to estimate the quality of the complete library based on the sequenced subset. As shown in FIGs.3G and 3H, the manufactured variational synthesis library matches the held-out data only slightly worse than the in silico samples from the variational synthesis model. For example, the BEAR test Bayes factor increases from 1.31 to 1.39. Overall, the quality of the in vitro, manufactured samples is not only substantially better than that of NNK libraries, but is even comparable to the quality of in silico samples from IgLM: in vitro samples from the variational synthesis model have lower MMD than in silico samples drawn from IgLM. As shown in FIG.3I, according to the witness function, 39% of the manufactured sequences are clearly realistic antibody sequences, compared to IgLM’s in silico fraction of 49%. By contrast, in the NNK library, only 0.08% of sequences are clearly realistic. In summary, the experimental data indicates that variational synthesis models are able to create tens of quadrillions of samples from the pan-human antibody repertoire distribution, with in vitro sample quality similar to that of an in silico-only, deep generative antibody model. EXAMPLE 3. Synthesizing Nucleic Acid Library Encoding the HLA-A*02:01 Epitope Repertoire Peptide epitopes presented on HLA molecules are recognized by human T cells, leading to an adaptive immune response to pathogens and tumors. These linear epitopes form the basis for T cell vaccines against infectious disease and cancer. Large scale synthesis of epitopes allows for experimental determination of the antigens that human T cells respond to, and discovery of novel vaccine candidates. FIGs.4A-4J correlate to experiments illustrating generation of a library of peptide epitopes. FIGs.4A-4J illustrate an example where samples were synthesized from a generative model of antigens presented on HLA-A*02:01. FIG.4A provides an overview of the workflow. As shown, the training dataset consists of diverse antigens (human and non-human) predicted to bind HLA-A*02:01. To comprehensively explore the space of human T cell antigens and synthesize the Attorney Docket No. 063640-514001WO full range of peptides that bind a common HLA allele, HLA-A*02:01, a generative model utilizing variational synthesis was trained on the results of a virtual screen for HLA-A*02:01 binders. Approximately ∼1016samples from the results of the variational synthesis model were manufactured and the resulting library was validated by sequencing. A large dataset of natural peptides predicted to bind HLA-A*02:01, covering both human (neo)antigens and pathogen antigens was obtained. In silico peptides were sampled according to the natural amino acid frequencies and length range, 8-12 amino acids. These peptides were screened for binding against HLA-A*02:01, using NetMHCpan, a state-of-the-art machine learning model that returns a prediction of binding strength for a given HLA-epitope pair. The resulting dataset consists of 2 million samples from the conditional distribution of natural peptides predicted to bind HLA-A*02:01. A generative model utilizing variational synthesis was trained on 90% of this dataset, achieving a perplexity of 14.2 per amino acid on the held-out 10% of the data, versus 20.6 for NNK libraries. Since the data is generated synthetically, the true data perplexity is computed as being 11.4, modestly lower than that achieved by the synthesis model. Samples of epitopes were synthesized with lengths 8-12 amino acids, according to the parameters of the variational synthesis model. This resulted in a yield of roughly 75 nanomoles, corresponding to 4 × 1016molecules. The samples were verified by sequencing a subset (11 million reads after assembly and filtering) and applying summary statistic and nonparametric two-sample evaluations. FIG.4B provides a low-dimensional representation of sequences from the held-out data. FIG.4C provides a low-dimensional representation of sequences from the held-out data, together with sequences generated in silico by the variational synthesis model. FIG.4D provides a low-dimensional representation of sequences from the held-out data, together with sequences manufactured in vitro by variational synthesis. FIG.4E provides a low-dimensional representation of sequences from the held-out data set, together with sequences generated in silico with NNK codons. FIG.4F illustrates per-residue discrepancy between generated sequences (in silico or in vitro) and the held-out data. The discrepancy is based on total variation distance. FIG.4G illustrates the per-sequence Bayes factor of the BEAR two-sample test, comparing generated sequences (in silico or in vitro) and the held-out data. FIG.4H illustrates the maximum mean discrepancy (MMD) between generated sequences and the held- out data. FIG.4I illustrates the fraction of generated sequences that are considered clearly realistic based on a high MMD witness function score. FIG.4J illustrates the fraction of sequences classified as non-, weak, or strong binders by NetMHCpan-4.1. A close correspondence between the synthesized DNA and held-out data is Attorney Docket No. 063640-514001WO illustrated, in average amino acid usage and in low-dimensional representations as shown in FIG.4D and 4F. In particular, as shown in FIGs.4E and 4F, the manufactured variational synthesis libraries show dramatic improvements over NNK libraries, based on in silico sampled NNK sequences. Nonparametric two-sample test evaluations also show close correspondence between synthesized DNA and held-out data, again with large improvements over NNK libraries, as is shown in FIGs.4G and 4H. Indeed, the BEAR test suggests the synthesized sequences are statistically indistinguishable from held-out data: the Bayes factor is below one, favoring the hypothesis that the model and data distributions are the same. According to the witness function score, 31% of the sequences manufactured by variational synthesis are clearly realistic, versus 4.6% of NNK library sequences, which is shown in FIG. 4I. The quality of manufactured samples can be evaluated using the binding predictor. As illustrated in FIG.4I, 36% of the sequences manufactured in vitro by variational synthesis are predicted to strongly bind HLA-A*02:01 by NetMHCpan, with an additional 25% predicted to bind weakly. By contrast, 0.7% of NNK library sequences are predicted to bind strongly, and 2% weakly. Overall, then, variational synthesis enables accurate manufacturing of about 1 × 1016samples from the global distribution of T cell epitopes predicted to present strongly on HLA- A*02:01, in a library where half of all sequences are predicted to be either a strong or weak binder. As discussed above, the screen selected sequences that were predicted to be strong binders by NetMHCpan, version 4.1. This data can be formally described as samples from a conditional generative model of peptides. Let the function b(·) : X → {0, 1} denote the NetMHCpan model for HLA-A*02:01, which took in an amino acid sequence x and output a prediction for whether it was a strong binder for HLA-A*02:01 (b(x) = 1) or not (b(x) = 0). A global natural distribution over possible antigens q(x) was considered. In particular, let v be length 20 vector (on the simplex) denoting the fraction of each amino acid across proteins. The amino acid frequencies for the human proteome were used, but similar frequencies hold acrossevolution. Define q(x) = ∏^^ ^^=1 ^^^^^^as the distribution over sequences where L amino acids were each sampled independently from v. L was chosen to be uniform over the most frequent lengths for class I HLAs (8-12 amino acids). Note although q(x) was a simple generative model, which ignored correlation structure among nearby amino acids, it was observed in practice that it provided a reasonable approximation of plausible antigen distributions: there Attorney Docket No. 063640-514001WO were a vast array of possible antigens (including human, bacterial, and viral proteins), and at short length scales there was little common structure among their sequences. The goal was to synthesize antigens presented on HLA-A*02:01. Therefore, the aim was to synthesize samples from the conditional generative model describing the distribution of antigens predicted to bind HLA-A*02:01, Samples from this distribution can be drawn via rejection sampling: draw a sample X ∼ q(x) and accept it if b(X) = 1, otherwise reject and repeat. This rejection sampling procedure was, precisely, a virtual screen of samples from q(x) using NetMHCpan. Additionally, thenormalizing constant ∑^^∈^^ ^^(^^)^^(^^) = Eq[b(X)] can be estimated as the fraction of samplesthat were accepted. This allowed an estimation of the likelihood q(x|b = 1) of the true distribution, and compute its perplexity, etc. Finally, the generative model was trained on samples from q(x|b = 1) and ran the corresponding synthesis procedure. An estimate pθ*(x) of the distribution over antigens likely to be presented on HLA-A*02:01 was constructed and then synthesized samples from pθ*(x). Table 2. Percentage of sequences in each library that are predicted to bind HLA-A*02:01 by NetMHCpan Overall, then, generative model utilizing variational synthesis enabled accurate manufacturing of about 1 × 1016samples from the global distribution of T cell epitopes predicted to present strongly by HLA-A*02, in a library where one in every two samples was estimated to be either a strong or weak binder. EXAMPLE 4. Synthesizing Nucleic Acid Library Encoding DNA Polymerases Across Evolution DNA polymerases play an essential role in modern biotechnology, with major Attorney Docket No. 063640-514001WO applications including polymerase chain reaction (PCR) and next generation sequencing (sequencing by synthesis). Large scale synthesis of polymerase variants enables screening for enzymes with greater fidelity, higher thermostability, non-natural substrates, and other desirable properties. DNA polymerases are ancient and found across life, so novel DNA polymerases spanning the range of evolutionary possibility were synthesized herein. FIGs.5A-5I present experimental data from synthesizing samples from a generative model of DNA polymerases. FIG.5A provides an overview of the workflow. The data consists of completed Taq polymerase sequences generated by an evolutionary protein language model. A variational synthesis model is trained on this data, samples from the model are generated in silico and in vitro, and evaluated for sample quality. FIG.5B provides low-dimensional representation of sequences from the held-out data. FIG.5C provides low-dimensional representation of sequences from the held-out data, together with sequences generated in silico by the variational synthesis model. FIG.5D provides low-dimensional representation of sequences from the held-out data, together with sequences manufactured in vitro by variational synthesis. FIG.5E provides low-dimensional representation of sequences from the held-out data set, together with sequences generated in silico with NNK codons. FIG.5F provides per- residue discrepancy between generated sequences (in silico or in vitro) and the held-out data. The discrepancy is based on total variation distance. FIG.5G provides per-sequence Bayes factor of the BEAR two-sample test, comparing generated sequences (in silico or in vitro) and held-out data. FIG.5H provides maximum mean discrepancy (MMD) between generated sequences and the held-out data. FIG.5I provides an illustration of the fraction of generated sequences that are considered clearly realistic based on a high MMD witness function score. As shown in FIG.5A, a generative model that is manufacturing-aware and utilizes variational synthesis was trained on a dataset of DNA polymerase sequences predicted to be generated by evolution. The model was used to generate and manufactured ∼1016samples from the generative model. The resulting library comprising nucleic acids encoding the DNA polymerase sequences was validated by sequencing, and confirmed that the distribution of manufactured sequences corresponds closely to heldout data. A generative model was trained on samples from a generative evolutionary model, conditioned on an initial segment of the Taq polymerase. In particular, let q(x) denote the distribution of Progen2 (progen2-xlarge), a protein language model trained on a large and diverse dataset (including UniRef and the BFD metagenomic dataset). ProGen2 has been trained on protein data drawn from across evolution, and provides a global estimate of sequences’ evolutionary probability. These estimates correlate with various laboratory Attorney Docket No. 063640-514001WO measures of protein function and fitness. The model was prompted or conditioned with an initial segment of the Taq polymerase, a thermostable DNA polymerase widely used in PCR. Specifically, the model was conditioned on all but the last 50 amino acids of Taq, and then generated new sequences with lengths up to 60 amino acids for the remainder of the enzyme. Each generated design represents a predicted evolutionary possibility for a complete protein. 250,000 samples were drawn from ProGen2, and used as training data for a generative model utilizing variational synthesis. Trained on 90% of the data, the variational synthesis model achieves a perplexity of 3.54 per amino acid on the held-out 10% of the data, versus 20.44 for NNK libraries. Since the data is generated from a model with tractable likelihoods, it is estimated that the true perplexity as 2.57, modestly lower than the variational synthesis model. Samples of the final region of the Taq polymerase were synthesized, with length 40- 60, according to the parameters of the trained variational synthesis model. A yield of roughly 35 nanomoles, corresponding to 2 × 1016molecules was attained. A random subset was sequenced: 500,000 reads after assembly and filtering. The distribution of synthesized DNA and held-out data match closely, according to summary statistics and nonparametric two- sample evaluations, and is shown in FIGs.5D-5H. Large improvements over NNK libraries are observed. As illustrated in FIG.5I, according to the witness function analysis, 2.8% of the sequences manufactured by variational synthesis are clearly realistic, while we could not find any clearly realistic sequences among 200,000 sampled from the NNK library (<0.0005% clearly realistic). The example illustrates that variational synthesis models can be trained on data generated by a conventional, non-manufacturing-aware generative model, and the resulting quadrillion manufactured samples closely follow that model’s designs. EXAMPLE 5. CAR Candidates Against Intracellular Targets FIGs.6A-6D illustrate applications of a generative model using variational synthesis to generate one or more CAR candidates, and screening yields for CAR candidates against intracellular targets. HLA-presented intracellular antigens represent a promising class of cancer immunotherapy targets, which expand the universe of actionable targets beyond surface proteins. These targets leverage the natural process of intracellular protein degradation and peptide presentation by HLA molecules, which allows the immune system to surveil and respond to aberrant intracellular activities, such as those driven by mutations or dysregulated protein expression in cancer or infected cells. Development of immunotherapies that can Attorney Docket No. 063640-514001WO specifically interact with HLA-presented targets have been limited by low discovery rates from traditional approaches such as phage display and hybridoma. Newly proposed de novo generative machine learning methods have increased hit rates, but since these proteins are de novo rather than human, they are missing the key developability features of post-selection human antibodies captured in our variational synthesis model. FIG.6A provides an illustration of assembled second generation CAR construct with CDRH3 region (VS) generated from the output of a generative model utilizing variational synthesis and the resulting library. To demonstrate the effectiveness of the manufactured post- selection human CDRH3 variational synthesis library at rapidly discovering new therapeutic candidates against hard targets, a subset of the manufactured library was assembled into functional second generation scFv CAR constructs and expressed on human cells. These candidates were screened against a series of HLA-A*02:01 restricted dextramers, and achieved specific hits against a number of important targets. FIG.6B provides an overview of the multiplexed screening assay. The CAR library is screened against barcoded dextramers and hits are single cell sequenced, resulting in the raw data shown on the right. ep.: epitope. H-2Db WT-1 denotes a mouse HLA-presented epitope used as an additional control for off-target activity. In detail, the assembled scFv CAR constructs were expressed on the surface of HEK cells.54% of the library showed protein expression.22.5 million expressing cells were screened against a panel of fluorescently labeled and DNA barcoded dextramers. Each dextramer contained a different HLA-A*02:01 presented intracellular tumor antigen epitope. The fluorescently labeled cells were sorted and single cell of the designed CDRH3 were sequenced together with the bound dextramer barcodes. This multiplexed single cell strategy allows simultaneous readout of synthesized sequences and their binding activity against multiple targets presented on HLA-A*02:01, enabling evaluation of specificity as well as on- target activity. FIG.6C illustrates the fraction of clearly realistic sequences based on the witness function score with specific hits from the screen included (hits). FIG.6D illustrates the distribution of witness function score for held-out post-selection human antibody data (blue) versus specific hits from the screen (hits). Table 3 below shows the number of specific hits discovered against each target HLA-A*02:01 presented epitope (ep.). The screen yielded 10s of specific hits against each of six key tumor antigens presented on HLA-A*02:01. As shown in FIGs.6C-6D, the discovered hits closely resemble held-out human antibody sequences, with 67% judged clearly realistic by the witness function score. Moreover, 5% of hits were an Attorney Docket No. 063640-514001WO exact match to a post-selection human antibody sequence in the OAS database, and a further 37% were 1-2 amino acids away. Table 3: Number of Specific Hits Discovered Against Each Target HLA-A*02:01 Presented Epitope In summary, this example illustrates that a generative model utilizing a manufacturing-aware architecture and variational synthesis can generate libraries that enable the discovery of functional scFv CARs that specifically bind challenging intracellular tumor antigens. These hits are samples from a generative model of human antibodies, and hence closely resemble real post-selection human antibodies, making them promising candidates for therapeutic development. EXAMPLE 6. Sequencing Protocol and Screening Additional methods and techniques for sequencing protocols and screening used in connection with the examples provided herein are described. Sequencing protocols 250ng of DNA were sampled from each library. The sequencing library was prepared using xGenTMssDNA & Low-Input DNA Library Preparation Kit (IDT) following manufacturer’s instructions with the exception that post-extension cleanup was performed using MinElute PCR Purification Kit (Qia-gen). All other cleanups were performed using 1.8x SPRI bead ratio. Dual indexed library samples were quantified using Kapa Library Quantification Kit (Roche). Pooled library samples were sequenced in a 2x150bp configuration on a NextSeq 2000 using P1 Reagents, 300 cycle Kit (Illumina). The resulting paired reads were basecalled Attorney Docket No. 063640-514001WO with bcl2fastq v2.20 and assembled with PEAR. 5,706,832 sequences were obtained for the antibody variational synthesis library illustrated in the example below, 10,844,511 sequences for the epitope library illustrated in the example above, and 510,864 for the DNA polymerase library illustrated in the example above. Screening protocols The synthesized CDRH3 library was assembled into the antibody (scFv) expressing backbone through Golden Gate Assembly. Then lentivirus was produced using library assembly products together with the helper plasmids (Aldevron pALD-Rev-K; pALD-VSV-G- K; pALD-GagPol-K) in HEK293T cells (ATCC CRL-3216). The lentivirus was concentrated using Amicon Ultra - 15 Centrifugal Filter (Millipore Sigma UFC910024) into LV-MAX medium (Gibco A3583401). The HEK293 B2M KO suspension cells were generated from HEK293 suspension cells (Gibco A35347) by knocking out B2M gene using SpCas9 (IDT 1081059) and guide RNA (GGCCGAGATGTCTCGCTCCG; SEQ ID NO: 8).120M HEK293 B2M KO suspension cells were transduced with the concentrated lentivirus in the condition having 30% positive frequencies. Three days after transduction, the cells are stained with the PE-labeled Immudex DeCode dextramer pool (MAGE-A1, HLA-A-0201: KVLEYVIKV; SEQ ID NO: 5; PRAME, HLA-A-0201: KMILKMVQL; SEQ ID NO: 7; PRAME, HLA-A-0201: SLLQHLIGL; SEQ ID NO: 4; WT-1, HLA-A-0201: RMFPNAPYL; SEQ ID NO: 1; p53, HLA-A-0201: GLAPPQHLIRV; SEQ ID NO: 2; p53, HLA-A-0201: KLCPVQLWV; SEQ ID NO: 3; PSA P2, HLA-A-0201: KLQCVDLHV; SEQ ID NO: 6) according to the manufacturer’s instruction. The mouse HLA dextramer WT-1 H-2 Db: YMFPNAPYL; SEQ ID NO: 9 was included as an HLA control. One dextramer that failed to manufacture (PSMA, HLA-A-0201: ALFDIESKV; SEQ ID NO: 10) was excluded from the results. Right after the staining, the dextramer reactive (PE+) cells were sorted on BD FACS Aria. Sorted cells were sequenced with single cell RNA sequencing (Chromium Single Cell 5’ Reagent Kits v2 with Feature Barcode technology, 10x Genomics). As CDRH3 amplicons including the 10x barcodes were too long for Illumina sequencing, they were sequenced on a PromethION instead (Oxford Nanopore Technologies). 10x cell barcodes, UMIs, CDRH3 sequences, and dextramer feature barcodes were identified by aligning the flanking sequences with parasail. Reads for which both a valid cell barcode, UMI, and CDRH3 (or cell barcode, UMI, and feature barcode) were present were selected and the rest were discarded. Finally, sequencing errors were corrected with an EM- based approach and UMI counts for each cell-CDRH3 sequence and cell-feature barcode pair were extracted. CDRH3 sequences with less than 2 UMIs were assumed to originate from PCR Attorney Docket No. 063640-514001WO errors and were discarded. A sequence was judged to be a specific hit if it had ≥ 10 on-target counts and < 10 total off-target counts across all other dextramers. EXAMPLE 7. Evaluations of Designed and Synthesized Libraries Additional methods and techniques for the evaluation of outputs generated by a generative model utilizing variational synthesis are described. 1. Average Amino Acid and Nucleotide Usage Per-Position Error The match between real data and manufactured sequences were evaluated in terms of average amino acid usage per position. For each position, the absolute error in amino acid frequency was quantified using an L1 discrepancy: the total variation distance between the marginal distribution over amino acids at each position. To account for length differences among sequences, each sequence on the right was pad with a stop character, $, and this was included in the alphabet. Formally, let Ɓ denote the amino acid alphabet. Let p(xj) denote the marginal likelihood of a sequence distribution p(x) at position j, such that p(xj = b) for b ∈ Ɓ was the probability of amino acid b appearing at position j, and p(xj = $) was the probability of asequence being shorter than length j. Hence, ∑^^∈Ɓ ∪{$} p(^^^^ = ^^) = 1. Now, the discrepancy inaverage amino acid usage at position j, between distributions p and q, is In practice, the per-position discrepancy is estimated from samples by taking theempirical frequency = ^^)where n is the number of samples and II(·) is the indicator function that takes value one when the statement is true and zero otherwise. For FIG.3F, FIG.4F, and FIG.5F, the data distribution was compared to the manufactured distribution generated by the generative model utilizing variational synthesis byevaluating D^^(p,̂ p^̂^∗), and / or comparing sequences from the heldout data to sequencesmanufactured by the generative model. For in silico libraries (such as in silico variational synthesis models or IgLM) 200,000 samples were used. Overall Accuracy To compute an overall measure of accuracy, the discrepancy across positions were averaged. The accuracy was one minus the error, Attorney Docket No. 063640-514001WO This was estimated empirically from sample as The maximum length L was set to the typical length of sequences in each dataset, rather than the maximum length, since the aim was to evaluate the ability to accurately match amino acid or nucleotide usage, rather than length. In practice, the length distributions was unimodal in each dataset, and L was set to the mode: L = 15 for the antibody dataset, L = 9 for the epitope dataset, L = 49 for the DNA polymerase dataset. This was the notion of accuracy reported in Table 1 above. This notion of accuracy appropriately generalized the standard notion of accuracy used when the goal was to synthesize a specific sequence. For example, consider a situation where the goal is to synthesize a sequence ACG, but only 95% of the sequences manufactured have an A at position one, 95% have a C at position two, and 95% have a G at position three. Then, the average accuracy would be 95%. Mathematically, the goal of synthesizing ACG corresponds to a target distribution of p(x) = δACG(x), a delta distribution at ACG. This corresponds to a target marginal distribution of p(x1= A) = 1 and p(x1= b) = 0 for b ≠ A. Meanwhile, the realized marginal distribution at this position was q̂(x1 = A) = 0.95 and∑^^≠^^ p̂ (^^1 = ^^) = 0.05. This gave D1(p,̂ q̂) = 0.05, and, extending the same logic, A(p,̂ q̂) =0.95. So, the accuracy metric reduced to the standard accuracy metric used when the goal was to synthesize a specific sequence. 2. Low Dimensional Representations To qualitatively evaluate the areas of sequence space covered by the target data and the libraries generated herein, two-dimensional sequence representations were computed. To do this (1) a matrix of pairwise Hamming distances was computed (with all sequences padded on the right to the same maximum length) between all sequences from the heldout data, the designed and manufactured generative libraries, and the NNK library, (2) this distance matrix was transformed into Gram matrix using the kernel, and (3) the dimension of this matrix was reduced using UMAP. Then, only the heldout set was shown (FIG.3B; FIG.4B, and FIG. 5B), the heldout set overlapped with the generative library (FIG.3C; FIG.4C, and FIG.5C) and the heldout set overlapped with the NNK library (FIG.3D; FIG.4D, and FIG.5D), to illustrate how the much of the target data distribution was covered by each library. For the sake Attorney Docket No. 063640-514001WO of visual interpretability, a sub-sample of 5,000 sequences from the heldout data and 50,000 from each library were used; more sequences from the libraries were taken since the manufactured libraries were much larger than the heldout data set. 3. Overlap Analysis In some implementations, overlap analysis can be used. To estimate the probability that a sequence manufactured by generative matched a sequence from the training or heldout data set, let Y1:^^denote the real data sequences (either training or heldout), and let X1:^^denote the manufactured generative sequences. Then, the probability that a manufactured sequence matched a real sequence was computed, where t is the mismatch tolerance; t = 0 for exact matches and t = 1 for off-by-one-matches. To compute this quantity efficiency, CompAIRR was used. 4. Nonparametric Tests (BEAR and MMD) In some implementations, distribution overlap analysis can be used. Set Up The basic goal of generative modeling such as in the generative models utilizing variational synthesis described above was to develop an estimate pθ*(x) of the distribution p0(x) underlying the data. In particular, the aim was for samples X ∼ pθ*(x) from the model to be indistinguishable from samples X ∼ p0(x) from the data distribution. In a generative model utilizing variational synthesis, the samples can be actual physical molecules, and a goal is to have the distribution of the complete library p(x) match the data distribution p0(x). A variety of metrics have been developed to evaluate generative biological sequence models. Some methods start with a function of sequences F(·) : X → Rd, which output a scalar or vector summary of the sequence’s features. A simple example is F(x) = xi, which outputs a one-hot encoding of the amino acid at position i. A more complex example is F(x) = pLLDT(x), the prediction confidence of a folding algorithm (ESMFold) applied to x. More broadly, F(x) can be any hand-crafted feature of x or the output of a machine learning model such as a supervised predictor or unsupervised embedding. Once F(x) is chosen, it can be used to compare samples from the generative model to (heldout) samples from the data, e.g. by examining the error Attorney Docket No. 063640-514001WO where X1:n∼ pθ*(x) are samples from the model and Y1:m∼ p0(x) are data. In sum, typical generative model evaluations check that samples from the model look like real data according to a chosen metric F(x). Indeed, the analysis of average amino acid usage, length distribution, and low-dimensional representations can each be seen as instantiations of this general idea (note also that popular metrics for generative image models, such as the Frèchet Inception Distance and the Inception Score, can also be framed in a similar way). These approaches were limited, however, because F(x) can only capture certain aspects of the data. Even if the model samples look similar to the data according to chosen metrics, they may differ dramatically along other dimensions. Adding additional metrics, i.e., alternative values of F(x), can reduce this problem but does not eliminate it. Moreover, since some generative models were optimized for certain metrics and not others, performance can vary widely among different models, and results were incomparable. To overcome these limitations and pursue a more unbiased evaluation, nonparametric two-sample tests were used. Rather than focus on specific features of sequences, these methods evaluated how well the entire model distribution pθ*(x) matched the data distribution p0(x). To accomplish this, they implicitly consider all possible summary functions F(x), making use of an infinite number of features. Hence, they resolve the problem of relying on a finite set of hand-picked metrics F(x). Nonparametric tests provided more global judgments of generative model quality than hand-picked features, but different nonparametric tests can still place different emphasis on different features of the data. This can result in different conclusions in practice, when the tests were applied to finite datasets. To mitigate this issue, two different nonparametric two- sample tests were utilized, both designed for biological sequence data. The BEAR test was based on an autoregressive model, while the MMD used a kernel involving the Hamming distance between sequences. As a result, the BEAR test placed somewhat greater focus on local correlations between nearby amino acids, while the MMD considered more global structure. For each set of sequences, a subsample of 200,000 sequences were chosen – a sufficiently large number for which computing pairwise distance matrix was feasible. Each evaluation was repeated across ten independent subsets of 200,000 sequences, drawn uniformly without replacement from each set of sequences, and reported the mean and standard error as error bars. To validate variational synthesis in vitro, a subset of the library was sequenced. Standard high throughput sequencing measures a random subset of sequences from the library. Attorney Docket No. 063640-514001WO This implies that this model this sequencing data as i.i.d. samples from an underlying distribution, X1:n∼ p(x), where p(x) is the distribution of the full library. BEAR and MMD, as well as the accuracy metric are designed to make inferences about the error in p(x) based on samples. So, from the sequenced subset of the library, this metric can be used to estimate the error in the full ∼ 1016library. Maximum Mean Discrepancy (MMD) First, maximum mean discrepancy was used to evaluate the mismatch between model and data. It is represented as where Hkis a reproducing kernel Hilbert space with kernel k(x, y). The space Hkcan contain an infinite number of functions, allowing the MMD to check all possible directions in which themodel and data distributions may differ. Each ^^ ∈ ^^^^ can be thought of as a possible sequence-to-phenotype mapping, so the MMD can be interpreted as the worst-case difference in average phenotype between the two distributions. The MMD can be estimated using samples from the model samples X1:n∼ pθ*(x) and heldout data Y1:m ∼ p0(x) as This value was a measure of model-data mismatch in FIG.2G. In practice, the heavy-tailed Hamming kernel was used where β = 0.5, C = 1, and dH is the Hamming distance, i.e., the number of mismatches between x and y. β < 1 was chosen to allow for a lower penalty between distant sequences, avoiding “diagonal dominance.” The sequences were padded to the same length on the right with stop symbols, so that when x and y have different length, the difference in lengths was added to the Hamming distance. Using this kernel, the MMD was guaranteed to be greater than zero if and only if pθ*≠ p0. Formally, this was guaranteed by the fact that the kernel has discrete masses, and hence was characteristic. Bayesian Embedded AutoRegressive (BEAR) The BEAR two-sample test was applied. The BEAR two-sample test is a nonparametric Bayesian test, which computed a Bayes factor comparing the hypothesis that the model did not match the data to the hypothesis that it did. That is, it compared the hypothesis Attorney Docket No. 063640-514001WO pθ*≠ p0to pθ*≠ p0. Given model samples X1:n∼ pθ*(x) and heldout data Y1:m∼ p0(x), the Bayes factor was . Here, π(p) was a prior distribution that covered all possible sequence distributions, namely the BEAR model distribution. The denominator computed the likelihood of the data assuming X1:n and Y1:n came from the same distribution, integrating over all possible distributions. The numerator computed the likelihood of the data assuming X1:n and Y1:n came from different distributions, again integrating over all possible distributions. The ratio was the amount of evidence in favor of the hypothesis that pθ*≠ p0. In detail, the hyperparameters of the BEAR test were set as follows. First, we set the maximum considered kmer length to the maximum sequence length in the dataset. Second, for convenience, we used a uniform autoregressive base model and a Jeffreys prior on the transition probability, corresponding to f(θ) = 1 and h = 2 in the notation of Amin et al., bioRxiv 2021.05.30.446360 (2021). To make the Bayes factor more interpretable and comparable across models and 1datasets, the geometric mean Bayes factor per sequence ^^^^(^^^^∗ , ^^0)^^+^^in FIGs.3G, 4G, and 5G. In practice n=m, so model and data contribute equally. Intuitively, this can be interpreted as the evidence provided by each sequence that the model and data distributions were distinct. 5. Witness function To estimate the fraction of in silico or in vitro designs that are realistic, a witness function of the MMD was used. The witness function is the function with the maximum difference in expected value under the model and data distributions, i.e. the function that attains the supremum in Equation (1). It is given by It can be estimated empirically from model samples X1:n ∼ pθ⋆(x) and held-out data Y1:m ∼ p0(x) as, The witness function measures how much a sequence x looks like it came from the data distribution p0(x) as compared to the model pθ*(x). Sequences that look like outliers with Attorney Docket No. 063640-514001WO respect to the data will receive low scores according to the witness function. Sequences that are heavily over-represented in the model samples, as compared to the data, will also receive low witness function scores. Note that as a result, the witness function takes into account some notion of diversity as well as realism, and will not overly favor models that just produce a limited set of points that are inliers for the data distribution. This distinguishes the witness function from other methods for outlier detection, and makes it particularly appropriate for evaluating generative models. The witness function is used to evaluate the realism of our generated samples. Specifically, let ^̂^α denote the α-quantile of the empirical distribution of witness function scores, ^^(Y1),..., ^^(Ym). Then, the fraction of “clearly realistic” sequences is computed as the fraction of generated sequences above this threshold, . If the model closely matched the data distribution, we would expect that the fraction of sequences above ^̂^α would be 1 − α on expectation. In practice it is lower, due to model-data mismatch. A threshold of α = 0.33 was used in the examples provided above, but additional values τ̂αfor all values of α are envisioned. 6. NNK Libraries To benchmark variational synthesis against conventional high-diversity library synthesis techniques, results from examples above were compared to NNK libraries. In an NNK library each codon is generated randomly, such that the first two positions (N) can be any base while the final position (K) is a G or T. To help evaluate NNK libraries, the distribution of sequences it produces formally can be specified using a synthesis model. Typically NNK libraries are used with one or a few different lengths. Let ℒ ⊆ {1, 2,...} be the set of lengths, a vector describing their relative concentrations. Then the generative model can be expressed as follows: Attorney Docket No. 063640-514001WO Equation (2) describes a distribution over DNA, with each column of the matrix giving the probability of A, C, G and T respectively. Each nucleotide is sampled independently according to these probabilities. Equation (3) samples the length of the sequences, choosing from among the lengths in ℒ according to their relative concentrations vL. Equation (4) translates the DNA sequences into protein: the function T(·) maps DNA sequences to amino acid sequences according to the standard genetic code. The final result X is a sample amino acid sequence. Note we can also calculate the likelihood of data under this model. In the above presented examples, the results from the generated variational synthesis libraries were compared to the following NNK libraries, with lengths chosen based on typical values in each training dataset (there is no established procedure for choosing lengths in degenerate codon libraries except to match the typical lengths of real sequences). The relative concentration vLwas compared to the relative frequency of each of the lengths in £ in the training data. 1. Antibodies: An in vitro library with lengths ℒ = {11, 15, 19} was created. For evaluations, a random subset of the library was sequenced : 4,494,385 sequences after assembly and filtering. 2. Epitopes: Samples were drawn in silico from the generative NNK library model above, with ℒ = {9, 10, 11}. 3. DNA Polymerases: Samples were drawn in silico from the generative NNK library model above, with ℒ = {48, 49, 50}. To compute finite perplexities describing how well these models match real sequence data, only data sequences with lengths in ℒ were used. Hence the perplexities reported herein should be taken as a conservative estimate of how well the NNK libraries match the real data; in fact they perform worse than the perplexity would suggest, since they do not match the Attorney Docket No. 063640-514001WO length distribution. 7. Details on Antibodies In this section details of a procedure for synthesizing CDRH3 sequences are provided. Dataset The described synthesis model is trained on the human CDRH3 sequences from the Observed Antibody Space database. All human repertoires available in the database on 2024- 06-07 were collected: 13,265 repertoires total, from 756 donors. Then, all repertoires labeled with diseases that may lead to production of potentially harmful antibodies were excluded: the full list in the OAS nomenclature is ‘HEALTHY / CELIAC-DISEASE’, ‘ALLERGY / NOSIT’, ‘ALLERGY / SIT’, ‘MS’, MUSK-MG’, ‘ACHR-MG’, SLE, ‘LIGHT-CHAIN- AMYLOIDOSIS’, ‘ASTHMA’, ‘ALLERGIC-RHINITIS-OUT-OF-SEASON’, ‘ALLERGIC- RHINITIS-IN-SEASON’. This left 11,271 repertoires from individuals aged 3 to 92 (42 ± 19) that contained 327,341,295 unique productive CDR3 amino acid sequences (CDR3 was determined as in OAS, i.e., positions 105 to 117 in the IMGT numbering scheme). As shown in FIG.7, the distribution of lengths of these sequences was evaluated and extreme cases were excluded. For example, sequences with length below 4 or above 40 were excluded, which resulted in 325,596,608 unique sequences. This set of sequences was split into 90% training and 10% held-out data. FIG.7 shows length distribution in the antibody OAS dataset. Sequences with lengths between the vertical lines were included in the training and held-out data for variational synthesis. Note the y-axis is on a log scale. Antibody Language Model A state-of-the-art generative antibody language model IgLM was used to generate sequences in silico. IgLM was trained on the OAS database as well, although in this case, the authors used full-length sequences from all species and chains, and added chain and species tokens to the model input.200000 full-length heavy chain human antibody sequences were generated using tokens [HEAVY] and [HUMAN], and then extracted CDR3s from the resulting sequences. Note that it is not possible to evaluate likelihoods of arbitrary CDR3 sequences using IgLM since it was trained to generate or infill full-length sequences. Details on HLA-A*02:01 Epitopes The synthesis model was trained on the results of a virtual screen for HLA-A*02:01 binders. The screen selects sequences that are predicted to be strong binders by NetMHCpan, version 4.1. This data can be described as samples from a conditional generative model of peptides. Let the function b(∙) : X → {0, 1} denote the NetMHCpan model for HLA-A*02:01, Attorney Docket No. 063640-514001WO which takes in an amino acid sequence x and outputs a prediction for whether it is a strong binder for HLA-A*02:01 (b(x) = 1) or not (b(x) = 0). A global natural distribution over possible antigens q(x) is considered. In particular, let v be length 20 vector (on the simplex) denoting the frequency of each amino acid across proteins. the amino acid frequencies for the human proteome, but similar frequencies hold across evolution. Then, , is the distribution over sequences where L amino acids are each sampled independently from v. The global distribution over possible antigens is defined as a uniform mixture across the most frequent lengths for class I HLAs (8-12 amino acids): where II(|x| = L) takes value one when x has length L. Note although q(x) is a simple generative model, which ignores correlation structure among nearby amino acids, observed in practice that it provides a reasonable approximation of plausible antigen distributions: there are a vast array of possible antigens (including human, bacterial, and viral proteins), and at short length scales there is little common structure among their sequences. To synthesize antigens presented on HLA-A*02:01, aim to synthesize samples from the conditional generative model describing the distribution of antigens predicted to bind HLA- A*02:01, Samples from this distribution can be drawn via rejection sampling: draw a sample X ∼ q(x) and accept it if b(X) = 1, otherwise reject and repeat. This rejection sampling procedure is, precisely, a virtual screen of samples from q(x) using NetMHCpan. Note the normalizing constant Σx∈X b(x)q(x) = ^^q[b(X)] can be estimated as the fraction of samples that are accepted. This allows estimation of the likelihood q(x|b = 1) of the true distribution, and hence compute its perplexity. Finally, the synthesis model is trained on samples from q(x|b = 1) and the corresponding synthesis procedure is run. G. Details on DNA polymerases The synthesis model can be trained on samples from a generative evolutionary model, Attorney Docket No. 063640-514001WO conditioned on an initial segment of the Taq polymerase. In particular, let q(x) denote the distribution of Progen2 (progen2-xlarge), a protein language model trained on a large and diverse dataset (including UniRef and the BFD metagenomic dataset). The model can be prompted with the first ℓ = 782 amino acids of Taq (Uniprot ID DPO1_THEAQ), removing the last 50 amino acids. The conditional distribution of the model can then be written q(x | x1:782 = taq1:782). This is a prediction of proteins likely to be generated by evolution, given that their initial segment matches Taq. Samples can be drawn from this distribution by generating from the model with the temperature parameter and nucleus sampling probability parameter both set to one. In practice many of the generated samples were quite long, in some cases with no stop codon at all. To ensure oligosynthesis was tractable without assembly, a length cutoff of T = 60 amino acids was set, rejecting samples from the model that go beyond this length before they reach a stop codon. This gives a sample from q(x | x1:782 = taq1:782, |x| − 782 ≤ 60). This can be used as a target distribution, and the synthesis model can be trained on samples from this distribution. Note our analysis of these samples focuses just on the completion, ignoring the constant region. To compute the likelihood and perplexity of sequences under this distribution, Bayes’ rule can be used, where II(·) is the indicator function which takes value 1 if the statement is true and 0 otherwise. The likelihood q(x) can be evaluated for ProGen2 since it is an autoregressive model. The term q(x1:t = taq1:782) can be evaluated similarly. Finally, q(|x| − 782 ≤ 60 | x1:t = taq1:782) can be estimated as the fraction of samples that were accepted during rejection sampling. H. Additional Example Data FIGs.8A – 8C illustrate a comparison of amino-acid usage and length distributions in the in silico and in vitro variational synthesis libraries. In particular, FIGs.8A-8C illustrate the distribution of sequence lengths for the antibody dataset, epitope dataset, and DNA polymerase dataset, respectively. FIGs.9A-9C compares the length distributions in the held-out data versus the in vitro variational synthesis libraries and versus the NNK libraries. FIGs.9A-9C compares held-out data to variational synthesis data for antibodies, epitopes, and DNA polymerases. FIGs.9D-9F compares held-out data to NKK library data for antibodies, epitopes and DNA polymerases, where FIG.9D is for an in vitro NKK library, and FIG.9E and FIG.9F are for an in silico Attorney Docket No. 063640-514001WO NKK library. FIGs.10A-10E illustrate witness function analysis of antibodies. FIG.10A shows the percentage of generated sequences (in silico or in vitro) with witness function values above the real-data percentile threshold given on the x-axis (VS: variational synthesis). Formally, this is a plot of ˆτα as a function of α. FIG.10B shows the distribution of witness function scores ˆ f(x) for in silico samples from IgLM versus held-out data. FIG.10C shows the distribution of witness function scores for variational synthesis in silico samples versus held-out data. FIG. 10D shows the distribution of witness function scores for variational synthesis in vitro manufactured samples versus held-out data. FIG.10E shows the distribution of witness function scores for NNK in vitro manufactured sequences versus held-out data. FIGs.11A and 11B illustrate witness function analysis of epitopes and DNA polymerases. Each plot shows the percentage of generated sequences (in silico or in vitro) with witness function values above the real data percentile threshold given on the x-axis. Formally, these are plots of ˆτα as a function of alpha. FIG.11A shows results for epitope models and libraries. FIG.11B shows results for DNA polymerase models and libraries. While the disclosure has been particularly shown and described with reference to specific embodiments (some of which are preferred embodiments), it should be understood by those having skill in the art that various changes in form and detail can be made therein without departing from the spirit and scope of the present disclosure as disclosed herein.
Claims
Attorney Docket No. 063640-514001WO WHAT IS CLAIMED IS:
1. A method of designing a library of nucleic acids or peptides in silico and synthesizing the library in vitro, optionally at petascale, the method comprising: (a) assembling a training dataset comprising nucleotide or amino acid sequences; (b) training a generative model on the training dataset, wherein an architecture of the generative model is a manufacturing-aware architecture suited to an in vitro synthesis platform configured to execute stochastic chemical reactions of synthesizing nucleic acids or peptides; (c) generating sequences of the nucleic acids or peptides in silico; (d) mapping in silico parameters of the generative model onto parameters of the in vitro synthesis platform and generating in vitro synthesis protocols; and (e) designing and synthesizing the library of the nucleic acids or peptides in silico and in vitro, optionally at petascale, based at least in part on learned parameters of the generative model and the in vitro synthesis protocols by controlling the stochastic chemical reactions with the in silico parameters.
2. A method of designing one or more nucleic acid, peptide, or polypeptide sequence(s) in silico for in vitro production, the method comprising: (a) assembling a training dataset comprising nucleotide or amino acid sequences; (b) training a generative model on the training dataset, wherein an architecture of the generative model is a manufacturing-aware architecture configured to execute stochastic chemical reactions of synthesizing the nucleic acids or peptides; optionally wherein the manufacturing-aware architecture is suited to an in vitro synthesis platform; and (c) generating the sequence(s) of the one or more nucleic acids, peptides, or polypeptides in silico; optionally wherein the sequences of the nucleic acids, the peptides, or the polypeptides form a library, which optionally is at petascale.
3. The method of claim 2, further comprising (d) mapping in silico parameters of the generative model onto the parameters of the in vitro synthesis platform and generating in vitro synthesis protocols.
4. The method of claim 3, wherein the one or more nucleic acids, peptides, or polypeptides are synthesized based at least in part on learned parameters of the generative model and the in vitro synthesis protocols by controlling the stochasticAttorney Docket No. 063640-514001WO chemical reactions with the in silico parameters.
5. The method of any one of claim 1-4, wherein the training dataset comprises a validation dataset for tuning the generative model parameters to prevent overfitting.
6. The method of any one of claim 1-5, wherein the training further comprises optimizing the in vitro parameters to match a distribution of the nucleotide or amino acid sequences in the training dataset.
7. The method of any one of claims 1-6, wherein the sequences of the in silico and in vitro synthesized nucleic acids, peptides, or optionally polypeptides are fed back into the generative model for iterative model improvement.
8. The method of any one of claims 1 and 5-7, further comprising (f) validating the library.
9. The method of claim 8, wherein the validating comprises sequencing a sample of the in vitro synthesized nucleic acids or peptides.
10. The method of any one of claims 1-9, wherein a subset of the nucleotide or amino acid sequences is held out as a heldout dataset.
11. The method of claim 10, wherein the generative model’s predictive performance is evaluated on the heldout dataset.
12. The method of claim 10 or claim 11, further comprising assessing distributional match between the samples of the in vitro or in silico synthesized nucleic acids or peptides and the heldout dataset.
13. The method of claim 12, the distributional match is assessed with summary statistic evaluations and nonparametric evaluations.
14. The method of claim 13, wherein the nonparametric evaluations are Bayesian Embedded AutoRegressive (BEAR) and Maximum Mean Discrepancy (MMD) tests.
15. The method of any one of claims 1-14, wherein the generative model generates a large, high quality dataset for training a machine learning model that maximizes distributional overlaps with target data and minimizes generalization errors.
16. The method of any one of claims 1-15, wherein the training dataset comprises amino acid sequences.
17. The method of any one of claims 1-16, wherein the nucleotide or amino acid sequences are generated by another generative machine learning model, a supervised machine learning model, a virtual screen or predictor, or an experimental screen or assay.
18. The method of any one of claims 1-17, wherein the nucleic acid or amino acidAttorney Docket No. 063640-514001WO sequences of the training dataset comprise non-natural sequences.
19. The method of claim 18, wherein the non-natural sequences represent non-canonical nucleic acids or peptides.
20. The method of claim 18 or claim 19, wherein the in vitro synthesis and the in silico design occur concurrently.
21. The method of any one of claims 18-20, wherein the nucleic acid library or the peptide library comprise non-natural nucleic acids or non-natural peptides, which optionally arenon-canonical nucleic acids or peptides.
22. The method of any one of claims 1-21, wherein the in silico designed and in vitro synthesized library comprises nucleic acids, which are DNAs or RNAs.
23. The method of any one of claims 1-22, wherein the in vitro synthesized library is a nucleic acid library.
24. The method of claim 23, wherein the nucleic acid library comprises nucleic acids encoding antibodies, antigenic epitopes, enzymes, neoantigens, or therapeutic proteins.
25. A system for synthesizing a library of nucleic acids or peptides in vitro optionally at petascale, the system comprising: (a) a generative model with a manufacturing-aware architecture trained on a training dataset comprising nucleotide or amino acid sequences; and (b) an in vitro synthesis platform configured to execute stochastic chemical reactions of synthesizing the library of nucleic acids or peptides in vitro optionally at petascale based at least in part on learned parameters of the generative model and in vitro synthesis protocols; wherein in silico parameters of the generative model map onto the parameters of the in vitro synthesis platform.
26. A system for designing one or more nucleic acid, peptide, or polypeptide sequences in silico for production of the one or more nucleic acids, peptides, or polypeptides, the system comprising: a generative model with a manufacturing-aware architecture trained on a training dataset comprising biological sequences; optionally wherein the system is for designing a library of nucleic acids, peptides, or polypeptides, optionally at petascale.
27. The system of claim 26, further comprising an in vitro synthesis platform configured to execute stochastic chemical reactions of synthesizing the library of nucleic acids or peptides in vitro optionally at petascale based at least in part on learned parameters ofAttorney Docket No. 063640-514001WO the generative model and in vitro synthesis protocols.
28. The system of claim 27, wherein in silico parameters of the generative model map onto the parameters of the in vitro synthesis platform.
29. The system of any one of claims 25-28, wherein the training dataset comprises a validation dataset for tuning the generative model parameters to prevent overfitting.
30. The system of any one of claims 25-29, wherein the training of the generative model with the training dataset further comprises optimizing the parameters of the in vitro synthesis platform to match a distribution of the biological sequences in the training dataset.
31. The system of any one of claims 25-30, the sequences of the in silico and in vitro synthesized nucleic acids and peptides are fed back into the generative model for iterative model improvement.
32. The system of any one of claims 25-31, further comprising a validation system.
33. The system of claim 25, wherein the validation system comprises sequencing a sample of the in vitro synthesized nucleic acids or peptides.
34. The system of any one of claims 25-33, wherein a subset of the nucleotide or amino acid sequences is held out as a heldout dataset.
35. The system of claim 34, wherein the generative model’s predictive performance is evaluated on the heldout dataset.
36. The system of any one of claims 32-35, the validation system comprises a system for assessing distributional match between the samples of the in vitro or in silico synthesized nucleic acids or peptides and the heldout dataset or the training dataset.
37. The system of claim 36, wherein the distributional match is assessed with summary statistic evaluations and nonparametric evaluations.
38. The system of claim 37, wherein the nonparametric evaluations are Bayesian Embedded AutoRegressive (BEAR) and Maximum Mean Discrepancy (MMD) tests.
39. The system of any one of claims 25-38, wherein the trained generative model generates a large, high quality dataset for training a machine learning model that maximizes distributional overlaps with a target data and minimizes generalization errors.
40. The system of any one of claims 25-39, wherein the nucleotide or amino acid sequences of the training dataset are nucleotide or peptide sequences.
41. The system of any one of claims 25-40, wherein the nucleotide or amino acid sequences are generated by another generative machine learning model, a supervisedAttorney Docket No. 063640-514001WO machine learning model, a virtual screen or predictor, or an experimental screen or assay.
42. The system of any one of claims 25-41, wherein the nucleotide or amino acid sequences of the training dataset comprise non-natural nucleotide or amino acid sequences.
43. The system of claim 42, wherein the non-natural nucleotide or amino acid sequences represents non-canonical nucleic acids or peptides.
44. The system of any one of claims 25-43, wherein the library is a library of nucleic acids, which optionally are DNAs.
45. The system of claim 44, wherein the nucleic acids encoding antibodies, antigenic epitopes, enzymes, neoantigens, or therapeutic proteins.
46. A method comprising: obtaining a training dataset comprising biological sequences; training a generative model with manufacturing-aware architecture to output one or more biological sequences and one or more instructions for synthesis of the one or more biological sequences; and generating a library of biological sequences based on the training dataset by synthesizing the output biological sequences output by the generative model by executing the output synthesis instructions.
Citation Information
Patent Citations
Systems and methods for artificial intelligence-guided biomolecule design and assessment
US20230022022A1
Computer representations of peptides for efficient design of drug candidates
WO2023200866A1