Method for assembling gene libraries from a pool of oligonucleotides
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- THE BROAD INST INC
- Filing Date
- 2024-07-12
- Publication Date
- 2026-05-20
AI Technical Summary
Current methods for assembling gene libraries are limited by high costs, slow turnaround times, and a lack of control over mutation and assembly, leading to a high percentage of undesired mutations.
The method involves a 'one pot' assembly of genetic libraries from a pool of oligonucleotides, where oligonucleotides attached to beads hybridize with nucleic acids to generate polynucleotides, allowing for directed mutagenesis and high-throughput assembly of variant libraries.
This approach enables the generation of gene libraries with a high percentage of useful mutations, improving the efficiency and cost-effectiveness of the gene library assembly process, while reducing the processing of undesired mutants.
Smart Images

Figure US2024037800_23012025_PF_FP_ABST
Abstract
Description
METHOD FOR ASSEMBLING GENE LIBRARIES FROM A POOL OF OLIGONUCLEOTIDESI. Field of Disclosure
[0001] The disclosure relates to the methods and compositions for assembling gene libraries from samples having a nucleic acid with a region with a sequence of interest.II. Priority and Related Applications
[0002] This application claims priority to and the benefit of U.S. Application No. 63 / 526,744, which was filed on July 14, 2023, the content of which is incorporated herein by reference for all purposes.III. Government Support
[0003] This invention was made with government support under Grant Nos. 2019-19081900002 awarded by the Intelligence Advanced Research Projects Activity, S4599-001 awarded by the Department of Defense, and AI142780 awarded by the National Institutes of Health. The government has certain rights in the invention.IV. Background
[0004] The cornerstone of synthetic biology is the design, build, and test process — an iterative process that requires DNA to be made accessible for rapid and affordable generation and optimization of these custom pathways and organisms. In the design phase, the A, C, T and G nucleotides that constitute DNA are formulated into the various gene sequences that would comprise the locus or the pathway of interest, with each sequence variant representing a specific hypothesis that will be tested. These variant gene sequences represent subsets of sequence space, a concept that originated in evolutionary biology and pertains to the totality of sequences that make up genes, genomes, transcriptome and proteome.
[0005] Many different variants are typically designed for each design-build-test cycle to enable adequate sampling of sequence space and maximize the probability of an optimized design. Though straightforward in concept, process bottlenecks around speed, throughput and quality of conventional synthesis methods dampen the pace at which this cycle advances, extendingdevelopment time. The inability to sufficiently explore sequence space due to the high cost of acutely accurate DNA and the limited throughput of current synthesis technologies remains the rate-limiting step.
[0006] Conventionally, synthesis of different gene variants was accomplished through random mutagenesis, often using large libraries that were subject to broad, random mutations in the hopes of identifying a desired phenotype. This may then require recursively identifying what mutation caused the desired phenotype, which can be a difficult task when the mutation arises from a randomized mutagenesis.
[0007] Polymerase chain reaction (PCR) mutagenesis techniques are widely-used, and represent the established methods for assembling mutagenesis libraries. These methods are favored for their comparatively high-throughput in generating libraries. In such methods, regions of a gene can be PCR amplified incorporating mutations and then cloned into backbones. This can be done for multiple regions. However, PCR mutagenesis is error prone, and aspects of the process can be difficult to direct.
[0008] In an attempt to improve upon mutagenesis, early chemical gene synthesis efforts focused on producing a large number of polynucleotides with overlapping sequence homology. These were then pooled and subjected to multiple rounds of polymerase chain reaction (PCR), enabling concatenation of the overlapping polynucleotides into a full length double stranded gene. A number of factors hinder this method, including time-consuming and labor-intensive construction, requirement of high volumes of phosphorami di tes, an expensive raw material, and production of nanomole amounts of the final product, significantly less than required for downstream steps, and a large number of separate polynucleotides required one 96 well plate to set up the synthesis of one gene.
[0009] Synthesizing of polynucleotides on microarrays provided a significant increase in the throughput of assembling gene libraries. A large number of polynucleotides could be synthesized on the microarray surface, then cleaved off and pooled together. Each polynucleotide destined for a specific gene contains a unique barcode sequence that enabled that specific subpopulation of polynucleotides to be depooled and assembled into the gene of interest. In this phase of the process, each subpool is transferred into one well in a well plate, increasing throughput to the number of available wells in the plate. Typically, this process is conducted using a 96 well plate. While this is two orders of magnitude higher in throughput than the classical method, it still does notadequately support the design, build, test cycles that require thousands of sequences at one time due to a lack of cost efficiency and slow turnaround times.V. Summary of the Disclosure
[0010] The methods of the disclosure provide more efficient and cost-effective methods for assembling gene libraries. In certain aspects, methods of the disclosure provide a method for generation of a “variant” sequence library in a single pot, also referred to as “one pot” assembly of genetic library from a pool of oligonucleotides. Beneficially, the methods of the disclosure provide a high-scale and high-throughput means for assembling gene libraries. Yet, unlike the comparatively high-throughput, random mutagenesis PCR techniques, the presently disclosed methods allow for more control and direction over the mutation and assembly of variant libraries. Consequently, the methods disclosed herein provide a higher percentage of useful or desired mutations than prior methods. This is a tremendous improvement over random PCR mutagenesis, especially when dealing with multiple mutations located at different positions across a gene(s). Although random PCR mutagenesis can produce libraries with upwards of 109unique variants in an assay, upwards of 90% will be undesired mutations. In this way, high-throughput becomes as problematic as it is beneficial — processing the sheer number of unwanted mutants can dramatically increase costs and lower overall efficiency.
[0011] Instead of whole regions being subject to random mutation, and thus production of unwanted mutants (e.g., having mutations at regions not intended to be mutated and / or lacking mutations where desired), the presently disclosed methods are able to direct mutagenesis and assembly where desired, and at a high-throughput, multiplex scale.
[0012] In addition to generation of gene libraries, the methods of the disclosure are beneficial for assembly of genetic circuits (e.g., as disclosed in WO 2015 / 184016, which is herein incorporated by reference), transfection and transformation of genes in the gene library. In certain aspects, the genes in the gene libraries are transferred on beads for clonal amplification in the droplet.
[0013] In certain aspects, the methods of the disclosure comprise providing a reaction mixture in a vessel. The mixture includes a plurality of nucleic acids each with a sequence corresponding to a different consistent section of a sequence of interest, one or more oligonucleotides comprising one or more nucleic acid sequences corresponding to a portion of the sequence of interest and comprising at least one variant section with a different sequence relative to the sequence of interest,wherein the one or more oligonucleotides are each attached to one of a plurality of beads, and reagents for an amplification reaction. The one or more oligonucleotides hybridize with a portion of one or more of the plurality of nucleic acids and generate polynucleotides in the vessel. These polynucleotides are assembled thereby producing variant nucleic acids comprising one of the variant sections (which may be referred to herein as variant domains, variant regions, variant sequences, and the like).
[0014] Preferably, the oligonucleotides are each attached to one of a plurality of beads. In certain aspects, said oligonucleotides are synthesized on the plurality of beads. In certain preferred embodiments, the nucleic acid comprising a sequence of interest comprises a consistent region. In exemplary methods, the one or more oligonucleotides hybridize with a portion of one or more of the plurality of nucleic acids and generate polynucleotide fragments of nucleic acids in said vessel. In certain aspects, a first reaction is performed to assemble nucleic acids comprising a variant sequence from the fragments of the nucleic acids in the vessel.
[0015] In preferred aspects, the methods of the disclosure further include amplifying the assembled nucleic acids. In certain aspects, amplifying the assembled nucleic acids occurs on the plurality of beads. In certain methods, the plurality of beads are in fluidic droplets during amplification of the assembled nucleic acids. In some aspects, the amplification of the assembled nucleic acids is an emulsion polymerase chain reaction (PCR). In certain aspects, the emulsion PCR is generated using membrane emulsification. In some methods, the plurality of beads are in a bulk phase during the amplification of the assembled nucleic acids. In certain aspects, the assembled nucleic acids are circularized before amplification, and wherein said amplification is a whole genome amplification.
[0016] In preferred methods of the disclosure, at least 50% of the assembled nucleic acids comprise a variant sequence in the variant region.
[0017] In certain embodiments, the nucleic acid comprising a sequence of interest comprises one or more variant section. In some aspects, the one or more oligonucleotides have at least partial overlap with sequences corresponding to the conserved region or variant region of the nucleic acid comprising a sequence of interest. The one or more oligonucleotides hybridize with a portion the nucleic acid and generates polynucleotide fragments of nucleic acids in the vessel. In certain aspects, methods of the disclosure further include conducting a first reaction to assemble the fragments of generated nucleic acids in the vessel. In certain embodiments, the nucleic acidfragments generated in the first reaction are assembled by ligation. In certain embodiments, the nucleic acid fragments have a partial overlap with each other. In certain embodiments, the nucleic acid fragments have no overlap with each other.
[0018] Methods of the disclosure may further comprise performing a second reaction to amplify the assembled nucleic acids from the previous step. Beneficially, in certain embodiments of the disclosure, all these reactions are conducted in the same vessel without any steps involving the isolation of any products and / or reagents between different reactions. In certain embodiments, the methods of the disclosure further comprise expressing the proteins from the genetic variants in gene libraries.
[0019] In certain embodiments, the methods of the disclosure generate at least one sequence of interest corresponding to the conserved region of the nucleic acid. In certain embodiments, the methods of the disclosure generate multiple or a plurality of sequences of interest corresponding to the conserved region of the nucleic acid. In certain embodiments, the nucleic acid fragments correspond to a conserved region of the sequence of interest. In certain embodiments, the nucleic acid fragments correspond to a plurality of conserved regions of the sequence of interest. In certain embodiments, the nucleic acid fragments correspond to a variant region of the nucleic acid sequence of interest.
[0020] In certain embodiments, the one or more oligonucleotides are synthesized on a plurality of beads. In certain aspects of the disclosure, the one or more oligonucleotides are synthesized on the plurality of beads by a novel split pool synthesis. Methods of the disclosure include methods using constrained split pool synthesis, in which the available combinations at each step (e.g., codons) are controlled and specified to restrict the design space (the locations mutations are desired), to make library generation as efficient as possible. In certain embodiments, the synthesis is conducted by chemical reactions. In certain embodiments, the synthesis is conducted by enzymatic reactions. In certain embodiments, the enzymatic reactions include a polymerase reaction. In certain embodiments, the enzymatic reactions include an endonuclease reaction. In certain aspects, the enzymatic synthesis is a Type IIS restriction assembly synthesis, e.g., Golden Gate assembly. In certain embodiments, the synthesis is split pool combinatorial synthesis. In certain embodiments of the disclosure, the one or more oligonucleotides that are synthesized on the plurality of beads are cleaved off from the beads. In certain embodiments, the one or more oligonucleotides synthesized comprises, oligonucleotides that are variants of one or more domains of the nucleicacid corresponding to a protein of interest. In certain embodiments, the one or more oligonucleotides hybridize with the conserved region of the nucleic acid. Advantageously, in some aspects, the present methods provide scarless assembly, which can be critical when introducing mutations within genes. In certain embodiments, the one or more oligonucleotides hybridize with the variant region of the nucleic acid. The disclosure beneficially recognizes that using a constrained / directed split pool approach to generate oligonucleotides lead to a high number of genes in the genetic library. This leads to generation of a gene variant library in a single tube from a variety of oligonucleotides synthesized on beads.
[0021] In certain aspects of the disclosure, the sequence of interest corresponds to a variant region of the corresponding peptide or protein. In certain embodiments, the one or more oligonucleotides comprises non-variant (consistent) portions of the nucleic acid. In certain embodiments, the one or more oligonucleotides comprise sequences overlapping or partially corresponding to the variant region of the sequence of interest. In certain embodiments, one or more oligonucleotides have a length of from about 20 to about 500 nucleotides. In certain embodiments, one or more oligonucleotides have a length of from about 20 to about 400 nucleotides. In certain embodiments, one or more oligonucleotides have a length of from about 20 to about 300 nucleotides. In certain embodiments, one or more oligonucleotides have a length of from about 20 to about 200 nucleotides. In certain embodiments, one or more oligonucleotides have a length of from about 20 to about 100 nucleotides. In certain embodiments, one or more oligonucleotides have a length of from about 20 to about 80 nucleotides. In certain embodiments, one or more oligonucleotides have a length of about 60 nucleotides. In certain embodiments, one or more oligonucleotides have a length of about 40 nucleotides.
[0022] In certain aspects of the disclosure, the one or more oligonucleotides have an annealing temperature of from about 50 °C to about 70 °C. In certain embodiments, the one or more oligonucleotides have an annealing temperature of from about 55 °C to about 68 °C. In certain embodiments, the one or more oligonucleotides have an annealing temperature of from about 60 °C to about 65 °C. In certain embodiments, the one or more oligonucleotides have an annealing temperature of about 62 °C ± 1 °C.
[0023] In certain aspects of the disclosure, the fragments of nucleic acids produced in the first reaction have an overlap of a number of nucleotides. The overlapping fragments are subsequently assembled to produce the genes in the genetic library. In certain embodiments, the fragments ofnucleic acid generated in the first reaction have an overlap of at least about 10 nucleotides. In certain embodiments, the fragments of nucleic acid generated in the first reaction have an overlap of at least about 14 nucleotides. In certain embodiments, the fragments of nucleic acid generated in the first reaction have an overlap of at least about 18 nucleotides. In certain embodiments, the fragments of nucleic acid generated in the first reaction have an overlap of at least about 20 nucleotides.
[0024] In certain aspects of the disclosure, the individual nucleic acids generated in the assembled library results in genetic library with genetic variants. In certain aspects, the disclosure provides for isolating these generated nucleic acids / genetic variants in their respective droplets. In certain embodiments, these isolated nucleic acids / genetic variants are further amplified in their respective droplets. In certain embodiments, the isolated nucleic acids / genetic variants are expressed in their respective droplets. In certain embodiments, the respective droplets comprising the isolated nucleic acids / genetic variants further comprise a bead. In certain embodiments, the beads are magnetic beads. In certain embodiments, the isolated nucleic acids / genetic variants with beads are amplified. In certain embodiments, the amplified isolated nucleic acids / genetic variants are transferred on the beads. In certain embodiments, the beads comprising nucleic acids / genetic variants are collected and encapsulated for protein expression. In certain embodiments, the collected beads are pooled together prior to expressing the genes as proteins.
[0025] In certain embodiments, the amplified nucleic acids / genetic variants that are in their respective droplets are pooled together prior to the expression as proteins.
[0026] In certain aspects, the disclosure further provides for whole genome amplification of the genetic variants generated in the gene library.
[0027] In certain aspects of the disclosure, the genes from the library are further used in cell transfection, cell signal transduction, and / or cell transformation. In certain embodiments, the genes in the variant sequence library are cloned in plasmids. In certain embodiments, the plasmids comprising the variant sequence library are transfected and expressed in the cells.
[0028] In certain embodiments, the amplification reaction used in the methods of the disclosure is any amplification reaction used for amplifying nucleic acids. In certain embodiments, the amplification reaction is a polymerase chain reaction. In certain embodiments, the amplification reaction used in the methods of the disclosure is an isothermal amplification reaction.VI. Brief Description of Drawings
[0029] FIG. 1 provides a schematic overview of the methods of the disclosure describing the generation of a variant gene library.
[0030] FIG. 2 provides a schematic of the one-pot assembly of the methods of the disclosure.
[0031] FIG. 3 is a schematic of gene amplification and expression carried out in droplets from the members of the variant gene library generated by methods of the disclosure.
[0032] FIGS. 4A-4B provide an exemplary Anti-CD3 Single-chain fragment variable antibody and the variant library.
[0033] FIG. 5 demonstrates a schematic for estimation of the size of the gene library generated by the methods of the disclosure.
[0034] FIG. 6 provides a schematic of programmable split pool codon synthesis.
[0035] FIGS. 7A-7B provide an overview of designing and programming a 103library of peptides that each include a different variant section.
[0036] FIGS. 8A-8B provide the results of nucleic acid sequencing of variant domains produced according to the programmed parameters of FIGS. 7A-7B.
[0037] FIGS. 9A-9B provide the results of nucleic acid sequencing of variant domains produced according to the programmed parameters of FIGS. 7A-7B.
[0038] FIGS. 10A-10B show an overview of a 106library of nucleic acids encoding variant domains, and the sequencing results that produce results indicating expected expression of the variant peptide sequences from the encoded nucleic acids.
[0039] FIG. 11A-C shows the production and validation of a variant peptide library using the variant sequences of the library in FIGS. 7A-7B.
[0040] FIG. 12A-C shows the production and validation of a variant peptide library using the variant sequences of the library in FIGS. 7A-7B.VII. Detailed Description
[0041] Throughout this disclosure, numerical features are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of any embodiments. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range to the tenth of the unit of thelower limit unless the context clearly dictates otherwise. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual values within that range, for example, 1.1, 2, 2.3, 5, and 5.9. This applies regardless of the breadth of the range. The upper and lower limits of these intervening ranges may independently be included in the smaller ranges, and are also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure, unless the context clearly dictates otherwise.
[0042] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of any embodiment. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0043] The term oligonucleotide(s), oligo, and polynucleotide are defined to be synonymous throughout. Libraries of variant nucleic acids or variant genes described herein may comprise a plurality of polynucleotides collectively encoding for one or more genes or gene fragments. In some instances, the genetic library comprises coding or non-coding sequences. In some instances, the gene library encodes for a plurality of cDNA sequences. Reference gene sequences from which the cDNA sequences are based may contain introns, whereas cDNA sequences exclude introns. The nucleic acids in the variant gene libraries described herein may encode for genes or gene fragments from an organism. Exemplary organisms include, without limitation, prokaryotes (e.g., bacteria) and eukaryotes (e.g., mice, rabbits, humans, and non-human primates). Each nucleic acid in the variant gene library described herein may encode a different sequence, i.e., non-identical sequence. In some instances, each nucleic acid in the variant gene library described herein comprises at least one portion that is complementary to sequence of another nucleic acid within the library. Nucleic acid sequences described herein may be, unless stated otherwise, comprise DNA or RNA.
[0044] The disclosure provides methods of synthesizing variant libraries, e g., of nucleic acids, polynucleotides, biomolecules, and / or cells expressing the same. In preferred aspects, variant libraries are variant gene libraries.
[0045] “Variant libraries” include nucleic acids, or peptides expressed from such nucleic acids, in which individual members of the library include one or more variant section. With reference to peptides, a variant section comprises a stretch of amino acid sequences that include variations relative to an underlying amino acid sequence of interest. In contrast, the portions of peptides containing such a variant, but which are not the variant section, are referred to herein as “consistent” or “conserved”. For example, a variant peptide library of the present disclosure may include antibody peptide sequences in which a variant section has been programmed into the heavy or light chain variable region of the antibody. Different members of the library may have different variant sections, and thus different variable regions. Concurrently, the members include portions of the antibody peptide sequence that are not variant, i.e., “consistent sections”, which are identical between members of the library.
[0046] Likewise, a variant nucleic acid library may include members that possess a different variant section between members (which may encode a gene that expresses a variant peptide sequence) and one or more consistent section shared by all members. The variants are nucleic acid sequence variations that may result in the amino acid differences of a variant peptide library.
[0047] In some embodiments, oligonucleotides of the invention encode 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10, or more, amino acids that are different from the amino acids relative to an underlying peptide sequence of interest. The encoded, variant amino acids may be adjacent to each other or they may be separate by 1, 2, 3, 4 or more amino acids. The encoded amino acids may be naturally occurring or non-naturally occurring amino acids. The methods of the disclosure may be practiced using populations of oligonucleotides each having the same variant regions, populations of oligonucleotides each having different variant regions, or populations of oligonucleotides where some have the same variant region and some have different variant regions.
[0048] The methods of the disclosure include methods to prepare variant gene libraries from a pool of oligonucleotides. Such methods of the disclosure are advantageous because the variant gene libraries are prepared in a single vessel. In certain aspects of the disclosure, the oligonucleotides used for the preparation of the variant gene library are synthesized on beads. Themethods of the disclosure are also beneficial because they provide large scale and high throughput in the generation of variant gene libraries.
[0049] In certain aspects, the method of the disclosure comprises providing a reaction mixture in a vessel that comprises (i) polynucleotide sequences corresponding to a sequence of interest and (ii) oligonucleotides corresponding to variant sections of the sequence of interest.
[0050] FIG. 1 provides an overview of the methods of the disclosure. As shown in FIG. 1, one or more nucleic acid sequence 103, which may for example, correspond to a gene, are designed with one or more sections that contain different variant sections / variant sequences 105 relative to an underlying sequence of interest (e.g., a wildtype sequence). Thus, while the remainder of the nucleic acid sequence 103 remains consistent (107), the library includes nucleic acids with both the consistent portion(s) 107 and the variant domain(s) 105.
[0051] In certain aspects, the method includes first selecting a sequence of interest from which a variant library is prepared. As an example, reference can be made to the top portion of Fig. 1 which shows a nucleic acid sequence 103 which encodes a selected protein. In certain aspects, the two primary polynucleotides used in methods of the disclosure correspond to portions of the nucleic acid sequence 103 of interest, namely consistent regions 107 of the sequence and variant domains 103 of the sequence. The former are the “nucleic acids” discussed herein, and the latter are the “oligonucleotide” sequences discussed herein. The “nucleic acids” thus correspond to consistent sections of the sequence of interest. The “oligonucleotides” correspond to at least one variant section of the sequence of interest and are the source of the variation found in the variant libraries produced using the methods of the invention.
[0052] In certain preferred methods, these nucleic acids and oligonucleotides are combined into a vessel with reagents for an amplification reaction. The vessel is maintained under conditions wherein the oligonucleotides hybridize with a portion of the nucleic acids to generate hybridized nucleic acid sequences in the vessel. The hybridized nucleic acid sequences are assembled and optionally amplified to produce a library of nucleic acid sequences that share one or more consistent sections and each have one or more different variant sequence at the same location(s).
[0053] In order to produce the various members of the library, as described in FIG. 1, the oligonucleotides corresponding to the different variant sections used in the current methods may be synthesized on beads, for example via combinatorial synthesis. The oligonucleotides aresubsequently cleaved from the beads on which the synthesis was conducted and are used for preparation of the variant gene library.
[0054] In certain embodiments, the nucleic acid comprising a sequence of interest comprises one or more consistent region 107. In certain embodiments, the nucleic acid comprising a sequence of interest comprises one or more variant sections 105. In certain embodiments, the nucleic acid comprising a sequence of interest comprises at least one consistent region and one or more variant sections. The one or more oligonucleotides corresponding to the variant sections of the library may have at least partial overlap with sequences corresponding to the consistent region or variant region of the nucleic acid comprising a sequence of interest.
[0055] The methods of the disclosure provide that the reaction mixture comprises: a nucleic acid comprising a sequence of interest; one or more oligonucleotides comprising a plurality of nucleic acid sequences corresponding to a region of the nucleic acid comprising the sequence of interest; and reagents for an amplification reaction in a vessel. As provided in FIG. 1, the one or more oligonucleotides hybridizes with a portion of the nucleic acid and generates polynucleotide fragments of nucleic acids in the vessel. The method further comprises, conducting a reaction to assemble the fragments of generated nucleic acids in the vessel. The assembly of various fragments of generated nucleic acids results in the generation of a library comprising gene variants.
[0056] In certain embodiments, the nucleic acid fragments generated in the first reaction are assembled by ligation. Ligation refers to the attachment of at least two separate nucleic acid fragments to produce a larger nucleic acid fragment. Methods for joining two nucleic acid fragments are known in the art, and include without limitation, enzymatic and non-enzymatic (e.g., chemical) methods. Examples of ligation reactions that are non-enzymatic include the non- enzymatic ligation techniques described in U.S. Pat. Nos. 5,780,613 and 5,476,930, which are herein incorporated by reference. In some embodiments, an adaptor oligonucleotide is joined to a target polynucleotide by a ligase, for example a DNA ligase or RNA ligase. Multiple ligases, each having characterized reaction conditions, are known in the art, and include, without limitation NAD+-dependent ligases including tRNA ligase, Taq DNA ligase, Thermus / hformis TTEE ligase, Escherichia coh DNA ligase, Tth DNA ligase, Thermus scotoductus DNA ligase (I and II), thermostable ligase, Ampligase thermostable DNA ligase, VanC-type ligase, 9° N DNA Ligase, Tsp DNA ligase, and novel ligases discovered by bioprospecting; ATP-dependent ligases including T4 RNA ligase, T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Pfu DNA ligase, DNAligase 1, DNA ligase III, DNA ligase IV, and novel ligases discovered by bioprospecting; and wild-type, mutant isoforms, and genetically engineered variants thereof. Ligation can be between nucleic acid fragments having hybridizable sequences, such as complementary overhangs. Ligation can also be between two blunt ends. Generally, a 5' phosphate is utilized in a ligation reaction. The 5' phosphate can be provided by the target polynucleotide, the adaptor oligonucleotide, or both. 5' phosphates can be added to or removed from polynucleotides to be joined, as needed.
[0057] In certain embodiments, the generated fragments of nucleic acid are assembled using polymerase cycling assembly (PCA). PCA uses polymerase-mediated chain extension in combination with at least two oligonucleotides having complementary ends that can anneal such that at least one of the polynucleotides has a free 3 '-hydroxyl capable of polynucleotide chain elongation by a polymerase (e.g., a thermostable polymerase such as Taq polymerase, VENT™ polymerase (New England Biolabs), KOD (Novagen) and the like). Overlapping oligonucleotides may be mixed in a standard PCR reaction containing dNTPs, a polymerase, and buffer. The overlapping ends of the oligonucleotides, upon annealing, create regions of double-stranded nucleic acid sequences that serve as primers for the elongation by polymerase in a PCR reaction. Products of the elongation reaction serve as substrates for formation of longer double-strand nucleic acid sequences, eventually resulting in the synthesis of a full-length target sequence. The PCR conditions may be optimized to increase the yield of the target long DNA sequence.
[0058] In certain embodiments, the nucleic acid fragments have a partial overlap with each other. In certain embodiments, the nucleic acid fragments have no overlap with each other. In some aspects, the nucleic acid fragments have an overlap of at least 10 nucleotides. In some aspects, the nucleic acid fragments have an overlap of at least 14 nucleotides. In some aspects, the nucleic acid fragments have an overlap of at least 18 nucleotides. In some aspects, the nucleic acid fragments have an overlap of at least 20 nucleotides. In some aspects, the nucleic acids fragments have any overlap of at least 1 nucleotide. In some aspects, the nucleic acids fragments have any overlap of at least 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, or 25 nucleotides.
[0059] The method of the disclosure further comprises performing a second reaction to amplify the assembled nucleic acids from the previous step. Beneficially, in certain embodiments of the disclosure, all reactions are conducted in the same vessel without any steps involving the isolation of any products and / or reagents between different reactions. In certain embodiments, the methods of the disclosure further comprise expressing the proteins from the genetic variants in gene libraries.
[0060] In certain embodiments, the method of the disclosure generates at least one sequence of interest corresponding to the conserved region of the nucleic acid. In certain embodiments, the method of the disclosure generates multiple sequences of interest corresponding to the conserved region of the nucleic acid. In certain embodiments, the nucleic acid fragments correspond to a conserved region of the sequence of interest. In certain embodiments, the nucleic acid fragments correspond to a plurality of conserved regions of the sequence of interest. In certain embodiments, the nucleic acid fragments correspond to a variant region of the nucleic acid sequence of interest. In certain embodiments, the nucleic acid fragments correspond to a plurality of variant regions of the nucleic acid sequence of interest.
[0061] FIG. 2 provides a depiction of the reactions conducted in the methods of the disclosure. As demonstrated in FIG. 2, in the first reaction (labeled as PCR#1), the first step is conducted wherein the multiple fragments are generated by the one or more oligonucleotides (prepared by combinatorial synthesis), followed by extension of the generated fragments using PCA. The second step (labeled as PCR#2) involves the amplification of the full-length nucleic acid fragments generated in the first step. The combination of these steps results in the generation of a library comprising gene variants.
[0062] In certain embodiments, the one or more oligonucleotides are synthesized on a plurality of beads. In certain aspects of the disclosure, the one or more oligonucleotides are synthesized on the plurality of beads by a combinatorial synthesis. In certain embodiments, the combinatorial synthesis is conducted by chemical reactions. In certain embodiments, the combinatorial synthesis is conducted by enzymatic reactions.
[0063] In some embodiments, the combinatorial synthesis is a constrained / directed split pool combinatorial synthesis. Beneficially, constrained / directed combinatorial libraries of oligonucleotides on beads provides a greater diversity of oligonucleotides created. This in turn results in a higher degree of control over the gene variants in the library. Another benefit ofconducting a constrained / directed split pool combinatorial synthesis of oligonucleotides on beads is that it provides the option of programmable codon synthesis. The term “programmable,” as used herein, indicates that during the oligonucleotide synthesis on the beads, the synthesis results in the sequence of nucleotides in the desired order.
[0064] For example, as depicted in FIG. 6, the programmable split-pool codon synthesis results in minimal spacing in the oligonucleotide sequence and the codons are placed as close as possible for optimal design of the oligonucleotide being synthesized using these methods. As a result of the programmable split-pool codon synthesis, there is minimal spacing between the codons. Thus, the use of combinatorial synthesis for production of oligonucleotides attached to beads provides a higher degree of diversity and control for the eventual variant gene library.
[0065] Similarly, as shown in FIG. 7A, a number of different amino acid sequences are designed as a variant region in a larger peptide sequence. As shown, the variant sequences may be designed with programmed amino acid choices at each position of the variant region. The programmed choices for the amino acid sequence of the variant region are expanded and listed for each position. As shown, from left to right, these exemplary variant regions were designed such that each amino acid position had between 1 and 10 possible amino acids based on the programming. Using all programmed possibilities available in each position, the total number of variant sequences that fall within the designed parameters is 1,000 total variant sequences. A portion of the potential sequences using the exemplary choices for the amino acids of the variant region is shown in FIG. 7B. The resulting library, produced using methods disclosed herein, such as that outlined in FIG. 6, is made of peptides that include the consistent portion of the amino acid sequence and one designed variant region sequence.
[0066] The programmable nature of the presently disclosed methods dramatically reduces the need to produce unwanted library members (i.e., those without desired or designed variant sections). By constraining the library to the desired peptide variant sequences, in this example, only 1,000 members need to be produced, and by extension, the underlying nucleic acid library may be designed to likewise include only 1,000 unique members. In contrast, if no programming was used along the six amino acid variant region of FIG. 7A, using only the 20 natural amino acids, the library could include upwards of 64,000,000 distinct members, with the underlying nucleic acid library including even more potential members due to codon wobble and non-coding library members.
[0067] In some embodiments, beads are used to produce the variant sequences. In some embodiments, a plurality of beads are used. In some aspects, the plurality of beads comprise beads having the same or similar properties. In some embodiments, the plurality of beads comprise beads having different properties. For example, in certain embodiments, the plurality of beads used in the disclosure may comprise gel beads. The gel bead may be a hydrogel bead. In certain embodiments, the beads may be magnetic beads. A gel bead may be formed from molecular precursors, such as a polymeric or monomeric species. In certain aspects, the gel bead may be formed from one or more acrylic polymer. In some embodiments, the acrylic polymer is a methacrylic polymer. In some embodiments, the polymer is a hydroxylated methyacrylic polymer such as in Toy opearl HW-65S beads from Tosoh Bioscience, LLC (Pennsylvania, United States of America). Examples of hydrogels that may be used to form beads used in the methods disclosed herein include, but are not limited to, collagen, hyaluronan, chitosan, fibrin, gelatin, alginate, agarose, chondroitin sulfate, polyacrylamide, polyethylene glycol (PEG), polyvinyl alcohol (PVA), acrylamide / bisacrylamide copolymer matrix, polyacrylamide / poly(acrylic acid) (PAA), hydroxyethyl methacrylate (HEMA), poly N- isopropyl acrylamide (NIP AM), and polyanhydrides, polypropylene fumarate) (PPF). In some cases, the beads may be rigid. In some cases, the beads may be flexible.
[0068] In some embodiments, the plurality of beads may comprise molecular precursors ( .g., monomers or polymers), which may form a polymer network via polymerization of the precursors. In some cases, a precursor may be an already polymerized species capable of undergoing further polymerization via, for example, a chemical cross-linkage. In certain embodiments, a precursor comprises one or more of an acrylamide or a methacrylamide monomer, oligomer, or polymer. In some cases, the bead may comprise prepolymers, which are oligomers capable of further polymerization. For example, polyurethane beads may be prepared using prepolymers. In certain embodiments, the beads may contain individual polymers that may be further polymerized together. In certain embodiments, beads may be generated via polymerization of different precursors, such that they comprise mixed polymers, co-polymers, and / or block co-polymers.
[0069] In certain embodiments, the plurality of beads may comprise natural and / or synthetic materials, including natural and synthetic polymers. Examples of natural polymers include proteins and sugars such as deoxyribonucleic acid, rubber, cellulose, starch (e.g., amylose, amylopectin), proteins, enzymes, polysaccharides, silks, polyhydroxyalkanoates, chitosan, dextran, collagen,carrageenan, ispaghula, acacia, agar, gelatin, shellac, sterculiagum, xanthan gum, Corn sugar gum, guar gum, gum karaya, agarose, alginic acid, alginate, or natural polymers thereof. Examples of synthetic polymers include acrylics, nylons, silicones, spandex, viscose rayon, polycarboxylic acids, polyvinyl acetate, polyacrylamide, polyacrylate, polyethylene glycol, polyurethanes, polylactic acid, silica, polystyrene, polyacrylonitrile, polybutadiene, polycarbonate, polyethylene, polyethylene terephthalate, poly(chlorotrifluoroethylene), poly(ethylene oxide), polyethylene terephthalate), polyethylene, polyisobutylene, poly(methyl methacrylate), poly(oxymethylene), polyformaldehyde, polypropylene, polystyrene, poly(tetrafluoroethylene), poly(vinyl acetate), poly(vinyl alcohol), poly(vinyl chloride), poly(vinylidene dichloride), poly(vinylidene difluoride), poly(vinyl fluoride) and combinations (e.g., co-polymers) thereof.
[0070] In certain embodiments, a chemical cross-linker may be a precursor used to cross-link monomers during polymerization of the monomers and / or may be used to functionalize a bead with a species. In some cases, polymers may be further polymerized with a cross-linker species or other type of monomer to generate a further polymeric network. Non-limiting examples of chemical cross-linkers (also referred to as a "crosslinker" or a "crosslinker agent" herein) include cystamine, gluteraldehyde, dimethyl suberimidate, N-Hydroxy succinimide crosslinker BS3, formaldehyde, carbodiimide (EDC), SMCC, Sulfo-SMCC, vinylsilance, N,N'diallyltartardiamide (DATD), N,N'-Bis(acryloyl)cystamine (BAC), or homologs thereof. In some cases, the crosslinker used in the present disclosure contains cystamine.
[0071] In certain embodiments, any beads described in WO 2014 / 210353 (incorporated herein by reference in its entirety) may be used in the methods of the disclosure.
[0072] In certain embodiments, the beads may be functionalized with a moiety to attach a nucleotide to the beads. The nucleotide attached to the bead may become the precursor to start assembling the oligonucleotide on the beads. In certain embodiments, the oligonucleotides are connected to an acrydite moiety that becomes cross-linked to the bead during the polymerization process. In some embodiments, the oligonucleotides are attached to the acrydite moiety by a disulfide linkage. In certain embodiments, the nucleotides may be assembled using a combinatorial approach.
[0073] In certain embodiments, the one or more oligonucleotides that are synthesized on the plurality of beads are cleaved off from the beads prior to generation of nucleic acid fragments. The oligonucleotides that are cleaved from the bead may be attached to the beads using a cleavablelinkage. In certain embodiments, the cleavable linkage may comprise, for example, a chemically cleavable linkage, a photocl eavable linkage, and / or a thermally cleavable linkage. In certain embodiments, the cleavable linkage is a disulfide linkage.
[0074] In certain embodiments, the one or more oligonucleotides synthesized using a combinatorial approach comprises oligonucleotides that are variants of one or more domains of the nucleic acid corresponding to a protein of interest. In certain embodiments, the one or more oligonucleotides hybridize with the conserved region of the nucleic acid. In certain embodiments, the one or more oligonucleotides hybridize with the variant region of the nucleic acid. The methods of the present disclosure beneficially recognizes that using a combinatorial approach to generate oligonucleotides lead to a high number of variant genes in the genetic library. This leads to generation of a gene variant library in a single tube from a variety of oligonucleotides synthesized on a plurality of beads.
[0075] In certain aspects of the disclosure, a sequence of interest corresponds to a variant region of the corresponding peptide or protein. In certain embodiments, the one or more oligonucleotides comprise non-variant portions of the nucleic acid. In certain embodiments, the one or more oligonucleotides comprise sequences overlapping or partially corresponding to the variant region of the sequence of interest.
[0076] In certain embodiments, one or more oligonucleotides have a length of from about 20 to about 500 nucleotides. In certain embodiments, one or more oligonucleotides have a length of from about 20 to about 400 nucleotides. In certain embodiments, one or more oligonucleotides have a length of from about 20 to about 300 nucleotides. In certain embodiments, one or more oligonucleotides have a length of from about 20 to about 200 nucleotides. In certain embodiments, one or more oligonucleotides have a length of from about 20 to about 100 nucleotides. In certain embodiments, one or more oligonucleotides have a length of from about 20 to about 80 nucleotides. In certain embodiments, one or more oligonucleotides have a length of about 60 nucleotides. In certain embodiments, one or more oligonucleotides have a length of about 40 nucleotides.
[0077] In certain aspects of the disclosure, the one or more oligonucleotides have an annealing temperature of from about 50°C to about 70°C. In certain embodiments, the one or more oligonucleotides have an annealing temperature of from about 55°C to about 68°C. The one ormore oligonucleotides have an annealing temperature of from about 60°C to about 65°C. The one or more oligonucleotides have an annealing temperature of about 62°C ± 1°C.
[0078] In certain aspects of the disclosure, the fragments of nucleic acids produced in the first reaction have an overlap of a number of nucleotides. The overlapping fragments are subsequently assembled to produce the genes in the genetic library. In certain embodiments, the fragments of nucleic acid generated in the first reaction have an overlap of at least about 10 nucleotides. In certain embodiments, the fragments of nucleic acid generated in the first reaction have an overlap of at least about 14 nucleotides. In certain embodiments, the fragments of nucleic acid generated in the first reaction have an overlap of at least about 18 nucleotides. In certain embodiments, the fragments of nucleic acid generated in the first reaction have an overlap of at least about 20 nucleotides.
[0079] In certain aspects of the disclosure, the individual nucleic acids generated in the assembled library results in a genetic library with genetic variants. In certain aspects, the disclosure provides for isolating these generated nucleic acids / genetic variants in their respective droplets. In certain embodiments, these isolated nucleic acids / genetic variants are further amplified in their respective droplets. In certain embodiments, the isolated nucleic acids / genetic variants are expressed in their respective droplets.
[0080] In certain aspects, the methods of the disclosure include partial “one pot” methods and methods that are not “one pot”. This may be required, for example, because of differing annealing temperatures. Thus, aspects, of the disclosure may also include those that use, for example, hierarchical assembly, or performing a bridge ligation of consistent regions.
[0081] The droplets may be a fluidic droplet. In certain embodiments, the droplet may be generated through a droplet generation device, e.g., a microfluidic device or a membrane. Droplets may also be formed by other methods known in the art, e.g., vortexing, extrusion, etc. In some aspects, the droplets generally range from about 0.1 to about 1000pm in diameter or largest dimension, and may have a variation in diameter or largest dimension of less than a factor of 10, e.g., less than a factor of 5, less than a factor of 4, less than a factor of 3, less than a factor of 2, less than a factor of 1.5, less than a factor of 1.4, less than a factor of 1.3, less than a factor of 1.2, less than a factor of 1.1, less than a factor of 1.05, or less than a factor of 1.01, in diameter or the largest dimension. In some embodiments, droplets have a variation in diameter or largest dimension such that at least 50% or more, e.g., 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, or 99%or more of the droplets, vary in diameter or largest dimension by less than a factor of 10, e.g., less than a factor of 5, less than a factor of 4, less than a factor of 3, less than a factor of 2, less than a factor of 1.5, less than a factor of 1.4, less than a factor of 1.3, less than a factor of 1.2, less than a factor of 1.1, less than a factor of 1.05, or less than a factor of 1.01. In some embodiments, droplets have a diameter of about 1.0 pm to 1000 pm, inclusive, such as about 1.0 pm to about 750 pm, about 1.0 pm to about 500 pm, about 1.0 pm to about 250 pm, about 1.0 pm to about 200 pm, about 1.0 pm to about 150 pm, about 1.0 pm to about 100 pm, about 1.0 pm to about 10 pm, or about 1.0 pm to about 5 pm, inclusive.
[0082] In practicing the methods as described herein, the composition and nature of the droplets, e.g., single-emulsion and multiple-emulsion droplets, may vary. For instance, in certain aspects, a surfactant may be used to stabilize the droplets. Accordingly, a droplet may involve a surfactant stabilized emulsion, e.g., a surfactant stabilized single emulsion or a surfactant stabilized double emulsion.
[0083] The droplets described herein may be prepared as emulsions, e.g., as an aqueous phase fluid dispersed in an immiscible phase carrier fluid (e.g., a fluorocarbon oil, silicone oil, or a hydrocarbon oil) or vice versa. For example, multiple-emulsion droplets as described herein may be provided as double-emulsions, e.g., as an aqueous phase fluid in an immiscible phase fluid, dispersed in an aqueous phase carrier fluid; quadruple emulsions, e.g., an aqueous phase fluid in an immiscible phase fluid, in an aqueous phase fluid, in an immiscible phase fluid, dispersed in an aqueous phase carrier fluid; and so on. Generating a single-emulsion droplet or a multipleemulsion droplet as described herein may be performed with or without microfluidic control, e.g., using membrane emulsification. In alternative embodiments, a single-emulsion may be prepared without the use of a microfluidic device, but then modified using a microfluidic device to provide a multiple emulsion, e.g., a double emulsion.
[0084] As depicted in FIG. 3, in certain embodiments, the respective droplets comprising the isolated nucleic acids / genetic variants further comprise a bead or a plurality of beads. In certain embodiments, the beads are magnetic beads. In certain embodiments, the beads comprising nucleic acids / genetic variants are collected and encapsulated for protein expression. Encapsulation in droplets beads and / or reagents, e.g., nucleic acids and / or nucleic acid synthesis reagents (e.g., isothermal nucleic acid amplification reagents and / or nucleic acid amplification reagents), barcode labels, and the like can be achieved via a number of methods, including microfluidic and non-microfluidic methods. In the context of microfluidic methods, there are a number of techniques that can be applied, including glass microcapillary emulsification or emulsification using sequential droplet generation in wettability patterned devices.
[0085] Microcapillary techniques form droplets by generating coaxial jets of the immiscible phase fluids that are induced to break into droplets via coaxial flow focusing through a nozzle. Sequential drop formation in spatially patterned droplet generation junctions can be achieved in devices fabricated lithographically. In some aspects, the present disclosure provides methods for generating an emulsion that encapsulates the bead with genetic variants and / or reagents, e.g., nucleic acids and / or nucleic acid synthesis reagents (e.g., isothermal nucleic acid amplification reagents and / or nucleic acid amplification reagents) without the use of a microfluidic device. In some embodiments described herein, the first fluid is an aqueous phase fluid; the second fluid is a fluid which is immiscible with the first fluid, such as a non-aqueous phase, e.g., a fluorocarbon, silicone oil, oil, or a hydrocarbon oil, or a combination thereof; and the third fluid may be an aqueous phase fluid. Alternatively, in some embodiments, the first fluid is a non-aqueous phase, e.g., a fluorocarbon oil, silicone oil, or a hydrocarbon oil, or a combination thereof; the second fluid is a fluid which is immiscible with the first fluid, e.g., an aqueous phase fluid; and the third fluid is a fluorocarbon oil, silicone oil, or a hydrocarbon oil or a combination thereof.
[0086] The non-aqueous phase fluid may serve as a carrier fluid forming a continuous phase fluid that is immiscible with water, or the non-aqueous phase fluid may be a dispersed phase fluid. The non- aqueous phase fluid may be referred to as an oil phase fluid including at least one oil, but may include any liquid (or liquefiable) compound or mixture of liquid compounds that is immiscible with water. The oil may be synthetic or naturally occurring. The oil may or may not include carbon and / or silicon, and may or may not include hydrogen and / or fluorine. The oil may be lipophilic or lipophobic. In other words, the oil may be generally miscible or immiscible with organic solvents. Exemplary oils may include at least one silicone oil, mineral oil, fluorocarbon oil, vegetable oil, or a combination thereof, among others. In exemplary embodiments, the oil is a fluorinated oil, such as a fluorocarbon oil, which may be a perfluorinated organic solvent. In certain aspects, the first droplets include one bead. In certain aspects, the droplets include a plurality of beads.
[0087] In certain aspects, a surfactant may be included in the first fluid, second fluid, and / or third fluid in preparing the droplets. Accordingly, a droplet may involve a surfactant stabilizedemulsion, e.g., a surfactant stabilized single emulsion or a surfactant stabilized double emulsion, where the surfactant is soluble in the first fluid, second fluid, and / or third fluid. Any convenient surfactant that allows for the desired reactions to be performed in the droplets may be used, including, but not limited to, octylphenol ethoxylate (Triton X-100), polyethylene glycol (PEG), C26H50010 (Tween 20) and / or octylphenoxypolyethoxyethanol (IGEPAL). In other aspects, a droplet is not stabilized by surfactants.
[0088] The surfactant used depends on a number of factors such as the oil and aqueous phases (or other suitable immiscible phases, e.g., any suitable hydrophobic and hydrophilic phases) used for the emulsions. For example, when using aqueous droplets in a fluorocarbon oil, the surfactant may have a hydrophilic block (PEG-PPO) and a hydrophobic fluorinated block (Krytox® FSH). If, however, the oil was switched to a hydrocarbon oil, for example, the surfactant may instead be chosen such that it had a hydrophobic hydrocarbon block, like the surfactant AB IL EM90. Other surfactants can also be envisioned, including ionic surfactants. Other additives can also be included in the oil to stabilize the droplets, including polymers that increase droplet stability at temperatures above 35°C.
[0089] Exemplary surfactants which may be utilized to provide thermostable emulsions are the “biocompatible” surfactants that include PEG-PFPE (polyethyleneglycol-perfluoropolyether) block copolymers, e.g., PEG-Krytox®, surfactants that include ionic Krytox® in the oil phase and Jeffamine® (polyetheramine) in the aqueous phase. Additional and / or alternative surfactants may be used provided they form stable interfaces. Many suitable surfactants will thus be block copolymer surfactants (like PEG-Krytox®) that have a high molecular weight. These examples include fluorinated molecules and solvents, but it is likely that non-fluorinated molecules can be utilized as well. The presently disclosed methods are not limited to a particular surfactant. A variety of surfactants are contemplated including, but not limited to, nonionic and ionic surfactants (e.g., TRITON X-100; TWEEN 20; and TYLOXAPOL) or combinations thereof.
[0090] As depicted in FIG. 3, in certain embodiments, the isolated nucleic acids / genetic variants in the droplets with beads are amplified. In certain embodiments, the amplification reaction is conducted in the droplets. In certain aspects, amplification is performed outside of the droplets. In certain aspects, amplification occurs in a bulk phase. In certain aspects, the amplification is a whole genome amplification (WGA) in the bulk phase. In certain aspects, the isolated nucleic acids / genetic variants are circularized and subjected to WGA in a bulk phase on the beads. Incertain embodiments, the amplified isolated nucleic acids / genetic variants are transferred on the beads. Bead-bound nucleic acid molecules may advantageously be amplified prior to sequencing or protein variant expression (e.g., increasing signal over noise) preferably within the droplets when bound to the beads. Amplification may comprise methods for creating copies of nucleic acids by using thermal cycling to expose reactants to repeated cycles of heating and cooling, and to permit different temperature-dependent reactions (e.g., by PCR). Any suitable PCR method known in the art may be used in connection with the presently described methods.
[0091] The amplified nucleic acid / gene variants may further be expressed in the droplets. The expression may be conducted in the droplets using the systems provided for cell-free protein synthesis (CFPS). The reagents necessary for CFPS are also provided in the droplets. Examples of CFPS are provided in Mei et al., Cell-Free Protein Synthesis in Microfluidic Array Devices, Biotechnol Prog, 2007 Nov-Dec; 23(6): 1305-11; Rolf et al., Application of Cell-Free Protein Synthesis for Faster Biocatalyst Development, Catalysts 2019, 9(2), 190; incorporated herein by reference.
[0092] In certain embodiments, the collected beads are pooled together. In certain embodiments, the nucleic acids and / or beads of the disclosure may be further amplified after they are pooled together. The disclosure further provides that the amplified nucleic acids / genetic variants, which were amplified in their respective droplets, are pooled together prior to the expression.
[0093] In certain aspects, the disclosure further provides for whole genome amplification of the genetic variants generated in the gene library.
[0094] The disclosure further provides methods for generating a library of nucleic acids / gene variants, wherein the library encodes for at least a portion of a variable region and / ora constant region of an antibody peptide of interest. In certain aspects, the library encodes for at least a portion of a variable region and a constant region of an antibody peptide sequence of interest. In certain aspects, members of the library include one or more consistent section that corresponds to a constant region of the antibody and different variant sequences that correspond to all or part of a variable region of the antibody. In certain embodiments, the protein of interest is an antibody. In certain embodiments, the gene variants encode for at least one CDR region of the antibody. Further provided herein are methods for forming a library of nucleic acids, wherein the gene variants in the library encodes for a CDR1, a CDR2, and / or a CDR3 on a heavy chain of the antibody and / or a CDR1, a CDR2, and / or a CDR3 on a light chain of the antibody.
[0095] In certain embodiments, the amplification reaction used in the methods of the disclosure is any amplification reaction used for amplifying nucleic acids. In certain embodiments, the amplification reaction is a polymerase chain reaction. In certain embodiments, the amplification reaction used in the methods of the disclosure is an isothermal amplification reaction.
[0096] Beneficially, the methods of the disclosure provide a high scale and high throughput for assembling gene libraries as compared to the conventional methods for generation of gene libraries. In addition to generation of gene libraries, the methods of the disclosure are beneficial for assembly of genetic circuits, transfection and transformation of genes in the gene library. In certain aspects, the genes in the gene libraries are transferred on beads for clonal amplification in the droplet.
[0097] In certain embodiments, the disclosure provides for sequencing of gene variants in the nucleic acid library. Sequencing nucleic acid molecules may be performed by methods known in the art. For example, see, generally, Quail, et al., 2012, A tale of three next generation sequencing platforms: comparison of Ion Torrent, Pacific Biosciences and Illumina MiSeq sequencers, BMC Genomics 13:341. Nucleic acid molecule sequencing techniques include classic dideoxy sequencing reactions (Sanger method) using labeled terminators or primers and gel separation in slab or capillary, or preferably, next generation sequencing methods. For example, sequencing may be performed according to technologies described in U.S. Pub. 2011 / 0009278, U.S. Pub. 2007 / 0114362, U.S. Pub. 2006 / 0024681, U.S. Pub. 2006 / 0292611, U.S. Pat. 7,960,120, U.S. Pat. 7,835,871, U.S. Pat. 7,232,656, U.S. Pat. 7,598,035, U.S. Pat. 6,306,597, U.S. Pat. 6,210,891, U.S. Pat. 6,828,100, U.S. Pat. 6,833,246, and U.S. Pat. 6,911,345, each incorporated by reference.
[0098] In certain embodiments, the disclosure provides methods for determining whether a specific nucleic acid is present in the nucleic acid library. A feature of certain methods as described herein is the use of a PCR-based assay to detect the presence of certain nucleic acids or binder barcodes. Examples of PCR-based assays of interest include, but are not limited to, quantitative PCR (qPCR), quantitative fluorescent PCR (QF-PCR), multiplex fluorescent PCR (MF-PCR), digital droplet PCR (ddPCR) single cell PCR, PCR-RFLP / real time-PCR-RFLP, hot start PCR, nested PCR, in situ polony PCR, in situ rolling circle amplification (RCA), bridge PCR, picotiter PCR, emulsion PCR, touchdown PCR, and reverse transcriptase PCR (RT-PCR). Other suitable amplification methods include the ligase chain reaction (LCR), transcription amplification, selfsustained sequence replication, selective amplification of target polynucleotide sequences,consensus sequence primed polymerase chain reaction (CP-PCR), arbitrarily primed polymerase chain reaction (AP-PCR), degenerate oligonucleotide-primed PCR (DOP- PCR) and nucleic acid based sequence amplification (NABSA).
[0099] A PCR-based assay may be used to detect the presence of certain nucleic acids (e.g., genes, mRNA, viral-associated nucleic acids) and / or protein barcodes. In such assays, one or more primers specific to each sequence of interest are reacted with the nucleic acids captured by the plurality of barcoded beads in each droplet. These primers have sequences specific to a sequence of interest, so that they will only hybridize and initiate PCR when they are complementary to the sequence of the cell / barcode. If the sequence of interest is present and the primer is a match, many copies of the sequence are created. To determine whether a particular sequence is present, the PCR products may be detected through an assay probing the liquid of the droplet, such as by staining the solution with an intercalating dye, like SybrGreen or ethidium bromide, or detecting them through an intermolecular reaction, such as FRET. These dyes, beads, and the like are each example of a “detection component,” a term that is used broadly and generically herein to refer to any component that is used to detect the presence or absence of nucleic acid amplification products, e.g., PCR products.
[0100] In certain aspects of the disclosure, the genes from the library are further used in cell transfection, cell signal transduction, and / or cell transformation. In certain embodiments, the genes in the variant sequence library are cloned in plasmids. In certain embodiments, the plasmids comprising the variant sequence library are transfected and expressed in the cells. Nucleic acid libraries prepared by methods described herein may be expressed in various cell types. Exemplary cell types include prokaryotes (e.g., bacteria and fungi) and eukaryotes (e.g., plants and animals). Exemplary animals include, without limitation, mice, rabbits, primates, fish, and insects. Exemplary plants include, without limitation, a monocot and dicot. Exemplary plants also include, without limitation, microalgae, kelp, cyanobacteria, and green, brown and red algae, wheat, tobacco, and com, rice, cotton, vegetables, and fruit.
[0101] Nucleic acid libraries prepared by methods described herein may be expressed in various cells associated with a disease state. Cells associated with a disease state include cell lines, tissue samples, primary cells from a subject, cultured cells expanded from a subject, or cells in a model system. Exemplary model systems include, without limitation, plant and animal models of a disease state.
[0102] Nucleic acid libraries prepared by methods described herein may be expressed in various cell types and used to assess a change in cellular activity. Exemplary cellular activities include, without limitation, proliferation, cycle progression, cell death, adhesion, migration, reproduction, cell signaling, energy production, oxygen utilization, metabolic activity, and aging, response to free radical damage, or any combination thereof.
[0103] To identify a variant molecule associated with prevention, reduction, or treatment of a disease state, a variant nucleic acid library described herein is expressed in a cell associated with a disease state, or one in which a disease state can be induced. In some instances, an agent is used to induce a disease state in cells. Exemplary tools for disease state induction include, without limitation, a Cre / Lox recombination system, LPS inflammation induction, and streptozotocin to induce hypoglycemia. The cells associated with a disease state may be cells from a model system or cultured cells, as well as cells from a subject having a particular disease condition. Exemplary disease conditions include a bacterial, fungal, viral, autoimmune, or proliferative disorder (e.g., cancer). In some instances, the variant nucleic acid library is expressed in the model system, cell line, or primary cells derived from a subject, and screened for changes in at least one cellular activity. Exemplary cellular activities include, without limitation, proliferation, cycle progression, cell death, adhesion, migration, reproduction, cell signaling, energy production, oxygen utilization, metabolic activity, and aging, response to free radical damage, or any combination thereof.VIII. ExamplesExample 1:
[0104] This example provides a demonstration of the methods of the disclosure by preparation of variant gene libraries. FIG. 4A (top panel) provides an exemplary Anti-CD3 scFv, wherein the variable heavy chain and variable light chain each comprise CDR1, CDR2, and CDR3 respectively. FIG. 4B provides the data regarding the number of oligonucleotides, the length of oligonucleotides, the length of overlap of the oligonucleotides, and the melting range of the temperature of the oligonucleotides.
[0105] FIG. 5 provides a depiction of the generation of the variant gene library for the exemplary Anti-CD3 scFv. As demonstrated in FIG. 5, that shows results of using a combinatorial approach for gene assembly, as the library is being generated, a large number of gene variants are produced through the annealing of the different nucleic acid fragments. The estimated library size generatedin this process is about 1,048,576 nucleic acids in the variant gene library. Importantly, the gene library is generated using a “one pot” approach, wherein the entire library is generated using a single vessel without any purification and / or separation required in between. As a result, the methods of the disclosure provide for high throughput and low costs for preparation of variant gene libraries.Example 2
[0106] This example provides experimental results corresponding to formation of a library using an exemplary programmable method of the disclosure.
[0107] Initially, as shown in FIGS. 7A-7B. A peptide sequence of interest, e.g., a protein sequence, was designed with a variant region. As shown in FIG. 7A, a number of different amino acid sequences were designed as the variant region in the larger amino acid sequence of the peptide of interest. As shown, the variant sequences were designed with between 1 and 10 programmed amino acid choices at each position of the variant region. The programmed choices for the amino acid sequence of the variant region are expanded and listed for each position. Using all programmed possibilities available in each position, the total number of variant sequences that fall within the designed parameters is 1,000 total variant sequences. A portion of the designed variant sequences using the programmed choices for the amino acids of the variant region is shown in FIG. 7B
[0108] In order to produce the peptide library, an underlying library of nucleic acid sequences was assembled using methods as disclosed herein. Each assembled member of the nucleic acid sequence library included consistent regions of the sequence and a variant region encoding one of the variant amino acid sequences.
[0109] To assure that peptides produced from the nucleic acid sequence library encoded the programmed variant regions, the nucleic acid sequence library was analyzed using nucleic acid sequencing.
[0110] The sequencing results are summarized in FIGS. 8A-8B. FIG. 8A provides a histogram showing the distribution of the number of reads identified per unique sequence in the nucleic acid library. As shown, the library include a high number of reads corresponding to a fairly small set of unique nucleic acid sequences. To confirm that these reads corresponded to desired library members (i.e., those encoding a designed variant region), each unique nucleic acid was comparedto the designed library possibilities. As shown in FIG. 8B, the vast majority of reads corresponded to a sequence in the desired list. Further, as shown, there was a large gap in the number of reads for each unique nucleic acid between the 1,000 sequences with the most reads (i.e., corresponding to the 1,000 possibilities of the designed sequences) and all other sequences found, which each produced a far smaller number of reads. These results indicate that a set of nucleic acids having the size of the designed library encoding the desired variant sequences was produced using the methods of the disclosure.[0U1] Further, as shown in FIGS. 9A-9B, the members of the sequenced library were not only preferentially produced relative to unwanted sequences, they were also produced in fairly equal amounts across all library members. As shown in FIGS. 9A-9B, the synthesized library included members with nucleic acid sequences encoding one of the programmed amino acids at each position of the designed variant region, while essentially excluding all other possible amino acids. Further, there appeared to be little bias towards any particular library member, as all programmed amino acids appeared in roughly equal amounts for each position of the variant region.
[0112] In order to confirm the accuracy and control of the methods disclosed herein, a more complex variant region was designed.
[0113] As shown in FIG. 10A, the newly programmed variant region introduced more potential nucleic acids at each position of the variant region. Using this programmed variant region, the number of possible unique sequences was 1,050,000, which is still far smaller than the potential >1,280,000,000 possible members if the variant region was not programmed. As shown in FIG. 10B, much like the nucleic acid library produced with the more-constrained variant region programming, the resulting nucleic acid library for the complex variant region included members encoding the programmed amino acids at generally expected ratios. Accordingly, the methods disclosed herein are able to produce libraries with over a million desired members, while excluding the far larger number of potential undesired sequences.Example 3
[0114] In this example, the 103variant sequences programmed in Example 2 and referenced in FIGS. 7A-7B were used to produce a library of Green Fluorescent Protein (GFP) proteins into which a variant region was inserted containing one of the designed variant amino acid sequences.
[0115] An overview of the steps undertaken in this example is set forth in FIG. 11A-C. Briefly, as shown in FIG. 11 A, oligonucleotides with sequences encoding the designed variant sections were produced and attached to beads using split-pool combinatorial synthesis. Then, using a one pot gene assembly by PCA, nucleic acid sequences were assembled to incorporate a sequence encoding a designed variant section along with the remainder (i.e., consistent) portion of the GFP gene.
[0116] Briefly, the genes were assembled using multiple rounds of PCR, in which fragments are generated from the oligonucleotides (e.g., encoding the variant section sequences and consistent section sequences) in a first reaction in which multiple fragments are generated by the oligonucleotides, followed by extension of the generated fragments using PCA. The second step involves the amplification of the full-length nucleic acid fragments generated in the first step. The combination of these steps results in the generation of a library comprising gene variants.
[0117] As shown in FIG. 11B, running the resulting library of nucleic acid sequences through a gel electrophoresis assay revealed banding at a size associated with the desired GFP products, indicating the desired assembly of the PCR products from the variant section section oligonucleotides.
[0118] As shown in FIG. 11C, gene products expressed from the library produced fluorescence when exposed to UV light, as is expected of a functional GFP protein.
[0119] Similarly, as shown in FIG. 12A-C, the method used in FIG. 11 was streamlined using fluidics and cell-free protein expression.
[0120] Briefly, the nucleic acid library was produced as in FIG. 11 and the nucleic acids attached to beads and amplified to produce bead-bound amplicons encapsulated in droplets with the required reagents (e.g., for transcription and translation) to express the encoded GFP peptides. Each droplet contains nucleic acids encoding one GFP library peptide member. As shown in FIG. 12B, exposing the droplets to UV light causes those loaded with a bead to glow, indicating successful GFP expression of a library member. Due to Poisson loading of beads into droplets, some droplets are not loaded with beads, and thus lack any GFP expression. As shown in FIG. 12C, the fluorescence from each droplet may be quantified over time, which may, for example, be used to compare relative performance of each library member.
[0121] Thus, as shown, the methods of the disclosure may be used to produce a programmed library of variable peptide sequences, without producing the overwhelming number of unwantednucleic acids / peptide sequences produced using random mutational processes. Further, these examples show the ability of the present disclosure to produce such variable peptide libraries with a consistent and non-biased representation of desired variant sections at the appropriate locations, which allows for the relative properties of each library member to be quickly compared.Incorporation by Reference
[0122] References and citations to other documents, such as patents, patent applications, patent publications, journals, books, papers, web contents, publicly accessible databases, have been made throughout this disclosure. All such documents are hereby incorporated herein by reference in their entirety for all purposes.Equivalents
[0123] Various modifications of the disclosure and many further embodiments thereof, in addition to those shown and described herein, will become apparent to those skilled in the art from the full contents of this document, including references to the scientific and patent literature cited herein. The subject matter herein contains important information, exemplification and guidance that can be adapted to the practice of this disclosure in its various embodiments and equivalents thereof.Non-limiting example combinations
[0124] Features described above as well as those claimed below may be combined in various ways without departing from the scope thereof. The following examples illustrate some possible, non-limiting combinations:
[0125] (Al) A method for generating a variant sequence library, the method comprising: providing in a vessel a reaction mixture comprising: a plurality of nucleic acids each corresponding to a different consistent section of a sequence of interest, one or more oligonucleotides comprising one or more nucleic acid sequence corresponding to a portion of the sequence of interest and comprising at least one variant section with a different sequence relative to the sequence of interest, wherein the one or more oligonucleotides are each attached to one of a plurality of beads, and reagents for an amplification reaction; wherein the one or more oligonucleotides hybridize with a portion of one or more of the plurality of nucleic acids and generate polynucleotide fragments in the vessel; and conducting a reaction to assemble the polynucleotide fragments thereby producing variant nucleic acids comprising at least one of the variant sections.
[0126] (A2) For the method denoted as (Al), amplifying the assembled nucleic acids.
[0127] (A3) For the method denoted as (A2), amplifying the assembled nucleic acids occurs on the plurality of beads.
[0128] (A4) For the method denoted as (A3), the plurality of beads are in fluidic droplets during amplification of the assembled nucleic acids.
[0129] (A5) For the method denoted as (A3) or (A4), the amplification of the assembled nucleic acids is via an emulsion polymerase chain reaction (PCR).
[0130] (A6) For the method denoted as (A5), the emulsion PCR is performed using membrane emulsification.
[0131] (A7) For the method denoted as any of (Al) through (A6), wherein the plurality of beads are in a bulk phase during the amplification of the assembled nucleic acids.
[0132] (A8) For the method denoted as any of (Al) through (A7), the assembled nucleic acids are circularized before amplification, and the amplification is a whole genome amplification.
[0133] (A9) For the method denoted as any of (Al) through (A8), at least 50% of the assembled nucleic acids comprise a variant sequence in the variant region.
[0134] (A10) For the method denoted as any of (Al) through (A9), the conducting step generates at least one sequence of interest.
[0135] (Al l) For the method denoted as any of (Al) through (A10), the conducting step generates a plurality of nucleic acid sequences of interest.
[0136] (A12) For the method denoted as any of (Al) through (Al l), the one or more oligonucleotides are synthesized on the plurality of beads by combinatorial synthesis.
[0137] (A13) For the method denoted as (A12), the one or more oligonucleotides are cleaved from the plurality of beads prior to hybridization.
[0138] (A14) For the method denoted as (A12), the combinatorial synthesis is conducted by chemical reactions.
[0139] (Al 5) For the method denoted as (A12), the combinatorial synthesis is conducted by enzymatic reactions.
[0140] (Al 6) For the method denoted as any of (Al) through (Al 5), the variant region of the sequence of interest corresponds to a variable region of the corresponding peptide or protein.
[0141] (A17) For the method denoted as any of (Al) through (A16), the providing step, the conducting step, and the performing step are conducted in the vessel without any isolation between the steps.
[0142] (Al 8) For the method denoted as any of (Al) through (Al 7), the one or more oligonucleotides further comprise non-variable portions of the nucleic acid.
[0143] (A19) For the method denoted as any of (Al) through (A18), the reaction mixture further comprises one or more oligonucleotides corresponding to non-variable regions of the nucleic acid.
[0144] (A20) For the method denoted as any of (Al) through (Al 9), the one or more oligonucleotides have a length of from about 20 to about 500 nucleotides.
[0145] (A21) For the method denoted as any of (Al) through (A20), the one or more oligonucleotides have a length of from about 20 to about 300 nucleotides.
[0146] (A22) For the method denoted as any of (Al) through (A21), the one or more oligonucleotides have a length of from about 20 to about 200 nucleotides.
[0147] (A23) For the method denoted as any of (Al) through (A22), the one or more oligonucleotides have a nucleotide length of about 60 nucleotides.
[0148] (A24) For the method denoted as any of (Al) through (A23), the one or more oligonucleotides have an annealing temperature of from about 50 °C to about 70 °C.
[0149] (A25) For the method denoted as any of (Al) through (A24), the one or more oligonucleotides have an annealing temperature of from about 55 °C to about 68 °C.
[0150] (A26) For the method denoted as any of (Al) through (A25), the one or more oligonucleotides have an annealing temperature of from about 60 °C to about 65 °C.
[0151] (A27) For the method denoted as any of (Al) through (A26), the one or more oligonucleotides have an annealing temperature of 62 °C ± 1 °C.
[0152] (A28) For the method denoted as any of (Al) through (A27), the fragments of nucleic acids have an overlap of at least about 10 nucleotides.
[0153] (A29) For the method denoted as any of (Al) through (A28), wherein the fragments of nucleic acids have an overlap of at least about 14 nucleotides.
[0154] (A30) For the method denoted as any of (Al) through (A29), the variant sequence library comprises oligonucleotides that are variants of one or more domains of the nucleic acid corresponding to a protein of interest.
[0155] (A31) For the method denoted as any of (A4), isolating genes in the variant sequence library in their respective droplets.
[0156] (A32) For the method denoted as (A31), further comprising amplifying in the respective droplets.
[0157] (A33) For the method denoted as (A31) or (A32), further comprising protein expression in the respective droplets.
[0158] (A34) For the method denoted as (A32), the genes are isolated from their respective droplets and each droplet further comprises a bead.
[0159] (A35) For the method denoted as (A34), wherein the bead is a magnetic bead.
[0160] (A36) For the method denoted as (A35), further comprising amplifying the isolated genes in the respective droplets.
[0161] (A37) For the method denoted as (A36), further comprising transferring the amplified genes on the bead.
[0162] (A38) For the method denoted as (A37), collecting each bead and encapsulating each bead in droplets for protein expression.
[0163] (A39) For the method denoted as any of (A32) through (A38), conducting amplification of the genes isolated in the droplets.
[0164] (A40) For the method denoted as any of (Al) through (A39), further comprising protein expression after the conducting step.
[0165] (A41) For the method denoted as any of (Al) through (A40), the plurality of nucleic acids comprises one or more nucleic acid fragments corresponding to the consistent region of a sequence of interest.
[0166] (A42) For the method denoted as (A41), the one or more nucleic acid fragments correspond to a plurality of consistent regions of the sequence of interest.
[0167] (A43) For the method denoted as any of (Al) through (A42), genes in the variant sequence library are further used in cell transfection, cell signal transduction, and / or cell transformation.
[0168] (A44) For the method denoted as any of (Al) through (A43), genes in the variant sequence library are cloned in plasmids.
[0169] (A45) For the method denoted as any of (Al) through (A44), the one or more oligonucleotides are cleaved from the plurality of beads.
[0170] (A46) For the method denoted as any of (Al) through (A45), the amplification reaction is a polymerase chain reaction.
[0171] (A47) For the method denoted as any of (Al) through (A45), the amplification reaction is an isothermal amplification reaction.
[0172] (A48) For the method denoted as any of (Al) through (A47), the one or more oligonucleotides hybridize with the consistent region of one or more of the plurality of nucleic acids.
[0173] (A49) For the method denoted as (A12), wherein the combinatorial synthesis is split pool combinatorial synthesis.
[0174] (A50) For the method denoted as any of (Al) through (Al l) and any of (Al 6) through (A49), the fragments of nucleic acids generated in the providing step are assembled by ligation.
[0175] (A51) For the method denoted as any of (Al) through (A50), the fragments of nucleic acids have a partial overlap.
[0176] (A52) For the method denoted as any of (Al) through (A50), the fragments of nucleic acid have no overlap.
Claims
ClaimsWhat is claimed is:
1. A method for generating a variant sequence library, the method comprising: providing in a vessel a reaction mixture comprising: a plurality of nucleic acids each corresponding to a different consistent section of a sequence of interest, one or more oligonucleotides comprising one or more nucleic acid sequence corresponding to a portion of the sequence of interest and comprising at least one variant section with a different sequence relative to the sequence of interest, wherein the one or more oligonucleotides are each attached to one of a plurality of beads, and reagents for an amplification reaction; wherein the one or more oligonucleotides hybridize with a portion of one or more of the plurality of nucleic acids and generate polynucleotides in the vessel; and conducting a reaction to assemble the polynucleotides thereby producing variant nucleic acids comprising at least one of the variant sections.
2. The method of claim 1, wherein the method further includes amplifying the assembled variant nucleic acids.
3. The method of claim 2, wherein amplifying the assembled variant nucleic acids occurs on the plurality of beads.
4. The method of claim 3, wherein individual beads of the plurality of beads are isolated in fluidic droplets during amplification of the assembled variant nucleic acids.
5. The method of claim 3, wherein the plurality of beads is in a bulk phase during the amplification of the assembled variant nucleic acids.
6. The method of claim 2, wherein the assembled variant nucleic acids are circularized before amplification, and wherein said amplification is a whole genome amplification.
7. The method of claim 1, wherein the one or more oligonucleotides are synthesized on the plurality of beads by combinatorial synthesis.
8. The method of claim 7, wherein the one or more oligonucleotides are cleaved from the plurality of beads prior to hybridization.
9. The method of claim 1, wherein the providing step and the conducting step are conducted in the vessel without any isolation between the steps.
10. The method of claim 1, wherein the hybridized nucleic acid sequences have an overlap of at least about 10 nucleotides with the oligonucleotides.
11. The method of claim 4, further comprising isolating genes in the variant sequence library in the respective droplets.
12. The method of claim 1, further comprising expressing one or more peptides from the assembled variant nucleic acids after the conducting step.
13. The method of claim 7, wherein the combinatorial synthesis is split pool combinatorial synthesis.
14. The method of claim 1, wherein the variant nucleic acids are assembled by ligation.
15. The method of claim 1, wherein the variant section corresponds to all or part of an antibody heavy or light chain variable region encoded by the sequence of interest.
16. The method of claim 15, wherein the consistent section corresponds to all or part of an antibody heavy or light chain constant region encoded by the sequence of interest.
17. A method for generating a variant sequence library, the method comprising: providing in a vessel a reaction mixture comprising: a plurality of nucleic acids each encoding a portion of a consistent region of a sequence of interest, one or more oligonucleotides each comprising a nucleic acid sequence encoding a variant section relative to the sequence of interest and a consistentsection relative to the sequence of interest, wherein the variant section comprises at least one codon corresponding to an amino acid that is different from an amino acid encoded by the sequence of interest, and wherein the one or more oligonucleotides are each attached to one of a plurality of beads, and reagents for an amplification reaction; maintaining the vessel under conditions wherein the one or more oligonucleotides hybridize with a portion of one or more of the plurality of nucleic acids and generate polynucleotides in the vessel; and conducting a reaction to assemble the polynucleotides thereby producing variant nucleic acids comprising at least one of the variant sections.
18. The method of claim 17, wherein the variant section corresponds to all or part of an antibody heavy or light chain variable region.
19. The method of claim 18, wherein the consistent section corresponds to all or part of an antibody heavy or light chain constant region.
20. The method of claim 17, further comprising amplifying the assembled variant nucleic acids.
21. The method of claim, 20 wherein individual beads of the plurality of beads are in a bulk phase during the amplification of the assembled variant nucleic acids.
22. The method of claim, 20 wherein the plurality of beads is isolated in droplets during the amplification of the assembled variant nucleic acids.
23. The method of claim 17, wherein the one or more oligonucleotides are synthesized on the plurality of beads by combinatorial synthesis.