Method for assembling gene libraries from oligonucleotide pools

By synthesizing and hybridizing oligonucleotide pools on beads, the problems of slow speed and low throughput in traditional synthetic biology have been solved, enabling efficient and economical gene library generation and mutation site control.

CN121909286APending Publication Date: 2026-04-21BROAD INSTITUTE
View PDF 17 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BROAD INSTITUTE
Filing Date
2024-07-12
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Traditional synthetic biology methods are slow, have low throughput and high cost when generating and optimizing custom DNA pathways, and it is difficult to effectively control mutation sites, resulting in a prolonged design-build-test cycle.

Method used

An oligonucleotide pool assembly method is employed in a separate container. Nucleic acids and oligonucleotides are synthesized and hybridized on beads, and variant sequence libraries are generated using restricted split-mix synthesis technology to control mutation locations and perform high-throughput assembly.

Benefits of technology

It enables efficient and economical generation of gene libraries, improves the controllability of mutation locations and the throughput of gene libraries, reduces the generation of unnecessary mutants, and improves design efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121909286A_ABST
    Figure CN121909286A_ABST
Patent Text Reader

Abstract

The present disclosure provides methods, compositions, and systems for generating a library of variant genes comprising nucleic acids.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] I. This disclosure of the field This disclosure relates to methods and compositions for assembling gene libraries from samples containing nucleic acids having target sequence regions.

[0002] II. Priority and related applications This application claims priority and interest in U.S. Application US63 / 526,744, filed July 14, 2023, the contents of which are incorporated herein by reference for all purposes.

[0003] III. Government support This invention was completed with government funding from the Intelligence Advanced Research Projects Agency (IARPA) grant No. 2019-19081900002, the Department of Defense grant No. S4599-001, and the National Institutes of Health grant No. AI142780. The government holds certain rights to this invention.

[0004] IV. background The cornerstone of synthetic biology is the design, build, and test process—an iterative process that requires making DNA readily available so that these customized pathways and organisms can be generated and optimized rapidly and economically. In the design phase, the A, C, T, and G nucleotides that make up DNA are formulated into multiple gene sequences containing target loci or pathways, each sequence variant representing a specific hypothesis to be tested. These variant gene sequences represent a subset of the sequence space, a concept originating in evolutionary biology that refers to the sum of all sequences that make up genes, genomes, transcriptomes, and proteomes.

[0005] For each design-build-test cycle, numerous different variants are typically designed to fully sample the sequence space and maximize the probability of optimizing the design. Despite the conceptual clarity, bottlenecks in speed, throughput, and quality of traditional synthetic methods hinder the progress of this cycle, prolonging development time. This remains a rate-limiting step due to the high cost of extremely precise DNA and the limited throughput of current synthetic techniques, which prevent full exploration of the sequence space.

[0006] Traditionally, the synthesis of different gene variants is achieved through random mutagenesis, typically using large libraries for extensive random mutations in order to find the desired phenotype. This may require recursively identifying the mutations that lead to the desired phenotype, which can be a challenging task when the mutations are generated by random mutagenesis.

[0007] Polymerase chain reaction (PCR) mutagenesis is widely used and represents an established method for assembling mutagenic libraries. These methods are favored for their relatively high throughput in library generation. In such methods, gene regions can be amplified by PCR and mutations introduced, then cloned into a backbone. This can be performed on multiple regions. However, PCR mutagenesis is error-prone, and certain aspects of the process are difficult to control.

[0008] To improve mutagenesis, early chemical gene synthesis efforts focused on producing large quantities of polynucleotides with overlapping sequence homology. These fragments were then assembled and subjected to multiple rounds of polymerase chain reaction (PCR) to link the overlapping polynucleotides into a full-length double-stranded gene. Many factors have hampered this approach, including the time-consuming and labor-intensive construction, the requirement for large quantities of phosphorusamide, expensive raw materials, and the production of nanomolar amounts of the final product, far less than what is needed for downstream steps. Furthermore, the large number of individual polynucleotides required to synthesize a single gene using a 96-well plate is also a significant challenge.

[0009] Synthesizing polynucleotides on microarrays significantly improves the assembly throughput of gene libraries. Large quantities of polynucleotides can be synthesized on the surface of a microarray, then cleaved and pooled. Each polynucleotide targeting a specific gene contains a unique barcode sequence that allows polynucleotides from that specific subgroup to be isolated and assembled into the target gene. During this stage of the process, each subpool is transferred to one well of a plate, increasing the throughput to the number of available wells. Typically, this process is performed in a 96-well plate. While this method offers two orders of magnitude higher throughput than conventional methods, its lack of cost-effectiveness and slow turnaround time still make it insufficient for design, construction, and testing cycles requiring the processing of thousands of sequences at once.

[0010] V. Overview of the Invention The method disclosed herein provides a more efficient and economical approach to assembling gene libraries. In some respects, the method disclosed herein provides a way to generate “variant” sequence libraries in a single pot, also known as “one-pot” assembly of gene libraries from an oligonucleotide pool. Advantageously, the method disclosed herein provides a large-scale, high-throughput means for assembling gene libraries. However, unlike the relatively high-throughput random mutagenesis PCR technique, the currently disclosed method allows for greater control and guidance in the assembly of mutant and variant libraries. Therefore, the method disclosed herein provides a higher percentage of useful or desired mutations than previous methods. This is a significant improvement compared to random PCR mutagenesis, especially when dealing with multiple mutations located at different locations in one or more genes. While random PCR mutagenesis can produce more than 10 mutations in a single assay... 9A library of unique variants, but over 90% of them will be unwanted mutations. Thus, high throughput has both advantages and disadvantages—processing a large number of unwanted mutants can significantly increase costs and reduce overall efficiency.

[0011] The currently available methods can direct mutagenesis and assembly to the desired location and can be carried out at high throughput and multiple scales, rather than allowing random mutations across the entire region, which could result in unwanted mutants (e.g., mutations occurring in areas where they shouldn't be and / or a lack of mutations where they are needed).

[0012] In addition to generating gene libraries, the methods disclosed herein also facilitate the assembly of gene circuits (e.g., as disclosed in WO2015 / 184016, which is incorporated herein by reference), and the transfection and transformation of genes in gene libraries. In some respects, genes in the gene library are transferred onto beads for clonal amplification in droplets.

[0013] In some aspects, the method of this disclosure includes providing a reaction mixture in a container. The mixture comprises: a plurality of nucleic acids, each nucleic acid having a sequence corresponding to a different, consistent fragment of a target sequence; one or more oligonucleotides comprising one or more nucleic acid sequences corresponding to a portion of the target sequence and comprising at least one variant fragment having a different sequence relative to the target sequence; wherein each of the one or more oligonucleotides is linked to one of a plurality of beads; and reagents for an amplification reaction. The one or more oligonucleotides hybridize with a portion of one or more of the plurality of nucleic acids and generate a polynucleotide in the container. These polynucleotides are assembled together to produce a variant nucleic acid comprising one of the variant fragments (which may herein be referred to as a variant domain, variant region, variant sequence, etc.).

[0014] Preferably, each of the oligonucleotides is linked to one of a plurality of beads. In some aspects, the oligonucleotides are synthesized on a plurality of beads. In some preferred embodiments, the nucleic acid containing the target sequence comprises a consistent region. In an exemplary method, the one or more oligonucleotides hybridize with a portion of one or more of the plurality of nucleic acids and generate a polynucleotide fragment of the nucleic acid in the container. In some aspects, a first reaction is performed to assemble a nucleic acid containing a variant sequence from the nucleic acid fragment in the container.

[0015] In a preferred embodiment, the method of this disclosure further includes amplifying the assembled nucleic acid. In some aspects, the amplification of the assembled nucleic acid occurs on a plurality of beads. In some aspects, the plurality of beads are situated in a fluid droplet during the amplification of the assembled nucleic acid. In some aspects, the amplification of the assembled nucleic acid is an emulsion polymerase chain reaction (PCR). In some aspects, emulsion PCR is generated using membrane emulsification. In some methods, the plurality of beads are situated in a bulk phase during the amplification of the assembled nucleic acid. In some aspects, the assembled nucleic acid is circularized prior to amplification, and the amplification is a whole-genome amplification.

[0016] In the preferred method of this disclosure, at least 50% of the assembled nucleic acids contain variant sequences in the variant regions.

[0017] In some embodiments, the nucleic acid containing the target sequence comprises one or more variant fragments. In some aspects, one or more oligonucleotides at least partially overlap with sequences corresponding to conserved or variant regions of the nucleic acid containing the target sequence. One or more oligonucleotides hybridize with a portion of the nucleic acid and generate a polynucleotide fragment of the nucleic acid in a container. In some aspects, the method of this disclosure further includes performing a first reaction in a container to assemble the generated nucleic acid fragment. In some embodiments, the nucleic acid fragments generated in the first step reaction are assembled by a ligation reaction. In some embodiments, the nucleic acid fragments partially overlap each other. In some embodiments, the nucleic acid fragments do not overlap each other.

[0018] The method of this disclosure may further include performing a second reaction to amplify the assembled nucleic acid from the previous step. Advantageously, in some embodiments of this disclosure, all these reactions are carried out in the same container without any steps involving the separation of any products and / or reagents between different reactions. In some embodiments, the method of this disclosure also includes expressing proteins from gene variants in a gene library.

[0019] In some embodiments, the method of this disclosure generates at least one target sequence corresponding to a conserved region of nucleic acid. In some embodiments, the method of this disclosure generates multiple target sequences corresponding to conserved regions of nucleic acid. In some embodiments, the nucleic acid fragment corresponds to a conserved region of the target sequence. In some embodiments, the nucleic acid fragment corresponds to multiple conserved regions of the target sequence. In some embodiments, the nucleic acid fragment corresponds to a variant region of the target nucleic acid sequence.

[0020] In some embodiments, one or more oligonucleotides are synthesized on multiple beads. In some aspects of this disclosure, one or more oligonucleotides are synthesized on multiple beads using a novel split-mix synthesis method. The method of this disclosure includes a restricted split-mix synthesis approach, wherein the available combinations (e.g., codons) at each step are controlled and specified to restrict the design space (desired mutation sites), thereby making library generation as efficient as possible. In some embodiments, synthesis is performed by a chemical reaction. In some embodiments, synthesis is performed by an enzymatic reaction. In some embodiments, the enzymatic reaction includes a polymerase reaction. In some embodiments, the enzymatic reaction includes a nuclease reaction. In some aspects, the enzymatic synthesis is an IIS-type restricted assembly synthesis, such as Golden Gate assembly. In some embodiments, the synthesis is a split-mix combinatorial synthesis. In some embodiments of this disclosure, one or more oligonucleotides synthesized on multiple beads are cleaved from the beads. In some embodiments, the synthesized one or more oligonucleotides include oligonucleotides that are variants of one or more domains of a nucleic acid corresponding to a target protein. In some embodiments, one or more oligonucleotides hybridize with conserved regions of the nucleic acid. In some respects, advantageously, the current method provides traceless assembly, which is crucial when introducing mutations within genes. In some embodiments, one or more oligonucleotides hybridize with variant regions of nucleic acids. This disclosure advantageously recognizes that using a restricted / directed split-mixing method to generate oligonucleotides results in a large number of genes in the gene library. This leads to the generation of gene variant libraries from a variety of oligonucleotides synthesized from beads in a single tube.

[0021] In some aspects of this disclosure, the target sequence corresponds to a variant region of the corresponding peptide or protein. In some embodiments, the one or more oligonucleotides comprise a non-variant (consistent) portion of a nucleic acid. In some embodiments, the one or more oligonucleotides comprise a sequence that overlaps with or partially corresponds to a variant region of the target sequence. In some embodiments, the one or more oligonucleotides are about 20-500 nucleotides in length. In some embodiments, the one or more oligonucleotides are about 20-400 nucleotides in length. In some embodiments, the one or more oligonucleotides are about 20-300 nucleotides in length. In some embodiments, the one or more oligonucleotides are about 20-200 nucleotides in length. In some embodiments, the one or more oligonucleotides are about 20-100 nucleotides in length. In some embodiments, the one or more oligonucleotides are about 20-80 nucleotides in length. In some embodiments, the one or more oligonucleotides are about 60 nucleotides in length. In some embodiments, the one or more oligonucleotides are about 40 nucleotides in length.

[0022] In some aspects of this disclosure, the annealing temperature of one or more oligonucleotides is about 50-70°C. In some embodiments, the annealing temperature of one or more oligonucleotides is about 55-68°C. In some embodiments, the annealing temperature of one or more oligonucleotides is about 60-65°C. In some embodiments, the annealing temperature of one or more oligonucleotides is about 62°C ± 1°C.

[0023] In some aspects of this disclosure, the nucleic acid fragments produced in the first reaction have multiple nucleotide overlaps. The overlapping fragments are then assembled to produce genes in a gene library. In some embodiments, the nucleic acid fragments produced in the first reaction have at least about 10 nucleotide overlaps. In some embodiments, the nucleic acid fragments produced in the first reaction have at least about 14 nucleotide overlaps. In some embodiments, the nucleic acid fragments produced in the first reaction have at least about 18 nucleotide overlaps. In some embodiments, the nucleic acid fragments produced in the first reaction have at least about 20 nucleotide overlaps.

[0024] In some aspects of this disclosure, the individual nucleic acids generated in the assembled library result in a gene library with gene variants. In some aspects, this disclosure provides for isolating these generated nucleic acids / gene variants in their respective droplets. In some embodiments, these isolated nucleic acid / gene variants are further amplified in their respective droplets. In some embodiments, the isolated nucleic acid / gene variants are expressed in their respective droplets. In some embodiments, the individual droplets containing the isolated nucleic acid / gene variants also contain beads. In some embodiments, these beads are magnetic beads. In some embodiments, the isolated nucleic acid / gene variants with beads are amplified. In some embodiments, the amplified isolated nucleic acid / gene variants are transferred onto the beads. In some embodiments, the beads containing the nucleic acid / gene variants are collected and encapsulated for protein expression. In some embodiments, the collected beads are pooled together, and the gene is then expressed as a protein.

[0025] In some implementations, amplified nucleic acid / gene variants in their respective droplets are pooled together and then expressed as proteins.

[0026] In some respects, this disclosure also provides for whole-genome amplification of gene variants generated in a gene library.

[0027] In some aspects of this disclosure, genes from the library are further used for cell transfection, cell signal transduction, and / or cell transformation. In some embodiments, genes from the variant sequence library are cloned into plasmids. In some embodiments, plasmids containing the variant sequence library are transfected into cells and expressed.

[0028] In some embodiments, the amplification reaction used in the methods of this disclosure is any amplification reaction used to amplify nucleic acids. In some embodiments, the amplification reaction is a polymerase chain reaction. In some embodiments, the amplification reaction used in the methods of this disclosure is an isothermal amplification reaction.

[0029] VI. Description of the accompanying drawings Figure 1 A schematic diagram of the method disclosed herein is provided, illustrating the generation of variant gene libraries.

[0030] Figure 2 A schematic diagram of one-pot assembly of the method disclosed herein is provided.

[0031] Figure 3 This is a schematic diagram illustrating gene amplification and expression of variant gene library members generated by the method of this disclosure in a droplet.

[0032] Figures 4A-4B An exemplary library of anti-CD3 single-chain fragment variable antibodies and variants is provided.

[0033] Figure 5 A schematic diagram is shown illustrating the estimation of the size of the gene library generated by the method of this disclosure.

[0034] Figure 6 A schematic diagram of programmable split-hybrid codon synthesis is provided.

[0035] Figures 7A-7B Provided 10 3 An overview of the design and programming of peptide libraries, where each peptide library contains different variant fragments.

[0036] Figures 8A-8B Provided according to Figures 7A-7B The nucleic acid sequencing results of variant domains generated by the programming parameters.

[0037] Figures 9A-9B Provided according to Figures 7A-7B The nucleic acid sequencing results of variant domains generated by the program parameters.

[0038] Figures 10A-10B Showing 10 of the encoded variant fields 6 An overview of the nucleic acid library and sequencing results, which produce the following results: indicating that the variant peptide sequence is expected to be expressed by the encoding nucleic acid.

[0039] Figure 11A -C indicates the use of Figures 7A-7B Variant sequences from Chinese libraries are used to produce and validate variant peptide libraries.

[0040] Figure 12A -C indicates the use of Figures 7A-7BVariant sequences from Chinese libraries are used to produce and validate variant peptide libraries.

[0041] VII. Detailed Description of the Invention In this disclosure, numerical characteristics are expressed in range form. It should be understood that the range form is merely for convenience and brevity and should not be construed as a rigid limitation on the range of any embodiment. Therefore, a range description should be considered as specifically disclosing all possible subranges within that range, as well as individual values ​​within that range up to the decimal place of the lower limit unit, unless the context explicitly indicates otherwise. For example, a description of a range such as 1-6 should be considered as specifically disclosing subranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, etc., and individual values ​​within that range, such as 1.1, 2, 2.3, 5, and 5.9. This principle applies regardless of how broad the range may be. The upper and lower limits of these intermediate ranges may be independently included within smaller ranges and also covered in the disclosure, subject to any specifically excluded limits within the range. If the range contains one or two limits, the range excluding the included one or two limits is also included in this disclosure unless the context explicitly specifies otherwise.

[0042] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit any embodiments. As used herein, the singular forms “a / an” and “the” are also intended to include the plural forms unless the context clearly specifies otherwise. It should also be understood that the term “comprises / comprising” as used herein means the presence of the specified features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. As used herein, the term “and / or” includes any one and all combinations of one or more of the related listed items.

[0043] In this document, the terms oligonucleotide, oligomer, and polynucleotide are defined as synonyms throughout. The variant nucleic acid library or variant gene library described herein may include multiple polynucleotides that collectively encode one or more genes or gene segments. In some cases, the gene library contains coding or non-coding sequences. In some cases, the gene library encodes multiple cDNA sequences. The reference gene sequence on which the cDNA sequence is based may contain introns, while the cDNA sequence itself may not contain introns. The nucleic acids in the variant gene library described herein may encode genes or gene segments derived from an organism. Exemplary organisms include, but are not limited to, prokaryotes (e.g., bacteria) and eukaryotes (e.g., mice, rabbits, humans, and non-human primates). Each nucleic acid in the variant gene library described herein may encode a different sequence, i.e., a non-identical sequence. In some cases, each nucleic acid in the variant gene library described herein contains at least a portion complementary to another nucleic acid sequence in the library. Unless otherwise stated, the nucleic acid sequences described herein may comprise DNA or RNA.

[0044] This disclosure provides methods for synthesizing variant libraries, such as variant libraries of nucleic acids, polynucleotides, biomolecules, and / or cells expressing them. In a preferred aspect, the variant library is a variant gene library.

[0045] A “variant library” comprises nucleic acids or peptides expressed by such nucleic acids, wherein individual members of the library contain one or more variant fragments. With respect to peptides, a variant fragment comprises an amino acid sequence that contains a variant relative to a target base amino acid sequence. Conversely, portions of a peptide that contain such variants but are not variant fragments are referred to herein as “consistent” or “conserved.” For example, a variant peptide library of this disclosure may include antibody peptide sequences in which variant fragments have been programmed into the variable regions of the heavy or light chains of the antibody. Different members in the library may have different variant fragments and therefore different variable regions. Simultaneously, these members also include portions of the antibody peptide sequence that are not variants, i.e., “consistent fragments,” which are identical among members of the library.

[0046] Similarly, a variant nucleic acid library may contain members with different variant fragments (which may encode genes expressing variant peptide sequences) as well as members with one or more consistent fragments shared by all members. These variants are nucleic acid sequence variants that can lead to amino acid differences in the variant peptide library.

[0047] In some embodiments, the oligonucleotides of the present invention encode 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more amino acids that differ from the amino acids in the target basic peptide sequence. The encoded variant amino acids may be adjacent to each other or separated by 1, 2, 3, 4, or more amino acids. The encoded amino acids may be naturally occurring or non-naturally occurring. The methods of this disclosure can be implemented by using a group of oligonucleotides having the same variant region, using a group of oligonucleotides having different variant regions, or using a group of oligonucleotides having both the same and different variant regions.

[0048] The methods disclosed herein include methods for preparing variant gene libraries from an oligonucleotide pool. Such methods of the present disclosure have advantages due to the preparation of variant gene libraries in a single container. In some aspects of the present disclosure, the oligonucleotides used to prepare the variant gene libraries are synthesized on beads. The methods of the present disclosure also have advantages due to the large scale and high throughput they provide in the generation of variant gene libraries.

[0049] In some aspects, the methods of this disclosure include providing a reaction mixture in a container, the reaction mixture comprising: (i) a polynucleotide sequence corresponding to a target sequence, and (ii) an oligonucleotide corresponding to a variant fragment of the target sequence.

[0050] Figure 1 An overview of the methods disclosed herein is provided. For example... Figure 1 As shown, one or more nucleic acid sequences 103 (e.g., possibly corresponding to a gene) are designed to have one or more fragments containing different variant fragments / variant sequences 105 relative to the target base sequence (e.g., wild-type sequence). Thus, while the remainder of the nucleic acid sequence 103 remains unchanged (107), the library contains nucleic acids with a consistent portion 107 and a variant domain 105.

[0051] In some respects, this method includes first selecting a target sequence and then preparing a variant library from it. As an example, see [reference needed]. Figure 1 The upper half of the diagram shows the nucleic acid sequence 103 encoding the selected protein. In some respects, the two major polynucleotides used in the methods of this disclosure correspond to portions of the target nucleic acid sequence 103, namely, the consistent region 107 of the sequence and the variant domain 103 of the sequence. The former is the “nucleic acid” discussed herein, while the latter is the “oligonucleotide” sequence discussed herein. Thus, the “nucleic acid” corresponds to a consistent fragment of the target sequence. The “oligonucleotide” corresponds to at least one variant fragment of the target sequence and is the source of variants found in a variant library generated using the methods of this invention.

[0052] In some preferred methods, these nucleic acids and oligonucleotides are mixed with reagents for the amplification reaction into a container. The container is maintained under conditions that allow the oligonucleotides to hybridize with a portion of the nucleic acids to generate hybridized nucleic acid sequences within the container. The hybridized nucleic acid sequences are assembled and optionally amplified to produce a library of nucleic acid sequences that share one or more consistent fragments and each nucleic acid sequence has one or more distinct variant sequences at the same position.

[0053] In order to generate Figure 1 The library described herein contains multiple members, and oligonucleotides corresponding to different variant motifs used in the current method can be synthesized on beads, for example, through combinatorial synthesis. The oligonucleotides are then cleaved from the beads used for synthesis to prepare variant gene libraries.

[0054] In some embodiments, the nucleic acid containing the target sequence includes one or more congruent regions 107. In some embodiments, the nucleic acid containing the target sequence includes one or more variant fragments 105. In some embodiments, the nucleic acid containing the target sequence includes at least one congruent region and one or more variant fragments. One or more oligonucleotides corresponding to variant fragments of the library may have sequences that at least partially overlap with the sequences corresponding to the congruent or variant regions of the nucleic acid containing the target sequence.

[0055] The method disclosed herein provides a reaction mixture comprising: a nucleic acid containing a target sequence; one or more oligonucleotides containing a plurality of nucleic acid sequences corresponding to the nucleic acid region containing the target sequence; and reagents for performing an amplification reaction in a container. Figure 1 The provided method involves hybridizing one or more oligonucleotides with a portion of a nucleic acid and generating a polynucleotide fragment of the nucleic acid in a container. The method also includes performing a reaction in the container to assemble the generated nucleic acid fragments. The assembly of multiple fragments of the generated nucleic acid results in the generation of a library containing gene variants.

[0056] In some embodiments, the nucleic acid fragments generated in the first reaction are assembled via a ligation reaction. A ligation reaction refers to joining at least two separate nucleic acid fragments to produce a larger nucleic acid fragment. Methods for joining two nucleic acid fragments are known in the art, including but not limited to enzymatic and non-enzymatic methods (e.g., chemical methods). Examples of non-enzymatic ligation reactions include the non-enzymatic ligation techniques described in U.S. Patents US 5,780,613 and US 5,476,930 (incorporated herein by reference). In some embodiments, the linker oligonucleotide is ligated to the target polynucleotide using a ligase (e.g., DNA ligase or RNA ligase). Various ligases are known in the art, each with well-defined reaction conditions, including but not limited to: NAD... +--dependent ligases, including tRNA ligase, Taq DNA ligase, and filamentous thermobacters ( Thermus fliformis DNA ligase, Escherichia coli ( Escherichia coli DNA ligase, Tth DNA ligase, and *Aquaticus niger* (a type of bacteria) Thermus scotoductus DNA ligases (I and II), thermostable ligases, Ampligase thermostable DNA ligases, VanC-type ligases, 9°N DNA ligases, Tsp DNA ligases, and novel ligases discovered through bioprospecting; ATP-dependent ligases, including T4 RNA ligase, T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Pfu DNA ligase, DNA ligase 1, DNA ligase III, DNA ligase IV, and novel ligases discovered through bioprospecting; and wild-type, mutant isoforms, and their genetically engineered variants. Ligation reactions can occur between nucleic acid fragments with hybridizable sequences (such as complementary overhangs). Ligation reactions can also occur between two blunt ends. Generally, a 5' phosphate is used in the ligation reaction. The 5' phosphate can be provided by the target polynucleotide, the linker oligonucleotide, or both. The 5' phosphate can be added to or removed from the polynucleotide to be ligated, as needed.

[0057] In some implementations, the generated nucleic acid fragments are assembled using polymerase cycle assembly (PCA). PCA uses polymerase-mediated chain extension to combine at least two oligonucleotides with complementary ends. These oligonucleotides can be annealed so that at least one polynucleotide has a free 3'-hydroxyl group, which can be extended into a polynucleotide chain by a polymerase (e.g., thermostable polymerases such as Taq polymerase, VENT™ polymerase (New England Biolabs), KOD (Novagen), etc.). The overlapping oligonucleotides can be mixed in a standard PCR reaction containing dNTPs, polymerase, and buffer. After annealing, the overlapping ends of the oligonucleotides form double-stranded nucleic acid sequence regions that can serve as primers for polymerase extension in the PCR reaction. The product of the extension reaction serves as a substrate for forming a longer double-stranded nucleic acid sequence, ultimately leading to the synthesis of the full-length target sequence. PCR conditions can be optimized to increase the yield of long target DNA sequences.

[0058] In some embodiments, the nucleic acid fragments partially overlap each other. In some embodiments, the nucleic acid fragments do not overlap each other. In some aspects, the nucleic acid fragments overlap by at least 10 nucleotides. In some aspects, the nucleic acid fragments overlap by at least 14 nucleotides. In some aspects, the nucleic acid fragments overlap by at least 18 nucleotides. In some aspects, the nucleic acid fragments overlap by at least 20 nucleotides. In some aspects, the nucleic acid fragments overlap by at least 1 nucleotide. In some aspects, the nucleic acid fragments overlap by at least 2, 3, 4, 5, 6, 7, 8, 9, 11, 12, 13, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides.

[0059] The method of this disclosure also includes performing a second reaction to amplify the nucleic acid assembled in the previous step. Advantageously, in some embodiments of this disclosure, all reactions are performed in the same container, eliminating the need to separate any products and / or reagents between different reactions. In some embodiments, the method of this disclosure also includes expressing proteins derived from gene variants in a gene library.

[0060] In some embodiments, the method of this disclosure generates at least one target sequence corresponding to a conserved region of nucleic acid. In some embodiments, the method of this disclosure generates multiple target sequences corresponding to conserved regions of nucleic acid. In some embodiments, the nucleic acid fragment corresponds to a conserved region of the target sequence. In some embodiments, the nucleic acid fragment corresponds to multiple conserved regions of the target sequence. In some embodiments, the nucleic acid fragment corresponds to a variant region of the target nucleic acid sequence. In some embodiments, the nucleic acid fragment corresponds to multiple variant regions of the target nucleic acid sequence.

[0061] Figure 2 A schematic diagram of the reaction performed in the method of this disclosure is provided. For example... Figure 2 As shown, in the first reaction (labeled PCR#1), the first step involves generating multiple fragments from one or more oligonucleotides (prepared through combinatorial synthesis), followed by PCA to extend the generated fragments. The second step (labeled PCR#2) involves amplifying the full-length nucleic acid fragments generated in the first step. These steps combined result in the generation of a library containing gene variants.

[0062] In some embodiments, one or more oligonucleotides are synthesized on multiple beads. In some aspects of this disclosure, one or more oligonucleotides are synthesized on multiple beads by combinatorial synthesis. In some embodiments, combinatorial synthesis is carried out by a chemical reaction. In some embodiments, combinatorial synthesis is carried out by an enzymatic reaction.

[0063] In some implementations, combinatorial synthesis is restricted / directed split-hybrid combinatorial synthesis. Advantageously, restricted / directed oligonucleotide combinatorial libraries on beads provide a greater diversity of oligonucleotides. This, in turn, leads to a higher degree of control over gene variants in the library. Another benefit of restricted / directed split-hybrid combinatorial synthesis of oligonucleotides on beads is that it provides the option of programmable codon synthesis. As used herein, "programmable" means that when oligonucleotides are synthesized on beads, the desired sequence of nucleotides is synthesized.

[0064] For example, such as Figure 6 As shown, programmable split-mix codon synthesis minimizes the spacing between oligonucleotide sequences and brings codons as close together as possible, thereby optimizing the design of oligonucleotides synthesized using the above method. Due to the programmable split-mix codon synthesis, the spacing between codons is minimized. Therefore, using combinatorial synthesis to generate bead-attached oligonucleotides can provide a higher degree of diversity and control for the final variant gene library.

[0065] Similarly, such as Figure 7A As shown, many different amino acid sequences were designed as variant regions within a larger peptide sequence. As illustrated, variant sequences can be designed using programmable amino acid selection at each position within the variant region. The programmable selection of amino acid sequences for the variant regions has been extended, and each position is listed. As shown, from left to right, these exemplary variant regions were designed such that each amino acid position has 1-10 possible amino acids according to programming. Utilizing all available programming possibilities at each position, the total number of variant sequences conforming to the design parameters is a total of 1,000 variant sequences. Figure 7B A portion of a potential sequence is shown, illustrating exemplary selection of amino acids from variant regions. Utilizing the methods disclosed herein (e.g.) Figure 6 The resulting library (the method outlined in the document) consists of peptides containing a consistent portion of the amino acid sequence and a designed variant region sequence.

[0066] The programmability of the currently disclosed methods significantly reduces the need to generate unwanted library members (i.e., members without the desired or designed variant fragments). By restricting the library to the desired peptide variant sequences, in this example, only 1,000 members need to be generated, thus the basic nucleic acid library can also be designed to contain only 1,000 unique members. In contrast, without following... Figure 7A Using programming in the six amino acid variation regions, with only 20 natural amino acids, the library can contain up to 64,000,000 different members, and the basic nucleic acid library contains even more potential members due to codon wobble and non-coding library members.

[0067] In some embodiments, beads are used to generate variant sequences. In some embodiments, multiple beads are used. In some aspects, the multiple beads comprise beads having the same or similar properties. In some embodiments, the multiple beads comprise beads with different properties. For example, in some embodiments, the multiple beads used in this disclosure may comprise gel beads. The gel beads may be hydrogel beads. In some embodiments, these beads may be magnetic beads. Gel beads may be formed from molecular precursors, such as polymers or monomeric substances. In some aspects, gel beads may be formed from one or more acrylic polymers. In some embodiments, the acrylic polymer is a methacrylic acid polymer. In some embodiments, the polymer is a hydroxylated methacrylic acid polymer, such as Toyopearl HW-65S beads from Tosoh Bioscience, LLC (Pennsylvania, USA). Examples of hydrogels that can be used to form the beads used in the methods disclosed herein include, but are not limited to: collagen, hyaluronic acid, chitosan, fibrin, gelatin, alginate, agarose, chondroitin sulfate, polyacrylamide, polyethylene glycol (PEG), polyvinyl alcohol (PVA), acrylamide / bisacrylamide copolymer matrix, polyacrylamide / polyacrylic acid (PAA), hydroxyethyl methacrylate (HEMA), poly(N-isopropylacrylamide) (NIPAM), and polyanhydride, polyacrylate (PPF). In some cases, the beads may be rigid. In other cases, the beads may be flexible.

[0068] In some embodiments, multiple beads may include molecular precursors (e.g., monomers or polymers) that can be polymerized to form a polymer network. In some cases, the precursors may be already polymerized substances capable of further polymerization, such as through chemical crosslinking. In some embodiments, the precursors comprise one or more of acrylamide or methacrylamide monomers, oligomers, or polymers. In some cases, the beads may comprise prepolymers, which are oligomers capable of further polymerization. For example, prepolymers can be used to prepare polyurethane beads. In some embodiments, the beads may comprise a single polymer that can be further polymerized together. In some embodiments, the beads can be generated by polymerizing different precursors, such that they comprise mixed polymers, copolymers, and / or block copolymers.

[0069] In some embodiments, multiple beads may comprise natural and / or synthetic materials, including natural and synthetic polymers. Examples of natural polymers include proteins and carbohydrates such as deoxyribonucleic acid (DNA), rubber, cellulose, starch (e.g., amylose, amylopectin), proteins, enzymes, polysaccharides, silk fibroin, polyhydroxyalkanoates, chitosan, dextran, collagen, carrageenan, psyllium husk, gum arabic, agar, gelatin, shellac, tung oil gum, xanthan gum, guar gum, kalaya gum, agarose, alginic acid, alginate, or natural polymers thereof. Examples of synthetic polymers include acrylates, nylon, silicone, spandex, viscose fiber, polycarboxylic acid, polyvinyl acetate, polyacrylamide, polyacrylate, polyethylene glycol, polyurethane, polylactic acid, silica, polystyrene, polyacrylonitrile, polybutadiene, polycarbonate, polyethylene, polyethylene terephthalate, poly(trifluorochloroethylene), poly(ethylene oxide), polyethylene terephthalate, polyethylene, polyisobutylene, poly(methyl methacrylate), poly(formaldehyde), polyoxymethylene, polypropylene, polystyrene, poly(tetrafluoroethylene), poly(vinyl acetate), polyvinyl alcohol, poly(vinyl chloride), poly(vinylidene chloride), poly(vinylidene fluoride), poly(vinyl fluoride), and combinations thereof (e.g., copolymers).

[0070] In some embodiments, the chemical crosslinking agent may be used as a precursor for crosslinking monomers during monomer polymerization and / or may be used to functionalize the beads with a substance. In some cases, the polymer may be further polymerized with the crosslinking agent substance or other types of monomers to generate further polymer networks. Non-limiting examples of chemical crosslinking agents (also referred to herein as "crosslinking agents" or "crosslinking reagents") include cystamine, glutaraldehyde, dimethyl octanoimide, N-hydroxysuccinimide crosslinking agent BS3, formaldehyde, carbodiimide (EDC), SMCC, sulfonyl-SMCC, vinylsilane, N,N'-diallyl tartrate diamine (DATD), N,N'-bis(acryloyl)cystamine (BAC), or homologues thereof. In some cases, the crosslinking agents used in this disclosure contain cystamine.

[0071] In some implementations, any beads described in WO2014 / 210353 (which is incorporated herein by reference in its entirety) may be used in the methods of this disclosure.

[0072] In some embodiments, the beads can be functionalized with a moiety to attach nucleotides to the beads. The nucleotides attached to the beads can serve as precursors for initiating the assembly of oligonucleotides on the beads. In some embodiments, the oligonucleotides are attached to an acrylamide moiety, which crosslinks with the beads during polymerization. In some embodiments, the oligonucleotides are attached to the acrylamide moiety via disulfide bonds. In some embodiments, combinatorial methods can be used to assemble the nucleotides.

[0073] In some embodiments, one or more oligonucleotides synthesized on multiple beads are cleaved from the beads before generating the nucleic acid fragment. The cleaved oligonucleotides can be linked to the beads using cleavable ligands. In some embodiments, cleavable ligands may include, for example, chemically cleavable ligands, photocleavable ligands, and / or thermally cleavable ligands. In some embodiments, the cleavable ligand is a disulfide bond.

[0074] In some embodiments, one or more oligonucleotides synthesized using a combinatorial approach include oligonucleotides of variants of one or more domains of a nucleic acid corresponding to a target protein. In some embodiments, the one or more oligonucleotides hybridize with conserved regions of the nucleic acid. In some embodiments, the one or more oligonucleotides hybridize with variant regions of the nucleic acid. The method of this disclosure effectively recognizes that using a combinatorial approach to generate oligonucleotides results in a large number of variant genes in a gene library. This leads to the generation of a gene variant library from multiple oligonucleotides synthesized on multiple beads in a single tube.

[0075] In some aspects of this disclosure, the target sequence corresponds to a variant region of the corresponding peptide or protein. In some embodiments, the one or more oligonucleotides comprise a non-variant portion of a nucleic acid. In some embodiments, the one or more oligonucleotides comprise a sequence that overlaps with or partially corresponds to a variant region of the target sequence.

[0076] In some embodiments, the length of one or more oligonucleotides is about 20-500 nucleotides. In some embodiments, the length of one or more oligonucleotides is about 20-400 nucleotides. In some embodiments, the length of one or more oligonucleotides is about 20-300 nucleotides. In some embodiments, the length of one or more oligonucleotides is about 20-200 nucleotides. In some embodiments, the length of one or more oligonucleotides is about 20-100 nucleotides. In some embodiments, the length of one or more oligonucleotides is about 20-80 nucleotides. In some embodiments, the length of one or more oligonucleotides is about 60 nucleotides. In some embodiments, the length of one or more oligonucleotides is about 40 nucleotides.

[0077] In some aspects of this disclosure, the annealing temperature of one or more oligonucleotides is about 50-70°C. In some embodiments, the annealing temperature of one or more oligonucleotides is about 55-68°C. The annealing temperature of one or more oligonucleotides is about 60-65°C. The annealing temperature of one or more oligonucleotides is about 62°C ± 1°C.

[0078] In some aspects of this disclosure, the nucleic acid fragments generated in the first reaction have multiple nucleotide overlaps. The overlapping fragments are then assembled to generate genes in a gene library. In some embodiments, the nucleic acid fragments generated in the first reaction have at least about 10 nucleotide overlaps. In some embodiments, the nucleic acid fragments generated in the first reaction have at least about 14 nucleotide overlaps. In some embodiments, the nucleic acid fragments generated in the first reaction have at least about 18 nucleotide overlaps. In some embodiments, the nucleic acid fragments generated in the first reaction have at least about 20 nucleotide overlaps.

[0079] In some aspects of this disclosure, individual nucleic acids generated in the assembled library produce a gene library with gene variants. In some aspects, this disclosure provides for isolating these generated nucleic acids / gene variants into their respective droplets. In some embodiments, these isolated nucleic acid / gene variants are further amplified in their respective droplets. In some embodiments, the isolated nucleic acid / gene variants are expressed in their respective droplets.

[0080] In some respects, the disclosed methods include both partially "one-pot" and non-"one-pot" methods. This may be necessary, for example, due to varying annealing temperatures. Therefore, this disclosure may also include methods using layered assembly or bridging connections between consistent regions.

[0081] These droplets can be fluid droplets. In some embodiments, droplets can be generated by a droplet generating device (e.g., a microfluidic device or membrane). Droplets can also be formed by other methods known in the art, such as vortexing, extrusion, etc. In some aspects, the diameter or maximum size of the droplets is typically about 0.1-1000 µm, and the variation in diameter or maximum size can be less than 10 times, for example less than 5 times, less than 4 times, less than 3 times, less than 2 times, less than 1.5 times, less than 1.4 times, less than 1.3 times, less than 1.2 times, less than 1.1 times, less than 1.05 times, or less than 1.01 times. In one embodiment, the droplet diameter or maximum size varies by at least 50% or more, such as 60% or more, 70% or more, 80% or more, 90% or more, 95% or more, or 99% or more, and the variation in diameter or maximum size is less than 10 times, such as less than 5 times, less than 4 times, less than 3 times, less than 2 times, less than 1.5 times, less than 1.4 times, less than 1.3 times, less than 1.2 times, less than 1.1 times, less than 1.05 times, or less than 1.01 times. In some embodiments, the droplet diameter is about 1.0-1000 µm (including the endpoint), such as about 1.0-750 µm, about 1.0-500 µm, about 1.0-250 µm, about 1.0-200 µm, about 1.0-150 µm, about 1.0-100 µm, about 1.0-10 µm, or about 1.0-5 µm (all including the endpoint).

[0082] When practicing the methods described herein, the composition and properties of the droplets (e.g., monoemulsion and multiemulsion droplets) may vary. For example, in some aspects, surfactants may be used to stabilize the droplets. Thus, the droplets may involve surfactant-stabilized emulsions, such as surfactant-stabilized monoemulsions or surfactant-stabilized biemulsions.

[0083] The droplets described herein can be prepared as emulsions, for example, by dispersing an aqueous fluid in an immiscible carrier fluid (e.g., fluorocarbon oil, silicone oil, or hydrocarbon oil), and vice versa. For example, the multiple emulsion droplets described herein can be provided as dual emulsions, such as an aqueous fluid in an immiscible phase fluid, then dispersed in an aqueous carrier fluid; or as quadruple emulsions, such as an aqueous fluid in an immiscible phase fluid, then in an aqueous phase fluid, then in an immiscible phase fluid, then dispersed in an aqueous carrier fluid; and so on. The generation of single or multiple emulsion droplets described herein can be performed with or without microfluidics, for example, using membrane emulsification. In an alternative embodiment, a single emulsion can be prepared without a microfluidic device and then modified using a microfluidic device to provide multiple emulsions, such as dual emulsions.

[0084] like Figure 3As shown, in some embodiments, each droplet containing the isolated nucleic acid / gene variant also contains one or more beads. In some embodiments, these beads are magnetic beads. In some embodiments, the beads containing the nucleic acid / gene variant are collected and encapsulated for protein expression. Encapsulating beads and / or reagents such as nucleic acids and / or nucleic acid synthesis reagents (e.g., isothermal nucleic acid amplification reagents and / or nucleic acid amplification reagents), barcode tags, etc., in droplets can be achieved by a variety of methods, including microfluidic and non-microfluidic methods. In the context of microfluidic methods, a variety of techniques can be applied, including glass microcapillary emulsification or emulsification using sequential droplet generation in wettability patterning devices.

[0085] Microcapillary technology forms droplets by generating coaxial jets of immiscible phase fluids, which are then induced to break into droplets by the coaxial flow focusing effect of a nozzle. In devices fabricated using photolithography, continuous droplet formation at spatially patterned droplet generation junctions can be achieved. In some aspects, this disclosure provides methods for generating emulsions that encapsulate beads with gene variants and / or reagents such as nucleic acids and / or nucleic acid synthesis reagents (e.g., isothermal nucleic acid amplification reagents and / or nucleic acid amplification reagents) without the use of microfluidic devices. In some embodiments described herein, the first fluid is an aqueous phase fluid; the second fluid is a fluid immiscible with the first fluid, such as a non-aqueous phase, such as a fluorocarbon compound, silicone oil, oil, or hydrocarbon oil, or a combination thereof; and the third fluid may be an aqueous phase fluid. Alternatively, in some embodiments, the first fluid is a non-aqueous phase, such as a fluorocarbon oil, silicone oil, or hydrocarbon oil, or a combination thereof; the second fluid is a fluid immiscible with the first fluid, such as an aqueous phase fluid; and the third fluid is a fluorocarbon oil, silicone oil, or hydrocarbon oil, or a combination thereof.

[0086] A non-aqueous phase fluid can serve as a carrier fluid, forming a continuous phase fluid that is immiscible with water; alternatively, the non-aqueous phase fluid can be a dispersed phase fluid. The non-aqueous phase fluid can be referred to as an oil phase fluid, containing at least one oil, but may include any liquid (or liquefiable) compound or mixture of liquid compounds that is immiscible with water. This oil may be synthetic or naturally occurring. The oil may or may not contain carbon and / or silicon, and may or may not contain hydrogen and / or fluorine. The oil may be lipophilic or lipophobic. In other words, the oil is generally miscible with organic solvents or immiscible. Exemplary oils may include at least one silicone oil, mineral oil, fluorocarbon oil, vegetable oil, or combinations thereof. In an exemplary embodiment, the oil is a fluorinated oil, such as a fluorocarbon oil, which may be a perfluorinated organic solvent. In some aspects, the initial droplet comprises a single bead. In some aspects, the droplet comprises multiple beads.

[0087] In some respects, surfactants may be included in the first, second, and / or third fluids during the preparation of the droplets. Therefore, the droplets may involve surfactant-stabilized emulsions, such as surfactant-stabilized monoemulsions or surfactant-stabilized dual emulsions, wherein the surfactant is soluble in the first, second, and / or third fluids. Any convenient surfactant that allows the desired reaction to occur in the droplets can be used, including but not limited to: octylphenol ethoxylate (Triton X-100), polyethylene glycol (PEG), C26H50O10 (Tween 20), and / or octylphenoxypolyethoxyethanol (IGEPAL). In other respects, the droplets are not stabilized by the surfactant.

[0088] The surfactant used depends on a variety of factors, such as the oil and aqueous phases (or other suitable immiscible phases, such as any suitable hydrophobic and hydrophilic phases) used in the emulsion. For example, when using aqueous droplets in fluorocarbon oils, the surfactant may have a hydrophilic block (PEG-PPO) and a hydrophobic fluorinated block (Krytox® FSH). However, if the oil is replaced with a hydrocarbon oil, for example, the surfactant can be selected to have a hydrophobic hydrocarbon block, such as surfactant ABIL EM90. Other surfactants, including ionic surfactants, can also be considered. Other additives can also be added to the oil to stabilize the droplets, including polymers that improve the stability of the droplets at temperatures above 35°C.

[0089] Exemplary surfactants that can be used to provide heat-stable emulsions are "biocompatible" surfactants, including PEG-PFPE (polyethylene glycol-perfluoropolyether) block copolymers, such as PEG-Krytox®, and surfactants containing ionic Krytox® in the oil phase and Jeffamine® (polyetheramine) in the aqueous phase. Other surfactants and / or alternative surfactants can be used, provided they form a stable interface. Therefore, many suitable surfactants would be block copolymer surfactants with high molecular weights (such as PEG-Krytox®). These examples include fluorinated molecules and solvents, but non-fluorinated molecules may also be used. The methods disclosed herein are not limited to specific surfactants. The use of a variety of surfactants is contemplated, including, but not limited to, nonionic surfactants and ionic surfactants (e.g., TRITONX-100; TWEEN 20; and TYLOXAPOL) or combinations thereof.

[0090] like Figure 3As shown, in some embodiments, nucleic acid / gene variants isolated in a droplet containing beads are amplified. In some embodiments, the amplification reaction takes place within the droplet. In some aspects, amplification is performed outside the droplet. In some aspects, amplification occurs in the bulk phase. In some aspects, amplification is whole-genome amplification (WGA) in the bulk phase. In some aspects, the isolated nucleic acid / gene variant is circularized and subjected to bulk phase WGA on the beads. In some embodiments, the amplified isolated nucleic acid / gene variant is transferred to the beads. The nucleic acid molecules bound to the beads can advantageously be amplified prior to sequencing or protein variant expression (e.g., to improve the signal-to-noise ratio), preferably amplified within the droplet while bound to the beads. Amplification can include methods for creating nucleic acid copies by exposing reactants to repeated heating and cooling cycles using thermal cycling, and allows for different temperature-dependent reactions (e.g., by PCR). Any suitable PCR method known in the art can be used in combination with the methods described herein.

[0091] The amplified nucleic acid / gene variants can be further expressed in the droplet. Expression can be performed in the droplet using the provided cell-free protein synthesis (CFPS) system. The reagents required for CFPS are also provided in the droplet. An example of CFPS is provided in Meiet al., Cell-Free Protein Synthesis in Microfluidic Array Devices. Biotechnol Prog , 2007 Nov-Dec; 23(6):1305-11; Rolf et al., Application of Cell-FreeProtein Synthesis for Faster Biocatalyst Development, Catalysts 2019, 9 (2), 190 (incorporated into this paper by reference).

[0092] In some embodiments, the collected beads are pooled together. In some embodiments, the nucleic acids and / or beads of this disclosure can be further amplified after pooling. This disclosure also provides for pooling amplified nucleic acid / gene variants amplified in their respective droplets prior to expression.

[0093] In some respects, this disclosure also provides whole-genome amplification of gene variants generated in a gene library.

[0094] This disclosure also provides methods for generating libraries of nucleic acid / gene variants, wherein the library encodes at least a portion of a variable region and / or a constant region of a target antibody peptide. In some aspects, the library encodes at least a portion of a variable region and a constant region of a target antibody peptide sequence. In some aspects, members of the library include one or more congruent fragments corresponding to the antibody constant region, and different variant sequences corresponding to all or part of the antibody variable region. In some embodiments, the target protein is an antibody. In some embodiments, the gene variant encodes at least one CDR region of the antibody. Methods for forming nucleic acid libraries are also provided herein, wherein gene variants in the library encode CDR1, CDR2, and / or CDR3 on the antibody heavy chain and / or CDR1, CDR2, and / or CDR3 on the antibody light chain.

[0095] In some embodiments, the amplification reaction used in this disclosure is any amplification reaction used to amplify nucleic acids. In some embodiments, the amplification reaction is a polymerase chain reaction. In some embodiments, the amplification reaction used in this disclosure is an isothermal amplification reaction.

[0096] Advantageously, compared to traditional gene library generation methods, the method disclosed herein enables large-scale, high-throughput assembly of gene libraries. In addition to gene library generation, the method also facilitates the assembly of gene circuits, transfection, and transformation of genes within the gene library. In some aspects, genes from the gene library are transferred onto beads for clonal amplification within droplets.

[0097] In some embodiments, this disclosure provides for the sequencing of gene variants in a nucleic acid library. Sequencing of nucleic acid molecules can be performed using methods known in the art. For example, generally, see Quail, et al., 2012, A tale of three next-generation sequencing platforms: comparison of Ion Torrent, Pacific Biosciences and Illumina MiSeq sequencers, BMC Genomics 13:341. Nucleic acid molecule sequencing techniques include classical dideoxy sequencing reactions (Sanger sequencing) using labeled terminators or primers, as well as plate or capillary gel separation, or preferably next-generation sequencing methods. For example, sequencing can be performed using techniques described in the following documents: US Publication No. US2011 / 0009278, US Publication No. US2007 / 0114362, US Publication No. US2006 / 0024681, US Publication No. US2006 / 0292611, US Patent No. US7,960,120, US Patent No. US7,835,871, US Patent No. US7,232,656, US Patent No. US7,598,035, US Patent No. US6,306,597, US Patent No. US6,210,891, US Patent No. US6,828,100, US Patent No. US6,833,246 and US Patent No. US6,911,345 (all incorporated herein by reference).

[0098] In some embodiments, this disclosure provides methods for determining the presence of specific nucleic acids in a nucleic acid library. Some of the methods described herein are characterized by using PCR-based assays to detect the presence of certain nucleic acids or binding barcodes. Examples of PCR-based detection methods of interest include, but are not limited to: quantitative PCR (qPCR), quantitative fluorescent PCR (QF-PCR), multiplex fluorescent PCR (MF-PCR), digital droplet PCR (ddPCR), single-cell PCR, PCR-RFLP / real-time PCR-RFLP, hot-start PCR, nested PCR, in situ polymerase cloning (polony) PCR, in situ rolling circle amplification (RCA), bridge PCR, picoliter PCR, emulsion PCR, falling PCR, and reverse transcription PCR (RT-PCR). Other suitable amplification methods include: ligase chain reaction (LCR), transcriptional amplification, self-sustaining sequence replication, selective amplification of target polynucleotide sequences, concordant primer polymerase chain reaction (CP-PCR), random primer polymerase chain reaction (AP-PCR), degenerate oligonucleotide primer PCR (DOP-PCR), and nucleic acid-based sequence amplification (NABSA).

[0099] PCR-based assays can be used to detect the presence of certain nucleic acids (e.g., genes, mRNA, virus-associated nucleic acids) and / or protein barcodes. In such assays, one or more primers specific to each target sequence react with nucleic acids captured by multiple barcode beads in each droplet. These primers have sequences specific to the target sequence, so hybridization and PCR are only initiated when they are complementary to the sequence of the cell / barcode. If the target sequence is present and the primers match, many copies of that sequence are produced. To determine the presence of a specific sequence, PCR products can be detected by assays probing the liquid in the droplet, such as by staining the solution with an intercalating dye (e.g., SybrGreen or ethidium bromide), or by detection via intermolecular reactions (e.g., FRET). These dyes, beads, etc., are all examples of "detection components," and the term "detection component," used broadly and generally throughout this document, refers to any component used to detect the presence or absence of nucleic acid amplification products (e.g., PCR products).

[0100] In some aspects of this disclosure, the genes in the library are further used for cell transfection, cell signal transduction, and / or cell transformation. In some embodiments, the genes in the variant sequence library are cloned into plasmids. In some embodiments, plasmids containing the variant sequence library are transfected into cells and expressed. Nucleic acid libraries prepared by the methods described herein can be expressed in a variety of cell types. Exemplary cell types include prokaryotes (e.g., bacteria and fungi) and eukaryotes (e.g., plants and animals). Exemplary animals include, but are not limited to, mice, rabbits, primates, fish, and insects. Exemplary plants include, but are not limited to, monocotyledons and dicotyledons. Exemplary plants also include, but are not limited to, microalgae, kelp, cyanobacteria, green algae, brown algae and red algae, wheat, tobacco and corn, rice, cotton, vegetables and fruits.

[0101] Nucleic acid libraries prepared by the methods described herein can be expressed in a variety of cells associated with disease states. Cells associated with disease states include: cell lines, tissue samples, primary cells of subjects, cultured cells expanded from subjects, or cells in model systems. Exemplary model systems include, but are not limited to: plant and animal models of disease states.

[0102] Nucleic acid libraries prepared by the methods described herein can be expressed in a variety of cell types and used to assess changes in cell activity. Exemplary cell activities include, but are not limited to: proliferation, cell cycle progression, cell death, adhesion, migration, reproduction, cell signal transduction, energy production, oxygen utilization, metabolic activity and senescence, response to free radical damage, or any combination thereof.

[0103] To identify variant molecules associated with the prevention, reduction, or treatment of disease states, the variant nucleic acid libraries described herein are expressed in cells associated with the disease state or in cells in which the disease state can be induced. In some cases, reagents are used to induce the disease state in cells. Exemplary tools for inducing disease states include, but are not limited to, the Cre / Lox recombination system, LPS-induced inflammation, and streptozotocin to induce hypoglycemia. Cells associated with the disease state can be cells in a model system or cultured cells, or cells from a subject suffering from a specific disease. Exemplary disease states include: bacterial conditions, fungal conditions, viral conditions, autoimmune conditions, or proliferative conditions (e.g., cancer). In some cases, the variant nucleic acid library is expressed in a model system, cell line, or primary cells derived from a subject, and changes in at least one cellular activity are screened. Typical cellular activities include, but are not limited to: proliferation, cell cycle progression, cell death, adhesion, migration, reproduction, cell signal transduction, energy production, oxygen utilization, metabolic activity and senescence, response to free radical damage, or any combination thereof.

[0104] VIII. Examples Example 1: This embodiment demonstrates the method of the present disclosure by preparing a variant gene library. Figure 4A (The above figure) provides an exemplary anti-CD3 scFv, wherein the variable heavy chain and the variable light chain include CDR1, CDR2 and CDR3, respectively. Figure 4B Data on the number of oligonucleotides, the length of oligonucleotides, the overlap length of oligonucleotides, and the melting temperature range of oligonucleotides are provided.

[0105] Figure 5 A diagram showing the generation of an exemplary variant gene library of anti-CD3 scFv is provided. Figure 5 As shown, this illustrates the results of genome assembly using a combinatorial approach, where a large number of gene variants were generated through annealing of different nucleic acid fragments during library production. The estimated library size generated in this process is approximately 1,048,576 nucleic acids in the variant gene library. Importantly, the gene library is generated using a "one-pot" method, where the entire library is generated using a single container without any intermediate purification and / or separation. Therefore, the method disclosed herein provides a high-throughput and low-cost method for preparing variant gene libraries.

[0106] Example 2 This embodiment provides experimental results corresponding to the formation of a library using the exemplary programmable method of this disclosure.

[0107] Initially, such as Figures 7A-7B As shown, target peptide sequences (e.g., protein sequences) with variant regions were designed. Figure 7AAs shown, many different amino acid sequences were designed as variant regions within the larger amino acid sequence of the target peptide. As illustrated, the variant sequences were designed to allow for 1-10 programmable amino acid choices at each position within the variant region. The programmable choices for the amino acid sequences in the variant regions are expanded and listed position by position. Utilizing all available programming possibilities at each position, a total of 1,000 variant sequences conforming to the design parameters were generated. Figure 7B This shows a portion of the variant sequence designed using the selection of amino acid programming for the variant region.

[0108] To generate peptide libraries, we assembled basic nucleic acid sequence libraries using the methods disclosed herein. Each assembled member of the nucleic acid sequence library contains a consistent region of the sequence and a variant region encoding one of the variant amino acid sequences.

[0109] To ensure that the peptides generated from the nucleic acid sequence library encode programmed variant regions, the nucleic acid sequence library was analyzed using nucleic acid sequencing.

[0110] The sequencing results are summarized in Figures 8A-8B middle. Figure 8A A histogram is provided to show the distribution of reads identified for each unique sequence in the nucleic acid library. As shown, the library contains a large number of reads corresponding to a relatively small number of unique nucleic acid sequences. To confirm that these reads correspond to the desired library members (i.e., library members encoding variant regions of the design), we compared each unique nucleic acid with multiple possibilities for the designed library. Figure 8B As shown, the vast majority of reads correspond to sequences in the desired list. Furthermore, as illustrated, there is a significant difference in the number of reads per unique nucleic acid between the 1,000 sequences with the most reads (i.e., corresponding to the designed 1,000 sequence possibilities) and all other discovered sequences, where the number of reads per sequence is much smaller. These results demonstrate that the method of this disclosure can generate a set of nucleic acids with a designed library size encoding the desired variant sequences.

[0111] In addition, such as Figures 9A-9B As shown, members of the sequencing library are not only preferentially generated relative to unwanted sequences, but are also generated in fairly equal numbers across all library members. Figures 9A-9B As shown, the synthesized library contains members of nucleic acid sequences that encode a programmable amino acid at each position in the designed variant region, while essentially excluding all other possible amino acids. Furthermore, there appears to be no preference for any particular library member, as the number of all programmable amino acids at each position in the variant region is approximately equal.

[0112] To verify the accuracy and controllability of the method disclosed in this paper, a more complex variant region was designed.

[0113] like Figure 10A As shown, the newly programmed variant region introduces more potential nucleic acids at each position within the variant region. Using this programmed variant region, the number of possible unique sequences is 1,050,000, which is still far fewer than the more than 1,280,000,000 potential members if the variant region were not programmed. Figure 10B As shown, similar to nucleic acid libraries generated using more restricted variant regions for programming, the resulting nucleic acid libraries with complex variant regions contain members encoding the programmed amino acids in proportions roughly as expected. Therefore, the method disclosed herein can generate libraries with over one million desired members while excluding a much larger number of potentially unwanted sequences.

[0114] Example 3 In this embodiment, the 10 programmed in Embodiment 2 is used. 3 Variant sequences (such as) Figures 7A-7B (The referenced text) generates a green fluorescent protein (GFP) protein library, in which a variant region containing one of the designed variant amino acid sequences is inserted.

[0115] Figure 11A -C illustrates an overview of the steps taken in this embodiment. In short, as... Figure 11A As shown, oligonucleotide sequences encoding the designed variant fragments were generated using a split-mix combinatorial synthesis method and ligated onto beads. Then, using PCA one-pot genome assembly, the nucleic acid sequences were assembled to include the sequences encoding the designed variant fragments along with the remaining (i.e., consistent) portion of the GFP gene.

[0116] In short, genes are assembled using multi-round PCR, where fragments (e.g., encoding variant fragment sequences and consistent fragment sequences) are generated from oligonucleotides in the first reaction, with multiple fragments generated via oligonucleotides, and then the generated fragments are expanded using PCA. The second step involves amplifying the full-length nucleic acid fragments generated in the first step. The combination of these steps produces a library containing gene variants.

[0117] like Figure 11B As shown, the nucleic acid sequence library obtained by the run was determined by gel electrophoresis, and the band size was correlated with the desired GFP product, indicating the desired assembly of PCR products from variant fragment oligonucleotides.

[0118] like Figure 11C As shown, the gene product expressed from the library fluoresces when exposed to ultraviolet light, just as expected of a functional GFP protein.

[0119] Similarly, such as Figure 12A As shown in Figure 11, the method used in Figure 11 has been simplified using fluid dynamics and cell-free protein expression.

[0120] In short, as shown in Figure 11, a nucleic acid library is generated, and the nucleic acids are linked to beads and amplified to produce bead-binding amplicons. These amplicons are encapsulated in droplets and carry the required reagents (e.g., reagents for transcription and translation) to express the encoded GFP peptide. Each droplet contains a nucleic acid encoding one of the GFP library peptide members. Figure 12B As shown, exposing droplets to ultraviolet light causes the droplets loaded with beads to fluoresce, indicating successful GFP expression of the library members. Because the beads are Poisson loaded into the droplets, some droplets do not load beads and therefore lack any GFP expression. Figure 12C As shown, the fluorescence change of each droplet over time can be quantified, for example, it can be used to compare the relative performance of each library member.

[0121] Therefore, as shown, the methods of this disclosure can be used to generate programmed variable peptide sequence libraries without producing a large number of unwanted nucleic acid / peptide sequences generated by random mutation processes. Furthermore, these examples demonstrate that this disclosure is capable of generating such variable peptide libraries that consistently and unbiasedly represent the desired variant fragments at appropriate positions, allowing for rapid comparison of the relative properties of each library member.

[0122] By incorporating via reference This disclosure makes references and citations throughout to other sources, such as patents, patent applications, patent publications, journals, monographs, papers, online content, and publicly accessible databases. All such documents are incorporated herein by reference in their entirety for all purposes.

[0123] equivalent In addition to what is shown and described herein, various modifications and many other embodiments of the invention will be apparent to those skilled in the art from the entire contents of this document (including references to scientific and patent literature cited herein). The subject matter contains important information, examples, and guidance applicable to the practice of various embodiments of this disclosure and their equivalents.

[0124] Unrestricted instance composition The features described above and those described below can be combined in various ways without departing from their scope. The following examples illustrate some possible combinations, but not all: (A1) A method for generating a variant sequence library, the method comprising: The reaction mixture is provided in a container, the reaction mixture comprising: Multiple nucleic acids, each corresponding to a different, consistent segment of the target sequence; One or more oligonucleotides comprising one or more nucleic acid sequences corresponding to a portion of a target sequence and containing at least one variant fragment having a different sequence relative to the target sequence, wherein each of the one or more oligonucleotides is linked to one of a plurality of beads; and Reagents used in the amplification reaction; The one or more oligonucleotides hybridize with a portion of one or more of the plurality of nucleic acids and generate a polynucleotide in the container; and A reaction is carried out to assemble the polynucleotide fragment, thereby producing a variant nucleic acid containing at least one variant fragment.

[0125] (A2) For the method described in (A1), amplify the assembled nucleic acid.

[0126] (A3) For the method described in (A2), the assembled nucleic acid is amplified on multiple beads.

[0127] (A4) For the method described in (A3), multiple beads are in droplets during the amplification of the assembled nucleic acid.

[0128] (A5) For the method described in (A3) or (A4), the amplification of the assembled nucleic acid is performed by emulsion polymerase chain reaction (PCR).

[0129] (A6) For the method described in (A5), emulsion PCR is performed using membrane emulsification.

[0130] (A7) For any one of (A1) to (A6) the method wherein, during the amplification of the assembled nucleic acid, a plurality of beads are in the bulk phase.

[0131] (A8) For any one of (A1) to (A7), the assembled nucleic acid is circularized prior to amplification, wherein the amplification is a whole-genome amplification.

[0132] (A9) For any one of (A1) to (A8), at least 50% of the assembled nucleic acids contain variant sequences in the variant region.

[0133] (A10) For any one of (A1) to (A9), the step of generating at least one target sequence is performed.

[0134] (A11) For any one of (A1) to (A10), the steps are performed to generate multiple target nucleic acid sequences.

[0135] (A12) For any one of (A1) to (A11), one or more oligonucleotides are synthesized on a plurality of beads by combination synthesis.

[0136] (A13) For the method described in (A12), one or more oligonucleotides are cleaved from multiple beads prior to hybridization.

[0137] (A14) For the method described in (A12), the combinatorial synthesis is carried out through a chemical reaction.

[0138] (A15) For the method described in (A12), the combinatorial synthesis is carried out by an enzymatic reaction.

[0139] (A16) For any one of (A1) to (A15) the variant region of the target sequence corresponds to the variable region of the corresponding peptide or protein.

[0140] (A17) For any one of (A1) to (A16) the providing step, the performing step and the executing step are all performed in a container, and there is no separation between the steps.

[0141] (A18) For any one of (A1) to (A17) the oligonucleotide further comprises a non-variable portion of a nucleic acid.

[0142] (A19) For any one of (A1) to (A18) the reaction mixture further comprises one or more oligonucleotides corresponding to the non-variable region of the nucleic acid.

[0143] (A20) For any one of (A1) to (A19), the length of one or more oligonucleotides is about 20-500 nucleotides.

[0144] (A21) In any one of (A1) to (A20) the length of one or more oligonucleotides is about 20-300 nucleotides.

[0145] (A22) In any one of (A1) to (A21) the length of one or more oligonucleotides is about 20-200 nucleotides.

[0146] (A23) In any one of (A1) to (A22) the nucleotide length of one or more oligonucleotides is about 60 nucleotides.

[0147] (A24) For any one of (A1) to (A23) the annealing temperature of one or more oligonucleotides is about 50-70°C.

[0148] (A25) For any one of (A1) to (A24) the annealing temperature of one or more oligonucleotides is about 55-68°C.

[0149] (A26) For any one of (A1) to (A25) the annealing temperature of one or more oligonucleotides is about 60-65°C.

[0150] (A27) For any one of (A1) to (A26) the annealing temperature of one or more oligonucleotides is 62°C ± 1°C.

[0151] (A28) For any one of (A1) to (A27) the nucleic acid fragments have an overlap of at least about 10 nucleotides.

[0152] (A29) For any one of (A1) to (A28), the nucleic acid fragments have an overlap of at least about 14 nucleotides.

[0153] (A30) For any one of (A1) to (A29) the variant sequence library contains an oligonucleotide that is a variant of one or more domains of a nucleic acid corresponding to the target protein.

[0154] (A31) For the method described in (A4), the genes in the variant sequence library are isolated into their respective droplets.

[0155] (A32) The method described in (A31) further includes amplification in the corresponding droplet.

[0156] (A33) For the method described in (A31) or (A32), the method further includes protein expression in the respective droplet.

[0157] (A34) In the method described in (A32), the genes are separated from their respective droplets, and each droplet also contains a bead.

[0158] (A35) For the method described in (A34), the beads are magnetic beads.

[0159] (A36) The method described in (A35) further includes amplifying the genes isolated from the corresponding droplets.

[0160] (A37) The method described in (A36) further includes transferring the amplified gene onto the beads.

[0161] (A38) For the method described in (A37), each bead is collected and each bead is encapsulated in a droplet for protein expression.

[0162] (A39) For any one of (A32) to (A38), the amplification of the isolated gene is carried out in a droplet.

[0163] (A40) For any one of (A1) to (A39), the method further includes protein expression after the step.

[0164] (A41) For any one of (A1) to (A40) the plurality of nucleic acids includes one or more nucleic acid fragments corresponding to a consistent region of the target sequence.

[0165] (A42) For the method described in (A41), one or more nucleic acid fragments correspond to multiple consistent regions of the target sequence.

[0166] (A43) For any one of (A1) to (A42) the gene in the variant sequence library is further used for cell transfection, cell signal transduction and / or cell transformation.

[0167] (A44) For any one of (A1) to (A43) the gene in the variant sequence library is cloned into a plasmid.

[0168] (A45) For any one of (A1) to (A44), one or more oligonucleotides are cleaved from a plurality of beads.

[0169] (A46) For any one of (A1) to (A45), the amplification reaction is a polymerase chain reaction.

[0170] (A47) For any one of (A1) to (A45), the amplification reaction is an isothermal amplification reaction.

[0171] (A48) For any one of (A1) to (A47) the method, one or more oligonucleotides hybridize with a consistent region of one or more of a plurality of nucleic acids.

[0172] (A49) For the method described in (A12), the combinatorial synthesis is a split-mix combinatorial synthesis.

[0173] (A50) For any one of the methods in (A1) to (A11) and any one of the methods in (A16) to (A49), the nucleic acid fragments generated in the steps are provided to be assembled by a ligation reaction.

[0174] (A51) For any one of (A1) to (A50) the nucleic acid fragments have partial overlap.

[0175] (A52) For any one of the methods in (A1) to (A50), the nucleic acid fragments do not overlap.

Claims

1. A method for generating a variant sequence library, the method comprising: The reaction mixture is provided in a container, the reaction mixture comprising: Multiple nucleic acids, each corresponding to a different, consistent segment of the target sequence; One or more oligonucleotides comprising one or more nucleic acid sequences corresponding to a portion of a target sequence and comprising at least one variant fragment having a different sequence relative to the target sequence, wherein each of the one or more oligonucleotides is linked to one of a plurality of beads; and Reagents used in the amplification reaction; The one or more oligonucleotides hybridize with a portion of one or more of the plurality of nucleic acids and generate a polynucleotide in the container; as well as A reaction is carried out to assemble the polynucleotide, thereby producing a variant nucleic acid containing at least one variant fragment.

2. The method of claim 1, wherein the method further comprises amplifying the assembled variant nucleic acid.

3. The method of claim 2, wherein the variant nucleic acid assembled on the plurality of beads is amplified.

4. The method of claim 3, wherein during the amplification and assembly of the variant nucleic acid, individual beads from a plurality of beads are separated in a fluid droplet.

5. The method of claim 3, wherein the plurality of beads are in the bulk phase during the amplification and assembly of the variant nucleic acid.

6. The method of claim 2, wherein the assembled variant nucleic acid is circularized prior to amplification, and wherein the amplification is a whole-genome amplification.

7. The method of claim 1, wherein the one or more oligonucleotides are synthesized on the plurality of beads by combination synthesis.

8. The method of claim 7, wherein the one or more oligonucleotides are cleaved from the plurality of beads prior to hybridization.

9. The method of claim 1, wherein the providing step and the performing step are carried out in the container and there is no separation between the steps.

10. The method of claim 1, wherein the hybridized nucleic acid sequence overlaps with the oligonucleotide by at least about 10 nucleotides.

11. The method of claim 4, further comprising isolating the gene from the variant sequence library in the respective droplet.

12. The method of claim 1, further comprising, after performing the step, expressing one or more peptides from the assembled variant nucleic acid.

13. The method of claim 7, wherein the combination synthesis is a split-mix combination synthesis.

14. The method of claim 1, wherein the variant nucleic acid is assembled via a ligation reaction.

15. The method of claim 1, wherein the variant fragment corresponds to all or part of the variable region of the antibody heavy or light chain encoded by the target sequence.

16. The method of claim 15, wherein the consistent fragment corresponds to all or part of the constant region of the antibody heavy or light chain encoded by the target sequence.

17. A method for generating a variant sequence library, the method comprising: The reaction mixture is provided in a container, the reaction mixture comprising: Multiple nucleic acids, each encoding a portion of a consistent region of the target sequence; One or more oligonucleotides, each oligonucleotide comprising a nucleic acid sequence encoding a variant fragment relative to a target sequence and a uniform fragment relative to the target sequence, wherein the variant fragment contains at least one codon corresponding to an amino acid different from the amino acid encoded by the target sequence, and wherein each of the one or more oligonucleotides is attached to one of a plurality of beads, and Reagents used in the amplification reaction; The container is maintained under certain conditions such that the one or more oligonucleotides hybridize with a portion of one or more of the plurality of nucleic acids and generate a polynucleotide in the container; as well as A reaction is carried out to assemble the polynucleotide, thereby producing a variant nucleic acid containing at least one variant fragment.

18. The method of claim 17, wherein the variant fragment corresponds to all or part of the variable region of the antibody heavy chain or light chain.

19. The method of claim 18, wherein the consistent fragment corresponds to all or part of the constant region of the antibody heavy chain or light chain.

20. The method of claim 17, further comprising amplifying and assembling the variant nucleic acid.

21. The method of claim 20, wherein during the amplification and assembly of the variant nucleic acid, individual beads of the plurality of beads are in the bulk phase.

22. The method of claim 20, wherein the plurality of beads are separated in droplets during the amplification and assembly of the variant nucleic acid.

23. The method of claim 17, wherein the one or more oligonucleotides are synthesized on the plurality of beads by combination synthesis.

Citation Information

Patent Citations

  • Methods for producing a paired tag from a nucleic acid sequence and methods of use thereof

    US20060024681A1

  • Paired end sequencing

    US20060292611A1

  • Confocal imaging methods and apparatus

    US20070114362A1

  • Nucleic acid sequencing system and method

    US20110009278A1

  • Non-enzymatic ligation of oligonucleotides

    US5476930A