De novo synthesized combinatorial nucleic acid library

By developing methods to synthesize mutant nucleic acid libraries with specific codon distributions, the challenges of high costs and low throughput in current DNA synthesis technologies are addressed, enabling efficient exploration of sequence space and accelerating genetic pathway optimization.

JP7696937B2Active Publication Date: 2025-06-23TWIST BIOSCIENCE CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023011285
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-10-27
Filing Date
2023-01-27
Publication Date
2025-06-23
Estimated Expiration
2038-03-14

AI Technical Summary

Technical Problem

Current DNA synthesis technologies are limited by high costs, low throughput, and inefficiencies in exploring the sequence space, hindering the rapid design, construction, and testing of genetic pathways and organisms.

Method used

Methods for synthesizing mutant nucleic acid libraries are developed, involving the generation of predefined sequences encoding polynucleotides with specific codon distributions, followed by synthesis, analysis, and collection of results for a high percentage of sequences, aiming to represent a significant portion of the predicted diversity.

Benefits of technology

These methods enable the rapid and efficient generation of diverse nucleic acid libraries, allowing for the exploration of a large portion of the sequence space, thereby accelerating the design and optimization of genetic pathways and organisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696937000020
    Figure 0007696937000020
  • Figure 0007696937000021
    Figure 0007696937000021
  • Figure 0007696937000022
    Figure 0007696937000022
Patent Text Reader

Abstract

To provide methods for generating highly accurate nucleic acid libraries encoding predetermined variants of a nucleic acid sequence.SOLUTION: Provided herein is a method comprising the steps of: (a) providing predetermined sequences for a plurality of polynucleotides, where the predetermined sequences have a first distribution value that is preselected and is not uniform; b. providing a machine instruction to randomly generate a set of nucleic acid sequence from a single reference sequence; c. calculating a second distribution value from the randomly generated set of nucleic acid sequences; d. determining whether the second distribution value matches the first distribution value or not; e. synthesizing a variant nucleic acid library comprising a plurality of polynucleotides encoded by the randomly generated set of nucleic acid sequences when the first and second distribution values match.SELECTED DRAWING: Figure 22
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference This application claims the benefit of U.S. Provisional Patent Application No. 62 / 578,326, filed Oct. 27, 2017; and U.S. Provisional Patent Application No. 61 / 471,723, filed Mar. 15, 2017, each of which is hereby incorporated by reference in its entirety.

[0002] Sequence Listing This application includes a sequence listing, which is submitted electronically in ASCII format and is hereby incorporated by reference in its entirety. The ASCII copy created on Mar. 13, 2018 has the name 44854-729_601_SL.txt and a size of 18,419 bytes.

Background Art

[0003] The foundation of synthetic biology is the process of design, construction, and testing - an iterative process that requires ready access to DNA for the rapid and facile generation and optimization of pathways and organisms for specific applications. In the design stage, the nucleotides A, C, T, and G that make up DNA are assembled into various gene sequences that include the desired loci or pathways, where each sequence variant is tested against a particular hypothesis. Such variant gene sequences represent a subset of sequence space, a concept that begins with evolutionary biology and relates to the totality of sequences that build genes, genomes, transcriptomes, and proteomes.

[0004] A large number of various mutants are typically designed for each design-build-test cycle to enable appropriate sampling of the sequence space and to maximize the potential for optimized designs. Although conceptually straightforward, process obstacles related to the speed, throughput, and quality of conventional synthesis methods slow down the pace at which this cycle progresses and extend development time. The inability to fully explore the sequence space, due to the high cost of very accurate DNA and the limited throughput of current synthesis technologies, remains the rate-limiting step.

[0005] Starting from the construction phase, two processes are notable: nucleic acid synthesis and gene synthesis. Historically, the synthesis of various gene mutants was accomplished through molecular cloning. Although robust, this approach is not scalable. Early chemical gene synthesis attempts focused on generating many polynucleotides with overlapping sequence homologies. These were pooled and subjected to multiple polymerase chain reactions (PCR) to enable the ligation of the overlapping polynucleotides into full-length double-stranded genes. Many factors, including a time-consuming and labor-intensive structure, the need for large amounts of phosphoramidite, expensive starting materials, and the production of a nanomolar amount of final product, significantly less than that required for downstream processing, hampered this method, and many separate polynucleotides required one 96-well plate to set up the synthesis of one gene.

[0006] The synthesis of polynucleotides on a microarray has significantly increased the throughput of gene synthesis. Many polynucleotides can be synthesized on the microarray surface and then cleaved and pooled together. Each polynucleotide destined for a specific gene contains a unique barcode sequence that allows it to be depooled from that specific subpopulation of polynucleotides and assembled into the desired gene. At this stage of the process, each subpool is transferred to one well in a 96-well plate, increasing the throughput to 96 genes. This is two orders of magnitude higher throughput than classical methods, but it is not cost-effective and time-consuming, and thus not suitable for properly supporting design, construction, and testing cycles that require thousands of sequences at a time.

Summary of the Invention

[0007] Methods for synthesizing mutant nucleic acid libraries are provided herein, the methods comprising: (a) providing a predefined sequence encoding at least 500 polynucleotide sequences, wherein the at least 500 polynucleotide sequences have a predefined codon distribution; (b) synthesizing a plurality of polynucleotides encoding the at least 500 polynucleotide sequences; (c) analyzing the activity of the nucleic acids encoded by the plurality of polynucleotides, or the proteins translated based on the plurality of polynucleotides; and (d) collecting the results from the assay of step (c), wherein the collecting step includes collecting the results of predefined sequences associated with negative or null results. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein step (d) includes collecting results for at least 80% of the predefined sequences. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein step (d) includes collecting results for at least 90% of the predefined sequences. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein step (d) includes collecting results for at least 100% of the predefined sequences. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein at least about 70% of the predicted diversity is represented. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein at least about 90% of the predicted diversity is represented. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein at least about 95% of the predicted diversity is represented. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein at least 80% of the at least 500 polynucleotide sequences are of appropriate size. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein at least about 80% of the at least 500 polynucleotide sequences are present in the mutant nucleic acid library in an amount within two-fold of the average frequency for each of the polynucleotide sequences of the mutant nucleic acid library.Furthermore, provided herein is a method for synthesizing a mutant nucleic acid library, the method further comprising collecting results from the assay of step (c) for a predefined sequence associated with enhanced or reduced activity. Furthermore, provided herein is a method for synthesizing a mutant nucleic acid library, wherein the activity is cell activity. Furthermore, provided herein is a method for synthesizing a mutant nucleic acid library, wherein the cell activity includes proliferation, growth, adhesion, death, migration, energy production, oxygen utilization, metabolic activity, cell signaling, response to free radical damage, or any combination thereof. Furthermore, provided herein is a method for synthesizing a mutant nucleic acid library, wherein the mutant nucleic acid library encodes a sequence for a mutant gene or a fragment thereof. Furthermore, provided herein is a method for synthesizing a mutant nucleic acid library, wherein the mutant nucleic acid library encodes at least a portion of an antibody, an enzyme, or a peptide. Furthermore, provided herein is a method for synthesizing a mutant nucleic acid library, wherein the nucleic acid library encodes a guide RNA (gRNA). Furthermore, provided herein is a method for synthesizing a mutant nucleic acid library, wherein the nucleic acid library encodes siRNA, shRNA, RNAi, or miRNA.

[0008] Methods for generating combinatorial libraries of nucleic acids are provided herein, the methods comprising: (a) designing a predetermined sequence encoding: (i) a first plurality of polynucleotides, each polynucleotide of the first plurality of polynucleotides encoding a variant sequence as compared to a single reference sequence; and (ii) a second plurality of polynucleotides, each polynucleotide of the second plurality of polynucleotides encoding a variant sequence as compared to the single reference sequence; (b) synthesizing the first plurality of polynucleotides and the second plurality of polynucleotides; and (c) mixing the first plurality of polynucleotides and the second plurality of polynucleotides to form a combinatorial library of nucleic acids, wherein at least about 70% of the predicted diversity is represented. Further provided herein are methods for generating a combinatorial library of nucleic acids, wherein the combinatorial library is an unsaturated combinatorial library. Further provided herein are methods for generating a combinatorial library of nucleic acids, wherein the combinatorial library is a saturated combinatorial library. Further provided herein are methods for generating a combinatorial library of nucleic acids, wherein at least 10,000 polynucleotides are synthesized. Further provided herein are methods for generating a combinatorial library of nucleic acids, wherein the total number of polynucleotides for generating an unsaturated combinatorial library is less than at least 25% of the total number of polynucleotides for generating a saturated combinatorial library. Further provided herein are methods for generating a combinatorial library of nucleic acids, wherein at least 80% of the variants are of appropriate size. Further provided herein are methods for generating a combinatorial library of nucleic acids, wherein at least about 90% of the predicted diversity is represented. Further provided herein are methods for generating a combinatorial library of nucleic acids, wherein at least about 95% of the predicted diversity is represented.Furthermore, methods for generating a combinatorial library of nucleic acids are provided herein, wherein the combinatorial library encodes a first reference sequence or a second reference sequence. Furthermore, methods for generating a combinatorial library of nucleic acids are provided herein, wherein the combinatorial library at the time of translation encodes a protein library. Furthermore, methods for generating a combinatorial library of nucleic acids are provided herein, wherein the nucleic acids of the combinatorial library are inserted into a vector. Furthermore, methods for generating a combinatorial library of nucleic acids are provided herein, which further comprises the step of performing PCR mutagenesis of the nucleic acids using the combinatorial library as a primer for the PCR mutagenesis reaction. Furthermore, methods for generating a combinatorial library of nucleic acids are provided herein, wherein the combinatorial library encodes a sequence for a mutant gene or a fragment thereof. Furthermore, methods for generating a combinatorial library of nucleic acids are provided herein, wherein the combinatorial library encodes at least a portion of an antibody, an enzyme, or a peptide. Furthermore, methods for generating a combinatorial library of nucleic acids are provided herein, wherein the combinatorial library encodes at least a portion of a variable region or a constant region of an antibody. Furthermore, methods for generating a combinatorial library of nucleic acids are provided herein, wherein the combinatorial library encodes at least one CDR region of an antibody. Furthermore, methods for generating a combinatorial library of nucleic acids are provided herein, wherein the combinatorial library encodes CDR1, CDR2, and CDR3 on the heavy chain of an antibody and CDR1, CDR2, and CDR3 on the light chain of an antibody. Furthermore, methods for generating a combinatorial library of nucleic acids are provided herein, wherein the combinatorial library encodes a guide RNA (gRNA).

[0009] Methods for synthesizing mutant nucleic acid libraries are provided herein, the methods comprising: (a) providing a predefined sequence encoding a plurality of polynucleotides, the polynucleotides encoding a plurality of codons having mutant sequences as compared to a single reference sequence; (b) selecting distribution values for codons at predefined positions in a predefined nucleic acid reference sequence; (c) providing machine instructions for randomly generating a set of nucleic acid sequences having distribution values that match the selected distribution values, the set of nucleic acid sequences being less than the amount of nucleic acid sequences required to generate a saturated codon mutant library; (d) synthesizing a mutant nucleic acid library of a predefined distribution, wherein at least about 70% of the predicted diversity is represented. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein at least 80% of the mutants are of appropriate size. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein at least about 90% of the predicted diversity is represented. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein at least about 95% of the predicted diversity is represented. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein the mutant nucleic acid library upon translation encodes a protein library. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein the nucleic acids of the mutant nucleic acid library are inserted into a vector. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, the methods further comprising performing PCR mutagenesis of a nucleic acid using the mutant nucleic acid library as a primer for a PCR mutagenesis reaction. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein codon assignment is used to determine each codon of a plurality of codons having mutant sequences. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein codon assignment is based on the frequency of codon sequences in an organism. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, wherein the organism is at least one of an animal, a plant, a fungus, a protist, an archaeon, and a bacterium.Furthermore, methods for synthesizing mutant nucleic acid libraries are provided herein, and codon assignments are based on codon sequence diversity.

[0010] Methods for synthesizing mutant nucleic acid libraries are provided herein, the methods comprising: (a) providing a predefined sequence encoding a plurality of polynucleotides, wherein the polynucleotides encode codons having mutant sequences as compared to a single reference sequence; (b) dividing the plurality of polynucleotides into a 5' fragment of the polynucleotide and a 3' fragment of the polynucleotide; (c) selecting a distribution value for a codon at a predefined position in a predefined nucleic acid reference sequence; (d) providing machine instructions for randomly generating a set of nucleic acids having distribution values that match the selected distribution value, wherein the set of nucleic acids is less than the amount of nucleic acids required to generate a saturated nucleic acid library; (e) synthesizing the 5' fragment of the polynucleotide and the 3' fragment of the polynucleotide; (f) mixing the 5' fragment of the polynucleotide and the 3' fragment of the polynucleotide to form a mutant nucleic acid library, wherein at least about 70% of the predicted diversity is represented. Further provided herein are methods for synthesizing mutant nucleic acid libraries in which at least 10,000 polynucleotides are synthesized. Further provided herein are methods for synthesizing mutant nucleic acid libraries in which at least 80% of the mutants are of appropriate size. Further provided herein are methods for synthesizing mutant nucleic acid libraries in which at least about 90% of the predicted diversity is represented. Further provided herein are methods for synthesizing mutant nucleic acid libraries in which at least about 95% of the predicted diversity is represented. Further provided herein are methods for synthesizing mutant nucleic acid libraries in which the plurality of polynucleotides are divided into at least one 5' fragment and at least one 3' fragment, each having more than one. Further provided herein are methods for synthesizing mutant nucleic acid libraries in which the mutant nucleic acid library encodes a protein library upon translation. Further provided herein are methods for synthesizing mutant nucleic acid libraries in which the nucleic acids of the mutant nucleic acid library are inserted into vectors. Further provided herein are methods for synthesizing mutant nucleic acid libraries, the methods further comprising performing PCR mutagenesis of nucleic acids using the mutant nucleic acid library as a primer for a PCR mutagenesis reaction.Provided herein is a method for synthesizing a mutant nucleic acid library, further comprising the step of identifying mutant sequences having enhanced or reduced activity. Further provided herein is a method for synthesizing a mutant nucleic acid library, wherein the activity is cell activity. Further provided herein is a method for synthesizing a mutant nucleic acid library, wherein the cell activity includes proliferation, growth, adhesion, death, migration, energy production, oxygen utilization, metabolic activity, cell signaling, response to free radical damage, or any combination thereof. Further provided herein is a method for synthesizing a mutant nucleic acid library, wherein the mutant nucleic acid library encodes a sequence for a mutant gene or a fragment thereof. Further provided herein is a method for synthesizing a mutant nucleic acid library, wherein the mutant nucleic acid library encodes at least a portion of an antibody, an enzyme, or a peptide. Further provided herein is a method for synthesizing a mutant nucleic acid library, wherein the mutant nucleic acid library encodes at least a portion of a variable region or a constant region of an antibody. Further provided herein is a method for synthesizing a mutant nucleic acid library, wherein the mutant nucleic acid library encodes at least one CDR region of an antibody. Further provided herein is a method for synthesizing a mutant nucleic acid library, wherein the mutant nucleic acid library encodes CDR1, CDR2, and CDR3 on the heavy chain of an antibody and CDR1, CDR2, and CDR3 on the light chain of an antibody. Further provided herein is a method for synthesizing a mutant nucleic acid library, wherein many different sequences synthesized in the mutant nucleic acid library range from 50 to 1,000,000. Further provided herein is a method for synthesizing a mutant nucleic acid library, wherein many different sequences synthesized in the mutant nucleic acid library range from 500 to 25,000. Further provided herein is a method for synthesizing a mutant nucleic acid library, wherein many different sequences synthesized in the mutant nucleic acid library range from 1,000 to 15,000. Further provided herein is a method for synthesizing a mutant nucleic acid library, the method further comprising performing PCR mutagenesis of a nucleic acid using the mutant nucleic acid library as a primer for a PCR mutagenesis reaction.Furthermore, methods for synthesizing mutant nucleic acid libraries are provided herein, and the codon assignments are used to determine codons having mutant sequences. Furthermore, methods for synthesizing mutant nucleic acid libraries are provided herein, and the codon assignments are based on the frequency of codon sequences in organisms. Furthermore, methods for synthesizing mutant nucleic acid libraries are provided herein, and the organisms are at least one of animals, plants, fungi, protists, archaea, and bacteria. Furthermore, methods for synthesizing mutant nucleic acid libraries are provided herein, and the codon assignments are based on the diversity of codon sequences. Furthermore, methods for synthesizing mutant nucleic acid libraries are provided herein, and the mutant nucleic acid libraries encode guide RNAs (gRNAs).

[0011] Methods for generating combinatorial libraries of nucleic acids are provided herein, the methods comprising: (a) providing a predefined sequence encoding (i) a first plurality of polynucleotides, wherein each polynucleotide of the first plurality of polynucleotides encodes a variant sequence as compared to a single reference sequence, and (ii) a second plurality of polynucleotides, wherein each polynucleotide of the second plurality of polynucleotides encodes a variant sequence as compared to the single reference sequence; (b) providing a structure having a surface; (c) synthesizing a first plurality of polynucleotides, wherein each polynucleotide of the first plurality of polynucleotides extends from the surface; (d) synthesizing a second plurality of polynucleotides, wherein each polynucleotide of the second plurality of polynucleotides extends from the surface; (e) releasing the first plurality of polynucleotides and the second plurality of polynucleotides from the surface; and (f) mixing the first plurality of polynucleotides and the second plurality of polynucleotides to form a combinatorial library of nucleic acids, wherein at least about 70% of the predicted diversity is represented. Further provided herein are methods for generating combinatorial libraries of nucleic acids, wherein at least about 90% of the predicted diversity is represented. Further provided herein are methods for generating combinatorial libraries of nucleic acids, wherein at least about 95% of the predicted diversity is represented.

[0012] Methods for synthesizing mutant nucleic acid libraries are provided herein, the methods comprising: (a) designing a predefined sequence encoding a plurality of polynucleotides, wherein the polynucleotides encode a plurality of codons having mutant sequences as compared to a single reference sequence; (b) synthesizing a plurality of polynucleotides to generate a mutant nucleic acid library, wherein at least about 70% of the predicted diversity is represented; (c) expressing the mutant nucleic acid library; and (d) evaluating an activity associated with the mutant nucleic acid library. Further provided herein are methods for synthesizing mutant nucleic acid libraries, wherein at least about 90% of the predicted diversity is represented. Further provided herein are methods for synthesizing mutant nucleic acid libraries, wherein at least about 95% of the predicted diversity is represented.

[0013] Methods for generating combinatorial libraries of nucleic acids are provided herein, the methods comprising: (a) providing a predefined sequence encoding (i) a first plurality of non-identical polynucleotides, each non-identical polynucleotide of the first plurality of non-identical polynucleotides encoding a variant sequence as compared to a single reference sequence, and (ii) a second plurality of non-identical polynucleotides, each non-identical polynucleotide of the second plurality of non-identical polynucleotides encoding a variant sequence as compared to the single reference sequence; (b) providing a structure having a surface; (c) synthesizing a first plurality of non-identical polynucleotides, each non-identical polynucleotide of the first plurality of non-identical polynucleotides extending from the surface; (d) synthesizing a second plurality of non-identical polynucleotides, each non-identical polynucleotide of the second plurality of non-identical polynucleotides extending from the surface; (e) releasing the first plurality of non-identical polynucleotides and the second plurality of non-identical polynucleotides from the surface; and (f) mixing the first plurality of polynucleotides and the second plurality of polynucleotides to form a combinatorial library of nucleic acids, the step being such that at least about 70% of the predicted diversity is represented. Methods for generating combinatorial libraries of nucleic acids are provided herein, the combinatorial library being an unsaturated combinatorial library. Methods for generating combinatorial libraries of nucleic acids are provided herein, the combinatorial library being a saturated combinatorial library. Methods for generating combinatorial libraries of nucleic acids are provided herein, wherein at least 10,000 polynucleotides are synthesized. Methods for generating combinatorial libraries of nucleic acids are provided herein, wherein the total number of polynucleotides for generating an unsaturated combinatorial library is less than at least 25% of the total number of polynucleotides for generating a saturated combinatorial library.Methods for generating combinatorial libraries of nucleic acids are provided herein, wherein at least 80% of the variants are of appropriate size. Methods for generating combinatorial libraries of nucleic acids are provided herein, wherein the variant combinatorial library encodes a first reference sequence or a second reference sequence. Methods for generating combinatorial libraries of nucleic acids are provided herein, wherein the combinatorial library at the time of translation encodes a protein library. Methods for generating combinatorial libraries of nucleic acids are provided herein, wherein the nucleic acids of the combinatorial library are inserted into a vector. Further, methods for generating combinatorial libraries of nucleic acids are provided herein, the method further comprising performing PCR mutagenesis of the nucleic acids using the combinatorial library as a primer for the PCR mutagenesis reaction. Methods for generating combinatorial libraries of nucleic acids are provided herein, wherein the combinatorial library encodes a sequence for a variant gene or a fragment thereof. Methods for generating combinatorial libraries of nucleic acids are provided herein, wherein the combinatorial library encodes at least a portion of an antibody, an enzyme, or a peptide. Methods for generating combinatorial libraries of nucleic acids are provided herein, wherein the combinatorial library encodes at least a portion of a variable region or a constant region of an antibody. Methods for generating combinatorial libraries of nucleic acids are provided herein, wherein the combinatorial library encodes at least one CDR region of an antibody. Methods for generating combinatorial libraries of nucleic acids are provided herein, wherein the combinatorial library encodes CDR1, CDR2, and CDR3 on the heavy chain of an antibody and CDR1, CDR2, and CDR3 on the light chain. Methods for generating combinatorial libraries of nucleic acids are provided herein, wherein the combinatorial library encodes a guide RNA (gRNA).Methods for generating combinatorial libraries of nucleic acids are provided herein, where the combinatorial libraries have a total error rate of less than 1 in 1000 bases compared to a predefined sequence. Methods for generating combinatorial libraries of nucleic acids are provided herein, where the structure is a solid support, a gel, or a bead, and the solid support is a plate or a column.

[0014] Methods for synthesizing mutant nucleic acid libraries are provided herein, the methods comprising: (a) providing a predefined sequence encoding a plurality of non-identical polynucleotides, the non-identical polynucleotides encoding a plurality of codons having mutant sequences as compared to a single reference sequence; (b) selecting distribution values for codons at predefined positions in a predefined nucleic acid reference sequence; (c) providing machine instructions for randomly generating a set of nucleic acids, the set of nucleic acids being less than the amount of nucleic acids required to generate a saturated codon mutant library; and (d) synthesizing a nucleic acid library of a predefined distribution, representing at least about 70% of the predicted diversity. Methods for synthesizing mutant nucleic acid libraries are provided herein, wherein at least 80% of the mutants are of appropriate size. Methods for synthesizing mutant nucleic acid libraries are provided herein, wherein the combinatorial library upon translation encodes a protein library. Methods for synthesizing mutant nucleic acid libraries are provided herein, wherein the nucleic acids of the combinatorial library are inserted into vectors. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, the methods further comprising performing PCR mutagenesis of nucleic acids using the combinatorial library as primers for a PCR mutagenesis reaction. Methods for synthesizing mutant nucleic acid libraries are provided herein, wherein codon assignment is used to determine each codon of a plurality of codons having mutant sequences. Methods for synthesizing mutant nucleic acid libraries are provided herein, wherein codon assignment is based on the frequency of codon sequences in an organism. Methods for synthesizing mutant nucleic acid libraries are provided herein, wherein the organism is at least one of an animal, a plant, a fungus, a protist, an archaeon, and a bacterium. Methods for synthesizing mutant nucleic acid libraries are provided herein, wherein codon assignment is based on the diversity of codon sequences.

[0015] Methods for synthesizing mutant nucleic acid libraries are provided herein, the methods comprising: (a) providing a predefined sequence encoding a plurality of non-identical polynucleotides, wherein the non-identical polynucleotides encode codons having mutant sequences as compared to a single reference sequence; (b) dividing the plurality of non-identical polynucleotides into a 5′ fragment of the non-identical polynucleotide and a 3′ fragment of the non-identical polynucleotide; (c) selecting a distribution value for a codon at a predefined position in a predefined nucleic acid reference sequence; (d) providing a machine instruction to randomly generate a set of nucleic acids, wherein the set of nucleic acids is less than the amount of nucleic acids required to generate a saturated nucleic acid library; (e) synthesizing a 5′ fragment of the non-identical polynucleotide and a 3′ fragment of the non-identical polynucleotide; (f) mixing a 5′ fragment of the non-identical polynucleotide and a 3′ fragment of the non-identical polynucleotide to form a mutant nucleic acid library, wherein at least about 70% of the predicted diversity is represented. Methods for synthesizing mutant nucleic acid libraries are provided herein, wherein at least 10,000 non-identical polynucleotides are synthesized. Methods for synthesizing mutant nucleic acid libraries are provided herein, wherein at least 80% of the mutants are of appropriate size. Methods for synthesizing mutant nucleic acid libraries are provided herein, wherein the plurality of non-identical polynucleotides are divided into at least one 5′ fragment and at least one 3′ fragment, each having more than one. Methods for synthesizing mutant nucleic acid libraries are provided herein, wherein the combinatorial library upon translation encodes a protein library. Methods for synthesizing mutant nucleic acid libraries are provided herein, wherein the nucleic acids of the combinatorial library are inserted into a vector. Further, methods for synthesizing mutant nucleic acid libraries are provided herein, the methods further comprising performing PCR mutagenesis of the nucleic acids using the combinatorial library as a primer for a PCR mutagenesis reaction. Methods for synthesizing mutant nucleic acid libraries are provided herein, further comprising identifying mutant sequences having enhanced or reduced activity.Methods for synthesizing mutant nucleic acid libraries are provided herein, and the activity is cell activity. Methods for synthesizing mutant nucleic acid libraries are provided herein, and cell activity includes proliferation, growth, adhesion, death, migration, energy production, oxygen utilization, metabolic activity, cell signaling, response to free radical damage, or any combination thereof. Methods for synthesizing mutant nucleic acid libraries are provided herein, and the mutant nucleic acid library encodes a sequence for a mutant gene or a fragment thereof. Methods for synthesizing mutant nucleic acid libraries are provided herein, and the mutant nucleic acid library encodes at least a portion of an antibody, an enzyme, or a peptide. Methods for synthesizing mutant nucleic acid libraries are provided herein, and the mutant nucleic acid library encodes a guide RNA (gRNA). Methods for synthesizing mutant nucleic acid libraries are provided herein, and the mutant nucleic acid library encodes at least a portion of the variable or constant region of an antibody. Methods for synthesizing mutant nucleic acid libraries are provided herein, and the mutant nucleic acid library encodes at least one CDR region of an antibody. Methods for synthesizing mutant nucleic acid libraries are provided herein, and the nucleic acid library encodes CDR1, CDR2, and CDR3 on the heavy chain of an antibody and CDR1, CDR2, and CDR3 on the light chain. Methods for synthesizing mutant nucleic acid libraries are provided herein, and the nucleic acid library has a total error rate of less than 1 in 1000 bases compared to a predefined sequence for a plurality of non-identical polynucleotides. Methods for synthesizing mutant nucleic acid libraries are provided herein, and many of the various sequences synthesized in the nucleic acid library range from about 50 to about 1,000,000. Methods for synthesizing mutant nucleic acid libraries are provided herein, and many of the various sequences synthesized in the nucleic acid library range from about 500 to about 25,000. Methods for synthesizing mutant nucleic acid libraries are provided herein, and many of the various sequences synthesized in the nucleic acid library range from about 1000 to about 15,000.Furthermore, methods for synthesizing mutant nucleic acid libraries are provided herein, which further include performing PCR mutagenesis of nucleic acids using a combinatorial library as primers for the PCR mutagenesis reaction. Methods for synthesizing mutant nucleic acid libraries are provided herein, and the codon assignment is used to determine codons having mutant sequences. Methods for synthesizing mutant nucleic acid libraries are provided herein, and the codon assignment is based on the frequency of codon sequences in an organism. Methods for synthesizing mutant nucleic acid libraries are provided herein, and the organism is at least one of an animal, a plant, a fungus, a protist, an archaeon, and a bacterium. Methods for synthesizing mutant nucleic acid libraries are provided herein, and the codon assignment is based on the diversity of codon sequences.

[0016] Methods for synthesizing mutant nucleic acid libraries are provided herein, the method comprising: (a) designing a predefined sequence encoding a plurality of non-identical polynucleotides, wherein the non-identical polynucleotides encode a plurality of codons having mutant sequences as compared to a single reference sequence; (b) synthesizing a plurality of non-identical polynucleotides to generate a mutant nucleic acid library, wherein at least about 70% of the predicted diversity is represented; (c) expressing the mutant nucleic acid library; and (d) evaluating an activity associated with the mutant nucleic acid library. Incorporation by reference

[0017] All publications, patents, and patent applications mentioned herein are hereby incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 5E

Figure 5F

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26A

Figure 26B

Figure 26C

Figure 26D

Figure 27

Figure 28A

Figure 28B

[0019] Unless otherwise specified, the present disclosure employs conventional molecular biology techniques within the scope of the art. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0020] **Definitions**

[0021] Throughout the present disclosure, numerical features are presented in range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as a rigid limitation on the scope of any embodiment. Accordingly, a range description should be considered to specifically disclose all possible sub-ranges and individual numerical values within that range down to the second decimal place of the unit of the lower limit, unless the context clearly dictates otherwise. For example, a range description such as 1 to 6 should be considered to specifically disclose sub-ranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, and individual numerical values within that range such as 1.1, 2, 2.3, 5, and 5.9. This applies regardless of the breadth of the range. The upper and lower limits of these intervening ranges may be independently included within a smaller range and are also included within the present invention according to any specifically excluded limits defined within the given range. If the given range includes one or both of the upper and lower limits, ranges excluding either or both of these included upper and lower limits are also included within the present invention, unless the context clearly indicates otherwise.

[0022] The terms used in this specification are for the purpose of describing only specific embodiments and are not intended to limit any embodiment. As used in this specification, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprising" and / or "comprises", when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used in this specification, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0023] Unless otherwise defined or not apparent from the context, as used in this specification, the term "about" in relation to a number or range of numbers means the stated number plus or minus 10% thereof, or for an enumerated value of a range, 10% below the enumerated lower limit and 10% above the enumerated upper limit.

[0024] As used in this specification, the terms "preselected sequence", "predefined sequence", or "predetermined sequence" are used interchangeably. The terms mean that the polymer sequence is known and is selected prior to the synthesis or assembly of the polymer. In particular, various aspects of the present invention are described herein primarily with respect to the preparation of nucleic acid molecules, and the sequences of oligonucleotides or polynucleotides are known and are selected prior to the synthesis or assembly of the nucleic acid molecules.

[0025] Methods and compositions for the production of polynucleotides that are synthesized (i.e., synthesized de novo or chemically synthesized) are provided herein. The terms oligonucleotide, oligo, and polynucleotide are defined as synonymous throughout. The libraries of synthesized polynucleotides described herein may collectively include multiple polynucleotides encoding one or more genes or gene fragments. In some examples, the polynucleotide library includes coding or non-coding sequences. In some examples, the polynucleotide library encodes multiple cDNA sequences. The reference gene sequences based on the cDNA sequences may include introns, but the cDNA sequences exclude introns. The polynucleotides described herein may encode genes or gene fragments of an organism. Exemplary organisms include, but are not limited to, prokaryotes (e.g., bacteria) and eukaryotes (e.g., mouse, rabbit, human, and non-human primates). In some examples, the polynucleotide library includes one or more polynucleotides, and each of the one or more polynucleotides encodes the sequences of multiple exons. Each polynucleotide within the libraries described herein may encode a different sequence (i.e., a non-identical sequence). In some examples, each polynucleotide within the libraries described herein includes at least one portion that is complementary to the sequence of another polynucleotide within the library. The polynucleotide sequences described herein may include DNA or RNA, unless otherwise specified.

[0026] Methods and compositions for the production of synthetic (i.e., de novo synthesized) genes are provided herein. Libraries of synthetic genes are constructed by various methods described in detail elsewhere herein, such as PCA, non-PCA gene assembly methods, or hierarchical gene assembly, by combining two or more double-stranded polynucleotides (“stitching”) to generate larger DNA units (i.e., chassis). Libraries of large constructs may contain polynucleotides that are at least 1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 400, 500 kilobases in length or longer. The large constructs are ligatable with an independently selected upper limit of about 5000, 10000, 20000, or 50000 base pairs. Any number of syntheses of polypeptide-segments encoding nucleotide sequences may include sequences encoding non-ribosomal peptides (NRPs), sequences encoding non-ribosomal peptide synthetase (NRPS) modules and synthetic variants, polypeptide segments of other modular proteins such as antibodies, polypeptide segments from other protein families, non-coding DNA or RNA such as regulatory sequences (e.g., promoters, transcription factors, enhancers, siRNA, shRNA, RNAi, miRNA, small nucleolar RNAs derived from microRNAs, or any functional or structural DNA or RNA unit of interest).The following are non-limiting examples of polynucleotides: coding or non-coding regions of genes or gene fragments, intergenic DNA, loci (loci) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, small interfering RNA (siRNA), small hairpin RNA (shRNA), microRNA (miRNA), small nuclear RNA, ribozymes, complementary DNA (cDNA) which is a DNA representation of mRNA usually obtained by reverse transcription or amplification of messenger RNA (mRNA); synthetically or amplified DNA molecules, genomic DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. The cDNA encoding a gene or gene fragment referred to herein may also contain at least one region encoding an exon sequence without intervening intron sequences as seen in the corresponding genomic sequence. Alternatively, the genomic sequence corresponding to the cDNA may first lack intron sequences.

[0027] Mutant library synthesis

[0028] The methods described herein provide for the synthesis of a library of nucleic acids, each encoding a defined variant of at least one predefined reference nucleic acid sequence. Optionally, the predefined reference sequence is a nucleic acid sequence encoding a protein, and the variant library includes sequences encoding mutations of at least one codon such that multiple different variants of a single residue in the subsequent protein encoded by the synthesized nucleic acids are generated by standard translation processes. Specific changes synthesized in the nucleic acid sequence can be introduced by incorporating nucleotide changes into overlapping or blunt-ended polynucleotide primers. Alternatively, a population of polynucleotides may collectively encode a long nucleic acid (e.g., a gene) and its variants. In this arrangement, the population of polynucleotides can be hybridized and subjected to standard molecular biology techniques to form the long nucleic acid (e.g., a gene) and its variants. When the long nucleic acid (e.g., a gene) and its variants are expressed in a cell, a variant protein library can be generated. Similarly, methods for the synthesis of variant libraries encoding RNA sequences (e.g., miRNA, shRNA, and mRNA) or DNA sequences (e.g., enhancer, promoter, UTR, and terminator regions) are provided herein. In some examples, the sequence is an exon sequence or a coding sequence. In some examples, the sequence does not include intron sequences. Further, downstream applications of variants selected from libraries synthesized using the methods described herein are provided. Downstream applications include, for example, the identification of variant nucleic acid or protein sequences with enhanced biochemical affinity, enzyme activity, changes in cellular activity, and biologically relevant functions for the treatment or prevention of medical conditions.

[0029] Combinatorial nucleic acid library

[0030] Methods are described herein for an efficient system for synthesizing a high-precision mutant nucleic acid library. Further provided herein are methods for synthesizing a combination-based mutant library. An advantageous feature of the methods provided herein is that the products and frequencies of the assembled nucleic acids in the combinatorial library can be accurately predicted, and the combinatorial library can be screened with an accurate understanding of the combinatorial products associated with enhancements related to biochemical activity or cellular activity, as well as the combinatorial products associated with negative or null results. Such a system is advantageous over modern methods (i.e., phage display) that do not permit an efficient means of gathering information regarding negative or null results. Another advantageous feature of the methods provided herein is that when a representative combinatorial library is designed and tested, it requires less material and cost compared to a fully saturated library, while also enabling the rapid generation of second and third generation libraries with sophisticated spotting criteria based on information gathered from screening the products of first generation combinatorial libraries.

[0031] Methods as described herein for the efficient and accurate synthesis of mutant nucleic acid libraries may result in a uniform and diverse library. Libraries generated using the methods described herein are not random. Libraries generated using the methods described herein result in the accurate introduction of each intended mutant at the desired frequency. Libraries generated using the methods described herein provide high accuracy by a reduction in the dropout rate of presentation and an improvement in the uniformity among the polynucleotides or species of long nucleic acids within each library. In addition, such high accuracy advantages at the polynucleotide synthesis level enable high accuracy at the functional level of downstream applications such as the evaluation of protein activity from translation products that introduce a defined pre-determined variance encoded at the codon level. In some examples, methods as described herein for the generation of accurate libraries enable an improved design of subsequent libraries. Such subsequent libraries may be more focused at the design stage as a result of information gathered regarding negative or null results from a first library. For example, a first mutant nucleic acid library synthesized using the methods described herein may be used to generate a mutant library of functional RNAs or proteins that are screened for a particular activity. Based on the observation of both positive and negative results associated with a precisely defined non-random library, design choices are then made for a second mutant library that is used in a further screening step to screen for and select species associated with the designated activity. This process can be repeated 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more times. Methods for library design, construction, screening, and iteration can be performed to identify enhanced species related to a single activity or multiple activities (e.g., binding affinity, stability, and expression).

[0032] When using in-silico library generation, the arrays can be known and not random. In some examples, the library contains at least or about 10 1 10 2 10 3 10 4 10 5 10 6 10 7 10 8 10 9 10 10 or 10 10 or more variants. In some examples, at least or about 10 1 10 2 10 3 10 4 10 5 10 6 10 7 10 8 10 9 or 10 10 For each variant in a library containing variants, the array is known. In some examples, the library includes the predicted diversity of the variants. In some examples, the diversity represented by the library is at least or about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or higher than 95% of the predicted diversity. In some examples, the diversity represented by the library is at least or about 70% of the predicted diversity. In some examples, the diversity represented by the library is at least or about 80% of the predicted diversity. In some examples, the diversity represented by the library is at least or about 90% of the predicted diversity. In some examples, the diversity represented by the library is at least or about 99% of the predicted diversity. As described herein, the term "predicted diversity" refers to the total theoretical diversity in a population containing all possible variants.

[0033] By generating a highly uniform and diverse library as described herein, where the sequences of the variants are known, an accurate understanding can be obtained of combinatorial products associated with enhanced or reduced activity and combinatorial products associated with negative or null results. Knowing the products associated with enhanced or reduced activity and such combinatorial products associated with negative or null results can also enable efficient use of the library in subsequent assays. For example, when performing a large screening, the variant sequences that result in enhanced or reduced activity are known. When conducting a subsequent screening, the sequences that resulted in negative or null results are excluded so that only the variant sequences that result in enhanced or reduced activity are screened.

[0034] In some examples, the enhanced or reduced activity is associated with cellular activity. Cellular activity includes, but is not limited to, proliferation, growth, adhesion, death, migration, energy production, oxygen utilization, metabolic activity, cell signaling, response to free radical damage, or any combination thereof.

[0035] In a first exemplary process, an unsaturated combinatorial library is generated. Generation of an unsaturated combinatorial library can reduce the number of synthesis steps. Referring to FIG. 1, a first population of nucleic acids (110) exhibits diversity at positions 1, 2, 3, and 4. A second population of nucleic acids (120) exhibits diversity at positions 5, 6, 7, and 8. The first population of nucleic acids (110) is combined with the second population of nucleic acids (120) to yield 16 combinations of nucleic acid fragments. The first population of nucleic acids (110) can be combined with the second population of nucleic acids (120) by blunt-end ligation. In some examples, the first and second populations are designed to have complementary overlapping sequences that include restriction enzyme recognition regions such that after cleavage of the nucleic acids in each population, the first and second populations can anneal to each other.

[0036] In some cases, the nucleic acid library is synthesized using two or more nucleic acid fragments. The nucleic acid library can be synthesized using at least two fragments, at least three fragments, at least four fragments, at least five fragments, or more. The length of each nucleic acid fragment or the average length of the synthesized nucleic acid can be at least or at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 300, 400, 500, 2000 nucleotides, or more. The length of each nucleic acid fragment or the average length of the synthesized nucleic acid can be at most about 2000, 500, 400, 300, 200, 150, 100, 50, 45, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10 nucleotides, or less. The length of each nucleic acid fragment or the average length of the synthesized nucleic acid is in the range of 10-2000, 10-500, 9-400, 11-300, 12-200, 13-150, 14-100, 15-50, 16-45, 17-40, 18-35, 19-25.

[0037] Various mixing processes and reagents such as ligation are known in the art and can be useful for performing the methods provided herein. Blunt-end ligation can be used to join fragments from one population of nucleic acids to fragments from a second population of nucleic acids. Ligases that can be included are, but are not limited to, E. coli ligase, T4 ligase, mammalian ligases (e.g., DNA ligase I, DNA ligase II, DNA ligase III, DNA ligase IV), thermostable ligases, and fast ligase. In some examples, the PCR extension overlap method is used to form longer nucleic acids by annealing and ligating two fragments. In such a configuration, the first fragment has a region complementary to the second fragment, and in the presence of DNA polymerase and amplification reagents such as dNTPs, buffer, and ATP, each fragment serves as a primer for another fragment for an amplification reaction that extends from the annealing position. In some examples, fragments from one population of nucleic acids are joined to fragments of a second population of nucleic acids by ligation after cleavage of a restriction enzyme recognition region. In some examples, the restriction enzyme generates an overhang that is then joined by a ligase. A 1:1 molar ratio of one nucleic acid fragment to another nucleic acid fragment can be used. Optionally, the molar ratio is at least 1:1, at least 1:2, at least 1:3, at least 1:4, or more. Alternatively, the ratio can be at least 2:1, at least 3:1, at least 4:1, or more. The total molar mass of the ligated nucleic acid fragments, or the molar mass of each of the nucleic acid fragments, can be at least or at least about 1, 10, 20, 30, 40, 50, 100, 250, 500, 750, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 25000, 50000, 75000, 100000 picomoles or more.

[0038] In some cases, the nucleic acid fragments generated by the methods described herein are blunt-ended prior to ligation. The nucleic acids can be blunt-ended using T4 DNA polymerase or the Klenow fragment. Alternatively, enzymes that directly generate blunt ends (e.g., Sma, Dpn I, Pvu II, Eco RV I) are used. In some examples, DNA endonucleases or DNA exonucleases are used to generate blunt ends.

[0039] In a second exemplary workflow, a saturated combinatorial library is generated. Referring to FIG. 2, a first population of nucleic acids (210) exhibits diversity at positions 1, 2, 3, and 4. A second population of nucleic acids (220) exhibits diversity at positions 5, 6, 7, and 8. As seen in FIG. 2, the population of nucleic acids (210) on the "left side" of the gene fragment has 4 4 diversities. The population of nucleic acids (220) on the "right side" of the gene fragment has 4 4 diversities. Subsequently, the long gene fragment is combined with another fragment having the diversity of the "right half" of the desired gene and synthesized using the diversity of the "left half" of the desired gene, resulting in a total of 4 8 diversities. The length of each nucleic acid fragment or the average length of the nucleic acids synthesized can be at least or at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 300, 400, 500, 2000 nucleotides, or more. The length of each nucleic acid fragment or the average length of the nucleic acids synthesized can be at most about 2000, 500, 400, 300, 200, 150, 100, 50, 45, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10 nucleotides, or less. The length of each nucleic acid fragment or the average length of the nucleic acids synthesized is in the range of 10 - 2000, 10 - 500, 9 - 400, 11 - 300, 12 - 200, 13 - 150, 14 - 100, 15 - 50, 16 - 45, 17 - 40, 18 - 35, 19 - 25.

[0040] The resulting nucleic acid is demonstrable. In some cases, the nucleic acid is demonstrated by sequencing. In some examples, the nucleic acid is demonstrated by high-throughput sequencing such as next-generation sequencing. Sequencing of the sequencing library can be performed using any suitable sequencing technology including single molecule real-time (SMRT) sequencing, polony sequencing, ligation sequencing, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination reaction (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis.

[0041] Methods are provided herein for the synthesis of nucleic acid libraries that are unsaturated or saturated in terms of the degree of variance, and the methods are highly accurate. In some examples, about 70% of the nucleic acid has no insertions or deletions. In some examples, at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more than 99% of the nucleic acid has no insertions and deletions. In some examples, about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more than 99% of the nucleic acid has no insertions and deletions. In some examples, more than about 90% of the nucleic acid has no insertions and deletions. In some instances, at least 80% of the nucleic acid is error-free. In some examples, at least about 70%, 75%, 80%, 85%, 90%, 95%, or more than 99% of the nucleic acid is error-free.

[0042] Methods are provided herein for the synthesis of nucleic acid libraries that are unsaturated or saturated in terms of the degree of dispersion, and the methods are highly accurate. In some examples, more than 80% of the nucleic acids in the de novo synthesized nucleic acid libraries described herein are represented within at least about 1.5X of the average representation of the entire library after amplification. In some examples, more than 80% of the nucleic acids in the de novo synthesized nucleic acid libraries described herein are represented within at least about 1.5X, 2X, 3X, 3.5X, or 4X of the average representation of the entire library after amplification. In some examples, more than 90% of the nucleic acids in the de novo synthesized nucleic acid libraries described herein are represented within at least about 1.5X of the average representation of the entire library after amplification. In some examples, more than 90% of the nucleic acids in the de novo synthesized nucleic acid libraries described herein are represented within at least about 1.5X, 2X, 3X, 3.5X, or 4X of the average representation of the entire library after amplification. In some examples, more than 80% of the nucleic acids in the de novo synthesized nucleic acid libraries described herein are represented within at least about 2X of the average representation of the entire library after amplification. In some examples, more than 80% of the nucleic acids in the de novo synthesized nucleic acid libraries described herein are represented within at least about 2X of the average representation of the entire library after amplification.

[0043] Generation of Representative Nucleic Acid Libraries

[0044] Methods are described herein for synthesizing nucleic acid libraries having a preselected distribution of variant codon coding regions. Further, such libraries are undersaturated with respect to the preselected distribution, but may provide insight into representative distributions. Further provided herein are methods for generating nucleic acids that, once translated, result in a preselected distribution of amino acids at specific positions. By generating a random sample of the preselected distribution, an undersaturated nucleic acid library is designed to have a representative distribution close to the preselected population distribution. Nucleic acid libraries as described herein with a representative distribution close to the preselected population distribution may further include the accurate introduction of each intended variant of the desired preselected distribution.

[0045] The computational methods described herein include, but are not limited to, random sampling. In a first process, for a preselected distribution of codon variance at each position, a cumulative distribution value for each position is calculated. In some examples, the cumulative distribution values are mapped to probabilities of about 0.0 - 1.0. For a population of nucleic acids, the cumulative distribution values result in the determination of the likelihood of a particular position codon variant. For example, the number of times a codon variant appears at each position in a population of nucleic acids can be summed and the proportion of amino acids appearing at each position can then be determined. The proportion in the sample population of nucleic acids is then compared to the preselected distribution. If the number of nucleic acids in a population is sufficient, a distribution of samples that matches the preselected distribution values is generated. In some examples, the sampling implemented applies uniform random sampling and is in the form of Monte Carlo sampling.

[0046] In some examples, nucleic acid libraries designed and synthesized to have a preselected distribution encode about 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, or more than 60% of non-identical nucleic acids as compared to a saturated nucleic acid library. In some examples, nucleic acid libraries designed and synthesized to have a preselected distribution encode at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, or more than 60% of non-identical nucleic acids as compared to a saturated nucleic acid library.

[0047] In some examples, nucleic acid libraries designed and synthesized to have a preselected distribution encode about 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, or more than 60% of non-identical nucleic acids as compared to a large nucleic acid library. In some examples, nucleic acid libraries designed and synthesized to have a preselected distribution encode at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, or more than 60% of non-identical nucleic acids as compared to a large nucleic acid library.

[0048] In some examples, the number of designed and synthesized nucleic acids in a representative subpopulation from a larger variant nucleic acid library ranges from about 50 - 100000, 100 - 75000, 250 - 50000, 500 - 25000, and 1000 - 15000, 2000 - 10000, and 4000 - 8000 sequences. In some examples, the population of nucleic acids is 500 sequences. In some examples, the population of nucleic acids is 5000, 10000, or 15000 sequences. In some examples, the population of nucleic acids has at least 50, 100, 150, 500, 1000, 2000, 5000, 10000, 20000, 50000, 100000, 200000, 400000, 800000, 1000000, or more different sequences. In some examples, each population of nucleic acids is at most 50, 100, 500, 1000, 2000, 5000, 10000, 20000, 50000, 100000, 200000, 400000, 800000, or 1000000.

[0049] In some examples, the synthesis of nucleic acid libraries by combinatorial methods to reach a preselected distribution of variant codon coding regions represents from 70% to 99% of the predicted diversity. In some examples, the synthesis of nucleic acid libraries by combinatorial methods to reach a preselected distribution of variant codon coding regions represents at least 70% of the predicted diversity. In some examples, the synthesis of nucleic acid libraries by combinatorial methods to reach a preselected distribution of variant codon coding regions represents from 70% to 75%, from 70% to 80%, from 70% to 85%, from 70% to 90%, from 70% to 95%, from 70% to 97%, from 70% to 99%, from 75% to 80%, from 75% to 85%, from 75% to 90%, from 75% to 95%, from 75% to 97%, from 75% to 99%, from 80% to 85%, from 80% to 90%, from 80% to 95%, from 80% to 97%, from 80% to 99%, from 85% to 90%, from 85% to 95%, from 85% to 97%, from 85% to 99%, from 90% to 95%, from 90% to 97%, from 90% to 99%, from 95% to 97%, from 95% to 99%, or from 97% to 99% of the predicted diversity. In some examples, the represented diversity of a synthesized representative population of nucleic acids is at least or about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 95% or more of the predicted diversity. In some examples, the represented diversity of a synthesized representative population of nucleic acids is 99% of the predicted diversity.

[0050] Generation of Representative Nucleic Acid Libraries Using Combinatorial Methods

[0051] Methods are provided herein for the synthesis of nucleic acid libraries by combinatorial methods to reach a preselected distribution of variant codon coding regions. In some examples, a reference sequence that serves as a template for variants to synthesize a population of nucleic acids is divided such that a first portion serves as a reference sequence for a first variant population of nucleic acids and a second portion serves as a reference sequence for a second variant population of nucleic acids.

[0052] In some examples, the random sampling method as described herein is used to generate a representative mutant distribution of a portion from a larger mutant library. A first representative population of nucleic acids representing mutants for a first portion of the complete reference sequence and a second representative population of nucleic acids representing mutants for a second portion of the complete reference sequence are synthesized and then combined by ligation such as blunt-end ligation or by biochemical techniques known in the art. In some cases, the resulting nucleic acid library is saturated. In some cases, the resulting nucleic acid library is non-saturated.

[0053] In some cases, the nucleic acid library is synthesized using two or more mutant nucleic acid populations that, when combined, result in the desired longer nucleic acid mutant library. The nucleic acid library can be synthesized using at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 populations, each encoding a different region of the reference nucleic acid. In some examples, each nucleic acid population ranges from about 50 - 100000, 100 - 75000, 250 - 50000, 500 - 25000, and 1000 - 15000, 2000 - 10000, and 4000 - 8000 sequences. In some examples, each nucleic acid population is about 500, 1000, 5000, 10000, or 15000 or more sequences. In some examples, each nucleic acid population is at least 50, 100, 150, 500, 1000, 2000, 5000, 10000, 20000, 50000, 100000, 200000, 400000, 800000, 1000000, or more. In some examples, each nucleic acid population is at most 50, 100, 500, 1000, 2000, 5000, 10000, 20000, 50000, 100000, 200000, 400000, 800000, and 1000000.

[0054] In some examples, the synthesis of a nucleic acid library by combinatorial methods to reach a preselected distribution of variant codon coding regions represents from 70% to 99% of the predicted diversity. In some examples, the synthesis of a nucleic acid library by combinatorial methods to reach a preselected distribution of variant codon coding regions represents at least 70% of the predicted diversity. In some examples, the synthesis of a nucleic acid library by combinatorial methods to reach a preselected distribution of variant codon coding regions represents from 70% to 75%, from 70% to 80%, from 70% to 85%, from 70% to 90%, from 70% to 95%, from 70% to 97%, from 70% to 99%, from 75% to 80%, from 75% to 85%, from 75% to 90%, from 75% to 95%, from 75% to 97%, from 75% to 99%, from 80% to 85%, from 80% to 90%, from 80% to 95%, from 80% to 97%, from 80% to 99%, from 85% to 90%, from 85% to 95%, from 85% to 97%, from 85% to 99%, from 90% to 95%, from 90% to 97%, from 90% to 99%, from 95% to 97%, from 95% to 99%, or from 97% to 99% of the predicted diversity. In some examples, the synthesis of a nucleic acid library by combinatorial methods to reach a preselected distribution of variant codon coding regions is at least or about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 95% or more of the predicted diversity. In some examples, the represented diversity of a synthesized representative population of nucleic acids is 99% of the predicted diversity.

[0055] Synthesis and subsequent PCR mutagenesis

[0056] Nucleic acid libraries (e.g., saturated or unsaturated) generated by the combinatorial methods described herein can be used in PCR mutagenesis methods. In some cases, a representative nucleic acid library having a preselected distribution is used in the PCR mutagenesis method. In this workflow, multiple polynucleotides are synthesized, each polynucleotide encoding a defined sequence that is a defined variant of a reference polynucleotide sequence. Referring to the figure which is a typical workflow depicted in FIGS. 3A-3D, the polynucleotides are generated on a surface. FIG. 3A depicts an enlarged view of a single cluster of a surface having 121 loci. Each nucleic acid depicted in FIG. 3B is a primer that can be used for amplification from a reference nucleic acid sequence to generate a library of mutant long nucleic acids (FIG. 3C). The library of mutant long nucleic acids is then optionally subjected to transcription and / or translation to generate a mutant RNA or protein library (FIG. 3D). In this typical figure, an apparatus having a substantially planar surface used for de novo synthesis of polynucleotides is depicted (FIG. 3A). In some examples, the apparatus includes clusters of loci, each locus being a site for polynucleotide extension. In some examples, a single cluster contains all of the polynucleotide variants required to generate a desired library of variant sequences. In an alternative arrangement, the plate includes regions of loci that are not separated into clusters.

[0057] Methods are provided herein for the synthesis of polynucleotides within a cluster (e.g., as seen in FIG. 3) and subsequent amplification of the polynucleotides within a single cluster. Such configurations result in improved nucleic acid presentation as compared to the amplification of non-identical polynucleotides across a plate without a clustered configuration. In some examples, the amplification of polynucleotides synthesized on the surface of loci within a cluster overcomes a negative effect on presentation by repeated synthesis of a large polynucleotide population having polynucleotides with a high GC content. In some examples, the clusters described herein include about 50 - 1000, 75 - 900, 100 - 800, 125 - 700, 150 - 600, 200 - 500, 50 - 500, or 300 - 400 separate loci. In some examples, the loci are spots, wells, microwells, channels, or posts. In some examples, each cluster has at least 1X, 2X, 3X, 4X, 5X, 6X, 7X, 8X, 9X, 10X, or more excess of another feature that supports an extension of a polynucleotide having an identical sequence. In some examples, an excess of 1X means having no polynucleotides using the identical sequence.

[0058] The de novo synthesized polynucleotide libraries described herein may contain a plurality of polynucleotides, each having at least one variant sequence at a first position, position “X”, and each variant polynucleotide is used as a primer in the first round of PCR to generate a first extension product. In this example, position “x” in the first polynucleotide (420) encodes a variant codon sequence, i.e., one of 19 possible variants from a reference sequence. See A of FIG. 4. A second polynucleotide (425) containing a sequence overlapping the sequence of the first polynucleotide is used as a primer in another round of PCR to generate a second extension product. Further, external primers (415), (430) may be used for amplification of fragments from a long nucleic acid sequence. The resulting amplification products are fragments of the long nucleic acid sequences (435), (440). See B of FIG. 4. Subsequently, the fragments of the long nucleic acid sequences (435), (440) are hybridized and subjected to an extension reaction to form variants of the long nucleic acid (445). See C of FIG. 4. The overlapping ends of the first and second extension products may serve as primers in the second round of PCR, thereby generating a third extension product (FIG. 4D) containing variants. To increase the yield, variants of the long nucleic acid are amplified in a reaction containing DNA polymerase, amplification reagents, external primers (415), (430). In some examples, the second polynucleotide includes, but is not limited to, sequences adjacent to the variant site. In an alternative arrangement, a first polynucleotide having a region overlapping with the second polynucleotide is generated. In this scenario, the first nucleic acid is synthesized with mutations at a single codon for up to 19 variants. The second nucleic acid does not contain variant sequences. Optionally, the first population includes the first polynucleotide variant and additional polynucleotides encoding variants at different codon sites. Alternatively, the first and second polynucleotides may be designed for blunt-end ligation.

[0059] In alternative mutagenesis, the PCR method is depicted in FIGS. 5A-5F. In such a process, a template nucleic acid molecule (500) comprising first and second strands (505), (510) is amplified in a PCR reaction comprising a first primer (515) and a second primer (520) (FIG. 5A). The amplification reaction includes uracil as a nucleotide reagent. An extension product (525) labeled with uracil is generated (FIG. 5B), optionally purified, and serves as a template for a subsequent PCR reaction using a first polynucleotide (535) and a plurality of second polynucleotides (530) to generate first extension products (540 and 545) (FIGS. 5C-5D). In this process, the plurality of polynucleotides (530) includes polynucleotides encoding mutant sequences (depicted as X, Y, and Z in FIG. 5C). The uracil-labeled template nucleic acid is digested with an excision reagent specific for uracil (e.g., USER digest commercially available from New England Biolabs). Mutants (535) and various codons (530) with mutants X, Y, and Z are added, and a limited PCR step is performed to generate FIG. 5D. After the uracil-containing template is digested, the overlapping ends of the extension products serve to stimulate the PCR reaction, and the first extension products (540 and 545) are combined with a first external primer (550) and a second external primer (555) to act as primers, thereby generating a library of nucleic acid molecules (560) containing a plurality of mutants X, Y, and Z at the mutant site of FIG. 5F.

[0060] De novo synthesis of a population with mutant and non-mutant portions of long nucleic acids

[0061] Nucleic acid libraries generated by the combinatorial methods described herein (e.g., saturated or unsaturated) can be used for de novo synthesis of multiple fragments of long nucleic acids, with at least one of the fragments being synthesized in multiple versions, each version being a different variant sequence. Optionally, a representative nucleic acid library having a preselected distribution is used for de novo synthesis, with at least one of the fragments being synthesized in multiple versions, each version being a different variant sequence. In this arrangement, all of the fragments required to assemble a library of variant long-distance nucleic acids are de novo synthesized. The synthesized fragments may have overlapping sequences such that after synthesis, the fragment library is subjected to hybridization. After hybridization, an extension reaction may be performed to fill any complementary gaps.

[0062] Alternatively, the synthesized fragments may be amplified with primers and then subjected to either blunt-end ligation or overlapping hybridization. In some examples, the device includes a cluster of loci, each locus being a site for polynucleotide extension. In some examples, a single cluster includes all polynucleotide variants of a pre-determined long nucleic acid and other fragment sequences to generate a desired library of variant nucleic acid sequences. The cluster may include about 50 - 500 loci. In some arrangements, the cluster includes more than 500 loci.

[0063] Each individual polynucleotide in the first polynucleotide population may be generated at a separate individually addressable locus of the cluster. One polynucleotide variant may be represented by multiple individually addressable loci. Each variant in the first polynucleotide population may be represented 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times. In some examples, each variant in the first polynucleotide population is represented at 3 or fewer loci. In some examples, each variant in the first polynucleotide population is represented at 2 loci. In some examples, each variant in the first polynucleotide population is represented at only 1 locus.

[0064] Methods are provided herein for generating nucleic acid libraries with reduced redundancy. In some examples, variant nucleic acids may be generated without the need to synthesize the variant nucleic acid more than once to obtain the desired variant nucleic acid. In some examples, the disclosure provides methods for generating variant nucleic acids without the need to synthesize the variant nucleic acid 1, 2, 3, 4, 5, more than 5, 6, 7, 8, 9, 10, or more times to generate the desired variant nucleic acid.

[0065] Variant nucleic acids may be generated without the need to synthesize the variant nucleic acid at separate sites more than 1 to obtain the desired variant nucleic acid. The disclosure provides methods for generating variant nucleic acids without the need to synthesize the variant nucleic acid at 1, 2, 3, 4, 5, 6, 7, 8, 9, or more than 10 sites. In some examples, the nucleic acid is synthesized at at most 6, 5, 4, 3, 2, or 1 separate site. The same nucleic acid may be synthesized at 1, 2, or 3 separate loci on the surface.

[0066] In some examples, the amount of locus representing a single variant nucleic acid depends on the amount of nucleic acid material required for downstream processing (e.g., amplification reaction or cell assay). In some examples, the amount of locus representing a single variant nucleic acid depends on the available loci in a single cluster.

[0067] Methods are provided herein for generating a library of nucleic acids that contain variant nucleic acids that differ at multiple sites in a reference nucleic acid. In such cases, each variant library is generated at individually addressable loci within a cluster of loci. It will be appreciated that the number of variant sites represented by the nucleic acid library is determined by the number of individually addressable loci in the cluster and the number of desired variants at each site. In some examples, each cluster contains from about 50 to 500 loci. In some examples, each cluster contains 100 to 150 loci.

[0068] In a typical arrangement, 19 variants are represented at variant sites corresponding to codons encoding each of 19 possible variant amino acids. In another typical case, 61 variants are represented at variant sites corresponding to triplets encoding each of 19 possible variant amino acids. In a non-limiting example, the cluster contains 121 individually addressable loci. In this example, the nucleic acid population contains 6 replicates (each of single-site variants (6 replicates x 1 variant site x 19 variants = 114 loci)), 3 replicates (each of double-site variants (3 replicates x 2 variant sites x 19 variants = 114 loci), or 2 replicates (each of triple-site variants (2 replicates x 3 variant sites x 19 variants = 114 loci). In some examples, the nucleic acid population contains variants at 4, 5, 6, or more than 6 variant sites.

[0069] Methods and compositions for the production of synthetic (i.e., de novo synthesized or chemically synthesized) nucleic acids are provided herein. The libraries of synthetic nucleic acids described herein may collectively contain multiple nucleic acids encoding one or more genes or gene fragments. In some examples, the nucleic acid library contains coding or non-coding sequences. In some examples, the nucleic acid library encodes multiple cDNA sequences. In some examples, the nucleic acid library contains one or more nucleic acids, each of which encodes the sequences of multiple exons. Each nucleic acid within the libraries described herein may encode a different sequence (i.e., a non-identical sequence). In some examples, each nucleic acid within the libraries described herein contains at least one portion that is complementary to the sequence of another nucleic acid within the library. The nucleic acid sequences described herein may include DNA or RNA, unless otherwise specified.

[0070] Methods and compositions for the production of synthetic (i.e., de novo synthesized) genes are provided herein. Libraries containing synthetic genes are constructed by various methods described in detail elsewhere herein, such as PCA, non-PCA gene assembly methods, or hierarchical gene assembly, combining two or more double-stranded nucleic acids ( "stitching") to generate larger DNA units (i.e., chassis). Libraries of larger constructs may contain nucleic acids that are at least 1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 400, 500 kb in length or greater. Larger constructs may also be ligated by an independently selected upper limit of about 5000, 10000, 20000, or 50000 base pairs. Any number of syntheses of polypeptide segments encoding nucleotide sequences may include sequences encoding non-ribosomal peptides (NRPs), sequences encoding non-ribosomal peptide synthetase (NRPS) modules and synthetic variants, polypeptide segments of other modular proteins such as antibodies, polypeptide segments from other protein families, non-coding DNA or RNA such as regulatory sequences (e.g., promoters, transcription factors, enhancers, siRNA, shRNA, RNAi, miRNA, small nucleolar RNA derived from microRNA, or any functional or structural DNA or RNA unit of interest).The following are non-limiting examples of nucleic acids: the coding or non-coding regions of genes or gene fragments, intergenic DNA, loci (loci) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, small interfering RNA (siRNA), small hairpin RNA (shRNA), microRNA (miRNA), small nuclear RNA, ribozymes, cDNA which is the DNA representation of mRNA usually obtained by reverse transcription or amplification of messenger RNA (mRNA); DNA molecules synthesized or generated by amplification, genomic DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. In the context of cDNA, the terms gene or gene fragment refer to a DNA nucleic acid sequence that includes at least one region encoding an exon sequence without intervening intron sequences.

[0071] In various embodiments, the methods and compositions described herein relate to libraries of genes. A gene library can include a plurality of subsegments. In one or more subsegments, the genes of the library can be covalently linked together. In one or more subsegments, the genes of the library encode components of a first metabolic pathway with one or more metabolic end products. In one or more subsegments, the genes of the library can be selected based on the production process of one or more target metabolic end products. One or more metabolic end products may include biofuels. In one or more subsegments, the genes of the library encode components of a second metabolic pathway with two or more metabolic end products. One or more end products of the first and second metabolic pathways can include one or more shared end products. In some cases, the first metabolic pathway includes an end product that is operated within the second metabolic pathway.

[0072] Variant nucleic acid libraries of organisms

[0073] The mutant nucleic acid libraries generated by the methods described herein may encode at least one gene of an organism. In some cases, the nucleic acid library encodes a single gene, pathway, or entire genome of an organism. In some examples, the mutant nucleic acid library encodes at least one of a gene (e.g., 1000 base pairs), a portion (e.g., 3 - 10 genes), a pathway (e.g., 10 - 100 genes), or a chassis (e.g., 100 - 1000 genes) of an organism. A non-limiting exemplary list of model organisms is provided in Table 1.

[0074] [Table 1]

[0075] Codon Variations

[0076] The mutant nucleic acid libraries described herein may contain a plurality of nucleic acids, each of which encodes a mutant codon sequence as compared to a reference nucleic acid sequence. In some examples, each nucleic acid of a first nucleic acid population contains a mutant at a single mutant site. In some examples, the first nucleic acid population contains multiple mutants at a single mutant site such that more than one mutant is present at the same mutant site. The first nucleic acid population may contain nucleic acids that collectively encode multiple codon mutants at the same mutant site. The first nucleic acid population may contain nucleic acids that collectively encode up to 19 or more codons at the same position. The first nucleic acid population may contain nucleic acids that collectively encode up to 60 mutant triplets at the same position, or alternatively, the first nucleic acid population may contain nucleic acids that collectively encode up to 61 different triplets of codons at the same position. Each mutant may encode a codon that results in a different amino acid during translation. Table 2 provides a list of each possible codon (and representative amino acid) for different sites.

[0077] [Table 2-1]

[0078]

Table 2-2

[0079] A mutant nucleic acid library is provided herein that includes a nucleic acid encoding a mutant codon sequence, where the mutant codon sequence is selected based on codon assignment. Exemplary codon assignments are found in Table 3, where the mutant codon sequence is first selected from left to right. In some examples, the codon assignment is based on the frequency of codons in an organism. Exemplary organisms include, but are not limited to, animals, plants, fungi, protists, archaea, or bacteria. For example, the codon assignment is based on E. coli or humans.

[0080]

Table 3-1

[0081]

Table 3-2

[0082] Provided herein is a mutant nucleic acid library comprising a nucleic acid encoding a mutant codon sequence as compared to a reference nucleic acid sequence, wherein the mutant codon sequence based on codon assignment is determined by various factors. In some examples, the mutant codon sequence is selected based on the complexity or diversity of the codon sequence. For example, a codon sequence containing three different nucleobases is selected instead of a codon sequence containing two different nucleobases or a codon sequence containing the same nucleobase. In some examples, the codon sequence is selected based on downstream applications. Downstream applications include, but are not limited to, minimizing the effect on expression levels after protein translation or improving the detection of mutant codon sequences by next-generation sequencing. Improving the detection of mutant codon sequences by next-generation sequencing may include avoiding homopolymers with a high error rate. In some examples, the codon sequence is selected unless it creates a site that causes disruption of a sequence such as a restriction enzyme site.

[0083] The codon sequences for mutant sites based on codon assignment as described herein may be randomized. In some examples, the codon sequences are not randomized. For example, for a single mutant library where one mutation is selected per peptide, the codon sequences are not randomized. In some examples, multiple mutant libraries include codon sequences that are randomized.

[0084] The nucleic acid population may include various nucleic acids that collectively encode up to 20 codon mutations at multiple positions. In such cases, each nucleic acid in the population contains codon mutations at more than one position within the same nucleic acid. In some examples, the nucleic acids in the population each contain codon mutations in 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more codons within a single nucleic acid. In some examples, the long nucleic acids of each variant contain codon mutations in 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more codons within a single long nucleic acid. In some examples, the variant nucleic acid population contains codon mutations in 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more codons within a single nucleic acid. In some examples, the variant nucleic acid population contains codon mutations in at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 275, 300, or more codons within a single long nucleic acid.

[0085] Provided herein is a process for generating a second nucleic acid population on a second cluster that includes a plurality of individually addressable loci. The second nucleic acid population may include a plurality of second nucleic acids that are constant for each codon position (i.e., encode the same amino acid at each position). The second nucleic acids may overlap at least a portion of the first nucleic acids. In some examples, the second nucleic acids do not include variant sites represented on the first nucleic acids. Alternatively, the second nucleic acid population may include a plurality of second nucleic acids that include at least one variant for one or more codon positions.

[0086] Provided herein is a method for synthesizing a library of nucleic acids in which a single population of nucleic acids containing variants at multiple codon positions is generated. The first nucleic acid population may be generated on a first cluster containing a plurality of individually addressable loci. In such cases, the first nucleic acid population contains variants at different codon positions. In some examples, the various sites are contiguous (i.e., encoding contiguous amino acids). For example, the first nucleic acid population contains variants at two contiguous codon positions encoding up to 19 variants at one position. In some examples, the first nucleic acid population contains variants at two contiguous codon positions encoding from about 1 to about 19 variants at one position. In some examples, about 38 nucleic acids are synthesized. The first nucleic acid population may contain various nucleic acids that collectively encode up to 19 codon variants at the same or additional variant sites. The first nucleic acid population may contain a plurality of first nucleic acids containing up to 19 variants at position x, up to 19 variants at position y, and up to 19 variants at position z. In such an arrangement, the variants each encode a different amino acid such that up to 19 amino acid variants are encoded at each of the various variant sites. In additional examples, a second nucleic acid population is generated on a second cluster containing a plurality of individually addressable loci. The second nucleic acid population may contain a plurality of second nucleic acids that are constant for each codon position (i.e., encoding the same amino acid at each position). The second nucleic acid may overlap at least a portion of the first nucleic acid. The second nucleic acid may not contain the variant sites represented on the first nucleic acid.

[0087] The mutant nucleic acid library generated by the process described herein results in the generation of a mutant protein library. In a first exemplary arrangement, the template nucleic acid encodes a sequence that, upon transcription and translation, results in a reference amino acid sequence having a number of codon positions indicated by a single circle (A of FIG. 6). Nucleic acid mutants of the template can be generated using the methods described herein. In some examples, a single mutant is present in the nucleic acid, resulting in a single amino acid sequence (B of FIG. 6). In some examples, more than one mutant is present in the nucleic acid, the mutants are separated by one or more codons, resulting in a protein with a spacing between the mutant residues (C of FIG. 6). In some examples, more than one mutant is present in the nucleic acid, the mutants are sequential and either adjacent or contiguous to each other, resulting in a run of mutants with a residue spacing (D of FIG. 6). In some examples, two runs of mutants are present in the nucleic acid, each run of mutants containing sequential, adjacent, or contiguous mutants (E of FIG. 6).

[0088] Methods for generating libraries of nucleic acid variants are provided herein, where each variant comprises a single-site codon variant. In one example, the template nucleic acid has many codon positions, and typical amino acid residues are indicated by circles using their respective one-letter code protein codons (Figure 7A). Figure 7B depicts a library of amino acid variants encoded by a library of variant nucleic acids, where each variant comprises a single-site variant, indicated by "X", located at a different one site. The variant at the first position has any codon for exchange with alanine, a second variant having any codon encoded by the library of variant nucleic acids for exchange with tryptophan, a third variant having any codon for exchange with isoleucine, a fourth variant having any codon for exchange with lysine, a fifth variant having any codon for exchange with arginine, a sixth variant having any codon for exchange with glutamic acid, and a seventh variant having any codon for exchange with glutamine. All, or fewer than all, codon variants are encoded by the variant nucleic acid library, and the resulting corresponding population of amino acid sequence variants is generated after protein expression (i.e., the standard cellular events of DNA transcription and subsequent translation and processing events).

[0089] In some arrangements, the library is generated at multiple sites of single-site variants. As depicted in Figure 8A, a wild-type template is provided. Figure 8B depicts the resulting amino acid sequence having two sites of single-site codon variants, where each codon variant encoding a different amino acid is indicated by a different pattern of circles.

[0090] A method for generating a library having a series of variants at a single position of multiple sites is provided herein. Each series of nucleic acids may have 1, 2, 3, 4, 5, or more variants. Each series of nucleic acids may have at least 1 variant. Each series of nucleic acids may have at least 2 variants. Each series of nucleic acids may have at least 3 variants. For example, a series of 5 nucleic acids may have 1 variant. A series of 5 nucleic acids may have 2 variants. A series of 5 nucleic acids may have 3 variants. A series of 5 nucleic acids may have 4 variants. For example, a series of 4 nucleic acids may have 1 variant. A series of 4 nucleic acids may have 2 variants. A series of 4 nucleic acids may have 3 variants. A series of 4 nucleic acids may have 4 variants.

[0091] In some examples, all of the variants at a single position may encode the same amino acid, for example, histidine. As shown in FIG. 9A, a reference amino acid sequence is provided. In this arrangement, a series of nucleic acids encodes multiple sites of variants at a single position and, upon expression, results in an amino acid sequence having all of the variants at a single position that encode histidine (FIG. 9B). In some embodiments, the variant library synthesized by the methods described herein does not encode more than 4 histidine residues in the resulting amino acid sequence.

[0092] In some examples, the variant library of nucleic acids generated by the methods described herein results in the expression of an amino acid sequence having another series of mutations. The template amino acid sequence is shown in FIG. 10A. A series of nucleic acids may contain only 1 variant codon in 2 series and, upon expression, results in the amino acid sequence shown in FIG. 10B. To show the mutations of amino acids at different positions in one series, the variants are shown in FIG. 10B by different patterns of circles.

[0093] Provided herein are methods and apparatuses for synthesizing nucleic acid libraries having one, two, three, or more codon variants, where the variants for each site are selectively controlled. The ratio of two amino acids for a single-site variant can be about 1:100, 1:50, 1:10, 1:5, 1:3, 1:2, 1:1. The ratio of three amino acids for a single-site variant can be about 1:1:100, 1:1:50, 1:1:20, 1:1:10, 1:1:5, 1:1:3, 1:1:2, 1:1:1, 1:10:10, 1:5:5, 1:3:3, or 1:2:2. Panel A of FIG. 11 shows the wild-type reference amino acid sequence encoded by the wild-type nucleic acid sequence. Panel B of FIG. 11 shows a library of amino acid variants, where each variant contains a run of the sequence (indicated by the patterned circles), and each position may have a specific ratio of amino acids in the resulting variant protein library. The resulting variant protein library is encoded by the variant nucleic acid library generated by the methods described herein. In this illustration, five positions are varied: the first position (1100) has a 50 / 50 K / R ratio; the second position (1110) has a 50 / 25 / 25 V / L / S ratio, the third position (1120) has a 50 / 25 / 25 Y / R / D ratio, the fourth position (1130) has an equal ratio for all 20 amino acids, and the fifth position (1140) has a 75 / 25 ratio for G / P. The ratios described herein are merely examples.

[0094] In some examples, a synthesized variant library is generated that encodes nucleic acid sequences that are ultimately translated into the amino acid sequences of proteins. Typical amino acid sequences include sequences that encode at least a portion of large peptides in addition to small peptides, such as antibody sequences. In some examples, each synthesized oligonucleic acid encodes variant codons in a portion of an antibody sequence. Typical antibody sequences encoded by a portion of the variant-synthesized nucleic acid include antigen-binding regions or their variable regions, or fragments thereof. Examples of antibody fragments partially encoded by the nucleic acids described herein include, but are not limited to, Fab, Fab’, F(ab’)2, and Fv fragments, bispecific antibodies, linear antibodies, single-chain antibody molecules, and multispecific antibodies formed from antibody fragments. Examples of antibody regions partially encoded by the oligonucleic acids described herein include, but are not limited to, the Fc region, Fab region, variable region of the Fab region, constant region of the Fab region, variable domains of the heavy or light chain (V H or V L ), or specific complementarity-determining regions (CDRs) of V H or V L . Variant libraries generated by the methods disclosed herein can result in one or more mutations in the antibody regions described herein. In one typical process, a variant library is generated for nucleic acids encoding multiple CDRs. Referring to Figure 12. A template nucleic acid encoding an antibody having regions of CDR1 (1210), CDR2 (1220), and CDR3 (1230) is modified by the methods described herein, and each CDR region contains multiple sites for mutation. Mutations are generated for each of the three CDRs (1215, 1225, and 1235) in a single variable domain of the heavy or light chain. Each site indicated by a star may include a single position, a run of multiple consecutive positions, or both, that are exchangeable with codon sequences different from the template nucleic acid sequence. The diversity of the variant library can be dramatically increased, using the methods provided herein, to a maximum diversity of ~10 10 or more.

[0095] In some examples, the variant library comprises one or more variants of the variable domain of the heavy or light chain (V H or V L ). In some examples, the variant library comprises one or more variants of the V H region. Exemplary V H regions include, but are not limited to, IGHV1, IGHV2, IGHV3, IGHV4, IGHV5, IGHV6, and IGHV7. In some examples, the variant library comprises one or more variants of the V L region. Exemplary V L regions include, but are not limited to, IGKV1, IGKV2, IGKV3, IGKV4, IGKV5, IGLV1, IGLV2, and IGLV3.

[0096] Mutations in the expression cassette

[0097] In some examples, a synthetic mutant library encoding part of an expression construct is generated. Typical parts of an expression construct include a promoter, an open reading frame, and a termination region. In some examples, the expression construct encodes 1, 2, 3 or more expression cassettes. A nucleic acid library is generated, which encodes codon mutations in separate regions of a single site or multiple sites that make up part of the expression construct cassette, as shown in FIG. 14. To generate two construct expression cassettes, a mutant nucleic acid encoding at least a part of a mutant sequence of a first promoter (1410), a first open reading frame (1420), a first terminator (1430), a second promoter (1440), a second open reading frame (1450), or a second terminator sequence (1460) is synthesized. After rounds of amplification, a library of 1,024 expression constructs was generated as described in the previous example. FIG. 14 provides an example layout. In some examples, additional regulatory sequences, such as untranslated control regions (UTRs) or enhancer regions, are also included in the expression cassettes referred to herein. The expression cassette may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more components for which mutant sequences are generated by the methods described herein. In some examples, the expression construct contains more than one gene in a multicistronic vector. In one example, the synthesized DNA nucleic acid is inserted into a viral vector (e.g., a lentivirus) and then packaged for transduction into cells or inserted into a non-viral vector for introduction into cells, and then screened and analyzed.

[0098] Expression vectors for inserting the nucleic acids disclosed herein include eukaryotic vectors (e.g., bacterial and fungal) and prokaryotic vectors (e.g., mammalian, plant, and insect expression vectors). Typical expression vectors include, but are not limited to, mammalian expression vectors: pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His ("6His" disclosed as SEQ ID NO:32), pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 Vector, pEF1a-tdTomato Vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), and pSF-CMV-PURO-NH2-CMYC; bacterial expression vectors: pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, and pSF-Tac; plant expression vectors: pRI 101-AN DNA and pCambia2301; and yeast expression vectors: pTYB21 and pKLAC2, and insect vectors: pAc5.1 / V5-His A and pDEST8. Typical cells include, but are not limited to, prokaryotic and eukaryotic cells. Typical eukaryotic cells include, but are not limited to, animal, plant, and fungal cells. Typical animal cells include, but are not limited to, insect, fish, and mammalian cells. Typical mammalian cells include mouse, human, and primate cells. The nucleic acids synthesized by the methods described herein are transferred into cells, which is performed by various methods known in the art including, but not limited to, transfection, transduction, and electroporation. Typical cell functions tested include, but are not limited to, changes in cell proliferation, migration / adhesion, metabolic activity, and cell signaling activity.

[0099] High-throughput nucleic acid synthesis

[0100] This specification provides a platform approach that utilizes miniaturization, parallelization, and vertical integration of processes between the ends from polynucleotide synthesis to gene assembly within nanowells on silicon to create an innovative synthetic platform. The devices described herein provide a silicon synthesis platform that can increase throughput by up to 1,000-fold or more compared to conventional synthesis methods, with the same footprint as a 96-well plate, and can produce up to approximately 1,000,000 or more polynucleotides, or 10,000 or more genes, in a single highly parallelized run.

[0101] With the advent of next-generation sequencing, high-resolution genomic data has become an important factor in studies that deeply explore the biological roles of various genes in both normal ecology and etiology. Central to this research are the central dogma of molecular biology and the concept of "sequential transfer of information residue by residue." Genomic information encoded in DNA is transcribed into a message, which is then translated into a protein, the active product within a given biological pathway.

[0102] Another exciting area of research relates to the discovery, development, and manufacture of therapeutic molecules that target highly specific cellular targets. DNA sequence libraries of high diversity lie at the heart of the development pipeline for targeted therapeutics. Gene mutants are used to express proteins in protein engineering cycles for design, structure, and testing of genes that are ideally optimized for high expression of proteins with high affinity for therapeutic targets. As an example, consider the binding pocket of a receptor. The ability to simultaneously test all permutations of all sequences of all residues within the binding pocket allows for a thorough examination and increases the likelihood of success. Saturation mutagenesis, in which researchers attempt to generate every possible mutation at a specific site within a receptor, represents one approach to this development challenge. Although expensive and time-consuming, this allows each mutant to be introduced at each position. In contrast, combinatorial mutagenesis, in which a few selected positions or short stretches of DNA can be extensively modified, generates an incomplete repertoire of mutants with a biased representation.

[0103] To facilitate the drug development pipeline, libraries with desirable mutants available at the correct positions and at the intended frequencies for testing, in other words, precision libraries, enable not only cost reduction but also a shortening of the screening time. Provided herein is a method for synthesizing a nucleic acid synthesis mutant library that results in the accurate introduction of each intended mutant at the desired frequency. For the end user, this translates into the ability to not only thoroughly sample the sequence space but also to interrogate these hypotheses in an efficient manner, reducing cost and screening time. Genome-wide editing can elucidate libraries in which important pathways, each mutant, and permutations of sequences can be tested for optimal functionality, can use thousands of genes to reconstruct entire pathways, and can use the genome to redesign biological systems for drug discovery.

[0104] In the first embodiment, the drug itself can be optimized using the methods described herein. For example, to improve the designated function of an antibody, a mutant nucleic acid library encoding a portion of the antibody is designed and synthesized. Thereafter, a mutant nucleic acid library against the antibody can be generated by the processes described herein (e.g., insertion into a vector following PCR mutagenesis). The antibody is then expressed in a production cell line and screened for enhanced activity. Examples of screening include examining the binding affinity for an antigen, stability, or modulation of effector functions (e.g., ADCC, complement, or apoptosis). Typical regions for optimizing an antibody include, but are not limited to, the Fc region, Fab region, variable region of the Fab region, constant region of the Fab region, variable domains of the heavy or light chains (V H or V L ), and the specific complementarity-determining regions (CDRs) of V H or V L .

[0105] Alternatively, the molecule for optimization is a receptor-binding epitope used as an activator or a competitive inhibitor. Following synthesis of a library of nucleic acid variants, the library of nucleic acid variants is inserted into a vector sequence and can then be expressed in cells. The receptor antigen can be expressed in cells (e.g., insect cells, mammalian cells, or bacterial cells) and then purified, or can be expressed in cells (e.g., mammalian cells) to examine the functional consequences from sequence variation. Functional consequences include, but are not limited to, changes in protein expression, binding affinity, and stability. Cellular functional consequences include, but are not limited to, changes in proliferation, growth, adhesion, death, migration, energy production, oxygen utilization, metabolic activity, cell signaling, aging, response to free radical damage, or any combination thereof. In some embodiments, the types of proteins selected for optimization are enzymes, transport proteins, G protein-coupled receptors, voltage-gated ion channels, transcription factors, polymerases, adapter proteins (proteins without enzymatic activity that serve to bring two other proteins together), and cytoskeletal proteins. Typical types of enzymes include, but are not limited to, signaling enzymes such as protein kinases, protein phosphatases, phosphodiesterases, histone deacetylases, and GTPases.

[0106] This specification provides a mutant nucleic acid library comprising mutants for molecules involved in all pathways or the entire genome. Exemplary pathways include, but are not limited to, metabolic, cell death, cell cycle progression, immune cell activation, inflammatory response, angiogenesis, lymphopoiesis, hypoxic stress response, oxidative stress response, or cell adhesion / migration pathways. Exemplary proteins in the cell death pathway include, but are not limited to, Fas, Cadd, Caspase 3, Caspase 6, Caspase 8, Caspase 9, Caspase 10, IAP, TNFR1, TNF, TNFR2, NF-kB, TRAF, ASK, BAD, and Akt. Exemplary proteins in the cell cycle pathway include, but are not limited to, NFkB, E2F, Rb, p53, p21, Cyclin A, Cyclin B, Cyclin D, Cyclin E, and cdc 25. Exemplary proteins in the cell migration pathway include, but are not limited to, Ras, Raf, PLC, cofilin, MEK, ERK, MLP, LIMK, ROCK, RhoA, Src, Rac, Myosin II, ARP2 / 3, MAPK, PIP2, integrin, talin, kindlin, migfilin, and filamin.

[0107] The nucleic acid libraries synthesized by the methods described herein can be expressed in a variety of cell types. Exemplary cell types include prokaryotic cells (e.g., bacterial and fungal cells) and eukaryotic cells (e.g., plant and animal cells). Exemplary animals include, but are not limited to, mice, rabbits, primates, fish, and insects. Exemplary plants include, but are not limited to, monocots and dicots. Exemplary plants also include, but are not limited to, microalgae, kelp, cyanobacteria, and green, brown, and red algae, wheat, tobacco, and corn, rice, cotton, vegetables, and fruits.

[0108] The nucleic acid libraries synthesized by the methods described herein can be expressed in various cells associated with disease states. Cells associated with disease states include cell lines, tissue samples, primary cells, cultured cells grown from a subject, or cells in a model system, obtained from a subject. Typical model systems include, but are not limited to, models of diseased plants and animals.

[0109] The nucleic acid libraries synthesized by the methods described herein can be expressed in various cell types, and changes in cell activity are evaluated. Typical cell activities include, but are not limited to, proliferation, cell cycle progression, cell death, adhesion, migration, reproduction, cell signaling, energy production, oxygen utilization, metabolic activity, and senescence, response to free radical damage, or any combination thereof.

[0110] To identify variant molecules associated with the prevention, reduction, or treatment of a disease state, the variant nucleic acid libraries described herein are expressed in cells associated with a disease state, or cells in which a disease state can be induced. In some examples, agents are used to induce a disease state in the cells. Typical tools for inducing a disease state include, but are not limited to, the Cre / Lox recombination system, LPS-induced inflammation, and streptozotocin, which induces hypoglycemia. Cells associated with a disease state can be cells from a subject with a particular medical condition, in addition to cells from a model system or cultured cells. Typical medical conditions include bacterial, fungal, viral, autoimmune, or proliferative disorders (e.g., cancer). In some examples, the variant nucleic acid libraries are expressed in a model system, cell line, or primary cells from a subject and screened for changes in at least one cell activity. Typical cell activities include, but are not limited to, proliferation, cell cycle progression, cell death, adhesion, migration, reproduction, cell signaling, energy production, oxygen utilization, metabolic activity, and senescence, response to free radical damage, or any combination thereof.

[0111] substrate

[0112] This specification provides a substrate comprising a plurality of clusters, where each cluster comprises a plurality of loci that support the binding and synthesis of polynucleotides. As used herein, the term "locus" refers to a discrete region of structure that provides support to a polynucleotide encoding a single predefined sequence for extension from a surface. In some examples, the locus is on a two-dimensional surface, such as a substantially flat surface. In some examples, the locus refers to a discrete raised or sunken site on the surface, such as a well, micro-well, channel, or post. In some examples, the surface of the locus comprises a substance that is actively functionalized to bind at least one nucleotide for polynucleotide synthesis, or, preferably, a population of the same nucleotides for the synthesis of a population of polynucleotides. In some examples, the polynucleotide refers to a population of polynucleotides encoding the same nucleic acid sequence. In some examples, the surface of the device encompasses one or more surfaces of the substrate.

[0113] The average error rate for polynucleotides synthesized within a library using the provided systems and methods is less than 1 in 1000, less than 1 in 1250, less than 1 in 1500, less than 1 in 2000, less than 1 in 3000, or less frequent than that. In some examples, the average error rate for polynucleotides synthesized within a library using the provided systems and methods is less than 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1100, 1 / 1200, 1 / 1250, 1 / 1300, 1 / 1400, 1 / 1500, 1 / 1600, 1 / 1700, 1 / 1800, 1 / 1900, 1 / 2000, 1 / 3000, or less than that. In some examples, the average error rate for polynucleotides synthesized within a library using the provided systems and methods is less than 1 / 1000.

[0114] In some examples, the overall error rate for polynucleotides synthesized in a library using the provided systems and methods is less than 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1100, 1 / 1200, 1 / 1250, 1 / 1300, 1 / 1400, 1 / 1500, 1 / 1600, 1 / 1700, 1 / 1800, 1 / 1900, 1 / 2000, less than 1 / 3000, or lower as compared to a predefined sequence. In some examples, the overall error rate for polynucleotides synthesized in a library using the provided systems and methods is less than 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, or less than 1 / 1000. In some examples, the overall error rate for polynucleotides synthesized in a library using the systems and methods provided herein is less than 1 / 500 or lower as compared to a predefined sequence.

[0115] In some examples, error correction enzymes can be used on polynucleotides synthesized in a library using the provided methods and systems. In some examples, the overall error rate for polynucleotides with error correction can be less than 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1100, 1 / 1200, 1 / 1300, 1 / 1400, 1 / 1500, 1 / 1600, 1 / 1700, 1 / 1800, 1 / 1900, 1 / 2000, less than 1 / 3000, or less than or equal to that as compared to a predefined sequence. In some examples, the overall error rate with error correction for polynucleotides synthesized in a library using the provided systems and methods can be less than 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, or less than 1 / 1000. In some examples, the overall error rate with error correction for polynucleotides synthesized in a library using the provided systems and methods can be less than 1 / 1000.

[0116] The error rate can limit the value of gene synthesis for the generation of libraries of gene variants. At an error rate of 1 / 300, approximately 0.7% of the clones in a 1500 base pair gene are correct. Since most of the errors from polynucleotide synthesis result in frameshift mutations, more than 99% of the clones in such a library do not produce full-length proteins. By reducing the error rate by 75%, the fraction of correct clones increases 40-fold. The methods and compositions of the present disclosure enable the rapid de novo synthesis of large nucleic acids and gene libraries at error rates lower than those generally observed for gene synthesis methods, thanks to both improved synthesis quality and the applicability of error correction methods made possible by ultra-parallel and time-efficient methods. Thus, the library can be synthesized with insertions, deletions, substitutions of bases, or with an overall error rate of less than 1 / 300, 1 / 400, 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1250, 1 / 1500, 1 / 2000, 1 / 2500, 1 / 3000, 1 / 4000, 1 / 5000, 1 / 6000, 1 / 7000, 1 / 8000, 1 / 9000, 1 / 10000, 1 / 12000, 1 / 15000, 1 / 20000, 1 / 25000, 1 / 30000, 1 / 40000, 1 / 50000, 1 / 60000, 1 / 70000, 1 / 80000, 1 / 90000, 1 / 100000, 1 / 125000, 1 / 150000, 1 / 200000, 1 / 300000, 1 / 400000, 1 / 500000, 1 / 600000, 1 / 700000, 1 / 800000, 1 / 900000, 1 / 1000000, or less, or with an overall error rate of 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99%, or more across the library.The methods and compositions of the present disclosure further relate to libraries of synthetic nucleic acids and genes on a large scale with low error rates associated with arrays that are error-free, compared to a predefined / preselected array, in at least a subset of the library, at least 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99%, or more of the polynucleotides or genes. In some examples, at least 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99%, or more of the polynucleotides or genes in isolated amounts within the library have the same sequence. In some examples, at least 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99%, or more of any polynucleotide or gene related to similarity or identity of 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or more have the same sequence. In some examples, the error rate associated with a specified locus on a polynucleotide or gene is optimized.Accordingly, each of the predetermined loci of one or more polynucleotides or genes as part of a large-scale library at a plurality of selected loci may have an error rate of 1 / 300, 1 / 400, 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1250, 1 / 1500, 1 / 2000, 1 / 2500, 1 / 3000, 1 / 4000, 1 / 5000, 1 / 6000, 1 / 7000, 1 / 8000, 1 / 9000, 1 / 10000, 1 / 12000, 1 / 15000, 1 / 20000, 1 / 25000, 1 / 30000, 1 / 40000, 1 / 50000, 1 / 60000, 1 / 70000, 1 / 80000, 1 / 90000, 1 / 100000, 1 / 125000, 1 / 150000, 1 / 200000, 1 / 300000, 1 / 400000, 1 / 500000, 1 / 600000, 1 / 700000, 1 / 800000, 1 / 900000, 1 / 1000000 or less. In various examples, loci optimized for such errors may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 30000, 50000, 75000, 100000, 500000, 1000000, 2000000, 3000000, or more loci. The loci optimized for the error may be distributed among at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 30000, 75000, 100000, 500000, 1000000, 2000000, 3000000, or more polynucleotides or genes.

[0117] Error rates can be achieved with or without error correction. Error rates can be achieved across the entire library or over 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99%, or more of the library.

[0118] Provided herein are structures that can include a surface that supports the synthesis of a plurality of polynucleotides having different predetermined arrays at addressable positions on a common support. In some examples, the device provides support for the synthesis of 2,000; 5,000; 10,000; 20,000; 30,000; 50,000; 75,000; 100,000; 200,000; 300,000; 400,000; 500,000; 600,000; 700,000; 800,000; 900,000; 1,000,000; 1,200,000; 1,400,000; 1,600,000; 1,800,000; 2,000,000; 2,500,000; 3,000,000; 3,500,000; 4,000,000; 4,500,000; 5,000,000; more than 10,000,000, or more non-identical polynucleotides. In some examples, the device provides support for the synthesis of 2,000; 5,000; 10,000; 20,000; 30,000; 50,000; 75,000; 100,000; 200,000; 300,000; 400,000; 500,000; 600,000; 700,000; 800,000; 900,000; 1,000,000; 1,200,000; 1,400,000; 1,600,000; 1,800,000; 2,000,000; 2,500,000; 3,000,000; 3,500,000; 4,000,000; 4,500,000; 5,000,000; more than 10,000,000, or more polynucleotides encoding another array. In some examples, at least a portion of the polynucleotides are configured to have the same array or be synthesized with the same array.

[0119] Methods and apparatuses for the manufacture and growth of polynucleotides having a length of about 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, or 2000 base lengths are provided herein. In some examples, the length of the polynucleotide formed is about 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, or 225 base lengths. The polynucleotide can be at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 base lengths. The polynucleotide can be 10 to 225 base lengths, 12 to 100 base lengths, 20 to 150 base lengths, 20 to 130 base lengths, or 30 to 100 base lengths.

[0120] In some examples, the polynucleotides are synthesized at separate loci of the substrate, where each locus supports the synthesis of a population of polynucleotides. In some examples, each locus supports the synthesis of a population of polynucleotides having a different sequence from a population of polynucleotides grown on a different locus. In some examples, the loci of the device are located within multiple clusters. In some examples, the device includes at least 10, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 11000, 12000, 13000, 14000, 15000, 20000, 30000, 40000, 50000 or more clusters. In some examples, the device includes 2,000; 5,000; 10,000; 100,000; 200,000; 300,000; 400,000; 500,000; 600,000; 700,000; 800,000; 900,000; 1,000,000; 1,100,000; 1,200,000; 1,300,000; 1,400,000; 1,500,000; 1,600,000; 1,700,000; 1,800,000; 1,900,000; 2,000,000; 300,000; 400,000; 500,000; 600,000; 700,000; 800,000; 900,000; 1,000,000; 1,200,000; 1,400,000; 1,600,000; 1,800,000; 2,000,000; 2,500,000; 3,000,000; 3,500,000; 4,000,000; 4,500,000; 5,000,000; or more than 10,000,000, or more separate loci. In some examples, the device includes approximately 10,000 separate loci. The amount of loci within a single cluster varies in different examples. In some examples, each cluster includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 130, 150, 200, 300, 400, 500 1000 or more loci. In some examples, each cluster includes approximately 50 - 500 loci. In some examples, each cluster includes approximately 100 - 200 loci.In some examples, each cluster contains from about 100 to 150 loci. In some examples, each cluster contains about 109, 121, 130, or 137 loci. In some examples, each cluster contains about 19, 20, 61, 64, or more loci.

[0121] The number of separate polynucleotides synthesized on the device can depend on the number of other loci available on the substrate. In some examples, the density of loci within a cluster of the device is at least or about 1 locus per mm 2 or about 1 locus per mm 2 10 loci per mm 2 25 loci per mm 2 50 loci per mm 2 65 loci per mm 2 75 loci per mm 2 100 loci per mm 2 130 loci per mm 2 150 loci per mm 2 175 loci per mm 2 200 loci per mm 2 300 loci per mm 2 400 loci per mm 2 500 loci per mm 2 1,000 loci per mm, or more. In some examples, the device has from about 500 loci per mm 2 to about 500 mm 2 about 10 loci per mm 2 to about 400 mm 2 about 25 loci per mm 2 to about 500 mm 2 about 50 loci per mm 2 to about 500 mm 2 about 100 loci per mm 2 to about 500 mm 2 about 150 loci per mm 2 to about 250 mm 2 about 10 loci per mm2 from about 250 mm 2 about 50 loci per hit, 1 mm 2 from about 200 mm 2 about 10 loci per hit, 1 mm 2 from about 200 mm 2 includes about 50 loci per hit. In some examples, the distance from the center of two adjacent loci within a cluster is from about 10 μm to about 500 μm, from about 10 μm to about 200 μm, or from about 10 μm to about 100 μm. In some examples, the distance from the centers of two adjacent loci is longer than about 10 μm, 20 μm, 30 μm, 40 μm, 50 μm, 60 μm, 70 μm, 80 μm, 90 μm or 100 μm. In some examples, the distance from the centers of two adjacent loci is less than about 200 μm, 150 μm, 100 μm, 80 μm, 70 μm, 60 μm, 50 μm, 40 μm, 30 μm, 20 μm or 10 μm. In some examples, each locus has a width of about 0.5 μm, 1 μm, 2 μm, 3 μm, 4 μm, 5 μm, 6 μm, 7 μm, 8 μm, 9 μm, 10 μm, 20 μm, 30 μm, 40 μm, 50 μm, 60 μm, 70 μm, 80 μm, 90 μm or 100 μm. In some examples, each locus has a width from about 0.5 μm to 100 μm, from about 0.5 μm to 50 μm, from about 10 μm to 75 μm, or from about 0.5 μm to 50 μm.

[0122] In some examples, the density of clusters within the device is at least or about 1 cluster per 100 mm 2 1 cluster per 10 mm 2 1 cluster per 5 mm 2 1 cluster per 4 mm 2 1 cluster per 3 mm 2 1 cluster per 2 mm 2 1 cluster per 1 mm 2 1 cluster per 1 mm 2 2 clusters per 1 mm 2 3 clusters per 1 mm 2 4 clusters per 1 mm 2 5 clusters per 1 mm 210 clusters per hit, 1 mm 2 50 clusters per hit, or more. In some examples, the device is 10 mm 2 from about 1 cluster per hit to 1 mm 2 includes about 10 clusters per hit. In some examples, the distance from the centers of two adjacent clusters is less than about 50 μm, 100 μm, 200 μm, 500 μm, 1000 μm, 2000 μm, or 5000 μm. In some examples, the distance from the centers of two adjacent clusters is from about 50 μm to about 100 μm, from about 50 μm to about 200 μm, from about 50 μm to about 300 μm, from about 50 μm to about 500 μm, and from about 100 μm to about 2000 μm. In some examples, the distance between the centers of two adjacent clusters is from about 0.05 mm to about 50 mm, from about 0.05 mm to about 10 mm, from about 0.05 mm to about 5 mm, from about 0.05 mm to about 4 mm, from about 0.05 mm to about 3 mm, from about 0.05 mm to about 2 mm, from about 0.1 mm to about 10 mm, from about 0.2 mm to about 10 mm, from about 0.3 mm to about 10 mm, from about 0.4 mm to about 10 mm, from about 0.5 mm to about 10 mm, from about 0.5 mm to about 5 mm, or from about 0.5 mm to about 2 mm. In some examples, each cluster has a diameter or width along one dimension of about 0.5 - 2 mm, about 0.5 - 1 mm, or about 1 - 2 mm. In some examples, each cluster has a diameter or width along one dimension of about 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 or 2 mm. In some examples, it has an internal diameter or width along one dimension of about 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.15, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 or 2.

[0123] The device may be approximately the size of a standard 96-well plate, e.g., about 100 mm and 200 mm by about 50 mm and 150 mm. In some examples, the device has a diameter less than about 1000 mm, 500 mm, 450 mm, 400 mm, 300 mm, 250 nm, 200 mm, 150 mm, 100 mm or 50 mm. In some examples, the diameter of the device is from about 25 mm to 1000 mm, from about 25 mm to about 800 mm, from about 25 mm to about 600 mm, from about 25 mm to about 500 mm, from about 25 mm to about 400 mm, from about 25 mm to about 300 mm, or from about 25 mm to about 200 mm. Non-limiting examples of the size of the device include about 300 mm, 200 mm, 150 mm, 130 mm, 100 mm, 76 mm, 51 mm and 25 mm. In some examples, the device is at least about 100 mm 2 ; 200 mm 2 ; 500 mm 2 ; 1,000 mm 2 ; 2,000 mm 2 ; 5,000 mm 2 ; 10,000 mm 2 ; 12,000 mm 2 ; 15,000 mm 2 ; 20,000 mm 2 ; 30,000 mm 2 ; 40,000 mm 2 ; 50,000 mm 2 or has a planar surface area of more than that. In some examples, the thickness of the device is from about 50 mm to about 2000 mm, from about 50 mm to about 1000 mm, from about 100 mm to about 1000 mm, from about 200 mm to about 1000 mm, or from about 250 mm to about 1000 mm. Non-limiting examples of the thickness of the device include 275 mm, 375 mm, 525 mm, 625 mm, 675 mm, 725 mm, 775 mm and 925 mm. In some examples, the thickness of the device varies with the diameter and depends on the composition of the substrate. For example, a device containing a substance other than silicon has a different thickness from a silicon device of the same diameter. The thickness of the device is determined by the mechanical strength of the material used and must be thick enough to support its own weight without cracking during handling. In some examples, the structure includes a plurality of devices described herein.

[0124] Surface substance

[0125] Devices that include a surface are provided herein, where the surface is modified at predetermined positions and as a result support low error rates, low dropout rates, high yields, and polynucleotide synthesis with high oligonucleotide representation. In some embodiments, the surface of the device provided herein for polynucleotide synthesis is made from a variety of substances that can be modified to support de novo polynucleotide synthesis reactions. In some cases, the device is sufficiently conductive such that, for example, a uniform electric field can be formed across all or part of the device. The devices described herein may include a flexible material. Exemplary flexible materials include, but are not limited to, modified nylon, unmodified nylon, nitrocellulose, and polypropylene. The devices described herein may include a rigid material. Exemplary rigid materials include, but are not limited to, glass, fused silica, silicon, silicon dioxide, silicon nitride, plastics (e.g., polytetrafluoroethylene, polypropylene, polystyrene, polycarbonate, and mixtures thereof), and metals (e.g., gold, platinum). The devices disclosed herein may be made from materials including silicon, polystyrene, agarose, dextran, cellulose-based polymers, polyacrylamide, polydimethylsiloxane (PDMS), glass, or any combination thereof. In some cases, the devices disclosed herein are manufactured from combinations of the materials listed herein or other suitable materials known in the art.

[0126] A list of the tensile strengths of exemplary materials described herein is as follows: nylon (70 MPa), nitrocellulose (1.5 MPa), polypropylene (40 MPa), silicon (268 MPa), polystyrene (40 MPa), agarose (1 - 10 MPa), polyacrylamide (1 - 10 MPa), polydimethylsiloxane (PDMS) (3.9 - 10.8 MPa). The solid supports described herein can have a tensile strength of 1 - 300, 1 - 40, 1 - 10, 1 - 5, or 3 - 11 MPa. The solid supports described herein can have a tensile strength of about 1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 20, 25, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 270, or more MPa. In some examples, the devices described herein include solid supports for polynucleotide synthesis in the form of a flexible material that can be stored in a continuous loop or reel, such as a tape or flexible sheet.

[0127] Young's modulus measures the resistance of a material to elastic (recoverable) deformation under load. A list of the Young's moduli of the stiffness of exemplary materials described herein is as follows: nylon (3 GPa), nitrocellulose (1.5 GPa), polypropylene (2 GPa), silicon (150 GPa), polystyrene (3 GPa), agarose (1 - 10 GPa), polyacrylamide (1 - 10 GPa), polydimethylsiloxane (PDMS) (1 - 10 GPa). The solid supports described herein can have a Young's modulus of 1 - 500, 1 - 40, 1 - 10, 1 - 5, or 3 - 11 GPa. The solid supports described herein can have a Young's modulus of about 1, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 20, 25, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 400, 500 GPa, or more. Since the relationship between softness and stiffness is inverse to each other, flexible materials have a low Young's modulus and change their shape greatly under load. In some examples, the solid supports described herein have a surface with at least the flexibility of nylon.

[0128] In some cases, the devices disclosed herein include a silicon dioxide base and a silicon dioxide surface layer. Alternatively, the device may have a silicon dioxide base. The surface of the device provided herein may be textured, resulting in an increase in the overall surface area for polynucleotide synthesis. The devices disclosed herein may comprise at least 5%, 10%, 25%, 50%, 80%, 90%, 95%, or 99% silicon. The devices disclosed herein may be made from a silicon-on-insulator (SOI) wafer.

[0129] Surface architecture

[0130] Devices are provided herein that include raised and / or recessed features. One advantage of having such features is an increase in the surface area that supports polynucleotide synthesis. In some examples, devices having raised and / or recessed features are referred to as three-dimensional substrates. In some examples, the three-dimensional device includes one or more channels. In some examples, one or more loci include a channel. In some examples, the channel is available for deposition of reagents by a deposition device such as a material deposition device. In some examples, reagents and / or fluids collect in a larger well that is in fluid communication with one or more channels. For example, the device includes a plurality of channels corresponding to a plurality of loci having clusters, and the plurality of channels are in fluid communication with one well of the cluster. In some methods, a library of polynucleotides is synthesized at a plurality of loci of a cluster.

[0131] In some examples, the structure is configured to enable control of the flow related to polynucleotide synthesis on the surface and control of the mass transfer pathway. In some examples, the configuration of the device enables control of the mass transfer pathway, chemical exposure time, and / or washing effect during polynucleotide synthesis and a uniform distribution. In some examples, the configuration of the device provides a volume sufficient for polynucleotide growth such that, for example, the volume excluded by the growing polynucleotide does not occupy 50, 45, 40, 35, 30, 25, 20, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1%, or less, of the volume available or the first available volume suitable for the growth of the polynucleotide, thereby enabling an increase in the scavenging rate. In some examples, the three-dimensional structure enables management of the fluid flow to allow for rapid exchange of chemical exposure.

[0132] Methods are provided herein for synthesizing DNA in amounts of 1 fM, 5 fM, 10 fM, 25 fM, 50 fM, 75 fM, 100 fM, 200 fM, 300 fM, 400 fM, 500 fM, 600 fM, 700 fM, 800 fM, 900 fM, 1 pM, 5 pM, 10 pM, 25 pM, 50 pM, 75 pM, 100 pM, 200 pM, 300 pM, 400 pM, 500 pM, 600 pM, 700 pM, 800 pM, 900 pM, or more. In some examples, the polynucleotide library may span about 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the length of the gene. The gene can vary up to about 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 100%.

[0133] Non-identical polynucleotides together may encode sequences for at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 100% of a gene. In some examples, a polynucleotide may encode 50%, 60%, 70%, 80%, 85%, 90%, 95%, or more of the sequence of a gene. In some examples, a polynucleotide may encode 80%, 85%, 90%, 95%, or more of the sequence of a gene.

[0134] In some examples, isolation is achieved by physical structure. In some examples, isolation is achieved by differential functionalization of a surface to create active and passive regions for polynucleotide synthesis. Differential functionalization is also achieved by alternating the hydrophobicity of the device surface, thereby creating a water contact angle effect that causes beading or wetting of the deposited reagent. Utilizing larger structures can reduce splashing and cross-contamination of separate polynucleotide synthesis sites with reagents in adjacent spots. In some examples, a device such as a polynucleotide synthesizer is used to deposit reagents at separate polynucleotide synthesis sites. A substrate having three-dimensional features is configured in a way that enables synthesis of many polynucleotides (e.g., more than about 10,000) with a low error rate (e.g., less than about 1:500, 1:1000, 1:1500, 1:2,000, 1:3,000, 1:5,000, or 1:10,000). In some examples, the device has a density of features of about 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, or 500, or more, per mm 2 and includes features having a density of features of about 1, 5, 10, 20, 30, 40, 50, 60, 70, 80, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, or 500, or more, per mm.

[0135] The wells of the device may have the same or different widths, heights, and / or volumes as other wells of the substrate. The channels of the device may have the same or different widths, heights, and / or volumes as other channels of the substrate. In some examples, the width of the cluster is from about 0.05 mm to about 50 mm, from about 0.05 mm to about 10 mm, from about 0.05 mm to about 5 mm, from about 0.05 mm to about 4 mm, from about 0.05 mm to about 3 mm, from about 0.05 mm to about 2 mm, from about 0.05 mm to about 1 mm, from about 0.05 mm to about 0.5 mm, from about 0.05 mm to about 0.1 mm, from about 0.1 mm to about 10 mm, from about 0.2 mm to about 10 mm, from about 0.3 mm to about 10 mm, from about 0.4 mm to about 10 mm, from about 0.5 mm to about 10 mm, from about 0.5 mm to about 5 mm, or from about 0.5 mm to about 2 mm. In some examples, the width of the well containing the cluster is from about 0.05 mm to about 50 mm, from about 0.05 mm to about 10 mm, from about 0.05 mm to about 5 mm, from about 0.05 mm to about 4 mm, from about 0.05 mm to about 3 mm, from about 0.05 mm to about 2 mm, from about 0.05 mm to about 1 mm, from about 0.05 mm to about 0.5 mm, from about 0.05 mm to about 0.1 mm, from about 0.1 mm to about 10 mm, from about 0.2 mm to about 10 mm, from about 0.3 mm to about 10 mm, from about 0.4 mm to about 10 mm, from about 0.5 mm to about 10 mm, from about 0.5 mm to about 5 mm, or from about 0.5 mm to about 2 mm. In some examples, the width of the cluster is less than 5 mm, 4 mm, 3 mm, 2 mm, 1 mm, 0.5 mm, 0.1 mm, 0.09 mm, 0.08 mm, 0.07 mm, 0.06 mm, or 0.05 mm. In some examples, the width of the cluster is from about 1.0 to about 1.3 mm. In some examples, the width of the cluster is about 1.150 mm. In some examples, the width of the well is less than 5 mm, 4 mm, 3 mm, 2 mm, 1 mm, 0.5 mm, 0.1 mm, 0.09 mm, 0.08 mm, 0.07 mm, 0.06 mm, or 0.05 mm. In some examples, the width of the well is from about 1.0 and 1.3 mm. In some examples, the width of the well is about 1.150 mm. In some examples, the width of the cluster is about 0.08 mm. In some examples, the width of the well is about 0.08 mm. The width of the cluster may refer to the cluster within a two-dimensional or three-dimensional substrate.

[0136] In some examples, the height of the well is from about 20 μm to about 1000 μm, from about 50 μm to about 1000 μm, from about 100 μm to about 1000 μm, from about 200 μm to about 1000 μm, from about 300 μm to about 1000 μm, from about 400 μm to about 1000 μm, or from about 500 μm to about 1000 μm. In some examples, the height of the well is less than about 1000 μm, less than about 900 μm, less than about 800 μm, less than about 700 μm, or less than about 600 μm.

[0137] In some examples, the device includes a plurality of channels corresponding to a plurality of loci within the cluster, where the height or depth of the channel is from about 5 μm to about 500 μm, from about 5 μm to about 400 μm, from about 5 μm to about 300 μm, from about 5 μm to about 200 μm, from about 5 μm to about 100 μm, from about 5 μm to about 50 μm, or from about 10 μm to about 50 μm. In some examples, the height of the channel is less than 100 μm, less than 80 μm, less than 60 μm, less than 40 μm, or less than 20 μm.

[0138] In some examples, the diameter of a channel, a locus (e.g., in a substantially planar substrate), or both a channel and a locus (e.g., in a three-dimensional structure device where the locus corresponds to the channel) is from about 1 μm to about 1000 μm, from about 1 μm to about 500 μm, from about 1 μm to about 200 μm, from about 1 μm to about 100 μm, from about 5 μm to about 100 μm, or from about 10 μm to about 100 μm, such as about 90 μm, 80 μm, 70 μm, 60 μm, 50 μm, 40 μm, 30 μm, 20 μm or 10 μm. In some examples, the diameter of a channel, a locus, or both a channel and a locus is less than about 100 μm, 90 μm, 80 μm, 70 μm, 60 μm, 50 μm, 40 μm, 30 μm, 20 μm, or 10 μm. In some examples, the distance from the center of two adjacent channels, loci, or a channel and a locus is from about 1 μm to about 500 μm, from about 1 μm to about 200 μm, from about 1 μm to about 100 μm, from about 5 μm to about 200 μm, from about 5 μm to about 100 μm, from about 5 μm to about 50 μm, or from about 5 μm to about 30 μm, such as about 20 μm.

[0139] Surface modification

[0140] In various examples, surface modification is utilized to effect chemical and / or physical changes to the surface of the device, or selected sites or regions of the device surface, by addition or subtraction to alter one or more chemical and / or physical properties of the surface. For example, surface modification includes, but is not limited to, (1) changing the wettability of the surface, (2) functionalizing the surface, i.e., providing, modifying, or substituting surface functional groups, (3) de-functionalizing the surface, i.e., removing surface functional groups, (4) otherwise changing the chemical composition of the surface, e.g., by etching, (5) increasing or decreasing the surface roughness, (6) providing a coating on the surface, e.g., a coating exhibiting a wettability different from that of the surface, and / or (7) depositing particles on the surface.

[0141] In some examples, the addition of a surface chemical layer (referred to as an adhesion promoter) facilitates the structured patterning of loci on the surface of the substrate. Exemplary surfaces for the application of adhesion promotion include, but are not limited to, glass, silicon, silicon dioxide, and silicon nitride. In some examples, the adhesion promoter is a chemical substance having a high surface energy. In some examples, a second chemical layer is deposited on the surface of the substrate. In some examples, the second chemical layer has a low surface energy. In some examples, the surface energy of the chemical layer coated on the surface supports the localization of droplets on the surface. Depending on the selected patterning arrangement, the region of locus proximity and / or fluid contact at the locus can be altered.

[0142] In some examples, for example, for polynucleotide synthesis, the device surface or the locus where the polynucleotide or other moiety is deposited or degraded is smooth, substantially planar (e.g., two-dimensional), or has irregularities such as raised or recessed features (e.g., three-dimensional features). In some examples, the device surface is modified with one or more different layers of compounds. Such modification of the layer of interest includes, but is not limited to, inorganic and organic layers such as metals, metal oxides, polymers, small organic molecules, etc. Non-limiting polymer layers include peptides, proteins, nucleic acids or their mimetics (e.g., peptide nucleic acids, etc.), polysaccharides, phospholipids, polyurethanes, polyesters, polycarbonates, polyureas, polyamides, polyethyleneamines, polyarylene sulfides, polysiloxanes, polyimides, polyacetates, and other suitable compounds described herein or otherwise known in the art. In some examples, the polymer is a heteropolymer. In some examples, the polymer is a homopolymer. In some examples, the polymer contains or is conjugated to a functional moiety.

[0143] In some examples, the resolved loci of the device are functionalized with one or more moieties that increase and / or reduce the surface energy. In some examples, a moiety is chemically inert. In some examples, a moiety is configured to support one or more processes in a desired chemical reaction, such as a polynucleotide synthesis reaction. The surface energy of the surface, i.e., the hydrophobicity, is a factor for determining the affinity of nucleotides binding onto the surface. In some examples, a method for functionalizing a device includes: (a) providing a device having a surface containing silicon dioxide; and (b) silanizing the surface using a suitable silanizing agent described herein or otherwise known in the art, such as an organofunctional alkoxysilane molecule.

[0144] In some examples, the organofunctional alkoxysilane molecule includes dimethylchloro-octadecyl-silane, methyldichloro-octadecyl-silane, trichloro-octadecyl-silane, trimethyl-octadecyl-silane, triethyl-octadecyl-silane, or any combination thereof. In some examples, the surface of the device is functionalized with polyethylene / polypropylene (functionalized by gamma-ray irradiation or chromic acid oxidation and reduction to a hydroxyalkyl surface), highly cross-linked polystyrene-divinylbenzene (derivatized by chloromethylation and aminated to a benzylamine functional surface), nylon (the terminal aminohexyl group is directly reactive), or etched with reduced polytetrafluoroethylene. Other methods and functionalizing agents are described in U.S. Patent No. 5,474,796, which is hereby incorporated by reference in its entirety.

[0145] In some examples, the surface of the device is functionalized by contact with a derivatization composition containing a mixture of silanes under reaction conditions effective to bind the silanes to the surface of the device via reactive hydrophilic moieties typically present on the surface of the device. The silanization generally covers the surface with organofunctional alkoxysilane molecules via self-assembly.

[0146] As is currently known in the art, for example, various reagents for functionalizing siloxanes can be further used to reduce or increase surface energy. Organofunctional alkoxysilanes can be classified according to their organic functions.

[0147] Provided herein are devices that can include patterning of agents that can bind to nucleosides. In some examples, the device may be coated with an active agent. In some examples, the device may be coated with a passive agent. Exemplary active agents included in the coating materials described herein include, but are not limited to, N-(3-triethoxysilylpropyl)-4-hydroxybutylamide (HAPS), 11-acetoxyundecyltriethoxysilane, n-decyltriethoxysilane, (3-aminopropyl)trimethoxysilane, (3-aminopropyl)triethoxysilane, 3-glycidoxypropyltrimethoxysilane (GOPS), 3-iodo-propyltrimethoxysilane, butyl-aldehyde-trimethoxysilane, dimer secondary aminoalkylsiloxane, (3-aminopropyl)-diethoxy-methylsilane, (3-aminopropyl)-dimethyl-ethoxysilane, and, (3-aminopropyl)-trimethoxysilane, (3-glycidoxypropyl)-dimethyl-ethoxysilane, glycidoxy-trimethoxysilane, (3-mercaptopropyl)-trimethoxysilane, 3-4 epoxycyclohexyl-ethyltrimethoxysilane, and, (3-mercaptopropyl)-methyl-dimethoxysilane, allyltrichlorosilane, 7-octa-1-enyltrichlorosilane, or bis(3-trimethoxysilylpropylamine).

[0148] Typical passivating agents included in the coating materials described herein include, but are not limited to, perfluorooctyltrichlorosilane; tridecafluoro-1,1,2,2-tetrahydrooctyl)trichlorosilane; 1H,1H,2H,2H-fluorooctyltriethoxysilane (FOS); trichloro(1H,1H,2H,2H-perfluorooctyl)silane; tert-butyl-[5-fluoro-4-(4,4,5,5-tetramethyl-1,3,2-dioxaborolan-2-yl)indol-1-yl]-dimethyl-silane; CYTOP (trademark); Fluorinate (trademark); perfluorooctyltrichlorosilane (PFOTCS); perfluorooctyldimethylchlorosilane (PFODCS); perfluorodecyltriethoxysilane (PFDTES); pentafluorophenyl-dimethylpropylchloro-silane (PFPTES); perfluorooctyltriethoxysilane; perfluorooctyltrimethoxysilane; octylchlorosilane; dimethylchloro-octadecyl-silane; methyldichloro-octadecyl-silane; trichloro-octadecyl-silane; trimethyl-octadecyl-silane; triethyl-octadecyl-silane; or octadecyltrichlorosilane.

[0149] In some examples, the functionalizing agent includes a hydrocarbon silane such as octadecyltrichlorosilane. In some examples, the functionalizing agent includes 11-acetoxyundecyltriethoxysilane, n-decyltriethoxysilane, (3-aminopropyl)trimethoxysilane, (3-aminopropyl)triethoxysilane, glycidyloxypropyl / trimethoxysilane, and N-(3-triethoxysilylpropyl)-4-hydroxybutylamide.

[0150] Polynucleotide synthesis

[0151] The methods of the present disclosure for polynucleotide synthesis can include processes involving phosphoramidite chemistry. In some examples, polynucleotide synthesis includes coupling a base with a phosphoramidite. Polynucleotide synthesis may include coupling a base by deposition of a phosphoramidite under coupling conditions, where the same base is optionally deposited more than once, i.e., with a phosphoramidite in a double bond. Polynucleotide synthesis may include capping of unreacted sites. In some examples, capping is optional. Polynucleotide synthesis may also include an oxidation or oxidation step. Polynucleotide synthesis may include deblocking, detritylation, and sulfurization. In some examples, polynucleotide synthesis includes either oxidation or sulfurization. In some examples, between one or each step during the polynucleotide synthesis reaction, the apparatus is washed, for example, using tetrazole or acetonitrile. The time frame for any one step in the phosphoramidite synthesis method may be shorter than about 2 minutes, 1 minute, 50 seconds, 40 seconds, 30 seconds, 20 seconds, and 10 seconds.

[0152] Polynucleotide synthesis using the phosphoramidite method may involve subsequent addition of the basic elements of the phosphoramidite (e.g., nucleoside phosphoramidite) to the growing polynucleotide chain for formation of the phosphite triester bond. Phosphoramidite polynucleotide synthesis proceeds in the 3’ to 5’ direction. Phosphoramidite polynucleotide synthesis enables controlled addition of one nucleotide to the growing polynucleotide chain per synthesis cycle. In some examples, each synthesis cycle includes a coupling step. The phosphoramidite bond formation includes formation of a phosphite triester bond between an activated nucleoside phosphoramidite and a nucleoside attached to the substrate, e.g., via a linker. In some examples, the nucleoside phosphoramidite is provided to the activated device. In some examples, the nucleoside phosphoramidite is provided to the device together with an activator. In some examples, the nucleoside phosphoramidite is provided to the device in an amount that is 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100-fold or more in excess of the nucleoside attached to the substrate. In some examples, the addition of the nucleoside phosphoramidite is performed in an anhydrous environment, e.g., in anhydrous acetonitrile. Following the addition of the nucleoside phosphoramidite, the device is optionally washed. In some examples, the coupling step is optionally repeated one or more additional times, together with a washing step between additions of the nucleoside phosphoramidite to the substrate. In some examples, the polynucleotide synthesis method used herein includes 1, 2, 3, or more consecutive coupling steps. In many cases, prior to coupling, the nucleoside attached to the device is deprotected by removal of a protecting group, where the protecting group functions to prevent polymerization. A common protecting group is 4,4’-dimethoxytrityl (DMT).

[0153] After coupling, the phosphoramidite polynucleotide synthesis method optionally includes a capping step. In the capping step, the growing polynucleotide is treated with a capping agent. The capping step is useful for blocking the 5'-OH groups attached to unreacted substrates after coupling from further chain elongation and preventing the formation of polynucleotides with internal base deletions. Further, the phosphoramidite activated with 1H-tetrazole may react slightly with the O6 position of guanosine. Without being bound by theory, when oxidized with I2 / water, this byproduct may undergo depurination, perhaps via O6-N7 migration. The depurinated site is ultimately cleaved during the final deprotection of the polynucleotide, thus reducing the yield of the full-length product. The O6 modification can be removed by treatment with a capping agent prior to oxidation with I2 / water. In some examples, including the capping step during polynucleotide synthesis results in a reduced error rate compared to synthesis without capping. As an example, the capping step includes treating the polynucleotide attached to the substrate with a mixture of acetic anhydride and 1-methylimidazole. Following the capping step, the apparatus is optionally washed.

[0154] In some examples, after the addition of nucleoside phosphoramidites and optionally after capping and one or more wash steps, the growing polynucleotide attached to the device is oxidized. The oxidation step oxidizes the phosphite triester to a tetracoordinate phosphate triester, which is a protected precursor of the naturally occurring phosphodiester bond between nucleosides. In some examples, oxidation of the growing polynucleotide is achieved by treatment with iodine and water, optionally in the presence of a weak base (e.g., pyridine, lutidine, collidine). Oxidation can be carried out under anhydrous conditions using, for example, tert-butyl hydroperoxide or (1S)-(+)-(10-camphorsulfonyl)-oxaziridine (CSO). In some methods, the capping step is carried out following oxidation. A second capping step allows for drying of the device since residual water from the ongoing oxidation can inhibit subsequent bonding. After oxidation, the device and the growing polynucleotide are optionally washed. In some examples, the oxidation step is replaced by a sulfurization step to obtain a polynucleotide phosphorothioate, where any capping step can be carried out after sulfurization. Many reagents, including but not limited to 3-(dimethylaminomethylene)amino)-3H-1,2,4-dithiazole-3-thione, DDTT, 3H-1,2-benzodithiol-3-one 1,1-dioxide, also known as the Beaucage reagent, and N,N,N’N’tetraethylthiuram disulfide (TETD), can effect efficient sulfur transfer.

[0155] To allow subsequent cycles of nucleoside incorporation to occur via ligation, the protected 5' end of the growing polynucleotide attached to the device is removed, such that the primary hydroxyl group reacts with the next nucleoside phosphoramidite. In some examples, the protecting group is DMT and deblocking occurs with trichloroacetic acid in dichloromethane. Performing detritylation for an extended period of time or with a more potent acid solution than recommended can lead to an increase in depurination of the polynucleotide attached to the solid support and thus may reduce the yield of the desired full-length product. The disclosed methods and compositions described herein provide controlled deblocking conditions that limit unwanted depurination reactions. In some examples, the polynucleotide attached to the device is washed after deblocking. In some examples, efficient washing after deblocking contributes to a synthesized polynucleotide having a low error rate.

[0156] Methods for the synthesis of polynucleotides typically include an iterating sequence of the following steps: application of a protected monomer to an actively functionalized surface (e.g., locus) of a protected monomer to bind to either an activated surface, a linker, or a previously deprotected monomer; deprotection of the applied monomer to react with a subsequently applied protected monomer; and application of another protected monomer for ligation. One or more intermediate steps include oxidation or sulfurization. In some examples, one or more washing steps precede or follow one or all of the steps.

[0157] Methods for phosphoramidite-based polynucleotide synthesis include a series of chemical steps. In some examples, one or more steps of the synthesis method include cycling of reagents, where one or more steps of the method include application of reagents to the device useful for the step. For example, the reagents are cycled by a series of liquid deposition and vacuum drying steps. For substrates containing three-dimensional features such as wells, microwells, channels, etc., the reagents are optionally passed through one or more regions of the device via wells and / or channels.

[0158] The methods and systems described herein relate to a polynucleotide synthesizer for the synthesis of polynucleotides. The synthesis can be performed in parallel. For example, at least or approximately at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 1000, 10000, 50000, 75000, 100000, or more polynucleotides can be synthesized in parallel. The total number of polynucleotides that can be synthesized in parallel can be 2 - 100000, 3 - 50000, 4 - 10000, 5 - 1000, 6 - 900, 7 - 850, 8 - 800, 9 - 750, 10 - 700, 11 - 650, 12 - 600, 13 - 550, 14 - 500, 15 - 450, 16 - 400, 17 - 350, 18 - 300, 19 - 250, 20 - 200, 21 - 150, 22 - 100, 23 - 50, 24 - 45, 25 - 40, 30 - 35. One skilled in the art will recognize that the total number of polynucleotides synthesized in parallel can be included within any range (e.g., 25 - 100) constrained by any of these values. The total number of polynucleotides synthesized in parallel can be included within any range defined by any of the values that function as endpoints of the range. The total molar mass of the polynucleotides synthesized in the device or the molar mass of each polynucleotide can be at least or at least about 10, 20, 30, 40, 50, 100, 250, 500, 750, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 25000, 50000, 75000, 100000 picomoles, or more. The length of each polynucleotide or the average length of the polynucleotides in the device can be at least or at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 300, 400, 500 nucleotides, or more.The length of each polynucleotide or the average length of the polynucleotides within the device can be at most or at least about 500, 400, 300, 200, 150, 100, 50, 45, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10 nucleotides, or less. The length of each polynucleotide or the average length of the polynucleotides within the device can be between 10 - 500, 9 - 400, 11 - 300, 12 - 200, 13 - 150, 14 - 100, 15 - 50, 16 - 45, 17 - 40, 18 - 35, 19 - 25. One of ordinary skill in the art will recognize that the length of each polynucleotide or the average length of the polynucleotides within the device can be within any range constrained by any of these values (e.g., 100 - 300). The length of each polynucleotide or the average length of the polynucleotides within the device can be within any range defined by any of the values that function as endpoints of the range.

[0159] The methods of polynucleotide synthesis on a surface provided herein enable high-speed synthesis. As an example, at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 125, 150, 175, 200 nucleotides, or more, are synthesized per hour. The nucleotides include the building blocks of adenine, guanine, thymine, cytosine, uridine, or analogs / modified versions thereof. In some examples, a library of polynucleotides is synthesized in parallel on a substrate. For example, a device containing approximately or at least approximately 100; 1,000; 10,000; 30,000; 75,000; 100,000; 1,000,000; 2,000,000; 3,000,000; 4,000,000; or 5,000,000 resolved loci can support the synthesis of at least the same number of distinct polynucleotides, where the polynucleotides encoding distinct sequences are synthesized at the resolved loci. In some examples, a library of polynucleotides is synthesized on a device at the low error rates described herein in about 3 months, 2 months, 1 month, 3 weeks, 15 days, 14 days, 13 days, 12 days, 11 days, 10 days, 9 days, 8 days, 7 days, 6 days, 5 days, 4 days, 3 days, 2 days, less than 24 hours, or less. In some examples, large nucleic acids assembled from a library of polynucleotides synthesized at low error rates using the substrates and methods described herein are prepared in about 3 months, 2 months, 1 month, 3 weeks, 15 days, 14 days, 13 days, 12 days, 11 days, 10 days, 9 days, 8 days, 7 days, 6 days, 5 days, 4 days, 3 days, 2 days, less than 24 hours, or less.

[0160] In some examples, the methods described herein result in the generation of a library of nucleic acids that contain variant nucleic acids at multiple codon positions. In some examples, the nucleic acid can have one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, twenty, thirty, forty, fifty, or more positions of variant codon sites.

[0161] In some examples, one or more positions of the variant codon sites may be adjacent. One or more positions of the variant codon sites may not be adjacent and may be separated by one, two, three, four, five, six, seven, eight, nine, ten, or more codons.

[0162] In some examples, the nucleic acid may contain multiple positions of variant codon sites, where all of the variant codon sites are adjacent to each other, forming an extension of the variant codon site. In some examples, the nucleic acid may contain multiple positions of variant codon sites, where the variant codon sites are not adjacent to each other. In some examples, the nucleic acid may contain multiple positions of variant codon sites, where some of the variant codon sites are adjacent to each other, forming an extension of the variant codon site, and some of the variant codon sites are not adjacent to each other.

[0163] Referring to the drawings, FIG. 15 shows an exemplary process workflow for the synthesis of nucleic acids (e.g., genes) from shorter polynucleotides. The workflow is typically divided into the following stages: (1) de novo synthesis of a single-stranded polynucleotide acid library, (2) ligation of polynucleotides to form larger fragments, (3) error correction, (4) quality control, and (5) transport. Prior to de novo synthesis, the intended nucleic acid sequence, or group of nucleic acid sequences, is preselected. For example, a group of genes is preselected for generation.

[0164] Once a larger nucleic acid is selected for generation, a predefined library of polynucleotides is designed for de novo synthesis. Various suitable methods for generating high-density polynucleotide arrays are known. In an example of a workflow, a surface layer (1501) of a device is provided. In this example, the chemistry of the surface is altered to improve the polynucleotide synthesis process. Regions of low surface energy are created to repel liquid, while regions of high surface energy are created to attract liquid. The surface itself may be in the form of a flat surface or may include deformations in shape such as protrusions or microwells that increase the surface area. In an example of a workflow, as disclosed in International Patent Application Publication WO / 2015 / 021080, which is hereby incorporated by reference in its entirety, selected high surface energy molecules serve a dual function that supports the chemistry of DNA.

[0165] In situ preparation of a polynucleotide array is generated on a solid support and utilizes a single nucleotide extension process to extend a plurality of oligomers in parallel. A deposition device, such as a material deposition device, is designed to release reagents in a stepwise manner such that a plurality of polynucleotides are extended one residue at a time in parallel to generate oligomers having a predefined nucleic acid sequence (1502). In some examples, the polynucleotide is cleaved from the surface at this stage. Cleavage includes, for example, gas cleavage with ammonia or methylamine.

[0166] The generated polynucleotide library is placed in a reaction chamber. In this exemplary workflow, the reaction chamber (also referred to as a "nanoreactor") is a well coated with silicon, which contains PCR reagents and is dispensed onto the polynucleotide library (1503). Before or after sealing of the polynucleotide (1504), reagents are added to release the polynucleotide from the substrate. In the exemplary workflow, the polynucleotide is released after sealing of the nanoreactor (1505). Once released, the single-stranded polynucleotide fragments hybridize to span the sequence over the full-length range of the DNA. Each synthesized polynucleotide is designed to have a small portion that overlaps with at least one other polynucleotide in the population, enabling partial hybridization (1505).

[0167] After hybridization, the PCA reaction begins. During the polymerase cycle, the polynucleotides anneal to complementary fragments and the gaps are filled by the polymerase. Each cycle increases the length of the various fragments randomly depending on which polynucleotides find each other. Complementarity between the fragments enables the formation of the full large full-length of double-stranded DNA (1506).

[0168] After PCA is completed, the nanoreactor is separated from the device (1507) and positioned for interaction with the device having primers for PCR (1508). After sealing, the nanoreactor is subjected to PCR (1509) and larger nucleic acids are amplified. After PCR (1510), the nanochamber is opened (1511), error correction reagents are added (1512), the chamber is sealed (1513), and an error correction reaction occurs to remove mismatched base pairs and / or strands with poor complementarity from the double-stranded PCR amplification product (1514). The nanoreactor is opened and separated (1515). The error-corrected product is then subjected to additional processing steps such as PCR and molecular barcoding and then packaged (1522) for transport (1523).

[0169] In some examples, quality control measures are taken. After error correction, the quality control process includes, for example, an interaction (1516) with a wafer having sequencing primers for amplification of the error-corrected product, sealing the wafer in a chamber containing the error-corrected amplification product (1517), and performing additional rounds of amplification (1518). The nano-reactor is opened (1519), the products are pooled (1520), and sequenced (1521). After an acceptable quality control decision is made, the packaged product (1522) is approved for shipping (1523).

[0170] In some examples, the polynucleotides generated by a workflow such as that of FIG. 15 are subjected to mutagenesis using the overlapping primers disclosed herein. In some examples, a library of primers is generated by in situ preparation on a solid support and utilizes a single nucleotide extension process to extend multiple oligomers in parallel. A deposition device, such as a material deposition device, is designed to release reagents in a stepwise manner such that multiple polynucleotides are extended one residue at a time in parallel to generate oligomers having a predetermined sequence (1502).

[0171] Computer system

[0172] Any of the systems described herein can be operably linked to a computer and can be automated locally or remotely via the computer. In various examples, the methods and systems of the present disclosure can further include software programs on a computer system and their use. Accordingly, computer control for synchronization of dispensing / decompression / refilling functions, such as orchestrating and synchronizing the operation of a material deposition device, dispensing actions, and decompression operations, is within the scope of the present disclosure. The computer system is programmed to interfere between a base sequence specified by a user and the position of a material deposition device to deliver the correct reagent to a specified region of a substrate.

[0173] The computer system (1600) illustrated in FIG. 16 can be understood as a logical device capable of reading instructions from a network port (1605) optionally connected to a server (1609) having a medium (1611) and / or a fixed medium (1612). A system such as that shown in FIG. 16 can include a CPU (1601), a disk drive (1603), optional input devices such as a keyboard (1615) and / or a mouse (1616), and an optional monitor (1607). Data communication can be achieved via a designated communication medium to a server at a local or remote location. The communication medium can include any means for transmitting and / or receiving data. For example, the communication medium can be a network connection, a wireless connection, or an Internet connection. Such connections can provide communication on the World Wide Web. Data related to the present disclosure can be transmitted by such a network or connection for reception and / or review by a party (1622) as illustrated in FIG. 16.

[0174] FIG. 17 is a block diagram illustrating an architecture of a first example of a computer system (1700) that can be used in connection with an example of the present disclosure. As represented in FIG. 17, an example of a computer system may include a processor (1702) for processing instructions. Non-limiting examples of processors include: Intel Xeon™ processors, AMD Opteron™ processors, Samsung 32-bit RISC ARM 1176JZ(F)-S v1.0™ processors, ARM Cortex-A8 Samsung S5PC100™ processors, ARM Cortex-A8 Apple A4™ processors, Marvell PXA 930™ processors, or functionally equivalent processors. Multiple threads of execution may be available for parallel processing. In some examples, multiple processors, or a processor having multiple cores, may also be used, whether in a single computer system, in a cluster, or distributed across a network of systems including multiple computers, cellular phones, and / or personal digital assistants.

[0175] As illustrated in FIG. 17, a high-speed cache (1704) can be connected to or incorporated within a processor (1702) to provide high-speed memory for instructions or data that have been recently used or are frequently used by the processor (1702). The processor (1702) is connected to a north bridge (1706) by a processor bus (1708). The north bridge (1706) is connected to a random access memory (RAM) (1710) by a memory bus (1712) and manages access to the RAM (1710) by the processor (1702). The north bridge (1706) is also connected to a south bridge (1714) by a chipset bus (1716). The south bridge (1714) is in turn connected to a peripheral bus (1718). The peripheral bus can be, for example, PCI, PCI-X, PCI Express, or other peripheral buses. The north bridge and the south bridge are often referred to as a processor chipset and manage data transfer between the processor, the RAM, and peripheral components on the peripheral bus (1718). In some alternative architectures, the functions of the north bridge can be incorporated into the processor instead of using a separate north bridge chip. In some examples, the system (1700) can include an accelerator card (1722) attached to the peripheral bus (1718). The accelerator can include a field programmable gate array (FPGA) or other hardware for facilitating specific processing. For example, the accelerator can be used for adaptive data restructuring or for evaluating algebraic expressions used in an extended setup process.

[0176] Software and data are stored in an external storage device (1724) and can be loaded into the RAM (1710) and / or cache (1704) for use by a processor. The system (1700) includes an operating system for managing system resources; non-limiting examples of operating systems include: Linux®, Windows™, MACOS™, BlackBerry OS™, iOS™, and other functionally equivalent operating systems, as well as application software that runs on an operating system for managing storage and optimization of data according to examples of the present disclosure. In this example, the system (1700) also includes a network interface card (NIC) (1720) and (1721) connected to a peripheral bus to provide a network interface to an external storage device such as a network attached storage (NAS), and other computer systems that can be used for distributed parallel processing.

[0177] FIG. 18 is a schematic diagram showing a network (1800) comprising a plurality of computer systems (1802a) and (1802b), a plurality of mobile phones and personal digital assistants (1802c), and network attached storage (NAS) (1804a) and (1804b). In an example, systems (1802a), (1802b), and (1802c) can manage data storage and optimize data access to data stored in network attached storage (NAS) (1804a) and (1804b). A mathematical model can be used for this data and evaluated using distributed parallel processing across computer systems (1802a) and (1802b), as well as mobile phones and personal digital assistant systems (1802c). Computer systems (1802a) and (1802b), as well as mobile phones and personal digital assistant systems (1802c) can also provide parallel processing for adaptive data reconstruction of data stored in network attached storage (NAS) (1804a) and (1804b). FIG. 18 illustrates only one example, and various other computer architectures and systems can be used in conjunction with the various examples of the present disclosure. For example, blade servers can be used to provide parallel processing. Processor blades can be connected via a backplane to provide parallel processing. Storage can also be connected to the backplane or connected as network attached storage (NAS) via another network interface. In some examples, a processor can maintain separate memory spaces and transfer data via a network interface, a backplane, or other connectors for parallel processing by other processors. In other examples, some or all of the processors can use a shared virtual address memory space.

[0178] FIG. 19 is a block diagram of a multiprocessor computer system (1900) using a shared virtual address memory space according to an example. The system includes a plurality of processors (1902a)-(1902f) that can access a shared memory subsystem (1904). The system incorporates a plurality of programmable hardware memory algorithm processors (MAP) (1906a)-(1906f) into the shared memory subsystem (1904). Each of the MAPs (1906a)-(1906f) can include memories (1908a)-(1908f) and one or more field programmable gate arrays (FPGA) (1910a)-(1910f). The MAP provides configurable functional units, and specific algorithms or parts of algorithms can be provided to the FPGAs (1910a)-(1910f) for tightly collaborative processing with respective processors. For example, the MAP can be used to evaluate algebraic expressions related to a data model and perform adaptive data reconstruction in the example. In this example, for such purposes, all of the processors can access each MAP globally. In one configuration, each MAP can use direct memory access (DMA) to access the associated memories (1908a)-(1908f), thereby enabling tasks to be executed separately and asynchronously from each microprocessor (1902a)-(1902f). In this configuration, the MAP can supply results directly to other MAPs for pipelining and parallel execution of algorithms.

[0179] The above computer architectures and systems are merely examples, and various other computer, mobile phone, and personal digital assistant architectures and systems can be used with systems that use any combination of general processors, coprocessors, FPGAs, and other programmable logic circuits, system-on-chips (SOCs), application-specific integrated circuits (ASICs), and other processing and logic elements. In some examples, all or part of a computer system may be implemented in software or hardware. Various data storage media can be used with examples including random access memory, hard drives, flash memory, tape drives, disk arrays, network-attached storage (NAS), and other local or distributed data storage devices and systems.

[0180] In an example, a computer system can be implemented using software modules executed on any of the above or other computer architectures and systems. In other examples, the functionality of the system can be partially or fully implemented in programmable logic circuits such as field-programmable gate arrays (FPGAs) as referred to in FIG. 19, system-on-chips (SOCs), application-specific integrated circuits (ASICs), or other processing and logic elements. For example, a set processor and optimizer can be implemented using hardware acceleration via a hardware accelerator card such as the accelerator card (1722) illustrated in FIG. 17.

[0181] The following examples are described to more clearly illustrate to those skilled in the art the principles and practice of the embodiments disclosed herein and are not to be construed as limiting the scope of any claimed embodiment. Unless otherwise specified, all parts and percentages are by weight.

Example

[0182] The following examples are provided to illustrate various embodiments of the present disclosure and are not intended to limit the present disclosure in any way. Together with the methods described herein, these examples illustrate preferred and typical embodiments and are not intended to be limiting of the scope of the present disclosure. Modifications and other uses within the spirit of the present disclosure as defined by the scope of the claims will be apparent to those skilled in the art.

[0183] Example 1: Functionalization of the Device Surface

[0184] The device was functionalized to support the binding and synthesis of a library of polynucleotides. The device surface was first wet cleaned using a piranha solution containing 90% H2SO4 and 10% H2O2 for 20 minutes. The device was rinsed in several beakers containing DI water and held under the gooseneck-shaped faucet of DI water for 5 minutes and dried with N2. Then, the device was immersed in NH4OH (1:100; 3 mL:300 mL) for 5 minutes, rinsed with DI water using a handgun, immersed in each of three consecutive beakers containing DI water for 1 minute each, and then rinsed again with DI water using a handgun. Thereafter, the device was plasma cleaned by exposing the device surface to O2. Using a SAMCO PC-300 instrument, O2 was plasma etched at 250 watts for 1 minute in the downstream mode.

[0185] The cleaned device surface was actively functionalized with a solution containing N-(3-triethoxysilylpropyl)-4-hydroxybutylamide using a YES-1224P evaporation oven system with the following parameters: 0.5 to 1 torr, 60 minutes, 70 °C, 135 °C vaporizer. The device surface was resist coated using a Brewer Science 200X spin coater. SPR™ 3612 photoresist was spin coated onto the device at 2500 rpm for 40 seconds. The device was pre-baked on a Brewer hot plate at 90 °C for 30 minutes. The device was subjected to photolithography using a Karl Suss MA6 mask aligner instrument. The device was exposed for 2.2 seconds and developed in MSF 26A for 1 minute. The remaining developer was rinsed with a hand-held gun and the device was immersed in water for 5 minutes. The device was baked in an oven at 100 °C for 30 minutes and then visually inspected for lithography defects using a Nikon L200. A cleaning process was used to remove the remaining resist using a SAMCO PC-300 instrument and O2 plasma etching was performed at 250 watts for 1 minute.

[0186] The device surface was passively functionalized with a 100 μL solution of perfluorooctyltrichlorosilane mixed with 10 μL of light oil. The device was placed in the chamber, pumped out for 10 minutes, then the valve was closed and the pump was stopped and it was left for 10 minutes. The chamber was vented. The device was resist stripped by sonicating at 70 °C for 5 minutes in 500 mL of NMP with two immersions by sonication at maximum power (9 on a Crest system). The device was then sonicated at maximum power for 5 minutes in 500 mL of isopropanol at room temperature. The device was immersed in 300 mL of 200 proof ethanol and blown dry with N2. The functionalized surface was activated to function as a support for polynucleotide synthesis.

[0187] Example 2: Synthesis of a 50-mer sequence

[0188] The two-dimensional oligonucleotide synthesizer was incorporated into a flow cell and connected to the flow cell (Applied Biosystems (ABI394 DNA Synthesizer)). The two-dimensional oligonucleotide synthesizer was uniformly functionalized with N-(3-triethoxysilylpropyl)-4-hydroxybutylamide (Gelest), and using this, an exemplary 50-bp polynucleotide (the "50-mer polynucleotide") was synthesized using the polynucleotide synthesis method described herein.

[0189] The sequence of the 50-mer is as set forth in SEQ ID NO.:20. 5’AGACAATCAACCATTTGGGGTGGACAGCCTTGACCTCTAGACTTCGGCAT##TTTTTTTTTT3’ (SEQ ID NO.: 20), where # represents thymidine-succinylhexamide CED phosphoramidite (CLP-2244 from ChemGenes), which is a cleavable linker that allows for the release of the polynucleotide from the surface during deprotection.

[0190] Synthesis was performed using standard DNA synthesis chemistry (coupling, capping, oxidation, and deblocking) according to the protocol of Table 4 and the ABI synthesizer.

[0191]

Table 4-1

[0192]

Table 4-2

[0193]

Table 4-3

[0194] The combination of phosphoramidite / activator was delivered in the same manner as the delivery of bulk reagents through a flow cell. Since the environment remained much "wetter" by the reagents, no drying step was performed.

[0195] The flow restrictor was removed from the ABI 394 synthesizer to allow for a faster flow. Without the flow restrictor, the flow rates of amidite (0.1 M in ACN), activator (0.25 M benzoylthiotetrazole ("BTT"; GlenResearch 30 - 3070 - xx) in ACN), and Ox (0.02 M I2 in 20% pyridine, 10% water, and 70% THF) were approximately ~100 μL / sec, for acetonitrile ("ACN") and capping reagent (1:1 mixture of CapA and CapB, where CapA is acetic anhydride in THF / pyridine and CapB is 16% 1 - methylimidizole in THF) approximately ~200 μL / sec, and for Deblock (3% dichloroacetic acid in toluene) approximately ~300 μL / sec (compared to ~50 μL / sec for all reagents with the flow restrictor). The time to completely extrude the oxidizer was observed, the timing of the chemical flow was adjusted as appropriate, and an extra ACN wash was introduced between the various chemicals. After polynucleotide synthesis, the chip was deprotected overnight in gaseous ammonia at 75 psi. Five drops of water were added to the surface to regenerate the polynucleotide. The regenerated polynucleotide was then analyzed with a small RNA chip of the BioAnalyzer (data not shown).

[0196] Example 3: Synthesis of a 100 - mer sequence

[0197] The same process described in Example 2 for the synthesis of 50-mer arrays was used for the synthesis of 100-mer polynucleotides (the "100-mer polynucleotides"; 5’CGGGATCCTTATCGTCATCGTCGTACAGATCCCGACCCATTTGCTGTCCACCAGTCATGCTAGCCATACCATGATGATGATGATGATGAGAACCCCGCAT##TTTTTTTTTT3’, where # represents thymidine-succinylhexamide CED phosphoramidite (CLP-2244 from ChemGenes); SEQ ID NO:21) on two different silicon chips. One silicon chip was uniformly functionalized with N-(3-triethoxysilylpropyl)-4-hydroxybutylamide, and the other silicon chip was functionalized with a 5 / 95 mixture of 11-acetoxyundecyltriethoxysilane and n-decyltriethoxysilane. Also, the polynucleotides extracted from the surface were analyzed on a BioAnalyzer instrument (data not shown).

[0198] Using the following thermal cycling program, all 10 samples from the two chips were further amplified in a 50 μL PCR mixture (25 μL of NEB Q5 mastermix, 2.5 μL of 10 μM forward primer, 2.5 μL of 10 μM reverse primer, 1 μL of polynucleotide extracted from the surface, and up to 50 μL of water) using a forward primer (5’ATGCGGGGTTCTCATCATC3’; SEQ ID NO.:22) and a reverse primer (5’CGGGATCCTTATCGTCATCG3’; SEQ ID NO.:23): 98°C, 30 seconds 98°C, 10 seconds; 63°C, 10 seconds; 72°C, 10 seconds; repeat 12 cycles 72°C, 2 minutes.

[0199] The PCR products were also run on a BioAnalyzer (data not shown), showing a sharp peak at the position of the 100-mer. Next, the PCR amplified samples were cloned and Sanger sequenced. Table 5 summarizes the results from Sanger sequencing for the samples obtained from spots 1-5 on chip 1 and spots 6-10 on chip 2.

[0200]

Table 5

[0201] Therefore, the synthesized polynucleotides of high quality and high uniformity were repeated on two chips with different surface chemical properties. Overall, 89% corresponding to 233 out of 262 sequenced 100-mers had perfect error-free sequences. Finally, Table 6 summarizes the error characteristics for the sequences obtained from the polynucleotide samples of spots 1-10.

[0202]

Table 6

[0203] Example 4: Generation of a nucleic acid library by single-site, single-position mutagenesis

[0204] Polynucleotide primers for a series of PCR reactions were de novo synthesized to generate a library of nucleic acid variants of the template nucleic acid (see FIGS. 2A-4D). Four types of primers were generated in FIG. 4A: an outer 5' primer (415), an outer 3' primer (430), an inner 5' primer (425), and an inner 3' primer (420). Using a polynucleotide synthesis method as generally outlined in Table 4, the inner 5' primer / first polynucleotide (420) and the inner 3' primer / second polynucleotide (425) were generated. The inner 5' primer / first polynucleotide (420) represents a set of up to 19 primers of a predetermined sequence, and each primer in the set differs from the others at one codon at one site of the sequence.

[0205] Perform polynucleotide synthesis on a device having at least two clusters, each cluster having 121 individually addressable loci.

[0206] The inner 5' primer (425) and the inner 3' primer (420) were synthesized in separate clusters. The inner 5' primer (425) was replicated 121 times and extended at 121 loci within a single cluster. For the inner 3' primer (420), 19 of the mutant sequences were each extended at 6 different loci, resulting in the extension of 114 polynucleotides at 114 different loci.

[0207] The synthesized polynucleotides were cleaved from the surface of the device and transferred to plastic vials. As illustrated in FIG. 4B, a first PCR reaction was performed using fragments of the long nucleic acid sequences (435), (440) to amplify the template nucleic acid. As illustrated in FIGS. 4C-4D, a second PCR reaction was performed using a combination of primers and the product of the first PCR reaction as a template. As shown in the trace of FIG. 20, the analysis of the second PCR product was performed on a BioAnalyzer.

[0208] Example 5: Generation of a nucleic acid library containing 96 different sets of variants at one position

[0209] As shown generally in FIG. 4A and discussed in Example 2, four sets of primers were generated using de novo polynucleotide synthesis. For the inner 5' primer (420), 96 different sets of primers were generated, and each set of primers targeted a different one codon located within one site of the template nucleic acid. For each set of primers, 19 different variants were generated, and each variant contained a codon encoding a different amino acid at one site. As shown generally in FIGS. 4A-4D and described in Example 2, two rounds of PCR were performed using the generated primers. In the electropherogram used to calculate the 100% amplification success rate, 96 sets of amplification products were visualized (FIG. 21).

[0210] Example 6: Generation of a nucleic acid library containing 500 different sets of variants at one position

[0211] As shown generally in FIG. 4A and discussed in Example 2, four sets of primers were generated using de novo polynucleotide synthesis. For the inner 5' primer (420), 500 different sets of primers were generated, and each set of primers targeted a different one codon located within one site of the template nucleic acid. For each set of primers, 19 different variants were generated, and each variant contained a codon encoding a different amino acid at one site. As shown generally in FIG. 4A and described in Example 2, two rounds of PCR were performed using the generated primers. The electropherogram displays each of the 500 sets of PCR products having a population of nucleic acids with 19 variants at a different single site (data not shown). Analysis of comprehensive sequencing of the library showed a success rate higher than 99% across mutations of the preselected codons (sequence trace and analysis data not shown).

[0212] Example 7: Single-site mutagenesis primer for one position

[0213] An example of codon mutation design is provided in Table 7 for Yellow Fluorescent Protein. In this case, one codon from the 50-mer sequence is mutated 19 times. Various nucleic acid sequences are shown in bold. The wild-type primer sequence is ATG GTG AGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCAT (SEQ ID NO.: 1). In this case, the wild-type codon encodes valine, which is indicated by the underline in SEQ ID NO:1. Therefore, the following 19 mutants exclude the codon encoding valine. In an alternative embodiment, when all triplets are considered, then 60 mutants are all generated containing alternative sequences to the wild-type codon.

[0214] [Table 7]

[0215] Example 8: Nucleic acid mutants at one site, double positions

[0216] De novo polynucleotide synthesis was carried out under the same conditions as described in Example 2. One cluster on the device was generated, and the cluster contained pre-determined synthetic mutants of nucleic acids for the positions of two consecutive codons at one site, and each position was a codon encoding an amino acid. In this arrangement, 19 mutants per position were generated for two positions with three replicates of each nucleic acid, and 114 synthesized nucleic acids were obtained as a result.

[0217] Example 9: Nucleic acid mutants at multiple sites, double positions

[0218] De novo polynucleotide synthesis was performed under the same conditions as described in Example 2. One cluster on the device was generated, and the cluster contained pre-determined synthetic variants of the nucleic acid for the positions of two non-consecutive codons, with each position being a codon encoding an amino acid. In this arrangement, 19 variants per position were generated for the two positions.

[0219] Example 10: Nucleic acid variants at a series of triple positions

[0220] De novo polynucleotide synthesis was performed under the same conditions as described in Example 2. One cluster on the device was generated, and the cluster contained pre-determined synthetic variants of the reference nucleic acid for three consecutive codon positions. In the arrangement of three consecutive codon positions, 19 variants per position were generated for the three positions with two replicates of each nucleic acid, resulting in 114 synthesized nucleic acids.

[0221] Example 11: Nucleic acid variants at multiple sites and triple positions

[0222] De novo polynucleotide synthesis was performed under the same conditions as described in Example 2. One cluster on the device was generated, and the cluster contained pre-determined synthetic variants of the reference nucleic acid for at least three non-consecutive codon positions. Within the pre-determined region, the positions of the codons encoding three histidine residues were changed.

[0223] Example 12: Nucleic acid variants at multiple sites and multiple positions

[0224] De novo polynucleotide synthesis was performed under conditions similar to those described in Example 2. One cluster on the device was generated, and the cluster contained pre-defined synthetic variants of the reference nucleic acid for one or more codon positions in a contiguous manner. Five positions were varied in the library. The first position encoded a codon for a resulting 50 / 50 K / R ratio in the expressed protein; the second position encoded a codon for a resulting 50 / 25 / 25 V / L / S ratio in the expressed protein, the third position encoded a codon for a resulting 50 / 25 / 25 Y / R / D ratio in the expressed protein, the fourth position encoded a codon for an equal ratio for all amino acids resulting in the expressed protein, and the fifth position encoded a codon for a resulting 75 / 25 G / P ratio in the expressed protein.

[0225] Example 13: Generation of Nucleic Acid Libraries by Sampling

[0226] Computational techniques were used to generate a population of nucleic acids having a preselected distribution. An exemplary preselected distribution is provided in Table 8 below, where the numbers represent the desired proportion of each amino acid at each position. Cumulative distribution values were first calculated, yielding values from 0.0 to 1.0 as seen in Table 9. In a program such as Excel, a uniform random number generator was used to create values between 0 and 1 for each of the 10 amino acid positions for 500 nucleic acids used as the sampling population. For example, in the case of position 1, a uniform random value of "0.95" belongs to the "S" bucket, so the amino acid "S" is displayed. This technique is called "roulette wheel" selection. Ten random numbers were generated from 10 discrete distributions for each designed oligonucleotide; this process was repeated 500 times to generate a sample population of 500 nucleic acids. Subsequently, to verify the generated sample population, the sum across the population of the frequency with which each amino acid appears at that position was determined and expressed as a proportion. For example, the proportion of amino acid C appearing at position 1 in a sample of 500 nucleic acids was calculated. The value represents the approximate distribution in the population. Using a sufficient number of nucleic acids in the population, the sample distribution was close to the preselected distribution.

[0227]

Table 8

[0228]

Table 9

[0229] Example 14: Generation of a Nucleic Acid Library by Filtered Sampling

[0230] Using the method described in Example 13, resampling of the population was performed to remove unwanted combinations and filter them out of the population. For example, combinations having four "H" (histidine) amino acids at any position were considered unsuitable for biological use. Thus, in this example, if the 500th oligonucleotide was generated as "HHHCCHHCHH (SEQ ID NO:55)", the combination was not desirable because it had eight Hs. As a result, other randomly generated combinations were generated in that location according to the method described in Example 13. Any number of criteria were used to generate a preselected distribution. For example, the population was generated such that each oligonucleotide at any position contained at least one "A" (alanine) amino acid. Further, the population was generated such that the generated combinations did not have two "M" (methionine) amino acids adjacent to each other. Thus, random sampling was performed until a preselected distribution and specific criteria were met.

[0231] Example 15: Combinatorial Library with Uniform Distribution

[0232] De novo polynucleotide synthesis was performed under conditions similar to those described in Example 2. As in Examples 4-6 and 8-12, nucleic acid populations were generated in which variants were preselected at each position and codon mutations were encoded at one or more sites having a preselected distribution.

[0233] To generate a uniform mutant distribution library by a combinatorial method, the reference sequence of the mutant library was split into two parts. As used herein, a uniform mutant distribution is intended to mean that each mutant is synthesized in approximately equal amounts. One of the split sides was called the 5'-side and the other split side was called the 3'-side. The sequences were designed and synthesized for each side of the reference sequence such that when annealed, the desired nucleic acid library was synthesized. For a uniform library with variations similar to Table 10, the diversity on the 5'-side is 2548 (14×14x13). On the 3'-side, the diversity is 546 (3×13x14). The 5'- and 3'-sides were synthesized by annealing, resulting in a total diversity of 1,391,208 (2548×546). The mutants were analyzed by next-generation sequencing (data not shown).

[0234]

Table 10

[0235] Example 16: Combinatorial Library with Non-uniform Distribution

[0236] De novo polynucleotide synthesis was performed under conditions similar to those described in Example 2. As in Examples 4-6 and 8-12, mutants were pre-selected at each position and a nucleic acid population encoding codon mutations at one or more sites with a pre-selected distribution was generated.

[0237] Libraries with non-uniform mutant distributions were also generated with a pre-selected distribution similar to that seen in Table 11. The reference sequence was again split in half, and mutants were generated for each part. One of the split sides was called the 5' side, and the other split side was called the 3' side. The expected probabilities of the 5' mutants and 3' mutants were calculated by multiplying the theoretical frequency of substitution of those mutants. For example, for the 5' mutants of the NRS sequence, the expected probability was 0.0677% (9.9% x 7.6% x 9.0%). For the 5' mutants and 3' mutants, some of the mutants had the same probability, and they were grouped (i.e., in the "bins" of the same probability). Thus, all mutants within the same bin have the same theoretical frequency of occurrence. For the total of 1,391,208 theoretical mutants, there were 162 different probabilities, and thus 162 different probability bins.

[0238]

Table 11

[0239] Subsequently, next-generation sequencing (NGS) was performed to determine to what extent the theoretical diversity was represented in the generated mutants. Since the sequencing was performed with 10^6 reads, only 30% of the actual diversity was observed. Thus, the total actual diversity represented at the desired frequencies was determined.

[0240] The 162 different probability bins representing the number of mutants with the same frequency were used to analyze the NGS data. For the 162 different probability bins, the reads from NGS were grouped by their expected probabilities of occurrence (dotted lines) as seen in Figure 22. Subsequently, the observed frequencies (solid lines) were compared to the expected probabilities. For each of the 162 bins, the observed frequency was determined by the total number of mutants divided by the number of mutants in that bin. This value was calculated for each bin and represented as an average as seen in Figure 23. As seen in Figure 22, these values were graphed as observed frequencies and compared to the expected probabilities.

[0241] The comparison of the observed frequency of mutants (solid line) and the expected probability of mutants (dotted line) as shown in Fig. 22 indicates whether the observed diversity was represented at the desired frequency. As seen in Fig. 22, the observed diversity was in good agreement with the expected probability, and more than 99% of the theoretical diversity was represented.

[0242] In addition, high-frequency combinations were observed in the same way as pre-determined low-frequency combinations. 89.9% of the NGS reads spanning the 39 base pair region of diversity were of appropriate size, and it was estimated that more than 70% of the complete 126 base pair constructs contained no insertions and deletions. Referring to Fig. 24, a high proportion of full-length fragments was generated as indicated by a single peak.

[0243] Example 17: Combinatorial library containing 144 single codon mutants and 9072 overlapping codon mutants at each of 8 positions

[0244] De novo polynucleotide synthesis was carried out under the same conditions as described in Example 2. A nucleic acid population was generated in the same way as in Examples 4 - 6 and 8 - 12. The nucleic acid population contained 144 single codon mutants and 9072 overlapping codon mutants (9216 diversities), and the mutants were pre-selected at 8 positions.

[0245] Subsequently, next-generation sequencing (NGS) was carried out to determine the distribution of mutants of the observed combinations. Sequencing was carried out in an application range exceeding 10^5 reads. As seen in Fig. 25, more than 99% of the mutants observed by NGS with a uniform distribution were detected. More than 90% of the observed mutants contained no insertions and deletions, and less than 5% of off-target sequences were detected. Less than 1% of the wild-type sequence was observed.

[0246] Example 18: Generation of a representative mutant library using an array-based method

[0247] A mutant library was de novo synthesized using an array-based method similar to Examples 1-3. Subsequently, the mutant library generated using the array-based method was compared with the mutant library generated using the PCR-based method.

[0248] After constructing the mutant library, colonies from the two libraries were sampled and sequenced. The data are shown in Table 12. The number of failed sequencing (“number of failed sequencing”) was determined as the number of colonies that could not be sequenced. The percentage of diversity (diversity (%)) was determined from the ratio of the number of mutants obtained after sequencing to the number of theoretically possible mutants expected. The percentage of accuracy (“accuracy (%)) was determined by the ratio of the number of mutants with the correct DNA sequence to the number of mutants used for sequencing. From Table 12, the mutant library generated using the array-based method showed higher “accuracy” and was correlated with the improvement of diversity and quality.

[0249] The two libraries were also compared at the protein level by sampling. The mutant library generated using the array-based method had a more representative mutant population, and the number of generated mutants was increased compared to that theoretically expected from the mutant library generated using the PCR-based method.

[0250]

Table 12

[0251] Example 19: Codon Assignment Scheme

[0252] A polynucleotide library was designed using codon assignment. Codon assignment was used to determine the codon sequence designed at each site.

[0253] Codon mutations were generated for human tumor protein p53 (TP53) having a wild-type (WT) amino acid sequence and WT DNA sequence as listed in Table 13. When generating codon mutations, the designed mutant codon sequences were based on the codon assignments in Table 3 above. Specifically, when generating mutant amino acids from the WT amino acids, the mutant codon sequences encoding the mutant amino acids were first selected from the codon sequences listed in Table 3 from left to right.

[0254] Referring to Table 13, the WT amino acid at position 2 of the peptide is "F" (bold). To generate a mutation at position 2, a mutant of the WT sequence was designed to change "F" to any of the other 19 amino acids. Then, using the codon assignments in Table 3, it was determined which mutant codon sequences to design to generate the mutant amino acids at that position. To generate a mutant in which "F" is changed to "A", the first selected mutant codon sequence in Table 3 was "GCT" instead of "GCA", "GCC", or "GCG" which encode "A". Table 14 lists all possible mutant amino acids for "F" at position 2 and which mutant codon sequences were designated to generate the mutant amino acids.

[0255]

Table 13-1

[0256]

Table 13-2

[0257]

Table 14

[0258] Example 20: A series in CDRs with multiple mutant sites

[0259] A nucleic acid library is generated as in Examples 4-6 and 8-12, and the variants encode codon mutations at one or more sites where the variant is preselected at each position. The region of the variant encodes at least a portion of the CDR. For example, refer to FIG. 12. The synthesized nucleic acid is released from the surface of the device and used as a primer to generate a nucleic acid library, which is expressed in cells to generate a mutant protein library. The mutant antibodies are evaluated for increased binding affinity to the epitope.

[0260] Example 21: Generation of a Mutant Antibody Library

[0261] As in the above examples, a nucleic acid library is generated. For the nucleic acids encoding representative CDRs from FIG. 12, a mutant library was generated. The representative CDRs were modified to include multiple positions for mutations such that the CDR region is as seen in FIG. 13. As shown in FIG. 13, different numbers of codon variants and positions of the variants are selected. In FIG. 13, the diversity of the mutant library that can be created is 1,152. Analysis by next-generation sequencing indicates the presence of the intended variants at the correct fraction and correct positions.

[0262] Example 22: Modular Plasmid Components for Expressing Diverse Peptides

[0263] As depicted in FIG. 14, a nucleic acid library is generated as in Examples 4-6 and 8-12, and for each of the separate regions that make up a part of the expression construct cassette, codon mutations are encoded at one or more sites. To generate two construct expression cassettes, mutant nucleic acids were synthesized that encode at least a portion of the variant sequences of the first promoter (1410), the first open reading frame (1420), the first terminator (1430), the second promoter (1440), the second open reading frame (1450), or the second terminator sequence (1460). After repeated amplification, a library of 1,024 expression constructs is generated as described in the previous examples.

[0264] Example 23: Variants at Multiple Sites, One Position

[0265] A nucleic acid library is generated as in Examples 4-6 and 8-12, encoding codon mutations at one or more sites in a region encoding at least a portion of the nucleic acid. A library of nucleic acid variants is generated, the library consisting of variants at multiple sites, one position. See, for example, FIG. 8B.

[0266] Example 24: Synthesis of Variant Libraries

[0267] De novo polynucleotide synthesis is performed under conditions similar to those described in Example 2. At least about 30,000 non-identical polynucleotides are synthesized de novo, where each of the non-identical polynucleotides encodes a codon mutation with a different amino acid sequence. The at least 30,000 non-identical polynucleotides synthesized are compared to a pre-determined sequence for each of the at least about 30,000 non-identical polynucleotides and have a total error rate of less than 1 in 1,000 bases. The library is used for PCR mutagenesis of long nucleic acids to form at least about 30,000 non-identical variant polynucleotides.

[0268] Example 25: Cluster-Based Variant Library Synthesis

[0269] Perform de novo polynucleotide synthesis under the same conditions as described in Example 2. Generate one cluster on the device, and the cluster contained predefined synthetic variants of the reference nucleic acid for two codon positions. In the arrangement of two consecutive codon positions, 19 variants per position were generated for two positions with two replicates of each nucleic acid, resulting in 38 synthesized nucleic acids. Each variant sequence is 40 bases in length. In the same cluster, additional non-variant and variant nucleic acids generate an additional non-variant nucleic acid sequence that collectively encodes the 38 variants of the gene's coding sequence. Each of the nucleic acids has at least one region that is complementary to another nucleic acid. Release the nucleic acids in the cluster by gaseous ammonia cleavage. A water-containing pin contacts the cluster, picks up the nucleic acids, and transfers the nucleic acids to a small vial. The vial further contains a DNA polymerase reagent for a polymerase cycling assembly (PCA) reaction. Anneal the nucleic acids, fill the gaps by an extension reaction, and form the resulting double-stranded DNA molecules to form a variant nucleic acid library. Optionally expose the variant nucleic acid library to a restriction enzyme and then ligate it to an expression vector.

[0270] Example 26: Screening of Variant Nucleic Acid Libraries for Changes in Protein Binding Affinity

[0271] Generate a plurality of expression vectors as described in Examples 13-16. In this example, the expression vector is a HIS-tagged bacterial expression vector. Electroporate the vector library into bacterial cells and then select clones for the expression and purification of the HIS-tagged variant protein. Screen the variant protein for changes in binding affinity to the target molecule.

[0272] The affinity is examined by methods using immobilized metal ion affinity chromatography (IMAC), etc., in which a metal ion-coated resin (e.g., IDA agarose or NTA agarose) is used to isolate the HIS-tagged protein. A string of histidine residues binds to various types of immobilized metal ions including nickel, cobalt, and copper under specific buffer conditions, enabling the purification and detection of the expressed His-tagged protein. An example binding / washing buffer consists of Tris-buffered saline (TBS) at pH 7.2 containing 10 - 25 mM imidazole. Elution and recovery of the captured HIS-tagged protein from the IMAC column are achieved with high concentrations of imidazole (at least 200 mM) (eluent), low pH (e.g., 0.1 M glycine-HCl, pH 2.5), or an excess of a strong chelating agent (e.g., EDTA).

[0273] Alternatively, anti-His tag antibodies are commercially available for use in assays for His-tagged proteins, such as pull-down assays to isolate His-tagged proteins or immunoblotting assays to detect His-tagged proteins.

[0274] Example 27: Screening of a mutant nucleic acid library for changes in activity against regulators of cell adhesion and migration

[0275] Insert the mutant nucleic acid library generated as described in Examples 13-16 into a mammalian expression vector tagged with GFP. Clone isolated from the library is transiently transfected into mammalian cells. Alternatively, express the protein, isolate from the cells containing the expression construct, and then deliver the protein to the cells for further measurement. Perform an immunofluorescence assay to evaluate changes in the cellular localization of the GFP-tagged mutant expression product. Perform a FACS assay to evaluate changes in the three-dimensional structure of a transmembrane protein that interacts with the non-mutant version of the GFP-tagged mutant protein expression product. Perform a wound healing assay to evaluate changes in the ability of cells expressing the GFP-tagged mutant protein to invade the space created by a scratch on a cell culture dish. Identify cells expressing the GFP-tagged protein and track them using a fluorescence source and a camera.

[0276] Example 28: Screening of a Mutant Nucleic Acid Library for Peptides that Inhibit Viral Progression

[0277] Insert the mutant nucleic acid library generated as described in Examples 13-16 into a mammalian expression vector tagged with FLAG, and the mutant nucleic acid library encodes a peptide sequence. Obtain primary mammalian cells from a subject suffering from a viral disorder. Alternatively, infect primary cells obtained from a healthy subject with a virus. Seed the cells onto a series of microwell dishes. Clone isolated from the mutant library is transiently transfected into the cells. Alternatively, express the protein, isolate from the cells containing the expression construct, and then deliver the protein to the cells for further measurement. Perform a cell survival assay to evaluate the infected cells for enhanced survival associated with the mutant peptide. Exemplary viruses include, but are not limited to, avian influenza, Zika virus, hantavirus, hepatitis C, and smallpox.

[0278] One exemplary assay is the neutral red cytotoxicity assay that uses neutral red dye, which diffuses into the plasma membrane when added to cells and accumulates in acidic lysosomal compartments due to the mild cationic nature of neutral red. Virus-induced cytopathic effects cause membrane fragmentation and loss of lysosomal ATP-driven proton translocation activity. The resulting decrease in neutral red within cells can be evaluated by spectrophotometry in the format of a multiwell plate. Cells expressing mutant peptides are scored by an increase in intracellular neutral red in a gain-of-signal color assay. Cells are evaluated for peptides that inhibit virus-induced cytopathic effects.

[0279] Example 29: Screening for Mutant Proteins that Increase or Decrease Cellular Metabolic Activity

[0280] To identify expression products that result in changes in cellular metabolic activity, multiple expression vectors are generated as described in Examples 13-16. In this example, the expression vectors are transferred (e.g., via transfection or transduction) to cells seeded on a series of microwell dishes. The cells are then screened for one or more changes in metabolic activity. Alternatively, the protein is expressed, isolated from cells containing the expression construct, and then delivered to cells to measure metabolic activity. Optionally, cells for measuring metabolic activity are treated with a toxin and then screened for one or more changes in metabolic activity. Exemplary toxins administered include, but are not limited to, botulinum toxins (including immunological types: A, B, C1, C2, D, E, F, and G), staphylococcal enterotoxin B, Yersinia pestis, hepatitis C, mustard agent, heavy metals, cyanide, endotoxin, Bacillus anthracis, Zika virus, avian influenza, herbicides, pesticides, mercury, organophosphates, and ricin.

[0281] The basal energy requirement is obtained from the oxidation of metabolic substrates (e.g., glucose) by either oxidative phosphorylation, which includes aerobic tricarboxylic acid (TCA) or the Krebs cycle, or anaerobic glycolysis. When glycolysis is the main energy source, the metabolic activity of cells can be estimated by monitoring the rate at which cells secrete acidic metabolites (e.g., lactate and CO2). In the case of aerobic metabolism, extracellular oxygen consumption and the production of oxidative free radicals reflect the energy requirement of the cells. The intracellular redox potential can be measured by the autofluorescence measurement of NADH and NAD + The amount of energy released by cells (e.g., heat) can be obtained from the analytical values for substances that are generated and / or consumed during metabolism, which can be predicted from the amount of oxygen consumed (e.g., 4.8 kcal / l of O2) under standard settings. The coupling between heat production and oxygen utilization can be disrupted by toxins. Direct microcalorimetry measures the temperature increase of a sample isolated by heat. Therefore, when combined with the measurement of oxygen consumption, calorimetry can be used to detect the uncoupling activity of toxins.

[0282] Various methods and apparatuses for measuring changes in various markers of metabolic activity are known in the art. For example, such methods, apparatuses, and markers are discussed in U.S. Patent No. 7,704,745, which is hereby incorporated by reference in its entirety. Briefly, the measurement of any of the following characteristics is recorded for each cell population: glucose, lactate, CO2, the ratio of NADH to NAD + Heat, O2 consumption, and free radical production. The cells to be screened may include hepatocytes, macrophages, or neuroblastoma cells. The cells to be screened may be cell lines, primary cells from a subject, or cells from a model system (e.g., a mouse model).

[0283] For measuring the oxygen consumption rate of a single cell or a population of cells in a multi-well chamber, various techniques are available. For example, the chamber containing the cells can be equipped with sensors for recording changes in temperature, current, or fluorescence, as well as an optical system (e.g., a fiber-coupled optical system) coupled to each chamber for monitoring fluorescence. In this example, each chamber has a window through which an illumination source stimulates molecules inside the chamber. The fiber-coupled optical system can detect autofluorescence for measuring the NADH / NAD ratio and voltage inside the cells, as well as calcium-sensitive dyes for determining membrane potential differences and intracellular calcium. In addition, changes in the fluorescence dye signal sensitive to CO2 and / or O2 are detected.

[0284] Example 30: Screening of a mutant nucleic acid library for selective targeting of cancer cells

[0285] A mutant nucleic acid library generated as described in Examples 13-16 is inserted into a FLAG-tagged mammalian expression vector, and the mutant nucleic acid library encodes a peptide sequence. Clones isolated from the mutant library are transiently transfected separately into cancer cells and non-cancer cells. Cell survival and cell death assays are performed on both cancer cells and non-cancer cells, each of which expresses a mutant peptide encoded by the mutant nucleic acid. The cells are evaluated for selective cancer cell death associated with the mutant peptide. The cancer cells are optionally cancer cell lines or primary cancer cells from a subject diagnosed with cancer. In the case of primary cancer cells from a subject diagnosed with cancer, the mutant peptide identified in the screening assay is optionally selected for administration to the subject. Alternatively, the protein is expressed, isolated from cells containing the protein expression construct, and then the protein is delivered to cancer cells and non-cancer cells for further measurement.

[0286] Example 31: Generation of a combinatorial library

[0287] Perform de novo polynucleotide synthesis under the same conditions as described in Example 2. Generate nucleic acid populations as in Examples 4-6 and 8-12, where the variants encode codon mutations at one or more sites where the variants are preselected at each position. Generate a combinatorial library by combining the nucleic acids of the first population with the nucleic acids from the second population. As shown in Figure 1, four populations of nucleic acids (110) are combined with another four populations of nucleic acids (120) to yield 16 combinations.

[0288] Anneal the nucleic acids by blunt-end ligation. Mix 50 ng of the DNA of one nucleic acid with 50 ng of the DNA of the other nucleic acid in a 1.5 mL vial. Next, add 1 μL of T4 DNA ligase (New England BioLabs) together with 20 μL of ligation buffer and 20 μL of nuclease-free water. Then, incubate the reaction mixture. After incubation, analyze the ligation products by sequencing.

[0289] Example 32: Generation of Combinatorial Library by Sampling

[0290] Perform de novo polynucleotide synthesis under the same conditions as described in Example 2. Generate nucleic acid populations as in Examples 4-6 and 8-12, where the variants encode codon mutations at one or more preselected sites at each position.

[0291] Referring to FIG. 26A, a library with a non-uniform mutant distribution was generated in a pre-selected distribution by a method similar to that described in Examples 13-16. Each pattern portion of the image represents one of four different pre-selected distributions of amino acids that differ at each position (A1, A2, A3, B1, B2, and B3). The black circles represent random selections within each position. Referring to FIG. 26B, five randomly generated samples for A and five randomly generated samples for B are generated independently. Then, as seen in FIG. 26C, the five randomly generated samples for A and the five randomly generated samples for B are annealed together, for example, by blunt-end ligation. This results in 25 combinations (n^2 = 5^2). Referring to FIG. 26D, a statistical comparison shows that the resulting distribution matches the pre-selected distribution.

[0292] Example 33: Generation of a Combinatorial Antibody Library

[0293] As in the above examples, a nucleic acid library is generated. A mutant library is generated for nucleic acids encoding one CDR region as seen in FIG. 27A, two CDR regions as seen in FIG. 27B, or multiple CDR regions as seen in FIG. 27C.

[0294] A mutant antibody library is also generated to include mutants in single or multiple heavy chain scaffolds and light chain scaffolds as seen in FIG. 28A, or mutants in single or multiple frameworks as seen in FIG. 28B.

[0295] Preferred embodiments of the present invention have been shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art will envision many changes, variations, and substitutions without departing from the present invention. It is to be understood that various alternatives to the embodiments of the invention described herein may be utilized in practicing the invention. The following claims define the scope of the present invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

1. A method for synthesizing a mutant nucleic acid library, said method comprising: a. providing a predefined nucleic acid sequence of a plurality of polynucleotides, wherein said plurality of polynucleotides encode a plurality of codons having mutant codon sequences as compared to a reference sequence of a single nucleic acid, and the codon assignment is used to determine each codon of the plurality of codons; b. selecting a distribution value for a codon at a predefined position in a predefined nucleic acid reference sequence; c. providing a machine instruction for randomly generating a set of nucleic acid sequences having a distribution value approximately equal to said distribution value, wherein the selected distribution value and the distribution value randomly generated from said set of nucleic acid sequences are each independently the probability that a codon occurs at a predefined position within the sequence, and the number of nucleic acids within the set of nucleic acid sequences is less than the number of nucleic acid sequences required to generate a saturated codon mutant library; d. synthesizing a mutant nucleic acid library based on said set of nucleic acid sequences, said step representing at least about 70% of the predicted diversity; A method comprising the above steps.

2. The method according to claim 1, wherein said mutant nucleic acid library is translated into a protein library.

3. The method according to claim 1, further comprising performing PCR mutagenesis of a nucleic acid using the mutant nucleic acid library as a primer for a PCR mutagenesis reaction.

4. The method according to claim 1, wherein said codon assignment is based on the frequency of codon sequences in an organism.

5. The method according to claim 4, wherein said organism is an animal, a plant, a fungus, a protist, an archaebacterium, or a bacterium.

6. The method according to claim 1, wherein the assignment of the codons used to determine the mutant codon array is based on the complexity of the codon array.

7. The method according to claim 1, wherein the mutant nucleic acid library encodes at least a part of an antibody, an enzyme, or a peptide.

8. The method according to claim 1, wherein the mutant nucleic acid library encodes at least a part of a variable region or a constant region of an antibody.

9. The method according to claim 1, wherein the mutant nucleic acid library encodes at least one CDR region of an antibody.

10. The method according to claim 1, wherein the mutant nucleic acid library encodes CDR1, CDR2, and CDR3 on the heavy chain of an antibody and CDR1, CDR2, and CDR3 on the light chain.

11. The method according to claim 1, wherein the probability is determined as the ratio at which each mutant codon of the plurality of codons appears at the previously selected position.

Citation Information

Patent Citations

  • JPP7335165B

  • Protein design automation for protein libraries

    US20030130827A1

  • Libraries and their design and assembly

    US20080287320A1