Enzymes for library preparation

Variant polypeptides with specific mutations are developed to address enzyme optimization challenges in sequencing applications, enhancing nucleic acid library preparation and sequencing performance.

JP2025541124APending Publication Date: 2025-12-18TWIST BIOSCIENCE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025532619
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-05
Filing Date
2023-12-05
Publication Date
2025-12-18

AI Technical Summary

Technical Problem

Enzyme design and implementation for chemical biology applications, such as sequencing, are challenging due to difficulties in optimizing enzyme properties.

Method used

Development of variant polypeptides with specific amino acid mutations and nucleic acids encoding these polypeptides, which are used to form covalent bonds between nucleotides and prepare nucleic acid libraries for sequencing, including methods for expression and purification.

Benefits of technology

The variant polypeptides enhance the efficiency and effectiveness of nucleic acid library preparation, improving sequencing performance through optimized enzyme activity and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025541124000001_ABST
    Figure 2025541124000001_ABST
Patent Text Reader

Abstract

Provided herein are methods and compositions relating to enzymes and libraries having nucleic acids encoding the enzymes that include modified sequences. Further provided herein are methods for enzyme optimization. Further provided herein are enzymes for sequencing library generation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 386,143, filed December 5, 2022, which is incorporated herein by reference in its entirety. All publications, patents, and patent applications mentioned herein are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Background technology]

[0002] Enzymes have the ability to catalyze a wide range of chemical reactions, including those used in chemical biology for sequencing applications. However, enzyme design and implementation can be difficult. Therefore, there is a need to develop compositions and methods for optimizing enzyme properties. Summary of the Invention

[0003] Provided herein are methods, compositions, and systems for enzyme optimization.

[0004] Provided herein is a variant polypeptide comprising at least one amino acid mutation relative to SEQ ID NO: 1. Further provided herein is a variant polypeptide, wherein the polypeptide comprises at least 80% similarity to any one of SEQ ID NOs: 2-3. Further provided herein is a variant polypeptide, wherein the polypeptide comprises at least 90% similarity to any one of SEQ ID NOs: 2-3. Further provided herein is a variant polypeptide, wherein the polypeptide comprises at least 95% similarity to any one of SEQ ID NOs: 2-3. Further provided herein is a variant polypeptide, wherein the polypeptide comprises at least 98% similarity to any one of SEQ ID NOs: 2-3. Further provided herein is a variant polypeptide, wherein the polypeptide comprises any one of SEQ ID NOs: 2-3. Further provided herein is a variant polypeptide, wherein the polypeptide comprises at least 10 consecutive amino acids of any one of SEQ ID NOs: 2-3. Further provided herein is a variant polypeptide, wherein the polypeptide comprises at least 20 consecutive amino acids of any one of SEQ ID NOs: 2-3. Further provided herein are variant polypeptides, wherein the polypeptide comprises 20 to 100 consecutive amino acids of any one of SEQ ID NOs: 2-3. Further provided herein are variant polypeptides, wherein the polypeptide comprises at least two amino acid mutations relative to SEQ ID NO: 1. Further provided herein are variant polypeptides, wherein the polypeptide comprises at least four amino acid mutations relative to SEQ ID NO: 1. Further provided herein are variant polypeptides, wherein the polypeptide comprises at least six amino acid mutations relative to SEQ ID NO: 1. Further provided herein are variant polypeptides, wherein the mutations are at one or more of positions E88, T91, V119, G128, E168, Q223, L231, L293, V372, E440, D448, and E483 relative to SEQ ID NO: 1. Further provided herein are variant polypeptides, wherein the mutations are at one or more of positions E88, V119, G128, E168, Q223, L231, L293, and E440 relative to SEQ ID NO: 1.Further provided herein is a variant polypeptide wherein the mutation is at one or more of positions E88, V119, Q223, L293, V372, and E483 relative to SEQ ID NO: 1. Further provided herein is a variant polypeptide wherein the mutation is selected from one or more of E88K, T91M, V119R, G128K, E168K, Q223K, L231A, L293E, V372I, E440K, D448W, D448P, and E483K relative to SEQ ID NO: 1. Further provided herein is a variant polypeptide wherein the mutation is selected from one or more of E88K, V119R, G128K, E168K, Q223K, L231A, L293E, and E440K relative to SEQ ID NO: 1. Further provided herein is a variant polypeptide, wherein the mutation is selected from one or more of E88K, V119R, Q223K, L293E, V372I, and E483K relative to SEQ ID NO: 1. Further provided herein is a variant polypeptide, wherein the polypeptide further comprises a purification tag.

[0005] Provided herein are nucleic acids encoding the polypeptides described herein. Further provided herein are nucleic acids comprising at least 80% similarity to any one of SEQ ID NOs: 4-5, provided that the polypeptide encodes the polypeptide of SEQ ID NO: 1. Further provided herein are nucleic acids comprising at least 90% similarity to any one of SEQ ID NOs: 4-5. Further provided herein are nucleic acids comprising at least 95% similarity to any one of SEQ ID NOs: 4-5.

[0006] Provided herein are vectors comprising the nucleic acids described herein. Further provided herein are vectors comprising plasmids. Provided herein are cells comprising the nucleic acids described herein. Further provided herein are cells, including bacterial cells.

[0007] Provided herein are methods for expressing the polypeptides described herein. Further provided herein are methods wherein expression comprises translation of the nucleic acid sequences provided herein. Further provided herein are methods including in vivo methods. Further provided herein are methods including cell-free methods.

[0008] Provided herein is a method for forming a covalent bond between two nucleotides, comprising contacting a first nucleotide and a second nucleotide with a polypeptide disclosed herein. Further provided herein is a method wherein the first nucleotide and the second nucleotide are present on the same nucleic acid. Further provided herein is a method wherein the covalent bond forms a circular nucleic acid. Further provided herein is a method wherein the first nucleotide is present on a first nucleic acid and the second nucleotide is present on a second nucleic acid. Further provided herein is a method wherein the first nucleic acid and / or the second nucleic acid comprises genomic DNA or a fragment thereof. Further provided herein is a method wherein the first nucleic acid and / or the second nucleic acid comprises cDNA. Further provided herein is a method wherein the first nucleic acid and / or the second nucleic acid comprises an adaptor. Further provided herein is a method wherein the first nucleic acid comprises a first adaptor and genomic DNA or cDNA. Further provided herein is a method wherein the second nucleic acid comprises a second adaptor. Further provided herein is a method wherein the adaptor comprises at least one barcode. Further provided herein is a method, wherein the barcode comprises one or more of a sample index, a plate index, a cell index, and a unique molecular identifier.

[0009] Provided herein is a method for preparing a nucleic acid library, comprising: (a) providing one or more sample nucleic acids; (b) contacting the one or more sample nucleic acids with a plurality of adaptors and polypeptides disclosed herein to form a nucleic acid sequencing library comprising adaptor-ligated nucleic acids; and (c) sequencing the nucleic acid library. Further provided herein is a method wherein the sample nucleic acid comprises a genomic fragment. Further provided herein is a method wherein the genomic fragment is obtained by genomic cleavage or amplification. Further provided herein is a method wherein the sample nucleic acid comprises cDNA. Further provided herein is a method wherein the sample nucleic acid comprises cfDNA. Further provided herein is a method further comprising one or more of end repair, A-tailing, and amplification. Further provided herein is a method further comprising concentrating the nucleic acid library prior to sequencing. [Brief explanation of the drawings]

[0010] [Figure 1] 1 depicts an automated workflow for optimizing ligation enzymes. [Figure 2A] Depicting a strategy for the design of ligase variants using MSA from high-entropy positions. Figure 2A depicts a plot of the cumulative probability of amino acids (0.0–1.0 in a unit interval of 0.2) versus their position in T4 ligase (left to right: 212–214, 222–224, 272–274, 296–298, 308–310). [Figure 2B] Depicting a strategy for the design of ligase variants using MSA from high-entropy positions. Figure 2B depicts a plot of Shannon entropy (0.0–3.0 in a unit interval of 0.5) versus position in T4 ligase (left to right: 212–214, 222–224, 272–274, 296–298, 308–310). [Figure 3A] Depicts the SYBR Green qPCR workflow for high-throughput cell-free screening and quantification of T4 ligases. [Figure 3B]Depict the amplification plot obtained from the workflow in Figure 3A. The y-axis is labeled RFU (0-4000 fluorescence units, in unit intervals of 1000), and the x-axis is labeled PCR cycles (0-40 in unit intervals of 10). [Figure 4] Plots from the first round of single variant screening are shown, from left to right: activity, thermostability, and salt. [Figure 5A] Heatmap from the first round of screening of single variants. Red indicates higher activity, and blue indicates lower activity. The legend indicates colors corresponding to activity from -2 to 3 in unit intervals of 1. [Figure 5B] Heatmap from the second round of screening using binary combinations of single variants to measure epistatic effects. Blue indicates higher activity, red indicates lower activity. The legend indicates colors corresponding to activity (units in ct values) from -4 to 6 in unit intervals of 1. [Figure 6A] Draw the plots obtained from rounds 4 / 5 using the raw addition of a single variant. The x-axis is labeled activity from 10 to 18 in a unit interval of 1, and the y-axis is labeled percentage from 0.0 to 1.0 in a unit interval of 0.2. [Figure 6B] Plots obtained from rounds 4 / 5 using raw addition of a single variant are depicted. The x-axis is the number of variants labeled (3, 4, 5, 6 from left to right), and the y-axis is labeled activity (1 / (2^ct)) from 0.0000 to 0.0014 in a unit interval of 0.0002. [Figure 7] SDS-PAGE gel used to prepare molecular biology-grade ligase from variants. Lanes: (1) ladder, (2) lysate, (3) flow-through, (4) blank, (5 to 0) ligase. [Figure 8A] Depicting structural information for the design of T4 ligase variants. Numerous lysine mutations (positive charges) were observed near the DNA substrate of the variants. [Figure 8B]1 depicts structural information for the design of T4 ligase variants. Residues that contact the DNA substrate are boxed and include positions 14, 15, 16, 39, 44, 46, 48, 49, 79, 82, 84, 116, 118, 119, 120, 121, 124, 157, 159, 164, 181, 182, 185, 217, 254, 258, 262, 263, 266, 268, 282, 361, 380, 382, ​​383, 384, 404, 406, 407, 410, 411, 412, 447, 448, 450, 455, 457, 458, 459, and 460. [Figure 9A] Plot of percent chimerism for a series of variant T4 ligases. The y-axis is labeled percent chimerism from 0.000 to 0.030 in unit intervals of 0.005. Variants from various selection rounds are labeled on the x-axis: 38, 6_1, 6_16, 6_8, 7_1, 7_10, 7_11, 7_12, 7_13, 7_14, 7_15, 7_16, 7_18, 7_19, 7_2, 7_20, 7_3, 7_6, 7_8, 7_9, AZ, E12, Qiagen, and WT. [Figure 9B] 1 depicts an example of a chimera formed from two biological sequences. [Figure 10A] Draw a 2D plot of chimeras versus activity. The y-axis is labeled 21-29 chimeras in a unit interval of 1. The x-axis is labeled 10.0-30.0 activity in a unit interval of 2.5. The legend is labeled data (blue), ngs sample (orange), green (low chimeras), red (seq38), wt (purple), and singleton (brown). [Figure 10B] A 2D plot of chimera versus activity is drawn. The y-axis is labeled as chimeras 21–26 in unit intervals of 1. The x-axis is labeled as activity 12–20 in unit intervals of 1. The heatmap legend labels adaptor-only CTs from 8 (dark blue) to 18 (light blue) in unit intervals of 2. Sequence 38 is shown. [Figure 10C] A single-site mutagenesis library of part of the T4 ligase sequence is shown. [Figure 11A]Plot of variant performance for sequence 38 across four NGS runs. The y-axis is labeled 0.0 to 1.0 variants / total reads for 38 with a unit interval of 0.2. The x-axis represents variants (left to right): r7r-18, r7r-24, r7r-12, r7r-8, r7r-21, 6-8, r7r-22, r7r-2, r7r-34, 6-1, r7r-5, r7r-1, 7-19, r7r-15, r7r-25, r7r-13, r7r-19, r7r-20, r7r-32 , r7r-4, r7r-29, r7r-6, r7r-17, r7r-26, r7r-30, r7r-3, r7r-23, r7r-9, r7r-33, r7r-11, 7-28, r7r-14, r7r-7, r7r-10, r7r-16, r7r-31, r7r-28, r7r-27, and wt are labeled. [Figure 11B] Plot of variant chimeras for sequence 38 across four NGS runs. The y-axis is labeled 0.0 to 1.0 variants / 38% chimeras in 0.2 unit intervals. The x-axis is variant ID (left to right): r7r-18, r7r-24, r7r-12, r7r-8, r7r-21, 6-8, r7r-22, r7r-2, r7r-34, 6-1, r7r-5, r7r-1, 7-19, r7r-15, r7r-25, r7r-13, r7r-19, r7r-20, r7r-3 2, r7r-4, r7r-29, r7r-6, r7r-17, r7r-26, r7r-30, r7r-3, r7r-23, r7r-9, r7r-33, r7r-11, 7-28, r7r-14, r7r-7, r7r-10, r7r-16, r7r-31, r7r-28, r7r-27, and wt are labeled. [Figure 12A]Draw a plot of total reads for all titrations of enzyme amount in the ligase experiment. The y-axis is labeled 0-800,000 total reads in unit intervals of 200,000. The x-axis is labeled variant and amount (ng) (left to right): 12-1000, 12-500, 12-250, 18-1000, 18-500, 18-250, 21-1000, 21-500, 21-500, 22-1000, 22-500, 22-250, 24-1000, 24-500, 24-250, 38-1000, 38-500, 38-250, 8-1000, 8-500, 8-250, wt-1000, wt-500, wt-250. [Figure 12B] Draw a plot of percent chimerism for all titrations of enzyme amount in the ligase experiment. The y-axis is labeled 0.000–0.016 total reads in unit intervals of 0.002. The x-axis is labeled variant and amount (ng) (left to right): 12–1000, 12–500, 12–250, 18–1000, 18–500, 18–250, 21–1000, 21–500, 21–500, 22–1000, 22–500, 22–250, 24–1000, 24–500, 24–250, 38–1000, 38–500, 38–250, 8–1000, 8–500, 8–250, wt–1000, wt–500, wt–250. [Figure 13] Plots of the sequencing performance of two variant T4 ligases and the wild type (from left to right: variant 21, variant 24, and wt) are drawn. The y-axis is labeled converted (normalized) reads from 0.00 to 2.00 in a unit interval of 0.25. Each set of three bars indicates enzyme mass / rxn (left to right: 250, 500, 1000 ng). DETAILED DESCRIPTION OF THE INVENTION

[0011] The present disclosure employs, unless otherwise indicated, conventional molecular biology techniques that are within the skill of the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0012] definition

[0013] Throughout this disclosure, various embodiments are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of any embodiment. Accordingly, the description of a range should be considered to have specifically disclosed all possible subranges and individual numerical values ​​within that range, to the tenth of the unit of the lower limit, unless the context clearly dictates otherwise. For example, description of a range such as 1 to 6 should be considered to have specifically disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual values ​​within that range, e.g., 1.1, 2, 2.3, 5, and 5.9. This applies regardless of the breadth of the range. The upper and lower limits of these intervening ranges may independently be included in the smaller ranges, which are also encompassed within the disclosure, subject to any specifically excluded limits in the stated ranges. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure, unless the context clearly dictates otherwise.

[0014] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit any embodiments. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly dictates otherwise. Furthermore, the terms "comprises" and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but are understood not to exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0015] Unless otherwise specified or clear from the context, as used herein, the term "about" in reference to a number or range of numbers is understood to mean the specified number and number + / - 10%, or 10% below the recited lower limit and 10% above the recited upper limit for the recited values ​​for a range.

[0016] Unless otherwise specified, the term "nucleic acid," as used herein, encompasses double- or triple-stranded nucleic acids as well as single-stranded molecules. In double- or triple-stranded nucleic acids, the nucleic acid strands need not be coextensive (i.e., a double-stranded nucleic acid need not be double-stranded along the entire length of both strands). Nucleic acid sequences, when provided, are listed in the 5' to 3' direction unless otherwise specified. The methods described herein provide for the production of isolated nucleic acids. The methods described herein further provide for the production of isolated and purified nucleic acids. A "nucleic acid" as referred to herein can comprise at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000 or more bases in length. Additionally, provided herein are methods for the synthesis of any number of polypeptide segments encoding nucleotide sequences, including sequences encoding nonribosomal peptides (NRPs), sequences encoding nonribosomal peptide synthetase (NRPS) modules and synthetic variants, polypeptide segments of other modular proteins such as antibodies, polypeptide segments from other protein families that contain regulatory sequences such as non-coding DNA or RNA, e.g., promoters, transcription factors, enhancers, siRNAs, shRNAs, RNAi, miRNAs, small nucleolar RNAs derived from microRNAs, or any functional or structural DNA or RNA unit of interest.The following are non-limiting examples of polynucleotides: coding or non-coding regions of genes or gene fragments, intergenic DNA, loci (locuses) defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), small nucleolar RNA, ribozymes, complementary DNA (cDNA), which is typically representative of the DNA of mRNA obtained by reverse transcription or amplification of messenger RNA (mRNA); DNA molecules produced synthetically or by amplification, genomic DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. The cDNA encoding the gene or gene fragment referred to herein may contain at least one region encoding an exon sequence without any intervening intron sequence in the genomic equivalent sequence.

[0017] Enzyme variants

[0018] Variants of enzymes are provided herein. In some examples, the enzyme comprises an enzyme for next-generation sequencing. In some examples, the enzyme comprises a ligase, polymerase, kinase, nuclease, phosphatase, methylase, topoisomerase, transferase, or other enzyme. In some examples, the enzyme comprises T4 ligase. In some examples, the T4 ligase is selected from Table 1. In some examples, the enzyme comprises a variant of SEQ ID NO: 1.

[0019] The enzymes provided herein may comprise one or more variants of SEQ ID NO: 1. In some examples, the variants comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, or at least 16 variant amino acid positions of SEQ ID NO: 1. In some examples, the variants comprise about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, or about 16 variant amino acid positions of SEQ ID NO: 1. In some examples, the enzyme comprises a mutation at one or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 relative to SEQ ID NO: 1. In some examples, the enzyme comprises a mutation at two or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 relative to SEQ ID NO: 1. In some examples, the enzyme comprises a mutation at three or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 relative to SEQ ID NO: 1. In some examples, the enzyme comprises a mutation at four or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 relative to SEQ ID NO: 1. In some examples, the enzyme comprises a mutation at five or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 relative to SEQ ID NO: 1. In some examples, the enzyme comprises a mutation at six or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 relative to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at seven or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 relative to SEQ ID NO:1.In some examples, the enzyme comprises a mutation at eight or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 relative to SEQ ID NO: 1. In some examples, the enzyme comprises a mutation at nine or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 relative to SEQ ID NO: 1. In some examples, the enzyme comprises a mutation at ten or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 relative to SEQ ID NO: 1.

[0020] In some examples, the enzymes provided herein comprise the amino acid sequence of any one of SEQ ID NOs: 2-3. In some examples, the enzymes provided herein comprise the nucleic acid sequence of any one of SEQ ID NOs: 5-6. The sequences provided herein, in some examples, comprise a purification tag. In some examples, the purification tag comprises a His6 tag.

[0021] [Table 1-1]

[0022] [Table 1-2] All sequences were expressed, in some cases, with a His6 tag (LEHHHHHH) at the C-terminus for purification purposes.

[0023] [Table 2-1]

[0024] [Table 2-2]

[0025] [Table 2-3]

[0026] [Table 2-4]

[0027] All sequences were expressed with a His6 tag (CTCGAGCACCACCACCACCACCACCAC) at the C-terminus for purification purposes in some cases.

[0028] The enzymes provided herein may comprise a sequence having homology or similarity to SEQ ID NO: 1. In some examples, the enzyme does not comprise SEQ ID NO: 1. In some examples, the enzymes provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO: 1. In some examples, at least 10 contiguous amino acids of the enzymes provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO: 1. In some examples, at least 50 contiguous amino acids of an enzyme provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO: 1. In some examples, at least 100 contiguous amino acids of an enzyme provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO: 1. In some examples, 20 to 100 contiguous amino acids of an enzyme provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO:1.

[0029] The enzymes provided herein may comprise a sequence having homology or similarity to SEQ ID NO: 2. In some examples, the enzymes provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO: 2. In some examples, at least 10 contiguous amino acids of the enzymes provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO: 2. In some examples, at least 50 contiguous amino acids of an enzyme provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO: 2. In some examples, at least 100 contiguous amino acids of an enzyme provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO: 2. In some examples, 20 to 100 contiguous amino acids of an enzyme provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO:2.

[0030] The enzymes provided herein may comprise a sequence having homology or similarity to SEQ ID NO: 3. In some examples, the enzymes provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO: 3. In some examples, at least 10 contiguous amino acids of the enzymes provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO: 3. In some examples, at least 50 contiguous amino acids of an enzyme provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO: 3. In some examples, at least 100 contiguous amino acids of an enzyme provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO: 3. In some examples, 20 to 100 contiguous amino acids of an enzyme provided herein comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or more similarity to SEQ ID NO:3.

[0031] The enzymes provided herein may include sequences having homology or similarity and mutations at one or more amino acid positions. In some examples, the enzyme includes a mutation at one or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 and at least 95% similarity to SEQ ID NO: 1. In some examples, the enzyme includes a mutation at two or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 and at least 90% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at three or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 90% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at four or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 90% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at five or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 90% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at six or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 90% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at seven or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 90% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at eight or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 90% similarity to SEQ ID NO: 1.In some examples, the enzyme comprises mutations at nine or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 90% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at ten or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 90% similarity to SEQ ID NO: 1.

[0032] In some examples, the enzyme comprises a mutation at one or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 80% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises a mutation at two or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 80% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises a mutation at three or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 80% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at four or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 80% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at five or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 80% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at six or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 80% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at seven or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 80% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at eight or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 80% similarity to SEQ ID NO: 1. In some examples, the enzyme comprises mutations at nine or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 with at least 80% similarity to SEQ ID NO: 1.In some examples, the enzyme comprises mutations at 10 or more positions selected from 88, 91, 119, 128, 168, 223, 231, 293, 372, 440, 448, or 483 relative to SEQ ID NO:1 and at least 80% similarity.

[0033] The enzymes provided herein may contain specific amino acid mutations. In some examples, the enzyme contains one or more mutations selected from E88K, T91M, V119R, G128K, E168K, Q223K, L231A, L293E, V372I, E440K, D448W, D448P, and E483K relative to SEQ ID NO: 1. In some examples, the enzyme contains two or more mutations selected from E88K, T91M, V119R, G128K, E168K, Q223K, L231A, L293E, V372I, E440K, D448W, D448P, and E483K relative to SEQ ID NO: 1. In some examples, the enzyme comprises three or more mutations selected from E88K, T91M, V119R, G128K, E168K, Q223K, L231A, L293E, V372I, E440K, D448W, D448P, and E483K relative to SEQ ID NO: 1. In some examples, the enzyme comprises five or more mutations selected from E88K, T91M, V119R, G128K, E168K, Q223K, L231A, L293E, V372I, E440K, D448W, D448P, and E483K relative to SEQ ID NO: 1. In some examples, the enzyme comprises one or more mutations selected from E88K, V119R, G128K, E168K, Q223K, L231A, L293E, and E440K relative to SEQ ID NO: 1. In some examples, the enzyme comprises one or more mutations selected from E88K, V119R, Q223K, L293E, V372I, and E483K relative to SEQ ID NO: 1. In some examples, the enzyme comprises two or more mutations selected from E88K, V119R, G128K, E168K, Q223K, L231A, L293E, and E440K relative to SEQ ID NO: 1. In some examples, the enzyme comprises two or more mutations selected from E88K, V119R, Q223K, L293E, V372I, and E483K relative to SEQ ID NO: 1. In some examples, the enzyme comprises four or more mutations selected from E88K, V119R, Q223K, L293E, V372I, and E483K relative to SEQ ID NO: 1. In some examples, the enzyme comprises two or more mutations selected from E88K, V119R, Q223K, L293E, E440K, and D448W relative to SEQ ID NO: 1.In some examples, the enzyme comprises three or more mutations selected from E88K, V119R, Q223K, L293E, E440K, and D448W relative to SEQ ID NO: 1. In some examples, the enzyme comprises four or more mutations selected from E88K, V119R, Q223K, L293E, E440K, and D448W relative to SEQ ID NO: 1. In some examples, the enzyme comprises two or more mutations selected from E88K, T91M, V119R, G128K, Q223K, L293E, and E440K relative to SEQ ID NO: 1. In some examples, the enzyme comprises three or more mutations selected from 88K, T91M, V119R, G128K, Q223K, L293E, and E440K relative to SEQ ID NO: 1. In some examples, the enzyme comprises four or more mutations selected from 88K, T91M, V119R, G128K, Q223K, L293E, and E440K relative to SEQ ID NO:1.

[0034] Enzyme optimization

[0035] Described herein are methods and systems for in silico library design. For example, an enzyme or enzyme fragment sequence is used as input. In some examples, any enzyme sequence is used as input to the methods and systems described herein. A database containing known mutations from an organism is queried, and a library of sequences containing combinations of these mutations is generated. In some examples, specific mutations or combinations of mutations are excluded from the library (e.g., known immunogenic sites, structural sites, etc.). In some examples, specific sites in the input sequence are systematically substituted with histidine, aspartic acid, glutamic acid, or combinations thereof. In some examples, a maximum or minimum number of mutations allowed for each region of the enzyme is specified. In some examples, mutations are described relative to the input sequence or the corresponding germline sequence of the input sequence. For example, sequences generated by optimization include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, or more than 16 mutations from the input sequence. In some examples, sequences generated by optimization include no more than 1, no more than 2, no more than 3, no more than 4, no more than 5, no more than 6, no more than 7, no more than 8, no more than 9, no more than 10, no more than 11, no more than 12, no more than 13, no more than 14, no more than 15, no more than 16, or no more than 18 mutations from the input sequence. In some examples, the sequences generated by optimization contain about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, or about 18 mutations relative to the input sequence. In silico enzyme libraries are, in some examples, synthesized, assembled, and / or enriched for desired sequences.

[0036] Also, germline sequences corresponding to input sequences may be modified to generate sequences in the library. For example, sequences generated by the optimization methods described herein contain at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, or more than 16 mutations from the germline sequence. In some examples, sequences generated by optimization contain no more than 1, no more than 2, no more than 3, no more than 4, no more than 5, no more than 6, no more than 7, no more than 8, no more than 9, no more than 10, no more than 11, no more than 12, no more than 13, no more than 14, no more than 15, no more than 16, or no more than 18 mutations from the germline sequence. In some examples, the sequence generated by optimization contains about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, or about 18 mutations relative to the germline sequence.

[0037] Machine Learning

[0038] Data from the preprocessing operations may be fed into one or more machine learning (ML) algorithms to identify a library containing one or more candidates with high affinity for the target and / or functional activity, as described herein. In some embodiments, the one or more candidates include one or more sequences encoding enzymes. In some examples, the library may be a synthetic library. In some embodiments, the ML algorithm may be integrated into a computational pipeline for intelligent decision-making and / or experimental validation. In some embodiments, the one or more ML algorithms may be supervised, semi-supervised, or unsupervised for training to identify anomalies. In some embodiments, the one or more ML algorithms may perform classification or clustering to identify anomalies or attacks. In some embodiments, the one or more ML algorithms may include a classical ML algorithm for performing clustering to identify outliers. Classical ML algorithms may consist of algorithms that learn from existing observations (i.e., known features) to predict outputs. In some cases, the classical ML algorithm for performing clustering may be k-means clustering, mean-shift clustering, density-based spatial clustering for applications with noise (DBSCAN), expectation-maximization (EM) clustering (e.g., using Gaussian mixture models (GMM)), agglomerative hierarchical clustering, or a combination thereof. In some embodiments, the one or more ML algorithms may include a classical ML algorithm for classification. In some cases, the classical ML algorithm may include logistic regression, naive Bayes, k-nearest neighbors, random forests or decision trees, gradient boosting, support vector machines (SVMs), or a combination thereof. In some embodiments, the one or more ML algorithms may employ deep learning. Deep learning algorithms may include algorithms that learn by extracting new features to predict outputs. Deep learning algorithms may be composed of layers, which may include neural networks.

[0039] Expression system

[0040] Provided herein are libraries containing nucleic acids encoding enzymes, which have improved specificity, stability, expression, folding, or downstream activity. In some examples, the libraries described herein are used for screening and analysis.

[0041] Libraries containing nucleic acids encoding enzymes are provided herein, and the nucleic acid libraries are used for screening and analysis. In some examples, the screening and analysis include in vitro, in vivo, or ex vivo assays. Cells for screening include primary cells or cell lines taken from a living subject. Cells can be derived from prokaryotes (e.g., bacteria and fungi) or eukaryotes (e.g., animals and plants). Exemplary animal cells include, but are not limited to, those derived from mice, rabbits, primates, and insects. In some examples, cells for screening include, but are not limited to, cell lines including Chinese hamster ovary (CHO) cell lines, human embryonic kidney (HEK) cell lines, or baby hamster kidney (BHK) cell lines. In some examples, the nucleic acid libraries described herein can also be delivered to multicellular organisms. Exemplary multicellular organisms include, but are not limited to, a plant, a mouse, a rat, a rabbit, a primate (e.g., a monkey or ape), a fish, a worm, a bird, a chicken, a camelid, a cat, a dog, a horse, a cow, a sheep, a goat, a frog, or an insect.

[0042] The nucleic acid libraries described herein can be screened for various pharmacological or pharmacokinetic properties. In some examples, the libraries are screened using in vitro assays, in vivo assays, or ex vivo assays. For example, the in vitro pharmacological or pharmacokinetic properties that are screened include, but are not limited to, binding affinity, binding specificity, and binding activity. Exemplary in vivo pharmacological or pharmacokinetic properties that are screened for the libraries described herein include, but are not limited to, therapeutic efficacy, activity, preclinical toxicity profile, clinical efficacy profile, clinical toxicity profile, immunogenicity, potency, and clinical safety profile.

[0043] Nucleic acid libraries are provided herein, and the nucleic acid libraries can be expressed in vectors. Expression vectors for inserting the nucleic acid libraries disclosed herein can include eukaryotic or prokaryotic expression vectors. Exemplary expression vectors include mammalian expression vectors: pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His, pCEP4 Examples of suitable vectors include, but are not limited to, pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 vector, pEF1a-tdTomato vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), and pSF-CMV-PURO-NH2-CMYC; bacterial expression vectors: pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, and pSF-Tac; plant expression vectors: pRI101-AN DNA and pCambia2301; and yeast expression vectors: pTYB21 and pKLAC2; and insect vectors: pAc5.1 / V5-His A and pDEST8. In some examples, the vector is pcDNA3 or pcDNA3.1.

[0044] To create constructs containing enzymes, nucleic acid libraries are described herein that are expressed in vectors.In some examples, the size of the constructs varies.In some examples, the constructs are at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 1100, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 2000, at least about 2400, at least about 2600, at least about 2800, At least or about 3000, at least or about 3200, at least or about 3400, at least or about 3600, at least or about 3800, at least or about 4000, at least or about 4200, at least or about 4400, at least or about 4600, at least or about 4800, at least or about 5000, at least or about 6000, at least or about 7000, at least or about 8000, at least or about 9000, at least or about 10000, or at least or more than 10000 bases.In some examples, the construct may be about 300-1,000, 300-2,000, 300-3,000, 300-4,000, 300-5,000, 300-6,000, 300-7,000, 300-8,000, 300-9,000, 300-10,000, 1,000-2,000, 1,000-3,000, 1,000-4,000, 1,000-5,000, 1, 000~6,000, 1,000~7,000, 1,000~8,000, 1,000~9,000, 1,000~10,000, 2,000~3,000, 2,000~4,000, 2,000~5,000, 2,000~6,000, 2,000~7,000, 2,000~8,000, 2,000~9,000, 2,000~10,000, 3,000~4,000, 3, 000~5,000, 3,000~6,000, 3,000~7,000, 3,000~8,000, 3,000~9,000, 3,000~10,000, 4,000~5,000, 4,000~6,000, 4,000~7,000, 4,000~8,000, 4,000~9,000, 4,000~10,000, 5,000~6,000, 5,000~7,000, 5, The ranges include 000-8,000, 5,000-9,000, 5,000-10,000, 6,000-7,000, 6,000-8,000, 6,000-9,000, 6,000-10,000, 7,000-8,000, 7,000-9,000, 7,000-10,000, 8,000-9,000, 8,000-10,000, or 9,000-10,000 bases.

[0045] A library containing nucleic acids encoding enzymes is provided herein, and the nucleic acid library is expressed in cells. In some examples, the library is synthesized to express a reporter gene. Exemplary reporter genes include, but are not limited to, acetohydroxyacid synthase (AHAS), alkaline phosphatase (AP), β-galactosidase (LacZ), β-glucoronidase (GUS), chloramphenicol acetyltransferase (CAT), green fluorescent protein (GFP), red fluorescent protein (RFP), yellow fluorescent protein (YFP), cyan fluorescent protein (CFP), cerulean fluorescent protein, citrine fluorescent protein, orange fluorescent protein, cherry fluorescent protein, turquoise fluorescent protein, blue fluorescent protein, horseradish peroxidase (HRP), luciferase (Luc), nopaline synthase (NOS), octopine synthase (OCS), luciferase, and derivatives thereof. Methods for determining the regulation of reporter genes are well known in the art and include, but are not limited to, fluorometry (e.g., fluorescence spectroscopy, fluorescence activated cell sorting (FACS), fluorescence microscopy), and antibiotic resistance determination.

[0046] The term "sequence identity" means that two polynucleotide sequences are identical (i.e., nucleotide-by-nucleotide) over a comparison window. The term "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over a comparison window, determining the number of positions where the same nucleic acid base (e.g., A, T, C, G, U, or I) occurs in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window (i.e., the window size), and multiplying the result by 100 to calculate the percentage of sequence identity.

[0047] The term "homology" or "similarity" between two proteins is determined by comparing the amino acid sequence and its conserved amino acid substitutes of one protein sequence with the sequence of a second protein. Similarity can be determined by procedures well known in the art, such as, for example, the BLAST program (Basic Local Alignment Search Tool at the National Center for Biological Information).

[0048] Provided herein are libraries containing nucleic acids encoding enzymes (e.g., ligases). The enzymes described herein allow for improved stability of a range of active site-encoding sequences. In some instances, the active site-encoding sequence is determined by the interaction between the substrate and the catalytic active site of the enzyme.

[0049] The sequence of the active site based on the surface interaction between the ligand / substrate and the enzyme described herein can be analyzed using various methods. For example, multiple types of computer analysis can be performed. In some examples, structural analysis can be performed. In some examples, sequence analysis can be performed. Sequence analysis can be performed using databases known in the art. Non-limiting examples of databases include, but are not limited to, NCBI BLAST (blast.ncbi.nlm.nih.gov / Blast.cgi), UCSC Genome Browser (genome.ucsc.edu / ), UniProt (www.uniprot.org / ), and IUPHAR / BPS Guide to PHARMACOLOGY (guidetopharmacology.org / ).

[0050] The active site described herein is designed based on sequence analysis between various organisms.For example, sequence analysis is carried out to identify homologous sequences in different organisms.Exemplary organisms include, but are not limited to, mouse, rat, horse, sheep, cow, primate (for example, chimpanzee, baboon, gorilla, orangutan, monkey), dog, cat, pig, donkey, rabbit, camelid, fish, fly or human.In some examples, homologous sequences are identified in the same organism across individuals.

[0051] After identifying the active site, a library can be generated that includes nucleic acids encoding the active site. In some examples, the active site library includes active site sequences designed based on conformational ligand / substrate interactions. The active site library can be translated to generate a protein library. In some examples, the active site library is translated to generate a peptide library, an immunoglobulin library, derivatives thereof, or combinations thereof. In some examples, the active site library is translated to generate a protein library, which is further modified to generate a peptidomimetic library. In some examples, the active site library is translated to generate a protein library that is used to generate small molecules.

[0052] The methods described herein provide for the synthesis of a library of active sites comprising nucleic acids, each encoding a predetermined variant of at least one predetermined reference nucleic acid sequence. In some examples, the predetermined reference sequence is a nucleic acid sequence encoding a protein, and the variant library comprises sequences encoding at least a single codon variation, such that multiple different variants of a single residue in the subsequent protein encoded by the synthesized nucleic acid are generated by standard translation processes. In some examples, the active site library comprises diverse nucleic acids that collectively encode variations at multiple positions. In some examples, the variant library comprises sequences encoding at least a single codon variation in the active site. In some examples, the variant library comprises sequences encoding multiple codon variations in the active site. Exemplary numbers of codons for variation include, but are not limited to, at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300.

[0053] The methods described herein provide for the synthesis of libraries containing nucleic acids encoding active sites, wherein the libraries contain sequences encoding variations in the length of the active site. In some examples, the libraries contain sequences encoding variations in length that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons smaller than a predetermined reference sequence. In some examples, the library contains sequences encoding mutations that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, or more than 300 codons long compared to a given reference sequence.

[0054] Following identification of the active site, an enzyme can be designed and synthesized to contain the active site. The enzyme containing the active site can be designed based on binding, specificity, stability, expression, folding, or downstream activity.

[0055] The methods described herein provide for the synthesis of a library of nucleic acids, each encoding a predetermined variant of at least one predetermined reference nucleic acid sequence. In some instances, the predetermined reference sequence is a nucleic acid sequence encoding a protein, and the variant library includes sequences encoding at least a single codon variation, such that multiple different variants of a single residue in the subsequent protein encoded by the synthesized nucleic acid are generated by standard translation processes. In some instances, the library includes diverse nucleic acids that collectively encode variations at multiple positions. In some instances, the variant library includes sequences encoding at least a single codon variation in the active site. For example, at least one single codon of an enzyme is changed. Exemplary numbers of codons for variation include, but are not limited to, at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300.

[0056] The methods described herein provide for the synthesis of a library of nucleic acids, each encoding a predetermined variant of at least one predetermined reference nucleic acid sequence, wherein the library comprises sequences encoding variations in the length of a domain in an enzyme. In some examples, the library comprises sequences encoding variations in length that are at least about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 225, 250, 275, 300, or more than 300 codons shorter than the predetermined reference sequence. In some examples, the library contains sequences encoding mutations that are at least or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, or more than 300 codons long compared to a given reference sequence.

[0057] Subsequently, an enzyme library is synthesized for screening and analysis. For example, the library is assayed for library displayability, screening, and / or panning. In some examples, displayability is assayed using a selectable tag. Exemplary tags include, but are not limited to, radioactive labels, fluorescent labels, enzymes, chemiluminescent tags, colorimetric tags, affinity tags, or other labels or tags known in the art. In some examples, the tag is histidine, polyhistidine, myc, hemagglutinin (HA), or FLAG. In some examples, libraries are assayed by sequencing using various methods, including, but not limited to, single molecule real-time (SMRT) sequencing, polony sequencing, sequencing by ligation, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis. In multiple examples, libraries are assayed for ligase activity or stability.

[0058] Variant Library

[0059] Codon mutation

[0060] The variant nucleic acid libraries described herein can include multiple nucleic acids, each encoding a variant codon sequence compared to a reference nucleic acid sequence. In some examples, each nucleic acid in a first nucleic acid population contains a variant at a single variant site. In some examples, a first nucleic acid population contains multiple variants at a single variant site, and the first nucleic acid population contains multiple variants at the same variant site. A first nucleic acid population can include nucleic acids that collectively encode multiple codon variants at the same variant site. A first nucleic acid population can include nucleic acids that collectively encode up to 19 or more codons at the same position. A first nucleic acid population can include nucleic acids that collectively encode up to 60 variant triplets at the same position, or a first nucleic acid population can include nucleic acids that collectively encode up to 61 different triplets of codons at the same position. Each variant can encode a codon that results in a different amino acid during translation. Table 3 provides a list of each possible codon (and representative amino acid) for the variant site.

[0061] [Table 3]

[0062] A nucleic acid population can include a variety of nucleic acids that collectively encode up to 20 codon variations at multiple positions. In such cases, each nucleic acid in the population contains codon variations at multiple positions within the same nucleic acid. In some examples, each nucleic acid in the population contains codon variations at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more codons within a single nucleic acid. In some examples, each variant length nucleic acid contains codon variations at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more codons within a single length nucleic acid. In some examples, the variant nucleic acid population comprises codon variations at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more codons within a single nucleic acid, hi some examples, the variant nucleic acid population comprises codon variations at at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more codons within a single length of nucleic acid.

[0063] Highly parallel nucleic acid synthesis

[0064] Provided herein is a platform approach that utilizes miniaturization, parallelization, and vertical integration of the end-to-end process from polynucleotide synthesis to gene assembly in silicon nanowells to create an innovative synthesis platform. The device described herein provides a silicon synthesis platform that increases throughput by up to 1,000-fold or more compared to traditional synthesis methods, producing up to approximately 1,000,000 or more polynucleotides or 10,000 or more genes in a single highly parallel run, in the same footprint as a 96-well plate.

[0065] With the advent of next-generation sequencing, high-resolution genomic data has become a key component of research that dissects the biological roles of various genes in both normal biology and disease development. At the heart of this research is the central dogma of molecular biology and the concept of "sequential information transmission, residue by residue." Genomic information encoded in DNA is transcribed into messages that are then translated into proteins, the active products within specific biological pathways.

[0066] Another exciting area of ​​research is the discovery, development, and production of therapeutic molecules focused on highly specific cellular targets. Highly diverse DNA sequence libraries are at the core of the targeted therapeutic drug development pipeline. Gene variants are used to express proteins in the design, build, and test cycle of protein engineering, which ideally results in genes optimized for high expression of proteins with high affinity for their therapeutic target. As an example, consider the binding pocket of a receptor. The ability to simultaneously test all sequence permutations of all residues within the binding pocket allows for thorough exploration and increases the likelihood of success. Saturation mutagenesis, in which researchers attempt to generate all possible mutations at a specific site within the receptor, is one approach to this development challenge. Although costly, time-consuming, and labor-intensive, it allows for the introduction of each variant at every position. In contrast, combinatorial mutagenesis allows extensive alterations to be made at a few selected positions or short stretches of DNA, resulting in an incomplete repertoire of variants due to biased representation.

[0067] To accelerate the drug development pipeline, libraries containing desired variants available at the intended frequency in the appropriate locations available for testing—i.e., precision libraries—can reduce screening costs and turnaround times. Provided herein are methods for synthesizing nucleic acid synthetic variant libraries that provide precise introduction of each intended variant at the desired frequency. For end users, this translates to the ability not only to exhaustively sample sequence space, but also to efficiently query these hypotheses, reducing costs and screening time. Genome-wide editing can uncover libraries in which critical pathways, individual variants, and sequence permutations can be tested for optimal functionality, enabling the reconstruction of pathways and entire genomes using thousands of genes to redesign biological systems for drug discovery.

[0068] In a first example, the enzyme itself can be optimized using the methods described herein. For example, to improve a specific function of the enzyme, a variant polynucleotide library encoding a portion of the enzyme is designed and synthesized. A nucleic acid library of variants of the enzyme can then be created by the processes described herein (e.g., PCR mutagenesis followed by insertion into a vector). The enzyme is then expressed in a production cell line and screened for enhanced activity. Examples of screening include testing binding affinity to substrates, stability (e.g., heat, salt), or modulation of function (e.g., substrate range, rate).

[0069] The nucleic acid library synthesized by the method described herein can be expressed in various cells related to disease state.Cells related to disease state include cell lines, tissue samples, primary cells from a subject, cultured cells grown from a subject, or cells in a model system.Exemplary model systems include, but are not limited to, plant models and animal models of disease state.

[0070] To identify variant molecules relevant to the prevention, alleviation, or treatment of a disease state, the variant nucleic acid libraries described herein are expressed in cells associated with the disease state or cells capable of inducing the disease state. In some examples, a drug is used to induce the disease state in cells. Exemplary tools for inducing a disease state include, but are not limited to, the Cre / Lox recombination system, LPS inflammation induction, and streptozotocin to induce hypoglycemia. Cells associated with a disease state can be cells from a model system or cultured cells, as well as cells from a subject with a particular disease state. Exemplary disease states include bacterial, fungal, viral, autoimmune, or proliferative diseases (e.g., cancer). In some examples, the variant nucleic acid library is expressed in a model system, cell line, or primary cells obtained from a subject and screened for changes in at least one cellular activity. Exemplary cellular activities include, but are not limited to, proliferation, cell cycle progression, cell death, adhesion, migration, regeneration, cell signaling, energy production, oxygen utilization, metabolic activity, aging, response to free radical damage, or any combination thereof.

[0071] In some examples, the methods described herein provide for the creation of a library of nucleic acids comprising variant nucleic acids that differ at multiple codon sites. In some examples, the nucleic acids may have one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, twenty, thirty, forty, fifty, or more variant codon sites. In some examples, one or more of the variant codon sites may be adjacent. In some examples, one or more of the variant codon sites may not be adjacent, but may be separated by one, two, three, four, five, six, seven, eight, nine, ten, or more codons. In some examples, a nucleic acid may contain multiple variant codon sites, where all of the variant codon sites are adjacent to one another to form a contiguous sequence of variant codon sites. In some examples, a nucleic acid may contain multiple variant codon sites, where none of the variant codon sites are adjacent to one another. In some examples, a nucleic acid may contain multiple variant codon sites, where some of the variant codon sites are adjacent to one another to form a contiguous sequence of variant codon sites, and some of the variant codon sites are not adjacent to one another.

[0072] Sequencing

[0073] The enzymes provided herein can be used in a variety of downstream applications. In some examples, the enzyme comprises a ligase. In some examples, a sample is obtained from one or more sources, and a population of sample polynucleotides is isolated. The sample is obtained from a biological source, such as saliva, blood, tissue, skin, or a completely synthetic source (non-limiting examples). Multiple polynucleotides obtained from the sample are fragmented, end-repaired, and adenylated to form double-stranded sample nucleic acid fragments. In some cases, end-repair is achieved by treatment with one or more enzymes, such as T4 DNA polymerase or a variant thereof, Klenow enzyme, and T4 polynucleotide kinase, in an appropriate buffer. Nucleotide overhangs to facilitate ligation to adapters are optionally added, along with a 3'-5' exo-minus Klenow fragment and dATP.

[0074] Adapters (e.g., universal adapters) may be ligated to both ends of sample polynucleotide fragments using a ligase such as T4 ligase to generate a library of adapter-tagged polynucleotide strands, and the adapter-tagged polynucleotide library is amplified using primers such as universal primers. In some examples, the adapters are Y-shaped adapters containing one or more primer binding sites, one or more graft regions, and one or more index (or barcode) regions. In some examples, one or more index regions are present on each strand of the adapter. In some cases, the graft regions are complementary to the flow cell surface, facilitating next-generation sequencing of the sample library. In some examples, the Y-shaped adapters contain partially complementary sequences. In some examples, the Y-shaped adapters contain a single thymidine overhang that hybridizes to the overhanging adenine of the double-stranded adapter-tagged polynucleotide strands. The Y-shaped adapters may contain modified nucleic acids that are resistant to cleavage. For example, a phosphorothioate backbone is used to attach the overhanging thymidine to the 3' end of the adapter. If universal primers are used, library amplification is performed to add barcoded primers to the adapters.

[0075] A plurality of nucleic acids (i.e., genomic sequences) are obtained from a sample, fragmented, optionally end-repaired, and adenylated. Adapters are ligated to both ends of the polynucleotide fragments to generate a library of adapter-tagged polynucleotide strands, and the adapter-tagged polynucleotide library is amplified. The adapter-tagged polynucleotide library is then denatured in the presence of an adapter blocker at elevated temperatures, preferably 96°C. A polynucleotide target library (probe library) is denatured in a hybridization solution at elevated temperatures, preferably about 90-99°C, and mixed with the denatured tagged polynucleotide library in the hybridization solution at about 45-80°C for about 10-24 hours. A binding buffer is then added to the hybridized tagged polynucleotide probes, and a solid support containing a capture moiety is used to selectively bind the hybridized adapter-tagged polynucleotide probes. The solid support is washed with buffer one or more times, preferably about two and five times, to remove unbound polynucleotides, after which an elution buffer is added to release the enriched adapter-tagged polynucleotide fragments from the solid support. The enriched library of adaptor-tagged polynucleotide fragments is amplified, and then the library is sequenced. Alternative variables such as incubation time, temperature, reaction volume / concentration, number of washes, or other variables consistent with this specification may also be used in this method.

[0076] In either case, detection or quantitative analysis of the oligonucleotides can be achieved by sequencing. Subunits or the entire synthesized oligonucleotide can be detected through complete sequencing of all oligonucleotides by any suitable method known in the art, such as Illumina sequencing by synthesis, PacBio nanopore sequencing, or BGI / MGI nanoball sequencing, including the sequencing methods described herein.

[0077] Sequencing can be achieved by the classical Sanger sequencing method, which is well known in the art.Sequencing can also be achieved using high-throughput systems, some of which allow the detection of sequenced nucleotide immediately after its incorporation into growing chain or at the time of its incorporation, that is, the detection of sequence in red time or substantially real time.In some cases, high-throughput sequencing generates at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000 or at least 500,000 sequence reads per hour, and each read is at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120 or at least 150 bases / read.

[0078] In some instances, high-throughput sequencing involves the use of technology available from Illumina's Genome Analyzer IIX, MiSeq Personal Sequencer, or HiSeq systems, such as the HiSeq2500, HiSeq1500, HiSeq2000, HiSeq1000, iSeq100, MiniSeq, MiSeq, NextSeq550, NextSeq2000, NextSeq550, or NovaSeq6000. These machines use reversible terminator-based sequencing by synthesis chemistry. These machines can generate over 6,000 Gb of reads in 13–44 hours. Smaller systems can be utilized for runs in 3, 2, 1, or less days. Short synthesis cycles can be used to minimize the time it takes to obtain sequencing results.

[0079] In some instances, high-throughput sequencing involves the use of technology available from the ABI Solid Systems. This genetic analysis platform allows for massively parallel sequencing of clonally amplified DNA fragments linked to beads. The sequencing method is based on sequential ligation with dye-labeled oligonucleotides.

[0080] Next-generation sequencing can include ion semiconductor sequencing (e.g., using technology from Life Technologies (Ion Torrent)). Ion semiconductor sequencing can take advantage of the fact that ions can be released when a nucleotide is incorporated into a strand of DNA. To perform ion semiconductor sequencing, a high-density array of microfabricated wells can be formed. Each well can hold a single DNA template. There can be an ion-sensitive layer beneath the well, and there can be an ion sensor beneath the ion-sensitive layer. When a nucleotide is added to the DNA, H+ can be released, which can be measured as a change in pH. The H+ ions can be converted to a voltage and recorded by the semiconductor sensor. The array chip can be flooded sequentially with nucleotides one after the other. No scanning, light, or camera is required. In some cases, an IONPROTON™ sequencer is used to sequence nucleic acids. In some cases, an IONPGM™ sequencer is used. The Ion Torrent Personal Genome Machine (PGM) can generate 10 million reads in two hours.

[0081] In some instances, high-throughput sequencing involves the use of technology available from Helicos BioSciences Corporation (Cambridge, Massachusetts), such as Single Molecule Sequencing by Synthesis (SMSS). SMSS is unique because it allows sequencing of the entire human genome in up to 24 hours. Finally, SMSS is powerful because, like MW technology, it does not require a pre-amplification step before hybridization. In fact, SMSS does not require amplification at all. SMSS is partially described in U.S. Patent Application Publication Nos. 2006002471 I, 20060024678, 20060012793, 20060012784, and 20050100932.

[0082] In some cases, high-throughput sequencing involves the use of technology available from 454 Lifesciences, Inc. (Branford, Connecticut), such as the Pico Titer Plate device, which contains a fiber optic plate that transmits the chemiluminescent signal generated by the sequencing reaction and is recorded by a CCD camera within the instrument. The use of this fiber optic allows for the detection of at least 20 million base pairs in 4.5 hours.

[0083] Methods using bead amplification followed by fiber optic detection are described in Marguiles, M., et al. "Genome sequencing in microfabricated high-density picolitre reactors", Nature, doi:10.1038 / nature03959, and U.S. Patent Application Publication Nos. 20020012930, 20030058629, 20030100102, 20030148344, 20040248161, 20050079510, 20050124022, and 20060078909.

[0084] In some examples, high-throughput sequencing is carried out using Clonal Single Molecule Array (Solexa, Inc.) or sequencing by synthesis (SBS) using reversible terminator chemistry.These technologies are partially described in U.S. Patent Nos. 6,969,488, 6,897,023, 6,833,246, 6,787,308, and U.S. Patent Application Publication Nos. 20040106130, 20030064398, 20030022207, and Constants, A., The Scientist 2003, 17(13):36.High-throughput sequencing of oligonucleotides can be achieved using any suitable sequencing method known in the art, such as those commercially available from Pacific Biosciences, Complete Genomics, Genia Technologies, Halcyon Molecular, Oxford Nanopore Technologies, etc. For other high-throughput sequencing systems, see Venter, J., et al., Science, February 16, 2001; Adams, M. et al., Science, March 24, 2000; and MJ, Levene, et al., Science, 299:682-686, January 2003; and US Patent Application Nos. 20030044781 and 2006 / 0078937. Overall, such systems involve sequencing a target oligonucleotide molecule having multiple bases by measuring the time-dependent addition of bases through a polymerization reaction on the oligonucleotide molecule, i.e., the activity of nucleic acid polymerase on the template oligonucleotide molecule to be sequenced is tracked in real time.Then, by identifying which bases are incorporated into the growing complementary strand of the target oligonucleotide through the catalytic activity of nucleic acid polymerase at each step in the sequence of base addition, the sequence can be deduced. A polymerase on the target oligonucleotide molecule complex is provided in a suitable position to move along the target oligonucleotide molecule and extend the oligonucleotide primer at the active site.A plurality of labeled types of nucleotide analogs are provided adjacent to the active site, with each distinct type of nucleotide analog being complementary to a different nucleotide in the target oligonucleotide sequence. The growing oligonucleotide chain is extended by using a polymerase to add nucleotide analogs to the oligonucleotide chain at the active site, with the added nucleotide analogs being complementary to nucleotides of the target oligonucleotide at the active site. The nucleotide analogs added to the oligonucleotide primer as a result of the polymerization step are identified. The steps of providing labeled nucleotide analogs, polymerizing the growing oligonucleotide chain, and identifying the added nucleotide analogs are repeated to further extend the oligonucleotide chain and determine the sequence of the target oligonucleotide.

[0085] Next-generation sequencing technologies include Pacific Biosciences' SMRT™ technology. In SMRT, each of the four DNA bases can be attached to one of four different fluorescent dyes. These dyes can be phosphate-linked. A single DNA polymerase can be immobilized on a single molecule of template single-stranded DNA at the bottom of a zero-mode waveguide (ZMW). The ZMW can be a confining structure that allows observation of the incorporation of a single nucleotide by the DNA polymerase against a background of fluorescent nucleotides that can rapidly diffuse out of the ZMW (in microseconds). It can take several milliseconds for the nucleotide to be incorporated into the growing strand. During this time, the fluorescent label can be excited to generate a fluorescent signal, and the fluorescent tag can be cleaved. The ZMW can be illuminated from below. Attenuated light from the excitation beam can penetrate the bottom 20–30 nm of each ZMW. A microscope with a detection limit of 20 zeptoliters (10 in. liters) can be created. The small detection volume can provide a 1000-fold improvement in background noise reduction. Detection of the corresponding fluorescence of the dye can indicate which base has been incorporated. This process can be repeated.

[0086] In some cases, next-generation sequencing is nanopore sequencing (see, e.g., Soni GV and Meller A. (2007) Clin Chem 53:1996-2001). A nanopore can be a small hole with a diameter on the order of about 1 nanometer. When a nanopore is immersed in a conductive fluid and a potential is applied across it, a small current can be generated due to the conduction of ions through the nanopore. The amount of current that flows is sensitive to the size of the nanopore. When a DNA molecule passes through the nanopore, each nucleotide on the DNA molecule can block the nanopore to a different extent. Therefore, changes in the current passing through the nanopore as the DNA molecule passes through the nanopore can represent a readout of the DNA sequence. The nanopore sequencing technology can be, for example, the GridION system manufactured by Oxford Nanopore Technologies. A single nanopore can be inserted into the polymer membrane across the top of a microwell. Each microwell can have an electrode for individual sensing. The microwells can be fabricated into array chips with 100,000 or more microwells per chip (e.g., greater than 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000). Instruments (or nodes) can be used to analyze the chip. Data can be analyzed in real time. One or more instruments can be operated at a time. The nanopore can be a protein nanopore, e.g., protein alpha hemolysin, a heptameric protein pore. The nanopore can be a fabricated solid-state nanopore, e.g., a nanometer-sized pore formed in a synthetic membrane (e.g., SiN xor SiO2). The nanopore can be a hybrid pore (e.g., a protein pore integrated into a solid membrane). The nanopore can be a nanopore with an integrated sensor (e.g., a tunneling electrode detector, a capacitive detector, or a graphene-based nanogap or edge-state detector (see, e.g., Garaj et al. (2010) Nature vol. 67, doi:10.1038 / nature09379)). The nanopore can be functionalized to analyze specific types of molecules (e.g., DNA, RNA, or proteins). Nanopore sequencing can include "strand sequencing," in which an intact DNA polymer is threaded through a protein nanopore and sequencing can be performed in real time as the DNA moves through the pore. An enzyme can separate the strands of double-stranded DNA and feed the strand through the nanopore. The DNA can have a hairpin at one end, allowing the system to read both strands. In some cases, nanopore sequencing is "exonuclease sequencing," in which individual nucleotides can be cleaved from a DNA strand by a processive exonuclease and passed through a protein nanopore. The nucleotides can be transiently bound to a molecule (e.g., cyclodextran) in the pore. The characteristic disruption of the current can be used to identify the base.

[0087] Nanopore sequencing technology from GENIA can be used. Engineered protein pores can be embedded in lipid bilayer membranes. "Active control" techniques can be used to enable efficient nanopore-membrane assembly and control of DNA translocation through the channel. In some cases, the nanopore sequencing technology is from NABsys. Genomic DNA can be fragmented into strands with an average length of approximately 100 kb. The 100 kb fragments can be single-stranded and subsequently hybridized with hexamer probes. The probe-bearing genome fragments can be driven through a nanopore, which can produce current versus time tracings. The current traces can provide the location of the probes on each genome fragment. The genome fragments can be aligned to create a genome probe map. This process can be performed in parallel for a library of probes. A genome-length probe map for each probe can be generated. Errors can be fixed with a process called "moving window sequencing by hybridization (mwSBH)." In some cases, the nanopore sequencing technology is from IBM / Roche. Nanopore-sized openings can be created in microchips using electron beams. Electric fields can be used to pull or thread DNA through the nanopore. A DNA transistor device within the nanopore can contain alternating nanometer-sized layers of metal and dielectric. Discrete charges in the DNA backbone can be trapped by the electric field inside the DNA nanopore. By turning the gate voltage off and on, the DNA sequence can be read.

[0088] Next-generation sequencing can include DNA nanoball sequencing (e.g., as performed by Complete Genomics; see, e.g., Drmanac et al. (2010) Science 327: 78-81). DNA can be isolated, fragmented, and size-selected. For example, DNA can be fragmented (e.g., by sonication) to an average length of approximately 500 bp. Adapters (Adl) can be attached to the ends of the fragments. The adapters can be used to hybridize to anchors for sequencing reactions. DNA with adapters attached to each end can be PCR-amplified. The adapter sequences can be modified so that complementary single-stranded ends can ligate to each other to form circular DNA. DNA can be methylated to protect it from cleavage by type IIS restriction enzymes used in subsequent steps. The adapter (e.g., the right adapter) can have a restriction recognition site, or the restriction recognition site can remain unmethylated. The unmethylated restriction recognition site in the adapter can be recognized by a restriction enzyme (e.g., Acul), and the DNA can be cleaved 13 bp to the right of the right adapter by Acul to form a linear double-stranded DNA. Second-round left and right adapters (Ad2) can be ligated to either end of the linear DNA, and all DNA bound by both adapters can be PCR-amplified (e.g., by PCR). The Ad2 sequence can be modified to allow them to ligate together to form circular DNA. The DNA can be methylated, but the restriction enzyme recognition site can remain unmethylated on the left Ad1 adapter. A restriction enzyme (e.g., Acul) can be applied, and the DNA can be cleaved 13 bp to the left of Ad1 to form a linear DNA fragment. Third-round left and right adapters (Ad3) can be ligated to the left and right flanks of the linear DNA, and the resulting fragments can be PCR-amplified. The adapters can be modified to allow them to ligate together to form circular DNA. A type III restriction enzyme (eg, EcoP15) can be added, which can cut the DNA 26 bp to the left of Ad3 and 26 bp to the right of Ad2.This cleavage removes a large segment of DNA and can re-linearize the DNA. Fourth-round left and right adapters (Ad4) can be ligated to the DNA, and the DNA can be amplified (e.g., by PCR) and modified so that they can recombine with each other to form the completed circular DNA template.

[0089] Rolling circle replication (e.g., using Phi 29 DNA polymerase) can be used to amplify small fragments of DNA. Four adapter sequences can contain palindromic sequences that can hybridize, allowing the single strand to fold back on itself to form DNA nanoballs (DNBs™), which can have an average diameter of approximately 200–300 nanometers. The DNA nanoballs can be attached (e.g., by adsorption) to a microarray (sequencing flow cell). The flow cell can be a silicon wafer coated with silicon dioxide, titanium, and hexamethyldisilazane (HMDS), and a photoresist material. Sequencing can be performed by non-chain sequencing by ligating fluorescent probes to DNA. The fluorescent color of the interrogated position can be visualized with a high-resolution camera. The identity of the nucleotide sequence between the adapter sequences can be determined.

[0090] Provided herein is a method for preparing a nucleic acid library, the method comprising one or more steps of: providing one or more sample nucleic acids; contacting the one or more sample nucleic acids with a plurality of adaptors and a T4 ligase variant described herein to form a nucleic acid sequencing library comprising adaptor-ligated nucleic acids; and sequencing the nucleic acid library. In some examples, the sample nucleic acid comprises a genome fragment. In some examples, the genome fragment is obtained by genomic cleavage. In some examples, the genome fragment is obtained by genomic amplification. In some examples, the sample nucleic acid comprises cDNA. In some examples, the sample nucleic acid comprises cfDNA. In some examples, the method further comprises one or more steps for preparing the nucleic acid library (e.g., end repair, a-tailing, and amplification). In some examples, the method further comprises concentrating the nucleic acid library before sequencing.

[0091] The following examples are presented to more clearly illustrate to those skilled in the art the principles and practice of the embodiments disclosed herein, and should not be construed as limiting the scope of any claimed embodiments. Unless otherwise specified, all parts and percentages are by weight. [Example]

[0092] The following examples are provided for the purpose of illustrating various embodiments of the present disclosure and are not intended to limit the disclosure in any manner. The examples, along with the methods described herein, are representative of currently preferred embodiments and are exemplary, not limiting the scope of the disclosure. Modifications thereof and other uses encompassed within the spirit of the disclosure as defined by the claims will occur to those skilled in the art.

[0093] Example 1: T4 ligase high-throughput assay

[0094] T4 variants were tested using the general protocol outlined in Figure 1. DNA encoding the T4 ligase variants was dispensed into a 384-well plate using an Echo Liquid Handler system (Beckman). Each fragment was diluted to a concentration of 20 ng / microliter. 1 / 40 of a microliter (1 drop, 0.5 ng) from each well was transferred to a new 384-well plate. PCR was then performed in the 384-well microplate to biotin-block the F side and phosphorylate the R side. After amplification, 20 microliters were transferred to a new plate, and amplicons were isolated using Felix SPRI. The amplicons were then eluted in 28 microliters for spectrophotometric quantification. The concentration was at least 67 ng / microliter. 20 ng (2x the plateau range of 10 ng of DNA input to generate a curve, 0.3 microliters) was then dispensed into a new plate as well as TXTL master mix (NEB) for cell-free protein expression reactions. Plates containing TXTL reactions were incubated in a thermal cycler at 37°C for 2 hours and then cooled to 4°C. The wells were then spiked with 39 microliters of adapter master mix at a 1:40 dilution to reduce TXTL inhibition. Ligations were incubated at 15°C for 10 minutes and then inactivated at 65°C for 10 minutes. Samples from each well were pooled, purified by manual SPRI beads, and then split into parallel PCR barcode amplifications and library preparations (one to count ligations and the other to count construct abundance). Both pools were then sequenced.

[0095] Example 2: T4 ligase optimization

[0096] Following the general procedure of Example 1, T4 ligase variants were generated using multiple rounds of optimization / selection. Variants from the wild-type sequence (SEQ ID NO: 1) were selected based in part on high-entropy positions (Figures 2A-2B) and screened using a high-throughput qPCR assay (Figures 3A-3B). In the first round, single variants were tested for ligation performance metrics, including activity, thermostability, and salt tolerance (Figure 4). In screening rounds 2 / 3, all binary combinations of variants were evaluated for epistatic relationships (Figures 5A-5B). Raw additions were also used (Figures 6A-6B). Structural information, including the location of lysine mutations near the DNA substrate, was also fed into the design using an iterative process (Figures 8A-8B). Beneficial lysine mutations were found to cluster near DNA contact regions. Positions near the DNA substrate were repeatedly mutated to lysines to test for improved activity. The revised design from this approach was expressed as a His6-tagged construct and subjected to molecular biology-grade protein purification for evaluation in a round 6 / 7 cfDNA performance comparison NGS assay (Figure 7). Briefly, 10 ng of input cfDNA (cfDNA reference number 104549, Twist Bioscience) was end-repaired / tailed (mechanical fragmentation number 104177, Twist Bioscience), ligated with adapters (UDI adapter number 101308, Twist Bioscience) and variant ligase (250 ng), amplified (mechanical fragmentation number 104177, Twist Bioscience), and sequenced on a NextSeq 500 / 550 instrument (Illumina). The standard ligation protocol used was 10x DNA ligase buffer (2 microliters), 40% PEG (2.5 microliters), adapter (1 microliter), ligase diluent (125 ng / microliter), ER / AT cfDNA sample (10 microliters), and water (2.5 microliters). The reaction was incubated in a thermal cycler (heated lid, 70°C) with the following program: 20°C for 15 minutes, 65°C for 10 minutes, and a 4°C hold.Each experiment was the result of two replicates and included T4 wild type as a control.

[0097] These constructs were evaluated using pass-filter reads, as shown in Figure 9A. The addition of further mutations introduced a negative phenotype into the protein (Figure 9B). Then, for round 8, an orthogonal assay was performed to determine chimeras by rescreening round 6 hits using site-directed mutagenesis of chimera-prone sites (Figures 10A-10B).

[0098] Round 9 involved generating single-site mutations to select variants that did not increase chimerism (Figure 10C), and then screening hits using the NGS assay described for rounds 6 and 7. Round 9 included a 9x design, 10x 384w plates, and 24x assays using 47 purified ligases. These 47 ligases were finally narrowed down, and six variants (18, 24, 12, 8, 21, and 22) were selected for further analysis.

[0099] The results are shown in Figures 11A-13. All variants had similar amounts of chimeras, within 0.1-0.2x of the best variant. The mutations in these variants are shown in Table 4.

[0100] [Table 4]

[0101] These six variants were then subjected to additional NGS screening with varying amounts of enzyme (250, 500, and 1000 ng conditions). Two control reactions were included, with each condition run in quadruplicate. Enzyme variants generally performed better at lower mass loads (Figures 12A-12B). Data relative to the wild-type control are shown for variants 21 and 24. Variants r7r-21 and r7r-24 contain protein SEQ ID NOs: 2 and 3 and were expressed from nucleic acid SEQ ID NOs: 5 and 6, respectively.

[0102] The present disclosure is further illustrated by the following non-limiting sections.

[0103] Item 1. A variant polypeptide comprising at least one amino acid mutation relative to SEQ ID NO:1.

[0104] Item 2. The variant polypeptide of Item 1, wherein the polypeptide comprises at least 80% similarity to any one of SEQ ID NOs: 2-3.

[0105] Item 3. The variant polypeptide of Item 1, wherein the polypeptide comprises at least 90% similarity to any one of SEQ ID NOs: 2-3.

[0106] Item 4. The variant polypeptide of Item 1, wherein the polypeptide comprises at least 95% similarity to any one of SEQ ID NOs: 2 to 3.

[0107] Item 5. The variant polypeptide of Item 1, wherein the polypeptide comprises at least 98% similarity to any one of SEQ ID NOs: 2-3.

[0108] Item 6. The variant polypeptide according to Item 1, wherein the polypeptide comprises any one of SEQ ID NOs: 2 to 3.

[0109] Item 7. The variant polypeptide of Item 1, wherein the polypeptide comprises at least 10 consecutive amino acids of any one of SEQ ID NOs: 2 to 3.

[0110] Item 8. The variant polypeptide of Item 1, wherein the polypeptide comprises at least 20 consecutive amino acids of any one of SEQ ID NOs: 2 to 3.

[0111] Item 9. The variant polypeptide according to Item 1, wherein the polypeptide comprises 20 to 100 consecutive amino acids of any one of SEQ ID NOs: 2 to 3.

[0112] Item 10. The variant polypeptide of item 1, wherein the polypeptide comprises at least two amino acid mutations relative to SEQ ID NO:1.

[0113] Item 11. The variant polypeptide of item 1, wherein the polypeptide comprises at least four amino acid mutations relative to SEQ ID NO:1.

[0114] Item 12. The variant polypeptide of item 1, wherein the polypeptide comprises at least six amino acid mutations relative to SEQ ID NO:1.

[0115] Item 13. The variant polypeptide of any one of items 1 to 12, wherein the mutation is at one or more of positions E88, T91, V119, G128, E168, Q223, L231, L293, V372, E440, D448, and E483 relative to SEQ ID NO: 1.

[0116] Item 14. The variant polypeptide of item 13, wherein the mutation is at one or more of positions E88, V119, G128, E168, Q223, L231, L293, and E440 relative to SEQ ID NO:1.

[0117] Item 15. The variant polypeptide of item 13, wherein the mutation is at one or more of positions E88, V119, Q223, L293, V372, and E483 relative to SEQ ID NO:1.

[0118] Item 16. The variant polypeptide of Item 13, wherein the mutation is selected from one or more of E88K, T91M, V119R, G128K, E168K, Q223K, L231A, L293E, V372I, E440K, D448W, D448P, and E483K relative to SEQ ID NO: 1.

[0119] Item 17. The variant polypeptide of item 13, wherein the mutation is selected from one or more of E88K, V119R, G128K, E168K, Q223K, L231A, L293E, and E440K relative to SEQ ID NO: 1.

[0120] Item 18. The variant polypeptide of item 13, wherein the mutation is selected from one or more of E88K, V119R, Q223K, L293E, V372I, and E483K relative to SEQ ID NO: 1.

[0121] Item 19. The variant polypeptide of any one of items 1 to 18, wherein the polypeptide further comprises a purification tag.

[0122] Item 20. A nucleic acid encoding the variant polypeptide according to any one of items 1 to 19.

[0123] Item 21. A nucleic acid comprising at least 80% similarity to any one of SEQ ID NOs: 4 to 5, provided that the polypeptide encodes the polypeptide of SEQ ID NO: 1.

[0124] Item 22. The nucleic acid according to Item 21, comprising at least 90% similarity to any one of SEQ ID NOs: 4 to 5.

[0125] Item 23. The nucleic acid according to Item 21, comprising at least 95% similarity to any one of SEQ ID NOs: 4 to 5.

[0126] Item 24. A vector comprising the nucleic acid according to any one of Items 20 to 23.

[0127] Item 25. The vector according to Item 24, comprising a plasmid.

[0128] Item 26. A cell containing the nucleic acid according to any one of Items 20 to 23.

[0129] Item 27. The cell according to Item 26, comprising a bacterial cell.

[0130] Item 28. A method for expressing the variant polypeptide according to any one of Items 1 to 19.

[0131] Item 29. The method of Item 25, wherein expression comprises translation of the nucleic acid sequence of any one of Items 20 to 23.

[0132] Item 30. The method according to Item 28 or 29, including in vivo methods.

[0133] Item 31. The method of items 28 or 29, including cell-free methods.

[0134] Item 32. A method for forming a covalent bond between two nucleotides, comprising the step of contacting a first nucleotide and a second nucleotide with the polypeptide according to any one of items 1 to 19.

[0135] Item 33. The method of Item 32, wherein the first nucleotide and the second nucleotide are present on the same nucleic acid.

[0136] Item 34. The method of Item 32, wherein the covalent bond forms a circular nucleic acid.

[0137] Item 35. The method of Item 32, wherein the first nucleotide is present on a first nucleic acid and the second nucleotide is present on a second nucleic acid.

[0138] Item 36. The method of Item 32, wherein the first nucleic acid and / or the second nucleic acid comprises genomic DNA or a fragment thereof.

[0139] Item 37. The method of Item 32, wherein the first nucleic acid and / or the second nucleic acid comprises cDNA.

[0140] Item 38. The method of Item 32, wherein the first nucleic acid and / or the second nucleic acid comprises an adaptor.

[0141] Item 39. The method of Item 38, wherein the first nucleic acid comprises a first adaptor and genomic DNA or cDNA.

[0142] Item 40. The method of Item 39, wherein the second nucleic acid comprises a second adaptor.

[0143] Item 41. The method of any one of Items 38 to 40, wherein the adapter comprises at least one barcode.

[0144] Item 42. The method of item 39, wherein the barcode includes one or more of a sample index, a plate index, a cell index, and a unique molecular identifier.

[0145] Item 43. A method for preparing a nucleic acid library, comprising: (a) providing one or more sample nucleic acids; (b) contacting one or more sample nucleic acids with a plurality of adaptors and the polypeptide of any one of items 1 to 19 to form a nucleic acid sequencing library comprising adaptor-ligated nucleic acids; (c) sequencing the nucleic acid library. The method of claim 43, wherein the sample nucleic acid comprises a genomic fragment.

[0146] Item 45. The method of Item 43, wherein the genomic fragment is obtained by cleavage or amplification of the genome.

[0147] Item 46. The method of Item 43, wherein the sample nucleic acid comprises cDNA.

[0148] Item 47. The method of Item 43, wherein the sample nucleic acid comprises cfDNA.

[0149] Item 48. The method according to any one of Items 43 to 47, further comprising one or more steps of end repair, A-tailing, and amplification.

[0150] Item 49. The method according to any one of Items 43 to 48, further comprising a step of concentrating the nucleic acid library prior to sequencing.

[0151] While preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the present disclosure. It is understood that various alternatives to the embodiments of the present disclosure described herein may be employed in practicing the present disclosure. The following claims define the scope of the disclosure, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

1. A variant polypeptide comprising at least one amino acid mutation relative to SEQ ID NO:

1.

2. 2. The variant polypeptide of claim 1, wherein the polypeptide comprises at least 80% similarity to any one of SEQ ID NOs: 2-3.

3. 2. The variant polypeptide of claim 1, wherein the polypeptide comprises at least 90% similarity to any one of SEQ ID NOs: 2-3.

4. 2. The variant polypeptide of claim 1, wherein the polypeptide comprises at least 95% similarity to any one of SEQ ID NOs: 2-3.

5. 2. The variant polypeptide of claim 1, wherein the polypeptide comprises at least 98% similarity to any one of SEQ ID NOs: 2-3.

6. The variant polypeptide of claim 1, wherein the polypeptide comprises any one of SEQ ID NOs: 2-3.

7. 2. The variant polypeptide of claim 1, wherein the polypeptide comprises at least 10 consecutive amino acids of any one of SEQ ID NOs: 2-3.

8. 2. The variant polypeptide of claim 1, wherein the polypeptide comprises at least 20 consecutive amino acids of any one of SEQ ID NOs: 2-3.

9. 2. The variant polypeptide of claim 1, wherein the polypeptide comprises 20 to 100 consecutive amino acids of any one of SEQ ID NOs: 2-3.

10. The variant polypeptide of claim 1 , wherein the polypeptide comprises at least two amino acid mutations relative to SEQ ID NO:

1.

11. The variant polypeptide of claim 1 , wherein the polypeptide comprises at least four amino acid mutations relative to SEQ ID NO:

1.

12. The variant polypeptide of claim 1 , wherein the polypeptide comprises at least six amino acid mutations relative to SEQ ID NO:

1.

13. 13. The variant polypeptide of any one of claims 1 to 12, wherein the mutations are at one or more of positions E88, T91, V119, G128, E168, Q223, L231, L293, V372, E440, D448, and E483 relative to SEQ ID NO:

1.

14. 14. The variant polypeptide of claim 13, wherein the mutation is at one or more of positions E88, V119, G128, E168, Q223, L231, L293, and E440 relative to SEQ ID NO:

1.

15. 14. The variant polypeptide of claim 13, wherein the mutation is at one or more of positions E88, V119, Q223, L293, V372, and E483 relative to SEQ ID NO:

1.

16. 14. The variant polypeptide of claim 13, wherein the mutations are selected from one or more of E88K, T91M, V119R, G128K, E168K, Q223K, L231A, L293E, V372I, E440K, D448W, D448P, and E483K relative to SEQ ID NO:

1.

17. 14. The variant polypeptide of claim 13, wherein the mutations are selected from one or more of E88K, V119R, G128K, E168K, Q223K, L231A, L293E, and E440K relative to SEQ ID NO:

1.

18. 14. The variant polypeptide of claim 13, wherein the mutations are selected from one or more of E88K, V119R, Q223K, L293E, V372I, and E483K relative to SEQ ID NO:

1.

19. The variant polypeptide of any one of claims 1 to 18, wherein the polypeptide further comprises a purification tag.

20. A nucleic acid encoding a variant polypeptide according to any one of claims 1 to 19.