High-throughput enzyme library sequencing using long-read amplicon sequencing

Long-read sequencing with single-round PCR and barcoded primers addresses the high cost and time issues of conventional methods, enabling efficient and uniform sequencing of enzyme libraries, thereby enhancing enzyme engineering efficiency.

WO2025264705A1PCT designated stage Publication Date: 2025-12-26SOLUGEN INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/034020
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-17
Filing Date
2025-06-17
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Conventional DNA sequencing technologies, such as Sanger sequencing, are cost-prohibitive and time-intensive, making it impractical to sequence every enzyme mutant in a library, which is necessary for directed enzyme evolution, especially for enzymes requiring properties like solvent tolerance and high-rate conversion.

Method used

The use of long-read sequencing technology, specifically Oxford Nanopore sequencing, combined with single-round PCR amplification and barcoded primers, allows for high-throughput sequencing of enzyme libraries by incorporating barcode sequences directly into PCR primers, enabling efficient amplification and identification of large numbers of enzyme mutants.

Benefits of technology

This approach significantly reduces costs and time while ensuring consistent and uniform sequencing coverage across enzyme variants, improving the reliability of mutation identification and accelerating enzyme engineering processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025034020_26122025_PF_FP_ABST
    Figure US2025034020_26122025_PF_FP_ABST
Patent Text Reader

Abstract

A method for high-throughput sequencing of enzyme libraries. The method may include amplifying a nucleic acid region of interest using forward and reverse primers, each primer having a template-binding region, a barcode sequence, and spacer regions, to produce barcoded amplicons. The method may also include pooling the barcoded amplicons. The method may also include sequencing the pooled barcoded amplicons using long-read sequencing technology.
Need to check novelty before this filing date? Find Prior Art

Description

HIGH-THROUGHPUT ENZYME LIBRARY SEQUENCING USING LONG-READ AMPLICON SEQUENCINGCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application Serial No. 63 / 660,643 filed June 17, 2024, and entitle “High-Throughput Enzyme Library Sequencing Using Long-Read Amplicon Sequencing,” which is hereby incorporated herein by reference in its entirety for all purposes.CROSS-REFERENCE TO SEQUENCE LISTING

[0002] The instant application contains a Sequence Listing which has been submitted electronically in XML file format and is hereby incorporated by reference in its entirety. Said XML file, created on June 16, 2025, is named “3416-20801_Sequence_Listing.xml” and is 128 kilobytes in size.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0003] Not applicable.TECHNICAL FIELD

[0004] The present disclosure relates generally to methods for high-throughput enzyme library sequencing. The methods for high-throughput enzyme library sequencing may use long-read amplicon sequencing for determining DNA sequences of large numbers of enzyme mutants, for example, using long-read sequencing technology.BACKGROUND

[0005] Enzymes are nature’s catalysts, capable of catalyzing a remarkable array of chemical transformations in various compounds. Scientists have long understood the utility of enzymes not only with respect to biochemistry but, also, to industrial chemical synthesis.

[0006] Although much progress has been made in applying naturally obtained enzymes to production of various chemical compounds, for many desirable chemical reactions, naturally-occurring enzymes, that is, those found in nature, do not have the properties required to catalyze their reactions feasibly on an industrial scale. For instance, enzymes exhibiting characteristics such as tolerance to organic solvents, tolerance of biologically-irrelevant pHs, or high-rate conversion at substrate concentrations much greater than what would be present in a biological environment, generally, are not foundin nature.

[0007] Evolutionary processes can be mimicked in a laboratory to direct accumulation of desired parameters over successive iterations of engineering through a process referred to as directed evolution. Directed evolution involves creation of a library of enzyme mutants, evaluating their activity, and selecting a top performing mutant as a parent enzyme for the next round of directed evolution. Over the course of several generations, an enzyme can acquire superior properties making them fit for purpose in industrial chemical production.

[0008] Typically, DNA sequencing is not performed on every single mutant in an enzyme library, since conventional sequencing technologies, like Sanger sequencing, can be cost-prohibitive and time-intensive.

[0009] As such, there is a need for improved methods for DNA sequencing that will enable higher-throughput, such as is necessary to sequence all or portions of an enzyme library.BRIEF SUMMARY OF THE DISCLOSURE

[0010] Disclosed herein are various embodiments of methods for high-throughput sequencing of enzyme libraries using long-read sequencing technology.

[0011] For example, in some embodiments is a method for high-throughput sequencing of enzyme libraries. The method may comprise amplifying a nucleic acid region of interest using forward and reverse primers, each primer comprising a template-binding region, a barcode sequence, and spacer regions, to produce barcoded amplicons. The method may also comprise pooling the barcoded amplicons. The method may also comprise sequencing the pooled barcoded amplicons using long-read sequencing technology.

[0012] Also disclosed herein is a system for high-throughput sequencing of enzyme libraries. The system may comprise PCR primers comprising template-binding regions, barcode sequences, and spacer regions. The system may also comprise automated liquid handling equipment for transferring primers and templates. The system may also comprise thermal cycling equipment for PCR amplification. The system may also comprise data processing software for demultiplexing sequencing data.

[0013] Also disclosed herein is a kit for high-throughput enzyme library sequencing. The kit may comprise a plurality of forward primers, each comprising a template-binding region, a unique barcode sequence, and AT spacer regions. The kit may also comprisea plurality of reverse primers, each comprising a template-binding region, a unique barcode sequence, and AT spacer regions. The kit may also comprise instructions for performing single-round PCR amplification and long-read sequencing.BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Fora detailed description of exemplary embodiments of the disclosure, reference will now be made to the accompanying drawings in which:

[0015] Fig. 1 is an overview of high-throughput long-read sequencing method according to one or more embodiments as disclosed herein;

[0016] Fig. 2 is a diagram of high-throughput long-read sequencing methods according to one of more embodiments as disclosed herein;

[0017] Fig. 3 is: A) a diagrammatic illustration of sequences of example forward and reverse primers used in high-throughput long-read sequencing methods according to one of more embodiments as disclosed herein; and B) binding of the example forward and reverse primers to the vector backbone of a pET-derived vector;

[0018] Fig. 4 is a diagrammatic illustration of preparation of dual primer plates from forward and reverse primer plates according to one or more embodiments as disclosed herein;

[0019] Fig. 5 is a diagrammatic illustration of pooling barcoded amplicons resulting from amplification according to one or more embodiments as disclosed herein;

[0020] Fig. 6 is an illustration of the results of agarose gel electrophoresis of barcoded amplicons according to one or more embodiments as disclosed herein;

[0021] Fig. 7 is an illustration of the results of agarose gel electrophoresis of barcoded amplicons according to one or more embodiments as disclosed herein;

[0022] Fig. 8 is a graph showing an assay standard curve for DNA quantitation as may be employed according to one or more embodiments as disclosed herein;

[0023] Fig. 9 is a histogram of read length for certain high-throughput long-read sequencing methods data according to one or more embodiments as disclosed herein;

[0024] Fig. 10 is a diagrammatic illustration of an overview of a data processing step associated with high-throughput long-read sequencing methods according to one or more embodiments as disclosed herein;

[0025] Fig. 11 is a graph a showing a summary of barcodes identified in high-throughput long-read sequencing methods according to one or more embodiments as disclosed herein; and

[0026] Fig. 12 is a graph showing an overall outcome of sequencing data analysis for high-throughput long-read sequencing methods according to one or more embodiments as disclosed herein.DETAILED DESCRIPTION OF THE DISCLOSED EMBODIMENTS

[0027] The following discussion is directed to various exemplary embodiments. However, one skilled in the art will understand that the examples and embodiments disclosed herein have broad application, and that the discussion of any embodiment is meant only to be exemplary of that embodiment, and not intended to suggest that the scope of the disclosure, including the claims, is limited to that embodiment.

[0028] Certain terms are used throughout the following description and claims to refer to particular features or components. As one skilled in the art will appreciate, different persons may refer to the same feature or component by different names. This document does not intend to distinguish between components or features that differ in name but not function. The drawing figures are not necessarily to scale. Certain features and components herein may be shown exaggerated in scale or in somewhat schematic form and some details of conventional elements may not be shown in interest of clarity and conciseness.

[0029] In the following discussion and in the claims, the terms “including” and “comprising” are used in an open-ended fashion, and thus should be interpreted to mean “including, but not limited to....” Also, the term “couple” or “couples” is intended to mean either an indirect or direct connection. Thus, if a first device couples to a second device, that connection may be through a direct connection, or through an indirect connection via other devices, components, and connections.

[0030] In some embodiments, the disclosed methods for high-throughput sequencing of enzyme libraries the sequencing workflow described herein, enable obtaining DNA sequence data for large numbers of enzyme mutants. A major challenge in directed enzyme evolution is determination of a link between differences in enzyme performance, such as thermostability, expression yield, specific activity, catalyst lifetime, solvent tolerance, pH tolerance, etc., and DNA sequences. This challenge exists because enzyme libraries are typically prepared in a randomized manner, such that each enzyme mutant's sequence is not known.

[0031] DNA sequence data for the enzymes of an enzyme mutant libraries would valuable for training machine learning models, for example, which may be helpful toimprove successive iterations of enzyme design and accelerate the enzyme engineering process. Conventionally, DNA sequencing is not performed on every single mutant from an enzyme library, since conventional technology, such as Sanger sequencing, is cost- prohibitive and time-intensive. For instance, a typical enzyme library may have 1 ,000 unique mutants, and Sanger sequencing requires an experiment for each sequence at a cost of about $5-$ 15 per experiment, resulting in a total cost of about $5,000-$15,000 per library as well as substantial labor cost in sample preparation.

[0032] Approaches to DNA sequencing differ, generally, based upon how the sample preparation is performed and the particular sequencing technology employed with respect to those samples. In terms of sequencing technology, conventionally, sequencing-by-synthesis methods have been used most-commonly; generally, sequencing-by-synthesis methods produce only short DNA reads, with a maximum of about 600 bp, making sequencing of longer genes challenging. More recently, long-read sequencing technology, such as Oxford Nanopore sequencing commercially available from Oxford Nanopore Technologies, has been increasing in popularity as a long-read sequencing technology capable of routine sequencing of fragments 10s to 100s of kilobase pairs (kb) in length. Some prior attempts for high-throughput sequencing methods have relied upon the use of polymerase chain reaction (PCR) to provide each enzyme mutant gene with a “barcode” unique to that DNA sequence in the sample. Another prior attempt for high-throughput sequencing involves using a transposase to simultaneously cut a DNA sequence and attach a unique barcode, in a process often called tagmentation.

[0033] In some embodiments, the present disclosure provides methods for determining sequences for large numbers of enzyme mutants that overcomes the disadvantages of prior methods. In the disclosed subject matter, the disclosed methods, which may be referred to as High-Throughput Long-Read Sequencing (HTLRS) methods, use long- read sequencing rather than short-read sequencing technologies. More particularly, disclosed herein are various embodiments of methods for high-throughput sequencing of enzyme libraries, more particularly, for sequencing a nucleic acid region of interest, using long-read sequencing technology, such as Oxford Nanopore sequencing commercially available from Oxford Nanopore Technologies.

[0034] Generally, the term "nucleic acid region of interest" refers to a specific DNA sequence that encodes an enzyme or protein of interest, for example, which may be the subject of directed evolution or functional analysis. In some embodiments, the nucleicacid region of interest may comprise the complete open reading frame (ORF) of an enzyme-encoding gene, including the start and stop codons, and may additionally include regulatory sequences such as promoter regions, ribosome binding sites, or terminator sequences.

[0035] In some embodiments, the nucleic acid region of interest may be contained within a plasmid vector, such as a pET-derived expression vector that is commonly used for protein expression in bacterial systems. The region of interest may span, for example, from about 1000 to about 2000 base pairs in length, which is typical for enzymeencoding genes, though the method may be employed for regions (e.g., genes) of different lengths. Because the read lengths may span typical about 1000 to about 2000 base-pairs for enzyme-encoding genes, primers incorporating barcode sequences can be designed to bind to sections of a common protein expression vector such as pET that flank the gene of interest, as will be disclosed herein.

[0036] In some embodiments, the nucleic acid region of interest may contain one or more mutations relative to a parent or wild-type sequence. These mutations may be introduced through various mutagenesis techniques including but not limited to site- directed mutagenesis, saturation mutagenesis, random mutagenesis, or combinatorial mutagenesis approaches. In some embodiments, the mutations may be characterized as affecting enzyme properties such as thermostability, expression yield, activity, catalyst lifetime, solvent tolerance, pH tolerance, or combinations thereof.

[0037] In some embodiments, the nucleic acid region of interest may be flanked by known sequences on the plasmid backbone that serve as binding sites for the PCR primers used in the amplification step. These flanking regions may include sequences such as Factor Xa cleavage sites, His-tag sequences, or other vector-specific sequences that are conserved across members of the enzyme library.

[0038] In some embodiments, the nucleic acid region of interest may be part of an enzyme library comprising a plurality of variants, wherein each variant contains a different combination of mutations. For example, the library may contain 10,000 or more unique variants, with each variant representing a different sequence of nucleic acids in the region of interest to be sequenced and analyzed.

[0039] Referring to the embodiment of Fig. 1 , the HTLRS methods may generally comprise the steps of amplifying a nucleic acid region of interest, for example, associated with an enzyme, to produce barcoded amplicons; pooling the barcoded amplicons; and sequencing the pooled barcoded amplicons.

[0040] In some embodiments, the step of amplifying the nucleic acid region of interest generally may comprise providing PCR primers, preparing PCR reaction mixtures, and performing thermal cycling under controlled conditions to generate barcoded amplicons containing the target gene sequences. Referring to Fig. 2, an overview of the step of amplifying the nucleic acid region of interest is shown.

[0041] In some embodiments, the amplification step may utilize PCR primers each that include a unique combination of nucleic acid so as to be uniquely identifiable. Generally, the PCR primers may each comprise an about 15 basepair to about 30 basepair segment, for example, an about 17 to about 19 basepair segment. Each of the PCR primers may be complementary to a region flanking the gene of interest, for example, a region on the plasmid backbone of a suitable vector, such as a pET-derived vector. These segments have melting temperature of about 62° C to about 64°C, as the OligoAnalyzer tool commercially available from Integrated DNA Technologies.

[0042] In some embodiments, the amplification step may utilize forward primers, reverse primers, or both forward and reverse primers, for example, that may work in concert to define the boundaries of the nucleic acid region of interest and enable selective amplification of the target sequence. The forward primer may bind to the 5' end of the target region on one strand of the DNA template, while the reverse primer may bind to the 3' end of the target region on the complementary strand, thereby allowing new DNA strands to be polymerized in the 5' to 3' direction, thereby amplifying the specific region of interest that lies between the two primer binding sites. In some embodiments, the forward primer and reverse primer pair may define the exact boundaries of the amplified product, may ensure that only the desired nucleic acid region of interest is amplified, and may incorporate the necessary barcode sequences for subsequent identification and demultiplexing during sequencing analysis.

[0043] In some embodiments, the PCR primers comprise an about 5 basepair to about 10 basepair, for example, an about 7 bp segment comprising a barcoding sequence. In some embodiments, the basepair sequences used in the barcodes may be derived from previously validated and publicly-available sequencing protocols, for example, which may been proven effective in multiplexed DNA sequencing applications. For example, the sequences of the barcoding segments may comprise an established set of barcode sequences exhibiting low cross-reactivity and accurate demultiplexing during sequencing analysis. Additionally, in some embodiments, the barcode sequences may also be characterized as having been optimized to avoid secondary structure formation,to minimize sequence similarity between different barcodes, and to reduce the likelihood of sequencing errors that could lead to misassignment of reads during the demultiplexing process. In some embodiments, the barcode sequences may be characterized as exhibiting balanced GC content, an absence of homopolymer runs, and sufficient Hamming distance between different barcode sequences to enable error correction during data analysis. Additionally or alternatively, in some embodiments, the barcode sequences may be selected from commercially available barcode sets that have been validated for use with long-read sequencing technology, such as Oxford Nanopore sequencing commercially available from Oxford Nanopore Technologies.

[0044] Additionally, in some embodiments, the use of established barcode sequences may allow for compatibility with existing bioinformatics pipelines and data analysis software that have been developed and validated for processing multiplexed sequencing data. For example, the use of established barcode sequences approach thereby may enables researchers to leverage existing computational tools and protocols rather than requiring development of new analysis methods.

[0045] In some embodiments, the forward primers utilized may match or substantially match (e.g, be identical or substantially identical) for portions of the forward primer sequence outside the barcoding segment; likewise, the reverse primers utilized may match or substantially match (e.g, be identical or substantially identical) for portions of the reverse primer sequence outside the barcoding segment. For example, referring to the embodiment of Fig. 3, each of the forward and reverse primers may comprise three distinct regions, particularly, an AT residue spacer region, a barcode sequence region, and a template-binding region. The template-binding region may be complementary to regions flanking the gene of interest on the plasmid backbone. Because each of the forward primer and each of the reverse primers may share an identical or substantially identical template-binding sequence, for example, an identical or substantially identical sequence outside the respective barcoding segments, the forward and reverse primers may bind to the same complementary region on the opposite strand of the target DNA, regardless of which specific barcode sequence is incorporated into each primer. Not intending to be bound by theory, the identical template-binding regions may ensure consistent amplification efficiency across all primer pairs, as each primer may have the same binding affinity and specificity for its target sequence. Additionally, the standardized sequences outside the barcoding segments may allow for the use of the same PCR conditions and thermal cycling parameters for all reactions, therebysimplifying the experimental protocol and improving reproducibility.

[0046] As such, in some embodiments, the only variable element between different primers of the same type (for example, between all forward primers and / or between all reverse primers) may be the barcode sequence, which may serve as a unique identifier for each sample during the subsequent sequencing and data analysis steps. This approach may allow for the multiplexing of hundreds or thousands of different samples while maintaining consistent amplification performance across the entire library. In some embodiments, suitable PCR primers may be ordered from commercial suppliers standardized to 100 pM in suitable buffer solutions. Examples of suitable PCR primers are set forth in Table 1 .TABLE 1

[0047] In some embodiments, forward and reverse barcoding primers may be combined and put into the wells of a plate, with each well containing a unique combination of forward and reverse primer. In some embodiments, the forward and reverse primers may be added directly to the wells of a PCR reaction plate or, alternatively, may be added to an intermediate plate for later transfer to a PCR reaction plate, for example, a 96-well plate, referred to herein as a “96-well dual barcode plate”.

[0048] Referring to Fig. 4, as an example, the following procedure may describe preparation of 1096-well plates, although those of ordinary skill in the art will appreciate that any suitable number of plates, each having a suitable number of wells may be employed. In this example, each 96-well dual barcode plate may be characterized as having eight (8) rows designed A, B, C, D, E, F, G, and H, respectively, and twelve (12) columns designed 1 through 12, respectively. One full plate of forward primers is ordered, and, in this example, five 96-well plate columns of reverse primers are ordered. The forward and reverse primers may be diluted to about 2 pM by addition of about 980 pL sterile water and about 20 pL of 100 pM forward primer or reverse primer, respectively. Next, about 50 pL from the forward primer plate may be placed into each of the first five 96-well PCR plates, and then to the second five dual barcode 96-well PCR plates rotated 180° relative to the forward primer plate such that well at position A1 of the forward primer plate may be transferred to H12 of the dual barcode plate, as shown in Fig. 4.

[0049] In some embodiments, after transfer of the forward barcode plates, about 50 pL of column 1 of the 2 pM reverse primer plate may be transferred to each column of dual barcode plates 1 and 6, with the orientation of plate 6 matching that of the source reverse primer plate. This may be repeated for the remaining dual barcode plates, iterating the reverse primer plate column by 1 for each subsequent dual index plate.

[0050] In some embodiments, this full process may be automated, such as by the operation of suitable equipment such as the Fluent liquid handler commercially available from Tecan. In various embodiments, the sequence of operations may be designed such that liquid transfer tasks may map well to manual multichannel pipetting.

[0051] In some embodiments, a PCR master mix may be prepared for use in the amplification of the DNA. Generally, the PCR master mix may generally comprise solution that may contain all the essential components necessary for PCR amplification, with the exception of the specific primers and template DNA that may be addedseparately for each individual reaction, for example, within each well.

[0052] In some embodiments, the PCR master mix may generally comprise water, buffer solution, deoxynucleotide triphosphates (dNTPs), and DNA polymerase enzyme. For example, the PCR master mix may be prepared with about 6.75 pL water, about 2 pL of buffer solution, such as 5x OneTaq® buffer commercially available from New England Biolabs, about 0.2 pL of 10 mM dNTPs, and about 0.05 pL per 10 pL reaction of DNA polymerase, such as OneTaq® DNA Polymerase commercially available from New England Biolabs. Those of ordinary skill in the art will appreciate, with the aid of this disclosure, that a suitable PCR master mix.

[0053] In some embodiments, the use of a PCR master mix may provide several advantages in high-throughput applications. For example, the master mix may ensure consistent reaction conditions across multiple samples, may reduce pipetting errors by minimizing the number of individual components that need to be added to each reaction, and may improve reproducibility by standardizing the concentrations of critical reagents.

[0054] In some embodiments, the master mix may be dispensed into wells of a PCR plate matching the number of samples to be processed, for example, via either manual multichannel pipetting or automated liquid handling equipment. The primers and PCR master mix may be transferred to a PCR plate, such as a 384-well PCR plate. For example, the fluids may be transferred using a QPix 420 instrument commercially available from Molecular Devices equipped with Yeast Picker+ deep-well rearraying pins, with three multi-dips performed on both source and destination plates, with washing for about 5 seconds in 10% bleach, sterile water, and 70% ethanol followed by a 10 second drying step between liquid transfers. The rearraying pins may be effective to transfer precise quantities of liquid, for example, about 1 .52 pL.

[0055] In some embodiments, enzyme library DNA template in the form of precultures are transferred to the 384-well PCR plates, for example, via either manual multichannel pipetting or automated liquid handling equipment. The precultures may comprise, for example, E. coli bacterial cultures that have been grown overnight in about 250 pL lysogeny broth medium while shaking at about 1000 rpm at about 30°C. In some embodiments, these precultures may contain plasmid DNA encoding each library member. The E. coli culture may be from the preculture grown to inoculate an expression subculture for downstream enzyme functional screening, thereby allowing sequencing and screening to be performed in parallel.

[0056] The precultures together with the PCR master mix, for example, in eachrespective well, may form the PCR reaction mixture. In some embodiments, the PGR reaction mixture may be subjected to thermal cycling in order to amplify the gene of interest in each of the sample, for example, thereby generating barcoded amplicons containing the target gene sequences. Generally, the thermal cycling conditions used may be those conditions associated with a standard PCR protocol, although the conditions may also be optimized for the specific requirements of the disclosed methods.

[0057] For example, the thermal cycling may each comprise three distinct phases that may work together to exponentially amplify the target sequence. An initial denaturation step may take place at about 95 °C for about 5 minutes to completely denature the double-stranded DNA template, thereby ensuring that all hydrogen bonds between complementary base pairs are broken and making the template strands available for primer binding. The denaturation step may be followed by an annealing step which may take place at about 60 °C for about 30 seconds, allowing the forward and reverse primers to bind to their complementary sequences on the template strands. The annealing step may be followed by an extension step, which may take place at about 68 °C for about 2 minutes, providing conditions for the DNA polymerase to synthesize new DNA strands from the primer binding sites. In some embodiments, the time for the extension step, for example, about 2 minutes, may be specifically chosen to accommodate the particular length of the target amplicons, which may typically be about 2 kb in length for enzyme-encoding genes. The temperature associated with the extension step, for example, a temperature of about 68 °C may represent the optimal temperature for the particular DNA polymerase employed, for example, OneTaq® DNA Polymerase commercially available from New England Biolabs.

[0058] In some embodiments, a suitable number of amplification cycles may be performed, for example, at least about 10 cycles, at least about 15 cycles, at least about 20 cycles, at least about 25 cycles, or at least about 30 cycles. In some embodiments, the final extension step at 68°C may be extended, for example, to about 5 minutes, to ensure that any incomplete DNA synthesis is completed, thereby maximizing the yield of full-length amplicons. Upon completion of the desired number of amplification cycles, the samples may be held at a low temperature, for example, about 8°C, to maintain the samples and preserve the integrity of the amplified products, for example, the barcoded amplicons, until further processing.

[0059] In some embodiments, the amplification step employed in the HTLRS methodsmay be characterized as a single round of PCR, for example, in that only one PGR amplification step is needed to generate the final barcoded amplicons, as opposed to conventional methods that may require multiple sequential PCR steps. More particularly, some conventional two-step approaches may require a first PCR step where gene-specific primers amplify the region of the gene to be sequenced followed by a second PCR step where barcoding primers attach unique barcodes to the amplicons from the first reaction. These primers in the first step may also harbor a constant binding region, to which the barcoding primers in the second round of PCR bind.

[0060] In some embodiments, the amplification step employed in the HTLRS methods may eliminate the need for this two-step process by incorporating the barcode sequences directly into the initial PCR primers. As disclosed herein, the forward and reverse primers used in amplification step employed in the HTLRS methods may contain both the template-binding sequences necessary for amplifying the gene of interest and the unique barcode sequences required for sample identification, thereby allowing both functions to be accomplished in a single PCR reaction.

[0061] In some embodiments, this single-round approach may provide significant advantages over conventional two-step methods. For example, such conventional the two-step PCR protocols may tend to yield product containing excessive primer dimer, which may contaminate the sequencing data even after agarose gel purification. By contrast, the single-round amplification step employed in the HTLRS methods may produce generally cleaner amplicons, for example, as demonstrated by the lack of prominent primer dimer bands when purified via agarose gel electrophoresis. Additionally, in some embodiments, the ability to complete the entire amplification and barcoding process in a single round may also simplify the experimental workflow, reduce the potential for contamination between steps, and decrease the overall time and labor required for sample preparation compared to multi-step approaches.

[0062] In some embodiments, the cleaner PCR products resulting from the single-step approach may translate to several practical benefits during the sequencing phase. For example, the single-round PCR approach may yield improved sequencing coverage consistency as a result of the reduced primer dimer contamination. More particularly, when primer dimers are present in significant quantities, they may compete with the target amplicons for sequencing resources, leading to uneven distribution of sequencing reads across different samples. The cleaner amplicons produced by the HTLRSmethods may therefore allow for more predictable and uniform sequencing coverage across all enzyme variants in the library. Also, relatively more even sequencing coverage across the enzyme library may mean that each variant receives a more consistent number of sequencing reads, which may improve the reliability of mutation identification. Also, fewer enzyme mutants with very few reads may reduce the number of samples that fail to provide sufficient sequencing data for analysis, thereby improving the overall success rate of the sequencing workflow.

[0063] In some embodiments, following the amplification to produce the barcoded amplicons, the barcoded amplicons generated via amplification may be combined, for example, so that the amplicons can be purified prior to sequencing. The pooling and purification process may generally comprise several sequential steps designed to consolidate the amplicons from multiple plates while removing contaminants that could interfere with downstream sequencing.

[0064] In some embodiments, for example, referring to the embodiment of Fig. 5, the pooling process may begin with quenching the PCR reactions to halt any residual polymerase activity. About 5 pL of 100 mM ethylenediaminetetraacetic acid (EDTA) at about pH 8.0 may be placed into the wells of the PCR plates to halt any residual polymerase activity from the PCR mixtures, which could cause undesired amplification of template with pools of combined barcode primers.

[0065] In some embodiments, the pooling may be performed in a hierarchical manner to systematically consolidate samples from multiple plates. For example, each 384-well PCR plate may be condensed into a separate 96-well PCR plate, and each 96-well PCR plate may be condensed into a single Eppendorf tube. From each quadrant of each 384- well PCR plate, about 5 pL of PCR product mixture may be stamped to the corresponding 96-well PCR plate.

[0066] In some embodiments, the pooling process may continue with further consolidation steps. For example, about 10 pL of each well of the 96-well PCR plates may be further condensed into separate columns of a separate 96-well PCR plate. Subsequently, about 60 pL of each well from each column may be transferred to a separate 1.5 ml_ Eppendorf tube. In an example with 10 source plates, the result may be three 96-well PCR plates filled with quenched PCR mixture.

[0067] In some embodiments, the pooling may be performed using automated liquid handling equipment such as the Fluent liquid handler commercially available from Tecan, but because every enzyme mutant gene amplicon has been uniquely barcoded,transfer can be performed by hand via multichannel using one set of pipette tips without worry of cross contamination.

[0068] In some embodiments, the combined PCR amplicons may be purified to remove EDTA, PCR enzymes, and low molecular weight contaminants such as primers and primer dimers. In some embodiments, the combined PCR amplicons may be purified via agarose gel purification. In some embodiments, referring to the embodiment of Fig. 6, the purification may be performed using agarose gel electrophoresis. Gel bands corresponding to intended PCR product, typically about 2 kb in length, may be excised and purified, for example, using the Gel DNA Recovery Kits commercially available from Zymoclean. An example of the results obtained from agarose gel electrophoresis is shown in Fig. 7,

[0069] In some embodiments, after purification, the amount of purified DNA in each of the purified amplicon mixtures may be quantified. For example, the amount of DNA in each of the purified amplicon mixtures may be quantified based upon fluorescence such as by using AccuClear Ultra High Sensitivity dsDNA Quantitation Kit commercially availability from Biotium. Additionally, in some embodiments, one of ordinary skill in the art, with the aid of this disclosure, will be able to adapt the manufacturer's protocol for use of the AccuClear Ultra High Sensitivity dsDNA Quantitation Kit for use with a NanoQuant plate read on an Infinite M200 plate reader, both commercially available from T ecan. As an example, referring to the embodiment of Fig. 8, serial dilutions of the included DNA quantitation standard may be prepared at about 5, 1.25, 0.3125, and 0 ng / pL using the included 1x DNA quantitation buffer, and about 2 pL of each sample may be transferred in duplicate to a NanoQuant plate. Samples may be diluted 20x using 1x DNA quantitation buffer, and about 2 pL of each may be transferred in duplicate to the NanoQuant plate. The

[0070] In some embodiments, the fluorescence measurement process may utilize specific wavelengths, for example, that may be optimized for the AccuClear Ultra High Sensitivity dsDNA Quantitation Kit. Generally, the excitation wavelength of 468 nm may represent the optimal wavelength for exciting the fluorescent dye molecules bound to double-stranded DNA, while the emission wavelength of 507 nm may correspond to the peak emission wavelength of the fluorescent signal produced by the dye-DNA complex.

[0071] In some embodiments, a standard curve may be generated using the serial dilutions of the included DNA quantitation standard prepared at known concentrations of about 5, 1.25, 0.3125, and 0 ng / pL. The fluorescence intensity measured at eachknown concentration may be plotted against the concentration values to determine a linear relationship that may allow for accurate determination of unknown sample concentrations based on their fluorescence readings. For example, referring to the embodiment of Fig. 8, the standard curve may demonstrate a strong linear correlation with an R2value of about 0.994, indicating high reliability for concentration determination. The equation derived from the standard curve may be used to calculate the DNA concentration of each purified amplicon sample based on its measured fluorescence intensity.

[0072] In some embodiments, the process of combining purified amplicon samples to equal concentrations may be effective for ensuring balanced representation of all enzyme variants during the subsequent sequencing step. For example, once the concentration of each purified amplicon mixture has been determined using the standard curve, the volumes may be calculated such that equal amounts of DNA from each sample are combined into a single tube.

[0073] In some embodiments, the sample submission requirements of third-party nanopore sequencing providers, an example, of which is Plasmidsaurus, may specify particular concentration ranges, total DNA amounts, and volume requirements for optimal sequencing performance. As such, the equimolar pooling of samples may ensure that each enzyme variant has an approximately equal probability of being sequenced, thereby preventing bias toward variants that might otherwise be overrepresented due to differences in amplification efficiency or DNA recovery during purification.

[0074] In some embodiments, typically about 1 Gb of sequencing data may be requested per ten 96-well plates to be analyzed, and the final pooled sample may be prepared according to the specific requirements provided by the sequencing service provider to ensure optimal sequencing results.

[0075] In some embodiments, the step of sequencing the pooled barcoded amplicons may generally comprise subjecting the purified and quantified amplicon samples to a sequencing protocol, for example, as may be provided by a third-party long-read sequencing service provider. The step of sequencing the pooled barcoded amplicons may further comprise receiving raw sequencing data and processing the data to assign sequences to individual enzyme variants.

[0076] In some embodiments, once the samples have been purified and combined at approximately equal concentrations, the samples may be sent to a third-party nanoporesequencing provider such as Plasmidsaurus. Typically, about 1 Gb of sequencing data may be obtained per ten 96-well plates to be analyzed. The sequencing may be performed using Oxford Nanopore Technology, which may be capable of routine sequencing of fragments 10s to 100s of kb in length.

[0077] In some embodiments, referring to the embodiment of Fig. 9, the nanopore sequencing data may be received as a single fastq.gz file containing all reads and quality data. Typical raw read quality may be excellent, and length histograms may show very large spikes at the amplicon's size range, which is typically about 2 kbp. For example, about 73.8% of reads may lie within the range of 1900 and 2100 bp, with about 97.5% of reads outside this range being below this range.

[0078] In some embodiments, referring to the embodiment of Fig. 10, the read data obtained from sequencing the samples may be processed to demultiplex the full set of read data, for example, in order to assign each read to a particular barcode plate and well coordinate. In some embodiments, this bioinformatics workflow may be implemented within an R package containing functions for executing different parts of the data processing, as disclosed herein.

[0079] In some embodiments, the data processing may begin with filtering read data for length to remove reads that are too short or too long to contain the full-length amplicon. Typically, reads within about 100 bp of the intended amplicon length may be retained. Next, a portion of the template-binding portion of the forward primers and of the reverse primers may be identified and aligned to each read to determine the direction of the read relative to the enzyme-encoding gene.

[0080] In some embodiments, reads that do not have a single detected read direction may be discarded. Additionally, reverse reads may be reverse-complemented so that all reads may be from the same strand to simplify further data processing. The barcode segments derived from both the forward and reverse barcode primers may be found for each read by extracting the barcode-character string relative to the site of the detected aligned position of the template-binding segment of either primer.

[0081] In some embodiments, referring to the embodiment of Fig. 11 , typically both barcodes can be identified for about one half of total reads filtered for length. For about half of the remaining reads for which at least one barcode was identified, the only barcode identified may be the barcode closest to the start of the read, that is, only the forward barcode for reads in the forward direction. Despite the large fraction of reads that may be unable to be used, sequencing coverage may be typically still far more thansufficient for accurate mutation identification.

[0082] In some embodiments, a dataframe containing every primer pair used in the PCR reactions may be joined with the read dataframe to identify the dual barcode primer plates and wells that amplified each read. This dataframe may be exported in a suitable file formate, for example, as a .csv file. Likewise, read data may be exported in a suitable file format, for example, as .fastq files, with each file containing all reads for a particular source dual barcode plate and well.

[0083] In some embodiments, once the reads have been demultiplexed to determine the source preculture plate and well from which each read originated, the DNA sequences may be analyzed, for example, using a suitable program such as Geneious. For example, the .fastq files grouped by dual barcode plate and well may be imported into Geneious, and a Geneious workflow may be executed that maps each read for each file to the parent gene sequence using, for example, using an alignment program such as MiniMap2.

[0084] In some embodiments, these alignments may be used to generate a consensus sequence, which represents the most likely or accurate sequence for a particular sample based on multiple sequencing reads that have been aligned and analyzed together. Any bases that do not match the parent gene may be annotated as variations. The consensus sequences may be exported as a single .fasta file, and any variation annotations may be exported as variant call .vcf files, with one file per dual barcode plate and well.

[0085] In some embodiments, the processed files from Geneious may be used as input for the continuation of the bioinformatics workflow in R. For example, the .vcf files may be imported to generate a dataframe of mutations for each enzyme library mutant. The consensus .fasta file may be imported and joined with the mutation dataframe. This full summary dataframe may be further processed, exported as a .csv file, and several summary plots may be generated.

[0086] In some embodiments, for example, referring to the embodiment of Fig. 12, analysis of the exported data may reveal a high sequencing success rate. For example, in some embodiments, about 96.5% of library wells may exhibit an anticipated sequencing result, such that sequence data may be found for wells intended to contain mutant or parent sequences, and no sequence data may be mapped to negative control wells. By comparison, using conventional protocols, typical success rates achieved may be closer to about 75%.

[0087] In some embodiments, a system for high-throughput sequencing of enzyme libraries may generally comprise PCR primers, automated liquid handling equipment, thermal cycling equipment, and data processing software that work together to implement the HTLRS methods, as disclosed herein. The system may be configured to provide an integrated platform for processing large numbers of enzyme variants efficiently and accurately.

[0088] In some embodiments, the automated liquid handling equipment may comprise a colony picker with rearraying pins specifically designed for precise liquid transfer operations. For example, the system may utilize a QPix 420 instrument commercially available from Molecular Devices equipped with Yeast Picker+ deep-well rearraying pins (commercially available as product number X4321 ) that may be configured to transfer small amounts of liquid to set up PCR reactions using an array of 96 metal pins. Additionally or alternatively, in some embodiments, the automated liquid handling equipment may implement a Library Compression QPix routine with three multi-dips performed on both source and destination plates.

[0089] Additionally or alternatively, on some embodiments, the system may comprise other automated liquid handling platforms such as a Fluent liquid handler commercially available from Tecan, which may be used for primer plate preparation and sample pooling operations. In various embodiments, the automated liquid handling equipment may be configured to process multiple plate formats, including transfer from 96-well primer plates to 384-well PCR plates, and consolidation of samples from 384-well plates to 96-well plates and ultimately to individual tubes for pooling operations.

[0090] In some embodiments, the data processing software may comprise an R package containing functions for executing different parts of the sequencing data analysis workflow. The software may be specifically designed to handle the demultiplexing and analysis of sequencing data generated from the HTLRS methods. In some embodiments, the R package may include functions for read filtering that may remove reads that are too short or too long to contain the full-length amplicon. In some embodiments, the data processing software may include functions for determining read direction relative to the enzyme-encoding gene by aligning portions of the templatebinding segments of forward and reverse primers to each read. The software may utilize specific primer sequences such as are disclosed herein to determine read orientation. In some embodiments, the software may include barcode extraction functions that may identify and extract the barcode segments from both forward and reverse barcodeprimers by analyzing the 7-character string relative to the detected aligned position of the template-binding segment of either primer. In some embodiments, the data processing software may include functions to identify the dual barcode primer plates and wells that amplified each read. The software may export this information as .csv files and may export read data as .fastq files, with each file containing all reads for a particular source dual barcode plate and well. In some embodiments, the software may be configured to interface with external sequence analysis programs such as Geneious, which may be used for mapping reads to parent gene sequences using MiniMap2, generating consensus sequences, and annotating variations. The software may import .vcf files to generate dataframes of mutations for each enzyme library mutant and may import and join consensus .fasta files with mutation dataframes.

[0091] In some embodiments, the thermal cycling equipment may comprise 384-well PCR thermocyclers that may be configured to accommodate the high-throughput nature of the HTLRS methods. The equipment may be programmed to execute the specific thermal cycling conditions required for the single-round PCR amplification protocol. For example, the thermal cycling equipment may be configured to perform the a program comprising the steps of: as initial denaturation at 95°C for 5 minutes, followed by about 30 cycles of denaturation at 95°C for 15 seconds, annealing at 60°C for 30 seconds, and extension at 68°C for 2 minutes, followed by a final extension at 68°C for 5 minutes, and a hold at 8°C.

[0092] In some embodiments, a kit for high-throughput enzyme library sequencing may generally comprise a plurality of forward primers, a plurality of reverse primers, and instructions for performing single-round PCR amplification and long-read sequencing. The kit may be configured to provide all necessary primer components for implementing the HTLRS methods described herein.

[0093] In some embodiments, the plurality of forward primers may each comprise three distinct structural elements: a template-binding region, a unique barcode sequence, and AT spacer regions. Each forward primer may be designed such that the templatebinding region is identical across all forward primers, while the barcode sequence may be unique to each individual primer. Specific examples of forward primers that may be included in the kit are provided in Table 1. In some embodiments, the kit may include many different forward primers, each with a unique barcode sequence, thereby allowing for multiplexing of up to 96 different samples in a single sequencing run. Likewise, in some embodiments, the plurality of reverse primers may be structured similarly to theforward primers, with each comprising a template-binding region, a unique barcode sequence, and AT spacer regions. Specific examples of reverse primers that may be included in the kit are provided in Table 1 .

[0094] In some embodiments, the instructions included with the kit may provide detailed protocols for preparing dual barcode plates, wherein forward and reverse primers may be combined to create unique primer combinations for each sample.

[0095] In some embodiments, the kit instructions may detail the single-round PGR amplification protocol, including the specific thermal cycling conditions: initial denaturation at 95°C for 5 minutes, followed by 30 cycles of denaturation at 95°C for 15 seconds, annealing at 60°C for 30 seconds, and extension at 68°C for 2 minutes, followed by a final extension at 68°C for 5 minutes.

[0096] In some embodiments, the instructions may also provide guidance for sample pooling, purification using agarose gel electrophoresis, DNA quantification, and preparation of samples for submission to suitable sequencing providers. The kit may thereby enable researchers to implement the complete HTLRS methods for high- throughput enzyme library sequencing.

[0097] Compared with conventional approaches, the present disclosure offers several advantages. For instance, because the sequencing data can span an entire gene, enzyme library design can encompass mutations across the entire gene without affecting the strategy for acquiring sequencing data. By comparison, conventional sequencing methods are usually only effective for sequencing mutant libraries only if the mutagenic region is confined within a few hundred base pairs.

[0098] Additionally, the HTLRS methods may be carried out using automation equipment common in laboratories performing routine enzyme engineering. For example, rather than transferring primers and PCR template cultures to the PGR plate by pipette, the HTLRS methods may be implemented using a QPix colony picker equipped with rearraying pins to transfer liquid to the PCR reaction plate. This method substantially reduces the consumption of pipette tips, thereby enabling savings of hundreds of dollars in sample preparation cost.

[0099] The following are additional embodiments of the disclosed subject matter.

[0100] A 1stembodiment is a method for high-throughput sequencing of enzyme libraries comprising a) amplifying a nucleic acid region of interest using forward and reverse primers, each primer comprising a template-binding region, a barcode sequence, and spacer regions, to produce barcoded amplicons; b) pooling thebarcoded amplicons; and c) sequencing the pooled barcoded amplicons using long- read sequencing technology.

[0101] A 2ndembodiment is the method of the 1stembodiment, wherein the long-read sequencing technology comprises Oxford Nanopore sequencing.

[0102] A 3rdembodiment is the method of one of the ist-2ndembodiments, wherein the template-binding region comprises about 17 to about 19 base pairs complementary to regions flanking a gene of interest on a plasmid backbone.

[0103] A 4thembodiment is the method of one of the 1st-3rdembodiments, wherein the barcode sequence comprises about 7 base pairs.

[0104] A 5thembodiment is the method of one of the 1st-4thembodiments, wherein the spacer regions comprise AT residue spacers flanking the barcode sequence.[OO1O5]A 6thembodiment is the method of one of the 1st-5thembodiments, wherein all forward primers have identical template-binding regions and all reverse primers have identical template-binding regions.

[0106] A 7thembodiment is the method of one of the 1st-6thembodiments, wherein the amplifying the nucleic acid region of interest comprises a single round of PGR amplification.[00i07]An 8thembodiment is the method of the 7thembodiment, wherein the single round of PCR amplification comprises: a) an initial denaturation at about 95°C for about 5 minutes; b) about 30 cycles of a denaturation at about 95°C for about 15 seconds, an annealing at about 60°C for about 30 seconds, and an extension at about 68°C for about 2 minutes; c) a final extension at about 68°C for about 5 minutes; and d) a hold at about 8°C.

[0108] A 9thembodiment is the method of one of the 1st-8thembodiments, wherein the pooling the barcoded amplicons comprises: a) quenching PCR reactions with EDTA; b) consolidating amplicons from multiple plates; and c) purifying the pooled barcoded amplicons using agarose gel electrophoresis.

[0109] A 10thembodiment is the method of one of the 1st-9thembodiments, further comprising quantifying DNA concentration of purified amplicons using fluorescence measurement.[00H0]An 11thembodiment is the method of the 10thembodiment, wherein the fluorescence measurement uses 468 nm excitation and 507 nm emission wavelengths.[OO111]A 12thembodiment is the method of one of the 1st-11thembodiments, furthercomprising demultiplexing sequencing data to assign sequences to individual enzyme variants based on barcode sequences.

[0112] A 13thembodiment is the method of the 12thembodiment, wherein the demultiplexing comprises: a) filtering reads by length; b) determining read direction relative to an enzyme-encoding gene; c) extracting barcode sequences from reads; and d) assigning reads to source samples based on barcode combinations.

[0113] A 14thembodiment is the method of one of the 1st-13thembodiments, wherein the nucleic acid region of interest comprises an enzyme-encoding gene of 1000-2000 base pairs in length.[OO114]A 15thembodiment is the method of one of the 1st-14thembodiments, wherein the enzyme library comprises enzyme variants with mutations affecting thermostability, expression yield, specific activity, catalyst lifetime, solvent tolerance, or pH tolerance.[OO115]A 16thembodiment is the method of one of the 1st-15thembodiments, wherein the primers are transferred using automated liquid handling equipment comprising a colony picker with rearraying pins.

[0116] A 17thembodiment is the method of the 16thembodiment, wherein the rearraying pins transfer approximately 1 .52 pL of liquid.[00H7]An 18thembodiment is the method of one of the 1st-17thembodiments, further comprising generating consensus sequences for each enzyme variant using sequence alignment software.[OO118]A 19thembodiment is the method of the 18thembodiment, further comprising identifying mutations by comparing consensus sequences to a parent gene sequence. [00H9]A 20thembodiment is the method of one of the 1st-19thembodiments, wherein the method achieves a sequencing success rate of at least about 90%.[OO12O]A 21stembodiment is a system for high-throughput sequencing of enzyme libraries comprising: a) PCR primers comprising template-binding regions, barcode sequences, and spacer regions; b) automated liquid handling equipment for transferring primers and templates; c) thermal cycling equipment for PCR amplification; and d) data processing software for demultiplexing sequencing data.

[0121] A 22ndembodiment is the system of the 21stembodiment, wherein the automated liquid handling equipment comprises a colony picker with rearraying pins.

[0122] A 23rdembodiment is the system of one of the 21st-22ndembodiments, wherein the data processing software comprises an R package with functions for read filtering,barcode extraction, and sequence assignment.[OO123]A 24thembodiment is a kit for high-throughput enzyme library sequencing comprising: a) a plurality of forward primers, each comprising a template-binding region, a unique barcode sequence, and AT spacer regions; b) a plurality of reverse primers, each comprising a template-binding region, a unique barcode sequence, and AT spacer regions; and c) instructions for performing single-round PCR amplification and long-read sequencing.[OO124]A 25thembodiment is the kit of the 24thembodiment, wherein the templatebinding regions are complementary to sequences flanking genes in pET-derived expression vectors.

[0125] While embodiments of the disclosure have been shown and described, modifications thereof can be made by one skilled in the art without departing from the scope or teachings herein. The embodiments described herein are exemplary only and are not limiting. Many variations and modifications of the systems, apparatus, and processes described herein are possible and are within the scope of the disclosure. For example, the relative dimensions of various parts, the materials from which the various parts are made, and other parameters can be varied. Accordingly, the scope of protection is not limited to the embodiments described herein, but is only limited by the claims that follow, the scope of which shall include all equivalents of the subject matter of the claims. Unless expressly stated otherwise, the steps in a method claim may be performed in any order. The recitation of identifiers such as (a), (b), (c) or (1 ), (2), (3) before steps in a method claim are not intended to and do not specify a particular order to the steps, but rather are used to simplify subsequent reference to such steps.

Claims

CLAIMSWhat is claimed is:

1. A method for high-throughput sequencing of enzyme libraries, the method comprising: a) amplifying a nucleic acid region of interest using forward and reverse primers, each primer comprising a template-binding region, a barcode sequence, and spacer regions, to produce barcoded amplicons; b) pooling the barcoded amplicons; and c) sequencing the pooled barcoded amplicons using long-read sequencing technology.

2. The method of claim 1 , wherein the long-read sequencing technology comprises Oxford Nanopore sequencing.

3. The method of claim 1 , wherein the template-binding region comprises about 17 to about 19 base pairs complementary to regions flanking a gene of interest on a plasmid backbone.

4. The method of claim 1 , wherein the barcode sequence comprises about 7 base pairs.

5. The method of claim 1 , wherein the spacer regions comprise AT residue spacers flanking the barcode sequence.

6. The method of claim 1 , wherein all forward primers have identical templatebinding regions and all reverse primers have identical template-binding regions.

7. The method of claim 1 , wherein the amplifying the nucleic acid region of interest comprises a single round of PCR amplification.

8. The method of claim 7, wherein the single round of PCR amplification comprises: a) an initial denaturation at about 95°C for about 5 minutes;b) about 30 cycles of a denaturation at about 95°C for about 15 seconds, an annealing at about 60°C for about 30 seconds, and an extension at about 68°C for about 2 minutes; c) a final extension at about 68°C for about 5 minutes; and d) a hold at about 8°C.

9. The method of claim 1 , wherein the pooling the barcoded amplicons comprises: a) quenching PCR reactions with EDTA; b) consolidating amplicons from multiple plates; and c) purifying the pooled barcoded amplicons using agarose gel electrophoresis.

10. The method of claim 1 , further comprising quantifying DNA concentration of purified amplicons using fluorescence measurement.11 . The method of claim 10, wherein the fluorescence measurement uses 468 nm excitation and 507 nm emission wavelengths.

12. The method of claim 1 , further comprising demultiplexing sequencing data to assign sequences to individual enzyme variants based on barcode sequences.

13. The method of claim 12, wherein the demultiplexing comprises: a) filtering reads by length; b) determining read direction relative to an enzyme-encoding gene; c) extracting barcode sequences from reads; and d) assigning reads to source samples based on barcode combinations.

14. The method of claim 1 , wherein the nucleic acid region of interest comprises an enzyme-encoding gene of 1000-2000 base pairs in length.

15. The method of claim 1 , wherein the enzyme library comprises enzyme variants with mutations affecting thermostability, expression yield, specific activity, catalyst lifetime, solvent tolerance, or pH tolerance.

16. The method of claim 1 , wherein the primers are transferred using automatedliquid handling equipment comprising a colony picker with rearraying pins.

17. The method of claim 16, wherein the rearraying pins transfer approximately 1.52 L of liquid.

18. The method of claim 1 , further comprising generating consensus sequences for each enzyme variant using sequence alignment software.

19. The method of claim 18, further comprising identifying mutations by comparing consensus sequences to a parent gene sequence.

20. The method of claim 1 , wherein the method achieves a sequencing success rate of at least about 90%.

21. A system for high-throughput sequencing of enzyme libraries, the system comprising: a) PCR primers comprising template-binding regions, barcode sequences, and spacer regions; b) automated liquid handling equipment for transferring primers and templates; c) thermal cycling equipment for PCR amplification; and d) data processing software for demultiplexing sequencing data.

22. The system of claim 21 , wherein the automated liquid handling equipment comprises a colony picker with rearraying pins.

23. The system of claim 21 , wherein the data processing software comprises an R package with functions for read filtering, barcode extraction, and sequence assignment.

24. A kit for high-throughput enzyme library sequencing, the kit comprising: a) a plurality of forward primers, each comprising a template-binding region, a unique barcode sequence, and AT spacer regions; b) a plurality of reverse primers, each comprising a template-binding region, a unique barcode sequence, and AT spacer regions; andc) instructions for performing single-round PCR amplification and long-read sequencing.

25. The kit of claim 24, wherein the template-binding regions are complementary to sequences flanking genes in pET-derived expression vectors.

Citation Information

Patent Citations

  • Methods of sample preparation

    US20150265995A1

  • Improved high-throughput combinatorial genetic modification system and optimized CAS9 enzyme variants

    US20230193251A1