Assays for measuring nucleic acid modifying enzyme activity
The method of compartmentalized in vitro expression and single-molecule sequencing addresses the limitations of current enzyme engineering by efficiently screening and identifying functional variants of nucleic acid-modifying enzymes, providing a scalable and accurate assessment of enzyme activity.
Patent Information
- Application Number
- JP2022522592
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-15
- Filing Date
- 2020-10-15
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2040-10-15
AI Technical Summary
Current methods for engineering nucleic acid-modifying enzymes, such as CRISPR-Cas proteins, are limited by the inability to efficiently screen and identify functional variants from large libraries, lack scalability, and do not provide information on the degree of protein activity, leading to biased selection towards active variants without distinguishing between highly active and inactive proteins.
A method involving compartmentalization of polynucleotide constructs in isolated compartments, followed by in vitro expression and modification of DNA/RNA targets, and subsequent single-molecule sequencing to detect and count modified and unmodified molecules, enabling high-throughput screening of enzyme variants, guide RNAs, and DNA/RNA targets.
Enables the direct and scalable detection of enzyme activity, allowing for the rapid determination of sequence-function relationships and identification of functional variants, thereby enhancing the engineering of nucleic acid-modifying enzymes and optimizing enzyme activity.
Smart Images

Figure 0007680440000019 
Figure 0007680440000020 
Figure 0007680440000021
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to Singapore Provisional Patent Application No. 10201909632P, filed on October 15, 2019, the entire contents of which are incorporated herein by reference for all purposes.
[0002] FIELD OF THEINVENTION The present invention relates to the field of biotechnology, and in particular to the development of multiplex assays suitable for measuring enzyme activity. [Background technology]
[0003] 2. Background of the Invention Nucleic acid-modifying enzymes, such as zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and clustered regularly interspaced short palindromic repeats (CRISPR)-associated nucleases, have become invaluable tools in both biopharmaceutical research and the biotechnology industry. As therapeutic modalities, these nucleic acid-modifying enzymes have enabled the treatment of previously untreatable genetic diseases by directly modifying DNA or RNA. To realize the enormous industrial and medical potential of nucleic acid-modifying enzymes, the limitations of the natural components must be addressed. These limitations include targeting efficiency and targeting specificity, immunogenicity issues, and compatibility with delivery vectors and functionalizing protein fusion moieties. To address these limitations, enzymes must be modified and their enzymatic activity assayed in a process called protein engineering. Enzymes such as CRISPR-Cas can also be engineered to have enhanced function (e.g., greater specificity for a target, more efficient targeting) and / or novel functions (e.g., base editing, immune evasion, epigenetic modification) by either altering the amino acid sequence of the protein or fusing / co-localizing function-conferring protein domains to the Cas protein or CRISPR complex.
[0004] A general approach to engineering enzymes starts by (i) designing and generating a library of DNA variants that encode many different sequences of the enzyme (with amino acid changes compared to the natural wild type), (ii) expressing these variants in a compartment, e.g., in cells or in vitro, and (iii) measuring or correlating enzyme activity with downstream biochemical reactions or cellular phenotypes, followed by either "screening" (so no selective pressure is applied to segregate active from inactive variants) or "selecting" (so selective pressure is applied to segregate active from inactive variants). Engineering of proteins, especially programmable endonucleases like CRISPR-Cas, has mostly been done by the latter "selection" approach. This approach is dually biased towards active variants (cells live if they express an active version of the protein and die if they express an inactive version of the protein; also called positive selection), does not provide information on the degree of protein activity (e.g. does not distinguish between highly active proteins and half-active proteins), nor does it consider / provide information on inactive protein variants. Negative selection may also be performed, so that only inactive variants are retained and identified, while active variants are depleted and not measured directly. In both cases, activity testing is linked to and manifested by enrichment / depletion of library members. "Screening" approaches are not scalable, because it requires increasing resources to maintain and measure both active and inactive variants. Thus, engineering and assaying nucleic acid modifying enzymes, such as CRISPR-Cas proteins, is limited both in terms of the number of variants that can be tested and in terms of the number of possible mutations per variant. CRISPR-Cas proteins can be engineered to work better, faster, and safer by incorporating multiple amino acid substitutions into the proteins, but current approaches cannot explore this functional space.
[0005] Thus, there is a need for high-throughput screening technology to detect and identify functional variants that can still recognize, cleave or modify their nucleic acid target accurately and efficiently from millions to billions of candidates in enzyme library.Such technology will enable the screening and engineering of novel nucleic acid modifying enzymes, and will also enable the screening and optimization of other factors that affect enzyme activity, such as guide RNA and target sequence.Therefore, the object of the present invention is to provide an improved method that meets the above needs. Summary of the Invention
[0006] In one aspect, the present disclosure provides a method for producing a method for manufacturing a semiconductor device comprising: a) isolating a plurality of polynucleotide constructs in compartments, each compartment containing one polynucleotide construct, each polynucleotide construct comprising: i) a first polynucleotide sequence encoding a nucleic acid modifying enzyme or a variant thereof, operably linked to a first promoter; and ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, where if the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target is driven by the first promoter and co-expressed contiguously with the nucleic acid modifying enzyme as a single RNA transcript; wherein the plurality of polynucleotide constructs encode different variants of the nucleic acid modifying enzyme and / or different DNA or RNA targets; b) subjecting said compartments to conditions that allow in vitro expression of RNA and protein; c) subjecting said plurality of compartments to conditions that allow modification of said DNA / RNA targets by a nucleic acid modifying enzyme having modifying activity against a DNA or RNA target, i. a polynucleotide construct and / or an RNA transcript or fragment thereof modified by said nucleic acid modifying enzyme; ii. polynucleotide constructs and / or RNA transcripts that have not been modified by the nucleic acid modifying enzyme Producing a population of DNA / RNA molecules comprising one or more of: d) recovering the population of DNA / RNA molecules produced in step (c) and subjecting it to single molecule sequencing; e) detecting and counting the DNA / RNA molecules according to steps c)i and c)ii based on the sequencing results; The present invention relates to a method comprising the steps of:
[0007] In another aspect, the present disclosure provides a method for producing a method for manufacturing a semiconductor device comprising: a) isolating a plurality of polynucleotide constructs in compartments, each compartment containing one polynucleotide construct, each polynucleotide construct comprising: i) a first polynucleotide sequence encoding a guide RNA (gRNA), operably linked to a first promoter; ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, where if the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target is driven by the first promoter and co-expressed contiguously with the gRNA as one RNA transcript; wherein the plurality of polynucleotide constructs encode different gRNAs and / or different DNA or RNA targets; and each compartment further comprises an RNA-guided nucleic acid modifying enzyme or a variant thereof, or a nucleotide template encoding same; b) subjecting said compartment to conditions that allow in vitro transcription and / or translation of RNA and protein; c) subjecting said compartment to conditions that allow modification of said DNA and / or RNA targets by an RNA-guided nucleic acid modifying enzyme having functional activity against a DNA or RNA target in the presence of a gRNA; i. a polynucleotide construct and / or an RNA transcript or fragment thereof modified by said nucleic acid modifying enzyme; ii. polynucleotide constructs and / or RNA transcripts that have not been modified by the nucleic acid modifying enzyme Producing a population of DNA / RNA molecules comprising one or more of: d) recovering the population of DNA / RNA molecules produced in step (c) and subjecting it to single molecule long-read sequencing; e) detecting and counting the DNA / RNA molecules according to steps c)i and / or c)ii based on the sequencing results; The present invention relates to a method comprising the steps of:
[0008] In another aspect, the present disclosure relates to a polynucleotide construct comprising a first polynucleotide sequence encoding a nucleic acid modifying enzyme or a variant thereof operably linked to a first promoter, and a second polynucleotide sequence comprising a DNA target.
[0009] In another aspect, the present disclosure relates to a polynucleotide construct comprising a first polynucleotide sequence encoding a nucleic acid modifying enzyme or a variant thereof operably linked to a first promoter, and a second polynucleotide sequence comprising a DNA template encoding an RNA target, wherein the RNA target is driven by the first promoter and co-expressed contiguously with the nucleic acid modifying enzyme as a single RNA transcript.
[0010] In yet another aspect, the present disclosure relates to a construct library comprising a plurality of polynucleotide constructs disclosed herein, the library being characterized by one or more of the following: a) the plurality of polynucleotide constructs encoding different variants of a nucleic acid modifying enzyme; b) the plurality of polynucleotide constructs encoding different DNA or RNA targets.
[0011] In a further aspect, the present disclosure relates to a construct library comprising a plurality of polynucleotide constructs disclosed herein, wherein the library is characterized by one or more of the following: a) the plurality of polynucleotide constructs encode different variants of a nucleic acid modifying enzyme; b) the plurality of polynucleotide constructs encode different DNA or RNA targets; c) the plurality of polynucleotide constructs encode different gRNAs.
[0012] In another aspect, the present disclosure relates to a polynucleotide construct comprising a first polynucleotide sequence encoding a guide RNA (gRNA) operably linked to a first promoter, and a second polynucleotide sequence comprising a DNA target.
[0013] In another aspect, the present disclosure relates to a polynucleotide construct comprising: a first polynucleotide sequence encoding a guide RNA (gRNA) operably linked to a first promoter; and a second polynucleotide sequence comprising a DNA template encoding an RNA target, wherein expression of the RNA target is driven by the first promoter and co-expressed contiguous with the gRNA as one RNA transcript.
[0014] In yet another aspect, the present disclosure relates to a construct library comprising a plurality of polynucleotide constructs disclosed herein, the library being characterized by one or more of the following: a) the plurality of polynucleotide constructs encoding different DNA or RNA targets; b) the plurality of polynucleotide constructs encoding different gRNAs.
[0015] In another aspect, the present disclosure relates to one or more compartments, each of which contains a polynucleotide construct disclosed herein, wherein the compartments are isolated from one another. [The present invention 1001] (a) isolating a plurality of polynucleotide constructs into compartments, each compartment containing one polynucleotide construct, each polynucleotide construct comprising: (i) a first polynucleotide sequence encoding a nucleic acid modifying enzyme or a variant thereof, operably linked to a first promoter; and (ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, wherein if the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target is co-expressed contiguously with the nucleic acid modifying enzyme as a single RNA transcript driven by the first promoter; wherein the plurality of polynucleotide constructs encode different variants of the nucleic acid modifying enzyme and / or different DNA or RNA targets; (b) subjecting said compartments to conditions that allow for in vitro expression of RNA and protein; (c) subjecting said plurality of compartments to conditions that allow modification of said DNA / RNA targets by a nucleic acid modifying enzyme having modification activity against a DNA target or an RNA target, (v) a polynucleotide construct and / or an RNA transcript or fragment thereof modified by said nucleic acid modifying enzyme; (vi) a polynucleotide construct and / or an RNA transcript that has not been modified by the nucleic acid modifying enzyme. Producing a population of DNA / RNA molecules comprising one or more of: (d) recovering the population of DNA / RNA molecules produced in step (c) and subjecting it to single molecule sequencing; (e) detecting and counting the DNA / RNA molecules described in steps (c)(i) and (c)(ii) based on the sequencing results; A method comprising: [The present invention 1002] 1001. The method of claim 1001, wherein the nucleic acid modifying enzyme is an RNA-guided nucleic acid modifying enzyme and each compartment further comprises a guide RNA or a nucleotide template encoding the same. [The present invention 1003] The method of claim 1001, wherein the nucleic acid modifying enzyme is an RNA-guided nucleic acid modifying enzyme, and each polynucleotide further comprises a third polynucleotide sequence encoding a variant guide RNA (gRNA), and the multiple polynucleotide constructs encode different variants of the nucleic acid modifying enzyme, and / or different DNA or RNA targets, and / or different gRNAs. [The present invention 1004] (a) isolating a plurality of polynucleotide constructs into compartments, each compartment containing one polynucleotide construct, each polynucleotide construct comprising: (i) a first polynucleotide sequence encoding a guide RNA (gRNA), operably linked to a first promoter; (ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, wherein if the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target is driven by the first promoter and co-expressed contiguously with the gRNA as a single RNA transcript; wherein the plurality of polynucleotide constructs encode different gRNAs and / or different DNA or RNA targets; and each compartment further comprises an RNA-guided nucleic acid modifying enzyme or a variant thereof, or a nucleotide template encoding same; (b) subjecting said compartment to conditions that allow in vitro transcription and / or translation of RNA and protein; (c) subjecting said compartment to conditions that allow modification of said DNA and / or RNA targets by an RNA-guided nucleic acid modifying enzyme having functional activity against a DNA or RNA target in the presence of a gRNA, (iii) a polynucleotide construct and / or an RNA transcript or fragment thereof modified by said nucleic acid modifying enzyme; (iv) a polynucleotide construct and / or an RNA transcript that has not been modified by the nucleic acid modifying enzyme. Producing a population of DNA / RNA molecules comprising one or more of: (d) recovering the population of DNA / RNA molecules produced in step (c) and subjecting it to single-molecule long-read sequencing; (e) detecting and counting the DNA / RNA molecules described in step (c)(i) and / or (c)(ii) based on the sequencing results; A method comprising: [The present invention 1005] Number of polynucleotide constructs and / or RNA transcripts modified by a nucleic acid-modifying enzyme (Σ count 修飾 ) and multiplied it by the number of polynucleotide constructs and / or RNA transcripts that were not modified by the nucleic acid modifying enzyme (ΣCount 無修飾 ) or the total number of polynucleotide constructs and / or RNA transcripts (ΣCount 修飾 + 無修飾 ) assessing the modification activity of one or more nucleic acid modifying enzymes on one or more of the DNA / RNA targets by The method according to any one of claims 1001 to 1004, further comprising: [The present invention 1006] The enzyme activity is represented by the following formula: TIFF0007680440000001.tif22128 The method of the present invention 1005 is represented by a value calculated using any one of the following: [The present invention 1007] The method of any one of 1001 to 1006, wherein step (d) further comprises disrupting the compartment by physical or chemical methods. [The present invention 1008] The method of any of claims 1001 to 1007, wherein step (d) further comprises purifying the recovered DNA / RNA molecules to remove excess DNA, RNA, and / or protein from the reaction. [The present invention 1009] The method of any of claims 1001 to 1008, wherein said recovered population of DNA / RNA molecules is not subjected to any further modifications other than those required for single molecule sequencing before being subjected to a single molecule sequencing reaction. [The present invention 1010] Any of the methods of claims 1001 to 1009, wherein the detection and counting of DNA / RNA molecules modified or unmodified by the nucleic acid modifying enzyme is based solely on data generated during single molecule sequencing and does not require further modification or processing of the DNA / RNA molecules. [The present invention 1011] the modification activity is a cleavage activity, and the detection and counting of the modified or unmodified polynucleotide construct or RNA transcript is carried out by aligning the sequencing reads of the DNA / RNA molecule to a reference sequence that includes a window of cleavage sites of the nucleic acid modifying enzyme; (i) if the 3' end of the DNA / RNA molecule maps to a region 3' downstream of the cleavage site window, then the DNA / RNA molecule is an unmodified polynucleotide construct or an RNA target; (ii) if the 3' end of the DNA / RNA molecule maps to a region within the window of the cleavage site, then the DNA / RNA molecule is a modified polynucleotide construct or an RNA target; (iii) if the 3' end of a DNA / RNA molecule maps to a region 5' upstream of a cleavage site window, that DNA / RNA molecule is uninformative and is not used to measure modification activity; Any of the methods according to the present invention 1005 to 1010. [The present invention 1012] a first polynucleotide sequence encoding a nucleic acid modifying enzyme or a variant thereof, operably linked to a first promoter; and A second polynucleotide sequence comprising a DNA target A polynucleotide construct comprising: [The present invention 1013] a first polynucleotide sequence encoding a nucleic acid modifying enzyme or a variant thereof, operably linked to a first promoter; and A second polynucleotide sequence comprising a DNA template encoding the RNA target. A polynucleotide construct comprising: The polynucleotide construct, wherein the RNA target is driven by the first promoter and is co-expressed contiguously with the nucleic acid modifying enzyme as one RNA transcript. [The present invention 1014] A construct library comprising a plurality of 10 or 10 polynucleotide constructs of the invention, (a) the plurality of polynucleotide constructs encode different variants of a nucleic acid modifying enzyme; (b) the multiple polynucleotide constructs encode different DNA or RNA targets. The library is characterized by one or more of the following: [The present invention 1015] The polynucleotide construct of any one of claims 1012 to 1013, further comprising a third polynucleotide sequence encoding a guide RNA (gRNA). [The present invention 1016] A construct library comprising a plurality of polynucleotide constructs of the present invention, (a) the plurality of polynucleotide constructs encode different variants of a nucleic acid modifying enzyme; (b) the multiple polynucleotide constructs encode different DNA or RNA targets; (c) the multiple polynucleotide constructs encode different gRNAs. The library is characterized by one or more of the following: [The present invention 1017] a first polynucleotide sequence encoding a guide RNA (gRNA), operably linked to a first promoter; and A second polynucleotide sequence comprising a DNA target A polynucleotide construct comprising: [The present invention 1018] a first polynucleotide sequence encoding a guide RNA (gRNA), operably linked to a first promoter; and A second polynucleotide sequence comprising a DNA template encoding the RNA target. A polynucleotide construct comprising: The polynucleotide construct, wherein expression of the RNA target is driven by the first promoter and is co-expressed contiguous with the gRNA as one RNA transcript. [The present invention 1019] A construct library comprising a plurality of polynucleotide constructs of the present invention 1021 or 1022, (a) the multiple polynucleotide constructs encode different DNA or RNA targets; (b) the multiple polynucleotide constructs encode different gRNAs. The library is characterized by one or more of the following: [The present invention 1020] The method of any one of claims 1001 to 1011 or the polynucleotide construct of any one of claims 1012, 1013, 1015, 1017 and 1018, wherein the first polynucleotide sequence and the second polynucleotide sequence overlap completely or partially. [The present invention 1021] The method of any one of claims 1002 to 1011 or the polynucleotide construct of any one of claims 1015, 1017 and 1018, wherein the DNA target or the RNA target comprises a protospacer that is at least partially complementary to the guide RNA. [The present invention 1022] Any of the methods of 1002 to 1012, any of the polynucleotide constructs of 1012, 1015 and 1017, or any of the methods or polynucleotide constructs of 1020 and 1021, wherein the DNA target also comprises a proximal protospacer adjacent motif (PAM) sequence. [The present invention 1023] Any of the methods of 1002 to 1011, any of the polynucleotide constructs of 1013, 1015 and 1018, or any of the methods or polynucleotide constructs of 1021 to 1023, wherein when the polynucleotide construct comprises a DNA template encoding an RNA target, the RNA target further comprises a proximal protospacer flanking sequence (PFS). [The present invention 1024] The method of any one of claims 1001 to 1011, the polynucleotide construct of any one of claims 1012, 1013, and 1015, or the method or polynucleotide construct of any one of claims 1021 to 1024, wherein the nucleic acid modifying enzyme is a CRISPR-associated protein (Cas). [The present invention 1025] The method of any one of claims 1001 to 1011, the polynucleotide construct of any one of claims 1012, 1013, 1015, or the method or polynucleotide construct of any one of claims 1020 to 1024, wherein the variant nucleic acid modifying enzyme comprises one or more inactivation catalytic sites and is capable of binding to and inhibiting expression of a DNA target without modifying the DNA target. [The present invention 1026] The method of any of claims 1001 to 1011, the polynucleotide construct of any of claims 1012, 1013, 1015, or the method or polynucleotide construct of any of claims 1020 to 1025, wherein the variant nucleic acid modifying enzyme is fused to one or more additional functional domains capable of modifying DNA or RNA. [The present invention 1027] One or more compartments each containing a polynucleotide construct of any of the present invention 1012, 1013, 1015, 1017 and 1018, and isolated from each other. [The present invention 1028] One or more compartments of any of methods 1001 to 1011, methods 1020 to 1026, or methods 1027, wherein each compartment further comprises in vitro transcription and translation (IVTT) reagents, said IVTT reagents enabling in vitro transcription and / or translation of protein and / or RNA. [The present invention 1029] The method of any one of claims 1001 to 1011, one or more of claims 1020 to 1026, or one or more of claims 1027 or 1028, wherein said compartments are emulsion droplets. [The present invention 1030] The method of any one of claims 1001 to 1011, one or more compartments of claims 1020 to 1026, or one or more compartments of claims 1027 to 1029, wherein said isolation is achieved using microfluidics, hydrogel-limited diffusion, or partitioned wells. [Brief description of the drawings]
[0016] The invention will be better understood with reference to the detailed description, taken in conjunction with the non-limiting examples and the accompanying drawings, in which: [Figure 1] FIG. 1 illustrates a non-limiting list of the main concepts and steps of the present invention. Compartmentalization occurs through the creation of water-in-oil emulsion droplets. [Diagram 2]
[0023] Figure 1 illustrates non-limiting examples of polynucleotide constructs disclosed in the present disclosure. "Cas nuclease" may be substituted for any nucleic acid modifying enzyme and may refer to a Cas variant, such as an inactivated Cas nuclease or a Cas protein fused or conjugated to a functional domain. [Diagram 3]Schematic showing how DNA / RNA molecule reads are tallied to calculate enzyme activity in an example where the enzyme is a Cas nuclease and the modification is DNA cleavage. In this example, the DNA target site is 3' of the encoded Cas variant. Nanopore-sequencing reads whose aligned 3' end maps to a site 3' downstream of the window of predicted Cas cleavage sites in the reference sequence (aligned against the reference sequence) are considered uncleaved (dark grey bars in "nanopore-seq aligned reads"; Figure 3), read alignments whose 3' end falls within the window of Cas cleavage sites are considered cleaved (light grey bars in "nanopore-seq aligned reads"; Figure 3), and reads that do not meet either criterion are discarded as uninformative since it is not possible to experimentally determine whether they are cleaved or uncleaved (white bars in "nanopore-seq aligned reads"; Figure 3). [Figure 4]Gel visualization of purified IVTT Sp Cas9 and dCas9 DNA constructs from compartmentalized (with emulsion) versus bulk IVTT reactions. 750 ng of Sp Cas9 construct was mixed with IVTT reagent (New England Biolabs PURExpress #E6800) on ice to make 75 μL of IVTT aqueous mixture. 50 μL of the aqueous mixture was added in five 10 μL portions to the oil and surfactant mixture in ice with a stir bar rotating at 1150 rpm for 2 minutes to make an emulsion mixture. The emulsion mixture was subsequently mixed for an additional minute in ice. In one example, the emulsion mixture was then subjected to homogenization (8000 rpm for 3 minutes; IKA Ultraturrax T10 homogenizer) to result in a more monodisperse distribution of emulsion droplet sizes. The remaining 25 μL of the aqueous mixture was kept in ice for the bulk IVTT reaction as a control. This was repeated with the Sp dCas9 construct. The emulsion-bulk IVTT mixture was then incubated for 4 hours at 37°C to allow IVTT to proceed, followed by 15 minutes at 65°C to inactivate the protein. DNA from all IVTT reactions was then purified separately and aliquots were size separated by gel electrophoresis and then visualized on an agarose gel. This data indicates that the IVTT reagent successfully transcribes and translates the protein in both the bulk reaction and the emulsion droplets. [Diagram 5]Nanopore sequencing reads from an emulsion IVTT autocleavage assay with high input concentration of Sp Cas9 construct. A small subset of reads from this sublibrary mapped to Sp dCas9 and are classified as misassigned in the plot (light grey section; Figure 5) because this emulsion IVTT reaction only received Sp Cas9 DNA as input. Sp Cas9 emulsion IVTT nanopore sequencing reads show that a mixture of cleaved and uncleaved construct fragments was detected (white and black sections, respectively; Figure 5). This data therefore indicates that nanopore single molecule sequencing can detect both modified and unmodified polynucleotide constructs (enzymatically active or inactive products) from emulsion IVTT reactions. [Figure 6] Nanopore sequencing reads from an emulsion IVTT self-cleavage assay with a high input concentration of Sp dCas9 construct. Reads that failed to pass the alignment quality filter were labeled as such. Sp dCas9 emulsion IVTT nanopore sequencing reads appear predominantly as uncleaved construct fragments, as expected (striped grey section; Figure 6). A small subset of reads from this sublibrary mapped to Sp Cas9 and were labeled as misassigned in the plot (light grey section; Figure 6) because only Sp dCas9 DNA was provided as input to this emulsion IVTT reaction. This result supports the robustness of the method, as sequencing reads accurately detect and measure inactivity of Sp dCas9. [Figure 7]Figure 7 shows an exemplary workflow of a time course experiment of bulk IVTT and self-cleavage assay with the resulting nanopore sequencing reads. In this example, bulk IVTT reactions were set up in ice for different CRISPR-Cas constructs (e.g. Sp Cas9, Sa Cas9, As Cpf1, Lb Cpf1), all of which shared a similar arrangement of components as described in the nucleic acid template sequence above. These were then split equally into five corresponding aliquots for each time point (Figure 7 part 1). These bulk IVTT aliquots were then incubated at 37°C and removed at the indicated time points and quenched with EDTA inhibitor and enzyme to stop the IVTT reaction and Cas cleavage of the coding DNA construct (Figure 7 part 2). The quenched IVTT reactions were then processed with SPRIselect bead cleanup to purify the DNA fragments (Figure 7 part 3). Small aliquots of DNA fragments of different Cas orthologs at these different IVTT time points were then visualized on agarose gels after size separation by gel electrophoresis, as shown in Figure 8. The remaining aliquots of purified DNA fragments were then pooled together for each time point, but regardless of the species of Cas at each time point, i.e., Sp Cas9, Sa Cas9, or other DNA fragments, and barcoded individually using the ONT EXP-NBD104 PCR-Free native barcoding expansion kit (Figure 7, part 4), thereby multiplexing these pooled sub-libraries for one run of nanopore sequencing (Figure 7, part 5). The nanopore sequencing results were then filtered to enhance quality and analyzed using publicly available bioinformatics tools, followed by the analytical approach disclosed in this invention. [Figure 8] Gel visualization of purified IVTT constructs of different CRISPR-Cas orthologs from bulk IVTT reactions after steps shown in Figure 7 part 3. This data shows that different Cas proteins (variants or orthologs) are successfully transcribed and translated in bulk reactions. [Figure 9] Plots of Cas-encoding DNA fragments detected by nanopore sequencing from a time course experiment of bulk IVTT and autocleavage assay. The data demonstrate that single-molecule sequencing can detect enzyme products and measure the enzymatic activity of different nucleic acid-modifying enzymes in a multiplexed manner. [Figure 10] Gel visualization of purified IVTT Sp Cas9 and dCas9 DNA constructs from bulk IVTT reactions. 500 ng of Sp Cas9 (sequences as above) was added to IVTT reagent (New England Biolabs PURExpress #E6800) on ice to make 50 μL of IVTT aqueous mixture. The same was done for the Sp dCas9 construct. The Sp dCas9 construct essentially contains the same DNA sequence as the Sp Cas9 construct, but differs in that the Sp dCas9 gene has two inactivating mutations (D10A and H840A) in the Sp Cas9 gene. These 50 μL bulk IVTT reactions were incubated at 37°C for 4 hours to allow IVTT to proceed, followed by protein inactivation at 65°C for 15 minutes. 20 mM EDTA (pH 8.0) inhibitor was added to the bulk IVTT reactions along with RNase cocktail and proteinase K to remove excess RNA and protein from the IVTT reactions at 37°C for 30 minutes. The DNA (polynucleotide constructs) from both bulk IVTT reactions were then purified separately with SPRIselect paramagnetic beads, and aliquots were size separated by gel electrophoresis and visualized on agarose gels. This data indicates that Cas proteins are successfully transcribed and translated in bulk IVTT reactions. [Figure 11]Demonstration of direct detection and counting of polynucleotides modified or not modified by nucleic acid modifying enzymes. Purified Sp Cas9 DNA constructs from bulk IVTT reactions were mixed with Sp dCas9 DNA constructs in different ratios, as shown in gel visualization in Figure 10. Mixtures of these purified DNA constructs were then prepared for nanopore sequencing. By aligning all nanopore sequencing reads to the Sp dCas9 construct reference sequence, the presence of truncated Sp Cas9 reads is detected using bioinformatics tools, which can be performed by those skilled in the art of sequencing data analysis. This workflow allows for the detection of variations (indels (insertions and deletions) or SNPs (single nucleotide polymorphisms)) in the sequenced reads aligned to the reference sequence. Of particular interest was the detection of expected sequence differences between otherwise identical Sp dCas9 constructs and Sp Cas9 constructs, namely SNPs representing Sp dCas9 carrying catalytic inactivation mutations D10A and H840A relative to Sp Cas9. Alignments of raw nanopore sequencing reads were classified into cleavage versus non-cleavage by sequence mapping to the Sp dCas9 reference sequence as shown in Figure 3, and then processed for SNP detection that resulted in amino acid residue changes. The Sp Cas9 sequence of each filtered read alignment was translated to its corresponding amino acid sequence, and SNPs that resulted in amino acid changes from the Sp dCas9 reference amino acid sequence were detected and tabulated. In the plot above, the detected SNPs are represented as a heat map of selected regions of interest in the Sp dCas9 reference that contain the D10A and H840A catalytic inactivation mutations in Sp dCas9. Reads classified as cleavage (left two subplots; Figure 11) were enriched for SNPs corresponding to having the D10 and H840 residues (dark gray boxes in the heat map; Figure 11), i.e., these cleavage reads contained catalytically active Sp Cas9 sequences. The other SNPs resulting in detected amino acid mutations, represented in the plot above as much lighter grey boxes in the heatmap, are false positives arising from raw sequencing errors inherent in currently available nanopore sequencing technologies.This data demonstrates detection of cleaved and uncleaved Sp Cas9 DNA fragments that can be distinguished from detection of uncleaved Sp dCas9 DNA fragments in live nanopore sequencing data. Notably, the method is able to detect cleaved Sp Cas9 DNA fragments even in a 1:10-5 mixture of purified Sp dCas9 bulk IVTT DNA product versus purified Sp Cas9 bulk IVTT DNA product (Figure 11). [Figure 12] Nanopore sequencing reads from an emulsion IVTT self-cleavage assay with limiting input concentration of Sp Cas9 construct. As expected for the Sp Cas9 enzyme, the emulsion IVTT nanopore sequencing reads show that a mixture of cleaved (white section; FIG. 12) and uncleaved (black section; FIG. 12) construct fragments are detected. Thus, this data supports the robustness of the assay in which the IVTT and enzyme reactions are performed in emulsion droplets. A small subset of reads from this sublibrary were mapped to Sp dCas9 and are therefore classified as misassigned in the plot (light grey section; FIG. 12), because only Sp Cas9 DNA was provided as input to this emulsion IVTT reaction. [Figure 13] Nanopore sequencing reads from an emulsion IVTT self-cleavage assay with limiting input concentration of Sp dCas9 construct. As expected, Sp dCas9 emulsion IVTT nanopore sequencing reads appear predominantly as uncleaved construct fragments (striped grey section; Figure 13), demonstrating that Sp dCas9 is mostly inactive. Thus, this data also supports the robustness of the assay in which IVTT and enzymatic reactions are performed in emulsion droplets. A small subset of reads from this sublibrary were mapped to Sp Cas9 and are therefore classified as misassigned in the plot (light grey section; Figure 13), because only Sp dCas9 DNA was provided as input to this emulsion IVTT reaction. [Figure 14]Nanopore sequencing reads from emulsion IVTT self-cleavage assay with limiting input concentrations of Sp Cas9 and Sp dCas9 constructs given in equimolar ratio. Nanopore sequencing reads show roughly equal distribution of Sp Cas9 and Sp dCas9 mapped reads, as expected. Moreover, the Sp Cas9 mapped reads are roughly evenly split between cleaved and uncleaved fragments (white and black sections, respectively; FIG. 14), while the majority of Sp dCas9 mapped reads are classified as uncleaved (striped grey section; FIG. 14). Thus, this data further demonstrates that the method disclosed herein can measure the enzymatic activity of different variants (Cas variants in this example, but the method can also be used to screen variants of other components of the enzymatic reaction, such as targets or gRNAs). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0017] definition Several terms used throughout this specification are defined in the following paragraphs. Additional definitions may be found throughout the body of the specification.
[0018] As used herein, the terms "about" and "approximately" in reference to numbers are used to include numbers within 20%, 10%, 5%, 2.5%, 2%, 1.5%, or 1% in either direction (greater or less) of that number, unless otherwise stated or clear from the context (except where such number would exceed 100% of any possible value).
[0019] The terms "polynucleotide", "nucleic acid", and "oligonucleotide" are used interchangeably and refer to a polymeric form of nucleotides of any length, whether deoxyribonucleotides, ribonucleotides, or analogs thereof. Polynucleotides may have any three-dimensional structure and may perform any function, known or unknown. The following are non-limiting examples of polynucleotides: genes or gene fragments (e.g., probes, primers, ESTs, or SAGE tags), exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides may contain modified nucleotides, such as methylated nucleotides and nucleotide analogs. Modifications to the nucleotide structure, if present, may be imparted before or after formation of the polynucleotide. The sequence of nucleotides may be interrupted by non-nucleotide components. Polynucleotides may be further modified after polymerization, such as by attachment of a labeling component. This term also refers to both double-stranded and single-stranded molecules.Unless otherwise stated or specified, polynucleotide includes the double-stranded form and each of the two complementary single-stranded forms known or predicted to constitute said double-stranded form.As used herein, the term "polypeptide" has a broad, art-recognized meaning of a polymer of amino acids.This term also refers to certain functional classes of polypeptides, such as nucleases, antibodies, etc.
[0020] The term "operably linked," as used herein, refers to a juxtaposition, where the components described herein are in a relationship that allows them to function in their intended manner. A functional element and a control element, e.g., a promoter, that are "operably linked" are joined such that expression and / or activity of the functional element is achieved under conditions compatible with the control element. In some embodiments, an "operably linked" control element is contiguous (e.g., covalently linked) with a coding element of interest, and in some embodiments, the control element acts in trans with respect to or otherwise from the functional element of interest.
[0021] The term "nucleic acid modifying enzyme" refers to a macromolecular biological catalyst, which may be a natural protein or nucleic acid, and is capable of modifying nucleic acid. The term "RNA-guided nucleic acid modifying enzyme" refers broadly to an enzyme that interacts with or forms a complex with a guide RNA, and can specifically target or bind to a polynucleotide of a specific sequence, which usually contains a sequence complementary to the target domain of the gRNA. When bound to a target polynucleotide, the RNA-guided nucleic acid modifying enzyme can remain bound to the target polynucleotide, or cleave the target polynucleotide if the RNA-guided nucleic acid modifying enzyme is a nuclease, or modify the polynucleotide in other ways if it has a functional domain that modifies the polynucleotide in other ways. In one example, the RNA-guided nucleic acid modifying enzyme is a CRISPR-associated protein (Cas). Many Cas proteins have endonuclease activity, and are also called Cas nucleases. In particular examples, the RNA-guided nucleic acid modifying enzyme is selected from the group consisting of Cas3, Cas9, Cas10, Cas12a (also known as Cpf1), Cas13a (also known as C2c2), Cas13b, Cas13c, Cas13d, Cas14, CasX, CasΦ, and variants thereof.
[0022] The terms "guide RNA" and "gRNA" refer to any nucleic acid that aids an RNA-guided nucleic acid modifying enzyme in specifically binding (or "targeting") a target sequence, whether within a cell or in a cell-free environment. gRNAs can be unimolecular (comprising one RNA molecule, also called chimeric) or modular (comprising two or more, typically two separate RNA molecules, such as crRNA and tracrRNA, which are usually linked together, e.g., by duplexing).
[0023] As used herein, the term "target" (or "target site") refers to a nucleic acid sequence that defines a portion of a nucleic acid (or polynucleotide) to which a binding molecule binds when sufficient conditions for binding exist. In some embodiments, a target site is a nucleic acid sequence to which a nucleic acid modifying enzyme described herein binds and / or is modified by such a nucleic acid modifying enzyme. In some embodiments, a target is a nucleic acid sequence to which a guide RNA described herein binds. A target can be single-stranded or double-stranded. The nucleic acid modifying enzymes disclosed herein can modify DNA or RNA. Thus, a "target" can be a DNA sequence or an RNA sequence, and are referred to as a "DNA target" and an "RNA target," respectively. In the context of a dimerizing nuclease, such as a nuclease that contains a FokI DNA cleavage domain, a target typically includes a left half site (where one monomer of the nuclease binds), a right half site (where a second monomer of the nuclease binds), and a spacer sequence where cleavage occurs between the half sites. In some embodiments, the left half site and / or the right half site are 10-18 nucleotides in length. In some embodiments, one or both of the half sites are shorter or longer. In some embodiments, the right half site and the left half site comprise different nucleic acid sequences. In the context of zinc finger nucleases, the target may comprise two half sites, each 6-18 bases in length, flanked on either side by a non-specific spacer region, in some embodiments, 4-8 bases in length. In the context of TALENs, the target may comprise two half sites, each 10-23 bases in length, flanked on either side by a non-specific spacer region, in some embodiments, 10-30 bases in length. In the context of RNA-guided (e.g., RNA-programmable) nucleic acid modifying enzymes, the target typically comprises a nucleotide sequence complementary to a guide RNA (gRNA) (e.g., a "protospacer" for CRISPR-Cas) and a protospacer adjacent motif (PAM) at the 3' or 5' end adjacent to the guide RNA complementary sequence. In the case of RNA-targeting CRISPR-Cas enzymes (e.g., the Cas13 family), the RNA target may contain a protospacer adjacent sequence (PFS) instead of the PAM sequence.The DNA or RNA target of the Cas enzyme, in some embodiments, may comprise a 16-24 nucleotide length complementary to the gRNA, and a 3-6 base pair PAM / PFS (e.g., NNN, where N represents any nucleotide).
[0024] As used herein, "binding" refers to a non-covalent interaction between macromolecules (eg, between a protein and a polynucleotide).
[0025] "Modifying" a polynucleotide refers to any chemical or physical alteration of the components or structure of a polynucleotide, including breaking / cleaving the polynucleotide, creating a nick (single-stranded break) in a double-stranded polynucleotide, substituting one or more nucleotide bases, inserting or deleting one or more nucleotide bases, or covalently modifying nucleotide bases with chemical and epigenetic markers (e.g., methylation and hydroxymethylation of cytosine).
[0026] As used herein, the term "variant" refers to any entity that exhibits significant structural identity with a reference, but is structurally different from the reference in the presence or level of one or more chemical moieties compared to the reference. In many embodiments, a variant also differs functionally from its reference. In general, the basis for whether a particular entity is properly considered to be a "variant" of a reference is the degree of structural identity with the reference. As those skilled in the art will understand, any biological or chemical reference has some characteristic structural element. A variant is defined as a distinct chemical entity that shares one or more of such characteristic structural elements. To give just a few examples, a polypeptide may have characteristic sequence elements that include multiple amino acids that have a defined position relative to each other in linear or three-dimensional space and / or contribute to a specific biological function, and a nucleic acid may have characteristic sequence elements that are composed of multiple nucleotide residues that have a defined position relative to each other in linear or three-dimensional space. For example, a variant polypeptide may differ from a reference polypeptide as a result of one or more differences in the amino acid sequence and / or one or more differences in the chemical moieties (e.g., sugars, lipids, etc.) covalently attached to the polypeptide backbone. In some embodiments, the variant polypeptide exhibits an overall sequence identity with a reference polypeptide (e.g., a nucleic acid modifying enzyme described herein) that is at least 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99%. Alternatively, or in addition, in some embodiments, the variant polypeptide does not share at least one characteristic sequence element with the reference polypeptide. In some embodiments, the reference polypeptide has one or more biological activities. In some embodiments, the variant polypeptide shares one or more of the biological activities of the reference polypeptide, such as an enzymatic activity. In some embodiments, the variant polypeptide does not have one or more of the biological activities of the reference polypeptide.In some embodiments, a variant polypeptide exhibits a reduced level of one or more biological activities (e.g., enzymatic activity) compared to a reference polypeptide. In some embodiments, a polypeptide of interest is considered a "variant" of a parent or reference polypeptide when the polypeptide of interest has an amino acid sequence identical to the parent amino acid sequence except for a small number of sequence changes at specific positions. Typically, fewer than 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% of the residues in the variant are substituted compared to the parent. In some embodiments, a variant has 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 substituted residues compared to the parent. Often, a variant has a very small number (e.g., fewer than 5, 4, 3, 2, or 1) of functional residues (i.e., residues that contribute to a specific biological activity) substituted. Furthermore, a variant typically has no more than 5, 4, 3, 2, or 1 additions or deletions compared to the parent, and often has no additions or deletions. Moreover, any additions or deletions are typically fewer than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, about 6 residues, and generally fewer than about 5, about 4, about 3, or about 2 residues. In some embodiments, the parent or reference polypeptide is one found in nature.
[0027] The term "library" as used herein, in the context of nucleic acids or proteins, refers to a population of two or more different polynucleotide constructs or proteins, respectively. In some embodiments, a library of polynucleotide constructs includes at least two polynucleotide constructs comprising different sequences encoding nucleic acid modifying enzymes, at least two polynucleotide constructs comprising different sequences encoding guide RNAs, at least two polynucleotide constructs comprising different PAMs, and / or at least two nucleic acid molecules comprising different target sites. In some examples, a library includes at least 10 1 , at least 10 2 , at least 10 3 , at least 10 4 , at least 105 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11 , at least 10 12 , at least 10 13 , at least 10 14 , or at least 10 15 In some embodiments, the library members include different nucleic acid templates. In some embodiments, the library members can include randomized sequences, e.g., fully or partially randomized sequences. In some embodiments, the library includes nucleic acid molecules that are unrelated to one another, e.g., nucleic acids that include fully randomized sequences. In other embodiments, at least some members of the library can be related, e.g., they can be variants or derivatives of a particular sequence.
[0028] As used herein, the term "expression" of a nucleic acid sequence refers to the production of any gene product from the nucleic acid sequence. In some examples, the gene product can be an RNA transcript. In some embodiments, the gene product can be a polypeptide. In some embodiments, expression of a nucleic acid sequence includes one or more of the following: (1) production of an RNA template from a DNA sequence (e.g., by transcription); (2) processing of the RNA transcript (e.g., by splicing, editing, 5' capping, and / or 3' end formation); (3) translation of the RNA into a polypeptide or protein; and / or (4) post-translational modification of the polypeptide or protein.
[0029] The term "compartment," as used herein, in the context of isolating polynucleotide constructs within a compartment, can refer to any physical or virtual compartment, such as emulsion droplets and nanowells, as well as the segregation of reagents and reactions enabled by virtual compartments, such as microfluidics or hydrogels.
[0030] The term "promoter" as used herein refers to a transcription promoter that provides precise transcription initiation. Promoters as used herein include any promoter that can be used to produce mRNAs that code for proteins (e.g., Cas proteins) or RNA transcripts (e.g., guide RNAs). In some examples, the promoter is compatible with cell-free in vitro transcription and translation reactions. Examples of promoters that can be used in the context of the present invention include, but are not limited to, T7 promoter, SP6 promoter, Lac promoter, and others. The term "terminator" as used herein refers to a transcription terminator that defines the end of a transcription unit (e.g., a gene) and initiates the process of releasing newly synthesized RNA from the transcription machinery. Examples of terminators include, but are not limited to, the T7 terminator and the rrnB terminator.
[0031] Detailed Description of the Invention The inventors of the present invention have developed, inter alia, a multiplex method for measuring the activity of nucleic acid modifying enzymes and screening one or more variables of the enzymatic reaction.For example, the method can physically associate the activity of nucleic acid modifying enzyme variants with their own coding DNA / RNA and their DNA / RNA target molecules, while simultaneously directly molecular measuring both the enzymatic activity and inactivity of each variant with respect to each individual target molecule, regardless of activity level (i.e., quantifying how active an active variant is, while inactive variants can also be measured as "inactive").This allows a direct route (through active variants) toward engineering nucleic acid modifying enzymes (e.g., CRISPR-Cas) to have enhanced or novel functionality, while at the same time creating a fitness landscape map of currently unproductive sequence variations (through inactive or weakly active variants).Similarly, variants of guide RNA and / or DNA / RNA target can also be screened using the methods disclosed herein.
[0032] A non-limiting and non-exclusive list of the key concepts of the present invention are given below: (i) a polynucleotide construct encoding the DNA / RNA target site and variable elements (e.g., nucleic acid modifying enzyme variants) to be tested; (ii) mixing the DNA with any of the commonly available RNA and protein expression reagents (so-called cell-free transcription-translation (TXTL) / in vitro transcription-translation (IVTT) reactions); (iii) encapsulation or compartmentalization of a single copy of the DNA construct variant with the IVTT reagents; (iv) running the IVTT reaction in an individual compartment, isolated from multiple other compartments, to express the nucleic acid modifying enzymes and (enzymes) in each compartment. (v) expressing the sgRNA (and in some embodiments, the RNA target co-transcribed as part of the Cas transcript) depending on the function of the encoded nucleic acid modifying enzyme; (vi) directly identifying and directly quantifying the enzymatic activity associated with each variable element in a molecular parallel manner by quantitating the cleaved, intact, or modified polynucleotide constructs (or in some embodiments, the RNA target) in parallel, e.g., by single molecule long read sequencing. This technology directly relates the phenotype of an encoded variable element (e.g., a nucleic acid modifying enzyme variant) to its coding sequence, allowing for the rapid determination of sequence-function relationships for large variant libraries. Figure 1 shows a non-limiting list of the main concepts of the invention.
[0033] method The method disclosed herein can be characterized as a method for measuring enzyme activity. Because the method is highly scalable and can screen a large number of variant polynucleotides, it can also be characterized as a method for screening nucleic acid modifying enzymes, and / or DNA / RNA targets (of the nucleic acid modifying enzymes), and / or guide RNAs, and / or other components of enzymatic reactions that may be encoded on polynucleotide constructs. Thus, in one aspect, the present disclosure provides: a) isolating a plurality of polynucleotide constructs in compartments, each compartment containing one polynucleotide construct, each polynucleotide construct comprising: i) a first polynucleotide sequence encoding a nucleic acid modifying enzyme or a variant thereof, operably linked to a first promoter; and ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, wherein if the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target is co-expressed as a continuous transcript with the nucleic acid modifying enzyme driven by the first promoter; wherein the plurality of polynucleotide constructs encode different variants of the nucleic acid modifying enzyme and / or different DNA or RNA targets; b) subjecting said compartments to conditions that allow in vitro expression of RNA and protein; c) subjecting said plurality of compartments to conditions that allow modification of said DNA / RNA targets by a nucleic acid modifying enzyme having modifying activity against a DNA or RNA target, iii. a polynucleotide construct and / or an RNA target or fragment thereof modified by said nucleic acid modifying enzyme; iv. A polynucleotide construct and / or an RNA target that has not been modified by the nucleic acid modifying enzyme. Producing a population of DNA / RNA molecules comprising one or more of: d) recovering the population of DNA / RNA molecules produced in step (c) and subjecting it to single molecule sequencing; e) detecting and counting the DNA / RNA molecules according to steps c)i and c)ii based on the sequencing results; The present invention relates to a method comprising the steps of:
[0034] In this first aspect, the polynucleotide construct encodes both the nucleic acid modifying enzyme (or its variant) and the DNA / RNA target. Thus, the nucleic acid modifying enzyme or the DNA / RNA target can be tested or screened as a variable. In some examples where the method is used to test or measure the activity of different nucleic acid modifying enzymes on a specific target (i.e., screening enzymes), multiple polynucleotide constructs can encode the same DNA / RNA target but different nucleic acid modifying enzymes (or different variants of the same nucleic acid modifying enzyme). In some examples where the method is used to test or measure the activity of a specific nucleic acid modifying enzyme on different DNA / RNA targets (i.e., screening DNA / RNA targets), multiple polynucleotide constructs can encode the same nucleic acid modifying enzyme but different DNA / RNA targets. In the context of CRISPR-Cas targets, the phrase "different DNA / RNA targets" may refer to DNA / RNA targets that differ in protospacer (sequence complementary to guide RNA) or PAM / PFS sequence.
[0035] In some examples, the nucleic acid modifying enzyme encoded by each polynucleotide construct is an RNA-guided nucleic acid modifying enzyme (e.g., CRISPR-Cas nuclease or its variant), the nucleic acid modifying enzyme may require a guide RNA (gRNA) to bind to and / or modify the DNA / RNA target. In some examples, the gRNA is directly provided to each compartment in the form of a gRNA or in the form of a DNA template that codes for the gRNA. Thus, in one example, the nucleic acid modifying enzyme is an RNA-guided nucleic acid modifying enzyme, each compartment further comprises a guide RNA or a nucleotide template that codes for the guide RNA.
[0036] In some other examples, the gRNA can be encoded on the same polynucleotide construct that encodes the enzyme and the DNA / RNA target. Thus, in one example, the nucleic acid modifying enzyme is an RNA-guided nucleic acid modifying enzyme, and each polynucleotide further comprises a third polynucleotide sequence encoding a variant guide RNA (gRNA), and the multiple polynucleotide constructs encode different variants of the nucleic acid modifying enzyme, and / or different DNA or RNA targets, and / or different gRNAs. In this example, the method disclosed herein comprises: a) isolating a plurality of polynucleotide constructs in compartments, each compartment containing one polynucleotide construct, each polynucleotide construct comprising: i) a first polynucleotide sequence encoding a nucleic acid modifying enzyme or a variant thereof operably linked to a first promoter, wherein the nucleic acid modifying enzyme is an RNA-guided nucleic acid modifying enzyme; ii) a second polynucleotide sequence comprising a DNA target or comprising a DNA template encoding an RNA target, where if the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target is driven by the first promoter and co-expressed contiguously with the nucleic acid modifying enzyme as one transcript; and iii) a third polynucleotide sequence encoding a variant guide RNA (gRNA) wherein the plurality of polynucleotide constructs encode different variants of the nucleic acid modifying enzyme and / or different DNA or RNA targets, and / or different gRNAs; b) subjecting said compartments to conditions that allow in vitro expression of RNA and protein; c) subjecting said plurality of compartments to conditions that allow modification of said DNA / RNA targets by a nucleic acid modifying enzyme having functional activity against a DNA or RNA target in the presence of a gRNA; i. a polynucleotide construct and / or an RNA target or fragment thereof modified by said nucleic acid modifying enzyme; ii. a polynucleotide construct and / or an RNA target that has not been modified by said nucleic acid modifying enzyme; Producing a population of DNA / RNA molecules comprising one or more of: d) recovering the population of DNA / RNA molecules produced in step (c) and subjecting it to single molecule long-read sequencing; e) detecting and counting the DNA / RNA molecules according to steps c)i and c)ii based on the sequencing results; Includes.
[0037] In the above examples, the gRNA (or a sequence encoding the gRNA) is physically linked to the nucleic acid modifying enzyme and the DNA / RNA target, so that any of the gRNA, DNA / RNA target, and nucleic acid modifying enzyme can be tested and screened as variables. Since the encoded gRNA will be expressed from a polynucleotide construct (e.g., in a compartment), the polynucleotide construct may contain other elements that facilitate expression of the gRNA, which are generally known to those skilled in the art. In some examples, the third polynucleotide sequence is operably linked to a second promoter. In some examples, the second promoter is a T7 promoter.
[0038] Screening for nucleic acid modifying enzymes In some examples where the methods are used to test or measure the activity of different nucleic acid modifying enzymes against a particular target, the multiple polynucleotide constructs may encode the same DNA / RNA target and gRNA, but different nucleic acid modifying enzymes (or different variants of the same nucleic acid modifying enzyme).
[0039] One example of nucleic acid modifying enzymes that can be tested, screened and optimized using this method is the Cas family of nucleases. In recent years, various CRISPR-Cas systems have been developed for DNA and RNA editing, enabling a wide range of applications that impact all areas of medicine and biotechnology. Of particular interest are class 2 Cas (CRISPR-associated) proteins, including Cas9, Cas12 (previously known as Cpf1), Cas13 and Cas14 nucleases, which have been well characterized in the literature. These Cas proteins are single-component nuclease effectors (i.e., single Cas protein, not a multimeric complex of different proteins), and typically utilize RNA oligonucleotides (guide RNA, gRNA; variants also known as single guide RNA, sgRNA; used interchangeably) to program Cas proteins to colocalize to specific locations on DNA and / or RNA, after which enzymatic activity such as cleavage (breakage by nucleotide strand cleavage in DNA / RNA) can occur. A segment of the gRNA sequence (spacer) is complementary to the DNA / RNA target sequence (protospacer). For functional targeting, another short sequence (typically 2-6 nt long) adjacent to the protospacer is required, also known as the protospacer adjacent motif (PAM; for DNA) or protospacer adjacent sequence (PFS; for RNA). Each Cas-gRNA system can recognize a unique PAM / PFS site and has different gRNA:protospacer requirements. Cas proteins have been and can be engineered to recognize new PAM / PFS sites, to have less stringent gRNA lengths or structures, and to be more specific and highly efficient. To minimize adverse immune responses while using Cas nucleases as therapeutics, immunogenic epitopes of Cas proteins can also be removed or masked, specifically by deleting or altering amino acid sequences while maintaining Cas function.New functions can also be engineered into Cas proteins or Cas fusion proteins to provide, for example, base editing (changing a targeted nucleotide to another), epigenetic modifications, or many other modifications yet to be demonstrated. These efforts typically involve some form of directed evolution, protein engineering, selection, and screening of Cas variant libraries. The methods disclosed herein are useful for measuring and screening the activity of large libraries of enzyme (e.g., Cas) variants, because they are: i) highly scalable, >10 9 IVTT reaction droplets / mL in parallel and ii) can accommodate a relatively large sequence space, the latter being particularly useful for large proteins (>10 3 This is important and useful when investigating the length of aa.
[0040] Screening DNA / RNA targets of nucleic acid modifying enzymes In some examples where the method is used to test or measure the activity of a particular nucleic acid modifying enzyme on different DNA / RNA targets in the presence of a particular gRNA, the multiple polynucleotide constructs may encode the same nucleic acid modifying enzyme and gRNA, but different DNA / RNA targets.
[0041] In these examples, the methods disclosed herein can be used to evaluate the ability of PAM or PFS variants to direct the binding or modification of DNA / RNA targets by RNA-guided nucleic acid modifying enzymes. The methods disclosed herein allow for the simultaneous assessment of multiple PAM / PFS variants for any given target site. Thus, data obtained from such methods can be used to compile a list of PAM variants that modify (e.g., cleave) a particular DNA / RNA target. As will be readily apparent to those skilled in the art, any non-PAM / PFS sequence on the target site that may have an effect on the activity of the enzyme can also be tested and screened using this method.
[0042] Guide RNA screening In some examples where the method is used to test or measure the activity of a specific nucleic acid modifying enzyme on a specific DNA / RNA target in the presence of different specific gRNAs, the multiple polynucleotide constructs may encode the same nucleic acid modifying enzyme and DNA / RNA target, but different gRNAs.
[0043] In these examples, the present disclosure provides methods to assess the ability of different gRNAs to mediate the binding and / or modification of a nucleic acid modifying enzyme to a specific DNA / RNA target. Thus, the results obtained from this method can be used to compile a list of guide RNA variants that mediate the modification of a specific target by a specific nucleic acid modifying enzyme.
[0044] In another aspect, the present disclosure provides a method for producing a method for manufacturing a semiconductor device comprising: a) isolating a plurality of polynucleotide constructs in compartments, each compartment containing one polynucleotide construct, each polynucleotide construct comprising: i) a first polynucleotide sequence encoding a guide RNA (gRNA), operably linked to a first promoter; ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, where if the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target is driven by the first promoter and co-expressed contiguously with the gRNA as one RNA transcript; wherein the plurality of polynucleotide constructs encode different gRNAs and / or different DNA or RNA targets; and each compartment further comprises an RNA-guided nucleic acid modifying enzyme or a variant thereof, or a nucleotide template encoding same; b) subjecting said compartment to conditions that allow in vitro transcription and / or translation of RNA and protein; c) subjecting said compartment to conditions that allow modification of said DNA and / or RNA targets by an RNA-guided nucleic acid modifying enzyme having functional activity against a DNA or RNA target in the presence of a gRNA; i) a polynucleotide construct and / or an RNA transcript or fragment thereof modified by said nucleic acid modifying enzyme; ii) polynucleotide constructs and / or RNA transcripts that have not been modified by the nucleic acid modifying enzyme Producing a population of DNA / RNA molecules comprising one or more of: d) recovering the population of DNA / RNA molecules produced in step (c) and subjecting it to single molecule long-read sequencing; e) detecting and counting the DNA / RNA molecules according to steps c)i and c)ii based on the sequencing results; The present invention relates to a method comprising the steps of:
[0045] In this aspect, the polynucleotide construct encodes a guide RNA (gRNA) and a DNA / RNA target, but the RNA-guided nucleic acid modifying enzyme is provided separately in each compartment. Thus, the gRNA or the DNA / RNA target can be tested or screened as a variable element. In some examples where the method is used to test or measure the activity of a specific nucleic acid modifying enzyme on a specific target in the presence of different gRNAs (i.e., screening gRNAs), the multiple polynucleotide constructs can encode the same DNA / RNA target, but different nucleic acid modifying enzymes (or different variants of the same nucleic acid modifying enzyme). In some examples where the method is used to test or measure the activity of a specific nucleic acid modifying enzyme on different DNA / RNA targets (i.e., screening DNA / RNA targets), the multiple polynucleotide constructs can encode the same gRNA, but different DNA / RNA targets.
[0046] Intracompartmental Sequestration of Polynucleotide Constructs There are several methods known to those skilled in the art to sequester polynucleotide constructs in compartments. In one example, polynucleotide constructs are sequestered in emulsion droplets by emulsification methods well known in the art. In general, emulsions can be produced from any suitable combination of immiscible liquids. In one typical example, emulsions include an aqueous phase, which contains (a) the components necessary for in vitro transcription and translation, and (b) the library of nucleic acid templates described herein. In emulsions, the aqueous phase is present in the form of finely divided droplets (dispersed, internal, or discontinuous phase). Emulsions further include a hydrophobic immiscible liquid ("oil") as a matrix (non-dispersed, continuous, or external phase) that suspends the droplets. Such emulsions are called "water-in-oil" (W / O) and the droplets are called "water-in-oil droplets". Many oils and many emulsifiers are known in the art and can be used to create water-in-oil emulsions. Suitable emulsifiers include, for example, light white mineral oil, and a surfactant, such as sorbitan monooleate (Span 80; ICI) and polyoxyethylene sorbitan monooleate (Tween 80; ICI), or any combination thereof. In one example, the emulsifier includes mineral oil, Span 80, and a surfactant, such as Tween 80; for example, mineral oil + 4.5% (v / v) Span 80 + 0.5% (v / v) Tween 80). Testing different emulsifiers is within the knowledge of one of ordinary skill in the art. In some examples, emulsions are produced by forcing the phases together with mechanical energy. A variety of methods can be used, including but not limited to the use of mechanical devices such as stirrers (e.g., magnetic stir bars, propeller and turbine stirrers, paddle devices, and whisks), homogenizers (e.g., rotor-stator homogenizers, high pressure valve homogenizers, and jet homogenizers), colloid mills, and ultrasonic and "membrane emulsification" devices. The size of the emulsion droplets (compartments) can be varied by one of skill in the art by adjusting the emulsion conditions used to form the emulsion according to the requirements of the system of choice.
[0047] A non-limiting example is now described: the creation of water-in-oil (w / o) emulsion droplets using the following steps, or other methods known to those skilled in the art. Briefly, 950 μL of oil and surfactant mixture (mineral oil + 4.5% (v / v) Span 80 + 0.5% (v / v) Tween 80) is added to a cryovial with a 3 x 8 mm magnetic stir bar and placed on ice. ≦1.66 fmol of DNA library is mixed with IVTT reagent (New England Biolabs PURExpress #E6800) on ice to make 50 μL of IVTT aqueous mixture. This 50 μL aqueous mixture is added in five 10 μL portions to the oil and surfactant mixture over a period of 2 minutes on ice with the stir bar rotating at 1150 rpm to make the emulsion mixture. The emulsion mixture is then allowed to mix for an additional minute on ice. In one example, the stirred emulsion mixture is mixed with a homogenizer (eg, an IKA Ultraturrax T10 homogenizer) for an additional 3 minutes at 8000 rpm to obtain a more monodisperse distribution of emulsion droplet diameters.
[0048] Other methods of creating emulsion droplets are possible and will be appreciated by those skilled in the art, including vortexing a mixture of water and oil, or using a microfluidic device such as Dolomite-Bio's microencapsulator to control the flow rates of water and oil inputs fed to a microfluidic chip junction to encapsulate the aqueous solution in oil into emulsion droplets.
[0049] Other compartmentalization methods are known to those skilled in the art. Both virtual and physical compartmentalization are encompassed by the term "compartment" as used herein, as long as the compartmentalization allows for the isolation of polynucleotide constructs, reagents, and reactions without creating physical encapsulation. In one example, compartmental isolation is achieved using microfluidics, hydrogel-limited diffusion, or partitioned wells (or nanowells).
[0050] IVTT system In some examples of the methods disclosed herein, each compartment comprises an in vitro transcription and translation (IVTT) reagent, which allows for in vitro transcription and / or translation of proteins and / or RNA. When an IVTT is included in a compartment, the use of cells in the assay is omitted. In some embodiments, the IVTT system comprises a cell extract, for example from bacteria, rabbit reticulocytes, or wheat germ. Many suitable systems are commercially available (for example, from ThermoFisher, Promega, and New England Biolabs). In one example, the system can be emulsified with the polynucleotide construct. Suitable conditions for in vitro transcription and translation as described in step b) will be clear or accessible to those skilled in the art by referring to the literature or the manual of a commercially available kit. In one non-limiting example, suitable conditions are incubation at 37° C. for 4 hours. The IVTT reaction can be stopped by methods well known in the art or described in the manual of a commercially available kit. In one example, the compartment containing the IVTT is incubated at 65° C. for 15 minutes to heat inactivate the IVTT reagent and any expressed nucleic acid modifying enzymes. In another example, 20 mM EDTA (pH 8.0) inhibitor is added to the compartment (e.g., emulsion droplets) and mixed.
[0051] Controlling the compartmentalization conditions of the IVTT reagent with the DNA ensures that only one copy of the polynucleotide construct is packaged in each compartment along with the IVTT reagent in volumes in the femtoliter to nanoliter range. This allows copies of each variant of DNA (and also the IVTT RNA and protein products) to be physically isolated within each compartment, thus allowing the user to physically confine the expressed RNA and protein along with the DNA that encodes them.
[0052] The conditions that allow modification of DNA / RNA targets by known nucleic acid modifying enzymes are generally known in the art and / or can be easily discovered or optimized. For newly discovered enzymes, such conditions can generally be approximated using information about closely related nucleases (e.g., homologs and orthologs) that are better characterized. Modification can refer to any chemical or physical alteration of the components or structure of the target, including breaking / cleaving polynucleotides, creating nicks (single-strand breaks) in double-stranded polynucleotides, substituting one or more nucleotide bases, inserting or deleting one or more nucleotide bases, or covalently modifying nucleotide bases with chemical and epigenetic markers (e.g., methylation and hydroxymethylation of cytosine).
[0053] Since each compartment contains one copy of the polynucleotide construct, the DNA / RNA target (contained on or expressed from the polynucleotide construct) and the nucleic acid modifying enzyme are also confined to the compartment. The activity (or lack thereof) of the nucleic acid modifying enzyme encoded on a particular construct against a DNA / RNA target encoded on the same construct is manifested as a modification (or lack thereof) of the DNA / RNA target. Since multiple compartments collectively contain multiple different polynucleotide constructs, step c) produces a population of DNA / RNA molecules, which can be: i. a polynucleotide construct and / or an RNA transcript or fragment thereof modified by said nucleic acid modifying enzyme; ii. polynucleotide constructs and / or RNA transcripts that have not been modified by the nucleic acid modifying enzyme one or more of In this case, if the polynucleotide construct comprises a DNA target and the encoded nucleic acid modifying enzyme has activity against the DNA target, the polynucleotide construct will be modified. Thus, the state of the polynucleotide construct (modified or unmodified) is associated with the enzyme by the enzyme-specific sequence contained on the same construct. If the polynucleotide construct comprises a DNA template encoding an RNA target, the RNA target will be contained on the transcript RNA expressed from the DNA template. Since the RNA target is co-expressed contiguously with the nucleic acid modifying enzyme as a single transcript, the state of the RNA target (modified or unmodified) is also associated with the enzyme by the enzyme-specific sequence contained on the RNA transcript.
[0054] Recovery of DNA / RNA molecules To measure the activity of the nucleic acid modifying enzyme on the DNA / RNA target, the population of DNA / RNA molecules produced in step c) is collected and then subjected to sequencing. In some examples, the collection of DNA / RNA molecules requires the destruction of the compartment. Thus, in one example of the method disclosed herein, step d) further comprises the destruction of the compartment by physical or chemical methods. In the example where the compartment is an emulsion droplet, the collection of DNA / RNA molecules comprises the destruction of the emulsion droplet.
[0055] Methods for breaking emulsion droplets are known to those skilled in the art. One non-limiting example of the method is as follows: transfer the emulsion mixture into a 2 mL centrifuge tube and centrifuge at 13000 g for 5 minutes at room temperature. Discard the upper oil layer. Add 1 mL of water-saturated diethyl ether to the remaining aqueous layer, vortex, and remove the upper solvent layer. Repeat this step once. Centrifuge the remaining aqueous layer in vacuum at room temperature for 5 minutes. In one example, the step of recovering DNA / RNA molecules also includes a step of quenching the IVTT. The step of quenching the IVTT can be carried out, for example, by treating the remaining aqueous layer with RNase cocktail and proteinase K to remove excess RNA and protein from the IVTT reaction. In some examples, the step of recovering DNA / RNA molecules also includes a clean-up step to purify the DNA / RNA molecules. Methods for cleaning up DNA / RNA are well known to those of skill in the art and there are many commercially available kits for this process, such as DNA Clean & Concentrator-5 (Zymo Research) or SPRIselect bead cleanup (Beckman Coulter).
[0056] In some examples, recovering the DNA / RNA molecules requires purifying the recovered DNA / RNA molecules to remove excess or unwanted DNA, RNA, and / or proteins from the reaction. Thus, in one example of the method disclosed herein, step d) further comprises purifying the recovered DNA / RNA molecules to remove excess DNA, RNA, and / or proteins from the reaction. In some examples, the excess DNA, RNA, and / or proteins may include, but are not limited to, gRNA, nucleic acid modifying enzymes, and IVTT reagents. In some examples, the term "excess" describes molecules that are subjected to sequencing.
[0057] Sequencing In a preferred example, the sequencing is single molecule sequencing. "Single molecule sequencing" refers to a technique that can directly read the base sequence from individual DNA or RNA strands present in a sample. At least two types of single molecule sequencing are commercially available: (a) Pacific Biosciences' real-time single molecule sequencing (SMRT), based on the detection and identification of fluorophore-labeled nucleotides in a subwavelength waveguide (ZMW), and (b) label-free sequencing used by Oxford Nanopore Technologies, which uses electronic means to read the signal when a nucleic acid (DNA / RNA) fragment is passed through a nanopore. Single molecule sequencing is facilitated by the long read length, and therefore may also be referred to as "long-read sequencing" or "single molecule long-read sequencing". The use of single molecule sequencing provides direct identification of variant sequences, thanks to which (i) the attachment of oligonucleotides to predetermined DNA / RNA ends and (ii) PCR amplification are omitted.
[0058] It is an important feature of the present invention to "directly" detect the enzyme products of individual variants and to quantify the molecular activity of individual variants by molecular counting of modified:unmodified DNA / RNA targets (or polynucleotide constructs / RNA transcripts). The term "directly" can refer to the direct detection of reaction products or the direct measurement of the enzyme activity of individual variants. In the latter sense, the expression "direct measurement of enzyme function" is in the context of directly calculating the phenotypic activity of a variant molecule (in a large-scale variant survey) with the associated genotypic information (also encoded in a molecule (either polynucleotide construct or RNA transcript)). Thus, the enzyme activity is measured directly on the molecule that is currently interacting. Based on the methods disclosed herein, the exact level of enzyme activity can be measured directly based on the modified count relative to the unmodified (or total) count. In one example, a particular variant associated with a 1:1 modified:unmodified target site is determined to be active at that target site 50% of the time.
[0059] Thus, in some instances, the population of recovered DNA / RNA molecules is not subjected to further modifications other than those required for single molecule sequencing before being subjected to a single molecule sequencing reaction. These modifications may include those required for conventional sequencing, such as adapter ligation of cut ends, barcode addition, PCR amplification of DNA / RNA molecules, etc.
[0060] In one example, sequencing is performed using the Oxford Nanopore Technologies platform. A non-limiting example of the sequencing process is described below.
[0061] Purified DNA is prepared for long-read sequencing according to the library preparation protocol recommended by the sequencing device manufacturer, for example, Oxford Nanopore Technologies (ONT) for the MinION Mk1B device, and the library is sequenced accordingly. In some examples, this may include multiplexing barcoded DNA sub-libraries using the ONT SQK-LSK109 ligation sequencing kit for general DNA library preparation together with the ONT EXP-NBD104 PCR-free native barcoding extension kit.
[0062] Long-read sequencing data can be collated using publicly available bioinformatics tools from repositories, such as minimap2 (Li, H. (2018). Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics, 34:3094-3100. doi:10.1093 / bioinformatics / bty191), NanoPack (De Coster, W. et al., (2018). NanoPack: visualizing and processing long-read sequencing data. Bioinformatics, 34:2666-2699. doi: 10.1093 / bioinformatics / bty149), samtools (Li, H. et al., (2009). The Sequence Alignment / Map format and SAMtools. Bioinformatics, 25:2078-9. doi: 10.1093 / bioinformatics / btp352), VarScan 2 (Koboldt, DC et al., (2012). VarScan 2: Somatic mutation and copy number alteration discovery in cancer by exome sequencing. Genome Research, 22: 568-576. doi: 10.1101 / gr.129684.111), or a custom script that can be generated by one skilled in the art of sequencing data analysis. For example, in some examples, one skilled in the art of sequencing analysis processes and analyzes raw nanopore sequencing reads generated from an ONT sequencing device by the following steps: 1. Process raw nanopore sequencing reads using the guppy toolkit provided by ONT (https: / / community.nanoporetech.com / protocols / Guppy-protocol / v / gpb_2003_v1_revm_14dec2018), using the base extraction and multiplexing algorithms of the toolkit as needed. Those skilled in the art of sequencing analysis may wish to adjust certain parameters, such as the filtering threshold of the multiplexed barcode quality score, as needed. These parameters are usually described in the manual of each tool. 2. Those skilled in the art of sequencing analysis may wish to further filter and process the reads using software tools such as NanoPack based on parameters such as read length and read quality score. 3. These processed reads can then be aligned to a (set of) reference sequences using minimap2 or other sequence alignment tools to generate a data set of read alignments. Similarly, those skilled in the art of sequencing analysis may wish to adjust read alignment parameters, such as alignment scoring matrices, as needed. These parameters are typically described in the manuals of the respective tools. 4. The user can then parse the generated read alignment file to calculate the number of unmodified and modified reads. In some embodiments, this can be done by using other alignment processing tools such as samtools or VarScan2 to detect and identify sequencing variations between the aligned sequencing reads and the reference sequence to which the sequencing reads are aligned. Similarly, a person skilled in the art of sequencing analysis can determine which parameters of these tools should be adjusted as necessary, for example, to set a minimum read count threshold to detect and identify true sequencing variations from background level sequencing errors. These parameters are usually described in the manual of each tool.
[0063] Molecular detection and counting Thus, enzymatic activity can be detected directly by detecting and counting polynucleotide constructs and / or RNA transcripts that have been modified by the nucleic acid modifying enzyme and polynucleotide constructs and / or RNA transcripts that have not been modified by the nucleic acid modifying enzyme. The polynucleotide constructs can include DNA targets and the RNA transcripts can include RNA targets.
[0064] Thus, in some examples, the method includes determining the number of polynucleotide constructs and / or RNA transcripts modified by a nucleic acid modifying enzyme (Σ count 修飾 ) and multiplied it by the number of polynucleotide constructs and / or RNA transcripts that were not modified by the nucleic acid modifying enzyme (ΣCount 無修飾 ) or the total number of polynucleotide constructs and / or RNA transcripts (ΣCount 修飾 + 無修飾 The method further comprises evaluating the modification activity of the one or more nucleic acid modifying enzymes on one or more of the DNA / RNA targets by comparing the modified activity of the one or more nucleic acid modifying enzymes with the modified activity of the one or more DNA / RNA targets.
[0065] In one example, the enzyme activity is determined according to the following formula: It is represented by a value calculated using one of the following: TIFF0007680440000002.tif17128.
[0066] Polynucleotide constructs and / or RNA transcripts or fragments thereof modified or unmodified by nucleic acid modifying enzymes can be detected and counted using sequencing data generated by sequencing platforms available to those skilled in the art.Since DNA / RNA molecules are directly sequenced by single molecule sequencing, in one example, detection and counting of DNA / RNA molecules modified or unmodified by nucleic acid modifying enzymes is based only on the data generated during single molecule sequencing and does not require further modification or processing of DNA / RNA molecules.
[0067] In one example, the modification activity is a cleavage activity, and the detection and calculation of modified or unmodified polynucleotide constructs or RNA targets is performed by aligning the sequencing reads of the DNA / RNA molecules to a reference sequence that includes a window of cleavage sites for the nucleic acid modifying enzyme; i) if the 3' end of the DNA / RNA molecule maps to the 3' downstream region of the cleavage site window, then the DNA / RNA molecule is an unmodified polynucleotide construct or RNA target; ii) if the 3' end of the DNA / RNA molecule maps to a region within the window of the cleavage site, then the DNA / RNA molecule is a modified polynucleotide construct or an RNA target; iii) If the 3' end of a DNA / RNA molecule is mapped to a region 5' upstream of the cleavage site window, then that DNA / RNA molecule is non-informative and is not used to measure modification activity.
[0068] In one example, sequencing reads are determined by their mapped endpoints (i.e., where the sequencing read ends) and whether they fall within a small window of predicted cleavage sites (grey triangles and dotted lines on the "Cas reference seq"; FIG. 3). One non-limiting example is described below, where the DNA target site is 3' of the encoded Cas nuclease variant. According to this metric, read alignments whose 3' ends are mapped to sites 3' downstream of the window of predicted Cas cleavage sites on the reference sequence are considered to be uncleaved (dark grey; FIG. 3), read alignments whose 3' ends fall within the window of Cas cleavage sites are considered to be cleaved (light grey; FIG. 3), and finally, reads that do not meet either criterion are discarded as uninformative, since it is not possible to experimentally determine whether they are cleaved or uncleaved (white; FIG. 3). In these examples, each cleaved Cas cleavage site represents one polynucleotide construct / RNA transcript that was modified. Similarly, each uncleaved Cas cleavage site represents one polynucleotide construct / RNA transcript that was not modified.
[0069] In some instances, the sequencing technology can detect or detect the chemical and sequence identity of a target site to determine whether the target has been modified by a Cas variant. For example, chemical modifications to nucleotides, such as methylation, can be detected using publicly available bioinformatics tools designed to extract chemically modified nucleotides in nanopore sequencing reads (Liu, Q., et al. (2019). Detection of DNA base modifications by deep recurrent neural network on Oxford Nanopore sequencing data. Nat Commun 10(1): 2449, doi : 10.1038 / s41467-019-10168-2; Liu, Q., et al. (2019). NanoMod: a computational tool to detect DNA modifications using Nanopore long-read sequencing data. BMC Genomics 20(Suppl 1): 78, doi: 10.1186 / s12864-018-5372-8; Rand, AC, et al. (2017). Mapping DNA methylation with high-throughput nanopore sequencing. Nat Methods 14(4): 411-413, doi: 10.1038 / nmeth.4189; Simpson, JT, et al. (2017). Detecting DNA cytosine methylation using nanopore sequencing. Nat Methods 14(4): 407-410, doi: 10.1038 / nmeth.4184).In other examples where the encoded Cas nuclease targets an RNA construct, the RNA molecules can be recovered and purified after the IVTT reaction using commercially available kits, such as RNA Clean & Concentrator-5 (Zymo Research), while Oxford Nanopore Technologies offers direct nanopore sequencing of the recovered RNA molecules with their SQK-RNA002 Direct RNA Sequencing Kit. Thus, this sequencing technology can be used to detect various types of modifications, including but not limited to strand breaks, sequence alterations, and epigenetic biochemical marks.
[0070] Polynucleotide constructs and libraries The present disclosure also relates to various polynucleotide constructs, construct libraries, and compartments (see, eg, FIG. 2 for some examples of polynucleotide constructs).
[0071] Construct 1: In one aspect, the present disclosure relates to a polynucleotide construct comprising a first polynucleotide sequence encoding a nucleic acid modifying enzyme or a variant thereof operably linked to a first promoter, and a second polynucleotide sequence comprising a DNA target.
[0072] Construct 2: In another aspect, the present disclosure relates to a polynucleotide construct comprising a first polynucleotide sequence encoding a nucleic acid modifying enzyme or a variant thereof operably linked to a first promoter, and a second polynucleotide sequence comprising a DNA template encoding an RNA target, wherein the RNA target is co-expressed contiguously with the nucleic acid modifying enzyme as one RNA transcript driven by the first promoter.
[0073] In another aspect, the present disclosure relates to a construct library comprising a plurality of polynucleotide constructs disclosed herein as construct 1 or construct 2, the library being characterized by one or more of the following: a. the plurality of polynucleotide constructs encode different variants of a nucleic acid modifying enzyme; b. The multiple polynucleotide constructs encode different DNA or RNA targets.
[0074] Construct 3: In one example, the present disclosure also relates to a polynucleotide construct of construct 1 or construct 2, wherein the polynucleotide construct further comprises a third polynucleotide sequence encoding a guide RNA (gRNA). Since the encoded gRNA will be expressed from the polynucleotide construct (e.g., in a compartment), the polynucleotide construct may comprise other elements that facilitate the expression of the gRNA, which are generally known to those skilled in the art. In some examples, the third polynucleotide sequence is operably linked to a second promoter. In some examples, the second promoter is a T7 promoter.
[0075] In another aspect, the present disclosure relates to a construct library comprising a plurality of the polynucleotide constructs disclosed herein as construct 3, the library being characterized by one or more of the following: a. the plurality of polynucleotide constructs encode different variants of a nucleic acid modifying enzyme; b. the multiple polynucleotide constructs encode different DNA or RNA targets; c. The multiple polynucleotide constructs encode different gRNAs.
[0076] Construct 4: In yet another aspect, the present disclosure relates to a polynucleotide construct comprising a first polynucleotide sequence encoding a guide RNA (gRNA) operably linked to a first promoter, and a second polynucleotide sequence comprising a DNA target.
[0077] Construct 5: In yet another aspect, the present disclosure relates to a polynucleotide construct comprising a first polynucleotide sequence encoding a guide RNA (gRNA) operably linked to a first promoter, and a second polynucleotide sequence comprising a DNA template encoding an RNA target, wherein expression of the RNA target is driven by the first promoter and co-expressed contiguous with the gRNA as one RNA transcript.
[0078] In another aspect, the present disclosure relates to a construct library comprising a plurality of polynucleotide constructs disclosed herein as construct 4, the library being characterized by one or more of the following: a. the multiple polynucleotide constructs encode different DNA or RNA targets; b. The multiple polynucleotide constructs encode different gRNAs.
[0079] In some examples of the methods or polynucleotide constructs disclosed herein, the first and second polynucleotide sequences overlap, either completely or partially. For example, a DNA / RNA target ("second polynucleotide") may be encoded within the coding sequence of a nucleic acid modifying enzyme ("first polynucleotide").
[0080] In some examples of the methods or polynucleotide constructs disclosed herein, the DNA target or RNA target comprises a protospacer that is at least partially complementary to the guide RNA.In some examples of the methods or polynucleotide constructs disclosed herein, the DNA target also comprises a proximal protospacer adjacent motif (PAM) sequence.In some examples of the methods or polynucleotide constructs disclosed herein, when the polynucleotide construct comprises a DNA template that encodes an RNA target, the RNA target further comprises a proximal protospacer adjacent sequence (PFS).
[0081] In some examples of the methods or polynucleotide constructs disclosed herein, the RNA-guided nucleic acid modifying enzyme is a CRISPR-associated (Cas) protein.In some examples, the RNA-guided nucleic acid modifying enzyme is selected from the group consisting of Cas3, Cas9, Cas10, Cas12a (also known as Cpf1), Cas13a (also known as C2c2), Cas13b, Cas13c, Cas13d, Cas14, CasX, CasΦ, and variants thereof.
[0082] In some examples of the methods or polynucleotide constructs disclosed herein, the variant nucleic acid modifying enzyme contains one or more inactivated catalytic sites and is capable of binding to and inhibiting expression of a DNA target without modifying the DNA target.
[0083] In some examples of the methods or polynucleotide constructs disclosed herein, the variant nucleic acid modifying enzyme is fused with one or more additional functional domains that can modify DNA or RNA. In some specific examples, the additional functional domains include, but are not limited to, cytidine deaminase domain, de novo DNA methyltransferase 3A (DNMT3A) domain, cytosine-5 methyltransferase domain, ten-eleven translocation dioxygenase 1 (TET1) catalytic domain, adenosine deaminase acting on RNA (ADAR2) deaminase domain, and DNA deoxyadenosine deaminase domain, etc.
[0084] For purposes of explanation and illustration, the sequences of the polynucleotide constructs are provided below. TIFF0007680440000003.tif98160TIFF0007680440000004.tif232160TIFF0007680440000005.tif231160TIFF0007680440000006.tif165159Sequences in bold and underlined format refer to elements that are separately annotated below. T7 / lacO promoter: TIFF0007680440000007.tif4146RBS(Ribosome binding site):AAGGAG(SEQ ID NO: 2) SpCas9 gene (coding sequence): TIFF0007680440000008.tif10159TIFF0007680440000009.tif231159TIFF0007680440000010.tif225159Synthetic terminator sequence (L3S1P52): TIFF0007680440000011.tif5167T7 promoter: TIFF0007680440000012.tif5128gRNA target sequence (protospacer): TIFF0007680440000013.tif4128SpCas9 gRNA scaffold: TIFF0007680440000014.tif11158T7 Terminator: TIFF0007680440000015.tif11129 Target region (DNA target): TIFF0007680440000016.tif11128
[0085] In some instances, such as those exemplified above, a protospacer adjacent motif (PAM) is found adjacent to the target sequence (known as a protospacer), e.g., the 5' PAM site TTTV (SEQ ID NO: 11) for Cpfl-type (also known as Cas12) Cas proteins, the 3' PAM site NGG for Sp Cas9, and the 3' PAM site NNGRRT (SEQ ID NO: 12) for Sa Cas9 protein, adjacent to the protospacer sequence. Adjacent to TIFF0007680440000017.tif4128. Standard IUPAC nucleic acid notation is used herein and throughout the specification.
[0086] compartment In one aspect, the present disclosure relates to one or more compartments, each of which comprises a polynucleotide construct disclosed herein, and the compartments are isolated from each other. In some examples, each compartment further comprises an in vitro transcription and translation (IVTT) reagent, which allows for in vitro transcription and / or translation of protein and / or RNA. In some examples, the compartment has a volume of 1000 μm 3 , 100 μm 3 , 10 μm 3 , or 1 μm 3 In some instances, the compartments are water-in-oil emulsion droplets. In some instances, the isolation is achieved using microfluidics, hydrogel-limited diffusion, or partitioned wells. EXAMPLES
[0087] Example 1: IVTT and cleavage of Sp Cas9 constructs in emulsion droplets Water-in-oil (w / o) emulsion droplets were made following the steps outlined in the protocol above. Briefly, 950 μL of oil surfactant mixture (mineral oil + 4.5% (v / v) Span 80 + 0.5% (v / v) Tween 80) was added to a cryovial equipped with a 3 x 8 mm magnetic stir bar and placed on ice.
[0088] High DNA input: >1 sequence copy encapsulated per emulsion droplet - for Cas expressed from bulk IVTT reactions, demonstrating that emulsification maintains Cas activity.
[0089] For this experiment, approximately 750 ng of Sp Cas9 construct (sequence as above) was mixed with IVTT reagent (New England Biolabs PURExpress #E6800) in ice to make 75 μL of IVTT aqueous mixture. 50 μL of the aqueous mixture was added in five 10 μL portions to the oil and surfactant mixture in ice with a stir bar rotating at 1150 rpm for 2 minutes to create an emulsion mixture. The emulsion mixture was subsequently mixed in ice for an additional minute. The emulsion mixture was then subjected to homogenization (8000 rpm for 3 minutes; IKA Ultraturrax T10 homogenizer) to result in a more monodisperse distribution of emulsion droplet sizes. The remaining 25 μL of the aqueous mixture was kept in ice for the bulk IVTT reaction as a control. This was repeated with the Sp dCas9 construct.
[0090] The emulsion and bulk IVTT mixtures were then incubated for 4 hours at 37°C to allow IVTT to proceed, followed by protein inactivation at 65°C for 15 minutes.
[0091] The emulsion IVTT mixture was then treated as above to break the emulsion. 20 mM EDTA (pH 8.0) inhibitor was added to the emulsion and mixed briefly by vortexing. The emulsion mixture was then centrifuged at 13000 g for 5 min at room temperature. The top oil layer was removed. 1 mL of water-saturated diethyl ether was added to the remaining aqueous layer, vortexed, and the top solvent layer was removed. This process was repeated once. The remaining aqueous layer was centrifuged in vacuum at room temperature for 5 min and then treated with RNase cocktail and proteinase K to remove excess RNA and protein from the IVTT reaction for 30 min at 37°C. The bulk IVTT reaction was also treated with 20 mM EDTA (pH 8.0) and a mixture of RNase cocktail and proteinase K to remove excess RNA and protein from the IVTT reaction for 30 min at 37°C. DNA from all IVTT reactions was then individually purified with SPRIselect paramagnetic beads, and aliquots were size-separated by gel electrophoresis and then visualized on an agarose gel (Figure 4). In emulsion IVTT reactions, constructs encoding active Sp Cas9 were cleaved (presence of a smaller band), whereas constructs encoding inactive Sp dCas9 were not cleaved (absence of a smaller band). This demonstrates that the CRISPR-Cas IVTT self-cleavage assay works whether the encoding DNA construct is compartmentalized in emulsion droplets or floating freely in bulk solution.
[0092] The purified DNA from the emulsion IVTT reaction was also processed for nanopore sequencing using the SQK-LSK109 ligation sequencing kit from Oxford Nanopore Technologies (ONT) and barcoded using the ONT EXP-NBD104 PCR-free native barcoding extension kit so that they could be optionally multiplexed into one pooled DNA library. Single-molecule long-read nanopore sequencing was then performed on this pooled DNA library using an ONT MinION Mk1B sequencing device. The nanopore sequencing results were then filtered to enhance quality and analyzed using publicly available bioinformatics tools. SpCas9 emulsion IVTT nanopore sequencing reads show that a mixture of cleaved and uncleaved construct fragments was detected (Figure 5). Sp dCas9 emulsion IVTT nanopore sequencing reads appeared overwhelmingly as uncleaved construct fragments, as expected (Figure 6), with the few reads classified as "cleaved" Sp dCas9 construct fragments likely the result of truncated / incomplete reads during nanopore sequencing, and / or the result of random DNA fragmentation events, and / or the result of errors in the sequencing device. Some reads in each sublibrary mapped to missequences, e.g., in the Sp Cas9-only sublibrary, reads mapped to Sp dCas9 rather than Sp Cas9. These were likely the result of random sequencing errors in the sequencing device, or misassignment of barcoded nanopore sequencing reads to their respective sublibraries during demultiplexing, and were therefore classified as misassigned and shown as such in the plots.
[0093] Example 2: Quantification of CRISPR-Cas cleavage activity by multiplexed single-molecule long-read sequencing of DNA constructs following bulk IVTT reactions Bulk IVTT reactions were set up in ice for different CRISPR-Cas constructs (Sp Cas9, Sa Cas9, As Cpf1, Lb Cpf1), all of which shared a similar arrangement of components as described in the nucleic acid template sequence above. These were then divided equally into five corresponding aliquots for each time point (Figure 7 part 1). These bulk IVTT aliquots were then incubated at 37°C and removed at the indicated time points to quench with EDTA inhibitor and enzyme to stop the IVTT reaction and Cas cleavage of the coding DNA construct (Figure 7 part 2). The quenched IVTT reactions were then processed with SPRIselect bead cleanup to purify DNA fragments (Figure 7 part 3).
[0094] Small aliquots of DNA fragments of different Cas orthologs at these different IVTT time points were then visualized on agarose gels after size separation by gel electrophoresis, as shown in Figure 8.
[0095] Aliquots of the remaining purified DNA fragments were then pooled together for each time point, but regardless of the Cas species at each time point, i.e., Sp Cas9, Sa Cas9, or other DNA fragments, and barcoded individually using the ONT EXP-NBD104 PCR-free native barcoding extension kit (Figure 7, part 4), whereby these pooled sub-libraries were multiplexed for one run of nanopore sequencing (Figure 7, part 5). The nanopore sequencing results were then filtered to enhance quality and analyzed using publicly available bioinformatics tools, followed by the analytical approach disclosed in this invention.
[0096] Figure 9 shows the number of cleaved DNA fragments encoding each active Cas construct normalized to the total number of cleaved and uncleaved DNA fragments encoding each Cas construct over five selected time points (0-4 hours) of IVTT incubation. With increasing duration of IVTT incubation, the expressed Cas protein has more time to cleave more coding DNA constructs, resulting in a higher incidence of cleaved fragments for each species at later time points. The results of the nanopore sequencing analysis plotted in Figure 9 show qualitative agreement with the gel image of purified IVTT DNA fragments in Figure 8, with both assays sharing the same purified DNA input obtained from the workflow steps shown in Figure 7 part 3. This example substantiates our claim that our workflow interrogates the nucleic acid products from individual IVTT reactions of multiple CRISPR-Cas self-cleavage assays.
[0097] Example 3: Demonstration of sensitivity of nanopore sequencing assays by titration ratio of purified CRISPR-Cas DNA end products from bulk IVTT reactions For this experiment, 500 ng of Sp Cas9 (sequence as above) was added to IVTT reagent (New England Biolabs PURExpress #E6800) on ice to make 50 μL of IVTT aqueous mixture. The same was done for the Sp dCas9 construct. The Sp dCas9 construct essentially contains the same DNA sequence as the Sp Cas9 construct, but differs in that the Sp dCas9 gene has two inactivating mutations (D10A and H840A) in the Sp Cas9 gene. These 50 μL bulk IVTT reactions were incubated at 37°C for 4 hours to allow IVTT to proceed, followed by 15 minutes at 65°C to inactivate the protein. 20 mM EDTA (pH 8.0) inhibitor was added to the bulk IVTT reaction along with RNase cocktail and proteinase K to remove excess RNA and protein from the IVTT reaction at 37°C for 30 minutes. DNA from both bulk IVTT reactions was then purified separately with SPRIselect paramagnetic beads and aliquots were visualized on agarose gels after size separation by gel electrophoresis as shown in FIG.
[0098] Quantify the concentration of purified DNA from bulk IVTT reactions of Sp dCas9 and Sp Cas9, and use a mass ratio of 1:1, 1:10, or 1:20. -1 , 1:10 -2 , 1:10 -3 , 1:10 -4 , 1:10 -5, and mixed 1:0. Seven mixtures with these titration ratios of purified Sp dCas9 bulk IVTT DNA product to purified Sp Cas9 bulk IVTT DNA product were then processed for nanopore sequencing using the ONT SQK-LSK109 ligation sequencing kit, and each of the seven mixtures was individually barcoded using the ONT EXP-NBD104 PCR-free native barcoding extension kit. The DNA library was then subjected to single molecule long-read nanopore sequencing using an ONT MinION Mk1B sequencing device. The nanopore sequencing results were then filtered to enhance quality and analyzed using publicly available bioinformatics tools, followed by the analytical approach disclosed in this invention.
[0099] The purpose of this assay was to assess the sensitivity of nanopore sequencing assays used for the large-scale investigation of DNA / RNA modification events, a capability claimed in our invention. Specifically, in this example, the self-cleavage events of Sp Cas9 IVTT constructs were titrated against uncleaved Sp dCas9 IVTT constructs. Using a combination of the above bioinformatics approaches, we demonstrated the detection of cleaved and uncleaved Sp Cas9 DNA fragments in raw nanopore sequencing data that can be distinguished from the detection of uncleaved Sp dCas9 DNA fragments. In particular, we demonstrated that the 1:10 ratio of purified Sp dCas9 bulk IVTT DNA products to purified Sp Cas9 bulk IVTT DNA products was significantly higher than that of uncleaved Sp dCas9 DNA fragments. -5 The cleaved Sp Cas9 DNA fragments could also be detected in this mixture (Figure 11).
[0100] Example 4: IVTT and cleavage of Sp Cas9 constructs in emulsion droplets Limit DNA input: encapsulate ≦1 copy of sequence per emulsion droplet – measure emulsification efficiency of single copy DNA constructs.
[0101] For this experiment, ≦1.66 fmol of Sp Cas9 construct (sequence as above) was added to IVTT reagent (New England Biolabs PURExpress #E6800) in ice to make 50 μL of IVTT aqueous mixture. This 50 μL aqueous mixture was added in five 10 μL portions to the oil and surfactant mixture in ice with a stir bar rotating at 1150 rpm over a period of 2 minutes to create an emulsion mixture. This emulsion mixture was subsequently mixed in ice for an additional minute. The emulsion mixture was then subjected to homogenization (8000 rpm for 3 minutes; IKA Ultraturrax T10 homogenizer) to result in a more monodisperse distribution of emulsion droplet sizes. This was repeated with the Sp dCas9 construct and with a 1:1 equimolar mixture of Sp Cas9 and Sp dCas9 constructs.
[0102] It should be noted that the use of a mixture of Sp Cas9 and Sp dCas9 DNA constructs measures the efficiency of encapsulating ≦1 DNA construct per emulsion droplet. With perfect efficiency of encapsulating ≦1 DNA construct per droplet, none of the Sp dCas9 sequences detected by nanopore sequencing at the end of the assay for the mixed DNA input condition should be cleaved. With non-perfect efficiency, some Sp dCas9 DNA constructs may be cleaved, as some may be exposed to active Sp Cas9 in the same droplet. Thus, if the detection rate of cleaved Sp dCas9 constructs in the assay of mixed Sp Cas9 and Sp dCas9 constructs is very low, comparable to the expected random sequencing error rate for long-read nanopore sequencing, the data indicates that ≦1 sequence copy was encapsulated in each emulsion droplet under these conditions. This example also demonstrates the entire workflow of the present invention as shown in FIG. 1.
[0103] The resulting emulsion IVTT mixture was then incubated for 4 hours at 37°C to allow IVTT to proceed, followed by protein inactivation at 65°C for 15 minutes.
[0104] The emulsion IVTT mixture was then treated as above to break the emulsion. 20 mM EDTA (pH 8.0) inhibitor was added to the emulsion and mixed briefly by vortexing. The emulsion mixture was then centrifuged at 13000 g for 5 min at room temperature. The upper oil layer was removed. 1 mL of water-saturated diethyl ether was added to the remaining aqueous layer, vortexed, and the upper solvent layer was removed. This process was repeated once. The remaining aqueous layer was centrifuged in vacuum at room temperature for 5 min and then treated with RNase cocktail and proteinase K to remove excess RNA and protein from the IVTT reactions for 30 min at 37°C. DNA from all IVTT reactions was then purified separately with a commercially available column purification kit (DNA Clean and Concentrator-5, Zymo Research) according to the manufacturer's instructions.
[0105] The purified DNA from the IVTT reactions was then processed for nanopore sequencing using the ONT SQK-LSK109 Ligation Sequencing Kit and individually barcoded using the ONT EXP-NBD104 PCR-free Native Barcoding Extension Kit. The DNA library was then subjected to single-molecule long-read nanopore sequencing using an ONT MinION Mk1B sequencing device. The nanopore sequencing results were then filtered to enhance quality and analyzed using publicly available bioinformatics tools, followed by the analytical approach disclosed in this invention.
[0106] Sp Cas9 emulsion IVTT nanopore sequencing reads showed that a mixture of cleaved and uncleaved construct fragments were detected (Figure 12), demonstrating that Sp Cas9 is active against a portion of the target (as demonstrated in the bulk reaction). Sp dCas9 emulsion IVTT nanopore sequencing reads, as expected, appeared overwhelmingly as uncleaved construct fragments, demonstrating that Sp dCas9 is mostly inactive (Figure 13), with the small number of reads classified as "cleaved" Sp dCas9 construct fragments likely the result of truncated / incomplete reads during nanopore sequencing and / or random DNA fragmentation events.
[0107] Note that some reads in the Sp Cas9-only and Sp dCas9-only sub-libraries mapped to incorrect sequences, e.g., in the Sp Cas9-only sub-library, reads mapped to Sp dCas9 instead of Sp Cas9, likely the result of random sequencing errors in the sequencing device or errors in demultiplexing of barcoded nanopore sequencing reads, and were therefore classified as misassigned and shown as such in the plots.
[0108] Nanopore sequencing reads generated from emulsion IVTT reactions with a 1:1 mixture of Sp Cas9 and Sp dCas9 constructs added at limiting concentrations show roughly equal distribution of Sp Cas9 and Sp dCas9 mapped reads as expected (Figure 14). The Sp Cas9 mapped reads are roughly evenly split into cleaved and uncleaved fragments, but the majority of the Sp dCas9 mapped reads are classified as uncleaved. A small number of Sp dCas9 mapped reads are classified as cleaved, which may be partially due to errors in sequencing or multiplexing, because these sequencing errors are known to occur on sequencing devices when sequencing a mixture of fragments or due to cross-contamination errors of enzyme complexes, but can be further reduced by technical optimization within the concept of the present invention. Taken together, this example embodies and demonstrates the disclosed invention, where the enzyme activity levels of quantified variants can be directly counted and determined on a single molecule basis.
Claims
1. (a) isolating a plurality of polynucleotide constructs into compartments, each compartment containing one polynucleotide construct, each polynucleotide construct comprising: (i) a first polynucleotide sequence encoding a nucleic acid modifying enzyme or a variant thereof, operably linked to a first promoter; and (ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, wherein if the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target is co-expressed contiguously with the nucleic acid modifying enzyme as a single RNA transcript driven by the first promoter; wherein the plurality of polynucleotide constructs encode different variants of the nucleic acid modifying enzyme and / or different DNA or RNA targets; (b) subjecting the compartment to conditions that allow for in vitro expression of RNA and protein; (c) subjecting said plurality of compartments to conditions that allow modification of said DNA / RNA targets by a nucleic acid modifying enzyme having modification activity against a DNA target or an RNA target, i. a polynucleotide construct and / or an RNA transcript or fragment thereof modified by said nucleic acid modifying enzyme; ii. polynucleotide constructs and / or RNA transcripts that have not been modified by said nucleic acid modifying enzymes Producing a population of DNA / RNA molecules comprising one or more of: (d) recovering the population of DNA / RNA molecules produced in step (c) and subjecting it to single molecule sequencing; (e) detecting and counting the DNA / RNA molecules produced in step (c) based on the sequencing results; A method comprising:
2. 2. The method of claim 1, wherein the nucleic acid modifying enzyme is a RNA-guided nucleic acid modifying enzyme and each compartment further comprises a guide RNA or a nucleotide template encoding same.
3. 2. The method of claim 1, wherein the nucleic acid modifying enzyme is a RNA-guided nucleic acid modifying enzyme, and each polynucleotide further comprises a third polynucleotide sequence encoding a variant guide RNA (gRNA).
4. (a) isolating a plurality of polynucleotide constructs into compartments, each compartment containing one polynucleotide construct, each polynucleotide construct comprising: (i) a first polynucleotide sequence encoding a guide RNA (gRNA), operably linked to a first promoter; (ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, wherein if the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target is driven by the first promoter and co-expressed contiguously with the gRNA as a single RNA transcript; wherein the plurality of polynucleotide constructs encode different gRNAs and / or different DNA or RNA targets; and each compartment further comprises an RNA-guided nucleic acid modifying enzyme or a variant thereof, or a nucleotide template encoding same; (b) subjecting the compartment to conditions that allow in vitro transcription and / or translation of RNA and protein; (c) subjecting said compartment to conditions that allow modification of said DNA and / or RNA targets by an RNA-guided nucleic acid modifying enzyme having functional activity against a DNA or RNA target in the presence of a gRNA, i. a polynucleotide construct and / or an RNA transcript or fragment thereof modified by said nucleic acid modifying enzyme; ii. polynucleotide constructs and / or RNA transcripts that have not been modified by said nucleic acid modifying enzymes Producing a population of DNA / RNA molecules comprising one or more of: (d) recovering the population of DNA / RNA molecules produced in step (c) and subjecting it to single-molecule long-read sequencing; (e) detecting and counting the DNA / RNA molecules produced in step (c) based on the sequencing results; A method comprising:
5. Number of polynucleotide constructs and / or RNA transcripts modified by a nucleic acid-modifying enzyme (Σ count 修飾 ) and multiplied it by the number of polynucleotide constructs and / or RNA transcripts that were not modified by the nucleic acid modifying enzyme (ΣCount 無修飾 ) or the total number of polynucleotide constructs and / or RNA transcripts (ΣCount 修飾 + 無修飾 ) to compare with assessing the modification activity of one or more nucleic acid modifying enzymes on one or more of the DNA / RNA targets by The method of any one of claims 1 to 4, further comprising:
6. The enzyme activity is represented by the following formula:
6. The method of claim 5, wherein the value is represented by a value calculated using any one of the following:
7. The method of any one of claims 1 to 6, wherein step (d) further comprises disrupting the compartments by physical or chemical methods.
8. The method of any one of claims 1 to 7, wherein step (d) further comprises purifying the recovered DNA / RNA molecules to remove excess DNA, RNA, and / or protein from the reaction.
9. 9. The method of any one of claims 1 to 8, wherein the population of recovered DNA / RNA molecules is not subjected to any further modifications other than those required for single molecule sequencing before being subjected to a single molecule sequencing reaction.
10. A method described in any one of claims 1 to 9, wherein detection and counting of the DNA / RNA molecules produced in step (c) is based solely on data generated during single molecule sequencing and does not require further modification or processing of the DNA / RNA molecules.
11. the modification activity is a cleavage activity, and the detection and counting of the modified or unmodified polynucleotide construct or RNA transcript is performed by aligning the sequencing reads of the DNA / RNA molecule to a reference sequence that includes a target site of the nucleic acid modifying enzyme; (i) if the 3' end of the DNA / RNA molecule maps to a region 3' downstream of the target site, then the DNA / RNA molecule is an unmodified polynucleotide construct or an RNA target; (ii) if the 3' end of the DNA / RNA molecule maps to a region within the target site, then the DNA / RNA molecule is a modified polynucleotide construct or an RNA target; (iii) if the 3' end of a DNA / RNA molecule maps to a 5' upstream region of the target site, that DNA / RNA molecule is uninformative and is not used to measure modification activity; The method according to any one of claims 5 to 10.
12. 12. The method of any one of claims 1 to 11, wherein the first polynucleotide sequence and the second polynucleotide sequence overlap completely or partially.
13. The method of any one of claims 2 to 11, wherein the DNA or RNA target comprises a protospacer that is at least partially complementary to the guide RNA.
14. The method of any one of claims 2 to 13, wherein the DNA target also comprises a proximal protospacer adjacent motif (PAM) sequence.
15. 15. The method of any one of claims 2 to 11, or the method of claim 13 or 14, wherein when the polynucleotide construct comprises a DNA template encoding an RNA target, the RNA target further comprises a proximal protospacer adjacent sequence (PFS).
16. The method of any one of claims 1 to 11 or any one of claims 13 to 15, wherein the nucleic acid modifying enzyme is a CRISPR associated protein (Cas).
17. 17. The method of any one of claims 1 to 16, wherein the variant nucleic acid modifying enzyme comprises one or more inactivated catalytic sites and is capable of binding to and inhibiting expression of a DNA target without modifying the DNA target.
18. 18. The method of any one of claims 1 to 17, wherein said variant nucleic acid modifying enzyme is fused to one or more additional functional domains capable of modifying DNA or RNA.
19. 19. The method of any one of claims 1 to 18, wherein each compartment further comprises in vitro transcription and translation (IVTT) reagents, said IVTT reagents allowing in vitro transcription and / or translation of proteins and / or RNA.
20. The method of any one of claims 1 to 18, wherein the compartments are emulsion droplets.
21. The method of any one of claims 1 to 18, wherein said isolation is achieved using microfluidics, hydrogel-restricted diffusion, or partitioned wells.
Citation Information
Patent Citations
Method for fragmenting genomic DNA using cas9
US20170044592A1
Novel crispr RNA targeting enzymes and systems and uses thereof
US20190002889A1
Nanopore sequencing complexes
WO2017125565A1
Emulsion-based screening methods
WO2018118968A1