Assays to measure nucleic acid-modifying enzyme activity
By separating the polynucleotide constructs in the compartment for in vitro expression and modification, and performing single-molecular sequencing, the problem of inefficient screening of nucleic acid modified enzyme variants in the prior art is solved, high-throughput screening and identification of the activity of nucleic acid modified enzyme variants is achieved, and the efficiency of the engineering process is improved.
Patent Information
- Application Number
- CN202080078223.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-15
- Filing Date
- 2020-10-15
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2040-10-15
AI Technical Summary
The prior art is difficult to screen and identify the functions of millions to billions of nucleic acid-modifying enzyme variants at high throughput, and cannot accurately identify and modify their nucleic acid targets, resulting in waste of resources and inefficiency in engineering processes.
High-throughput screening and identification of the activity of nucleic acid-modified enzyme variants is achieved by separating multiple polynucleotide constructs into compartments, in vitro expression and modification, followed by single-molecular sequencing and detection.
It has achieved efficient screening and identification of the activities of nucleic acid modified enzyme variants, can accurately identify and modify nucleic acid targets, and improves the efficiency of the engineering process and resource utilization rate.
Smart Images

Figure CN114651067B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of priority from Singapore Provisional Application No. 10201909632P, filed on October 15, 2019, the contents of which are incorporated herein by reference in their entirety for all purposes. Technical Field
[0003] The present invention relates to the field of biotechnology, and in particular to the development of multiplex assays suitable for measuring enzyme activity. Background Art
[0004] Nucleic acid-modifying enzymes, such as zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and clustered regularly interspaced short palindromic repeats (CRISPR)-associated nucleases, have become extremely useful tools in biomedical research and the biotechnology industry. As therapeutic modalities, these nucleic acid-modifying enzymes can treat previously incurable genetic diseases by directly modifying DNA or RNA. To realize the enormous industrial and medical potential of nucleic acid-modifying enzymes, the limitations of naturally occurring components must be addressed. These limitations include targeting efficiency, targeting specificity, immunogenicity, and compatibility issues with delivery vehicles and function-conferring protein fusion moieties. To address these limitations, enzymes must be modified and their activity measured in a process known as protein engineering. Enzymes such as CRISPR-Cas can also be engineered to have enhanced functionality (e.g., more specific for their targets, more efficient targeting) and / or new functions (e.g., base editing, immune evasion, epigenetic modification) by altering the protein amino acid sequence or fusing / co-localizing function-conferring protein domains to Cas proteins or CRISPR complexes.
[0005] Conventional methods for engineering enzymes begin with (i) designing and creating a library of DNA variants encoding many different enzyme sequences (with amino acid changes compared to the naturally occurring wild type), (ii) expressing these variants in a compartment, such as in cells or in vitro; (iii) measuring enzyme activity or associated enzyme activity through downstream biochemical reactions or cell phenotypes, followed by "screening" (where no selection pressure is applied to separate active variants from inactive variants) or "selection" (where selection pressure is applied to separate active variants from inactive variants). Protein engineering, particularly programmable nucleases like CRISPR-Cas, is primarily performed through the latter "selection" approach. This approach favors active variants in a binary manner (cells survive when expressing the active form of the protein and die when expressing the inactive form of the protein; also known as positive selection), without providing information about the degree of protein activity (e.g., it does not distinguish between highly active proteins and half-active proteins), nor does it consider / provide information about inactive protein variants. Negative selection can also be performed, whereby only inactive variants are retained and identified while active variants are depleted, and no direct measurement is performed. In both cases, activity testing is associated with and performed by enrichment / depletion of library members. This "screening" approach is not scalable due to the increased resources required to maintain and measure active and inactive variants. Consequently, the engineering and assaying of nucleic acid modifying enzymes such as CRISPR-Cas proteins is limited both in the number of variants that can be tested and the number of possible mutations per variant. While CRISPR-Cas proteins can be engineered to better, faster, and more safely engineer multiple amino acid substitutions into proteins, current methods do not allow exploration of this functional space.
[0006] Therefore, there is a need for a high-throughput screening technology to detect and identify millions to billions of functional variants exceeding candidate enzyme libraries that can still accurately and effectively recognize, cleave or modify their nucleic acid targets. Such technology will enable the screening and engineering of novel nucleic acid-modifying enzymes, as well as the screening and optimization of other factors that affect enzyme activity, such as guide RNA and target sequences. Therefore, the object of the present invention is to provide an improved method for addressing the above needs. SUMMARY OF THE INVENTION
[0008] In one aspect, the present disclosure relates to a method comprising the steps of:
[0009] a) separating the plurality of polynucleotide constructs into compartments, wherein each compartment comprises a single polynucleotide construct, wherein each polynucleotide construct comprises
[0010] i) a first polynucleotide sequence encoding a nucleic acid modification enzyme or a variant thereof operably linked to a first promoter; and
[0011] ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, wherein when the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target and the nucleic acid modification enzyme are continuously co-expressed as a single RNA transcript driven by the first promoter;
[0012] and wherein the plurality of polynucleotide constructs encode different variants of the nucleic acid modification enzyme, and / or different DNA or RNA targets;
[0013] b) subjecting the compartment to conditions allowing in vitro expression of RNA and protein;
[0014] c) subjecting a plurality of said compartments to conditions that allow modification of said DNA / RNA target by a nucleic acid modifying enzyme having modification activity on said DNA or RNA target, thereby producing a population of DNA / RNA molecules comprising one or more of:
[0015] i. a polynucleotide construct and / or RNA transcript or a fragment thereof that has been modified by the nucleic acid modifying enzyme;
[0016] ii. a polynucleotide construct and / or RNA transcript that has not been modified by the nucleic acid modifying enzyme;
[0017] d) harvesting the DNA / RNA molecule population produced in step (c) and performing single molecule sequencing on it;
[0018] e) detecting and counting the DNA / RNA molecules mentioned in steps c)i and c)ii based on the sequencing results.
[0019] In another aspect, the present disclosure relates to a method comprising the steps of:
[0020] a) separating the plurality of polynucleotide constructs into compartments, wherein each compartment comprises a single polynucleotide construct, wherein each polynucleotide construct comprises:
[0021] i) a first polynucleotide sequence encoding a guide RNA (gRNA) operably linked to a first promoter;
[0022] ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, wherein when the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target and the gRNA are continuously co-expressed as a single RNA transcript driven by the first promoter;
[0023] wherein the plurality of polynucleotide constructs encode different gRNAs, and / or different DNA or RNA targets; and wherein each compartment further comprises an RNA-guided nucleic acid modification enzyme or a variant thereof or a nucleotide template encoding the same;
[0024] b) subjecting the compartment to conditions allowing in vitro transcription and / or translation of RNA and protein;
[0025] c) subjecting the compartment to conditions that allow modification of the DNA and / or RNA target by an RNA-guided nucleic acid modification enzyme functionally active against the DNA or RNA target in the presence of a gRNA, thereby producing a population of DNA / RNA molecules comprising one or more of:
[0026] i. a polynucleotide construct and / or RNA transcript or a fragment thereof that has been modified by the nucleic acid modifying enzyme;
[0027] ii. a polynucleotide construct and / or RNA transcript that has not been modified by the nucleic acid modifying enzyme;
[0028] d) harvesting the DNA / RNA molecule population produced in step (c), and performing single-molecule long-read sequencing on it;
[0029] e) detecting and counting the DNA / RNA molecules mentioned in step c)i and / or c)ii based on the sequencing results.
[0030] In another aspect, the present disclosure relates to a polynucleotide construct comprising: a first polynucleotide sequence encoding a nucleic acid modification enzyme or a variant thereof operably linked to a first promoter; and a second polynucleotide sequence comprising a DNA target.
[0031] In another aspect, the present disclosure relates to a polynucleotide construct comprising: a first polynucleotide sequence encoding a nucleic acid modification enzyme or a variant thereof operably linked to a first promoter; and a second polynucleotide sequence comprising a DNA template encoding an RNA target; and wherein the RNA target and the nucleic acid modification enzyme are continuously co-expressed as a single RNA transcript driven by the first promoter.
[0032] In yet another aspect, the present disclosure relates to a construct library comprising a plurality of polynucleotide constructs as disclosed herein, wherein the library is characterized in that one or more of the following: a) the plurality of polynucleotide constructs encode different variants of nucleic acid modification enzymes; b) the plurality of polynucleotide constructs encode different DNA or RNA targets.
[0033] In yet another aspect, the present disclosure relates to a construct library comprising a plurality of polynucleotide constructs as disclosed herein, wherein the library is characterized in that one or more of the following: a) the plurality of polynucleotide constructs encode different variants of nucleic acid modification enzymes; b) the plurality of polynucleotide constructs encode different DNA or RNA targets; c) the plurality of polynucleotide constructs encode different gRNAs.
[0034] In another aspect, the present disclosure relates to a polynucleotide construct comprising: a first polynucleotide sequence encoding a guide RNA (gRNA) operably linked to a first promoter; and a second polynucleotide sequence comprising a DNA target.
[0035] In another aspect, the present disclosure relates to a polynucleotide construct comprising: a first polynucleotide sequence encoding a guide RNA (gRNA) operably linked to a first promoter; and a second polynucleotide sequence comprising a DNA template encoding an RNA target; wherein the RNA target and the gRNA are continuously co-expressed as a single RNA transcript driven by the first promoter.
[0036] In yet another aspect, the present disclosure relates to a construct library comprising a plurality of polynucleotide constructs as disclosed herein, wherein the library is characterized by one or more of the following: a) the plurality of polynucleotide constructs encode different DNA or RNA targets; b) the plurality of polynucleotide constructs encode different gRNAs.
[0037] In another aspect, the present disclosure relates to one or more compartments, each compartment comprising a polynucleotide construct as disclosed herein, wherein the compartments are separated from each other. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The invention will be better understood with reference to the detailed description when considered in conjunction with the non-limiting examples and accompanying drawings, in which:
[0039] Figure 1 . Schematic diagram illustrating a non-limiting list of key concepts and steps of the present invention, wherein compartmentalization is performed by generating water-in-oil emulsion droplets.
[0040] Figure 2 Schematic diagram illustrating non-limiting examples of polynucleotide constructs as disclosed in the present disclosure. Note that 'Cas nuclease' can be replaced by any nucleic acid modification enzyme and can also refer to Cas variants, such as inactivated Cas nucleases, or Cas proteins fused or associated with function-conferring domains.
[0041] Figure 3.A graphical representation of how DNA / RNA molecule reads are counted to calculate enzyme activity, in one embodiment, where the enzyme is a Cas nuclease and the modification is DNA cleavage. In this embodiment, the DNA target site is located 3' of the encoded Cas variant. Nanopore sequencing reads (aligned to the reference sequence) with aligned 3' ends that map to the reference sequence at a site 3' downstream of the expected Cas cleavage site window are considered uncleaved ("Nanopore sequencing aligned reads" for dark grey bars); Figure 3 ), while aligned reads with their 3′ ends located within the Cas cleavage site window were considered cleaved (light grey bars “nanopore sequencing aligned reads”; Figure 3 ), and reads that did not meet either criterion were discarded as uninformative reads, as it was not possible to empirically determine whether these reads were cleaved (white bars for “nanopore sequencing-aligned reads”; Figure 3 ).
[0042] Figure 4 .Gel visualization of purified IVTT SpCas9 and dCas9 DNA constructs from compartmentalized (via emulsion) IVTT reactions and bulk IVTT reactions. 750 ng of Sp Cas9 construct was mixed with IVTT reagent (New England Biolabs PURExpress #E6800) on ice to produce a 75 μL IVTT aqueous mixture. 50 μL of this aqueous mixture was added to the oil-surfactant mixture on ice in 5 portions of 10 μL over 2 minutes while the stirring bar was rotated at 1150 rpm to produce an emulsion mixture. The emulsion mixture was allowed to continue mixing on ice for another minute. In one embodiment, the emulsion mixture was then homogenized (8000 rpm for 3 minutes; IKA Ultraturrax T10 homogenizer) to produce a more monodisperse distribution of emulsion droplet size. The remaining 25 μL of the aqueous mixture was kept on ice for bulk IVTT reaction as a control. This operation was also repeated for the Sp dCas9 construct. The emulsion and bulk IVTT mixtures were then incubated at 37°C for 4 hours to perform IVTT, followed by incubation at 65°C for 15 minutes to inactivate the proteins. DNA from all IVTT reactions was then individually purified, and aliquots were visualized after size separation by gel electrophoresis on agarose gels. This data demonstrates that the IVTT reagents successfully transcribed and translated proteins in both the bulk reactions and the emulsion droplets.
[0043] Figure 5Nanopore sequencing reads from an emulsion IVTT self-cleavage assay using a high input concentration of the SpCas9 construct. A small subset of reads in this sublibrary mapped to Sp dCas9 and are therefore classified as misassigned on the plot (light grey; Figure 5 ), as only Sp Cas9 DNA was provided as input for this emulsion IVTT reaction. Sp Cas9 emulsion IVTT nanopore sequencing reads showing detected cleaved and uncleaved construct fragments (white and black, respectively; Figure 5 Thus, this data demonstrates that nanopore single-molecule sequencing can detect both modified and unmodified polynucleotide constructs (enzymatically active or inactive products) from emulsion IVTT reactions.
[0044] Figure 6 Nanopore sequencing reads from an emulsion IVTT self-cleavage assay using a high input concentration of the SpdCas9 construct. Reads that failed the alignment quality filter were categorized accordingly. The SpdCas9 emulsion IVTT nanopore sequencing reads overwhelmingly appear as uncleaved construct fragments, as expected (striped gray; Figure 6 A small subset of reads in this sublibrary mapped to Sp Cas9 and are therefore classified as misassigned on the graph (light grey; Figure 6 ), because only Sp dCas9 DNA was provided as input for the emulsion IVTT reaction. This result supports the robustness of this method, as the inactivation of Sp dCas9 was accurately detected and measured by sequencing reads.
[0045] Figure 7 A diagram depicting an exemplary workflow for large-volume IVTT and self-cleavage assay time course experiments with readout of nanopore sequencing results. In this example, large-volume IVTT reactions were set up on ice for different CRISPR-Cas constructs (e.g., Sp Cas9, SaCas9, As Cpf1, Lb Cpf1), all with components arranged similarly to those described for the nucleic acid template sequences above. They were then divided equally into five corresponding aliquots for each time point ( Figure 7 Part 1). These large-volume IVTT aliquots were then incubated at 37°C and removed at each designated time point, quenched with EDTA inhibitor and enzyme to stop the IVTT reaction and Cas cleavage of the encoding DNA construct ( Figure 7 Part 2). The quenched IVTT reaction was then cleaned up using SPRIselect beads to purify the DNA fragments ( Figure 7Part 3). Small aliquots of these DNA fragments from different Cas orthologs at different IVTT time points were then visualized after size sorting by gel electrophoresis on agarose gels, as follows Figure 8 The remaining aliquots of purified DNA fragments were then pooled according to their respective time points, regardless of the Cas species, i.e., DNA fragments of Sp Cas9, Sa Cas9, etc. were mixed together at each time point and barcoded individually using the ONT EXP-NBD104 PCR-Free Amplification-Free Barcoding Extension Kit ( Figure 7 Part 4) to multiplex these pooled sublibraries for single nanopore sequencing runs ( Figure 7 Part 5). The nanopore sequencing results were then quality filtered and analyzed using publicly available bioinformatics tools followed by the analysis methods disclosed herein.
[0046] Figure 8 .From Figure 7 Gel visualization of purified IVTT constructs of different CRISPR-Cas orthologs of the bulk IVTT reaction following the steps shown in Section 3. This data demonstrates that different Cas proteins (variants or orthologs) were successfully transcribed and translated in the bulk reaction.
[0047] Figure 9 Graph of Cas-encoding DNA fragments detected by nanopore sequencing from large-volume IVTT and self-cleavage assay time-course experiments. This data demonstrates that single-molecule sequencing can detect enzyme products and measure the enzymatic activities of different nucleic acid-modifying enzymes in a multiplexed manner.
[0048] Figure 10.Gel visualization of purified IVTT Sp Cas9 and dCas9 DNA constructs from large-volume IVTT reactions. 500 ng of Sp Cas9 (sequence as described above) was mixed with IVTT reagent (New England Biolabs PURExpress # E6800) on ice to produce a 50 μL IVTT aqueous mixture. The same operation was performed on the Sp dCas9 construct; the Sp dCas9 construct contained a DNA sequence substantially identical to that of the Sp Cas9 construct, except that two inactivating mutations (D10A and H840A) in the Sp Cas9 gene produced the Sp dCas9 gene. These 50 μL large-volume IVTT reactions were incubated at 37°C for 4 hours for IVTT and then incubated at 65°C for 15 minutes to inactivate the protein. 20 mM EDTA (pH 8.0) inhibitor was added to the large-volume IVTT reaction with RNase mixture and proteinase K for 30 minutes at 37°C to remove excess RNA and protein from the IVTT reaction. The DNA (polynucleotide constructs) from these two large-volume IVTT reactions were then purified separately using SPRIselect paramagnetic beads, and aliquots of the DNA were then visualized after size sorting by gel electrophoresis on agarose gels. This data demonstrates that the Cas protein was successfully transcribed and translated in the large-volume IVTT reaction.
[0049] Figure 11 Schematic diagram of direct detection and counting of polynucleotides that have been modified or not modified by nucleic acid modifying enzymes. Sp Cas9 and Sp dCas9 DNA constructs purified from large-volume IVTT reactions (gel visualization depicted in Figure 10 (in) are mixed together in different ratios. These mixtures of purified DNA constructs are then prepared for nanopore sequencing. By aligning all nanopore sequencing reads with the Sp dCas9 construct reference sequence, the presence of cleaved Sp Cas9 reads is detected using bioinformatics tools accessible to those of ordinary skill in the art of sequencing data analysis. This workflow is capable of detecting variations (indels - insertions and deletions or SNPs - single nucleotide polymorphisms) in sequencing reads aligned to the reference sequence; with particular attention paid to the detection of SNPs that represent expected sequence differences between otherwise identical Sp dCas9 and Sp Cas9 constructs, i.e., D10A and H840A catalytically inactivating mutations in Sp dCas9 relative to Sp Cas9. As Figure 3As shown, the aligned raw nanopore sequencing reads are classified as cleaved and uncleaved by sequence mapping for the Sp dCas9 reference sequence, and then SNP detection is performed, and SNP causes amino acid residue changes. The Sp Cas9 sequence on each filtered alignment read is translated into its corresponding amino acid sequence, and the detected SNPs that cause amino acid changes with the Sp dCas9 reference amino acid sequence are counted. In the above chart, the detected SNPs are shown in the heat map in the selected target region containing D10A and H840A catalytic inactivation mutations in Sp dCas9 relative to the Sp dCas9 reference. Reads classified as cleavage (2 subgraphs on the left; Figure 11 ) are enriched in SNPs corresponding to residues D10 and H840 (dark grey squares in the heat map; Figure 11 ), i.e., these cleaved reads contain catalytically active Sp Cas9 sequences. The other detected SNPs shown in the above graphs that result in amino acid mutations and have lighter grey squares in the heat map are false positives, which are caused by raw sequencing errors inherent in currently available nanopore sequencing technologies. This data illustrates the detection of cleaved and uncleaved Sp Cas9 DNA fragments, which can be distinguished by detecting uncleaved Sp dCas9 DNA fragments in raw nanopore sequencing data. Notably, this method was even able to detect 1:10 of the purified Sp dCas9 and Sp Cas9 bulk IVTT DNA products, respectively. -5 The presence of cleaved Sp Cas9 DNA fragments was detected in the mixture ( Figure 11 ).
[0050] Figure 12 Nanopore sequencing reads from an emulsion IVTT self-cleavage assay using a limited input concentration of the SpCas9 construct. As expected for the SpCas9 enzyme, emulsion IVTT nanopore sequencing reads show detected cleavage (white; Figure 12 ) and uncleaved (black part; Figure 12 ) construct fragments. Thus, this data supports the robustness of the assay in which IVTT and enzymatic reactions were performed in emulsion droplets. A small subset of reads in this sublibrary mapped to Sp dCas9 and are therefore classified as misassigned on the graph (light gray; Figure 12 ), because only Sp Cas9 DNA was provided as input for this emulsion IVTT reaction.
[0051] Figure 13Nanopore sequencing reads from an emulsion IVTT self-cleavage assay using a limited input concentration of the SpdCas9 construct. The SpdCas9 emulsion IVTT nanopore sequencing reads appear mostly uncleaved (striped grey; Figure 13 ) construct fragments, indicating that Sp dCas9 is inactive most of the time, as expected. Therefore, this data also supports the robustness of the assay in which IVTT and enzymatic reactions are performed in emulsion droplets. A small subset of reads in this sublibrary mapped to Sp Cas9 and are therefore classified as misassigned on the graph (light gray portion; Figure 13 ), because only Sp dCas9 DNA was provided as input for this emulsion IVTT reaction.
[0052] Figure 14 Nanopore sequencing reads from an emulsion IVTT self-cleavage assay, in which limited input concentrations of Sp Cas9 and Sp dCas9 constructs were provided in an equimolar ratio. Nanopore sequencing reads show a roughly equal distribution of Sp Cas9 and Sp dCas9 mapped reads, as expected. In addition, Sp Cas9 mapped reads appear to be almost equally divided into cleaved and uncleaved fragments (white and black fractions, respectively; Figure 14 ), while the vast majority of Sp dCas9 mapped reads were classified as uncleaved (gray with stripes; Figure 14 Thus, this data further demonstrates that the methods disclosed herein can measure the enzymatic activity of different variants (in this example, Cas variants, but the method can also be used to screen variants of other components in the enzymatic reaction, such as targets or gRNAs).
[0053] definition
[0054] Several terms used throughout the specification are defined in the following paragraphs. Additional definitions can also be found in the text of the specification.
[0055] As used herein, the terms "about" and "approximately" with respect to numbers are used herein to include numbers that fall within 20%, 10%, 5%, 2.5%, 2%, 1.5% or 1% in either direction (greater than or less than) of that number unless otherwise stated or obvious from the context (except where the number would exceed 100% of the possible value).
[0056] The terms "polynucleotide," "nucleic acid," and "oligonucleotide" are used interchangeably and refer to a polymeric form of nucleotides (deoxyribonucleotides or ribonucleotides or their analogs) of any length. A polynucleotide can have any three-dimensional structure and can perform any known or unknown function. The following are non-limiting examples of polynucleotides: genes or gene fragments (e.g., probes, primers, EST or SAGE tags), exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes and primers. A polynucleotide can comprise modified nucleotides, such as methylated nucleotides and nucleotide analogs. If there is a modification to the nucleotide structure, the modification can be imparted before or after the polynucleotide is assembled. The nucleotide sequence can be interrupted by non-nucleotide components. A polynucleotide can be further modified after polymerization, for example, by conjugation with a labeling component. The term also refers to double-stranded and single-stranded molecules. Unless otherwise specified or required, a polynucleotide includes a double-stranded form and each of the two complementary single-stranded forms known or predicted to constitute the double-stranded form. As used herein, the term "polypeptide" generally has the meaning of an amino acid polymer recognized in the art. The term is also used to refer to specific functional classes of polypeptides, such as nucleases, antibodies, etc.
[0057] As used herein, the term "operably linked" refers to a juxtaposition in which the described components are in a relationship that allows them to function in their intended manner. A control element, such as a promoter, that is "operably linked" to a functional element is associated in a manner that achieves expression and / or activity of the functional element under conditions compatible with the control element. In some embodiments, an "operably linked" control element is adjacent to (e.g., covalently linked to) a coding element of interest; in some embodiments, the control element acts in trans or otherwise on the functional element of interest.
[0058] The term "nucleic acid modifying enzyme" refers to a macromolecular biocatalyst that can be protein or nucleic acid in nature and is capable of modifying nucleic acids. The term "RNA-guided nucleic acid modifying enzyme" generally refers to an enzyme that interacts with or forms a complex with a guide RNA and can specifically target or bind to a polynucleotide having a specific sequence, which typically comprises a sequence complementary to the targeting domain of the gRNA. Upon binding to the target polynucleotide, the RNA-guided nucleic acid modifying enzyme can remain bound to the target polynucleotide, or if the RNA-guided nucleic acid modifying enzyme is a nuclease, it can cleave the target polynucleotide; or if the RNA-guided nucleic acid modifying enzyme has a functional domain, it can modify the polynucleotide in other ways. In one example, the RNA-guided nucleic acid modifying enzyme is a CRISPR-associated protein (Cas). Many Cas proteins have endonuclease activity, also known as Cas nucleases. In one embodiment, the RNA-guided nucleic acid modification enzyme is selected from the group consisting of Cas3, Cas9, Cas10, Cas12a (also known as Cpf1), Cas13a (also known as C2c2), Cas13b, Cas13c, Cas13d, Cas14, CasX, CasΦ, and variants thereof.
[0059] The terms "guide RNA" and "gRNA" refer to any nucleic acid that facilitates the specific binding (or "targeting") of an RNA-guided nucleic acid-modifying enzyme to a target sequence in a cell or a cell-free environment. gRNAs can be unimolecular (comprising a single RNA molecule, or a chimera) or modular (comprising one or more, typically two, separate RNA molecules, such as crRNA and tracrRNA, which are typically associated with each other, for example, by forming a duplex).
[0060] As used herein, the term "target" (or "target site") refers to a nucleic acid sequence that defines a portion of a nucleic acid (or polynucleotide) to which a binding molecule will bind, provided that sufficient binding conditions exist. In some embodiments, a target site is a nucleic acid sequence that is bound by and / or modified by a nucleic acid modification enzyme as described herein. In some embodiments, a target is a nucleic acid sequence that is bound by a guide RNA as described herein. The target can be single-stranded or double-stranded. Nucleic acid modification enzymes as disclosed herein can modify DNA or RNA. Thus, a "target" can be a DNA sequence or an RNA sequence, referred to as a "DNA target" and an "RNA target," respectively. In the case of a dimeric nuclease, such as a nuclease comprising a Fok1 DNA cleavage domain, the target typically comprises a left half-site (bound by one monomer of the nuclease), a right half-site (bound by the second monomer of the nuclease), and a spacer sequence between the half-sites that are cut. In some embodiments, the length of the left half-site and / or the right half-site is between 10-18 nucleotides. In some embodiments, one or both half-sites are shorter or longer. In some embodiments, the left half site and the right half site include different nucleic acid sequences. In the case of zinc finger nucleases, in some embodiments, the target can include two half sites, and the length of each half site is 6-18bp, which is located on both sides of the non-specified spacer of 4-8bp length. In the case of TALEN, in some embodiments, the target can include two half sites, and the length of each half site is 10-23bp, which is located on both sides of the non-specified spacer of 10-30bp length. In the case of RNA-guided (e.g., RNA programmable) nucleic acid modification enzymes, the target generally includes a nucleotide sequence (e.g., "protospacer" in CRISPR-Cas) complementary to guide RNA (gRNA), and a protospacer adjacent motif (PAM) adjacent to the guide RNA complementary sequence at 3' end or 5' end. For CRISPR-Cas enzymes (e.g., Cas13 family) targeting RNA, the RNA target can include a protospacer flanking sequence (PFS) instead of a PAM sequence. In some embodiments, the DNA or RNA target of the Cas enzyme can comprise 16-24 nucleotides in length complementary to the gRNA, and 3-6 base pairs of PAM / PFS (e.g., NNN, where N represents any nucleotide).
[0061] As used herein, "binding" refers to a non-covalent interaction between macromolecules (eg, between a protein and a polynucleotide).
[0062] "Modification" of a polynucleotide refers to any chemical or physical change in the composition or structure of the polynucleotide, including fragmentation / cleavage of the polynucleotide, nicking (single-strand breaks) in a double-stranded polynucleotide, substitution of one or more nucleotide bases, insertion or deletion of one or more nucleotide bases, or covalent modification of nucleotide bases with chemical and epigenetic marks (e.g., cytosine methylation and hydroxymethylation).
[0063] As used herein, the term "variant" refers to an entity that shows great structural identity to a reference entity, but is structurally different from the reference entity in terms of the presence or level of one or more chemical moieties compared to the reference entity. In many embodiments, the variant is also functionally different from its reference entity. In general, whether a particular entity is properly considered a "variant" of a reference entity is based on the degree of structural identity it has with the reference entity. As will be understood by those skilled in the art, any biological or chemical reference entity has certain characteristic structural elements. By definition, a variant is a unique chemical entity that shares one or more such characteristic structural elements. To give just a few examples, a polypeptide can have a characteristic sequence element comprising a plurality of amino acids that have a specified position relative to each other in linear or three-dimensional space and / or contribute to a specific biological function; a nucleic acid can have a characteristic sequence element consisting of a plurality of nucleotide residues that have a specified position relative to each other in linear or three-dimensional space. For example, a variant polypeptide may differ from a reference polypeptide due to one or more differences in amino acid sequence and / or one or more differences in chemical moieties (e.g., carbohydrates, lipids, etc.) covalently attached to the polypeptide backbone. In some embodiments, the variant polypeptide shows an overall sequence identity of at least 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97% or 99% to a reference polypeptide (e.g., a nucleic acid modification enzyme as described herein). Alternatively or in addition, in some embodiments, the variant polypeptide does not share at least one characteristic sequence element with the reference polypeptide. In some embodiments, the reference polypeptide has one or more biological activities. In some embodiments, the variant polypeptide shares one or more biological activities of the reference polypeptide, such as enzymatic activity. In some embodiments, the variant polypeptide lacks one or more biological activities of the reference polypeptide. In some embodiments, compared to the reference polypeptide, the variant polypeptide shows reduced levels of one or more biological activities (e.g., enzymatic activity). In some embodiments, if the target polypeptide has an amino acid sequence identical to that of the parent but has an amino acid sequence with a small amount of sequence changes at a specific position, the target polypeptide is considered to be a "variant" of the parent or reference polypeptide. Typically, less than 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% of the residues in the variant are substituted compared to the parent. In some embodiments, the variant has 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 substituted residues compared to the parent. Typically, the variant has a very small amount (e.g., less than 5, 4, 3, 2 or 1) of substituted functional residues (i.e., residues involved in a specific biological activity). In addition, compared to the parent, the variant typically has no more than 5, 4, 3, 2 or 1 additions or deletions, and typically no additions or deletions.Furthermore, any additions or deletions are typically less than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, about 6, and often less than about 5, about 4, about 3, or about 2 residues. In some embodiments, the parent or reference polypeptide is one found in nature.
[0064] As used herein in the context of nucleic acids or proteins, the term "library" refers to a population of two or more different polynucleotide constructs or proteins, respectively. In some embodiments, the polynucleotide construct library comprises at least two polynucleotide constructs comprising different sequences encoding nucleic acid modification enzymes, at least two polynucleotide constructs comprising different sequences encoding guide RNAs, at least two polynucleotide constructs comprising different PAMs, and / or at least two nucleic acid molecules comprising different target sites. In some examples, the library comprises at least 10 1 , at least 10 2 , at least 10 3 , at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11 , at least 10 12 , at least 10 13 , at least 10 14 or at least 10 15 In some embodiments, the members of the library may comprise randomized sequences, e.g., completely or partially randomized sequences. In some embodiments, the library comprises nucleic acid molecules that are unrelated to each other, e.g., nucleic acids comprising completely randomized sequences. In other embodiments, at least some members of the library may be related, e.g., they may be variants or derivatives of a particular sequence.
[0065] As used herein, the term "expression" of a nucleic acid sequence refers to the production of any gene product from the nucleic acid sequence. In some examples, the gene product can be an RNA transcript. In some embodiments, the gene product can be a polypeptide. In some embodiments, expression of a nucleic acid sequence involves one or more of the following: (1) generation of an RNA template from a DNA sequence (e.g., by transcription); (2) processing of the RNA transcript (e.g., by splicing, editing, 5' cap formation, and / or 3' end formation); (3) translation of the RNA into a polypeptide or protein; and / or (4) post-translational modification of the polypeptide or protein.
[0066] As used herein in the context of partitioning polynucleotide constructs into compartments, the term "compartment" can refer to any physical or virtual compartment, such as emulsion droplets and nanowells, as well as virtual compartments such as microfluidics or hydrogels that can separate reagents and reaction systems.
[0067] As used herein, the term "promoter" refers to a transcription promoter that confers accurate transcription initiation. Promoters as used herein include any promoter that can be used to produce mRNA encoding proteins (e.g., Cas proteins) or RNA transcripts (e.g., guide RNAs). In some instances, the promoter is compatible with cell-free in vitro transcription and translation reactions. Examples of promoters that can be used in the context of the present invention include, but are not limited to, T7 promoters, SP6 promoters, Lac promoters, and the like. As used herein, the term "terminator" refers to a transcription terminator that defines the end of a transcription unit (e.g., a gene) and initiates the process of releasing newly synthesized RNA from the transcription machinery. Examples of terminators include, but are not limited to, T7 terminators and rrnB terminators.
[0068] Detailed Description of the Invention
[0069] In particular, the inventors of the present invention have developed a multiplexing method to measure the activity of nucleic acid modifying enzymes and screen one or more variable elements of enzymatic reactions. For example, the method can physically link the activity of nucleic acid modifying enzyme variants to their own encoding DNA / RNA and their DNA / RNA target molecules, while directly measuring the enzyme activity and inactivation of each variant to individual target molecules (i.e., the degree of activity of the active variant is quantitative, or the inactive variant can be measured as 'inactivation') on the molecule regardless of the activity level. This achieves a direct approach to engineered nucleic acid modifying enzymes (e.g., CRISPR-Cas) to provide them with enhanced or novel functions (through active variants), while constructing a fitness landscape map (fitness landscape map) of currently useless sequence variations (through inactive or less active variants). Similarly, variants of guide RNA and / or DNA / RNA targets can also be screened using the methods disclosed herein.
[0070] A non-limiting and non-exhaustive list of the key concepts of the present invention is described below: (i) a polynucleotide construct encoding a DNA / RNA target site and a variable element to be tested (e.g., a nucleic acid modifying enzyme variant), (ii) mixing the DNA with any commonly used RNA and protein expression reagents (also known as a cell-free transcription-translation (TXTL) / in vitro transcription-translation (IVTT) reaction), (iii) encapsulating or compartmentalizing a single copy of the DNA construct variant along with the IVTT reagents, (iv) allowing IVTT reactions in separate compartments, expressing the nucleic acid modifying enzyme and sg in each compartment isolated from the other compartments. RNA (if the enzyme is an RNA-guided enzyme) (and in some embodiments, the RNA target is co-transcribed as part of the Cas transcript), (v) depending on the functionality of the encoded nucleic acid modification enzyme, a single polynucleotide construct (or in some embodiments, an RNA target transcribed from the construct) is cleaved, left intact, or otherwise modified, and (vi) the cleaved, intact, or modified polynucleotide constructs (or in some embodiments, RNA targets) are quantified in parallel, for example by single-molecule long-read sequencing, thereby directly identifying and directly quantifying the enzymatic activity associated with each variable element in a molecularly parallel manner. This technology directly links the phenotype of the encoded variable element (e.g., nucleic acid modification enzyme variant) to its coding sequence, allowing rapid determination of sequence-function relationships for large variant libraries. Figure 1 A non-limiting list of key concepts of the invention is depicted.
[0071] method
[0072] The method disclosed herein can be characterized as a method for measuring enzyme activity. Since the method is highly scalable and can screen a large number of variant polynucleotides, the method can also be characterized as a method for screening nucleic acid modifying enzymes and / or (nucleic acid modifying enzymes) DNA / RNA targets and / or guide RNAs and / or other components of enzymatic reactions that can be encoded on polynucleotide constructs. Therefore, in one aspect, the present disclosure relates to a method comprising the following steps:
[0073] a) separating the plurality of polynucleotide constructs into compartments, wherein each compartment comprises a single polynucleotide construct, wherein each polynucleotide construct comprises:
[0074] i) a first polynucleotide sequence encoding a nucleic acid modification enzyme or a variant thereof operably linked to a first promoter; and
[0075] ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, wherein when the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target and the nucleic acid modification enzyme are continuously co-expressed as a single transcript driven by the first promoter;
[0076] and wherein the plurality of polynucleotide constructs encode different variants of the nucleic acid modification enzyme, and / or different DNA or RNA targets;
[0077] b) subjecting the compartment to conditions allowing in vitro expression of RNA and protein;
[0078] c) subjecting a plurality of said compartments to conditions that allow modification of said DNA / RNA target by a nucleic acid modifying enzyme having modification activity on said DNA or RNA target, thereby producing a population of DNA / RNA molecules comprising one or more of:
[0079] iii. a polynucleotide construct and / or RNA target or fragment thereof that has been modified by the nucleic acid modifying enzyme;
[0080] iv. a polynucleotide construct and / or RNA target that is not modified by the nucleic acid modifying enzyme;
[0081] d) harvesting the DNA / RNA molecule population produced in step (c) and performing single molecule sequencing on it;
[0082] e) detecting and counting the DNA / RNA molecules mentioned in steps c)i and c)ii based on the sequencing results.
[0083] In this first aspect, polynucleotide constructs encode nucleic acid modifying enzyme (or its variant) and DNA / RNA target.Therefore, nucleic acid modifying enzyme or DNA / RNA target can be tested or screened as variable element.In some examples of the method for testing or measuring the activity (i.e. screening enzyme) of different nucleic acid modifying enzymes to specific targets, multiple polynucleotide constructs can encode identical DNA / RNA target, but encode different nucleic acid modifying enzymes (or different variants of same nucleic acid modifying enzyme).In some examples of the method for testing or measuring the activity (i.e. screening DNA / RNA target) of specific nucleic acid modifying enzymes to different DNA / RNA targets, multiple polynucleotide constructs can encode identical nucleic acid modifying enzyme, but encode different DNA / RNA targets.In the case of CRISPR-Cas target, statement " different DNA / RNA targets " can refer to the DNA / RNA targets different from prototype spacer (sequence complementary to guide RNA) or PAM / PFS sequences.
[0084] In some examples where the nucleic acid modifying enzyme encoded by each polynucleotide construct is an RNA-guided nucleic acid modifying enzyme (e.g., CRISPR-Cas nuclease or variants thereof), the nucleic acid modifying enzyme may require a guide RNA (gRNA) to bind and / or modify the DNA / RNA target. In some examples, the gRNA is directly provided to each compartment in the form of a gRNA or in the form of a DNA template encoding the gRNA. Therefore, in an example where the nucleic acid modifying enzyme is an RNA-guided nucleic acid modifying enzyme, each compartment also includes a guide RNA or a nucleotide template encoding it.
[0085] In some other examples, the gRNA can be encoded on the same polynucleotide construct encoding the enzyme and the DNA / RNA target. Thus, in one example, the nucleic acid modifying enzyme is an RNA-guided nucleic acid modifying enzyme, wherein each polynucleotide further comprises a third polynucleotide sequence encoding a variant guide RNA (gRNA); and wherein the multiple polynucleotide constructs encode different variants of the nucleic acid modifying enzyme, and / or different DNA or RNA targets, and / or different gRNAs. In this example, the method disclosed herein comprises the following steps:
[0086] a) separating the plurality of polynucleotide constructs into compartments, wherein each compartment comprises a single polynucleotide construct, wherein each polynucleotide construct comprises:
[0087] i) a first polynucleotide sequence encoding a nucleic acid modification enzyme or a variant thereof operably linked to a first promoter, wherein the nucleic acid modification enzyme is an RNA-guided nucleic acid modification enzyme;
[0088] ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, wherein when the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target and the nucleic acid modification enzyme are continuously co-expressed as a single transcript driven by the first promoter; and
[0089] iii) a third polynucleotide sequence encoding a variant guide RNA (gRNA);
[0090] wherein the multiple polynucleotide constructs encode different variants of nucleic acid modifying enzymes, and / or different DNA or RNA targets, and / or different gRNAs;
[0091] b) subjecting the compartment to conditions allowing in vitro expression of RNA and protein;
[0092] c) subjecting a plurality of said compartments to conditions that allow modification of said DNA / RNA target by a nucleic acid modifying enzyme functionally active against said DNA or RNA target in the presence of a gRNA, thereby producing a population of DNA / RNA molecules comprising one or more of:
[0093] i. a polynucleotide construct and / or RNA target or a fragment thereof that has been modified by the nucleic acid modifying enzyme;
[0094] ii. a polynucleotide construct and / or RNA target that has not been modified by the nucleic acid modifying enzyme;
[0095] d) harvesting the DNA / RNA molecule population produced in step (c), and performing single-molecule long-read sequencing on it;
[0096] e) detecting and counting the DNA / RNA molecules mentioned in steps c)i and c)ii based on the sequencing results.
[0097] In the above examples, because gRNA (or the sequence encoding the gRNA) is physically connected to nucleic acid modifying enzyme and DNA / RNA target, any one of gRNA, DNA / RNA target and nucleic acid modifying enzyme can be tested and screened as variable element. Encoded gRNA will be expressed (such as in a compartment) from a polynucleotide construct, so the polynucleotide construct can include other elements promoting gRNA expression, and these other elements are generally known to those skilled in the art. In some instances, the third polynucleotide sequence is operably connected to the second promoter. In some instances, the second promoter is a T7 promoter.
[0098] Screening for nucleic acid modifying enzymes
[0099] In some examples where the method is used to test or measure the activity of different nucleic acid modification enzymes on a particular target, multiple polynucleotide constructs can encode the same DNA / RNA target and gRNA, but encode different nucleic acid modification enzymes (or different variants of the same nucleic acid modification enzyme).
[0100] One example of a nucleic acid-modifying enzyme that can be tested, screened, and optimized using this approach is the Cas family of nucleases. In recent years, a variety of CRISPR-Cas systems for DNA and RNA editing have been developed, enabling widespread application across all areas of medicine and biotechnology. Of particular interest are Class 2 Cas (CRISPR-associated) proteins, including the Cas9, Cas12 (formerly known as Cpf1), Cas13, and Cas14 nucleases that have been well characterized in the literature. These Cas proteins are single-component nuclease effectors (i.e., a single Cas protein, rather than a multimeric complex of different proteins); they typically utilize RNA oligonucleotides (guide RNA, gRNA; also known as engineered forms of single guide RNA, sgRNA; used interchangeably) to program and colocalize the Cas protein to a specific site on the DNA and / or RNA, where enzymatic activity can then occur, such as cleavage (endonucleolytic cleavage in the DNA / RNA). A segment of the gRNA sequence (the spacer) is complementary to the target sequence (the protospacer) of the DNA / RNA. Another short sequence (typically 2-6 nt in length) adjacent to the protospacer is required for functional targeting, and is also referred to as a protospacer adjacent motif (PAM; when on DNA) or a protospacer flanking sequence (PFS; when on RNA). Each Cas-gRNA system can recognize a unique PAM / PFS site and has different gRNA:protospacer requirements. Cas proteins have been and can be further engineered to recognize new PAM / PFS sites, have less stringent gRNA lengths or structures, and are more specific and effective. In order to use Cas nucleases as therapeutic agents while minimizing adverse immune responses, immunogenic epitopes in Cas proteins can also be removed or masked, particularly by deleting or changing the amino acid sequence while maintaining Cas function. New functions can also be engineered into Cas proteins or Cas fusion proteins, for example to achieve base editing (changing the target nucleotide to another), epigenetic modifications, or many other modifications that have not yet been demonstrated. These efforts typically require some form of directed evolution, protein engineering, selection, and screening of Cas variant libraries. The method disclosed herein can be used to measure and screen the activity of large libraries of enzyme (e.g., Cas) mutants because it is i) highly scalable, with >10 9 ii) are compatible with a larger sequence space, which is particularly important when working with large proteins (>10 3 aa long), such as CRISPR-Cas proteins, are particularly important and useful when working together.
[0101] Screening DNA / RNA targets of nucleic acid-modifying enzymes
[0102] In some examples where the method is used to test or measure the activity of a specific nucleic acid modification enzyme on different DNA / RNA targets in the presence of a specific gRNA, multiple polynucleotide constructs can encode the same nucleic acid modification enzyme and gRNA, but encode different DNA / RNA targets.
[0103] In these examples, the methods disclosed herein can be used to evaluate the ability of PAM or PFS variants to guide the binding or modification of DNA / RNA targets by RNA-guided nucleic acid modifying enzymes. The methods disclosed herein allow for the simultaneous evaluation of multiple PAM / PFS variants for any given target site. Therefore, the data obtained from such methods can be used to compile a list of PAM variants that modify (e.g., cleave) a specific DNA / RNA target. It will be apparent to those skilled in the art that this method can also be used to test and screen for any non-PAM / PFS sequences at the target site that may have an impact on enzyme activity.
[0104] Screening guide RNA
[0105] In some examples where the method is used to test or measure the activity of a specific nucleic acid modification enzyme on a specific DNA / RNA target in the presence of different specific gRNAs, multiple polynucleotide constructs can encode the same nucleic acid modification enzyme and DNA / RNA target, but encode different gRNAs.
[0106] In these examples, the present disclosure provides methods for evaluating the ability of different gRNAs to mediate binding and / or modification of a specific DNA / RNA target by a nucleic acid modifying enzyme. Thus, the results obtained from these methods can be used to compile a list of guide RNA variants that mediate modification of a specific target by a specific nucleic acid modifying enzyme.
[0107] In another aspect, the present disclosure relates to a method comprising the steps of:
[0108] a) separating the plurality of polynucleotide constructs into compartments, wherein each compartment comprises a single polynucleotide construct, wherein each polynucleotide construct comprises:
[0109] i) a first polynucleotide sequence encoding a guide RNA (gRNA) operably linked to a first promoter;
[0110] ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, wherein when the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target and the gRNA are continuously co-expressed as a single RNA transcript driven by the first promoter;
[0111] wherein the plurality of polynucleotide constructs encode different gRNAs, and / or different DNA or RNA targets; and wherein each compartment further comprises an RNA-guided nucleic acid modification enzyme or a variant thereof or a nucleotide template encoding the same;
[0112] b) subjecting the compartment to conditions allowing in vitro transcription and / or translation of RNA and protein;
[0113] c) subjecting the compartment to conditions that allow modification of the DNA / RNA target by an RNA-guided nucleic acid modification enzyme functionally active against the DNA or RNA target in the presence of a gRNA, thereby producing a population of DNA / RNA molecules comprising one or more of:
[0114] i) a polynucleotide construct and / or RNA transcript or a fragment thereof that has been modified by said nucleic acid modifying enzyme;
[0115] ii) a polynucleotide construct and / or RNA transcript that has not been modified by the nucleic acid modifying enzyme;
[0116] d) harvesting the DNA / RNA molecule population produced in step (c), and performing single-molecule long-read sequencing on it;
[0117] e) detecting and counting the DNA / RNA molecules mentioned in steps c)i and c)ii based on the sequencing results.
[0118] In this respect, polynucleotide constructs encode guide RNA (gRNA) and DNA / RNA target, and the nucleic acid modifying enzyme that RNA guides is provided to each compartment respectively.Therefore, gRNA or DNA / RNA target can be tested or screened as variable element.In the method, for testing or measuring specific nucleic acid modifying enzyme to some examples of the activity (i.e. screening gRNA) of specific target in the presence of different gRNA, a plurality of polynucleotide constructs can encode identical DNA / RNA target, but encode different nucleic acid modifying enzymes (or different variants of same nucleic acid modifying enzyme).In the method, for testing or measuring specific nucleic acid modifying enzyme to some examples of the activity (i.e. screening DNA / RNA target) of different DNA / RNA targets, a plurality of polynucleotide constructs can encode identical gRNA, but encode different DNA / RNA targets.
[0119] Compartmentalization of polynucleotide constructs into compartments
[0120] A variety of methods for separating polynucleotide constructs into compartments are known to those skilled in the art. In one example, polynucleotide constructs are separated into emulsion droplets by an emulsification method commonly known in the art. Generally, emulsions can be prepared by any suitable combination of immiscible liquids. In a typical example, the emulsion comprises an aqueous phase comprising (a) components required for in vitro transcription and translation; and (b) a nucleic acid template library as described herein. In the emulsion, the aqueous phase exists in the form of fine droplets (dispersed phase, internal phase, or discontinuous phase). The emulsion also comprises a hydrophobic immiscible liquid ("oil") as a matrix (non-dispersed phase, continuous phase, or external phase) for the droplet suspension. This type of emulsion is called "water-in-oil" (W / O), and the droplets are called "water-in-oil droplets." Many oils and many emulsifiers are known in the art and can be used to produce water-in-oil emulsions. Suitable emulsifiers include, for example, light white mineral oil and surfactants such as sorbitan monooleate (Span 80; ICI) and polyoxyethylene sorbitan monooleate (Tween 80; ICI), or any combination thereof. In one example, the emulsifier comprises mineral oil, Span 80, and a surfactant such as Tween 80; for example, mineral oil + 4.5% (v / v) Span 80 + 0.5% (v / v) Tween 80. Testing different emulsifiers is within the knowledge of those skilled in the art. In some examples, mechanical energy is used to force the phases together to create an emulsion. Various methods can be employed, including, but not limited to, the use of mechanical devices, including stirrers (e.g., magnetic stir bars, propeller and turbine stirrers, paddle devices, and whisks), homogenizers (including rotor-stator homogenizers, high-pressure valve homogenizers, and jet homogenizers), colloid mills, ultrasound, and "membrane emulsification" devices. One skilled in the art can vary the size of the emulsion droplets (compartments) by tailoring the emulsion conditions used to form the emulsion according to the requirements of the chosen system.
[0121] A non-limiting example is described herein: a water-in-oil (w / o) emulsion droplet is generated using the following steps or other methods known to those of ordinary skill in the art. In summary, 950 μL of an oil-surfactant mixture (mineral oil + 4.5% (v / v) Span 80 + 0.5% (v / v) Tween 80) is added to a cryovial with a 3×8 mm magnetic stir bar; placed on ice. ≤1.66 fmol of DNA library is mixed with IVTT reagent (New England Biolabs PURExpress #E6800) on ice to produce 50 μL of IVTT aqueous mixture. 50 μL of this aqueous mixture is added to the oil-surfactant mixture on ice in five 10 μL portions over 2 minutes while the stir bar is rotated at 1150 rpm to produce the emulsion mixture. The emulsion mixture is allowed to mix on ice for another minute. In one example, the stirred emulsion mixture is mixed for an additional 3 minutes at 8000 rpm using a homogenizer (eg, an IKA Ultraturrax T10 homogenizer) to achieve a more monodisperse distribution of emulsion droplet diameters.
[0122] Other methods of generating emulsion droplets are also possible and known to those skilled in the art, including vortexing the water and oil mixture, or using a microfluidic device (such as Dolomite-Bio's μ-encapsulator) to control the flow rate of water and oil inputs fed to the connection of the microfluidic chip, thereby encapsulating the aqueous solution with the oil in the form of emulsion droplets.
[0123] Other compartmentalization methods are also known to those of ordinary skill in the art. As used herein, the term "compartment" encompasses both virtual and physical compartmentalization, as long as compartmentalization can separate polynucleotide constructs, reagents, and reaction systems without producing physical encapsulation. In one example, the separation of compartments is achieved using microfluidics, hydrogels to restrict diffusion, or separation wells (or nanowells).
[0124] IVTT system
[0125] In some examples of the methods disclosed herein, each compartment contains in vitro transcription and translation (IVTT) reagents that are capable of achieving in vitro transcription and / or translation of proteins and / or RNA. Including IVTT in the compartment bypasses the use of cells in the assay. In some embodiments, the IVTT system includes a cell extract, such as a cell extract from bacteria, rabbit reticulocytes, or wheat germ. Many suitable systems are commercially available (e.g., from ThermoFisher, Promega, and New England Biolabs). In one example, the system can be emulsified with the polynucleotide construct. Conditions suitable for the in vitro transcription and translation mentioned in step b) will be apparent or available to those skilled in the art through reference or the manual of a commercial kit. In one non-limiting example, suitable conditions are a 4-hour incubation at 37°C. The IVTT reaction can be terminated by methods well known in the art or described in the manual of a commercial kit. In one example, the compartment containing the IVTT is incubated at 65°C for 15 minutes to heat inactivate the IVTT reagents and any expressed nucleic acid modification enzymes. In another example, 20 mM EDTA (pH 8.0) inhibitor was added to a compartment (eg, emulsion droplet) and mixed.
[0126] By controlling the compartmentalization conditions of IVTT reagents and DNA, it is possible to ensure that no more than one copy of a polynucleotide construct is encapsulated with IVTT reagents in each compartment, with volumes ranging from femtoliters to nanoliters. This enables each variant copy of DNA (and IVTT RNA and protein products) within each compartment to be physically isolated, allowing the user to physically confine the expressed RNA and protein to their respective encoding DNA.
[0127] Conditions allowing modification of DNA / RNA targets by known nucleic acid modifying enzymes are generally known in the art and / or can be easily found or optimized. For newly discovered enzymes, these conditions can generally be simulated using information about the related nucleic acid enzymes (e.g., homologues and orthologues) that are better characterized. Modification can refer to any chemical or physical change in the component or structure of the target, including fracture / cracking polynucleotides, creating nicks (single-strand breaks) in double-stranded polynucleotides, replacing one or more nucleotide bases, inserting or lacking one or more nucleotide bases, or covalently modifying nucleotide bases (e.g., cytosine methylation and hydroxymethylation) with chemical and epigenetic markers.
[0128] Since each compartment contains a single copy of the polynucleotide construct, the DNA / RNA target (which is contained on the polynucleotide construct or expressed from the construct) and the nucleic acid modifying enzyme are also confined to the compartment. The activity (or inactivity) of the nucleic acid modifying enzyme encoded on a particular construct on the DNA / RNA target encoded on the same construct will be reflected as the modification (or lack of modification) of the DNA / RNA target. Since these compartments as a whole contain a plurality of different polynucleotide constructs, step c) produces a population of DNA / RNA molecules comprising one or more of the following:
[0129] i. a polynucleotide construct and / or RNA transcript or a fragment thereof that has been modified by the nucleic acid modifying enzyme;
[0130] ii. a polynucleotide construct and / or RNA transcript that has not been modified by the nucleic acid modifying enzyme;
[0131] Wherein the polynucleotide construct comprises a DNA target, if the nucleic acid modifying enzyme of the encoding is active on the DNA target, the polynucleotide construct will be modified. Therefore, the state of the polynucleotide construct (modified or unmodified) is associated with the enzyme by the enzyme-specific sequence contained on the same construct. Wherein the polynucleotide construct comprises a DNA template encoding an RNA target, the RNA target will be included on the transcript RNA expressed from the DNA template. Since the RNA target and the nucleic acid modifying enzyme are continuously co-expressed as a single transcript, the state of the RNA target (modified or unmodified) is also associated with the enzyme by the enzyme-specific sequence contained on the RNA transcript.
[0132] Harvesting DNA / RNA molecules
[0133] To measure the activity of the nucleic acid modifying enzyme on the DNA / RNA target, the population of DNA / RNA molecules produced in step c) is harvested and then sequenced. In some instances, harvesting the DNA / RNA molecules requires disrupting the compartment. Therefore, in one embodiment of the method disclosed herein, step d) further comprises disrupting the compartment by physical or chemical means. In instances where the compartment is an emulsion droplet, harvesting the DNA / RNA molecules comprises disrupting the emulsion droplet.
[0134] Methods for disrupting emulsion droplets are known to those of ordinary skill in the art. A non-limiting example of this method is as follows: transfer the emulsion mixture to a 2 mL centrifuge tube and centrifuge at 13,000 g for 5 minutes at room temperature. Discard the upper oil layer. Add 1 mL of water-saturated diethyl ether to the remaining aqueous layer, vortex, and remove the upper solvent layer; repeat this step once. Centrifuge the remaining aqueous layer under vacuum for 5 minutes at room temperature. In one embodiment, the step of harvesting the DNA / RNA molecules also includes an IVTT quenching step. For example, the IVTT quenching step can be performed by treating the remaining aqueous layer with an RNase cocktail and proteinase K to remove excess RNA and protein from the IVTT reaction. In some embodiments, the step of harvesting the DNA / RNA molecules also includes a cleanup step to purify the DNA / RNA molecules. Methods for DNA / RNA cleanup are well known to those of ordinary skill in the art, and many commercial kits for this method are available, such as DNA Clean & Concentrator-5 (Zymo Research) or SPRIselect bead cleanup (Beckman Coulter).
[0135] In some instances, the harvesting of DNA / RNA molecules requires purification of the harvested DNA / RNA molecules to remove excess or unwanted DNA, RNA, and / or protein from the reaction. Thus, in one example of the methods disclosed herein, step d) further comprises purifying the harvested DNA / RNA molecules to remove excess DNA, RNA, and / or protein from the reaction. In some instances, the excess DNA, RNA, and / or protein may include, but is not limited to, gRNA, nucleic acid modifying enzymes, and IVTT reagents. In some instances, the term "excess" describes the molecules to be sequenced.
[0136] Sequencing
[0137] In a preferred embodiment, sequencing is single molecule sequencing." single molecule sequencing " refers to the technology that can read the base sequence directly from the single chain of DNA or RNA present in the sample. At least two types of single molecule sequencing are commercially available: (a) single molecule real-time sequencing (SMRT) of Pacific Biosciences, based on detecting and identifying fluorophore-labeled nucleotides with a waveguide (ZMW) smaller than the wavelength, and (b) a label-free sequencing method of the signal when an electronic device reads the nucleic acid (DNA / RNA) fragment through the nanopore used by Oxford Nanopore Technologies. Long reads contribute to single molecule sequencing, which may also be referred to as "long read sequencing" or "single molecule long read sequencing". The use of single molecule sequencing provides direct identification of variant sequences, which bypasses (i) oligonucleotides being connected to predetermined DNA / RNA ends, and (ii) PCR amplification.
[0138] "Directly" detecting the enzyme product of a single variant and performing molecular counting on modified: unmodified DNA / RNA targets (or polynucleotide constructs / RNA transcripts) to quantify the molecular activity of each variant is an important feature of the present invention. The term "directly" can refer to directly detecting the reaction product, or directly measuring the enzymatic activity of a single variant. In the latter meaning, the expression "direct measurement of enzymatic function" is based on the phenotypic activity (in a large-scale survey of variants) of the variant molecules directly calculated using the associated genotype information (also encoded within a single molecule (polynucleotide construct or RNA transcript)). Therefore, enzyme activity is measured directly on the molecules that actually interact. Based on the methods disclosed herein, the precise level of enzyme activity can be directly measured based on the count of modified and unmodified (or total). In one example, a specific variant associated with a 1:1 modified: unmodified target site was determined to be active on the target site within 50% of the time.
[0139] Thus, in some instances, the harvested population of DNA / RNA molecules is not further modified prior to the single molecule sequencing reaction, except for modifications required for single molecule sequencing. These modifications can include modifications required for conventional sequencing, such as ligation of cleaved ends to adapters, attachment of barcodes, amplification of DNA / RNA molecules by PCR, and the like.
[0140] In one example, sequencing is performed using an Oxford Nanopore Technologies platform. A non-limiting example of a sequencing process is described below.
[0141] Prepare purified DNA for long-read sequencing according to the library preparation protocol recommended by the sequencing device manufacturer (e.g., Oxford Nanopore Technologies (ONT) MinION Mk1B device) and sequence the library accordingly. In some examples, this may involve using the ONT SQK-LSK109 Sequencing by Ligation Kit for universal DNA library preparation, along with the ONT EXP-NBD104 PCR-Free Native Barcode Extension Kit to obtain multiplexed barcoded DNA sublibraries.
[0142] The sequences were aligned using bioinformatics tools available on public repositories (e.g., minimap2 (Li, H. (2018). Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics, 34: 3094-3100. doi: 10.1093 / bioinformatics / bty191), NanoPack (De Coster, W. et al., (2018). NanoPack: visualizing and processing long-read sequencing data. Bioinformatics, 34: 2666-2699. doi: 10.1093 / bioinformatics / bty149, samtools (Li, H. et al., (2009). The SequenceAlignment / Map format and SAMtools. Bioinformatics, 25: 2078-9. doi: 10.1093 / bioinformatics / btp352, VarScan 2 (Koboldt, DC et al., (2012). VarScan 2: Somatic mutation and copy number alteration discovery in cancer by exomesequencing. Genome Research, 22: 568-576. doi: 10.1101 / gr.129684.111)) or a custom script created by a person skilled in the art of sequencing data analysis to process and analyze long read sequencing data. For example, in some examples, a person skilled in the art of sequencing analysis can take the following steps to process and analyze raw nanopore sequencing reads generated from an ONT sequencing device:
[0143] 1. Use the guppy toolkit provided by ONT (https: / / community.nanoporetech.com / protocols / Guppy-protocol / v / gpb_2003_v1_revm_14dec2018) to process raw nanopore sequencing reads using the base calling and demultiplexing algorithms in the toolkit, if necessary. Those skilled in the art of sequencing analysis may wish to adjust certain parameters, such as the filtering threshold for barcode quality scores used for multiplexing, based on their needs. These parameters are generally described in the corresponding tool manuals.
[0144] 2. One of ordinary skill in the art of sequencing analysis may wish to use software tools such as NanoPack to further filter and process reads based on parameters such as read length and read quality score.
[0145] 3. These processed reads can then be aligned to a reference sequence (set) using minimap2 or other sequence alignment tools to generate a read alignment dataset. Similarly, those skilled in the art of sequencing analysis may wish to adjust read alignment parameters, such as the alignment scoring matrix, as needed. These parameters are typically described in the corresponding tool manual.
[0146] 4. The user can then parse the generated read alignment file to calculate the counts of unmodified and modified reads. In some embodiments, this can be done by using other alignment processing tools such as samtools or VarScan2 to detect and identify sequencing changes between the aligned sequencing reads and the reference sequence to which the sequencing reads are compared. Similarly, those skilled in the art of sequencing analysis can determine which parameters in these tools should be adjusted as needed, for example, setting a minimum read count threshold to detect and identify true sequencing changes from background levels of sequencing errors. These parameters are generally described in the corresponding tool manuals.
[0147] Molecular detection and counting
[0148] Thus, enzyme activity can be directly detected by detecting and counting polynucleotide constructs and / or RNA transcripts or fragments thereof that have been modified by the nucleic acid modifying enzyme and polynucleotide constructs and / or RNA transcripts that have not been modified by the nucleic acid modifying enzyme. The polynucleotide construct may comprise a DNA target, and the RNA transcript may comprise an RNA target.
[0149] Thus, in some examples, the method further comprises evaluating the modification activity of one or more nucleic acid modification enzymes on one or more DNA / RNA targets by counting the number of polynucleotide constructs and / or RNA transcripts modified by the nucleic acid modification enzyme (∑ count 修饰 ), and compare it to the number of polynucleotide constructs and / or RNA transcripts that have not been modified by the nucleic acid modifying enzyme (∑ count 未修饰 ) or the total number of polynucleotide constructs and / or RNA transcripts (∑ count 修饰+未修饰 ) for comparison.
[0150] In one example, enzyme activity is represented by a value calculated using any of the following formulas:
[0151] 1) Enzyme activity ≈ ∑ counts 修饰 / ∑ count 未修饰
[0152] 2) Enzyme activity ≈ ∑ counts 修饰 / ∑ count 修饰+未修饰 .
[0153] Polynucleotide constructs and / or RNA transcripts or fragments thereof that have or have not been modified by a nucleic acid modification enzyme can be detected and counted using sequencing data generated by sequencing platforms available to those skilled in the art. Since DNA / RNA molecules are directly sequenced by single-molecule sequencing, in one embodiment, the detection and counting of DNA / RNA molecules that have or have not been modified by a nucleic acid modification enzyme is based solely on the data generated during single-molecule sequencing, and no further modification or processing of the DNA / RNA molecules is required.
[0154] In one example, wherein the modifying activity is a cleavage activity, and the detection and enumeration of modified and unmodified polynucleotide constructs or RNA targets is achieved by aligning sequencing reads of the DNA / RNA molecules to a reference sequence that includes a cleavage site window for the nucleic acid modifying enzyme, wherein
[0155] i) when the 3' end of the DNA / RNA molecule is mapped to the region 3' downstream of the cleavage site window, the DNA / RNA molecule is an unmodified polynucleotide construct or RNA target;
[0156] ii) when the 3' end of the DNA / RNA molecule maps to a region within the cleavage site window, the DNA / RNA molecule is a modified polynucleotide construct or RNA target;
[0157] iii) When the 3' end of a DNA / RNA molecule maps to a region 5' upstream of the cleavage site window, the DNA / RNA molecule is non-informative and is not used to measure modification activity.
[0158] In one example, sequenced reads are identified by their respective mapping endpoints (i.e., the location where the sequenced read ends) to determine whether the endpoint is within a small window of the expected cleavage site (grey triangle and dashed line on the "Cas Reference Sequence"); Figure 3 ). A non-limiting example is described below, where the DNA target site is located 3' of the encoded Cas nuclease variant. In this way, reads with alignments that map to the 3' end of the reference sequence at a site 3' downstream of the expected Cas cleavage site window (aligned to the reference sequence) are considered uncleaved (dark grey; Figure 3 ), while aligned reads with 3′ ends located within the Cas cleavage site window were considered cleaved (light grey; Figure 3 ), and reads that do not meet either criterion are ultimately discarded as uninformative reads, because it is impossible to empirically determine whether these reads are cleaved (white; Figure 3). In these examples, each cleaved Cas cleavage site represents a modified polynucleotide construct / RNA transcript. Similarly, each uncleaved Cas cleavage site represents an unmodified polynucleotide construct / RNA transcript.
[0159] In some examples, sequencing technologies can detect or determine the chemical and sequence identity of a target site to determine whether the target is modified by a Cas variant. For example, chemical modifications of nucleotides, such as methylation, can be detected using publicly available bioinformatics tools designed to pick out chemically modified nucleotides in nanopore sequencing reads (Liu, Q. et al. (2019). Detection of DNA base modifications by deep recurrent neural network on Oxford Nanopore sequencing data. Nat Commun 10(1):2449, doi:10.1038 / s41467-019-10168-2; Liu, Q. et al. (2019). NanoMod: a computational tool to detect DNA modifications using Nanopore long-read sequencing data. BMC Genomics 20(Suppl 1):78, doi:10.1186 / s12864-018-5372-8; Rand, AC et al. (2017). Mapping DNA methylation with high-throughput nanopore sequencing.) Nat Methods 14(4):411-413, doi:10.1038 / nmeth.4189; Simpson, JT et al. (2017). Detecting DNA cytosine methylation using nanopore sequencing. Nat Methods 14(4):407-410, doi:10.1038 / nmeth.4184). In other examples of encoded Cas nuclease targeting RNA constructs, RNA molecules can be harvested and purified using commercial kits such as RNA Clean & Concentrator-5 (Zymo Research) after the IVTT reaction, while Oxford Nanopore Technologies uses its commercially available SQK-RNA002 RNA direct sequencing kit to perform direct nanopore sequencing on the harvested RNA molecules. Therefore, sequencing technology can be used to detect various types of modifications, including but not limited to strand breaks, sequence changes, and epigenetic biochemical markers.
[0160] Polynucleotide constructs and libraries
[0161] The present disclosure also relates to various polynucleotide constructs, libraries of constructs, and compartments (for some examples of polynucleotide constructs, see e.g., Figure 2 ).
[0162] Construct 1: In one aspect, the present disclosure relates to a polynucleotide construct comprising: a first polynucleotide sequence encoding a nucleic acid modification enzyme or a variant thereof operably linked to a first promoter; and a second polynucleotide sequence comprising a DNA target.
[0163] Construct 2: In another aspect, the present disclosure relates to a polynucleotide construct comprising: a first polynucleotide sequence encoding a nucleic acid modification enzyme or a variant thereof operably linked to a first promoter; and a second polynucleotide sequence comprising a DNA template encoding an RNA target; and wherein the RNA target and the nucleic acid modification enzyme are continuously co-expressed as a single RNA transcript driven by the first promoter.
[0164] In another aspect, the present disclosure relates to a construct library comprising a plurality of polynucleotide constructs as disclosed herein as Construct 1 or Construct 2, wherein the library is characterized by one or more of the following:
[0165] a. the plurality of polynucleotide constructs encoding different variants of nucleic acid modification enzymes;
[0166] b. The multiple polynucleotide constructs encode different DNA or RNA targets.
[0167] Construct 3: In one example, the present disclosure further relates to a polynucleotide construct according to construct 1 or construct 2, wherein the polynucleotide construct further comprises a third polynucleotide sequence encoding a guide RNA (gRNA). The encoded gRNA will be expressed (e.g., in a compartment) from the polynucleotide construct, so the polynucleotide construct may include other elements that promote gRNA expression, which are generally known to those skilled in the art. In some instances, the third polynucleotide sequence is operably linked to a second promoter. In some instances, the second promoter is a T7 promoter.
[0168] In another aspect, the present disclosure relates to a construct library comprising a plurality of polynucleotide constructs as disclosed herein as Construct 3, wherein the library is characterized by one or more of the following:
[0169] a. the plurality of polynucleotide constructs encoding different variants of nucleic acid modification enzymes;
[0170] b. the plurality of polynucleotide constructs encoding different DNA or RNA targets;
[0171] c. The multiple polynucleotide constructs encode different gRNAs.
[0172] Construct 4: In yet another aspect, the present disclosure relates to a polynucleotide construct comprising: a first polynucleotide sequence encoding a guide RNA (gRNA) operably linked to a first promoter; and a second polynucleotide sequence comprising a DNA target.
[0173] Construct 5: In yet another aspect, the present disclosure relates to a polynucleotide construct comprising: a first polynucleotide sequence encoding a guide RNA (gRNA) operably linked to a first promoter; and a second polynucleotide sequence comprising a DNA template encoding an RNA target; wherein the RNA target and the gRNA are continuously co-expressed as a single RNA transcript driven by the first promoter.
[0174] In another aspect, the present disclosure relates to a construct library comprising a plurality of polynucleotide constructs as disclosed herein as construct 4, wherein the library is characterized by one or more of the following:
[0175] a. the multiple polynucleotide constructs encode different DNA or RNA targets;
[0176] b. The multiple polynucleotide constructs encode different gRNAs.
[0177] In some examples of the methods or polynucleotide constructs as disclosed herein, the first and second polynucleotide sequences overlap completely or in part. For example, a DNA / RNA target ("second polynucleotide") can be encoded within the coding sequence ("first polynucleotide") of a nucleic acid modifying enzyme.
[0178] In some examples of the methods or polynucleotide constructs disclosed herein, the DNA or RNA target comprises a protospacer that is at least partially complementary to the guide RNA. In some examples of the methods or polynucleotide constructs disclosed herein, the DNA target further comprises a proximal protospacer adjacent motif (PAM) sequence. In some examples of the methods or polynucleotide constructs disclosed herein, when the polynucleotide construct comprises a DNA template encoding an RNA target, the RNA target further comprises a proximal protospacer flanking sequence (PFS).
[0179] In some examples of the methods or polynucleotide constructs disclosed herein, the RNA-guided nucleic acid modification enzyme is a CRISPR-associated (Cas) protein. In a specific example, the RNA-guided nucleic acid modification enzyme is selected from the group consisting of Cas3, Cas9, Cas10, Cas12a (also known as Cpf1), Cas13a (also known as C2c2), Cas13b, Cas13c, Cas13d, Cas14, CasX, CasΦ, and variants thereof.
[0180] In some examples of the methods or polynucleotide constructs as disclosed herein, the variant nucleic acid modification enzyme comprises one or more inactivated catalytic sites and is capable of binding to and inhibiting expression of a DNA target without modifying the DNA target.
[0181] In some examples of the methods or polynucleotide constructs disclosed herein, the variant nucleic acid modification enzyme is fused to one or more additional functional domains capable of modifying DNA or RNA. In some specific examples, the additional functional domains include, but are not limited to, a cytidine deaminase domain, a de novo DNA methyltransferase 3A (DNMT3A) domain, a cytosine-5 methyltransferase domain, a ten-eleven translocation dioxygenase 1 (TET1) catalytic domain, an adenosine deaminase acting on RNA (ADAR2) deaminase domain, and a DNA deoxyadenosine deaminase domain.
[0182] For illustrative and exemplary purposes, the sequences of the polynucleotide constructs are provided below.
[0183]
[0184]
[0185]
[0186]
[0187] Bold and underlined sequences refer to elements that will be individually annotated below:
[0188] T7 / lacO promoter:
[0189]
[0190] RBS (ribosome binding site): AAGGAG (SEQ ID NO: 2)
[0191] Sq Cas9 gene (coding sequence):
[0192]
[0193]
[0194]
[0195] Synthetic terminator sequence (L3S1P52):
[0196]
[0197] T7 promoter: TAATACGACTCACTATAG (SEQ ID NO: 5)
[0198] gRNA target sequence (protospacer): TCTGACAGCAGACGTGCACTGGCCAG (SEQ ID NO: 6)
[0199] SpCas9 gRNA scaffold:
[0200]
[0201] T7 terminator:
[0202]
[0203] Target region (DNA target):
[0204]
[0205] In some examples, such as the one exemplified above, a protospacer adjacent motif (PAM) is present next to the target sequence (also known as a protospacer), for example: 5' PAM site TTTV (SEQ ID NO: 11) for Cpf1-type (also known as Cas12) Cas proteins, 3' PAM site NGG for Sp Cas9, and 3' PAM site NNGRRT (SEQ ID NO: 12) for Sa Cas9 proteins, flanked by TCTGACAGCAGACGTGCACTGGCCAG (SEQ ID NO: 6) protospacer sequences. Standard IUPAC nucleic acid symbols are used herein and throughout the specification.
[0206] compartment
[0207] In one aspect, the present disclosure relates to one or more compartments, each compartment comprising a polynucleotide construct as disclosed herein, wherein the compartments are separated from each other. In some instances, each compartment further comprises an in vitro transcription and translation (IVTT) reagent capable of achieving in vitro transcription and / or translation of protein and / or RNA. In some instances, the volume of the compartment is less than 1000 μm 3 , 100μm 3 , 10μm 3 or 1 μm 3 In some instances, the compartments are water-in-oil emulsion droplets. In some instances, separation is achieved using microfluidics, hydrogel restricted diffusion, or partitioned wells. Example
[0208] Example 1: IVTT and cleavage of SpCas9 constructs in emulsion droplets
[0209] Water-in-oil (w / o) emulsion droplets were generated following the steps outlined in the protocol provided above. Briefly, 950 μL of an oil-surfactant mixture (mineral oil + 4.5% (v / v) Span 80 + 0.5% (v / v) Tween 80) was added to a cryovial with a 3×8 mm magnetic stir bar and placed on ice.
[0210] High DNA input: >1 sequence copy encapsulated per emulsion droplet—demonstrating that emulsification preserves Cas activity, as seen with Cas expressed from bulk IVTT reactions.
[0211] For this experiment, about 750ng of Sp Cas9 construct (sequence as shown above) was mixed with IVTT reagent (New England Biolabs PURExpress # E6800) on ice to produce 75 μL IVTT aqueous mixture. Within 2 minutes, 50 μL of this aqueous mixture was added to the oil-surfactant mixture on ice in 5 portions of 10 μL, while the stirring bar was rotated at 1150 rpm to produce an emulsion mixture. The emulsion mixture was allowed to mix on ice for another minute. The emulsion mixture was then homogenized (8000 rpm for 3 minutes; IKA Ultraturrax T10 homogenizer) to produce a more monodisperse distribution of emulsion droplet size. The remaining 25 μL aqueous mixture was kept on ice and a large volume IVTT reaction was performed as a control. This operation was also repeated for the Sp dCas9 construct.
[0212] The emulsion and large-volume IVTT mixture were then incubated at 37°C for 4 h for IVTT and then at 65°C for 15 min to inactivate the protein.
[0213] The emulsion IVTT mixture was then treated as described above to break the emulsion. 20 mM EDTA (pH 8.0) inhibitor was added to the emulsion and mixed briefly by vortexing. The emulsion mixture was then centrifuged at 13,000 g for 5 minutes at room temperature. The upper oil layer was removed. 1 mL of water-saturated diethyl ether was added to the remaining aqueous layer, vortexed, and the upper solvent layer was removed; this step was repeated once. The remaining aqueous layer was vacuum centrifuged at room temperature for 5 minutes and then treated with RNase mix and proteinase K at 37°C for 30 minutes to remove excess RNA and protein from the IVTT reaction. The large volume IVTT reaction system was also treated with 20 mM EDTA (pH 8.0) and a mixture of RNase mix and proteinase K at 37°C for 30 minutes to remove excess RNA and protein from the IVTT reaction. DNA from all IVTT reactions was then individually purified using SPRIselect paramagnetic beads, and aliquots were visualized after size sorting by gel electrophoresis on agarose gels ( Figure 4). In the emulsion IVTT reaction, the construct encoding active Sp Cas9 was cleaved (smaller bands were present), while the construct encoding inactive Sp dCas9 was not cleaved (smaller bands were absent). This shows that the CRISPR-Cas IVTT self-cleavage assay is effective regardless of whether the encoding DNA construct is compartmentalized in emulsion droplets or freely floating in a bulk solution.
[0214] Purified DNA from the emulsion IVTT reaction was also processed for nanopore sequencing using the commercially available SQK-LSK109 ligation sequencing kit from Oxford Nanopore Technologies (ONT) and barcoded using the ONT EXP-NBD104 PCR-Free amplification-free barcode extension kit so that they could be optionally multiplexed in a single pooled DNA library. The pooled DNA library was then subjected to single-molecule long-read nanopore sequencing using the ONT MinION Mk1B sequencing device. The nanopore sequencing results were then quality filtered and analyzed using publicly available bioinformatics tools. Sp Cas9 emulsion IVTT nanopore sequencing reads appear as a mixture of detected cleaved and uncleaved construct fragments ( Figure 5 The vast majority of Sp dCas9 emulsion IVTT nanopore sequencing reads appeared as uncleaved construct fragments, as expected ( Figure 6 ); The very few reads of the Sp dCas9 construct fragments classified as "cleaved" may be the result of truncated / incomplete reads and / or random DNA shearing events and / or errors on the sequencing device during nanopore sequencing. Some reads in each sublibrary were mapped to the wrong sequence, for example, for the Sp Cas9-only sublibrary, reads were mapped to Sp dCas9 instead of Sp Cas9. These may be the result of random sequencing errors on the sequencing device or incorrectly assigned to their respective sublibraries during the demultiplexing process of the barcoded nanopore sequencing reads; therefore, these reads are classified as misassigned and depicted as such in the figure.
[0215] Example 2: Quantification of CRISPR-Cas cleavage activity by multiplexed single-molecule long-read sequencing of DNA constructs after large-volume IVTT reactions
[0216] Large volume IVTT reactions were set up on ice for different CRISPR-Cas constructs (Sp Cas9, Sa Cas9, As Cpf1, Lb Cpf1) with components arranged similarly to those described for the nucleic acid template sequences above. They were then divided equally into five corresponding aliquots for each time point ( Figure 7 Part 1). Aliquots of these large-volume IVTTs were then incubated at 37°C and removed at each designated time point, quenched with EDTA inhibitor and enzyme to stop the IVTT reaction and Cas cleavage of the encoding DNA construct ( Figure 7 Part 2). The quenched IVTT reaction was then treated with SPRIselect bead cleanup to purify the DNA fragments ( Figure 7 Part 3).
[0217] Small aliquots of these DNA fragments from different Cas orthologs at different IVTT time points were then visualized after size sorting by gel electrophoresis on agarose gels, e.g. Figure 8 shown.
[0218] The remaining aliquots of purified DNA fragments were then pooled according to their respective time points, but regardless of the Cas species, i.e., DNA fragments of Sp Cas9, Sa Cas9, etc. were mixed together at each time point and barcoded individually using the ONT EXP-NBD104 PCR-Free Amplification-Free Barcoding Extension Kit ( Figure 7 Part 4) to multiplex these pooled sublibraries for a single nanopore sequencing run ( Figure 7 Part 5). The nanopore sequencing results were then quality filtered and analyzed using publicly available bioinformatics tools and then the analysis methods disclosed in this invention.
[0219] Figure 9 The counts of each active Cas construct encoding cleaved DNA fragments are normalized to the total counts of each Cas construct encoding cleaved and uncleaved DNA fragments at 5 selected time points of IVTT incubation (between 0 and 4 hours). As the IVTT incubation time increases, the expressed Cas protein has more time to cleave more encoding DNA constructs, resulting in a higher incidence of each type of cleaved fragment at later time points. Figure 9 The nanopore sequencing analysis results shown in Figure 8 The gel images of the purified IVTT DNA fragments in the assay were qualitatively consistent, and both assays had the same purified DNA input, which was obtained from Figure 7 obtained during the workflow steps outlined in Section 3. This example demonstrates the concept of our workflow for studying nucleic acid products from separate IVTT reactions of multiple CRISPR-Cas self-cleavage assays.
[0220] Example 3: Demonstration of the sensitivity of nanopore sequencing assays by titrating the ratio of purified CRISPR-Cas DNA final products from large-volume IVTT reactions
[0221] For this experiment, 500 ng of Sp Cas9 (sequence shown above) was mixed with IVTT reagent (New England Biolabs PURExpress #E6800) on ice to produce a 50 μL IVTT aqueous mixture. The same operation was performed on the Sp dCas9 construct; the Sp dCas9 construct contains a DNA sequence essentially identical to that of the Sp Cas9 construct, except that two inactivating mutations (D10A and H840A) in the Sp Cas9 gene are present to produce the Sp dCas9 gene. These 50 μL large-volume IVTT reactions were incubated at 37°C for 4 hours to perform IVTT, followed by incubation at 65°C for 15 minutes to inactivate the protein. 20 mM EDTA (pH 8.0) inhibitor was added to the large-volume IVTT reaction with RNase mix and proteinase K for 30 minutes at 37°C to remove excess RNA and protein from the IVTT reaction. DNA from these two large-volume IVTT reactions was then individually purified using SPRIselect paramagnetic beads, and aliquots of the DNA were visualized after size sorting by gel electrophoresis on agarose gels, e.g. Figure 10 shown.
[0222] The concentrations of purified DNA from large-volume IVTT reactions of Sp dCas9 and Sp Cas9 were quantified and then mixed at the following mass ratios: 1:1, 1:10 -1 , 1:10 -2 , 1:10 -3 , 1:10 -4 , 1:10 -5 , 1:0. These mixtures, at seven ratios, with titrated purified Sp dCas9 and Sp Cas9 bulk IVTT DNA products were then processed for nanopore sequencing using the ONT SQK-LSK109 Ligation Sequencing Kit, while each of the seven mixtures was individually barcoded using the ONT EXP-NBD104 PCR-Free Amplification-Free Barcoding Extension Kit. The DNA libraries were then subjected to single-molecule long-read nanopore sequencing using the ONT MinION Mk1B sequencing device. The nanopore sequencing results were then quality filtered and analyzed using publicly available bioinformatics tools, followed by the analytical methods disclosed in this invention.
[0223] The purpose of this assay is to evaluate the sensitivity of nanopore sequencing assays for large-scale surveys of DNA / RNA modification events, a capability claimed in our invention. Specifically, in this case, self-cleavage events of the Sp Cas9 IVTT construct were titrated against non-cleavage of the Sp dCas9 IVTT construct. Using a combination of the above-described bioinformatics approaches, the inventors demonstrated the detection of cleaved and uncleaved Sp Cas9 DNA fragments, which could be distinguished from the detection of uncleaved Sp dCas9 DNA fragments in the raw nanopore sequencing data. Remarkably, the inventors were even able to detect cleaved and uncleaved Sp Cas9 DNA fragments at a 1:10 ratio of purified Sp dCas9 and Sp Cas9 bulk IVTT DNA products, respectively. -5 The presence of cleaved Sp Cas9 DNA fragments was detected in the mixture ( Figure 11 ).
[0224] Example 4: IVTT and cleavage of SpCas9 constructs in emulsion droplets
[0225] Limit DNA input: encapsulate ≤1 sequence copy per emulsion droplet - measure the efficiency of emulsifying a single copy of the DNA construct.
[0226] In this experiment, ≤1.66fmol of Sp Cas9 construct (sequence as shown above) was mixed with IVTT reagent (New England Biolabs PURExpress # E6800) on ice to produce 50 μL IVTT aqueous mixture. Within 2 minutes, 50 μL of this aqueous mixture was added to the oil-surfactant mixture on ice in 5 portions of 10 μL, while the stirring bar was rotated at 1150 rpm to produce an emulsion mixture. The emulsion mixture was allowed to continue mixing on ice for another minute. The emulsion mixture was then homogenized (8000 rpm for 3 minutes; IKA Ultraturrax T10 homogenizer) to produce a more monodisperse distribution of emulsion droplet size. This operation was repeated for Sp dCas9 constructs and 1: 1 equimolar mixtures of Sp Cas9 and Sp dCas9 constructs.
[0227] Note that the efficiency of encapsulating ≤1 DNA construct per emulsion droplet was measured using a mixture of Sp Cas9 and Sp dCas9 DNA constructs. Under perfect efficiency, where only ≤1 DNA construct is encapsulated per droplet, none of the Sp dCas9 sequences detected by nanopore sequencing should be cleaved at the end of the assay for mixed DNA input conditions. Under imperfect efficiency, some Sp dCas9 DNA constructs may be cleaved because some constructs are exposed to active Sp Cas9 in the same droplet. Therefore, if the detection rate of cleaved Sp dCas9 constructs in an assay with mixed Sp Cas9 and Sp dCas9 constructs occurs at a very low rate comparable to the random sequencing error rate of long-read nanopore sequencing, the data would indicate that under these conditions, ≤1 sequence copy is encapsulated per emulsion droplet. This example demonstrates the complete workflow of our invention, as shown in Figure 2. Figure 1 shown.
[0228] The resulting emulsion IVTT mixture was then incubated at 37°C for 4 h to perform IVTT and then at 65°C for 15 min to inactivate the protein.
[0229] The emulsion IVTT mixture was then treated as described above to break the emulsion. 20 mM EDTA (pH 8.0) inhibitor was added to the emulsion and mixed briefly by vortexing. The emulsion mixture was then centrifuged at 13,000 g for 5 minutes at room temperature. The upper oil layer was removed. 1 mL of water-saturated diethyl ether was added to the remaining aqueous layer, vortexed and the upper solvent layer was removed; this step was repeated once. The remaining aqueous layer was vacuum centrifuged for 5 minutes at room temperature and then treated with RNase cocktail and proteinase K at 37°C for 30 minutes to remove excess RNA and protein from the IVTT reaction. DNA from all IVTT reactions was then individually purified using a commercial column purification kit (DNA Clean and Concentrator-5, Zymo Research) according to the manufacturer's instructions.
[0230] Purified DNA from the IVTT reaction was then processed for nanopore sequencing using the ONT SQK-LSK109 Sequencing by Ligation Kit and barcoded using the ONT EXP-NBD104 PCR-Free Amplification-Free Barcode Extension Kit. The DNA library was then subjected to single-molecule, long-read nanopore sequencing using the ONT MinION Mk1B sequencing device. The nanopore sequencing results were then quality-filtered and analyzed using publicly available bioinformatics tools, followed by the analytical methods disclosed herein.
[0231] Sp Cas9 emulsion IVTT nanopore sequencing reads showing a mixture of cleaved and uncleaved construct fragments detected ( Figure 12 ), indicating that Sp Cas9 is active against a small fraction of targets (as demonstrated in the bulk reaction). The vast majority of SpdCas9 emulsion IVTT nanopore sequencing reads appeared as uncleaved construct fragments, demonstrating that Sp dCas9 is inactive the vast majority of the time, as expected ( Figure 13 ); The very few reads of the Sp dCas9 construct fragment classified as “cleaved” may be the result of truncated / incomplete reads and / or random DNA shearing events during nanopore sequencing.
[0232] Note that some reads in the Sp Cas9-only and Sp dCas9-only sublibraries were mapped to incorrect sequences, e.g., in the Sp Cas9-only sublibrary, reads mapped to Sp dCas9 instead of Sp Cas9, which could be the result of sequencing errors on the sequencing device or demultiplexing errors of the barcoded nanopore sequencing reads and are therefore classified as misassignments and depicted as such in the figure.
[0233] Nanopore sequencing reads generated from emulsion IVTT reactions of a 1:1 mixture of Sp Cas9 and Sp dCas9 constructs added at limiting concentrations showed a roughly equal distribution of Sp Cas9 and Sp dCas9 mapped reads, as expected ( Figure 14 ). Sp Cas9 mapped reads showed an almost equal division of cleaved and uncleaved fragments, while the majority of Sp dCas9 mapped reads were classified as uncleaved. As may occur due to sequencing or demultiplexing errors, a minority of Sp dCas9 mapped reads were classified as cleaved, as these sequencing errors are known to occur on sequencing devices when sequencing a mixture of fragments or through errors of cross-contamination of the enzyme complex, which can be further reduced by technical optimization within the inventive concept. In summary, this example embodies and demonstrates the disclosed invention, in which the level of enzymatic activity of variants can be directly counted and determined based on single molecules. Sequence Listing <110> Agency for Science, Technology and Research (A*STAR) <120> Assays to measure enzyme activity <130> 9869SG5983 <160> 12 <170> PatentIn Version 3.5 <210> 1 <211> 44 <212> DNA <213> Artificial sequence <220> <223> T7 / lacO promoter <400> 1 taatacgact cactataggg gaattgtgag cggataacaa ttcc 44 <210> 2 <211> 6 <212> DNA <213> Artificial sequence <220> <223> RBS (ribosome binding site) <400> 2 aaggag 6 <210> 3 <211> 4137 <212> DNA <213> Artificial sequence <220> <223> Sp Cas9 coding sequence <400> 3 atggacaaga agtactccat tgggctcgat atcggcacaa acagcgtcgg ctgggccgtc 60 attacggacg agtacaaggt gccgagcaaa aaattcaaag ttctgggcaa taccgatcgc 120 cacagcataa agaagaacct cattggcgcc ctcctgttcg actccgggga aacggccgaa 180 gccacgcggc tcaaaagaac agcacggcgc agatataccc gcagaaagaa tcggatctgc 240 tacctgcagg agatctttag taatgagatg gctaaggtgg atgactcttt cttccatagg 300 ctggaggagt cctttttggt ggaggaggat aaaaagcacg agcgccaccc aatctttggc 360 420 aagcttgtag acagtactga taaggctgac ttgcggttga tctatctcgc gctggcgcat 480 atgatcaaat ttcggggca cttcctcatc gaggggacc tgaacccaga caacagcgat 540 gtcgacaaac tctttatcca actggttcag acttacaatc agcttttcga agagaacccg 600 atcaacgcat ccggagttga cgccaaagca atcctgagcg ctaggctgtc caaatcccgg 660 cggctcgaaa acctcatcgc acagctccct gggggagaagaacggcct gtttggtaat 720 cttatcgccc tgtcactcgg gctgaccccc aactttaaat ctaacttcga cctggccgaa 780 gatgccaagc ttcaactgag caaagacacc tacgatgatg atctcgacaa tctgctggcc 840 cagatcggcg accagtacgc agaccttttt ttggcggcaa agaacctgtc agacgcatt 900 ctgctgagtg atattctgcg agtgaaccg gagatcacca aagctccgct gagcgctagt 960 atgatcaagc gctatgatga gcaccaccaa gacttgactt tgctgaaggc ccttgtcaga 1020 cagcaactgc ctgaagaagta caaggaaatt ttcttcgatc agtctaaaaa tggctacgcc 1080 ggatacattg acggcggagc aagccaggag gaattttaca aatttattaa gcccatcttg 1140 gaaaaaatgg acggcaccga ggagctgctg gtaaagctta acagagaaga tctgttgcgc 1200 aaacagcgca ctttcgacaa tggaagcatc ccccaccaga ttcacctggg cgaactgcac 1260 gctatcctca ggcggcaaga ggatttctac cccttttga aagataacag ggaaaagatt 1320 gagaaaatcc tcacatttcg gataccctac tatgtaggcc ccctcgcccg gggaaattcc 1380 agattcgcgt ggatgactcg caaatcagaa gagaccatca ctccctggaa cttcgaggaa 1440 gtcgtggata aggggcctc tgcccagtcc ttcatcgaaa ggatgactaa ctttgataaa 1500 aatctgccta acgaaaaggt gcttcctaaa cactctctgc tgtacgagta cttcacagtt 1560 tataacgagc tcaccaaggt caaatacgtc acagaaggga tgagaaagcc agcattcctg 1620 tctggagagc agaagaaagc tatcgtggac ctcctcttca agacgaaccg gaaagttacc 1680 gtgaaacagc tcaaagaaga ctatttcaaa aagattgaat gtttcgactc tgttgaaatc 1740 agcggagtgg aggatcgctt caacgcatcc ctgggaacgt atcacgatct cctgaaaatc 1800 attaagaca aggactcct ggacaatgag gagaacgagg acatcttga ggacattgtc 1860 ctcaccctta cgttgtttga agatagggag atgattgaag aacgcttga aacttacgct 1920 catctcttcg acgacaagt catgaacag ctcaagaggc gccgatatac aggatggggg 1980 cggctgtcaa gaaactgat caatgggatc cgagacaagc agagtggaaa gatacctg 2040 gattttctta agtccgatgg atttgccac cggaacttca tgcagttgat ccatgatgac 2100 tctctcacct ttaggagga catccagaaa gcacaagttt ctggccaggg ggacagtctt 2160 cacgagcaca tcgctaatct tgcaggtagc ccagctatca aaaagggaat actgcagacc 2220 gttaaggtcg tggatgaact cgtcaaagta atgggaggc ataagcccga gatatcgtt 2280 atcgagatgg cccgagagaa ccaactacc cagaagggac agagacag tagggaagg 2340 atgaagagga ttgaagagggg tataaaagaa ctggggtccc aaatccttaa ggacaccca 2400 gttgaaaaca cccagcttca gatgagaag cttacctgt actacctgca gaacggcagg 2460 gatagtacg tggatcagga actggacatc aatcggctct ccgactacga cgtggatcat 2520 atcgtgcccc agtcttttct caaagatgat tctattgata ataaagtgtt gacaagatcc 2580 gataaaaata gagggaagag tgataacgtc ccctcagaag aagttgtcaa gaaaatgaaa 2640 aattattggc ggcagctgct gaacgccaaa ctgatcacac aacggaagtt cgataatctg 2700 actaaggctg aacgaggtgg cctgtctgag ttggataaag ccggcttcat caaaaggcag 2760 cttgttgaga cacgccagat caccaagcac gtggcccaaa ttctcgattc acgcatgaac 2820 accaagtacg atgaaaatga caaactgatt cgagaggtga aagttattac tctgaagtct 2880 aagctggtct cagatttcag aaaggacttt cagttttata aggtgagaga gatcaacaat 2940 taccaccatg cgcatgatgc ctacctgaat gcagtggtag gcactgcact tatcaaaaaa 3000 tatcccaagc ttgaatctga atttgtttac ggagactata aagtgtacga tgttaggaaa 3060 atgatcgcaa agtctgagca ggaaataggc aaggccaccg ctaagtactt cttttacagc 3120 aatattatga attttttcaa gaccgagatt acactggcca atggagagat tcggaagcga 3180 ccacttatcg aaacaaacgg agaaacagga gaaatcgtgt gggacaaggg tagggatttc 3240 gcgacagtcc ggaggtcct gtccatgccg caggtgaaca tcgttaaaaa gaccgaagta 3300 cagaccggag gctctccaa ggaagtatc ctcccgaaaa ggaacagcga caagctgatc 3360 gcacgcaaaa aagattggga cccaagaaa tacggcggat tcgattctcc tacagtcgct 3420 tacagtgtac tggttgtggc aaagtggag aaagggaagt ctaaaaaact caaagcgtc 3480 aaggaactgc tgggcatcac atcatggag cgatcaagct tcgaaaaaaa cccatcgac 3540 ttctcgagg cgaaggata taaagaggtc aaaaaagacc tcatcattaa gctcccaag 3600 tactctctct ttgagcttga aaacggccgg aaacgaatgc tcgctgc gggcgagctg 3660 cagaaaggta acgagctggc actgccctct aaatacgtta atttctgta tctggccagc 3720 cactatgaaa agctcaagg gtctcccgaa gataatgagc agaagcagct gttcgtggaa 3780 storm actaccttga tgagatcatc gagaataa gcgaattctc siaagagtg 3840 atcctcgccg acgctaacct cgataaggtg ctttctgctt acataagca caggtaag 3900 cccatcaggg agcaggcaga aaacattatc cacttgttta ctctgaccaa cttggggcgcg 3960 cctgcagcct tcaagtactt cgacaccacc atagacagaa agcggtacac ctctacaaag 4020 gaggtcctgg acgccacact gattcatcag tcaattacgg ggctctatga aacaagaatc 4080 gacctctctc agctcggtgg agacagcagg gctgacccca agaagaagag gaaggtg 4137 <210> 4 <211> 52 <212> DNA <213> Artificial sequence <220> <223> Synthetic terminator sequence (L3S1P52) <400> 4 tctaactaaa aaggcctccc aaatcggggg gcctttttta ttgataacaa aa 52 <210> 5 <211> 18 <212> DNA <213> Artificial sequence <220> <223> T7 promoter <400> 5 taatacgact cactatag 18 <210> 6 <211> 26 <212> DNA <213> Artificial sequence <220> <223> gRNA target sequence (protospacer) <400> 6 tctgacagca gacgtgcact ggccag 26 <210> 7 <211> 77 <212> DNA <213> Artificial sequence <220> <223> SpCas9 gRNA scaffold <400> 7 gttttagagc tagaaatagc aagttaaaat aaggctagtc cgttatcaac ttgaaaaagt 60 ggcaccgagt cggtgct 77 <210> 8 <211> 48 <212> DNA <213> Artificial sequence <220> <223> T7 terminator <400> 8 ctagcataac cccttggggc ctctaaacgg gtcttgaggg gttttttg 48 <210> 9 <211> 36 <212> DNA <213> Artificial sequence <220> <223> Target region (DNA target): <400> 9 tttatctgac agcagacgtg cactggccag ggggat 36 <210> 10 <211> 6416 <212> DNA <213> Artificial sequence <220> <223> Sequences of exemplary polynucleotide constructs <400> 10 gcgacatcgt ataacgttac tggtttcaca ttcaccaccc tgaattgact ctcttccggg 60 cgctatcatg ccataccgcg aaaggttttg cgccattcga tggtgtccgg gatctcgacg 120 ctctccctta tgcgactcct gcattaggaa attaatacga ctcactatag gggaattgtg 180 agcggataac aattcccctg tagaataat ttgtttaac tttaatagg agatatacca 240 tggacaagaa gtactccatt gggctcgata tcggcacaa cagcgtcggc tggccgtca 300 ttacggacga gtacaagtg ccgagcaaa aattcaagt tctgggcaat accgatcgcc 360 acagcataaa gaagaacctc attggcgccc tcctgttcga ctccggggaa acggccgaag 420 ccacgcggct caaagaaca gcacggcgca gatatacccg cagaaagaat cggatctgct 480 acctgcagga gatctttagt atgagatgg ctaggtgga tgactttc ttccataggc 540 tggaggagtc ctttttgtg gaggaggata aaagcacga gcgccaccca atctttggca 600 atatcgtgga cgaggtggcg taccatgaaa agtacccaac catatatcat ctgaggaaga 660 agcttgtaga cagtactgat aaggctgact tgcggttgat ctatctcgcg ctggcgcata 720 tgatcaaatt tcgggacac ttcctcatcg aggggacct gaacccagac aacagcgatg 780 tcgacaact ctttatccaa ctggttcaga cttacaatca gctttcgaa gagaacccga 840 tcaacgcatc cggagttgac gccaaagcaa tcctgagcgc taggctgtcc aaatcccggc 900 ggctcgaaaa cctcatcgca cagctccctg gggagaagaa gaacggcctg tttggtaatc 960 ttatcgccct gtcactcggg ctgaccccca actttaaatc taacttgac ctggccgaag 1020 atgccaagct tcaactgagc aaagacacct acgatgatga tctcgacaat ctgctggccc 1080 agatcggcga ccagtacgca gacctttttt tggcggcaaa gaacctgtca gacgccattc 1140 tgctgagtga tattctgcga gtgaacacgg agatcaccaa agctccgctg agcgctagta 1200 tgatcaagcg ctatgatgag caccaccaag acttgacttt gctgaaggcc cttgtcagac 1260 agcaactgcc tgagaagtac aaggaaattt tcttcgatca gtctaaaaat ggctacgccg 1320 gatacattga cggcggagca agccaggagg aattttacaa atttattaag cccatcttgg 1380 aaaaaatgga cggcaccgag gagctgctgg taaagcttaa cagagaagat ctgttgcgca 1440 aacagcgcac tttcgacaat ggaagcatcc cccaccagat tcacctgggc gaactgcacg 1500 ctatcctcag gcggcaagag gatttctacc cctttttgaa agataacagg gaaagattg 1560 agaaaatcct cacatttcgg ataccctact atgtaggccc cctcgccccgg ggaaattcca 1620 gattcgcgtg gatgactcgc aaatcagaag agaccatcac tccctggaac ttcgaggaag 1680 tcgtggataa gggggcctct gcccagtcct tcatcgaaag gatgactaac tttgataaaa 1740 atctgcctaa cgaaaaggtg cttcctaaac actctctgct gtacgagtac ttcacagttt 1800 ataacgagct caccaaggtc aaatacgtca cagaagggat gagaaagcca gcattcctgt 1860 ctggagagca gaaaagct atcgtggacc tcctcttcaa gacgaaccgg aaagttaccg 1920 tgaaacagct caaagaac tatttcaaaa agattgaatg tttcgactct gttgaaatca 1980 gcggagtgga ggatcgcttc aacgcatccc tgggaacgta tcacgatctc ctgaaaatca 2040 2100 tcacccttac gttgtttgaa gatagggaga tgattgaaga acgcttgaaa acttacgctc 2160 atctcttcga cgacaaagtc atgaaacagc tcaagaggcg ccgatataca ggatggggc 2220 ggctgtcaag aaaactgatc aatgggatcc gagacaagca gagtggaaag acaatcctgg 2280 attttcttaa gtccgatgga tttgccaacc ggaacttcat gcagttgatc catgatgact 2340 ctctcacctt tagggagac atccagaaag cacaagtttc tggccagggg gacagtcttc 2400 acgagcacat cgctaatctt gcaggtagcc cagctatcaa aagggataa ctgcagaccg 2460 ttaaggtcgt ggatgaactc gtcaaagtaa tgggaaggca taagcccgag aatatcgtta 2520 tcgagatggc ccgagagaac caaactaccc agaagggaca gaacagt agggaaagga 2580 tgaagaggat tgaagagggt ataaaagaac tggggtccca aatccttaag gaacacccag 2640 ttgaaaacac ccagcttcag aatgagaagc tctacctgta ctacctgcag aacggcaggg 2700 acatgtacgt ggatcagaa ctggacatca atcggctctc cgactacgac gtggatcata 2760 tcgtgcccca gtcttttctc aaagatgatt ctattgataa taaagtgttg aaagatccg 2820 2880 attattggcg gcagctgctg aacgccaaac tgatcacaca acggaagttc gataatctga 2940 ctaaggctga acgaggtggc ctgtctgagt tggataaagc cggcttcatc aaaaggcagc 3000 ttgttgagac acgccagatc accaagcacg tggcccaaat tctcgattca cgcatgaaca 3060 ccaagtacga tgaaatgac aaactgattc gagaggtgaa agttattact ctgaagtcta 3120 agctggtctc agatttcaga aaggacttc agttttatata ggtgagagag atcaacatt 3180 accaccatgc gcatgatgcc tacctgaatg cagtgtagg cactgcactt atcaaaaat 3240 atcccaagct tgaatctgaa ttgtttacg gagactataa agtgtacgat gttaggaaaa 3300 tgatcgcaaa gtctgagcag gaataggca aggccaccgc taagtactc ttttacagca 3360 atattatgaa tttttcaag accgagatta cactggccaa tggagatt cggaagcgac 3420 cacttatcga aaaaacgga gaacaggag aaatcgtgtg ggacaagggt agggattcg 3480 cgacagtccg gaagtcctg tccatgccgc aggtgaacat cgttaaaaag accgaagtac 3540 agaccggagg cttctccaag gaagtatcc tcccgaaag gacagcgac aagctgatcg 3600 cacgcaaaaa agatgggac cccaagaat acggcggatt cgattctcct acagtcgctt 3660 acagtgtact gttgtggcc aaagtggaga aagggaagtc taaaaaactc aaagcgtca 3720 aggaactgct gggcatcaca atcatggagc gatcaagctt cgaaaaaaac cccatcgact 3780 ttctcgaggc gaaggatat aaagaggtca aaaaagacct catcattaag ctcccaagt 3840 actctctctt tgagcttga aacggccgga aacgaatgct cgctagctgcg ggcgagctgc 3900 agaaaggtaa cgagctggca ctgccctcta atacgttaa ttcttgtat ctggccagcc 3960 actatgaaaa gctcaaaggg tctcccgaag atatgagca gaagcagctg ttcgtggaac 4020 aacaaaca ctaccttgat gagatcatcg agcaataag cgaattctcc aaagagtga 4080 tcctcgccga cgctaacctc gataaggtgc ttctgctta caataagcac agggataagc 4140 ccatcaggga gcaggcagaa aacattatcc acttgtttac tctgaccac ttggcggcgc 4200 ctgcagcctt caagtacttc vakaccacca tagacagaaa gcggtacacc tctacaagg 4260 aggtcctgga cgccacactg attcatcagt cattacggg gctctatgaa acagaatcg 4320 acctctctca gctcggtgga gagagcaggg ctgaccccaa gagagagg aaggtggatc 4380 aaaaaaaaaaaaaaaaaaaaaaaaaa out from outletcgagc gattachaag 4440 accatgacgg tgatttaaa gatcatgaca tcgattaca ggatgacgat tgaagtgag 4500 ctttctaact aaaaaggcct cccaaatcgg ggggcctttt ttattgataa caaaacgcta 4560 gcggccgcat aatgcttaag tcgaacagaa agtaatcgta ttgtacacgg ccgcataatc 4620 gaaattaata cgactcacta taggtctgac agcagacgtg cactggccag gttttagagc 4680 tagaaatagc aagttaaaat aaggctagtc cgttatcaac ttgaaaaagt ggcaccgagt 4740 cggtgctccg ctgagcaata actagcataa ccccttgggg cctctaaacg ggtcttgagg 4800 ggttttttgt taccctttat ctgacagcag acgtgcactg gccaggggga tggtttgggc 4860 ctcacgtgac atgtgagcaa aagctgaaac ctcaggcatt tgagaagcac acggtcacac 4920 tgcttccggt agtcaataaa ccggtaaacc agcaatagac ataagcggct atttaacgac 4980 cctgccctga accgacgacc gggtcgaatt tgctttcgaa tttctgccat tcatccgctt 5040 attatcactt attcaggcgt agcaaccagg cgtttaaggg caccaataac tgccttaaaa 5100 aaattacgcc ccgccctgcc actcatcgca gtactgttgt aattcattaa gcattctgcc 5160 gacatggaag ccatcacaaa cggcatgatg aacctgaatc gccagcggca tcagcacctt 5220 gtcgccttgc gtataatatt tgcccatagt gaaaacgggg gcgaagaagt tgtccatatt 5280 ggccacgttt aaatcaaaac tggtgaaact cacccaggga ttggctgaga cgaaaaacat 5340 attctcaata aaccctttag ggaaataggc caggttttca ccgtaacacg ccacatcttg 5400 cgaatatatg tgtagaaact gccggaaatc gtcgtggtat tcactccaga gcgatgaaaa 5460 cgtttcagtt tgctcatgga aaacggtgta acaagggtga acactatccc atatcaccag 5520 ctcaccgtct ttcattgcca tacggaactc cggatgagca ttcatcaggc gggcaagaat 5580 gtgaataaag gccggataaa acttgtgctt atttttcttt acggtcttta aaaaggccgt 5640 aatatccagc tgaacggtct ggttataggt acattgagca actgactgaa atgcctcaaa 5700 atgttcttta cgatgccatt gggatatatc aacggtggta tatccagtga tttttttctc 5760 cattttagct tccttagctc ctgaaaatct cgataactca aaaaatacgc ccggtagtga 5820 tcttatttca ttatggtgaa agttggaacc tcttacgtgc cgatcaacgt ctcattttcg 5880 ccaaaagttg gcccagggct tcccggtatc aacagggaca ccaggattta tttattctgc 5940 gaagtgatct tccgtcacag gtatttattc ggcgcaaagt gcgtcgggtg atgctgccaa 6000 cttactgatt tagtgtatga tggtgttttt gaggtgctcc agtggcttct gtttctatca 6060 gctgtccctc ctgttcagct actgacgggg tggtgcgtaa cggcaaaagc accgccggac 6120 atcagcgcta gcggagtgta tactggctta ctatgttggc actgatgagg gtgtcagtga 6180 agtgcttcat gtggcaggag aaaaaaggct gcaccggtgc gtcagcagaa tatgtgatac 6240 aggatatatt ccgcttcctc gctcactgac tcgctacgct cggtcgttcg actgcggcga 6300 gcggaaatgg cttacgaacg gggcggagat ttcctggaag atgccaggaa gatacttaac 6360 agggaagtga gagggccgcg gcaaagccgt ttttccatag gctccgcccc cctgac 6416 <210> 11 <211> 4 <212> DNA <213> Artificial Sequence <220> <223> Exemplary 5' PAM Site <400> 11 tttv 4 <210> 12 <211> 6 <212> DNA <213> Artificial Sequence <220> <223> Exemplary 3' PAM Site <220> <221> Feature Unspecified <222> (1)..(2) <223> n is a, c, g, or t <400> 12 nngrrt 6
Claims
1. A method for measuring the activity of a nucleic acid modifying enzyme, comprising the following steps: a) separating the plurality of polynucleotide constructs into compartments, wherein each compartment comprises a single polynucleotide construct, wherein each polynucleotide construct comprises i) a first polynucleotide sequence encoding a nucleic acid modification enzyme or a variant thereof operably linked to a first promoter; and ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, wherein when the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target and the nucleic acid modification enzyme are continuously co-expressed as a single RNA transcript driven by the first promoter; and wherein the plurality of polynucleotide constructs encode different variants of the nucleic acid modification enzyme, and / or different DNA or RNA targets; b) subjecting the compartment to conditions allowing in vitro expression of RNA and protein; c) subjecting a plurality of said compartments to conditions that allow modification of said DNA / RNA target by a nucleic acid modifying enzyme having modification activity on said DNA or RNA target, thereby producing a population of DNA / RNA molecules comprising one or more of: i. a polynucleotide construct and / or RNA transcript or a fragment thereof that has been modified by the nucleic acid modifying enzyme; ii. a polynucleotide construct and / or RNA transcript that has not been modified by the nucleic acid modifying enzyme; d) harvesting the DNA / RNA molecule population produced in step (c) and performing single molecule sequencing on it; e) detecting and counting the DNA / RNA molecules mentioned in steps c)i and c)ii based on the sequencing results.
2. The method of claim 1, wherein the nucleic acid modifying enzyme is an RNA-guided nucleic acid modifying enzyme, and each compartment further comprises a guide RNA or a nucleotide template encoding the same.
3. The method of claim 1, wherein the nucleic acid modifying enzyme is an RNA-guided nucleic acid modifying enzyme, and each polynucleotide further comprises a third polynucleotide sequence encoding a variant guide RNA (gRNA).
4. A method for measuring the activity of a nucleic acid modifying enzyme, comprising the following steps: a) separating the plurality of polynucleotide constructs into compartments, wherein each compartment comprises a single polynucleotide construct, wherein each polynucleotide construct comprises: i) a first polynucleotide sequence encoding a guide RNA (gRNA) operably linked to a first promoter; ii) a second polynucleotide sequence comprising a DNA target or a DNA template encoding an RNA target, wherein when the second polynucleotide sequence comprises a DNA template encoding an RNA target, the RNA target and the gRNA are continuously co-expressed as a single RNA transcript driven by the first promoter; wherein the plurality of polynucleotide constructs encode different gRNAs, and / or different DNA or RNA targets; and wherein each compartment further comprises an RNA-guided nucleic acid modification enzyme or a variant thereof or a nucleotide template encoding the same; b) subjecting the compartment to conditions allowing in vitro transcription and / or translation of RNA and protein; c) subjecting the compartment to conditions that allow modification of the DNA and / or RNA target by an RNA-guided nucleic acid modification enzyme functionally active against the DNA or RNA target in the presence of a gRNA, thereby producing a population of DNA / RNA molecules comprising one or more of: i. a polynucleotide construct and / or RNA transcript or a fragment thereof that has been modified by the nucleic acid modifying enzyme; ii. a polynucleotide construct and / or RNA transcript that has not been modified by the nucleic acid modifying enzyme; d) harvesting the DNA / RNA molecule population produced in step (c), and performing single-molecule long-read sequencing on it; e) detecting and counting the DNA / RNA molecules mentioned in step c)i and / or c)ii based on the sequencing results.
5. The method according to any one of claims 1 to 4, wherein the method further comprises evaluating the enzymatic activity of one or more nucleic acid modifying enzymes on one or more of the DNA / RNA targets, the evaluation being performed by calculating the number of polynucleotide constructs and / or RNA transcripts modified by the nucleic acid modifying enzyme (∑ count 修饰 ), and compare it with the number of polynucleotide constructs and / or RNA transcripts not modified by the nucleic acid modifying enzyme (∑ count 未修饰 ) or the total number of polynucleotide constructs and / or RNA transcripts (∑ count 修饰+未修饰 ) for comparison.
6. The method according to claim 5, wherein the enzyme activity is represented by a value calculated using any one of the following formulas: 1) Enzyme activity ≈ ∑ counts 修饰 / ∑ count 未修饰 ; 2) Enzyme activity ≈ ∑ counts 修饰 / ∑ count 修饰+未修饰 .
7. The method according to any one of claims 1 to 4, wherein step d) further comprises destroying the compartment by physical or chemical means.
8. The method according to any one of claims 1 to 4, wherein step d) further comprises purifying the harvested DNA / RNA molecules to remove excess DNA, RNA and / or protein from the reaction.
9. The method according to any one of claims 1 to 4, wherein no further modifications other than those required for sequencing are performed on the harvested DNA / RNA molecule population before the sequencing reaction.
10. The method according to any one of claims 1 to 4, wherein the detection and counting of DNA / RNA molecules that have or have not been modified by the nucleic acid modifying enzyme is based solely on the data generated during the sequencing and does not require further modification or processing of the DNA / RNA molecules.
11. The method of claim 5, wherein the enzymatic activity is a cleavage activity, and the detection and enumeration of modified and unmodified polynucleotide constructs or RNA transcripts is achieved by aligning sequencing reads of the DNA / RNA molecule with a reference sequence, the reference sequence comprising a cleavage site window for the nucleic acid modification enzyme, wherein i) when the 3' end of the DNA / RNA molecule is mapped to the region 3' downstream of the cleavage site window, the DNA / RNA molecule is an unmodified polynucleotide construct or RNA target; ii) when the 3' end of the DNA / RNA molecule maps to a region within the cleavage site window, the DNA / RNA molecule is a modified polynucleotide construct or RNA target; iii) When the 3' end of a DNA / RNA molecule maps to a region 5' upstream of the cleavage site window, the DNA / RNA molecule is non-informative and is not used to measure modification activity.
12. The method according to any one of claims 1 to 4, wherein each compartment further comprises in vitro transcription and translation (IVTT) reagents capable of achieving in vitro transcription and / or translation of proteins and / or RNA.
13. The method according to any one of claims 1 to 4, wherein the compartments are emulsion droplets.
14. The method of any one of claims 1 to 4, wherein the compartmentalization is achieved using microfluidics, hydrogel restricted diffusion, or compartmentalized wells.
15. The method of any one of claims 1 to 4, wherein the nucleic acid modifying enzyme is a CRISPR-associated protein (Cas).
16. The method of any one of claims 1 to 4, wherein the variant of the nucleic acid modifying enzyme contains one or more inactivated catalytic sites and is capable of binding to and inhibiting the expression of a DNA target without modifying the DNA target.
17. The method according to any one of claims 1 to 4, wherein the variant of the nucleic acid modifying enzyme is fused to one or more additional functional domains capable of modifying DNA or RNA.
18. The method according to any one of claims 1 to 4, wherein the first polynucleotide sequence and the second polynucleotide sequence completely overlap or partially overlap.
19. The method of any one of claims 2 to 4, wherein the DNA or RNA target comprises a protospacer that is at least partially complementary to the guide RNA.
20. The method of any one of claims 2 to 4, wherein when the polynucleotide construct comprises a DNA template encoding an RNA target, the RNA target further comprises a protospacer-flanking sequence (PFS).
Citation Information
Patent Citations
Emulsion-based screening methods
WO2018118968A1