A DEAR nucleic acid manipulation system based on RNA ribozyme and its application
Through the RNA ribozyme-based DEAR nucleic acid manipulation system, bacterial C-type intron RNA molecules are used for programmable targeted cutting, which solves the off-target effects and immune response problems of the CRISPR-Cas system and achieves efficient DNA and RNA cutting.
Patent Information
- Application Number
- CN202310424082.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-19
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2043-04-19
AI Technical Summary
The existing CRISPR-Cas nuclease system has problems such as off-target effects, excessive protein affecting transfection efficiency and potential immune responses, which limit the application of gene editing.
To develop an RNA ribozyme-based DEAR nucleic acid manipulation system that utilizes bacterial C-type group II intron-derived RNA molecules containing programmable substrate recognition regions for targeting and cleaving DNA or RNA.
It achieves efficient DNA and RNA cutting in Escherichia coli and mammalian eukaryotic cells, avoids the problems of large protein molecules and immune response, and provides more efficient gene editing capabilities.
Smart Images

Figure CN118813611B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biotechnology, and specifically relates to a DEAR nucleic acid manipulation system based on RNA ribozymes and its application, and more specifically to a programmable DEAR nucleic acid manipulation system with broad-spectrum RNA and DNA cutting capabilities and gene editing capabilities. Background Art
[0002] Currently, the commonly used gene editing technologies in my country are all developed based on RNA-guided CRISPR-Cas nucleases.
[0003] However, the CRISPR-Cas system still has some problems: first, the CRISPR-Cas system has off-target effects, and editing of Cas proteins in non-target areas may cause uncontrollable harmful mutations; second, the CRISPR-Cas system has the problem of too large proteins. The protein size of the currently used CRISPR-Cas editing tool molecules SpyCas9 and AsCas12a both exceeds 1,300 amino acids. The excessive molecular weight affects the transfection efficiency of the CRISPR-Cas system tools; at the same time, the CRISPR-Cas system has potential immune responses. The currently used SpyCas9 and AsCas12a proteins are derived from pathogenic bacteria that humans have come into contact with, which may cause human immune responses.
[0004] Therefore, the CRISPR-Cas nuclease system is limited by the limitations of its protein components. If a new generation of RNA-based nucleic acid targeting technology that combines gene sequence-specific targeting and catalytic activity can be developed, it is expected to overcome the limitations of the application of protease-based gene editing systems. Summary of the Invention
[0005] Problems to be solved by the invention
[0006] Based on the various problems existing in the CRISPR-Cas nuclease system in the prior art, the purpose of the present invention is to provide a DEAR nucleic acid manipulation system based on RNA ribozymes, and to apply it to the targeted modification (e.g., cutting) of nucleic acids (DNA, RNA).
[0007] Solutions for solving problems
[0008] The first aspect of the present invention provides a DEAR nucleic acid manipulation system, wherein the DEAR nucleic acid manipulation system comprises an RNA molecule of a C-type second group intron derived from bacteria, wherein the RNA molecule comprises a substrate recognition region that hybridizes with a target sequence in a target nucleic acid, and the C-type second group intron is a C-type second group intron in which there is no open reading frame encoding an intron-encoded protein in the IV domain.
[0009] In some embodiments, the C-type second group intron consists of domains I to VI, each domain exists in the form of a stem-loop structure and is separated from each other, and the substrate recognition region is located in the top ring region of domain I.
[0010] In some embodiments, the nucleotide sequence of the RNA molecule is selected from any one of the following:
[0011] (i) comprising a nucleotide sequence as shown in any one of SEQ ID NOs: 1 to 9;
[0012] (ii) a nucleotide sequence comprising the reverse complementary sequence of any one of SEQ ID NOs: 1 to 9;
[0013] (iii) a reverse complementary sequence of a sequence that can hybridize to the nucleotide sequence shown in (i) or (ii) under high stringency hybridization conditions or very high stringency hybridization conditions;
[0014] (iv) a sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, and most preferably at least 99% sequence identity with the nucleotide sequence shown in (i) or (ii).
[0015] In some embodiments, the substrate recognition region has a length of 6 nucleotides.
[0016] In some preferred embodiments, the substrate recognition region is programmable to hybridize to different target sequences.
[0017] In some embodiments, the target nucleic acid is DNA or RNA.
[0018] In some embodiments, the primary cleavage site of the DEAR nucleic acid manipulation system is 0-1 nt downstream of the 3' end of the target sequence in the target nucleic acid.
[0019] The second aspect of the present invention provides an isolated polynucleotide, wherein the polynucleotide comprises a nucleotide sequence encoding the DEAR nucleic acid manipulation system described in the first aspect of the present invention.
[0020] The third aspect of the present invention provides a nucleic acid construct, wherein the nucleic acid construct comprises the isolated polynucleotide described in the second aspect of the present invention.
[0021] The fourth aspect of the present invention provides a vector, wherein the vector comprises the isolated polynucleotide described in the second aspect of the present invention, or the nucleic acid construct described in the third aspect of the present invention.
[0022] The fifth aspect of the present invention provides a cell, wherein the cell comprises the DEAR nucleic acid manipulation system described in the first aspect of the present invention, the isolated polynucleotide described in the second aspect of the present invention, the nucleic acid construct described in the third aspect of the present invention, or the vector described in the fourth aspect of the present invention.
[0023] The sixth aspect of the present invention provides a reagent or kit, wherein the reagent or kit comprises the DEAR nucleic acid manipulation system described in the first aspect of the present invention, the isolated polynucleotide described in the second aspect of the present invention, the nucleic acid construct described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, or the cell described in the fifth aspect of the present invention.
[0024] The seventh aspect of the present invention provides a pharmaceutical composition, wherein the pharmaceutical composition comprises the DEAR nucleic acid manipulation system described in the first aspect of the present invention, the isolated polynucleotide described in the second aspect of the present invention, the nucleic acid construct described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention or the cell described in the fifth aspect of the present invention; and, optionally, a pharmaceutically acceptable carrier.
[0025] The eighth aspect of the present invention provides a method for modifying a target nucleic acid, which comprises the step of contacting the target nucleic acid with the DEAR nucleic acid manipulation system described in the first aspect of the present invention, the isolated polynucleotide described in the second aspect of the present invention, the nucleic acid construct described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, the cell described in the fifth aspect of the present invention, or the reagent or kit described in the sixth aspect of the present invention.
[0026] Use of the DEAR nucleic acid manipulation system described in the first aspect of the present invention, the isolated polynucleotide described in the second aspect of the present invention, the nucleic acid construct described in the third aspect of the present invention, the vector described in the fourth aspect of the present invention, and the cell described in the fifth aspect of the present invention in modifying target nucleic acids or preparing reagents or kits for modifying target nucleic acids.
[0027] Effects of the Invention
[0028] The RNA ribozyme-based DEAR nucleic acid manipulation system provided by the present invention is based on bacterial C-type second intron RNA molecules, avoiding the problems of the CRISPR-Cas system, such as the large protein molecule affecting transfection efficiency, and the potential immunogenicity of the Cas protein. The RNA ribozyme-based DEAR nucleic acid manipulation system provided by the present invention can achieve DNA and RNA cleavage and also has DNA cleavage ability in Escherichia coli and mammalian eukaryotic cells. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figures 1A to 1I: Display of the secondary structures of DEAR1~9. Figure 1A-Figure 1I The secondary structure prediction results for DEAR1-9 are shown. RNA fold was used for prediction, and domains I-VI and TRS are marked in the figure.
[0030] Figure 2 : Quality identification of RNA ribozyme molecules DEAR1~9.
[0031] Figure 3 :Verification of RNA cleavage activity of RNA ribozyme molecules DEAR1~9.
[0032] Figure 4 :Verification of the cleavage activity of RNA ribozyme molecules DEAR1~6 on unpaired RNA substrates.
[0033] Figure 5 :Verification of ssDNA cleavage activity of RNA ribozyme molecules DEAR1~6.
[0034] Figure 6 :Comparison of cleavage of paired and unpaired DNA substrates by RNA ribozyme molecules DEAR1-6.
[0035] Figure 7 : Verification of the ssDNA cleavage sites of RNA ribozyme molecules DEAR1~6.
[0036] Figure 8 :Results of reaction condition optimization for RNA ribozyme molecule DEAR1.
[0037] Figure 9 :Efficiency curve of reaction condition optimization of RNA ribozyme molecule DEAR1.
[0038] Figure 10 : Comparison of DNA cleavage activity between DEAR1 and RNA-guided protein nucleases.
[0039] Figure 11 :Verification of plasmid cleavage activity of RNA ribozyme molecule DEAR1.
[0040] Figure 12 :Verification of plasmid cleavage activity of RNA ribozyme molecules DEAR1~3 in Escherichia coli.
[0041] Figure 13 : Further verification of the plasmid cleavage activity of RNA ribozyme molecule DEAR1 in Escherichia coli.
[0042] Figure 14 :Verification of plasmid cleavage activity of RNA ribozyme molecules DEAR4~9 in bacteria.
[0043] Figure 15:Verification of the ssDNA cleavage activity of RNA ribozyme molecules DEAR1~6 that reprogram the TRS region.
[0044] Figure 16 : Schematic diagram of the survival of cells stably transfected with DEAR1 stable transfection plasmid and DEAR-NT stable transfection plasmid in Example 7.
[0045] Figure 17 A~ Figure 17 B is a schematic diagram of the sequencing results analysis of Example 7. DETAILED DESCRIPTION
[0046] In order to make the present invention more easily understood, certain technical and scientific terms are specifically defined below. Unless otherwise expressly defined herein, all other technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which the present invention belongs.
[0047] In this specification, the numerical range expressed using "a numerical value A to a numerical value B" means a range including the endpoints A and B.
[0048] In this specification, the use of “substantially” or “essentially” means that the standard deviation from a theoretical model or theoretical data is within a range of 5%, preferably 3%, and more preferably 1%.
[0049] In this specification, the use of "may" includes both the meaning of performing a certain process and the meaning of not performing a certain process.
[0050] As used herein, "optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and that the description includes instances where the event occurs and instances where it does not.
[0051] In this specification, references to "some specific / preferred embodiments," "other specific / preferred embodiments," "embodiments," etc., mean that the specific elements (e.g., features, structures, properties, and / or characteristics) described in connection with the embodiments are included in at least one embodiment described herein, and may or may not be present in other embodiments. In addition, it should be understood that the elements may be combined in various embodiments in any suitable manner.
[0052] The terms "comprise," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps is not limited to the listed steps or modules but may optionally include steps not listed, or other steps inherent to the process, method, product, or device.
[0053] In this application, "plurality" refers to two or more. "And / or" describes the relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the related objects are in an "or" relationship.
[0054] As used herein, the terms "polynucleotide" and "nucleic acid" are used interchangeably to refer to a polymeric form of nucleotides (ribonucleotides or deoxyribonucleotides) of any length. Thus, the term includes, but is not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derived nucleotide bases.
[0055] In the art, "G", "C", "A", "T" and "U" generally represent the bases guanine, cytosine, adenine, thymine and uracil, respectively, but it is also generally known in the art that "G", "C", "A", "T" and "U" each generally represent nucleotides containing guanine, cytosine, adenine, thymine and uracil as bases, respectively, which is a common way to represent deoxyribonucleic acid sequences and / or ribonucleic acid sequences. Therefore, in the context of the present invention, the meanings represented by "G", "C", "A", "T", and "U" include the above-mentioned various possible situations. However, it should be understood that the term "ribonucleotide" or "nucleotide" can also refer to a modified nucleotide or an alternative replacement part. Those skilled in the art will recognize that guanine, cytosine, adenine and uracil can be replaced by other parts without substantially changing the base pairing properties of an oligonucleotide (including a nucleotide having such a replacement part).
[0056] As used herein, the term "nucleic acid manipulation" includes binding, nicking one strand, or cleaving (i.e., cutting) two strands of a nucleic acid, or includes modifying or editing a nucleic acid. Nucleic acid manipulation can silence, activate, or modulate (increase or decrease) the expression of an RNA or polypeptide encoded by the nucleic acid.
[0057] As used herein, "hybridizable," "complementary," or "substantially complementary" means that a nucleic acid (e.g., RNA, DNA) comprises a nucleotide sequence that enables the nucleic acid to non-covalently bind (i.e., form Watson-Crick base pairs and / or G / U base pairs), "anneal," or "hybridize" with another nucleic acid in a sequence-specific, antiparallel manner (i.e., the nucleic acid specifically binds to the complementary nucleic acid) under appropriate in vitro and / or in vivo temperature and solution ionic strength conditions. Standard Watson-Crick base pairing includes: adenine (A) pairs with thymidine (T), adenine (A) pairs with uracil (U), and guanine (G) pairs with cytosine (C). In addition, for hybridization between two RNA molecules (e.g., dsRNA), and for hybridization between a DNA molecule and an RNA molecule (e.g., when a DNA or RNA target nucleic acid base pairs with the substrate recognition region of the DEAR nucleic acid manipulation system): guanine (G) can also pair with uracil (U). For example, in the case of tRNA anticodons base pairing with codons in mRNA, G / U base pairing is at least partially responsible for the degeneracy of the genetic code.
[0058] Hybridization and washing conditions are well known and are exemplified in Sambrook, J., Fritsch, E.F. and Maniatis, T. Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), particularly Chapter 11 and Table 11.1 of this reference; and Sambrook, J. and Russell, W., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor (2001). Conditions of temperature and ionic strength determine the "stringency" of hybridization.
[0059] As used herein, "moderately stringent conditions," "medium-high stringency conditions," "high stringency conditions," or "very high stringency conditions" describe conditions for nucleic acid hybridization and washing. Guidance for conducting hybridization reactions can be found in Current Protocols in Molecular Biology, John Wiley & Sons, NY (1989), 6.3.1-6.3.6, which is incorporated herein by reference. Both aqueous and nonaqueous methods are described in this document, and either method can be used. For example, specific hybridization conditions are as follows: (1) low stringency hybridization conditions are in 6× sodium chloride / sodium citrate (SSC) at about 45°C, followed by two washes in 0.2×SSC, 0.1% SDS at at least 50°C (the wash temperature can be increased to 55°C for low stringency conditions); (2) moderate stringency hybridization conditions are in 6×SSC at about 45°C, followed by one or more washes in 0.2×SSC, 0.1% SDS at 60°C; (3) high stringency hybridization conditions are in 6×SSC at about 45°C, followed by one or more washes in 0.2×SSC, 0.1% SDS at 65°C, and preferably; (4) very high stringency hybridization conditions are 0.5 M sodium phosphate, 7% SDS at 65°C, followed by one or more washes in 0.2×SSC, 1% SDS at 65°C.
[0060] Hybridization requires that the two nucleic acids contain complementary sequences, but mismatches between bases are possible. Conditions suitable for hybridization between two nucleic acids depend on the length of the nucleic acids and the degree of complementarity, which are variables well known in the art.
[0061] In the present invention, a DNA sequence that "encodes" a specific RNA is a DNA nucleotide sequence that is transcribed into RNA. A DNA polynucleotide may encode RNA (mRNA) that is converted into protein (thus, both DNA and mRNA encode protein), or a DNA polynucleotide may encode RNA that is not translated into protein (e.g., tRNA, rRNA, microRNA (miRNA), "non-coding" RNA (ncRNA), and the DEAR nucleic acid manipulation system provided herein).
[0062] As used herein, the terms "naturally occurring," "unmodified," or "wild-type" as applied to nucleic acids, polypeptides, cells, or organisms refer to nucleic acids, polypeptides, cells, or organisms as they occur in nature. For example, a polypeptide or polynucleotide sequence that is present in an organism and that can be isolated from a source in nature is naturally occurring.
[0063] In the present invention, "recombinant" means that a specific nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, polymerase chain reaction (PCR) and / or ligation steps, which produce a construct having a structural coding sequence or non-coding sequence that can be distinguished from the endogenous nucleic acids present in the natural system. The DNA sequence encoding the polypeptide can be assembled from cDNA fragments or from a series of synthetic oligonucleotides to provide a synthetic nucleic acid that can be expressed by a recombinant transcription unit contained in a cell or in a cell-free transcription and translation system. Genomic DNA containing related sequences can also be used in the formation of recombinant genes or transcription units. Sequences of non-translated DNA can be present at the 5' end or 3' end of the open reading frame, wherein such sequences do not interfere with the manipulation or expression of the coding region and can actually play a role in regulating the production of the desired product by various mechanisms (see "DNA regulatory sequences"). Alternatively, DNA sequences encoding untranslated RNA (e.g., the DEAR nucleic acid manipulation system provided by the present invention) can also be considered to be recombinant. Therefore, for example, the term "recombinant" nucleic acid refers to a non-naturally occurring polynucleotide or nucleic acid, such as a polynucleotide or nucleic acid made by the artificial combination of two otherwise separated segments of the sequence through human intervention. This artificial combination is often accomplished by chemical synthesis means or by artificially manipulating the isolated segments of nucleic acid (e.g., by genetic engineering techniques). This operation is usually performed to replace codons with codons encoding the same amino acid, conservative amino acids, or non-conservative amino acids. Alternatively, this operation is performed to link together nucleic acid segments with the desired function to produce the desired functional combination. This artificial combination is often accomplished by chemical synthesis means or by artificially manipulating the isolated segments of nucleic acid (e.g., by genetic engineering techniques).
[0064] As used in this disclosure, the term "isolated" means a substance that is in a form or environment not found in nature. Non-limiting examples of isolated substances include (1) any non-naturally occurring substance, (2) any substance, including but not limited to any enzyme, mutant, nucleic acid, protein, peptide, or cofactor, that is at least partially removed from one or more or all of the naturally occurring components with which it is essentially associated; (3) any substance that has been artificially modified relative to the substance found in nature; or (4) any substance that has been modified by increasing the amount of the substance relative to other components with which it is naturally associated (e.g., recombinant production in a host cell; multiple copies of a gene encoding the substance; and use of a stronger promoter than the promoter naturally associated with the gene encoding the substance).
[0065] As used herein, the term "nucleic acid construct" comprises a polynucleotide encoding a polypeptide, domain, or module operably linked to suitable regulatory sequences necessary for expression of the polynucleotide in a selected cell or strain. In the present disclosure, transcriptional regulatory elements include promoters and, further, may include enhancers, silencers, insulators, and other elements.
[0066] The term "vector" refers to a genetic element, such as a plasmid, cosmid, bacmid, phage, or virus, to which another genetic sequence or element (DNA or RNA) can be attached. A vector can be a replicon, thereby causing the replication of the attached sequence or element. An "expression vector" is a vector that facilitates the expression of a nucleic acid or a nucleic acid sequence encoding a polypeptide in a host cell or organism.
[0067] In the present invention, the terms "recombinant expression vector" or "DNA construct" are used interchangeably herein to refer to a DNA molecule comprising a vector and an insert. Recombinant expression vectors are generally produced for the purpose of expressing and / or propagating one or more inserts, or for the purpose of constructing other recombinant nucleotide sequences. The one or more inserts may or may not be operably linked to a promoter sequence and may or may not be operably linked to a DNA regulatory sequence.
[0068] As used herein, the term "operably linked" refers to a nucleic acid sequence that is placed into a functional relationship with another nucleic acid sequence. Examples of nucleic acid sequences that can be operably linked include, but are not limited to, promoters, transcription terminators, enhancers or activators, and heterologous genes that, when transcribed and, if appropriate, translated, produce a functional product, such as a protein, ribozyme, or RNA molecule.
[0069] As used herein, the term "derived" refers to origin or source and may include naturally occurring, recombinant, unpurified or purified molecules. A nucleic acid derived from an original nucleic acid may partially or completely comprise the original nucleic acid and may be a fragment or variant of the original nucleic acid.
[0070] In the present invention, the term "ribozyme" refers to an RNA molecule that can catalyze a specific biochemical reaction. Common examples of such reactions include the cleavage or ligation and modification of RNA and DNA.
[0071] In the present invention, a "target nucleic acid" is a polynucleotide (e.g., DNA such as genomic DNA, RNA, etc.) that includes a site ("target site" or "target sequence") targeted by the DEAR nucleic acid manipulation system provided herein. The target sequence is the sequence to which the substrate recognition region of the DEAR nucleic acid manipulation system will hybridize. For example, the target site (or target sequence) 5'-UGUCUU-3' or 5'-TGTCTT-3' within the target nucleic acid is targeted (or bound by, hybridizes to, or is complementary to) the sequence 5'-AAGACA-3'. Suitable hybridization conditions include physiological conditions normally present in cells.
[0072] In the present invention, "cleavage" means the breakage of the covalent backbone of a target nucleic acid molecule (e.g., RNA, DNA). Both single-stranded and double-stranded cleavage are possible, and double-stranded cleavage can occur as a result of two distinct single-stranded cleavage events. A "primary cleavage site" refers to the DNA / RNA breakage site corresponding to a cleavage product with distinct bands. A "secondary cleavage site" refers to the DNA / RNA breakage site corresponding to a cleavage product with less distinct bands.
[0073] The technical solution of the present invention is described in detail below.
[0074] The present invention constructs a DEAR nucleic acid manipulation system based on RNA ribozymes based on bacterial class II intron elements. Class II introns are composed of two parts: RNA ribozymes and intron-encoded proteins (IEPs). RNA ribozymes can catalyze the self-splicing maturation of primary transcripts, while protein IEPs play an auxiliary role. The RNA ribozyme portion includes six domains, I to VI. Domain I is the largest of all domains and plays an important stabilizing role in the formation of the overall structure of the intron. It contains an exon binding site (EBS) for binding to exons. Domains II and III also participate in the formation of the ribozyme structure. Domain IV contains an open reading frame (ORF), and the protein it encodes is IEP. Domain V is the catalytic center of the RNA ribozyme, while domain VI performs auxiliary catalytic functions. According to the primary sequence and secondary structure characteristics of RNA, the second class of introns can be divided into A, B and C classes, among which C class is considered to be a more ancient class of introns (DM Simon et al., Group II introns in eubacteria and archaea: ORF-less introns and new varieties. RNA 14, 1704-1713 (2008); AM Lambowitz, S. Zimmerly, Mobile group II introns. Annu Rev Genet 38, 1-35 (2004); JSRest, DP Mindell, Retroids in archaea: phylogeny and lateral origins. Mol Biol Evol 20, 1134-1142 (2003).). The present invention focuses on the C-type second class introns in which there is no open reading frame encoding IEP in the IV domain.
[0075] The present invention discovered that the EBS of a C-type second-class intron and its surrounding sequences can be used as substrate recognition elements for the target nucleic acid of a ribozyme (referred to as the target recognition site (TRS) in the present invention). The programmability of the TRS was discovered and demonstrated, and the target nucleic acid (RNA, DNA) was hydrolyzed and cleaved with the help of the V domain of the RNA intron ribozyme. Therefore, the present invention refers to the system with programmable nucleic acid recognition and cleavage capabilities constructed based on a C-type second-class intron derived from bacteria, whose IV domain does not encode an open reading frame encoding an intron-encoded protein, as the RNA ribozyme-based DEAR nucleic acid manipulation system.
[0076] <DEAR Nucleic Acid Manipulation System>
[0077] In some embodiments of the present invention, a DEAR nucleic acid manipulation system based on RNA ribozyme is provided, which comprises an (isolated) RNA molecule derived from group II intron type C of bacteria, and the RNA molecule contains a substrate recognition region that hybridizes with a target sequence in a target nucleic acid.
[0078] In some embodiments of the present invention, the group II intron type C is a group II intron type C in which there is no open reading frame encoding IEP in the IV domain.
[0079] The DEAR nucleic acid manipulation system provided by the present invention acts as an endonuclease, which catalyzes nucleic acid cleavage at a specific sequence in a targeted target nucleic acid (such as DNA, RNA). As will be detailed later, the sequence specificity is provided by the substrate recognition region in the DEAR nucleic acid manipulation system, and this substrate recognition region hybridizes with the target sequence in the target nucleic acid. Therefore, the DEAR nucleic acid manipulation system binds to the target nucleic acid through the hybridization of the substrate recognition region with the target sequence in the target nucleic acid. In other words, the position where specific binding (and / or cleavage) of the target nucleic acid occurs is determined by the base pairing complementarity between the substrate recognition region and the target nucleic acid.
[0080] In some specific embodiments, the main cleavage site of the DEAR nucleic acid manipulation system is 0 - 1 nt downstream of the 3' end of the target sequence in the target nucleic acid, that is, the main cleavage site is located at 0 - 1 nt downstream of the 3' end of the region paired with the substrate recognition region on the target nucleic acid.
[0081] In some embodiments of the present invention, the nucleotide sequence of the RNA molecule of the DEAR nucleic acid manipulation system is selected from any one of the following:
[0082] (i) comprising a nucleotide sequence as shown in any one of SEQ ID NOs: 1 - 9;
[0083] (ii) comprising a nucleotide sequence of the reverse complementary sequence of the sequence as shown in any one of SEQ ID NOs: 1 - 9;
[0084] (iii) under high stringency hybridization conditions or very high stringency hybridization conditions, the reverse complementary sequence of a sequence that can hybridize with the nucleotide sequence shown in (i) or (ii);
[0085] (iv) a sequence having at least 90%, optionally at least 95%, preferably at least 97%, more preferably at least 98%, most preferably at least 99% sequence identity with the nucleotide sequence shown in (i) or (ii).
[0086] SEQ ID NOs: 1 - 9 are as follows:
[0087]
[0088]
[0089]
[0090]
[0091] Among them, "NNNNNN" is the substrate recognition region, where N is A, U, G or C.
[0092] In some specific embodiments, the DEAR nucleic acid manipulation system comprises an RNA molecule, the nucleotide sequence of which is a nucleotide sequence as shown in any one of SEQ ID NOs: 1 to 9.
[0093] In some preferred embodiments, the DEAR nucleic acid manipulation system comprises an RNA molecule, the nucleotide sequence of which is a nucleotide sequence as shown in any one of SEQ ID NOs: 1 to 6.
[0094] In some further preferred embodiments, the DEAR nucleic acid manipulation system comprises an RNA molecule, and the nucleotide sequence of the RNA molecule is a nucleotide sequence as shown in any one of SEQ ID NOs: 1 to 3, and 5.
[0095] (Substrate recognition region)
[0096] In some embodiments, the substrate recognition region of the DEAR nucleic acid manipulation system is a nucleotide sequence that is complementary to a sequence in a target nucleic acid (target sequence). In other words, the substrate recognition region of the DEAR nucleic acid manipulation system can interact with a target nucleic acid (e.g., DNA, RNA) in a sequence-specific manner through hybridization (i.e., base pairing). The substrate recognition region can be modified (e.g., by genetic engineering) / designed to hybridize with any desired target sequence within a target nucleic acid (e.g., a prokaryotic target nucleic acid, a eukaryotic target nucleic acid, an isolated target nucleic acid).
[0097] In some embodiments, the substrate recognition region is programmable in that it can be designed or engineered to recognize and bind different target sequences.
[0098] In some embodiments, the complementarity percentage between the substrate recognition region and the target sequence of the target nucleic acid is 60% or more (e.g., 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some embodiments, the complementarity percentage between the substrate recognition region and the target sequence of the target nucleic acid is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some embodiments, the complementarity percentage between the substrate recognition region and the target sequence of the target nucleic acid is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some embodiments, the complementarity percentage between the substrate recognition region and the target sequence of the target nucleic acid is 100%.
[0099] In some embodiments, the substrate recognition region has a length of 6 nucleotides (nt). In some specific embodiments, the sequence of the substrate recognition region is: N1N2N3N4N5N6; wherein N1 to N6 are A, G, C or U respectively.
[0100] In some embodiments, at least four nucleotides in the substrate recognition region of the DEAR nucleic acid manipulation system are complementary to the target sequence of the target nucleic acid. In some preferred embodiments, at least five nucleotides in the substrate recognition region of the DEAR nucleic acid manipulation system are complementary to the target sequence of the target nucleic acid. In some more preferred embodiments, six nucleotides in the substrate recognition region of the DEAR nucleic acid manipulation system are complementary to the target sequence of the target nucleic acid.
[0101] In some specific embodiments, the sequence of the substrate recognition region is selected from, but not limited to:
[0102] (a)AAGACA;
[0103] (b) UAGGCA;
[0104] (c)CAGACA;
[0105] (d)AAUGAA;
[0106] (e)AUAACA;
[0107] (f)ACAUCA;
[0108] (g) CACUCA;
[0109] (h)AUUACA.
[0110] In some specific embodiments, the DEAR nucleic acid manipulation system comprises an RNA molecule, the nucleotide sequence of which is a nucleotide sequence as shown in any one of SEQ ID NOs: 10 to 18.
[0111] In some specific embodiments, the sequence of the substrate recognition region is:
[0112] (i) CGAUAG.
[0113] (Target nucleic acid)
[0114] In the present invention, the DEAR nucleic acid manipulation system can bind to and cleave a target nucleic acid. In the present invention, the target nucleic acid can be any nucleic acid (e.g., DNA, RNA), can be any type of nucleic acid (e.g., chromosomal (genomic DNA), derived from a chromosome, chromosomal DNA, plasmid, viral, extracellular, intracellular, mitochondrial, chloroplast, linear, circular, etc.), and can be from any organism (e.g., as long as the DEAR nucleic acid manipulation system comprises a nucleotide sequence that hybridizes to a target sequence in the target nucleic acid, such that the target nucleic acid can be targeted).
[0115] Specifically, in the present invention, the target nucleic acid can be DNA or RNA. In some exemplary embodiments, the target nucleic acid is selected from: mRNA, rRNA, tRNA, non-coding RNA (ncRNA), long non-coding RNA (lncRNA) and microRNA (miRNA). In some exemplary embodiments, the target nucleic acid is viral DNA, plasmid DNA. The target nucleic acid can be located anywhere, for example, outside of a cell in vitro, inside a cell in vitro, inside a cell in vivo, inside a cell in vitro.
[0116] <Biological Materials>
[0117] (Isolated Polynucleotide)
[0118] In some embodiments of the present invention, an isolated polynucleotide is provided, wherein the polynucleotide comprises a nucleotide sequence encoding the DEAR nucleic acid manipulation system of the present invention.
[0119] (Nucleic acid construct)
[0120] In some embodiments of the present invention, a nucleic acid construct is provided, wherein the nucleic acid construct comprises the isolated polynucleotide of the present invention.
[0121] In some optional embodiments, the polynucleotide is operably linked to one or more regulatory sequences, which are nucleotide sequences comprising a promoter and / or a ribosome binding site, and the regulatory sequences direct the expression of the genes of the DEAR nucleic acid manipulation system in the host cell.
[0122] (Carrier)
[0123] In some embodiments of the present invention, a vector is provided, wherein the vector comprises the isolated polynucleotide of the present invention, or the nucleic acid construct of the present invention.
[0124] In some specific embodiments, the vector is a recombinant expression vector.
[0125] Suitable recombinant expression vectors include viral expression vectors (e.g., viral vectors based on viruses such as vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retroviral vectors (e.g., murine leukemia virus, spleen necrosis virus, and vectors derived from retroviruses such as Rous sarcoma virus, Harvey sarcoma virus, avian leukemia virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus), and the like.
[0126] (cell)
[0127] In some embodiments of the present invention, the present invention provides a cell comprising the DEAR nucleic acid manipulation system of the present invention, the isolated polynucleotide of the present invention, the nucleic acid construct of the present invention, or the vector of the present invention.
[0128] The cell can be any of a variety of cells, including, for example, in vitro cells, in vivo cells, ex vivo cells, primary cells, cancer cells, animal cells, plant cells, algae cells, fungal cells, and the like.
[0129] In some embodiments, the cell is a recipient of the DEAR nucleic acid manipulation system, isolated polynucleotide, nucleic acid construct, or vector provided herein, and may also be referred to as a "host cell" or "target cell." A host cell or target cell can be a recipient of the DEAR nucleic acid manipulation system, isolated polynucleotide, nucleic acid construct, or vector provided herein.
[0130] In some specific embodiments, non-limiting examples of cells include: prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, cells of unicellular eukaryotic organisms, protozoan cells, cells from plants, algal cells, fungal cells, animal cells, cells from invertebrates, cells from vertebrates, cells from mammals (e.g., ungulates; rodents; non-human primates; humans; felines; dogs, etc.), etc. In some cases, the cell is a cell that is not derived from a natural organism (e.g., the cell can be a synthetic cell; also known as an artificial cell).
[0131] Depending on the host / vector system utilized, any of a number of suitable transcription and / or translation control elements, including constitutive and inducible promoters, transcription enhancer elements, transcription terminators, and the like, may be used in the recombinant expression vector.
[0132] Methods for introducing nucleic acids into host cells are known in the art, and any convenient method can be used to introduce nucleic acids (e.g., recombinant expression vectors, isolated polynucleotides, nucleic acid constructs, DEAR nucleic acid manipulation systems provided by the present invention) into cells. Suitable methods include, for example, viral infection, transfection, liposome transfection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, direct microinjection, nanoparticle-mediated nucleic acid delivery, etc.
[0133] <Reagents, Kits, Pharmaceutical Compositions>
[0134] In some embodiments of the present invention, the present invention provides a reagent or kit comprising the DEAR nucleic acid manipulation system of the present invention, the isolated polynucleotide of the present invention, the nucleic acid construct of the present invention, the vector of the present invention or the cell of the present invention.
[0135] In some embodiments of the present invention, the present invention provides a pharmaceutical composition comprising the DEAR nucleic acid manipulation system of the present invention, the isolated polynucleotide of the present invention, the nucleic acid construct of the present invention, the vector of the present invention, or the cell of the present invention, and optionally, a pharmaceutically acceptable carrier.
[0136] <Methods and Uses for Modifying Target Nucleic Acids>
[0137] The present invention provides a method for modifying a target nucleic acid, comprising contacting the target nucleic acid with the DEAR nucleic acid manipulation system, the isolated polynucleotide, the nucleic acid construct, the vector, the cell, the reagent, or the kit described herein. In some embodiments, the contacting results in modification of the target nucleic acid by the DEAR nucleic acid manipulation system.
[0138] The present invention provides uses of the DEAR nucleic acid manipulation system, the isolated polynucleotide, the nucleic acid construct, the vector, and the cell of the present invention in modifying target nucleic acids or preparing reagents or kits for modifying target nucleic acids.
[0139] In some specific embodiments, the modification is cleavage of the target nucleic acid.In some specific embodiments, the target nucleic acid is selected from the group consisting of: DNA, RNA, genomic DNA, and extrachromosomal DNA.
[0140] In some specific embodiments, the contacting occurs in vitro or in vivo. In some specific embodiments, the contacting occurs inside a cell or outside a cell.
[0141] In some specific embodiments, the cell is a eukaryotic cell or a prokaryotic cell.
[0142] In some more specific embodiments, the cell is selected from the group consisting of: a plant cell, a fungal cell, a mammalian cell, a reptile cell, an insect cell, an avian cell, a fish cell, a parasite cell, an arthropod cell, an invertebrate cell, a vertebrate cell, a rodent cell, a mouse cell, a rat cell, a primate cell, a non-human primate cell, and a human cell.
[0143] In some more specific embodiments, said contacting results in genome editing.
[0144] In some embodiments, the contacting comprises introducing the DEAR nucleic acid manipulation system into the cell.
[0145] Example
[0146] The present invention will be further described in detail below in conjunction with specific embodiments. The examples provided are only for illustrating the present invention and are not intended to limit the scope of the present invention. The examples provided below can serve as a guide for further improvements by those skilled in the art and are not intended to limit the present invention in any way.
[0147] Unless otherwise specified, the experimental methods in the following examples are conventional methods and were performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials and reagents used in the following examples, unless otherwise specified, were all commercially available.
[0148] Example 1. Screening of RNA sequences for C-type group II introns
[0149] To screen for C-type Class II introns, this example used 92 C-type Class II introns from public databases to construct sequence and structure covariance models for the RNA sequences of the conserved I-III and V-VI domains, respectively. Furthermore, a hidden Markov model of the amino acid sequence characteristics of potential IEP proteins was constructed. Typically, the length of a C-type Class II intron does not exceed 4000 nt, so a 4000 bp recognition window was set in this example. Potential C-type Class II introns needed to meet the requirements of both the I-III and V-VI domains with high confidence within a 4000 bp range. If no IEP protein could be identified in the IV domain, it was considered a C-type Class II intron without an ORF.
[0150] Based on the aforementioned multiple covariance model, this example identified 5,684 C-type group II introns in the Earth metagenome dataset. Active C-type group II introns should have multiple, highly similar copies within the genome of the same strain. Therefore, this example clustered highly similar candidate C-type group II introns within the metagenomes of the same species, identifying a total of 469 potentially active C-type group II introns with multiple copies.
[0151] In order to screen stable ORF-free ribozymes, this example ranked candidate C-type second-class introns (GIIC introns) according to the predicted thermal stability of the secondary structure. At the same time, this example also used RNA secondary structure prediction to further screen candidate C-type second-class introns with conserved secondary structures in the substrate recognition region (TRS). Finally, DEAR1~9 were selected as the DEAR nucleic acid manipulation system, and the substrate cleavage activity was verified. The secondary structure prediction of DEAR1~9 obtained by screening (using RNAfold WebServer to predict RNA secondary structure: http: / / rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi) and its domain annotation are shown in Figures 1A to 1I As can be seen, the secondary structures of DEAR1-9 are quite similar, all consisting of domains I-VI. Each domain exists as a stem-loop structure and is naturally separated. The programmable TRS region is located in the top loop region of domain I and is used to recognize nucleic acid substrates. The sequences of DEAR1-9 are shown in Table 1 below, where the underlined and bolded parts are the TRS.
[0152] Table 1:
[0153]
[0154]
[0155]
[0156] Example 2. RNA Preparation Method
[0157] First, the DNA sequences corresponding to DEAR1-9 screened in Example 1 were synthesized, and a T7 promoter (TAATACGACTCACTATA; SEQ ID NO: 19) was added upstream of each DEAR by PCR. The PCR amplification products were purified using DNA purification magnetic beads (VAHTS DNA CleanBeads, Vazyme, Cat. No. N411-01) and used as templates for in vitro transcription (IVT). The IVT reaction was performed in 30 mM Tris pH 8.1, 25 mM MgCl2, 0.01% Triton X-100, 2 mM spermidine, and 5 mM DTT, with 5 mM each NTP. RNase inhibitor (Promega, Cat. No. N2111) and T7 RNA polymerase (NEB, Cat. No. M0251S) were added according to the reagent supplier's instructions. After 4 hours of reaction at 37°C, DNase I (Promega, Catalog No. M6101) and Proteinase K (Biyuntian, Catalog No. ST533) digestion were performed sequentially to remove DNA template and protein. The transcripts were then washed and concentrated using a 100 kDa molecular weight cutoff concentrator. 8% Urea-PAGE electrophoresis was used to test RNA quality. Figure 2 As shown, nine ribozyme RNAs were successfully prepared. Compared with RNAs of known lengths, it was found that the size of each RNA ribozyme was consistent with its theoretical length.
[0158] Example 3. DEAR Nucleic Acid Manipulation System for Cleavage of Single-Stranded RNA, DNA, and Plasmids in Vitro
[0159] 1. DEAR nucleic acid manipulation system for in vitro single-stranded RNA cleavage
[0160] According to the sequence shown in Table 2 below, single-stranded RNA (ssRNA) substrates with DEAR target sequences were synthesized (the underlined and bold parts are the target sequences recognized by DEAR), and the 3' end of each single-stranded RNA substrate was labeled with -Cy5. Each DEAR (1.5 μM) was incubated with single-stranded RNA (100 nM) substrate under the conditions of 500 mM NH4Cl, 125 mMMgCl2, 40 mM MOPS 7.5, and 50°C for 1 hour for reaction. After terminating the reaction, Urea-PAGE electrophoresis was performed and the gel fluorescence signal was scanned on a fluorescence imager. The results are shown in Figure 3 ,like Figure 3As shown, the products obtained by cutting ssRNA are at the bottom, and it can be seen that DEAR1~DEAR9 can all cut single-stranded RNA. Figure 3 In the figure, I represents the input ssRNA and C represents the cleavage product.
[0161] Table 2:
[0162]
[0163] 2. Verification of ssRNA targeting region of DEAR nucleic acid manipulation system
[0164] DEAR1 to 6 (1.5 μM each) were incubated with a single-stranded RNA (100 nM) substrate that cannot pair with the TRS region (the substrate used for DEAR1 is the sequence shown in SEQ ID NO: 21; the substrate used for DEAR2 is the sequence shown in SEQ ID NO: 23; the substrate used for DEAR3 is the sequence shown in SEQ ID NO: 23; the substrate used for DEAR4 is the sequence shown in SEQ ID NO: 22; the substrate used for DEAR5 is the sequence shown in SEQ ID NO: 21; the substrate used for DEAR6 is the sequence shown in SEQ ID NO: 23) under the conditions of 10 mM KCl, 50 mM MgCl2, 40 mM MOPS 7.5, and 37°C. Samples were taken at different time points (0 min, 5 min, 10 min, 30 min, 60 min, and 120 min). After terminating the reaction, Urea-PAGE electrophoresis was performed, and the gel fluorescence signal was scanned on a fluorescence imager. Gel image results are shown in Figure 2. Figure 4 ,like Figure 4 As shown, the product obtained by cleaving ssRNA is below the substrate and cannot be cleaved when the substrate cannot pair with the TRS region.
[0165] 3. DEAR nucleic acid manipulation system for in vitro single-stranded DNA cleavage
[0166] According to the sequences shown in Table 3 below, single-stranded DNA (ssDNA) substrates with the corresponding target sequences of DEAR1 to 6 were synthesized (the underlined and bold parts are the target sequences recognized by DEAR), and their 3' ends were labeled with -Cy5. Subsequently, each DEAR (1.5μM) and the corresponding single-stranded DNA (100nM) substrate were incubated under the conditions of 500mM NH4Cl, 125mM MgCl2, 40mM MOPS 7.5, and 50°C for reaction, and samples were taken at different time points (0min, 5min, 10min, 20min, 40min, 60min, 120min). After terminating the reaction, Urea-PAGE electrophoresis was used, and the gel fluorescence signal was scanned on a fluorescence imager. For gel images and efficiency curve results, see Figure 5,like Figure 5 As shown, the product obtained by cutting ssDNA is below the substrate, which shows that DEAR1~DEAR6 can all cut single-stranded DNA.
[0167] Table 3:
[0168]
[0169] 4. Verification of ssDNA targeting region of DEAR nucleic acid manipulation system
[0170] Each of DEAR1 to 6 (1.5 μM) was incubated with single-stranded DNA (100 nM) substrates that can and cannot pair with the TRS region (the substrates used for DEAR1 are sequences shown in SEQ ID NO: 29 and SEQ ID NO: 30, respectively; the substrates used for DEAR2 are sequences shown in SEQ ID NO: 30 and SEQ ID NO: 32, respectively; the substrates used for DEAR3 are sequences shown in SEQ ID NO: 31 and SEQ ID NO: 32, respectively; the substrates used for DEAR4 are sequences shown in SEQ ID NO: 32 and SEQ ID NO: 31, respectively; the substrates used for DEAR5 are sequences shown in SEQ ID NO: 33 and SEQ ID NO: 30, respectively; and the substrates used for DEAR6 are sequences shown in SEQ ID NO: 34 and SEQ ID NO: 32, respectively) under the conditions of 500 mM NH4Cl, 125 mM MgCl2, 40 mM MOPS 7.5, and 50°C for 1 h. After terminating the reaction, perform Urea-PAGE electrophoresis and scan the gel fluorescence signal on a fluorescence imager. Figure 6 ,like Figure 6 As shown, T represents the paired substrate, T* represents the cleavage product of the paired substrate, N represents the unpaired substrate, N* represents the cleavage product of the unpaired substrate, M represents the marker, and the product obtained by cleaving ssDNA is below the substrate. It can be seen that DEAR1~DEAR6 can all cause cleavage of single-stranded DNA, and cannot be cleaved when the substrate cannot pair with the TRS region.
[0171] 5. Verification of ssDNA cleavage sites in the DEAR nucleic acid manipulation system
[0172] According to the sequences shown in Table 4 below, single-stranded DNA (ssDNA) substrates with target sequences corresponding to DEAR1 to 6 were synthesized (the underlined and bold parts are the target sequences recognized by DEAR), and their 3' ends were labeled with -Cy5. Subsequently, each DEAR (1.5 μM) and the corresponding single-stranded DNA (100 nM) substrate were incubated for 24 hours under the conditions of 500 mM NH4Cl, 125 mMMgCl2, 40 mM MOPS 7.5, and 50°C for reaction. After terminating the reaction, Urea-PAGE electrophoresis was used, and the gel fluorescence signal was scanned on a fluorescence imager. For gel images, see Figure 7 ,like Figure 7 As shown, the larger triangle indicates the primary cleavage site, the smaller triangle indicates the secondary cleavage site, I indicates the input substrate, Dr1-6 indicates the cleavage products of DEAR1-6, L is a ladder generated by random digestion of ssDNA using DNase I (Promega, Cat. No. M6101), which is used to indicate the product length, M indicates a marker, and the product obtained by cleaving ssDNA is below the substrate. It can be seen that the primary cleavage site is located 0-1 nt downstream of the 3' end of the TRS pairing region.
[0173] Table 4:
[0174]
[0175] 6. Optimization of DNA cleavage conditions for DEAR1
[0176] DEAR1 (1.5 μM) and single-stranded DNA 1X-DEAR1 (SEQ ID NO: 35; 100 nM) substrate were incubated at different concentrations of monovalent and divalent ions and temperatures, and samples were taken at different time points (0 min, 5 min, 10 min, 20 min, 40 min, 1 h, 2 h, 4 h, 8 h, 12 h, 24 h). After terminating the reaction, Urea-PAGE electrophoresis was performed and the gel fluorescence signal was scanned on a fluorescence imager. Gel images and efficiency curve results are shown in the following table. Figures 8 and 9 ,like Figures 8 and 9 As shown, the product obtained by cutting DNA is below the substrate, which shows that DEAR1 prefers K +, and the lower the concentration, the stronger the activity, which is different from the previously reported second-class intron (N.Toor, KSKeating, SDTaylor, AMPyle, Crystal structure of a self-spliced group II intron. Science 320, 77-82 (2008); C.Quiroga, PHRoy, D.Centron, The S.ma.I2 class C group II intron inserts at integron attCsites. Microbiology (Reading) 154, 1341-1353 (2008).). In addition, DEAR1 is activated at higher Mg 2+ The higher the reaction temperature, the stronger the activity. It is active at 25-50℃, and the activity is better at 37-42℃. Specific reaction conditions: Figure 8 A, 150 mM KCl, 10 / 50 / 125 mM MgCl2, 40 mM MOPS 7.5, 37 °C; Figure 8 Medium B, 10 / 150 / 500 mM KCl, 50 mM MgCl2, 40 mM MOPS 7.5, 37 °C; Figure 8 C, 10 / 150 / 500 mM NH4Cl, 50 mM MgCl2, 40 mM MOPS 7.5, 37 °C; Figure 8 D in, 10 / 150 / 500 mM NaCl, 50 mM MgCl2, 40 mM MOPS7.5, 37 °C; Figure 8 E in, 10 mM KCl, 50 mM MgCl2, 40 mM MOPS 7.5, 25 / 37 / 42 / 50 / 60℃.
[0177] 7. Comparison of DNA cleavage efficiency between DEAR1 and RNA-guided proteases
[0178] DEAR1 (1.5 μM) and single-stranded DNA 1X-DEAR1 (SEQ ID NO: 35; 100 nM) substrate were incubated under the conditions of 10 mM KCl, 50 mM MgCl2, 40 mM MOPS 7.5, and 37°C for reaction, and samples were taken at different time points (0 min, 10 min, 30 min, 1 h, 2 h, 4 h, 8 h, 16 h). For the CRISPR-Cas nuclease system used, the reaction system was configured with a ratio of RNP:DNA=15:1 according to the method described in A.Sun et al.,Thecompact Caspi(Cas12l)'bracelet'provides a unique structural platform for DNA manipulation.Cell Res 33,229-244(2023), CATsuchida et al.,Chimeric CRISPR-CasX enzymes and guide RNAs for improved genome editing activity.Mol Cell 82,1199-1209e1196(2022) and M.Jinek et al.,A programmable dual-RNA-guided DNAendonuclease in adaptive bacterial immunity.Science 337,816-821(2012)., and samples were taken at different time points (0 min, 10 min, 30 min, 1 h, 2 h, 4 h, 8 h, 16 h). After terminating the reaction, perform Urea-PAGE electrophoresis and scan the gel fluorescence signal on a fluorescence imager. For gel images and efficiency curve results, see Figure 10 ,like Figure 10 As shown, the product obtained by cutting DNA is below the substrate. It can be seen that the cutting efficiency of DEAR1 is close to that of SpyCas9 and AbCasπ1, and higher than that of PlmCasX.
[0179] 8. DEAR nucleic acid manipulation system in vitro plasmid cutting
[0180] A plasmid with the target sequence corresponding to DEAR1 (TGTCTTAAGACA; SEQ ID NO: 41) was designed and synthesized (the backbone was the pUC19 plasmid purchased from Addgene, Plasmid #50005). DEAR1 (1.5 μM) and plasmid substrate (0.03 μM) were incubated under the conditions of 150 mM KCl, 10 mM MgCl2, 40 mM MOPS 7.5, and 37°C for reaction, and samples were taken at different time points (0 h, 3 h, 8 h, and 24 h). After terminating the reaction, agarose gel electrophoresis was performed and the gel image was obtained by taking a photo on a UV imager. The results are shown in FIG. Figure 11 As shown, L represents: the plasmid treated with EcoRI (NEB, Catalog No. R0101V), showing a linear double-stranded state; OC represents: the plasmid treated with Nt.BspQI (NEB, Catalog No. R0644S), showing an open circular state; SC represents: the untreated plasmid, showing a supercoiled state; 0, 3, 8, 24 represent the time (in hours) for cutting the plasmid with DEAR1. Figure 11 The substrate plasmid is in a supercoiled state, and the product obtained by the cleavage reaction is in an open ring state. Above the substrate, it can be seen that DEAR1 can cut the plasmid.
[0181] Example 4. Plasmid interference in E. coli cells
[0182] 1. Construction of targeting plasmid
[0183] The ccdB toxic gene inducible expression plasmid with the corresponding target sequences of DEAR1-3 at the plasmid replication origin (ori) was used as the targeting plasmid (addgene sequence number: 69056).
[0184] 2. Construction of DEAR expression plasmid
[0185] The DEAR expression plasmid uses the J23119 promoter (SEQ ID NO: 42: TTGACAGCTAGCTCAGTCCTAGGTATAATACTAGT) to drive expression of each DEAR sequence. The DEAR expression plasmid was constructed as follows: the J23119 promoter was linked to DEAR1-3, and then the entire DEAR1-3 sequence linked to the J23119 promoter was inserted into the pCDFDuet1 plasmid vector (Novagen Catalog No. 71340-3) via homologous recombination, completely replacing the sequence between 410 and 3765 of the plasmid.
[0186] 3. Construction of CRISPR-Cas nuclease system
[0187] In the CRISPR-Cas nuclease expression plasmid, the Trc promoter (its specific sequence is (SEQ ID NO: 43): TTGACAATTAATCATCCGGCTCGTATAATG) was used to drive the expression of Cas9 nuclease, and the J23119 promoter (its specific sequence is the same as above) was used to drive the expression of its corresponding guide nucleic acid (sgRNA) sequence. Figure 12 The sgRNA expressed by the negative control group (labeled as PC, i.e., PC group) contained a 20-base target sequence (its specific sequence is (SEQ ID NO: 44): GCGATAAGTCGTGTCTTACC), and the targeted plasmid was cut under the guidance of the sgRNA; the negative control group ( Figure 12 The sgRNA expressed by the target plasmid (labeled NC in the NC group) does not contain the 20-base target sequence and cannot cleave the targeting plasmid. Its construction method is as follows: After connecting the Trc promoter to the Cas9 sequence, the J23119 promoter and sgRNA are connected. The Trc-Cas9-J23119-sgRNA sequence is then inserted into the pCDFDuet1 plasmid vector (Novagen Catalog No. 71340-3) by homologous recombination, completely replacing the sequence between 410 and 3765 of the plasmid.
[0188] 4. Plasmid interference detection in E. coli cells
[0189] The targeting plasmid in step 1 was combined with the different expression plasmids constructed in steps 2 and 3 (DEAR1-3 expression plasmids, CRISPR-Cas nuclease system expression plasmids) and co-introduced into Escherichia coli BW25141 strain (CGSC strain deposit number: 7635). After a certain period of incubation, bacterial samples were taken and cultured on plates containing ccdB inducer (10mM arabinose, Biotechnology Catalog No.: A610071) and targeting plasmid resistance plates (ampicillin). Figure 12 In Figure A, when bacteria contain only the targeting plasmid, they survive and form plaques on ampicillin plates, but fail to grow on ccdB-induced expression plates (BC group). This pattern is essentially identical to that observed when the expression plasmid is transfected but does not cleave the targeting plasmid (NC group, expressing Cas9 without cleaving ccdB). When the expression plasmid cleaves the targeting plasmid (PC group, expressing Cas9 that cleaves ccdB; DEAR1-DEAR3 groups, expressing the corresponding intronic RNA sequences), the ccdB toxic gene is unable to express normally, allowing the bacteria to survive on ccdB-induced expression plates. Simultaneously, due to cleavage of the targeting plasmid, the bacteria lose ampicillin resistance and die on ampicillin plates. Bacterial plating results and analysis of ccdB gene expression levels indicate that DEAR1-DEAR3 can all cleave the plasmid within E. coli cells.
[0190] 5. Further verification of plasmid interference in DEAR1 E. coli cells
[0191] The DEAR1 expression plasmid was subjected to PCR using primers GGATGAGTTTGCAAACAAAGTCCTTTCTGCCG (SEQ ID NO: 45) and AGGACTTTGTTTGCAAACTCATCCAATGATACCTAGC (SEQ ID NO: 46) and constructed by homologous recombination with a ΔTRS mutant expression plasmid, denoted as Dr1_ΔTRS (expression plasmid constructed by deleting 6 nucleotides of the TRS sequence of DEAR1), as one of the expression plasmids;
[0192] Use primers for the plasmid used in the NC group in step 3:
[0193] AGTACAGCATCGGCCTGGCCATCGGCACCAACTCTGTGG (SEQ ID NO: 47); and GGCCAGGCCGATGCTGTACTTCTTGTCAGAACCGTGGTGA (SEQ ID NO: 48) were used for PCR. The PCR products were further PCR-polymerized with primers:
[0194] CCGACTACGATGTGGACGCCATCGTGCCTCAGAGCTTTC (SEQ ID NO: 49) and GGCGTCCACATCGTAGTCGGACAGCCGGTTGATGTCC (SEQ ID NO: 50) were subjected to PCR and homologous recombination to construct a dCas9 expression plasmid (all two enzyme cleavage active sites of Cas9 were mutated and inactivated), denoted as dCas9, as one of the expression plasmids;
[0195] Use primers for the plasmids used in the PC group in step 3:
[0196] CCGACTACGATGTGGACGCCATCGTGCCTCAGAGCTTTC (SEQ ID NO: 49) and GGCGTCCACATCGTAGTCGGACAGCCGGTTGATGTCC (SEQ ID NO: 50) were subjected to PCR, and an nCas9 expression plasmid (one of the enzyme cleavage active sites of Cas9, H840, was mutated and inactivated) was constructed by homologous recombination, denoted as nCas9, as one of the expression plasmids;
[0197] The plasmid used in the PC group in step 3 was used as the wtCas9 expression plasmid without modification. The targeting plasmid constructed in step 1 and the above-mentioned different expression plasmids (dCas9, nCas9, wtCas9, Dr1_ΔTRS, and the DEAR1 expression plasmid used in step 2) were combined and introduced into Escherichia coli BW25141 strain (CGSC strain deposit number: 7635). After a certain period of incubation, a bacterial sample was taken and cultured on a targeting plasmid resistance plate (ampicillin). See Figure 13 When bacteria contained only the targeting plasmid, they survived and formed plaques on ampicillin plates (Blank group). This pattern was largely consistent with the survival and death patterns of bacteria transfected with the DEAR expression plasmid Dr1_ΔTRS, which lacks the TRS region, and expressing dCas9. The DEAR1 expression plasmid cleaved the targeting plasmid, demonstrating similar results to those expressing nCas9 and wtCas9. The bacteria lost ampicillin resistance due to cleavage of the targeting plasmid and died on the ampicillin plates. Bacterial plating results and analysis of AmpR gene expression levels demonstrated that DEAR1 can cleave the plasmid within E. coli cells, targeting the TRS region.
[0198] Example 5. Plasmid Interference Detection of DEAR4-9 in E. coli Cells
[0199] 1. Construction of targeting plasmid
[0200] A ccdB toxic gene inducible expression plasmid with the corresponding intron RNA targeting sequences of DEAR1 and DEAR4-9 at the plasmid replication origin (ori) was used as the targeting plasmid (Addgene sequence number: 69056).
[0201] 2. Construction of other DEAR expression plasmids
[0202] Basically the same method as in Example 4, the DEAR expression plasmid uses the J23119 promoter (its specific sequence is (SEQ ID NO: 42): TTGACAGCTAGCTCAGTCCTAGGTATAATACTAGT) to drive the expression of each DEAR sequence. The construction method is as follows: the J23119 promoter is connected to Dr1_ΔTRS, DEAR1, DEAR4-9 (wherein, Dr1_ΔTRS and DEAR1 are the same as in Example 4), and then the entire sequence of Dr1_ΔTRS, DEAR1, and DEAR4-9 connected to the J23119 promoter is inserted into the pCDFDuet1 plasmid vector (Novagen Catalog No.: 71340-3) by homologous recombination, respectively, replacing the entire sequence between 410 and 3765 of the plasmid.
[0203] 3. Plasmid interference detection in E. coli cells
[0204] Basically the same method as in Example 4, the targeting plasmid in step 1 and the different expression plasmids constructed in step 2 (Dr1_ΔTRS, DEAR1, DEAR4-9 expression plasmids) were combined and introduced into Escherichia coli BW25141 strain (CGSC strain deposit number: 7635). After a certain period of incubation, bacterial samples were taken and cultured on resistance plates (ampicillin) containing the targeting plasmid. Figure 14 When the expression plasmids that did not cleave the targeting plasmid were transfected (Dr1_△TRS, DEAR4, DEAR6, DEAR7, DEAR8, and DEAR9 groups: expressing the corresponding intronic RNA sequences, respectively), the plaques survived. However, when the expression plasmids cleaved the targeting plasmid (DEAR1 and DEAR5 groups: expressing the corresponding intronic RNA sequences, respectively), the bacteria died on the ampicillin plates due to loss of ampicillin resistance due to cleavage of the targeting plasmid. Bacterial plating results and analysis of Amp gene expression levels showed that DEAR1 and DEAR5 could cleave the plasmid in E. coli cells, while DEAR4, 6, 7, 8, and 9 had no plasmid cleavage activity in E. coli cells.
[0205] Example 6. Reprogramming TRS DEAR to Cleave New DNA Sites
[0206] According to the sequence shown in Table 5 below, a single-stranded DNA substrate with a new DEAR target sequence was synthesized (the underlined and bold parts are the target sequences recognized by DEAR), and its 3' end was labeled with -Cy5. DEAR1~6 (1.5μM) with the TRS sequence changed to CGAUAG were incubated with this single-stranded DNA (100nM) substrate under the conditions of 50mM MgCl2, 10mM KCl, 40mM MOPS 7.5, and 37℃ for 8h to react. After terminating the reaction, Urea-PAGE electrophoresis was performed and the gel fluorescence signal was scanned on a fluorescence imager. The results are shown in Figure 15 ,like Figure 15 As shown, the product obtained by cutting ssDNA is below the substrate, which shows that DEAR1~DEAR6 can all cut the new single-stranded DNA. Figure 15 In the figure, I represents the input ssDNA substrate, and Dr1*-Dr6* represent the cleavage products of ssDNA by DEAR1~6 after changing TRS.
[0207] Table 5:
[0208]
[0209] Example 7. Genomic DNA cleavage in mammalian cells
[0210] 1. Construction of stable transfection plasmid
[0211] Using PiggyBacTM Transposon Vector System (from System Biosciences) constructs DEAR1 targeting sequence stable transfection plasmid and DEAR stable transfection plasmid:
[0212] (1) Construction of a stable transfection plasmid containing the DEAR1 targeting sequence: The sequence of the puromycin resistance (PuroR) gene with a frameshift containing the DEAR1 targeting sequence at the N-terminus was inserted into the plasmid.
[0213] The capital letters represent the DEAR1 targeting sequence, the bold and underlined letters represent the sites that DEAR1 can specifically recognize and cleave, and DEAR1 can cleave the sense and antisense strands at this site, resulting in double-strand breaks. The lowercase letters represent the PuroR gene). The XbaI restriction site in the multiple cloning site of the PiggyBac Dual promoter PB513B-1 plasmid was inserted by homologous recombination, and the blasticidin resistance (Blasticidine S-deaminase) gene was also inserted by homologous recombination.
[0214] atggccaagcctttgtctcaagaagaatccaccctcattgaaagagcaacggctacaatcaacagcatccccatctctgaagactacagc
[0215] gtcgccagcgcagctctctctagcgacggccgcatcttcactggtgtcaatgtatatcattttactgggggaccttgtgcagaactcgtggt
[0216] gctgggcactgctgctgctgcggcagctggcaacctgacttgtatcgtcgcgatcggaaatgagaacaggggcatcttgagcccctgcg
[0217] gacggtgccgacaggtgcttctcgatctgcatcctgggatcaaagccatagtgaaggacagtgatggacagccgacggcagttgggattcgtgaattgctgccctctggttatgtgtgggagggctaa (SEQ ID NO: 53) was inserted between the NcoI and SalI restriction sites.
[0218] (2) Construction of DEAR stable transfection plasmid: The DEAR1 sequence initiated by U6 promoter and terminated by TTTTTTTT signal were respectively:
[0219]
[0220] The uppercase, bold, and underlined letters represent the U6 promoter sequence, the uppercase, non-bold and non-underlined letters represent the corresponding DNA sequence of DEAR1, and the lowercase letters represent the transcription termination signal and part of the vector backbone sequence), or the DEAR-NT sequence initiated by the U6 promoter and terminated by the TTTTTTTT signal:
[0221] The uppercase, bold, and underlined letters are the U6 promoter sequence, the uppercase, non-bold and non-underlined letters are the DEAR-NT (DNA sequence corresponding to DEAR2) sequence, and the lowercase letters are the transcription termination signal and part of the vector backbone sequence) were inserted into the PiggyBac Dual promoter PB513B-1 plasmid between the SfiI and MluI restriction sites by homologous recombination, and the hygromycin resistance (HygBR) gene was also inserted by homologous recombination:
[0222]
[0223] 2. Stable transfection and resistance screening and enrichment of DEAR targeting sequence and DEAR
[0224] HEK-293T (ATCC CRL-11268) cells were cultured in DMEM high glucose medium containing 10% fetal bovine serum at 37°C and 5% CO2 until the logarithmic phase, digested with 0.25% trypsin, washed twice with PBS (pH 7.0-7.2), and resuspended in Opti-MEM. TM (Gibco, Catalog No.: 31985070) culture medium, and adjust the cell density to 5×10 4 / μL, add 2μg Integration PB transposase plasmid (System Biosciences) and 2μg DEAR1 targeting sequence stable transfection plasmid to 20μL cell suspension, and electroporate the cell suspension at 450V (Celetrix biotechnologies, model: LE+). The electroporated cells are added to DMEM high-glucose medium containing 10% fetal bovine serum. After 24 hours of electroporation, the medium is replaced with a medium containing 10μg / mL Blasticidin. Drug selection is carried out for one week, during which the cells are passaged according to their growth status. After the cells are stable, a stably transfected cell line containing the DEAR1 targeting sequence is obtained. The same method was used to electroporate 2 μg of Integration PB transposase plasmid (SystemBiosciences) and 2 μg of DEAR stable plasmid (DEAR1 stable plasmid or DEAR-NT stable plasmid) into the stably transfected cell line containing the DEAR1 targeting sequence. Twenty-four hours after electroporation, the culture medium was replaced with a medium containing 50 μg / mL Hygromycin B and the cells were selectively conditioned for one week. During this period, the cells were passaged according to their growth status. After the cells stabilized, the culture medium was replaced with a medium containing 10 μg / mL Puromycin and selectively conditioned for one week.
[0225] For stable cell lines containing the DEAR1 targeting sequence, the PuroR gene integrated in the cells is in a frameshift state and cannot express the correct protein, thus not having resistance to Puromycin. DEAR1 (DEAR1 stable plasmid) can cut the DEAR1 targeting sequence to cause DNA double-strand breaks. The insertion or deletion mutation introduced by the break repair can restore the frameshifted PuroR gene to normal expression, resulting in cell survival under Puromycin screening; while DEAR-NT (DEAR-NT stable plasmid) cannot cut the DEAR1 targeting sequence, and the cells cannot express the correct PuroR gene, resulting in cell death under Puromycin screening. Figure 16 As shown, cells stably transfected with DEAR1 (DEAR1 stable plasmid) survived, whereas cells stably transfected with DEAR-NT (DEAR-NT stable plasmid) died.
[0226] 3. Next-generation sequencing verifies that DEAR1 cleaves genomic DNA in mammalian cells
[0227] for Figure 16 The genome of the cells that survived (stable transfection with DEAR1, i.e., stable transfection with DEAR1 plasmid) was extracted and the DEAR1 target sequence was sequenced using the TIANSeq Fast DNA Library Kit (Illumina) for next-generation sequencing. The next-generation sequencing was performed by Novogene. The next-generation sequencing data were analyzed online using the CRISPResso2 website. The results are as follows: Figure 17 As shown in A, 47.14% of the reads have mutations. Figure 17 B shows a sequence alignment of reads near the first and second cleavage sites in the DEAR1 targeted sequence. Both insertions and deletions occur near the DEAR1 cleavage site (dashed line). The sequencing data demonstrate that DEAR1 specifically cleaves double-stranded genomic DNA in mammalian cells.
Claims
1. A DEAR nucleic acid manipulation system, wherein: The DEAR nucleic acid manipulation system comprises an RNA molecule, the nucleotide sequence of which is the nucleotide sequence shown in any one of SEQ ID NOs: 1 to 9, wherein "NNNNNN" in the nucleotide sequence is a substrate recognition region, wherein N is A, U, G or C.
2. A DEAR nucleic acid manipulation system, wherein: The DEAR nucleic acid manipulation system comprises an RNA molecule of a C-type second intron derived from bacteria, wherein the RNA molecule comprises a substrate recognition region that hybridizes with a target sequence in a target nucleic acid, wherein the C-type second intron is a C-type second intron in which there is no open reading frame encoding an intron-encoded protein in the IV domain. Wherein, the nucleotide sequence of the RNA molecule is selected from any one of the following: (i) a nucleotide sequence as shown in any one of SEQ ID NOs: 1 to 9; (ii) a nucleotide sequence that is the reverse complement of the sequence shown in any one of SEQ ID NOs: 1 to 9; The "NNNNNN" in the nucleotide sequence is the substrate recognition region, where N is A, U, G or C.
3. The DEAR nucleic acid manipulation system according to claim 2, wherein: The substrate recognition region has a length of 6 nucleotides.
4. The DEAR nucleic acid manipulation system according to claim 3, wherein: The substrate recognition region is programmable to hybridize to different target sequences.
5. The DEAR nucleic acid manipulation system according to claim 2, wherein: The target nucleic acid is DNA or RNA.
6. The DEAR nucleic acid manipulation system according to any one of claims 2 to 5, wherein The primary cleavage site of the DEAR nucleic acid manipulation system is 0-1 nt downstream of the 3' end of the target sequence in the target nucleic acid.
7. An isolated polynucleotide, wherein The polynucleotide comprises a nucleotide sequence encoding the DEAR nucleic acid manipulation system according to any one of claims 1 to 6.
8. A nucleic acid construct, wherein The nucleic acid construct comprises the isolated polynucleotide of claim 7.
9. A vector, wherein The vector comprises the isolated polynucleotide according to claim 7 or the nucleic acid construct according to claim 8.
10. A cell, wherein: The cell comprises the DEAR nucleic acid manipulation system according to any one of claims 1 to 6, the isolated polynucleotide according to claim 7, the nucleic acid construct according to claim 8, or the vector according to claim 9.
11. A reagent, wherein The reagent comprises the DEAR nucleic acid manipulation system according to any one of claims 1 to 6, the isolated polynucleotide according to claim 7, the nucleic acid construct according to claim 8, the vector according to claim 9 or the cell according to claim 10.
12. A kit, wherein: The kit comprises the DEAR nucleic acid manipulation system according to any one of claims 1 to 6, the isolated polynucleotide according to claim 7, the nucleic acid construct according to claim 8, the vector according to claim 9 or the cell according to claim 10.
13. A pharmaceutical composition, wherein The pharmaceutical composition comprises the DEAR nucleic acid manipulation system according to any one of claims 1 to 6, the isolated polynucleotide according to claim 7, the nucleic acid construct according to claim 8, the vector according to claim 9 or the cell according to claim 10; and a pharmaceutically acceptable carrier.
14. A method for modifying a target nucleic acid, the method comprising the step of contacting the target nucleic acid with the DEAR nucleic acid manipulation system according to any one of claims 1 to 6, the isolated polynucleotide according to claim 7, the nucleic acid construct according to claim 8, the vector according to claim 9, the cell according to claim 10, the reagent according to claim 11, or the kit according to claim 12.
15. Use of the DEAR nucleic acid manipulation system according to any one of claims 1 to 6, the isolated polynucleotide according to claim 7, the nucleic acid construct according to claim 8, the vector according to claim 9, or the cell according to claim 10 in modifying a target nucleic acid or preparing a reagent or kit for modifying a target nucleic acid.