Reverse transcriptase, DNA editing system, method for editing target DNA by using the same, and method for producing cell in which target DNA is edited
A split reverse transcriptase system for prime editing, active only when its fragments are combined, addresses inefficiencies in existing technologies by enhancing editing specificity and efficiency, reducing off-target events.
Patent Information
- Application Number
- JP2024081112
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-17
- Publication Date
- 2025-11-28
AI Technical Summary
Existing prime editing technologies require improvements in editing efficiency and specificity, particularly in the reverse transcriptase component, to reduce off-target events.
Development of a split-type reverse transcriptase that is only active when its N-terminal and C-terminal fragments are present, integrated into a DNA editing system with a prime editing guide RNA and Cas protein, enhancing editing specificity and efficiency.
The split reverse transcriptase system significantly improves the specificity and efficiency of prime editing by ensuring activity only when all fragments are present, reducing off-target effects and enhancing DNA editing accuracy.
Smart Images

Figure 2025174633000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a reverse transcriptase, a DNA editing system, and a method for editing target DNA using the same and a method for producing a cell in which the target DNA has been edited. [Background technology]
[0002] Genome editing is a technique that specifically cuts target DNA using artificial DNA-cleaving enzymes (e.g., zinc-finger nucleases (ZFNs) and transcription activator-like effector nucleases (TALENs)) or RNA-guided DNA-cleaving enzymes (e.g., clustered regularly interspaced short palindromic repeats and CRISPR-associated protein (CRISPR-Cas)), and then utilizes the organism's DNA repair system to modify genes. Genome editing can theoretically be applied to all biological species and cell lines that contain DNA, and it is possible to edit (substitution, deletion, and insertion) anything from a single base to long chains. Therefore, it is used in a variety of fields, not only in agriculture and industry, but also in medical fields such as gene therapy.
[0003] Among the many genome editing techniques, prime editing (PE) is a technique that specifically enables the substitution of all bases, the insertion / deletion (indel) of short strands, or a combination of these (Anzalone, AV et al., Nature 576, pp. 149-157, 2019 (Non-Patent Document 1)).
[0004] Prime editing typically involves the use of a Prime Editor, which is a fusion of a partially nuclease-depleted nickase-type Cas protein (nCas, e.g., nicking Cas9 (nCas9)) and a reverse transcriptase (RTase), and a prime editing guide RNA (pegRNA). Target DNA editing is thought to occur via the mechanism shown in Figures 1A-1C. Specifically, the spacer sequence of the pegRNA first recognizes the pegRNA recognition sequence on the target DNA, binds to the target DNA, and then guides a Cas protein (nCas9 in Figures 1A-1C) to the target site. The guided Cas protein (nCas9) then recognizes a protospacer adjacent motif (PAM) and creates a nick in the target strand where the PAM is located. The 3'-terminal sequence of the target strand released by the binding and nick of the pegRNA binds to the primer binding site (PBS) of the pegRNA (Figure 1B). Then, a reverse transcriptase (RTase) fused to a Cas protein (nCas9) performs reverse transcription using the template sequence (reverse transcription template (RTT)) contained in the 5' end of the pegRNA following the PBS. This creates an equilibrium state in the target DNA, where the edited strand (the strand synthesized by the reverse transcriptase) is released at the 3' flap (Figure 1C (a)) and the unedited strand is released at the 5' flap (Figure 1C (b)). The 5' flap is excised by intracellular enzymes, repairing the nick. The edited strand is then incorporated into the target DNA, replacing the target site; however, the incorporated edited strand is not complementary. Therefore, DNA repair using the edited strand as a template allows the desired edited DNA to be obtained (Figure 1C (c)).
[0005] Non-Patent Document 1 also describes improved prime editing systems, such as PE2, which introduces five mutations into the reverse transcriptase to increase its processing and binding capacity, and PE3, which introduces an additional nicking guide RNA (ngRNA) to create a nick in the non-edited strand, increasing the probability that the edited strand will be adopted during the DNA repair stage, and reports that these systems improve editing efficiency.
[0006] Furthermore, in an effort to improve the editing efficiency of prime editing, it has been reported that adding an 8-nt linker and a 51-nt RNA motif called tmpknot (trimmed mpknot) to the 3' end of pegRNA prevents the free 3' end from being cleaved by exonuclease (Nelson, JW et al., Nat Biotechnol 40, p. 402-410, 2022 (Non-Patent Document 2)).
[0007] Another technology known as PE3max has been reported, in which two mutations are introduced into nCas9, a double nuclear localization signal (NLS) is placed between nCas9 and reverse transcriptase, the codons of the reverse transcriptase are optimized, and a c-Myc NLS is placed at the C-terminus of Prime Editor (Chen, PJ et al., Nat Rev Genet 24, pp. 161-177, 2023 (Non-Patent Document 3)). Furthermore, it has been found that the efficiency can be further improved by using this in combination with the pegRNA described in Non-Patent Document 2.
[0008] It has also been reported that target DNA can be edited by prime editing even when nCas9 and reverse transcriptase are not fused but are expressed separately (Liu, B. et al., Nat Biotechnol 40, pp. 1388-1393, 2022 (Non-Patent Document 4)). Furthermore, in addition to nickase-type Cas9 proteins, it has also been reported that target DNA can be edited by a mechanism essentially similar to that of prime editing, for example, by double strand break (DSB) using a Cas9 protein that does not lose nuclease activity (having nuclease activity) (Li, X. et al., Nature Communications 14, Article number: 305, 2023, https: / / doi.org / 10.1038 / s41467-023-35870-0 (Non-Patent Document 5)). [Prior art documents] [Non-patent literature]
[0009] [Non-Patent Document 1] Anzalone,AV et al.,Nature 576,p.149-157,2019 [Non-patent document 2] Nelson, JW et al., Nat Biotechnol 40, p. 402-410, 2022 [Non-patent document 3] Chen,PJ et. al.,Nat Rev Genet 24,p.161-177,2023 [Non-patent document 4] Liu, B. et. al.,Nat Biotechnol 40,p.1388-1393,2022 [Non-Patent Document 5] Li, X. et. al.,Nature Communications 14,Article number:305,2023,https: / / doi.org / 10.1038 / s41467-023-35870-0 Summary of the Invention [Problem to be solved by the invention]
[0010] However, research into prime editing, which has only recently been developed, has not yet progressed sufficiently, and further improvements in editing efficiency and systems to suppress off-target events are needed. In particular, improvements to reverse transcriptase have not been sufficiently investigated.
[0011] The present invention has been made in consideration of the problems associated with the above-mentioned conventional technology, and aims to provide a novel reverse transcriptase that is useful for prime editing and that is active only under specific conditions, more specifically, a novel split-type reverse transcriptase that is active only when all of its constituent fragments are present, a DNA editing system containing the same, and methods for editing target DNA using these and methods for producing cells in which target DNA has been edited. [Means for solving the problem]
[0012] As a result of extensive research aimed at achieving the above-mentioned objective, the present inventors have designed a new reverse transcriptase that is split into N- and C-terminal fragments and whose activity is restored only when all of the resulting fragments are present. Using this as the reverse transcriptase in a prime editing system is expected to improve the specificity of prime editing. The present inventors first split the reverse transcriptase into two fragments, an N-terminal fragment and a C-terminal fragment, at multiple positions, and compared the activity of each fragment when used alone with the activity of both fragments in the presence of a split GFP. As a result, they found that reverse transcription activity was not observed in the N- or C-terminal fragments alone, but was specifically restored only when both fragments were present, leading to the completion of the present invention.
[0013] The present invention is provided based on these findings in the following aspects. [1] The reverse transcriptase is a split reverse transcriptase containing two split N-terminal fragments and C-terminal fragments, The N-terminal fragment and the C-terminal fragment are the following (a) to (c): (a) an N-terminal fragment and a C-terminal fragment obtained by splitting a polypeptide comprising the amino acid sequence set forth in SEQ ID NO: 1 at any one position within the region of positions 114 to 133 or any one position within the region of positions 167 to 186 of the amino acid sequence set forth in SEQ ID NO: 1; (b) an N-terminal fragment and a C-terminal fragment obtained by dividing a polypeptide comprising an amino acid sequence in which one or more amino acid residues in the amino acid sequence set forth in SEQ ID NO: 1 have been substituted, deleted, inserted, and / or added at any one position within the region corresponding to positions 114 to 133 or any one position within the region corresponding to positions 167 to 186 of the amino acid sequence set forth in SEQ ID NO: 1; (c) an N-terminal fragment and a C-terminal fragment obtained by splitting a polypeptide comprising an amino acid sequence having 90% or more homology to the amino acid sequence set forth in SEQ ID NO: 1 at any one position within a region corresponding to positions 114 to 133 or any one position within a region corresponding to positions 167 to 186 of the amino acid sequence set forth in SEQ ID NO: 1; At least one selected from the group consisting of the reverse transcriptase activity is regenerated by the presence of the N-terminal fragment and the C-terminal fragment; Reverse transcriptase. [2] (A) to (C) below: (A) At least one selected from the group consisting of the reverse transcriptase according to [1], an expression vector for the reverse transcriptase, and a polynucleotide encoding the reverse transcriptase; (B) at least one Cas protein selected from the group consisting of Cas proteins (Cas) having nuclease activity and nickase-type Cas proteins (nCas) that are partially deficient in nuclease activity, an expression vector for the Cas protein, and at least one polynucleotide encoding the Cas protein; (C) at least one selected from the group consisting of a prime editing guide RNA (pegRNA) for the Cas protein, an expression vector for the pegRNA, and a polynucleotide encoding the pegRNA; A DNA editing system comprising: [3] (D) at least one nicking guide RNA (ngRNA) for the Cas protein, wherein the ngRNA binds to an ngRNA recognition sequence that is a sequence different from the pegRNA recognition sequence to which the pegRNA binds; an expression vector for the ngRNA; and a polynucleotide encoding the ngRNA; The DNA editing system described in [2], further comprising: [4] A method for editing target DNA, [2] or [3], contacting the DNA editing system with target DNA and editing the target site of the target DNA; the pegRNA comprises, in order from the 5' side, a spacer sequence, a template sequence, and a primer binding sequence, the spacer sequence is a sequence that binds to a pegRNA recognition sequence that includes a complementary sequence of the target site, the template sequence is a sequence that serves as a template for the reverse transcriptase, and the primer binding sequence is a sequence that binds to a primer sequence on the 5' side of the target site. [5] A method for producing cells with edited target DNA,
[0023] The DNA editing system according to [2] or [3] is introduced into a cell, contacted with a target DNA, and edits the target site of the target DNA. the pegRNA comprises, in order from the 5' side, a spacer sequence, a template sequence, and a primer binding sequence, the spacer sequence is a sequence that binds to a pegRNA recognition sequence that includes a complementary sequence of the target site, the template sequence is a sequence that serves as a template for the reverse transcriptase, and the primer binding sequence is a sequence that binds to a primer sequence on the 5' side of the target site. [Effects of the Invention]
[0014] According to the present invention, it is possible to provide a novel split reverse transcriptase that is active only when all of its constituent fragments are present, a DNA editing system containing the same, a method for editing target DNA using these, and a method for producing cells in which target DNA has been edited. [Brief explanation of the drawings]
[0015] [Figure 1A] FIG. 1 is a schematic diagram showing one embodiment of the conventional prime editing mechanism, showing (a) target DNA, (b) Prime Editor, and (b) pegRNA contained in the prime editing system. [Figure 1B] FIG. 1 is a schematic diagram showing one aspect of the conventional prime editing mechanism, illustrating the bound state of target DNA, Prime Editor, and pegRNA. [Figure 1C] This is a schematic diagram showing one aspect of the conventional prime editing mechanism, showing (a) the 3' flap where the edited strand is released in the target DNA, (b) the 5' flap where the non-edited strand is released in the target DNA, and their equilibrium state; and (c) edited DNA obtained by repair using the edited strand as a template after the 5' flap has been removed and the edited strand has been replaced with the target site in the target DNA. [Figure 2] This is a schematic diagram showing the structure of the vector "KS2" that expresses nCas9 under the control of the EF1a promoter, which was prepared in <Preparation of vector>. [Figure 3] FIG. 1 is a schematic diagram showing the structure of the vector "KS15" that expresses RTase under the control of a CMV enhancer / promoter, prepared in <Preparation of a vector>. [Figure 4] FIG. 1 is a schematic diagram showing the structure of the vector "#1373," which was prepared in <Preparation of vector> and expresses Prime Editor (a fusion protein of nCas9 and RTase) under the control of the EF1a promoter. [Figure 5]FIG. 1 is a schematic diagram showing the positional relationship of PAM, pegRNA recognition sequence (comp_pegRNA), and ngRNA recognition sequence (comp_ngRNA) present on the target gene RNF2 used in <Vector construction>. [Figure 6] FIG. 1 is a schematic diagram showing the structure of pegRNA used in <Vector construction>. [Figure 7] FIG. 1 is a schematic diagram showing the structures of (a) "KS2-1+ng+peg" and (b) "KS25-1+ng+peg" prepared in <Preparation of vectors>. [Figure 8] 1 is a graph showing the results of comparing the editing efficiency measured in <Vector production> between "KS25-1+ng+peg (all-in-one)" and the combination of "KS25+ng" and a vector expressing pegRNA (separate). [Figure 9] FIG. 1 is a schematic diagram showing the mechanism of DNA editing by the cooperation of split bodies (fusion proteins with split GFP), nCas9, pegRNA, and ngRNA designed in <Test Example 1>. [Figure 10] FIG. 1 is a schematic diagram showing the method of (construction of N-terminal vector) in <Test Example 1> and <Test Example 2>. [Figure 11] FIG. 1 is a schematic diagram showing the method of (preparation of C-terminal vector) in <Test Example 1> and <Test Example 2>. [Figure 12] 1 is a graph showing the editing efficiency (%) measured in <Test Example 1>. [Figure 13] 1 is a graph showing the editing efficiency (%) measured in <Test Example 2>. [Figure 14A] FIG. 1 shows the results of alignment of the amino acid sequence shown in SEQ ID NO: 1 with the amino acid sequences of reverse transcriptases derived from various organisms (positions 3 to 62 of the amino acid sequence of SEQ ID NO: 1). [Figure 14B] FIG. 1 shows the results of alignment of the amino acid sequence shown in SEQ ID NO: 1 with the amino acid sequences of reverse transcriptases derived from various organisms (positions 63 to 122 / 123 to 182 of the amino acid sequence of SEQ ID NO: 1). [Figure 14C]FIG. 1 shows the results of an alignment of the amino acid sequence shown in SEQ ID NO: 1 with the amino acid sequences of reverse transcriptases derived from various organisms (positions 183-242 / 243-302 of the amino acid sequence of SEQ ID NO: 1). [Figure 14D] FIG. 1 shows the results of an alignment of the amino acid sequence shown in SEQ ID NO: 1 with the amino acid sequences of reverse transcriptases derived from various organisms (positions 303-362 and 363-422 of the amino acid sequence of SEQ ID NO: 1). [Figure 14E] FIG. 1 shows the results of an alignment of the amino acid sequence shown in SEQ ID NO: 1 with the amino acid sequences of reverse transcriptases derived from various organisms (positions 423-482 and 483-542 of the amino acid sequence of SEQ ID NO: 1). [Figure 14F] FIG. 1 shows the results of an alignment of the amino acid sequence shown in SEQ ID NO: 1 with the amino acid sequences of reverse transcriptases derived from various organisms (positions 543-602 and 603-662 of the amino acid sequence of SEQ ID NO: 1). DETAILED DESCRIPTION OF THE INVENTION
[0016] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the drawings, but the present invention is not limited thereto. In the following description and drawings, the same or corresponding elements are designated by the same reference numerals, and redundant description will be omitted.
[0017] <Reverse transcriptase> Reverse transcriptase (sometimes referred to as "RTase" herein) is an RNA-dependent DNA polymerase that catalyzes the reverse transcription reaction that synthesizes DNA using single-stranded RNA as a template.
[0018] The reverse transcriptase of the present invention is a split-type reverse transcriptase that is split into an N-terminal fragment and a C-terminal fragment, and the presence of both the N-terminal fragment and the C-terminal fragment restores the reverse transcriptase activity that it possesses when unsplit. The N-terminal fragment refers to the N-terminal fragment of the unsplit reverse transcriptase, and the C-terminal fragment refers to the C-terminal fragment of the unsplit reverse transcriptase. The "N-terminus (also referred to as the amino terminus, NH2-terminus, N-terminal portion, or amine terminus)" refers to the free amine (-NH2) group of the first amino acid residue of a polypeptide, and the "C-terminus" (also referred to as the carboxy terminus, COOH-terminus, or C-terminal portion) refers to the free carboxy group (-COOH) of the last amino acid residue of a polypeptide.
[0019] In the reverse transcriptase of the present invention, the N-terminal fragment and the C-terminal fragment are typically N-terminal fragment and C-terminal fragment obtained by cleaving (a) a polypeptide comprising the amino acid sequence of SEQ ID NO: 1 at any one position within the region of positions 114 to 133 or any one position within the region of positions 167 to 186 of the amino acid sequence of SEQ ID NO: 1. The amino acid sequence of SEQ ID NO: 1 is the amino acid sequence of a reverse transcriptase derived from a murine leukemia virus (Moloney Murine Leukemia Virus Reverse Transcriptase (M-MLV RTase)).
[0020] As used herein, "split at any one position" refers to cleavage between adjacent amino acids (to generate the C-terminus and N-terminus). For example, combinations of an N-terminal fragment and a C-terminal fragment (N-terminal fragment / C-terminal fragment) "split at any one position within the region of positions 114 to 133" include T1 (the first T in the amino acid sequence of SEQ ID NO: 1; the same applies below) to D114 / L115 to E691, T1 to L115 / R116 to E691, T1 to R116 / E117 to E691, T1 to E117 / V118 to E691, T1 to V118 / N119 to E691, T1 to N119 / K12 0~E691, T1~K120 / R121~E691, T1~R121 / V122~E691, T1~V122 / E123~E691, T1~E 123 / D124~E691, T1~D124 / I125~E691, T1~I125 / H126~E691, T1~H126 / P127~E6 91, T1-P127 / T128-E691, T1-T128 / V129-E691, T1-V129 / P130-E691, T1-P130 / N131-E691, T1-N131 / P132-E691, or T1-P132 / Y133-E691, and similarly in other regions.
[0021] In the reverse transcriptase of the present invention, the division position of the N-terminal fragment and the C-terminal fragment is preferably any one position within the region of positions 115 to 132 of the amino acid sequence of SEQ ID NO: 1, more preferably any one position within the region of positions 118 to 129, even more preferably any one position within the region of positions 120 to 127, particularly preferably any one position within the region of positions 122 to 125, or any one position within the region of positions 168 to 185 of the amino acid sequence of SEQ ID NO: 1, more preferably any one position within the region of positions 171 to 182, even more preferably any one position within the region of positions 173 to 180, and particularly preferably any one position within the region of positions 175 to 178.
[0022] The amino acid sequence of reverse transcriptase is highly conserved among biological species, as shown in Figures 14A to 14F, which show the results of alignment of the amino acid sequences of reverse transcriptase derived from each biological (viral) species listed in Table 1 below with the amino acid sequence set forth in SEQ ID NO: 1.
[0023] [Table 1]
[0024] Furthermore, as shown in the Examples below, even split products of reverse transcriptase variants (e.g., S92R) can be cleaved at the above positions to restore reverse transcriptase activity in the presence of both the N-terminal and C-terminal fragments. Therefore, the reverse transcriptase of the present invention can be reverse transcriptases derived from other biological species or variants of reverse transcriptase (natural mutants and artificial mutants). In other words, the reverse transcriptase of the present invention also includes split products of reverse transcriptases consisting of an amino acid sequence highly homologous to the amino acid sequence set forth in SEQ ID NO: 1, as long as the reverse transcriptase activity is restored in the presence of both the N-terminal and C-terminal fragments when split at the above positions.
[0025] Thus, embodiments of the reverse transcriptase of the present invention also include an embodiment in which the N-terminal fragment and the C-terminal fragment are (b) N-terminal fragments and C-terminal fragments obtained by splitting a polypeptide comprising an amino acid sequence in which one or more amino acid residues have been substituted, deleted, inserted, and / or added in the amino acid sequence of SEQ ID NO: 1 at any one position within the region corresponding to positions 114 to 133 or any one position within the region corresponding to positions 167 to 186 of the amino acid sequence of SEQ ID NO: 1. In this embodiment (b), preferred aspects of the split positions into the N-terminal fragment and the C-terminal fragment correspond to the preferred aspects in (a), respectively.
[0026] Here, in the amino acid sequence, "an amino acid sequence in which amino acid residues have been substituted, deleted, inserted, and / or added" refers to an amino acid sequence in which amino acid residues in the amino acid sequence have been substituted, deleted, inserted, or added, or an amino acid sequence in which two or more of these have been combined. Furthermore, "multiple" refers to an integer of 30, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2. In the amino acid sequence of the polypeptide (b), "one or more" preferably refers to 1 to 30 or 1 to 20 amino acid residues (e.g., 1 to 10, 1 to 5, 1 to 3, 2 or less).
[0027] Further embodiments of the reverse transcriptase of the present invention include (c) an embodiment in which the N-terminal fragment and the C-terminal fragment are N-terminal fragments and C-terminal fragments obtained by splitting a polypeptide comprising an amino acid sequence having 90% or more homology to the amino acid sequence of SEQ ID NO: 1 at any one position within the region corresponding to positions 114 to 133 or any one position within the region corresponding to positions 167 to 186 of the amino acid sequence of SEQ ID NO: 1. In this embodiment (c), preferred aspects of the split positions into the N-terminal fragment and the C-terminal fragment correspond to the preferred aspects in (a), respectively.
[0028] Here, when referring to amino acid sequence homology, the amino acid in the target amino acid sequence that is located at the same position as the amino acid in the target amino acid sequence (the reference amino acid) when the two sequences are aligned using amino acid sequence analysis software or the like may be the same amino acid as the reference amino acid, or may have the same properties as the reference amino acid. When referring to amino acid sequence identity, the target amino acid is the same amino acid as the reference amino acid. Groups of amino acids with similar properties are well known in the art to which this invention pertains, and can be classified into, for example, acidic amino acids (aspartic acid and glutamic acid); basic amino acids (lysine, arginine, histidine); and neutral amino acids with hydrocarbon chains (glycine, alanine, valine, leucine, isoleucine, proline), hydroxyl groups (serine, threonine), sulfur-containing amino acids (cysteine, methionine), amide groups (asparagine, glutamine), imino groups (proline), and aromatic groups (phenylalanine, tyrosine, tryptophan).
[0029] Such amino acid sequence homology and identity are determined by comparing two sequences aligned to maximize sequence identity. Methods for determining sequence homology or identity (%) are known to those skilled in the art. Algorithms for obtaining optimal alignment and sequence identity can be any algorithm known to those skilled in the art (e.g., BLAST algorithm, FASTA algorithm, etc.), and sequence homology or identity of amino acid sequences can be determined using sequence analysis software such as BLASTP or FASTA. Furthermore, the amino acid sequence of the polypeptide (c) may have a homology of 90% or more with the amino acid sequence of (a) (SEQ ID NO: 1), preferably 95% or more (e.g., 96% or more, 97% or more, 98% or more, or 99% or more). Identity is more preferably 90% or more, 95% or more (e.g., 96% or more, 97% or more, 98% or more, or 99% or more).
[0030] Furthermore, in an amino acid sequence, a region that "corresponds to" a region consisting of a specific amino acid sequence refers to a region consisting of an amino acid sequence that is located at the same position as the specific amino acid sequence (a control amino acid sequence, for example, positions 114 to 133 or positions 167 to 186 of the amino acid sequence set forth in SEQ ID NO: 1) when the amino acid sequences are aligned using amino acid sequence analysis software (for example, GENETYX-MAC, Sequencher, ClustalW, etc.) (for example, parameters: default values (i.e., initial settings)).
[0031] In the present invention, the "presence" of both the N-terminal fragment and the C-terminal fragment means that they are both present in the same reaction field, for example, within the same cell (more preferably, for example, within the same nucleus, the same mitochondrion, or the same chloroplast), and it is preferable that they are in close proximity to each other. In this case, "close proximity" means that they are close enough to interact with each other, and in this case, they may be associated by bonding (for example, bonding via intermolecular forces, hydrogen bonds, etc.).
[0032] In the present invention, the regeneration of reverse transcriptase activity in the presence of both the N-terminal fragment and the C-terminal fragment can be appropriately confirmed by methods known to those skilled in the art. For example, as described in the Examples below, split GFP (GFP1-10 and GFP11, Berrios, KN, et al., Nat Chem Biol 17, pp. 1262-1270, 2021 (Reference I)) was used to prepare fusion proteins of the N-terminal fragment and GFP1-10, and fusion proteins of the C-terminal fragment and GFP11. These were then introduced into cells together with the Cas protein (preferably nCas) and pegRNA described below. DNA editing alone did not occur or the editing efficiency was significantly low, but when the two were combined, DNA editing occurred and the editing efficiency was high. The occurrence and efficiency of DNA editing can be measured by conventional methods, such as the EditR analysis described in the Examples below. In this method, for example, if the editing efficiency of the target DNA when using each single entity (N-terminal fragment or C-terminal fragment) is 10% or less compared to the editing efficiency (standard efficiency) of the target DNA when using a reverse transcriptase (unsplit reverse transcriptase) consisting of a polypeptide with the amino acid sequence set forth in SEQ ID NO: 1, it can be determined that DNA editing does not occur or that the editing efficiency is significantly low.On the other hand, if the editing efficiency when both are combined is 15% or more compared to the standard efficiency, it can be determined that DNA editing occurs and the editing efficiency is high.
[0033] The reverse transcriptase of the present invention may be tagged with an epitope tag (e.g., a flag tag, an HA tag, etc.) for purification or detection, or with various transport signals (e.g., a nuclear transport signal, a mitochondrial transport signal, or a plastid transport signal).
[0034] The N-terminal fragment and C-terminal fragment of the reverse transcriptase of the present invention can be obtained appropriately by conventional methods. For example, they can be prepared by artificial synthesis based on amino acid sequence information, or they can also be prepared at the nucleic acid level. That is, the N-terminal fragment and C-terminal fragment of the reverse transcriptase of the present invention can be obtained in an appropriate host cell by introducing and expressing polynucleotides encoding the N-terminal fragment and C-terminal fragment, respectively, or an expression vector containing the polynucleotide into the cell. When using expression vectors, the polynucleotides encoding the N-terminal fragment and the C-terminal fragment may be contained in separate vectors or in a single vector. Therefore, embodiments of the reverse transcriptase of the present invention also include expression vectors for the reverse transcriptase (i.e., N-terminal fragment expression vector, C-terminal fragment expression vector, and N- and C-terminal fragment expression vectors) and polynucleotides encoding the reverse transcriptase (i.e., polynucleotides encoding the N-terminal fragment and the C-terminal fragment, respectively).
[0035] The expression vector can be appropriately selected from vectors used in the art, including, for example, plasmid vectors, viral vectors, phage vectors, phagemid vectors, BAC vectors, YAC vectors, MAC vectors, and HAC vectors. Furthermore, the expression vector preferably contains a regulatory element operably linked to the polynucleotide to be expressed. Examples of the regulatory element include a promoter, an enhancer, an internal ribosome entry site (IRES), and other expression control elements (e.g., a transcription termination signal, e.g., a polyadenylation signal and a polyU sequence). Furthermore, the expression vector is preferably one that can stably express the encoded protein without being integrated into the host genome. Such an expression vector can be prepared according to a conventionally known method.
[0036] As the host cell into which the expression vector is introduced, it can be appropriately selected in consideration of its compatibility with the expression vector. For example, prokaryotes such as Escherichia coli, Actinomycetes, and Archaea, and cells of eukaryotes such as yeast, sea urchin, silkworm, zebrafish, mouse, rat, frog, tobacco, Arabidopsis thaliana, and rice can be mentioned. Note that "polynucleotide" includes both DNA and RNA (mRNA), and in each polynucleotide, codon optimization may be performed for the purpose of enhancing the expression efficiency in cells.
[0037] <DNA Editing System> The DNA editing system of the present invention comprises the following (A) to (C): (A) At least one selected from the group consisting of the reverse transcriptase of the present invention, the expression vector of the reverse transcriptase, and the polynucleotide encoding the reverse transcriptase, (B) At least one Cas protein selected from the group consisting of a Cas protein (Cas) having nuclease activity and a nickase-type Cas protein (nCas) with partially defective nuclease activity, the expression vector of the Cas protein, and at least one selected from the group consisting of the polynucleotide encoding the Cas protein, (C) At least one selected from the group consisting of the prime editing guide RNA (pegRNA) of the Cas protein, the expression vector of the pegRNA, and the polynucleotide encoding the pegRNA, is included. Further, as the DNA editing system of the present invention, (D) At least one selected from the group consisting of the nicking guide RNA (ngRNA) of the Cas protein, the expression vector of the ngRNA, and the polynucleotide encoding the ngRNA is preferably further included. The reverse transcriptase, the expression vector of the reverse transcriptase, and the polynucleotide encoding the reverse transcriptase are as described above, including their preferred embodiments.
[0038] (Target DNA) In the present invention, DNA containing a target site to be subjected to DNA editing is referred to as "target DNA." Furthermore, when the target site is contained in a region encoding a specific gene, the target DNA is also referred to as a "target gene." The target DNA of the present invention is double-stranded DNA, and for convenience, the "target site" refers to the region of one of the strands that is replaced by the strand (edited strand) synthesized by reverse transcriptase. For convenience, to illustrate the correspondence with the DNA editing system of the present invention, the target DNA of the present invention is defined as having a structure in which one strand (first strand, target strand) contains, from the 5' end, a primer sequence, a target site, and a PAM (PAM_1), and the other strand (second strand) contains a pegRNA recognition sequence containing a sequence complementary to the target site. Furthermore, when using the ngRNA described below, an ngRNA recognition sequence is further present on the 3' side of PAM_1 of the first strand, and a PAM (PAM_2) is also present on the 5' side of the second strand. Here, when different Cas proteins are used for pegRNA and ngRNA, PAM_1 and PAM_2 may be the same or different.
[0039] A "PAM (protospacer adjacent motif)" is a sequence recognized by a Cas protein. Its length varies depending on the type of Cas protein, but typically consists of 2 to 5 bases adjacent to the 3' side of the complementary sequence of the pegRNA recognition sequence (and the ngRNA recognition sequence when an ngRNA is used). The nucleotide sequence of the PAM also varies depending on the type of Cas protein, but typically, when the Cas protein is Cas9 or nCas9, the nucleotide sequence of the PAM recognized by this protein is 5'-NGG. However, the nucleotide sequence of such a PAM can also be changed by modifying the Cas protein (e.g., by introducing a mutation), as described below, thereby expanding the range of target sites available.
[0040] As described above, the "target site" according to the present invention refers to a region of one strand of double-stranded DNA that is replaced by the edited strand. However, since the replaced edited strand is not complementary, when repair is performed by the DNA repair mechanism using the edited strand as a template, both strands are ultimately edited (FIG. 1C (c)). The target site according to the present invention is a site that is set according to the position of the PAM and is designed to include at least one base of the site that is subject to cleavage by a Cas protein (preferably single-strand cleavage by nCas). The site that is subject to cleavage by a Cas protein is usually between the third and fourth bases from the 5' side of the PAM, although this varies depending on the type of Cas protein. The length of such a target site is preferably 1 to 100 bases, more preferably 1 to 80 bases.
[0041] In the present invention, the "primer sequence" is a sequence located adjacent to the 5' side of the target site, to which the primer binding sequence of the pegRNA described below binds, and its 3' end serves as the initiation point for reverse transcription by reverse transcriptase. Therefore, the primer sequence is designed to include the 3' end of the target strand generated upon cleavage by the Cas protein (preferably single-strand cleavage by nCas). The length of such a primer sequence is preferably 4 to 30 bases, more preferably 6 to 25 bases, and even more preferably 8 to 20 bases.
[0042] In the present invention, a "pegRNA recognition sequence" is a sequence that is recognized and bound by the spacer sequence of the pegRNA described below. Furthermore, an "ngRNA recognition sequence" is a sequence that is recognized and bound by the spacer sequence of the ngRNA. These pegRNA recognition sequences and ngRNA recognition sequences are set depending on the positions of the PAM and target site described above. The lengths of the pegRNA recognition sequence and ngRNA recognition sequence according to the present invention are each independently dependent on the type of Cas protein, but in the case of Cas9 or nCas9, they are preferably 12 to 25 bases, more preferably 14 to 23 bases, and even more preferably 15 to 22 bases.
[0043] When the following ngRNA is used, the pegRNA recognition sequence and the ngRNA recognition sequence according to the present invention are sequences located at different positions. The length between the pegRNA recognition sequence and the ngRNA recognition sequence is preferably 0 (adjacent to each other) to 200 bases, more preferably 0 to 150 bases, and even more preferably 0 to 100 bases.
[0044] The nucleotide sequence of such target DNA is not particularly limited, and by designing the DNA editing system of the present invention (particularly pegRNA or ngRNA) to match that sequence, it can become the target of DNA editing in the present invention.
[0045] The target DNA according to the present invention is preferably DNA present in cells (endogenous DNA). DNA present in cells may be endogenous DNA or exogenous DNA. Examples of endogenous DNA include genomic DNA in chromosomes, mitochondria, or chloroplasts, and examples of exogenous DNA include DNA introduced into cells (reporter genes, marker genes, genes of viruses that infect hosts, bacteria, or protozoa, etc.).
[0046] (Cas protein) In the present invention, the term "Cas protein" refers to a CRISPR (clustered regularly interspaced short palindromic repeat)-associated enzyme, and is at least one selected from the group consisting of Cas proteins having nuclease activity and nickase-type Cas proteins partially deficient in nuclease activity. The Cas protein may be a class 1-CRISPR-associated enzyme or a class 2-CRISPR-associated enzyme. Among these, the Cas protein according to the present invention is preferably a class 2-CRISPR-associated enzyme, and more preferably a nickase-type Cas protein partially deficient in nuclease activity.
[0047] Among the Cas proteins, the "Cas protein having nuclease activity" is not particularly limited, and examples thereof include a Cas9 protein (Cas9, for example, a Cas9 protein derived from Streptococcus pyogenes (SpCas9), a Cas9 protein derived from Staphylococcus aureus (SaCas9), and a Cas9 protein derived from Campylobacter jejuni (CjCas9)), a Cas12a protein (Cas12, for example, a Cas12a (Cpf1) protein derived from Francisella novicida (FnCas12a), a Cas12a protein derived from Lachnospiraceae bacterium (LbCas12a), and a Cas12a protein derived from Acidaminococcus Examples of Cas proteins include Cas12a protein (AsCas12a) derived from Caspase, Cas12b (C2c1) protein (Cas12b), Cas12e (CasX) protein (Cas12e), Cas14 protein (Cas14), Cas13 protein (Cas13), and Cas3 protein (Cas3). Among these, the Cas according to the present invention is preferably a class 2-CRISPR-associated enzyme, and more preferably Cas9, because it does not require a complex (cascade).
[0048] Furthermore, in the present invention, a "nickase-type Cas protein partially lacking nuclease activity" refers to a nickase-type Cas protein in which a portion of the nuclease activity of the Cas protein (Cas) is missing and the Cas protein has nickase activity (single-stranded DNA cleavage activity), and in this specification, this term is sometimes referred to as "nCas (nicking Cas)."
[0049] Such nCas is not particularly limited, but examples thereof include nickase-type Cas9 proteins (nCas9, for example, nSpCas9 (nickase-type SpCas9) in which at least one selected from the group consisting of D10, E762, D839, H983, D986, H840, and N863 of SpCas9 is mutated; nSaCas9 (nickase-type SaCas9) in which at least one selected from the group consisting of D10, D556, H557, and N580 of SaCas9 is mutated). Examples of such Cas12a proteins include those obtained by the addition of a nickase-type Cas12a protein (nCas12a, for example, R1138A for nLbCas12a (a nickase-type of LbCas12a); R1226A for nAsCas12a (a nickase-type of AsCas12a), etc.), a nickase-type Cas12b protein (nCas12b), a nickase-type Cas14 protein (nCas14), a nickase-type Cas3 protein (nCas3), and a nickase-type Cas10d protein (nCas10d). Among these, the nCas according to the present invention is preferably a nickase-type Class 2-CRISPR-associated Cas, and more preferably nCas9.
[0050] The amino acid sequences encoding these Cas proteins are known and can be obtained from publicly available databases, such as GenBank (http: / / www.ncbi.nlm.nih.gov) and published literature.
[0051] Furthermore, the Cas proteins of the present invention are not limited to these, and further mutations, for example, mutations to modify PAM recognition, may be introduced (Benjamin, P. et al., Nature 523, pp. 481-485, 2015; Hirano, S. et al., Molecular Cell 61, pp. 886-894, 2016).
[0052] Furthermore, the Cas protein of the present invention may be one that has been appropriately modified (e.g., by substitution, deletion, insertion, and / or addition of amino acid residues) based on the amino acid sequence of a known Cas protein, as long as its functions (e.g., the ability to form a complex with pegRNA and, if necessary, ngRNA, the nuclease activity, or the nickase activity) are not inhibited. The ability of a Cas protein to form a complex and the nuclease activity or the nickase activity can be confirmed, for example, by showing, when introduced into cells together with pegRNA and reverse transcriptase, a target DNA editing efficiency that is equal to or greater than that of an unmodified Cas protein (e.g., a Cas protein consisting of a polypeptide with a known amino acid sequence) (e.g., 50% or greater, preferably 80% or greater).
[0053] When using the following ngRNA, the corresponding Cas protein is preferably nCas. In this case, the DNA editing system of the present invention may contain multiple Cas proteins, each of which is different for the pegRNA and the ngRNA. However, from the viewpoint of simplicity, it is preferable to contain one Cas protein common to both.
[0054] The Cas protein of the present invention may be tagged with an epitope tag (e.g., a flag tag, an HA tag, etc.) for purification or detection, or with various localization signals (e.g., a nuclear localization signal, a mitochondrial localization signal, or a plastid localization signal).
[0055] The Cas protein of the present invention can be obtained by conventional methods. For example, it can be prepared by artificial synthesis based on amino acid sequence information, or it can be prepared at the nucleic acid level. That is, the Cas protein can be obtained in an appropriate host cell by introducing a polynucleotide encoding the Cas protein or an expression vector for the Cas protein containing the polynucleotide into the cell and expressing the polynucleotide. The expression vector and host cell are the same as those described for the reverse transcriptase of the present invention. Therefore, embodiments of the Cas protein of the present invention also include the expression vector for the Cas protein and the polynucleotide encoding the Cas protein.
[0056] (prime editing guide RNA) In the present invention, a "prime editing guide RNA" is an RNA comprising, in order from the 5' end, a spacer sequence, a template sequence, and a primer binding sequence. It may be in the form of a single nucleotide chain (RNA chain), or may be in the form of two or more nucleotide chains joined together via bonds between complementary sequences. In this specification, "prime editing guide RNA" may also be referred to as "pegRNA (prime editing guide RNA)" in some cases.
[0057] A "spacer sequence," also known as a "protospacer," is originally a sequence derived from exogenous DNA incorporated during the adaptation process into the CRISPR structure of the bacterial genome from which the CRISPR-Cas system is derived. However, the spacer sequence of the present invention is a sequence designed to recognize and bind to the pegRNA recognition sequence. Such a spacer sequence is preferably a complementary sequence to the pegRNA recognition sequence, but any nucleotide sequence capable of hybridizing with the pegRNA recognition sequence to form a double-stranded chain is sufficient; it need not be completely complementary. In the present invention, the sequence complementarity between the spacer sequence and the pegRNA recognition sequence is preferably 80% or more, 85% or more, 90% or more, or 95% or more (e.g., 96% or more, 97% or more, 98% or more, or 99% or more). Those skilled in the art can appropriately calculate the sequence complementarity using known methods (e.g., BLAST (NCBI)).
[0058] The spacer sequence of the pegRNA of the present invention can be designed to correspond to the pegRNA recognition sequence corresponding to the desired target site, and its length, although depending on the type of Cas protein, is preferably 10 to 40 bases. For example, in the case of Cas9 or nCas9, it is more preferably 12 to 25 bases, even more preferably 14 to 23 bases, and even more preferably 15 to 22 bases.
[0059] In the present invention, the "template sequence," also referred to as a "reverse transcription template (RTT)," is a sequence that serves as a template for the reverse transcriptase. The template sequence according to the present invention can be appropriately designed depending on the nucleotide sequence desired to be replaced with the target site. The length of the template sequence according to the present invention is preferably 1 to 100 bases, more preferably 5 to 90 bases, and even more preferably 10 to 80 bases.
[0060] In the present invention, a "primer binding sequence," also referred to as a "primer binding site (PBS)," is a sequence designed to recognize and bind to a primer sequence at the 3' end generated on the target strand of a target DNA by a Cas protein (preferably nCas). Such a primer binding sequence is preferably a complementary sequence to the primer sequence, but it need not be completely complementary as long as it is a nucleotide sequence that can hybridize with the primer sequence to form a double strand. In the present invention, the sequence complementarity between such a primer binding sequence and the primer sequence is preferably 80% or more, 85% or more, 90% or more, or 95% or more (e.g., 96% or more, 97% or more, 98% or more, or 99% or more).
[0061] The length of the primer binding sequence according to the present invention is preferably 4 to 30 bases, more preferably 7 to 23 bases, and even more preferably 10 to 17 bases.
[0062] The pegRNA of the present invention preferably also contains a scaffold sequence as a tracrRNA structure that forms a complex with the Cas protein. Conventionally known scaffold sequences or sequences similar thereto can be appropriately used as the scaffold sequence. The pegRNA of the present invention may further contain a reverse transcriptase termination sequence (e.g., a sequence that forms a hairpin loop structure), a linker, an RNA motif (e.g., tmpknot, tevopreQ1, xrRNA), an RNA aptamer (e.g., MS2, PP7, boxB, com), etc.
[0063] The pegRNA of the present invention can be obtained appropriately by conventionally known methods. For example, it can be prepared by artificial synthesis based on sequence information designed according to the pegRNA recognition sequence, primer sequence, and desired template sequence, or it can be prepared at the nucleic acid level. That is, pegRNA can be obtained in suitable host cells by introducing a polynucleotide encoding the pegRNA or an expression vector for pegRNA containing the polynucleotide into the cells and expressing the polynucleotide. The expression vector and host cell are the same as those described for the reverse transcriptase of the present invention. Therefore, aspects of the pegRNA of the present invention also include the expression vector for the pegRNA and the polynucleotide encoding the pegRNA.
[0064] (nicking guide RNA) In the present invention, a "nicking guide RNA" is a guide RNA used to nick the unedited strand, i.e., the strand opposite the target site, with the Cas protein (preferably nCas), for the purpose of increasing the probability of DNA repair occurring in the target DNA using the strand synthesized by reverse transcriptase (edited strand) as a template (Non-Patent Document 1), and is also referred to as "ngRNA (nicking guide RNA)" in this specification.
[0065] The ngRNA of the present invention includes at least a spacer sequence designed to recognize and bind to the ngRNA recognition sequence. Such a spacer sequence is preferably a complementary sequence to the ngRNA recognition sequence, but it need not be completely complementary as long as it is a nucleotide sequence that can hybridize with the ngRNA recognition sequence to form a double strand. In the present invention, the sequence complementarity between the spacer sequence and the ngRNA recognition sequence is preferably 80% or more, 85% or more, 90% or more, or 95% or more (e.g., 96% or more, 97% or more, 98% or more, or 99% or more).
[0066] The spacer sequence of the ngRNA of the present invention can be designed to correspond to the ngRNA recognition sequence corresponding to the desired target site, and its length, although depending on the type of Cas protein, is preferably 10 to 40 bases. For example, in the case of nCas9, it is more preferably 12 to 25 bases, even more preferably 14 to 23 bases, and even more preferably 15 to 22 bases.
[0067] The ngRNA of the present invention preferably also contains a scaffold sequence as a tracrRNA structure that forms a complex with the Cas protein. The scaffold sequence may be a conventionally known scaffold sequence or a sequence similar thereto.
[0068] The ngRNA of the present invention can be obtained appropriately by conventionally known methods. For example, it can be prepared by artificial synthesis based on sequence information designed according to the ngRNA recognition sequence, or it can be prepared at the nucleic acid level. That is, ngRNA can be obtained in an appropriate host cell by introducing and expressing a polynucleotide encoding the ngRNA or an expression vector for the ngRNA containing the polynucleotide into the cell. The expression vector and host cell are the same as those described for the reverse transcriptase of the present invention. Therefore, embodiments of the ngRNA of the present invention also include the expression vector for the ngRNA and the polynucleotide encoding the ngRNA.
[0069] <Method for editing target DNA> The present invention also provides a method for editing target DNA, comprising the steps of contacting the target DNA with the DNA editing system of the present invention and editing a target site in the target DNA. In the method for editing target DNA of the present invention, the DNA editing system and target DNA are as described above for the DNA editing system of the present invention, including preferred embodiments thereof.
[0070] When the DNA editing system, i.e., the reverse transcriptase (N-terminal fragment, C-terminal fragment), Cas protein, pegRNA, and optionally ngRNA of the present invention, is contacted with the target DNA, the spacer sequence of the pegRNA recognizes and binds to the corresponding pegRNA recognition sequence on the target DNA, guiding the Cas protein to the vicinity of the target site on the target DNA. This allows the Cas protein to recognize the PAM and introduce a nicking (a double-strand break (cut) or nick, preferably a nick) at the target site. Furthermore, the presence of both the N-terminal fragment and the C-terminal fragment restores reverse transcriptase activity. The primer binding sequence of the pegRNA recognizes the primer sequence at the 3' end of the target DNA (first strand, target strand) resulting from the introduction of the nicking, and the reverse transcriptase then uses the subsequent template sequence as a template to synthesize a DNA strand (edited strand) starting from the primer sequence. In the target DNA, the 5' flap (see (b) in Figure 1C) from which the non-edited strand is released, which is in equilibrium with the 3' flap (see (a) in Figure 1C) from which the edited strand is released, is cut by an intracellular enzyme, repairing the nicking and replacing the target site with the edited strand, resulting in DNA editing according to the present invention. Furthermore, since the incorporated edited strand is not complementary and is in a mismatched state of double-stranded DNA, repair using the edited strand as a template by the intracellular DNA repair mechanism can result in edited DNA in which both strands are edited. In this case, if ngRNA is also used, a nick is introduced into the opposite strand (second strand) of the target site by a Cas protein (preferably nCas) induced by the ngRNA, thereby improving the probability of DNA repair using the edited strand as a template (Non-Patent Document 1). This results in mainly base substitutions at the target site and its complementary strand, but does not preclude the introduction of various mutations due to further base substitutions or deletions or insertions of one or several bases during mismatch repair of double-stranded DNA.Furthermore, for example, by using a pair of DNA editing systems of the present invention, it is possible to perform editing over a wide range (e.g., 16 megabases or less), such as deletion or substitution of a long chain between the two cleavage sites (see, for example, Choi J. et al., Nature Biotechnology 40, pp. 218-226, 2022; Tao R. et al., Signal Transduction and Targeted Therapy, 2022 Apr 20, 7(1):108. doi:10.1038 / s41392-022-00936-w). It has also been suggested that, for example, chromosomal translocations, inversions, and duplications are also possible. Therefore, DNA editing according to the present invention also includes deletion of one or more bases, substitution with one or more other bases, or insertion of one or more bases, or a combination of these mutations, in the vicinity of the target site or in the long chain including the target site.
[0071] The method for editing target DNA of the present invention is preferably carried out intracellularly. The "intracellular" where the method for editing target DNA of the present invention is carried out may be a eukaryotic cell or a prokaryotic cell, but is preferably a eukaryotic cell. Examples of the eukaryotic cells include animal cells (such as cells of mammals, fish, birds, reptiles, amphibians, and insects), plant cells, algae cells, and yeast. Examples of the prokaryotic cells include Escherichia coli, Salmonella, Bacillus subtilis, lactic acid bacteria, and extreme thermophiles.
[0072] "Animal cells" include, for example, cells constituting an individual animal, cells constituting organs or tissues extracted from an animal, and cultured cells derived from animal tissues. Specific examples include germ cells such as oocytes and sperm; germ cells of various stages of embryos (e.g., 1-cell, 2-cell, 4-cell, 8-cell, 16-cell, and morula stages); stem cells such as induced pluripotent stem (iPS) cells and embryonic stem (ES) cells; and somatic cells such as fibroblasts, hematopoietic cells, neurons, muscle cells, bone cells, liver cells, pancreatic cells, brain cells, and kidney cells. Pre- and post-fertilization oocytes can be used as the oocytes, but fertilized oocytes, i.e., fertilized eggs, are preferred. Pronuclear stage embryos are particularly preferred. Oocytes can be used by thawing cryopreserved oocytes.
[0073] "Plant cells" include, for example, cells that constitute an individual plant, cells that constitute organs or tissues separated from a plant, cultured cells derived from plant tissue, etc. Examples of plant organs and tissues include leaves, stems, shoot tips (growing points), roots, tubers, calli, etc.
[0074] The method for contacting the DNA editing system with the target DNA is not particularly limited, and examples thereof include a method for introducing the DNA editing system into cells containing the target DNA, such as the method for producing cells with edited target DNA described below.
[0075] <Method for producing cells with targeted DNA editing> The present invention also provides a method for producing a cell in which target DNA has been edited, comprising the steps of introducing the DNA editing system of the present invention into a cell, contacting it with target DNA, and editing the target site of the target DNA.
[0076] In the method of the present invention for producing cells in which target DNA has been edited (hereinafter sometimes simply referred to as the "production method"), the DNA editing system and target DNA, including preferred embodiments thereof, are as described above for the method of the present invention for editing target DNA. The target DNA in the production method of the present invention is genomic DNA, and pegRNA and, if necessary, ngRNA can be designed depending on the purpose of editing the genomic DNA.
[0077] Furthermore, in the production method of the present invention, the DNA editing system, i.e., the reverse transcriptase of the present invention, the Cas protein, the pegRNA, and, if necessary, the ngRNA, can be contacted with the target DNA by introducing the DNA editing system into a cell in the form of a protein or RNA, or by introducing the DNA editing system into a cell in the form of a polynucleotide encoding the protein or RNA, and / or by introducing the DNA editing system into a cell in the form of an expression vector and expressing it within the cell. Thus, the DNA editing systems may be independently introduced into a cell in the form of a protein or RNA, or may be introduced into a cell in the form of DNA or RNA (polynucleotide) encoding these and expressed within the cell, or may be introduced into a cell in the form of an expression vector and expressed within the cell.
[0078] Furthermore, when the DNA editing system is introduced into a cell in the form of an expression vector, for example, a vector that expresses each protein or RNA separately may be introduced into the cell, or a vector that expresses a combination of these may be introduced into the cell.
[0079] Methods for introducing the proteins or RNAs, polynucleotides encoding the proteins or RNAs, and expression vectors into cells can be appropriately selected from known methods for introducing proteins or polynucleotides into cells depending on the cell type. Examples of such methods include electroporation, microinjection, particle gun technology, calcium phosphate transfection, polyethyleneimine (PEI) transfection, liposome transfection, DEAE-dextran transfection, cationic lipid-mediated transfection, viruses (adenovirus, lentivirus, adeno-associated virus, baculovirus, etc.), Agrobacterium transfection, lithium acetate transfection, spheroplast transfection, and heat shock transfection (calcium chloride transfection, rubidium chloride transfection). These methods are described in many standard laboratory manuals, such as Davis et al., Basic Methods in Molecular Biology, New York: Elsevier, 1986.
[0080] When the DNA editing system is introduced into a cell, each introduced component or each component expressed in the cell comes into contact with the target DNA in the cell, and the DNA editing described in the method for editing target DNA of the present invention above causes the desired base to be replaced at the target site, resulting in the production of a cell in which the target DNA has been edited.
[0081] The present invention also provides a method for producing a non-human individual containing cells in which the target DNA has been edited. This method includes a step of producing a non-human individual from cells obtained by the above-mentioned production method. Examples of the non-human individual include non-human animals and plants. Examples of the non-human animal include mammals (e.g., mice, rats, guinea pigs, hamsters, rabbits, monkeys, pigs, cows, goats, and sheep), fish, birds, reptiles, amphibians, and insects. When producing a model animal, the mammal is preferably a rodent such as a mouse, rat, guinea pig, or hamster, and particularly preferably a mouse. Examples of the plant include grains, oilseed crops, forage crops, fruits, and vegetables. Specific examples of crops include rice, corn, banana, peanut, sunflower, tomato, rapeseed, tobacco, wheat, barley, potato, soybean, cotton, and carnation.
[0082] Known methods can be used to create non-human individuals from cells in which the target DNA has been edited. When creating non-human individuals from cells in animals, germ cells or pluripotent stem cells are typically used. For example, the DNA editing system is microinjected into oocytes, and the resulting oocytes are implanted into the uterus of a pseudopregnant female non-human mammal, after which offspring can be obtained. It has long been known that somatic cells of plants possess totipotency. For example, a plant in which the desired DNA has been edited can be obtained by microinjecting the DNA editing system into plant cells and regenerating the plant from the resulting plant cells. Furthermore, offspring or clones in which the desired DNA has been edited can also be obtained from the resulting non-human individuals.
[0083] Confirmation of the presence or absence of targeted DNA editing and determination of the genotype can be performed based on conventionally known techniques, such as PCR, sequencing, Southern blotting, etc.
[0084] <Kit> The present invention relates to a kit for use in the method for editing target DNA of the present invention, the production method of the present invention, or the method for producing a non-human individual of the present invention, (A) at least one member selected from the group consisting of the reverse transcriptase of the present invention, an expression vector for the reverse transcriptase, and a polynucleotide encoding the reverse transcriptase; The kit of the present invention also includes the following (B) to (C): (B) at least one selected from the group consisting of a Cas protein, an expression vector for the Cas protein, and a polynucleotide encoding the Cas protein; (C) at least one selected from the group consisting of a pegRNA, an expression vector for the pegRNA, and a polynucleotide encoding the pegRNA; It is preferable that the composition further contains at least one selected from the group consisting of: (D) at least one selected from the group consisting of ngRNA, an expression vector for the ngRNA, and a polynucleotide encoding the ngRNA More preferably, it further comprises:
[0085] When the kit of the present invention includes (B) to (C) or (B) to (D), and when the above (A) to (D) are in the form of an expression vector, these may be in an embodiment in which two or more components are combined and contained in one vector. Furthermore, when the above (C) and (D) are in the form of an expression vector, these may be in a form in which the user can design the pegRNA and ngRNA according to the target site of the target DNA, and the expression vector for the pegRNA is (C') A pegRNA expression vector containing an insertion site for a spacer sequence, an insertion site for a template sequence, and an insertion site for a primer binding sequence. The ngRNA expression vector may be (D') ngRNA expression vector containing an insertion site for a spacer sequence It may also be possible to use the following.
[0086] The kit of the present invention may further comprise one or more additional reagents. Examples of such additional reagents include, but are not limited to, a dilution buffer, a nucleic acid introduction reagent, a protein introduction reagent, and a control reagent (e.g., a complete reverse transcriptase consisting of a polypeptide having the amino acid sequence set forth in SEQ ID NO: 1). The kit may also further comprise instructions for carrying out the method of the present invention.
[0087] Each component included in the kit of the present invention may be contained in a separate container or may be contained in the same container. Each component may be contained in a container in a single-use amount, or in a single container in amounts sufficient for multiple uses. Each component may be contained in a container in a dry form, or in a form dissolved in an appropriate solvent (a solvent containing a buffer, stabilizer, preservative, antiseptic, etc.). [Example]
[0088] The present invention will be explained in more detail below based on test examples including working examples, but the present invention is not limited to the following test examples.
[0089] <Basic operations> The basic operations performed in this test example are shown below.
[0090] 1. Agarose Electrophoresis Agarose S was added to 1x TAE at a ratio of 1-3% by mass and heated to dissolve. This solution was poured into a gel maker, a comb inserted, and left to set at room temperature. The solidified gel was placed in the electrophoresis chamber, and an appropriate amount of 1x TAE was poured in. Loading buffer was added to the sample to a concentration of 1x or higher, and the gel was applied to the gel wells along with a DNA size marker. If the gel was to be excised after electrophoresis and DNA extracted, it was electrophoresed at 50V for approximately 1 hour; if not, it was electrophoresed at 100V for approximately 30 minutes. The electrophoresed gel was placed in ion-exchanged water, 2 μL of ethidium bromide was added, and the gel was shaken for approximately 20 minutes. The gel was photographed using a UV transilluminator to confirm the bands, and if DNA extraction was to be performed, the gel was excised at the position of the desired band.
[0091] 2. DNA extraction using Wizard SV Gel and PCR Clean-Up System (Promega) The weight of the gel excised in step 1 above was measured, and 100 μL of Membrane Binding Solution was added to 0.1 g of 1% agarose gel. The gel was vortexed for approximately 15 seconds and heated in a 55°C heat block for approximately 10–15 minutes. After confirming that each gel was dissolved, the solution was transferred to a collection tube containing a column and allowed to stand at room temperature for 1 minute. The column was centrifuged at 20°C and 16,000 g for 1 minute, and the filtrate was discarded. 700 μL of Membrane Wash Solution was added to the column, which was then centrifuged again at 20°C and 16,000 g for 1 minute, and the filtrate was discarded. 500 μL of Membrane Wash Solution was then added to the column, which was then centrifuged at 20°C and 16,000 g for 5 minutes, and the filtrate was discarded. The column was transferred to a new 1.5 mL tube, and 15 μL of nuclease-free water was added. The column was left at room temperature for 1 minute, and then centrifuged at 20°C and 16,000 g for 1 minute. Another 15 μL of nuclease-free water was added, and the column was left at room temperature for 1 minute, and then centrifuged at 20°C and 16,000 g for 1 minute. The concentration of the filtrate was measured using a Nanodrop (Thermo Fisher Scientific), adjusted to 200 ng / μL, and stored at -20°C.
[0092] 3. PCR using PrimeSTAR Max DNA Polymerase (Takara Bio Inc.) 5 μL of PrimeSTAR Max Premix (2x), 0.3 μL of 10 μM forward primer, 0.3 μL of 10 μM reverse primer, 0.2 μL of 1 ng / μL template, and 4.2 μL of autoclaved ultrapure water were mixed in a PCR tube. After preheating at 94°C for 2 minutes, the PCR tube underwent 35 cycles of denaturation at 98°C for 10 seconds, annealing at 60–68°C (depending on the primer Tm) for 5 seconds, and extension at 72°C for 5–70 seconds (depending on the number of bases to be amplified), and then stored at 4°C.
[0093] 4. PCR using PrimeSTAR GXL DNA Polymerase (Takara Bio Inc.) A PCR tube was mixed with 0.5 μL of PrimeSTAR GXL DNA Polymerase, 5 μL of 5x PrimeSTAR GXL Buffer, 2 μL of 2.5 mM each dNTP mixture, 0.75 μL of 10 μM forward primer, 0.75 μL of 10 μM reverse primer, 1.25 μL of 1–10 ng / μL template, and 14.75 μL of autoclaved ultrapure water. The tube was preheated to 94°C for 2 minutes, followed by 45 cycles of denaturation at 98°C for 10 seconds, annealing at 60–68°C (depending on the primer Tm) for 15 seconds, and extension at 68°C for 300–350 seconds (depending on the number of bases to be amplified), and then stored at 4°C.
[0094] 5. In-Fusion Reaction 0.8 μL of insert DNA and 0.8 μL of vector DNA, or 1.6 μL of vector DNA alone, were mixed with 0.4 μL of 5x In-Fusion Premix (Takara Bio Inc.) in a PCR tube. The mixture was reacted at 50°C for 15 minutes using a thermal cycler and then stored at 4°C.
[0095] 6. Restriction Enzyme Digestion A 1.5 mL tube was mixed with 200 ng / μL of plasmid, 0.3 μL of restriction enzyme (0.3 μL each if two restriction enzymes were used), 1 μL of the buffer corresponding to the restriction enzyme, and 7.7 μL of autoclaved ultrapure water (7.4 μL if two restriction enzymes were used). This mixture was left in an incubator at 37°C for at least one hour to allow the restriction enzyme reaction to proceed. If ligation (see step 7 below) was to be performed, 0.5 μL of rAPid Alkaline Phosphatase 2 (ROCHE) was added, and the mixture was again left in an incubator at 37°C for one hour to allow phosphorylation.
[0096] 7.Ligation 1 μL of insert DNA, 1 μL of vector DNA, and 1.0 μL of 2× Ligation Mix (manufactured by Takara Bio Inc.) were mixed in a PCR tube, and then reacted at 16° C. for 30 minutes using a thermal cycler.
[0097] 8. Sequencing analysis using the SeqStudio Genetic Analyzer (Thermo Fisher Scientific) 8.1 Cycle sequencing reactions One microliter of 200 ng / μL or 100 ng / μL plasmid, 1 μL of BigDye Terminator v3.1 Ready Reaction Mix (Applied Biosystems), 1.5 μL of 5x Sequencing Buffer, 0.32 μL of 10 μM primer, and 6.18 μL of autoclaved ultrapure water were mixed. Using a thermal cycler, the mixture was preheated to 96°C for 2 minutes, followed by 25 cycles of thermal denaturation at 96°C for 10 seconds, annealing at 50°C for 5 seconds, and extension at 60°C for 4 minutes. The mixture was then stored at 4°C.
[0098] 8.2 Purification of cycle sequencing products using Optima DTR 8-Well Strip Kit (EdgeBio) The 8-well strip was removed from the 96-well holder, placed in a flat-bottom waste plate, and centrifuged at 2600 rpm for 3 minutes. The 8-well strip was then placed in a 96-well v-bottom collection plate, and 10 μL of cycle sequencing product was added to the 8-well strip and centrifuged at 2600 rpm for 5 minutes.
[0099] 8.3 Sequence Analysis 10 μL of the product purified in 8.2 above was added to a sequencing plate, spun down, and subjected to sequencing analysis using a SeqStudio Genetic Analyzer (Applied Biosystems).
[0100] 9. Transformation into E. coli The sample (all transformed samples in the test examples below contain the ampicillin resistance gene) was mixed with at least 10 times the amount of competent cells (XL10-Gold) in a 1.5 mL tube. After leaving it to stand on ice for at least 10 minutes, it was incubated at 42°C for 30 seconds using a heat block, and then cooled on ice for approximately 2 minutes. In a safety cabinet, the bacterial suspension was spread on an LB+Ampicillin plate, which was then placed in a 37°C incubator and cultured for 16-18 hours.
[0101] 10. Colony PCR A PCR reaction mixture was prepared in a PCR tube by mixing 4 μL of 2× SapphireAmp Fast PCR Master Mix (Takara Bio Inc.), 0.16 μL of 10 μM forward primer, 0.16 μL of 10 μM reverse primer, and 3.68 μL of autoclaved ultrapure water. A colony on the plate was picked with a 200 μL pipette tip and placed on a replica LB+Ampicillin plate. The tip was then immersed in the PCR reaction mixture. The tip was discarded, and the tube was preheated at 94°C for 2 minutes using a thermal cycler. This tube was then subjected to 27 cycles of thermal denaturation at 98°C for 5 seconds, annealing at 52–68°C (depending on the primer Tm) for 5 seconds, and extension at 72°C for 10–40 seconds (depending on the number of bases to be amplified), followed by storage at 4°C. The presence of the target band was confirmed by agarose electrophoresis as described above.
[0102] 11. Plasmid extraction using GenElute Plasmid miniprep Kit (Sigma-Aldrich) 11.1 Small culture 15 μL of 25 mg / mL ampicillin was added to 3 mL of LB liquid medium in a test tube. An E. coli colony was picked with a 200 μL pipette tip and dropped into the test tube, followed by shaking culture at 37°C for at least 16 hours.
[0103] 11.2 Extraction of Plasmids from E. coli Approximately half of the culture medium after small culture (11.1) was decanted into collection tubes and centrifuged at 20°C and 10,000g for 1 minute, after which the supernatant was discarded. The remaining culture medium was also transferred to a tube and centrifuged at 20°C and 10,000g for 1 minute, after which the supernatant was discarded. 200 μL of Resuspension Solution stored at 4°C was added to each collection tube, and the precipitate was suspended by vortexing. 200 μL of Lysis Solution was added to each collection tube, mixed by inversion, and then the tubes were left at room temperature for 3 minutes with the lids open. Next, 350 μL of Neutralization / Binding Buffer was added, mixed by inversion, and centrifuged at 20°C and 12,000g for 10 minutes. Meanwhile, a column was placed in the collection tube, 500 μL of Column Preparation Solution was added, and the tubes were centrifuged at 20°C and 12,000g for 1 minute. The filtrate was discarded, and the lysate supernatant was transferred to a column and centrifuged at 20°C and 12,000g for 1 minute. The filtrate was discarded, and 500μL of Wash solution 1 was added. This column was then centrifuged at 20°C and 12,000g for 1 minute. The filtrate was then discarded, and 750μL of Wash solution 2 was added. This column was then centrifuged at 20°C and 12,000g for 1 minute. The filtrate was discarded, and this column was centrifuged again at 20°C and 12,000g for 1 minute. The column was placed in a DNA LoBind tube, and 15μL of Elution solution was added. This column was left for 1 minute, and then centrifuged at 20°C and 12,000g for 1 minute. Another 15μL of Elution solution was added, left for 1 minute, and then centrifuged at 20°C and 12,000g for 1 minute. The DNA concentration of the filtrate was measured using NanoDrop2000 (Thermo Fisher Scientific) and diluted with elution solution to 200 ng / μL (100 ng / μL if the measured concentration was below 200 ng / μL).
[0104] 12. Annealing of Oligo DNA 0.5 μL each of sense and antisense oligo DNA (100 μM) synthesized for the template of the ngRNA target sequence was mixed with 1 μL of 10x Annealing Buffer (400 μM Tris-HCl (pH 8.0), 200 μM MgCl2, 500 μM NaCl) and 8 μL of autoclaved ultrapure water. After incubating at 95°C for 5 minutes using a thermal cycler, the temperature was lowered to 25°C over 90 minutes to anneal the oligo DNA.
[0105] 13. Insertion into a plasmid by Golden Gate reaction A mixture of 0.3 μL of 25 ng / μL plasmid, 0.5 μL of the oligo DNA annealed in step 12 above, 0.1 μL of BpiI, 0.1 μL of Quick Ligase (New England Biolabs), 0.2 μL of 10× T4 DNA Ligase Buffer (Takara Bio Inc.), and 0.8 μL of autoclaved ultrapure water was used. Using a thermal cycler, three cycles of restriction enzyme reaction at 37°C for 5 minutes followed by ligation reaction at 16°C for 10 minutes were performed, and the mixture was then stored at 4°C.
[0106] <Cell experiments> The cell experiments carried out in this test example were carried out using the following procedures: All cell experiment procedures were carried out in a clean bench, and the instruments used were previously treated with UV light for 15 minutes or more in the clean bench.
[0107] 1. Preparation of Complete Medium To 500 mL of D-MEM medium (High Glucose) containing L-glutamine and phenol red, 5.5 mL of 10x Non-Essential Amino Acid (NEAA), 5.5 mL of Penicillin-Streptomycin (ST-PN), and 55 mL of Fetal Bovine Serum (FBS) inactivated at 56°C for 30 minutes were added, mixed well, and stored at 4°C.
[0108] 2. Culturing HEK293T Cells The culture medium in the 100mm dish was removed using an aspirator, and 10mL of new complete medium was added. The cells were detached from the dish by pipetting with an electric pipettor, and a cell suspension was prepared. Approximately 9mL of complete medium and 1mL of cell suspension were added to a new 100mm dish, and the dish was shaken to homogenize the cells before being incubated at 37°C and 5% CO2. The amount of complete medium and cell suspension was adjusted according to the cell growth rate, and this procedure was repeated every 2-3 days to prevent the cells from overpopulating the dish.
[0109] 3. Lipofection using Lipofectamin LTX (Thermo Fisher Scientific) The day before or the day of lipofection, the plasmid to be transfected was diluted with autoclaved ultrapure water to a concentration of 90 ng / 6 μL or 200 ng / 6 μL, with a volume of 7.2 μL. 25 μL of D-MEM and 6 μL of the pre-adjusted plasmid were added to each well of a 96-well plate. Lipofectamin LTX and D-MEM were mixed in a reservoir to give 0.7 μL of each reagent, and then added to each well. The mixture was then incubated at room temperature for 30 minutes.
[0110] During this time, cells were prepared as follows: the culture medium of HEK293T cells cultured in a 100 mm dish was removed using an aspirator, and 3 mL of TrypLE TM Express was added to the dish and incubated for 30 seconds. 7 mL of complete medium was added to the dish, and the cells were detached using an electric pipettor. 3 mL of the cell suspension was transferred to a 50 mL tube. This tube was centrifuged at 20°C and 1000 rpm for 3 minutes, the supernatant was removed using an aspirator, and 1.5 mL of complete medium was added to resuspend the cells. 10 μL of the cell suspension was placed in a LUNA automated cell counter (Logos Biosystems) to measure the cell concentration, which was 1.5 × 10 5 The cells were diluted with complete medium to give a total of 100 cells / mL.
[0111] At the timing when the above 30-minute incubation ended, 100 μL of the cell suspension was added to each well, and the cells were cultured for 72 hours under the conditions of 37 °C and 5% CO2.
[0112] <Measurement of Editing Efficiency Using EditR> In this test example, the editing efficiency was measured by the following EditR.
[0113] 1. Cell Recovery When a GFP expression vector was introduced into the cells, before cell recovery, the 96-well plate was observed with a fluorescence microscope to confirm whether the plasmid was introduced into the cells. Then, the cell medium was removed, and 50 μL of TrypLE TM Express was added to each well and incubated at 37 °C for 5 minutes. 150 μL of 1×PBS(-) was added to all wells, and the cells were detached by pipetting. After transferring the entire cell suspension to a PCR tube, centrifugation was performed at 20 °C and 13,000 rpm for 3 minutes. The supernatant was discarded, 150 μL of 1×PBS(-) was added again, and centrifugation was performed at 20 °C and 13,000 rpm for 3 minutes. The supernatant was discarded, and 20 μL of Cell Lysis Buffer (manufactured by Invitrogen) and 0.8 μL of Protein Degrader (manufactured by Invitorogen) were added to each tube. A program of 60 °C for 15 minutes → 95 °C for 10 minutes was executed using a thermal cycler and stored at -20 °C.
[0114] 2. Genomic PCR and ExoSAP-IT Treatment Following the procedure described in step 3 above, PCR amplification of the region surrounding the target site was performed using Prime STAR Max DNA Polymerase (Takara Bio Inc.). In experiments targeting the RNF2 locus, primers RNF2-KN-F1 (SEQ ID NO: 2, Tm: 62.1°C) and RNF2-KN-R1 (SEQ ID NO: 3, Tm: 63.9°C) were used. Of the 10 μL of PCR product, 3 μL was used for agarose gel electrophoresis, and the remaining 7 μL was used for analysis. 1 μL of ExoSAP-IT (Applied Biosystems) was added to the 7 μL PCR product, and the thermal cycler was run at 37°C for at least 15 minutes followed by 80°C for 15 minutes, after which the product was stored at -20°C.
[0115] 3. Sequence Analysis According to the above basic operation 8, sequence analysis was performed using a SeqStudio Genetic Analyzer (manufactured by Thermo Fisher Scientific) to obtain waveform data of the nucleotide sequence including the edited portion.
[0116] 4. Measuring genome editing efficiency using EditR The waveform data obtained in step 3 above was entered into EditR (https: / / moriaritylab.shinyapps.io / editr_v10 / , Kluesner MG et al., 2018) and the genome editing efficiency (editing efficiency (%)) was calculated.
[0117] <Vector construction> 1. Construction of nCas9 vector, RTase vector, and PE vector First, we constructed "KS2" (nCas9 vector, SEQ ID NO: 4), a vector that expresses nCas9 under the control of the EF1a promoter, and "KS15" (RTase vector, SEQ ID NO: 5), a vector that expresses RTase under the control of a CMV enhancer / promoter, following the procedure described in Non-Patent Document 4. Schematic diagrams of the structures of "KS2" and "KS15" are shown in Figures 2 and 3, respectively. Here, "KS2" has an ngRNA promoter and a scaffold sequence, and the amino acid sequence of the RTase encoded by "KS15" is the amino acid sequence described in SEQ ID NO: 1, with S at position 92 replaced by R (S92R).
[0118] Additionally, for use as a positive control, a vector "#1373" (PE vector, SEQ ID NO: 6) expressing Prime Editor (a fusion protein of nCas9 and RTase) under the EF1a promoter was also constructed following Non-Patent Document 4. A schematic diagram of the structure of "#1373" is shown in Figure 4. Here, "#1373" has an ngRNA promoter and a scaffold sequence.
[0119] The target gene was RNF2, the effectiveness of which was verified in Non-Patent Document 1, the original paper on prime editing. Figure 5 shows a schematic diagram of the relative positions of the PAM, pegRNA recognition sequence (comp_pegRNA), and ngRNA recognition sequence (comp_ngRNA) to which the ngRNA binds, which are present on RNF2.
[0120] 2. Construction of ngRNA and insertion into PE vector To base the ngRNA on PE3max, as described in Non-Patent Document 3, we designed a separate ngRNA. First, an ngRNA expression cassette encoding ngRNA (SEQ ID NO: 9) was obtained by annealing the oligo DNA "RNF2-nick-s (sense strand, SEQ ID NO: 7)" and "RNF2-nick-as (antisense strand, SEQ ID NO: 8)" as described in Basic Procedure 12 above. The resulting ngRNA expression cassette was inserted into the nCas9 vector "KS2" and the PE vector "#1373" using the Golden Gate reaction described in Basic Procedure 13 above. Next, 0.1 μL of BpiI and 0.1 μL of 10× Buffer G were added to the total sample. The restriction enzyme reaction was carried out at 37°C for 1 hour using a thermal cycler, followed by heat inactivation of the restriction enzyme at 80°C for 5 minutes, digesting the plasmid without the ngRNA expression cassette.
[0121] Next, transformation was performed as in Basic Procedure 9 above, followed by plasmid extraction as in Basic Procedure 11, and the extracted plasmid was subjected to restriction enzyme treatment with BpiI as in Basic Procedure 6. The sample after restriction enzyme treatment was subjected to agarose electrophoresis as in Basic Procedure 1 above. The BpiI recognition sequence was lost, preventing cleavage and resulting in the appearance of multiple bands, confirming that the ngRNA expression cassette had been inserted.
[0122] The vector in which the ngRNA expression cassette was inserted into the nCas9 vector "KS2" was designated "KS2+ng," and the vector in which the ngRNA expression cassette was inserted into the PE vector "#1373" was designated "KS25+ng."
[0123] 3. Design of pegRNA and insertion into PE vector (creation of all-in-one vector) pegRNA is expressed from the U6 promoter and contains, from the 5' end, a 20-nt spacer sequence (protospacer), a scaffold sequence (gRNA scaffold), a 14-nt template sequence (RTT), a 15-nt primer binding sequence (PBS), a linker (pegLIT linker), tmpknot, and a U6 terminator. A schematic diagram of the structure of pegRNA is shown in Figure 6, and the nucleotide sequence of pegRNA is shown in SEQ ID NO: 10. Of the nucleotide sequence (191 nt) shown in SEQ ID NO: 10, bases 1 to 20 represent the spacer sequence, bases 21 to 96 represent the scaffold sequence, bases 97 to 110 represent the template sequence, bases 111 to 125 represent the primer binding sequence, and bases 134 to 184 represent tmpknot.
[0124] First, a vector expressing pegRNA targeting RNF2 was digested with Esp3I as described in step 6 above. The digested sample was subjected to agarose electrophoresis as described in step 1 above. The gel at the position of the band for the pegRNA expression cassette (SEQ ID NO: 11) encoding the pegRNA was excised, and DNA was extracted as described in step 2 above. Using the "KS2+ng" or "KS25+ng" prepared in step 2 above as vector DNA, ligation was performed as described in step 7 above. After transformation as described in step 9 above, colony PCR was performed as described in step 10 above using primers pegRNA-ColP-R (SEQ ID NO: 12, Tm: 60.7°C) and EF1α-proseq-R (SEQ ID NO: 13, Tm: 59.2°C). Next, after extracting the plasmid according to the above-mentioned basic procedure 11, restriction enzyme treatment was carried out using HindIII and PmeI according to the above-mentioned basic procedure 6. The restriction enzyme-treated sample was subjected to agarose electrophoresis according to the same procedure 1. It was confirmed that the pegRNA expression cassette had been inserted into the plasmid, a HindIII recognition site had appeared, and cleavage had occurred at two locations within the plasmid.
[0125] The vector in which a pegRNA expression cassette was inserted into "KS2+ng" was designated "KS2-1+ng+peg (SEQ ID NO: 14)," and the vector in which a pegRNA expression cassette was inserted into "KS25+ng" was designated "KS25-1+ng+peg (SEQ ID NO: 15)." Schematic diagrams of the structures of "KS2-1+ng+peg" and "KS25-1+ng+peg," respectively, are shown in Figure 7(a) and (b).
[0126] 4. PE vector efficiency verification The "KS25-1+ng+peg (all-in-one)" prepared in 3. above or the combination of the "KS25+ng" prepared in 2. above and a vector expressing pegRNA (separate) was introduced into HEK293T cells by Lipofection as described in Cell Experiment 3. above, and the editing efficiency was measured using EditR as described above. The editing efficiency was also measured using the pcDNA plasmid, which does not act on cells, and the pGFP plasmid, which expresses only GFP. The results are shown in Figure 8. As shown in Figure 8, there was no significant difference in editing efficiency between the all-in-one and separate types.
[0127] <Test Example 1> (1) Split body design We designed a split RTase (hereinafter sometimes referred to as "split RTase") in which RTase functions and primes editing only when all of the constituent fragments are present. Based on the description in Berrios, KN, et al., Nat Chem Biol 17, pp. 1262-1270, 2021 (Reference I), we first tested the regeneration of split RTase activity by utilizing the association of two halves of GFP, GFP1-10 (the N-terminal end) and GFP11 (the C-terminal end). Specifically, we constructed a split RTase consisting of a fusion protein in which the N-terminal fragment of RTase and GFP1-10 were linked via a GSAGSAAGSG linker (SEQ ID NO: 16), and a fusion protein in which the remaining C-terminal fragment of RTase and GFP11 were linked via a (GGGGS)3 linker (SEQ ID NO: 17). By cooperating this split body (a fusion protein with split GFP in this test example) with nCas9 and pegRNA, target DNA is edited by prime editing when both fusion proteins are present, as shown in Figure 9. In this test, ngRNA was also cooperating to improve the probability of adopting the edited strand during the DNA repair stage (Non-Patent Documents 1 and 3).
[0128] The split bodies are obtained by splitting a polypeptide consisting of the amino acid sequence set forth in SEQ ID NO: 1 or a polypeptide consisting of the amino acid sequence set forth in SEQ ID NO: 1 in which S at position 92 is substituted with R (S92R) as an RTase at the following seven positions: Split between E123 and D124: "SPL RT1" Split between E176 and M177 (S92R): "SPL RT2" Split between E233 and L234 (S92R): "SPL RT3" Split between T293 and P294 (S92R): "SPL RT4" Split between G405 and W406 (S92R): "SPL RT5" Split between L21 and G22 (S92R): "SPL RT+N" Split between E488 and E489 (S92R): "SPL RT+C" We verified the following.
[0129] (Construction of N-terminal vector) First, inverse-PCR was performed using PrimeSTAR Max DNA Polymerase (Takara Bio Inc.) with the RTase-expressing plasmid "KS15(S92R)" or "KS16 (KS15 with R92 reverted to S92)" as a template, according to the above-mentioned basic procedure 3. This performed amplification of the N-terminal fragment of the split RTase, excluding the unnecessary portion (here, the C-terminal fragment). The primer combinations used in each inverse-PCR, along with their Tm values and SEQ ID NOs. 2, are shown in Table 2. Furthermore, the region encoding GFP1-10 on the plasmid (SEQ ID NO: 18) was used as a template, and the region encoding GFP1-10 was amplified according to the above-mentioned basic procedure 3. The primer combinations used in this amplification, along with their Tm values and SEQ ID NOs. 2, are also shown in Table 2. The primers were designed so that after amplification, a 15-base homologous sequence would be generated at the end of each fragment, and the N-terminal fragment of RTase and GFP1-10 would be connected by the GSAGSAAGSG linker (Figure 10).
[0130] [Table 2]
[0131] The procedures after PCR were the same as those for inserting the ngRNA expression cassette into the PE vector in Vector Preparation 2 above, except that the insertion of GFP1-10 was confirmed using EcoRI and AgeI during restriction enzyme treatment. Expression vectors for fusion protein (N) in which the N-terminal fragment of each split RTase and GFP1-10 were linked with a GSAGSAAGSG linker were prepared: "SPL RT1(N)-GFP1-10 (SEQ ID NO: 35)," "SPL RT2(N)-GFP1-10 (SEQ ID NO: 36)," "SPL RT3(N)-GFP1-10 (SEQ ID NO: 37)," "SPL RT4(N)-GFP1-10 (SEQ ID NO: 38)," "SPL RT5(N)-GFP1-10 (SEQ ID NO: 39)," "SPL RT+N(N)-GFP1-10 (SEQ ID NO: 40)," and "SPL RT+C(N)-GFP1-10 (SEQ ID NO: 41)."
[0132] (C-terminal vector construction) First, inverse-PCR was performed using PrimeSTAR Max DNA Polymerase (Takara Bio Inc.) or PrimeSTAR GXL DNA Polymerase (Takara Bio Inc.) with the RTase-expressing plasmid "KS15" or "KS16" as a template, according to the above-mentioned basic procedure 3 or 4., to amplify the C-terminal fragment of the split RTase, excluding the unnecessary portion (here, the N-terminal fragment). The primer combinations used in each inverse-PCR, along with their Tm values and SEQ ID NOs. (Figure 3) show the primer combinations and their nucleotide sequences. Furthermore, the region encoding GFP11 (SEQ ID NO: 42) on the plasmid was used as a template, and the region encoding GFP11 was amplified according to the above-mentioned basic procedure 3 or 4. The primer combinations used in this amplification, along with their Tm values and SEQ ID NOs. (Figure 3). The primers were designed so that after amplification, a 15-base homologous sequence was generated at the end of each fragment, and the C-terminal fragment of RTase and GFP11 were connected by the (GGGGS)3 linker (Figure 11).
[0133] [Table 3]
[0134] The procedures after PCR were the same as for the preparation of the N-terminal vector described above, and expression vectors for fusion proteins (C) in which the C-terminal fragments of each split form of RTase and GFP11 were linked with a (GGGGS)3 linker were prepared: "GFP11-SPL RT1(C) (SEQ ID NO: 59)," "GFP11-SPL RT2(C) (SEQ ID NO: 60)," "GFP11-SPL RT3(C) (SEQ ID NO: 61)," "GFP11-SPL RT4(C) (SEQ ID NO: 62)," "GFP11-SPL RT5(C) (SEQ ID NO: 63)," "GFP11-SPL RT+N(C) (SEQ ID NO: 64)," and "GFP11-SPL RT+C(C) (SEQ ID NO: 65)."
[0135] (2) Verification of cell introduction and editing efficiency The expression vectors for each split body prepared in (1) above (the expression vector for fusion protein (N), the expression vector for fusion protein (C), or a combination of these) were introduced into HEK293T cells by Lipofection as described in Cell Experiment 3 above, along with the "KS2-1+ng+peg (vector expressing nCas9, ngRNA, and pegRNA)" prepared in Vector Preparation 3 above. The editing efficiency was measured using EditR as described above. As positive controls, the "KS25+ng" combination prepared in (2) above and a vector expressing pegRNA (separate), the non-cytosinophilic plasmid pcDNA, and the GFP-only plasmid pGFP were also used to measure the editing efficiency. The results are shown in Figure 12.
[0136] As shown in Figure 12, for the split forms other than "SPL RT+N," which was split near the N-terminus, and "SPL RT+C," which was split near the C-terminus, the editing efficiency of the N-terminal fragment (fusion protein (N)) alone or the C-terminal fragment (fusion protein (C)) alone was significantly low, and it was confirmed that each fragment alone did not exhibit reverse transcription activity ("(N)" and "(C)" in Figure 12). Furthermore, for three of the seven split forms, "SPL RT3," "SPL RT4," and "SPL RT5," the editing efficiency remained low even when the N-terminal fragment and the C-terminal fragment were combined ("(N) + (C)" in Figure 12). On the other hand, for two split forms, "SPL RT1" and "SPL RT2," the editing efficiency was improved by combining the N-terminal fragment (fusion protein (N)) and the C-terminal fragment (fusion protein (C)) ("(N) + (C)" in Figure 12). These results confirmed that "SPL RT1" and "SPL RT2" specifically do not exhibit reverse transcription activity on their own, and that reverse transcription activity can be restored only when both the N-terminal fragment and the C-terminal fragment are present.
[0137] <Test Example 2> (1) Design of additional split bodies The split body was prepared by using an RTase consisting of a polypeptide (S92R) consisting of an amino acid sequence in which S at position 92 of the amino acid sequence set forth in SEQ ID NO: 1 was substituted with R, and the RTase was further split at the following three positions: Split between P141 and S142 (S92R): "SPL RT1.1" Split between Q165 and P166 (S92R): "SPL RT1.2" Split between R205 and D206 (S92R): "SPL RT2.1" was also further verified.
[0138] (Construction of N-terminal vector) In the same manner as in the preparation of the N-terminal vector in Test Example 1 above, a region of the N-terminal fragment of the split RTase that did not contain the unnecessary portion (here, the C-terminal fragment) was amplified. The primer combinations used in each inverse-PCR, as well as their Tm values and SEQ ID NOs. indicating their nucleotide sequences, are shown in Table 4 below. Similarly, a region encoding GFP1-10 was amplified. The primer combinations used in this amplification, as well as their Tm values and SEQ ID NOs. indicating their nucleotide sequences, are also shown in Table 4 below. The primers were designed so that, after amplification, a 15-base homologous sequence would be generated at the end of each fragment, and the N-terminal fragment of RTase and GFP1-10 would be linked by the GSAGSAAGSG linker (Figure 10).
[0139] [Table 4]
[0140] Furthermore, after PCR, expression vectors for fusion protein (N) in which the N-terminal fragments of each split RTase and GFP1-10 were linked with a GSAGSAAGSG linker were prepared in the same manner as in the preparation of the N-terminal vector in Test Example 1 above: "SPL RT1.1(N)-GFP1-10 (SEQ ID NO: 74)," "SPL RT1.2(N)-GFP1-10 (SEQ ID NO: 75)," and "SPL RT2.1(N)-GFP1-10 (SEQ ID NO: 76)."
[0141] (C-terminal vector construction) In the same manner as in the preparation of the C-terminal vector in Test Example 1 above, a region of the C-terminal fragment of the split RTase that does not contain the unnecessary portion (here, the N-terminal fragment) was amplified. The primer combinations used in each inverse-PCR, as well as their Tm values and SEQ ID NOs indicating their nucleotide sequences, are shown in Table 5 below. Similarly, a region encoding GFP11 was amplified. The primer combinations used in this amplification, as well as their Tm values and SEQ ID NOs indicating their nucleotide sequences, are also shown in Table 5 below. The primers were designed so that, after amplification, a 15-base homologous sequence would be generated at the end of each fragment, and the C-terminal fragment of RTase and GFP11 would be linked by the (GGGGS)3 linker (Figure 11).
[0142] [Table 5]
[0143] Furthermore, after PCR, expression vectors for fusion proteins (C) in which the C-terminal fragments of each split form of RTase and GFP11 are linked with a (GGGGS)3 linker were prepared in the same manner as in the preparation of the C-terminal vector in Test Example 1 above: "GFP11-SPL RT1.1(C) (SEQ ID NO: 85)," "GFP11-SPL RT1.2(C) (SEQ ID NO: 86)," and "GFP11-SPL RT2.1(C) (SEQ ID NO: 87)."
[0144] (2) Introduction into cells The split expression vectors (fusion protein (N) expression vector, fusion protein (C) expression vector, or a combination of these) prepared in (1) above were introduced into HEK293T cells by Lipofection as described in Cell Experiment 3 above, along with the "KS2-1+ng+peg (vector expressing nCas9, ngRNA, and pegRNA)" prepared in Vector Preparation 3 above. Editing efficiency was measured using EditR as described above. As positive controls, the "KS25+ng" combination prepared in (2) above and a vector expressing pegRNA (separate), the non-cytosinase plasmid pcDNA, and the plasmid expressing GFP alone, pGFP, were also used to measure editing efficiency. The results are shown in Figure 13.
[0145] As shown in Figure 13, for all of the additionally prepared split bodies, the editing efficiency was significantly low when the N-terminal fragment (fusion protein (N)) or the C-terminal fragment (fusion protein (C)) was used alone ("(N)" and "(C)" in Figure 13), and the editing efficiency did not improve when the N-terminal fragment and the C-terminal fragment were combined ("(N) + (C)" in Figure 13). Therefore, it was confirmed that reverse transcription activity was specifically restored in "SPL RT1" and "SPL RT2" only when both the N-terminal fragment and the C-terminal fragment were present. [Industrial Applicability]
[0146] As described above, the present invention provides a novel split reverse transcriptase that is active only when all of its constituent fragments are present, a DNA editing system containing the same, a method for editing target DNA using the same, and a method for producing cells with edited target DNA. Such a reverse transcriptase is useful as a reverse transcriptase for prime editing and can improve the specificity of the prime editing.
Claims
1. The reverse transcriptase is a split reverse transcriptase containing two split N-terminal fragments and a C-terminal fragment, The N-terminal fragment and the C-terminal fragment are the following (a) to (c): (a) an N-terminal fragment and a C-terminal fragment obtained by splitting a polypeptide comprising the amino acid sequence set forth in SEQ ID NO: 1 at any one position within the region of positions 114 to 133 or any one position within the region of positions 167 to 186 of the amino acid sequence set forth in SEQ ID NO: 1; (b) an N-terminal fragment and a C-terminal fragment obtained by dividing a polypeptide comprising an amino acid sequence in which one or more amino acid residues in the amino acid sequence set forth in SEQ ID NO: 1 have been substituted, deleted, inserted and / or added at any one position within the region corresponding to positions 114 to 133 or any one position within the region corresponding to positions 167 to 186 of the amino acid sequence set forth in SEQ ID NO: 1; (c) an N-terminal fragment and a C-terminal fragment obtained by dividing a polypeptide comprising an amino acid sequence having 90% or more homology with the amino acid sequence set forth in SEQ ID NO: 1 at any one position within a region corresponding to positions 114 to 133 of the amino acid sequence set forth in SEQ ID NO: 1 or at any one position within a region corresponding to positions 167 to 186 of the amino acid sequence set forth in SEQ ID NO: 1; At least one selected from the group consisting of The reverse transcriptase activity is regenerated by the presence of the N-terminal fragment and the C-terminal fragment. Reverse transcriptase.
2. The following (A) to (C): (A) at least one selected from the group consisting of the reverse transcriptase according to claim 1, an expression vector for the reverse transcriptase, and a polynucleotide encoding the reverse transcriptase; (B) at least one Cas protein selected from the group consisting of a Cas protein (Cas) having nuclease activity and a nickase-type Cas protein (nCas) partially lacking nuclease activity; an expression vector for the Cas protein; and at least one polynucleotide encoding the Cas protein; (C) at least one selected from the group consisting of a prime editing guide RNA (pegRNA) of the Cas protein, an expression vector of the pegRNA, and a polynucleotide encoding the pegRNA; A DNA editing system comprising:
3. (D) a nicking guide RNA (ngRNA) for the Cas protein, wherein the ngRNA binds to an ngRNA recognition sequence that is a sequence different from the pegRNA recognition sequence to which the pegRNA binds; at least one selected from the group consisting of an expression vector for the ngRNA; and a polynucleotide encoding the ngRNA; The DNA editing system of claim 2, further comprising:
4. 1. A method for editing target DNA, comprising: contacting a target DNA with the DNA editing system of claim 2 or 3 to edit a target site in the target DNA; the pegRNA comprises, in order from the 5' side, a spacer sequence, a template sequence, and a primer binding sequence, the spacer sequence is a sequence that binds to a pegRNA recognition sequence that includes a complementary sequence of the target site, the template sequence is a sequence that serves as a template for the reverse transcriptase, and the primer binding sequence is a sequence that binds to a primer sequence on the 5' side of the target site.
5. A method for producing a cell with edited target DNA, comprising: a step of introducing the DNA editing system according to claim 2 or 3 into a cell, contacting the cell with a target DNA, and editing a target site in the target DNA; the pegRNA comprises, in order from the 5' side, a spacer sequence, a template sequence, and a primer binding sequence, the spacer sequence is a sequence that binds to a pegRNA recognition sequence that includes a complementary sequence of the target site, the template sequence is a sequence that serves as a template for the reverse transcriptase, and the primer binding sequence is a sequence that binds to a primer sequence on the 5' side of the target site.