Gene editing mediated by reverse transcription of modular RNA templates tethered by tagged gRNA, and its use
The gene editing complex with an RNA guide nickase, reverse transcriptase, and tethered RNA template addresses PAM motif and template size constraints, enabling precise and efficient editing of target sites with reduced bystander effects.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- RUTGERS THE STATE UNIV
- Filing Date
- 2024-07-01
- Publication Date
- 2026-07-29
AI Technical Summary
Existing gene editing technologies, such as CRISPR base editing and prime editing, face limitations in target site accessibility due to PAM motif requirements and template size constraints, leading to inefficiencies and potential bystander editing, especially when GC content is suboptimal.
A gene editing complex comprising an RNA guide nickase, reverse transcriptase, and a modular RNA template tethered by a matching gRNA (magRNA) that includes a tethering tag to recruit and anchor the RNA template, allowing for precise editing without DSBs and overcoming PAM motif limitations.
Enables versatile and efficient editing of target sites with reduced bystander effects, achieving precise point mutations, insertions, and deletions by utilizing a modular RNA template system that enhances editing efficiency and flexibility.
Smart Images

Figure 2026525246000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims the benefit of the earlier filing date of U.S. Provisional Patent Application No. 63 / 511,710, filed on July 3, 2023, under 119(e) of the U.S. Patent Act. The contents of that application are incorporated herein by reference in their entirety.
[0002] Government interests This invention was made with government support under MD200088, granted by the Department of Defense. The government has certain rights to this invention.
[0003] Reference to electronic sequence listings The contents of the electronic sequence listing (096738.00780SeqList.xml; size: 162,826 bytes; and creation date: June 28, 2024) are incorporated herein by reference in their entirety.
[0004] This disclosure relates to gene editing, related systems, and their use. [Background technology]
[0005] Targeted gene editing has been used for genetic manipulation of eukaryotic cells, embryos, and animals. Sequence-specific nucleases, including Talen, zinc finger nucleases, and RNA-guided nucleases such as CRISPR / Cas9, have provided tools for precise genome editing. After binding to the target sequence, the native nuclease generates DNA double-strand breaks (DSBs), triggering cellular DNA repair pathways, including non-homologous end joining (NHEJ) and homology-directed recombination (HDR). As a result, the desired sequence alteration or gene editing can be achieved.
[0006] During the editing process, on-target and off-target DNA DSB intermediates can cause chromosomal translocations and other mutagenic events with potential carcinogenic tendencies. To avoid the requirements of DNA DSBs, base editing platforms have been developed. CRISPR base editing factors utilize nuclease-null or nickase versions of the CRISPR protein. Mutant CRISPR proteins have superior DNA sequence recognition and do not cause DSBs. Instead, the mutant CRISPR protein-gRNA complex recruits cytidine deaminase or adenine deaminase, which then convert C to U or A to G at the target, respectively, resulting in sequence-specific point mutations. By avoiding DSBs, base editing reduces the carcinogenic tendency and is widely used for precision genome editing for both basic research and therapeutic development.
[0007] Because nucleotide deamination is the underlying mechanism for base changes, base editing factors edit transition site mutations but not transversion mutations. In addition, base editing factors require the target editing nucleotide to be in the R-loop adjacent to the PAM motif. This requirement excludes the possibility of reaching target sites that do not have an adjacent PAM motif. Furthermore, within the editing activity window, deamination is generally messy, which leads to bystander base editing within the window. While bystander base editing may be harmless in the development of some therapies, such as the correction of loss-of-function mutations, precise editing without bystander editing is generally preferred in many other situations.
[0008] Prime editing is another precision gene editing platform that does not require double-strand DNA breaks (Anzalone, AV, et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019)). The Prime editing factor complex includes a nickase version of the CRISPR protein fused with reverse transcriptase (RT) and a modified gRNA called pegRNA, which typically contains an RNA template and primer-binding sequence at the 3' end of the gRNA. After binding to the target DNA sequence, the Prime editing factor creates a DNA nick. The nicked DNA strand binds to the primer-binding sequence in the pegRNA, which acts as a primer that synthesizes a new DNA sequence using the RNA template in the pegRNA. The sequence information from the template RNA is then copied from RNA to DNA by reverse transcriptase. Subsequently, the DNA sequence is further incorporated into the target site. By copying the desired RNA template at the 3' end of pegRNA, prime editing can achieve bystander-free, precise genome editing of base pairs. These base pairs can be both transitions and transversions. Furthermore, by designing insertions and deletions in the pegRNA template, prime editing can also create insertions and deletions at target locations.
[0009] At the heart of the prime editing system is a modified gRNA, named pegRNA, which contains both the gRNA scaffold, primer-binding sequence, and editing template within the same RNA molecule. However, this configuration limits the size of the editing template because the long RNA template sequence within the same RNA molecule of the gRNA can create secondary structures that may interfere with, for example, the secondary structure of the gRNA CRISPR protein-binding scaffold. The longest template attempted within a pegRNA molecule was 34 nt in length, as reported by Anzalone, AV, et al. Nature 576, 149-157 (2019). This template size limitation prevents prime editing factors from reaching certain target sites in the genome due to the lack of nearby suitable PAM motifs. Even when a PAM motif is within the limits of the RT (reverse transcriptase) template size limit, the template size limit can sometimes restrict the optionality of RT template selection. This is because sequences adjacent to the PAM may have suboptimal GC content, and this unfavorable GC content for priming and extension may make certain target sites in the genome impractical for prime editing.
[0010] One variant configuration of prime editing separates pegRNA into two molecules: an unmodified gRNA and an RT template containing an exogenous accessory RNA aptamer structure, e.g., MS2 (WO2020 / 191248A1). In this configuration, the CRISPR protein requires the inclusion of an additional fusion partner to interact with the RNA aptamer within the RT template; for example, an MCP protein is fused to the Cas9 protein to recruit the MS2 RNA aptamer-containing RT template. Within the system, the gRNA is responsible for complex formation with the CRISPR protein for target site recognition. RNA aptamers such as MS2 are recruited into the CRISPR complex by the additional fusion portion (e.g., MCP). Prime editing factors with this configuration exhibit a fraction of the editing efficiency of their pegRNA prime editing factor counterparts (Figure 73 in WO2020 / 191248A1). Furthermore, the ratio of unwanted insertions / deletions to correct edits is higher than that of pegRNA prime editing factors. WO2020 / 191248A1 also showed that including an accompanying RNA aptamer in the RT template, particularly at the 5' end, could potentially result in unwanted copies of exogenous genetic information (RNA aptamer) at the target site. [Overview of the project]
[0011] There is a need for more versatile and effective systems for editing target nucleic acid molecules.
[0012] This disclosure addresses, in many ways, the needs mentioned above.
[0013] In one embodiment, the disclosure provides a gene editing complex for editing a target site in a target DNA molecule. The gene editing complex comprises (A) an RNA guide nickase; (B) a reverse transcriptase; (C) an RNA template molecule; and (D) a matching gRNA (magRNA) molecule. The RNA template molecule comprises (1) a template segment containing a template sequence complementary to the desired DNA sequence or its complement to be introduced into the target site, and (2) a priming segment complementary to the 3' end of the nicking strand of the target DNA molecule. The magRNA molecule comprises (1) a guide or spacer sequence complementary to the sequence on the target strand of the target DNA molecule; (2) an RNA scaffold capable of binding to the RNA guide nickase; and (3) a tethering tag sequence complementary to the segment in the RNA template.
[0014] The terms "tethering tag" and "matching tag" are used interchangeably in this application. Both are referred to as tag sequences attached to gRNA that complementarily match a region in a modular RNA template in order to recruit and tether the RNA template.
[0015] The reverse transcriptase can be incorporated into the complex by any preferred means. In one embodiment, the reverse transcriptase is ligated to or fused to an RNA guide nickase (by covalent or non-covalent bond). In another embodiment, the magRNA further comprises a protein-binding motif (e.g., MS2 or PP7) that can bind to an RNA-interacting protein, and the reverse transcriptase is ligated to or fused to an RNA-interacting protein (e.g., MCP or PCP) (by covalent or non-covalent bond).
[0016] In one embodiment, the tethering tag in magRNA is not poly(N) where N is a repeating nucleotide A, C, U, or G.
[0017] In one embodiment, the RNA template does not include any additional sequences other than the priming segment and the template segment.
[0018] In one embodiment, the RNA guide nickase is a nickase variant of a Cas protein or a nickase variant of an RNA guide endonuclease protein encoded by a transposon. In one embodiment, the Cas proteins include Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cpf1 (Cas12a), C2c1 (Cas12b), C2c3 (Cas12c), CasY (Cas12d), CasX (Cas12e), Cas14 (Cas12f), CasPhi (Cas12j), Cas13a, Cas13b, Cas13c, Cas13d, Cas13x, CasF, and CasG. Selected from CasH, Csy1, Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3 (or CasE), Cse4 (or CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1, Csx15, Csf1, Csf2, Csf3, Csf4, Cu1966, and their orthologs. In one embodiment, the endonuclease protein encoded by the transposon is an IscB, IsrB, or TnpB endonuclease protein and its ortholog.
[0019] In one embodiment, the RNA guide nickase is a nickase variant of an ortholog of a Cas protein listed above, or an ortholog of an RNA guide endonuclease protein encoded by a transposon. In one embodiment, the Cas protein is Cas9. In one embodiment, the nickase variant of the Cas9 ortholog protein is Streptococcus pyogenes(sp)nCas9(H840A) or Staphylococcus aureus(sa)nCas9(N580A). In one embodiment, the nickase variant of the Cas9 ortholog protein is Staphylococcus aureus nCas9(N580A E782K / N968K / R1015H).
[0020] In one embodiment, the reverse transcriptase is a naturally occurring reverse transcriptase derived from a retrovirus or retrotransposon, or a variant thereof. In one embodiment, the reverse transcriptase is selected from the group consisting of Moloney's mouse leukemia virus (M-MLV), human immunodeficiency virus (HIV) reverse transcriptase, avian sarcoma-leukemia virus (ASLV) reverse transcriptase, Rous sarcoma virus (RSV) reverse transcriptase, avian myeloblastosis virus (AMV) reverse transcriptase, avian erythroblastosis virus (AEV) helper virus MCAV reverse transcriptase, avian myelocytomatosis virus MC29 helper virus MCAV reverse transcriptase, avian reticuloendotheliosis virus (REV-T) helper virus REV-A reverse transcriptase, avian sarcoma virus UR2 helper virus UR2AV reverse transcriptase, avian sarcoma virus Y73 helper virus YAV reverse transcriptase, Rous-associated virus (RAV) reverse transcriptase, and myeloblastosis-associated virus (MAV) reverse transcriptase, Line 1 ORF2, R2Bm, and R2Ol. In one embodiment, the reverse transcriptase is MMLV-RT or a variant thereof.
[0021] In one embodiment, retrovirus- and retrotransposon-derived reverse transcriptases have intrinsic DNA-dependent DNA polymerase activity, such as Line 1 ORF2 and HIV RT. In one embodiment, the retrotransposon-derived reverse transcriptase is R2Bm.
[0022] In one embodiment, the tethering tag sequence may be present at the 3' or 5' end of the magRNA. In one embodiment, the tethering tag sequence is covalently linked to the magRNA via a polynucleotide linker. In one embodiment, the tethering tag sequence in the magRNA is 6 to 24 nucleotides in length.
[0023] In one embodiment, the tethering tag sequence of the magRNA is complementary to the 5' end of the RNA template molecule. In one embodiment, the tethering tag sequence of the magRNA is complementary to an internal region of the RNA template.
[0024] In a second aspect, the present disclosure features a system for editing a target site in a target DNA molecule. The system includes (I) a first gene editing complex of any one of those described above, and (II) a second gene editing complex.
[0025] In one embodiment, the second gene editing complex includes (A) a second RNA-guided nickase, and (B) a gRNA molecule containing a second guide or spacer sequence, with or without an RNA aptamer (e.g., MS2).
[0026] In one embodiment, the second gene editing complex includes (A) a second RNA-guided nickase, and (B) a second magRNA molecule containing (1) a second guide or spacer sequence, (2) a second RNA scaffold capable of binding to the second RNA-guided nickase, and (3) a second tethering tag sequence.
[0027] In one embodiment, the second gene editing complex includes (C) a second reverse transcriptase.
[0028] In one embodiment, the two guide sequences in the two gene editing complexes (i.e., the guide sequence for magRNA in the first gene editing complex, and the second guide sequence for gRNA or the second magRNA in the second gene editing complex) are complementary to two target sequences of the target DNA molecule.
[0029] In one embodiment, the second tethering tag sequence of the second magRNA is complementary to the second segment in the RNA template.
[0030] In one embodiment, the 5' end of the RNA template contains the same sequence as the 3' end of the strand of the target DNA molecule that is nicked by a second RNA guide nickase.
[0031] In one embodiment, the second gene editing complex further comprises (C) a second reverse transcriptase or (D) a second RNA template, or both.
[0032] The second reverse transcriptase can be incorporated into the second gene editing complex or system by any preferred means. In one embodiment, the second reverse transcriptase can be ligated (covalently or noncovalently) to or fused to a second RNA guide nickase. In one embodiment, the second magRNA or gRNA molecule comprises a second protein-binding motif capable of binding to a second RNA-interacting protein, and the second reverse transcriptase is ligated (covalently or noncovalently) to or fused to the second RNA-interacting protein.
[0033] In one embodiment, the second gRNA or magRNA in the second gene editing complex contains an RNA aptamer (e.g., MS2 or PP7) that recruits a non-reverse transcriptase effector fused with an RNA aptamer-binding protein (e.g., MCP or PCP). In one embodiment, the non-reverse transcriptase effector is a 50-kDa RNase inhibitor 1 protein (RNH1) or its ortholog. In one embodiment, the non-reverse transcriptase effector is a dominant-negative protein of the MMR pathway proteins, including the MLH1 protein, or the 5' DNA nuclease Fen1 protein.
[0034] In one embodiment, the non-reverse transcriptase effector is a pioneer transcription factor. In one embodiment, the pioneer transcription factor is p65.
[0035] In one embodiment, the non-reverse transcriptase effector is a chromatin-modulating protein or protein complex. In one embodiment, the chromatin-modulating protein has histone acetylation activity.
[0036] Methods for modifying a target DNA molecule in a cell are also provided in this disclosure. The method includes the step of contacting the target DNA molecule with the gene editing complex or system described above. In one embodiment, the modification results in a point mutation, insertion, deletion, or combination thereof to the target DNA molecule. In one embodiment, the cell is selected from the group consisting of archaeal cells, bacterial cells, eukaryotic cells, eukaryotic unicellular organisms, somatic cells, germ cells, stem cells, plant cells, algal cells, animal cells, invertebrate cells, vertebrate cells, fish cells, frog cells, avian cells, mammalian cells, pig cells, bovine cells, goat cells, sheep cells, rodent cells, rat cells, mouse cells, non-human primate cells, and human cells. In one embodiment, the cell is present in, isolated from, or derived from a human or non-human subject.
[0037] In further embodiments, the disclosure provides genetically engineered cells or their offspring obtained according to the methods described above. In one embodiment, the cells are selected from the group consisting of archaeal cells, bacterial cells, eukaryotic cells, eukaryotic unicellular organisms, somatic cells, germ cells, stem cells, plant cells, algal cells, animal cells, invertebrate cells, vertebrate cells, fish cells, frog cells, avian cells, mammalian cells, pig cells, bovine cells, goat cells, sheep cells, rodent cells, rat cells, mouse cells, non-human primate cells, and human cells. In one embodiment, the cells are selected from the group consisting of pluripotent stem cells (PSCs), adult stem cells (ASCs), hematopoietic stem cells (HSCs), fibroblasts, chondrocytes, keratinocytes, hepatocytes, islet cells, and immune cells including T cells, dendritic cells (DCs), natural killer (NK) cells, and macrophages, isolated from or derived from human or non-human subjects.
[0038] In another embodiment, the present disclosure features the magRNA described above. The magRNA comprises (1) a guide or spacer sequence complementary to a target sequence on the target strand of a target DNA molecule, (2) an RNA scaffold capable of binding to an RNA guide nickase, and (3) a tethering tag sequence complementary to a segment in an RNA template. The magRNA may further include a protein-binding motif capable of binding to an RNA-interacting protein, and a reverse transcriptase is linked to the RNA-interacting protein. In one embodiment, the tethering tag in the magRNA is not poly(N) where N is a repeating nucleotide A, C, U, or G. In one embodiment, the tethering tag sequence is located at the 3' or 5' end of the magRNA. In one embodiment, the tethering tag sequence is covalently linked to the magRNA via a polynucleotide linker. In one embodiment, the tethering tag sequence in the magRNA is about 6 to 24 nucleotides long.
[0039] In another embodiment, the present disclosure features an RNA complex comprising the magRNA described above and an RNA template tethered to the magRNA described above.
[0040] The tethered RNA template does not contain any extra sequences other than the desired sequence designed to be present within the edited genome.
[0041] In another aspect, the disclosure provides nucleic acids encoding (i) the magRNA molecules described above, and (ii) one or two of the RNA complexes described above.
[0042] In another embodiment, the disclosure features a vector containing nucleic acids.
[0043] In another embodiment, the disclosure features (i) a packaging material and (ii) a kit comprising one, two or more of the gene editing complex, the system, the cells, the magRNA molecule, the RNA complex, the nucleic acid, and the vector described above.
[0044] In another embodiment, the disclosure provides a pharmaceutical composition comprising (i) a pharmaceutically acceptable carrier and (ii) one, two or more of the gene editing complex, the system, the cell, the magRNA molecule, the RNA complex, the nucleic acid, and the vector described above.
[0045] Details of one or more embodiments of this disclosure are set forth below. Other features, objectives, and advantages of this disclosure will become apparent from this specification and the claims. [Brief explanation of the drawing]
[0046] [Figure 1]Figures 1A, 1B, and 1C illustrate a single-module match editing system using CRISPR / gRNA for RNA-guided sequence recognition. The system comprises (Figure 1A) a modular RNA template that does not contain any exogenous RNA recruiting elements, e.g., RNA aptamers (e.g., MS2 or PP7); (Figure 1B) a magRNA containing a 5' spacer sequence for target DNA recognition, an RNA scaffold for CRISPR / Cas complex formation, and a 3' match RNA sequence complementary to the region in the modular template; and (Figure 1C) a nickase CRISPR protein, nCRISPR / nCas9 (H840A), fused with reverse transcriptase (MMLV-RT) via a polypeptide linker. [Figure 2] Figure 2 illustrates how the match editing system works. nCRISPR / nCas9-RT forms a complex with magRNA. The 5' spacer sequence of magRNA recognizes and complements the target sequence in the target DNA; while the 3' tag sequence of magRNA complements the region in the RNA template, anchoring the RNA template to the CRISPR-RT / magRNA complex. The niccasse activity of nCRISPR / nCas9 causes a single-strand DNA break (nicking) upstream of the PAM motif. The nicking DNA then binds to the 3' end of the RNA template by sequence complementarity, further strengthening the complex. The free 3' strand of the nicking DNA then acts as a primer for reverse transcriptase, using the complex-forming RNA as a nucleotide elongation template. [Figure 3]Figure 3 illustrates how the cell's repair mechanism results in the incorporation of a newly synthesized DNA strand into the target DNA locus. Figure 3A shows the relative arrangement of the RNA-guided nicking site (indicated by "nicking" and shown as a star on the upper DNA strand) and the desired mutation site (indicated by "mutation" and shown as a complementary arrow on the DNA strand). Figures 3B–3D show the synthesis of a new DNA strand extending the 3' nicking strand, followed by the removal of the original DNA fragment at the 5' flap, resulting in a new upper strand with the desired sequence change and nicks. Figure 3E shows the resolution of mismatches and filling of the nicking gap by the cell's repair mechanism, resulting in either a DNA molecule with the desired edits (by removing mismatches on the lower strand) or DNA with the original sequence (by removing mismatches on the upper strand). [Figure 4] Figure 4 shows an example of a magRNA-RNA template interaction in which the 3'-matching tag sequence of magRNA is specifically complementary to the 5' end sequence of the RNA template. [Figure 5] Figure 5 shows a magRNA with a 5' match tag that is complementary to the RNA template region and located at the 5' end of the gRNA. A linker polynucleotide sequence is positioned between the 5' match tag sequence and the spacer sequence. The editing mechanism is similar to that described in Figure 2, except for the interaction site between the magRNA tag sequence and the template RNA. [Figure 6] Figures 6A, 6B, 6C, and 6D illustrate a variant-match editing system in which the reverse transcriptase (RT) is provided in a fragmented state rather than being directly fused to an RNA guide sequence-specific nickase.
[0047] Figure 6A shows the modular RNA template.
[0048] Figure 6B shows the manipulated magRNA. The magRNA contains a 5' spacer sequence for target sequence recognition, a 3' tag sequence for recruiting an RNA template (Figure 6A), and an RNA aptamer (e.g., MS2) in the gRNA stem-loop region (SL) for individually recruiting reverse transcriptase.
[0049] Figure 6C shows RNA guide niccas without directly ligated reverse transcriptase.
[0050] Figure 6D shows a protein containing a reverse transcriptase fused via a polypeptide linker to a cognitive protein (e.g., MCP) that binds to an aptamer (e.g., MS2) within a gRNA scaffold. [Figure 7] Figures 7A, 7B, 7C, and 7D illustrate how reverse transcriptase functions in the divided system.
[0051] Figure 7A shows the modular RNA template.
[0052] Figure 7B shows the manipulated magRNA.
[0053] Figure 7C shows RNA guide nickase.
[0054] Figure 7D shows a protein containing reverse transcriptase fused via a polypeptide linker to a cognitive protein (e.g., MCP) that binds to an aptamer (e.g., MS2) within the gRNA scaffold. The process is similar to the representation in Figure 2, except that RT is recruited in a split state. [Figure 8] Figure 8 shows a variant configuration of the split-state RT system in which the 3' magRNA elongation sequence is specifically complementary to the sequence at the 5' end of the RT template. [Figure 9A]Figures 9A and 9B show examples of dual ME systems. Figure 9A shows a dual ME system with fused RT and magRNA-gRNA pair formation. In this configuration, the first ME contains typical ME components, including RT and magRNA fused with nCRISPR / nCas9. The second module is for creating a second nick to increase editing efficiency. The gRNA in the second module does not have a 3' tag sequence for RNA template interaction. [Figure 9B] Figures 9A and 9B show examples of dual ME systems. Figure 9B shows a dual ME system having trans RT and a magRNA with MS2-gRNA 0xMS2. In this configuration, the first ME contains a magRNA with an aptamer for recruiting a split reverse transcriptase. The second module is for creating a second nick to increase editing efficiency. The gRNA in the second ME module does not contain a 3' matching tag for RNA template recruitment. [Figure 10] Figure 10 shows that a second nickeling in the nearby opposite strand, mediated by the dual ME system, significantly enhances the efficiency of targeted editing. Figures 10A–10C are identical to Figures 3A–3C. Figure 10D shows that a second nickeling site is created in the lower strand, which favors the removal of DNA flaps containing a free 5' nickeling end over DNA flaps containing a free 3' nickeling end. As a result, it becomes unfavorable to return to the original DNA sequence (Figure 10E, right, with stop sign), and the product is primarily a DNA molecule with the desired edit (Figure 10E, left). [Figure 11A] Figures 11A and 11B show dual ME systems coupled to a single template for gene editing using large DNA fragments. Figure 11A shows a coupled dual ME system with fused RTs. [Figure 11B]Figures 11A and 11B show a dual ME system coupled to a single template for gene editing with a large DNA fragment. Figure 11B shows a coupled dual ME system with a split state RT. A large reverse transcriptase RNA template is supplied in this configuration to enable editing of the desired sequence at a location further away from the RNA guide-nicking site generated by the first ME module. In this configuration, the second ME module contains a magRNA with a 3' tag sequence that interacts with the RNA template at different segments. The difference between Figures 11A and 11B is that in Figure 11A, the reverse transcriptase is covalently linked to nCRISPR / nCas9, while in Figure 11B, the RT is supplied individually. [Figure 12A] Figures 12A and 12B show exemplary versions of a coupled dual ME system with a second priming mechanism. Figure 12A shows an RT fusion version designed so that the 5' end of the RNA template is identical to the 3' end of the second nickeling strand. As a result, the second nickeling strand functions as a primer for the synthesis of the second DNA strand, using the newly synthesized DNA strand as a template. [Figure 12B] Figures 12A and 12B show exemplary versions of a coupled dual ME system with a second priming mechanism. Figure 12B shows a split-state RT version. [Figure 13]Figure 13 shows a scheme for large fragment insertion involving second-strand DNA-dependent DNA synthesis by a dual ME system coupled to a single template. Figure 13A shows the first RNA-guided nicking by the first ME module. Figure 13B shows the first DNA strand synthesis by the first ME module using a modular RNA template. Figures 13C-13D show the introduction of the second nicking by the second ME module. Figure 13E shows priming of the newly synthesized DNA strand and the free 3' end of the second nicking strand. Figure 13F shows strand elongation and synthesis of the second DNA strand by reverse transcriptase DNA-dependent DNA polymerase activity or by the cell's DNA-dependent DNA polymerase. Figures 13G-13H show the completion of insertion of the newly synthesized DNA fragment by removing two DNA flaps containing free 5' ends and then ligation by the cell's repair mechanism. [Figure 14] Figure 14 shows a dual template dual-module system for insertion and / or deletion. In this configuration, both ME modules contain magRNA and RNA templates. The reverse transcriptase may be supplied fused to nCRISPR / nCas9 or in a fragmented state. [Figure 15] Figures 15A, 15B, 15C, 15D, 15E, and 15F illustrate how dual-template dual-module MEs produce de novo gene compositions (insertions) or large fragment gene deletions.
[0055] Figures 15A and 15B show that module 1 ME and module 2 ME mediate nicking at their corresponding target sites, leading to first-strand DNA synthesis via reverse transcription.
[0056] Figure 15C shows the annealing of the 3' ends of the two newly synthesized first DNA strands via complementary sequences.
[0057] Figure 15D shows the synthesis of second-strand DNA by reverse transcriptase DNA-dependent DNA polymerase activity or endogenous cellular DNA polymerase activity.
[0058] Figures 15E and 15F illustrate the removal and ligation of the 5' flap. This mechanism can result in both insertions and deletions (de novo gene composition), depending on the two RNA templates and the content and length of the sequence being removed. [Figure 16] Figures 16A to 16D show an ME system with a second effect pedal.
[0059] Figure 16A shows the modular RNA template.
[0060] Figure 16B shows a magRNA having a 5' spacer sequence for target site recognition, a 3' tag sequence for reverse transcription RNA template recruitment, and an aptamer sequence (e.g., MS2) in a stem-loop configuration for the recruitment of a second effector.
[0061] Figure 16C shows an RNA guide niccas (e.g., nCRISPR / nCas9) fused with a reverse transcriptase (e.g., MMLV-RT) via a linker peptide.
[0062] Figure 16D shows a second effector protein (e.g., an RNase inhibitor or FEN1) fused to an aptamer-binding protein (e.g., MCP) via a peptide linker. [Figure 17] Figures 17A, 17B, 17C, and 17D show the components of a match editing system for correcting the point mutation A200G in the nfEGFP gene.
[0063] Figure 17A shows the components of the match editing system and the gRNA control expression plasmid.
[0064] Figure 17B shows the nfEGFP gene to be edited.
[0065] Figure 17C shows the sequences of the target nicking sites (sequences 1 and 2), with PAM (underlined) and nicking sites (arrows) indicated; mutation G200 is enclosed in a square.
[0066] Figure 17D shows the primer binding sequence (P14, SEQ ID NO: 3) and the match editing template having RT extensions of 29 nt or 57 nt in length. [Figure 18] Figures 18A, 18B, and 18C demonstrate the effectiveness of the match editing system in correcting the point mutations shown in Figure 17.
[0067] Figure 18A shows the results of a functional assay for the match editing system in correcting point mutations.
[0068] Figures 18B and 18C show the Sanger sequencing results (SEQ ID NO: 4) from lanes 2 and 4 of Figure 18A, respectively. As indicated, a shorter template (Ex29) with a 29nt RT extension was used for lanes 3 and 4 with gRNA or magRNA. As indicated, a longer template was used for lanes 5 and 6 with gRNA or magRNA. The tag sequence of the magRNA is 12nt long and matches the 5' end sequence of the corresponding template. PEG RNA-mediated prime editing (lane 7) was used for comparison. [Figure 19] Figures 19A and 19B show a comparison of the editing effectiveness of magRNAs matching the internal sequence (12M29) versus the 5'-sequence (12M57) of the long template Ex57.
[0069] Figure 19A shows the results of fluorescence microscopy observation.
[0070] Figure 19B shows the percentage of cells expressing fluorescent EGFP, as determined by flow cytometry. Panels 1–7 in Figure 19A are the same corresponding panels in Figure 19B. For each panel, the treatment is indicated in Figure 19B. [Figure 20] Figures 20A, 20B, and 20C demonstrate the effectiveness of ME in correcting deletion mutations. Panels 1-5 in Figure 20A correspond to the same panels in Figure 20B.
[0071] Figure 20A shows the results of fluorescence microscopy observation.
[0072] Figure 20B shows the percentage of cells expressing fluorescent EGFP as determined by flow cytometry. Panel 1, untreated cells; Panel 2, electroporated cells with EGFP deletion mutants (equomolar mixture of Δ4A, Δ20C, and Δ35C); Panels 3, 4, and 5, cells expressing ME components with indicated EGFP deletion, template, and magRNA(12M29).
[0073] Figure 20C shows the DNA sequencing results (SEQ ID NO: 5) from panel 3 cells. [Figure 21] Figure 21 shows a comparison of the editing effectiveness of magRNA matching the internal (12M29) versus 5' end sequence (12M57) of the long template Ex57 for insertion editing. Panel 1, untreated cells; Panel 2, electroporated cells with equal mixtures of EGFP deletion plasmids Δ35C and Δ50C; Panels 3-6, cells expressing ME components with the indicated EGFP plasmids, template, and magRNA. [Figure 22] Figure 22 shows the insertion editing rate of ME at a position 100 nucleotides from the target nicking site (SEQ ID NO: 6) using a template with a RT elongation length of 111 nt and a magRNA (12M29) matched to the internal sequence. [Figure 23] Figure 23A compares the match editing efficiency in correcting deletion mutations in the EGPF gene using magRNAs with varying matching lengths (9nt to 24nt).
[0074] Figure 23B compares the match editing efficiency in correcting deletion mutations in the EGFP gene using magRNAs of the same length (12 nt) but with varying matching positions (a 3' tag position complementary to the template's 23 nt to 41 nt upstream (5') of the RT start site). [Figure 24] Figures 24A, 24B, and 24C show the transRT match editing system, its components, and its efficiency in correcting deletions in EGFP.
[0075] Figure 24A shows the trans-RT ME system and its components, where the reverse transcriptase is not fused with nCas9. The magRNA contains the MS2 aptamer at the stem-loop position. RT is fused to MCP via a linker peptide.
[0076] Figure 24B shows the EGFP gene, which is the gene to be edited, and contains a C deletion 50 nucleotides upstream of the 200-L nicking site.
[0077] Figure 24C shows the effectiveness of the transRT match editing system in correcting deletions compared to the direct fusion ME system. [Figure 25] Figures 25A, 25B, and 25C demonstrate the effectiveness of the dual ME system and its editing capabilities through magRNA-gRNA pair formation.
[0078] Figure 25A shows the components of the dual ME system.
[0079] Figure 25B shows the two second module regions (SEQ ID NOs: 7 and 8) that were examined. The second module ME contains gRNA instead of magRNA.
[0080] Figure 25C shows the editing effectiveness of single-module ME and dual ME. [Figure 26]Figures 26A, 26B, and 26C show that match editing effectively transformed G to A at the intrinsic site and the HEK4 site.
[0081] Figure 26A shows the results from untreated cells (SEQ ID NO: 9).
[0082] Figure 26B shows the results from electroporated cells of an ME system containing HEK4 target sites with unmatched gRNA (SEQ ID NO: 9).
[0083] Figure 26C shows the results from electroporated cells of an ME system containing a HEK4 target site with a magRNA having a matching tag for the template (SEQ ID NO: 9). [Figure 27] Figure 27 shows that match editing was effective in deletion editing. The mutant EGFP gene contains a 4-nt insertion starting at position 180. An ME template with a 29nt RT elongation length was used for ME editing together with 12M29 magRNA containing a 12-nt matching tag to the 5' end of the template. Prime editing experiments with pegRNA were performed as a control. [Figure 28] Figures 28A and 28B compare the efficiency of deletion editing using single-module ME, dual-module ME, single-module primer editing factor (PE2), and dual-module PE (PE3) in a reporter construct that results in exon 23 skipping at the splicing donor site (SDS) of mouse dystrophin gene exon 23.
[0084] Figure 28A shows a schematic diagram of a GFP-based splicing reporter.
[0085] The upper expression construct contains an EGFP gene split into 5' and 3' halves, interrupted by artificial introns containing splicing donor sequences (SDS) and splicing acceptor sequences (SAS). Transcription and splicing of the splicing reporter gene will result in the production of functional EGFP.
[0086] The expression vector at the bottom is constructed by inserting mouse dystrophin gene exon 23, which contains intron 22 and intron 23 splicing regulatory sequences. Transcription and splicing of the reporter gene will result in non-fluorescent production of the EGFP protein due to the inserted peptide encoded by exon 23. Deletion of Dmd exon 23 SAS or SDS will result in exon 23 skipping and lead to the production of fluorescent EGFP.
[0087] Single-module MEs and dual-module MEs with magRNA-gRNA pairing were designed to delete the 29nt sequence at the Dmd exon 23-intron 23 junction and eliminate the SDS sequence, as indicated. Primer editing factors PE2 and PE3, with the same RT elongation length and primer binding site, were designed for comparison. The nicking site for the first ME or PE2 (first nicking) and the second nicking site for the dual ME or PE3 are indicated by arrows.
[0088] Figure 28B shows the results from cells that were either untreated or electroporated with an mDmd exon 23 splicing reporter alone, a single-module ME (magLow), a dual-module ME (magLow+sg), or a reporter having PE2 or PE3. The splicing reporter was used as a positive control. [Figure 29]Figure 29 compares the editing efficiency of single-module ME, dual-module ME with magRNA-gRNA pairing, and dual-module ME with magRNA-magRNA pairing. Except for the last panel experiment in which the second gRNA also contains a second matching tag at its 3' end, the ME systems are similar to those illustrated in Figure 25. Results show cells that were either untreated or electroporated with a mutant EGFP (deletion of one nucleotide at position 50 relative to the nicking site) expression vector (EGFP Δ50) or a mutant EGFP expression vector, plus a single-module ME (magLow) or dual-module ME with magRNA-gRNA pairing (magLow+sgUp119) or dual-module ME with magRNA-magRNA pairing (magLow+magUp119) (all with RNA templates having an elongation length of 111 nt). In MagRNA-magRNA pairing, the matching tag sequences in the two magRNAs are complementary to two separate sequences in the RNA RT template. [Figure 30] Figures 30A, 30B, and 30C show the editing efficiency at the endogenous HEK3 site in HEK293T cells using dual ME with varying ratios of RNA templates, magRNA, and a second gRNA.
[0089] Figure 30A illustrates the target site, HEK3, sequence (5'TGGGGCCCAGACTGAGCACGTGATGGCAGAGGAAAGGAAGCCCTGCTTCCTCCAGAGGGCGTCGCAGGACAGCTTTTCCTAGACAGGGGCTAGTATGTGCAGC, SEQ ID NO: 10) and its complementary sequence (5'GCTGCACATACTAGCCCCTGTCTAGGAAAAGCTGTCCT GCGACGCCCTCTGGAGGAAGCAGGGCTTCCTTTCCTCTGCCATCACGTGCTCAGTCTGGGCCCCA, SEQ ID NO: 11) annotated by dual magRNA-gRNA ME system components, including a magRNA(12M42) guide, a second nicking gRNA guide, and desired point mutations, 5G>T and 12G>C encoded in RNA template T43. +63 indicates the distance between the two nicks. For illustrative purposes, only the 5' and 3' end sequences of the target site are shown.
[0090] Figure 30B shows representative Sanger sequencing chromatographs of untreated (UT) and edited (2mut) HEK3 genomic DNA (SEQ ID NO: 12).
[0091] Figure 30C compares the editing efficiency (5G>T) at the endogenous genomic target HEK3 under varying ratios of RNA template (T43), magRNA, and a second gRNA. [Figure 31] Figures 31A, 31B, and 31C show the insertion editing efficiency at the endogenous HEK3 site in HEK 293T cells using the dual ME system.
[0092] Figure 31A shows the HEK3 sequence and its complementary sequences (SEQ ID NOs: 10 and 11) with annotations for the magRNA(12M42) guide, second nicking guide, nicking site, PAM, and insertion site (+5). RNA templates (T46, T50, and T79) were designed to insert a 3nt (CTT), 7nt AP1 site (ACTCAGT), or 36nt TCR variable region into an endogenous HEK3 genomic site. For illustrative purposes, only the 5' and 3' end sequences of the target site are shown.
[0093] Figure 31B shows representative Sanger sequencing chromatographs of untreated (UT) and edited genomic DNA (SEQ ID NO: 13).
[0094] Figure 31C shows the efficiency of 3nt, 7nt, and 36nt insertions. Even when nCas9-RT fusion was used in this experiment, the second gRNA construct also contained MS2 in the stem-loop region. [Figure 32] Figures 32A, 32B, and 32C show the editing efficiency of the transdual ME system in correcting point mutations in the episomal nfEGFP reporter gene in HEK293 cells when various polymerases, including variants of MMLV-RT, R2Bm (retrotransposon RT), and Helraiser (polymerase for DNA transposon Heliton), were recruited. When Helraiser was used, RNA and DNA templates were examined. The transME system configuration was previously shown in Figure 24A. Figure 32A shows the components of the transME system. The gene to be edited is the nfEGFP target site with the A200G mutation (Figure 32B). Figure 32C shows the editing effectiveness of the transdual ME under various conditions, quantified by the percentage of cells with edited fluorescent EGFP. [Figure 33]Figures 33A and 33B show the effects of recruiting the non-RT effector p65 and the presence of MS2 in the second gRNA on editing efficiency in HEK293T cells. Figure 33A shows the dual magRNA-gRNA ME components. The endogenous target site, HEK3, was previously shown in Figure 30A. Figure 33B compares the editing effectiveness (5G>T, Sanger sequencing) of dual ME systems with and without p65, with and without the second gRNA having MS2. [Figure 34] Figures 34A, 34B, 34C, and 34D show the results of an assay comparing the editing efficacy of the split-dual ME system and its counterpart, the split-prime editing system (PE), at the HEK3 site in HEK293T cells. The target HEK3 sequence was previously shown in Figure 30A.
[0095] Figures 34A to 34C show the components of ME and PE. Figures 34A to 34B show the ME and PE-specific components, respectively. The pegRNA contains the same RNA template sequence as ME, covalently linked to the 3' of the gRNA. Figure 34C shows the common components of the trans-ME and PE systems. Figure 34D shows the editing efficacy of the split ME and split PE under their corresponding optimal conditions. nCas9-RT splitting means that nCas9 and RT are provided as two individual proteins, rather than a fused protein. For split ME, the optimal conditions for producing the best editing efficiency for the RNA template and magRNA were 1000 ng RNA template vector and 250 ng magRNA vector. For trans-PE, the optimal condition for pegRNA producing the best editing efficiency was 1000 ng of the RNA template expression vector molar equivalent. The molar ratio of RNA template in the PE system: pegRNA scaffold in PE: RNA template in ME: magRNA scaffold in ME was roughly equal to 1:1:1:0.25. The conditions for common components, namely MCP-RT (500 ng), sp-nCas9 (500 ng), and second gRNA (250 ng), were the same between the ME and PE systems. [Figure 35] Figures 35–37 show the results of assays comparing the PE3 system and its dual ME counterpart in editing the endogenous HEK3 site by installing five different types of mutations in HEK293T cells. Both editing efficacy and insertion of gRNA scaffold components were tested. Figures 35A, 35B, and 35C show the components of the dual ME system and its PE3 counterpart. Figure 35A shows the ME-specific components, the RNA template and magRNA isolation module. Figure 35B shows the PE-specific component, the pegRNA with gRNA and RT template covalently linked. Figure 35C shows the common components of ME and PE. [Figure 36] Figures 35–37 show the results of assays comparing the PE3 system and its dual ME counterpart in editing the endogenous HEK3 site by installing five different types of mutations in HEK293T cells. Both editing efficacy and insertion of gRNA scaffold components were tested. Figure 36A shows the HEK3 site and its complement sequences (SEQ ID NOs: 10 and 11) with gRNA guide annotations. The RT template, T43-P13-5mut, in both the ME and PE systems contains five mutations across the template. For illustrative purposes, only the 5' and 3' end sequences of the target site are shown.
[0096] Figure 36B shows representative Sanger sequencing chromatographs of untreated (UT) and edited HEK3 genomic DNA (SEQ ID NO: 14). The arrow, 5mut, in the edited sample indicates the positions of five edited bases.
[0097] Figure 36C shows the quantification of editing effectiveness detected by Sanger sequencing at five locations under each editing condition.
[0098] Figure 36D shows the results of an assay comparing the editing efficiency between the PE3 system and its ME counterpart at the 5G>T position. *P<0.05, **P<0.01, n=3. [Figure 37] Figures 35–37 show the results of assays comparing the PE3 system and its dual ME counterpart in editing the endogenous HEK3 site by installing five different types of mutations in HEK293T cells. Both editing efficacy and insertion of gRNA scaffold components were tested. Figures 37A, 37B, and 37C show the results of assays comparing the ME and PE systems in terms of editing efficiency and insertions induced by RT synthesis (reverse transcription) using the gRNA scaffold as a template, reanalyzing representative samples from Figure 36 by next-generation sequencing (NGS). Figure 37A shows the editing efficacy at five locations by ME and PE editing. Figure 37B shows the number of reads containing insertions at the end of template sequences having three or more base pairs identical to the gRNA scaffold 3' sequence. The number of edited reads and the ratio of inserted reads to edited reads are also shown in the table. ME-250, PE3-250, ME-500, and PE3-500 were conditions shown in Figure 37A, with varying amounts of the second gRNA (250 ng vs. 500 ng). Figure 37C shows the ratio of reads containing gRNA scaffold template insertions to edited reads. [Figure 38] Figures 38A and 38B show the results of an assay comparing a dual ME system and its PE3 counterpart in editing the endogenous HEK3 site by deletion or insertion of a 3nt sequence in HEK293T cells.
[0099] Figure 38A shows the HEK3 site and its complement (SEQ ID NOs: 10 and 11), as well as edited sites with intended insertions (CTT) (SEQ ID NOs: 15-16) or deletions (GCA) (SEQ ID NOs: 17-18).
[0100] Figure 38B shows the editing effectiveness of ME and PE3 in insertion and deletion installation, as analyzed by Sanger sequencing. *P<0.05, n=3. [Figure 39]Figure 39 shows the results of an assay comparing the editing efficiency of a dual ME system and its counterpart PE in the installation of single nucleotide insertions in an episomal EGFP plasmid containing 1nt deletions at various positions in K562 cells. Δ4A, Δ20C, and Δ35C are the positions relative to the EGFP nick site in G202 (Figure 17C). The ME components were the same as in Figure 25. [Figure 40] Figures 40A and 40B show the results of an assay comparing the editing efficiency of a dual ME system and its counterpart PE in creating 29nt deletions in a Dmd exon 23 skipping reporter plasmid in K562 cells. The ME and PE components were the same as those used in Figure 28, when the experiment was performed in HEK293T cells. Figure 40A shows editing efficiency quantified by the percentage of GFP-positive cells, while Figure 40B shows editing efficiency quantified by Sanger sequencing. [Figure 41] Figures 41A, 41B, and 41C show the results of assays comparing a dual ME system and its PE3 counterpart in editing endogenous HEK3 sites by deletion or insertion of 3nt sequences in K562 cells. Both editing efficacy and insertion rates caused by RT template extension to the gRNA scaffold were compared. The ME components were the same as those described in Figure 38, where the experiment was performed in HEK293T cells. Figures 41A and 41B show the efficacy of CTT insertion and GCA deletion analyzed by Sanger sequencing and next-generation sequencing (NGS), respectively. Figure 41C shows the ratio of NGS gRNA scaffold insertion reads to edited reads. [Figure 42]Figures 42A and 42B demonstrate the gene editing efficacy of the dual ME system in installing a point mutation in the endogenous HBB (hemoglobin beta) near the sickle cell anemia E6V mutation site in K562 cells. Figure 42A shows the HBB sequence exon 1 and its complement (SEQ ID NOs: 19 and 20) with annotations of the magRNA guide, second nick guide, nicking site, and the intended +5G>T mutation (circled in the first PAM) encoded in the RNA template. The adenine base immediately 5' relative to the +5G base was mutated in sickle cell anemia E6V. Figure 42B demonstrates the efficacy of ME in installing the +5G>T mutation in K562 cells. [Figure 43] Figures 43A, 43B, 43C, and 43D show the results of assays comparing the editing efficacy of the saCas9 dual ME system and its counterpart, the saCas9 PE system. In this study, nickase saCas9(N580A)-RT fusion and saCas9 gRNA scaffolds were used in both systems. Efficiency was compared in the installation of 1nt insertions and silent point mutations in the EGFPΔ100G plasmid in HEK293.
[0101] Figure 43A shows the EGFPΔ100G target sequence (5'tgagcaagggcgaggagctgttcaccggggtggtgcccatcctggtcgagctggacggcgacgaaacggccacaagttcagcgtgtccggcgaggcgagggcgatgccacctacggcaagctgaccctgaagttcatctgcaccaccggcaag3', SEQ ID NO: 21) and its complement (5'cttgccggtggtgcagatgaacttcagggtcagcttgccgtaggtggcatcgccctcgcctcgccggacacacgctgaacttgtggccgtttacgtcgccgtccagctcgaccaggatgggcaccaccccggtgaacagctcctcgcccttgctca, SEQ ID NO: 22), along with the magRNA guide, the second nicking guide, and saCas9. PAM, and the intended editing sequence (5'tgagcaagggcgaggagctgttcaccggggtggtgcccatcctggtcgagctggacggcgacgtaaacggccacaagttcagcgtgtccggcgag G gcgagggcgatgccacctacggcaagctgacc T tgaagttcatctgcaccaccggcaag3', SEQ ID NO: 23) and its complement (5'cttgccggtggtgcagatgaacttca A ggtcagcttgccgtaggtggcatcgccctcgc C The sequence ctcgccggacacgctgaacttgtggccgtttacgtcgccgtccagctcgaccaggatgggcaccaccccggtgaacagctcctcgcccttgctca3' (sequence number 24) is annotated. For illustrative purposes, only the 5' end, deletions or insertions, and 3' end sequences are shown.
[0102] Figure 43B illustrates the conversion from non-fluorescent GFP to fluorescent GFP. Figures 43C-43D compare the editing efficiencies of ME and PE, quantified by GFP conversion (Figure 43C) and Sanger sequencing (Figure 43D, insert). [Modes for carrying out the invention]
[0103] This disclosure relates to a novel RNA-guided gene editing system called a match editing system or ME system, which utilizes reverse transcriptase to copy desired sequence information from an RNA template to DNA and insert this DNA into a target site in genomic or extrachromosomal DNA. The gene editing system includes various novel features, such as (1) an editing RNA template that can be a completely independent RNA molecule and does not require an associated recruiting aptamer sequence such as an MS2 aptamer, (2) a gRNA that can be modified to contain a short polyribonucleotide tag complementary to a region in the template, and as a result, (3) the template can be anchored to the editing configuration via a Watson-Crick base pairing interaction between the editing RNA template and its matching gRNA. No previous design has used modified gRNA to directly recruit an RNA template.
[0104] Designed and configured differently from previously reported systems, the novel system described herein provides an effective platform for convenient gene editing and de novo gene construction at target loci in genomic or extrachromosomal DNA. The novel design enables gene editing at target sequences further away from PAM motifs, can edit all types of point mutations, and excels at insertions and deletions. Furthermore, the modular features of the design allow for cooperative multiplexing, enhancing functionality. Thus, the novel system and associated platform add a unique and versatile tool to gene editing technology.
[0105] As disclosed herein, a feature of the ME system invention is the ability to manipulate gRNA for direct RNA template recruitment by enabling Watson-Crick sequence pairing with a native template sequence. Firstly, the prior art did not teach the manipulation of gRNA sequences for the direct recruitment of other gene editing components (e.g., RNA templates) by Watson-Crick sequence complementarity. Secondly, the prior art in reverse transcriptase-mediated gene editing, including prime editing, did not teach the manipulation of gRNA sequences for the direct recruitment of RNA templates by Watson-Crick sequence complementarity. The match editing feature was neither a general RNA-guided gene editing operation nor reverse transcriptase-mediated gene editing, such as prime editing, and was more unexpected and non-obvious than any prior art.
[0106] Furthermore, the separate tethering mechanism between the manipulated gRNA, named magRNA in this disclosure, and the RNA template enables numerous novel features not found in any other reverse transcriptase-mediated gene editing system.
[0107] Firstly, this recruitment mechanism involves base pairing between a native template sequence and its complementary sequence attached to magRNA. Exogenous or accessory sequences, such as the MS2 aptamer, are not required in the RNA template. On the other hand, transactive elements, such as MCP fusions, are not required for template recruitment. This design ensures effective gene editing; at the same time, the design is streamlined and compact compared to other similar reverse transcriptase-mediated genome editing systems, which is an advantage for effective delivery.
[0108] Secondly, modular and versatile tethering mechanisms simplify the manipulation of complex systems with multiplexed modules that work synergistically to enhance effectiveness or add novel functions. For example, two modules, each containing magRNAs with tags that match different sequences within the same RNA template, can cooperatively recruit the RNA template and thus dramatically increase editing effectiveness. In addition to the cooperative recruitment of the RNA template by two separate modules, each module can also provide distinct effectors that work together synergistically. For example, one module may contain an RNA aptamer for recruiting a reverse transcriptase fused to an aptamer-binding protein, and the other module may contain another RNA aptamer for recruiting a chromatin-modifying enzyme fused to a corresponding aptamer-binding protein. The physical and functional interactions between the two modules are re-enforced by interactions to the same RNA template recruited by the two magRNAs.
[0109] Thirdly, the absence of requirements for exogenous sequences in RNA templates mechanistically eliminates all sources of “scars” currently affecting reverse transcriptase-mediated gene editing. One of the most challenging problems in sequence-specific introduction of genes or genetic fragments into the genome is the simultaneous introduction of unwanted exogenous sequences, i.e., “scars.” To date, no effective “scar-free” solution has been available when various enzymes, such as reverse transcriptase, recombinase, and integrase, are used to generate transgene insertions in non-germline mammalian cells. The underlying cause of “scars” is that the template, whether an RNA or DNA molecule, needs to contain not only the genetic information to be inserted into the target site, but also handles for specific interactions and recruitment between the template molecule and the mechanism of genetic manipulation. Examples of such accessory sequences for interaction with the genetic manipulation system are flanking LTRs, ITRs, and RNA aptamers. Such accessory sequences are sources of genetic “scars.” In particular in relation to this disclosure, it has been well reported that prime editing factors often result in the insertion of undesired pegRNA components, such as part or all of the gRNA CRISPR interaction scaffold, into the target editing site, in addition to the desired RNA template within the pegRNA (Peter J. Chen et al, Enhanced prime editing systems by manipulating cellular determinants of editing outcomes. Cell 184, 5635-5652, October 28, 2021., e.g., Figure 2 in the said document).
[0110] Fourth, the system design separates gRNA from the RNA template, while simultaneously installing different template recruitment functions in the gRNA via a non-covalent Watson-Crick base pairing mechanism. The combination allows for the use of long RNA templates for effective gene editing without the inherent problem that long RNA templates can impair the secondary structure of gRNA when the template and gRNA are covalently linked.
[0111] Fifth, structural separation of the RNA template from the gRNA scaffold allows for a variable ratio between the RNA template and gRNA / nCas9 / reverse transcriptase for optimal editing outcomes. In other words, in pegRNA design, the ratio between the RNA template and gRNA molecules is always 1:1. As a result, at the editing site, one nCas9-RT molecule associates with one pegRNA (one gRNA and a covalently linked template). In contrast, ME allows for variation of the RNA template to gRNA ratio to increase editing effectiveness and reduce nonspecific editing. For example, instead of only achieving a 1:1 ratio with pegRNA, a 1:0.25 ratio can be performed with ME as desired.
[0112] As a result, by employing a unique magRNA design for directly tethering and recruiting separate “scarless” modular RNA template molecules, the ME system disclosed herein overcomes numerous significant existing obstacles in the field of gene editing (e.g., inherent “scarring” tendencies) while increasing flexibility and versatility to create optimal conditions for genome editing (e.g., high efficacy and low to negligible off-target effects) (e.g., by allowing changes in the ratio between RNA template and gRNA and reverse transcriptase). Thus, ME is a unique system that possesses distinctive features and improvements in both functionality and elegance in its design compared to existing systems.
[0113] Match Editing (ME) System As described herein, the ME system includes numerous distinctive features. For example, the desired genetic sequence can be planned by an independent RNA template molecule that is separate from the gRNA and does not contain any protein recruitment motifs, such as MS2 aptamers, or any other sequences not intended to be present in the final product. In other words, ME gives the freedom to design a mechanistically scar-free template. Furthermore, the separate matching gRNA molecule contains a short RNA sequence or matching tag that is complementary to the region in the template RNA molecule. In addition, template recruitment is mediated by Watson-Crick pair formation between the template RNA and the matching sequence in the gRNA. The RNA templates described herein are often referred to as modular RNA templates because they are separate from and independent of the sequence-specific DNA-binding module. A gRNA with a matching tag complementary to the region in the template is called a magRNA.
[0114] The separation of DNA sequence recognition modules (gRNAs) from the genetic sequence information (RNA templates) to be incorporated into the genome makes the system highly modular. As a result, longer templates can be designed independently without concerns about interfering with the secondary structure and function of the gRNAs. Furthermore, the Watson-Crick base pairing tethering mechanism between RNA molecules eliminates unnecessary components, such as recruiting RNA aptamer sequences in the template and aptamer-binding fusion sites in CRISPR proteins, in contrast to affinity binding between RNA motifs (e.g., MS2 RNA aptamers) and proteins (e.g., aptamer-binding proteins MCPs). This not only makes the system compact but also eliminates sources of editing "scars." More importantly, the modular design of the system and the simple base pairing recruitment mechanism allows for the construction of complex systems by simply and synergistically multiplexing individual modules linked ("clicked") by nucleotide sequence complementarity to various regions of the same RNA template.
[0115] In some embodiments, the disclosure provides examples of dual-module ME configurations achieved by "clicking" two single-module MEs, along with various basic single-module ME configurations for achieving versatile gene editing and gene de novo composition functions.
[0116] As an example, Figure 1 shows a single-module ME basic configuration having a reverse transcriptase fused with an RNA-guided nickase using the CRISPR / Cas9 protein. This single-module match editing system uses a CRISPR / gRNA complex for RNA-guided sequence recognition. The system includes (a) a modular RNA template that does not contain any RNA recruiting elements, e.g., RNA aptamers (e.g., MS2 or PP7), (b) a magRNA containing a 5' spacer sequence for target DNA recognition, an RNA scaffold for CRISPR / Cas complex formation, and a 3' match RNA sequence complementary to the region in the modular template, and (c) a nickase CRISPR protein, nCRISPR / nCas9 (H840A), fused with a reverse transcriptase (MMLV-RT) variant via a polypeptide linker. As illustrated in Figure 1, the single-module basic system is: (1) An RNA template for reverse transcription (or a nucleotide encoding template expression) having a 3' end sequence complementary to the nickel DNA strand at the target site, and containing a nucleotide sequence to be reverse transcribed to replace a desired sequence at the target site; (2) a. Spacer sequences for sequence-specific recognition at target sites; b. RNA scaffolds for complex formation with RNA guide nickase proteins; and c. RNA sequence tags (6-24 nucleotides in size) complementary to the sequence in the RNA template (1) Manipulated magRNA containing; and (3) RNA guide nickas fused with reverse transcriptase via a polypeptide linker (e.g., nCRISPR / nCas9 H840A) It can include...
[0117] As illustrated in Figure 1, the matching sequence is located at 3' of the gRNA, but the tag can be localized to other locations within the gRNA, such as the stem-loop region, tetra-loop region, or 5'.
[0118] Figure 2 illustrates how the match editing system works. As illustrated in Figure 2, the match editing system recognizes a DNA target site via a spacer sequence at the 5' end of the manipulated gRNA and forms an R-loop at the target DNA locus. The tag sequence at the 3' end of the magRNA binds to a segment in the RNA template complementary to the tag, tethering the template to that site. Meanwhile, nickase activity causes a single-strand DNA break in the non-target strand (e.g., 3 nucleotides upstream of the PAM motif if nCas9 H840A nickase is used). Next, the 3' end of the RNA template anneals with the nicked complementary non-target DNA strand in the R-loop. The nicked non-target DNA end then uses the RNA template to function as a primer for reverse transcriptase. As a result, the sequence information in the modular RNA template is copied to the DNA strand, which is then incorporated into the target DNA molecule.
[0119] Figure 3 illustrates the cellular repair mechanism that results in the incorporation of newly synthesized DNA strands into the target DNA locus. More specifically, Figure 3 illustrates the process by which match editing results in precise genome editing. The process includes: A. Nicking is introduced into the non-target chain within the target site by nickase; B. Reverse transcriptase copies the RNA template sequence to the DNA sequence primed by the nickel DNA strand; C. The newly synthesized 3'-DNA flap competes with the 5'-DNA flap; The D.5'-flap is preferentially removed in the cell by a nuclease, and the newly synthesized DNA strand anneals to the opposite strand; E. DNA repair either removes mismatched sequences in the "old" strand, resulting in editing, or removes mismatched sequences in the newly synthesized strand, without resulting in editing events.
[0120] The newly synthesized strand can contain desired point mutations (either transversions, transitions, or both), insertions, or deletions.
[0121] A specific example of magRNA-RNA template interaction is shown in Figure 4. In this example, the 3'-matching tag sequence of the magRNA is specifically complementary to the 5' end sequence of the RNA template.
[0122] Another special example with a magRNA having a 5' match tag is shown in Figure 5. In this example, the tag, complementary to the RNA template region, is located at the 5' end of the gRNA. The linker polynucleotide sequence is positioned between the 5' match tag sequence and the spacer sequence.
[0123] Figure 6 shows an exemplary variant of a match editing system in which the reverse transcriptase is provided in a split state rather than being directly fused to an RNA guide sequence-specific nickase. As illustrated in Figure 6, the split-state RT system includes: (1) An RNA template for reverse transcription containing a polynucleotide sequence to be reverse transcribed to replace a desired sequence at a target site, and a 3' end sequence complementary to the nickel strand at the target site; (2) RNA guide nickases for sequence recognition; (3) a. 5' guide for sequence-specific recognition at the target site b. RNA scaffold for complex formation with RNA guide nickase protein c. RNA aptamer sequences in RNA scaffold loops such as stem-loops or tetra-loops d. A 3' elongated sequence (6-24 nucleotides in size) complementary to the sequence segment located at the 5' end or center of the RNA template (1). Manipulated magRNA containing; and (4) A protein containing a reverse transcriptase that is fused via a polypeptide linker to a cognitive protein (e.g., MCP) that binds to an aptamer (e.g., MS2) within a gRNA scaffold.
[0124] Figure 7 illustrates how the split-state RT system works. As illustrated in Figure 7, the split-state RT system works in a manner very similar to the schematic diagram illustrated in Figure 2, except that the reverse transcriptase (RT) is not fused to an RNA guide sequence-specific complex protein, such as nCas9. Instead, the gRNA contains an MS2 aptamer in the stem-loop (or other configuration) of the gRNA that recruits the MCP-RT fusion protein.
[0125] Figure 8 shows a variant configuration of the split-state RT system in which the 3' magRNA elongation sequence is complementary to the 5' end sequence of the RT template.
[0126] Figure 9 shows two dual ME systems, each having a nicking module that creates a single-strand DNA break in a near downstream position on the opposite strand. A dual ME system with fusion RT and magRNA-gRNA pairing is shown in Figure 9A. The second module is for creating a second nick to increase editing efficiency. A dual ME system with split RT and magRNA with MS2-gRNA 0xMS2 is shown in Figure 9B. The second module is for creating a second nick to increase editing efficiency. Each dual ME system increases editing efficiency by the mechanism illustrated in Figure 10.
[0127] More specifically, as shown in Figure 10, mechanistically, the second nicking prefers to remove mismatched sequences in the nicking strand, resulting in higher editing efficiency. Consequently, the second nicking on the nearby opposite strand, mediated by dual ME, significantly enhances the target editing efficiency.
[0128] A dual ME system coupled to a single template for editing a target site using a long RNA template is disclosed herein. Two examples, a coupled dual ME with fused RT (Figure 11A) and a coupled dual ME with split state RT (Figure 11B), are illustrated. Each module of the system contains magRNAs with complementary match tags in separate regions of the RNA template. The tethering of the RNA template by two magRNAs enhances the interaction between the editing complex and the template, thereby increasing editing efficiency or enhancing the functionality of the individual modules.
[0129] Figure 11 also illustrates the principle of how two or more individual ME modules are physically and functionally linked or "clicked" together to enhance activity or create new synergistic functionality through simple base-pairing complementarity to different regions of the same RNA template.
[0130] Coupled dual MEs for enhanced functionality are also disclosed herein. Several examples (RT fusion version and split-state RT version) are illustrated in Figure 12, and their mechanisms are described in Figure 13. In these examples, the individual modules are designed to function cooperatively. For example, in the RT fusion version, the 3' end of the RNA template is designed to be identical to the 3' end of the second nickeling strand. As a result, the second nickeling strand functions as a primer for synthesizing the second DNA strand using the newly synthesized DNA strand as a template.
[0131] Figure 13 illustrates the scheme of cooperative interaction and second strand DNA synthesis involving a long RNA template by a dual ME system coupled to a single template: (A)-(B) First nicking and reverse transcription of the long template by the first module; (C)-(D) Second nicking by the second module; (E)-(F) Second DNA strand synthesis mediated by DNA-dependent DNA polymerase activity of the reverse transcriptase, facilitated by complementarity between the 3' end of the second nicking site and the 5' end of the RNA template; (G)-(H) Removal of the 5' flap.
[0132] Examples of systems that can be used for insertion or deletion are also disclosed herein. As illustrated in Figure 14, Module 1 and Module 2 are separate ME modules, each having the separate components described above, including separate RNA templates. In this configuration, the two modules may be designed to cooperate functionally, but not physically interact with each other via a common RNA template. For example, the RNA templates may be designed such that the reverse-transcribed DNA strands are partially complementary to each other.
[0133] The two illustrated modules create a trans nick on the opposite strand. Two modules, or multiplex modules, can be functional in both cis (by PAM and guide sequences on the same strand) or trans configurations to achieve various functionalities, such as inserting small fragments in tandem to achieve insertion of a larger fragment, or correcting multiple mutations for various deletions.
[0134] Figure 15 is a schematic diagram illustrating how a dual-template dual-module ME system can produce de novo gene composition (insertion) or large fragment gene deletion. As illustrated in Figures 15A and 15B, module 1 ME and module 2 ME mediate nicking at their corresponding target sites, resulting in first-strand DNA synthesis by reverse transcription of their corresponding RNA templates. This then leads to annealing of the 3' ends of the two first DNA strands by complementary sequences, as shown in Figure 15C. This further enables second-strand DNA synthesis by DNA-dependent DNA polymerase activity or endogenous DNA polymerase activity of the reverse transcriptase, as shown in Figure 15D. After removal of the 5' flap and ligation (Figure 15E), insertion or deletion is achieved (Figure 15F). This mechanism can produce both insertions and deletions (de novo gene composition), depending on the content and length of the two RNA templates and the sequence to be removed.
[0135] Various second effectors may be included in any of the systems described herein. As shown in Figure 16, the second effectors may be provided in a split state or by fusion. Examples include proteins that facilitate the editing process or increase editing efficiency, such as chromatin modifying enzymes, mismatch repair enzyme inhibitors, and proteins that can stabilize template RNA and magRNA. Examples of effectors include human RNase inhibitor proteins, dominant-negative proteins of MMR pathway proteins including MLH1 protein, 5' DNA nuclease Fen1 protein, or others that can enhance editing efficiency. The modules illustrated in Figure 16 are intended to be “clicked” or multiplexed into dual or complex systems illustrated in Figures 9–15. RNA guide niccas A crucial second component of the gene editing complex or system disclosed herein is an RNA-guided nickase. A nickase is an enzyme that creates a single-strand break (also known as a "nick") in double-stranded DNA, that is, it cuts one strand of the DNA double helix but not the other.
[0136] As used herein, “RNA-guided nickase” means a polypeptide or polypeptide complex having DNA nickase activity, where DNA nickase activity may be sequence-specific and dependent on the RNA sequence. Exemplary RNA-guided nickases include Cas nickases. Cas nickases include, but are not limited to, the Csm or Cmr complex of the type III CRISPR system, CaslO, Csml or its Cmr2 subunit, the cascade complex of the type I CRISPR system, its Cas3 subunit, and the nickase forms of class 2 Cas nucleases. Class 2 Cas nickases include class 2 Cas nuclease variants having RNA-guided DNA nickase activity in which only one of the two catalytic domains is inactivated. Class 2 Cas nickases include, for example, Cas9 (e.g., SpyCas9 H840A, D10A, or N863A variants), Cpfl, C2cl, C2c2, C2c3, HF Cas9 (e.g., N497A, R661A, Q695A, Q926A variants), HypaCas9 (e.g., N692A, M694A, Q695A, H698A variants), eSPCas9(1.0) (e.g., K810A, K1003A, R1060A variants) and eSPCas9(II) (e.g., K848A, K1003A, R1060A variants) proteins, as well as their modifications. The Cpfl protein, Zetsche et al., Cell, 163: 1-13 (2015), is homologous to Cas9 and contains a RuvC-like protein domain. Zetsche's Cpfl sequence is incorporated herein by reference in its entirety. See, for example, Zetsche, Tables SI and S3. "Cas9" includes 5.pyogenes(Spy)Cas9, the Cas9 variants listed herein, and their equivalents.For example, see Makarova et al., Nat Rev Microbiol, 13(11): 722-36 (2015); Shmakov et al., Molecular Cell, 60:385-397 (2015); and Makarova et al., NAT. REV. MICROBIOL, 18:67-83 (2020).
[0137] In some embodiments, the RNA guide nickase disclosed herein is a Cas nickase. In some embodiments, the RNA guide nickase is derived from a specific Cas nuclease in which its catalytic domain(s) are inactivated. In some embodiments, the RNA guide nickase is a class 2 Cas nickase, e.g., type II Cas9 nickase or Cpfl nickase. In some embodiments, the RNA guide nickase is S. pyogenes Cas9 nickase. In some embodiments, the RNA guide nickase is Neisseria meningitidis Cas9 nickase. In some embodiments, the RNA guide nickase is Staphylococcus aureus Cas9 nickase.
[0138] In some embodiments, class 2 Cas is a type V Cas protein, such as Cas12. In some embodiments, Cas12 is Cpf1 (Cas12a), C2c1 (Cas12b), C2c3 (Cas12c), CasY (Cas12d), CasX (Cas12e), Cas14 (Cas12f), CasPhi (Cas12j), and their orthologues and variants.
[0139] In some embodiments, the Cas protein is a type VI Cas protein, such as Cas13. In some embodiments, Cas13 is Cas13a, Cas13b, Cas13c, Cas13d, Cas13x, and their orthologues and variants.
[0140] In some embodiments, the RNA guide nickases disclosed herein are variants of transposon-encoded IscB, IsrB, or TnpB family endonuclease proteins.
[0141] In some embodiments, the RNA guide nickase is a modified class 2 Cas protein or derived from a class 2 Cas protein. In some embodiments, the RNA guide nickase is modified from or derived from a Cas protein, e.g., a class 2 Cas nuclease (e.g., a type II, V, or VI Cas nuclease). Class 2 Cas nucleases include, for example, Cas9, Cpfl, C2cl, C2c2, and C2c3 proteins, and their modifications. Examples of Cas9 nucleases include the Cas9 nucleases of the type II CRISPR system in S. pyogenes, S. aureus, and other prokaryotes (see, for example, the list in the following paragraph), and their modified (e.g., engineered or mutant) versions. See, for example, US2016 / 0312198Al and US2016 / 0312199Al, which are incorporated herein by reference in their entirety. Other examples of Cas nucleases include the Csm or Cmr complex or Cas10, Csml, or their Cmr2 subunit in the type III CRISPR system; and the cascade complex or its Cas3 subunit in the type I CRISPR system. In some embodiments, the Cas nuclease may be derived from the type IIA, type IIB, or type IIC system. For descriptions of various CRISPR systems and Cas nucleases, see, for example, Makarova et al., NAT. REV. MICROBIOL. 9:467-477 (2011); Makarova et al., NAT. REV. MICROBIOL, 13: 722-36 (2015); Shmakov et al., MOLECULAR CELL, 60:385-397 (2015); and Makarova et al., NAT. REV. MICROBIOL, 18:67-83 (2020).
[0142] The Casnickases described herein are Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus species, Staphylococcus aureus, Listeria innocua, Lactobacillus gasseri, Francisella novicida, Wolinella succinogenes, Sutterella wadsworthensis, Gammaproteobacterium, Neisseria meningitidis, Campylobacter jejuni, Pasteurella multocida, Fibrobacter succinogene, Rhodospirillum rubrum, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Lactobacillus buchneri, Treponema denticola, Microscilla marina, Burkholderiales bacteria, Polaromonas naphthalenivorans, Polaromonas species, Crocosphaera watsonii, Cyanothece species, Microcystis aeruginosa, Synechococcus species, Acetohalobium arabaticum, Ammonifex degensii, Caldicelluliruptor becscii, Candidates Desulforudis, Clostridiumbotulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter species, Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Streptococcus This may be a nickas form of Cas nuclease derived from species including, but not limited to, pasteurianus, Neisseria cinerea, Campylobacter lari, Parvibaculum lavamentivorans, Corynebacterium diphtheria, Acidaminococcus species, Lachnospiraceae bacteria ND2006, or Acaryochloris marina.
[0143] In some embodiments, Cas nickase is the nickase form of Cas9 nuclease derived from Streptococcus pyogenes. In some embodiments, Cas nickase is the nickase form of Cas9 nuclease derived from Streptococcus thermophilus. In some embodiments, Cas nickase is the nickase form of Cas9 nuclease derived from Neisseria meningitidis. See, for example, WO / 2020081568, which describes Nme2Cas9 D16A nickase. In some embodiments, Cas nickase is the nickase form of Cas9 nuclease derived from Staphylococcus aureus. In some embodiments, Cas nickase is the nickase form of Cpfl nuclease derived from Francisella novicida. In some embodiments, Cas nickase is the nickase form of Cpfl nuclease derived from species of the genus Acidaminococcus. In some embodiments, Cas nickase is the nickase form of Cpfl nuclease derived from the Lachnospiraceae bacterium ND2006. In further embodiments, Cas nickase is the nickase form of Cpfl nuclease derived from Francisella tularensis, Lachnospiraceae bacteria, Butyrivibrio proteoclasticus, Peregrinibacteria bacteria, Parcubacteria bacteria, Smithella, Acidaminococcus, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi, Leptospira inadai, Porphyromonas crevioricanis, Prevotella disiens, or Porphyromonas macacae. In certain embodiments, Cas nickase is the nickase form of Cpfl nuclease derived from Act daminococcus or Lachnospiraceae.As described elsewhere, nickase can be derived from (i.e., related to) a specific Cas nuclease, in that it is a form of nuclease in which one of its two catalytic domains is inactivated by mutation of an active site residue essential for nucleolysis, such as D10, H840, or N863 in Spy Cas9. Those skilled in the art will be familiar with techniques for easily identifying corresponding residues in other Cas proteins, such as sequence alignment and structural alignment, which are described in detail below.
[0144] In other embodiments, Cas nickase may be involved in the type I CRISPR / Cas system. In some embodiments, Cas nickase may be a component of the cascade complex of the type I CRISPR / Cas system. In some embodiments, Cas nickase may be a Cas3 protein. In some embodiments, Cas nickase may originate from the type II CRISPR / Cas system.
[0145] In some embodiments, Cas nickase is a nickase form of a modified Cas nuclease in which the Cas nuclease or endonuclease-type nucleolytic active site is inactivated by one or more modifications (e.g., point mutations) in the catalytic domain. For a description of Cas nickase and exemplary catalytic domain modifications, see, for example, U.S. Patent No. 8,889,356.
[0146] Wild-type S. pyogenes Cas9 has two catalytic domains: RuvC and HNH. The RuvC domain cleaves non-target DNA strands, and the HNH domain cleaves target DNA strands. In some embodiments, the Cas nuclease may contain amino acid substitutions in the RuvC or RuvC-like nuclease domain. An exemplary amino acid substitution in the RuvC or RuvC-like nuclease domain is D10A (based on the S. pyogenes Cas9 protein). See, for example, Zetsche et al. (2015) Cell Oct 22:163(3): 759-771. In some embodiments, the Cas nuclease may contain amino acid substitutions in the HNH or HNH-like nuclease domain. Exemplary amino acid substitutions in the HNH or HNH-like nuclease domain are E762A, H840A, N863A, H983A, and D986A (based on the S. pyogenes Cas9 protein). See, for example, Zetsche et al. (2015). Further exemplary amino acid substitutions include D917A, E1006A, and D1255A (based on the Francisella novicida U112 Cpfl (FnCpfl) sequence (UniProtKB - A0Q7Q2 (CPF1 FRATN))).
[0147] In some embodiments, Cas nickase, such as Cas9 nickase, has an inactivated RuvC or HNH domain. In some embodiments, a nickase having a RuvC domain with reduced activity is used. In some embodiments, a nickase having an inactive RuvC domain is used. In some embodiments, a nickase having an HNH domain with reduced activity is used. In some embodiments, a nickase having an inactive HNH domain is used. In some embodiments, Cas9 nickase has an active HNH nuclease domain and can cleave the untargeted strand of DNA, i.e., the gRNA-bound strand, and has an inactive RuvC nuclease domain and cannot cleave the targeted strand of DNA, i.e., the strand to which base editing by deaminase is desired.
[0148] An example Cas9 nickase amino acid sequence is shown below:
[0149] [ka]
[0150] Other examples include the nickase versions of the following Cas variants: SpCas9 D1135E variant, SpCas9 VRER variant, SpCas9 EQR variant, SpCas9 VQR variant, SpCas9-NG, xCas9, SpCas9-NG; Staphylococcus aureus (SA); SaCas9, Acidaminococcus species (AsCpf1) and Lachnospiraceae bacteria (LbCpf1); AsCpf1 RR variant, LbCpf1 RR variant, AsCpf1 RVR variant, Campylobacter jejuni (CJ)Cas9, Neisseria meningitidis (NM)Cas9, Streptococcus thermophilus (ST)Cas9, and Treponema denticola (TD)Cas9.
[0151] In some embodiments, the RNA guide nickase includes an amino acid sequence that is at least 80%, 90%, 95%, 98%, or 99% identical to the sequence described above.
[0152] reverse transcriptase The gene editing complexes or systems disclosed herein include reverse transcriptases. As used herein, “reverse transcriptase” refers to a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require primers to synthesize a DNA transcript from an RNA template. Historically, reverse transcriptases have been used primarily to transcribe mRNA into cDNA, which can then be cloned into a vector for further manipulation.
[0153] The reverse transcriptase of avian myeloidosis virus (AMV) was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473:1 (1977)). This enzyme possesses 5'-3' RNA-directed DNA polymerase activity, 5'-3' DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a processive 5' and 3' ribonuclease specific to the RNA strand for RNA-DNA hybrids (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Detailed studies of AMV reverse transcriptase activity and its associated RNase H activity were published by Berger et al., Biochemistry 22:2365-2372 (1983). Another reverse transcriptase widely used in molecular biology is the reverse transcriptase derived from Moloney's mouse leukemia virus (M-MLV). See, for example, Gerard, GR, DNA 5:271-279 (1986) and Kotewicz, ML, et al., Gene 35:249-258 (1985). M-MLV reverse transcriptases substantially lacking RNase H activity have also been described. See, for example, U.S. Patent No. 5,244,797. This disclosure considers the use of any such reverse transcriptase or its variants or mutants.
[0154] Reverse transcriptases can also originate from non-viral sources, including retrotransposons. Examples include LTR retrotransposons, endogenous retroviruses, non-LTR retrotransposons, and reverse transcriptases derived from LINE (Line 1) molecules. Examples of such proteins include Line 1 ORF2, R2Bm, and R2Ol.
[0155] This disclosure considers any wild-type reverse transcriptase obtained from any naturally occurring organism or virus, or from any commercial or non-commercial source. In addition, the reverse transcriptases usable in the gene editing complexes or systems of this disclosure may include any naturally occurring mutant RT, engineered mutant RT, or other variant RT, including functionally retained truncated variants. The RT may be engineered to include specific amino acid substitutions, such as those specifically disclosed herein.
[0156] Reverse transcriptase is a multifunctional enzyme that typically possesses three enzymatic activities, including RNA and DNA-dependent DNA polymerization activity, as well as RNaseH activity, which catalyzes the cleavage of RNA in RNA-DNA hybrids. Some mutants of reverse transcriptase have disabled the RNaseH moiety to prevent unintended damage to RNA. Such enzymes, which synthesize complementary DNA (cDNA) using mRNA as a template, were first identified in RNA viruses. Subsequently, reverse transcriptase was isolated and purified directly from viral particles, cells, or tissues (e.g., each incorporated herein by reference: Kacian et al., 1971, Biochim. Biophys. Acta 46: 365-83; Yang et al., 1972, Biochem. Biophys. Res. Comm. 47: 505-11; Gerard et al., 1975, J. Virol. 15: 785-97; Liu et al., 1977, Arch. Virol. 55187-200; Kato et al., 1984, J. Virol. Methods 9: 325-39; Luke et al., 1990, Biochem. 29: 1764-69 and Le Grice et al., 1991, J. Virol. 65: (See 7004-07). More recently, mutants and fusion proteins have been created in the search for improved properties such as thermal stability, fidelity, and activity. Any wild-type, variant, and / or mutant form of reverse transcriptase known in the art or that can be produced using methods known in the art is considered herein.
[0157] Examples of sources of reverse transcriptase include, but are not limited to, Moloney mouse leukemia virus (M-MLV or MLVRT); human T-cell leukemia virus type 1 (HTLV-1); bovine leukemia virus (BLV); Rous sarcoma virus (RSV); human immunodeficiency virus (HIV); yeast including Saccharomyces and Neurospora; Drosophila; primates; and rodents. For example, Weiss, et al., U.S. Patent No. 4,663,290 (1987); Gerard, GR, DNA:271-79 (1986); Kotewicz, ML, et al., Gene 35:249-58 (1985); Tanese, N., et al., Proc. Natl. Acad. Sci. (USA):4944-48 (1985); Roth, MJ, et al., J. Biol. Chem. 260:9326-35 (1985); Michel, F., et al., Nature 316:641-43 (1985); Akins, RA, et al., Cell 47:505-16 (1986), EMBO J.4:1267-75 (1985); and Fawcett, DF, Cell See 47:1007-15 (1986) (each of these in its entirety is incorporated herein by reference).
[0158] Exemplary enzymes include, but are not limited to, M-MLV reverse transcriptase and RSV reverse transcriptase. Enzymes with reverse transcriptase activity are commercially available. In certain embodiments, the reverse transcriptase is supplied trans to other components of the ME system; that is, the reverse transcriptase is expressed as an individual component or otherwise supplied.
[0159] Those skilled in the art will recognize that reverse transcriptases including, but not limited to, Rous sarcoma virus (RSV) reverse transcriptase, avian myeloblastosis virus (AMV) reverse transcriptase, avian erythroblastosis virus (AEV) helper virus MCAV reverse transcriptase, avian myelocytomatosis virus MC29 helper virus MCAV reverse transcriptase, avian reticuloendotheliopathy virus (REV-T) helper virus REV-A reverse transcriptase, avian sarcoma virus UR2 helper virus UR2AV reverse transcriptase, avian sarcoma virus Y73 helper virus YAV reverse transcriptase, Rous-associated virus (RAV) reverse transcriptase, and myeloblastosis-associated virus (MAV) reverse transcriptase, as well as wild-type reverse transcriptases including, but not limited to, Moloney's mouse leukemia virus (M-MLV); human immunodeficiency virus (HIV) reverse transcriptase and avian sarcoma-leukemia virus (ASLV) reverse transcriptase, can be suitably used in the methods and compositions described herein. Additional examples are provided in Martin-Alonso, S. et al., Trends in Biotechnology 39 (2): 194-210 (2021) and WO2020191248, the contents of which are incorporated herein.
[0160] [ka]
[0161] Nuclear localization signals In some embodiments, ME component proteins contain one or more nuclear localization signal (NLS) peptides. NLS peptides can localize to the N-terminus, C-terminus, or internal region of a protein. NLS peptides are generally short peptides that act as signal fragments mediating protein transport from the cytoplasm to the nucleus. A list of commonly used NLS peptides can be found in Lu, J., Wu, T., Zhang, B. et al. Cell Commun Signal 19, 60 (2021). A complete list of NLSs has been manually selected in the database genome.unmc.edu / LocSigDB / .
[0162] In some embodiments, the NLS is a classical afractional or bifractional nuclear localization signal (cNLS). Afractional NLS consists of 4 to 8 basic amino acids and generally contains 4 or more positively charged arginine (R) and lysine (K) residues. Bifractional NLS contains two clusters of basic amino acids separated by spacers of about 10 amino acids. In some embodiments, the bifractional NLS is the nucleoplasmin NLS peptide KR[PAATKKAGQA]KKKK (SEQ ID NO: 28).
[0163] In some embodiments, the non-fractionated NLS is the SV40 large T antigen NLS peptide Pro-Lys-Lys-Arg-Lys-Val (PKKKRKV, SEQ ID NO: 29). In some embodiments, the SV40 NLS is located at the N-terminus of the ME component protein. In some embodiments, the SV40 NLS is located at the C-terminus of the ME component protein.
[0164] Linker In some embodiments, the reverse transcriptase and the RNA guide nickase described herein are conjugated via linkers, including but not limited to chemical modifications, peptide linkers, chemical linkers, covalent or non-covalent bonds, or protein fusion, or by any means known to those skilled in the art. The conjugation may be permanent or reversible. See, for example, U.S. Patents 4,625014, 5057301, and 5514363, U.S. Patent Applications 20150182596 and 20100063258, and WO2012142515, whose entire contents are incorporated herein by reference. In some embodiments, several linkers may be included to take advantage of the desired properties of each linker and each protein domain in the conjugate. For example, a mobile linker and a linker that increases the solubility of the conjugate may be considered for use alone or in combination with other linkers. Peptide linkers can be linked to one or more protein domains in a conjugate by expressing DNA that encodes the linker. The linkers may be acid-cleavable, photocleavable, and thermosensitive. Methods for conjugation are well known to those skilled in the art and are incorporated for use in this disclosure.
[0165] In some embodiments, the linker may be an organic molecule, polymer, or chemical part. In some embodiments, the linker may be a peptide linker. In some embodiments, the peptide linker may be a stretch of any amino acid having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, or at least 50 or more amino acids.
[0166] In some embodiments, the peptide linker is a 16-residue "XTEN" linker or a variant thereof (see, for example, Schellenberger et al. A recombinant polypeptide extends the in vivo half-life of peptides and proteins in a tunable manner. Nat. Biotechnol. 27, 1186-1190 (2009)).
[0167] In some embodiments, the peptide linker is (GGGGS)n (SEQ ID NO: 30), (G)n, (EAAAK) n (SEQ ID NO: 31), (GGS)n, SGSETPGTSESATPES motif (SEQ ID NO: 32) (see, for example, Guilinger JP, Thompson DB, Liu D R. Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014; 32(6): 577-82, the entire content of which is incorporated herein by reference) or (XP) n It includes motifs, or any combination thereof (where n is an independent integer between 1 and 30). See WO2015089406, the entire content of which is incorporated herein by reference.
[0168] Matching gRNA (magRNA) molecule Another essential component of the gene editing complex or system disclosed herein is an RNA molecule called matching gRNA or magRNA. This magRNA may contain three subcomponents: (1) a programmable guide RNA sequence or spacer sequence for target DNA recognition, (2) an RNA scaffold that can bind to an RNA guide nickase, and (3) a tethering tag sequence complementary to a segment in the RNA template. This magRNA may be either a single RNA molecule or a complex of multiple RNA molecules.
[0169] As disclosed in some embodiments, magRNA and RNA guide nickase (e.g., Cas protein) together form a complex (e.g., a CRISPR / Cas-based module) for sequence targeting and recognition. The tag sequence carries a template to the target site to be edited, while the reverse transcriptase may be recruited to the site by different means. In some examples, the reverse transcriptase may be recruited directly by fusion to the nickase. In other examples, the reverse transcriptase may be recruited (individually) in a fragmented state by a recruiting RNA motif that recruits the reverse transcriptase via an RNA-protein binding pair.
[0170] In some embodiments, the match tag sequence in the gRNA does not contain polyN (wherein N is one of the repeating bases of the ribonucleotide of the deoxyribonucleotide).
[0171] Programmable Guide RNA / Spacer One of the components is a programmable guide RNA. Its simplicity and efficiency allow the CRISPR-Cas system to be used for genome editing in cells of various organisms. The specificity of this system is directed by base pairing between the target DNA and the custom-designed guide RNA. By manipulating and adjusting the base pairing properties of the guide RNA, any desired sequence can be targeted, provided that the target sequence contains a PAM sequence.
[0172] Of the magRNA subcomponents disclosed herein, the guide sequence provides targeting specificity. It includes a region that is complementary to and can hybridize to a pre-selected target site of interest. In various embodiments, this guide sequence may contain about 10 to more than 25 nucleotides. For example, the base-pairing region between the guide sequence and the corresponding target site sequence may be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or more than 25 nucleotides in length. In exemplary embodiments, the guide sequence is about 17 to 20 nucleotides in length, for example, 20 nucleotides.
[0173] In some embodiments, one requirement for selecting a suitable target nucleic acid is that it has a 3' PAM site / sequence. In some examples, each target sequence and its corresponding PAM site / sequence are referred to herein as the Cas targeting site. One of the best-characterized systems, the Type II CRISPR system, requires only the Cas9 protein and a guide RNA complementary to the target sequence to influence the target cleavage. The Type II CRISPR system of S. pyogenes uses a target site having N12-20NGG, where NGG represents the PAM site derived from S. pyogenes, and N12-20 represents 12-20 nucleotides directly 5' relative to the PAM site. Additional PAM site sequences from other bacterial species include NGGNG, NNNNGATT, NNAGAA, NNAGAAW, and NAAAAC. For example, US20140273233, WO2013176772, Cong et al., (2012), Science 339 (6121): 819-823, Jinek et al., (2012), Science 337 (6096): 816-821, Mali et al, (2013), Science 339 (6121): 823-826, Gasiunas et al., (2012), Proc Natl Acad Sci US A. 109 (39): E2579-E2586, Cho et al., (2013) Nature Biotechnology 31, 230-232, Hou et al., Proc Natl Acad Sci US A. 2013 Sep 24;110(39):15644-9, Mojica et See al., Microbiology. 2009 Mar;155(Pt 3):733-40 and www.addgene.org / CRISPR / . The contents of these documents are incorporated herein by reference in their entirety.
[0174] The target nucleic acid strand can be either of the two strands of genomic DNA in the host cell. Examples of such genomic dsDNA include, but are not limited to, host cell chromosomes, mitochondrial DNA, and stably maintained extrachromosomal DNA. However, it should be understood that this method can be performed on other dsDNA present in the host cell, such as unstable plasmid DNA, viral DNA, and phagemid DNA, as long as a Cas targeting site is present, regardless of the properties of the host cell dsDNA. This method can also be performed on RNA.
[0175] RNA scaffolds that can bind to RNA guide nickases. In addition to the guide sequences described above, magRNAs may contain additional active or inactive subcomponents. For example, a magRNA may have an RNA motif with tracrRNA activity (e.g., a CRISPR motif). For instance, a magRNA may be a hybrid RNA molecule in which the aforementioned programmable guide RNA is fused to tracrRNA to mimic a native crRNA:tracrRNA double helix.
[0176] In some embodiments, magRNA includes both crRNA and tracrRNA, and matching sequencing follows tracrRNA. The matching tag may be present in either crRNA or tracrRNA.
[0177] Various tracrRNA sequences are known in the art, including, for example, the following tracrRNAs and their active regions. In one example, the active region of a tracrRNA retains the ability to form a complex with Cas proteins such as Cas9 or dCas9. See, for example, WO2014144592. Methods for generating crRNA-tracrRNA hybrid RNAs are known in the art. See, for example, WO2014099750, US20140179006, and US20140273226. The contents of these documents are incorporated herein by reference in their entirety. In some embodiments, the tracrRNA activity and the guide sequence are two separate RNA molecules that combine to form the guide RNA and associated scaffold. In this case, the molecule having tracrRNA activity should be able to interact with the molecule having the guide sequence (usually by base pairing).
[0178] Matching tags or mooring tags One of the distinctive features of ME is the magRNA containing a matching tag (also called a tethering tag) that is complementary to the segment in the modular RNA template. In some embodiments, the tag in the magRNA is located at the 3' end of the magRNA. In other embodiments, the tag is located at the 5' end of the magRNA. In some embodiments, the tag and spacer are located at opposite ends of the RNA scaffold; in some embodiments, the tethering tag and spacer are located at the same end of the RNA scaffold.
[0179] In some embodiments, the tag lengths are 6nt, 9nt, 12nt, 15nt, 18nt, 21nt, and 24nt. In some embodiments, the tag lengths are any size from 6nt to 24nt.
[0180] In some embodiments, the tag is complementary to the 5' end of the RNA template. In some embodiments, the tag matches an internal sequence in the RNA template. In some embodiments, the 5' end complementary sequence in the RNA template is located 23nt, 29nt, 35nt, 41nt, 50nt, and 100nt upstream (5') of the reverse transcription start site. In some embodiments, the complementary sequence in the RNA template sequence is located at a distance of 20nt to 10kb upstream of the reverse transcription start site.
[0181] In some embodiments, when a dual-module ME having two magRNAs is used, the two tag sequences in the two different magRNAs are complementary to two different segments in the RNA template.
[0182] In some embodiments, when a single-module ME is used, a match tag that matches the sequence 29nt upstream of the reverse transcription start site shows high effectiveness.
[0183] In some embodiments, two magRNAs having two different tags are used to recruit two separate RNA templates.
[0184] The selection of magRNA tags and RNA template complementary region pairs is determined by Watson-Crick base pairing. GC content, length, and melting point should be taken into consideration when designing strong tethering magRNA tags.
[0185] Examples of RNA scaffolds and magRNAs are listed below:
[0186] [ka]
[0187] MagRNA match sequences for EGFP-200L guides
[0188] [Table 1]
[0189] Recruit RNA motif In some embodiments, magRNA may include one or more recruit RNA motifs that ligate or recruit reverse transcriptase or effector proteins to a target site.
[0190] One way to recruit reverse transcriptase or effector proteins to a target sequence is by direct fusion to an RNA guide nickase such as dCas9. While direct fusion of effectors to proteins required for sequence recognition (such as dCas9) has achieved success in sequence-specific transcriptional activation or repression, protein-protein fusion designs can result in undesirable spatial interference for proteins (e.g., enzymes) that require the formation of a multimeric complex for their own activity. In such cases, it is advantageous to recruit reverse transcriptase or effector proteins to the target site in a split-type manner, as shown in Figures 6, 7, 8, 10B, 11B, 12B, 14, and 16.
[0191] More specifically, this splitting system utilizes various RNA motif / RNA-binding protein binding pairs. For this purpose, magRNA can be designed so that an RNA motif (e.g., an MS2 operator motif) that specifically binds to an RNA-binding protein (e.g., MS2 coat protein, MCP) can be ligated to or incorporated into it. The recruited RNA motif can be fused to any suitable location in the magRNA. For example, this can replace loops within the RNA scaffold, particularly tetraloops and / or stemloops.
[0192] As a result, the RNA scaffold components of the magRNA disclosed herein are engineered RNA molecules that contain not only a gRNA motif for specific DNA / RNA sequence recognition (e.g., a CRISPR RNA motif for dCas9 binding) but also a recruiting RNA motif for effector recruitment. In this way, the recruited RT or effector protein fusion can be recruited to the target site by its ability to bind to the recruiting RNA motif. Due to the flexibility of RNA scaffold-mediated recruitment, dimers, tetramers, or oligomers, along with functional monomers, can be relatively easily formed near the target DNA or RNA sequence. Such RNA recruiting motif / binding protein pairs may originate from naturally occurring sources (e.g., RNA phage or yeast telomerase) or may be artificially designed (e.g., RNA aptamers and their corresponding binding protein ligands). A non-comprehensive list of examples of recruiting RNA motif / RNA binding protein pairs that may be used in the systems described herein is summarized in Table 2.
[0193] [Table 2]
[0194] The sequences of the aforementioned conjugate pairs are listed below.
[0195] [ka]
[0196] [ka]
[0197] [ka]
[0198] [ka]
[0199] [ka]
[0200] [ka]
[0201] [ka]
[0202] RNA template molecule One of the key features of the ME system is the separation of template RNA from gRNA and the freedom in template RNA design. The modular RNA template only needs to contain a short track polyribonucleotide sequence complementary to the 3' end of a nickeling DNA strand with a free 3' end, allowing the nickeling DNA to function as a primer for new DNA strand synthesis. The length of the primer binding site (PBS) is typically 9 to 18 nt. Primer binding sites longer than 18 nt are also acceptable.
[0203] The tag sequence in magRNA is dependent on the RNA template sequence, not the other way around; therefore, it is not necessary to consider including any exogenous sequences in the template for recruitment. The design is simply to introduce desired new sequence changes to the targeted placement of the genome. In some embodiments, point mutations, which are either transitions or transversions, are included in the RNA template to introduce point mutations. In some embodiments, deletions are included in the RNA template to introduce deletions. In some embodiments, insertions are included to introduce insertions. In some embodiments, 1nt, 4nt, 29nt, and 52nt changes are introduced, including point mutations, deletions, and insertions. Nucleotide changes of other lengths from 1nt to 52nt may be made. In some embodiments, mutations, insertions, and deletions of size from 53nt to 100nt are introduced. In some embodiments, insertions and deletions of size from 101nt to 50kb are programmed by the RNA template.
[0204] In some embodiments, the distance between the sequence change and the nicking site is 1 nt, 4 nt, 20 nt, 35 nt, 50 nt, or 100 nt. Other distances between 1 nt and 100 nt may be selected. In some embodiments, the distance is between 100 nt and 1000 nt.
[0205] In some embodiments, RNA templates of 32nt, 41nt, 47nt, 72nt, 131nt, and 135nt are effective for ME. In some embodiments, modular RNA templates are 15nt to 200nt in size. In some embodiments, RNA templates are 201nt to 2kb in size. In some embodiments, RNA templates are 2kb to 10kb in size.
[0206] Typically, RNA templates contain stretches (one or more) of sequences that are identical to or homologous to sequences near the target site, in order to facilitate cell repair and increase editing efficiency.
[0207] The gene editing complexes or systems described herein can transcribe RNA sequence templates to host target DNA sites by reverse transcription primed to a target. By writing DNA sequences by direct reverse transcription of RNA sequence templates into the host genome, gene editing complexes or systems can insert target sequences into the target genome and eliminate the need for exogenous DNA sequences to be introduced into the host cell. Thus, gene editing complexes or systems provide a platform for the use of customized RNA sequence templates containing target sequences, such as sequences containing heterologous gene coding and / or functional information.
[0208] In some embodiments, the RNA template may include an open reading frame or its reverse complement. In some embodiments, the RNA template may include a coding sequence for a bacterial or viral antigen epitope. In some embodiments, the RNA template may include a sequence encoding the antigen-recognition variable region of an antibody. In some embodiments, the RNA template may include a sequence encoding the T cell receptor (TCR) antigen-recognition variable region. In some embodiments, the RNA template may include a signal peptide for efflux or a nuclear localization signal for nuclear transport or a peptide for intracellular transport. In some embodiments, the RNA template may be converted to double-stranded DNA (e.g., by reverse transcription) before the open reading frame can be transcribed and translated. In some embodiments, the RNA may include homology to a DNA target site.
[0209] In certain embodiments, RNA templates may be identified, designed, manipulated, and constructed to contain sequences that alter or specify host genome function, for example, by introducing heterologous coding regions into the genome; influencing or inducing exon structure / alternative splicing; causing disruption of endogenous genes; causing transcriptional activation of endogenous genes; causing epigenetic regulation of endogenous DNA; or causing upregulation or downregulation of operably linked genes. In certain embodiments, RNA templates may be manipulated to contain sequences encoding exons and / or transgenes that provide binding sites to transcription factor activators, repressors, enhancers, etc., and combinations thereof. In other embodiments, coding sequences may be further customized by splice acceptor sites, poly-A tails.
[0210] The RNA template can have some degree of homology to the target DNA. In some embodiments, the RNA template has at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 175, or 200 or more bases at the 3' end of the RNA template that are precisely homologous to the target DNA. In some embodiments, the RNA template has, for example, at the 5' end of the RNA template at least 2, 3, 4, 5, 6, 7, 8, 9, 90, 95, 97, 98, 99, or 100% homology to the target DNA, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 175, 180, or 200 or more bases.
[0211] The RNA template components described herein are typically capable of binding to magRNA. In some embodiments, the RNA template has a region capable of binding to magRNA. This region can be located at any preferred position within the RNA template.
[0212] An RNA template typically contains a target sequence for insertion into target DNA. The target sequence may be coding or non-coding. In some embodiments, the system or method described herein includes a single RNA template. In some embodiments, the system or method described herein includes multiple RNA templates.
[0213] In some embodiments, the template contains mutations intended to alter the PAM or protospacer sequence required by the magRNA or gRNA. As a result, the CRISPR complex can edit the original DNA sequence, but the edited DNA sequence is not exposed to recognition by the CRISPR complex, thus avoiding unproductive repetitive editing. In some embodiments, if desired, the template contains mutations to generate a new PAM or protospacer, allowing the CRISPR complex to target only the edited sequence, rather than the unedited sequence.
[0214] In some embodiments, the target sequence may contain an open reading frame. In some embodiments, the RNA template has a Kozak sequence. In some embodiments, the RNA template has an intra-sequence ribosome entry site. In some embodiments, the RNA template has a self-cleaving peptide such as a T2A or P2A site. In some embodiments, the RNA template has a start codon. In some embodiments, the RNA template has a splice acceptor site. In some embodiments, the RNA template has a splice donor site. In some embodiments, the RNA template has a stop codon. In some embodiments, the RNA template has a microRNA binding site downstream of the stop codon. In some embodiments, the RNA template has a poly-A tail downstream of the stop codon in the open reading frame. In some embodiments, the RNA template contains one or more exons. In some embodiments, the RNA template contains one or more introns. In some embodiments, the RNA template contains a eukaryotic transcription terminator. In some embodiments, the RNA template contains an enhanced translation element or translation enhancement element.
[0215] In some embodiments, the RNA template includes a coding sequence for a functional small non-protein-coding RNA. Examples of such functional non-coding RNAs (ncRNAs) include transfer RNA (tRNA), ribosomal RNA (rRNA), and small RNAs such as microRNAs, siRNAs, piRNAs, snoRNAs, snRNAs, exRNAs, and scaRNAs. In some embodiments, the non-coding RNA is an engineered small regulatory RNA that enhances or inhibits the function of its endogenous ncRNA counterpart.
[0216] In some embodiments, the target sequence may include a non-coding sequence. For example, the RNA template may include a promoter or enhancer sequence. In some embodiments, the RNA template includes a tissue-specific promoter or enhancer, which may be unidirectional or bidirectional. In some embodiments, the promoter is an RNA polymerase I promoter, an RNA polymerase II promoter, or an RNA polymerase III promoter. In some embodiments, the promoter includes a TATA element. In some embodiments, the promoter has one or more binding sites for transcription factors.
[0217] In some embodiments, the RNA template may include a promoter sequence, such as a tissue-specific promoter or enhancer sequence. In some embodiments, the tissue-specific promoter or enhancer is used to increase the target cell specificity of the gene. For example, the promoter or enhancer may be selected based on its activity in the target cell type but inactivity (or activity at a lower level) in the non-target cell type. Thus, even if the promoter or enhancer is integrated into the genome of a non-target cell, it will not drive (or only drive low-level) expression of the integrated gene. Systems having a tissue-specific promoter or enhancer sequence in the RNA template can also be used in combination with a microRNA binding site in the RNA template, for example, as described herein. In some embodiments, the RNA template may include a microRNA sequence, an siRNA sequence, a guide RNA sequence, or a piwi RNA sequence. In some embodiments, a tissue-specific silencer or repressor sequence is used to silence or suppress the expression of a target gene in target cells and tissues.
[0218] In some embodiments, the RNA template may include sites that modulate epigenetic modifications. In some embodiments, the RNA template may include elements that inhibit, for example, prevent, epigenetic silencing. In some embodiments, the RNA template may include chromatin insulators. For example, the RNA template may include CTCF sites or sites that are targeted for DNA methylation.
[0219] To promote higher levels or more stable gene expression, the RNA template may include features that prevent or inhibit gene silencing. In some embodiments, these features prevent or inhibit DNA methylation. In some embodiments, the RNA template includes sequences that mediate transcriptional repression. In some embodiments, these features promote DNA demethylation. In some embodiments, these features prevent or inhibit histone deacetylation. In some embodiments, these features prevent or inhibit histone methylation. In some embodiments, these features promote histone acetylation. In some embodiments, these features promote histone demethylation. In some embodiments, multiple features may be incorporated into the RNA template to promote one or more of these modifications. CpG dinucleotides are subjected to methylation by host methyltransferases. In some embodiments, the RNA template is CpG dinucleotide depleted, for example, does not contain CpG dinucleotides or contains a reduced number of CpG dinucleotides compared to the corresponding unmodified sequence. In some embodiments, the promoter driving the transgene expression from the incorporated DNA is CpG dinucleotide-depleted.
[0220] In some embodiments, the RNA template includes a gene expression unit comprising at least one regulatory region operably ligated to an effector sequence. The effector sequence may be a sequence transcribed to RNA (e.g., a coding sequence or a non-coding sequence, e.g., a sequence encoding a microRNA).
[0221] In some embodiments, the target sequence of the RNA template may be, for example, 50 to 50,000 base pairs (e.g., 50 to 40,000 bp, 500 to 30,000 bp, 500 to 20,000 bp, 100 to 15,000 bp, 500 to 10,000 bp, 50 to 10,000 bp, 50 to 5,000 bp). In some embodiments, the heterologous target sequence may be less than 1,000, 1,300, 1,500, 2,000, 3,000, 4,000, 5,000, or 7,500 nucleotides in length.
[0222] qualification The RNA molecules described herein (e.g., magRNA or RNA templates) may include one or more modifications. Such modifications may include at least one non-naturally occurring nucleotide or a modified nucleotide or analog thereof. The modified nucleotide may be modified in the ribose, phosphate, and / or base moieties. The modified nucleotide may include a 2'-O-methyl analog, a 2'-deoxy analog, or a 2'-fluoro analog. The nucleic acid backbone may be modified, for example, a phosphorothioate backbone may be used. The use of locked nucleic acid (LNA) or cross-linked nucleic acid (BNA) may also be possible. Yet another example of modified bases includes, but is not limited to, 2-aminopurine, 5-bromouridine, pseudouridine, inosine, and 7-methylguanosine. These modifications may be applied to any component of the systems described herein. In preferred embodiments, these modifications may be applied to RNA components, for example, a guide RNA sequence.
[0223] In some embodiments, the RNA molecule or its subsection described above may include one or more modifications, such as base modifications or skeletal modifications, to provide the nucleic acid with novel or enhanced features (e.g., improved stability).
[0224] Modified skeleton and modified nucleoside linkages Examples of suitable nucleic acids containing modifications include nucleic acids containing a modified backbone or non-natural internucleoside linkages. Nucleic acids (having a modified backbone) include nucleic acids that retain a phosphorus atom in the backbone and nucleic acids that do not have a phosphorus atom in the backbone.
[0225] Suitable modified oligonucleotide backbones containing a phosphorus atom therein include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkyl phosphotriesters, 3'-alkylene phosphonates, 5'-alkylene phosphonates and chiral phosphonates including methyl and other alkyl phosphonates, phosphinates, phosphoramidates including 3'-aminophosphoramidates and aminoalkyl phosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkyl phosphonates, thionoalkyl phosphotriesters, selenophosphates and boranophosphates, which have a normal 3'-5' linkage, or their 2'-5' linkage analogs, and those having an inverted polarity, and one or more internucleoside linkages are 3'-to-3', 5'-to-5' or 2'-to-2' linkages. Suitable oligonucleotides having an inverted polarity include a single 3'-to-3' linkage in the most 3'-internucleoside linkage, i.e., including a single inverted nucleoside residue that can be basic (lacking a nucleobase or having a hydroxyl group instead). Various salts (e.g., potassium or sodium, etc.), mixed salts and free acid forms are also included.
[0226] In some embodiments, the target nucleic acid comprises one or more phosphorothioate and / or heteroatom nucleoside linkages, particularly, -CH2-NH-O-CH2-, -CH2-N(CH3)-O-CH2- (known as the methylene(methylimino) or MMI backbone), -CH2-O-N(CH3)-CH2-, -CH2-N(CH3)-N(CH3)-CH2- and -O-N(CH3)-CH2-CH2- (in the sequence, the native phosphodiester nucleotide linkage is represented as -O-P(-O)(OH)-O-CH2-). The MMI type nucleoside linkage is disclosed in U.S. Patent No. 5,489,677, which is referred to above. Suitable amide nucleoside linkages are disclosed in U.S. Patent No. 5,602,240.
[0227] For example, nucleic acids having a morpholino backbone structure as described in U.S. Patent No. 5,034,506 are also suitable. For example, in some embodiments, the target nucleic acid comprises a 6-membered morpholino ring instead of the ribose ring. In some of these embodiments, phosphorodiamidate or other non-phosphodiester nucleoside linkages replace the phosphodiester linkages.
[0228] Suitable modified polynucleotide backbones that do not contain phosphorus atoms therein have a backbone formed by short chain alkyl or cycloalkyl nucleoside linkages, mixed heteroatom and alkyl or cycloalkyl nucleoside linkages, or one or more short chain heteroatom or heterocyclic nucleoside linkages. Such include those having morpholino linkages (formed in part from the sugar portion of the nucleoside); siloxane backbones; sulfide, sulfoxide and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl and thioformacetyl backbones; riboacetyl backbones; alkene-containing backbones; sulfamate backbones; methyleneimino and methylenehydrazino backbones; sulfonate and sulfonamide backbones; amide backbones; and others having mixed N, O, S and CH2 component parts.
[0229] Mimetic The target nucleic acids disclosed herein (e.g., RNA templates, magRNA, etc.) may be nucleic acid mimetic. The term “mimetic,” when applied to polynucleotides, is intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups; substitution of only the furanose ring is also referred to in the art as a sugar substitute. The heterocyclic base moiety or modified heterocyclic base moiety is maintained for hybridization with a suitable target nucleic acid. One such nucleic acid, a polynucleotide mimetic shown to have excellent hybridization properties, is called a peptide nucleic acid (PNA). In PNAs, the sugar backbone of the polynucleotide is replaced with an amide-containing backbone, in particular, an aminoethylglycine backbone. The nucleotide is retained and directly or indirectly bonded to the aza nitrogen atom of the amide moiety of the backbone.
[0230] One polynucleotide mimetic reported to possess excellent hybridization properties is peptide nucleic acid (PNA). The backbone in PNA compounds consists of two or more linked aminoethylglycine units, which give PNA an amide-containing backbone. The heterocyclic base moiety is directly or indirectly bonded to the aza nitrogen atom of the amide moiety of the backbone. Representative U.S. patents describing the preparation of PNA compounds include, but are not limited to, U.S. Patents 5,539,082; 5,714,331; and 5,719,262.
[0231] In some embodiments, a DNA template is used instead of an RNA template, and DNA polymerase is used instead of reverse transcriptase.
[0232] Another class of polynucleotide mimetic compounds being studied is based on linked morpholino units (morpholino nucleic acids), which have heterocyclic bases attached to a morpholino ring. Numerous linking groups have been reported for linking morpholino monomer units in morpholino nucleic acids. One class of linking groups was selected to obtain nonionic oligomeric compounds. Oligomer compounds based on nonionic morpholino are less likely to have undesirable interactions with cellular proteins. Morpholino-based polynucleotides are nonionic mimics of oligonucleotides that are less likely to form undesirable interactions with cellular proteins (Dwaine A. Braasch and David R. Corey, Biochemistry, 2002, 41(14), 4503-4510). Morpholino-based polynucleotides are disclosed in U.S. Patent No. 5,034,506. Various compounds within the morpholino class of polynucleotides, each having different linking groups for linking monomer subunits, have been prepared.
[0233] Another class of polynucleotide mimetic structures is called cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in DNA / RNA molecules is replaced by a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers are prepared and used for oligomer compound synthesis after classical phosphoramidite chemistry. Fully modified CeNA oligomer compounds and oligonucleotides with specific CeNA-modified positions have been prepared and studied (see Wang et al., J. Am. Chem. Soc., 2000, 122, 8595-8602). In general, the incorporation of CeNA monomers into DNA strands increases the stability of the DNA / RNA hybrid. CeNA oligoadenylates complexed with RNA and DNA complements with stability similar to that of native complexes. Studies of the incorporation of CeNA structures into native nucleic acid structures have shown, by NMR and circular dichroism, to facilitate conformational adaptation.
[0234] Further modifications include locked nucleic acids (LNA) in which a 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene linkage, which in turn forms a bicyclic sugar moiety. The linkage can be a methylene (-CH2-) group bridging the 2' oxygen atom and the 4' carbon atom, with n being 1 or 2 (Singh et al., Chem. Commun., 1998, 4, 455-456). LNA and LNA analogs exhibit very high double-chain thermal stability with complementary DNA and RNA (Tm=+3~+10°C), stability against 3'-exonuclease degradation, and excellent solubility properties. Potent and non-toxic antisense oligonucleotides containing LNA have been described (Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 2000, 97, 5633-5638).
[0235] The synthesis and preparation of LNA monomers adenine, cytosine, guanine, 5-methylcytosine, thymine, and uracil are described along with their oligomerization and nucleic acid recognition properties (Koshkin et al., Tetrahedron, 1998, 54, 3607-3630). LNA and its preparation are also described in WO98 / 39352 and WO99 / 14226.
[0236] Modified sugar portion The target nucleic acid may also contain one or more substituted sugar moieties. Preferred polynucleotides contain sugar substituents selected from OH;F;O-, S- or N-alkyl;O-, S- or N-alkenyl;O-, S- or N-alkynyl; or O-alkyl-Co-alkyl, wherein alkyl, alkenyl and alkynyl are substituted or unsubstituted C1-C 10 Alkyl or C2-C 10 It can be an alkenyl or alkinyl. O((CH2) n O) m CH3, O(CH2) n OCH3, O(CH2) n NH2, O(CH2)n CH3, O(CH2) n ONH2 and O(CH2) n ON((CH2) n CH3)2 (in the sequence, n and m are from 1 to about 10) are particularly preferred. Other preferred polynucleotides are C1-C 10 lower alkyl, substituted lower alkyl, alkenyl, alkynyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, a group for improving the pharmacokinetic properties of the oligonucleotide, or a group for improving the pharmacodynamic properties of the oligonucleotide, and sugar substituents selected from other substituents having similar properties. Preferred modifications include 2'-methoxyethoxy (2'-O-CH2CH2OCH3, which is also known as 2'-O-(2-methoxyethyl) or 2'-MOE) (Martin et al., Helv. Chim. Acta, 1995, 78, 486-504), i.e., an alkoxyalkoxy group. Further preferred modifications include 2'-dimethylaminooxyethoxy, i.e., the O(CH2)2ON(CH3)2 group (also known as 2'-DMAOE and described in the examples below) and 2'-dimethylaminoethoxyethoxy (also known in the art as 2'-O-dimethyl-amino-ethoxy-ethyl or 2'-DMAEOE), i.e., 2'-O-CH2-O-CH2-N(CH3)2.
[0237] Other suitable sugar substituents include methoxy(-O-CH3), aminopropoxy(-OCH2CH2NH2), allyl(-CH2-CH-CH2, -O-allylCH2-CH-CH2), and fluoro(F). The 2'-sugar substituent may be located at the arabino(upper) or ribo(lower) position. A preferred 2'-arabino modification is 2'-F. Similar modifications may be made at other positions in the oligomeric compound, particularly at the 3' position of the sugar in the 3'-terminal nucleoside or 2'-5' linked oligonucleotide, and at the 5' position of the 5'-terminal nucleotide. The oligomeric compound may also have a sugar mimetic, such as a cyclobutyl moiety, instead of a pentofuranosyl sugar.
[0238] Base modification and substitution The nucleic acids in question may also include nucleobase modifications or substitutions (often simply referred to in the art as "bases"). As used herein, "unmodified" or "natural" nucleobases include the purine bases adenine (A) and guanine (G), and the pyrimidine bases thymine (T), cytosine (C), and uracil (U). Modified nucleobases include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl(-CC-CH3)uracil and cytosine, and other alkynyl derivatives of pyrimidine bases, 6-azo(azo) This includes uracil, cytosine and thymine, 5-uracil (pseudracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, in particular 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine. Further modified nucleobases include tricyclic pyrimidines, e.g., phenoxadinecytidine (1H-pyrimido(5,4-b)(1,4)benzoxadine-2(3H)-one), phenothiazinecytidine (1H-pyrimido(5,4-b)(1,4)benzothiadin-2(3H)-one), G-clamps, e.g., substituted phenoxadinecytidine (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxadine-2(3H)-one), carbazolecytidine (2H-pyrimido(4,5-b)indole-2-one), and pyridoindolecytidine (H-pyrimido(3',2':4,5)pyrrolo(2,3-d)pyrimidine-2-one).
[0239] The heterocyclic base moiety may also include moieties in which the purine or pyrimidine base is replaced by other heterocyclic bases, such as 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine, and 2-pyridone. Further nucleobases include those disclosed in U.S. Patent No. 3,687,808, those disclosed in The Concise Encyclopedia Of Polymer Science And Engineering, pages 858-859, Kroschwitz, JI, ed. John Wiley & Sons, 1990, those disclosed by Englisch et al., Angewandte Chemie, International Edition, 1991, 30, 613, and those disclosed by Sanghvi, YS, Chapter 15, Antisense Research and Applications, pages 289-302, Crooke, ST and Lebleu, B., ed., CRC Press, 1993. Some of these nucleobases are useful for increasing the binding affinity of oligomeric compounds. Examples of such substitutions include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-methylcytosine substitution has been shown to increase nucleic acid double-strand stability by 0.6–1.2°C (Sanghvi et al., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276–278), and is a suitable base substitution, for example, when combined with 2'-O-methoxyethyl sugar modification.
[0240] Expression system To use the complexes, systems, or platforms described above, it may be desirable to express one or more of the protein and RNA components from the nucleic acids encoding them. This can be done in various ways. For example, nucleic acids encoding RNA or proteins can be cloned into one or more intermediate vectors for introduction into prokaryotic or eukaryotic cells for replication and / or transcription. Intermediate vectors are typically prokaryotic vectors, such as plasmids or shuttle vectors or insect vectors, for the production of RNA or proteins, or for the storage or manipulation of nucleic acids encoding RNA or proteins. The nucleic acids can also be cloned into one or more expression vectors for administration to plant cells, animal cells, preferably mammalian or human cells, fungal cells, bacterial cells, or protist cells. Accordingly, this disclosure provides nucleic acids encoding either the RNA or proteins mentioned above. Preferably, the nucleic acids are isolated and / or purified.
[0241] This disclosure also provides recombinant constructs or vectors having sequences encoding one or more of the RNAs or proteins described above. Examples of constructs include vectors, such as plasmids or viral vectors, in which the nucleic acid sequences of this disclosure are inserted in a forward or reverse orientation. In preferred embodiments, the construct further includes a regulatory sequence comprising a promoter operably ligated to the sequence. Numerous suitable vectors and promoters are known to those skilled in the art and are commercially available. Cloning and expression vectors suitable for use with prokaryotic and eukaryotic hosts are also described, for example, in Sambrook et al. (2001, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press).
[0242] A vector refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is ligated. A vector may be capable of autonomous replication or integration into host DNA. Examples of vectors include plasmids, cosmids, or viral vectors. The vectors of this disclosure contain nucleic acids in a form suitable for nucleic acid expression in host cells. Preferably, a vector contains one or more regulatory sequences operably ligated to the nucleic acid sequence to be expressed. "Regulatory sequences" include promoters, enhancers, and other expression regulatory elements (e.g., polyadenylation signals). Regulatory sequences include inducible regulatory sequences, along with sequences that lead to the constitutive expression of nucleotide sequences. The design of an expression vector may depend on factors such as the selection of host cells to be transformed, transfected, or transduced, the desired level of RNA or protein expression, and others.
[0243] Examples of expression vectors include chromosomes, non-chromosomal and synthetic DNA sequences, bacterial plasmids, phage DNA, baculoviruses, yeast plasmids, vectors derived from combinations of plasmids and phage DNA, and viral DNA, such as vaccinia, adenovirus, fowlpox virus, and pseudorabies. However, any other vector may be used, provided that it is replicable and viable in the host. Suitable nucleic acid sequences can be inserted into vectors by various procedures. In general, nucleic acid sequences encoding one of the RNAs or proteins described above can be inserted into suitable restriction endonuclease sites by procedures known in the art. Such procedures and related subcloning procedures are within the scope of the skill of those skilled in the art.
[0244] The vector may contain appropriate sequences for amplifying expression. In addition, the expression vector preferably contains one or more selectable marker genes for resulting in phenotypic traits for the selection of transformed host cells, such as dihydrofolate reductase or neomycin resistance for eukaryotic cell culture, or tetracycline or ampicillin resistance in E. coli.
[0245] Vectors for expressing RNA can include an RNA Pol III promoter, such as the HI, U6 or 7SK promoter, to drive the expression of RNA. Vectors for expressing RNA can include a Pol I promoter, such as an influenza virus promoter, to drive the expression of the template RNA. These human promoters enable the expression of RNA in mammalian cells after vector introduction. Alternatively, for example, a T7 promoter can be used for in vitro transcription, and the RNA can be transcribed and purified in vitro.
[0246] Using a vector containing an appropriate promoter or control sequence together with the appropriate nucleic acid sequence described above, an appropriate host can be transformed, transfected or infected to enable the host to express the RNA or protein described above. Examples of suitable expression hosts include bacterial cells (e.g., E. coli, Streptomyces, Salmonella typhimurium), fungal cells (yeast), insect cells (e.g., Drosophila and Spodoptera frugiperda (Sf9)), animal cells (e.g., CHO, COS and HEK 293), adenoviruses and plant cells. The selection of an appropriate host is within the skill of those skilled in the art. In some embodiments, the present disclosure provides a method for producing the RNA or protein mentioned above by transforming, transfecting or infecting a host cell with an expression vector having a nucleotide sequence encoding one of an RNA or a polypeptide or a protein. Next, the host cell is cultured under suitable conditions that enable the expression of the RNA or protein.
[0247] Any procedure known in the art for introducing exogenous nucleotide sequences and proteins into host cells can be used. Examples include calcium phosphate transfection, polyblens, protoplast fusion, electroporation, nucleofection, liposomes, microinjection, naked DNA, plasmid vectors, viral vectors (both for episomes and integration), and any other well-known method for introducing cloned genomic DNA, cDNA, synthetic DNA, or other exogenous genetic material into host cells.
[0248] In some embodiments, the components of the match editing system may be delivered in the form of RNA (e.g., protein-coding mRNA, magRNA, and template RNA), DNA (DNA expression vector), or protein (e.g., purified CRISPR-RT fusion protein), or any combination of the above forms (e.g., ribonucleoprotein complex).
[0249] Any procedure known in the art for systemically delivering exogenous nucleotide sequences and proteins to a host can be used. Examples include the use of viral vectors (both episome and embedding), such as adeno-associated virus (AAV), adenovirus vector (AD), lentiviral vector, retrovirus vector, embedding-deficient lentiviral vector, herpes simplex virus vector (HSV), stomatitis virus vector (VSV), modified vaccinia virus Ankara (MVA), arenavirus vector, Sendai virus vector, and parvovirus vector. Examples also include the use of non-viral vectors, such as liposomes, lipid particles, lipid nanoparticles, and genome-free virus-like particles (VLPs).
[0250] method Another aspect of this disclosure comprises methods for modifying target DNA sequences (e.g., chromosomal sequences) or target RNA sequences in cells, embryos, human or non-human animals. Sequence modifications may be base changes, insertions, deletions, or combinations thereof. In one embodiment, the embryo is a non-human animal embryo. The method comprises the steps of introducing (A) RNA guide nickase; (B) reverse transcriptase; (C) RNA template molecule and (D) magRNA molecule into a cell or embryo. The magRNA guides the other components to the target polynucleotide at the target site, and the reverse transcriptase writes the sequence to the target site based on the sequence of the RNA template molecule described herein. The RNA template contains the desired nucleotide sequence modification information.
[0251] In this specification, a target polynucleotide is a sequence that an RNA guide nickase creates a single-strand DNA break or nicking. There are no sequence restrictions on the target polynucleotide nicking site, except that a PAM sequence is immediately following (downstream or 3') the sequence when a CRISPR-type nickase is used. Examples of PAMs include, but are not limited to, NG, NGG, NGGNG, and NNAGAAW (wherein N is defined as any nucleotide and W is defined as either A or T). Other examples of PAM sequences are shown above, and those skilled in the art will be able to identify yet another PAM sequence for use with a given nickase (e.g., a CRISPR protein). In some cases, an RNA guide nickase may not require a PAM motif (e.g., an artificially engineered CRISPR protein without a PAM). Target nucleotide modifications can be located in the coding region of a gene, introns of a gene, transcriptional regulatory regions of a gene, inter-gene regulatory regions, etc. The gene may be a protein-coding gene or an RNA-coding gene.
[0252] The desired nucleotide sequence changes and their arrangement are directed by the designed RNA template. The RNA template is an individual RNA molecule not covalently linked to the guide RNA. The RNA template can be freely designed according to the desired sequence changes to be produced, without concern for interference with the guide RNA structure. There are no distance requirements between the desired DNA modification and the target nickeling site. Any distance within a 200-nucleotide range exhibits high editing efficacy. There are no content requirements for the RNA template design, except that it includes a primer-binding sequence at the 3' end complementary to the nickeled DNA strand. There is no requirement for the RNA template to include any exogenous accessory sequences for recruitment. This freedom in RNA template design mechanistically eliminates the underlying source of "scarring," which is often incorporated into genome editing sites when other reverse transcriptase, recombinase, integrase, and transposase editing tools are used.
[0253] After an RNA template is designed for a desired DNA modification at a desired target site, an RNA tag complementary to the sequence segment within the RNA template is then designed. To recruit and anchor the RNA template to the target nicking site, the tag is typically attached to the gRNA at its 3' end. The size of the RNA tag can range from 6 to 24 nucleotides. The location within the RNA template where the tag RNA is complementary can be at the 5' end or within the RNA template itself.
[0254] In another example, the present disclosure provides systems and related methods for recruiting one or more additional different functional effectors to the same target sequence. The effectors work synergistically together to facilitate genetic transformation. Examples include proteins that facilitate or increase the editing process, such as inhibitors of chromatin modifying enzymes or mismatch repair enzymes. Examples of effectors may include human RNase inhibitor proteins (RNH1), dominant-negative proteins of MMR pathway proteins including MLH1 protein, 5' DNA nuclease Fen1 protein, or others that can enhance editing efficiency.
[0255] [ka]
[0256] Thus, the match editing system provides an extremely versatile method for editing polynucleotides in cells. Cells can be single-celled organisms, cells isolated from human or non-human organisms, cells derived from or manipulated from human or non-human organisms, or cells in human or non-human organisms.
[0257] The target polynucleotide to be edited may be either endogenous or exogenous to the cell. For example, the target polynucleotide may be a DNA molecule present in the nucleus of a eukaryotic cell. The target polynucleotide may be a sequence that codes for a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide). The DNA molecule may be endogenous to the cell or may be from an infecting virus or microorganism.
[0258] The protein components of the System of the Disclosure may be introduced into cells or embryos as isolated proteins. Alternatively, the components may be introduced via nucleic acids encoding such components, such as DNA or RNA (e.g., in vitro transcribed RNA). In one embodiment, each protein may include at least one cell-permeable domain to facilitate the uptake of the protein into cells. In other embodiments, an mRNA molecule or DNA molecule encoding one or more proteins may be introduced into cells or embryos. Generally, a DNA sequence encoding a protein is operably ligated to a promoter sequence that will function in the cell or embryo of interest. The DNA sequence may be linear, or the DNA sequence may be part of a vector. In yet another embodiment, the protein may be introduced into cells or embryos as an RNA-protein complex comprising the protein, RNA template, and magRNA described above.
[0259] In alternative embodiments, the DNA encoding a protein(s) may further include one or more sequences encoding components of RNA components (e.g., magRNA and / or RNA templates). Generally, the DNA sequences encoding proteins and RNA are operably ligated to appropriate promoter control sequences that enable the expression of proteins and RNA, respectively, in cells or embryos. The DNA sequences encoding proteins and RNA may further include additional expression control, regulatory, and / or processing sequences. The DNA sequences encoding proteins and RNA may be linear or part of a vector.
[0260] In embodiments where RNA is introduced into cells via an RNA-coding DNA molecule, the RNA coding sequence may be operably ligated to a promoter control sequence for RNA expression in eukaryotic cells. For example, the RNA coding sequence may be operably ligated to a promoter sequence recognized by RNA polymerase III (Pol III). The RNA coding sequence may be operably ligated to a promoter sequence recognized by RNA polymerase I (Pol I). Examples of suitable Pol III promoters include, but are not limited to, mammalian U6 or H1 promoters. In exemplary embodiments, the RNA coding sequence is ligated to a mouse or human U6 promoter. In other exemplary embodiments, the RNA coding sequence is ligated to a mouse or human H1 promoter. In other embodiments, the RNA coding sequence is ligated to a viral Pol I promoter, for example, an influenza Pol I promoter.
[0261] The DNA molecules encoding proteins and / or RNA described herein may be linear or circular. In some embodiments, the DNA sequence may be part of a vector, such as a multicistronic vector. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, and viral vectors. In exemplary embodiments, the DNA encoding proteins and / or RNA resides within the plasmid vector. Non-limiting examples of suitable plasmid vectors include pUC, pBR322, pET, pBluescript, and their variants. The vector may include additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, and others.
[0262] The protein components (or nucleic acids(s) encoding them) and RNA components (or DNA encoding them) of the System of this Disclosure can be introduced into cells or embryos by various means. In one embodiment, the embryo is a non-human animal embryo. Typically, the embryo is a fertilized one-cell stage embryo of the species of interest. In some embodiments, the cells or embryo are transfected. Suitable transfection methods include calcium phosphate-mediated transfection, nucleofection (or electroporation), cationic polymer transfection (e.g., DEAE-dextran or polyethyleneimine), viral transfection, virosomal transfection, virion transfection, liposome transfection, cationic liposome transfection, immunoliposome transfection, non-liposomal lipid transfection, dendrimer transfection, heat shock transfection, magnetofection, lipofection, gene gun delivery, impalefection, sonoporation, optical transfection, gold nanoparticle-mediated transfection, and nucleic acid uptake enhanced with a proprietary agent. Transfection methods are well known in the art (see, for example, “Current Protocols in Molecular Biology” Ausubel et al., John Wiley & Sons, New York, 2003 or “Molecular Cloning: A Laboratory Manual” Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001). In other embodiments, the molecule is introduced into cells or embryos by microinjection. For example, the molecule may be injected into the pronucleus of a single-cell embryo.
[0263] The protein components (or nucleic acids(s) encoding them) and RNA components (or DNA encoding them) of the system disclosed herein may be introduced into cells or embryos simultaneously or sequentially. The ratio of protein (or its encoding nucleic acid) to RNA (or RNA-encoding DNA) will generally be approximately stoichiometric, so that the two can form RNA-protein complexes. Similarly, the ratio of two different proteins (or encoding nucleic acids) will be approximately stoichiometric. In one embodiment, the protein components and RNA components (or DNA sequences encoding both) are delivered together within the same nucleic acid or vector.
[0264] The method further includes the step of maintaining cells or embryos under appropriate conditions such that a guide RNA guides an effector protein to a targeting site in the target sequence, and the effector domain modifies the target sequence.
[0265] Generally, cells can be maintained under conditions suitable for cell growth and / or maintenance. Suitable cell culture conditions are well known in the art, for example, “Current Protocols in Molecular Biology” Ausubel et al., John Wiley & Sons, New York, 2003 or “Molecular Cloning: A Laboratory Manual” Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001), Santiago et al. (2008) PNAS 105:5809-5814; Moehle et al. (2007) PNAS 104:3055-3060; Urnov et al. (2005) Nature 435:646-651; and Lombardo et al. (2007) Nat. Biotechnology. This is described in 25:1298-1306. Those skilled in the art will acknowledge that methods for culturing cells are known in the art and may and will vary depending on the cell type. In all cases, routine optimization can be used to determine the best technique for a particular cell type.
[0266] Embryos can be cultured in vitro (e.g., in cell culture). Typically, embryos are cultured in appropriate temperatures and media with the necessary O2 / CO2 ratio to allow for the expression of proteins and RNA scaffolds, as needed. Suitable non-limiting examples of media include M2, M16, KSOM, BMOC, and HTF media. Those skilled in the art will recognize that culture conditions can and will vary depending on the embryo species. In all cases, routine optimization can be used to determine the best culture conditions for a particular embryo species. In some cases, cell lines may be derived from in vitro cultured embryos (e.g., embryonic stem cell lines).
[0267] Alternatively, embryos can be cultured in vivo by transferring them into the uterus of a female host. Generally speaking, the female host is of the same or similar species as the embryo. Preferably, the female host is pseudopregnant. Methods for preparing pseudopregnant female hosts are known in the art. Furthermore, methods for transferring embryos into female hosts are known. In vivo culture of embryos can allow the embryo to develop and result in the birth of live offspring of animals derived from the embryo. Such animals will contain modified chromosome sequences in all cells of their bodies.
[0268] Various eukaryotic cells are suitable for use in this method. For example, cells may be human cells, non-human mammalian cells, non-mammalian vertebrate cells, invertebrate cells, insect cells, plant cells, yeast cells, or single-celled eukaryotes. Various embryos are suitable for use in this method. For example, embryos may be 1-cell, 2-cell, or 4-cell human or non-human mammalian embryos. Exemplary mammalian embryos, including 1-cell embryos, include, but are not limited to, mouse, rat, hamster, rodent, rabbit, cat, dog, sheep, pig, cattle, horse, and primate embryos. In yet other embodiments, cells may be stem cells. Suitable stem cells include, but are not limited to, embryonic stem cells, ES-like stem cells, fetal stem cells, adult stem cells, pluripotent stem cells, induced pluripotent stem cells, polypotent stem cells, oligopotent stem cells, unipotent stem cells, and others. In exemplary embodiments, cells are mammalian cells, or embryos are mammalian embryos.
[0269] Various mammalian cells suitable for match editing include cells isolated from or derived from human or non-human subjects, as well as immune cells including NK cells, pluripotent stem cells (PSCs), adult stem cells (ASCs), fibroblasts, chondrocytes, keratinocytes, hepatocytes, pancreatic islet cells, and immune cells including T cells, dendritic cells (DCs), natural killer (NK) cells, and macrophages.
[0270] Cells suitable for match editing include cells from both human and non-human subjects. In some cases, match editing components can be delivered systemically using viral or non-viral vectors. Any procedure known in the art for systemic delivery of exogenous nucleotide sequences and proteins to a host can be used. Examples include the use of viral vectors (both episome and integration), e.g., adeno-associated virus (AAV), adenovirus vector (AD), lentiviral vector, retroviral vector, integration-deficient lentiviral vector, herpes simplex virus vector (HSV), stomatitis virus vector (VSV), modified vaccinia virus Ankara vector (MVA), arenavirus vector, Sendai virus vector, and parvovirus vector. Examples include the use of non-viral vectors, e.g., liposomes, lipid particles, lipid nanoparticles, and genome-free virus-like particles (VLPs).
[0271] Usefulness and Application The systems and methods disclosed herein have a wide range of applications, including modification and editing (e.g., inactivation and activation) of targeted polynucleotides in numerous cell types. Therefore, these systems and methods have broad applications, for example, in research and therapeutics. For instance, multiple systems with different guide RNAs can be used for high-throughput screening targeting multiple different loci to obtain and screen for multiple different phenotypic outcomes (e.g., better proliferation or lethal screening in cell lines). In another example, these systems and methods can be used in mutagenesis (similar to CRISPR tiling) or in genes to create novel proteins.
[0272] In some embodiments, the genome editing complexes, systems, compositions, or methods described above can be used to generate point mutations (both transversions and transitions), insertions, and deletions in cells derived from plant organisms, animal organisms, and humans.
[0273] In some embodiments, the genome editing complexes, systems, compositions, or methods described above can be used to generate point mutations (both transversions and transitions), insertions, and deletions in cells within plant organisms, animal organisms, and humans.
[0274] The genome editing complexes, systems, compositions, or methods described above can be used in cell manipulation, cell therapy, bioengineering, gene therapy, agricultural improvement, veterinary medicine, and cell and animal models for research.
[0275] When used to create mutations, insertions, and deletions, the match editing techniques disclosed herein enable precise genetic manipulation of DNA by copying a desired sequence from an RNA molecule and inserting it at a target location within the genome. This technique can effectively introduce genetic changes including single nucleotide changes, both transitions and transversions, insertions and deletions, or combinations of these genetic modifications. Importantly, a novel feature of match editing techniques is the independent modular design of the RNA template and magRNA. Consequently, this design avoids potential interference between gRNA and template secondary structures, eliminates sources of genetic "scars" in the template, and enables and simplifies synergistic multiplexing.
[0276] In some embodiments, the match editing techniques disclosed herein can correct mutations that cause genetic disease by installing single nucleotide changes, insertions, or deletions. In some embodiments, match editing can introduce stop codons, generate missense frameshifts, or eliminate splicing sites by installing single nucleotide changes, insertions, or deletions. As a result, the expression of target genes can be effectively silenced. In a broader sense, this technique enables the rewiring of cellular regulatory networks.
[0277] In one example, the match editing techniques disclosed herein can alter or install transcriptional regulatory sequences, such as transcription factor binding sites, to eliminate or introduce specific transcriptional activation or inhibition. In another example, match editing can alter protein modification sites, such as phosphorylation or acetylation sites, to alter cellular signaling. In yet another example, match editing can alter or install amino acid interfaces involved in protein-protein or protein-nucleic acid interactions, to eliminate or install specific protein interactions with other macromolecules, and to rewire cellular signaling pathways.
[0278] In some embodiments, ME can insert bacterial or viral antigen epitopes into cellular proteins. In some embodiments, ME can insert antibody antigen-recognition variable regions into proteins. In some embodiments, ME can replace T cell receptor (TCR) antigen-recognition variable regions. In some embodiments, ME can insert signal peptides for efflux or nuclear localization signals or peptides for nuclear transport.
[0279] In some embodiments, the match editing disclosed herein can specifically "label" intracellular or membrane proteins by creating a tag through the insertion or replacement of a short stretch of a peptide coding sequence within a target gene.
[0280] In some embodiments, the match editing disclosed herein can alter cell surface protein properties and generate novel cell-cell interaction patterns by inserting or replacing intracellular peptide sequences. For example, the variable sequence of a T cell receptor (TCR) protein can be altered to change the targeting of T cells to different antigen-presenting cells. In another example, the variable sequence of a B cell receptor (BCR) protein can be altered to change the targeting of B cells.
[0281] In some embodiments, the sequence of a secretory protein can be altered using match editing as disclosed herein, so that the protein is modified to produce cells that can secrete altered hormones, cytokines, and growth factors. In some examples, the secretory protein may be an antibody, in which case the variable region of the antibody gene may be altered to act as an antibody production switch.
[0282] In some embodiments, nerve cells can be manipulated using the match editing disclosed herein, thereby allowing for differential tagging of cells and redesigning and monitoring of their interactions and wiring.
[0283] The match editing techniques disclosed herein can be widely used for a number of important fields.
[0284] Match editing can be used for research, therapeutic development, agricultural development, and industrial bioengineering.
[0285] For the development of therapeutic methods, match editing can be used for the development of ex vivo therapies by manipulating therapeutic cells, which are either autologous cells or allogeneic interstellar cells.
[0286] For example, match editing can be used to manipulate autologous hematopoietic cells derived from patients with a disease. For instance, hematopoietic cells from patients with sickle cell anemia or beta-thalassemia can be manipulated ex vivo to correct the underlying disease-causing mutations or to reactivate fetal hemoglobin expression by altering the regulatory sequences of genes. The manipulated autologous cells are then injected back into the patient.
[0287] Similar strategies involving the introduction of point mutations or deletions into transcriptional regulatory elements of genes can generally be applied to the treatment of other diseases by eliminating transcriptional repression, introducing transcriptional activation, or introducing transcriptional repression (all by altering transcription factor binding sites).
[0288] For example, match editing can inactivate genes for cell therapy operations by introducing stop codons and frameshift insertions / deletions. For instance, it can inactivate genes involved in graft-versus-host disease (GvHD) and host-versus-graft response (HvGR), as well as genes that negatively affect CAR-T functionality, in order to generate allogeneic CAR-T cells.
[0289] For example, GvHD and HvGR genes, as well as genes that negatively affect therapeutic cell function, may be inactivated in other types of therapeutic cells, including NK cells, pluripotent stem cells (PSCs), adult stem cells (ASCs), fibroblasts, chondrocytes, keratinocytes, hepatocytes, pancreatic islet cells, and immune cells, including T cells, dendritic cells (DCs), natural killer (NK) cells, and macrophages, derived from human or non-human subjects, in order to generate allogeneic therapeutic cells and more potent therapeutic cells.
[0290] For example, match editing technology can be used to replace the MHC-antigen complex recognition sequence of a T cell receptor with a sequence that recognizes an MHC-cancer antigen complex sequence in order to generate TCR-T cells for cancer cell therapy. Furthermore, allogeneic cells may be generated by the strategies mentioned above for CAR-T cells. In addition, replacement of various regions of the TCR by match editing may be achieved for in vivo therapy.
[0291] In some cases, match editing techniques can generate regulatory T cells (Tregs) by installing an active cis-transcriptional enhancer element into the FOXP3 gene in T cells. In some cases, the TCR in Treg cells is further manipulated by match editing methods for tissue and cell-specific targeting. In some cases, Treg cells are further manipulated by match editing at HLA loci and MHC genes to generate allogeneic Tregs. In some cases, the Treg cell manipulation step is performed simultaneously with multiplexed match editing or in combination with other gene editing approaches.
[0292] For the development of therapeutic methods, expression vectors for match-editing components or genes encoding such components may be delivered by appropriate delivery media, including viral and nonviral delivery systems for in vivo gene therapy.
[0293] In one example, in vivo therapy involves delivering a match-editing system in vivo to correct a delta 508 deletion in the CFTR gene in order to treat cystic fibrosis.
[0294] One example of in vivo therapy involves delivering a match-editing system in vivo for exon skipping of mutated dystrophin genes to partially restore dystrophin function for the treatment of Duchenne muscular dystrophy.
[0295] One example of in vivo therapy involves delivering a match-editing system in vivo to reduce the expression of amyloid-beta (Aβ), tau, or alpha-synuclein in order to treat Alzheimer's disease.
[0296] The versatility of the match editing gene editing strategy allows it to encompass point mutations, insertions, deletions, and combinations of these mutations, so that similar principles of ex vivo cell manipulation for cell therapy and in vivo gene therapy can be applied to the majority of human genetic diseases. Furthermore, the mutation site does not need to be immediately adjacent to the PAM.
[0297] Similar principles of ex vivo cell manipulation for cell therapy and in vivo gene therapy can be applied to scavenging or inactivating the expression of disease-causing proteins or RNAs in order to treat cancer, autoimmune diseases, and neurodegenerative diseases.
[0298] In one example, autologous therapeutic cells can be manipulated by inactivating the GvHD and HvGR genes; in another example, allogeneic therapeutic cells can be manipulated.
[0299] Match editing can be used for the treatment of human and veterinary diseases, and for the manipulation of iPSCs of human or non-human origin.
[0300] Match editing can be used to manipulate germline cells of non-human mammals, non-human animals, plants, and agricultural plants. Living organisms can be produced from these manipulated germline cells.
[0301] Match editing can be used for industrial and non-industrial microbial genome manipulation.
[0302] Many devastating human diseases have one common cause: genetic modification or mutation. The mutations that cause disease in patients are acquired through inheritance from parents or are caused by environmental factors. Such diseases include, but are not limited to, the following categories: Firstly, some genetic disorders are caused by germline mutations. One example is cystic fibrosis, caused by a mutation in the CFTR gene inherited from a parent. A second suppressor mutation in mutant CFTR can partially restore the function of the CFTR protein in somatic cell tissues. Other examples of genetic diseases caused by correctable point genetic mutations include, to name a few, Gaucher disease, alpha-trypsin deficiency disease, and sickle cell anemia. Secondly, some diseases, such as chronic viral infectious diseases, are caused by exogenous environmental factors and the resulting genetic modifications. One example is AIDS, caused by the insertion of the human HIV virus genome into the genome of infected T cells. Thirdly, some neurodegenerative diseases involve genetic modifications. One example is Huntington's disease, caused by an expansion of the CAG trinucleotide in the huntingtin gene of affected patients. Other examples include lysosomal storage disorders, epidermolysis bullosa, and retinal degeneration. Finally, cancer is caused by various somatic mutations that accumulate in cancer cells. Therefore, correcting the genetic mutations that cause these diseases or functional corrections of their sequences offer an attractive therapeutic opportunity to treat these diseases.
[0303] Somatic genetic editing is an attractive therapeutic strategy for many human diseases. Three key factors are considered essential for successful therapeutic genetic editing: (i) a method for achieving sequence-specific recognition ("sequence recognition module"); (ii) a method for correcting the underlying mutation ("correction module"); and (iii) a method for integrating the "correction module" with the "sequence recognition module" to achieve sequence-specific correction. Numerous methods exist for achieving each of these individual tasks. However, none of the currently available platforms or technologies can achieve optimal and practical somatic genetic editing. More specifically, current gene-specific editing technologies are largely based on nuclease-induced DNA DSBs and resulting DSB-induced homologous recombination, whose activity is low or absent in most somatic cells. Therefore, these technologies have limited applications for the therapeutic correction of pathological genetic mutations in somatic tissues for most diseases.
[0304] In contrast, the systems and methods disclosed herein enable DNA sequence-directed editing of genes that do not rely on nuclease activity. These systems and methods do not generate DSBs or rely on DSB-mediated homologous recombination. Furthermore, the system design is modular, which allows for an extremely flexible and simple way to target any desired DNA or RNA sequence. Essentially, this approach also allows for guiding DNA or RNA editing enzymes to substantially any DNA or RNA sequence in somatic cells, including stem cells. Precise editing of target DNA or RNA sequences allows the enzymes to correct mutated genes in genetic disorders, inactivate viral genomes in infected cells, generate stop codons to eliminate the expression of disease-causing proteins in diseases, including neurodegenerative diseases, silence oncogenic proteins in cancer, mutate splicing consensus sites to eliminate disease-causing exons, or mutate regulatory sequences to restore therapeutic gene expression / inactivation. Therefore, the systems and methods disclosed herein can be used in correcting underlying genetic alterations in diseases including the genetic disorders, chronic infectious diseases, neurodegenerative diseases, and cancers mentioned above. Importantly, cells can be manipulated using the systems and methods disclosed herein for both the generation of research tools or the generation of cell-based therapies.
[0305] Genetic disorders It is estimated that over 6,000 genetic disorders are caused by known genetic mutations. Correcting the mutation that causes the underlying disease in the affected tissue / organ can lead to disease mitigation or cure. For example, cystic fibrosis affects 1 in 3,000 people in the United States. Cystic fibrosis is caused by inheritance of a mutated CFTR gene, and 70% of patients have the same mutation, which is a trinucleotide deletion that results in a deletion of phenylalanine at position 508 (called ΔPhe508). ΔPhe508 results in mislocalization and degradation of CFTR. Using the systems and methods disclosed herein, the Val509 residue (GTT) in affected tissue (lung) can be converted to Phe509 (TTT), thereby functionally correcting the ΔPhe508 mutation. In addition, a second suppressor mutation (such as R553Q, R553M, or V510D) in the mutant ΔPhe508 CFTR can partially restore the function of the CFTR protein in somatic cell tissues.
[0306] chronic infectious diseases The systems and methods disclosed herein can also be used to specifically inactivate any gene in a viral genome incorporated into human cells / tissues. For example, the systems and methods disclosed herein can create stop codons for the early termination of translation of essential viral genes, thereby enabling the correction or cure of chronic debilitating infectious diseases. For example, current AIDS treatments can reduce the viral load, but they cannot completely eliminate resting HIV from positive T cells. The systems and methods disclosed herein can be used to permanently inactivate the expression of essential HIV genes in the incorporated HIV genome in human T cells by introducing one or more stop codons. Another example is hepatitis B virus (HBV). The systems and methods disclosed herein can be used to specifically inactivate essential HBV genes incorporated into the human genome, thereby silencing the HBV life cycle.
[0307] Neurodegenerative diseases Some neurodegenerative diseases are caused by gain-of-function mutations. For example, SOD1G93A leads to the development of amyotrophic lateral sclerosis (ALS). Using the systems and methods disclosed herein, mutations can be corrected or mutant protein expression can be eliminated by introducing stop codons or altering splice sites. For example, the alternative splicing form of tau protein containing exon 10 plays a causal role in Alzheimer's disease. Altering the CG base pair at the consensus exon 10 splice site would eliminate the alternative splicing version of tau.
[0308] cancer Many genes (including tumor suppressor genes, oncogenes, and DNA repair genes) contribute to the development of cancer. Mutations in these genes often result in various cancers. The systems and methods disclosed herein can be used to specifically target and correct these mutations. As a result, causative oncogenic proteins can be functionally suppressed or their expression eliminated by introducing point mutations in either the catalytic or splicing site.
[0309] Somatic gene knockout In some embodiments, gene protein expression in somatic cells in human and non-human organisms can be eliminated by generating stop codons. This approach can be used for therapeutic purposes or to generate research tools.
[0310] Change of adjustment element This method can be used to alter the sequences of regulatory elements in DNA and RNA. Consequently, this method provides an approach to modify, silence, or activate gene expression by altering various mechanisms involved in gene expression. This can be used for therapeutic purposes and to generate research tools.
[0311] Stem cell genetic modification In some embodiments, cells reprogrammed to become different cell types can be genetically modified using the systems and methods disclosed herein. Suitable cells include, for example, stem cells (adult stem cells, embryonic stem cells, induced pluripotent stem cells, mesenchymal stem cells, etc., as referenced in Stem cells: past, present, and future. Zakrzewski et al. Stem Cell Res Ther. 2019 Feb 26;10(1):68.) and progenitor cells (e.g., cardiac progenitor cells, neural progenitor cells, etc.) or mature cells used for conversion to different cell types (e.g., using the algorithm as referenced in Molecular Interaction Networks to Select Factors for Cell Conversion. Ouyang JF et al., Methods Mol Biol. 2019;1975:333-361). Suitable cells may originate from any multicellular organism, including, for example, mammals (e.g., rodents, humans, horses, camels, and pigs), insects, and birds (e.g., chickens and ducks). Suitable host cells include in vitro or ex vivo host cells, such as isolated host cells.
[0312] In some embodiments, the complexes, systems, and methods disclosed herein can be used for targeted and precise genetic modification of cells or tissues ex vivo to correct underlying genetic defects. After ex vivo correction, the tissue can be returned to the patient. Furthermore, this technology can be widely used in cell-based therapies to correct genetic disorders.
[0313] The term “stem cell” as used herein refers to a cell that, under favorable conditions, can differentiate into a diverse range of specialized cell types, but under other favorable conditions, can self-regenerate and maintain an essentially undifferentiated pluripotent state. The term “stem cell” also encompasses pluripotent cells, compound pluripotent cells, progenitor cells, and precursor cells. Exemplary human stem cells can be obtained from hematopoietic or mesenchymal stem cells obtained from bone marrow tissue, embryonic stem cells obtained from embryonic tissue, or embryonic germ cells obtained from fetal reproductive tissue. Exemplary pluripotent stem cells can also be produced from somatic cells by reprogramming them into a pluripotent state through the expression of certain transcription factors associated with pluripotency; such cells are called “induced pluripotent stem cells” or “iPSc or iPS cells.”
[0314] Embryonic stem (ES) cells are undifferentiated pluripotent cells obtained from early embryos, such as the inner cell mass at the blastocyst stage, or produced by artificial means (e.g., nuclear transfer), and can give rise to any differentiated cell type in the embryo or adult, including germ cells (e.g., sperm and eggs).
[0315] Induced pluripotent stem cells (iPSc or iPS cells) are cells produced by reprogramming somatic cells by expressing or inducing the expression of a combination of factors (referred herein to as reprogramming factors). iPS cells can be produced using fetal, postnatal, neonatal, juvenile, or adult somatic cells. Factors that may be used to reprogram somatic cells into pluripotent stem cells include, for example, Oct4 (sometimes referred to as Oct3 / 4), Sox2, c-Myc, Klf4, Nanog, and Lin28. In some embodiments, to reprogram somatic cells into pluripotent stem cells, somatic cells are reprogrammed by expressing at least two, at least three, at least four, at least five, at least six, or at least seven reprogramming factors.
[0316] "Hematopoietic progenitor cells" or "hematopoietic precursor cells" refer to cells committed to the hematopoietic lineage but capable of further hematopoietic differentiation, and include hematopoietic stem cells, multipotential hematopoietic stem cells, myeloid common precursors, megakaryocyte precursors, erythrocyte precursors, and lymphoid precursors. Hematopoietic stem cells (HSCs) are multipotential stem cells that give rise to all blood cell types, including myeloid (monocytes and macrophages, granulocytes (neutrophils, basophils, eosinophils, and mast cells), erythrocytes, megakaryocytes / platelets, and dendritic cells) and lymphoid lineages (T cells, B cells, and NK cells).
[0317] "Pluripotent stem cells" refer to stem cells that have the potential to differentiate into any cell that constitutes one or more tissues or organs, or preferably any of three germ layers: endoderm (inner lining of the stomach, gastrointestinal tract, lungs), mesoderm (muscle, bone, blood, genitourinary tract), or ectoderm (epithelial tissue and nervous system).
[0318] As used herein, the term “somatic cell” refers to any cell other than a germ cell, such as an egg, sperm, or other, that does not directly transmit its own DNA to the next generation. Typically, somatic cells have limited or no pluripotency. As used herein, somatic cells may be naturally occurring or genetically modified.
[0319] Cell therapy and ex vivo therapy Various embodiments of this disclosure also provide cell lines produced or used according to any other embodiment of this disclosure for use in therapeutics. In one embodiment, this disclosure relates to a method for generating therapeutic cells, such as T cells engineered to express a chimeric antigen receptor (CAR-T) or T cell receptor (TCR-T). In one embodiment, this disclosure relates to a method for generating therapeutic regulatory T cells (Treg). CAR-T / TCR-T cells may be derived from primary T cells or differentiated from stem cells. Preferred stem cells include, but are not limited to, hematopoietic, neural, embryonic, induced pluripotent stem cells (iPSCs), mesenchymal, mesodermal, liver, pancreatic, muscle, and retinal stem cells, as well as mammalian stem cells such as human stem cells. Other stem cells include, but are not limited to, mouse stem cells, such as mouse embryonic stem cells.
[0320] In various embodiments, the complexes, systems, and methods disclosed herein can be used to knock down, modify, or increase the expression of a single gene or multiple genes in various types of cells or cell lines, including but not limited to mammalian-derived cells. The technology can be used for many applications, including but not limited to gene knockdown to prevent graft-versus-host disease by making non-host cells non-immunogenic to the host, or to prevent host-versus-graft disease by making non-host cells resistant to host attack. These approaches are also relevant to the generation of therapeutic agents based on allogeneic (off-the-shelf) or autologous (patient-specific) cells. Such genes include T cell receptors (TRAC), major histocompatibility complex (MHC class I and class II) genes including B2M, co-receptors (HLA-F, HLA-G), genes involved in innate immune responses (MICA, MICB, HCP5), inflammation (NKBBiL, LTA, TNF, LTB, LST1, NCR3, AIF1), immune receptors (LY6), heat shock proteins (HSPA1L, HSPA1A, HSPA1B), complement cascades, regulatory receptors (NOTCH4), antigen processing (TAP, HLA-DM, HLA-DO), peptide transport (RING1), increased potency or persistence (PD-1, CTLA-4, FOXP3, and B7, etc.), and genes involved in T cell interactions with the tumor microenvironment. This includes, but is not limited to, donor genes (receptors for cytokines such as TGFB, interleukin (IL)-4, IL-7, IL-2, IL-4, and repressors for IL-15, IL-12, IL-18, IL-2, and IFN-gamma), genes involved in contributions to cytokine release syndrome (including, but not limited to, GMCSF), genes encoding antigens targeted by CAR / TCR (e.g., endogenous CS1 when the CAR is designed for CS1), or other genes found to be beneficial to CAR-T / TCR-T or other cell-based therapeutics, including, but not limited to, CAR-NK, CAR-B, etc.For example, see DeRenzo et al., Genetic Modification Strategies to Enhance CAR T Cell Persistence for Patients With Solid Tumors. Front. Immunol., 15 February 2019.
[0321] This technology can also be used to knock down or modify genes involved in fratricide (killing siblings) of immune cells such as T cells and NK cells, or genes that alert the immune system of a patient or animal when foreign cells, particles, or molecules enter the patient or animal, or proteins that are currently therapeutic targets used to impair or boost the immune response, such as genes encoding CD52 and PD1, respectively.
[0322] One application involves manipulating HLA alleles in bone marrow cells to increase haplotype matching. The manipulated cells can then be used for bone marrow transplantation to treat leukemia. Another application involves manipulating the negative regulatory element of the fetal hemoglobin gene in hematopoietic stem cells to treat sickle cell anemia and beta-thalassemia. The negative regulatory element is mutated, reactivating the expression of the fetal hemoglobin gene in hematopoietic stem cells and compensating for the loss of function caused by mutations in the adult alpha or beta hemoglobin gene. Yet another application involves manipulating iPS cells to generate allogeneic therapeutic cells for various degenerative diseases, including Parkinson's disease (neuronal cell loss) and type 1 diabetes (pancreatic beta cell loss). Other exemplary applications include manipulating HIV-resistant T cells by inactivating the CCR5 gene and other genes encoding receptors required by HIV to enter the cell.
[0323] This technology can also be used to generate transgenic animals that can be used as disease models or for gene function studies.
[0324] As used herein, the term “immune cells” generally includes leukocytes (white blood cells) derived from hematopoietic stem cells (HSCs) produced in the bone marrow. Examples of immune cells include, but are not limited to, lymphocytes (T cells, B cells, and natural killer (NK) cells) and myeloid-derived cells (neutrophils, eosinophils, basophils, monocytes, macrophages, and dendritic cells).
[0325] Immune cells can be isolated from subjects, particularly human subjects. Immune cells can be obtained from subjects of interest, such as subjects suspected of having a specific disease or condition, subjects suspected of being predisposed to a specific disease or condition, or subjects undergoing treatment for a specific disease or condition. Immune cells can be collected from any location in which they exist in the subject, including but not limited to blood, umbilical cord blood, spleen, thymus, lymph nodes, and bone marrow. Isolated immune cells can be used immediately or stored for a period of time by freezing or other means.
[0326] Immune cells can be enriched / purified from any tissue in which they exist, including but not limited to blood (including blood collected by a blood bank or umbilical cord blood bank), spleen, bone marrow, tissues removed and / or exposed during surgical procedures, and tissues obtained by biopsy procedures. The tissues / organs from which immune cells are enriched, isolated and / or purified can be isolated from both living and non-living subjects, the non-living subject being an organ donor. In certain embodiments, immune cells are isolated from blood, e.g., peripheral blood or umbilical cord blood. In some embodiments, immune cells isolated from umbilical cord blood have enhanced immunomodulatory capabilities, such as those measured by CD4 or CD8-positive T cell suppression. In specific embodiments, immune cells are isolated from pooled blood, particularly pooled umbilical cord blood, due to their enhanced immunomodulatory capabilities. Pooled blood may originate from two or more sources, e.g., 3, 4, 5, 6, 7, 8, 9, 10 or more sources (e.g., donor subjects).
[0327] A population of immune cells can be obtained from a subject who requires treatment for a disease associated with reduced immune cell activity or who suffers from such a disease. Therefore, the cells may be autologous to the subject requiring treatment. Alternatively, a population of immune cells can be obtained from a donor, preferably a tissue-matched donor. The immune cell population can be collected from peripheral blood, umbilical cord blood, bone marrow, spleen, or any other organ / tissue where immune cells are present in the subject or donor. Immune cells can be isolated from a pool of subjects and / or donors, such as pooled umbilical cord blood.
[0328] When a population of immune cells is obtained from a donor separate from the target, the donor is preferably allogeneic, provided that the obtained cells are target-compatible in that they can be introduced into the target. Allogeneic donor cells may or may not be human leukocyte antigen (HLA) compatible. To make them target-compatible, allogeneic cells may be treated to reduce their immunogenicity.
[0329] In some embodiments, immune cells are T cells (e.g., regulatory T cells, CD4 + T cells, CD 8 The immune cells may be T cells or gamma-delta T cells, NK cells, invariant NK cells, NKT cells, or stem cells (e.g., mesenchymal stem cells (MSCs) or induced pluripotent stem (iPSC) cells). In some embodiments, the cells are monocytes or granulocytes, such as myeloid cells, macrophages, neutrophils, dendritic cells, mast cells, eosinophils, and / or basophils. Methods for producing and manipulating immune cells, as well as methods for using and administering cells for adoptive cell therapy, are also provided herein, in which case the cells may be autologous or allogeneic. Thus, immune cells can be used as immunotherapies, for example, to target cancer cells.
[0330] Genetic editing in animals and plants Using the systems and methods described above, transgenic non-human animals or plants having one or more desired genetic modifications can be generated. In some embodiments, the transgenic non-human animals are homozygous with respect to the genetic modifications. In some embodiments, the transgenic non-human animals are heterozygous with respect to the genetic modifications. In some embodiments, the transgenic non-human animals are vertebrates, e.g., fish (e.g., zebrafish, goldfish, pufferfish, cave fish, etc.), amphibians (e.g., frogs, salamanders, etc.), birds (e.g., chickens, turkeys, etc.), reptiles (e.g., snakes, lizards, etc.), mammals (e.g., ungulates, e.g., pigs, cows, goats, sheep, etc.; rabbits (e.g., rabbits); rodents (e.g., rats, mice); non-human primates.
[0331] The gene editing complexes, systems, and methods disclosed herein can be used to treat diseases in animals in a manner similar to those used to treat diseases in humans described above. Alternatively, they can be used to generate knock-in animal disease models with specific genetic mutations for research, drug discovery, and target validation purposes. The systems and methods described above can also be used to introduce point mutations into ES cells or embryos of various organisms for breeding and for improving animal stock and crop quality.
[0332] Methods for introducing exogenous nucleic acids into plant cells are well known in the art. Preferred methods include viral infection (e.g., double-stranded DNA viruses), transfection, conjugation, protoplast fusion, electroporation, particle gun technology, calcium phosphate precipitation, direct microinjection, silicon carbide whisker technology, Agrobacterium-mediated transformation, and others. The choice of method generally depends on the type of cells being transformed and the circumstances under which the transformation takes place (i.e., in vitro, ex vivo, or in vivo).
[0333] kit This disclosure further provides kits containing reagents for carrying out the methods described above, for example, CRISPR / Cas-guided target binding or correction reactions. To that end, one or more reaction components for the methods disclosed herein, such as RNA, RNA-guided nickase protein, reverse transcriptase protein, fusion protein, and associated nucleic acid, can be supplied in the form of a kit for use. In one embodiment, the kit includes a nickase protein, a reverse transcriptase protein, or a nucleic acid encoding such protein, an effector protein, one or more of the RNAs described above, and a set of RNA molecules described above. In other embodiments, the kit may include one or more other reaction components. In such a kit, appropriate amounts of one or more reaction components are provided in one or more containers or held on a substrate.
[0334] Examples of additional components of the kit include, but are not limited to, one or more host cells, one or more reagents for introducing exogenous nucleotide sequences into host cells, one or more reagents for detecting RNA or protein expression or verifying the state of a target nucleic acid (e.g., probes or PCR primers), and buffers or culture media for the reaction (in 1× or concentrated form). The kit may also include one or more of the following components: supports, termination, modification or digestion reagents, osmoregulators, and devices for detection.
[0335] The reaction components used can be provided in various forms. For example, components (e.g., enzymes, RNA, probes and / or primers) can be suspended in aqueous solution or provided as freeze-dried or lyophilized powder, pellets or beads. In the latter case, the components, when restored, form a complete mixture of components for use in the assay. The kits disclosed herein can be provided at any preferred temperature. For example, for storage of kits containing protein components or complexes in liquid, it is preferable to provide and maintain them below 0°C, preferably at -20°C or below -20°C, or in other frozen states.
[0336] A kit or system may contain any combination of the components described herein in an amount sufficient for at least one assay. In some applications, one or more reaction components may be supplied in pre-measured single-use amounts in individual, typically disposable, tubes or equivalent containers. Such arrangement allows the RNA-guided reaction to be carried out by directly adding the target nucleic acid or a sample or cells containing the target nucleic acid to the individual tubes. The amount of components supplied in a kit may be any appropriate amount and may depend on the target market the product is intended for. The container(s) in which the components are supplied may be any conventional container capable of retaining the supplied form, e.g., microcentrifuge tubes, microtiter plates, ampoules, bottles, or integral testing devices, e.g., fluid devices, cartridges, lateral flow devices, or other similar devices.
[0337] The kit may also include packaging materials for holding containers or combinations of containers. Typical packaging materials for such kits and systems include a solid matrix (e.g., glass, plastic, paper, foil, fine particles, and others) for holding reaction components or detection probes in any of the various configurations (e.g., in vials, microtiter plate wells, microarrays, and others). The kit may further include instructions recorded in tangible form for the use of the components.
[0338] definition Nucleic acids or polynucleotides refer to DNA molecules (e.g., cDNA or genomic DNA, but not limited to these) or RNA molecules (e.g., mRNA, but not limited to these), and include DNA or RNA analogs. DNA or RNA analogs can be synthesized from nucleotide analogs. DNA or RNA molecules may contain non-naturally occurring parts, such as modified bases, modified backbones, or deoxyribonucleotides in RNA. Nucleic acid molecules can be single-stranded or double-stranded.
[0339] When the term “isolated” refers to a nucleic acid molecule or polypeptide, it means that the nucleic acid molecule or polypeptide substantially does not contain at least one other component that is associated with or found together in nature.
[0340] As used herein, the term “guide RNA” generally refers to an RNA molecule (or group of RNA molecules collectively) that can bind to an RNA guide nickase (e.g., a CRISPR protein) and target the RNA guide nickase protein to a specific position within target DNA. Guide RNA may include two segments: a DNA targeting guide segment and a protein-binding segment. The DNA targeting segment contains a nucleotide sequence that is complementary to (or at least hybridizable under stringent conditions to) the target sequence. The protein-binding segment interacts with an RNA guide nickase (e.g., a CRISPR protein), e.g., Cas9 or a Cas9-related polypeptide. These two segments can be located on the same RNA molecule or on two or more separate RNA molecules. When the two segments are on separate RNA molecules, the molecule containing the DNA targeting guide segment is sometimes referred to as CRISPR RNA (crRNA), while the molecule containing the protein-binding segment is referred to as transactivating RNA (tracrRNA).
[0341] As used herein, the terms “target nucleic acid” or “target” refer to a nucleic acid containing a target nucleic acid sequence. Target nucleic acids can be single-stranded or double-stranded, and are often double-stranded DNA. “Target nucleic acid sequence,” “target sequence,” or “target region” as used herein mean a specific sequence or its complement that one wishes to bind or modify using the systems disclosed herein. Target sequences may be in any form of single-stranded or double-stranded nucleic acid and may be present in nucleic acids in vitro or in vivo within the genome of a cell.
[0342] The “target nucleic acid strand” refers to the strand of target nucleic acid used for base pairing with the guide RNA disclosed herein. That is, the strand of target nucleic acid that hybridizes with the crRNA and the guide sequence is referred to as the “target nucleic acid strand.” The other strand of target nucleic acid that is not complementary to the guide sequence is referred to as the “non-complementary strand.” In the case of a double-stranded target nucleic acid (e.g., DNA), each strand may be a “target nucleic acid strand” for designing the crRNA and the guide RNA and may be used in carrying out the methods of this disclosure, insofar as suitable PAM sites are present.
[0343] As used herein, the term “derived from” means a process in which a first component (e.g., a first molecule) or information derived from this first component is used to isolate, derive, or construct a different second component (e.g., a second molecule different from the first molecule). For example, mammalian codon-optimized Cas9 polynucleotides are derived from the amino acid sequence of the wild-type Cas9 protein. Also, variant mammalian codon-optimized Cas9 polynucleotides, including Cas9 single mutant nickase (nCas9, e.g., nCas9D10A) and Cas9 double mutant null-nuclease (dCas9, e.g., dCas9 D10A H840A), are derived from polynucleotides encoding the wild-type mammalian codon-optimized Cas9 protein.
[0344] As used herein, the term “wild type” is a term of the art as understood by those skilled in the art, and means the typical form of an organism, lineage, gene or feature as it occurred in nature, distinguished from mutant or variant forms.
[0345] As used herein, the term “variant” refers to a first composition (e.g., a first molecule) relating to a second composition (e.g., a second molecule, also named the “parent” molecule). A variant molecule may be derived from, isolated from, based on, or homologous to the parent molecule. For example, mutant forms of mammalian codon-optimized Cas9 (hspCas9), including Cas9 single mutant nickase and Cas9 double mutant null-nuclease, are variants of mammalian codon-optimized wild-type Cas9 (hspCas9). The term “variant” can be used to represent either a polynucleotide or a polypeptide.
[0346] When applied to polynucleotides, a variant molecule may have complete nucleotide sequence identity with the original parent molecule, or conversely, it may have less than 100% nucleotide sequence identity with the parent molecule. For example, a variant of a gene nucleotide sequence may be a second nucleotide sequence whose nucleotide sequence is at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% or more identical to the original nucleotide sequence. A polynucleotide variant also includes a polynucleotide that includes the entire parent polynucleotide and further includes additional fused nucleotide sequences. A polynucleotide variant also includes a polynucleotide that is a part or subsequence of the parent polynucleotide, and specific subsequences of polynucleotides disclosed herein (determined, for example, by standard sequence comparison and alignment techniques) are also covered by this disclosure.
[0347] In another embodiment, a polynucleotide variant includes a nucleotide sequence containing a trace, trivial, or insignificant change to the parent nucleotide sequence. For example, trace, trivial, or insignificant changes include changes to the nucleotide sequence that (i) do not change the amino acid sequence of the corresponding polypeptide, (ii) occur outside the protein-coding open reading frame of the polynucleotide, (iii) result in a deletion or insertion that may affect the corresponding amino acid sequence but have little or no effect on the biological activity of the polypeptide, or (iv) result in an amino acid substitution with a chemically similar amino acid. If the polynucleotide does not encode a protein (e.g., tRNA or crRNA or tracrRNA), the variant of the polynucleotide may include nucleotide changes that do not result in a loss of function of the polynucleotide. In another embodiment, conserved variants of the disclosed nucleotide sequences that result in functionally identical nucleotide sequences are encompassed by this disclosure. Those skilled in the art will recognize that many variants of the disclosed nucleotide sequences are encompassed by this disclosure.
[0348] When applied to proteins, a variant polypeptide may have complete amino acid sequence identity with the original parent polypeptide, or conversely, it may have less than 100% amino acid identity with the parent protein. For example, an amino acid sequence variant may be a second amino acid sequence that is at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the original amino acid sequence.
[0349] Polypeptide variants include polypeptides comprising the entire parent polypeptide and further comprising additional fused amino acid sequences. Polypeptide variants also include polypeptides that are a part or partial sequence of the parent polypeptide, and specific partial sequences of polypeptides disclosed herein (determined, for example, by standard sequence comparison and alignment techniques) are also included in this disclosure.
[0350] In another embodiment, polypeptide variants include polypeptides containing trace, trivial, or insignificant changes to the parent amino acid sequence. For example, trace, trivial, or insignificant changes include amino acid changes (including substitutions, deletions, and insertions) that have little or no effect on the biological activity of the polypeptide and result in a functionally identical polypeptide, including the addition of a non-functional peptide sequence. In yet another embodiment, variant polypeptides of the Disclosure include mutant variants of Cas9 polypeptides in which the biological activity of the parent molecule is altered, for example, in which nuclease activity is modified or lost. Those skilled in the art will recognize that many variants of the disclosed polypeptides are encompassed by the Disclosure.
[0351] In some embodiments, the polynucleotide or polypeptide variants of the present disclosure may include variant molecules that modify, add, or delete a small percentage of nucleotide or amino acid positions, for example, typically less than about 10%, less than about 5%, less than 4%, less than 2%, or less than 1%.
[0352] As used herein, the term “conservative substitution” in nucleotide or amino acid sequences refers to a change in a nucleotide sequence that (i) results in no corresponding change to the amino acid sequence due to the redundancy of the triplet codon code, or (ii) results in the substitution of the original parent amino acid by an amino acid having a chemically similar structure. Conservative substitution tables presenting functionally similar amino acids are well known in the art, in which one amino acid residue is substituted for another amino acid residue having similar chemical properties (e.g., aromatic or positively charged side chains), and thus does not substantially alter the functional properties of the resulting polypeptide molecule.
[0353] Next, we present a classification of naturally occurring amino acids that possess similar chemical properties, where substitutions within the group are "conservative" amino acid substitutions. This classification, indicated below, is not rigid, as these naturally occurring amino acids may be placed in different groups if different functional properties are considered. Amino acids with nonpolar and / or aliphatic side chains include glycine, alanine, valine, leucine, isoleucine, and proline. Amino acids with polar, uncharged side chains include serine, threonine, cysteine, methionine, asparagine, and glutamine. Amino acids with aromatic side chains include phenylalanine, tyrosine, and tryptophan. Amino acids with positively charged side chains include lysine, arginine, and histidine. Amino acids with negatively charged side chains include aspartic acid and glutamic acid.
[0354] A "Cas9 mutant" or "Cas9 variant" refers to a protein or polypeptide derivative of the wild-type Cas9 protein, such as the S. pyogenes Cas9 protein, which has one or more point mutations, insertions, deletions, truncations, fusion proteins, or combinations thereof. It substantially retains the RNA targeting activity of the Cas9 protein. The protein or polypeptide may contain, consist of, or be essentially composed of fragments of the wild-type protein. Generally, the mutant / variant is at least 50% (e.g., any number from 50% to 100% (including both ends)) identical to the protein. The mutant / variant may be able to bind to RNA molecules, be targeted to specific DNA sequences via RNA molecules, and may also possess nuclease activity. Examples of these domains include the RuvC-like motif (aa.7-22, 759-766, and 982-989) and the HNH motif (aa.837-863). See Gasiunas et al., Proc Natl Acad Sci US A. 2012 September 25; 109(39): E2579-E2586 and WO2013176772.
[0355] "Complementarity" refers to the ability of a nucleic acid to form hydrogen bonds (or more) with another nucleic acid sequence, either through traditional Watson-Crick base pairing or other non-traditional methods. Percent complementarity refers to the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, and 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary, respectively). "Perfectly complementary" means that all adjacent residues in one nucleic acid sequence will form hydrogen bonds with the same number of adjacent residues in the second nucleic acid sequence. "Substantially complementary," as used herein, refers to a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% across regions of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids that hybridize under stringent conditions.
[0356] As used herein, “stringent conditions” for hybridization refer to conditions under which a nucleic acid complementary to the target sequence hybridizes primarily with the target sequence and substantially not with the non-target sequence. Stringent conditions are generally sequence-dependent and vary depending on numerous factors. Generally, the longer the sequence, the higher the temperature at which the sequence specifically hybridizes with its target sequence. Non-limiting examples of stringent conditions are described in detail in Tijssen (1993), Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes Part I, Second Chapter “Overview of principles of hybridization and the strategy of nucleic acid probe assay”, Elsevier, NY.
[0357] Hybridization, or the act of hybridizing, refers to the process by which completely or partially complementary nucleic acid strands unite under specified hybridization conditions to form a double-stranded structure or region in which the two constituent strands are linked by hydrogen bonds. Hydrogen bonds are typically formed between adenine and thymine or uracil (A and T or U) or cytidine and guanine (C and G), but other base pairs may also form them (e.g., Adams et al., The Biochemistry of the Nucleic Acids, 11th ed., 1992).
[0358] As used herein, “expression” refers to the process by which polynucleotides are transcribed from a DNA template (to mRNA or other RNA transcripts, etc.) and / or the process by which the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and the polypeptides they encode may be collectively referred to as “gene products.” If the polynucleotides originate from genomic DNA, expression may include the splicing of mRNA in eukaryotic cells.
[0359] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to polymers of amino acids of any length. Polymers may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acid groups. These terms also encompass modified amino acid polymers; e.g., disulfide bond formation, glycosylation, lipid addition, acetylation, phosphorylation, pegylation, or any other manipulation, e.g., conjugation with a labeling component. Where used herein, the term “amino acid” includes glycine and both D or L optical isomers, as well as natural and / or unnatural or synthetic amino acids, including amino acid analogs and peptidomimetic compounds.
[0360] The terms “fusion polypeptide” or “fusion protein” refer to a protein created by integrating two or more polypeptide sequences. Fusion polypeptides as encompassed in this disclosure include the translation product of a chimeric gene construct in which a first polypeptide, e.g., a nucleic acid sequence encoding an RNA-binding domain, is fused with a second polypeptide, e.g., a nucleic acid sequence encoding an effector domain, to form a single open reading frame. In other words, a “fusion polypeptide” or “fusion protein” is a recombinant protein of two or more proteins linked by peptide bonds or via several peptides. A fusion protein may also include a peptide linker between the two domains.
[0361] The term "linker" refers to any means, entity, or part used to conjugate two or more entities. A linker may be a covalent or non-covalent linker. An example of a covalent linker is a linker moiety covalently attached to one or more proteins or domains to be linked. A non-covalent linker may be an organometallic bond via a metal center, such as a platinum atom. For covalent linking, various functional groups can be used, including amide groups containing carbon dioxide derivatives, ethers, esters containing organic and inorganic esters, aminos, urethanes, ureas, and others. To provide conjugation, domains can be modified by oxidation, hydroxylation, substitution, reduction, etc., to provide a site for coupling. Methods for conjugation are well known to those skilled in the art and are incorporated for use in this disclosure. The linker moiety includes, but is not limited to, a chemical linker moiety or, for example, a peptide linker moiety (linker sequence). Modifications that do not significantly reduce the function of the RNA-binding domain and effector domain are preferred.
[0362] As used herein, the terms “conjugate,” “conjugation,” or “linked” refer to the act of attaching two or more entities together to form a single entity. Conjugates encompass both peptide-small molecule conjugates and peptide-protein / peptide conjugates.
[0363] The terms “subject” and “patient” are used interchangeably herein to refer to vertebrates, preferably mammals, more preferably humans. Mammals include, but are not limited to, mice, monkeys, humans, livestock, sport animals, and pets. Tissues, cells, and their offspring of biological entities obtained in vivo or cultured in vitro are also included. In some embodiments, the subject may be an invertebrate, such as an insect or nematode; in other embodiments, the subject may be a plant or a fungus.
[0364] As used herein, “treatment,” “to treat,” “to alleviate,” and “to improve” are interchangeable. These terms refer to an approach to obtain a beneficial or desired outcome, including but not limited to therapeutic and / or preventive benefits. Therapeutic benefit means any therapeutically relevant improvement or effect in one or more diseases, conditions, or symptoms during treatment. For preventive benefit, a composition may be administered to subjects at risk of developing a particular disease, condition, or symptom, or to subjects reporting one or more physiological symptoms of a disease, even if the disease, condition, or symptom may not yet be present.
[0365] The phrase "pharmaceutically or pharmacologically acceptable" refers, where appropriate, to molecular entities and compositions that, when administered to animals such as humans, do not produce harmful, allergic, or other troubling reactions. The preparation of therapeutic agents, e.g., pharmaceutical compositions containing cells or additional active ingredients, will be apparent to those skilled in the art in light of this disclosure. Furthermore, for administration to animals (e.g., humans), it will be understood that the preparations should meet the sterility, pyrogenicity, general safety, and purity standards required by the FDA Office of Biological Standards. As used herein, “pharmaceutically acceptable carriers” include, as is well known to those skilled in the art, all kinds of aqueous solvents (e.g., water, alcoholic / aqueous solutions, saline solutions, parenteral media, e.g., sodium chloride, Ringer's dextrose), non-aqueous solvents (e.g., propylene glycol, polyethylene glycol, vegetable oils and organic esters for injection, e.g., ethyl oleate), dispersions, coatings, surfactants, antioxidants, preservatives (e.g., antibacterial or antifungal agents, antioxidants, chelating agents and inert gases), isotonic agents, absorption retarders, salts, drugs, drug stabilizers, gels, binders, excipients, disintegrants, lubricants, sweeteners, flavorings, pigments, fluids and nutritional supplements, similar materials, and combinations thereof. The pH and precise concentrations of the various components in a pharmaceutical composition are adjusted according to well known parameters.
[0366] As used herein, the term “to bring into contact” includes any process in which the components to be brought into contact are mixed into the same mixture (e.g., added to the same compartment or solution) and does not necessarily require actual physical contact between the enumerated components. The enumerated components may be brought into contact in any order or any combination (or subcombination) and may include situations in which one or more of the enumerated components are subsequently removed from the mixture, as may be done, prior to the addition of other enumerated components. For example, “to bring A into contact with B and C” includes any of the following situations: (i) A is mixed with C, and then B is added to the mixture; (ii) A and B are mixed into the mixture; B is removed from the mixture, and then C is added to the mixture; and (iii) A is added to the mixture of B and C. "Contacting" a target nucleic acid or cell with one or more reaction components, such as a Cas protein or guide RNA, includes any of the following situations: (i) the target or cell is contacted with a first component of the reaction mixture to create the mixture; then other components of the reaction mixture are added to the mixture in any order or combination; and (ii) the reaction mixture is fully formed prior to mixing with the target or cell.
[0367] The term "split" or "split state," when referring to the composition of a CRISPR-effector protein(s), means that CRISPR and the effector polypeptide are separated into two individual polypeptides. In other words, the CRISPR and effector proteins are not covalently fused to form a single polypeptide.
[0368] As used herein, the term “mixture” refers to a combination of elements that are scattered and not in any particular order. A mixture is heterogeneous and spatially inseparable from its different components. Examples of a mixture of elements include a number of different elements dissolved in the same aqueous solution, or a number of different elements attached randomly or in no particular order to a solid support, in which case the different elements are not spatially distinct. In other words, a mixture is not addressable.
[0369] As disclosed herein, numerous ranges of values are provided. It is understood that, to the extent of a tenth of a unit of the lower limit, each of the values intervening between the upper and lower limits of such range is also specifically disclosed, unless the context clearly indicates otherwise. Each smaller range between any described value or intervening value in a described range and any other described value or intervening value in that described range is encompassed within this disclosure. Such upper and lower limits of smaller ranges may be independently included in or excluded from the range, and each range that includes either or both limits, or neither limits, is also encompassed within this disclosure and is subject to the influence of any specifically excluded limit in the described range. Where the described range includes one or both limits, any range that excludes either or both of those included limits is also included in this disclosure. The term “about” generally refers to plus or minus 10% of the number indicated. For example, "approximately 10%" can refer to a range of 9% to 11%, and "approximately 20" can mean 18 to 22. Other meanings of "approximately" can be clear from the context, such as rounding; therefore, for example, "approximately 1" can mean 0.5 to 1.4. [Examples]
[0370] material and method Cells and cell culture conditions HEK293T and K562 cells were purchased from ATCC (HEK293T, CRL-3216; K562, CCL-243). HEK293T cells were grown and maintained at 37°C and 5% CO2 in Dulbecco's modified Eagle medium (Thermo Fisher Scientific) supplemented with 10% fetal bovine serum, 1× glutamine (Thermo Fisher Scientific), and 1× antibiotic antifungal solution (Thermo Fisher Scientific). K562 cells were grown in RPMI 1640 medium containing 10% fetal bovine serum and 1× glutamine.
[0371] Transfection Transfection was performed by electroporation using a Neon transfection system (Thermo Fisher Scientific) according to the manufacturer's instructions. In short, 2 × 10 5Approximately 2 μg of DNA from Cas9 / RT, gRNA / magRNA, template RNA, and episome-targeted plasmid constructs was transfected into individual cells. For the experiments in Figures 17–29, the amounts of each expression vector DNA used in each electroporation were: 250 ng of gRNA or magRNA or molar equivalent of pegRNA expression vector, 250 ng of RNA template expression vector, 500 ng of nCas9(H840A)-RT fusion protein expression vector, 250 ng of MCP-RT fusion expression vector, and 100 ng of targeted episome-targeted plasmid DNA (or sometimes 500 ng when targeting the episome EGFP sequence as indicated). For the experiments in Figure 30 and thereafter, the amounts of DNA used for nCas9(H840A)-RT, MCP-RT fusion expression vector, and episome-targeted plasmid DNA remained the same as described above. The gRNA, pegRNA, magRNA, and RNA template expression vectors varied and are described in each experiment. Unless otherwise specified, the amounts of magRNA and the second nicking gRNA expression vectors were 250 ng each. The RNA template expression vector was 1000 ng, and the pegRNA expression vector was the molar equivalent to 1000 ng of RNA template expression vector. Under these most commonly used conditions shown in Figures 30–43, the molar ratio of RNA template in the PE system: pegRNA scaffold in PE: RNA template in ME: magRNA scaffold in ME was roughly equal to 1:1:1:0.25.
[0372] U6 expression cassettes for pegRNA, RNA template, magRNA, and gRNA were cloned into pBS KS(+) vectors (2964 bp), resulting in plasmid sizes of approximately 3500 bp in length. Since these plasmids also share the U6 promoter and transcriptional terminator, the size difference between the pegRNA, gRNA, magRNA, and RNA template expression vectors is typically within 100 bp (of the approximately 3500 bp total length). CMV expression cassettes for proteins such as nCas9, nCas9-RT fusion, MCP-RT, and MCP-p65 are cloned into pcDNA3.1(+) (5.4 kb).
[0373] When comparing the activity between the ME and PE systems, the molar ratio of RNA template (in pegRNA) in the PE system: pegRNA scaffold in PE: RNA template in ME: magRNA scaffold in ME was roughly equal to 1:1:1:0.25, as described above. Student's t-tests were performed for the comparison, where * indicates P<0.05, **P<0.01, and ***P<0.001.
[0374] After electroporation, cells were seeded in 12-well plates. Three days after electroporation, EGFP-containing cells were observed under a fluorescence microscope, and representative images were captured. To quantify the percentage of fluorescent EGFP cells, cells were subjected to flow cytometry analysis. All experiments were performed with two or three biologically independent replications. To directly determine base changes in target sequences, DNA sequences containing the target site (either EGFP or endogenous gene) were amplified by PCR, analyzed by Sanger sequencing, and quantified using the EditR analysis tool. Some samples were further analyzed by next-generation sequencing (NGS). Specific conditions for each experiment are further described in the legend in the figure.
[0375] Design of magRNA and RNA templates for match editing To design magRNA and RNA templates for protein-coding or RNA-coding targets, the RNA template is typically a sense sequence. This minimizes the potential formation of double-stranded RNA in the cell. In this scenario, the non-coding strand (or bottom strand) is nicked and extended using the sense RNA template sequence. This is named the bottom strand nicking and extension mode.
[0376] For lower strand nicking and extension modes, the following are example steps used to design RNA templates and magRNAs: a. Align two complementary target DNA strands, each containing a sense upper strand DNA (5'-3' orientation, right to left) and its complementary antisense lower strand DNA (3'-5' orientation, right to left).
[0377] b. Identify the PAM motif in the lower chain and identify the nicking site in the lower chain (the lower chain is the non-target chain, and the upper chain is the target chain complementary to the guide or spacer sequence, with the guide being immediately 5' relative to the PAM).
[0378] c. As an example, to design an RNA template with a 13-nucleotide prime-binding sequence (PBS) length and a 35-nucleotide RT extension length: i. Identify the lower chain cutting site (located 3nt 5' relative to PAM); ii. Identify the upper strand PBS: Starting with the nucleotide complementary to the cutting nucleotide (lower strand) (upper strand), count approximately 13 nt toward the 3' end of the upper strand (ideally ending in C or G); iii. Upper chain RT extension: Starting from the complementary nucleotide of the cutting nucleotide, count 35nt toward the 5' end of the upper chain (ideally ending with G); or, depending on the nature of the deletion / insertion, select an appropriate region of homologous sequence for RT extension.
[0379] iv. Copy the complete 5'-RT extension (35nt) + PBS (13nt)-3' sequence from the top strand to a new file.
[0380] v. Modify the template by adding the desired mutation to the RT extension sequence in the above-mentioned file.
[0381] vi. The resulting sequence is the match-edited RNA template.
[0382] To design the d.magRNA sequence, we will use a 12nt tethering tag as an example: i. Position the PAM and guide from the lower chain; ii. Copy the guide (lower chain, approximately 20 nt, 5' of PAM, copied in 5'-3' direction).
[0383] iii. Paste the guide / spacer sequence onto the 5' end of the gRNA scaffold (using the spCas9 gRNA scaffold as an example).
[0384] iv. From the template, find the 12-nucleotide sequence to be tethered (for example, the sequence starting at 29nt 5' of the PBS sequence). Copy the lower strand (5'-3') that is complementary to the v.12 nucleotide sequence.
[0385] vi. Then paste it onto the 3' end of the gRNA (using spCas9 gRNA scaffolds as an example, which have a tethering tag attached to the 3' end of the gRNA without a linker).
[0386] vii. A magRNA having a 5'-spacer / guide + gRNA scaffold + a 12nt 3' template tethering tag sequence.
[0387] To design RNA templates and magRNAs with upper strand nicking, an example of upper strand nicking and extension mode design steps is described below: a. As an example, to design an RNA template having a 13-nucleotide prime binding sequence and a 35-nucleotide RT elongation length: i. Align two complementary target DNA strands, each containing an upper strand DNA (5'-3' orientation, right to left) and a complementary lower strand DNA (3'-5' orientation, right to left).
[0388] ii. Identify the upper strand PAM motif and nicking site (the upper strand is the non-target strand, and the lower strand is the target strand containing sequences complementary to the guide / spacer sequences in the gRNA / magRNA, with the guide located immediately 5' relative to the PAM and the nicking site located 3nt 5' relative to the PAM); iii. Reading the lower strand: For PBS with a length of 13 nucleotides in the design: Starting from the complementary nucleotide (lower) of the cutting site (upper), count 13 nt toward the 3' end of the lower strand; iv. 35nt lower-chain RT elongation length: Count 35nt towards the 5' end of the lower chain, starting from the complementary nucleotide of the cutting site; or, depending on the nature of the deletion / insertion, select an appropriate region of homologous sequence for RT elongation.
[0389] v. Copy the sequence below containing the PBS and RT extension sequences mentioned above (copy in the 5'-3' orientation) and paste it onto the template plasmid.
[0390] The template is modified by adding a desired mutation to the vi.RT extension sequence.
[0391] vii. The resulting sequence is the match-edited RNA template.
[0392] To design the b.magRNA sequence, use the size of a 12nt tethering tag as an example: i. Identify the upper strand PAM motif and nicking site (the upper strand is the non-target strand, and the lower strand is the target strand containing sequences complementary to the guide / spacer sequences in the gRNA / magRNA, with the guide being immediately 5' relative to the PAM); ii. Copy the guide (upper chain, 20 nt, 5' of PAM, 5'-3' direction).
[0393] iii. Paste the guide sequence onto the 5' end of the gRNA scaffold (using the spCas9 gRNA scaffold as an example). iv. From the template, find the 12-nucleotide sequence to be tethered (for example, the sequence starting at 29nt 5' of the PBS sequence).
[0394] v. Copy the underlying strand sequence complementary to the 12-nucleotide sequence to be tethered (copy in the 5'-3' direction).
[0395] vi. Paste this onto the 3' end of the gRNA (using a spCas9 gRNA scaffold as an example, which has a tethering tag attached to the 3' end of the gRNA without a linker).
[0396] vii. A magRNA having a 5'-spacer / guide + gRNA scaffold + a 12nt 3' template tethering tag sequence.
[0397] For cell expression, an RNA template or magRNA sequence was inserted into the expression plasmid between the U6 promoter sequence and the U6 transcription terminator. It is important to note that this description is intended to provide some examples, but not to limit the way in which the system may be designed.
[0398] array An example sequence used in the embodiment is shown below.
[0399] [ka]
[0400] [ka]
[0401] [ka]
[0402] [ka]
[0403] The niccase saCas9(N580A)-KKH variant sequence is identical to the saCas9(N580A) sequence except for three mutations (E782K / N968K / R1015H).
[0404] [ka]
[0405] Note: In the example above, the nuclear localization signal (NLS) is attached to the N-terminus or C-terminus of the fusion protein, and the linker peptide is attached between the fusion regions.
[0406] [ka]
[0407] Note: a. The sequence encodes wild-type EGFP. b.nfEGFP contains the A200G point mutation (A200G, Tyr66Cys). The A200 position in the wild-type EGFP gene is marked and underlined. c. In the deletion constructs EGFPΔ4A, Δ20C, Δ35C, Δ50C, and Δ100G, the deleted nucleotides are labeled and marked (underlined). The relative numbers 4, 20, 35, 50, and 100 are relative distances toward the lower strand NGG PAM (in the sequence, N is labeled as +1, corresponding to position T199). d. EGFP insertion construct, EGFP-4nt-insertion, contains four nucleotides (ATAG) between nucleotides 179 and 180. e.EGFP123Δ52nt contains a deletion from nucleotides 124-175 (a total of 52 nucleotides).
[0408] [ka]
[0409] Note: Sequences with a single underline are U6 promoter sequences, and sequences with a double underline are used as U6 terminators. Template RNA, gRNA, or magRNA sequences are inserted between the promoter and terminator sequences for expression.
[0410] [ka]
[0411] Note: The underlined sequence is the ribozyme HDV sequence. The magRNA is inserted between the U6 promoter and the HDV sequence.
[0412] [ka]
[0413] Note: The first single-underlined sequence at the beginning is a 5' half-GFP sequence, and the last single-underlined sequence is a 3' half-GFP sequence. The sequence in between is an artificial intron with SAS and SDS sites at the intron-exon junction.
[0414] [ka]
[0415] Note: The first single-underlined sequence at the beginning is the 5' half-GFP sequence, and the last single-underlined sequence is the 3' half-GFP sequence. The double-underlined sequence is the mouse Dmd exon 23 sequence. The sequence between the 5' half-GFP and the exon 23 sequence contains mouse Dmd intron 22, while the sequence between exon 23 and the 3' half-GFP sequence contains Dmd intron 23.
[0416] [Table 3]
[0417] [Table 4]
[0418] The magRNA matching sequences for the EGFP-200L guide have been previously included in the matching gRNA examples section.
[0419] Some additional magRNA information is included below:
[0420] [Table 5]
[0421] [Table 6]
[0422] [Table 7]
[0423] [Table 8-1]
[0424] [Table 8-2]
[0425] [Table 9]
[0426] [Table 10] [Examples]
[0427] This embodiment describes a specific match-editing (ME) system targeting EGFP with a point mutation at position 200 (A200G). The system was delivered by a plasmid containing an RNA template, a U6 promoter for transcription of magRNA and control gRNA (Figure 17A), and a CMV promoter for nCas9-RT fusion protein and EGFP expression (Figures 17A and 17B). The target site of EGFP 200L is also included in the figures and analyzed in detail, showing an underlined lower-chain PAM (3'-gga-5') followed by a guide sequence. Furthermore, the mutant base pairs, nicking site, and primer-binding sequence are all indicated in Figure 17C. In addition, the RNA template design is illustrated in Figure 17D, which includes a 3'-primer-binding site (P14) that is exactly the same as the sequence illustrated at the target site, except for the nucleotide(s) to be changed, and a 5'RT extension that is exactly the same as the sequence at 5' relative to the primer-binding site at the target site. In experiments to correct the A200G point mutation, two templates were used, one with a 29-nucleotide 5'RT extension and the other with a 57-nucleotide extension. [Examples]
[0428] This example describes experimental results using the ME system described in Example 1 to correct point mutations.
[0429] In short, cells were electroporated with plasmids expressing the indicated ME component or control component. The percentage of cells expressing fluorescent EGFP was determined by flow cytometry (Figure 18A), and target EGFP DNA fragments were amplified from control (Figure 18B) or ME (Figure 18C) treated cells, and the complementary strands were sequenced. The results are shown in Figure 18A (lane 1, untreated cells; lane 2, nfEGFP-expressing cells; lanes 3-6, cells expressing various ME components as indicated; lane 7, prime-edited control). Figures 18B and 18C show the sequencing results from lanes 2 and 4, respectively.
[0430] As shown in Figure 18A, ME with magRNA demonstrated high gene editing efficiency, with a functional editing efficiency of 60% for mutations in EGFP (lane 4). Furthermore, compared to gRNA without a complementary match tag, magRNA dramatically increased editing effectiveness (more than a 3-fold increase, lanes 3 and 4). The effectiveness was comparable to that of the primed editing counterpart (no statistically significant difference between lanes 4 and 7).
[0431] Figures 18B and 18C demonstrate confirmation that the conversion from non-fluorescent EGFP to fluorescent EGFP was indeed induced by a complementary C-to-T point mutation. Importantly, sequencing results (Figure 18C) clearly demonstrated the absence of bystander editing; i.e., only the target C was converted to T, while the adjacent C was not. The absence of bystander editing is one of the unique advantages over base editing. The discrepancy in editing efficiency between direct Sanger sequencing (26%) and the functional EGFP assay with fluorescent cell segregation (60%) was attributed to the fact that each cell possesses multiple copies of the EGFP gene. The results clearly demonstrate the functionality, usefulness, and effectiveness of the match editing technique by tethering the RNA template with a complementary matching tag attached to magRNA. [Examples]
[0432] This example compared the editing effectiveness of long RNA templates (Ex57-P14) and short RNA templates (Ex29-P14) in correcting the A200G mutation (3 nucleotides downstream of the cutting site).
[0433] Cells were electroporated with plasmids expressing the indicated ME component or control component, and the cells were examined by fluorescence microscopy. The results are shown in Figure 19A (fluorescence and brightfield views are shown). The percentage of cells expressing fluorescent EGFP was determined by flow cytometry. The results are shown in Figure 19B (Panel 1, untreated cells; Panel 2, nfEGFP-expressing cells; Panels 4, 6, and 7, cells expressing ME components with different templates and magRNAs; Lane 3 and Lane 5, cells expressing gRNA instead of magRNA).
[0434] As shown in this figure, both templates were effective, but the shorter template consistently showed higher efficiency in correcting point mutations (Panel 4 vs. Panels 6-7 in Figures 19A and 19B). Furthermore, the efficiency of two different magRNAs matching either the central part of the template (12M29) or the 5' end (12M57) of the longer template (Ex57-P14) was compared. The data showed that the magRNA matching the internal sequence (12M29) was more effective than the magRNA matching the distal 5' end sequence (12M57) (Panel 7 vs. Panel 6). In most subsequent experiments, magRNA 12M29 matching the internal region of the RNA template was selected. [Examples]
[0435] This example demonstrates the effectiveness of ME in correcting deletion mutations by inserting the correct base pairs.
[0436] In short, cells were electroporated with plasmids expressing the indicated ME component (using 12M29 magRNA) or a control component, and tested by fluorescence microscopy. The percentage of cells expressing fluorescent EGFP was determined by flow cytometry. The results are shown in Figures 20A and 20B. Target EGFP DNA fragments were amplified from ME-treated cells, and the complementary strands were sequenced. DNA sequencing results from panel 3 cells are shown in Figure 20C.
[0437] As shown in these figures, using various correction RNA templates and 12M29 magRNA, ME effectively corrected deletion mutations 4, 20, and 35 nucleotides downstream of the nicking site using templates that covered these distances (Panels 3, 4, and 5 in Figures 20A and 20B). The correction of deletion mutations was confirmed by direct Sanger sequencing (Figure 20C). This study demonstrates that the ME system can be used for insertion editing. [Examples]
[0438] This example compared the effectiveness of two magRNAs in correcting deletions at positions 35 or 50 downstream of the nicking site using a long RNA template (Ex57-P14).
[0439] Cells were electroporated with either a plasmid expressing ME or a control plasmid, and the percentage of cells expressing fluorescent EGFP was determined by flow cytometry. The results are shown in Figure 21. In both cases, magRNA (12M29), which matches the proximal sequence of the RNA template (closer to the 3' priming sequence), was found to exhibit higher editing efficacy than magRNA (12M57), which matches the distal 5' end of the RNA template. In other experiments, magRNA 12M29 was similarly used in the majority of deletion / insertion editing experiments. [Examples]
[0440] This example demonstrates the effectiveness of ME in correcting deletion mutations located 100 base pairs downstream of the nicking site. In this experiment, an RNA template with a 111nt elongated sequence (Ex111-P14) and 12M29 magRNA were used.
[0441] Cells were electroporated with EGFP Δ100G plasmid and ME plasmid. Target EGFP DNA fragments were amplified from ME-treated cells, and the complementary strands were sequenced. The results are shown in Figure 22. The results clearly demonstrated the effectiveness of ME in editing distant mutations relative to the target site using an independent long RNA template. [Examples]
[0442] This example compared the effects of magRNA matching tag length and matching position in the template on ME editing efficiency.
[0443] In short, cells were electroporated with either a plasmid expressing ME or a control plasmid, and the percentage of cells expressing fluorescent EGFP was determined by flow cytometry. The results are shown in Figures 23A and 23B.
[0444] As indicated in Figure 23A, cells were electroporated using different ME systems with either EGFP-deleting plasmid Δ50C alone (EGFPΔ50) or EGFPΔ50 plus magRNA with matching lengths varying in 3nt increments from 9 nucleotides (9M) to 24 nucleotides (24M).
[0445] As indicated in Figure 23B, cells were electroporated either untreated (NT), with EGFP-deleting plasmid Δ35C alone (EGFPΔ35), or with EGFPΔ35 plus different ME systems (12M23, 12M29, 12M35, and 12M41, respectively) that had magRNAs with varying matching positions in the template at the first matching nucleotide at positions 23, 29, 35, or 41 from the RT elongation site. The matching tag length remained constant at 12nt due to the varying matching to complementary positions. The ME template used in both studies was Ex57-P14.
[0446] By varying the matching length from 9 nucleotides to 24 nucleotides, it was found that 9, 12, and 15 nucleotide matching tags showed the highest effectiveness, while longer matching tags of 18, 21, and 24 nucleotides had dramatically reduced effectiveness. See Figure 23A.
[0447] By varying the matching position from the priming start site to +23 (matching to a sequence between +12 and +23 (12nt)), +29, and from +35 to +41, it was found that the greatest effectiveness was achieved at the matching position in the template upstream of the extension start site at +29 (matching to the template from +18 to +29 from the first template extension nucleotide). See Figure 23B. [Examples]
[0448] This example demonstrates the effectiveness of ME components and configurations in which the reverse transcriptase is not directly fused with the RNA guide nickase, but is provided in a fragmented state by recruiting the MCP-RT fusion protein via the MS2 aptamer in the magRNA stem-loop. The experimental design and results are shown in Figures 24A to 24C.
[0449] As shown in Figure 24A, in the split RT ME system, the reverse transcriptase was not fused with nCas9. The magRNA contained the MS2 aptamer at the stem-loop position. RT was fused to MCP via a linker peptide. As shown in Figure 24B, the gene to be edited was the EGFP gene containing a C deletion 50 nucleotides upstream of the 200-L nicking site. As indicated, cells were electroporated with plasmids expressing split RT ME or fused RT ME, or with EGFPΔ50 alone, or left untreated. The percentage of cells expressing fluorescent EGFP was determined by flow cytometry. The ME template was Ex57-P14; the match tag was 12M29.
[0450] As shown in Figure 24C, ME demonstrated lower but comparable effectiveness as a direct fusion ME in correcting deletions 50 bp downstream of PAM. [Examples]
[0451] This embodiment demonstrates that a second ME module, which provided only nicking at the transformer position downstream of the target site of the first ME module, dramatically enhanced the editing effectiveness of the first ME module. The experimental design and results are shown in Figures 25A to 25C.
[0452] Figure 25A shows a specific dual-module ME system containing one template (Ex57-P14), a first magRNA of a first ME module targeting a first target site, and a second gRNA (and therefore gRNA, not magRNA) of a second ME module targeting a second target site but lacking a match tag. Functionally, the second ME module provides only a second nicking. Two second target sites, 119-U and 151-U (Figure 25B), were examined against the first target site 200-L, which were approximately 80 nt or 50 nt downstream of the first target site.
[0453] As indicated, cells were electroporated with plasmids expressing a single module or dual ME, or with EGFPΔ50 alone, or left untreated. The percentage of cells expressing fluorescent EGFP was determined by flow cytometry. The ME template was Ex57-P14, and the match tag was 12M29. As shown in Figure 25C, the results clearly demonstrated that a second nicking module dramatically increased editing efficacy, and modules 80 nt downstream showed better efficacy. [Examples]
[0454] This example demonstrates that match editing is effective in editing the endogenous gene locus HEK4 site. The experimental design and results are shown in Figures 26A to 26C.
[0455] In short, cells were electroporated using an ME system containing HEK4 target sites, either untreated (Figure 26A) or gRNA without a matching tag (Figure 26B) or magRNA with a suitable matching tag for the template (Figure 26C). Genomic DNA was extracted three days after electroporation. PCR-amplified target fragments were sequenced, and mutations were quantified using the EditR program.
[0456] As shown in Figures 26A-26C, ME produced a 10% G-to-A base change using magRNA targeting HEK4 and a template intended to change G to A at the +5 position (upstream nicking site) (Figure 26C). As a control, no editing was observed when gRNA was used instead of magRNA (Figure 26B). Figure 26A shows the HEK4 site sequence of untreated cells. In this figure, arrows indicate the editing rate in percentage units. It is important to note that, due to endogenous gene editing, editing with gRNA (lack of matching tag) showed no editing effect, while ME with magRNA showed a 10% editing efficiency, highlighting the crucial functional role of the matching tag for match editing. [Examples]
[0457] This embodiment demonstrated that ME is effective in deletion editing.
[0458] An EGFP reporter gene containing a 4-nucleotide frameshift insertion starting at position 180 was generated for this experiment. Cells were either untreated or electroporated with an EGFP expression vector containing only the 4-nucleotide frameshift insertion at position 180 (EGFP 180 Ins 4), or with an EGFP 180 Ins 4 plus match editing factor (Ex29M12, 12M29 magRNA, and Ex29-P14 template), or a prime editing factor (peg_Ex35). The percentage of fluorescent cells was determined by flow cytometry. The results are shown in Figure 27.
[0459] As shown in Figure 27, when cells were electroporated with the EGFP insertion construct, only minimal levels of fluorescent GFP were observed due to frameshift mutations. In contrast, when cells were electroporated with the construct concurrently with either the ME system or the prime editing system, removal of frameshift deletions and detection of cells expressing high levels of EGFP were detected. These results demonstrate the effectiveness of ME in deletion editing. [Examples]
[0460] This example demonstrates that both single-ME and dual-ME systems are effective in treating deletions of splicing regulatory sequences and splicing donor sites (SDS) that result in exon skipping. The dual-ME system with magRNA-gRNA pairing is more effective than the single-ME system. The experimental design and results are shown in Figures 28A and 28B.
[0461] A splicing reporter construct was generated by splitting the EGFP coding sequence into a 5' half gene and a 3' half gene linked by an artificial intron containing a splicing donor site (SDS) at the 3' end of the 5' half gene and a splicing acceptor site (SAS) at the 5' end of the 3' half gene. See Figure 28A, which shows a schematic diagram of the GFP-based splicing reporter. The splicing reporter contains a 5' half GFP and a 3' half GFP linked by an intron having the indicated SDS and SAS.
[0462] Further mouse Duchenne muscular dystrophy gene (Dmd) exon 23 skipping reporter constructs were generated by inserting the mDmd exon 23 and its intron SAS and SDS sites into the artificial intron sequence of the splicing reporter. As a result, the mDmd exon 23 skipping reporter construct produces a non-fluorescent protein, a protein having the amino acid sequence encoded by exon 23 between the 5' and 3' halves of the EGFP gene. Deletion of splicing regulatory sites such as exon 23 SDS would induce splicing skipping of exon 23, resulting in functional EGFP.
[0463] The ME single-module and ME dual-module systems were designed to delete a 29 bp sequence containing the exon 23 / intron 23 SDS region. Similar designs were performed using prime editing factors in both the single-module PE2 and dual-module PE3 configurations. The results are shown in Figure 28B.
[0464] As shown in Figure 28B, a single ME module produced 40% EGFP-expressing cells (indicated as magLow), and the addition of a second ME module (the second module contains only gRNA instead of magRNA, causing a second nicking) further increased the percentage of EGFP-producing cells to 70% (indicated as magLow+sgUp).
[0465] PE2 produced outcomes similar to a single-module ME system, while PE3 (which introduced a second nicking in addition to the first nicking caused by PE2) was more effective than PE2 but less effective than a dual-module ME. [Examples]
[0466] This example (Figure 29) compared the editing effectiveness of single ME, dual-ME with magRNA-gRNA pairing, and dual-ME with magRNA-magRNA in correcting deletion mutations when complexed with a long RNA template (Ex111-P14). The ME systems were similar to those illustrated in Figure 25, except for the last panel experiment in which the second gRNA also contained a second matching tag at its 3' end.
[0467] The results show cells that were either untreated or electroporated with a mutant EGFP (deletion of one nucleotide at position 50 relative to the nickeling site) expression vector (EGFP Δ50), or a mutant EGFP expression vector plus one module ME (magLow), or a dual module ME with magRNA-gRNA pairing (magLow+sgUp119), or a dual module ME with magRNA-magRNA pairing (magLow+magUp119) (all having RNA templates with an elongation length of 111 nt). In MagRNA-magRNA pairing, the match-tag sequences in the two magRNAs are complementary to two separate sequences in the RNA RT template.
[0468] The results showed that a dual-ME system with a second magRNA exhibited the highest editing efficacy, followed by a dual-ME system with a magRNA paired with a gRNA, and then a single-module ME system. [Examples]
[0469] This example demonstrates a dual ME system for editing the endogenous site, HEK3, in HEK293T cells. Importantly, this example also shows the distinct individual contributions of RNA templates and magRNAs to gene editing efficiency.
[0470] As shown in Figure 30A, templates intended to install two point mutations, 5G>T and 12G>C, were designed to target the HEK3 site. The template, T43-P13-(5GT-12GC), contains a 43nt RT extension sequence and a 13nt primer-binding sequence. Both point mutations are transversion mutations, with (5G>T) intended to disrupt the PAM motif to avoid repetitive, unproductive editing. Based on the template sequence, magRNA_12M42 was designed with an upper guide as shown in Figure 30A and a 12nt complementary to the 5' end of the T43 template. The expression vectors for the template, magRNA, second nicking gRNA, and guide are shown in Figure 30A. These plasmids and the nCas9(H840A)-RT fusion plasmid were introduced into HEK293 cells by electroporation. Genomic DNA was extracted from treated and untreated cells after 3 days. DNA fragments containing HEK3 were amplified by PCR and analyzed by Sanger sequencing. Figure 30B shows clear evidence of transversion editing. Furthermore, the editing efficiency was similar at 5G>T and 12G>C. The editing efficiency at 5G>T was used as the quantifiable value (Figure 30C).
[0471] To determine the individual contributions of RNA template and magRNA to editing efficiency, the amounts of RNA template and magRNA vectors were varied. As shown in Figure 30C, maintaining the magRNA vector at 250 ng while increasing the RNA template vector from 250 ng to 1000 ng dramatically increased editing efficiency. On the other hand, further increasing the magRNA vector from 250 ng to 500 ng had little effect on editing effectiveness. Consistently, when the RNA template vector was maintained at 250 ng, increasing the magRNA vector to 500 ng did not significantly affect editing efficiency (data not shown).
[0472] This embodiment demonstrates that a dual ME system is effective in installing transversion point mutations at endogenous HEK3 sites. Importantly, this embodiment demonstrates that RNA templates and gRNAs contribute differently to editing efficiency. Optimizing the ratio between RNA templates and gRNAs can optimize genome editing outcomes.
[0473] The ratio of 250 ng of magRNA vector, 1000 ng of template vector, and 250 ng of a second nicking gRNA vector yielded excellent editing outcomes in HEK293 cells; therefore, most of the subsequent studies by the applicants described below adopted this ratio unless otherwise specified. [Examples]
[0474] This example demonstrates a dual ME system for installing insertion mutations at the endogenous HEK3 site in HEK293 cells. The ME component design was similar to that of Example 14, except that the templates, T46, T50, and T79, contained 3nt, 7nt, and 36nt sequences, respectively, to introduce insertions at the +5 position, as shown in Figure 31A. The 3nt of CTT is a trinucleotide deleted in the cystic fibrosis gene CFTR delta 508 allele. The 7nt of ACTCAGT is the consensus binding site of the transcription factor AP-1. The 36nt encodes a peptide in the TCR variable region for recognizing the influenza virus epitope.
[0475] As shown in Figures 31B and 31C, the effectiveness of insertion editing at the HEK3 site was evident. In this experiment, 250 ng magRNA and 1000 ng RNA template vectors were used. Furthermore, an extra MS2 was inserted into the gRNA of the second nicking module, which was found to increase editing efficiency for reasons that are unclear. This example demonstrates that dual ME is effective in installing insertions at target endogenous genomic sites. [Examples]
[0476] This example demonstrates that polymerases other than MMLV-RT can be used in the ME system, and that DNA templates can be recruited by magRNA for match editing. A split dual ME system was constructed using the components illustrated in Figure 32A. An nfEGFP with the A200G mutation was used as the reporter (Figure 32B). The RNA or DNA template, T29-P14-3GA, was designed to correct for the mutation. The magRNA contained an MS2 aptamer in its stem-loop region to recruit the polymerase fused to the MCP. In addition to MMLV-RT, variants of R2Bm (retrotransposon RT) and Helraiser (polymerase for DNA transposon Heliton) were fused to the MCP. The transdual ME expression vector was introduced into HEK293 cells by electroporation, and the editing efficiency was determined by flow cytometry for the edited fluorescent EGFP-expressing cells. When a DNA template was used, a single-stranded DNA oligonucleotide was used instead of the expression vector.
[0477] As shown in Figure 32C, the experiment demonstrated that R2Bm and Helraiser also exhibited significant editing activity, albeit lower than MMLV-RT. Furthermore, when Helraiser was used, the DNA template was effective in facilitating editing, but the efficiency was significantly lower for both MMLV-RT and R2Bm when combined with the RNA template. [Examples]
[0478] This embodiment demonstrated that editing efficiency can be enhanced by recruiting additional effectors to the dual ME system. As shown in Figure 33A, the dual ME system was constructed to target the HEK3 site to install the two mutations previously described in Figure 30. In addition, a second nicking gRNA containing the MS2 aptamer was constructed to recruit the pioneer transcription factor, p65, fused to MCP. The ME system was introduced into HEK293T cells, and genome editing efficacy was determined by Sanger sequencing.
[0479] As quantified in Figure 33C, this example demonstrated that recruitment of p65 by a second gRNA to the dual ME system significantly increased editing efficiency. Interestingly, in the absence of MCP-p65, the presence of MS2 in the second gRNA alone also significantly increased editing efficiency. This phenome was observed in other editing sites and cell lines (data not shown herein). Although the mechanism is unclear, MS2 may have an effect in stabilizing the structure of the second CRISPR complex. [Examples]
[0480] This example compared the editing effectiveness of a split-dual ME system and its counterpart, a split-prime editing (PE) system, at the HEK3 site of HEK293T cells.
[0481] As shown in Figure 34A, a split dual ME system for installing two point mutations at the HEK3 site was constructed using the same template and magRNA as described in Figure 30. In this dual ME system, the MCP-RT fusion was recruited by a second nicking gRNA containing the MS2 aptamer, as shown in Figure 34C. The corresponding split PE3 was constructed using the ME magRNA and pegRNA containing the same guide and RT priming / extension sequences found in the RNA template, respectively, as shown in Figures 34B and 34C.
[0482] The BE system was introduced into cells using a 1000 ng RNA template vector and a 250 ng magRNA vector. In the same experiment, the corresponding PE system was introduced into HEK293 cells using a molar equivalent pegRNA vector (equivalent to the 1000 ng RNA template vector). The molar ratio of RNA template in the PE system: pegRNA scaffold in PE: RNA template in ME: magRNA scaffold in ME was roughly equal to 1:1:1:0.25. The editing efficiency was determined and illustrated in Figure 34D.
[0483] The split dual ME system demonstrated higher editing efficacy than its counterpart PE system for HEK3 site editing. To rule out the possibility that high pegRNA vector levels were unfavorable for the PE system, the pegRNA level was reduced to approximately 250 ng for dose setting. It was found that the PE system with a molar equivalent of pegRNA vector relative to a 1000 ng RNA template vector showed higher genome editing efficacy than lower levels of pegRNA (data not shown).
[0484] This embodiment demonstrates that the split dual ME system has higher editing effectiveness than its corresponding split PE3 system when editing HEK3 sites under its own preferred conditions. [Examples]
[0485] This example, illustrated in Figures 35-37, compares a dual ME system and its PE3 counterpart in the installation of five point mutations at the endogenous HEK3 site in HEK293T cells. In this example, editing efficiency was determined by Sanger sequencing and next-generation sequencing (NGS). Furthermore, unintended insertion formation at the target site was also analyzed using NGS data.
[0486] Figure 35 illustrates how to construct a dual ME targeting HEK3. The magRNA has the same guide and matching tag as described in Figure 30A. However, as illustrated in Figure 36A, the template, T43-P13-(5G>T-12G>C-18A>C-24T>C-30C>T), is different, intended to install five point mutations across a 43nt RT elongation sequence. Its counterpart, PE3, has a pegRNA containing the same guide and 3' elongation sequence as the ME RNA template. Common components of ME and PE include a second nicking gRNA and a sp-nCas9-RT fusion expression vector.
[0487] ME and PE systems with the same ratios as described in Example 18 (Figure 34) were separately introduced into HEK293T cells. The second gRNA concentration was also varied in both the ME and PE systems (250 ng vs. 500 ng). Three days after electroporation, genomic HEK3 DNA was amplified by PCR and analyzed by Sanger sequencing and NGS.
[0488] As shown in Figure 36B, genome editing by ME was evident in the installation of five point mutations. Furthermore, quantification of Sanger sequencing results in Figure 36C showed that the mutation patterns of the five mutations were similar under all conditions. The editing efficiency of the ME system was higher than that of the PE system under both second gRNA conditions (250 ng and 500 ng). The statistical significance between ME and PE editing efficiencies at the 5G>T editing site is shown in Figure 36D. The same statistical significance was observed between the ME and PE systems at the other four mutation sites (data not shown).
[0489] NGS of representative samples under these conditions provided similar but more precise images, as shown in Figure 37A. Of the five point mutations, editing efficiency decreased as the distance from the priming site increased, which is consistent with the editing mechanism. Nevertheless, the NGS results, as with conclusions drawn from Sanger sequencing analysis, showed that the ME system exhibited higher editing efficiency than the PE system under both conditions with varying amounts of the second gRNA.
[0490] One of the distinctive features of ME is that the modular ME template is inherently "unscarred," while the PE peg template mechanistically tends to incorporate the 3' end gRNA scaffold sequence into the edited end when reverse transcription follows into the gRNA scaffold. This has been experimentally demonstrated herein by NGS data. The inventors analyzed the insertion of edited reads at the ends of the template sequence and counted reads containing insertions of three or more nucleotide bases identical to the 3' end of the gRNA scaffold. As summarized in the table in Figure 37B, under conditions of 250 ng of a second gRNA vector, the dual ME system showed only background gRNA scaffold sequence insertions (same as untreated cells). However, the PE system generated approximately 350 times more reads containing insertions identical to the 3' end of the gRNA scaffold. The ratio of these reads to edited reads was approximately 350 times higher in the PE system compared to the ME system (Figure 37B). When the second gRNA vector was increased to 500 ng, the insertion rate in the ME system increased above the background level, which may be due to the increase in DSBs. Nevertheless, the ratio of these reads to edited reads was about 1 / 45th of that observed using its counterpart PE3 system under the same conditions (Figure 37B). The striking contrast between the ME and PE systems in the generation of unintended gRNA scaffold sequence insertions at the edges of the editing site is illustrated in Figure 37C.
[0491] This embodiment demonstrates that the described dual ME system is more effective than its counterpart PE3 system in the simultaneous installation of multiple point mutations, including both transversion and transition point mutations, at scattered locations within the HEK3 site. Furthermore, this embodiment directly and undeniably demonstrates that the ME and PE systems have a fundamental difference in “scar” formation at the edges of the edited site. This embodiment confirms that the PE system tends to incorporate the 3' end gRNA scaffold sequence at the edges of the edited site, whereas the ME system does not. [Examples]
[0492] This example compared a dual ME system and its PE3 counterpart in editing the endogenous HEK3 site by deletion or insertion of a 3nt sequence in HEK293T cells. A dual ME system targeting HEK3 was constructed in the same manner as previously described in Figure 30, except for a template containing a 3nt(CTT) insertion or 3nt(GCA) deletion as illustrated in Figure 38A. Its corresponding PE3 system was similarly constructed using the same template sequence covalently linked to the 3' end of the pegRNA.
[0493] The dual ME system or its counterpart PE3 was introduced into HEK293T cells, and their editing efficacy in installing deletions or insertions was analyzed and compared using Sanger sequencing. The results in Figure 38B demonstrate that the dual ME system exhibits higher editing efficacy than its counterpart PE3 system in both the installation of 3nt insertions and 3nt deletions at HEK3 sites. [Examples]
[0494] This example demonstrates the efficacy of the ME system in different mammalian cell types, K562 cells. This example also compares the editing efficiency of the dual ME system in installing nucleotide insertions in an episomal EGFP plasmid containing 1nt deletions at various positions relative to the EGFP nick site in G202, namely Δ4A, Δ20C, and Δ35C, with that of its counterpart, the PE3 system (Figure 17C). The dual ME system was constructed as described in Figure 25 (200L + 119U) (a similar experiment was performed in HEK293T cells in Figure 25). The counterpart PE3 system contains pegRNA with the same guide and template as the ME system. The ME or PE system, along with a 1nt deletion frameshift EGFP expression plasmid, was electroporated into K562 cells. Gene editing efficiency was quantified by flow cytometry, as shown in Figure 39, determined by the percentage of cells with fluorescent EGFP expression.
[0495] Figure 39 demonstrates that the dual ME system effectively installs single nucleotide insertions at various deletion sites in K562 cells. Furthermore, Figure 39 shows that the ME system is more effective than its PE counterpart in installing these insertions at the Δ20 and Δ35 positions. When installing insertions at the Δ4 position, the effectiveness of ME and PE was almost identical. [Examples]
[0496] This example demonstrates that the dual ME system effectively generates 29nt deletions in episomal plasmids in K562 cells. In addition, this example compares the editing efficiency of the dual ME system with its counterpart, the PE system. The components of ME and PE and the experiments are shown in Figure 28, where the same experiments were performed in HEK293T cells. This example replicates the previously performed experiments in HEK293T cells (Figure 28) in a different mammalian cell type, K562 cells. Editing efficiency was quantified by the percentage of GFP-positive cells (Figure 40A) and Sanger sequencing (Figure 40B). Figures 40A and 40B again demonstrate that the dual ME system is more effective than its counterpart, the PE3 system, in generating 29nt deletions in K562 cells. Similar observations in HEK293T cells are shown in Figure 28. [Examples]
[0497] This example compared a dual ME system and its PE3 counterpart in K562 cells for editing endogenous HEK3 sites by deletion or insertion of 3nt sequences. Both editing effectiveness and insertion rates induced by RT template extension to the gRNA scaffold were compared. The ME components were the same as those described in Figure 38, where the experiment was performed in HEK293T cells. The edited genomic DNA was subjected to Sanger sequencing and next-generation sequencing (NGS).
[0498] Figures 41A and 41B show that the effectiveness of CTT insertion and GCA deletion at the HEK3 site in K562 cells is higher when using the ME system than when using the ME system. Furthermore, Figure 41C shows that the PE system generates two orders of magnitude more gRNA scaffold template insertions at the ends of the HEK3 editing site than the ME system.
[0499] This embodiment demonstrated that the dual ME system effectively installed deletions and insertions in K562 cells. Furthermore, this embodiment demonstrated that the dual ME system was more effective than its counterpart, the PE3 system, in installing 3nt deletions and insertions at this HEK3 site. Importantly, this embodiment again directly and undeniably demonstrated that the ME and PE systems have fundamental differences in “scar” formation at the end of the edited site in different mammalian cell types. The PE system’s tendency to incorporate the 3’ end gRNA scaffold sequence at the edited end is approximately two orders of magnitude higher than that of the ME system. [Examples]
[0500] This embodiment demonstrates that the dual ME system is effective in installing mutations in the endogenous HBB (hemoglobin beta) gene site near the sickle cell anemia E6V mutation, a therapeutically relevant genomic site, in K562 cells. Figure 42A shows the HBB sequence near the E6V(GAG>GTG) mutation site, annotated with the intended +5G>T mutation encoded in the RNA template, along with the magRNA guide, second nick guide, and nicking site. The A base immediately 5' relative to +5G is mutated in sickle cell anemia E6V.
[0501] The dual ME system was electroporated into K562 cells, and the edits were analyzed using Sanger sequencing. Figure 42B demonstrates that the dual ME system can efficiently install +5G>T transversion base changes into the intrinsic therapeutic site of the HBB in K562 cells. [Examples]
[0502] This example investigated whether ME can use CRISPR complexes other than the spCas9 system. This example demonstrated that the Cas9 ortholog, saCas9, is effective when used in a dual ME system. Furthermore, this example demonstrated that the dual saCas9 ME system is more effective than its saPE3 counterpart in editing EGFP mutations.
[0503] Dual ME was constructed by Cas9 orthologue, nickase saCas9(N580A)-RT fusion. Therefore, magRNA and a second nicking gRNA were constructed using the saCas9 gRNA scaffold. The gene to be edited is EGFP with a single nucleotide deletion, Δ100G, at the EGFP coding position 100. Figure 43A illustrates the targeting sequence, saCas9 PAM, magRNA guide, and second nicking gRNA guide. The template was designed to install the missing G at position 100 and separate silencing mutations. For the PE3 system, the pegRNA uses the same guide as the magRNA, and the 3'RT extension and primer binding sequences are identical to the RNA template sequence of the ME.
[0504] ME and PE systems were electroporated into HEK293 cells along with an EGFP (Δ100G) expression vector. After 3 days, cells were subjected to flow cytometry assays to quantify the percentage of cells expressing fluorescent EGFP (Figures 43B-43C). EGFP plasmid DNA surrounding the deletion region was amplified by PCR and analyzed by Sanger sequencing (Figure 43D). Figure 43B shows clear editing by the saCas9 dual ME system in correcting the deletion mutation. Figures 43C and 43D demonstrate that the dual saCas9 ME is more effective than its saPE3 counterpart.
[0505] The above-mentioned examples and preferred embodiments should be construed as descriptive rather than limiting of the disclosure as defined by the claims. As readily apparent, numerous variations and combinations of the features described above can be utilized without departing from the disclosure as set forth in the claims. Such variations will not be considered departures from the scope of the disclosure, and all such variations are intended to be included within the following claims. All references cited herein are incorporated herein by reference in their entirety.
Claims
1. A gene editing complex for editing a target site in a target DNA molecule, (A) RNA guide nickas; (B) Reverse transcriptase; (C) (1) A template segment containing a template sequence complementary to the desired DNA sequence to be introduced into the target site, and (2) A priming segment complementary to the 3' end of the nickel strand of the target DNA molecule. RNA template molecules containing; and (D) (1) A guide or spacer sequence complementary to the target sequence on the target strand of the target DNA molecule, (2) An RNA scaffold capable of binding to the RNA guide nickase, and (3) A tethering tag sequence complementary to the segment in the RNA template. Matching gRNA (magRNA) molecules containing A gene editing complex that includes [the specified element].
2. The complex according to claim 1, wherein the reverse transcriptase is covalently linked to the RNA guide niccas.
3. The complex according to claim 1, wherein the magRNA further comprises a protein-binding motif capable of binding to an RNA-interacting protein, and the reverse transcriptase is linked to the RNA-interacting protein.
4. The complex according to any one of claims 1 to 3, wherein the tethering tag in the magRNA is not poly-N, where N is a repeating nucleotide A, C, U, or G.
5. The complex according to any one of claims 1 to 4, wherein the RNA template does not include any additional sequences other than the priming segment and the template segment.
6. The complex according to claim 1, wherein the RNA guide nickase is a nickase variant of a Cas protein, or a nickase variant of an IscB, IsrB, or TnpB family endonuclease protein encoded by a transposon.
7. The Cas protein is Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7 , Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cpf1 (Cas12a), C2c1 (C as12b), C2c3 (Cas12c), CasY (Cas12d), CasX (Cas12e), Cas14 (Cas12f), Cas Phi (Cas12j), Cas13a, Cas13b, Cas13c, Cas13d, Cas13x, CasF, CasG, CasH, Cs The composite according to claim 6, selected from y1, Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3 (or CasE), Cse4 (or CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1, Csx15, Csf1, Csf2, Csf3, Csf4, Cu1966 and its orthologues.
8. The complex according to claim 6, wherein the Cas protein is Cas9 and its ortholog.
9. The complex according to claim 6, wherein the nickas variant of the Cas protein is Streptococcus pyogenes nCas9 (H840A) or Staphylococcus aureus nCas9 (N580A).
10. The complex according to any one of claims 1 to 9, wherein the reverse transcriptase is a naturally occurring reverse transcriptase derived from a retrovirus or retrotransposon, or a variant thereof.
11. The complex according to claim 10, wherein the reverse transcriptase is selected from the group consisting of Moloney's mouse leukemia virus (M-MLV), human immunodeficiency virus (HIV) reverse transcriptase, avian sarcoma / leukemia virus (ASLV) reverse transcriptase, Rous sarcoma virus (RSV) reverse transcriptase, avian myeloblastosis virus (AMV) reverse transcriptase, avian erythroblastosis virus (AEV) helper virus MCA reverse transcriptase, avian myelocytomatosis virus MC29 helper virus MCA reverse transcriptase, avian reticuloendotheliopathy virus (REV-T) helper virus REV-A reverse transcriptase, avian sarcoma virus UR2 helper virus UR2AV reverse transcriptase, avian sarcoma virus Y73 helper virus YAV reverse transcriptase, Rous-associated virus (RAV) reverse transcriptase, myeloblastosis-associated virus (MAV) reverse transcriptase, Line 1 ORF2, R2Bm, and R2Ol.
12. The complex according to claim 10 or 11, wherein the reverse transcriptase is MMLV-RT or a variant thereof.
13. The complex according to any one of claims 1 to 12, wherein the tethering tag sequence is located at the 3' or 5' end of the magRNA.
14. The complex according to claim 13, wherein the tethering tag sequence is covalently linked to the magRNA via a polynucleotide linker.
15. The complex according to any one of claims 1 to 14, wherein the tethering tag sequence in the magRNA is approximately 6 to 24 nt in length.
16. A system for editing a target site in a target DNA molecule, (I) The first gene editing complex according to any one of claims 1 to 15, and (II) Second gene editing complex A system that includes this.
17. The second gene editing complex described above, (A) A second RNA guide nickas, and (B) A gRNA molecule comprising a second guide or spacer sequence and a second RNA scaffold capable of binding to the second RNA guide nickase. The system according to claim 16, including the system described in claim 16.
18. The second gene editing complex described above, (A) A second RNA guide nickase; and (B) (1) Second guide or spacer arrangement, (2) A second RNA scaffold capable of binding to the second RNA guide nickase, and (3) Second tethering tag array A second magRNA molecule containing The system according to claim 16, including the system described in claim 16.
19. The system according to claims 17 and 18, wherein the second gene editing complex comprises (C) a second reverse transcriptase.
20. The system according to any one of claims 16 to 19, wherein the guide sequence of the magRNA in the first gene editing complex and the second guide sequence of the gRNA or second magRNA in the second gene editing complex are complementary to two target sequences of the target DNA molecule.
21. The system according to any one of claims 18 to 20, wherein the second tethering tag sequence of the second magRNA is complementary to the second segment in the RNA template.
22. The system according to any one of claims 16 to 21, wherein the 5' end of the RNA template contains the same sequence as the 3' end of the strand of the target DNA molecule that is nicked by the second RNA guide nickase.
23. The system according to any one of claims 19 to 20, wherein the second gene editing complex further comprises (D) a second RNA template.
24. The system according to any one of claims 19 to 23, wherein the second reverse transcriptase is linked to the second RNA guide nickase.
25. The system according to any one of claims 19 to 23, wherein the second magRNA molecule or the gRNA molecule comprises a second protein-binding motif capable of binding to a second RNA-interacting protein, and the second reverse transcriptase is linked to the second RNA-interacting protein.
26. A method for modifying a target DNA molecule in a cell, comprising contacting the target DNA molecule with a gene editing complex according to any one of claims 1 to 15 or a system according to any one of claims 16 to 25.
27. The method according to claim 26, wherein the modification results in a point mutation in the target DNA molecule, an insertion in the target DNA molecule, a deletion in the target DNA molecule, or a combination thereof.
28. The method according to any one of claims 26 to 27, wherein the cells are selected from the group consisting of archaeal cells, bacterial cells, eukaryotic cells, eukaryotic unicellular organisms, somatic cells, germ cells, stem cells, plant cells, algal cells, animal cells, invertebrate cells, vertebrate cells, fish cells, frog cells, bird cells, mammalian cells, pig cells, bovine cells, goat cells, sheep cells, rodent cells, rat cells, mouse cells, non-human primate cells, and human cells.
29. The method according to any one of claims 26 to 28, wherein the cells are present in or derived from a human or non-human subject.
30. Genetically modified cells or their offspring obtained according to the method described in any one of claims 26 to 29.
31. The cell according to claim 30, selected from the group consisting of archaeal cells, bacterial cells, eukaryotic cells, eukaryotic unicellular organisms, somatic cells, germ cells, stem cells, plant cells, algal cells, animal cells, invertebrate cells, vertebrate cells, fish cells, frog cells, bird cells, mammalian cells, pig cells, bovine cells, goat cells, sheep cells, rodent cells, rat cells, mouse cells, non-human primate cells, and human cells.
32. Cells according to claims 30 to 31, selected from the group consisting of pluripotent stem cells (PSCs), adult stem cells (ASCs), fibroblasts, chondrocytes, keratinocytes, hepatocytes, pancreatic islet cells, and immune cells including T cells, dendritic cells (DCs), natural killer (NK) cells, and macrophages, derived from human or non-human subjects.
33. The magRNA according to claim 1, 3, 4, 13, 14, or 15.
34. An RNA complex comprising the magRNA described in claim 33 and an RNA template tethered to the magRNA described in claim 1, 5, 21, 22, or 23.
35. (i) The magRNA molecule according to claim 33, and (ii) RNA complex according to claim 34 A nucleic acid that codes for one or two of the following.
36. A vector comprising the nucleic acid described in claim 35.
37. (i) Packaging materials and (ii) The gene editing complex according to any one of claims 1 to 15, The system according to any one of claims 16 to 25, The cell according to claim 30, 31, or 32 The magRNA molecule according to claim 33, The RNA complex according to claim 34, The nucleic acid according to claim 35, and The vector according to claim 36 One, two or more of the above A kit that includes this.
38. (i) A pharmaceutically acceptable carrier, (ii) The gene editing complex according to any one of claims 1 to 15, The system according to any one of claims 16 to 25, The cell according to claim 30, 31, or 32 The magRNA molecule according to claim 33, The RNA complex according to claim 34, The nucleic acid according to claim 35, and The vector according to claim 36 One, two or more of the above A pharmaceutical composition containing the above.