Reverse transcription mediated gene editing of modular RNA templates anchored by tag gRNA and uses thereof

By using a gene editing complex containing RNA-guided nickase, reverse transcriptase, and magRNA, the limitations of template size and flanking editing in existing technologies have been overcome, enabling broader target site editing and efficient, precise genome editing.

CN121773208APending Publication Date: 2026-03-31RUTGERS THE STATE UNIV
View PDF 28 Cites 0 Cited by

Patent Information

Application Number
CN202480056339.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-03
Filing Date
2024-07-01
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing gene editing technologies, such as CRISPR base editors and leader editors, have limitations in template size and non-specificity of flanking editing when editing target sites. This results in some target sites being unusable or having low editing efficiency, and may introduce unwanted foreign genetic information.

Method used

The gene editing complex uses an RNA-guided nicking enzyme, reverse transcriptase, and matching gRNA (magRNA). The magRNA contains a template fragment and a promoter fragment. The reverse transcriptase integrates the RNA template sequence into the target DNA, avoiding double-strand DNA breaks and achieving precise editing.

Benefits of technology

It enables broader target site editability and efficient genome editing, reduces the risk of cancer, improves the accuracy and efficiency of editing, and reduces the occurrence of flanking editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121773208A_ABST
    Figure CN121773208A_ABST
Patent Text Reader

Abstract

The present disclosure relates to gene editing, related systems, and uses thereof. The gene editing is mediated by reverse transcription of a modular RNA template anchored by a complementary sequence in the modified guide RNA.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of the earlier filing date of U.S. Provisional Application No. 63 / 511,710, filed July 3, 2023, pursuant to 35 USC §119(e). The contents of that application are incorporated herein by reference in their entirety.

[0003] Government interests

[0004] This invention was completed with government funding granted by the U.S. Department of Defense under license number MD200088. The government holds certain rights to this invention.

[0005] Reference to electronic sequence listing

[0006] The contents of the electronic sequence list (096738.00780SeqList.xml, 162,826 bytes in size, created on June 28, 2024) are incorporated herein by reference in their entirety.

[0007] Invention Field

[0008] This disclosure relates to gene editing, related systems, and their uses. Background of the Invention

[0010] Targeted gene editing has been used for gene manipulation in eukaryotic cells, embryos, and animals. Sequence-specific nucleases, including Talen, zinc finger nucleases, and RNA-guided nucleases (such as CRISPR / Cas9), provide tools for precise genome editing. When natural nucleases bind to a target sequence, they create a DNA double-strand break (DSB), thereby initiating cellular DNA repair pathways, including non-homologous end joining (NHEJ) and homologous directed repair (HDR). This allows for the achievement of desired sequence alterations or gene editing.

[0011] During gene editing, on-target and off-target DNA DSB intermediates can lead to chromosomal translocations and other mutagenic events with potential oncogenic risks. To avoid the need for DNA DSBs, base editing platforms have been developed. CRISPR base editors utilize nuclease-inactivated or nicking enzyme forms of CRISPR proteins. These mutant CRISPR proteins can efficiently recognize DNA sequences without generating DSBs. Furthermore, the mutant CRISPR protein-gRNA complex recruits cytosine deaminases or adenine deaminases, converting C to U or A to G at the target site, respectively, resulting in sequence-specific point mutations. By avoiding DSBs, base editing reduces the risk of oncogenicity and is widely used in precision genome editing for basic research and the development of therapeutic drugs.

[0012] Because nucleotide deamination is the fundamental mechanism of base alteration, base editors can edit transition point mutations but not transversion mutations. Furthermore, base editors require the target nucleotide to be located within an R-loop adjacent to the PAM motif. This requirement excludes the accessibility of target sites without an adjacent PAM motif. Moreover, deamination is typically nonspecific within the editing activity window, leading to flanking base editing within the window. While flanking base editing may be harmless in some therapeutic developments, such as correcting loss-of-function mutations, precise editing without flanking editing is preferred in many other cases.

[0013] Prime Editing is another precise gene editing platform that does not require double-strand DNA breaks (Anzalone, AV, et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149–157 (2019)). The PrimeEditor complex contains a CRISPR protein in the form of a nicking enzyme fused to reverse transcriptase (RT) and a modified gRNA called pegRNA, which contains an RNA template and a primer-binding sequence, typically located at the 3' end of the gRNA. Upon binding to the target DNA sequence, the PrimeEditor creates a DNA nick. The nicked DNA strand acts as a primer, binding to the primer-binding sequence within the pegRNA and using the RNA template within the pegRNA to synthesize a new DNA sequence. The reverse transcriptase then copies the sequence information of the template RNA from the RNA to the DNA. Subsequently, this DNA sequence is further integrated into the target site. By replicating the desired RNA template at the 3' end of the pegRNA, PrimeEditor enables precise editing of genomic base pairs without side-side effects. Base changes can be either transition or transversion. Furthermore, by designing insertions and deletions in the pegRNA template, PrimeEditor can also generate insertions and deletions at the target location.

[0014] The core of the lead editing system is a modified gRNA, called pegRNA, which contains the gRNA scaffold, primer binding sequence, and editing template within the same RNA molecule. However, this configuration limits the size of the editing template because long RNA template sequences within the same gRNA molecule can form secondary structures that may, for example, interfere with the secondary structure of the gRNA-CRISPR protein binding scaffold. (Anzalone, AV) et al. NatureThe longest template length among pegRNA molecules attempted in 576, 149–157 (2019) was 34 nt. Template size limitations prevent some target sites in the genome from being utilized by the leader editor due to the lack of suitable nearby PAM motifs. Even when the PAM motif is within the limits of RT (reverse transcriptase) template size, the template size limitation can sometimes restrict RT template selectivity because adjacent sequences to the PAM motif may have suboptimal GC content, resulting in some target sites in the genome being unsuitable for leader editing due to difficulties in initiation and elongation.

[0015] One variant configuration of the leader editor splits the pegRNA into two molecules: unmodified gRNA and an RT template with a foreign accessory RNA aptamer structure, such as MS2 (WO2020 / 191248A1). In this configuration, the CRISPR protein needs to include an additional fusion chaperone to interact with the RNA aptamer within the RT template; for example, the MCP protein is fused with the Cas9 protein to recruit the RT template containing the MS2 RNA aptamer. In this system, the gRNA is responsible for complexing with the CRISPR protein to recognize the target site. The RNA aptamer (e.g., MS2) is then recruited to the CRISPR complex via an additional fusion moiety (e.g., MCP). The leader editor with this configuration is several orders of magnitude less efficient than its pegRNA leader editor counterpart (see Figure 73 of WO2020 / 191248A1). Furthermore, its unwanted insertion / deletion to correct editing ratio is higher than that of the pegRNA leader editor. WO2020 / 191248A1 also indicates that adding accessory RNA aptamers to the RT template, especially at the 5' end, may result in the copying of unwanted exogenous genetic information (RNA aptamers) to the target site.

[0016] More versatile and efficient systems are needed to edit target nucleic acid molecules. Summary of the Invention

[0017] This disclosure addresses the aforementioned needs in several ways.

[0018] On one hand, this disclosure provides a gene editing complex for editing target sites in a target DNA molecule. The gene editing complex comprises: (A) an RNA-guided nicking enzyme; (B) a reverse transcriptase; (C) an RNA template molecule; and (D) a matching gRNA (magRNA) molecule. The RNA template molecule comprises: (1) a template fragment containing a template sequence complementary to the target DNA sequence to be introduced into the target site or its complementary sequence; and (2) a priming segment complementary to the 3' end of the nicking strand of the target DNA molecule. The magRNA molecule comprises: (1) a guide sequence or spacer sequence complementary to a sequence on the target strand of the target DNA molecule; (2) an RNA scaffold capable of binding to the RNA-guided nicking enzyme; and (3) an anchoring tag sequence complementary to a fragment in the RNA template.

[0019] In this application, anchor tags and matching tags are used interchangeably. Both refer to tag sequences added to the gRNA that complementarily match regions in the modular RNA template for recruiting and anchoring the RNA template.

[0020] Reverse transcriptase can be introduced into the complex in any suitable manner. In one embodiment, the reverse transcriptase is linked (covalently or non-covalently) or fused with an RNA-guided nicking enzyme. In another embodiment, the magRNA also contains a protein-binding motif (e.g., MS2 or PP7) capable of binding to an RNA-interacting protein, and the reverse transcriptase is linked (covalently or non-covalently) or fused with an RNA-interacting protein (e.g., MCP or PCP).

[0021] In one implementation, the anchor tag in the magRNA is not polyN, where N is a repeating nucleotide A, C, U, or G.

[0022] In one implementation, the RNA template does not contain any additional sequences other than the initiation fragment and the template fragment.

[0023] In one embodiment, the RNA-guided nickase is a nickase variant of the Cas protein, or a nickase variant of an RNA-guided endonuclease protein encoded by a transposon. In one embodiment, the Cas protein is selected from Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cpf1 (Cas12a), C2c1 (Cas12b), C2c3 (Cas12c), CasY (Cas12d), CasX (Cas12e), Cas14 (Cas12f), CasPhi (Cas12j), Cas13a, Cas13b, Cas13c, Cas13d, Cas13x, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3 (or CasE), Cse4 (or CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1, Csx15, Csf1, Csf2, Csf3, Csf4, Cu1966, and their orthologs. In one embodiment, the transposon encodes an endonuclease protein that is an IscB, IsrB, or TnpB endonuclease protein or its ortholog.

[0024] In one embodiment, the RNA-guided nickase is a nickase variant of an ortholog of the Cas protein, or a nickase variant of an RNA-guided endonuclease protein encoded by the transposon described above. In one embodiment, the Cas protein is Cas9. In one embodiment, a nickase variant of the Cas9 ortholog is *Streptococcus pyogenes* (…). Streptococcus pyogenes, sp) nCas9(H840A) or Staphylococcus aureus ( Staphylococcus aureus, sa) nCas9 (N580A). In one embodiment, the cleavage enzyme variant of the Cas9 ortholog is Staphylococcus aureus nCas9 (N580A E782K / N968K / R1015H).

[0025] In one embodiment, the reverse transcriptase is a naturally occurring reverse transcriptase derived from a retrovirus or a retrotransposon, or a variant thereof. In one embodiment, the reverse transcriptase is selected from Moloney murine leukosis virus (M-MLV), human immunodeficiency virus (HIV) reverse transcriptase, avian sarcoma leukosis virus (ASLV) reverse transcriptase, Rous sarcoma virus (RSV) reverse transcriptase, avian myeloblastosis virus (AMV) reverse transcriptase, avian erythroblastosis virus (AEV) helper virus MCAV reverse transcriptase, avian myelomavirus MC29 helper virus MCAV reverse transcriptase, avian reticuloendotheliosis virus (REV-T) helper virus REV-A reverse transcriptase, avian sarcoma virus UR2 helper virus UR2AV reverse transcriptase, avian sarcoma virus Y73 helper virus YAV reverse transcriptase, Rous-associated virus (RAV) reverse transcriptase, and myeloblastosis-associated virus (MAV) reverse transcriptase, Line 1 ORF2, R2Bm, and R2Ol. In one implementation, the reverse transcriptase is MMLV-RT or a variant thereof.

[0026] In one embodiment, the retrovirus and the retrotransposon-derived reverse transcriptase have inherent DNA-dependent DNA polymerase activity, such as Line 1 ORF 2 and HIV RT. In one embodiment, the retrotransposon-derived reverse transcriptase is R2Bm.

[0027] In one embodiment, the anchor tag sequence may be located at the 3' or 5' end of the magRNA. In one embodiment, the anchor tag sequence is covalently linked to the magRNA via a polynucleotide linker. In one embodiment, the anchor tag sequence in the magRNA is 6-24 nucleotides in length.

[0028] In one embodiment, the anchor tag sequence of the magRNA is complementary to the 5' end of the RNA template molecule. In another embodiment, the anchor tag sequence of the magRNA is complementary to an internal region of the RNA template.

[0029] Secondly, this disclosure relates to a system for editing target sites in a target DNA molecule. The system comprises: (I) any of the above-described first gene editing complexes; and (II) a second gene editing complex.

[0030] In one embodiment, the second gene editing complex comprises (A) a second RNA-guided nicking enzyme and (B) a gRNA molecule containing a second guide sequence or spacer sequence, the gRNA molecule may or may not contain an RNA aptamer (e.g., MS2).

[0031] In one embodiment, the second gene editing complex comprises (A) a second RNA-guided nicking enzyme; and (B) a second magRNA molecule comprising (1) a second guide or spacer sequence, (2) a second RNA scaffold capable of binding to the second RNA-guided nicking enzyme, and (3) a second anchoring tag sequence.

[0032] In one implementation, the second gene editing complex comprises (C) a second reverse transcriptase.

[0033] In one implementation, the two guide sequences (i.e., in the two gene editing complexes) , The guide sequence of the magRNA in the first gene editing complex and the guide sequence of the gRNA or the second magRNA in the second gene editing complex are complementary to the two target sequences of the target DNA molecule, respectively.

[0034] In one implementation, the second anchor tag sequence of the second magRNA is complementary to a second fragment within the RNA template.

[0035] In one embodiment, the 5' end of the RNA template contains the same sequence as the 3' end of the strand of the target DNA molecule cleaved by a nicking enzyme guided by a second RNA.

[0036] In one implementation, the second gene editing complex further comprises (C) a second reverse transcriptase, or (D) a second RNA template, or both.

[0037] The second reverse transcriptase can be introduced into the second gene editing complex or system in any suitable manner. In one embodiment, the second reverse transcriptase can be (covalently or non-covalently) linked or fused to a second RNA-guided nicking enzyme. In one embodiment, the second magRNA molecule or gRNA molecule contains a second protein-binding motif capable of binding to a second RNA-interacting protein, and the second reverse transcriptase is (covalently or non-covalently) linked or fused to the second RNA-interacting protein.

[0038] In one embodiment, the second gRNA or magRNA in the second gene editing complex comprises an RNA aptamer (e.g., MS2 or PP7) that recruits a non-reverse transcriptase effector fused to an RNA aptamer-binding protein (e.g., MCP or PCP). In one embodiment, the non-reverse transcriptase effector is a 50 kDa RNase inhibitor 1 protein (RNH1) or its ortholog. In one embodiment, the non-reverse transcriptase effector is a dominant-negative protein of the MMR pathway (including MLH1 protein) or a 5' DNA nuclease Fen1 protein.

[0039] In one implementation, the non-reverse transcriptase effector is a pioneer transcription factor. In one implementation, the pioneer transcription factor is p65.

[0040] In one embodiment, the non-reverse transcriptase effector is a chromatin regulatory protein or a protein complex. In one embodiment, the chromatin regulatory protein has histone acetylation activity.

[0041] This disclosure also provides a method for modifying intracellular target DNA molecules. The method includes contacting the target DNA molecule with the aforementioned gene-editing complex or system. In one embodiment, the modification results in a point mutation, insertion into the target DNA molecule, deletion of the target DNA molecule, or a combination thereof. In one embodiment, the cell is selected from archaea cells, bacterial cells, eukaryotic cells, eukaryotic unicellular organisms, somatic cells, germ cells, stem cells, plant cells, algal cells, animal cells, invertebrate cells, vertebrate cells, fish cells, frog cells, bird cells, mammalian cells, pig cells, bovine cells, goat cells, sheep cells, rodent cells, rat cells, mouse cells, non-human primate cells, and human cells. In one embodiment, the cell is located in, isolated from, or derived from a human or non-human subject.

[0042] On the other hand, this disclosure provides a genetically engineered cell or its progeny obtained according to the above method. In one embodiment, the cell is selected from archaea cells, bacterial cells, eukaryotic cells, eukaryotic unicellular organisms, somatic cells, germ cells, stem cells, plant cells, algal cells, animal cells, invertebrate cells, vertebrate cells, fish cells, frog cells, bird cells, mammalian cells, pig cells, bovine cells, goat cells, sheep cells, rodent cells, rat cells, mouse cells, non-human primate cells, and human cells. In one embodiment, the cell is selected from pluripotent stem cells (PSCs), adult stem cells (ASCs), hematopoietic stem cells (HSCs), fibroblasts, chondrocytes, keratinocytes, hepatocytes, pancreatic islet cells, and immune cells, including T cells, dendritic cells (DCs), natural killer (NK) cells, and macrophages, isolated or derived from human or non-human subjects.

[0043] On the other hand, this disclosure relates to magRNA as described above. The magRNA comprises: (1) a guide sequence or spacer sequence complementary to a target sequence on the target strand of a target DNA molecule; (2) an RNA scaffold capable of binding to an RNA-guided nicking enzyme; and (3) an anchor tag sequence complementary to a fragment in an RNA template. The magRNA may also comprise a protein-binding motif capable of binding to an RNA-interacting protein, and a reverse transcriptase is linked to the RNA-interacting protein. In one embodiment, the anchor tag in the magRNA is not a polyN, where N is a repeating nucleotide A, C, U, or G. In one embodiment, the anchor tag sequence is located at the 3' or 5' end of the magRNA. In one embodiment, the anchor tag sequence is covalently linked to the magRNA via a polynucleotide linker. In one embodiment, the anchor tag sequence in the magRNA is approximately 6-24 nucleotides in length.

[0044] On the other hand, this disclosure relates to an RNA complex comprising the aforementioned magRNA and the RNA template anchored by the aforementioned magRNA.

[0045] The anchored RNA template contains no other sequences besides those designed to be present in the target sequence for genome editing.

[0046] On the other hand, this disclosure provides a nucleic acid that encodes one or both of the following: (i) the above-described magRNA molecule, and (ii) the above-described RNA complex.

[0047] On the other hand, this disclosure relates to vectors containing the said nucleic acid.

[0048] On the other hand, this disclosure relates to a kit comprising (i) packaging materials, and (ii) one, two or more of the following: the gene editing complex described above, the system described above, the cells described above, the magRNA molecule described above, the RNA complex described above, the nucleic acid described above, and the vector described above.

[0049] On the other hand, this disclosure provides a pharmaceutical composition comprising: (i) a pharmaceutically acceptable carrier, and (ii) one, two or more of the following: the gene editing complex, the system, the cell, the magRNA molecule, the RNA complex, the nucleic acid and the carrier.

[0050] The following specification describes in detail one or more embodiments of this disclosure. Other features, objects, and advantages of this disclosure will be apparent from the specification and claims.

[0051] Brief description of the attached figures

[0052] Figure 1Images A, 1B, and 1C illustrate a single-module matching editing system for RNA-guided sequence recognition using CRISPR / gRNA. The system comprises: ( Figure 1 A) A modular RNA template that does not contain any exogenous RNA recruitment elements, such as RNA aptamers (e.g., MS2 or PP7); Figure 1 B) magRNA, which contains a 5' spacer sequence for target DNA recognition, an RNA scaffold for CRISPR / Cas complexation, and a 3' matching RNA sequence complementary to a region in the modular template; and ( Figure 1 C) The nicking enzyme CRISPR protein nCRISPR / nCas9 (H840A) is fused to reverse transcriptase (MMLV-RT) via a polypeptide linker.

[0053] Figure 2 This demonstrates the working principle of the matching editing system. nCRISPR / nCas9-RT forms a complex with magRNA. The 5' spacer sequence of the magRNA recognizes and complements the target sequence in the target DNA; simultaneously, the 3' tag sequence of the magRNA is complementary to a region in the RNA template, thereby anchoring the RNA template to the CRISPR-RT / magRNA complex. The nicking enzyme activity of nCRISPR / nCas9 causes a single-strand DNA break (nick) upstream of the PAM motif. The nicked DNA binds to the 3' end of the RNA template through sequence complementarity, further strengthening the complex. The free 3' strand of the nicked DNA then serves as a primer for reverse transcriptase, using the complexed RNA as a nucleotide extension template.

[0054] Figure 3 This demonstrates how cellular repair mechanisms introduce newly synthesized DNA strands into target DNA loci. Figure 3 A shows the relative positions of the RNA-guided nick site (indicated by an asterisk in the DNA strand and labeled "nick") and the target mutation site (indicated by a complementary arrow in the DNA strand and labeled "mutation"). Figure 3 BD shows the process of synthesizing a new DNA strand that extends the 3' nick chain, followed by the removal of the wobbling 5' original DNA fragment, resulting in a new upper strand with the target sequence alteration and nick. Figure 3 E shows how cellular repair mechanisms resolve mismatches and fill nicks, ultimately producing DNA molecules with the target edit (by removing mismatches on the lower strand) or DNA molecules with the original sequence (by removing mismatches on the upper strand).

[0055] Figure 4 An example of magRNA-RNA template interaction is shown, in which the 3' end matching tag sequence of the magRNA is specifically complementary to the 5' end sequence of the RNA template.

[0056] Figure 5 The displayed magRNA has a 5' matching tag that is complementary to the RNA template region and located at the 5' end of the gRNA. The linker polynucleotide sequence lies between the 5' matching tag sequence and the spacer sequence. Its editing mechanism is similar to... Figure 2 The difference lies in the interaction position between the magRNA tag sequence and the template RNA.

[0057] Figure 6 A, 6B, 6C, and 6D show a variant of a matching editing system in which reverse transcriptase (RT) is provided in the form of splitting rather than directly fusing with an RNA-guided sequence-specific nicking enzyme.

[0058] Figure 6 A shows a modular RNA template.

[0059] Figure 6 B shows the engineered magRNA. This magRNA contains a 5' spacer sequence for target sequence recognition and a 3' tag sequence for recruiting RNA templates. Figure 6 A) and RNA aptamers (e.g., MS2) located in the stem-loop region (SL) of the gRNA for recruiting reverse transcriptase independently.

[0060] Figure 6 C shows an RNA-guided nicking enzyme that does not contain a directly linked reverse transcriptase.

[0061] Figure 6 The protein shown in D contains a reverse transcriptase fused with an associated protein (e.g., MCP) via a polypeptide linker, which binds to an aptamer (e.g., MS2) within the gRNA scaffold.

[0062] Figure 7 A, 7B, 7C, and 7D illustrate the working principle of reverse transcriptase in the splitting system.

[0063] Figure 7 A shows a modular RNA template.

[0064] Figure 7 B shows engineered magRNA.

[0065] Figure 7 C shows an RNA-guided nicking enzyme.

[0066] Figure 7 The protein shown in D contains a reverse transcriptase and an associated protein (e.g., MCP) fused via a polypeptide linker, which binds to an aptamer within a gRNA scaffold (e.g., MS2). This process is related to... Figure 2The description is similar to that in [the previous text], except that RT is raised in a split format.

[0067] Figure 8 A variant configuration of RT in the splitting system is shown, in which the 3' magRNA extension sequence is sequence-specifically complementary to the 5' end of the RT template.

[0068] Figure 9 A and 9B show examples of dual-ME systems.

[0069] Figure 9 A illustrates a dual ME system that fuses a RT and features magRNA-gRNA pairing. In this configuration, the first ME has classic ME components, including the RT fused with nCRISPR / nCas9 and the magRNA. The second module is used to generate a second cut to improve editing efficiency. The gRNA in the second module does not have a 3' tag sequence for interacting with the RNA template.

[0070] Figure 9 B illustrates a dual-ME system comprising a trans-RT and a magRNA with MS2-gRNA 0xMS2. In this configuration, the first ME contains the magRNA with an aptamer for recruiting the split form of reverse transcriptase. The second module is used to generate a second nick to improve editing efficiency. The gRNA of the second ME module does not contain a 3' matching tag for RNA template recruitment.

[0071] Figure 10 The results show that the second cut on the adjacent relative chain mediated by the dual ME system significantly improves target editing efficiency. Figure 10 A to 10C and Figure 3 A through 3C are the same. Figure 10 The D pattern indicates a second nick site on the lower strand, which facilitates the removal of DNA flaps containing free 5' nick ends rather than those containing free 3' nick ends. Therefore, it is detrimental to the recovery of the original DNA sequence. Figure 10 E, on the right, marked with a stop sign), and the product is mainly a DNA molecule with the target editing sequence ( Figure 10 E, left side).

[0072] Figure 11 A and 11B show a single-template coupled dual-ME system for gene editing that uses large DNA fragments. Figure 11 A shows a coupled dual ME system with fused RT. Figure 11B illustrates a coupled dual-ME system with a split RT configuration. This configuration provides a large reverse transcriptase RNA template for editing the target sequence at a location remote from the RNA-guided nick site generated by the first ME module. In this configuration, the second ME module contains magRNA with a 3' tag sequence that interacts with the RNA template at a different fragment. Figure 11 The difference between 11A and 11B is that in 11A, the reverse transcriptase is covalently linked to nCRISPR / nCas9, while in 11B, the reverse transcriptase is provided separately.

[0073] Figure 12 A and 12B show exemplary versions of coupled dual-ME systems with a second start-up mechanism.

[0074] Figure 12 A shows an RT fusion version in which the 5' end of the RNA template is designed to be identical to the 3' end of the second nick strand. Therefore, the second nick strand can be used as a primer to synthesize a second DNA strand using the newly synthesized DNA strand as a template.

[0075] Figure 12 B shows the RT-split version.

[0076] Figure 13 A scheme for large fragment insertion using a single-template coupled dual-ME system is shown, which involves second-strand DNA-dependent DNA synthesis. Figure 13 A shows the first RNA-guided cut generated by the first ME module. Figure 13 B shows that the first ME module synthesizes the first DNA strand using a modular RNA template. Figure 13 CD shows that the second ME module introduces a second cut. Figure 13 E shows that the free 3' end of the second nick strand is initiated using the newly synthesized DNA strand. Figure 13 F represents the strand extension and synthesis of the second DNA strand via DNA-dependent DNA polymerase activity through reverse transcriptase or via cellular DNA-dependent DNA polymerase. Figure 13 GH demonstrated that the insertion of a newly synthesized DNA fragment was accomplished by removing two DNA fragments containing a free 5' end and then ligating them through cellular repair mechanisms.

[0077] Figure 14 A dual-template-dual-module system for insertion and / or deletion is shown. In this configuration, both ME modules contain magRNA and an RNA template. The reverse transcriptase can be provided in fusion with nCRISPR / nCas9 or as a split.

[0078] Figure 15A, 15B, 15C, 15D, 15E, and 15F show how a dual-template, dual-module ME leads to de novo gene construction. de novo Gene composition (insertion) or large gene deletion.

[0079] Figure 15 A and 15B show that modules 1 (ME) and 2 (ME) mediate nicks at their respective target sites and generate first-strand DNA synthesis via reverse transcription.

[0080] Figure 15 C shows that the 3' ends of the two newly synthesized first DNA strands annealed through complementary sequences.

[0081] Figure 15 D shows the synthesis of second-strand DNA via reverse transcriptase DNA-dependent DNA polymerase activity or endogenous cellular DNA polymerase activity.

[0082] Figure 15 E and 15F show the removal and ligation of the 5' fragment. This mechanism can produce insertions and deletions (de novo gene construction) depending on the composition and length of the two RNA templates and the sequence being removed.

[0083] Figure 16 AD shows an ME system with a second effector.

[0084] Figure 16 A shows a modular RNA template.

[0085] Figure 16 The magRNA shown in B has a 5' spacer sequence for target site recognition, a 3' tag sequence for reverse transcription RNA template recruitment, and an aptamer sequence (e.g., MS2) at the stem-loop position for recruiting second effectors.

[0086] Figure 16 C demonstrates the fusion of RNA-guided nickases (e.g., nCRISPR / nCas9) with reverse transcriptases (e.g., MMLV-RT) via adaptor peptides.

[0087] Figure 16 D shows that the second effector protein (e.g., an RNase inhibitor or FEN1) fuses with an aptamer-binding protein (e.g., MCP) via a peptide linker.

[0088] Figure 17 A, 17B, 17C, and 17D show the components of a matching editing system used to correct the A200G point mutation in the nfEGFP gene.

[0089] Figure 17A shows the components of the matching editing system and the expression plasmids for the gRNA control.

[0090] Figure 17 B shows the nfEGFP gene to be edited.

[0091] Figure 17 C shows the sequence of the target incision site (SEQ ID NO: 1 and 2), where PAM (underlined) and incision location (arrow) are shown, and the mutation G200 is marked with a box.

[0092] Figure 17 D shows the matching edit template, which has a primer binding sequence (P14, SEQ ID NO:3) and an RT extension of 29 nt or 57 nt in length.

[0093] Figure 18 A, 18B, and 18C show the matching editing system's... Figure 17 The effect of correcting point mutations is shown.

[0094] Figure 18 A shows the results of correcting point mutations using a matched editing system derived from functional assays.

[0095] Figure 18 B and 18C respectively show Figure 18 Sanger sequencing results for lanes 2 and 4 in A (SEQ ID NO:4). Lanes 3 and 4 used shorter templates with a 29 nt RT extension (Ex29) and either gRNA or magRNA, as shown. Lanes 5 and 6 used longer templates and either gRNA or magRNA, as shown. The magRNA tag sequence was 12 nt in length and matched the 5' end sequence of the corresponding template. PEG RNA-mediated leader editing (lane 7) was used for comparison.

[0096] Figure 19 A and 19B show a comparison of the editing efficiency of magRNAs that match the internal sequence (12M29) and 5' sequence (12M57) of the long template Ex57.

[0097] Figure 19 A shows the results of fluorescence microscopy.

[0098] Figure 19 B shows the percentage of cells expressing fluorescent EGFP, as determined by flow cytometry. Figure 19 A, numbers 1 to 7, and Figure 19 The corresponding parts in B are the same. The processing methods for each part are as follows: Figure 19 All of these are explained in section B.

[0099] Figure 20A, 20B, and 20C demonstrate the efficiency of ME in correcting deletion mutations. Figure 20 A, numbers 1 to 5, and Figure 20 The corresponding parts in B are the same.

[0100] Figure 20 A shows the results of fluorescence microscopy.

[0101] Figure 20 B shows the percentage of cells expressing fluorescent EGFP as determined by flow cytometry. 1 represents untreated cells; 2 represents cells electroporated with an equimolar mixture of EGFP deletion mutants (Δ4A, Δ20C, and Δ35C); 3, 4, and 5 represent cells expressing the ME component with the EGFP deletion, template, and magRNA (12M29) shown.

[0102] Figure 20 C shows the DNA sequencing results from cell 3 (SEQ ID NO: 5).

[0103] Figure 21 The editing efficiency of magRNA matching the inner sequence (12M29) and 5' end sequence (12M57) of the long template Ex57 is compared in insertion editing. 1 is untreated cells; 2 is cells electroporated with an equal mixture of EGFP-deleted plasmids Δ35C and Δ50C; 3-6 are cells expressing ME components with the EGFP plasmid, template, and magRNA shown.

[0104] Figure 22 The ME insertion editing rate at a position 100 nucleotides from the target cleavage site (SEQ ID NO: 6) is shown using a template with a 111 nt RT extension length and magRNA that matches the internal sequence (12M29).

[0105] Figure 23 A comparison was made of the matching editing efficiency of magRNAs with different matching lengths (9nt to 24nt) in correcting EGPF gene deletion mutations.

[0106] Figure 23 B compared the matching editing efficiency of magRNAs with the same length (12 nt) but different matching positions (3' tag position is complementary to the position on the template 23 nt to 41 nt upstream of the RT start site (5')) in correcting EGFP gene deletion mutations.

[0107] Figure 24 A, 24B, and 24C show the Trans RT matching editing system, its components, and its efficiency in correcting EGFP deletions.

[0108] Figure 24 A shows the Trans RT ME system and its components, in which the reverse transcriptase is not fused with nCas9. The magRNA contains the MS2 aptamer at the stem-loop position. The RT is fused with the MCP via an adaptor peptide.

[0109] Figure 24 B shows a gene to be edited, namely the EGFP gene, which has a C deletion 50 nucleotides upstream of the 200-L nick site.

[0110] Figure 24 C demonstrates the efficiency of the Trans RT matching editing system in correcting missing parts and compares it with the Direct Fusion ME system.

[0111] Figure 25 A, 25B, and 25C show the dual ME system with magRNA-gRNA pairing and its editing efficiency.

[0112] Figure 25 A shows the components of the dual-ME system.

[0113] Figure 25 B shows the two second module sites tested (SEQ ID No: 7 and 8). The second module ME contains gRNA instead of magRNA.

[0114] Figure 25 C demonstrates the editing efficiency of single-module ME and dual-module ME.

[0115] Figure 26 A, 26B, and 26C show that the matching edit efficiently changed G to A at the endogenous site HEK4.

[0116] Figure 26 A shows the results for untreated cells (SEQ ID NO: 9).

[0117] Figure 26 B shows the results of electroporation of cells using the ME system with HEK4 as the target site, which contains gRNA without a matching tag (SEQ ID NO: 9).

[0118] Figure 26 C shows the results of electroporation of cells using the ME system with HEK4 as the target site, which contains magRNA (SEQ ID NO: 9) with a tag that matches the template.

[0119] Figure 27Matching edits were shown to be effective in deletion edits. The mutant EGFP gene contained a 4-nt insertion starting at position 180. ME editing was performed using an ME template with a 29-nt RT extension, along with 12M29 magRNA containing a 12-nt tag matching the 5' end of the template. A lead editing experiment using pegRNA served as a control.

[0120] Figure 28 A and 28B compared the deletion editing efficiency of single-module ME, dual-module ME, single-module leader editor (PE2), and dual-module PE (PE3) in deleting the splice donor site (SDS) of exon 23 of the mouse dystrophin gene in the reporter construct and thereby generating exon 23 skip reads.

[0121] Figure 28 A shows a schematic diagram of a GFP-based splicing reporter.

[0122] The expression construct above contains the EGFP gene, which is divided into a 5' half and a 3' half, separated by an artificial intron containing a splice donor sequence (SDS) and a splice acceptor sequence (SAS). Transcription and splicing of the splice reporter gene will produce a functional EGFP.

[0123] The expression vector below was constructed by inserting exon 23 of the mouse dystrophin gene and introns 22 and 23, which are splice regulatory sequences. Transcription and splicing of the reporter gene will result in the production of a non-fluorescent EGFP protein with an insert peptide encoded by exon 23. Deletion of the Dmd exon 23 SAS or SDS will cause exon 23 to skip, thus producing fluorescent EGFP.

[0124] Single-module MEs and dual-module MEs (with magRNA-gRNA pairing) were designed to delete the 29 nt sequence at the Dmd exon 23-intron 23 junction, thereby eliminating the SDS sequence, as shown. For comparison, leader editors PE2 and PE3 with the same RT extension length and primer binding sites were also designed. Arrows indicate the nick site (first nick) of the first ME or PE2 and the second nick site of the dual-module ME or PE3.

[0125] Figure 28 B shows the results for untreated cells, or cells treated by electroporation with only the mDmd exon 23 spliced ​​reporter, or a single-module ME and the reporter (magLow), or a dual-module ME (magLow+sg), or cells treated with PE2 or PE3. The spliced ​​reporter was used as a positive control.

[0126] Figure 29The editing efficiencies of single-module ME, dual-module ME with magRNA-gRNA pairing, and dual-module ME with magRNA-magRNA pairing were compared. These ME systems are compared with... Figure 25 The systems shown are similar, differing only in the experiment in the last figure, where the second gRNA also contains a second matching tag at the 3' end. Results show untreated cells, or cells expressed using a mutant EGFP (with a nucleotide deleted at position 50 relative to the nick site) expression vector (EGFP). 50) Electroporated cells, or cells electroporated with a mutant EGFP expression vector plus a single-module ME (magLow), or dual-module ME (magLow+sgUp119) with magRNA-gRNA pairing, or dual-module ME (magLow+magUp119) with magRNA-magRNA pairing, all of which have an RNA template with an extension length of 111 nt. In magRNA-magRNA pairing, the matching tag sequences in the two magRNAs are complementary to the two independent sequences in the RNA RT template.

[0127] Figure 30 A, 30B, and 30C show the editing efficiency of endogenous HEK3 sites in HEK293T cells using dual ME with different ratios of RNA template, magRNA, and second gRNA.

[0128] Figure 30 Figure A shows the target site HEK3 sequence (5'TGGGGCCCAGACTGAGCACGTGATGGCAGAGGAAAGGAAGCCCTGCTTCCTCCAGAGGGCGTCGCAGGACAGCTTTTCCTAGACAGGGGCTAGTATGTGCAGC, SEQ ID NO: 10) and its complementary sequence (5'GCTGCACATACTAGCCCCTGTCTAGGAAAAGCTGTCCTGCGACGCCCTCTGGAGGAAGCAGGGCTTCCTTTCCTCTGCCATCACGTGCTCAGTCTGGGCCCCA, SEQ ID NO: 11), annotated with the components of the dual magRNA-gRNA ME system, including the magRNA (12M42) guide, the second nick gRNA guide, and the target point mutations 5G>T and 12G>C encoded in the RNA template T43. +63 indicates the distance between the two nicks. For clarity, only the 5' and 3' ends of the target site are shown in the figure.

[0129] Figure 30B shows representative Sanger sequencing chromatograms of unprocessed (UT) and edited (2 mut) HEK3 genomic DNA (SEQ ID NO:12).

[0130] Figure 30 C compared the editing efficiency at the endogenous genome target HEK3 under different ratios of RNA template (T43), magRNA, and second gRNA (5G>T).

[0131] Figure 31 A, 31B, and 31C demonstrate the efficiency of insertion editing at the endogenous HEK3 site in HEK 293T cells via the dual ME system.

[0132] Figure 31 Figure A shows the HEK3 sequence and its complementary sequences (SEQ ID NO: 10 and 11), labeled with the magRNA (12M42) guide, second nick guide, nick site, PAM, and insertion position (+5). RNA templates (T46, T50, and T79) were designed to insert 3-nt (CTT), 7-nt AP1 (ACTCAGT), or 36-nt TCR variable regions into endogenous HEK3 genomic sites. For clarity, only the 5' and 3' ends of the target sites are shown in the figure.

[0133] Figure 31 B shows representative Sanger sequencing chromatograms of untreated (UT) and edited genomic DNA (SEQ ID NO: 13).

[0134] Figure 31 C shows the efficiency of 3nt, 7nt, and 36nt insertions. Although the nCas9-RT fusion was used in this experiment, the second gRNA construct also contained MS2 in the stem-loop region.

[0135] Figure 32 A, 32B, and 32C demonstrate the editing efficiency of the trans-dual ME system in correcting point mutations in the free nfEGFP reporter gene in HEK293 cells when multiple polymerases, including MMLV-RT, R2Bm (a retrotransposon RT), and a Helraiser variant (a polymerase of the DNA transposon Heliton), were recruited. RNA and DNA templates were tested separately when using Helraiser. The configuration of the trans-ME system has been described in [details omitted]. Figure 24 Displayed in A. Figure 32 A shows the components of the trans-ME system. The gene to be edited is the nfEGFP target site with the A200G mutation. Figure 32 B). Figure 32C shows the editing efficiency of trans-double ME under different conditions, quantified as the percentage of cells with edited fluorescent EGFP.

[0136] Figure 33 A and 33B show the effects of recruitment of the non-RT effector p65 and the presence of MS2 in the second gRNA on editing efficiency in HEK293T cells. Figure 33 A shows a dual magRNA-gRNA ME assembly. The endogenous target site HEK3 has already been identified. Figure 30 Displayed in A. Figure 33 B compared the editing efficiency of the dual-ME system with or without p65, and with or without MS2 in the second gRNA (5G>T, Sanger sequencing).

[0137] Figure 34 A, 34B, 34C, and 34D show the results of comparing the editing efficiency of the split-type dual-ME system and its corresponding split-type starter editing system (PE) on the HEK3 site in HEK293T cells. The target HEK3 sequence has been... Figure 30 Displayed in A.

[0138] Figure 34 AC displays the components of ME and PE. Figure 34 A and B show the components specific to ME and PE, respectively. The pegRNA contains the same RNA template sequence as ME and is covalently linked to the 3' end of the gRNA. Figure 34 C shows the common components of the trans ME and PE systems. Figure 34 D shows the editing efficiency of ME and PE splitting under their respective optimal conditions. nCas9-RT splitting refers to nCas9 and RT being provided as two separate proteins rather than fusion proteins. For ME splitting, the optimal conditions for the highest editing efficiency of RNA template and magRNA are 1000 ng RNA template vector and 250 ng magRNA vector. For trans PE, the optimal conditions for the highest editing efficiency of pegRNA are 1000 ng of RNA template expression vector molar equivalents. The molar ratio of RNA template in the PE system: pegRNA scaffold in PE: RNA template in ME: magRNA scaffold in ME is approximately 1:1:1:0.25. The conditions are the same for common components, namely MCP-RT (500 ng), sp-nCas9 (500 ng), and the second gRNA (250 ng), between the ME and PE systems.

[0139] Figures 35 to 37This study compares the PE3 system and its dual-ME counterpart in editing endogenous HEK3 sites by introducing five different types of mutations into HEK293T cells. Editing efficiency and gRNA scaffold assembly insertion were also assessed.

[0140] Figure 35 A, 35B, and 35C show the components of the dual ME system and its PE3 counterpart. Figure 35 A shows the ME-specific components, namely the RNA template and the magRNA separation module. Figure 35 B shows a component unique to PE, namely pegRNA, in which the gRNA and RT template are covalently linked to each other. Figure 35 C shows the common components of ME and PE.

[0141] Figure 36 Figure A shows the HEK3 site sequence and its complementary sequences (SEQ ID NO: 10 and 11), with gRNA guidance annotated. The RT template T43-P13-5mut in the ME and PE systems contains 5 mutations. For clarity, only the 5' and 3' ends of the target site are shown in the figure.

[0142] Figure 36 B shows representative Sanger sequencing chromatograms of untreated (UT) and edited HEK3 genomic DNA (SEQ ID NO: 14). The arrows in the edited sample 5mut indicate the positions of the 5 edited bases.

[0143] Figure 36 C shows the quantification of editing efficiency detected at five locations by Sanger sequencing under each editing condition.

[0144] Figure 36 D shows the results of the comparison of editing efficiency measurements of the PE3 system and its ME counterpart at the 5G>T position. P<0.05, P<0.01, n=3.

[0145] Figure 37 The assay results shown in A, 37B, and 37C were obtained using next-generation sequencing (NGS). Figure 36 Representative samples were reanalyzed, and the differences between the ME and PE systems in editing efficiency and insertion caused by RT synthesis (reverse transcription) using gRNA scaffolds as templates were compared. Figure 37 A shows the editing efficiency of ME and PE edits at five sites. Figure 37B shows the number of reads containing the insertion at the end of the template sequence, which is the same as the 3' end sequence of the gRNA scaffold having 3 or more base pairs. The table also shows the number of edited reads and the ratio of inserted to edited reads. Figure 37 The dosage of the second gRNA was varied in ME-250, PE3-250, ME-500, and PE3-500 conditions in A (250 ng vs. 500 ng). 37C shows the ratio of reads containing gRNA scaffold templated insertions to edited reads.

[0146] Figure 38 A and 38B show the assay results in HEK293T cells comparing the editing of endogenous HEK3 sites by deleting or inserting 3-nt sequences using the dual ME system and its PE3 counterpart.

[0147] Figure 38 A shows the HEK3 site and its complementary site (SEQ ID NO: 10 and 11) as well as edit sites with expected insertion (CTT) (SEQ ID NO: 15-16) or deletion (GCA) (SEQ ID NO: 17-18).

[0148] Figure 38 B shows the editing efficiency of ME and PE3 in terms of import insertions and deletions, analyzed by Sanger sequencing. P<0.05, n=3.

[0149] Figure 39 The results show the editing efficiency of the dual-ME system and its corresponding PE system in K562 cells, comparing the efficiency of single nucleotide insertions at different sites on free EGFP plasmids containing a 1-nt deletion. Δ4A, Δ20C, and Δ35C are relative to the EGFP cleavage site at G202. Figure 17 C). ME components and Figure 25 The same as in.

[0150] Figure 40 A and 40B show the results of measurements comparing the editing efficiency of the dual ME system and its corresponding PE system in producing a 29-nt deletion in the Dmd exon 23 skip-read reporter plasmid in K562 cells. ME and PE components are compared with... Figure 28 same, Figure 28 The experiments were conducted in HEK293T cells. Figure 40 A shows the editing efficiency quantified by the percentage of GFP-positive cells, while Figure 40 B shows the editing efficiency quantified by Sanger sequencing.

[0151] Figure 41A, 41B, and 41C show assay results in K562 cells comparing the editing of endogenous HEK3 sites by deletion or insertion of 3-nt sequences using the dual ME system and its PE3 counterpart. Editing efficiency and insertion rate induced by RT template extension to the gRNA scaffold were compared. ME components and... Figure 38 The same as described in the text. Figure 38 The experiments were conducted in HEK293T cells. Figure 41 A and 41B show the efficiency of CTT insertions and GCA deletions analyzed by Sanger sequencing and next-generation sequencing (NGS), respectively. Figure 41 C shows the ratio of inserted reads to edited reads on the NGS gRNA scaffold.

[0152] Figure 42 A and 42B show the gene editing efficiency of the dual-ME system in introducing point mutations at endogenous HBB (hemoglobin β) near the E6V mutation site in sickle cell anemia in K562 cells. Figure 42 Image A shows exon 1 of the HBB sequence and its complementary sequences (SEQ ID NO: 19 and 20), and labels the magRNA guide, the second nick guide, the nick site, and the target +5G>T mutation encoded in the RNA template (marked with a box in the first PAM). In sickle cell anemia E6V, the mutation occurs at the adenine base immediately adjacent to the 5' end of the +5G base. Figure 42 B shows the efficiency of ME in introducing the +5G>T mutation in K562 cells.

[0153] Figure 43 A, 43B, 43C, and 43D show the results of the editing efficiency measurements comparing the saCas9 dual-ME system and its corresponding saCas9 PE system. In this study, both systems used the saCas9 (N580A)-RT fusion enzyme and the saCas9 gRNA scaffold. The efficiency of 1-nt insertion and silencing point mutations of the EGFP Δ100G plasmid in HEK293 cells was compared.

[0154] Figure 43A shows the EGFPΔ100G target sequence (5'tgagcaagggcgaggagctgttcaccggggtggtgcccatcctggtcgagctggacggcgacgtaaacggccacaagttcagcgtgtccggcgaggcgagggcgatgccacctacggcaagctgaccctgaagttcatctgcaccaccggcaag 3', SEQ ID NO: 21) and its complementary sequence (5'cttgccggtggtgcagatgaacttcagggtcagcttgccgtaggtggcatcgccctcgcctcgccggacacgctgaacttgtggccgtttacgtcgccgtccagctcgaccaggatgggcaccaccccggtgaacagctcctcgcccttgctca, SEQ ID NO: 22), and annotates magRNA guidance, second nick guidance, and saCas9. PAM and the expected edit sequence (5'tgagcaagggcgaggagctgttcaccggggtggtgcccatcctggtcgagctggacggcgacgtaaacggccacaagttcagcgtgtccggcgag) G gcgagggcgatgccacctacggcaag

[0155] ctgacc T tgaagttcatctgcaccaccggcaag 3'SEQ ID NO: 23) and its complementary sequence (5'cttgccggtggtgcagatgaacttca A ggtcagcttgccgtaggtggcatcgccctcgc C ctcgccggacacgctgaacttgtggccgtttacgtcgccgtccagctcgaccaggatgggcaccaccccggtgaacagctcctcgcccttgctca3', SEQ ID NO: 24). For clarity, only the 5' end, deletions or insertions, and the 3' end sequence are shown in the figure.

[0156] Figure 43 B shows the transition from non-fluorescent GFP to fluorescent GFP. Figure 43 CD compared the conversion efficiency via GFP ( Figure 43 C) and Sanger sequencing ( Figure 43D, Insert) to perform quantitative editing efficiency of ME and PE. Invention Details

[0158] This invention relates to a novel RNA-guided gene editing system, called a Match Editing (ME) system, which uses reverse transcriptase to copy the target sequence information from an RNA template onto DNA and inserts that DNA into a target site in the genome or extrachromosomal DNA. This gene editing system incorporates several novel features, such as: (1) the RNA template can be a completely independent RNA molecule and does not contain associated recruitment aptamer sequences such as MS2 aptamers; (2) the gRNA can be modified to contain short polynucleotide tags complementary to regions within the template; and (3) the template can be anchored to the editing site through Watson-Crick base pairing interactions between the edited RNA template and its matching gRNA. Previously, no design had been developed to directly recruit RNA templates using modified gRNA.

[0159] Unlike previously reported systems in terms of design and configuration, the novel system described in this paper provides an efficient platform for convenient gene editing and de novo gene construction at target loci in the genome or extrachromosomal DNA. This novel design allows for gene editing of target sequences far from PAM motifs, enabling the editing of all types of point mutations and efficient insertion and deletion. Furthermore, the modular nature of this design allows for synergistic multiple operations, thereby enhancing functionality. Therefore, this new system and its associated platform add a unique and versatile tool to gene editing technology.

[0160] As disclosed herein, an innovative feature of the ME system is that the gRNA can be engineered to guide the recruitment of the RNA template through its pairing with the Watson-Crick sequence of the natural template sequence. First, no existing technology teaches the engineering of gRNA sequences to directly recruit other gene-editing components (e.g., the RNA template) through Watson-Crick sequence complementarity. Second, no existing reverse transcriptase-mediated gene editing technology, including leader editing, teaches the engineering of gRNA sequences to directly recruit the RNA template through Watson-Crick sequence complementarity. This matching editing property is unexpected and not readily apparent compared to any existing technology (whether conventional RNA-guided gene editing or reverse transcriptase-mediated gene editing (e.g., leader editing)).

[0161] Furthermore, the unique anchoring mechanism between the engineered gRNA (referred to as magRNA in this disclosure) and the RNA template disclosed in this invention enables many novel properties that do not exist in any other reverse transcriptase-mediated gene editing system.

[0162] First, this recruitment mechanism is achieved through base pairing between the natural template sequence and its complementary sequence added to the magRNA. No exogenous or additional sequences, such as the MS2 aptamer, are required in the RNA template. Furthermore, template recruitment does not require trans-acting elements, such as MCP fusions. This design ensures efficient gene editing; moreover, compared to other similar reverse transcriptase-mediated genome editing systems, it is more streamlined and compact, facilitating efficient delivery.

[0163] Secondly, the modular and versatile anchoring mechanism facilitates the engineering of complex systems with multiple modules that work synergistically to enhance efficiency or add new functionalities. For example, two modules, each containing a tagged magRNA that matches different sequences in the same RNA template, can jointly recruit the RNA template, significantly improving editing efficiency. In addition to two independent modules jointly recruiting the RNA template, each module can also provide different effector molecules that work synergistically. For instance, one module might contain an RNA aptamer for recruiting reverse transcriptases that fuse with aptamer-binding proteins; another module might contain a different RNA aptamer for recruiting chromatin-modifying enzymes that fuse with the corresponding aptamer-binding protein. The physical and functional interactions between the two modules are reinforced by their interaction with the same RNA template recruited by the two magRNAs.

[0164] Third, the RNA template does not require exogenous sequences, thus mechanistically eliminating all current sources of "scarring" that affect reverse transcriptase-mediated gene editing. One of the most challenging problems in sequence-specific gene or gene fragment insertion into the genome is the simultaneous introduction of unwanted exogenous sequences, i.e., "scarring." To date, there is no effective "scar-free" solution when generating transgene inserts in mammalian non-germline cells using various enzymes such as reverse transcriptase, recombinase, and integrase. The root cause of "scarring" is that the template (whether RNA or DNA molecule) needs to contain not only the genetic information of the target site but also handles for specific interactions and recruitment between the template molecule and the gene manipulation mechanism. Examples of these auxiliary sequences for interaction with the gene manipulation system include flanking LTRs, ITRs, and RNA aptamers. These auxiliary sequences are sources of gene "scarring." Particularly relevant to this disclosure is the well-documented fact that leader editors often result in the insertion of unwanted pegRNA components (e.g., partial or complete gRNA CRISPR interaction scaffolds) into the target editing site, not just the target RNA template within the pegRNA (Peter J. Chen). et al Enhanced prime editing systems by manipulating cellular determinants of editing outcomes. Cell 184, 5635–5652, October 28, 2021. For example, among them... Figure 2 ).

[0165] Fourth, the system design separates the gRNA from the RNA template and, through a non-covalent Watson-Crick base pairing mechanism, endows the gRNA with different template recruitment functions. This combination makes it possible to perform efficient gene editing using long RNA templates without worrying about the inherent problem that the template might disrupt the secondary structure of the gRNA when covalently linked to it.

[0166] Fifth, structurally separating the RNA template and gRNA scaffold allows for alteration of the ratio between the RNA template and gRNA / nCas9 / reverse transcriptase, thereby achieving optimal editing results. In other words, in pegRNA design, the ratio of RNA template to gRNA molecules is always 1:1. Therefore, at the editing site, one nCas9-RT molecule binds to one pegRNA (a gRNA and a covalently linked template). In contrast, in ME, the ratio of RNA template to gRNA can be varied to improve editing efficiency and reduce non-specific editing. For example, unlike pegRNA, which can only achieve a 1:1 ratio, ME can achieve a 1:0.25 ratio as needed.

[0167] Therefore, by employing a unique magRNA design that directly anchors to and recruits independent, “scar-free” modular RNA template molecules, the ME system disclosed in this paper overcomes some of the existing key obstacles in the field of gene editing (e.g., the inherent risk of “scarring”), while improving flexibility and versatility (e.g., by allowing changes in the ratio of RNA template and gRNA to reverse transcriptase), thus creating conditions for optimal genome editing (e.g., high efficiency with negligible off-target effects). Therefore, the ME system exhibits unique characteristics and improvements in both functionality and design sophistication compared to existing systems.

[0168] Match Editing (ME) System

[0169] As described herein, the ME system possesses several unique characteristics. For example, the target gene sequence can be encoded by an independent RNA template molecule, independent of the gRNA, and free of protein recruitment motifs (e.g., MS2 aptamers) or any other sequences undesirable in the final product. In other words, ME allows for the free design of mechanistically scarless templates. Furthermore, the independent matching gRNA molecule contains a short RNA sequence or matching tag complementary to a region in the template RNA molecule. Further, template recruitment is mediated by Watson-Crick pairing between the template RNA and the matching sequence in the gRNA. The RNA template used herein is often referred to as a modular RNA template because it is a separate and independent sequence-specific DNA-binding module. The gRNA with a matching tag complementary to a region within the template is called magRNA.

[0170] Separating the DNA sequence recognition module (gRNA) from the gene sequence information (RNA template) to be integrated into the genome makes the system highly modular. Therefore, longer templates can be designed independently without concern for interfering with the secondary structure and function of the gRNA. Furthermore, compared to affinity binding between RNA motifs (e.g., MS2 RNA aptamers) and proteins (e.g., aptamer-binding proteins MCP), the anchoring mechanism of intermolecular Watson-Crick base pairing eliminates unnecessary components, such as RNA aptamer recruitment sequences within the template and aptamer-binding fusion portions in CRISPR proteins. This not only makes the system compact but also eliminates sources of editing “scars.” More importantly, the modular design and simple base pairing recruitment mechanism of the system enable the construction of complex systems through simple and synergistic multiplexing of individual modules connected (“clicks”) by nucleotide sequence complementarity with different regions of the same RNA template.

[0171] In some implementations, this disclosure provides examples of various basic single-module ME configurations and dual-module ME configurations by “clicking” two single-module MEs, enabling versatile gene editing and de novo gene construction functions.

[0172] Figure 1 The diagram shows the basic configuration of a single-module ME, in which reverse transcriptase is fused with an RNA-guided nicking enzyme (using CRISPR / Cas9 protein as an example). This single-module matching editing system utilizes a CRISPR / gRNA complex to recognize RNA guide sequences. The system comprises: (a) a modular RNA template without any RNA recruitment elements (e.g., RNA aptamers, such as MS2 or PP7); (b) a magRNA containing a 5' spacer sequence for target DNA recognition, an RNA scaffold for CRISPR / Cas complexation, and a 3' matching RNA sequence complementary to a region in the modular template; and (c) a nicking enzyme CRISPR protein nCRISPR / nCas9 (H840A) fused to a reverse transcriptase (MMLV-RT) variant via a polypeptide linker. Figure 1 As shown, this single-module basic system may include: (1) An RNA template (or nucleotide encoding the template expression) for reverse transcription, which contains a nucleotide sequence to be reverse transcribed to replace the target sequence at the target site, and whose 3' end sequence is complementary to the cleavage DNA strand at the target site. (2) Engineered magRNA, which contains a. Spacer sequences used for sequence-specific identification of target sites; b. An RNA scaffold for complexing with RNA-guided nicking enzyme proteins; and c. An RNA sequence tag (6-24 nucleotides in size) complementary to the sequence within the RNA template (1); and (3) RNA-guided nicking enzymes fused to reverse transcriptase via polypeptide linkers (e.g., nCRISPR / nCas9H840A).

[0173] Although Figure 1 As shown, the matching sequence is located at the 3' end of the gRNA, but the tag can be located at other locations within the gRNA, such as the stem-loop or tetra-loop region, or the 5' end of the gRNA.

[0174] Figure 2 This demonstrates how the matching editing system works. For example... Figure 2As shown, the matching editing system identifies the DNA target site through the spacer sequence at the 5' end of the engineered gRNA, forming an R-loop at the target DNA locus. The tag sequence at the 3' end of the gRNA binds to a complementary fragment in the RNA template, anchoring the template to the site. Simultaneously, nicking enzyme activity causes a single-strand DNA break on the non-target strand (e.g., if the nCas9 H840A nicking enzyme is used, the break occurs 3 nucleotides upstream of the PAM motif). The 3' end of the RNA template then anneals to the nicked complementary non-target DNA strand in the R-loop. The nicked end of the non-target DNA then acts as a primer for reverse transcriptase, utilizing the RNA template. Thus, the sequence information in the modular RNA template is copied onto the DNA strand, which is subsequently introduced into the target DNA molecule.

[0175] Figure 3 This demonstrates a cellular repair mechanism that introduces newly synthesized DNA strands to target DNA loci. More specifically, Figure 3 This explains how matched editing enables precise genome editing. The process includes the following steps: A. The nicking enzyme introduces a nick on the non-target strand within the target site; B. Reverse transcriptase copies the RNA template sequence into the DNA sequence initiated by a nick in the DNA strand; C. Newly synthesized 3'-DNA fragments compete with 5'-DNA fragments; D. Nucleases preferentially remove 5' fragments from the cell, and the newly synthesized DNA strand anneals with the opposite strand; E. DNA repair removes mismatched sequences from the "old" strand, resulting in editing, or removes mismatched sequences from the newly synthesized strand, resulting in no editing.

[0176] The newly synthesized strand can carry target point mutations (transversion, transformation, or both), insertions, or deletions.

[0177] Figure 4 This illustrates a specific example of magRNA-RNA template interaction. In this example, the 3' end matching tag sequence of the magRNA is specifically complementary to the 5' end sequence of the RNA template.

[0178] Figure 5 Another specific example is shown where the magRNA carries a 5' matching tag. In this example, the tag, complementary to the region of the RNA template, is located at the 5' end of the gRNA. The linker polynucleotide sequence is located between the 5' matching tag sequence and the spacer sequence.

[0179] Figure 6An exemplary variant of the matching editing system is shown, in which the reverse transcriptase is provided in a split form, rather than fused directly with an RNA-guided sequence-specific nicking enzyme. (See also...) Figure 6 As shown, the RT splitting system includes the following components: (1) An RNA template for reverse transcription, comprising a polynucleotide sequence to be reverse transcribed to replace the target sequence at the target site, and a 3' end sequence complementary to the cleavage strand at the target site. (2) RNA-guided nicking enzymes for sequence recognition; (3) Engineered magRNA, containing a. 5' guide for sequence-specific identification at the target site b. RNA scaffolds for complexing with RNA-guided nicking enzyme proteins c. RNA aptamer sequences at the circular structures of the RNA scaffold (e.g., stem loop or quadruple loop). d. A 3' extension sequence (6-24 nucleotides in size) complementary to the 5' end or middle sequence fragment of the RNA template (1); and (4) Proteins that contain reverse transcriptase and associated proteins (e.g., MCP) fused via polypeptide linkers, which bind to aptamers (e.g., MS2) within the gRNA scaffold.

[0180] Figure 7 This demonstrates the working principle of the RT splitting system. For example... Figure 7 As shown, the working principle of the RT splitting system is similar to... Figure 2 The diagrams shown are very similar, except that the reverse transcriptase (RT) is not fused to an RNA-guided sequence-specific complex protein (such as nCas9). Instead, the gRNA contains the MS2 aptamer at the gRNA stem loop (or other location), which recruits the MCP-RT fusion protein.

[0181] Figure 8 A variant configuration of the RT splitting system is shown, in which the 3' magRNA extension sequence is complementary to the 5' end sequence of the RT template.

[0182] Figure 9 Two dual-ME systems are shown, each with a nicking module that generates single-strand DNA breaks at a downstream position adjacent to the opposite strand. Figure 9 Figure A shows a dual ME system with fused RT and magRNA-gRNA pairing. The second module is used to generate a second notch to improve editing efficiency. Figure 9 Figure B shows a dual-ME system with splitting RT and magRNA containing MS2-gRNA 0xMS2. A second module is used to generate a second cut to improve editing efficiency. Each dual-ME system is... Figure 10 The mechanism shown improves editing efficiency.

[0183] More specifically, such as Figure 10 As shown, mechanistically, the second cut facilitates the removal of mismatched sequences in the cut chain, thereby improving editing efficiency. Therefore, the second cut on the adjacent opposite chain mediated by double ME significantly improves target editing efficiency.

[0184] This article discloses a single-template-coupled dual-ME system that utilizes a long RNA template to edit target sites. Two examples are shown: a dual-ME system with a fusion RT. Figure 11 A) and coupled dual ME with split RT ( Figure 11 (B) Each module of this system contains a magRNA with a matching tag complementary to a separate region of the RNA template. Anchoring the RNA template with two magRNAs enhances the interaction between the editing complex and the template, thereby improving editing efficiency or enhancing the function of individual modules.

[0185] Figure 11 It also illustrates how two (or more) individual ME modules can physically and functionally connect or “click” together with different regions of the same RNA template through simple base pairing complementarity to enhance activity or generate new synergistic functions.

[0186] This paper also discloses a coupling dual ME for enhancing functionality. Figure 12 Several examples are shown (RT merged version and RT split version). Figure 13 The mechanisms behind these processes are explained. In these examples, the individual modules are designed to work together. For instance, in the RT fusion version, the 3' end of the RNA template is designed to be identical to the 3' end of the second nick strand. Therefore, the second nick strand acts as a primer, using the newly synthesized DNA strand as a template to synthesize the second DNA strand.

[0187] Figure 13 The diagram illustrates the co-interactions involving the long RNA template and the protocol for second-strand DNA synthesis via a dual ME system coupled to a single template: (A)-(B) The first module performs the first nick and reverse transcription on the long template; (C)-(D) The second module performs the second nick; (E)-(F) The DNA-dependent DNA polymerase activity of the reverse transcriptase mediates the synthesis of the second DNA strand, thanks to the complementary interaction between the 3' end of the second nick site and the 5' end of the RNA template; (G)-(H) Removal of the 5' fragment.

[0188] This paper also discloses an example system that can be used for insertion or deletion. For example... Figure 14As shown, Module 1 and Module 2 are separate ME modules, each containing individual components as described above, including a separate RNA template. In this configuration, although the two modules do not physically interact via a common RNA template, they can be designed to functionally cooperate. For example, the RNA template can be designed such that the reverse-transcribed DNA strands are partially complementary to each other.

[0189] The two modules shown generate trans cuts on opposite strands. Two or more modules can function in cis (PAM sequence and guide sequence are on the same strand) or trans, thereby achieving a variety of functions, such as tandem insertion of small fragments to achieve insertion of large fragments, or correction of multiple mutations with various defects.

[0190] Figure 15 The diagram illustrates how the dual-template, dual-module ME system enables de novo gene construction (insertion) or large-fragment gene deletion. For example... Figure 15 As shown in A and 15B, modules 1 (ME) and 2 (ME) mediate cleavage at their respective target sites and synthesize first-strand DNA via reverse transcription of their respective RNA templates. Subsequently, the 3' ends of the two first-strand DNAs anneal through complementary sequences, as shown in Figures 1 and 2. Figure 15 As shown in C. This further enables the synthesis of second-strand DNA via DNA-dependent DNA polymerase activity of reverse transcriptase or endogenous DNA polymerase activity, such as... Figure 15 As shown in D. After removing the 5' strip and connecting it ( Figure 15 E), which can achieve insertion or deletion ( Figure 15 F). This mechanism can lead to insertions and deletions (de novo gene construction), depending on the composition and length of the two RNA templates and the sequence being removed.

[0191] Any system described herein may contain multiple second effectors. These effectors can be provided in an insplit or fusion manner, such as... Figure 16 As shown. Examples can include proteins that facilitate the editing process or improve editing efficiency, such as chromatin-modifying enzymes, mismatch repair enzyme inhibitors, and proteins that can stabilize template RNA and magRNA. Examples of effectors include human RNase inhibitor proteins that enhance editing efficiency, dominant and negative MMR pathway proteins (including MLH1 protein), and the 5' DNA nuclease Fen1 protein. Figure 16 The modules shown are designed for "click" or multiplication to Figures 9 to 15 In the dual-system or complex system shown.

[0192] RNA-guided nickase

[0193] The key second component of the gene editing complex or system disclosed in this article is an RNA-guided nicking enzyme. A nicking enzyme is an enzyme that creates single-strand breaks (also known as "nicks") in double-stranded DNA, that is, it cuts one strand of the DNA double helix without cutting the other strand.

[0194] As used herein, the term "RNA-guided nickase" refers to a polypeptide or polypeptide complex with DNA nickase activity, wherein the DNA nickase activity may be sequence-specific and dependent on the RNA sequence. Exemplary RNA-guided nickases include Cas nickases. Cas nickases include, but are not limited to, the nickase form of the Csm or Cmr complex of the type III CRISPR system, its Cas1O, Csml, or Cmr2 subunits, the Cascade complex of the type I CRISPR system, its Cas3 subunit, and class 2 Cas nucleases. Class 2 Cas nickases include two classes of Cas nuclease variants in which only one of the two catalytic domains is inactivated, and these variants possess RNA-guided DNA nickase activity. Class 2 Cas nickases include, for example, Cas9 (e.g., SpyCas9 variants H840A, D10A, or N863A), Cpfl, C2cl, C2c2, C2c3, HF Cas9 (e.g., N497A, R661A, Q695A, Q926A variants), HypaCas9 (e.g., N692A, M694A, Q695A, H698A variants), eSPCas9(1.0) (e.g., K810A, KI003A, R1060A variants), and eSPCas9(ll) (e.g., K848A, K1003A, R1060A variants) proteins and their modifications. Cpfl protein (Zetsche...) et al The Cpfl sequence (Cell, 163: 1-13 (2015)) is homologous to Cas9 and contains a RuvC-like protein domain. Zetsche's Cpfl sequence is incorporated in its entirety by reference. See, for example, Zetsche's Tables S1 and S3. "Cas9" includes 5.pyogenes (Spy) Cas9, the Cas9 variants listed herein, and their equivalents. See, for example, Makarova. et al ., NatRev Microbiol, 13(11): 722-36 (2015); Shmakov et al ., Molecular Cell, 60:385-397 (2015); Makarova et al ., NAT. REV. MICROBIOL, 18:67-83 (2020).

[0195] In some embodiments, the RNA-guided nickase disclosed herein is a Cas nickase. In some embodiments, the RNA-guided nickase is derived from a specific Cas nuclease whose catalytic domain has been inactivated. In some embodiments, the RNA-guided nickase is a class 2 Cas nickase, such as a type II Cas9 nickase or a Cpfl nickase. In some embodiments, the RNA-guided nickase is a Streptococcus pyogenes Cas9 nickase. In some embodiments, the RNA-guided nickase is a Neisseria meningitidis Cas9 nickase. In some embodiments, the RNA-guided nickase is a Staphylococcus aureus Cas9 nickase.

[0196] In some embodiments, class 2 Cas is a type V Cas protein, such as Cas12. In some embodiments, Cas12 is Cpf1 (Cas12a), C2c1 (Cas12b), C2c3 (Cas12c), CasY (Cas12d), CasX (Cas12e), Cas14 (Cas12f), CasPhi (Cas12j), and their orthologs and variants.

[0197] In some embodiments, the Cas protein is a type VI Cas protein, such as Cas 13. In one embodiment, Cas13 is Cas13a, Cas13b, Cas13c, Cas13d, Cas13x, and their orthologs and variants.

[0198] In some implementations, the RNA-guided nickase disclosed herein is a variant of a transposon-encoded IscB, IsrB, or TnpB family endonuclease protein.

[0199] In some embodiments, the RNA-guided nicking enzyme is a modified class 2 Cas protein or a Cas protein derived from a class 2 Cas protein. In some embodiments, the RNA-guided nicking enzyme is a modified Cas protein or a Cas protein derived from a Cas protein, such as a class 2 Cas nuclease (e.g., it may be a type II, type V, or type VI Cas nuclease). Class II Cas nucleases include, for example, Cas9, Cpfl, C2cl, C2c2, and C2c3 proteins and their modifications. Examples of Cas9 nucleases include type II CRISPR systems of Streptococcus pyogenes, Staphylococcus aureus, and other prokaryotes (e.g., see the list in the next paragraph) and their modified (e.g., engineered or mutated) versions. See, for example, US2016 / 0312198 A1 and US 2016 / 0312199 A1, the entire contents of which are incorporated herein by reference. Other examples of Cas nucleases include the Csm or Cmr complex of type III CRISPR systems or their Cas10, Csml, or Cmr2 subunits; and the Cascade complex of type I CRISPR systems or their Cas3 subunits. In some embodiments, the Cas nuclease may be derived from type IIA, type IIB, or type IIC systems. For a discussion of various CRISPR systems and Cas nucleases, see, for example, Makarova. et al ., NAT. REV. MICROBIOL. 9:467-477 (2011); Makarova et al ., NAT.REV. MICROBIOL, 13: 722-36 (2015); Shmakov et al ., MOLECULAR CELL, 60:385-397(2015); Makarova et al ., NAT. REV. MICROBIOL, 18:67-83 (2020).

[0200] The Cas nickases described in this article can be nickase forms of Cas nucleases from various bacterial species, including but not limited to Streptococcus pyogenes (Streptococcus pyogenes). Streptococcus pyogenes Streptococcus thermophilus ( Streptococcus thermophilus Streptococcus spp. Streptococcus sp. Staphylococcus aureus Staphylococcus aureus Listeria monocytogenes ( ), harmless Listeria monocytogenes ( Listeria innocua Lactobacillus gasseri ( Lactobacillus gasseri ), the new culprit, Francisella ( Francisella novicida ), succinic acid-producing Woring bacteria ( Wolinella succinogenes ), Sartorius vulgaris ( Sutterella wadsworthensis ), γ-Proteobacteria ( Gammaproteobacterium ), Neisseria meningitidis ( Neisseria meningitidis Campylobacter jejuni ( Campylobacter jejuni ), Pasteurella multocida ( Pasteurella multocida ), Succinic acid-producing filamentous bacteria ( Fibrobacter succinogenes ), Rhodospirillum rubrum ( Rhodospirillum rubrum ), Nocardia dassonvillei ( Nocardiopsis dassonvillei ), Streptomyces coccidioides ( Streptomyces ancient spiral ), Streptomyces greeni Streptomyces viridochromogenes ), Streptomyces greeni Streptomyces viridochromogenes ), Rose cysts ( Streptosporangium roseum ), Rose cysts ( Streptosporangium roseum ), Bacillus acidophilus ( Alicyclobacillus acid-caloric ), Pseudomycium-like Bacillus ( Bacillus pseudomycoides ), selenium-reducing Bacillus ( Bacillus selenite reducing ), Siberian microbacteria ( Exiguobacterium sibiricum Lactobacillus delbrueckii (), Lactobacillus delbrueckii ), Lactobacillus salivarius ( Lactobacillus salivarius Lactobacillus bruneri ( Lactobacillus buchneri ), dental spirochetes ( Treponema denticola Marine microoscillator bacteria ( Marine microscilla Burkholderia ( ) Burkholderiales bacteria ), Naphthylazine-eating aeromonas ( Polaromonas naphthalenivorans ), *Extreme Monoclonalella* spp. Polaromonas sp. ), Chlorella vulgaris ( Crocosphaera watsonii ), Cyanobacteria ( Cyanothece sp. Microcystis aeruginosa ( Microcystis aeruginosa Synechocybe ( Synechococcus sp. ), Arachidonic acid bacteria ( Acetohalobium Arabic ), ammonia-producing bacteria ( Ammonite of the Degens ), Caldicellosiruptor becscii , Desulfurous Candidate Clostridium botulinum ( Clostridium botulinum Clostridium difficile ( Clostridium difficile ), Griffon's bacterium ( Finegoldia magna ), thermophilic anaerobic bacteria ( Natranaerobius thermophilus ), Bacillus pyrolyticus ( Pelotomaculum thermopropionic ), Thiobacillus acidophilus ( Acidithiobacillus caldus ), Thiobacillus ferrooxidans ( Acidithiobacillus ferrooxidans ), wine-coloring bacteria ( Allochromatium vinous ), Marinebacteria ( Marinobacter sp. ), halophilic nitrosococci ( Nitrosococcus halophilus ), Nitrostrophus warwicki ( Nitrosococcus watsoni ), Salt-treated pseudoalternomonas ( Pseudoalteromonas haloplanktis ), racemic fibroblasts ( Ktedonobacter racemifer ), trace methanophiles ( Methanohalobium discovered ), and variable anemones ( Anabaena variabilis ), Foamy Glomerula ( Nodular foamy ), Nostoc ( Nostoc sp. ), Spirulina macrophylla ( Arthrospira maxima ), *Spiralella obtusifolia* ( Arthrospira platensis ), genus Arthrospira ( Arthrospira sp. ), genus *Cyclophora* ( Lyngbya sp. ), Earth microcoleopteran ( Microcoleus chthonoplastes ), Oscillatoria ( Oscillatoria sp. ), active fossil fungi ( Petrotoga mobilis ), African thermob Thermosipho africanus ), Pasteurella multocida ( Streptococcus pasteurianus ), Neisseria grayi ( Neisseria cinerea ), Campylobacter rubrum ( Campylobacter lari ), Parabacterium washing ( Parvibaculum lavamentivorans Corynebacterium diphtheriae Corynebacterium diphtheriae ), amino acid cocci ( Acidaminococcus sp. ), Styloides bacterium ND2006 ( Lachnospiraceae bacterium ND2006 ) or marine prochlorococcus ( Acaryochloris marina ).

[0201] In some embodiments, the Cas nickase is a nickase form of the Cas9 nuclease from *Streptococcus pyogenes*. In some embodiments, the Cas nickase is a nickase form of the Cas9 nuclease from *Streptococcus thermophilus*. In some embodiments, the Cas nickase is a nickase form of the Cas9 nuclease from *Neisseria meningitidis*. See, for example, WO / 2020081568, which describes an Nme2Cas9 D16A nickase. In some embodiments, the Cas nickase is a nickase form of the Cas9 nuclease from *Staphylococcus aureus*. In some embodiments, the Cas nickase is a nickase form of the Cpfl nuclease from *Neotrichomoniasis*. In some embodiments, the Cas nickase is a nickase form of the Cpfl nuclease from *Aminococcus* spp. In one embodiment, the Cas nickase is from bacteria of the family *Trichophyton* (…). Lachnospiraceae bacterium The Cas nickase is a cleavage enzyme form of the Cpfl nuclease from ND2006. In other embodiments, the Cas nickase is derived from *Tulafrancsis* (…). Francisella tularensis ), bacteria of the family Trichophyceae ( Lachnospiraceae bacterium ), Vibrio butyricum ( Butyrivibrio proteoclasticus), Wandering bacteria ( Peregrinibacteria bacterium ), Small Box Bacteria ( Parcubacteria bacterium ), genus Smith ( Smithella ), amino acid cocci ( Acidaminococcus ), termite candidate methane protoplasts ( Candidatus Methanoplasma termitum ), picky eubacterium ( Eubacterium eligens ), Moraxella calfii ( Moraxella bovoculi ), Leptospira in paddy fields ( Leptospira inadai ), Porphyromonas canis gingivalis ( Porphyromonas crevioricanis ), Prevotella difficile ( Prevotella disiens ) or Porphyromonas maculatus ( Porphyromonas macacae The cleavage enzyme form of Cpfl nuclease. In some embodiments, the Cas cleavage enzyme is derived from aminococci (C. spp.). Act daminococcus This can be a nicking enzyme form of a Cpfl nuclease from the family Trichophyceae. As previously mentioned, nicking enzymes can be derived from (i.e., associated with) a specific Cas nuclease because the nicking enzyme is an inactivated form of one of the two catalytic domains of that nuclease, for example, by mutating active site residues crucial for nucleic acid degradation, such as D10, H840, or N863 in SpyCas9. Techniques for easily identifying corresponding residues in other Cas proteins, such as sequence alignment and structure alignment, will be well known to those skilled in the art, and these techniques will be discussed in detail below.

[0202] In other embodiments, the Cas nickase may be associated with a type I CRISPR / Cas system. In some embodiments, the Cas nickase may be a component of the Cascade complex of a type I CRISPR / Cas system. In some embodiments, the Cas nickase may be the Cas3 protein. In some embodiments, the Cas nickase may be derived from a type III CRISPR / Cas system.

[0203] In some embodiments, the Cas nickase is a nickase form of a Cas nuclease, or a modified Cas nuclease in which the endonuclease active site is inactivated, for example, by alteration of one or more of the catalytic domains (e.g., point mutation). See, for example, U.S. Patent No. 8,889,356, which discusses Cas nickases and exemplary catalytic domain alterations.

[0204] Wild-type *Streptococcus pyogenes* Cas9 possesses two catalytic domains: RuvC and HNH. The RuvC domain cleaves the non-target DNA strand, while the HNH domain cleaves the target DNA strand. In some embodiments, the Cas nuclease may contain amino acid substitutions in the RuvC or RuvC-like nuclease domains. Exemplary amino acid substitutions in the RuvC or RuvC-like nuclease domains include D10A (based on the *Streptococcus pyogenes* Cas9 protein). See, for example, Zetsche... et al (2015) Cell Oct 22:163(3):759-771. In some embodiments, the Cas nuclease may contain amino acid substitutions in the HNH or HNH-like nuclease domain. Exemplary amino acid substitutions in the HNH or HNH-like nuclease domain include E762A, H840A, N863A, H983A, and D986A (based on the *Streptococcus pyogenes* Cas9 protein). See, for example, Zetsche et al. (2015). Other exemplary amino acid substitutions include D917A, E1006A, and D1255A (based on the *Streptococcus pyogenes* U112 Cpfl (FnCpfl) sequence (UniProtKB - A0Q7Q2 (CPF1 FRATN))).

[0205] In some embodiments, the Cas nickase (e.g., Cas9 nickase) has an inactivated RuvC or HNH domain. In some embodiments, a nickase with a reduced-activity RuvC domain is used. In some embodiments, a nickase with an inactivated RuvC domain is used. In some embodiments, a nickase with a reduced-activity HNH domain is used. In some embodiments, a nickase with an inactivated HNH domain is used. In some embodiments, the Cas9 nickase has an active HNH nuclease domain capable of cleaving the non-target strand of DNA (i.e., the gRNA-bound strand) and an inactivated RuvC nuclease domain, unable to cleave the target strand of DNA (i.e., the strand requiring base editing by deaminase).

[0206] The following provides an exemplary amino acid sequence of the Cas9 nickase: spCas9 (H840A) (SEQ ID NO:25)

[0207] saCas9(N580A) (SEQ ID NO: 26)

[0208] Other examples include schizokinase versions of the following Cas variants: SpCas9 D1135E variant, SpCas9 VRER variant, SpCas9 EQR variant, SpCas9 VQR variant, SpCas9-NG, xCas9, SpCas9-NG; Staphylococcus aureus (SA); SaCas9, Aminococcus spp. (AsCpf1) and Trichophyton spp. (LbCpf1); AsCpf1 RR variant, LbCpf1RR variant, AsCpf1 RVR variant, Campylobacter jejuni (CJ) Cas9, Neisseria meningitidis (NM) Cas9, Streptococcus thermophilus (ST) Cas9, and Treponema denticulatum (TD) Cas9.

[0209] In some embodiments, the RNA-guided nickase contains an amino acid sequence that is at least 80%, 90%, 95%, 98%, or 99% identical to the sequence described above.

[0210] reverse transcriptase

[0211] The gene editing complexes or systems disclosed herein contain reverse transcriptase. As used herein, “reverse transcriptase” describes a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require primers to synthesize DNA transcripts using RNA as a template. Historically, reverse transcriptases have primarily been used to transcribe mRNA into cDNA, which was then cloned into vectors for further manipulation.

[0212] Avian myoblastoma virus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473:1 (1977)). This enzyme possesses 5'-3' RNA-directed DNA polymerase activity, 5'-3' DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a persistent 5' and 3' ribonuclease that specifically acts on the RNA strand of the RNA-DNA hybrid (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Berger et alBiochemistry 22:2365-2372 (1983) provided a detailed study of AMV reverse transcriptase and its associated RNase H activity. Another widely used reverse transcriptase in molecular biology is the reverse transcriptase derived from Moloney mouse leukemia virus (M-MLV). See, for example, Gerard, GR, DNA 5:271-279 (1986) and Kotewicz, ML, et al Gene 35:249-258 (1985). Furthermore, M-MLV reverse transcriptases that are substantially lacking in RNase H activity have been reported. See, for example, U.S. Patent No. 5,244,797. This disclosure contemplates the use of any such reverse transcriptase, or a variant or mutant thereof.

[0213] Reverse transcriptases also originate from non-viral sources, including retrotransposons. Examples include reverse transcriptases from LTR retrotransposons, endogenous retroviruses, non-LTR retrotransposons, and LINE. Example proteins include Line 1ORF2, R2Bm, and R2Ol.

[0214] This disclosure covers any wild-type reverse transcriptase obtained from any naturally occurring organism or virus, or from commercial or non-commercial sources. Furthermore, the reverse transcriptases usable in the gene-editing complexes or systems of this disclosure may include any naturally occurring mutant RT, engineered mutant RT, or other variant RT, including truncated variants that retain function. RT may also be engineered to include specific amino acid substitutions, such as those specifically disclosed herein.

[0215] Reverse transcriptases are multifunctional enzymes, typically possessing three enzymatic activities: RNA- and DNA-dependent DNA polymerization activity, and RNase H activity catalyzing the cleavage of RNA in RNA-DNA hybrids. Some reverse transcriptase mutants partially inactivate RNase H to prevent accidental damage to RNA. These enzymes, which synthesize complementary DNA (cDNA) from mRNA templates, were initially discovered in RNA viruses. Subsequently, reverse transcriptases have been isolated and purified directly from viral particles, cells, or tissues. (See, for example, Kacian...) et al ., 1971, Biochim. Biophys. Acta 46: 365-83; Yang et al ., 1972,Biochem. Biophys. Res. Comm. 47: 505-11; Gerard et al ., 1975, J. Virol.15:785-97; Liu et al., 1977, Arch. Virol.55187- 200; Kato et al ., 1984, J.Virol. Methods 9: 325-39; Luke et al ., 1990, Biochem.29: 1764-69 and Le Grice et al (See references: ., 1991, J. Virol.65: 7004-07, all of which are incorporated herein by reference). In recent years, mutant and fusion proteins have been constructed to improve properties such as thermal stability, fidelity, and activity. This article covers any wild-type, variant, and / or mutant forms of reverse transcriptase known in the art or that can be prepared using methods known in the art.

[0216] Examples of sources of reverse transcriptase include, but are not limited to: Moloney murine leukemia virus (M-MLV or MLVRT); human T-cell leukemia virus type 1 (HTLV-1); bovine leukemia virus (BLV); Rous sarcoma virus (RSV); human immunodeficiency virus (HIV); yeast, including Saccharomyces cerevisiae, Neurospora, and fruit flies; primates; and rodents. See, for example, Weiss, et al ., US Pat. No. 4,663,290 (1987); Gerard, GR,DNA:271-79 (1986); Kotewicz, ML, et al ., Gene 35:249- 58 (1985); Tanese,N., et al ., Proc. Natl. Acad. Sci. (USA):4944-48 (1985); Roth, MJ, et al .,J. Biol. Chem.260:9326-35 (1985); Michel, F., et al ., Nature 316:641-43(1985); Akins, RA, et al ., Cell 47:505-16 (1986), EMBO J.4:1267-75 (1985); and Fawcett, DF, Cell 47:1007-15 (1986) (all of the above are incorporated herein by reference in their entirety).

[0217] Exemplary enzymes may include, but are not limited to, M-MLV reverse transcriptase and RSV reverse transcriptase. Enzymes with reverse transcriptase activity are commercially available. In some embodiments, the reverse transcriptase is provided in a trans-present manner to other components of the ME system. That is, the reverse transcriptase is expressed as a separate component or otherwise provided.

[0218] Those skilled in the art will recognize that wild-type reverse transcriptases, including but not limited to Moloney murine leukosis virus (M-MLV); human immunodeficiency virus (HIV) reverse transcriptases and avian sarcoma leukosis virus (ASLV) reverse transcriptases, including but not limited to Rous sarcoma virus (RSV) reverse transcriptases, avian myeloblastosis virus (AMV) reverse transcriptases, avian erythroblastosis virus (AEV) helper virus MCAV reverse transcriptases, avian myelomavirus MC29 helper virus MCAV reverse transcriptases, avian reticuloendotheliosis virus (REV-T) helper virus REV-A reverse transcriptases, avian sarcoma virus UR2 helper virus UR2AV reverse transcriptases, avian sarcoma virus Y73 helper virus YAV reverse transcriptases, Rous-associated virus (RAV) reverse transcriptases, and myeloblastosis-associated virus (MAV) reverse transcriptases, are all suitable for use in the methods and compositions described herein. Further examples are described at Martín-Alonso, S. et al The contents of Trends in Biotechnology 39 (2): 194-210 (2021) and WO2020191248 have been incorporated into this article.

[0219] Example reverse transcriptase variant: MMLV-RT variant

[0220] nuclear localization signal

[0221] In some implementations, ME component proteins contain one or more nuclear localization signal (NLS) peptides. NLS peptides can be located at the N-terminus, C-terminus, or internal region of the protein. NLS peptides are typically short peptides that act as signaling fragments mediating protein transport from the cytoplasm to the nucleus. A list of commonly used NLS peptides can be found in Lu, J., Wu, T., Zhang, B. et al Cell CommunSignal 19, 60 (2021). The complete NLS list was manually compiled from the database genome.unmc.edu / LocSigDB / .

[0222] In some embodiments, the NLS is a classic monopartite or bipartite nuclear localization signal (cNLS). A monopartite NLS consists of 4-8 basic amino acids, typically containing four or more positively charged residues of arginine (R) and lysine (K). A bipartite NLS consists of two basic amino acid segments separated by a spacer of approximately 10 amino acids. In some embodiments, the bipartite NLS is the nucleoplasmic protein NLS peptide KR[PAATKKAGQA]KKKK (SEQ ID NO: 28).

[0223] In some embodiments, the single NLS is the SV40 large T antigen NLS peptide Pro-Lys-Lys-Lys-Arg-Lys-Val (PKKKRKV, SEQ ID NO: 29). In some embodiments, the SV40 NLS is located at the N-terminus of the ME component protein. In some embodiments, the SV40 NLS is located at the C-terminus of the ME component protein.

[0224] connector

[0225] In some embodiments, the reverse transcriptase and RNA-guided nicking enzyme described herein are linked by a linker, such as, but not limited to, chemical modification, peptide linkers, chemical linkers, covalent or non-covalent bonds, protein fusion, or any method known to those skilled in the art. This linker may be permanent or reversible. See, for example, U.S. Patent Nos. 4,625,014, 5,057,301, and 5,514,363, U.S. Patent Application Nos. 20,150,182,596, and 20,100,063,258, and WO 2012,142,515, the contents of which are incorporated herein by reference in their entirety. In some embodiments, multiple linkers may be included to utilize the desired properties of each linker and each protein domain in the conjugate. For example, flexible linkers and linkers capable of improving the solubility of the conjugate may be considered, either alone or in combination with other linkers. The peptide linker may be linked to one or more protein domains in the conjugate by expressing DNA encoding the linker. The linker may be acid-cleaving, photocleaving, and thermosensitive linkers. These conjugation methods are well known to those skilled in the art and are included in this disclosure.

[0226] In some embodiments, the linker can be an organic molecule, a polymer, or a chemical component. In some embodiments, the linker can be a peptide linker. In some embodiments, the peptide linker can be any amino acid segment having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, or more amino acids.

[0227] In some implementations, the peptide linker is a 16-residue "XTEN" linker, or a variant thereof (see, for example, Schellenberger). et al . A recombinant polypeptide extends the in vivo half-lifeof peptides and proteins in a tunable manner. Nat. Biotechnol. 27, 1186-1190(2009)).

[0228] In some embodiments, the peptide linker comprises (GGGGS)n (SEQ ID NO: 30), (G)n, and (EAAAK). n (SEQ ID NO: 31), (GGS)n, SGSETPGTSESATPES motif (SEQ ID NO: 32) (see, for example, Guilinger JP, Thompson DB, Liu D R. Fusion of catalytically inactive Cas9 to FokInuclease improves the specificity of genome modification. Nat. Biotechnol. 2014; 32(6): 577-82; the entire contents of which are incorporated herein by reference), or (XP) n A base order, or any combination of the above base orders, where n is an integer between 1 and 30. See WO2015089406, the entire contents of which are incorporated herein by reference.

[0229] Matching gRNA (magRNA) molecules

[0230] Another key component of the gene editing complex or system disclosed herein is an RNA molecule called matching gRNA or magRNA. The magRNA may contain three sub-components: (1) a programmable guide RNA sequence or spacer sequence for recognizing target DNA; (2) an RNA scaffold capable of binding to an RNA-guided nicking enzyme; and (3) an anchoring tag sequence complementary to a fragment in the RNA template. The magRNA may be a single RNA molecule or a complex of multiple RNA molecules.

[0231] As disclosed in some embodiments, magRNA and an RNA-guided nicking enzyme (e.g., a Cas protein) together form a complex (e.g., a CRISPR / Cas-based module) for sequence targeting and recognition. The tag sequence carries the template to the target site to be edited, and the reverse transcriptase can be recruited to this site in various ways. In some examples, the reverse transcriptase can be recruited directly by fusing with a nicking enzyme. In other examples, the reverse transcriptase can be recruited in a split (individual) manner by recruiting an RNA motif that recruits the reverse transcriptase via an RNA-protein binding pair.

[0232] In some implementations, the matching tag sequence in the gRNA does not contain polyN, where N is any repeating base of a ribonucleotide or deoxyribonucleotide.

[0233] Programmable guide RNA / spacer region

[0234] One subcomponent is the programmable guide RNA. Due to its simplicity and efficiency, the CRISPR-Cas system can be used for genome editing in cells of various organisms. The system's specificity depends on the base pairing between the target DNA and the custom-designed guide RNA. By engineering and adjusting the base pairing characteristics of the guide RNA, any sequence of interest can be targeted, provided that the PAM sequence is present in the target sequence.

[0235] In the subcomponents of the magRNA disclosed herein, a guide sequence imparts target specificity. This guide sequence contains a region complementary to and capable of hybridizing with a pre-selected target site. In various embodiments, the guide sequence may contain from about 10 nucleotides to more than about 25 nucleotides. For example, the length of the base-pairing region between the guide sequence and the corresponding target site sequence may be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or greater than 25 nucleotides. In one exemplary embodiment, the guide sequence is about 17-20 nucleotides long, for example, 20 nucleotides.

[0236] In some implementations, a requirement for selecting a suitable target nucleic acid is that it has a 3' PAM site / sequence. In some examples, this document refers to each target sequence and its corresponding PAM site / sequence as the Cas-targeted site. The type II CRISPR system is one of the most thoroughly studied systems to date, requiring only the Cas9 protein and a guide RNA complementary to the target sequence to achieve target cleavage. The type II CRISPR system for *Streptococcus pyogenes* uses a target site with N12-20NGG, where NGG represents the *Streptococcus pyogenes* PAM site and N12-20 represents 12-20 nucleotides immediately adjacent to the 5' end of the PAM site. PAM site sequences for other bacterial species include NGNNG, NNNGATT, NNAGAAW, and NAAAAAC. See, for example, US20140273233, WO 2013176772, Cong et al. , (2012), Science 339 (6121): 819–823, Jinek et al. , (2012), Science 337 (6096): 816–821, Mali et al , (2013), Science339 (6121): 823–826, Gasiunas et al. , (2012), Proc Natl Acad Sci US A. 109(39): E2579–E2586, Cho et al. , (2013) Nature Biotechnology 31, 230–232, Hou et al., Proc Natl Acad Sci US A. 2013 Sep 24;110(39):15644-9, Mojica et al., Microbiology. 2009 Mar;155(Pt 3):733-40 and www.addgene.org / CRISPR / . The contents of these documents are incorporated herein by reference in their full text.

[0237] The target nucleic acid strand can be either of the two strands of the host cell's genomic DNA. Examples of such genomic dsDNA include, but are not limited to, host cell chromosomes, mitochondrial DNA, and stable extrachromosomal DNA. However, it is understood that this method can also be used for other dsDNAs present in the host cell, such as unstable plasmid DNA, viral DNA, and bacteriophage DNA, as long as the Cas-targeted site is present, regardless of the nature of the host cell dsDNA. This method can also be used for RNA.

[0238] RNA scaffolds capable of binding to RNA-guided nicking enzymes

[0239] In addition to the guide sequence described above, magRNA may also contain other active or inactive subcomponents. In one instance, magRNA may have an RNA motif (e.g., a CRISPR motif) carrying tracrRNA activity. For example, magRNA may be a hybrid RNA molecule in which the programmable guide RNA described above is fused with tracrRNA to mimic the natural crRNA:tracrRNA duplex.

[0240] In some implementations, magRNA comprises both crRNA and tracrRNA, with the tracrRNA followed by a matching sequence. The matching tag can be located within either the crRNA or the tracrRNA.

[0241] Various tracrRNA sequences are known in the art, including the following tracrRNAs and their active moieties. In one example, the active moieties of the tracrRNA retain the ability to form complexes with Cas proteins (e.g., Cas9 or dCas9). See, for example, WO2014144592. Methods for generating crRNA-tracrRNA hybrid RNAs are known in the art. See, for example, WO2014099750, US 20140179006, and US 20140273226. The entire contents of these documents are incorporated herein by reference. In some embodiments, the tracrRNA activity and the guide sequence are two separate RNA molecules that together form the guide RNA and associated scaffold. In this case, the molecule with tracrRNA activity should be able to interact with the molecule with the guide sequence (typically through base pairing).

[0242] Matching tags or anchoring tags

[0243] A unique feature of ME is the magRNA, which contains a matching tag (also known as an anchor tag) complementary to a fragment in the modular RNA template. In some embodiments, the tag in the magRNA is located at the 3' end of the magRNA. In other embodiments, the tag is located at the 5' end of the magRNA. In some embodiments, the tag and spacer region are located at opposite ends of the RNA scaffold; in some embodiments, the anchor tag and spacer region are located at the same end of the RNA scaffold.

[0244] In some implementations, the label length is 6nt, 9nt, 12nt, 15nt, 18nt, 21nt, and 24nt. In some implementations, the label length is any length between 6nt and 24nt.

[0245] In some implementations, the tag is complementary to the 5' end of the RNA template. In some implementations, the tag matches an internal sequence within the RNA template. In some implementations, the 5' complementary sequence in the RNA template is located 23 nt, 29 nt, 35 nt, 41 nt, 50 nt, or 100 nt upstream (5' direction) of the reverse transcription start site. In some implementations, the complementary sequence in the RNA template sequence is located 20 nt to 10 kb upstream of the reverse transcription start site.

[0246] In some implementations, when using a dual-module ME with two magRNAs, the two tag sequences in the two different magRNAs are complementary to two different fragments in the RNA template.

[0247] In some implementations, when using a single-module ME, matching tags that match sequences 29 nt upstream of the reverse transcription start site show high efficiency.

[0248] In some implementations, two magRNAs with two different tags are used to recruit two separate RNA templates.

[0249] The selection of complementary regions between the magRNA tag and the RNA template follows the Watson-Crick base pairing principle. GC content, length, and melting point are factors to consider when designing strong anchoring magRNA tags.

[0250] Examples of RNA scaffolds and magRNAs are as follows: The sgRNA scaffold coding sequence (T will be transcribed into U in RNA) (spCas9) 5'GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC-3' (SEQ ID NO: 33) Sometimes, "TT" is added to the 3' end to stabilize the magRNA.

[0251] Disassembled gRNA / magRNA sequences (containing crRNA and transRNA)

[0252] crRNA (spCas9): 5'- NNNNNNNNNNNNNNNNNNNN GUUUUAGAGCUAUGCU-3' (N20 is the guide / spacer sequence) (SEQ ID NO: 34)

[0253] tracrRNA-mag(spCas9)

[0254] 5'- AGCAUAGCA AGUUA AAAUA AGGCU AGUCC GUUAU CAACU UGAAA AAGUG GCACCGAGUC GGUGC-MMMMMMMMMMMM-3' (SEQ ID NO: 35) (M12 is the matching sequence). Sometimes UU is added before the first M.

[0255] saCas9 sgRNA scaffold

[0256] 5'GTTTTAGTACTCTGGAAACAGAATCTACTAAAACAAGGCAAAATGCCGTGTTTATCTCGTCAACTTGTTGGCGAGA-3' (SEQ ID NO: 36)

[0257] magRNA matching sequence for EGFP-200L guidance

[0258] Recruiting RNA motifs

[0259] In some implementations, magRNA may contain a recruitment RNA motif that links or recruits reverse transcriptases or effector proteins to a target site.

[0260] One approach to recruiting reverse transcriptases or effector proteins to target sequences is through direct fusion with RNA-guided nickases (e.g., dCas9). Direct fusion of effector proteins to proteins requiring sequence recognition (e.g., dCas9) has been successfully used for sequence-specific transcriptional activation or repression, but this protein-protein fusion design can introduce steric hindrance, which is not ideal for proteins that require the formation of multimeric complexes to exert their activity (e.g., enzymes). In such cases, as... Figure 6 , 7 As shown in 8, 10B, 11B, 12B, 14, and 16, recruiting reverse transcriptases or effector proteins to target sites in a split manner is more advantageous.

[0261] More specifically, this splitting system utilizes various RNA motif / RNA-binding protein binding pairs. For this purpose, magRNAs can be engineered to incorporate RNA motifs (e.g., the MS2 operon motif) that specifically bind RNA-binding proteins (e.g., the MS2 coat protein, MCP). The recruited RNA motif can be fused to any suitable location on the magRNA. For example, it can replace circular structures within the RNA scaffold, particularly tetraloop and / or stem-loop structures.

[0262] Therefore, the RNA scaffold assembly of the magRNA disclosed herein is a designed RNA molecule that contains not only gRNA motifs for specific DNA / RNA sequence recognition (e.g., CRISPR RNA motifs for dCas9 binding) but also recruitment RNA motifs for effector recruitment. In this way, recruited RT or effector protein fusions can be recruited to target sites by binding to the recruitment RNA motifs. Due to the flexibility of RNA scaffold-mediated recruitment, functional monomers, as well as dimers, tetramers, or oligomers, can be formed relatively readily near the target DNA or RNA sequence. These RNA recruitment motif / binding protein pairs can be derived from naturally occurring sources (e.g., RNA phages or yeast telomerases) or can be artificially designed (e.g., RNA aptamers and their corresponding binding protein ligands). Table 2 summarizes a non-exhaustive list of examples of recruitment RNA motif / RNA binding protein pairs that can be used in the systems described herein.

[0263] Table 2. Examples of recruiting RNA motifs that can be used in this disclosure, and their paired RNA-binding proteins / proteins. Structural domain.

[0264]

[0265]

[0266] The sequences of the above binding pairs are listed below.

[0267] 1. Telomerase Ku-binding motif / Ku heterodimer

[0268] a. Ku combined with hair clip

[0269] b. Ku heterodimer

[0270] 2. Telomerase Sm7 binding motif / Sm7 homoheptamer

[0271] a. Sm common sites (single strand)

[0272] 5'-AAUUUUUGGA-3' (SEQ ID NO: 56)

[0273] b. Monomeric Sm-like protein (archaea)

[0274] 3. MS2 phage operon stem loop / MS2 capsid protein

[0275] a. MS2 phage operon stem loop

[0276] 5'-GCGCACAUGAGGAUCACCCAUGUGC-3' (SEQ ID NO: 58)

[0277] b. MS2 capsid protein

[0278] 4. PP7 phage operon stem loop / PP7 capsid protein

[0279] a. PP7 phage operon stem loop

[0280] 5'-AUAAGGAGUUUAUAUGGAAAACCCUUA-3' (SEQ ID NO: 60)

[0281] b. PP7 capsid protein (PCP)

[0282] 5. SfMu Com stem-loop / SfMu Com binding protein

[0283] a. SfMu Com stem ring

[0284] 5'-CUGAAUGCCUGCGAGCAUC-3' (SEQ ID NO: 62)

[0285] b.SfMu Com binding protein

[0286] 6. sgRNA scaffold (spCas9) with MS2 inserted in the stem-loop.

[0287] 7. sgRNA scaffold (saCas9) with MS2 inserted in the stem-loop.

[0288] RNA template molecule

[0289] A key feature of the ME system is the separation of template RNA and gRNA, as well as the freedom in template RNA design. The modular RNA template only needs to include a short polynucleotide sequence at its 3' end that is complementary to a nick DNA strand with a free 3' end; the nick DNA can then be used as a primer for the synthesis of new DNA strands. Primer binding sites (PBSs) are typically 9 to 18 nucleotides in length. Primer binding sites longer than 18 nucleotides are also acceptable.

[0290] Because the tag sequence in the magRNA depends on the RNA template sequence, rather than the other way around, there is no need to consider adding any exogenous sequences to the template for recruitment. The purpose of this design is solely to introduce the target novel sequence alteration to the target site in the genome. In some embodiments, the RNA template used to introduce point mutations contains point mutations (transversions or inversions). In some embodiments, the RNA template used to introduce deletions contains deletions. In some embodiments, the RNA template used to introduce insertions contains insertions. In some embodiments, alterations of 1 nt, 4 nt, 29 nt, and 52 nt are introduced, including point mutations, deletions, and insertions. Nucleotide alterations of other lengths between 1 nt and 52 nt can also be made. In some embodiments, mutations, insertions, and deletions in sizes from 53 nt to 100 nt are introduced. In some embodiments, the RNA template is programmed to implement insertions and deletions in sizes from 101 nt to 50 kb.

[0291] In some implementations, the distance between the sequence alteration and the nick site is 1 nt, 4 nt, 20 nt, 35 nt, 50 nt, or 100 nt. Other distances between 1 nt and 100 nt can also be chosen. In some implementations, the distance is between 100 nt and 1000 nt.

[0292] In some embodiments, RNA templates of 32 nt, 41 nt, 47 nt, 72 nt, 131 nt, and 135 nt are effective for ME. In some embodiments, the size of the modular RNA template is from 15 nt to 200 nt. In some embodiments, the size of the RNA template is from 201 nt to 2 kb. In some embodiments, the size of the RNA template is from 2 kb to 10 kb.

[0293] Typically, RNA templates contain one or more sequences that are identical or homologous to sequences near the target site to promote cell repair and improve editing efficiency.

[0294] The gene editing complex or system described herein can transcribe an RNA sequence template into a host target DNA site via target-initiated reverse transcription. By directly writing the DNA sequence into the host genome through reverse transcription of the RNA sequence template, this gene editing complex or system can insert the target sequence into the target genome without introducing a foreign DNA sequence into the host cell, and can eliminate the need for foreign DNA insertion. Therefore, this gene editing complex or system provides a platform for using custom RNA sequence templates containing the target sequence (e.g., sequences containing heterologous gene coding and / or functional information).

[0295] In some embodiments, the RNA template may contain an open reading frame or its reverse complementary sequence. In some embodiments, the RNA template may contain a sequence encoding a bacterial or viral antigenic epitope. In some embodiments, the RNA template may contain a sequence encoding an antigen-recognition variable region of an antibody. In some embodiments, the RNA template may contain a sequence encoding a T-cell receptor (TCR) antigen-recognition variable region. In some embodiments, the RNA template may contain a signal peptide sequence for secretion, a nuclear localization signal sequence for nuclear transport, or a peptide sequence for intracellular transport. In some embodiments, the RNA template may be converted into double-stranded DNA (e.g., by reverse transcription) before the open reading frame is transcribed and translated. In some embodiments, the RNA may contain a sequence homologous to a DNA target site.

[0296] In some implementations, RNA templates can be identified, designed, engineered, and constructed to contain sequences that alter or specify host genome functions, such as by introducing heterologous coding regions into the genome; affecting or causing exon structure / alternative splicing; causing disruption of endogenous genes; causing transcriptional activation of endogenous genes; causing epigenetic regulation of endogenous DNA; or causing upregulation or downregulation of operable linker genes. In some implementations, the RNA template can be engineered to contain sequences encoding exons and / or transgenes, providing binding sites and combinations thereof with transcription factor activators, repressors, enhancers, etc. In other implementations, the coding sequence can be further customized via splice acceptor sites and poly-A tails.

[0297] The RNA template may have some homology with the target DNA. In some embodiments, the RNA template has complete homology with the target DNA at the 3′ end of the RNA template by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 175, 200 or more bases. In some embodiments, the RNA template has at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 100% homology with the target DNA at, for example, the 5′ end of the RNA template, of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 175, 180, or 200 or more bases.

[0298] The RNA template components described herein are typically capable of binding to magRNA. In some embodiments, the RNA template has a region capable of binding to magRNA. This region can be located anywhere suitable within the RNA template.

[0299] RNA templates typically contain a target sequence for insertion into target DNA. This target sequence can be a coding or non-coding sequence. In some embodiments, the systems or methods described herein comprise a single RNA template. In some embodiments, the systems or methods described herein comprise multiple RNA templates.

[0300] In some implementations, the template contains mutations in the PAM or protospacer sequence required to alter the magRNA or gRNA. Therefore, while the CRISPR complex can edit the original DNA sequence, the edited DNA sequence is not recognized by the CRISPR complex, thus avoiding ineffective repeated editing. In some implementations, the template contains mutations to generate new PAM or protospacer sequences, allowing the CRISPR complex to target only the edited sequence as needed, without targeting the unedited sequence.

[0301] In some embodiments, the target sequence may contain an open reading frame. In some embodiments, the RNA template has a Kozak sequence. In some embodiments, the RNA template has an internal ribosome entry site. In some embodiments, the RNA template has a self-cleaving peptide, such as a T2A or P2A site. In some embodiments, the RNA template has a start codon. In some embodiments, the RNA template has a splice acceptor site. In some embodiments, the RNA template has a splice donor site. In some embodiments, the RNA template has a stop codon. In some embodiments, the RNA template has a microRNA binding site downstream of the stop codon. In some embodiments, the RNA template has a polyA tail downstream of the stop codon in the open reading frame. In some embodiments, the RNA template contains one or more exons. In some embodiments, the RNA template contains one or more introns. In some embodiments, the RNA template contains a eukaryotic transcription terminator. In some embodiments, the RNA template contains an enhancing translation element or a translation-enhancing element.

[0302] In some embodiments, the RNA template comprises the coding sequence of a functional small non-protein coding RNA. Examples of these functional non-coding RNAs (ncRNAs) include transfer RNA (tRNA), ribosomal RNA (rRNA), and small RNAs such as microRNA, siRNA, piRNA, snoRNA, snRNA, exRNA, and scaRNA. In some embodiments, these non-coding RNAs are engineered small regulatory RNAs that enhance or inhibit the function of their endogenous ncRNA counterparts.

[0303] In some embodiments, the target sequence may comprise a non-coding sequence. For example, the RNA template may comprise a promoter or enhancer sequence. In some embodiments, the RNA template comprises a tissue-specific promoter or enhancer, each promoter or enhancer being unidirectional or bidirectional. In some embodiments, the promoter is an RNA polymerase I promoter, an RNA polymerase II promoter, or an RNA polymerase III promoter. In some embodiments, the promoter comprises a TATA element. In some embodiments, the promoter has one or more transcription factor binding sites.

[0304] In some embodiments, the RNA template may contain a promoter sequence, such as a tissue-specific promoter or enhancer sequence. In some embodiments, the tissue-specific promoter or enhancer is used to enhance the target cell specificity of the gene. For example, a promoter or enhancer may be selected based on the following criteria: it is active in target cell types but inactive (or has low activity) in non-target cell types. Therefore, even if the promoter or enhancer is integrated into the genome of a non-target cell, it will not drive the expression of the integrated gene (or will only drive low-level expression). Systems containing tissue-specific promoter or enhancer sequences in the RNA template can also be used in combination with, for example, microRNA binding sites in the RNA template, as described herein. In some embodiments, the RNA template may contain a microRNA sequence, siRNA sequence, guide RNA sequence, or piwi RNA sequence. In some embodiments, tissue-specific silencer or repressor sequences are used to silence or repress the expression of target genes in target cells and tissues.

[0305] In some embodiments, the RNA template may contain sites that coordinate epigenetic modifications. In some embodiments, the RNA template contains elements that repress (e.g., prevent) epigenetic silencing. In some embodiments, the RNA template contains chromatin insulators. For example, the RNA template may contain CTCF sites or sites targeted by DNA methylation.

[0306] To promote higher or more stable gene expression, the RNA template may contain features that prevent or inhibit gene silencing. In some embodiments, these features prevent or inhibit DNA methylation. In some embodiments, the RNA template contains sequences that mediate transcriptional repression. In some embodiments, these features promote DNA demethylation. In some embodiments, these features prevent or inhibit histone deacetylation. In some embodiments, these features prevent or inhibit histone methylation. In some embodiments, these features promote histone acetylation. In some embodiments, these features promote histone demethylation. In some embodiments, multiple features may be introduced into the RNA template to promote one or more such modifications. CpG dinucleotides are susceptible to methylation by host methyltransferases. In some embodiments, the CpG dinucleotides in the RNA template are depleted, for example, they do not contain CpG nucleotides, or their CpG dinucleotide content is reduced compared to the corresponding unaltered sequence. In some embodiments, the CpG dinucleotides in the promoter driving transgene expression that integrates DNA are depleted.

[0307] In some embodiments, the RNA template contains a gene expression unit comprising at least one regulatory region operatively linked to an effector sequence. The effector sequence may be a sequence transcribed into RNA (e.g., a coding sequence or a non-coding sequence, such as a sequence encoding microRNA).

[0308] In some embodiments, the target sequence of the RNA template can be, for example, 50-50,000 base pairs (e.g., 50-40,000 bp, 500-30,000 bp, 500-20,000 bp, 100-15,000 bp, 500-10,000 bp, 50-10,000 bp, or 50-5,000 bp). In some embodiments, the heterologous target sequence is less than 1,000, 1,300, 1,500, 2,000, 3,000, 4,000, 5,000, or 7,500 nucleotides in length.

[0309] Modification

[0310] The RNA molecules described herein (e.g., magRNA or RNA template) may contain one or more modifications. Such modifications may include the introduction of at least one non-naturally occurring nucleotide or a modified nucleotide or its analogue. Modified nucleotides may be modified at the ribose, phosphate, and / or base moieties. Modified nucleotides may include 2'-O-methyl analogues, 2'-deoxy analogues, or 2'-fluoro analogues. The nucleic acid backbone may also be modified; for example, a phosphate thioester backbone may be used. Locked nucleic acids (LNAs) or bridging nucleic acids (BNAs) may also be used. Other examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromouridine, pseudouridine, inosine, and 7-methylguanosine. These modifications may be applied to any component of the system described herein. In a preferred embodiment, these modifications may be applied to RNA components, such as guide RNA sequences.

[0311] In some implementations, the RNA molecule or its sub-parts may contain one or more modifications, such as base modifications, backbone modifications, etc., to endow the nucleic acid with new or enhanced properties (e.g., improved stability).

[0312] Linkage between the modified backbone and the modified nucleosides

[0313] Examples of suitable nucleic acids containing modifications include those with a modified backbone or non-natural nucleoside linkages. Nucleic acids with a modified backbone include those with a phosphorus atom retained in the backbone and those without a phosphorus atom in the backbone.

[0314] Suitable modified oligonucleotide backbones containing phosphorus atoms include, for example, thiophosphates, chiral thiophosphates, dithiophosphates, phosphate triesters, aminoalkyl phosphate triesters, methyl and other alkylphosphonates including 3'-alkylene phosphonates, 5'-alkylene phosphonates and chiral phosphonates, phosphinates, phosphoramides including 3'-aminophosphatamides and aminoalkylphosphatamides, phosphoridamides, thiophosphatamides, thioalkylphosphonates, thioalkyl phosphate triesters, selenophosphates and borophosphates (with normal 3'-5' linkages), 2'-5' linkage analogs of these compounds, and compounds with reverse polarity, wherein one or more nucleotide linkages are 3'-3', 5'-5', or 2'-2' linkages. Suitable oligonucleotides with reverse polarity contain a single 3'-3' linkage at the nucleotide linkage at the 3' end. , This refers to a single inverted nucleoside residue, which can be basic (with a nucleobase deletion or substitution by a hydroxyl group). Various salts (e.g., potassium or sodium salts), mixed salts, and free acid forms are also included.

[0315] In some embodiments, the target nucleic acid comprises one or more thiophosphate and / or heteroatom nucleoside links, particularly —CH2—NH—O—CH2—, —CH2—N(CH3)—O—CH2- (referred to as a methylene (methylimino) or MMI backbone), —CH2—O—N(CH3)—CH2—, —CH2—N(CH3)—N(CH3)—CH2—, and —O—N(CH3)—CH2—CH2— (wherein the native phosphodiester nucleoside link is represented as —O—P(═O)(OH)—O—CH2—). MMI-type nucleoside links are disclosed in the above-cited U.S. Patent No. 5,489,677. Suitable amide nucleoside links are disclosed in U.S. Patent No. 5,602,240.

[0316] Nucleic acids with a morpholine backbone structure are also applicable, such as those described in U.S. Patent No. 5,034,506. 。 For example, in some embodiments, the target nucleic acid comprises a six-membered morpholine ring to replace the ribose ring. In some of these embodiments, a phosphoryldiamine or other non-phosphodiester nucleoside linker replaces the phosphodiester linker.

[0317] Suitable modified polynucleotide backbones that do not contain phosphorus atoms have backbones formed by linkages between short-chain alkyl or cycloalkyl nucleosides, linkages between mixed heteroatoms and alkyl or cycloalkyl nucleosides, or linkages between one or more short-chain heteroatoms or heterocyclic nucleosides. These include those having: morpholino linkages (partially formed by the sugar moiety of the nucleoside); siloxane backbones; thioether, sulfoxide, and sulfone backbones; methylacetyl and thiomethylacetyl backbones; methylenemethylacetyl and thiomethylacetyl backbones; riboacetyl backbones; olefin-containing backbones; aminosulfonate backbones; methyleneimino and methylenehydrazine backbones; sulfonate and sulfonamide backbones; amide backbones; and others having mixed N, O, S, and C. The skeleton of the components.

[0318] Simulation

[0319] The target nucleic acids disclosed herein (e.g., RNA templates, magRNAs, etc.) can be nucleic acid mimics. When the term "mimic" is used for polynucleotides, it is intended to include polynucleotides in which only the furanose ring or both the furanose ring and the nucleotide linker are replaced by non-furanose groups; the replacement of only the furanose ring is also referred to in the art as a sugar substitute. The heterocyclic base moiety or modified heterocyclic base moiety is retained to hybridize with a suitable target nucleic acid. One such nucleic acid (polynucleotide mimic) that has been shown to have excellent hybridization performance is called a peptide nucleic acid (PNA). In PNAs, the sugar backbone of the polynucleotide is replaced by an amide-containing backbone (particularly an aminoethylglycine backbone). The nucleotide is retained and binds directly or indirectly to the nitrogen atom of the amide moiety of the backbone.

[0320] One type of polynucleotide mimic that has been reported to have excellent hybridization properties is peptide nucleic acid (PNA). The backbone of a PNA compound is two or more linked aminoethylglycine units, giving PNA an amide-containing backbone. The heterocyclic base moiety is directly or indirectly bonded to the aza-nitrogen atom of the amide moiety of the backbone. Representative U.S. patents describing methods for preparing PNA compounds include, but are not limited to, U.S. Patent Nos. 5,539,082, 5,714,331, and 5,719,262.

[0321] In some implementations, a DNA template is used instead of an RNA template, and a DNA polymerase is used instead of a reverse transcriptase.

[0322] Another class of polynucleotide mimics studied is based on linked morpholine units (morpholine nucleic acids), with heterocyclic bases linked to the morpholine ring. Various linker groups have been reported for linking morpholine monomer units in morpholine nucleic acids. One class of linker groups has been selected for constructing nonionic oligomers. Nonionic morpholine oligomers are less likely to interact undesirably with cellular proteins. Morpholinyl polynucleotides are nonionic mimics of oligonucleotides, which are less likely to interact undesirably with cellular proteins (Dwaine A. Braasch and David R. Corey, Biochemistry, 2002, 41(14), 4503-4510). Morpholinyl polynucleotides are disclosed in US Patent No. 5,034,506. Various morpholinoyl polynucleotide compounds with different linker groups have been prepared for linking monomer subunits.

[0323] Another class of polynucleotide mimics is called cyclohexenyl nucleic acid (CeNA). The furanyl ring, typically present in DNA / RNA molecules, is replaced by a cyclohexenyl ring. CeNA DMT-protected phosphoramide monomers have been prepared and used to synthesize oligomers according to classical phosphoramide chemistry methods. Fully modified CeNA oligomers, as well as oligonucleotides modified with CeNA at specific positions, have been prepared and studied (see Wang). et al (J. Am. Chem. Soc., 2000, 122, 8595-8602). Typically, incorporating CeNA monomers into DNA strands can improve the stability of DNA / RNA hybrids. Complexes formed by CeNA oligoadenylates with complementary sequences of RNA and DNA exhibit similar stability to native complexes. NMR and circular dichroism studies have shown that conformational adaptation can be easily achieved by incorporating CeNA structures into native nucleic acid structures.

[0324] Another modification involves nucleosyl nucleotides (LNAs), in which a 2′-hydroxyl group is attached to the 4′ carbon atom of the sugar ring, forming a 2′-C,4′-C-oxomethylene link, thereby forming a bicyclic sugar structure. This link can be a methylene (—CH2—) group bridging the 2′ oxygen atom and the 4′ carbon atom, where n is 1 or 2 (Singh et al (Chem. Commun., 1998, 4, 455-456). LNA and LNA analogs exhibit extremely high double-stranded thermal stability (Tm = +3 to +10°C), stability against 3′-exonuclease degradation, and good solubility with complementary DNA and RNA. Highly efficient and non-toxic antisense oligonucleotides containing LNA (Wahlestedt) have been described. et al ., Proc. Natl. Acad. Sci. USA, 2000, 97, 5633-5638).

[0325] The synthesis and preparation of LNA monomers adenine, cytosine, guanine, 5-methylcytosine, thymine, and uracil, as well as their oligomerization and nucleic acid recognition properties, have been described (Koshkin). et al., Tetrahedron, 1998, 54, 3607-3630). LNA and its preparation methods are also described in WO 98 / 39352 and WO 99 / 14226.

[0326] Modified sugar portion

[0327] The target nucleic acid may also contain one or more substituted sugar moieties. Suitable polynucleotides contain sugar substituents selected from the following: OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-ynyl; or O-alkyl-Co-alkyl, wherein the alkyl, alkenyl, and ynyl groups may be substituted or unsubstituted C1 to C2. 10 Alkyl or C2 to C 10 Alkenyl and ynyl groups. Particularly suitable are O((CH2)). n O) m CH3, O(CH2) n OCH3, O(CH2) n NH2, O(CH2) n CH3, O(CH2) n ONH2 and O(CH2) n ON((CH2) n CH3)2, where n and m are 1 to approximately 10. Other suitable polynucleotides include a sugar substituent selected from C1 to C2. 10Lower alkyl groups, substituted lower alkyl groups, alkenyl groups, alkynyl groups, aralkyl groups, O-alkaneyl groups or O-aralkyl groups, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocyclic alkyl groups, heterocyclic alkaneyl groups, aminoalkylamino groups, polyalkylamino groups, substituted silyl groups, RNA cleaving groups, reporter groups, intercalating agents, groups used to improve the pharmacokinetic properties of oligonucleotides or groups used to improve the pharmacodynamic properties of oligonucleotides, and other substituents with similar properties. Suitable modifications include 2′-methoxyethoxy (2′—O—CH2CH2OCH3, also known as 2′-O-(2-methoxyethyl) or 2′-MOE) (Martin) et al (Helv. Chim. Acta, 1995, 78, 486-504), i.e., alkoxyalkoxy group. Further suitable modifications include 2′-dimethylaminooxyethoxy, i.e., O(CH2)2ON(CH3)2 group, also known as 2′-DMAOE, as described in the examples below, and 2′-dimethylaminoethoxyethoxy (also known in the art as 2′-O-dimethyl-amino-ethoxy-ethyl or 2′-DMAEOE), i.e., 2′—O—CH2—O—CH2—N(CH3)2.

[0328] Other suitable sugar substituents include methoxy (-O-C) —aminopropoxy (—O CH2CH2NH2), allyl (—CH2—CH═CH2), —O-allyl (CH2—CH═CH2), and fluorine (F). The 2′-sugar substituent can be located at either the arabinose (top) or ribose (bottom) position. A suitable 2′-arabinose modification is 2′-F. Similar modifications can also be made at other positions in oligomers, particularly at the 3′ position of the sugar in 3′-terminal nucleotides or 2′-5′-linked oligonucleotides, and at the 5′ position of the 5′-terminal nucleotide. Oligomers can also be replaced with sugar mimics such as the cyclobutyl moiety instead of pentofuranose.

[0329] Base modification and substitution

[0330] The target nucleic acid may also include nucleobase (generally referred to in the art as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases include purine bases adenine (A) and guanine (G), and pyrimidine bases thymine (T), cytosine (C), and uracil (U). Modified nucleobases include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (—C═C—CH3)uracil and cytosine, and alkynyl derivatives of other pyrimidine bases, 6-azouracil. Cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halogenated, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxy and other 8-substituted adenine and guanine, 5-halogenated especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracil and cytosine, 7-methylguanine and 7-methyladenine, 2-fluoroadenine, 2-aminoadenine, 8-nitroguanine and 8-nitroadenine, 7-denitroguanine and 7-denitroadenine, and 3-denitroguanine and 3-denitroadenine. Further modified nucleobases include tricyclic pyrimidines, such as phenoxazincytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenthiazincytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamp sequences such as substituted phenoxazincytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazolecytidine (2H-pyrimido(4,5-b)indol-2-one), and pyridoindolcytidine (H-pyrido(3′,2′:4,5)pyrrolo(2,3-d)pyrimido-2-one).

[0331] The heterocyclic base moiety may also include those where the purine or pyrimidine base is substituted by another heterocycle, such as 7-deadenine, 7-deadenine, 2-aminopyridine, and 2-pyridone. Other nucleobases include those disclosed in U.S. Patent No. 3,687,808, *The Concise Encyclopedia of Polymer Science and Engineering*, pages 858-859, Kroschwitz, JI, ed. John Wiley & Sons, 1990, and Englisch. et alThose disclosed in Angewandte Chemie, International Edition, 1991, 30, 613, and those disclosed in Sanghvi, YS, Chapter 15, Antisense Research and Applications, pages 289-302, Crooke, ST and Lebleu, B., ed., CRC Press, 1993. Some of these nucleobases can be used to increase the binding affinity of oligomers. These include 5-substituted pyrimidines, 6-azines, and N-2, N-6, and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-methylcytosine substitution has been shown to increase the stability of nucleic acid duplexes by 0.6–1.2 °C (Sanghvi). et al (,eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp.276-278), and is a suitable base substitution, for example, when combined with 2'-O-methoxyethyl sugar modification.

[0332] Expression System

[0333] To use the aforementioned complexes, systems, or platforms, it may be necessary to express one or more of the described protein and RNA components from their encoded nucleic acids. This can be achieved in various ways. For example, the nucleic acid encoding RNA or a protein can be cloned into one or more intermediate vectors for introduction into prokaryotic or eukaryotic cells for replication and / or transcription. Intermediate vectors are typically prokaryotic vectors (e.g., plasmids or shuttle vectors) or insect vectors, used to store or manipulate the nucleic acid encoding RNA or a protein to produce RNA or a protein. The nucleic acid can also be cloned into one or more expression vectors for application to plant cells, animal cells (preferably mammalian or human cells), fungal cells, bacterial cells, or protozoan cells. Therefore, this disclosure provides nucleic acids encoding any of the aforementioned RNAs or proteins. Preferably, the nucleic acid is isolated and / or purified.

[0334] This disclosure also provides recombinant constructs or vectors having sequences encoding one or more of the aforementioned RNAs or proteins. Examples of constructs include vectors (e.g., plasmids or viral vectors) in which the nucleic acid sequences of this disclosure have been inserted, either forward or reverse. In a preferred embodiment, the construct further comprises a regulatory sequence (including a promoter) operably linked to said sequence. A large number of suitable vectors and promoters are known to those skilled in the art, and these vectors and promoters are commercially available. Cloning and expression vectors suitable for prokaryotic and eukaryotic hosts have also been developed, for example, in Sambrook. et al. It is described in (2001, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press).

[0335] A vector is a nucleic acid molecule capable of transporting another nucleic acid linked to it. Vectors can autonomously replicate or integrate into host DNA. Examples of vectors include plasmids, granules, or viral vectors. The vectors disclosed herein contain nucleic acids in a form suitable for expression in host cells. Preferably, the vector contains one or more regulatory sequences effectively linked to the nucleic acid sequence to be expressed. "Regulatory sequences" include promoters, enhancers, and other expression control elements (e.g., polyadenylation signals). Regulatory sequences include both constitutive regulatory sequences that guide constitutive expression of the nucleotide sequence and inducible regulatory sequences. The design of the expression vector can depend on factors such as the selection of host cells to be transformed, transfected, or transduced, and the desired expression level of the RNA or protein.

[0336] Examples of expression vectors include chromosomal DNA, non-chromosomal DNA and synthetic DNA sequences, bacterial plasmids, bacteriophage DNA, baculoviruses, yeast plasmids, vectors derived from combinations of plasmids and bacteriophage DNA, and viral DNA such as vaccinia virus, adenovirus, fowlpox virus, and pseudorabies virus. However, any other vector may be used, provided that it can replicate and survive in a host. Suitable nucleic acid sequences can be inserted into the vector using various methods. Typically, a nucleic acid sequence encoding one of the aforementioned RNAs or proteins can be inserted into a suitable restriction endonuclease site using methods known in the art. These methods and related subcloning methods are within the scope of those skilled in the art.

[0337] The vector may contain a suitable sequence for amplifying expression. Furthermore, the expression vector preferably contains one or more selectable marker genes to provide phenotypic characteristics for screening transformed host cells, such as dihydrofolate reductase or neomycin resistance in eukaryotic cell cultures, or tetracycline or ampicillin resistance, for example, in Escherichia coli.

[0338] Vectors for expressing RNA can contain RNA Pol III promoters to drive RNA expression, such as the HI, U6, or 7SK promoters. Vectors for expressing RNA can contain Pol I promoters to drive template RNA expression, such as the influenza virus promoter. These human promoters enable RNA to be expressed in mammalian cells after vector introduction. Alternatively, the T7 promoter can be used for, for example, in vitro transcription, in which the RNA can be transcribed and purified in vitro.

[0339] Vectors containing the aforementioned suitable nucleic acid sequences and suitable promoters or control sequences can be used to transform, transfect, or infect suitable hosts, thereby enabling the host to express the aforementioned RNA or protein. Examples of suitable expression hosts include bacterial cells (e.g., *Escherichia coli*, *Streptomyces*, *Salmonella typhimurium*), fungal cells (yeast), and insect cells (e.g., fruit flies and fall armyworms). Spodoptera frugiperda (Sf9)), animal cells (e.g., CHO, COS, and HEK 293), adenoviruses, and plant cells. The selection of a suitable host is within the scope of those skilled in the art. In some embodiments, this disclosure provides methods for producing the aforementioned RNA or protein by transforming, transfecting, or infecting host cells with an expression vector having a nucleotide sequence encoding one of RNA, a polypeptide, or a protein. The host cells are then cultured under suitable conditions to express the RNA or protein.

[0340] Any method known in the art for introducing exogenous nucleotide sequences and proteins into host cells may be used. Examples include transfection with calcium phosphate, polybrene, protoplast fusion, electroporation, nuclear transfection, liposomes, microinjection, naked DNA, plasmid vectors, viral vectors (including free and integrated types), and any other known method for introducing cloned genomic DNA, cDNA, synthetic DNA, or other exogenous genetic material into host cells.

[0341] In some implementations, components of the matching editing system may be delivered in the form of RNA (e.g., mRNA, magRNA, and template RNA encoding proteins), DNA (DNA expression vector), or protein (e.g., purified CRISPR-RT fusion protein), or any combination of the above forms (e.g., ribonucleoprotein complex).

[0342] Any method known in the art for the systematic delivery of exogenous nucleotide sequences and proteins to a host may be used. Examples include the use of viral vectors, including free and integrated viral vectors such as adeno-associated virus (AAV), adenovirus vector (AD), lentiviral vectors, retroviral vectors, integration-deficient lentiviral vectors, herpes simplex virus vectors (HSV), stomatitis virus vectors (VSV), modified vaccinia virus ankara vector (MVA), arenavirus vectors, Sendai virus vectors, and parvovirus vectors. Examples include the use of non-viral vectors such as liposomes, lipid particles, lipid nanoparticles, and genomeless versions of virus-like particles (VLPs).

[0343] method

[0344] Another aspect of this disclosure includes a method for modifying a target DNA sequence (e.g., a chromosomal sequence) or a target RNA sequence in a cell, embryo, human, or non-human animal. Sequence modification can be a base alteration, insertion, deletion, or a combination of these modifications. In one embodiment, the embryo is a non-human animal embryo. The method includes introducing the aforementioned (A) RNA-guided nickase; (B) reverse transcriptase; (C) RNA template molecule and (D) magRNA molecule into the cell or embryo. The magRNA guides other components to a target polynucleotide at the target site, and the reverse transcriptase writes the sequence into the target site according to the sequence of the RNA template molecule described herein. The RNA template contains the desired nucleotide sequence modification information.

[0345] The target polynucleotide sequence referred to in this article is the sequence at which an RNA-guided nicking enzyme produces a single-strand DNA break or nick. There are no sequence restrictions for the target polynucleotide nicking site, but if a CRISPR-type nicking enzyme is used, the sequence is immediately followed (downstream or at the 3' end) by a PAM sequence. Examples of PAM sequences include, but are not limited to, NG, NGG, NGNG, and NNAGAAW (where N is defined as any nucleotide and W as A or T). Other examples of PAM sequences have been given above, and those skilled in the art will be able to identify more PAM sequences suitable for a given nicking enzyme (e.g., a CRISPR protein). In some examples, RNA-guided nicking enzymes may not require a PAM motif (e.g., an engineered CRISPR protein without PAM). Target nucleotide modifications can be located in the coding region of a gene, introns of a gene, transcriptional regulatory regions of a gene, intergenetic control regions, etc. The gene can be a protein-coding gene or an RNA-coding gene.

[0346] The required nucleotide sequence alteration and its location are determined by the designed RNA template. This RNA template is an independent RNA molecule that is not covalently linked to the guide RNA. The RNA template can be freely designed according to the desired sequence alteration without concern for interfering with the structure of the guide RNA. There are no distance requirements between the desired DNA modification and the target nick site. High editing efficiency is observed at any distance within the range of 200 nucleotides. There are no content requirements for the RNA template design; only a primer-binding sequence complementary to the nick DNA strand needs to be included at the 3' end. No additional exogenous sequences for recruitment are required in the RNA template. This freedom in RNA template design mechanistically eliminates the root cause of the “scarring” often introduced at genome editing sites when using other reverse transcriptases, recombinases, integrases, and transposases.

[0347] After designing an RNA template for the desired DNA modification at the target site, the next step is to design an RNA tag complementary to a sequence fragment within the RNA template. This tag is added to the gRNA, typically at the 3' end, to recruit the RNA template and anchor it to the target cleavage site. The RNA tag can be 6 to 24 nucleotides in length. The complementary site within the RNA template to the tag RNA can be located at the 5' end or inside the RNA template.

[0348] In another example, this disclosure provides a system and related method for recruiting one or more additional, different functional effectors to the same target sequence. These effectors work synergistically to facilitate gene conversion. Examples may include proteins that promote the editing process or improve editing efficiency, such as chromatin-modifying enzymes or inhibitors of mismatch repair enzymes. Examples of effectors may be human RNase repressor protein (RNH1), dominant or negative MMR pathway proteins (including MLH1 protein), 5' DNA nuclease Fen1 protein, or similar molecules capable of improving editing efficiency.

[0349] Example of a human RNase inhibitor-MCP fusion protein sequence (NLS-RNH1-linker 25-MCP):

[0350] Therefore, the matching editing system provides a highly versatile method for editing intracellular polynucleotides. The cell can be a single-celled organism, a cell isolated from a human or non-human organism, a cell derived from or engineered from a human or non-human organism, or a cell present in a human or non-human organism.

[0351] The target polynucleotide to be edited can be any polynucleotide, whether endogenous or exogenous to the cell. For example, the target polynucleotide can be a DNA molecule within the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide). The DNA molecule can be endogenous to the cell, or a DNA molecule infected with a virus or microorganism.

[0352] The protein components of the system disclosed herein can be introduced into cells or embryos in the form of isolated proteins. Alternatively, the components can also be introduced via nucleic acids (e.g., DNA or RNA, such as in vitro transcribed RNA) encoding these components. In one embodiment, each protein may include at least one cell-penetrating domain that facilitates protein uptake by the cell. In other embodiments, mRNA or DNA molecules encoding the protein can be introduced into cells or embryos. Typically, the DNA sequence encoding the protein is operatively linked to a promoter sequence that functions in the target cell or embryo. The DNA sequence may be linear, or it may be part of a vector. In other embodiments, the protein can be introduced into cells or embryos as an RNA-protein complex comprising the aforementioned protein, an RNA template, and magRNA.

[0353] In other embodiments, the protein-coding DNA may also contain one or more sequences encoding RNA components (e.g., magRNA and / or RNA template). Typically, the protein- and RNA-coding DNA sequences are operatively linked to appropriate promoter control sequences that allow expression of the protein and RNA, respectively, in a cell or embryo. The protein- and RNA-coding DNA sequences may also contain one or more additional expression control, regulation, and / or processing sequences. The protein- and RNA-coding DNA sequences may be linear or may be part of a vector.

[0354] In some embodiments, when RNA is introduced into a cell via a DNA molecule encoding that RNA, the RNA-coding sequence can be operatively linked to a promoter control sequence for expressing that RNA in eukaryotic cells. For example, the RNA-coding sequence can be operatively linked to a promoter sequence recognized by RNA polymerase III (Pol III). The RNA-coding sequence can be operatively linked to a promoter sequence recognized by RNA polymerase I (Pol I). Examples of suitable Pol III promoters include, but are not limited to, mammalian U6 or H1 promoters. In an exemplary embodiment, the RNA-coding sequence is linked to a mouse or human U6 promoter. In other exemplary embodiments, the RNA-coding sequence is linked to a mouse or human H1 promoter. In other embodiments, the RNA-coding sequence is linked to a viral Pol I promoter (e.g., the influenza virus Pol I promoter).

[0355] The DNA molecules encoding proteins and / or RNA described herein can be linear or circular. In some embodiments, the DNA sequence may be part of a vector, such as a polycistronic vector. Suitable vectors include plasmid vectors, bacteriophages, granules, artificial / miniature chromosomes, transposons, and viral vectors. In one exemplary embodiment, the DNA encoding proteins and / or RNA is contained within a plasmid vector. Non-limiting examples of suitable plasmid vectors include pUC, pBR322, pET, pBluescript, and variants thereof. The vector may contain additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), optional marker sequences (e.g., antibiotic resistance genes), origin of replication, etc.

[0356] The protein components (or nucleic acids encoding them) and RNA components (or DNA encoding them) of the systems disclosed herein can be introduced into cells or embryos in a variety of ways. In one embodiment, the embryo is a non-human animal embryo. Typically, the embryo is a fertilized single-cell embryo of the target species. In some embodiments, the cells or embryos are transfected. Suitable transfection methods include calcium phosphate-mediated transfection, nuclear transfection (or electroporation), cationic polymer transfection (e.g., DEAE-glucan or polyethyleneimine), viral transduction, virion transfection, viral particle transfection, liposome transfection, cationic liposome transfection, immunoliposome transfection, non-liposomal lipid transfection, dendritic polymer transfection, heat shock transfection, magnetic transfection, lipofection, gene gun delivery, puncture transfection, acoustic perforation, optical transfection, gold nanoparticle-mediated transfection, and nucleic acid uptake enhanced by proprietary reagents. Transfection methods are well known in the art (see, for example, "Current Protocols in Molecular Biology" Ausubel). et al. (John Wiley & Sons, New York, 2003 or "Molecular Cloning: A Laboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001). In other embodiments, the molecule is introduced into cells or embryos via microinjection. For example, the molecule can be injected into the pronucleus of a single-celled embryo.

[0357] The protein components (or nucleic acids encoding them) and RNA components (or DNA encoding them) of the systems disclosed herein can be introduced into cells or embryos simultaneously or sequentially. The ratio of protein (or its encoding nucleic acid) to RNA (or DNA encoding RNA) is typically approximating a stoichiometric ratio, thereby enabling the formation of an RNA-protein complex. Similarly, the ratio of two different proteins (or their encoding nucleic acids) is also approximating a stoichiometric ratio. In one embodiment, the protein components and RNA components (or the DNA sequences encoding them) are co-delivered in the same nucleic acid or vector.

[0358] The method further includes maintaining cells or embryos under appropriate conditions such that guide RNA guides effector proteins to target sites in the target sequence, and the effector domain modifies the target sequence.

[0359] Generally, cells can be maintained under conditions suitable for cell growth and / or maintenance. Suitable cell culture conditions are well-known in the art, for example, in Current Protocols in Molecular Biology, Auspitz. et al. , John Wiley & Sons, New York, 2003 or "Molecular Cloning: ALaboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001), Santiago et al. (2008) PNAS 105:5809-5814;Moehle et al. (2007) PNAS 104:3055-3060; Urnov et al. (2005) Nature 435:646-651; and Lombardo et al. (2007) Nat. Biotechnology 25:1298-1306. Those skilled in the art will understand that cell culture methods are known in the art and vary depending on the cell type. In all cases, routine optimization can be performed to determine the best technique for a particular cell type.

[0360] Embryos can be cultured in vitro (e.g., cell culture). Typically, embryos are cultured in suitable temperatures and media, maintaining the necessary O2 / CO2 ratio to promote the expression of protein and RNA scaffolds (if necessary). Non-limiting examples of suitable media include M2, M16, KSOM, BMOC, and HTF media. Those skilled in the art understand that culture conditions will vary depending on the embryo species. In all cases, routine optimization can be performed to determine the optimal culture conditions for a specific embryo species. In some cases, cell lines can be derived from in vitro cultured embryos (e.g., embryonic stem cell lines).

[0361] Alternatively, embryos can be cultured in vivo by transferring them into the uterus of a female host. Typically, the female host and the embryo belong to the same or similar species. Preferably, the female host is in a state of pseudopregnancy. Methods for preparing pseudopregnant female hosts are known in the art. Furthermore, methods for transferring embryos into a female host are also known. In vivo culture of embryos allows them to develop and eventually give birth to a living animal derived from that embryo. Each cell in this animal contains a modified chromosome sequence.

[0362] Various eukaryotic cells are suitable for this method. For example, the cells can be human cells, non-human mammalian cells, non-mammal vertebrate cells, invertebrate cells, insect cells, plant cells, yeast cells, or single-celled eukaryotic organisms. Various embryos are suitable for this method. For example, the embryo can be a single-celled, two-celled, or four-celled human or non-human mammalian embryo. Exemplary mammalian embryos (including single-celled embryos) include, but are not limited to, mouse, rat, hamster, rodent, rabbit, cat, dog, sheep, pig, cattle, horse, and primate embryos. In other embodiments, the cells can be stem cells. Suitable stem cells include, but are not limited to, embryonic stem cells, ES-like stem cells, fetal stem cells, adult stem cells, pluripotent stem cells, induced pluripotent stem cells, multipotent stem cells, oligopotent stem cells, unipotent stem cells, etc. In exemplary embodiments, the cells are mammalian cells, or the embryo is a mammalian embryo.

[0363] A variety of mammalian cells are suitable for matching editing, including cells isolated from or engineered from human or non-human subjects, including NK cells, pluripotent stem cells (PSCs), adult stem cells (ASCs), fibroblasts, chondrocytes, keratinocytes, hepatocytes, pancreatic islet cells, and immune cells, including T cells, dendritic cells (DCs), natural killer (NK) cells, and macrophages.

[0364] Cells suitable for matching editing also include cells from human and non-human subjects. In some examples, the components to be matched for editing can be delivered systematically via viral or non-viral vectors. Any method known in the art for the systematic delivery of exogenous nucleotide sequences and proteins to a host may be used. Examples include the use of both viral vectors, free and integrated, such as adeno-associated virus (AAV), adenovirus vector (AD), lentiviral vectors, retroviral vectors, integration-deficient lentiviral vectors, herpes simplex virus vectors (HSV), stomatitis virus vectors (VSV), modified vaccinia virus ankara vector (MVA), arenavirus vectors, Sendai virus vectors, and parvovirus vectors. Examples include the use of non-viral vectors such as liposomes, lipid particles, lipid nanoparticles, and genomeless versions of virus-like particles (VLPs).

[0365] Uses and applications

[0366] The systems and methods disclosed herein have broad applications, including the modification and editing (e.g., inactivation and activation) of target polynucleotides in various cell types. Therefore, these systems and methods have wide applications in fields such as research and therapy. For example, these systems and methods can be used for high-throughput screening, where multiple systems with different guide RNAs target multiple different loci to obtain and screen for a variety of different phenotypic outcomes (e.g., better proliferation or lethal screening in cell lines). In another example, these systems and methods can be used for mutagenesis (similar to overlay CRISPR) or genes to produce novel proteins.

[0367] In some implementations, the aforementioned genome editing complexes, systems, compositions, or methods can be used to generate point mutations (including transversions and transformations), insertions, and deletions in cells of organisms derived from plant and animal organisms and humans.

[0368] In some implementations, the aforementioned genome editing complexes, systems, compositions, or methods can be used to generate point mutations (transversions and transitions), insertions, and deletions in the bodies of plant and animal organisms and in human cells.

[0369] The aforementioned genome editing complexes, systems, compositions, or methods can be used for cell engineering, cell therapy, organism engineering, gene therapy, agricultural improvement, veterinary medicine, and cell and animal models for research.

[0370] The matching editing technology disclosed in this paper enables precise genetic manipulation of DNA by copying the target sequence from an RNA molecule and inserting it into the target location in the genome when used to create mutations, insertions, and deletions. This technology can effectively introduce a variety of genetic alterations, including single nucleotide changes (transversions or inversions), insertions, deletions, or combinations of these alterations. Importantly, the novel characteristics of the matching editing technology allow for independent modular design of the RNA template and magRNA. Therefore, this design avoids potential interference between the secondary structures of the gRNA and the template, eliminates the source of genetic "scarring" in the template, and makes synergistic multiplexing possible and convenient.

[0371] In some implementations, the matching editing techniques disclosed herein can correct pathogenic mutations in hereditary diseases by introducing single-base alterations, insertions, or deletions. In some implementations, matching editing by introducing single-base alterations, insertions, or deletions can introduce stop codons, generate missense frameshifts, or eliminate splice sites. Therefore, the expression of target genes can be effectively silenced. More broadly, this technology can reconstruct cellular regulatory networks.

[0372] In one instance, the matching editing technology disclosed herein can alter or introduce transcriptional regulatory sequences, such as transcription factor binding sites, thereby eliminating or introducing specific transcriptional activation or repression. In another instance, it can alter protein modification sites, such as phosphorylation or acetylation sites, thereby altering cell signal transduction. In yet another instance, it can alter or introduce amino acid interfaces involved in protein-protein or protein-nucleic acid interactions, thereby eliminating or introducing specific interactions between proteins and other macromolecules and reconstructing cell signaling pathways.

[0373] In some embodiments, the ME can insert bacterial or viral antigenic epitopes into cellular proteins. In some embodiments, the ME can insert antigen-recognition variable regions of antibodies into proteins. In some embodiments, the ME can replace the antigen-recognition variable regions of T-cell receptors (TCRs). In some embodiments, the ME can insert signal peptides for secretion, nuclear localization signals for nuclear transport, or peptides for intracellular transport.

[0374] In some implementations, the matching editing disclosed herein can specifically “tag” intracellular or membrane proteins by inserting or replacing a small peptide-coding sequence within a target gene, thereby creating a tag.

[0375] In some implementations, the matching editing disclosed herein can alter the properties of cell surface proteins by inserting or replacing peptide sequences within the cell, thereby generating new patterns of intercellular interactions. In one instance, the variable sequence of the T cell receptor (TCR) protein can be altered to redirect T cells to different antigen-presenting cells. In another instance, the variable sequence of the B cell receptor (BCR) protein can be altered to redirect B cell targets.

[0376] In some implementations, the matching editing techniques disclosed herein can be used to alter the sequence of secretory proteins, thereby creating cells capable of secreting altered hormones, cytokines, and growth factors. In some instances, the secretory protein can be an antibody, and antibody production can be switched by altering the variable region of the antibody gene.

[0377] In some implementations, the matching edits disclosed herein can be used to engineer neuronal cells to differentially label the cells and redesign and monitor their interactions and connections.

[0378] The matching editing techniques disclosed in this article can be widely applied in many important fields.

[0379] It can be used for research, therapeutic drug development, agricultural development, and industrial organism engineering.

[0380] For therapeutic drug development, matching editing can be used to engineer therapeutic cells (autologous or allogeneic cells) for the development of in vitro therapeutic drugs.

[0381] In one instance, this technology can be used to engineer autologous hematopoietic cells from patients with certain diseases. For example, hematopoietic cells from patients with sickle cell anemia or β-thalassemia can be engineered in vitro to correct underlying pathogenic mutations or to reactivate fetal hemoglobin expression by altering gene regulatory sequences. The engineered autologous cells are then infused back into the patient.

[0382] Similar strategies, such as introducing point mutations or deletions into the transcriptional regulatory elements of genes, can be applied to treat other diseases by eliminating transcriptional repression or introducing transcriptional activation or repression (both by altering transcription factor binding sites).

[0383] In one instance, match editing can introduce stop codons and frameshift insertions / deletions that inactivate genes for cell therapy engineering. For example, in generating allogeneic CAR-T cells, genes involved in graft-versus-host disease (GvHD) and host resistance to graft reaction (HvGR), as well as genes that negatively impact CAR-T function, can be inactivated.

[0384] In one instance, the GvHD and HvGR genes, as well as genes that negatively affect the function of therapeutic cells, can be inactivated in other types of therapeutic cells, including NK cells derived from human or non-human subjects, pluripotent stem cells (PSCs), adult stem cells (ASCs), fibroblasts, chondrocytes, keratinocytes, hepatocytes, islet cells, and immune cells (including T cells, dendritic cells (DCs), natural killer (NK) cells, and macrophages), to generate allogeneic therapeutic cells and more effective therapeutic cells.

[0385] In one example, match-editing technology can replace the MHC-antigen complex recognition sequence of a T-cell receptor with a sequence that recognizes the MHC-cancer antigen complex, thereby generating TCR-T cells for cancer cell therapy. Furthermore, the same strategy used for CAR-T cells can also be used to generate allogeneic cells. Moreover, match-editing technology can also replace different regions of the TCR for in vivo therapy.

[0386] In some instances, matching editing techniques can introduce active cis-transcriptional enhancer elements into the FOXP3 gene of T cells, thereby generating regulatory T cells (Tregs). In some instances, the TCRs in Treg cells are further engineered using matching editing methods to achieve tissue- and cell-specific targeting. In some instances, Treg cells are further engineered using matching editing techniques to modify HLA loci and MHC genes, thereby generating allogeneic Tregs. In some instances, the engineering steps of Treg cells are performed simultaneously with multiple matching editing or in combination with other gene editing methods.

[0387] For therapeutic drug development, expression vectors that match editing components or the genes encoding these components can be delivered via appropriate delivery media (including viral and non-viral delivery systems) for in vivo gene therapy.

[0388] In one instance, in vivo therapy involves delivering a matching editing system in vivo to correct delta508 deletion in the CFTR gene, thereby treating cystic fibrosis.

[0389] In one instance, in vivo therapy involves delivering a matching editing system within the body to skip the exons of the mutated dystrophin gene, thereby partially restoring the function of dystrophin, for the treatment of Duchenne muscular dystrophy.

[0390] In one instance, in vivo therapy involves delivering a matching editing system in vivo to reduce the expression of amyloid-β (Aβ), tau, or α-synuclein, thereby treating Alzheimer's disease.

[0391] The similar principles of ex vivo cell engineering in cell therapy and in vivo gene therapy can be applied to the vast majority of human genetic diseases because the universality of matched editing gene editing strategies allows it to cover point mutations, insertions, deletions, and combinations of these mutations. Furthermore, the mutation site does not need to be adjacent to the PAM (Potentially Amputated Matrix Amplifier).

[0392] The similar principles of ex vivo cell engineering in cell therapy and in vivo gene therapy can be applied to eliminate the expression of pathogenic proteins or RNA or to inactivate them in order to treat cancer, autoimmune diseases and neurodegenerative diseases.

[0393] In one instance, autologous therapeutic cells can be engineered; in another, allogeneic therapeutic cells can be engineered by inactivating the GvHD and HvGR genes.

[0394] Matching editing can be used to engineer human or non-human iPSCs for the treatment of human and veterinary diseases.

[0395] Match editing can be used to engineer reproductive cells in non-human mammals, non-human animals, plants, and crops. These engineered reproductive cells can then be used to produce living organisms.

[0396] Matching editing can be used for the engineering modification of microbial genomes in both industrial and non-industrial settings.

[0397] Many serious human diseases share a common cause: gene alterations or mutations. The pathogenic mutations in a patient's body are either inherited from parents or caused by environmental factors. These diseases include, but are not limited to, the following categories: First, some hereditary diseases are caused by germ cell mutations. One example is cystic fibrosis, which is caused by a mutation in the CFTR gene inherited from parents. A second repressor mutation in CFTR can partially restore the function of the CFTR protein in somatic tissues. Other correctable hereditary diseases caused by point mutations include Gaucher disease, alpha-trypsin deficiency, sickle cell anemia, etc. Second, some diseases, such as chronic viral infections, are caused by exogenous environmental factors and the resulting gene alterations. One example is AIDS, which is caused by the insertion of the human HIV virus genome into the genome of infected T cells. Third, some neurodegenerative diseases also involve gene alterations. One example is Huntington's disease, which is caused by the amplification of the CAG trinucleotide in the Huntington gene of affected patients. Other examples include lysosomal storage diseases, epidermolysis bullosa, and retinal degeneration. Finally, cancer is caused by the accumulation of various somatic mutations within cancer cells. Therefore, correcting pathogenic gene mutations or functionally correcting sequences presents a highly attractive therapeutic opportunity for treating these diseases.

[0398] Somatic cell gene editing is a highly attractive therapeutic strategy for a wide range of human diseases. Three key factors are crucial for successful therapeutic gene editing: (i) how to achieve sequence-specific recognition (“sequence recognition module”); (ii) how to correct underlying mutations (“correction module”); and (iii) how to link the “correction module” with the “sequence recognition module” to achieve sequence-specific correction. There are many approaches to achieving each individual task. However, no single platform or technology currently offers optimal and practical somatic cell gene editing. More specifically, most existing gene-specific editing techniques are based on nuclease-induced DNA DSB and subsequent DSB-induced homologous recombination, which are inactive or nonexistent in most somatic cells. Therefore, these techniques are limited in their application for therapeutically correcting pathogenic gene mutations in somatic tissues in most diseases.

[0399] In contrast, the systems and methods disclosed herein allow for targeted DNA sequence editing of genes without relying on nuclease activity. These systems and methods do not generate DSBs and do not rely on DSB-mediated homologous recombination. Furthermore, the system employs a modular design, enabling targeting of any target DNA or RNA sequence in an extremely flexible and convenient manner. Essentially, this method can guide DNA or RNA editing enzymes to target virtually any DNA or RNA sequence in somatic cells, including stem cells. Through precise editing of the target DNA or RNA sequence, the enzymes can correct mutated genes in hereditary diseases, inactivate viral genomes in infected cells, generate stop codons to inactivate and eliminate the expression of pathogenic proteins in diseases (including neurodegenerative diseases), silence oncogenes in cancer, mutate shared splicing sites to eliminate pathogenic exons, or mutate regulatory sequences to restore therapeutic gene expression / inactivation. Therefore, the systems and methods disclosed herein can be used to correct potential gene alterations in a variety of diseases, including the aforementioned hereditary diseases, chronic infectious diseases, neurodegenerative diseases, and cancers. Importantly, the systems and methods disclosed herein can be used for cell engineering to generate research tools or cell-based therapies.

[0400] Hereditary diseases

[0401] It is estimated that over six thousand genetic diseases are caused by known gene mutations. Correcting the underlying mutation in the diseased tissue / organ can alleviate or cure the disease. For example, cystic fibrosis affects one in 3,000 people in the United States. It is caused by a genetic mutation in the CFTR gene, with 70% of patients having the same mutation, which results in the deletion of phenylalanine at position 508 (called...). Trinucleotide deletion in Phe 508. Val 509 residues (GTT) can cause CFTR misalignment and degradation. The systems and methods disclosed herein can be used to convert Val 509 residues (GTT) in affected tissues (lungs) to Phe509 (TTT), thereby functionally correcting... The Phe 508 mutation. Furthermore, the mutant... Second repressive mutations in Phe 508 CFTR (such as R553Q, R553M, or V510D) can partially restore the function of the CFTR protein in somatic tissues.

[0402] Chronic infectious diseases

[0403] The systems and methods disclosed herein can also be used to specifically inactivate any gene in a viral genome introduced into human cells / tissues. For example, the systems and methods disclosed herein can generate stop codons, thereby prematurely terminating the translation of key viral genes, and thus alleviating or curing chronic debilitating infectious diseases. For example, current HIV therapies can reduce viral load but cannot completely eliminate dormant HIV in positive T cells. The systems and methods disclosed herein can permanently inactivate the expression of key HIV genes in the integrated HIV genome of human T cells by introducing one or more stop codons. Another example is hepatitis B virus (HBV). The systems and methods disclosed herein can be used to specifically inactivate key HBV genes integrated into the human genome and silence the HBV life cycle.

[0404] Neurodegenerative diseases

[0405] Some neurodegenerative diseases are caused by gain-of-function mutations. For example, the SOD1G93A mutation leads to amyotrophic lateral sclerosis (ALS). The systems and methods disclosed in this disclosure can be used to correct mutations or eliminate mutant protein expression by introducing a stop codon or altering the splice site. For example, an alternative splice form of Tau protein containing exon 10 plays a pathogenic role in Alzheimer's disease. Altering the CG base pair at the shared exon 10 splice site will eliminate this alternative splice form of Tau protein.

[0406] cancer

[0407] Many genes, including tumor suppressor genes, oncogenes, and DNA repair genes, contribute to cancer development. Mutations in these genes often lead to various cancers. Using the systems and methods disclosed in this disclosure, these mutations can be specifically targeted and corrected. Therefore, by introducing point mutations at catalytic or splicing sites, oncogenes can be functionally inhibited or their expression eliminated.

[0408] Somatic cell gene knockout

[0409] In some implementations, protein expression of genes in somatic cells of human and non-human organisms can be eliminated by generating early stop codons. This method can be used for therapeutic purposes or to develop research tools.

[0410] Changes in control elements

[0411] The method described can be used to alter the sequence of regulatory elements in DNA and RNA. Therefore, it provides a pathway to alter, silence, or activate gene expression by modifying mechanisms related to gene expression. This can be used for therapeutic purposes as well as to develop research tools.

[0412] Stem cell gene modification

[0413] In some implementations, the systems and methods disclosed herein can be used to genetically modify cells reprogrammed into different cell types. Suitable cells include, for example, stem cells (adult stem cells, embryonic stem cells, induced pluripotent stem cells, mesenchymal stem cells, etc., see Stem cells: past, present, and future. Zakrzewski). et al Stem Cell Res Ther. 2019 Feb 26;10(1):68.) and progenitor cells (e.g., cardiac progenitor cells, neural progenitor cells, etc.), or mature cells used to convert to different cell types (e.g., using Molecular Interaction Networks to Select Factors for Cell Conversion). Ouyang JF et al. The algorithm described in Methods MolBiol. 2019;1975:333-361. Suitable cells can be derived from any multicellular organism, including, for example, mammals (including rodents, humans, horses, camels, pigs, etc.), insects, and birds (including chickens, ducks, etc.). Suitable host cells include in vitro or ex vivo host cells, such as isolated host cells.

[0414] In some implementations, the complexes, systems, and methods disclosed herein can be used for in vitro targeted and precise gene modification of cells or tissues to correct underlying genetic defects. After in vitro correction, the tissue can be reinfused into the patient. Furthermore, this technology has broad applications in cell-based therapies for correcting hereditary diseases.

[0415] In this article, the term "stem cell" refers to a pluripotent cell that, under suitable conditions, can differentiate into various specialized cell types, while under other suitable conditions, it can self-renew and remain essentially undifferentiated. The term "stem cell" also encompasses pluripotent cells, multipotent cells, progenitor cells, and predisposing cells. Exemplary human stem cells can be obtained from hematopoietic stem cells or mesenchymal stem cells derived from bone marrow tissue, embryonic stem cells derived from embryonic tissue, or embryonic germ cells derived from fetal reproductive tissue. Exemplary pluripotent stem cells can also be generated by reprogramming somatic cells into a pluripotent state by expressing certain transcription factors associated with pluripotency; these cells are referred to as "induced pluripotent stem cells" or "iPScs or iPS cells."

[0416] Embryonic stem cells (ES cells) are undifferentiated pluripotent stem cells that are obtained from early embryos, such as the inner cell mass of the blastocyst, or produced artificially (e.g., by nuclear transfer). They can produce any differentiated cell type in the embryo or adult, including germ cells (e.g., sperm and eggs).

[0417] "Induced pluripotent stem cells (iPScs or iPS cells)" are cells generated by reprogramming somatic cells through the expression or induction of the expression of multiple factors (referred to herein as reprogramming factors). iPS cells can be generated from fetal, postnatal, neonatal, juvenile, or adult somatic cells. Factors that can be used to reprogram somatic cells into pluripotent stem cells include, for example, Oct4 (sometimes referred to as Oct3 / 4), Sox2, c-Myc, Klf4, Nanog, and Lin28. In some embodiments, somatic cells can be reprogrammed into pluripotent stem cells by expressing at least two, at least three, at least four, at least five, at least six, or at least seven reprogramming factors.

[0418] "Hematopoietic progenitor cells" or "hematopoietic precursor cells" refer to cells that have been directed to differentiate into hematopoietic lineages but still possess the ability to further differentiate into hematopoiesis. These include hematopoietic stem cells, pluripotent hematopoietic stem cells, common myeloid progenitors, megakaryocyte progenitors, erythrocyte progenitors, and lymphoid progenitors. Hematopoietic stem cells (HSCs) are pluripotent stem cells capable of producing all types of blood cells, including myeloid (monocytes and macrophages, granulocytes (neutrophils, basophils, eosinophils, and mast cells), erythrocytes, megakaryocytes / platelets, dendritic cells) and lymphoid (T cells, B cells, NK cells).

[0419] "Pluripotent stem cells" refer to stem cells that have the potential to differentiate into all the cells that make up one or more tissues or organs, or preferably any one of the three germ layers: endoderm (gastric lining, gastrointestinal tract, lung), mesoderm (muscle, bone, blood, urogenital system), or ectoderm (epidermal tissue and nervous system).

[0420] The term "somatic cell" as used in this article refers to any cell other than germ cells (such as eggs, sperm, etc.) that does not directly pass on its DNA to the next generation. Somatic cells typically have limited or no pluripotency. The somatic cells used in this article can be naturally occurring or genetically modified.

[0421] Cell therapy and ex vivo therapy

[0422] Several embodiments of this disclosure also provide cell lines produced or used in any other embodiment of this disclosure for therapeutic purposes. In one embodiment, this disclosure relates to a method for generating therapeutic cells (e.g., T cells engineered to express chimeric antigen receptor (CAR-T) or T cell receptor (TCR-T)). In one embodiment, this disclosure relates to a method for generating therapeutic regulatory T cells (Tregs). CAR-T / TCR-T cells may be derived from primary T cells or differentiated from stem cells. Suitable stem cells include, but are not limited to, mammalian stem cells, such as human stem cells, including but not limited to hematopoietic stem cells, neural stem cells, embryonic stem cells, induced pluripotent stem cells (iPSCs), mesenchymal stem cells, mesodermal stem cells, liver stem cells, pancreatic stem cells, muscle stem cells, and retinal stem cells. Other stem cells include, but are not limited to, mammalian stem cells, such as mouse stem cells, for example, mouse embryonic stem cells.

[0423] In various embodiments, the complexes, systems, and methods disclosed herein can be used to knock down, modify, or enhance the expression of single or multiple genes in various types of cells or cell lines (including, but not limited to, mammalian cells). This technology can be used for a variety of applications, including, but not limited to, preventing graft-versus-host disease by knocking down genes to render non-host cells immunogenic to the host, or preventing host-versus-graft disease by making non-host cells resistant to host attack. These methods are also relevant to the generation of allogeneic (off-the-shelf) or autologous (patient-specific) cell-based therapies. These genes include, but are not limited to, T-cell receptors (TRAC), major histocompatibility complexes (MHC), and other similar genes. Class I and II genes (including B2M), co-receptors (HLA-F, HLA-G), genes involved in innate immune responses (MICA, MICB, HCP5), inflammation-related genes (NKBBiL, LTA, TNF, LTB, LST1, NCR3, AIF1), immune receptors (LY6), heat shock proteins (HSPA1L, HSPA1A, HSPA1B), complement cascades, regulatory receptors (NOTCH4), antigen processing (TAP, HLA-DM, HLA-DO), peptide transporters (RING1), and genes involved in increased potency or persistence (e.g., PD-1, CTLA-4, FOXP3, and B7). Genes that interact with the tumor microenvironment by T cells (including, but not limited to, cytokine receptors such as TGFβ, interleukin (IL)-4, IL-7, IL-2, IL-4, and repressors of IL-15, IL-12, IL-18, IL-2, and IFNγ; genes involved in promoting cytokine release syndrome (including, but not limited to, GMCSF); genes encoding antigens targeted by CAR / TCR (e.g., endogenous CS1 designed for CAR targeting CS1); or other genes beneficial to CAR-T / TCR-T or other cell-based therapies (including, but not limited to, CAR-NK, CAR-B, etc.). See, for example, DeRenzo et al., Genetic Modification Strategies to Enhance CAR T Cell Persistence for Patients With Solid Tumors. Front. Immunol., 15 February 2019.

[0424] This technology can also be used to knock down or modify genes involved in the self-destruction of immune cells (such as T cells and NK cells), or to alert the immune system of a patient or animal to the presence of foreign cells, particles or molecules entering the patient or animal's body, or to encode genes that currently target therapeutic proteins used to weaken or enhance immune responses, such as CD52 and PD1.

[0425] One application is engineering HLA alleles in bone marrow cells to improve haplotype matching. The engineered cells can then be used in bone marrow transplantation to treat leukemia. Another application is engineering negative regulatory elements of the fetal hemoglobin gene in hematopoietic stem cells for the treatment of sickle cell anemia and β-thalassemia. This negative regulatory element is mutated, and the expression of the fetal hemoglobin gene in hematopoietic stem cells is reactivated, compensating for the loss of function caused by mutations in the adult α or β hemoglobin gene. Yet another application is engineering iPS cells to generate allogeneic therapeutic cells for treating various degenerative diseases, including Parkinson's disease (loss of neurons) and type 1 diabetes (loss of pancreatic β cells). Other exemplary applications include engineering HIV-resistant T cells by inactivating the CCR5 gene and other genes encoding receptors required for HIV entry into cells.

[0426] This technology can also be used to produce transgenic animals, which can be used as disease models or for gene function research.

[0427] As used in this article, the term "immune cells" generally includes white blood cells (leukocytes) derived from hematopoietic stem cells (HSCs) produced in the bone marrow. Examples of immune cells include, but are not limited to, lymphocytes (T cells, B cells, and natural killer (NK) cells) and myeloid-derived cells (neutrophils, eosinophils, basophils, monocytes, macrophages, and dendritic cells).

[0428] Immune cells can be isolated from subjects, especially human subjects. These immune cells can be obtained from target subjects, such as subjects suspected of having a specific disease or condition, subjects suspected of being susceptible to a specific disease or condition, or subjects receiving treatment for a certain disease or condition. Immune cells can be collected from any site within the subject's body, including but not limited to blood, umbilical cord blood, spleen, thymus, lymph nodes, and bone marrow. The isolated immune cells can be used directly or stored for a period of time, such as by freezing.

[0429] Immune cells can be enriched / purified from any tissue in which they reside, including but not limited to blood (including blood collected from blood banks or cord blood banks), spleen, bone marrow, tissue removed and / or exposed during surgery, and tissue obtained through biopsy. The tissue / organ from which immune cells are enriched, isolated, and / or purified can be isolated from living and non-living subjects, where non-living subjects refer to organ donors. In some embodiments, immune cells are isolated from blood, such as peripheral blood or cord blood. In some aspects, immune cells isolated from cord blood have enhanced immunomodulatory capacity, measured, for example, by CD4 or CD8 positive T cell suppression. In some specific aspects, to enhance immunomodulatory capacity, immune cells are isolated from consorted blood (particularly consorted cord blood). The consorted blood may be from two or more sources, such as three, four, five, six, seven, eight, nine, ten, or more sources (e.g., donor subjects).

[0430] Immune cell populations can be obtained from subjects requiring treatment or suffering from diseases related to reduced immune cell activity. Therefore, these cells can be autologous cells from the subject requiring treatment. Alternatively, immune cell populations can be obtained from a donor, preferably a tissue-compatible donor. Immune cell populations can be collected from peripheral blood, umbilical cord blood, bone marrow, spleen, or any other organ / tissue containing immune cells in the subject or donor. Immune cells can also be isolated from a bank of the subject and / or donor, for example, from concomitant umbilical cord blood.

[0431] When the immune cell population is obtained from a donor different from the subject, the donor is preferably an allogeneic donor, provided that the obtained cells can be introduced into the subject and are compatible with the subject. Allogeneic donor cells may or may not be compatible with human leukocyte antigens (HLA). To make them compatible with the subject, the allogeneic cells can be treated to reduce their immunogenicity.

[0432] In some implementations, the immune cells may be T cells (e.g., regulatory T cells, CD4+). + T cells (CD8 T cells or γδ T cells), NK cells, invariant NK cells, NKT cells, and stem cells (e.g., mesenchymal stem cells (MSCs) or induced pluripotent stem cells (iPSCs)). In some embodiments, these cells are monocytes or granulocytes, such as myeloid cells, macrophages, neutrophils, dendritic cells, mast cells, eosinophils, and / or basophils. This document also provides methods for producing and engineering immune cells, as well as methods for using and administering cells for adoptive cell therapy, wherein the cells can be autologous or allogeneic. Therefore, immune cells can be used in immunotherapy, such as targeting cancer cells.

[0433] Gene editing in animals and plants

[0434] The systems and methods described above can be used to generate transgenic non-human animals or plants with one or more target gene modifications. In some embodiments, the transgenic non-human animal is a homozygous for the gene modification. In some embodiments, the transgenic non-human animal is a heterozygous for the gene modification. In some embodiments, the transgenic non-human animal is a vertebrate, such as fish (e.g., zebrafish, goldfish, pufferfish, cave fish, etc.), amphibians (e.g., frogs, salamanders, etc.), birds (e.g., chickens, turkeys, etc.), reptiles (e.g., snakes, lizards, etc.), mammals (e.g., ungulates, pigs, cattle, goats, sheep, etc.), lagomorphs (e.g., rabbits), rodents (e.g., rats, mice), or non-human primates.

[0435] The gene-editing complexes, systems, and methods disclosed herein can be used to treat animal diseases in a manner similar to those described for treating human diseases. Furthermore, they can be used to generate knock-in animal disease models carrying specific gene mutations for research, drug discovery, and target validation. The systems and methods described above can also be used to introduce point mutations into ES cells or embryos of various organisms for breeding and improving animal breeds and crop quality.

[0436] Methods for introducing exogenous nucleic acids into plant cells are well known in the art. Suitable methods include viral infection (e.g., double-stranded DNA viruses), transfection, conjugation, protoplast fusion, electroporation, gene gun technology, calcium phosphate precipitation, direct microinjection, silicon carbide whisker technology, Agrobacterium-mediated transformation, etc. The choice of method usually depends on the type of cells to be transformed and the conditions under which transformation occurs (i.e., in vitro, ex vivo, or in vivo).

[0437] Reagent test kit

[0438] This disclosure further provides kits comprising reagents for performing the methods described above, including, for example, CRISPR / Cas-guided target binding or correction reactions. To this end, one or more reaction components of the methods disclosed herein, such as RNA, RNA-guided nickase proteins, reverse transcriptase proteins, fusion proteins, and associated nucleic acids, can be provided in kit form. In one embodiment, the kit comprises a nickase protein, a reverse transcriptase protein, or a nucleic acid encoding that protein, an effector protein, one or more of the aforementioned RNAs, or a group of the aforementioned RNA molecules. In other embodiments, the kit may comprise one or more other reaction components. In such kits, appropriate amounts of one or more reaction components are provided in one or more containers or immobilized on a substrate.

[0439] Examples of other components of the kit include, but are not limited to: one or more host cells; one or more reagents for introducing exogenous nucleotide sequences into the host cells; one or more reagents (e.g., probes or PCR primers) for detecting RNA or protein expression or verifying the status of target nucleic acids; and buffers or culture media (1X or concentrated form) required for the reaction. The kit may also contain one or more of the following components: support, termination reagent, modification or digestion reagent, osmotic regulator, and detection device.

[0440] The reaction components used can be provided in various forms. For example, these components (e.g., enzymes, RNA, probes, and / or primers) can be suspended in an aqueous solution or exist as lyophilized or freeze-dried powders, granules, or beads. In the latter case, these components form a complete mixture of components upon reconstitution for assay. The kits disclosed herein can be provided at any suitable temperature. For example, for the storage of kits containing protein components or their complexes in liquid form, they are preferably provided and maintained below 0°C, preferably at -20°C or lower, or stored frozen.

[0441] The kit or system may contain any combination of the components described herein in sufficient quantities to perform at least one assay. In some applications, one or more reaction components may be provided in pre-measured single-use amounts in individual tubes or equivalent containers (typically for single use). With this setup, RNA-guided reactions can be performed by adding the target nucleic acid, or a sample or cells containing the target nucleic acid, directly to the individual tubes. The amounts of components provided in the kit may be any suitable amount and may depend on the target market for the product. The containers providing the components may be any conventional containers capable of containing this form of provision, such as microcentrifuge tubes, microplates, ampoules, vials, or integrated assay devices, such as fluidic devices, boxes, lateral flow devices, or other similar devices.

[0442] The kit may also include packaging materials for containing containers or combinations of containers. Typical packaging materials for such kits and systems include solid matrices (e.g., glass, plastic, paper, foil, microparticles, etc.) that hold reaction components or detection probes in various configurations (e.g., in vials, microtiter plate wells, microarrays, etc.). The kit may also include instructions for use of the components recorded in tangible form.

[0443] definition

[0444] Nucleic acids or polynucleotides refer to DNA molecules (e.g., but not limited to cDNA or genomic DNA) or RNA molecules (e.g., but not limited to mRNA), including DNA or RNA analogs. DNA or RNA analogs can be synthesized from nucleotide analogs. DNA or RNA molecules may contain non-naturally occurring parts, such as modified bases, modified backbones, deoxyribonucleotides in RNA, etc. Nucleic acid molecules can be single-stranded or double-stranded.

[0445] When referring to a nucleic acid molecule or polypeptide, the term "isolated" means that the nucleic acid molecule or polypeptide is essentially free of at least one other component that is bound to or coexists with it in nature.

[0446] In this article, the term "guide RNA" generally refers to an RNA molecule (or a group of RNA molecules) that can bind to an RNA-guided nicking enzyme (such as a CRISPR protein) and target that nicking enzyme to a specific location within the target DNA. Guide RNA can contain two segments: a DNA-targeting guide segment and a protein-binding segment. The DNA-targeting segment contains a nucleotide sequence complementary to (or at least capable of hybridizing with) the target sequence under stringent conditions. The protein-binding segment interacts with an RNA-guided nicking enzyme (such as a CRISPR protein), such as Cas9 or a Cas9-related polypeptide. These two segments can be located in the same RNA molecule or in two or more separate RNA molecules. When these two segments are located in separate RNA molecules, the molecule containing the DNA-targeting guide segment is sometimes called CRISPR RNA (crRNA), while the molecule containing the protein-binding segment is called trans-activating RNA (tracrRNA).

[0447] As used herein, the term "target nucleic acid" or "target" refers to a nucleic acid containing a target nucleic acid sequence. The target nucleic acid can be single-stranded or double-stranded, but is typically double-stranded DNA. The terms "target nucleic acid sequence," "target sequence," or "target region" as used herein refer to a specific sequence or its complementary sequence that is intended to be bound or modified using the systems disclosed herein. The target sequence can be present in the nucleic acid in vitro or in vivo within the cellular genome, and the nucleic acid can be any form of single-stranded or double-stranded nucleic acid.

[0448] "Target nucleic acid strand" refers to the strand of the target nucleic acid that pairs base-with the guide RNA disclosed herein. In other words, the strand of the target nucleic acid that hybridizes with the crRNA and the guide sequence is called the "target nucleic acid strand." The other strand of the target nucleic acid, i.e., the strand that is not complementary to the guide sequence, is called the "non-complementary strand." For double-stranded target nucleic acids (e.g., DNA), each strand can be used as a "target nucleic acid strand" to design crRNA and guide RNA, and to implement the methods disclosed herein, provided a suitable PAM site is present.

[0449] As used herein, the term "derived from" refers to the process of isolating, deriving, or preparing a different second component (e.g., a second molecule different from the first molecule) using a first component (e.g., a first molecule) or information derived from that first component. For example, mammalian codon-optimized Cas9 polynucleotides are derived from the amino acid sequence of the wild-type Cas9 protein. Furthermore, variant mammalian codon-optimized Cas9 polynucleotides, including Cas9 single-mutant nickases (nCas9, e.g., nCas9D10A) and Cas9 double-mutant nonfunctional nucleases (dCas9, e.g., dCas9 D10A H840A), are all derived from polynucleotides encoding wild-type mammalian codon-optimized Cas9 proteins.

[0450] As used herein, the term "wild type" is a term in the art as understood by those skilled in the art, referring to the typical form of an organism, strain, gene, or trait that appears in nature, as opposed to mutant or variant forms.

[0451] As used herein, the term “variant” refers to a first composition (e.g., a first molecule) associated with a second composition (e.g., a second molecule, also known as a “parent” molecule). Variant molecules can be derived from, isolated from, based on, or homologous to a parent molecule. For example, mutant forms of mammalian codon-optimized Cas9 (hspCas9), including Cas9 single-mutant nickases and Cas9 double-mutant nonfunctional nucleases, are variants of mammalian codon-optimized wild-type Cas9 (hspCas9). The term “variant” can be used to describe either polynucleotides or polypeptides.

[0452] When used with polynucleotides, variant molecules may have an identical nucleotide sequence to the original parent molecule, or they may have less than 100% nucleotide sequence identity with the parent molecule. For example, a variant of a gene nucleotide sequence may be a second nucleotide sequence having at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or higher nucleotide sequence identity compared to the original nucleotide sequence. Polynucleotide variants also include polynucleotides comprising the entire parent polynucleotide and further comprising additional fusion nucleotide sequences. Polynucleotide variants also include polynucleotides that are part of or subsequences of the parent polynucleotide, such as unique subsequences of the polynucleotides disclosed herein (e.g., subsequences determined by standard sequence comparison and alignment techniques).

[0453] On the other hand, polynucleotide variants include nucleotide sequences that contain minor, insignificant, or unimportant changes compared to the parental nucleotide sequence. For example, minor, insignificant, or unimportant changes include alterations to the following nucleotide sequences: (i) those that do not change the amino acid sequence of the corresponding polypeptide; (ii) those that occur outside the protein-coding open reading frame of the polynucleotide; (iii) those that result in deletions or insertions that may affect the corresponding amino acid sequence but have little or no effect on the biological activity of the polypeptide; and (iv) those where the nucleotide change results in the substitution of an amino acid with a chemically similar amino acid. If the polynucleotide does not encode a protein (e.g., tRNA, crRNA, or tracrRNA), the variant of the polynucleotide may contain nucleotide changes that do not result in loss of function of the polynucleotide. Furthermore, this disclosure also covers conserved variants of the disclosed nucleotide sequences that produce functionally identical nucleotide sequences. Those skilled in the art will understand that many variants of the disclosed nucleotide sequences are included in this disclosure.

[0454] When used for proteins, variant peptides can have the exact same amino acid sequence as the original parent peptide, or they can have less than 100% amino acid identity with the parent protein. For example, a variant amino acid sequence can be a second amino acid sequence that has at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or higher amino acid sequence identity compared to the original amino acid sequence.

[0455] Peptide variants include peptides comprising the entire parent peptide as well as peptides further comprising additional fusion amino acid sequences. Peptide variants also include peptides that are portions or subsequences of the parent peptide, for example, unique subsequences of the peptides disclosed herein (e.g., subsequences determined by standard sequence comparison and alignment techniques) are also included in this disclosure.

[0456] On the other hand, peptide variants include peptides that protect against minor, insignificant, or unimportant changes compared to the parental amino acid sequence. For example, minor, insignificant, or unimportant changes include amino acid alterations (including substitutions, deletions, and insertions) that have little or no effect on the peptide's biological activity and produce functionally identical peptides, including the addition of nonfunctional peptide sequences. In other respects, the variant peptides of this disclosure alter the biological activity of the parent molecule, for example, mutant variants of Cas9 peptides with modified or lost nuclease activity. Those skilled in the art will understand that this disclosure covers many variants of the disclosed peptides.

[0457] In some respects, the polynucleotide or polypeptide variants disclosed herein may include variant molecules that have altered, added to, or omitted a small percentage of nucleotide or amino acid positions, for example, typically less than about 10%, less than about 5%, less than 4%, less than 2%, or less than 1%.

[0458] As used herein, the term "conserved substitution" in a nucleotide or amino acid sequence refers to alterations in the nucleotide sequence that (i) do not result in any corresponding changes to the amino acid sequence due to the redundancy of triplet codons, or (ii) result in the substitution of the original parent amino acid by an amino acid with a similar chemical structure. Conserved substitutions of functionally similar amino acids are well known in the art, where one amino acid residue is substituted by another amino acid residue with similar chemical properties (e.g., aromatic or positively charged side chains) without substantially altering the functional properties of the resulting polypeptide molecule.

[0459] The following lists groups of natural amino acids with similar chemical properties, where the substitutions of amino acids within a group are "conserved" amino acid substitutions. This grouping is not absolute; these natural amino acids may be classified into different groups when considering different functional properties. Amino acids with nonpolar and / or aliphatic side chains include: glycine, alanine, valine, leucine, isoleucine, and proline. Amino acids with polar, uncharged side chains include: serine, threonine, cysteine, methionine, asparagine, and glutamine. Amino acids with aromatic side chains include: phenylalanine, tyrosine, and tryptophan. Amino acids with positively charged side chains include: lysine, arginine, and histidine. Amino acids with negatively charged side chains include: aspartic acid and glutamic acid.

[0460] A “Cas9 mutant” or “Cas9 variant” refers to a protein or polypeptide derivative of the wild-type Cas9 protein (e.g., the Streptococcus pyogenes Cas9 protein), such as a protein having one or more point mutations, insertions, deletions, truncations, fusion proteins, or combinations thereof. It substantially retains the RNA-targeting activity of the Cas9 protein. The protein or polypeptide may contain fragments of the wild-type protein, consist of fragments of the wild-type protein, or consist primarily of fragments of the wild-type protein. Typically, the mutant / variant has at least 50% (e.g., any value from 50% to 100%) identity with the protein. The mutant / variant can bind to an RNA molecule and target specific DNA sequences through that RNA molecule, and may also possess nuclease activity. Examples of these domains include RuvC-like motifs (amino acids 7-22, 759-766, and 982-989) and HNH motifs (amino acids 837-863). See Gasiunas et al., Proc Natl Acad Sci US A. 2012 September 25; 109(39): E2579–E2586 and WO2013176772.

[0461] "Complementarity" refers to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid sequence through traditional Watson-Crick base pairing or other non-traditional methods. The complementarity percentage indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., 50%, 60%, 70%, 80%, 90%, and 100% complementarity out of 10 residues). "Complete complementarity" means that all adjacent residues in a nucleic acid sequence can form hydrogen bonds with the same number of adjacent residues in a second nucleic acid sequence. As used in this article, "substantially complementary" means that the complementarity level in regions of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100%; or it means two nucleic acids that hybridize under strict conditions.

[0462] The “strict conditions” for hybridization used in this article refer to conditions under which the nucleic acid primarily hybridizes with the target sequence and minimally hybridizes with non-target sequences when the nucleic acid is complementary to the target sequence. Strict conditions are typically sequence-dependent and influenced by various factors. Generally, the longer the sequence, the higher the temperature required for specific hybridization with the target sequence. For a non-limiting example of strict conditions, see Tijssen (1993), Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes Part I, Second Chapter "Overview of principles of hybridization and the strategy of nucleic acidprobe assay", Elsevier, NY.

[0463] "Hybridization" refers to the process by which completely or partially complementary nucleic acid strands combine under specific hybridization conditions to form a double-stranded structure or a region where two of the constituent strands are linked by hydrogen bonds. While hydrogen bonds typically form between adenine and thymine or uracil (A and T or U), or between cytosine and guanine (C and G), other base pairs can also be formed (e.g., Adams). et al. , The Biochemistry of the Nucleic Acids, 11th ed., 1992).

[0464] As used in this article, “expression” refers to the process of transcribing polynucleotides from a DNA template (e.g., into mRNA or other RNA transcripts), and / or the subsequent translation of the transcribed mRNA into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides are collectively referred to as “gene products.” If the polynucleotides are derived from genomic DNA, expression may include the splicing of mRNA in eukaryotic cells.

[0465] In this document, the terms “polypeptide,” “peptide,” and “protein” are used interchangeably to refer to an amino acid polymer of any length. This polymer may be linear or branched, may contain modified amino acids, and may be separated by non-amino acid components. These terms also cover modified amino acid polymers; for example, those with disulfide bond formation, glycosylation, esterification, acetylation, phosphorylation, PEGylation, or any other modification, such as coupling with a labeled component. The term “amino acid” as used herein includes natural and / or non-natural or synthetic amino acids, including glycine and its D- or L-optical isomers, as well as amino acid analogs and peptides.

[0466] The terms "fusion polypeptide" or "fusion protein" refer to a protein composed of two or more polypeptide sequences linked together. Fusion polypeptides covered by this disclosure include the translational product of a chimeric gene construct that links a nucleic acid sequence encoding a first polypeptide (e.g., an RNA-binding domain) with a nucleic acid sequence encoding a second polypeptide (e.g., an effector domain) to form a single open reading frame. In other words, a "fusion polypeptide" or "fusion protein" is a recombinant protein composed of two or more proteins linked together by peptide bonds or multiple peptide segments. Fusion proteins may also contain peptide linkers between the two domains.

[0467] The term “connector” refers to any method, entity, or portion used to connect two or more entities. A connector can be covalent or non-covalent. Examples of covalent connectors include connector portions covalently bonded or covalently linked to one or more proteins or domains to be connected. Connectors can also be non-covalent, such as organometallic bonds through a metal center (e.g., platinum atoms). For covalent connections, various functional groups can be used, such as amide groups (including carbonate derivatives), ethers, esters (including organic and inorganic esters), amino groups, carbamates, ureas, etc. To achieve connection, domains can be modified by oxidation, hydroxylation, substitution, reduction, etc., to provide coupling sites. Various coupling methods are well known to those skilled in the art and are included in this disclosure. Connector groups include, but are not limited to, chemical connector portions, or, for example, peptide connector portions (connector sequences). It should be understood that modifications that do not significantly impair the function of RNA-binding domains and effector domains are preferred.

[0468] As used herein, the terms “conjugate,” “conjugate,” or “link” refer to the joining of two or more entities to form a single entity. Conjugates include peptide-small molecule conjugates and peptide-protein / peptide conjugates.

[0469] In this document, the terms "subject" and "patient" are used interchangeably and both refer to vertebrates, preferably mammals, and more preferably humans. Mammals include, but are not limited to, rodents, apes, humans, livestock, locomotor animals, and pets. Tissues, cells, and their progeny of biological entities obtained in vivo or cultured in vitro are also included. In some embodiments, the subject may be an invertebrate, such as an insect or nematode; while in other embodiments, the subject may be a plant or fungus.

[0470] The terms “treatment,” “relief,” or “mitigation” used herein are used interchangeably. These terms refer to methods of achieving a beneficial or anticipated outcome, including, but not limited to, therapeutic and / or preventative effects. A therapeutic effect refers to any therapeutically significant improvement or influence on one or more diseases, conditions, or symptoms to which treatment has been administered. For preventative effects, the composition may also be administered to subjects at risk of developing a particular disease, condition, or symptom, or to subjects who report the presence of one or more physiological symptoms of a disease, even if the disease, condition, or symptom has not yet manifested.

[0471] The phrase "pharmaceutically or pharmacologically acceptable" refers to molecular entities and compositions that will not produce adverse reactions, allergic reactions, or other undesirable reactions when administered to animals (e.g., humans). Based on this disclosure, those skilled in the art will understand methods for preparing pharmaceutical compositions comprising therapeutic agents (e.g., cells) or other active ingredients. Furthermore, it should be understood that, for use in animals (e.g., humans), the preparation should meet the sterility, pyrogenicity, general safety, and purity standards required by the FDA's Office of Biologics Standards. The term "pharmaceutically acceptable carrier" as used herein includes any and all aqueous solvents (e.g., water, alcohol / aqueous solutions, salt solutions, parenteral media such as sodium chloride, Ringer's glucose solution, etc.), non-aqueous solvents (e.g., propylene glycol, polyethylene glycol, vegetable oils, and injectable organic esters such as ethyl oleate), dispersion media, coatings, surfactants, antioxidants, preservatives (e.g., antibacterial or antifungal agents, antioxidants, chelating agents, and inert gases), isotonic agents, absorption retardants, salts, pharmaceuticals, pharmaceutical stabilizers, gels, binders, excipients, disintegrants, lubricants, sweeteners, flavorings, dyes, liquids, and nutritional supplements, as well as similar materials and combinations thereof known to those skilled in the art. The pH and precise concentration of each component in the pharmaceutical composition are adjusted according to known parameters.

[0472] As used herein, the term "contact" refers to any combination of components, including any process of mixing the components to be contacted into the same mixture (e.g., adding them to the same container or solution), and does not necessarily require actual physical contact between the components. The components may be contacted in any order or in any combination (or sub-combination), and may include cases where one or more of the components are subsequently removed from the mixture, optionally before the addition of other components. For example, "contacting A with B and C" includes any and all of the following: (i) mixing A with C and then adding B to the mixture; (ii) mixing A and B into a mixture; removing B from the mixture and then adding C to the mixture; (iii) adding A to a mixture of B and C. "Contacting" a target nucleic acid or cell with one or more reaction components (e.g., Cas protein or guide RNA) includes any and all of the following: (i) contacting the target or cell with a first component of the reaction mixture to form a mixture; then adding other components of the reaction mixture in any order or combination; and (ii) the reaction mixture is fully formed before being mixed with the target or cell.

[0473] When describing the configuration of CRISPR-effect protein, "split" or "split state" refers to the CRISPR and the effector peptide being separated into two separate polypeptide chains. In other words, the CRISPR and the effector protein are not covalently fused to form a single polypeptide.

[0474] As used in this article, the term "mixture" refers to a combination of multiple elements that are interspersed and not arranged in a specific order. Mixtures are heterogeneous and cannot be spatially separated into their distinct components. Examples of elemental mixtures include multiple different elements dissolved in the same aqueous solution, or multiple different elements randomly or without a specific order attached to a solid support and whose different elements are spatially indistinguishable. In other words, mixtures are not addressable.

[0475] As disclosed herein, several numerical ranges are provided. Unless the context explicitly specifies otherwise, each intermediate value (accurate to one-tenth of the lower limit unit) between the upper and lower limits of a range is also explicitly disclosed. Each smaller range between any stated numerical value or an intermediate value within a range and any other stated numerical value or intermediate value within that range is included within the scope of this disclosure. The upper and lower limits of these smaller ranges may be independently included within or excluded from the range; and each range is also included within the scope of this disclosure when one of two limits is included, neither is included, or both are included in the smaller range, but any explicitly excluded limit within the range is permitted. If a range contains one or both limits, the range excluding one or both limits is also included within the scope of this disclosure. The term “about” generally refers to plus or minus 10% of the stated numerical value. For example, “about 10%” could mean a range of 9% to 11%, and “about 20” could mean 18% to 22. Other meanings of “about” may be apparent from the context, such as rounding, so, for example, “about 1” could also mean 0.5 to 1.4.

[0476] Example

[0477] Materials and Methods

[0478] Cells and cell culture conditions

[0479] HEK293T and K562 cells were purchased from ATCC (HEK293T, CRL-3216; K562, CCL-243). HEK293T cells were incubated at 37°C with 5% C Under these conditions, K562 cells were cultured and maintained in Dulbecco modified Eagle medium (Thermo Fisher Scientific) supplemented with 10% fetal bovine serum, 1× glutamine (Thermo Fisher Scientific), and 1× antibiotic-antifungal solution (Thermo Fisher Scientific). K562 cells were cultured in RPMI 1640 medium supplemented with 10% fetal bovine serum and 1× glutamine.

[0480] transfection

[0481] Electroporation transfection was performed using the Neon transfection system (Thermo Fisher Scientific) according to the manufacturer's instructions. In short, 2×1 Each cell was transfected with a mixture of approximately 2 μg of Cas9 / RT, gRNA / magRNA, template RNA, and free target plasmid construct DNA. For Figures 17 to 29 In the experiments, the amount of DNA used for each expression vector in each electroporation was: 250 ng gRNA or magRNA or an equimolar amount of pegRNA expression vector, 250 ng RNA template expression vector, 500 ng nCas9(H840A)-RT fusion protein expression vector, 250 ng MCP-RT fusion expression vector, and 100 ng target free plasmid DNA (sometimes 500 ng if targeting a free EGFP sequence, as shown). Figure 30 As in the experiments shown in the following figures, the amounts of DNA used for the nCas9(H840A)-RT, MCP-RT fusion expression vectors, and free target plasmid DNA were the same as described above. The expression vectors for gRNA, pegRNA, magRNA, and RNA template differed and are described in detail in each experiment. Unless otherwise specified, the expression vectors for magRNA and the second nick gRNA were both 250 ng. The RNA template expression vector was 1000 ng, and the molar amount of the pegRNA expression vector was equivalent to 1000 ng of the RNA template expression vector. Figures 30-43 Under the most common conditions, the molar ratio of RNA template in the PE system to pegRNA scaffold in PE to RNA template in ME to magRNA scaffold in ME is approximately 1:1:1:0.25.

[0482] U6 expression cassettes of pegRNA, RNA template, magRNA, and gRNA were cloned into the pBS KS(+) vector (2964 bp), resulting in plasmids approximately 3500 bp in length. Since these plasmids share the U6 promoter and transcription terminator, the size differences between the pegRNA, gRNA, magRNA, and RNA template expression vectors are typically within 100 bp (total length approximately 3500 bp). CMV expression cassettes of proteins such as nCas9, nCas9-RT fusion, MCP-RT, and MCP-p65 were cloned into pcDNA3.1(+) (5.4 kb).

[0483] When comparing the activities of the ME and PE systems, the molar ratio of RNA template (in pegRNA) in the PE system: pegRNA scaffold in the PE system: RNA template in the ME system: magRNA scaffold in the ME system was approximately 1:1:1:0.25, as described above. A Student's t-test was used for comparison. This indicates that P < 0.05. This indicates that P < 0.01. This indicates that P < 0.001.

[0484] Following electroporation, cells were seeded into 12-well plates. Three days post-electroporation, EGFP-containing cells were observed under a fluorescence microscope, and representative images were taken. Flow cytometry analysis was performed to quantify the percentage of EGFP-containing cells. All experiments were conducted using two to three biologically independent replicates. To directly determine base changes in the target sequence, DNA sequences containing the target site (EGFP or endogenous gene) were amplified by PCR, analyzed by Sanger sequencing, and then quantified using EditR. Some samples underwent further next-generation sequencing (NGS) analysis. Specific experimental conditions are further described in the figure captions.

[0485] To match the design of magRNA and RNA templates

[0486] To design magRNAs and RNA templates targeting protein- or RNA-coding targets, the RNA template is typically a positive-strand sequence. This minimizes the likelihood of double-stranded RNA formation within the cell. In this case, the non-coding strand (or lower strand) is nicked and then extended using the positive-strand RNA template sequence. This is known as the lower strand nicking and extension pattern.

[0487] For the lower chain nick and extension pattern, the following are example steps for designing RNA templates and magRNA: a. Align two complementary target DNA strands, which contain the sense upper strand (5'-3' orientation, from right to left) and its complementary antisense lower strand (3'-5' orientation, from right to left). b. Identify the PAM motif on the lower strand and the cleavage site on the lower strand (the lower strand is a non-target strand, the upper strand is a target strand complementary to the guide sequence or spacer sequence, and the guide is located immediately at the 5' of the PAM). c. For example, to design an RNA template with a 13-nucleotide primer-binding sequence (PBS) and a 35-nucleotide RT extension: i. Identify the lower chain cleavage site (5' 3nt from PAM) ii. Identifying the upper PBS: Starting with the nucleotide complementary to the cleavage nucleotide (lower chain) (upper chain), count upwards to the 3' end for approximately 13 nt (ideally ending in C or G). iii. Up-chain RT extension: Starting from the complementary nucleotide of the cleavage nucleotide, count 35 nt upwards to the 5' end (ideally ending with G); or, depending on the nature of the deletion / insertion, select an appropriate region of the homologous sequence for RT extension. iv. Copy the complete 5'-RT extension sequence (35 nt) + PBS (13 nt) - 3' sequence from the upper strand into a new file. v. Modify the template by adding the desired mutation to the RT extension sequence in the above file. vi. The resulting sequence is a template for editing RNA. d. Using a 12nt anchor tag as an example, design a magRNA sequence: i. Locate PAM and bootstrap from the lower chain ii. Replication guide (lower chain, approximately 20 nucleotides, 5' end of PAM, replicating in the 5'-3' direction) iii. Attach the guide / spacer sequence to the 5' end of the gRNA scaffold (using the spCas9 gRNA scaffold as an example). iv. From the template, find a 12-nucleotide sequence to be anchored (e.g., the sequence starting 29 nt from the 5' end of the PBS sequence). v. Replicate the base strand (5'-3') that is complementary to the 12-nucleotide sequence. vi. Then attach it to the 3' end of the gRNA (using the spCas9 gRNA scaffold as an example, add the anchor tag to the 3' end of the gRNA without a linker). vii. The magRNA has a 5'-spacer / guide + gRNA scaffold + 3' template anchoring tag sequence (12 nt).

[0488] To design RNA templates and magRNAs with up-chain nicks, the following describes exemplary steps for designing up-chain nicks and extension patterns: a. For example, to design an RNA template with a 13-nucleotide primer-binding sequence and a 35-nucleotide RT extension length: i. Align two complementary target DNA strands, the two strands comprising an upper DNA strand (5'-3' orientation, from right to left) and a complementary lower DNA strand (3'-5' orientation, from right to left). ii. Identify the upper-strand PAM motif and cleavage site (the upper strand is a non-target strand, the lower strand is a target strand and contains a sequence complementary to the guide / spacer sequence in the gRNA / magRNA; the guide is located immediately at the 5' end of the PAM, and the cleavage site is 3 nt from the 5' end of the PAM). iii. Reading the lower chain: For PBS designed to be 13 nucleotides in length: start with the complementary nucleotide (lower chain) to the cleavage site (upper chain). Down to the 3' end Count 13nt iv. Down-chain RT extension length 35 nt: Starting from the complementary nucleotide at the cleavage site, count 35 nt down to the 5' end of the chain; or, depending on the nature of the deletion / insertion, select an appropriate region of the homologous sequence for RT extension. v. Copy the bottom sequence containing the PBS and RT extension sequences described above (copy in the 5'-3' direction) and paste it into the template plasmid. vi. Modify the template by adding the desired mutation to the RT extension sequence. vii. The obtained sequence is a template for editing RNA. b. For example, in designing magRNA sequences, a 12nt anchor tag is used: i. Identify the upper PAM motif and cleavage site (the upper strand is a non-target strand, the lower strand is a target strand and contains a sequence complementary to the guide / spacer sequence in the gRNA / magRNA, with the guide located immediately at the 5' of the PAM). ii. Replication bootstrapping (on-chain, 20nt, 5' end of PAM, replication in the 5'-3' direction) iii. Paste the guide sequence to the 5' end of the gRNA scaffold (using the spCas9 gRNA scaffold as an example). iv. From the template, find a 12-nucleotide sequence to be anchored (e.g., the sequence starting at 5' 29nt in the PBS sequence). v. Replicate the base strand sequence complementary to the 12 nucleotides to be anchored (replicating in the 5'-3' direction). vi. Attach it to the 3' end of the gRNA (using the spCas9 gRNA scaffold as an example, such as adding an anchor tag to the 3' end of the gRNA without a linker). vii. magRNA, with a 5'-spacer region / guide + gRNA scaffold + 3' template anchoring tag sequence (12 nt).

[0489] An RNA template or magRNA sequence is inserted between the U6 promoter sequence and the U6 transcription terminator of the expression plasmid to enable expression in cells. It should be noted that this description is intended to provide examples and not to limit how systems can be designed.

[0490] sequence

[0491] The following are exemplary sequences used in the examples.

[0492] The nicking enzyme spCas9-RT sequence (NLS-spCas9 (H840A)-MMTV-RT(mut)-NLS):

[0493] The saCas9-RT sequence of the cleavage enzyme (NLS-saCas9 cleavage enzyme (N580A)-MMTV-RT(mut)-NLS):

[0494] The cleavage enzyme spCas9 in the ME (NLS-spCas9 (H840A)-NLS) sequence

[0495] The cleavage enzyme SaCas9 in the ME (NLS-saCas9 (N580A)-NLS) sequence was split.

[0496] The sequence of the nickase saCas9(N580A)-KKH variant is the same as that of saCas9(N580A), the difference being the presence of three mutations (E782K / N968K / R1015H).

[0497] NLS-MCP-L25-MMLV(mut)-RT-NLS sequence:

[0498] Notice: In the above example, a nuclear localization signal (NLS) is added to the N-terminus or C-terminus of the fusion protein, and a linker peptide is added between the fusion moieties.

[0499] NLS-MCP-L25-p65-NLS sequence

[0500] wtEGFP and mutant EGFP genes:

[0501] Notice: a. The sequence encodes wild-type EGFP.

[0502] b. nfEGFP contains the A200G point mutation (A200G, Tyr66Cys). The A200 site in the wild-type EGFP gene is marked and underlined.

[0503] c. The nucleotides missing in the EGFP Δ4A, Δ20C, Δ35C, Δ50C, and Δ100G constructs are labeled (underlined). The relative numbers 4, 20, 35, 50, and 100 represent the relative distances from the lower chain NGG PAM (where N is marked as +1, corresponding to the T199 position).

[0504] d. EGFP insertion construct: EGFP-4nt-insertion, containing four nucleotides (ATAG) between nucleotides 179 and 180.

[0505] e. EGFP123Δ52nt contains deletions from nucleotides 124 to 175 (a deletion of 52 nucleotides).

[0506] U6 starter and terminator sequences:

[0507] Notice: A single underscore sequence is the U6 promoter sequence, and a double underscore sequence is used as the U6 terminator. Template RNA, gRNA, or magRNA sequences are inserted between the promoter and terminator sequences for expression.

[0508] U6, ribozyme HDV, U6 terminator:

[0509] Notice: The underlined sequence is the HDV ribozyme sequence. The magRNA is inserted between the U6 promoter and the HDV sequence.

[0510] HEK4 genome sequence, with the guide sequence underlined and PAM double-underlined:

[0511] HEK3 genome sequence, with the guide sequence underlined and PAM double-underlined:

[0512] The splice reporter contains a 5' half-EGFP-artificial intron-3' half-EGFP:

[0513] Notice: The first single-underlined sequence is the 5' half of the GFP sequence, and the last single-underlined sequence is the 3' half of the GFP sequence. The middle sequence is an artificial intron, with SAS and SDS sites at the intron-exon junctions.

[0514] The mouse Dmd exon 23 skip read reporter gene contains the 5' half of EGFP-Dmd intron 22-Dmd exon 23- The EGFP reporter sequence of the 23-3' half of the Dmd intron:

[0515] Notice: The first single underlined sequence is the 5' half of the GFP sequence, and the last single underlined sequence is the 3' half of the GFP sequence. The double underlined sequence is the mouse Dmd exon 23 sequence. The sequence between the 5' half of the GFP sequence and the exon 23 sequence contains mouse Dmd intron 22, while the sequence between the exon 23 sequence and the 3' half of the GFP sequence contains Dmd intron 23.

[0516]

[0517] The matching sequences for EGFP-200L-guided magRNAs have previously been listed in the section on matching gRNA examples.

[0518] The following is some additional information about magRNA:

[0519] HEK3 (site 3)

[0520] EGFP

[0521] HBB

[0522] Example 1

[0523] This embodiment describes a specific match-edit (ME) system that targets EGFP with a point mutation (A200G) at position 200. The system is delivered via a plasmid containing a U6 promoter for transcription of RNA template, magRNA, and control gRNA. Figure 17 A), and the CMV promoter for nCas9-RT fusion protein and EGFP expression ( Figure 17 A and 17B). The figure also lists and analyzes the target sites of EGFP 200L, with the lower chain PAM (3'-gga-5') underlined, followed by the guide sequence. Furthermore, Figure 17 C also indicates the mutated base pairs, cleavage sites, and primer binding sequences. Additionally, Figure 17D also shows the design of the RNA template, which includes a 3' primer-binding site (P14) with a sequence identical to that shown at the target site; and a 5' RT extension sequence, identical to the 5' end sequence of the primer-binding site at the target site, except for nucleotide changes. In the experiments correcting the A200G point mutation, two templates were used, each with a 5' RT extension of 29 or 57 nucleotides.

[0524] Example 2

[0525] This embodiment describes the experimental results of correcting point mutations using the ME system described in Example 1.

[0526] In short, cells were electroporated using plasmids expressing either the shown ME component or the control component. The percentage of cells expressing fluorescent EGFP was determined by flow cytometry. Figure 18 A), amplify the target EGFP DNA fragment, and compare it with the control ( Figure 18 B) or ME Figure 18 C) The complementary strands of the treated cells were sequenced. The results are as follows: Figure 18 As shown in A (lane 1: untreated cells; lane 2: cells expressing nfEGFP; lanes 3-6: cells expressing different ME components shown; lane 7: pilot editing control). Figure 18 B and 18C show the sequencing results for lanes 2 and 4, respectively.

[0527] like Figure 18 As shown in A, MEs with magRNA exhibited high gene editing efficiency, with a functional editing efficiency of 60% for EGFP mutations (lane 4). Furthermore, compared to gRNAs without a matching tag complementary to the template, magRNAs significantly improved editing efficiency (more than 3-fold improvement, lanes 3 and 4). Their efficiency was similar to that of the lead editing counterpart (no statistical difference between lanes 4 and 7).

[0528] Figure 18 B and 18C confirmed that the conversion from non-fluorescent EGFP to fluorescent EGFP was indeed caused by a point mutation of complementary base C to T. Importantly, the sequencing results ( Figure 18 (C) This clearly demonstrates the absence of flanking editing; that is, only the target C is converted to T, while nearby Cs remain unmutated. The lack of flanking editing is a unique advantage of this method compared to base editing. Direct Sanger sequencing yielded an editing efficiency of 26%, while functional EGFP detection via sorted fluorescent cells yielded 60%. The difference is due to the multiple EGFP gene copies per cell. The results clearly demonstrate the functionality, practicality, and effectiveness of the matching editing technique, which anchors the RNA template using a complementary matching tag added to the magRNA.

[0529] Example 3

[0530] This example compares the editing efficiency of long RNA templates (Ex57-P14) and short RNA templates (Ex29-P14) in correcting the A200G mutation (3 nucleotides downstream of the cleavage site).

[0531] Cells were electroporated using plasmids expressing the specified ME component or control component, and the cells were observed using a fluorescence microscope. Results are as follows: Figure 19 Figure A shows the fluorescent and bright fields of view. The percentage of cells expressing fluorescent EGFP was determined by flow cytometry. Results are as follows: Figure 19 As shown in B (1: Untreated cells; 2: Cells expressing nfEGFP; 4, 6 and 7: Cells expressing ME components with different templates and magRNAs; lanes 3 and 5: Cells expressing gRNA instead of magRNA).

[0532] As shown in the figure, although both templates are effective, the shorter template consistently demonstrates higher efficiency in correcting this point mutation. Figure 19 (4 in A and 19B compared to 6-7). Furthermore, we compared the efficiency of two different magRNAs that matched the middle region of the template (12M29) or the 5' end of the long template (Ex57-P14) (12M57), respectively. The data showed that the magRNA matching the inner sequence (12M29) was more efficient than the magRNA matching the distal 5' sequence (12M57) (7 compared to 6). In most subsequent experiments, we chose magRNA 12M29, which matched the inner region of the RNA template.

[0533] Example 4

[0534] This example demonstrates the effectiveness of ME in correcting deletion mutations by inserting the correct base pairs.

[0535] In summary, cells were electroporated with plasmids expressing the indicated ME component (using 12M29 magRNA) or the control component, and observed by fluorescence microscopy. The percentage of cells expressing fluorescent EGFP was determined by flow cytometry. Results are as follows: Figure 20 As shown in A and 20B. The target EGFP DNA fragment was amplified, and the complementary strand of the fragment was sequenced from ME-treated cells. Figure 20 C shows the DNA sequencing results from cells in group 3.

[0536] As shown in the figure, using different corrective RNA templates and 12M29 magRNA, ME effectively corrected deletion mutations at 4, 20, and 35 nucleotides downstream of the nick site, with the templates covering this distance ( Figure 20 3, 4, and 5 in A and 20B). Correction of deletion mutations was validated by direct Sanger sequencing. Figure 20 C). This study demonstrates that the ME system can be used for insert editing.

[0537] Example 5

[0538] This example compares the efficacy of two magRNAs in correcting deletions at positions 35 or 50 downstream of the nick site using a long RNA template (Ex57-P14).

[0539] Cells were electroporated using plasmids expressing ME or control plasmids, and the percentage of cells expressing fluorescent EGFP was determined by flow cytometry. Results are as follows: Figure 21 As shown in the figure. The results indicate that, in both cases, magRNA matching the proximal sequence of the RNA template (closer to the 3' promoter sequence, 12M29) exhibits higher editing efficiency than magRNA matching the distal 5' end of the RNA template (12M57). In other experiments, most deletion / insertion editing experiments also used magRNA 12M29.

[0540] Example 6

[0541] This example demonstrates the efficacy of ME in correcting deletion mutations 100 base pairs downstream of the nick site. This experiment used an RNA template (Ex111-P14) with a 111-nt extended sequence and 12M29 magRNA.

[0542] Cells were electroporated using EGFP Δ100G and ME plasmids. The target EGFP DNA fragment was amplified, and the complementary strand in ME-treated cells was sequenced. Results are as follows: Figure 22 As shown in the figure. The results clearly demonstrate that, using independent long RNA templates, ME effectively edits mutations far from the target site.

[0543] Example 7

[0544] This embodiment compares the effects of the matching tag length of magRNA and the matching position on the template on ME editing efficiency.

[0545] In short, cells were electroporated using a plasmid expressing ME or a control plasmid, and the percentage of cells expressing fluorescent EGFP was then determined by flow cytometry. Results are as follows: Figure 23 As shown in A and 23B.

[0546] like Figure 23As shown in A, cells were electroporated with either a single EGFP deletion plasmid Δ50C (EGFPΔ50) or EGFP Δ50 plus the ME system, wherein the matching length of the magRNA ranged from 9 nucleotides (9M) to 24 nucleotides (24M) in increments of 3 nt.

[0547] like Figure 23 As shown in B, cells were electroporated either untreated (NT), with EGFP-deleted plasmid Δ35C (EGFPΔ35) alone, or with EGFPΔ35 plus the ME system simultaneously. The matching sites of the magRNA in the template differed, with the first matching nucleotide located at positions 23, 29, 35, or 41 from the RT extension site (12M23, 12M29, 12M35, and 12M41, respectively). The matching tag length remained 12 nt, but the complementary positions of the matches differed. The ME template used in both studies was Ex57-P14.

[0548] The study found that matching tags of 9, 12, and 15 nucleotides exhibited the highest efficiency when the matching length varied from 9 to 24 nucleotides, while longer matching tags of 18, 21, and 24 nucleotides showed significantly reduced efficiency. See also Figure 23 A.

[0549] The study also found that by changing the matching position (from +23 to +23 (matching sequences between +12 and +23, 12-nt), +29, +35 to +41), the highest efficiency was achieved when the matching position in the template was +29 upstream of the extension start site (matching nucleotides +18 to +29 from the first template extension). See also Figure 23 B.

[0550] Example 8

[0551] This embodiment demonstrates the components and efficacy of an ME configuration in which the reverse transcriptase is not directly fused with an RNA-guided nicking enzyme, but is provided in a detached form, recruiting the MCP-RT fusion protein via the MS2 aptamer at the magRNA stem-loop. Experimental design and results are as follows. Figure 24 As shown in A-24C.

[0552] like Figure 24 As shown in Figure A, in the split-type RT-ME system, the reverse transcriptase is not fused with nCas9. The magRNA contains the MS2 aptamer at the stem-loop position. RT fuses with MCP via an adaptor peptide. Figure 24As shown in B, the gene to be edited is the EGFP gene, which contains a C deletion 50 nucleotides upstream of the 200-L nick site. Cells were electroporated using plasmids expressing split or fusion-type RTME, the EGFPΔ50 plasmid alone, or without treatment, as shown. The percentage of cells expressing fluorescent EGFP was determined by flow cytometry. The ME template was Ex57-P14; the matching tag was 12M29.

[0553] like Figure 24 As shown in C, ME exhibits lower but comparable efficacy to directly fused ME in correcting deletions 50 bp downstream of PAM.

[0554] Example 9

[0555] This embodiment demonstrates a second ME module that provides an incision only at a trans-position downstream of the target site of the first ME module, significantly improving the editing efficiency of the first ME module. Experimental design and results are as follows... Figure 25 As shown in A-25C.

[0556] Figure 25 A illustrates a specific bimodal ME system comprising a template (Ex57-P14), a first magRNA targeting a first target site, and a second gRNA (hence gRNA, not magRNA) targeting a second target site but without a matching tag. Functionally, the second ME module provides only a second nick. Two second target sites, 119-U and 151-U, were tested against the first target site 200-L. Figure 25 B), they are located approximately 80nt or 50nt downstream of the first target site, respectively.

[0557] Cells were electroporated with plasmids expressing a single module or dual ME, a single EGFP Δ50 plasmid, or without treatment, as shown. The percentage of cells expressing fluorescent EGFP was determined by flow cytometry. The ME template was Ex57-P14, and the matching tag was 12M29. Figure 25 As shown in Figure C, the results clearly demonstrate that the second cut-out module significantly improves editing efficiency, with the downstream 80 nt module performing even better.

[0558] Example 10

[0559] This embodiment demonstrates that matched editing technology effectively edits the endogenous locus HEK4. Experimental design and results are as follows: Figure 26 As shown in AC.

[0560] In short, the cells were untreated ( Figure 26A), or electroporation using an ME system with a HEK4 target site, wherein the ME system contains gRNA without a matching tag ( Figure 26 B) or contains magRNA with a suitable tag that matches the template. Figure 26 C). Three days after electroporation, genomic DNA was extracted. The target fragment amplified by PCR was sequenced, and mutations were quantified using the EditR program.

[0561] like Figure 26 As shown in AC, using a magRNA targeting HEK4 and a template designed to change a G base to A at the +5 site (upstream of the cleavage site), ME resulted in a 10% G-to-A base change. Figure 26 C). As a control, no editing was observed when gRNA was used instead of magRNA. Figure 26 B). Figure 26 Figure A shows the HEK4 site sequence of untreated cells. In the figure, the arrows indicate the percentage of editing efficiency. Notably, for endogenous gene editing, editing using gRNA (lacking a matching tag) showed no editing effect, while ME using magRNA showed a 10% editing efficiency, indicating that the matching tag plays a crucial functional role in matched editing.

[0562] Example 11

[0563] This example demonstrates that ME is effective in deletion and editing.

[0564] This experiment constructed an EGFP reporter gene containing a 4-nucleotide frameshift insertion starting at position 180. Cells were treated as follows: no treatment; electroporation using only the expression vector with the 4-nucleotide frameshift insertion at position 180 (EGFP 180 Ins 4); electroporation using EGFP 180 Ins 4 with a matching editor (Ex29M12, using 12M29 magRNA and Ex29-P14 template); or electroporation using a leader editor (peg_Ex35). The percentage of fluorescent cells was determined by flow cytometry. Results are as follows. Figure 27 As shown.

[0565] like Figure 27 As shown, when cells were electroporated with the EGFP-inserted construct, only extremely low levels of fluorescent GFP were observed due to frameshift mutations. In contrast, when the construct was co-electroplated into cells with either the ME system or the lead editing system, the frameshift deletion was eliminated, and high levels of EGFP expression were detected in the cells. These results indicate that ME is effective in deletion editing.

[0566] Example 12

[0567] This embodiment demonstrates that both the single-ME and dual-ME systems effectively delete splicing regulatory sequences and splicing donor sites (SDS), leading to exon skipping. The dual-ME system using magRNA-gRNA pairing is more effective than the single-ME system. Experimental design and results are as follows. Figure 28 As shown in A and 28B.

[0568] A splicing reporter construct was constructed by splitting the EGFP coding sequence into a 5' half-gene and a 3' half-gene and linking them with artificial introns. The artificial introns contain a splice donor site (SDS) at the 3' end of the 5' half-gene and a splice acceptor site (SAS) at the 5' end of the 3' half-gene. See also Figure 28 Figure A shows a schematic diagram of a GFP-based splicing reporter. This splicing reporter contains a 5'-half GFP and a 3'-half GFP, linked by introns. SDS and SAS are labeled in the figure.

[0569] A mouse Dmd exon 23 skip-read reporter construct was constructed by inserting exon 23 of the mouse Duchenne muscular dystrophy (Dmd) gene and its introns SAS and SDS sites into an artificial intron sequence of a splicing reporter. Results showed that the mDmd exon 23 skip-read reporter construct resulted in the expression of a protein in which the amino acid sequence encoding exon 23 is located between the 5' and 3' halves of the EGFP gene, and this protein was non-fluorescent. The deletion of splicing regulatory sites (e.g., exon 23 SDS) led to exon 23 splicing skip-read, resulting in functional EGFP.

[0570] The ME single-module and ME dual-module systems were designed with a 29bp sequence missing the SDS region containing exon 23 / intron 23. Similar designs were also performed using the leader editor in the single-module PE2 and dual-module PE3 configurations. Results are as follows... Figure 28 As shown in B.

[0571] like Figure 28 As shown in B, a single ME module resulted in 40% of EGFP-expressing cells (denoted as magLow), and the addition of a second ME module (which contains only gRNA and no magRNA, resulting in a second nick) further increased the number of EGFP-producing cells to 70% (denoted as magLow+sgUp).

[0572] PE2 produced results similar to those of the single-module ME system, while PE3 (which introduced a second cut in addition to the first cut caused by PE2) was more effective than PE2 but less effective than the dual-module ME.

[0573] Example 13

[0574] This embodiment ( Figure 29 This study compared the editing efficiency of single MEs, dual MEs with magRNA-gRNA pairing, and dual MEs with magRNA-magRNA pairing when complexed with a long RNA template (Ex111-P14) in correcting deletion mutations. These ME systems are compared with... Figure 25 The systems shown are similar, except that in the last small diagram experiment, the 3' end of the second gRNA also contains a second matching tag.

[0575] The results showed that cells were treated as follows: no treatment, or with a mutant EGFP expression vector (with a nucleotide deleted at position 50 relative to the nick site). 50) Electroporation was performed, either using a mutant EGFP expression vector with a single-module ME (magLow), a dual-module ME (magLow+sgUp119) with magRNA-gRNA pairing, or a dual-module ME (magLow+magUp119) with magRNA-magRNA pairing. All treatments used an RNA template with an extension length of 111 nt. In magRNA-magRNA pairing, the matching tag sequences in the two magRNAs were complementary to the two individual sequences in the RNA RT template.

[0576] The results showed that the dual-ME system with a second magRNA had the highest editing efficiency, followed by dual-ME systems with magRNAs paired with gRNAs, and finally single-module ME systems.

[0577] Example 14

[0578] This example demonstrates the editing of the endogenous site HEK3 in HEK293T cells using a dual-ME system. Importantly, it also shows the different contributions of RNA template and magRNA to gene editing efficiency.

[0579] like Figure 30 As shown in Figure A, a template designed to introduce two point mutations, 5G>T and 12G>C, targeting the HEK3 site was constructed. The template T43-P13-(5GT-12GC) contains a 43 nt RT extension sequence and a 13 nt primer-binding sequence. Both point mutations are transversion mutations, one of which (5G>T) aims to disrupt the PAM motif to avoid redundant editing. Based on this template sequence, magRNA_12M42 was designed, with its upstream guide as shown in Figure A. Figure 30 As shown in A, it has 12nt complementary to the 5' end of the T43 template. Figure 30Image A shows the template, magRNA, and expression vector with a guide second nick gRNA. These plasmids and the nCas9(H840A)-RT fusion plasmid were introduced into HEK293 cells via electroporation. Genomic DNA was extracted from treated and untreated cells three days later. DNA fragments containing HEK3 were amplified by PCR and analyzed by Sanger sequencing. Figure 30 B shows significant reversal editing. Furthermore, the editing efficiency of 5G>T and 12G>C is similar. The editing efficiency of 5G>T is used as a quantitative value ( Figure 30 C).

[0580] To determine the respective contributions of RNA template and magRNA to editing efficiency, the amounts of RNA template and magRNA vector were varied. For example... Figure 30 As shown in Figure C, maintaining the magRNA vector at 250 ng while increasing the RNA template vector from 250 ng to 1000 ng significantly improved editing efficiency. Conversely, further increasing the magRNA vector from 250 ng to 500 ng had minimal impact on editing efficiency. Similarly, when the RNA template vector remained at 250 ng, increasing the magRNA vector to 500 ng had no significant effect on editing efficiency (data not shown).

[0581] This example demonstrates that the dual-ME system effectively introduces transversion point mutations at the endogenous HEK3 site. Importantly, it also shows that RNA template and gRNA contribute differently to editing efficiency. Optimizing the ratio of RNA template to gRNA can optimize genome editing results.

[0582] Because the ratio of 250 ng magRNA vector, 1000 ng template vector, and 250 ng second-cut gRNA vector produced excellent editing results in HEK293 cells, we have used this ratio for most of our subsequent studies (described below) unless otherwise stated.

[0583] Example 15

[0584] This embodiment demonstrates a method for introducing an insertion mutation at the endogenous HEK3 site in HEK293 cells using a dual ME system. The design of the ME components is similar to that in Example 14, except that templates T46, T50, and T79 contain 3-nt, 7-nt, and 36-nt sequences, respectively, for introducing the insertion at the +5 position. Figure 31 As shown in Figure A, the 3-nt CTT is a missing trinucleotide in the delta508 allele of the cystic fibrosis gene CFTR. The 7-nt ACTCAGT is a shared binding site for the transcription factor AP-1. The 36-nt encodes a peptide in the variable region of the TCR, used to recognize influenza virus epitopes.

[0585] like Figure 31 B and Figure 31 As shown in Figure C, the insertion editing efficiency at the HEK3 site was significant. This experiment used 250 ng magRNA and 1000 ng RNA template vectors. Furthermore, inserting additional MS2 into the gRNA of the second nick module improved editing efficiency, but the reason is unclear. This example demonstrates that dual-ME efficiently delivers the insert fragment to the endogenous genomic target site.

[0586] Example 16

[0587] This example demonstrates that polymerases other than MMLV-RT can be used in the ME system, and that magRNA can recruit DNA templates for matching and editing. We constructed a split-type dual ME system, whose components are as follows: Figure 32 As shown in Figure A. nfEGFP with the A200G mutation was used as a reporter ( Figure 32 (B) We designed an RNA or DNA template, T29-P14-3GA, to correct this mutation. The magRNA contained the MS2 aptamer in the stem-loop region for recruiting the polymerase that fuses with MCP. In addition to MMLV-RT, R2Bm (retrotransposon RT) and the Helraiser variant (polymerase of DNA transposon Heliton) were also fused with MCP. The trans-double-ME expression vector was introduced into HEK293 cells by electroporation, and the editing efficiency of the edited fluorescent EGFP-expressing cells was detected by flow cytometry. When using a DNA template, single-stranded DNA oligonucleotides were used instead of the expression vector.

[0588] like Figure 32 As shown in Figure C, experiments indicate that R2Bm and Helraiser also possess significant editing activity, although lower than MMLV-RT. Furthermore, when using Helraiser, the DNA template effectively promotes editing, but its editing efficiency is significantly lower compared to MMLV-RT and R2Bm binding to the RNA template.

[0589] Example 17

[0590] This embodiment demonstrates that additional effectors can be recruited into a dual-ME system to improve editing efficiency. For example... Figure 33 As shown in Figure A, a dual-ME system targeting the HEK3 site was constructed to introduce two mutations, as described above. Figure 30 Furthermore, a second-cut gRNA containing the MS2 aptamer was constructed to recruit the pioneer transcription factor p65, which fuses with MCP. This ME system was introduced into HEK293T cells, and genome editing efficiency was determined by Sanger sequencing.

[0591] like Figure 33 As quantified by C, this embodiment demonstrates that recruiting p65 into the dual-ME system via a second gRNA significantly improves editing efficiency. Interestingly, even without MCP-p65, the presence of MS2 in the second gRNA alone significantly improves editing efficiency. This phenomenon has also been observed in other editing sites and cell lines (data not shown here). The mechanism is unclear, but MS2 may contribute to stabilizing the structure of the second CRISPR complex.

[0592] Example 18

[0593] This embodiment compares the editing efficiency of the split-type dual-ME system and its corresponding split-type leader editing (PE) system on the HEK3 site in HEK293T cells.

[0594] like Figure 34 As shown in A, using and Figure 30 Using the same template and magRNA described above, a split-type dual-ME system was constructed for introducing two point mutations at the HEK3 site. In this dual-ME system, the MCP-RT fusion product is recruited by a second nick gRNA containing the MS2 aptamer, as described above. Figure 34 As shown in C, the corresponding split PE3 was constructed using pegRNA, which contains the same guide sequence and RT initiation / extension sequence as in MEmagRNA and the RNA template, such as... Figure 34 As shown in B and 34C.

[0595] The BE system, containing 1000 ng of RNA template vector and 250 ng of magRNA vector, was introduced into cells. In the same experiment, the corresponding PE system, containing an equimolar amount of pegRNA vector (equivalent to 1000 ng of RNA template vector), was introduced into HEK293 cells. The molar ratio of RNA template in the PE system: pegRNA scaffold in the PE system: RNA template in the ME system: magRNA scaffold in the ME system was approximately 1:1:1:0.25. Editing efficiency was determined and displayed. Figure 34 D.

[0596] Clearly, the split-type dual-ME system exhibits higher editing efficiency than the corresponding PE system when editing the HEK3 site. To rule out the possibility that high-level pegRNA vectors might not be preferred for the PE system, we down-tied the pegRNA concentration to approximately 250 ng. The results showed that the PE system achieved a higher genome editing efficiency than lower-level pegRNA vectors when the pegRNA vector molar concentration was equivalent to 1000 ng of RNA template vector (data not shown).

[0597] This embodiment demonstrates that, under their respective advantages, the split-type dual-ME system has higher editing efficiency than its corresponding split-type PE3 system in editing HEK3 sites.

[0598] Example 19

[0599] This implementation example Figures 35 to 37 As shown, the ability of the dual-ME system and its PE3 counterpart to introduce five point mutations into the endogenous HEK3 site in HEK293T cells was compared. In this embodiment, editing efficiency was determined using Sanger sequencing and next-generation sequencing (NGS). Furthermore, NGS data were used to analyze the formation of unintended insertions at the target sites.

[0600] Figure 35 This demonstrates a method for constructing a dual ME targeting HEK3. The guide and matching tags of the magRNA are shown. Figure 30 The same as described in A. However, the template T43-P13-(5G>T-12G>C-18A>C-24T>C-30C>T) is different, as it is designed to introduce five point mutations into the 43nt RT extension sequence, such as... Figure 36 As shown in Figure A, the corresponding PE3 contains a pegRNA, which includes the same guide sequence as the magRNA and the same 3' extension sequence as the ME RNA template. Common components of ME and PE include a second nick gRNA and an sp-nCas9-RT fusion expression vector.

[0601] Example 18 ( Figure 34 HEK293T cells were introduced into ME and PE systems with the same proportions as described in [reference needed]. The concentration of the second gRNA in the ME and PE systems was also altered (250 ng vs 500 ng, respectively). Three days after electroporation, genomic HEK3 DNA was amplified by PCR and analyzed using Sanger sequencing and NGS.

[0602] like Figure 36 As shown in Figure B, the ME system introduced five point mutations, resulting in significant genome editing. Furthermore, Figure 36 Quantitative analysis of Sanger sequencing results in C showed that the patterns of these five mutations were similar under all conditions. Under both second gRNA conditions (250 ng and 500 ng), the editing efficiency of the ME system was higher than that of the PE system. Figure 36 D shows that the editing efficiency of the ME and PE systems at the 5G>T editing site is statistically significant. The same statistical significance was observed between the ME and PE systems at the other four mutation sites (data not shown).

[0603] Under these conditions, NGS of representative samples yielded similar but more accurate results, such as Figure 37 As shown in Figure A, among the five point mutations, editing efficiency decreased with increasing distance between the mutation site and the promoter site, which is consistent with the editing mechanism. However, similar to the conclusions of the Sanger sequencing analysis, the NGS results also showed that, under both conditions, regardless of the amount of the second gRNA, the editing efficiency of the ME system was higher than that of the PE system.

[0604] A unique characteristic of ME is that modular ME templates are inherently "scar-free," while PE peg templates, during reverse transcription into the gRNA scaffold, mechanistically tend to introduce the 3' end of the gRNA scaffold sequence into the edited end. This has been confirmed by NGS data presented in this paper. The inventors analyzed the insertion of edited reads at the ends of the template sequence and counted the number of reads containing three or more nucleotides identical to the 3' end of the gRNA scaffold. Figure 37 As summarized in Table B, under the condition of 250 ng of the second gRNA vector, the dual ME system only exhibited background gRNA scaffold sequence insertion (identical to untreated cells). However, the PE system generated approximately 350-fold more reads containing insertions identical to the 3' end of the gRNA scaffold. The ratio of these reads to edited reads was approximately 350-fold higher in the PE system compared to the ME system. Figure 37 B). When the second gRNA vector was increased to 500 ng, the insertion rate of the ME system increased to above the background level, likely due to the increase in DSBs. However, under the same conditions, the ratio of these reads to edited reads was approximately 45 times lower than that of its corresponding PE3 system. Figure 37 B). Figure 37 C shows a significant difference between the ME and PE systems in producing unexpected gRNA scaffold sequence insertions at the end of the editing site.

[0605] This embodiment demonstrates that the dual-ME system is more efficient than its corresponding PE3 system in simultaneously introducing multiple point mutations, including transversion and transition point mutations, at multiple dispersed locations within the HEK3 site. Furthermore, this embodiment directly and explicitly demonstrates a significant difference between the ME and PE systems in terms of "scarring" at the end of the edit site. It confirms that the PE system tends to introduce the 3' gRNA scaffold sequence to the end of the edit site, while the ME system does not exhibit this tendency.

[0606] Example 20

[0607] This embodiment compares the ability of the dual-ME system and its corresponding PE3 system to edit endogenous HEK3 sites in HEK293T cells by deleting or inserting 3nt sequences. The dual-ME system was constructed to... Figure 30The same approach is described in the text to target HEK3, the difference being that the template contains 3-nt (CTT) insertions or 3-nt (GCA) deletions, such as... Figure 38 As shown in Figure A, the corresponding PE3 system is constructed similarly, with the same template sequence covalently linked to the 3' end of the pegRNA.

[0608] The dual-ME system or its corresponding PE3 system was introduced into HEK293T cells, and their editing efficiency in terms of insertion or deletion was analyzed and compared using Sanger sequencing. Figure 38 The results of B indicate that the dual-ME system has higher editing efficiency than its corresponding PE3 system in both 3-nt insertion and 3-nt deletion at the HEK3 site.

[0609] Example 21

[0610] This example demonstrates the efficacy of the ME system in different mammalian cells (K562 cells). This example also compares the editing efficiency of the dual-ME system with its corresponding PE3 system in introducing nucleotide insertions into free EGFP plasmids containing 1-nt deletions at different sites (i.e., Δ4A, Δ20C, and Δ35C relative to the EGFP nick site at G202). Figure 17 C). The construction method of the dual-ME system is as follows: Figure 25 As shown (200L + 119U)( Figure 25 Similar experiments were performed in HEK293T cells. The corresponding PE3 system contains pegRNA with the same guide and template as the ME system. The ME or PE system, along with a 1-nt deleted frameshift EGFP expression plasmid, was electroporated into K562 cells. Gene editing efficiency was quantified by flow cytometry measuring the percentage of fluorescent EGFP-expressing cells, such as... Figure 39 As shown.

[0611] Figure 39 This indicates that the dual ME system effectively introduces a single nucleotide insertion at multiple deletion sites in K562 cells. Furthermore, the results show that the ME system is more efficient than its corresponding PE system when introducing these insertions at Δ20 and Δ35 sites. When introducing insertions at the Δ4 site, the efficiency of the ME and PE systems is approximately the same.

[0612] Example 22

[0613] This embodiment demonstrates that the dual-ME system effectively induced a 29-nt deletion in free plasmids of K562 cells. Furthermore, this embodiment compares the editing efficiency of this dual-ME system with its corresponding PE system. The components and experimental procedures of ME and PE are as follows... Figure 28As shown, the same experiment was performed in HEK293T cells. This example replicated the previous experiment performed in HEK293T cells in another mammalian cell type, K562 cells. Figure 28 Editing efficiency was measured by the percentage of GFP-positive cells. Figure 40 A) and Sanger sequencing ( Figure 40 B) Quantitative analysis. Figure 40 A and Figure 40 B further demonstrates that the dual-ME system is more effective than its counterpart, the PE3 system, in generating 29-nt deletions in K562 cells. Figure 28 Similar phenomena were also observed in HEK293T cells.

[0614] Example 23

[0615] This embodiment compares the effectiveness of the dual ME system and its corresponding PE3 system in editing the endogenous HEK3 site in K562 cells by deleting or inserting the 3-nt sequence. Editing efficiency and insertion rate resulting from RT template extension into the gRNA scaffold were compared. ME components and... Figure 38 The procedure was identical to that described in the study (which was performed in HEK293T cells). The edited genomic DNA was subjected to Sanger sequencing and next-generation sequencing (NGS).

[0616] Figure 41 A and 41B show that the ME system is more efficient than the other system in inducing CTT insertion and GCA deletion at the HEK3 site in K562 cells. Furthermore, Figure 41 C shows that the PE system produces two orders of magnitude more gRNA scaffold templated insert fragments at the HEK3 editing site end than the ME system.

[0617] This example demonstrates that the dual-ME system effectively introduces deletions and insertions in K562 cells. Furthermore, it shows that the dual-ME system is more effective than its counterpart, the PE3 system, in introducing 3-nt deletions and insertions at the HEK3 site. Importantly, this example again directly and definitively demonstrates, in different mammalian cell types, a significant difference in “scarring” formation at the edit site ends between the ME and PE systems. The PE system is approximately two orders of magnitude more prone to introducing 3' gRNA scaffold sequences to the edit ends than the ME system.

[0618] Example 24

[0619] This embodiment demonstrates that the dual ME system effectively introduces a mutation into K562 cells at the endogenous HBB (hemoglobin β) gene locus, a treatment-related gene locus near the E6V mutation in sickle cell anemia. Figure 42Image A shows the HBB sequence near the E6V (GAG>GTG) mutation site, with annotations for the magRNA guide sequence, second nick guide, nick site, and the expected +5G>T mutation encoded in the RNA template. In sickle cell anemia E6V, a mutation occurs at the A base immediately adjacent to the 5' end of +5G.

[0620] The dual-ME system was electroporated into K562 cells, and the editing results were analyzed using Sanger sequencing. Figure 42 B shows that the dual-ME system can efficiently introduce +5G>T transversion base mutations at the HBB endogenous therapeutic site in K562 cells.

[0621] Example 25

[0622] This example tested whether ME could use CRISPR complexes that are not part of the spCas9 system. This example demonstrates that the Cas9 ortholog saCas9 is effective when used in a dual ME system. Furthermore, this example shows that the dual saCas9 ME system is more effective than its corresponding saPE3 system in editing EGFP mutations.

[0623] A dual ME was constructed using the Cas9 ortholog nickase saCas9 (N580A)-RT fusion. Correspondingly, magRNA and a second nick gRNA were constructed using the saCas9 gRNA scaffold. The gene to be edited was EGFP, which contained a single nucleotide deletion (Δ100G) at position 100 of its coding sequence. Figure 43 A shows the targeting sequence, saCas9 PAM, magRNA guide, and second nick gRNA guide. The template was designed to introduce a deleted G at position 100 and an independent silencing mutation. For the PE3 system, the pegRNA uses the same guide as the magRNA, and the 3' RT extension sequence and primer-binding sequence are identical to the RNA template sequence of ME.

[0624] The ME and PE systems, along with the EGFP (Δ100G) expression vector, were electroporated into HEK293 cells. Three days later, flow cytometry analysis was performed to quantify the percentage of cells expressing fluorescent EGFP. Figure 43 BC). EGFP plasmid DNA surrounding the deletion region was amplified by PCR and analyzed using Sanger sequencing. Figure 43 D). Figure 43 B shows that the saCas9 dual-ME system significantly edited and corrected the deletion mutation. Figure 43 C and Figure 43 D indicates that the dual saCas9 ME is more efficient than its corresponding saPE3 system.

[0625] The descriptions of the above embodiments and preferred embodiments should be considered illustrative and not limiting of the present disclosure as defined by the claims. It should be understood that various variations and combinations of the above features may be employed without departing from the present disclosure as defined by the claims. Such variations are not considered to depart from the scope of the present disclosure, and all such variations are intended to be included within the scope of the following claims. All references cited herein are incorporated herein by reference in their entirety.

Claims

1. A gene editing complex for editing a target site in a target DNA molecule, comprising: (A) an RNA-guided nickase; (B) a reverse transcriptase; (C) an RNA template molecule comprising: (1) a template segment comprising a template sequence complementary to a target DNA sequence to be introduced into the target site, and (2) a priming segment complementary to the 3’ end of a nicked strand of the target DNA molecule; and (D) a matching gRNA (magRNA) molecule comprising (1) a guide sequence or spacer sequence complementary to a target sequence on a target strand of the target DNA molecule, (2) an RNA scaffold capable of binding to an RNA-guided nickase, and (3) an anchor tag sequence complementary to a segment in the RNA template.

2. The complex of claim 1, wherein the reverse transcriptase is covalently linked to the RNA-guided nickase.

3. The complex of claim 1, wherein the magRNA further comprises a protein binding motif capable of binding to an RNA interacting protein, and the reverse transcriptase is linked to the RNA interacting protein.

4. The complex of any one of claims 1-3, wherein the anchor tag in the magRNA is not polyN, wherein N is a repeating nucleotide A, C, U, or G.

5. The complex of any one of claims 1-4, wherein the RNA template does not comprise additional sequences other than the priming segment and template segment.

6. The complex of claim 1, wherein the RNA-guided nickase is a nickase variant of a Cas protein, or a nickase variant of an IscB, IsrB, or TnpB family endonuclease protein encoded by a transposon.

7. The complex of claim 6, wherein the Cas protein is selected from Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8al, Cas8a2, Cas8b, Cas8c, Cas9, CaslO, CaslOd, Cpf 1 (Casl2a), C2cl (Casl2b), C2c3 (Casl2c), CasY (Casl2d), CasX (Casl2e), Casl4 (Casl2f), CasPhi (Casl2j), Casl3a, Casl3b, Casl3c, Casl3d, Casl3x, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3 (or CasE), Cse4 (or CasC), Csl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cml, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Cszl, Csx15, Csf1, Csf2, Csf3, Csf4, Cu1966, and orthologs thereof.

8. The complex of claim 6, wherein the Cas protein is Cas9 and orthologs thereof.

9. The complex according to claim 6, wherein the cleavage enzyme variant of the Cas protein is Streptococcus pyogenes (Streptococcus pyogenes). Streptococcus pyogenes nCas9 (H840A) or Staphylococcus aureus ( Staphylococcus aureus )nCas9(N580A).

10. The complex of any preceding claim, wherein the reverse transcriptase is a naturally occurring reverse transcriptase from a retrovirus or a retrotransposon, or a variant thereof.

11. The complex of claim 10, wherein the reverse transcriptase is selected from Moloney murine leukemia virus (M-MLV), human immunodeficiency virus (HIV) reverse transcriptase, avian sarcoma leukosis virus (ASLV) reverse transcriptase, Rous sarcoma virus (RSV) reverse transcriptase, avian myeloblastosis virus (AMV) reverse transcriptase, avian erythroblastosis virus (AEV) helper virus MCAV reverse transcriptase, avian myelocytomatosis virus MC29 helper virus MCAV reverse transcriptase, avian reticuloendotheliosis virus (REV-T) helper virus REV-A reverse transcriptase, avian sarcoma virus UR2 helper virus UR2AV reverse transcriptase, avian sarcoma virus Y73 helper virus YAV reverse transcriptase, Rous associated virus (RAV) reverse transcriptase, and myeloblastosis associated virus (MAV) reverse transcriptase, Line 1 ORF2, R2Bm, and R20l.

12. The complex of claim 10 or 11, wherein the reverse transcriptase is MMLV-RT or a variant thereof.

13. The complex of any one of claims 1-12, wherein the anchor tag sequence is located at the 3’ end or the 5’ end of the magRNA.

14. The complex of claim 13, wherein the anchor tag sequence is covalently linked to the magRNA by a polynucleotide linker.

15. The complex of any preceding claim, wherein the anchor tag sequence in the magRNA is about 6-24 nt in length.

16. A system for editing a target site in a target DNA molecule, comprising (I) a first gene editing complex of any one of claims 1-15, and (II) a second gene editing complex.

17. The system of claim 16, wherein the second gene editing complex comprises (A) a second RNA-guided nickase, and (B) a gRNA molecule comprising a second guide sequence or spacer sequence and a second RNA scaffold capable of binding to the second RNA-guided nickase.

18. The system of claim 16, wherein the second gene editing complex comprises (A) a second RNA-guided nickase; and (B) a second magRNA molecule comprising (1) a second guide sequence or spacer sequence, (2) a second RNA scaffold capable of binding to the second RNA-guided nickase, and (3) a second anchor tag sequence.

19. The system of claims 17 and 18, wherein the second gene editing complex comprises (C) a second reverse transcriptase.

20. The system of any one of claims 16 to 19, wherein, the guide sequence of the magRNA in the first gene editing complex and the second guide sequence of the gRNA or the second magRNA in the second gene editing complex are complementary to two target sequences of the target DNA molecule, respectively.

21. The system of any one of claims 18-20, wherein, the second anchor tag sequence of the second magRNA is complementary to a second segment within the RNA template.

22. The system of any one of claims 16-21, wherein the 5’ end of the RNA template comprises the same sequence as the 3’ end of the strand of the target DNA molecule that is nicked by the second RNA-guided nickase.

23. The system of any one of claims 19-20, wherein, the second gene editing complex further comprises (D) a second RNA template.

24. The system of any one of claims 19 to 23, wherein, the second reverse transcriptase is linked to the second RNA-guided nickase.

25. The system of any one of claims 19 to 23, wherein, the second magRNA molecule or the gRNA molecule comprises a second protein binding motif capable of binding to a second RNA interacting protein, and the second reverse transcriptase is linked to the second RNA interacting protein.

26. A method of modifying a target DNA molecule in a cell, comprising contacting the target DNA molecule with the gene editing complex of any one of claims 1-15 or the system of any one of claims 16-25.

27. The method of claim 26, wherein the modification results in a point mutation in the target DNA molecule, an insertion in the target DNA molecule, a deletion in the target DNA molecule, or a combination thereof.

28. The method of any one of claims 26-27, wherein the cell is selected from the group consisting of an archaeal cell, a bacterial cell, a eukaryotic cell, a eukaryotic unicellular organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algal cell, an animal cell, an invertebrate animal cell, a vertebrate animal cell, a fish cell, a frog cell, a bird cell, a mammalian cell, a pig cell, a cow cell, a goat cell, a sheep cell, a rodent cell, a rat cell, a mouse cell, a non-human primate cell, and a human cell.

29. The method of any one of claims 26-28, wherein the cell is in or derived from a human or non-human subject.

30. A genetically engineered cell or progeny thereof obtained by the method of any one of claims 26-29.

31. The cell of claim 30, wherein the cell is selected from the group consisting of an archaeal cell, a bacterial cell, a eukaryotic cell, a eukaryotic unicellular organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algal cell, an animal cell, an invertebrate animal cell, a vertebrate animal cell, a fish cell, a frog cell, a bird cell, a mammalian cell, a pig cell, a cow cell, a goat cell, a sheep cell, a rodent cell, a rat cell, a mouse cell, a non-human primate cell, and a human cell.

32. The cell of claims 30-31, wherein the cell is selected from the group consisting of a pluripotent stem cell (PSC), an adult stem cell (ASC), a fibroblast cell, a chondrocyte cell, a keratinocyte cell, a hepatocyte cell, an islet cell, and an immune cell, including a T cell, a dendritic cell (DC), a natural killer (NK) cell, and a macrophage cell, derived from a human or non-human subject.

33. The magRNA of claims 1, 3, 4, 13, 14, or 15.

34. An RNA complex comprising the magRNA of claim 33 and the RNA template anchored by the magRNA of claims 1, 5, 21, 22, or 23.

35. A nucleic acid encoding one or both of: (i) the magRNA molecule of claim 33, and (ii) the RNA complex of claim 34.

36. A vector comprising the nucleic acid of claim 35.

37. A kit comprising: (i) packaging material, and (ii) one, two, or more of: the gene editing complex of any one of claims 1-15, the system of any one of claims 16-25, the cell of claims 30, 31, or 32, the magRNA molecule of claim 33, the RNA complex of claim 34, the nucleic acid of claim 35, and the vector of claim 36.

38. A pharmaceutical composition comprising: (i) a pharmaceutically acceptable carrier, and (ii) one, two, or more of: the gene editing complex of any one of claims 1-15, the system of any one of claims 16-25, the cell of claims 30, 31, or 32, the magRNA molecule of claim 33, the RNA complex of claim 34, the nucleic acid of claim 35, and the vector of claim 36. The magRNA molecule of claim 33, The RNA complex of claim 34, The nucleic acid of claim 35, and The vector of claim 36.

Citation Information

Patent Citations

  • Fusion protein constructs

    US20100063258A1

  • Crispr-CAS component systems, methods and compositions for sequence manipulation

    US20140179006A1

  • Crispr / CAS systems for genomic modification and gene modulation

    US20140273226A1

  • Crispr-based genome modification and regulation

    US20140273233A1

  • Targeted therapeutics

    US20150182596A1