Reverse transcription mediated gene editing system and uses thereof

By integrating CRISPR nuclease and reverse transcriptase into a gene editing system, and combining guide RNA and reverse transcription donor RNA, the accuracy and efficiency problems of gene editing in existing technologies have been solved, achieving efficient nucleotide substitution at targeted genomic sites, which is suitable for disease treatment and genome research.

CN122074087APending Publication Date: 2026-05-22ARBOR BIOTECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ARBOR BIOTECHNOLOGIES INC
Filing Date
2024-08-30
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems struggle to achieve efficient and accurate nucleotide substitutions in gene editing, especially in designed editing that targets specific genomic sites.

Method used

A fusion peptide system containing CRISPR nuclease and reverse transcriptase (RT) fragments, combined with guide RNA and reverse transcription donor RNA, was developed to achieve precise gene editing by genetically engineering the CRISPR nuclease to enhance its enzymatic activity and reduce the strictness of PAM recognition.

Benefits of technology

It enables highly efficient nucleotide substitution at targeted genomic sites, improving the accuracy and efficiency of gene editing, and is applicable to disease treatment, animal and plant reproduction, and genome function research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122074087A_ABST
    Figure CN122074087A_ABST
Patent Text Reader

Abstract

The present disclosure provides a gene editing system comprising (a) a fusion polypeptide or a nucleic acid encoding the fusion polypeptide, the fusion polypeptide comprising a CRISPR nuclease and a reverse transcriptase, and (b) an RNA molecule, the RNA molecule comprising a guide RNA and a reverse transcription donor RNA, or a nucleic acid encoding the RNA molecule. Also provided herein are methods of modifying a target gene of interest using the gene editing system.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 580,188, filed September 1, 2023; U.S. Provisional Application No. 63 / 580,168, filed September 1, 2023; U.S. Provisional Application No. 63 / 553,974, filed February 15, 2024; and U.S. Provisional Application No. 63 / 638,559, filed April 25, 2024. Each priority application in the priority applications is incorporated herein by reference in its entirety.

[0003] sequence list

[0004] This application contains a sequence list that has been electronically submitted in XML format and is hereby incorporated in its entirety by reference. The XML copy created on August 28, 2024, is named 063586-525001WO_SeqList_ST26.xml and has a size of 597.0 kilobytes. Background Technology

[0005] Clusters of regularly spaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) genes are collectively referred to as the CRISPR-Cas or CRISPR / Cas system, which is an adaptive immune system in archaea and bacteria that protects specific species from foreign genetic elements.

[0006] Reverse transcriptase (RT) is an enzyme that produces a DNA strand complementary to an RNA template. The combination of reverse transcriptase and the CRISPR / Cas system shows great potential in gene editing. CRISPR-guided reverse transcription allows the introduction of desired nucleotide substitutions at genomic sites.

[0007] Therefore, the development of efficient and accurate reverse transcriptase-CRISPR gene editing systems for disease treatment has attracted much attention. Summary of the Invention

[0008] This disclosure provides a reverse transcriptase-CRISPR-mediated gene editing system that successfully introduces designed nucleotide substitutions into target genetic sites. In some embodiments, the gene editing systems disclosed herein involve fusion peptides comprising a CRISPR nuclease fragment and a reverse transcriptase (RT) fragment, and optionally one or more nuclear localization sequences (NLS) and / or peptide linkers. The CRISPR nuclease peptide can be genetically engineered to possess favorable enzymatic activities (e.g., high indel activity and / or DNA cleavage activity, and precise gene editing as designed). Therefore, the gene editing systems provided herein are expected to demonstrate superior efficiency in inserting desired base substitutions at genomic sites of interest.

[0009] Therefore, one aspect of this disclosure is characterized by a gene editing system comprising: (a) a fusion polypeptide or a first nucleic acid encoding said fusion polypeptide, said fusion polypeptide comprising a CRISPR nuclease polypeptide and a reverse transcriptase (RT) polypeptide; and (b) an RNA molecule or a second nucleic acid encoding said RNA molecule, said RNA molecule comprising guide RNA (gRNA) and reverse transcription donor RNA (RT donor RNA). The CRISPR nuclease may be the reference nuclease of SEQ ID NO: 1 or a variant thereof. The gRNA comprises a scaffold sequence recognizable by the CRISPR nuclease and a spacer sequence specific to a target sequence within a genomic site of interest, said target sequence being located upstream of a protospacer adjacent motif (PAM). The RT donor RNA comprises a primer binding site (PBS) and a template sequence.

[0010] In some embodiments, the CRISPR nuclease polypeptide comprises the amino acid sequence of SEQ ID NO: 1. In other embodiments, the CRISPR nuclease polypeptide is a variant of SEQ ID NO: 1, the variant comprising: (i) one or more mutations in the HNH nuclease domain or RuvC nuclease domain of SEQ ID NO: 1, the one or more mutations reducing or eliminating the nuclease activity of the HNH nuclease domain or RuvC nuclease domain; (ii) one or more arginine and / or lysine substitutions, optionally one or more arginine substitutions; (iii) a combination of (i) and (ii);

[0011] In some cases, the fusion polypeptide comprises one or more nuclear localization signals (NLS) upstream or downstream of the CRISPR nuclease polypeptide, the RT polypeptide, or both. For example, the fusion polypeptide comprises a first NLS, the CRISPR nuclease polypeptide, the RT polypeptide, and a second NLS from the N-terminus to the C-terminus. In some instances, the fusion polypeptide may further comprise a peptide linker between the CRISPR nuclease polypeptide and the RT polypeptide.

[0012] In some instances, the fusion polypeptide includes a first peptide linker located between the CRISPR nuclease polypeptide and the RT polypeptide. Further, the fusion polypeptide may include a first NLS and a second NLS located at the N-terminus and / or C-terminus of the fusion polypeptide. In some cases, the fusion polypeptide may further include additional NLS, such as a third NLS and optionally a fourth NLS. Alternatively or additionally, the fusion polypeptide may further include a second peptide linker and optionally a third peptide linker. These peptide linkers may position (e.g., connect) the CRISPR nuclease polypeptide and / or the RT polypeptide to the first and / or second NLS. In some cases, these peptide linkers may position (e.g., connect) two NLSs.

[0013] The following provides specific exemplary configurations (from N-terminus to C-terminus) of the fusion peptides disclosed herein:

[0014] (i) the first NLS, the second NLS, the CRISPR nuclease polypeptide, the first peptide linker, and the RT polypeptide;

[0015] (ii) the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the first NLS, and the second NLS;

[0016] (iii) The first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the third NLS, the second peptide linker, and the second NLS;

[0017] (iv) the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the third NLS, the second NLS, the second peptide linker, and the first NLS;

[0018] (v) The first NLS, the second peptide linker, the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the third peptide linker, and the second NLS;

[0019] (vi) The first NLS, the second peptide linker, the CRISPR nuclease, the first peptide linker, the RT peptide, the third peptide linker, the third NLS, the fourth NLS, and the second NLS;

[0020] (vii) The first NLS, the second peptide linker, the RT polypeptide, the first peptide linker, the CRISPR nuclease polypeptide, the third peptide linker, the third NLS, the fourth NLS, and the second NLS;

[0021] (viii) the RT peptide, the first peptide linker, the CRISPR nuclease peptide, the third NLS, the second NLS, the second peptide linker, and the first NLS;

[0022] (ix) The first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the second peptide linker, the second NLS, and the RT polypeptide;

[0023] (x) The first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the third NLS, the second peptide linker, the RT polypeptide, the third peptide linker, and the second NLS;

[0024] (xi) The first NLS, the RT polypeptide, the first peptide linker, the second NLS, the second peptide linker, and the CRISPR nuclease polypeptide.

[0025] (xii) The first NLS, the RT peptide, the first peptide linker, the third NLS, the second peptide linker, the CRISPR nuclease peptide, the third peptide linker, and the second NLS;

[0026] (xiii) The first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, and the second NLS; or

[0027] (xiv) The first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the second peptide linker, the third peptide linker, and the second NLS.

[0028] In specific examples, fusion peptides can have (iv), (v), (vii), (ix), or (x) configurations (from the N-terminus to the C-terminus). See also Table 16 below.

[0029] In some cases, the length of the peptide linker between the CRISPR nuclease peptide and the RT peptide is about 20-80 amino acids.

[0030] In some embodiments, the CRISPR nuclease polypeptide comprises the variant of SEQ ID NO: 1. For example, the variant of SEQ ID NO: 1 may comprise one or more mutations in the HNH nuclease domain at positions D844, H845, and / or N868 relative to SEQ ID NO: 1. In one instance, the mutation is located at position H845 (e.g., an H845A substitution). In some instances, the mutation at D844 is an amino acid substitution of D844A, D844G, D844L, or D844S. In some instances, the mutation at H845 is an amino acid substitution of H845A, H845G, H845L, or H845S. In a specific example, the CRISPR nuclease polypeptide comprises the mutation at position H845 relative to SEQ ID NO: 1 (e.g., H845A) (e.g., comprising the amino acid sequence of SEQ ID NO: 32). In some instances, the mutation at N868 is an amino acid substitution of N868A, N868G, N868L, or N868S.

[0031] Alternatively or additionally, the CRISPR nuclease polypeptide comprises a bridged helix (BH) domain, a nucleic acid recognition (REC) domain, and a phosphate-locked ring (PLL), wedge (WED), and PAM interaction (PID) domains, wherein one or more arginine and / or lysine substitutions are optionally located in the BH domain, the REC domain, the PLL domain, the WED domain, the PID domain, or a combination thereof. For example, the CRISPR nuclease polypeptide contains up to 20 arginine and / or lysine substitutions relative to a reference CRISPR nuclease. In a specific example, the CRISPR nuclease polypeptide contains up to 15 arginine and / or lysine substitutions relative to a reference CRISPR nuclease. In a more specific example, the one or more arginine and / or lysine substitutions are located at positions K736, L784, Q812, N813, I857, and / or A919.

[0032] In some instances, the CRISPR nuclease polypeptide contains at least two arginine and / or lysine substitutions relative to a reference CRISPR nuclease. These two arginine and / or lysine substitutions are located at positions K736, L784, Q812, N813, I857, and / or A919. In some cases, the CRISPR nuclease polypeptide contains arginine and / or lysine substitutions at the following positions relative to a reference CRISPR nuclease:

[0033] (a) I857, L784 and K736 (e.g. I857R, L784R and K736R);

[0034] (b) I857, A919 and K736 (I857R, A919R and K736R);

[0035] (c) I857, N813 and L784 (I857R, N813R and L784R);

[0036] (d) I857, L784 and A919 (I857R, L784R and A919R);

[0037] (e) I857, N813 and K736 (I857R, N813R and K736R);

[0038] (f) I857 and N813 (I857R and N813R);

[0039] (g) L784, A919 and K736 (L784R, A919R and K736R);

[0040] (h) I857 and L784 (I857R and L784R); or

[0041] (i) I857 and A919 (I857R and A919R).

[0042] In a specific example, the CRISPR nuclease polypeptide may contain arginine substitutions relative to I857R, L784R, and K736R of SEQ ID NO: 1.

[0043] In some cases, the CRISPR nuclease polypeptides disclosed herein, in which mutations are introduced, may further comprise one or more mutations that enhance double-stranded nuclease activity relative to a reference CRISPR nuclease. Examples are provided in Table 19 below.

[0044] Alternatively or additionally, the engineered CRISPR nuclease peptides disclosed herein may contain or further contain one or more of the mutations described above for reducing the strictness of PAM recognition. In some cases, the one or more mutations for reducing the strictness of PAM recognition may be located at positions D61, A68, H494, L1117, D1144, S1145, G1227, E1228, S1327, A1332, R1343, R1345 and / or T1347 of SEQ ID NO: 1. In some instances, such mutations may comprise: (i) one or more arginine and / or lysine substitutions at positions D61, A68, H494, L1117, G1227, S1327, A1332 and / or T1347 of SEQ ID NO: 1, optionally with arginine substitution; (ii) one or more amino acid substitutions at positions D1144, S1145, E1228, R1343 and / or R1345 of SEQ ID NO: 1; or (iii) a combination of (i) and (ii). In specific instances, the one or more amino acid substitutions in (ii) may comprise optional substitutions at D1144L, S1145W, E1228Q, R1343P, R1345V and / or R1345Q relative to SEQ ID NO: 1.

[0045] In specific examples, engineered CRISPR nuclease peptides with lower PAM recognition strictness may contain the following combinations of mutations relative to SEQ ID NO: 1: L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and A68R. In other specific examples, engineered CRISPR nuclease peptides with lower PAM recognition strictness may contain the following combinations of mutations relative to SEQ ID NO: 1: L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and D61R. In other specific examples, engineered CRISPR nuclease peptides with lower PAM recognition strictness may comprise the following combinations of mutations relative to SEQ ID NO: 1: L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and H494R. Other exemplary engineered CRISPR nuclease peptides can be found in Table 20, each of which is within the scope of this disclosure.

[0046] The engineered CRISPR nuclease peptides disclosed herein recognize the 5'-NDR-3' PAM sequence, where N represents A, C, G, or U, D represents A, G, or T, and R represents G or A. In some cases, engineered CRISPR nuclease peptides with reduced PAM recognition strictness, as disclosed herein, can recognize the 5'-NGN-3' PAM sequence. See Example 10 below. In some examples, the PAM is 5'-NRG-3' or 5'-NRR-3', where N and R are as defined herein. In some specific examples, the PAM is 5'-NGG-3', where N represents any nucleotide. In other specific examples, the PAM can be 5'-TGC-3' or 5'-GGA-3'.

[0047] Any CRISPR nuclease polypeptide may contain the arginine and / or lysine substitutions disclosed herein, any nickase mutations in the HNH or RuvC nuclease domains also disclosed herein, any mutations that result in reduced PAM recognition strictness, or combinations thereof. For example, the CRISPR nuclease polypeptide may contain (a) one or more of the nickase mutations in the HNH nuclease domain at positions D844, H845, and / or N868 relative to SEQ ID NO: 1 (e.g., the mutation is located at position H845); and (b) one or more arginine and / or lysine substitutions relative to SEQ ID NO: 1 (e.g., located at positions I857, L784, and K736). For example, in some instances, the CRISPR nuclease polypeptide can contain... Contains (e.g., constitutes) a nickase mutation (e.g., H845A mutation) at position H845 relative to SEQ ID NO: 1. And arginine and / or lysine substitutions at position I857 (e.g., I857R substitution). In other instances, the CRISPR nuclease polypeptide may contain or further contain one or more mutations that cause a reduction in PAM recognition strictness (e.g., at positions L1117, D1144, G1227, E1228, A1332, R1345 and / or T1347 of SEQ ID NO: 1, and optionally at one or more positions D61, A68 and H494 of SEQ ID NO: 1).

[0048] Any CRISPR nuclease polypeptide provided herein may contain at least 90% of the same amino acid sequence as SEQ ID NO: 1. In some instances, the CRISPR nuclease polypeptide contains at least 95% of the same amino acid sequence as SEQ ID NO: 1. In still other instances, the CRISPR nuclease polypeptide contains at least 98% of the same amino acid sequence as SEQ ID NO: 1.

[0049] The RT polypeptide is Moloney murine leukemia virus (MMLV)-RT or a variant thereof. In one example, the MMLV-RT contains the amino acid sequence of SEQ ID NO: 53.

[0050] In some embodiments, the gene editing system disclosed herein comprises the fusion polypeptide. Alternatively, the system comprises the first nucleic acid encoding the fusion polypeptide. In some instances, the first nucleic acid is located on a vector, optionally a viral vector. In other instances, the first nucleic acid is a first messenger RNA (mRNA).

[0051] In any gene editing system provided herein, the spacer sequence in the gRNA of (b) may be 15-30 nucleotides in length. In one example, the spacer sequence may be 15-20 nucleotides in length. In a specific example, the spacer sequence may be about 17 nucleotides in length.

[0052] In some embodiments, the scaffold sequence comprises at least 85% identical nucleotide sequences to SEQ ID NO: 2. In one example, the scaffold sequence comprises the nucleotide sequence of SEQ ID NO: 2.

[0053] In some embodiments, the PBS in the RT donor RNA portion of the RNA molecule is 5-50 nucleotides long. In some instances, the PBS is 5-20 nucleotides long. In specific instances, the PBS may be 7-17 nucleotides long. In some embodiments, the PBS binds to a site of targeting PBS that is adjacent to or overlaps with the target sequence. For example, the site of targeting PBS is adjacent to or overlaps with the target sequence. In other instances, the site of targeting PBS is adjacent to the 5' of the PAM.

[0054] In some embodiments, the template sequence in the RT donor RNA portion of the RNA molecule may be 5-100 nucleotides in length. In some instances, the template sequence may be 15-25 nucleotides in length. In some embodiments, the template sequence in the RT donor RNA is homologous to the genomic site of interest and contains one or more nucleotide variations relative to the genomic site of interest. In some instances, at least one nucleotide variation may be localized within the target sequence. Alternatively, at least one nucleotide variation may be localized within the PAM.

[0055] In some embodiments, any RNA molecule provided herein (b) may further include a 3' extension. In some instances, the RNA molecule may further include a 5' protectant, a 3' protectant, or both, each of the 5' and 3' protectants forming a secondary structure, optionally a hairpin, circular, pseudoknot, or triplet structure.

[0056] In some instances, the RNA molecule in (b) comprises, from 5' to 3', the spacer sequence, the scaffold sequence, the template sequence, and the PBS. In other instances, the RNA molecule may comprise, from 5' to 3', the spacer sequence, the scaffold sequence, the template sequence, the PBS, and the 3' extension.

[0057] In some embodiments, the gene editing system disclosed herein may comprise the RNA molecule, which includes optionally one or more elements selected from gRNA, RT donor RNA, and other elements disclosed herein. Alternatively, the gene editing system may comprise the nucleic acid encoding the RNA molecule. In some instances, the nucleic acid is localized to a vector, optionally a viral vector.

[0058] In some embodiments, the gene editing system disclosed herein may comprise one or more lipid nanoparticles (LNPs) associated with one or more of elements (a)-(b). Alternatively, the gene editing system may comprise one or more viral vectors encoding one or more of elements (a)-(b), for example, optionally one or more adeno-associated virus (AAV) vectors.

[0059] This document also provides a pharmaceutical composition comprising any gene editing system provided herein, and a kit comprising elements (a) and (b) of a gene editing system as disclosed herein.

[0060] In other aspects, this disclosure features a gene-editing method comprising delivering a gene-editing system disclosed herein to a host cell to edit a genomic site targeted by the gRNA of the gene-editing system. In some embodiments, the host cell is cultured in vitro. In other embodiments, the host cell is located in a subject requiring gene editing.

[0061] Furthermore, this disclosure provides a fusion polypeptide comprising any CRISPR nuclease polypeptide described herein and any reverse transcriptase polypeptide also described herein. Such fusion polypeptides may comprise the amino acid sequence of SEQ ID NO: 55 or 57.

[0062] Additionally, this disclosure provides a nucleic acid encoding the fusion polypeptide disclosed herein. This nucleic acid may comprise the nucleotide sequence of SEQ ID NO: 54 or 56. In some instances, the nucleic acid is a vector, such as an expression vector.

[0063] Details of one or more embodiments of the invention are set forth in the following description. Other features or advantages of the invention will become apparent from the following drawings and the following detailed description of several embodiments, and also from the appended claims. Attached Figure Description

[0064] The following drawings form part of this specification and are included to further illustrate certain aspects of this disclosure. A better understanding of certain aspects of this disclosure can be achieved by referring to the drawings in combination with a detailed description of the specific embodiments presented herein.

[0065] Figure 1 This is a graph showing the gene editing efficiency of the reference CRISPR nuclease SEQ ID NO: 1 on exemplary target genes AAVS1, EMX1, and VEGFA.

[0066] Figures 2A-2D Includes quantitative gel images showing nuclease activity. Figure 2A Gel images captured using the 700 nm channel, showing the target strand (labeled with IR700 dye at the 5' end) of the target DNA substrate in vitro cleaved by a reference CRISPR nuclease, a putative HNH knockout cleavage enzyme, or a putative RuvC knockout cleavage enzyme. Figure 2B Gel images captured using the 800 nm channel, showing the non-target strand of the target DNA substrate cleaved in vitro by a reference CRISPR nuclease, a putative HNH knockout cleavage enzyme, or a putative RuvC knockout cleavage enzyme (labeled with IR800 at the 5' end). Figure 2C :use Figure 2A and Figure 2B Overlay images captured by the 700 nm and 800 nm channels. Figure 2D Quantification of the percentage of cleaved target DNA and non-target DNA produced by reference CRISPR nuclease, putative HNH knockout cleavage enzyme, and putative RuvC knockout cleavage enzyme tested in Example 4.

[0067] Figure 3A and 3B This includes a graph showing the efficiency of reverse transcription-mediated gene editing using a reference CRISPR nuclease-reverse transcriptase fusion peptide or a CRISPR nickase variant-reverse transcriptase fusion peptide. Figure 3A : Percentage of NGS reads containing indels (black bars) and edits encoded by the template RNA (gray bars) of the reference CRISPR nuclease-reverse transcriptase fusion peptide. Figure 3B : Percentage of NGS reads containing indels (black bars) and edits encoded by the editing template RNA (gray bars) of the CRISPR nicking enzyme variant-reverse transcriptase fusion peptide.

[0068] Figure 4A and Figure 4B Includes graphs showing the gene editing efficiency of CRISPR nuclease-reverse transcriptase fusion peptides with various configurations as indicated in Tables 16 and 17. Figure 4A Gene editing efficiency of fusion peptides containing the CRISPR nuclease of SEQ ID NO: 1. The constructs correspond to those listed in Table 17, except that the nickase sequence is replaced by the reference nuclease of SEQ ID NO: 1. Figure 4B The gene editing efficiencies of the fusion peptides listed in Table 17, each of which contains the CRISPR nickase of SEQ ID NO: 32.

[0069] Figure 5 This is a graph showing the gene editing efficiency of the exemplary CRISPR nickase-reverse transcriptase fusion peptides listed in Table 17, as indicated by the percentage of eGFP-positive cells relative to mCherry-positive cells.

[0070] Figure 6 This is a graph showing the gene editing efficiency of converting BFP to eGFP using the exemplary CRISPR nickase-reverse transcriptase fusion peptides listed in Table 17, relative to the percentage of eGFP-positive cells measured for each exemplary CRISPR nickase-reverse transcriptase fusion peptide.

[0071] Figure 7 This is a graph showing the gene editing efficiency at the EMX1_T2 site of the exemplary CRISPR nickase-reverse transcriptase fusion peptides listed in Table 17 relative to the percentage of eGFP-positive cells measured for each exemplary CRISPR nickase-reverse transcriptase fusion peptide. Detailed Implementation

[0072] This disclosure provides a gene editing system involving both a CRISPR nuclease peptide and a reverse transcriptase (RT) peptide, as well as a guide RNA (which guides gene editing at a desired genomic site) and an RT donor RNA (which acts as an RNA template for the RT peptide to synthesize a DNA strand carrying desired base substitutions). In some embodiments, the gene editing system provided herein may comprise a fusion peptide or a nucleic acid encoding a fusion peptide, the fusion peptide comprising a CRISPR nuclease fragment and an RT fragment. Alternatively or additionally, the gene editing system may comprise a single RNA molecule or a nucleic acid encoding a single RNA molecule comprising a guide RNA and an RT donor RNA.

[0073] The gene-editing system described in this paper has successfully replaced nucleotides at target sites. See the examples below. Such gene-editing systems are expected to be effective in introducing desired nucleotide substitutions at genetic sites of interest, thereby achieving desired therapeutic effects (e.g., correcting genetic defects). The gene-editing system described in this paper can also be used in other fields, such as animal and plant reproduction and genome function research.

[0074] I. RT-CRISPR-mediated gene editing systems

[0075] In some aspects, this document provides an RT-CRISPR-mediated gene editing system involving at least two protein components, namely a CRISPR nuclease peptide and an RT peptide, and at least two RNA components, namely a guide RNA and an RT donor RNA. In specific embodiments, the two protein components may be localized to a fusion peptide. Alternatively or additionally, the two protein components may be localized to a single RNA molecule. In some cases, the gene editing system may comprise protein components and / or RNA components. In other cases, the gene editing system may comprise nucleic acids encoding protein components and / or nucleic acids encoding RNA components.

[0076] In some cases, RT-CRISPR-mediated gene editing systems contain CRISPR nucleases with nicking enzyme activity as disclosed herein. Such gene editing systems are expected to achieve precise gene editing at desired genomic target sites.

[0077] A. Protein components

[0078] The gene editing system provided herein involves at least two enzymes, namely CRISPR nuclease and RT. In some embodiments, the gene editing system comprises both enzymes. In a specific example, the gene editing system may comprise a fusion polypeptide comprising the two enzyme components. Alternatively, the gene editing system may comprise one or more nucleic acids encoding the two enzyme components. For example, the gene editing system may comprise one or more expression vectors (e.g., viral vectors, such as retroviral vectors, adenoviral vectors, or adeno-associated virus vectors) capable of expressing CRISPR nuclease, RT, or fusion polypeptides containing such vectors. In other examples, the gene editing system may comprise one or more mRNA molecules encoding CRISPR nuclease, RT, or fusion polypeptides containing such vectors.

[0079] In some embodiments, the CRISPR nuclease peptide and RT peptide disclosed herein can form a complex, which can be a heterodimer of the two protein components via dimerization domains (e.g., leucine zippers), an antibody, a nanobody, or an aptamer.

[0080] (i) CRISPR nuclease polypeptide

[0081] The CRISPR nuclease polypeptide used in the gene editing system disclosed herein may be a CRISPR nuclease (refer to CRISPR nuclease) or a variant thereof containing the amino acid sequence of SEQ ID NO: 1.

[0082] As used herein, the term “CRISPR nuclease” refers to an effector that is an RNA guide capable of binding to nucleic acids and introducing single-strand or double-strand breaks. CRISPR nucleases typically contain multiple functional domains, such as nuclease domains (e.g., RuvC and HNH), bridged helix (BH) domains, nucleic acid recognition (REC) domains, phosphate-locked loop (PLL) domains, wedge-shaped domains (WED), PAM interaction domains (PID), or combinations thereof. As used herein, the term “domain” refers to a distinct functional and / or structural unit of a polypeptide. In some cases, functional domains may be linear. In others, functional domains may be discontinuous and conformational. In some embodiments, a domain may contain a conserved amino acid sequence across different CRISPR nucleases.

[0083] The reference CRISPR nuclease of SEQ ID NO: 1 (see Table 1 below) is a CRISPR nuclease comprising both a RuvC nuclease domain (located at residues 1-59, 722-771, and 927-1101 of SEQ ID NO: 1) and an HNH domain (located at residues 772-926 of SEQ ID NO: 1). The RuvC and HNH nuclease domains coordinate the cleavage of the DNA strand adjacent to the 5'-NDR-3' PAM motif, where N represents any nucleotide, D represents A, G, or T, and R represents G or A. In some instances, the PAM is 5'-NRG-3' or 5'-NRR'3', where N and R are as defined herein. In one specific instance, the PAM is 5'-NGG-3', where N represents any nucleotide. Positions D10, E763, and D991 are considered active sites in the RuvC domain, and positions D844, H845, and N868 are considered active sites in the HNH domain. In addition to the nuclease domains, the reference CRISPR nuclease of SEQ ID NO: 1 also includes the BH domain (residues 60-93 of SEQ ID NO: 1), the REC domain (residues 94-721 of SEQ ID NO: 1), the PLL domain (residues 1102-1148 of SEQ ID NO: 1), the WED domain (residues 1149-1208 of SEQ ID NO: 1), and the PID domain (residues 1209-1378 of SEQ ID NO: 1).

[0084] In some embodiments, the gene editing system disclosed herein comprises a variant CRISPR nuclease polypeptide derived from SEQ ID NO: 1, for example, a variant CRISPR nuclease polypeptide comprising one or more arginine or lysine substitutions, one or more mutations in one of the nuclease domains (such as the HNH nuclease domain or the RuvC nuclease domain), or combinations thereof, as opposed to the reference CRISPR nuclease of SEQ ID NO: 1. Such protein components may form complexes with gRNA in the same gene editing system.

[0085] CRISPR nuclease peptide variants

[0086] In some embodiments, the CRISPR nuclease peptide in the gene editing system disclosed herein is a variant of the reference CRISPR nuclease of SEQ ID NO: 1, for example, by introducing one or more mutations into the reference CRISPR nuclease to modulate (e.g., enhance or reduce) one or more activities of the nuclease. As used herein, the term "variant CRISPR nuclease peptide" refers to a CRISPR nuclease peptide containing alterations (e.g., substitution, insertion, deletion, and / or fusion) at one or more residue positions compared to the reference CRISPR nuclease (SEQ ID NO: 1).

[0087] Variant CRISPR nuclease peptides may contain one or more mutations (e.g., arginine substitutions) relative to a reference CRISPR nuclease. Alternatively or additionally, variant CRISPR nuclease peptides may contain one or more mutations in the RuvC nuclease domain or the HNH nuclease domain. Such mutations may reduce or eliminate the nuclease activity of the RuvC nuclease domain or the HNH nuclease domain, thereby producing variants exhibiting nickase activity. Variant CRISPR nuclease peptides may share high sequence homology with the reference CRISPR nuclease (e.g., at least 85% sequence identity).

[0088] As used herein, the term "nicking enzyme" refers to an enzyme that cleaves one strand of double-stranded DNA at a specific nucleotide sequence (e.g., the target sequence disclosed herein). A nicking enzyme can interact with one strand of a DNA duplex to produce a DNA molecule cleaved at one strand (also known as a nicked molecule). In some embodiments, the nicking enzyme is a variant of a CRISPR nuclease containing an inactivated HNH domain. In some embodiments, the nicking enzyme is a variant of a CRISPR nuclease containing an inactivated RuvC domain. In other embodiments, the variant CRISPR nuclease peptide may contain one or more mutations relative to SEQ ID NO: 1, such mutations resulting in reduced PAM recognition strictness compared to a corresponding CRISPR nuclease peptide without such mutations. The variant CRISPR nuclease peptide may share high sequence homology (e.g., at least 85% sequence identity) with a reference CRISPR nuclease.

[0089] The variant CRISPR nuclease peptides provided herein are expected to possess advantageous characteristics relative to the reference CRISPR nuclease, such as exhibiting nicking enzyme activity and / or higher nuclease activity. Thus, the variant CRISPR nuclease peptides disclosed herein are expected to exhibit enhanced activity in gene editing relative to the reference CRISPR nuclease, for example, greater efficiency and accuracy in gene editing involving strand substitution.

[0090] In some embodiments, relative to the reference CRISPR nuclease SEQ ID NO: 1, the variant CRISPR nuclease polypeptides provided herein contain one or more mutations in the RuvC nuclease domain or the HNH nuclease domain (e.g., the HNH nuclease domain) to reduce or eliminate nuclease activity, and / or contain one or more arginine and / or lysine substitutions to improve the nuclease characteristics suitable for gene editing.

[0091] The variant CRISPR nuclease peptides provided herein are expected to exhibit one or more modulated activities (e.g., enhanced or reduced) relative to a reference CRISPR nuclease. As used herein, the term "activity" refers to biological activity. In some embodiments, activity includes enzymatic activity, such as the catalytic ability of an effector. For example, activity may include nuclease activity. In some embodiments, activity includes cleavage enzyme activity. For example, the variant CRISPR nuclease peptide may cleave substantially at only one strand of the target DNA double helix.

[0092] In some embodiments, activity includes binding activity, such as the binding of an effector (e.g., a CRISPR nuclease) to an RNA guide and / or target nucleic acid. In some instances, the variant CRISPR nuclease peptides disclosed herein exhibit enhanced binding to homologous guide RNA (gRNA) compared to a reference CRISPR nuclease, for example, binding activity that is at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 2-fold, 2-fold, 5-fold, 10-fold, or higher than the binding activity of the reference CRISPR nuclease. Homologous gRNA refers to gRNA having a scaffold that can be recognized by a CRISPR nuclease.

[0093] In some instances, the variant CRISPR nuclease peptides disclosed herein exhibit enhanced enzymatic activity relative to a reference CRISPR nuclease, for example, at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 2-fold, 2-fold, 5-fold, 10-fold, or higher than the enzymatic activity of the reference CRISPR nuclease. In other instances, the variant CRISPR nuclease peptides disclosed herein exhibit reduced enzymatic activity relative to a reference CRISPR nuclease (e.g., enzymatic activity for cleaving both strands of a target DNA duplex), for example, at least 20%, 30%, 40%, 50%, 60%, or 70% lower than the enzymatic activity of the reference CRISPR nuclease. In some cases, the reduced enzymatic activity is achieved by reducing or attenuating the nuclease activity of the RuvC domain. In other cases, the reduced enzymatic activity is achieved by reducing or attenuating the nuclease activity of the HNH domain.

[0094] In some cases, the variant CRISPR nuclease peptides disclosed herein exhibit enhanced indel activity relative to a reference CRISPR nuclease. As used herein, the term "indel activity" refers to the ability of a CRISPR nuclease to introduce an indel (insertion / deletion) into a sequence (e.g., a genomic target). For example, in some embodiments, the CRISPR nuclease will double Chain breaks are introduced into sequences (e.g., genomic targets in cells) and, through DNA repair mechanisms, generate indels.

[0095] In some embodiments, the variant CRISPR nuclease peptides provided herein share high sequence homology with a reference CRISPR nuclease. For example, the variant CRISPR nuclease peptide may contain at least 70% (e.g., at least 80%, 85%, 90%, 95%, or higher) of the same amino acid sequence as SEQ ID NO: 1. In some cases, the variant CRISPR nuclease peptide may contain at least 90% of the same amino acid sequence as SEQ ID NO: 1. In some cases, the variant CRISPR nuclease peptide may contain at least 95% of the same amino acid sequence as SEQ ID NO: 1. In other cases, the variant CRISPR nuclease peptide may contain at least 97% (e.g., 98%, 99%, 99.5%, or higher) of the same amino acid sequence as SEQ ID NO: 1.

[0096] The "percentage of identity" (also known as sequence identity) between two nucleic acid or two amino acid sequences is determined using the algorithm described by Karlin and Altschul in *Proceedings of the National Academy of Sciences of the United States of America* (Proc. Natl. Acad. Sci. USA) 87:2264-68, 1990, modified as described by Karlin and Altschul in *Proceedings of the National Academy of Sciences of the United States of America* (Proc. Natl. Acad. Sci. USA) 90:5873-77, 1993. This algorithm is incorporated into the NBLAST and XBLAST programs (version 2.0) of Altschul et al., *Journal of Molecular Biology* 215:403-10, 1990. BLAST nucleotide searches can be performed using the NBLAST program (score = 100, word length = 12) to obtain nucleotide sequences homologous to the nucleic acid molecules of this invention. BLAST protein searches can be performed using the XBLAST program (score = 50, word length = 3) to obtain amino acid sequences homologous to the protein molecules of this invention. In cases where gaps exist between two sequences, gapped BLAST can be used, as described by Altschul et al., Nucleic Acids Res. 25(17):3389-3402, 1997. When using BLAST and gapped BLAST programs, the default parameters of the respective programs (e.g., XBLAST and NBLAST) can be used.

[0097] The variant CRISPR nuclease peptides provided herein may contain one or more alterations relative to the reference CRISPR nuclease of SEQ ID NO: 1, such as substitution of one or more amino acid residues, deletion of one or more, insertion of one or more, fusion of one or more, or combinations thereof. In some cases, alterations may be introduced into the BH domain, PLL domain, WED domain, PID domain, or combinations thereof.

[0098] In some embodiments, the variant CRISPR nuclease peptides provided herein may contain one or more arginine substitutions relative to SEQ ID NO: 1. "Arginine substitution" or "lysine substitution" means replacing a non-arginine or non-lysine residue in SEQ ID NO: 1 with an arginine residue or a lysine residue. In some instances, the variant CRISPR nuclease peptide may contain up to 20 arginine and / or lysine substitutions, for example, up to 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 arginine and / or lysine substitutions. In specific instances, the variant CRISPR nuclease peptide may contain 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 arginine and / or lysine substitutions.

[0099] In some cases, one or more substituted arginine residues may be replaced by conserved amino acid residues such as lysine or histidine. In some embodiments, the variant CRISPR nuclease polypeptides provided herein may contain one or more arginine substitutions, one or more lysine substitutions, or combinations thereof.

[0100] In some cases, arginine substitutions can be located in the BH domain, PLL domain, WED domain, PID domain, or any combination thereof. In some instances, the variant CRISPR nuclease polypeptide may contain one or more arginine and / or lysine substitutions at one or more of the following positions in SEQ ID NO: 1: I79, E331, Y348, S473, F501, I581, D720, A730, G731, Q741, V752, M753, Q809, Q840, Q849, S872, S898, E982, K918, D985, Y986, Y1015, E1037, K1091, S1094, P1096, N1 099, T1104, E1105, I1106, T1108, L1117, K1131, I1147, E1179, M1205, P1208E1214, A1226, Q1230, A1236, P1238, F1241, L1281, D1284, F1285, A1292, N1295, K1298, G1329, A1333, K1344, S1348, Q1360, and I1370. In specific examples, arginine and / or lysine substitutions may be located at one or more of the following positions in SEQ ID NO: 1: I857R, K736, L784, N813, Q812, I857, and A919.

[0101] In some specific instances, the variant CRISPR nuclease polypeptide may contain one or more of the following arginine substitutions relative to SEQ ID NO: 1: I79R, E331R, Y348R, S473R, F501R, I581R, D720R, A730R, G731R, Q741R, V752R, M753R, Q809R, Q840R, Q849R, S872R, S898R, E982R, K918R, D985R, Y986R, Y1015R, E1037R, K1091R, S1094R, P1096R, N1099R, T1104R E1105R, I1106R, T1108R, L1117R, K1131R, I1147R, E1179R, M1205R, P1208, E1214R, A1226R, Q1230R, A1236R, P1238R, F1241R, L1281R, D1284R, F1285R, A1292R, N1295R, K1298R, G1329R, A1333R, K1344R, S1348R, Q1360R, and I1370R.

[0102] In some instances, the variant CRISPR nuclease polypeptide may contain one or more arginine substitutions at one or more of the positions described above. Examples include I857R, N813R, L784R, K736R, A919R, Q812R, or combinations thereof. In other instances, the variant CRISPR nuclease polypeptide may contain one or more lysine substitutions at one or more of the positions described above. Examples include I857K, N813K, L784K, A919K, Q812K, or combinations thereof.

[0103] In other examples, the variant CRISPR nuclease polypeptide may contain a combination of arginine and / or lysine substitutions at the following positions in SEQ ID NO: 1: I79, E331, Y348, S473, F501, I581, D720, A730, G731, Q741, V752, M753, Q809, Q840, Q849, S872, S898, E982, K918, D985, Y986, Y1015, E1037, K1091, S1094, P1096, N1099, T110 4. E1105, I1106, T1108, L1117, K1131, I1147, E1179, M1205, P1208, E1214, A1226, Q1230, A1236, P1238, F1241, L1281, D1284, F1285, A1292, N1295, K1298, G1329, A1333, K1344, S1348, Q1360, and / or I1370. In specific examples, the variant CRISPR nuclease polypeptide may contain combinations of arginine and / or lysine substitutions at I857R, K736, L784, N813, Q812, I857, and / or A919 of SEQ ID NO: 1 (e.g., combinations of arginine substitutions).

[0104] In specific examples, the engineered CRISPR nuclease peptide may contain arginine and / or lysine substitutions at the following positions relative to SEQ ID NO: 1: (a) I857, L784, and K736; (b) I857, A919, and K736; (c) I857, N813, and L784; (d) I857, L784, and A919; (e) I857, N813, and K736; (f) I857 and N813; (g) L784, A919, and K736; (h) I857 and L784; or (i) I857 and A919. In some cases, the engineered CRISPR nuclease peptide may contain arginine substitutions at any combination of positions in SEQ ID NO: 1. In one specific example, the engineered CRISPR nuclease peptide may contain arginine substitutions of I857R, L784R, and K736R relative to SEQ ID NO: 1. Other examples of arginine and / or lysine substitutions can be found in Table 4 below.

[0105] Alternatively or additionally, the variant CRISPR nuclease peptides provided herein may contain one or more mutations within the RuvC nuclease domain or the HNH nuclease domain to reduce or eliminate the nuclease activity of the target domain, thereby producing variants with nicking enzyme activity. Such mutations (nicking enzyme mutations) may be deletions, insertions, amino acid substitutions, or combinations thereof. In some embodiments, the mutation within the RuvC nuclease domain or the HNH nuclease domain is an amino acid substitution, wherein the substituted amino acid residue is not a conserved substitution of the native amino acid residue located at the site of the mutation. For example, if the native amino acid residue is R, the substituted residue may be any amino acid residue other than K. Similarly, if the native amino acid residue is K, the substituted residue may be any amino acid residue other than R. A set of conserved amino acid residue substitutions is provided herein.

[0106] In some cases, the one or more nickase mutations may be located within the HNH nuclease domain, for example, at D844, H845, and / or N868 of SEQ ID NO: 1. In some instances, the mutation may be an amino acid residue substitution, and the native amino acid residue of SEQ ID NO: 1 may be replaced by an amino acid residue of a different type than the native residue. For example, a positively charged residue may be replaced by an uncharged amino acid residue, or vice versa. In some instances, the amino acid residue substitution at D844 may be D844G, D844A, D844L, or D844S. In one specific instance, the mutation may be D844A. In another instance, the amino acid residue substitution at H845 may be H845G, H845A, H845I, H845L, H845M, H845V, or H845S. In one specific instance, the mutation at position H845 may be H845A. Alternatively or additionally, amino acid residue substitutions may also be made at position N868, for example, N868G, N868A, N868L, or N868S. In one instance, the mutation at position N868 is N868A.

[0107] In some cases, the one or more mutations may be located within the RuvC nuclease domain, for example, at positions D10, E763, D991, or combinations thereof in SEQ ID NO: 1 (e.g., at positions E763 and / or D991). In some instances, the mutation may be an amino acid residue substitution, and the native amino acid residue of SEQ ID NO: 1 may be substituted with an amino acid residue of a different type than the native residue. For example, a positively charged residue may be substituted with an uncharged amino acid residue, or vice versa. In some instances, the amino acid residue substitution at D10 may be D10G, D10A, D10L, or D10S. In some instances, the amino acid residue substitution at E763 may be E763G, E763A, E765L, or E763S. Alternatively or additionally, the amino acid residue substitution at D991 may be D991G, D991A, D991L, or D991S.

[0108] In some instances, the variant CRISPR nuclease peptides provided herein can be nickase variants containing one or more mutations in a nuclease domain (e.g., an HNH nuclease domain, such as at position H845, e.g., H845A). Exemplary nickase variants are provided in Table 5 below. Such nickase variants may further contain one or more arginine or lysine substitutions (e.g., arginine substitution) to enhance certain characteristics, such as indel activity. Exemplary arginine and / or lysine substitutions are provided herein, for example, at one or more of positions I857R, K736, L784, N813, Q812, I857, and A919 in SEQ ID NO: 1 (e.g., arginine substitution at positions I857, L784, and K736). In In some instances, the CRISPR nuclease polypeptide may be contained (e.g., composed of) at a position relative to SEQ ID NO: 1. A nickase mutation at H845 (e.g., H845A mutation) and an arginine and / or lysine substitution at position I857 (e.g., (Replaced by I857R).

[0109] In some instances, the variant CRISPR nuclease peptides disclosed herein exhibit enhanced double-stranded nuclease activity. Table 19 below provides examples of such CRISPR nuclease peptides, each of which is within the scope of this disclosure.

[0110] In some embodiments, the variant CRISPR nuclease polypeptide disclosed herein may contain one or more mutations relative to a reference CRISPR nuclease, said one or more mutations reducing the stringency of PAM recognition. In some cases, said one or more mutations for reducing the stringency of PAM recognition may be located at the positions of L1117, D1144, S1145, G1227, E1228, S1327, A1332, R1343, R1345 and / or T1347 of SEQ ID NO: 1. In some instances, the one or more mutations may comprise: (i) one or more arginine and / or lysine substitutions at positions L1117, G1227, S1327, A1332 and / or T1347 of SEQ ID NO: 1, optionally with arginine substitution; (ii) one or more amino acid substitutions at positions D1144, S1145, E1228, R1343 and / or R1345 of SEQ ID NO: 1; or (iii) a combination of (i) and (ii).

[0111] In some instances, this variant CRISPR nuclease polypeptide may contain mutations (e.g., amino acid residue substitutions) at the following positions of SEQ ID NO: 1: D61 (e.g., D61R or D61K), L1117 (e.g., L1117R or L1117K), D1144 (e.g., D1144V, D1144A, D1144G, or D1144S), S1145 (e.g., S1145W, S1145Y, or S1145F), G1227 (e.g., G1227K or G1227R), E1228 (e.g., E12...). 28F, E1228Y or E1228W), A1327 (e.g., A1327R or A1332K), A1332 (e.g., A1332R or A1332K), R1345 (e.g., R1345Q or R1345N), R1345 (e.g., R1345V, R1345A, R1345G or R1345S or R1345Q, R1345N), T1347 (e.g., T1347R or T1347K) or combinations thereof.

[0112] In a specific example, the one or more amino acid substitutions in (ii) may comprise D1144L, S1145W, E1228Q, R1343P, R1345V, and / or R1345Q relative to SEQ ID NO: 1. Alternatively, the amino acid residues at one or more of the substitution positions D1114, S1145, E1228, R1343, and R1345 may be conservative substitutions of L, W, Q, P, V, and Q, respectively. For example, the substitution at position D1114 may be D1114M, D1114I, or D1114V. The substitution at position S1145 may be S1145F or S1145Y. The substitution at position E1228 may be E1228N. The substitution at position R1345 may be R1345M, R1345I, R1345L, or R1345N.

[0113] Table 20 provides specific examples of such variant CRISPR nuclease peptides, each of which is within the scope of this disclosure.

[0114] In some specific instances, the engineered CRISPR nuclease polypeptide comprises (or is composed of) L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, and T1347R relative to SEQ ID NO: 1. This CRISPR nuclease polypeptide has the amino acid sequence shown in SEQ ID NO: 236.

[0115] In some instances, the engineered CRISPR nuclease peptides disclosed herein (e.g., peptides with mutations at positions L1117, D1144, G1227, E1228, A1332, R1343, R1345 and / or T1347) may further comprise one or more single arginine substitutions (e.g., single arginine substitutions) at the following positions relative to SEQ ID NO: 236: D61, A68, H494, L64, S410, T67, Q849, G1110, F501, T659, L784, Y516, G55, E1037, N57, D720, A919, A1294, Q812, N700, H65 7. T73, Q899, I857, K751, D327, I581, D462, E331, A589, D471, I699, N1295, T470, I 1147, E130, S473, A353, K40, K334, A60, S1348, K367, A1118, K31, Q349, K341, Q83, K585, Q840, G660, K527, G727, Y42, L1281, L122, Q123, T1108, E41, K1131, K30, S87 2. I1206, D1132, K460, L80, E459, K1182, M696, K918, K126, N721, Q809, K1091, K73 6. K783, N498, K723, H1119, F463, L594, D472, K744, E365, G595, K45, Y348, K964, S1181, N813, D407, S839, Y658, E586, G754, A730, Y1015, D903, A1333, S461 or H1359.

[0116] In some specific examples, the engineered CRISPR nuclease peptide comprises L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and A68R relative to SEQ ID NO: 1. In other specific examples, the engineered CRISPR nuclease peptide comprises L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and D61R relative to SEQ ID NO: 1. In still other specific examples, the engineered CRISPR nuclease peptide comprises L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and H494R relative to SEQ ID NO: 1.

[0117] In some cases, variant CRISPR nuclease peptides that exhibit less stringent recognition of PAM sequences can recognize 5'-NGN-3', 5'-NRN-3', or 5'-NYN-3' PAM sequences, where N represents any nucleotide, R represents A or G, and Y represents C or T.

[0118] In some instances, the variant CRISPR nuclease peptide may contain one or more mutations among those disclosed herein (e.g., one or more arginine and / or lysine substitutions, one or more nickase mutations, and / or one or more mutations that result in lower PAM recognition stringency). Any variant CRISPR nuclease peptide disclosed herein may share at least 90% (e.g., 95%, 97%, 98%, 99%, 99.5% or higher) sequence identity with SEQ ID NO: 1.

[0119] In some cases, in addition to mutations in the HNH or RuvC nuclease domain, arginine / lysine substitutions, and / or mutations that reduce the strictness of PAM recognition, variant CRISPR nuclease peptides may contain one or more conserved amino acid residue substitutions.

[0120] As used herein, “conservative amino acid substitution” refers to an amino acid substitution that does not alter the relative charge or size properties of the protein to which the amino acid substitution is made. Variants can be prepared according to methods for altering polypeptide sequences known to those skilled in the art, as found in references that compile such methods, such as *Molecular Cloning: A Laboratory Manual*, edited by J. Sambrook et al., 2nd edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, 1989, or *Current Protocols in Molecular Biology*, edited by FMAusubel et al., John Wiley & Sons, Inc., New York. Conservative substitutions of amino acids include substitutions between amino acids in the following groups: (a) M, I, L, V; (b) F, Y, W; (c) K, R, H; (d) A, G; (e) S, T; (f) Q, N; and (g) E, D.

[0121] Exemplary CRISPR nuclease peptides for use in the gene editing systems provided herein are disclosed in Tables 1, 5, 19 and 20 below, each of which is within the scope of this disclosure.

[0122] In some embodiments, the CRISPR nuclease polypeptide in the gene editing system disclosed herein (e.g., the reference nuclease of SEQ ID NO: 1 or any variant thereof disclosed herein) may be a fusion polypeptide comprising a CRISPR nuclease and one or more additional functional parts. As used herein, the terms “fusion” and “fused” refer to the connection of at least two nucleotides or protein molecules. For example, “fusion” and “fused” may refer to the connection of at least two polypeptide domains encoded by a single gene in nature. Fusion may be N-terminal fusion, C-terminal fusion, or intramolecular fusion. In some aspects, the domains are transcribed and translated to produce a single polypeptide.

[0123] Exemplary functional components included in fusion peptides include peptide tags, fluorescent proteins, base editing domains, DNA methylation domains, histone residue modification domains, localization factors, transcriptional modification factors, light-gated control factors, chemically induced factors, chromatin visualization factors, or combinations thereof.

[0124] In some embodiments, additional functional portions may include nuclear localization signals (NLS), nuclear output signals (NES), or combinations thereof. In some instances, the fusion polypeptide may contain an NLS, which may be located at the N-terminus or the C-terminus. In specific instances, the fusion polypeptide may contain a first NLS located at the N-terminus and a second NLS located at the C-terminus. The first NLS fragment and the second NLS fragment may be identical. Alternatively, the two NLS fragments may be different. In some embodiments, the fusion polypeptide may contain an NLS near the N-terminus and / or near the C-terminus (e.g., within about 1, 2, 3, 4, or 5 amino acids of the first or last amino acid of the CRISPR nuclease). In some embodiments, the fusion polypeptide may contain an NLS located within the flexible loop of the CRISPR nuclease.

[0125] In some embodiments, the additional functional component may be a flexible peptide linker, such as an XTEN peptide linker or a G / S-rich peptide linker. An example of such a peptide linker is provided in Example 1 below.

[0126] In some embodiments, the gene editing system provided herein may comprise a CRISPR nuclease polypeptide that can form a ribonucleoprotein (RNP) complex with homologous guide RNA. As used herein, the term "complex" refers to a combination of two or more molecules. In some embodiments, the complex comprises polypeptide and nucleic acid molecules that interact with each other (e.g., bind, contact, adhere).

[0127] In other embodiments, the gene editing system provided herein may comprise a nucleic acid encoding a CRISPR nuclease polypeptide. In some instances, the nucleotide sequence encoding the CRISPR nuclease polypeptide described herein may be codon-optimized for use in a specific host cell or organism. For example, the nucleic acid may be codon-optimized for any non-human eukaryote, including mice, rats, rabbits, dogs, livestock, or non-human primates. Codon usage tables are readily available, for example, in a “codon usage database” at kazusa.orjp / codon / , and these tables can be adapted in various ways. See Nakamura et al., Nucleic Acids Res. 28:292 (2000), which is incorporated herein by reference in its entirety. Computer algorithms for codon-optimizing specific sequences for expression in specific host cells are also available, such as GeneForge (Aptagen, Inc.; Jacobus, PA). In some instances, the nucleic acid encoding a CRISPR nuclease polypeptide as disclosed herein may be an mRNA molecule, which may be codon-optimized. Exemplary codon-optimized nucleotide sequences encoding exemplary CRISPR nuclease polypeptides can be found in Tables 1 and 5 below, any of which are within the scope of this disclosure.

[0128] In some instances, gene editing systems may contain vectors encoding CRISPR nuclease peptides (e.g., viral vectors such as AAV vectors, AdV vectors, or retroviral vectors).

[0129] (ii) Reverse transcriptase polypeptide

[0130] The gene editing system disclosed herein also includes a reverse transcriptase (RT) peptide, which may be wild-type RT or a variant thereof. In some cases, the RT peptide disclosed herein and a CRISPR nuclease peptide may form a fusion protein. As used herein, the term "reverse transcriptase" or "RT" refers to a multifunctional enzyme that typically possesses three enzymatic activities, including RNA-dependent and DNA-dependent DNA polymerization activity and RNase H activity catalyzing the cleavage of RNA in an RNA-DNA hybrid. Reverse transcriptase can generate DNA from an RNA template.

[0131] In some embodiments, the reverse transcriptase polypeptide is obtained from any naturally occurring organism or virus, or from any wild-type reverse transcriptase obtained from commercial or non-commercial sources. The reverse transcriptase polypeptide may also be a variant reverse transcriptase polypeptide.

[0132] Reverse transcriptase peptides can be obtained from many different sources. For example, the gene can be obtained from eukaryotic cells infected with retroviruses or from plasmids containing part or all of the retroviral genome. Additionally, RNA containing the reverse transcriptase gene can be obtained from retroviruses. In some embodiments, the reverse transcriptase is expressed as a separate component or otherwise provided, i.e., not as a fusion protein with the CRISPR nuclease peptides provided herein.

[0133] Those skilled in the art will recognize that reverse transcriptases are known in the art, including but not limited to Moloney murine leukosis virus (MMLV) reverse transcriptase, human immunodeficiency virus (HIV) reverse transcriptase, and avian sarcoma-leukosis virus (ASLV) reverse transcriptase, which include, but are not limited to, Rouss sarcoma virus (RSV) reverse transcriptase, avian myeloblastoma virus (AMV) reverse transcriptase, avian erythroblastovirus (AEV) helper virus MCAV reverse transcriptase, avian myelomavirus MC29 helper virus MCAV reverse transcriptase, avian reticuloendotheliosis virus (REV-T) helper virus REV-A reverse transcriptase, avian sarcoma virus UR2 helper virus UR2AV reverse transcriptase, avian sarcoma virus Y73 helper virus YAV reverse transcriptase, Rouss-associated virus (RAV) reverse transcriptase, and myeloblastosis-associated virus (MAV) reverse transcriptase, which can be used in the compositions described herein.

[0134] In some embodiments, the reverse transcriptase is MMLV-RT, MarathonRT or RTX reverse transcriptase from Eubacterium rectale, or a variant of MMLV-RT, MarathonRT or RTX reverse transcriptase. In some embodiments, the reverse transcriptase is a sequence shown in Table 10, a variant thereof, or an ortholog thereof.

[0135] Table 10. Exemplary reverse transcriptase sequences

[0136]

[0137]

[0138]

[0139]

[0140] In some embodiments, the reverse transcriptase polypeptide is a "fallacious" reverse transcriptase variant. Fallacious reverse transcriptases known and / or available in the art can be used. It should be understood that reverse transcriptases naturally do not possess any proofreading function; therefore, the error rate of reverse transcriptases is generally higher than that of DNA polymerases containing proofreading activity. In some embodiments, a reverse transcriptase is considered "fallacious" if it has an error rate of less than one error per approximately 15,000 nucleotides synthesized.

[0141] In some embodiments, the reverse transcriptase peptide has one or more mutations in the RNase H domain. In some embodiments, the reverse transcriptase peptide does not contain the RNase H domain (e.g., the RNase H domain has been removed from the reverse transcriptase peptide). In some embodiments, the RNase H domain is truncated in the reverse transcriptase peptide. In some embodiments, the reverse transcriptase peptide has one or more mutations in the RNA-dependent DNA polymerase domain. In some embodiments, the reverse transcriptase peptide is a variant with altered thermostability characteristics. The thermostability of reverse transcriptase is an important aspect of cDNA synthesis. Increased reaction temperatures help denature RNA with strong secondary structures and / or high GC content, allowing reverse transcriptase to read the entire sequence. Therefore, reverse transcription at higher temperatures can synthesize full-length cDNA and achieve higher yields. Wild-type M-MLV reverse transcriptase typically has an optimal temperature range of 37-48°C; however, mutations can be introduced that allow reverse transcription activity at higher temperatures above 48°C, including 49°C, 50°C, 51°C, 52°C, 53°C, 54°C, 55°C, 56°C, 57°C, 58°C, 59°C, 60°C, 61°C, 62°C, 63°C, 64°C, 65°C, 66°C, and higher.

[0142] The variant reverse transcriptase peptides used herein may be at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to any reference reverse transcriptase peptide, including any wild-type reverse transcriptase, mutant reverse transcriptase, or fragment of reverse transcriptase disclosed or considered herein or known in the art, or any other reverse transcriptase variant. In some embodiments, the reverse transcriptase variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, or 29 variants compared to a reference reverse transcriptase. 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or up to 100, or up to 200, or up to 300, or up to 400, or up to 500 or more amino acid variations. In some embodiments, the reverse transcriptase variant comprises a fragment of a reference reverse transcriptase such that the fragment is at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to the corresponding fragment of the reference reverse transcriptase.

[0143] Reverse transcriptases, including error-prone reverse transcriptases, thermostable reverse transcriptases, and variants with increased continuous synthetic capacity, can be engineered using various conventional strategies, including mutagenesis or evolutionary processes. In some cases, variants can be generated by introducing a single mutation. In others, variants may require more than one mutation. For mutants containing more than one mutation, the effect of a given mutation can be assessed by introducing the identified mutation into the wild-type gene through site-directed mutagenesis, isolated from other mutations carried by the specific mutant. Screening assays of the resulting individual mutants will allow for the individual determination of the effect of the mutation.

[0144] In some embodiments, the reverse transcriptase peptide includes or is fused with a domain to improve the elongation rate and / or efficiency of the reverse transcriptase. In some embodiments, the reverse transcriptase peptide is fused with an Sso7d peptide, such as the Sso7d peptide from *Sulfolobus solfataricus*. See, for example, Wang et al., *Nucleic Acid Research* 32(3):1197-207 (2004).

[0145] In some embodiments, the reverse transcriptase in any of the embodiments described herein interacts with a ligase, integrase, and / or recombinase. In some embodiments, the reverse transcriptase in any of the embodiments described herein is fused with a ligase, integrase, and / or recombinase. In some embodiments, the ligase, integrase, and / or recombinase is fused to the N-terminus or C-terminus of the reverse transcriptase. In some embodiments, the ligase, integrase, and / or recombinase is internally fused with the reverse transcriptase. In some embodiments, the integrase is a serine integrase. In some embodiments, the integrase is a Bxb1, TP901, or PhiBT1 integrase. In some embodiments, the recombinase is a serine recombinase or a tyrosine recombinase. In some embodiments, the recombinase is a CRE recombinase. In some embodiments, the reverse transcriptase interacting with or fused to the ligase, integrase, and / or recombinase further interacts with or is fused to a CRISPR nuclease polypeptide disclosed herein.

[0146] In other embodiments, the gene editing system provided herein may comprise a nucleic acid encoding an RT polypeptide. In some instances, the nucleotide sequence encoding the RT polypeptide described herein may be codon-optimized for use in a specific host cell or organism. In some instances, the nucleic acid encoding the RT polypeptide may be an mRNA molecule, which may be codon-optimized. Exemplary codon-optimized nucleotide sequences encoding exemplary RT polypeptides can be found in Table 7 below, within the scope of this disclosure. In some instances, the gene editing system may comprise a vector encoding the RT polypeptide (e.g., a viral vector, such as an AAV vector, an Adv vector, or a retroviral vector).

[0147] (iii) Fusion Peptides

[0148] In some embodiments, the gene editing system provided herein comprises a fusion polypeptide, including both a CRISPR nuclease polypeptide and an RT polypeptide as disclosed herein. Alternatively, the gene editing system may comprise a nucleic acid encoding the fusion polypeptide (e.g., a vector such as an expression vector).

[0149] As used herein, the term "fusion" refers to the connection of at least two nucleotides or protein molecules. For example, "fusion / fused" can refer to the connection of at least two polypeptide domains encoded by individual genes (e.g., the CRISPR nuclease polypeptide and reverse transcriptase polypeptide presented herein). Fusions can be N-terminal fusions, C-terminal fusions, or intramolecular fusions.

[0150] In some embodiments, the fusion polypeptide may comprise a reverse transcriptase polypeptide at its N-terminus and a CRISPR nuclease polypeptide downstream of the RT polypeptide. In other embodiments, the fusion polypeptide may comprise a CRISPR nuclease polypeptide at its N-terminus and an RT polypeptide downstream of the CRISPR nuclease polypeptide. In some embodiments, the RT polypeptide may be fused to the CRISPR nuclease polypeptide at an intramolecular location within the RT polypeptide; for example, the CRISPR nuclease polypeptide may be located within a loop of the reverse transcriptase polypeptide.

[0151] Any CRISPR nuclease peptide disclosed herein and any RT peptide disclosed herein can be used to construct fusion peptides. In some cases, the CRISPR nuclease peptide may be a CRISPR nuclease, such as SEQ ID NO: 1. Alternatively, the CRISPR nuclease peptide may be a variant of the reference CRISPR nuclease of SEQ ID NO: 1. For example, the variant may be a nickase, such as the nickase of SEQ ID NO: 32, having any of the above-described mutations in the HNH domain relative to the reference CRISPR nuclease. In other instances, the variant may contain a combination of mutations in the HNH domain as disclosed herein and one or more arginine and / or lysine substitutions.

[0152] In some embodiments, any CRISPR nuclease-RT fusion peptide disclosed herein may include one or more additional functional elements, such as those provided herein. In some cases, the additional functional elements may be one or more NLS elements. In some instances, the fusion peptide may include an NLS at the N-terminus, C-terminus, or both. Alternatively or additionally, the additional functional element may be a flexible peptide linker that can be positioned between the CRISPR nuclease peptide and the RT peptide. Suitable peptide linkers include, but are not limited to, G / S-rich peptide linkers and XTEN peptide linkers. Examples of NLS and peptide linkers are provided in Tables 1, 5, and 15 below. See also Examples 1 and 7.

[0153] In some instances, the CRISPR nuclease-RT fusion peptides provided herein contain a peptide linker (first peptide linker) located between the CRISPR nuclease peptide and the RT peptide. In some cases, the CRISPR nuclease peptide is the N-terminus of the RT peptide. In other cases, the CRISPR nuclease peptide is the C-terminus of the RT peptide. In some cases, additional peptide linkers and / or one or more NLS signals may be located between the CRISPR nuclease peptide and the RT peptide. For example, in addition to the first peptide linker, additional peptide linkers and NLS signals may be located between the CRISPR nuclease peptide and the RT peptide. In some specific instances, the length of the peptide linker between the CRISPR nuclease peptide and the RT peptide is at least 20 amino acids, for example, in the range of about 20 to 100 amino acids.

[0154] Alternatively or additionally, the CRISPR nuclease-RT fusion peptide provided herein may comprise at least two NLSs (a first NLS and a second NLS), at least one of which is located at the N-terminus or C-terminus of the fusion peptide. In some instances, one of the two NLSs is located at the N-terminus and the other at the C-terminus. In other instances, the two NLSs are located at the N-terminus. In still other instances, the two NLSs are located at the C-terminus.

[0155] In some cases, the CRISPR nuclease-RT fusion peptides provided herein may include one or more additional NLSs (e.g., a third NLS and optionally a fourth NLS). Such additional NLSs may be located between the CRISPR nuclease peptide and the RT nuclease. In other instances, optionally via a peptide linker, additional NLSs may be located between the CRISPR nuclease / RT peptide and the terminal NLS.

[0156] In some instances, the CRISPR nuclease-RT fusion peptide disclosed herein may comprise, from the N-terminus to the C-terminus, a first NLS, a CRISPR nuclease, a peptide linker, an RT peptide, and a second NLS. In other instances, the CRISPR nuclease-RT peptide disclosed herein may comprise, from the N-terminus to the C-terminus, a first NLS, an RT nuclease, a peptide linker, a CRISPR nuclease peptide, and a second NLS (which may be identical to the first NLS), a CRISPR nuclease peptide, a first peptide linker, a second NLS (which may be identical to the first NLS), a second peptide linker, an RT nuclease, a third peptide linker (which may be identical to the first peptide linker), a third NLS (which may be identical to the first NLS), a fourth peptide linker, and a fourth NLS. Table 8 below provides examples of CRISPR nuclease-RT peptides, all of which are within the scope of this disclosure.

[0157] In other examples, the CRISPR nuclease-RT fusion peptides disclosed herein may have the configurations shown in Table 16 below (from N-terminus to C-terminus). In some cases, the CRISPR nuclease-RT fusion peptides do not include the FLAG motif in any of the configurations listed in Table 16. In one case, the CRISPR nuclease-RT fusion peptide has a configuration corresponding to configuration 4, except that the FLAG motif is removed. In another case, the CRISPR nuclease-RT fusion peptide has a configuration corresponding to configuration 5, except that the FLAG motif is removed. In yet another case, the CRISPR nuclease-RT fusion peptide has a configuration corresponding to configuration 7, except that the FLAG motif is removed. In yet another case, the CRISPR nuclease-RT fusion peptide has a configuration corresponding to configuration 9, except that the FLAG motif is removed. Alternatively, the CRISPR nuclease-RT fusion peptide has a configuration corresponding to configuration 10, except that the FLAG motif is removed. Further exemplary CRISPR nuclease-RT fusion peptides are provided in Table 17. Variants of these fusion peptides in which the FLAG motif has been removed are also within the scope of this disclosure.

[0158] In some instances, any CRISPR nuclease-RT fusion polypeptide may contain a CRISPR nuclease variant. In some specific instances, the CRISPR variant is a nickase, such as those provided herein (e.g., containing mutations as listed in Table 5). In some cases, the nickase contains one or more mutations in the HNH domain relative to the reference nuclease of SEQ ID NO: 1, for example, at position H845, such as H845A.

[0159] In other embodiments, the gene editing system provided herein may comprise a nucleic acid encoding a CRISPR nuclease-RT fusion polypeptide. In some instances, the nucleotide sequence encoding the fusion polypeptide described herein may be codon-optimized for use in a specific host cell or organism. In some instances, the nucleic acid encoding the fusion polypeptide disclosed herein may be an mRNA molecule, which may be codon-optimized. Exemplary codon-optimized nucleotide sequences encoding exemplary CRISPR nuclease-RT fusion polypeptides can be found in Tables 8 and 17 below, any of which are within the scope of this disclosure. See also Examples 5 and 7. The coding sequences of variants of the exemplary CRISPR nuclease-RT fusion polypeptide in which the FLAG motif is removed are also within the scope of this disclosure.

[0160] In some instances, gene editing systems may contain vectors encoding CRISPR nuclease-RT fusion peptides (e.g., viral vectors such as AAV vectors, AdV vectors, or retroviral vectors).

[0161] (iv) Preparation of protein components

[0162] The CRISPR nuclease peptides, RT peptides, or CRISPR nuclease-RT fusion peptides disclosed herein can be prepared using conventional methods or the methods disclosed herein. For example, CRISPR nuclease peptides, RT peptides, or CRISPR nuclease-RT fusion peptides can be prepared by culturing host cells (such as bacterial or mammalian cells) capable of producing nuclease peptides, isolating the nuclease peptides thus produced, and optionally purifying the nuclease peptides. CRISPR nuclease peptides, RT peptides, or CRISPR nuclease-RT fusion peptides can also be prepared using an in vitro coupled transcription-translation system.

[0163] There are no particular restrictions on the host cells that can be used to prepare CRISPR nuclease peptides, RT peptides, or CRISPR nuclease-RT fusion peptides, as long as they can produce CRISPR nuclease peptides, RT peptides, or CRISPR nuclease-RT fusion peptides. Some non-limiting examples of host cells include bacterial cells (e.g., E. coli cells), yeast cells, insect cells, or mammalian cells.

[0164] carrier

[0165] This disclosure provides vectors for expressing CRISPR nuclease peptides, RT peptides, or CRISPR nuclease-RT fusion peptides. In some embodiments, the vectors disclosed herein comprise a nucleotide sequence encoding a CRISPR nuclease peptide, RT peptide, or CRISPR nuclease-RT fusion peptide. In some embodiments, the vectors contain a Pol II promoter or a Pol III promoter.

[0166] Expression of natural or synthetic polynucleotides is typically achieved by operatively linking a polynucleotide encoding a CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide to a promoter and incorporating the construct into an expression vector. There are no specific limitations on the expression vector, as long as it includes a polynucleotide encoding the CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide of the present invention and is suitable for replication and integration in eukaryotic cells.

[0167] Typical expression vectors include transcription and translation terminators, initiation sequences, and promoters for the desired polynucleotide expression. For example, plasmid vectors carrying RNA polymerase recognition sequences (pSP64, pBluescript, etc.) can be used. Vectors derived from retroviruses, such as lentiviruses, are suitable tools for achieving long-term gene transfer because they allow for the long-term stable integration of transgenes and their propagation in daughter cells. Examples of vectors include expression vectors, replication vectors, probe-generating vectors, and sequencing vectors. Expression vectors can be delivered to cells in the form of viral vectors.

[0168] Viral vector technology is well-known in the field and is described in various virology and molecular biology handbooks. Viruses that can be used as vectors include, but are not limited to, bacteriophages, retroviruses, adenoviruses, adeno-associated viruses, herpesviruses, and lentiviruses. Typically, a suitable vector contains an origin of replication that is functional in at least one organism, a promoter sequence, a convenient restriction endonuclease site, and one or more selectable markers.

[0169] There are no specific restrictions on the type of vector, and vectors that can be expressed in host cells can be appropriately selected. More specifically, depending on the type of host cell, a promoter sequence that ensures the expression of the polypeptide from the polynucleotide is appropriately selected, and this promoter sequence and polynucleotide are inserted into any different plasmids, etc., for the preparation of the expression vector.

[0170] Additional promoter elements (e.g., enhancer sequences) regulate the frequency of transcription initiation. These are typically located in a region 30–110 bp upstream of the start site, but many promoters have recently been shown to also contain functional elements downstream of the start site. Depending on the promoter, these elements appear to act synergistically or independently to activate transcription.

[0171] Furthermore, this disclosure is not limited to the use of constitutive promoters. Inducible promoters are also contemplated as part of this disclosure. The use of inducible promoters provides a molecular switch capable of turning on the expression of a polynucleotide sequence operatively linked to it when such expression is desired, or turning off said expression when it is not desired. Examples of inducible promoters include, but are not limited to, metallothionein promoters, glucocorticoid promoters, progesterone promoters, and tetracycline promoters.

[0172] The expression vector to be introduced may also contain either or both of an optional marker gene or a reporter gene to facilitate the identification and selection of expressing cells from a cell population seeking transfection or infection via a viral vector. In other respects, the optional marker may be carried on a separate DNA fragment and used in co-transfection procedures. Both the optional marker and the reporter gene may be side-linked with appropriate transcriptional control sequences to enable expression in host cells. Examples of such markers include dihydrofolate reductase genes and neomycin resistance genes for eukaryotic cell culture, and tetracycline resistance genes and ampicillin resistance genes for culturing *E. coli* and other bacteria. By using such selection markers, it can be confirmed whether the polynucleotide encoding the polypeptide of the present invention has been transferred to the host cell and then successfully expressed.

[0173] The methods for preparing recombinant expression vectors are not particularly limited, and examples include methods using plasmids, phages, or granules.

[0174] Methods of expression

[0175] This disclosure includes a method for protein expression, the method comprising translating the CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide described herein.

[0176] In some embodiments, the host cells described herein are used to express CRISPR nuclease peptides, RT peptides, or CRISPR nuclease-RT fusion peptides. The host cells are not particularly limited, and a variety of known cells may preferably be used. Specific examples of host cells include bacteria such as *Escherichia coli*, yeasts (budding yeasts, *Saccharomyces cerevisiae*, and *Schizosaccharomyces pombe*), nematodes (*Caenorhabditis elegans*), Xenopus laevis oocytes, and animal cells (e.g., CHO cells, COS cells, and HEK293 cells). The methods used to transfer the expression vectors described above into host cells, i.e., the transformation methods, are not particularly limited, and known methods such as electroporation, calcium phosphate methods, liposome methods, and DEAE dextran methods may be used.

[0177] After transforming the host cell with the expression vector, the host cell can be cultured, propagated, or multiplied to produce a CRISPR nuclease peptide, an RT peptide, or a CRISPR nuclease-RT fusion peptide. Following expression, the host cells can be collected and the CRISPR nuclease peptide, RT peptide, or CRISPR nuclease-RT fusion peptide can be purified from the culture using conventional methods (e.g., filtration, centrifugation, cell disruption, gel filtration chromatography, ion exchange chromatography, etc.).

[0178] Various methods can be used to determine the production levels of mature CRISPR nuclease peptides, mature RT peptides, or mature CRISPR nuclease-RT fusion peptides in host cells. Such methods include, but are not limited to, methods utilizing, for example, polyclonal or monoclonal antibodies that are specific to proteins or to labeling tags as described elsewhere herein. Exemplary methods include, but are not limited to, enzyme-linked immunosorbent assay (ELISA), radioimmunoassay (RIA), fluorescence immunoassay (FIA), and fluorescence activated cell sorting (FACS). These and other assays are well known in the art (see, for example, Maddox et al., *Journal of Experimental Medicine* 158:1211

[1983] ).

[0179] This disclosure provides an in vivo expression method for a CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide (and optionally gRNA and / or RT donor RNA in the gene editing system disclosed herein). This method may comprise providing a polynucleotide encoding a CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide to host cells in a subject (e.g., a human subject), wherein the polynucleotide encodes a CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide, and expressing the CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide from the cells.

[0180] B. RNA components

[0181] The gene editing system provided herein also involves at least two RNA components: a guide RNA (gRNA) that guides gene editing at the desired genetic site, and a RT donor RNA that serves as an RNA template for the RT polypeptide during reverse transcription. The RT donor RNA contains the desired nucleotide substitution to be inserted into the genetic site of interest. In some embodiments, the gene editing system comprises said two RNA molecules. In a specific example, the gene editing system may comprise a single RNA molecule comprising both gRNA and RT donor RNA. Alternatively, the gene editing system may comprise one or more nucleic acids encoding said two RNA components. For example, the gene editing system may comprise one or more expression vectors (e.g., viral vectors, such as retroviral vectors, adenoviral vectors, or adeno-associated virus vectors) capable of producing gRNA, RT donor RNA, or single RNA molecules containing such vectors.

[0182] In some embodiments, gRNA and RT donor RNA, as disclosed herein, can form a complex.

[0183] (i) Guide RNA

[0184] The gene editing systems disclosed herein further comprise one or more gRNAs or nucleic acids encoding such RNA. As used herein, the terms “RNA guide,” “RNA guide sequence,” or “guide RNA (gRNA)” refer to an RNA molecule or modified RNA molecule that facilitates the targeting of a genomic site of interest by the CRISPR nuclease described herein. For example, an RNA guide may be a molecule comprising a spacer sequence and a scaffold sequence. The spacer sequence recognizes (e.g., binds to) a site in a non-PAM strand complementary to a target sequence in the PAM strand, for example, and is designed to be complementary to a specific nucleic acid sequence. The scaffold sequence contains a nuclease-binding sequence for binding to a CRISPR nuclease. In some embodiments, the scaffold is an RNA sequence.

[0185] In some cases, the gRNA disclosed herein may further include adapter sequences, 5' end protection fragments and / or 3' end protection fragments, or combinations thereof.

[0186] Interval subsequence

[0187] As used herein, the terms "spacer" and "spacer sequence" (also known as DNA-binding sequence) are part of an RNA guide for the RNA equivalent of a target sequence (DNA sequence). A spacer contains a sequence capable of binding to a non-PAM strand through base pairing at a site complementary to the target sequence (which is in the PAM strand). Such spacers are also referred to as being specific to the target sequence. In some cases, the spacer may be at least 75% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99%) identical to the target sequence, except for RNA-DNA sequence differences. In some cases, the spacer may be 100% identical to the target sequence, except for RNA-DNA sequence differences.

[0188] The gene editing systems disclosed herein comprise one or more gRNAs, each containing a spacer and a scaffold for targeting a genomic site of interest, the scaffold being recognizable by a variant CRISPR nuclease peptide contained in the gene editing system. The target sequence may be adjacent to a 5'-NDR-3' protospacer adjacent motif (PAM), where N represents any nucleotide, D represents A, G, or T, and R represents A or G. In some instances, the PAM is 5'-NRR-3', where N represents any nucleotide and R represents A or G; for example, the PAM may be 5'-NRG-3', where N represents any nucleotide and R represents A or G. In specific instances, the PAM motif is 5'-NGG-3', where N represents any nucleotide. In some instances, the PAM sequence recognizable by the CRISPR nuclease peptide may be 5'-NGN-3', 5'-NRN-3', or 5'-NYN-3', where N represents any nucleotide, R represents A or G, and Y represents C or T.

[0189] The PAM motif is located at the 3' end of the target sequence. As used herein, the terms "protospacer adjacent motif" or "PAM sequence" refer to the DNA sequence adjacent to the target sequence. In some embodiments, the PAM sequence is essential for binding CRISPR nucleases and / or indel activity. In a double-stranded DNA molecule, the strand containing the PAM motif is referred to as the "PAM strand," and the complementary strand is referred to as the "non-PAM strand." The gRNA binds to a site in the non-PAM strand complementary to the target sequence disclosed herein, and the PAM sequence as described herein is present in the PAM strand. The PAM motif may be located upstream of the target sequence.

[0190] As used herein, the term "adjacent to" refers to a nucleotide or amino acid sequence that is very close to another nucleotide or amino acid sequence. In some embodiments, if no nucleotide separates the two sequences, the nucleotide sequence is adjacent to the other nucleotide sequence (i.e., directly adjacent). In some embodiments, if a small number of nucleotides separate the two sequences (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides), the nucleotide sequence is adjacent to the other nucleotide sequence.

[0191] The length of the spacer sequence disclosed herein can be from about 15 nucleotides to about 30 nucleotides. For example, the length of the spacer sequence can be from about 15 nucleotides to about 20 nucleotides, from about 15 nucleotides to about 25 nucleotides, from about 20 nucleotides to about 25 nucleotides, or from about 20 nucleotides to about 30 nucleotides. In some embodiments, the spacer in the gRNA can typically be designed to be between 15 and 25 nucleotides in length (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, and 25 nucleotides) and complementary to the specific target sequence. In some embodiments, the spacer sequence can be designed to be between 18 and 22 nucleotides in length (e.g., 20 nucleotides).

[0192] In some embodiments, the spacer sequence may have at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.5% sequence identity with the target sequence as described herein, and may be able to bind to the complementary region of the target sequence via base pairing.

[0193] In some embodiments, the spacer sequence contains only RNA bases. In some embodiments, the spacer sequence contains DNA bases (e.g., the spacer contains at least one thymine). In some embodiments, the spacer sequence contains both RNA and DNA bases (e.g., the DNA-binding sequence contains at least one thymine and at least one uracil).

[0194] stent sequence

[0195] The scaffold sequence in the gRNA can also be recognized by variant CRISPR nuclease peptides in gene editing systems. In some cases, the scaffold sequence contains SEQ ID NO: 2, which is a homologous scaffold of the reference CRISPR nuclease of SEQ ID NO: 1.

[0196] GUUUUAGAGCUGUGCUGAAAAGCACAGCACGUUAAAAUAAGGCAGUGAUUGAAAAAUCCAGUCCGUAUUCAGCUUGAAAAAGUGAGCACCGAAUCGGUGCUU (SEQ ID NO: 2)

[0197] In other cases, the scaffold sequence may be a variant derived from SEQ ID NO: 2. Such variant scaffold sequences may contain at least 80% (e.g., at least 85%, 90%, 95%, 98%, or higher) the same nucleotide sequence as SEQ ID NO: 2. Alternatively or additionally, variant scaffold sequences may contain deletions, nucleotide substitutions, or combinations thereof. The binding of the variant CRISPR nuclease polypeptide to the variant scaffold sequence may be increased compared to the scaffold of SEQ ID NO: 2. In some instances, the variant scaffold may be a fragment of SEQ ID NO: 2 as disclosed herein, or a variant thereof. For example, the length of the variant scaffold used for the gRNA provided herein may be in the range of 100-150 nucleotides.

[0198] In gRNA, the scaffold can be positioned at the 3' end of the spacer. In some cases, the scaffold and spacer are directly connected. In others, they are connected via nucleotide linkers.

[0199] (ii) RT donor RNA

[0200] As used herein, the term "reverse transcription donor RNA" or "RT donor RNA" refers to an RNA molecule that contains a reverse transcription template sequence (RTT sequence) and a primer binding site (PBS). RT donor RNA can be fused to the RNA guide at the 5' or 3' end.

[0201] Any RT donor RNA disclosed herein contains: (i) a primer binding site (PBS); and (ii) an RTT sequence. In some cases, the RT donor RNA may further contain: (iii) a nucleotide linker sequence, (iv) a 5' and / or 3' protective fragment (see disclosure herein), or a combination thereof. In some instances, the 5' or 3' protective fragment (e.g., a 3' extension) may contain a pseudoknot motif to protect it from 3' exonuclease activity.

[0202] In some embodiments, the RT donor RNA comprises an aptamer. In some embodiments, the aptamer recruits a reverse transcriptase polypeptide.

[0203] Primer binding site (PBS)

[0204] In some embodiments, the PBS in the RT donor RNA disclosed herein is an RNA sequence capable of binding to a DNA strand via base pairing. The DNA strand has been, or can be, cleaved or cut by a CRISPR nuclease peptide of the gene editing system disclosed herein. In some embodiments, the PBS contains an RNA sequence capable of binding to a DNA strand (the site targeting the PBS) via base pairing. The DNA strand may have a free 3' end or a 3' free end that can be generated by cleavage by a CRISPR nuclease peptide contained in the same gene editing system. In some instances, the site targeting the PBS may be located on the same DNA strand as the PAM sequence (the PAM strand).

[0205] In some embodiments, the length of the PBS can be about 5-50 nucleotides. For example, the length of the PBS can be about 5-40, 5-30, or 5-20 nucleotides. In specific instances, the length of the PBS can be about 5-20 nucleotides (e.g., 7-17 nucleotides). In some instances, the PBS can contain 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides.

[0206] As used herein, the term "PBS-targeting site" refers to a PBS-binding region. The PBS-targeting site may be adjacent to (e.g., upstream of) the PAM. In gene editing systems comprising a CRISPR nuclease polypeptide as a nickase variant (e.g., comprising a disrupted HNH nuclease domain as disclosed herein), PBS in the RT donor RNA may bind to a region on the PAM strand (the PBS-targeting site). In some embodiments, the PBS-targeting site may partially or completely overlap with the target sequence. In some cases, the PBS-targeting site may be located upstream of the PAM sequence. For example, the PBS-targeting site may be located up to 100 nucleotides upstream of the PAM sequence, such as up to 50, 30, 25, 20, 15, 10, or 5 nucleotides upstream of the PAM sequence. In specific examples, the PBS targeting site can begin approximately 3 to 10 nucleotides upstream of the PAM sequence (i.e., the longest 5' nucleotide of PBS can bind to approximately 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides upstream of PAM). In specific examples, the PBS targeting site can begin 1, 1-2, 1-3, 1-4, or 1-5 nucleotides upstream of the PAM sequence. When the CRISPR nuclease peptide in the gene editing system generates a free 3' end within or near the target sequence, the binding of PBS to the PAM chain at a site upstream of the PAM sequence, starting from the free 3' end generated from the non-PAM chain, can efficiently promote DNA synthesis via RT peptides in the gene editing system.

[0207] Reverse transcription template (RTT) sequence

[0208] The reverse transcription template sequence (RTT sequence) serves as a template for RT-peptide-mediated reverse transcription in the gene editing systems disclosed herein. In some embodiments, the RTT sequence comprises a sequence having at least one encoded edit. In some embodiments, the RTT sequence comprises a sequence homology to a target sequence having at least one encoded edit or its complementary region. In some embodiments, the length of the RTT sequence is at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides. In some embodiments, the length of the RTT sequence is about 10 nucleotides, 20 nucleotides, 30 nucleotides, 40 nucleotides, 50 nucleotides, 60 nucleotides, 70 nucleotides, 80 nucleotides, 90 nucleotides, 100 nucleotides, 110 nucleotides, or 120 nucleotides or any length therein.

[0209] In some embodiments, the RTT sequence is about 10 nucleotides long. In some embodiments, the RTT sequence is about 11 nucleotides long. In some embodiments, the RTT sequence is about 12 nucleotides long. In some embodiments, the RTT sequence is about 13 nucleotides long. In some embodiments, the RTT sequence is about 14 nucleotides long. In some embodiments, the RTT sequence is about 15 nucleotides long. In some embodiments, the RTT sequence is about 16 nucleotides long. In some embodiments, the RTT sequence is about 17 nucleotides long. In some embodiments, the RTT sequence is about 18 nucleotides long. In some embodiments, the RTT sequence is about 19 nucleotides long. In some embodiments, the RTT sequence is about 20 nucleotides long. In some embodiments, the RTT sequence is about 21 nucleotides long. In some embodiments, the RTT sequence is about 22 nucleotides long. In some embodiments, the RTT sequence is about 23 nucleotides long. In some embodiments, the RTT sequence is about 24 nucleotides long. In some embodiments, the RTT sequence is about 25 nucleotides long. In some embodiments, the RTT sequence is about 26 nucleotides long. In some embodiments, the RTT sequence is about 27 nucleotides long. In some embodiments, the RTT sequence is about 28 nucleotides long. In some embodiments, the RTT sequence is about 29 nucleotides long. In some embodiments, the RTT sequence is about 30 nucleotides long.

[0210] In some embodiments, the reverse transcription template sequence contains at least one (e.g., at least two) encoded edits relative to the target sequence. In some embodiments, the at least one encoded edit contains at least one substitution, insertion, and / or deletion. In some embodiments, the edit in the target sequence contains a substitution, insertion, and / or deletion of a sequence relative to the target sequence. In some embodiments, the reverse transcription template sequence contains at least one LoxP site.

[0211] In some embodiments, the editing can be a mononucleotide or polynucleotide substitution, such as G to T substitution, G to A substitution, G to C substitution, T to G substitution, T to A substitution, T to C substitution, C to G substitution, C to T substitution, C to A substitution, A to T substitution, A to G substitution, or A to C substitution. In some embodiments, sequence changes can convert G:C base pairs to T:A base pairs, G:C base pairs to A:T base pairs, G:C base pairs to C:G base pairs, T:A base pairs to G:C base pairs, T:A base pairs to A:T base pairs, T:A base pairs to C:G base pairs, C:G base pairs to G:C base pairs, C:G base pairs to T:A base pairs, C:G base pairs to A:T base pairs, A:T base pairs to T:A base pairs, A:T base pairs to G:C base pairs, or A:T base pairs to C:G base pairs.

[0212] In some embodiments, the template sequence described herein may further introduce one or more silent mutations. As used herein, a silent mutation is a mutation that does not alter the amino acid residue encoded by the codon containing the mutation. The RTT sequence can be transcribed into DNA using the reverse transcriptase of the gene editing system described herein. In some embodiments, the RTT sequence is transcribed from the 5' to the 3' end into the DNA of the PAM strand.

[0213] In some embodiments, the RTT sequence is the 5' of PBS. In some embodiments, the RTT sequence is the 3' of PBS. In some cases, the PBS and RTT sequences in the RT donor RNA provided herein can be ligated using an adapter sequence. In some implementation In this example, the RTT and the end protection segment (e.g., the 3' end protection segment) can be connected via a connector sequence to avoid the two... Spatial hindrance between RNA components.

[0214] (iii) single RNA molecule

[0215] In some embodiments, the gene editing system provided herein comprises a single RNA molecule or a nucleic acid encoding a single RNA molecule, said single RNA molecule comprising both gRNA and RT donor RNA. This single RNA molecule is capable of mediating the cleavage of a CRISPR nuclease peptide at a target sequence within a genomic site of interest and synthesizing DNA fragments from the free 3' end of the free DNA strand generated by cleavage with the CRISPR nuclease peptide based on the RTT sequence in the single RNA molecule.

[0216] In some embodiments, a single RNA molecule may include an RNA guide that is linked (optionally via a adapter) to an RT donor RNA. In some instances, a single RNA molecule includes a spacer sequence, a scaffold sequence recognizable by a CRISPR nuclease peptide, an RTT sequence, and PBS from the 5' to the 3' end. In specific instances, a single RNA molecule may include a spacer sequence, a scaffold sequence, an RTT sequence, PBS, and a protective fragment from the 5' to the 3' end.

[0217] Any single RNA molecule presented herein may further include a adapter, which may be located between the scaffold sequence and the RTT or after PBS. In some instances, the adapter may include a hairpin structure. In some instances, the connector can It includes adaptive structural domains.

[0218] In some instances, the 5' and / or 3' ends of a single RNA molecule, or gRNA and / or RT donor RNA, may contain protective fragments that enhance the RNA molecule's resistance to exonuclease activity. In some cases, the end-protected fragment may contain nucleotide sequences capable of forming secondary structures such as hairpins, circularizations, pseudoknots, or triplet structures. In other cases, the end-protected fragment may contain sequences of exonuclease-resistant RNA (xrRNA), transfer RNA (tRNA), or truncated tRNA. In some embodiments, the modification is a Zika-like pseudoknot, a murine leukemia virus pseudoknot (MLV-PK) sequence, a red clover necrosis mosaic virus (RCNMV) sequence, an osmanthus necrosis mosaic virus (SCNMV) sequence, a carnation ringspot virus (CRSV) sequence, a preQ1 aptamer sequence, etc. Truncated preQ1 aptamer sequence, boxB RNA sequence RNA sequence or RNA phage MS2 sequence.

[0219] In some instances, the 5' end of a single RNA molecule (or, when using a separate RNA molecule, gRNA and / or RT) is supplied. The 5' end of the somatic RNA may contain a 5' extension motif, which may be any protected fragment disclosed herein. In some instances, the 3' end of a single RNA molecule (or, when using a single RNA molecule, the 3' end of gRNA and / or RT donor RNA) may contain a 3' extension motif, which may be any protected fragment disclosed herein. A specific example of a 3' extension motif is provided in Example 6 below.

[0220] (iv) Nucleic acid modification

[0221] Any RNA component in the gene editing systems disclosed herein, such as a single RNA molecule, gRNA, and / or RT donor RNA, may include one or more modifications. Exemplary modifications may include any modifications to sugars, nucleobases, nucleoside internucleotide bonds (e.g., to phosphate-linked / phosphodiester bonds / phosphodiester backbones), and any combinations thereof. Some exemplary modifications provided herein are described in detail below.

[0222] Any RNA component disclosed herein may include any available modifications such as to sugars, nucleobases, or nucleoside bonds (e.g., to phosphate linkages, phosphodiester bonds, or the phosphodiester backbone). One or more atoms of the pyrimidine nucleobase may be replaced or substituted with an optionally substituted amino group, an optionally substituted thiol, an optionally substituted alkyl group (e.g., methyl or ethyl), or a halogen (e.g., chlorine or fluorine). One or more atoms of a purine nucleobase can be optionally substituted with an amino group or an optionally substituted sulfur group. Alcohols, optionally substituted alkyl groups (e.g., methyl or ethyl) or halogens (e.g., chlorine or fluorine) may be used as substitutes or replacements. In some embodiments, the modification (e.g., one or more modifications) is present in each of the sugar and nucleoside bonds. The modification can be a modification of ribonucleic acid (RNA) to deoxyribonucleic acid (DNA), threonucleic acid (TNA), glycolic acid (GNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), or a hybrid thereof. Other modifications are described herein.

[0223] In some embodiments, any RNA component of the gene editing system disclosed herein contains debases. Site (i.e., a site that does not contain purines or pyrimidines). In some embodiments, a base-degrading site (also referred to as purine-free / pyrimidine-free) Pyridine sites exist in the template RNA for editing. For example, a base-degrading sites can be present in the RTT of the template RNA for editing. In some... In the examples, the activity of reverse transcriptase ceases at or near the abase site.

[0224] In some embodiments, modifications may include chemical modifications or cell-induced modifications. For example, Lewis and Pan describe some non-limiting examples of intracellular RNA modifications in “RNA modifications and structures cooperateto guide RNA-protein interactions” in Nature Reviews Molecular Cell Biology, 2017, 18:202-210.

[0225] Different sugar modifications, nucleotide modifications, and / or nucleotide inter-bonds (e.g., backbone structure) can be present at different positions in the sequence. Those skilled in the art will understand that nucleotide analogs or other modifications can be located anywhere in the sequence without substantially reducing the function of the sequence. The sequence may include about 1% to about 100% of modified nucleotides (relative to the total nucleotide content, or relative to one or more types of nucleotides, i.e., any one or more of A, G, U, or C) or any intermediate percentage (e.g., 1% to 20%, 1% to 25%, 1% to 50%, 1% to 60%, 1% to 70%, 1% to 80%, 1% to 90%, 1% to 95%, 10% to 20%, 10% to 25%, 10% to 50%, 10% to 60%, 10% to 70%, 10% to 80%, 10% to 90%, 10% to 95%, 10% to ... 100%, 20% to 25%, 20% to 50%, 20% to 60%, 20% to 70%, 20% to 80%, 20% to 90%, 20% to 95%, 20% to 100%, 50% to 60%, 50% to 70%, 50% to 80%, 50% to 90%, 50% to 95%, 50% to 100%, 70% to 80%, 70% to 90%, 70% to 95%, 70% to 100%, 80% to 90%, 80% to 95%, 80% to 100%, 90% to 95%, 90% to 100%, and 95% to 100%.

[0226] In some embodiments, sugar modifications (e.g., at the 2' or 4' position) or sugar substitution at one or more ribonucleotides in the sequence, as well as backbone modifications, may include modifications or substitutions of phosphodiester bonds. Specific examples of sequences include, but are not limited to, sequences comprising a modified backbone or sequences without natural internucleotide bonds, such as internucleotide modifications including modifications or substitutions of phosphodiester bonds. Furthermore, sequences having a modified backbone include those without a phosphorus atom in the backbone. For the purposes of this application, and as sometimes referred to in the art, modified RNA without a phosphorus atom in its internucleotide backbone may also be considered an oligonucleotide. In certain embodiments, the sequence will comprise ribonucleotides having a phosphorus atom in their internucleotide backbone.

[0227] The modified sequence backbone may include, for example, thiophosphates, chiral thiophosphates, dithiophosphates, phosphate triesters, aminoalkyl phosphate triesters, methyl and other alkylphosphonates, such as 3'-alkylene phosphonates and chiral phosphonates, hypophosphonates, phosphoramidites, such as 3'-aminophosphatidyl esters and aminoalkylphosphamide esters, thiocarbonylphosphatidyl esters, thiocarbonylalkylphosphonates, thiocarbonylalkyl phosphate triesters, and borane phosphates having normal 3'-5' bonds, 2'-5' linked analogs of these esters, and those esters having reverse polarity, wherein adjacent pairs of nucleoside units are linked in 3'-5' to 5'-3' or 2'-5' to 5'-2' configurations. Various salts, mixed salts, and free acid forms are also included. In some embodiments, the sequence may be negatively or positively charged.

[0228] Modified nucleotides that can be incorporated into the sequence can be modified at nucleoside internucleotides (e.g., phosphate backbones). In this document, the phrases “phosphate ester” and “phosphate diester” are used interchangeably in the context of a polynucleotide backbone. The backbone phosphate ester group can be modified by replacing one or more oxygen atoms with different substituents. Further, modified nucleosides and nucleotides can include large-scale substitution of the unmodified phosphate ester portion with another nucleoside internucleotide as described herein. Examples of modified phosphate ester groups include, but are not limited to, thiophosphates, selenophosphates, boranophosphates, boranophosphate esters, hydrogen phosphates, phosphoramide esters, phosphate diamide esters, alkyl or aryl phosphonates, and triphosphate esters. In dithiophosphates, both non-linked oxygen atoms are replaced by sulfur. The phosphate ester linker can also be modified by replacing the linking oxygen with nitrogen (bridged phosphoramide ester), sulfur (bridged thiophosphate), and carbon (bridged methylene phosphonate).

[0229] The α-thiosubstituted phosphate moiety imparts stability to RNA and DNA polymers via non-natural thiophosphate backbone bonds. Thiophosphate DNA and RNA exhibit increased nuclease resistance and thus a longer half-life in the cellular environment.

[0230] In specific embodiments, the modified nucleosides include α-thio-nucleosides (e.g., 5'-O-(1-thiophosphate)-adenosine, 5'-O-(1-thiophosphate)-cytidine (α-thiocytidine), 5'-O-(1-thiophosphate)-guanosine, 5'-O-(1-thiophosphate)-uridine, or 5'-O-(1-thiophosphate)-pseuuridine).

[0231] This document describes other nucleoside bonds that can be used according to the present invention, including nucleoside bonds that do not contain phosphorus atoms.

[0232] In some embodiments, the sequence may include one or more cytotoxic nucleosides. For example, cytotoxic nucleosides may be incorporated into the sequence, such as through bifunctional modifications. Cytotoxic nucleosides may include, but are not limited to, adenosine arabinoside, 5-azacytidine, 4'-thio-cytarabine, cyclopentenylcytosine, cladribine, clofarabine, cytarabine, cytosine arabinoside, 1-(2-C-cyano-2-deoxy-β-D-arabinose-furanopentosyl)-cytosine, decitabine, 5-fluorouracil, and fludarabine. Examples of cytarabine include fluxuridine, gemcitabine, tegafur, and combinations of uracil, tegafur ((RS)-5-fluoro-1-(tetrahydrofuran-2-yl)pyrimidin-2,4(1H,3H)-dione), troxacitabine, tezacitabine, 2'-deoxy-2'-methylenecytidine (DMDC), and 6-mercaptopurine. Other examples include fludarabine phosphate, N4-behenyl-1-β-D-arabinose-furanose cytosine, N4-octadecyl-1-β-D-arabinose-furanose cytosine, N4-palmitoyl-1-(2-C-cyano-2-deoxy-β-D-arabinose-furanose)cytosine, and P-4055 (cytarabine 5'-trans oleate).

[0233] In some embodiments, the sequence includes one or more post-transcriptional modifications (e.g., capping, cleavage, polyadenylation, splicing, poly-A sequence, methylation, acylation, phosphorylation, methylation of lysine and arginine residues, acetylation, and nitrosylation of thiol and tyrosine residues, etc.). The one or more post-transcriptional modifications can be any post-transcriptional modification, such as any of the more than one hundred different nucleoside modifications identified in RNA (Rozenski, J, Crain, P, and McCloskey, J. (1999). The RNA Modification Database: 1999 update. Nucleic Acid Research 27: 196-197). In some embodiments, the first isolated nucleic acid comprises messenger RNA (mRNA). In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of: pyridine-4-ketoribonucleotide, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyluridine, 1-carboxymethyl-pseudouridine, 5-propynyluridine, 1-propynyl-pseudouridine, 5-taurylmethyluridine, 1-taurylmethyl-pseudouridine, 5-taurylmethyl-2-thio- Urate, 1-taurylmethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-denitro-pseudouridine, 2-thio-1-methyl-1-denitro-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio-pseudouridine. In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of: 5-aza-cytidine, pseudocytidine, 3-methylcytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudocytidine, pyrrolo-cytidine, pyrrolo-pseudocytidine, 2-thio-cytidine, 2-thio-5-methylcytidine, 4-thio-pseudocytidine, 4-thio-1-methyl- Pseudoisocytidine, 4-thio-1-methyl-1-denitro-pseudoisocytidine, 1-methyl-1-denitro-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-1-methyl-pseudoisocytidine.In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of: 2-aminopurine, 2,6-diaminopurine, 7-deadenine, 7-deaden-8-aza-adenine, 7-deaden-2-aminopurine, 7-deaden-8-aza-2-aminopurine, 7-deaden-2,6-diaminopurine, 7-deaden-8-aza-2,6-diaminopurine, 1-methyladenosine, N... 6-Methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenosine, 2-methylthio-adenosine, and 2-methoxy-adenosine. In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of: inosine, 1-methyl-inosine, woyoside, woyoside, 7-deazoguanosine, 7-deazo-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deazo-guanosine, 6-thio-7-deazo-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methyl-guanosine, N2-methyl-guanosine, N2,N2-dimethyl-guanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine.

[0234] The sequence may or may not be uniformly modified along the entire length of the molecule. For example, one or more or all types of nucleotides (e.g., naturally occurring nucleotides, purines or pyrimidines, or any one or all of A, G, U, C, I, pU) may or may not be uniformly modified in the sequence or in its given predetermined sequence region. In some embodiments, the sequence includes pseudouridine. In some embodiments, the sequence includes inosine, which may help the immune system characterize the sequence as endogenous relative to viral RNA. Incorporation of inosine may also mediate increased RNA stability / reduced degradation. See, for example, Yu, Z. et al. (2015), RNA editing by ADAR1 marks dsRNA as “self”. Cell Research 25, 1283-1284, which is incorporated in its entirety by reference.

[0235] In some embodiments, any RNA sequence described herein, such as the template RNA for editing, may contain end modifications (e.g., 5' or 3' modifications). In some embodiments, the end modifications are chemical modifications. In some embodiments, the end modifications are structural modifications. See the disclosure herein.

[0236] When the gene editing systems disclosed herein contain nucleic acids (e.g., mRNA molecules) encoding CRISPR nucleases and / or RT peptides, such nucleic acid molecules may contain any modifications disclosed herein, where applicable.

[0237] C. Exemplary gene editing system

[0238] The exemplary gene editing system described herein is for illustrative purposes only and may include:

[0239] (a) A fusion polypeptide or a nucleic acid encoding such a polypeptide, wherein the fusion polypeptide comprises any CRISPR nuclease polypeptide, any RT polypeptide and optionally one or more NLS (which may be located at the N-terminus and / or C-terminus), one or more peptide linkers or combinations thereof; and

[0240] (b) A single RNA molecule comprising a guide RNA, an RT donor RNA and optionally one or more nucleotide linkers, one or more 5' or 3' protection elements or combinations thereof.

[0241] In some embodiments, the CRISPR nuclease peptide in the fusion peptide may be a nickase variant (e.g., those provided in Table 5 below, such as nickase variants containing the H845 mutation, such as H845A). Alternatively or additionally, the CRISPR nuclease peptide in the fusion peptide may contain one or more arginine and / or lysine substitutions, for example at the positions disclosed herein (e.g., K736, L784, Q812, N813, I857, and / or A919 in SEQ ID NO: 1). In some specific examples, the CRISPR nuclease peptide in the fusion peptide may contain a combination of arginine and / or lysine substitutions at the positions provided herein (e.g., I857, L784, and K736). Alternatively or additionally, the CRISPR nuclease peptide in the fusion peptide may contain one or more mutations for reducing the strictness of PAM recognition, for example, said one or more mutations located at positions D61, A68, H494, L1117, D1144, S1145, G1227, E1228, S1327, A1332, R1343, R1345 and / or T1347 in SEQ ID NO: 1. In some specific instances, such mutations may include: (i) one or more arginine and / or lysine substitutions at positions D61, A68, H494, L1117, G1227, S1327, A1332 and / or T1347 of SEQ ID NO: 1, optionally with arginine substitution; (ii) one or more amino acid substitutions at positions D1144, S1145, E1228, R1343 and / or R1345 of SEQ ID NO: 1; or (iii) a combination of (i) and (ii).

[0242] In some embodiments, the RT peptide in the fusion peptide may be an MMLV variant, such as SEQ ID NO: 53 provided in Example 5 below.

[0243] In some instances, the fusion peptides provided herein may comprise an N-terminal CRISPR nuclease peptide and a C-terminal RT peptide. Optionally, the fusion peptide may comprise a peptide linker (e.g., a G / S-rich linker or an XTEN peptide linker) between the CRISPR nuclease peptide and the RT peptide. Alternatively or additionally, the fusion peptide may comprise an NLS at the N-terminus and / or C-terminus. In some instances, the fusion peptide may comprise two distinct NLSs, one at the N-terminus and the other at the C-terminus.

[0244] In some instances, the fusion peptides provided herein may have any of the configurations disclosed herein, such as those shown in Table 16 below (e.g., configuration 4, configuration 5, configuration 7, configuration 9, or configuration 10), except that the FLAG motif is removed.

[0245] A single RNA molecule contained in an exemplary gene editing system may comprise guide RNA and RT donor RNA in any orientation. Optionally, the single RNA molecule may contain one or more nucleotide linkers between the gRNA and the RT donor RNA and / or between functional domains in the gRNA (e.g., between a spacer and a scaffold sequence) and / or in the RT donor RNA (e.g., between a PBS and an RTT sequence). In some instances, the single RNA molecule may further comprise protective fragments at the 5' and / or 3' ends (e.g., those disclosed herein).

[0246] In a specific example, a single RNA molecule contained in an exemplary gene editing system may include a spacer sequence, a scaffold sequence, an RTT, PBS, and a 3' extension from 5' to 3', the extension of which may have a pseudoknot motif.

[0247] In some instances, the exemplary gene editing systems provided herein may comprise any CRISPR-RT fusion peptide or variant thereof provided in Table 8 or Table 17 below, wherein the FLAG motif is removed, and a single RNA molecule provided herein (e.g., comprising a spacer sequence, a scaffold sequence, an RTT, PBS, and a 3' extension from 5' to 3', the extension of which may have a pseudoknot motif).

[0248] In other instances, the exemplary gene editing systems provided herein may comprise any CRISPR-RT fusion peptide or variant thereof provided in Table 8 or Table 17 below, wherein the FLAG motif is removed, and a nucleic acid (e.g., a vector, such as a viral vector) encoding a single RNA molecule provided herein (e.g., comprising a spacer sequence, a scaffold sequence, an RTT, PBS, and a 3' extension from 5' to 3', the extension of which may have a pseudoknot motif).

[0249] In other instances, the exemplary gene-editing systems provided herein may comprise nucleic acids encoding any of the CRISPR-RT fusion peptides provided in Table 8 below, and a single RNA molecule provided herein (e.g., comprising a spacer sequence, a scaffold sequence, an RTT, PBS, and a 3' extension from 5' to 3', the extension of which may have a pseudoknot motif). The nucleic acid may comprise codon-optimized coding nucleotide sequences, such as those provided in Table 8 or Table 17, or variants in which the FLAG motif is removed.

[0250] In other instances, the exemplary gene-editing systems provided herein may comprise nucleic acids encoding any CRISPR-RT fusion peptides or variants thereof with the FLAG motif removed, as provided in Tables 8 or 17 below, and nucleic acids (e.g., vectors, such as viral vectors) encoding a single RNA molecule provided herein (e.g., comprising a spacer sequence, a scaffold sequence, an RTT, PBS, and a 3' extension, which may have a pseudoknot motif). The nucleic acid encoding the fusion peptide may be codon-optimized, such as those provided in Tables 8 or 17 or variants thereof with the FLAG motif removed. In some cases, the gene-editing system may comprise two vectors, one encoding the fusion peptide and the other encoding a single RNA molecule. Alternatively, the gene-editing system may comprise one vector encoding both the fusion peptide and the single RNA molecule.

[0251] III. Methods for gene editing

[0252] Any gene editing system can be used to modify (edit) a target nucleic acid, which can be a gene site of interest, such as a gene site that needs to be edited, for example, to repair gene mutations, introduce protective mutations, or introduce modifications to regulate gene expression.

[0253] A. Delivering gene editing systems to target cells

[0254] Components of any gene-editing system disclosed herein can be formulated, for example, including carriers such as liposomes and / or polymer carriers, such as liposomes, and delivered to cells (e.g., mammalian cells, etc.) by known methods. These methods include, but are not limited to, transfection (e.g., lipid-mediated cationic polymers, calcium phosphate, dendritic structures); electroporation or other membrane disruption methods (e.g., nuclear transfection); viral delivery (e.g., lentiviruses, retroviruses, adenoviruses, adeno-associated viruses (AAVs)); microinjection; microparticle bombardment (“gene gun”); fugene; direct acoustic loading; cell extrusion; optical transfection; protoplast fusion; impale infection; magnetic transfection; exogenous body-mediated transfer; lipid nanoparticle-mediated transfer; and any combination thereof. In some instances, the delivery method involves using lipid nanoparticles to mediate the delivery of one or more components of the gene-editing system disclosed herein.

[0255] In some embodiments, the method includes delivering one or more nucleic acids (e.g., nucleic acids encoding a CRISPR nuclease peptide, an RT peptide, or a fusion peptide comprising both, an RNA guide, an RT donor RNA, or a single RNA molecule comprising both, etc.), one or more transcripts thereof, and / or a pre-formed RNA guide / CRISPR nuclease peptide / RT peptide complex into a cell, wherein a ternary complex is formed. In some embodiments, the RNA guide and / or RT donor RNA, or a fusion thereof, and RNA encoding a CRISPR nuclease peptide or an RT peptide, or a fusion peptide comprising both, are delivered together in a single composition. In some embodiments, the RNA guide and RNA encoding a CRISPR nuclease peptide are delivered in separate compositions. In some embodiments, the same delivery technique is used to deliver the RNA guide / RT donor RNA and the RNA encoding a CRISPR nuclease peptide / RT peptide delivered in separate compositions. In some embodiments, different delivery techniques are used to deliver the RNA guide / RT donor RNA and the RNA encoding a CRISPR nuclease peptide / RT peptide delivered in separate compositions.

[0256] In some embodiments, one or more protein components and one or more RNA components are delivered together. For example, a CRISPR nuclease and / or RT peptide, along with RNA guide and / or RT donor RNA, are packaged together in a single AAV particle. In another example, the CRISPR nuclease and / or RT peptide, along with RNA guide and / or RT donor RNA, are delivered together via lipid nanoparticles (LNPs). In some embodiments, the CRISPR nuclease and / or RT peptide, along with RNA guide and / or RT donor RNA, are delivered separately. For example, the CRISPR nuclease and / or RT peptide, along with RNA guide and / or RT donor RNA, are packaged into separate AAV particles. In another example, the CRISPR nuclease and / or RT peptide are delivered by a first delivery mechanism, and the RNA guide and / or RT donor RNA are delivered by a second delivery mechanism.

[0257] Exemplary intracellular delivery methods include, but are not limited to, viruses, such as AAV, or virus-like agents; chemical-based transfection methods, such as those using calcium phosphate, dendritic cells, liposomes, or cationic polymers (e.g., DEAE-glucan or polyethyleneimine); non-chemical methods, such as microinjection, electroporation, cell extrusion, sonoporosis, optical transfection, puncture transfection, protoplast fusion, bacterial conjugation, plasmid or transposon delivery; particle-based methods, such as those using gene guns, magnetic transfection or magnetically assisted transfection, particle bombardment; and hybridization methods, such as nuclear transfection. In some embodiments, the lipid nanoparticles comprise mRNA encoding a CRISPR nuclease-RT fusion polypeptide, an editable template RNA, or mRNA encoding such a polypeptide. In some embodiments, this application further provides cells generated by such methods, and organisms comprising or generated from such cells (e.g., animals, plants, or fungi).

[0258] B. Genetically modified cells

[0259] Any gene-editing system disclosed herein can be delivered to a variety of cells (e.g., mammalian cells such as mouse cells, non-human primate cells, or human cells). In some embodiments, the cells are in cell cultures or co-cultures of two or more cell types. In some embodiments, the cells are in vitro. In some embodiments, the cells are obtained from a living organism and maintained in a cell culture.

[0260] In some embodiments, the cells are derived from cell lines. A variety of cell lines are known in the art for tissue culture. Examples of cell lines include, but are not limited to, 293T, MF7, K562, HeLa, CHO, and their transgenic variants. Cell lines can be obtained from a variety of sources known to those skilled in the art (see, for example, the United States Type Culture Collection (ATCC) (Manassas, Virginia)). In some embodiments, the cells are immortalized cells or cells that proliferate indefinitely.

[0261] In some embodiments, the cell is a primary cell. In some embodiments, the cell is a stem cell, such as a totipotent stem cell (e.g., pluripotent), multipotent stem cell, unipotent stem cell, oligopotent stem cell, or unipotent stem cell. In some embodiments, the cell is an induced pluripotent stem cell (iPSC) or derived from iPSC. In some embodiments, the cell is a differentiated cell. In some embodiments, the cell is a mammalian cell, such as a human cell or a mouse cell. In some embodiments, the mouse cell is derived from a wild-type mouse, an immunosuppressed mouse, or a disease-specific mouse model. In some embodiments, the cell is a cell in a living tissue, organ, or organism.

[0262] Any genetically modified cells produced using any gene-editing system disclosed herein are also within the scope of this disclosure. Such modified cells may contain disrupted target genes.

[0263] Any gene editing system, composition comprising such a system, vector, nucleic acid, RNA guide, and cell disclosed herein can be used in therapy. The gene editing systems, compositions, vectors, nucleic acids, RNA guides, and cells disclosed herein can be used in methods of treating a subject's disease or condition. Any suitable delivery or administration method known in the art can be used to deliver the compositions, vectors, nucleic acids, RNA guides, and cells disclosed herein. Such methods may involve contacting a target sequence with the compositions, vectors, nucleic acids, or RNA guides disclosed herein. Such methods may involve methods of editing target sequences as disclosed herein. In some embodiments, cells engineered using the RNA guides disclosed herein are used for ex vivo gene therapy.

[0264] IV. Therapeutic applications

[0265] Any gene-editing system or modified cells produced using such a gene-editing system as disclosed herein can be used to treat diseases associated with a target gene, such as genetic defects in the target gene.

[0266] In some embodiments, this document provides a method for treating a target disease as disclosed herein, the method comprising administering any gene-editing system disclosed herein to a subject requiring treatment (e.g., a human patient). The gene-editing system can be delivered to a specific tissue or a specific type of cell requiring gene editing. The gene-editing system may comprise an LNP covering one or more of the components, one or more vectors (e.g., a viral vector) encoding one or more of the components, or a combination thereof. The components of the gene-editing system may be formulated to form a pharmaceutical composition, which may further comprise one or more pharmaceutically acceptable carriers.

[0267] In some embodiments, modified cells produced using any of the gene-editing systems disclosed herein can be administered to a subject requiring treatment (e.g., a human patient). Modified cells may contain substitutions, insertions, and / or deletions as described herein. In some instances, modified cells may comprise a cell line modified with a CRISPR nuclease peptide, an RT peptide, or a CRISPR nuclease-RT fusion peptide, and RNA guide and RT donor RNA, or a single RNA molecule containing both. In some cases, modified cells may be a heterologous population comprising cells with different types of gene edits. Alternatively, modified cells may comprise a substantially homologous cell population (e.g., at least 80% of the cells in the entire population) containing a specific gene edit of a target gene. In some instances, cells may be suspended in a suitable culture medium.

[0268] In some embodiments, this document provides compositions comprising a gene-editing system or a component thereof. Such compositions may be pharmaceutical compositions. Useful pharmaceutical compositions may be prepared, packaged, or marketed in formulations suitable for oral, rectal, vaginal, parenteral, topical, pulmonary, intranasal, intralesional, buccal, ocular, intravenous, intra-organ, or other routes of administration. The pharmaceutical compositions disclosed herein may be prepared, packaged, or marketed in bulk in the form of a single unit dose or multiple single unit doses. As used herein, a “unit dose” is a discrete amount of a pharmaceutical composition (e.g., a gene-editing system or a component thereof) that will be the amount administered to a subject or a convenient fraction of such a dose, such as half or one-third of such a dose.

[0269] Formulas of pharmaceutical compositions suitable for parenteral administration may contain an active agent (e.g., a gene-editing system or a component thereof, or modified cells) in combination with a pharmaceutically acceptable carrier (such as sterile water or sterile isotonic saline). Such formulas may be prepared, packaged, or marketed in a form suitable for bolus or continuous administration. Some injectable formulas may be prepared, packaged, or marketed in unit dosage forms, such as ampoules or multi-dose containers containing preservatives. Some formulas for parenteral administration include, but are not limited to, suspensions, solutions, emulsions in oily or aqueous media, pastes, and implantable, sustained-release, or biodegradable formulas. Some formulas may further contain one or more additional ingredients, including, but not limited to, suspending agents, stabilizers, or dispersants.

[0270] Pharmaceutical compositions may be in the form of sterile, injectable aqueous or oily suspensions or solutions. Such suspensions or solutions may be formulated according to known techniques and may contain additional components besides cells, such as dispersants, wetting agents, or suspending agents described herein. These sterile, injectable formulations may be prepared using non-toxic, parenteral-acceptable diluents or solvents, such as water or saline. Other acceptable diluents and solvents include, but are not limited to, Ringer's solution, isotonic sodium chloride solution, and fixed oils, such as synthetic monoglycerides or diglycerides. Other parenteral-application formulations available include those that may contain cells in packaged form, in liposome formulations, or as components of biodegradable polymer systems. Some compositions for sustained release or implantation may contain pharmaceutically acceptable polymers or hydrophobic materials, such as emulsions, ion exchange resins, microsoluble polymers, or microsoluble salts.

[0271] V. reagent kits and their uses

[0272] This disclosure also provides kits that can be used, for example, to perform gene modification methods for target genes described herein. In some embodiments, the kit includes an RNA guide and an RT donor RNA, or a single RNA molecule containing both a CRISPR nuclease polypeptide and an RT polypeptide, or a fusion polypeptide thereof. In some embodiments, the kit includes a single RNA molecule and a CRISPR nuclease-RT fusion polypeptide. In some embodiments, the kit includes a polynucleotide encoding a CRISPR nuclease polypeptide, an RT polypeptide, or a CRISPR nuclease-RT fusion polypeptide, and optionally the polynucleotide is contained within a vector, such as that described herein. In some embodiments, the kit includes a polynucleotide encoding an RNA component disclosed herein. The CRISPR nuclease polypeptide, the RT polypeptide, or a fusion polypeptide thereof (or a polynucleotide encoding such a polynucleotide) and the RNA component (e.g., as a ribonucleoprotein) may be packaged in the same or other container within the kit, or may be packaged in separate vials or other containers, the contents of which may be mixed prior to use.

[0273] The CRISPR nuclease peptide, RT peptide, and RNA component can be packaged in the same or other containers within the kit, or in separate vials or other containers, the contents of which can be mixed prior to use. The kit may additionally include instructions for use of the buffer and / or RNA component, CRISPR nuclease peptide, RT peptide, or their fusion peptide.

[0274] General technology

[0275] Unless otherwise indicated, the practice of this disclosure will employ conventional molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology techniques within the scope of the art. Such techniques are well explained in the following literature: *Molecular Cloning: A Laboratory Manual*, 2nd edition (Sambrook et al., 1989), Cold Spring Harbor Press; *Oligonucleotide Synthesis* (edited by MJ Gait, 1984); *Methods in Molecular Biology*, Humana Press; *Cell Biology: A Laboratory Notebook* (edited by JE, 1989), Academic Press; *Animal Cell Culture* (edited by RI Freshney, 1987); *Introduction to Cell and Tissue Culture* (JP Mather and PE Roberts, 1998), Plenum Press; *Cell and Tissue Culture: Laboratory Procedures* (A. Doyle, JB...). Edited by Griffiths and DG Newell, 1993-8) John Wiley & Son Publishing Co., Ltd.; *Methods in Enzymology* (Academic Press, Inc.); *Handbook of Experimental Immunology* (edited by D.M. Weir and C.C. Blackwell); *Gene Transfer Vectors for Mammalian Cells* (edited by J.M. Miller and MP. Calos, 1987); *Experimental Guide to Contemporary Molecular Biology* (FM)Ausubel et al. (eds., 1987); *PCR: The Polymerase Chain Reaction* (Mullis et al. (eds., 1994); *Current Protocols in Immunology* (JE Coligan et al. (eds., 1991); *Short Protocols in Molecular Biology* (John Willie & Son Publishing, 1999); *Immunobiology* (CA Janeway and P. Travers, 1997); *Antibodies* (P. Finch, 1997); *Antibodies: a practice approach* (D. Catty. (ed., IRL Press, 1988-1989); *Monoclonal antibodies: a practical approach* (P. Shepherd and C. Dean (ed.), Oxford University Press, 2000; Using antibodies: a laboratory manual (E. Harlow and D. Lane, Cold Spring Harbor Laboratory Press, 1999); The Antibodies (M. Zanetti and JD Capra, ed.), Harwood Academic Publishers, 1995; DNA Cloning: A practical approach, Volumes I and II (D. Glover, ed., 1985); Nucleic Acid Hybridization (BD Hames and SJ Higgins, ed., 1985); Transcription and Translation (BD Hames and SJ Higgins, ed., 1984); Animal Cell Culture (RI...Freshney, ed. (1986); *Immobilized Cells and Enzymes* (IRL Press, 1986); and B. Perbal, *A Practical Guide to Molecular Cloning* (1984); FM Ausubel et al. (eds.).

[0276] Without further elaboration, it is believed that those skilled in the art can utilize this disclosure to the fullest extent based on the foregoing description. Therefore, the following specific embodiments should be interpreted as illustrative only and do not limit the remainder of this disclosure in any way. All publications cited herein for the purposes or subject matter are incorporated herein by reference.

[0277] Example

[0278] The following examples are provided to further illustrate some embodiments of this disclosure, but are not intended to limit the scope of this disclosure; it will be understood by their exemplary nature that other procedures, methods or techniques known to those skilled in the art may be used alternatively.

[0279] Example 1: CRISPR nuclease-mediated human target gene editing in HEK293T cells

[0280] This example describes genome editing of exemplary target genes, including AAVS1, EMX1, and VEGFA genes, by introducing the CRISPR nuclease of SEQ ID NO: 1 into the HEK293T cell line via lipid-based transient transfection.

[0281] CRISPR nucleases were labeled with the N-terminal SV40 nuclear localization sequence (NLS) and the C-terminal XTEN adapter located directly upstream of the nucleoplasmic protein NLS, and their coding sequences were converted into human codon-optimized DNA sequences. These sequences were then synthesized and cloned into the pcDNA3.1 vector (Invitrogen) containing the CMV promoter for expression. The reference sequences and NLS-labeled sequences used are shown in Table 1. Plasmids were purified using the midiprep kit.

[0282] Table 1. Sequences of CRISPR nuclease constructs

[0283]

[0284]

[0285]

[0286]

[0287]

[0288] An RNA guide was designed and cloned into the pUC19 plasmid following the U6 PolIII promoter, terminated with a 6x polyT sequence. The RNA guide was designed to be specific for target sequences within the coding exons of AAVS1, EMX1, and VEGFA, which have a 5'-NGG-3' PAM sequence (with the PAM sequence located at the 3' end of the target sequence). The U6 PolIII promoter plays a crucial role in the transcription process. For more efficient transcription, a +1 G is used at the beginning (i.e., the 5' end of the RNA), which is excluded from the sequence described here. See all RNA guide sequences in Table 2. Purify plasmids using the midiprep kit.

[0289] Table 2. Target sequences and RNA guide sequences

[0290]

[0291] Approximately 16 hours prior to transfection, 25,000 HEK293T cells in DMEM / 10% FBS + Pen / Strep (D10 medium) were seeded into each well of a 96-well plate. On the day of transfection, cells were 70-90% confluent. For each well to be transfected, a mixture of Lipofectamine 2000™ (Thermo Fisher Scientific) and Opti-MEM™ (Thermo Fisher Scientific) was prepared and incubated at room temperature for 5 minutes (Solution 1). After incubation, the Lipofectamine 2000™:Opti-MEM™ mixture was added to a separate mixture containing a CRISPR nuclease plasmid (NLS-tagged), an RNA guide plasmid, and Opti-MEM™ (Solution 2). In the negative control case, the CRISPR nuclease plasmid was excluded. Solutions 1 and 2 were mixed by pipetting and then incubated at room temperature for 25 minutes. After incubation, the mixture of Solutions 1 and 2 was added dropwise to each well of a 96-well plate containing cells. Approximately 72 hours post-transfection, the cells were treated with trypsin by adding TrypLE™ (Thermo Fisher Scientific) to the center of each well and incubated at 37°C for approximately 5 minutes. D10 medium was then added to each well and mixed to resuspend the cells. The resuspended cells were centrifuged for 10 minutes to obtain a pellet, and the supernatant was discarded. The cell pellet was then resuspended in QuickExtract™ buffer (Lucigen®) and incubated at 65°C for 15 minutes, 68°C for 15 minutes, and 98°C for 10 minutes.

[0292] Next-generation sequencing (NGS) samples were prepared using two rounds of PCR. Three technical replicas of the reference and each target for each variant were analyzed. The first round (PCR1) was used to amplify target-specific genomic regions. A second round of PCR (PCR2) was performed to add Inmena adaptors and indexes. The reactions were then pooled and purified by column purification. Sequencing runs were performed using a 150-cycle NextSeq 500 / 550 medium or high output v2.5 kit or a 200-cycle NovaSeq 6000 SP or S1 kit v1.5.

[0293] For NGS analysis, the indel mapping function uses the sample's FASTQ file, amplicon reference sequence, and forward primer sequence. For each read, the KMER scanning algorithm is used to calculate the edit operations (matches, mismatches, insertions, deletions) between the read and the reference sequence. To remove a small number of primer dimers present in some samples, the first 30 nt of each read is required to match the reference, and reads with more than half of the mapped nucleotide mismatches are also filtered out. Up to 50,000 reads that pass through those filters are used for analysis, and if the reads contain insertions or deletions, they are counted as indel reads. The minimum QC criterion for the number of reads passing through the filters is 10,000.

[0294] For each target, the indel ratio was calculated for each sample and its homologous protein-free control; the indel ratio refers to the fraction of NGS reads containing indels. When a CRISPR nuclease is included in the transfection, targets containing a higher percentage of indels indicate the outcome of DNA editing in the cells.

[0295] like Figure 1 As shown, when the CRISPR nuclease plasmid was present, each of the six targets tested exhibited a high level of indel. Therefore, this example demonstrates that the CRISPR nuclease of SEQ ID NO: 1 edited a human gene.

[0296] Example 2: Effectiveness of variant CRISPR nucleases in targeting exemplary mammalian genes

[0297] This example describes the indel evaluation of an exemplary mammalian target using a CRISPR nuclease variant transfected into HEK293T cells.

[0298] Arginine scanning mutagenesis was performed to individually replace selected non-arginine residues of a reference CRISPR nuclease (SEQ ID NO: 1) with arginine. SEQ ID NO: 1 is referred to herein as the reference sequence. This yielded 372 single-arginine substitution variants. The nucleic acids encoding the reference and each CRISPR nuclease variant were then individually cloned into the pcDNA3.1 backbone (Invitrogen™), and plasmids were prepared and diluted. The plasmids contained a CMV promoter, a first NLS (MKRTADGSEFESPKKKRKV; SEQ ID NO: 3) upstream of the coding sequence, an XTEN adapter (SGGSSGGSSGSETPGTSESATPESSGGSSGGSS; SEQ ID NO: 8), and a second NLS (KRPAATKKAGQAKKKK; SEQ ID NO: 4) downstream of the coding sequence. See also Example 1 above.

[0299] This study used exemplary RNA guides VEGFA-T6 and EMX1-T7. Details of these gRNAs are provided in Table 2 above. The RNA guides were cloned into the pUC19 backbone (New England Biolabs). ® The plasmids were purified and diluted using the maxi-prep kit. Cells were transfected, and NGS samples were prepared as described in Example 1. The indel ratio for the reference and each variant was calculated, where the indel ratio refers to the fraction of NGS reads containing indels. The indel ratio used for fold change calculation was the average of two technical replicas. Then, to calculate the fold change of the indel ratio, the indel ratio of each variant was divided by the indel ratio of the reference. Table 3 shows the fold change of the indel ratio for each target tested. The reference nuclease (i.e., without NLS) is numbered relative to SEQ ID NO: 1.

[0300] As shown in Table 3, six of the 372 variants with a single arginine substitution (left column) are characterized by an indel ratio that is at least 1.5 times higher than the reference indel ratio when the average is taken across two targets (right column).

[0301] Table 3. Changes in Indel Ratio*

[0302]

[0303] Fifty-five variants with a single arginine substitution were analyzed, and their indel ratios were 1-1.5 times that of the reference indel ratio: G1329R, Q741R, P1238R, A1236R, Q1230R, E1214R, A730R, Q849R, S473R, D985R, M753R, K918R, I1106R, M1205R, Y1015R, Q1360R, K1344R, F501R, K1091R, G731R, N1295R, F1241R, V752R, and N1099R. The variants included are E1179R, D720R, F1285R, I1370R, E1037R, E982R, S1094R, S872R, P1096R, A1226R, Q840R, I1147R, L1117R, E331R, T1108R, K1298R, Y986R, S898R, A1333R, P1208R, E1105R, Q809R, L1281R, A1292R, K1131R, I581R, I79R, D1284R, T1104R, Y348R, and S1348R. The remaining variants (311 variants) with a single arginine substitution resulted in a lower indel ratio relative to the reference indel ratio (a fold change in the indel ratio of less than 1.0).

[0304] The following variants were selected to engineer the combined variants, wherein the indel ratio of the variants is increased by at least 1.5 times relative to the reference indel ratio: I857R, N813R, L784R, K736R, A919R, and Q812R.

[0305] Example 3: The effectiveness of combining CRISPR nuclease variants in targeting mammalian genes

[0306] This example describes the evaluation of indels on mammalian targets using CRISPR nuclease variants containing two or more substitutions identified in Example 2 as increasing indel activity. Thirty-five combined CRISPR variants were tested.

[0307] As described in Example 2, each CRISPR nuclease variant and RNA guide was cloned. Exemplary RNA guides VEGFA-T6 and EMX1-T7 were used in this study. Details of these gRNAs are provided in Table 2 above. HEK293T cells were further transfected as described in Example 2, followed by NGS analysis. For each target, the indel ratio of the reference CRISPR nuclease (SEQ ID NO: 1) and each variant CRISPR nuclease was calculated, where the indel ratio refers to the percentage of NGS reads containing indels. The indel ratios shown in Table 4 were calculated as the average of two biological replicas, each containing two technical replicas.

[0308] Table 4. Indel ratio of mammalian targets

[0309]

[0310]

[0311] As shown in Table 4, each of the CRISPR nuclease variants with amino acid substitution combinations exhibited higher indel activity than the reference CRISPR nuclease (SEQ ID NO: 1). When averaged across two targets, nine CRISPR nuclease variants resulted in an indel ratio exceeding 0.25, indicating that over 25% of the NGS reads contained indels. These nine CRISPR nuclease variants contain the following substitution combinations: a) I857R, L784R, K736R; b) I857R, A919R, K736R; c) I857R, N813R, L784R; d) I857R, L784R, A919R; e) I857R, N813R, K736R; f) I857R, N813R; g) L784R, A919R, K736R; h) I857R, L784R; and i) I857R, A919R. When averaging across two targets, eight CRISPR nuclease variants resulted in indel ratios between 0.2 and 0.24, indicating that between 20% and 24% of NGS reads contain indels. When averaged across two targets, the 18 CRISPR nuclease variants resulted in indel ratios ranging from 0.1 to 0.19, indicating that between 10% and 19% of NGS reads contained indels. For all variants tested, the average indel ratio across two targets exceeded the reference indel ratio.

[0312] Based on this experiment, the best-performing CRISPR nuclease variant containing substitutions for I857R, L784R, and K736R was selected for further testing. Compared to the reference CRISPR nuclease, this CRISPR nuclease variant exhibited a 2.5-fold increase in indel activity.

[0313] Example 4: Engineering and effectiveness of CRISPR nickase variants targeting mammalian genes

[0314] This example describes the introduction of a mutation into the CRISPR nuclease SEQ ID NO: 1, which disrupts either the HNH or RuvC domain to produce a functional nickase. D844, H845, and N868 were identified as putative catalytic residues of the HNH domain. Positions D10, E763, and D991 were identified as putative catalytic residues of the RuvC domain. These positions were identified by analyzing models of structural regions similar to known HNH and RuvC active sites generated using AlphaFold2 (Jumper et al., Nature 596: 583-9 (2021)) and / or by sequence alignment with other nucleases at previously identified candidate positions. Examples of reference structures used to identify the HNH active site and RuvC active site are represented by the following protein database (PDB) identifiers: 5h0m, 7eu9, 6ltu, 7odf, 7lys, 8dc2, 4cmp, 4oo8, 7z4j, 5axw, 5b2o, 6kc8, 7utn, 8csz, 8ctl, 8dmb.

[0315] The coding sequence of the reference CRISPR nuclease was converted into an E. coli codon-optimized DNA sequence, synthesized, and cloned into the pET-28a(+) vector (Novagen), which contains lac and T7 RNA polymerase promoters for gene expression. To test nickase activity, individual alanine mutants were cloned for each of the positions of the putative active site residues identified as HNH and RuvC domains. A leucine mutant at position H845 was also cloned. Research-grade plasmids were obtained from GenScript. The engineered nickase sequences are shown in Table 5. In the nucleotide sequences, codons encoding substituted residues are uppercase, bold, and underlined, and in the amino acid sequences, substituted residues are shown in bold and underlined. The putative HNH knockout nickase is expected to cleave the non-target strand but not the target strand. The putative RuvC knockout nickase is expected to cleave the target strand but not the non-target strand.

[0316] Table 5. CRISPR nuclease and nickase sequences

[0317]

[0318]

[0319]

[0320]

[0321]

[0322]

[0323]

[0324]

[0325]

[0326]

[0327]

[0328]

[0329]

[0330]

[0331]

[0332]

[0333] A linear DNA template encoding an RNA guide was designed with the T7 promoter upstream and the T7Te terminator sequence downstream. The RNA guide was designed to be specific to previously tested target sequences within the coding exon of EMX1 having a 5'-NGG-3' PAM sequence (PAM being the 3' of the target sequence), as described in Example 1 above and Table 2. T7 promoter in Using +1 G at the start of the transcript (i.e., the 5' end of the RNA) enables more efficient transcription, which is specific to SEQ ID NO: 44. Show. The sequence of the encoded RNA guide and its individual components are shown in Table 6.

[0334] Table 6. RNA Sequences

[0335]

[0336] DNA targets were designed and sequenced as synthetic linear DNA fragments. Target sequences, consisting of 10 bases upstream and downstream of EMX1 and exons, were flanked by 200 unrelated bases upstream and 100 unrelated bases downstream. Additional sequences were added to ensure good separation of cleaved and uncleaved products on a gel. Target and non-target strands were labeled with 5' IR700 and 5' IR800, respectively, using PCR amplification with labeled primers. The DNA target sequences, individual components of the DNA target, and labeled PCR primers are shown in Table 7.

[0337] Table 7. Target gBlock and primer sequences

[0338]

[0339] The cleavage activity of each of the reference CRISPR nuclease (SEQ ID NO: 1) and the putative cleavage enzyme was evaluated using an in vitro cleavage assay. The cleavage activity was assessed by in vitro cleavage assays using a plasmid encoding the protein of interest from Table 5 and a linear DNA template containing EMX1-T2 sgRNA transcribed from T7 from Table 6, in a PURExpress RNase inhibitor containing SUPERase•In™ (Ingenium Biotech). ® Each peptide was co-expressed individually with the RNA guide in vitro by incubation at 37°C for 2 hours in NEB buffer solution. The unpurified peptide / RNA solution was then diluted in 1X NEB buffer 2 (NEB) containing approximately 1 ng / µl of labeled DNA target amplicon. The solution was then incubated at 37°C for 1 hour. The reaction was stopped by incubation at 37°C with RNase Cocktail™ (Ingenieur; final concentration approximately 1 U / µl) for 15 minutes, followed by incubation at 55°C with proteinase K (NEB; final concentration approximately 0.04 U / µl) for 30 minutes. DNA was then purified using CleanNGS DNA and RNA cleaning magnetic beads (Bulldog Bio).

[0340] Samples were run on 10% TBE urea PAGE gels to separate cleaved and uncleaved products of the target and non-target strands. The gels were imaged using a LI-COR Odysssey M imaging system, and 5' IR700 and 5' IR800 labels on the target and non-target strands of the target DNA substrate were visualized using 700 nm and 800 nm channels. Band intensities were quantified using ImageJ software.

[0341] Gel images are shown in Figure 2A-2C The percentage of cleaved target chain and non-target chain is quantitatively shown in the figure. Figure 2D The text indicates uncutable chains, chains cut with HNH, and chains cut with RuvC. Figure 2A The image is a gel image captured using a 700 nm channel, showing the cleavage of the target chain. Figure 2B These are gel images captured using an 800 nm channel, showing the cleavage of the non-target chains. Figure 2C It comes from Figure 2A and Figure 2B Superimposed gel images. For example... Figures 2A-2D As shown, the reference CRISPR nuclease (SEQ ID NO: 1) cleaves both the target and non-target strands as expected. Three of the four HNH knockout cleavage enzyme constructs (H845A, H845L, and N868A) showed significantly reduced activity on the target strand, while retaining activity on the non-target strand. Each of the three RuvC knockout cleavage enzyme constructs (D10A, E763A, and D991A) showed significantly reduced activity on the non-target strand, while retaining activity on the target strand. Figures 2A-2D ).

[0342] Therefore, this example demonstrates that the HNH knockout nickase and the RuvC knockout nickase have been successfully engineered. As described in Example 5 below, the H845A variant was selected to install the edits into a human gene target.

[0343] Example 5: Fusion of CRISPR nuclease and CRISPR nickase with reverse transcriptase

[0344] In this example, the reverse transcriptase peptide is fused to the C-terminus of the CRISPR nuclease of SEQ ID NO: 1 or to the H845A nickase variant of the CRISPR nuclease (SEQ ID NO: 32). See Tables 1 and 5 above.

[0345] The sequence encoding the CRISPR nuclease-reverse transcriptase fusion polypeptide was cloned into the pcDNA3.1 vector (Ingenie) containing the CMV promoter. The fusion contained the following components arranged from the N-terminus to the C-terminus: 1) SV40 NLS, 2) the CRISPR nuclease of SEQ ID NO: 1, 3) the XTEN adapter, 4) a Moloney murine leukemia virus (MMLV) reverse transcriptase variant, and 5) the nucleoplasmic protein NLS. The nucleotide and amino acid sequences of SV40 NLS, the CRISPR nuclease of SEQ ID NO: 1, the XTEN adapter, and the nucleoplasmic protein NLS are shown in Table 1 of Example 1, and the nucleotide and amino acid sequences of the variant MMLV are shown in Table 9. The variant MMLV reverse transcriptase is a human codon-optimized DNA sequence. Research-grade plasmids were obtained from GenScript.

[0346] Table 9. Reverse transcriptase (RT) sequences

[0347]

[0348]

[0349] The H845A nickase mutation was installed using a site-directed mutagenesis kit (New England Biolabs®) and then using CRISPR nuclease-reverse transcriptase fusion peptide plasmid DNA (see Example 4). The nucleotide and amino acid sequences of the CRISPR nuclease-reverse transcriptase fusion peptide and the CRISPR nickase-reverse transcriptase fusion peptide are shown in Table 8. In the nucleotide sequences, codons encoding the substituted H845A residues are uppercase, bold, and underlined, and in the amino acid sequences, the substituted residues are shown in bold and underlined. The sequence-validated plasmid was then purified using the Qiagen Maxiprep kit.

[0350] Table 8. Fusion Peptides

[0351]

[0352]

[0353]

[0354]

[0355]

[0356]

[0357]

[0358]

[0359] This example describes how to clone CRISPR nuclease-reverse transcriptase fusion peptides and CRISPR nickase (H845A)-reverse transcriptase fusion peptide constructs. These constructs were used in Example 6 for installation and editing of human target genes.

[0360] Example 6: Using CRISPR nuclease-reverse transcriptase and CRISPR nickase-reverse transcriptase fusion peptides to... Editing of human genes in HEK293T cells using RNA as a template

[0361] This example demonstrates the genetic modification of a human gene using the CRISPR nuclease-reverse transcriptase and CRISPR nickase-reverse transcriptase fusion peptide constructed in Example 5. Specifically, the fusion peptide was used to insert a 6-nucleotide sequence substitution into a human target gene.

[0362] The editing template RNA was designed to be specific to the target sequences shown in Table 11 and cloned into the pUC19 plasmid containing the U6PolIII promoter and the 6x polyT terminator sequence. The editing template was synthesized by GenScript and contained the following five components from 5' to 3': 1) a spacer sequence, 2) a scaffold motif, 3) a reverse transcription template (RTT) encoding a 6-nucleotide substitution, 4) a primer binding site (PBS), and 5) a 3' extension motif. The RTT was at least 23 nucleotides in length. Two different PBS sequences of varying lengths were tested: 7 nucleotides (editing templates 1–5) and 9 nucleotides (encoding templates 6–10). The 3' extension motif contained a short adapter sequence and a pseudoknot. The adapter sequence was added to prevent spatial collisions between the PBS and the pseudoknot motif. The pseudoknot was added to prevent 3' exonuclease activity of the editing template RNA.

[0363] This example uses guides AAVS1-T3, EMX1-T2, EMX1-T7, VEGFA-T3, and VEGFA-T6. The PAM, target, spacer, and gRNA sequences are provided in Table 2 above. The sequences of each component and the full-length editing template RNA sequence are shown in Tables 11 and 12 below. The U6 PolIII promoter uses +1 G at the start of the transcript (i.e., the 5' end of the RNA) for more efficient transduction. This is excluded from the sequences described in Table 12.

[0364] Table 11. Sequences of Edited Template RNA Components

[0365]

[0366]

[0367]

[0368] Table 12. Full-length editing template RNA sequence

[0369]

[0370]

[0371] Approximately 16 hours prior to transfection, 25,000 HEK293T cells in DMEM / 10% FBS + Pen / Strep (D10 medium) were seeded into each well of a 96-well plate. On the day of transfection, cells were 50-70% confluent. For each well to be transfected, a mixture of Lipofectamine 2000™ (Thermo Fisher Scientific) and Opti-MEM™ (Thermo Fisher Scientific) was prepared and incubated at room temperature for 5 minutes (Solution 1). After incubation, the Lipofectamine 2000™:Opti-MEM™ mixture was added to a separate mixture containing either a CRISPR nuclease-reverse transcriptase fusion peptide, an edited template RNA, and Opti-MEM™ (Solution 2) or a CRISPR nickase-reverse transcriptase fusion peptide, an encoding template RNA, and Opti-MEM™ (Solution 2). Solutions 1 and 2 were mixed by pipetting and then incubated at room temperature for 25 minutes. After incubation, the mixture of Solutions 1 and 2 was added dropwise to each well of a 96-well plate containing cells. Approximately 72 hours post-transfection, the cells were treated with trypsin by adding TrypLE™ (Thermo Fisher Scientific) to the center of each well and incubated at 37°C for approximately 5 minutes. D10 medium was then added to each well and mixed to resuspend the cells. The resuspended cells were centrifuged for 10 minutes to obtain a pellet, and the supernatant was discarded. The cell pellet was then resuspended in Quick Extract™ buffer (Lucigen®) and incubated at 65°C for 15 minutes, 68°C for 15 minutes, and 98°C for 10 minutes.

[0372] Samples were prepared as described in Example 1 for NGS and analysis. For each target, the fraction of indel-containing NGS reads for each sample and its homologous protein-free control was calculated. To determine the percentage of edits installed in the target gene, sequencing reads containing 6-nucleotide substitutions encoded by the edited template RNA were analyzed and quantified. The percentages of NGS reads containing indels and 6-nucleotide edits are shown in Tables 13 and 14, and further described in... Figure 3A and Figure 3B In. Figure 3A and Figure 3B In the table, the percentage of NGS reads is shown on the y-axis, with total edits shown as black bars and 6-nucleotide edits shown as gray bars. (Tables 13 and 14) Figure 3A and Figure 3B The data in the figure is the average of three technical replicas.

[0373] Table 13. Editing efficiency in the case of CRISPR nuclease-reverse transcriptase fusion peptides

[0374]

[0375] Table 14. Editing efficiency in the case of CRISPR nickase-reverse transcriptase fusion peptides

[0376]

[0377] like Figure 3A and Figure 3B As shown, the CRISPR nuclease-reverse transcriptase fusion peptide and the CRISPR nickase-reverse transcriptase fusion peptide introduced substitutions encoded by the tested editing template RNA at the AAVS1, EMX1, and VEGFA target loci, respectively. For the CRISPR nuclease-reverse transcriptase fusion peptide, the average percentage of NGS reads containing indels ranged from 11.67% to 41.32%, while the average percentage of NGS reads containing edit codes ranged from 0.29% to 12.80% (Table 13 and...). Figure 3A Editing template RNA 3 showed the lowest indel and encoded edit incorporation, while editing template RNA 2 showed the highest indel and encoded edit installation. For CRISPR nickase-reverse transcriptase fusion peptides, the average percentage of NGS reads containing indels ranged from 0.29% to 0.74%, while the average percentage of NGS reads containing encoded edits ranged from 0.05% to 10.66%. (Table 14 and...) Figure 3B For the CRISPR nickase-reverse transcriptase fusion peptide, low indel incorporation (less than 1%) confirmed that H845A substitution converts the CRISPR nuclease into a CRISPR nickase. For both the CRISPR nuclease-reverse transcriptase fusion peptide and the CRISPR nickase-reverse transcriptase fusion peptide, the edited template RNA 2 showed the largest encoded edit installation.

[0378] Conversely, the control consisting of the CRISPR nuclease of SEQ ID NO: 1 and the RNA guide in Table 2 or the editing templates in Tables 13 and 14 induced indel formation but did not incorporate 6-nucleotide substitutions encoded by the editing template RNA. Furthermore, the control consisting of the CRISPR nuclease-reverse transcriptase fusion peptide or the CRISPR nickase-reverse transcriptase fusion peptide and the RNA guide in Table 2 did not induce the incorporation of 6-nucleotide substitutions encoded by the editing template RNA.

[0379] The editing efficiency mediated by the CRISPR nuclease-reverse transcriptase and CRISPR nickase-reverse transcriptase fusion peptides was further tested using editing template RNAs with PBS lengths of 11, 13, 15, and 17 nucleotides and RTT lengths of 18 and 23 nucleotides. The behavior of these editing template RNAs was found to be similar to that of editing template RNAs containing PBS lengths of 7 or 9 nucleotides and RTT lengths of 23 nucleotides.

[0380] In summary, this example demonstrates that CRISPR nuclease-reverse transcriptase fusion peptides and CRISPR nickase-reverse transcriptase fusion peptides can be incorporated into human genes by substitution encoded by the edited template RNA.

[0381] Example 7: Optimization of CRISPR nuclease and CRISPR nickase fusion into reverse transcriptase

[0382] This example describes the engineering of a further fusion of the CRISPR nuclease of SEQ ID NO: 1 or the H845A nickase variant of the CRISPR nuclease (SEQ ID NO: 32) with a reverse transcriptase peptide.

[0383] The plasmid library was designed to contain various combinations and orientations of NLS tags, flexible adapters, the CRISPR nuclease of SEQ ID NO: 1 or the CRISPR nickase of SEQ ID NO: 32, the variant reverse transcriptase of SEQ ID NO: 53, and FLAG tags, and was synthesized by GenScript. The sequences of the individual NLS, adapter, and FLAG tag components are shown in Table 15. The resulting configurations are shown in Table 16.

[0384] Table 15. Sequences of library components

[0385]

[0386]

[0387]

[0388]

[0389]

[0390] Table 16. Configuration of library components

[0391]

[0392]

[0393] Plasmid libraries were screened in HEK293T cells using the lipid-based transient transfection method described in Example 5. Each well of a 96-well plate was transfected with a plasmid encoding a unique CRISPR nuclease-reverse transcriptase fusion peptide or a CRISPR nickase-reverse transcriptase fusion peptide. The editing template RNA sequence of SEQ ID NO: 70, designed to introduce a 6-nucleotide substitution into the EMX1_T2 target, was also transfected into each well. Quantification of the edits was performed as described in Example 6.

[0394] Figure 4A and Figure 4B The CRISPR nuclease-reverse transcriptase fusion peptide and the CRISPR nickase-reverse transcriptase fusion peptide are shown separately, representing the highest percentage of edits encoded by the edit template RNA in all tested constructs. The results are the average of the two technical replicas. The dashed line depicts the constructs containing the control CRISPR nuclease-reverse transcriptase fusion peptide (…). Figure 4A ) or CRISPR nickase-reverse transcriptase fusion peptide ( Figure 4B The percentage of 6-nucleotide-substituted reads installed. Figure 4A and Figure 4B As shown, compared to the control CRISPR nuclease-reverse transcriptase and CRISPR nickase-reverse transcriptase fusion peptides, several fusion constructs resulted in increased installation of 6-nucleotide substitutions. Table 17 shows the sequence of the best-performing CRISPR nickase-reverse transcriptase fusion peptide.

[0395] Table 17. Configurations and sequences of the best-performing fusion constructs

[0396]

[0397]

[0398]

[0399]

[0400]

[0401]

[0402]

[0403]

[0404]

[0405]

[0406]

[0407]

[0408]

[0409]

[0410] Therefore, this example demonstrates that optimizing the conformation of CRISPR nuclease-reverse transcriptase and CRISPR nickase-reverse transcriptase fusion peptides can improve the efficiency of editing and incorporation into targets.

[0411] Example 8: Further screening of CRISPR nickase-reverse transcriptase fusion peptides

[0412] This example describes the design and implementation of a reporter gene-based HEK293T stable cell line for measuring the activity of a CRISPR nickase-reverse transcriptase fusion peptide. This system is an orthogonal reading of the NGS-based assay used in the previous example.

[0413] A modified version of the Traffic Light Reporter Gene (TLR) assay described in Glaser et al., *Molecular Ther Nucleic Acids* 5(7): e334 (2016), was used to create stable cell lines with an integrated blue fluorescent protein (BFP) reporter gene target. The percentage of eGFP intensity in the total cell population was measured by editing BFP to eGFP via a two-amino acid change (SH>TY). In this assay, cells containing the integrated reporter gene were mCherry-positive.

[0414] Table 18 shows the sequence of the BFP target and the sequence of the editing template RNA designed to convert BFP to eGFP. As described in Example 6, the editing template RNA was cloned and transfected into stable cell lines. The CRISPR nickase-reverse transcriptase fusion peptide library described in the preceding examples was also transfected.

[0415] Table 18. Sequences screened by TLR

[0416]

[0417] TLR screening was analyzed by imaging live cells on the Operetta CLS (PerkinElmer) and its Harmony software 72 hours post-transfection. eGFP was quantified and compared to the total mCherry-positive cell population. The mCherry population represents the total number of cells containing the integrated reporter gene. Imaging data were collected and quantified as the percentage of eGFP-positive cells relative to the mCherry-positive cell population.

[0418] The best-performing CRISPR nickase-reverse transcriptase fusion peptide, selected based on TLR screening, is shown in... Figure 5 The percentage of eGFP-positive cells in the mCherry-positive population is shown in the table. These hits are ranked by activity compared to the control CRISPR nickase-reverse transcriptase peptide, and are then compared with... Figure 4A and Figure 4B The NGS data are relevant. Another observed trend is the increasing length of the flexible linker between the CRISPR nicking enzyme component and the reverse transcriptase component (linker_16xGGGGS > linker_8xGGGGS > linker_4xGGGGS > linker_1xGGGGS), which makes BFP... Increased eGFP editing leads to an increase in eGFP-positive cells in reporter gene assays. Therefore, increasing the linker length between the CRISPR nickase and reverse transcriptase components of the CRISPR nickase-reverse transcriptase fusion peptide in Table 17 may be beneficial for further improving editing efficiency.

[0419] Next, to compare the robustness of reporter gene-based eGFP quantification selected according to TLR, the best-performing hits were plotted and compared with NGS-quantified editing to convert BFP to eGFP. Figure 6 The percentage of edited cells was positively correlated with the percentage of eGFP-positive cells; higher edit installations resulted in a higher percentage of eGFP-positive cells. Finally, edits at the EMX1_T2 target from the aforementioned examples were compared with BFP-to-eGFP edits from the best-performing CRISPR nickase-reverse transcriptase fusion peptide. Figure 7 As shown, for each CRISPR nickase-reverse transcriptase fusion peptide, there was a positive correlation between editing at the two target sites. Furthermore, the editing of each CRISPR nickase-reverse transcriptase fusion peptide exceeded that of the control CRISPR nickase-reverse transcriptase fusion peptide.

[0420] Therefore, this example demonstrates that the optimized CRISPR nickase-reverse transcriptase fusion peptide introduces editing at multiple target loci.

[0421] Example 9: Design of another CRISPR nuclease

[0422] This example illustrates the design of an additional CRISPR nuclease with the desired biological activity.

[0423] The other variant of the CRISPR nickase described herein is engineered. Each individual amino acid residue from the CRISPR nickase with the H845A substitution from Example 3 was replaced with the remaining nineteen available amino acids (except for the H845A residue). The nucleic acid encoding the CRISPR nickase variant was cloned into a pcDNA3.1 vector (Ingenieur Technologies) containing a CMV promoter. The expression vector was introduced into host cells to express the CRISPR nickase variant. The CRISPR nickase variant expressed in host cells was purified, and its nickase activity was evaluated according to the procedure provided in Example 3 above.

[0424] Additional CRISPR nuclease variants were engineered to enhance double-stranded nuclease activity. As described in Example 2, the variants in Table 19 were cloned and evaluated.

[0425] Table 19. Variant CRISPR nuclease sequences

[0426]

[0427]

[0428]

[0429]

[0430]

[0431]

[0432]

[0433]

[0434]

[0435]

[0436]

[0437]

[0438]

[0439]

[0440]

[0441]

[0442]

[0443]

[0444]

[0445]

[0446]

[0447]

[0448]

[0449]

[0450]

[0451]

[0452]

[0453]

[0454]

[0455]

[0456]

[0457]

[0458]

[0459]

[0460]

[0461]

[0462]

[0463]

[0464]

[0465]

[0466]

[0467]

[0468]

[0469]

[0470]

[0471]

[0472]

[0473]

[0474]

[0475]

[0476]

[0477]

[0478]

[0479]

[0480] Other CRISPR nuclease variants were engineered, and their ability to recognize less stringent PAM sequences was evaluated. As described in Example 2, variants in Table 20 were cloned and evaluated using target sequences adjacent to the 5'-NGN-3', 5'-NRN-3', or 5'-NYN-3' PAM sequences, where N represents any nucleotide, R represents G or A, and Y represents C or T.

[0481] Table 20. Variant CRISPR nuclease sequences

[0482]

[0483]

[0484]

[0485] Example 10: Variants of CRISPR nucleases with relaxed PAM strictness targeting exemplary mammalian genes Validity

[0486] This example describes the indel evaluation of an exemplary mammalian target using a variant of a CRISPR nuclease with relaxed PAM transfected into HEK293T cells.

[0487] Arginine scanning mutagenesis was performed to individually replace selected non-arginine residues of the CRISPR nuclease variant of SEQ ID NO: 236 with arginine. This yielded 372 single-arginine substitution variants. Variants were cloned and evaluated using target sequences adjacent to the 5'-NGN-3' PAM sequence, as summarized in Table 21, as described in Example 2.

[0488] Table 21. Mammalian targets and corresponding crRNAs

[0489]

[0490] As described in Example 2, HEK293T cells were further transfected, followed by NGS analysis. The Indel activity of the CRISPR nuclease variant of SEQ ID NO: 236 is shown in Table 22. The data in Table 22 are averages of ten control samples, each containing two biological replicas and two technical replicas.

[0491] Table 22. Percentage of NGS reads containing Indel

[0492]

[0493] Next, for each target, the indel ratio of the variant CRISPR nuclease (SEQ ID NO: 236) and each variant CRISPR nuclease was calculated, where the indel ratio refers to the percentage of NGS reads containing indels. Then, to calculate the fold change in the indel ratio, the indel ratio of each variant was divided by the indel ratio of the variant CRISPR nuclease of SEQ ID NO: 236. The indel ratio used for fold change calculation was the average of two technical replicas. As shown in Table 23, three variants (left column) out of 372 variants with a single arginine substitution are characterized by an indel ratio that, when averaged across two targets (right column), increases by at least 2-fold relative to the indel ratio of the variant CRISPR nuclease of SEQ ID NO: 236.

[0494] Table 23. Changes in Indel Ratio*

[0495]

[0496] Eleven variants with a single arginine substitution were analyzed, with indel ratios 1.5 to 2 times higher than the reference indel ratio: L64R, S410R, T67R, Q849R, G1110R, F501R, T659R, L784R, Y516R, G55R, and E1037R. Ninety-two variants showed indel ratios 1 to 1.4 times higher than the reference indel ratio: N57R, D720R, A919R, A1294R, Q812R, N700R, H657R, T73R, Q899R, T1347R, I857R, K751R, D327R, I581R, D462R, E331R, A589R, D471R, I699R, and N1... 295R, T470R, I1147R, E130R, S473R, A353R, K40R, K334R, A60R, S1348R, K367R, A1118R, K3 1R, Q349R, K341R, Q83R, K585R, Q840R, G660R, K527R, G727R, Y42R, L1281R, L122R, Q123R, T 1108R, E41R, K1131R, K30R, S872R, I1206R, D1132R, K460R, L80R, E459R, K1182R, L1117R, M696R, K918R, K126R, N721R, G1227R, Q809R, K1091R, K736R, A1332R, K783R, N498R, K723R, E1228R, H1119R, F463R, L594R, D472R, K744R, E365R, G595R, K45R, Y348R, K964R, S1181R, N813R, D407R, S839R, Y658R, E586R, G754R, A730R, Y1015R, D903R, A1333R, S461R, and H1359R. The remaining variants (266 variants) with a single arginine substitution exhibited a lower indel ratio (fold change less than 1.0) compared to the variant CRISPR nuclease of SEQ ID NO: 236.

[0497] Therefore, this example demonstrates that the CRISPR nuclease variant of SEQ ID NO: 236 is an active nuclease capable of editing the target sequence adjacent to 5'-NGN-3'PAM (where N represents A, C, G, or U), and that specific additional arginine substitutions (e.g., D61R, A68R, and / or H494R) increase nuclease activity.

[0498] Example 11: mRNA-mediated target sequence editing in primary human hepatocytes

[0499] This example describes genome editing of the EMX1 and VEGFA genes using mRNA encoding either a CRISPR nuclease-reverse transcriptase fusion peptide or a CRISPR nickase-reverse transcriptase fusion peptide.

[0500] The nucleic acids encoding the CRISPR nuclease-reverse transcriptase fusion peptide of SEQ ID NO: 55 and the CRISPR nickase-reverse transcriptase fusion peptide of SEQ ID NO: 57 were cloned into an in vitro transcription (IVT) backbone containing a T7 promoter. Research-grade and sequence-validation plasmids were obtained using the Maxi Prep kit (Kiagene). The mRNAs were generated by in vitro transcription onto the IVT backbone, with the addition of a 5' cap and a 3' polyA tail. Full-length mRNA sequences are shown in Table 24. Working solutions of each mRNA were prepared in water.

[0501] Table 24. mRNA Sequences

[0502]

[0503]

[0504]

[0505]

[0506]

[0507]

[0508] Editing template RNAs 1 and 2 (Table 25) were designed to insert 3-nucleotide insertions at the EMX1 and VEGFA target genes. The editing template RNAs were ordered from GenScript as desalting synthesis guides with the following chemical modifications: 2'-O-methyl groups on the first three and last three bases; and phosphate thioester bonds between the first three and last three bases, as shown in bold in the guide (RNA) column of Table 25.

[0509] Table 25. Edited template RNA sequences

[0510]

[0511]

[0512] PHH cells from human donors (Thermo Fisher Scientific or Lonza) were rapidly thawed from liquid nitrogen in a 37°C water bath. Cells were added to preheated hepatocyte recovery medium (Thermo Fisher Scientific, CM7000) and centrifuged. The cell pellet was resuspended in an appropriate volume of William's E medium (Thermo Fisher Scientific) supplemented with a hepatocyte plate seeding supplement (containing serum) (Thermo Fisher Scientific). Cells were counted using a trypan blue viability counter and a Vi-CELL BLU cell counter. The desired number of viable cells were then washed in PBS and resuspended in P3 buffer + supplement (Lonza, V4SP-3096) and transfection enhancer oligonucleotides. The resuspended cells were aliquoted into Lonza 96-well electroporation plates. mRNA effectors (1 mg / mL in water) were mixed with a synthetic RNA guide (1 mM in water) at a 1:1 volume ratio. The mRNA / guide RNA mixture was added to each reaction at a final mRNA concentration of 25 nM. The plates were electroporated using an electroporation apparatus (Program DS-150, Lonza 4D-Nuclear Transfection). After electroporation, preheated hepatocyte plate seeding medium was added to each well and mixed very gently. For each technical replica plate, 125,000 cells of diluted nuclear-infected cells were seeded into preheated collagen-coated 96-well plates (Thermo Fisher Scientific) containing hepatocyte plate seeding medium in each well. The cells were then incubated at 37°C. After 4 hours, the medium was replaced with hepatocyte maintenance medium (William's Medium E, Thermo Fisher Scientific) supplemented with William's E Medium Cell Maintenance Mixture (Thermo Fisher Scientific).

[0513] Three days after electroporation, cells were collected from the wells using Accutase (Thermo Fisher Scientific) and transferred to a 96-well twin.tec® PCR plate (Eppendorf) and centrifuged. The culture medium was gently removed, and the cells were resuspended in DNA extraction buffer (QuickExtract). TM The samples were cycled in a PCR machine at 65°C for 15 minutes, 68°C for 15 minutes, and 98°C for 10 minutes. The samples were then frozen at -20°C and subsequently analyzed by NGS as described in Example 1.

[0514] Tables 26 and 27 illustrate the editing incorporation in PHH using mRNA encoding the CRISPR nuclease-reverse transcriptase fusion peptide of SEQ ID NO: 55 (SEQ ID NO: 243) or the CRISPR nickase-reverse transcriptase fusion peptide of SEQ ID NO: 57 (SEQ ID NO: 244), respectively. For the CRISPR nuclease-reverse transcriptase fusion peptide, the average percentage of NGS reads containing indels ranged from 12.6% to 13.2%, while the average percentage of NGS reads containing 3-nucleotide inserts ranged from 0.57% to 0.95% (Table 26). For the CRISPR nickase-reverse transcriptase fusion peptide, the average percentage of NGS reads containing indels ranged from 0.38% to 2.71%, while the average percentage of NGS reads containing 3-nucleotide inserts ranged from 0.17% to 1.05% (Table 27). Overall, the use of mRNA encoding a CRISPR nickase-reverse transcriptase fusion peptide and mRNA encoding a CRISPR nuclease-reverse transcriptase fusion peptide resulted in similar levels of 3-nucleotide insertion at the target locus.

[0515] Table 26. Editing efficiency of mRNA encoding CRISPR nuclease-reverse transcriptase fusion peptide in PHH

[0516]

[0517] Table 27. Editing efficiency of mRNA encoding CRISPR nickase-reverse transcriptase fusion peptide in PHH

[0518]

[0519] Therefore, this example demonstrates that the combination of the aforementioned mRNA and the synthesis guide can edit non-dividing human cells.

[0520] Example 12: Using CRISPR nuclease-reverse transcriptase and CRISPR nickase-reverse transcriptase fusion peptides in mice Editing mouse genes using RNA as a template

[0521] This example demonstrates the genetic modification of the DNMT1 gene in mice using mRNA encoding a CRISPR nuclease-reverse transcriptase fusion peptide or a CRISPR nickase-reverse transcriptase fusion peptide.

[0522] The mRNA molecules provided in Table 24 of Example 11 were used. The editing template RNAs provided in Table 28 below were designed to install G>C substitutions or insert CCC into the DNMT1 locus. A 12-nucleotide RTT was used to install G>C substitutions (editing template RNA 1), and a 15-nucleotide RTT was used to install CCC insertions (editing template RNA 2). The editing template RNAs provided in Table 28 were tested in the presence of the RNA guides provided in Table 29. These editing template guide RNAs and RNA guides were ordered from GenScript as HPLC-purified synthetic guides with the following chemical modifications: 2'-O-methyl groups on the first three and last three bases; and phosphate thioester bonds between the first three and last three bases, as shown in bold in the guide (RNA) column of Table 28.

[0523] Table 28. Sequences of the edited template RNA and its components

[0524]

[0525]

[0526] Table 29. RNA Guide Sequences

[0527]

[0528] The following components were encapsulated in lipid nanoparticles (LNPs):

[0529] a) mRNA encoding CRISPR nuclease-reverse transcriptase (SEQ ID NO: 55) (SEQ ID NO: 243) or mRNA encoding CRISPR nickase-reverse transcriptase fusion polypeptide (SEQ ID NO: 57) (SEQ ID NO: 244).

[0530] b) The editing template RNAs provided in Table 28, and

[0531] c) The RNA guide provided in Table 29.

[0532] LNP contains 46.3% of the cationic lipid 6-((2-hexyldecanoyl)oxy)-N-(6-(((2-hexyldecanoyl)oxy)hexyl)-N-(4-hydroxybutyl)hex-1-ammonium, 9.4% of the phospholipid 1,2-distearate-sn-glycerol-3-phosphocholine (DSPC), 42.7% cholesterol, and 1.6% of the PEG lipid 2-[(polyethylene glycol)-2000]-N,N-tetracosylacetamide, and is formulated to have a molar N / P ratio of ~6. LNP was prepared according to the general procedure described in Schoenmaker, International Journal of Pharmaceutical Sciences (IJPharm), 601:120586, 2021, the relevant disclosure of which is incorporated herein by reference for the purposes and purposes described herein.

[0533] Male C57BL6 mice (6 weeks old, Jackson Laboratories, Bar Harbor, ME) were used for these studies. Animals were acclimatized in captivity for at least 3 days prior to study. Animals were weighed before administration. Study 1 was treated with mRNA encoding the CRISPR nickase-reverse transcriptase fusion peptide. Study 2 was treated with mRNA encoding the CRISPR nuclease-reverse transcriptase fusion peptide. The ratio columns in Tables 30 and 31 refer to the ratio of mRNA encoding the fusion peptide to the editing template RNA to the RNA guide. On day 0 of the study, the LNP-modified editing template was administered intravenously in a volume of 200 µl via retroorbital injection.

[0534] Table 30: Study 1 conditions (CRISPR nickase-reverse transcriptase fusion peptide)

[0535]

[0536] Table 31: Study 2 conditions (CRISPR nuclease-reverse transcriptase fusion peptide)

[0537]

[0538] Seven days after LNP administration, animals were euthanized and perfused with PBS. Left liver lobe fragments containing 30–50 mg of tissue were collected and frozen on dry ice for tissue processing. Tissues were treated with 5 mm beads in Quick Extract™ buffer (Lucigen®), with one bead per sample treated with TissueLyser II, and then resuspended in Quick Extract™ buffer (Lucigen®). Cells were incubated at 65°C for 15 min and then at 98°C for 2 min.

[0539] Samples were prepared as described in Example 1 for NGS and analysis. The percentage of NGS reads containing indels or edited template RNA incorporated into the DNMT1 locus was quantified. Results are shown in Tables 32 and 33.

[0540] Table 32: Results of Study 1 (CRISPR nickase-reverse transcriptase fusion peptide)

[0541]

[0542] Table 33: Results of Study 2 (CRISPR nuclease-reverse transcriptase fusion peptide)

[0543]

[0544] In Study 1, with the CRISPR nickase-reverse transcriptase fusion peptide, the average percentage of NGS reads containing indels ranged from 0.42% to 0.98%, while the average percentage of NGS reads containing edit mounts ranged from 1.13% to 5.41% (Table 32). For both types of editable template RNAs, the incorporation of encoded edits increased in a dose-dependent manner (Table 32). In Study 2, with the CRISPR nuclease-reverse transcriptase fusion peptide, the average percentage of NGS reads containing indels ranged from 5.92% to 13.02%, while the average percentage of NGS reads containing edit mounts ranged from 0.186% to 1.12% (Table 33). Overall, the use of the CRISPR nickase-reverse transcriptase fusion peptide resulted in a higher edit mount at the DNMT1 locus compared to the CRISPR nuclease-reverse transcriptase fusion peptide.

[0545] Therefore, this example demonstrates that CRISPR nuclease-reverse transcriptase and CRISPR nickase-reverse transcriptase fusion peptides can be used to install editing at the DNMT1 locus in mice.

[0546] Other embodiments

[0547] All features disclosed in this specification can be combined in any combination. Each feature disclosed in this specification can be replaced by alternative features for the same, equivalent, or similar purpose. Therefore, unless otherwise expressly stated, each disclosed feature is merely an example of an equivalent or similar feature in a general series.

[0548] From the above description, those skilled in the art can readily determine the essential characteristics of the present invention, and various changes and modifications can be made to adapt it to various uses and conditions without departing from the spirit and scope of the invention. Therefore, other embodiments are also within the scope of the claims.

[0549] Equivalent solution

[0550] Although several embodiments of the invention have been described and illustrated herein, those skilled in the art will readily conceive of various other means and / or structures for performing the functions described herein and / or obtaining these results and / or one or more of these advantages, and each such variation and / or modification is considered to be within the scope of the embodiments of the invention described herein. More generally, those skilled in the art will readily understand that all parameters, dimensions, materials, and configurations described herein are exemplary, and actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications used in applying the teachings of this invention. Those skilled in the art will recognize or be able to determine many equivalents of the specific embodiments of the invention described herein using only conventional experimentation. Therefore, it should be understood that the foregoing embodiments are presented by way of example only, and that the embodiments of the invention may be practiced in ways other than those specifically described and claimed within the scope of the appended claims and their equivalents. The embodiments of the invention disclosed herein relate to each individual feature, system, article of manufacture, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits and / or methods is included within the scope of this disclosure if such features, systems, articles, materials, kits and / or methods do not contradict each other.

[0551] All definitions defined and used in this document should be understood as controls over dictionary definitions, definitions incorporated by reference in other documents, and / or the general meaning of the defined terms.

[0552] All references, patents and patent applications disclosed herein are incorporated by reference in relation to their respective subjects, and in some cases may cover the entire document.

[0553] Unless otherwise expressly stated otherwise, the indefinite articles “a” and “an” as used herein in the specification and claims shall be understood to mean “at least one (type)”.

[0554] As used herein in the specification and claims, the phrase “and / or” should be understood to mean “any one or both” of the elements so combined, i.e., elements that coexist in some cases and exist separately in others. Multiple elements listed with “and / or” should be interpreted in the same way, i.e., “one or more” of the elements so combined. In addition to the elements specifically identified by the “and / or” clause, other elements may optionally exist, whether related to or unrelated to those specifically identified. Thus, as a non-limiting example, when used in conjunction with open-ended language such as “comprising,” a reference to “A and / or B” may in one embodiment refer only to A (optionally including elements other than B); in another embodiment, only to B (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); and so on.

[0555] As used herein in this specification and claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” should be interpreted as inclusive, i.e., including a plurality of elements or at least one element in a list of elements, but also including more than one element and optionally additional unlisted items. Terms that are clearly indicated only in reverse, such as “only one of…” or “exact one of…” or, when used in a claim, “consisting of…” will refer to a plurality of elements or exactly one element in a list of elements. Generally, when preceded by an exclusive term such as “any one,” “one of…,” “only one of…” or “exact one of…”, the term “or” as used herein should be interpreted only as indicating an exclusive alternative (i.e., “one or the other, but not both”). When used in a claim, “consisting substantially of…” should have the ordinary meaning as used in the field of patent law.

[0556] As used herein in this specification and claims, the phrase "at least one" relating to a list of one or more elements should be understood to mean at least one element selected from any one or more elements in the element list, but not necessarily including at least one element from every element specifically listed in the element list, and does not exclude any combination of elements in the element list. This definition also allows for the optional presence of elements other than those specifically identified in the element list referred to by the phrase "at least one," whether related to or unrelated to those specifically identified elements. Thus, as a non-limiting example, in one embodiment, "at least one of A and B" (or equivalently, "at least one of A or B," or equivalently, "at least one of A and / or B") may mean optionally including more than one A, with no B (and optionally including elements other than B); in another embodiment, it may mean optionally including more than one B, with no A (and optionally including elements other than A); in yet another embodiment, it may mean optionally including at least one more than one A, and optionally including more than one B (and optionally including other elements); and so on.

[0557] It should also be understood that, unless explicitly stated otherwise, in any method claimed herein that includes more than one step or action, the order of the steps or actions of the method is not necessarily limited to the order of the steps or actions of the described method.

Claims

1. A gene editing system comprising: (a) A fusion polypeptide or a first nucleic acid encoding the fusion polypeptide, the fusion polypeptide comprising a CRISPR nuclease polypeptide and a reverse transcriptase (RT) polypeptide; wherein the CRISPR nuclease polypeptide comprises the amino acid sequence of SEQ ID NO: 1, or a variant of SEQ ID NO: 1, the variant comprising: (i) One or more mutations in the HNH nuclease domain or RuvC nuclease domain of SEQ ID NO: 1, wherein the one or more mutations reduce or eliminate the nuclease activity of the HNH nuclease domain or RuvC nuclease domain; (ii) One or more arginine and / or lysine substitutions, optionally one or more arginine substitutions; (iii) One or more mutations used to reduce the strictness of PAM recognition; or (iv) any combination of (i), (ii) and (iii); (b) An RNA molecule or a second nucleic acid encoding the RNA molecule, the RNA molecule comprising guide RNA (gRNA) and reverse transcription donor RNA (RT donor RNA). The gRNA contains a scaffold sequence that can be recognized by CRISPR nucleases and a spacer sequence specific to a target sequence within the genomic locus of interest, the target sequence being located upstream of the original spacer adjacent motif (PAM); and The RT donor RNA contains a primer binding site (PBS) and a template sequence.

2. The gene editing system of claim 1, wherein the fusion polypeptide further comprises the CRISPR nuclease polypeptide, the RT polypeptide, or one or more nuclear localization signals (NLS) upstream or downstream of both.

3. The gene editing system of claim 2, wherein the fusion polypeptide comprises a first NLS, the CRISPR nuclease polypeptide, the RT polypeptide, and a second NLS from the N-terminus to the C-terminus; optionally, wherein the fusion polypeptide further comprises a peptide linker between the CRISPR nuclease polypeptide and the RT polypeptide.

4. The gene editing system of claim 2, wherein the fusion polypeptide comprises a first peptide linker located between the CRISPR nuclease polypeptide and the RT polypeptide, and wherein the fusion polypeptide comprises a first NLS and a second NLS located at the N-terminus and / or the C-terminus of the fusion polypeptide.

5. The gene editing system of claim 4, wherein the fusion polypeptide further comprises a third NLS and optionally a fourth NLS.

6. The gene editing system of claim 4, wherein the fusion polypeptide further comprises a second peptide linker connecting the CRISPR nuclease polypeptide or the RT polypeptide to the first or second NLS; optionally, wherein the fusion polypeptide further comprises a third peptide linker connecting the two NLSs.

7. The gene editing system according to any one of claims 4 to 6, wherein the fusion polypeptide comprises, from the N-terminus to the C-terminus: (i) the first NLS, the second NLS, the CRISPR nuclease polypeptide, the first peptide linker, and the RT polypeptide; (ii) the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the first NLS, and the second NLS; (iii) The first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the third NLS, the second peptide linker, and the second NLS; (iv) the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the third NLS, the second NLS, the second peptide linker, and the first NLS; (v) The first NLS, the second peptide linker, the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the third peptide linker, and the second NLS; (vi) The first NLS, the second peptide linker, the CRISPR nuclease, the first peptide linker, the RT peptide, the third peptide linker, the third NLS, the fourth NLS, and the second NLS; (vii) The first NLS, the second peptide linker, the RT polypeptide, the first peptide linker, the CRISPR nuclease polypeptide, the third peptide linker, the third NLS, the fourth NLS, and the second NLS; (viii) the RT peptide, the first peptide linker, the CRISPR nuclease peptide, the third NLS, the second NLS, the second peptide linker, and the first NLS; (ix) The first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the second peptide linker, the second NLS, and the RT polypeptide; (x) The first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the third NLS, the second peptide linker, the RT polypeptide, the third peptide linker, and the second NLS; (xi) The first NLS, the RT polypeptide, the first peptide linker, the second NLS, the second peptide linker, and the CRISPR nuclease polypeptide. (xii) The first NLS, the RT peptide, the first peptide linker, the third NLS, the second peptide linker, the CRISPR nuclease peptide, the third peptide linker, and the second NLS; (xiii) The first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, and the second NLS; or (xiv) The first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the second peptide linker, the third peptide linker, and the second NLS.

8. The gene editing system of claim 7, wherein the fusion polypeptide comprises (iv), (v), (vii), (ix), or (x) from the N-terminus to the C-terminus.

9. The gene editing system according to any one of claims 4 to 8, wherein the length of the peptide linker between the CRISPR nuclease polypeptide and the RT polypeptide is about 20-80 amino acids.

10. The gene editing system according to any one of claims 1 to 9, wherein the CRISPR nuclease polypeptide is the variant of SEQ ID NO:

1.

11. The gene editing system of claim 10, wherein the CRISPR nuclease polypeptide is a variant of SEQ ID NO:1, the variant comprising one or more mutations in the HNH nuclease domain at positions D844, H845 and / or N868 relative to SEQ ID NO:1; optionally wherein the mutation is located at position H845, the mutation optionally being H845A.

12. The gene editing system according to claim 11, wherein: (a) The mutation at D844 is an amino acid substitution of D844A, D844G, D844L or D844S; (b) The mutation at H845 is an amino acid substitution of H845A, H845G, H845L, or H845S; and (c) The mutation at N868 is an amino acid substitution of N868A, N868G, N868L or N868S.

13. The gene editing system according to any one of claims 10 to 12, wherein the CRISPR nuclease polypeptide comprises a bridged helix (BH) domain, a nucleic acid recognition (REC) domain, and a phosphate-locked ring (PLL), wedge (WED), and PAM interaction (PID) domains, and wherein one or more arginine and / or lysine residues are substituted, optionally with arginine substitutions located in the BH domain, the REC domain, the PLL domain, the WED domain, the PID domain, or a combination thereof.

14. The gene editing system according to 10, wherein the CRISPR nuclease variant is any one of the variants listed in Table 19 or Table 20.

15. The gene editing system according to any one of claims 10 to 14, wherein the CRISPR nuclease polypeptide contains up to 20 arginine and / or lysine substitutions relative to SEQ ID NO: 1; optionally, wherein the CRISPR nuclease polypeptide contains up to 15 arginine and / or lysine substitutions relative to SEQ ID NO: 1; preferably, wherein one or more arginine and / or lysine substitutions are located at positions K736, L784, Q812, N813, I857 and / or A919 of SEQ ID NO: 1, wherein the substitution is optionally located at position I857, wherein the substitution is optionally I857R.

16. The gene editing system of claim 15, wherein the CRISPR nuclease polypeptide contains at least two arginine and / or lysine substitutions relative to SEQ ID NO: 1, and wherein the at least two arginine and / or lysine substitutions are located at positions K736, L784, Q812, N813, I857 and / or A919 of SEQ ID NO:

1.

17. The gene editing system of claim 16, wherein the CRISPR nuclease polypeptide contains arginine and / or lysine substitutions at the following positions relative to SEQ ID NO: 1: (a) I857, L784 and K736; (b) I857, A919 and K736; (c) I857, N813 and L784; (d) I857, L784 and A919; (e) I857, N813 and K736; (f) I857 and N813; (g) L784, A919 and K736; (h) I857 and L784; and (i) I857 and A919.

18. The gene editing system of claim 17, wherein the CRISPR nuclease polypeptide comprises the following arginine substitutions relative to SEQ ID NO: 1: (a) I857R, L784R and K736R; (b) I857R, A919R and K736R; (c) I857R, N813R and L784R; (d) I857R, L784R, and A919R; (e) I857R, N813R and K736R; (f) I857R and N813R; (g) L784R, A919R and K736R; (h) I857R and L784R; and (i) I857R and A919R; Optionally, the CRISPR nuclease polypeptide comprises (a) the arginine substitution.

19. The gene editing system of claim 1, wherein the CRISPR nuclease polypeptide comprises a nicking enzyme mutation at position H845 relative to SEQ ID NO: 1, optionally H845A, and an arginine and / or lysine substitution at position I857, optionally I857R.

20. The gene editing system according to any one of claims 1 to 19, wherein the CRISPR nuclease polypeptide comprises one or more mutations for reducing PAM recognition strictness, optionally wherein the one or more mutations are located at positions D61, A68, H494, L1117, D1144, S1145, G1227, E1228, S1327, A1332, R1343, R1345 and / or T1347 of SEQ ID NO:

1.

21. The gene editing system of claim 20, wherein the one or more mutations comprise: (i) One or more arginine and / or lysine substitutions at positions D61, A68, H494, L1117, G1227, S1327, A1332 and / or T1347 of SEQ ID NO: 1, optionally with arginine substitution; (ii) One or more amino acid substitutions at positions D1144, S1145, E1228, R1343 and / or R1345 of SEQ ID NO: 1, optionally D1144L, S1145W, E1228Q, R1343P, R1345V and / or R1345Q; or (iii) Combinations of (i) and (ii).

22. The gene editing system of claim 21, comprising the following combination of mutations relative to SEQ ID NO: 1: (i) L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R and A68R; (ii) L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and D61R; or (iii) L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R and H494R.

23. The gene editing system according to any one of claims 10 to 22, wherein the CRISPR nuclease polypeptide comprises (a) one or more mutations in the HNH nuclease domain at positions D844, H845 and / or N868 relative to SEQ ID NO: 1; optionally wherein the mutation is located at position H845; and (b) one or more arginine and / or lysine substitutions relative to SEQ ID NO: 1, optionally wherein the arginine and / or lysine substitutions are located at positions I857, L784 and K736.

24. The gene editing system according to any one of claims 10 to 23, wherein the CRISPR nuclease polypeptide comprises at least 90% of the same amino acid sequence as SEQ ID NO:

1.

25. The gene editing system of claim 24, wherein the CRISPR nuclease polypeptide comprises at least 95% of the same amino acid sequence as SEQ ID NO:

1.

26. The gene editing system of claim 25, wherein the CRISPR nuclease polypeptide comprises at least 98% of the same amino acid sequence as SEQ ID NO:

1.

27. The gene editing system according to any one of claims 1 to 26, wherein the RT polypeptide is Moloney murine leukemia virus (MMLV)-RT, optionally the MMLV-RT comprising the amino acid sequence of SEQ ID NO:

53.

28. The gene editing system of claim 1, wherein the fusion polypeptide is shown in Table 8 or Table 17.

29. The gene editing system according to any one of claims 1 to 28, wherein the system comprises the fusion polypeptide.

30. The gene editing system according to any one of claims 1 to 29, wherein the system comprises the first nucleic acid encoding the fusion polypeptide.

31. The gene editing system of claim 30, wherein the first nucleic acid is located on a vector, wherein the vector is optionally a viral vector.

32. The gene editing system of claim 31, wherein the first nucleic acid is a first messenger RNA (mRNA).

33. The gene editing system according to any one of claims 1 to 32, wherein the length of the spacer sequence in the gRNA of (b) is 15-30 nucleotides; the length of the spacer sequence is optionally 15-20 nucleotides.

34. The gene editing system according to any one of claims 1 to 33, wherein the PAM is 5'-NDR-3' or 5'-NGN-3', wherein N represents any nucleotide, D represents A, G or T, and R represents G or A; optionally wherein the PAM is 5'-NRG-3' or 5'-NRR'3', wherein N and R are as defined herein; preferably wherein the PAM is 5'-NGG-3', wherein N represents any nucleotide.

35. The gene editing system according to any one of claims 1 to 34, wherein the scaffold sequence comprises at least 85% identical nucleotide sequences to SEQ ID NO:

2.

36. The gene editing system of claim 35, wherein the scaffold sequence comprises the nucleotide sequence of SEQ ID NO:

2.

37. The gene editing system according to any one of claims 1 to 36, wherein the length of the PBS in the RT donor RNA is 5-50 nucleotides; the length of the PBS is optionally 5-20 nucleotides.

38. The gene editing system according to any one of claims 1 to 37, wherein the PBS binds to a site of the target PBS that is adjacent to or overlaps with the target sequence.

39. The gene editing system according to any one of claims 1 to 38, wherein the site targeting PBS is adjacent to or overlaps with the target sequence.

40. The gene editing system of claim 39, wherein the site targeting PBS is adjacent to the 5' of the PAM, and optionally the 3' nucleotide of the site targeting PBS is located about 2-15 nucleotides upstream of the PAM.

41. The gene editing system according to any one of claims 1 to 40, wherein the template sequence in the RT donor RNA is 5-100 nucleotides in length; the template sequence is optionally 15-25 nucleotides in length.

42. The gene editing system according to any one of claims 1 to 41, wherein the template sequence in the RT donor RNA is homologous to the genomic site of interest and contains one or more nucleotide variations relative to the genomic site of interest.

43. The gene editing system of claim 42, wherein at least one nucleotide variant is located within the target sequence; and / or wherein at least one nucleotide variant is located within the PAM.

44. The gene editing system according to any one of claims 1 to 43, wherein the RNA molecule of (b) further comprises a 3' end extension.

45. The gene editing system according to any one of claims 1 to 44, wherein the RNA molecule of (b) further comprises a 5' end protector, a 3' end protector, or both, each of the 5' end protector and the 3' end protector forming a secondary structure, the secondary structure optionally being a hairpin, pseudoknot, circular, or triplet structure.

46. ​​The gene editing system according to any one of claims 1 to 45, wherein the RNA molecule of (b) comprises from 5' to 3' the following: (i) the spacer subsequence, the scaffold sequence, the template sequence, and the PBS; or (ii) the spacer subsequence, the scaffold sequence, the template sequence, the PBS, and the 3' extension.

47. The gene editing system according to any one of claims 1 to 46, wherein the system comprises the RNA molecule of (b).

48. The gene editing system according to any one of claims 1 to 46, wherein the system comprises the second nucleic acid encoding the RNA molecule.

49. The gene editing system of claim 48, wherein the nucleic acid is located on a vector, wherein the vector is optionally a viral vector.

50. The gene editing system according to any one of claims 1 to 49, wherein the system comprises one or more lipid nanoparticles (LNPs) associated with one or more of elements (a)-(b).

51. The gene editing system according to any one of claims 1 to 50, wherein the system comprises a viral vector of one or more coding elements (a)-(b), optionally one or more adeno-associated virus (AAV) vectors.

52. A pharmaceutical composition comprising the gene editing system according to any one of claims 1 to 51.

53. A kit comprising elements (a)-(b) of a gene editing system according to any one of claims 1 to 51.

54. A gene editing method comprising delivering a gene editing system according to any one of claims 1 to 51 to a host cell to edit a genomic site targeted by the gRNA of the gene editing system.

55. The gene editing method according to claim 54, wherein the host cell is cultured in vitro.

56. The gene editing method of claim 55, wherein the host cell is located in the subject to whom gene editing is required.

57. A fusion polypeptide comprising a CRISPR nuclease polypeptide as set forth in any one of claims 1 to 26 and a reverse transcriptase polypeptide as set forth in claim 1 or claim 27.

58. The fusion polypeptide according to claim 57, comprising the amino acid sequence of SEQ ID NO: 55 or 57.

59. A nucleic acid encoding a fusion polypeptide according to claim 57 or claim 58.

60. The nucleic acid according to claim 59, comprising the nucleotide sequence of SEQ ID NO: 54, 243, 56 or 244.

61. The nucleic acid according to claim 60, wherein it is a vector, optionally an expression vector.