Reverse transcription-mediated gene editing systems and their use

JP2026530078APending Publication Date: 2026-09-03ARBOR BIOTECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026513623
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-25
Filing Date
2024-08-30
Publication Date
2026-09-03

Smart Images

  • Figure 2026530078000001_ABST
    Figure 2026530078000001_ABST
Patent Text Reader

Abstract

A gene editing system comprising (a) a fusion polypeptide comprising a CRISPR nuclease and a reverse transcriptase, or a nucleic acid encoding a fusion polypeptide, and (b) an RNA molecule comprising a guide RNA and a reverse transcription donor RNA, or a nucleic acid encoding an RNA molecule. Methods of using a gene editing system to modify a target gene of interest are also provided herein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-Reference to Related Applications This application claims the benefit of the filing dates of U.S. Provisional Application No. 63 / 580,188 filed on September 1, 2023, U.S. Provisional Application No. 63 / 580,168 filed on September 1, 2023, U.S. Provisional Application No. 63 / 553,974 filed on February 15, 2024, and U.S. Provisional Application No. 63 / 638,559 filed on April 25, 2024. Each of the priority applications is incorporated herein by reference in their entireties.

[0002] Sequence Listing This application contains a Sequence Listing which has been electronically filed in XML format, and is incorporated herein by reference in its entirety. The XML copy, created on August 28, 2024, is named 063586-525001WO_SeqList_ST26.xml and has a size of 597.0 bytes. [Background Art]

[0003] Clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) genes, collectively known as CRISPR-Cas or CRISPR / Cas systems, are adaptive immune systems in archaea and bacteria that protect specific species against exogenous genetic elements.

[0004] Reverse transcriptase (RT) is an enzyme that generates a DNA strand complementary to an RNA template. The combination of reverse transcriptase and the CRISPR / Cas system has shown great potential in gene editing. CRISPR-guided reverse transcription enables the introduction of desired nucleotide substitutions at genomic loci.

[0005] Therefore, it is of great interest to develop an efficient and accurate reverse transcriptase-CRISPR gene editing system for use in disease treatment. [Summary of the Invention]

[0006] This disclosure provides a reverse transcriptase-CRISPR-mediated gene editing system that successfully introduces a designed nucleotide substitution into a target gene site. In some embodiments, the gene editing systems disclosed herein include a fusion polypeptide comprising a CRISPR nuclease fragment and a reverse transcriptase (RT) fragment, and optionally one or more nuclear localization sequences (NLS) and / or peptide linkers. The CRISPR nuclease polypeptide can be genetically modified to have advantageous enzymatic activity (e.g., high indel activity and / or DNA cleavage activity, as well as precise gene editing). Therefore, the gene editing systems provided herein are expected to exhibit superior effectiveness in inserting desired base substitutions into target genomic sites.

[0007] Accordingly, one aspect of the present disclosure features a gene editing system comprising (a) a fusion polypeptide comprising a CRISPR nuclease polypeptide and a reverse transcriptase (RT) polypeptide, or a first nucleic acid encoding the fusion polypeptide, and (b) an RNA molecule comprising a guide RNA (gRNA) and a reverse transcription donor RNA (RT donor RNA), or a second nucleic acid encoding the RNA molecule. The CRISPR nuclease may be the reference nuclease of Sequence ID No. 1 or a variant thereof. The gRNA comprises a scaffold sequence recognizable by the CRISPR nuclease and a spacer sequence specific to a target sequence within a genomic site of interest, the target sequence being upstream of a protospacer adjacent motif (PAM). The RT donor RNA comprises a primer binding site (PBS) and a template sequence.

[0008] In some embodiments, the CRISPR nuclease polypeptide comprises the amino acid sequence of SEQ ID NO: 1. In other embodiments, the CRISPR nuclease is a variant of SEQ ID NO: 1, the variant comprising (i) one or more mutations in the HNH nuclease domain or the RuvC nuclease domain of SEQ ID NO: 1 that reduce or eliminate its nuclease activity, (ii) one or more arginine and / or lysine substitutions, optionally one or more arginine substitutions, or (iii) a combination of (i) and (ii).

[0009] In some cases, the fusion polypeptide contains one or more nuclear localization signals (NLS) upstream or downstream of the CRISPR nuclease polypeptide, the RT polypeptide, or both. For example, the fusion polypeptide contains a first NLS, a CRISPR nuclease polypeptide, an RT polypeptide, and a second NLS, from the N-terminus to the C-terminus. In some cases, the fusion polypeptide may further contain a peptide linker between the CRISPR nuclease polypeptide and the RT polypeptide.

[0010] In some cases, the fusion polypeptide includes a first peptide linker located between the CRISPR nuclease polypeptide and the RT polypeptide. Furthermore, the fusion polypeptide may include a first NLS and a second NLS located at the N-terminus and / or C-terminus of the fusion polypeptide. In some cases, the fusion polypeptide may further include additional NLSs, e.g., a third NLS, and optionally a fourth NLS. Alternatively or in addition, the fusion polypeptide may further include a second peptide linker and optionally a third peptide linker. These peptide linkers may be located (e.g., linked) between the CRISPR nuclease polypeptide and / or the RT polypeptide, as well as between the first and / or second NLS. In some cases, these peptide linkers may be located between two NLSs (e.g., linked).

[0011] The following provides specific exemplary configurations (from N-terminus to C-terminus) of the fusion polypeptides disclosed herein. (i) First NLS, second NLS, CRISPR nuclease polypeptide, first peptide linker, and RT polypeptide, (ii) CRISPR nuclease polypeptide, first peptide linker, RT polypeptide, first NLS, and second NLS, (iii) First NLS, CRISPR nuclease polypeptide, first peptide linker, RT polypeptide, third NLS, second peptide linker, and second NLS, (iv) CRISPR nuclease polypeptide, first peptide linker, RT polypeptide, third NLS, second NLS, second peptide linker, and first NLS, (v) First NLS, second peptide linker, CRISPR nuclease polypeptide, first peptide linker, RT polypeptide, third peptide linker, and second NLS, (vi) First NLS, second peptide linker, CRISPR nuclease, first peptide linker, RT polypeptide, third peptide linker, third NLS, fourth NLS, and second NLS, (vii) First NLS, second peptide linker, RT polypeptide, first peptide linker, CRISPR nuclease polypeptide, third peptide linker, third NLS, fourth NLS, and second NLS, (viii) RT polypeptide, first peptide linker, CRISPR nuclease polypeptide, third NLS, second NLS, second peptide linker, and first NLS, (ix) First NLS, CRISPR nuclease polypeptide, first peptide linker, second peptide linker, second NLS, and RT polypeptide, (x) First NLS, CRISPR nuclease polypeptide, first peptide linker, third NLS, second peptide linker, RT polypeptide, third peptide linker, and second NLS, (xi) First NLS, RT polypeptide, first peptide linker, second NLS, second peptide linker, and CRISPR nuclease polypeptide, (xii) First NLS, RT polypeptide, first peptide linker, third NLS, second peptide linker, CRISPR nuclease polypeptide, third peptide linker, and second NLS, (xiii) the first NLS, CRISPR nuclease polypeptide, the first peptide linker, RT polypeptide, and the second NLS, or (xiv) First NLS, CRISPR nuclease polypeptide, first peptide linker, RT polypeptide, second peptide linker, third peptide linker, and second NLS.

[0012] In specific examples, fusion polypeptides may have the configurations (iv), (v), (vii), (ix), or (x) (from N-terminus to C-terminus). See, for example, Table 16 below.

[0013] In some cases, the peptide linker(s) between the CRISPR nuclease polypeptide and the RT polypeptide is approximately 20 to 80 amino acids long.

[0014] In some embodiments, the CRISPR nuclease polypeptide contains a variant of SEQ ID NO: 1. For example, the variant of SEQ ID NO: 1 may contain one or more mutations in the HNH nuclease domain at positions D844, H845, and / or N868 compared to SEQ ID NO: 1. In one example, the mutation is at position H845 (e.g., a mutation of H845A). In some examples, the mutation at D844 is an amino acid substitution of D844A, D844G, D844L, or D844S. In some examples, the mutation at H845 is an amino acid substitution of H845A, H845G, H845L, or H845S. In one specific example, the CRISPR nuclease polypeptide contains a mutation at position H845 (e.g., H845A) compared to SEQ ID NO: 1 (e.g., an amino acid sequence of SEQ ID NO: 32). In some cases, mutations in N868 are amino acid substitutions of N868A, N868G, N868L, or N868S.

[0015] Alternatively or in addition, CRISPR nuclease polypeptides include a bridge helix (BH) domain, a nucleic acid recognition (REC) domain, a phosphate-locked loop (PLL) domain, a wedge (WED) domain, and a PAM interaction (PID) domain, with one or more arginine and / or lysine substitutions, optionally arginine substitutions, located in the BH domain, the REC domain, the PLL domain, the WED domain, the PID domain, or a combination thereof. For example, a CRISPR nuclease polypeptide may contain up to 20 arginine and / or lysine substitutions relative to a reference CRISPR nuclease. In a more specific example, a CRISPR nuclease polypeptide may contain up to 15 arginine and / or lysine substitutions relative to a reference CRISPR nuclease. In a more specific example, one or more arginine and / or lysine substitutions may be located at positions K736, L784, Q812, N813, I857, and / or A919.

[0016] In some cases, the CRISPR nuclease polypeptide contains at least two arginine and / or lysine substitutions relative to the reference CRISPR nuclease. The two arginine and / or lysine substitutions are located at positions K736, L784, Q812, N813, I857, and / or A919. In some cases, the CRISPR nuclease polypeptide contains arginine and / or lysine substitutions relative to the reference CRISPR nuclease at the following positions: (a) I857, L784, and K736 (e.g., I857R, L784R, K736R), (b) I857, A919, and K736 (I857R, A919R, and K736R), (c) I857, N813, and L784 (I857R, N813R, and L784R), (d) I857, L784, and A919 (I857R, L784R, and A919R), (e) I857, N813, and K736 (I857R, N813R, and K736R), (f) I857 and N813 (I857R and N813R), (g) L784, A919, and K736 (L784R, A919R, and K736R), (h) I857 and L784 (I857R and L784R), or (i) Contained in I857 and A919 (I857R and A919R).

[0017] In a specific example, the CRISPR nuclease polypeptide includes the arginine substitutions I857R, L784R, and K736R in SEQ ID NO: 1.

[0018] In some cases, the CRISPR nuclease polypeptides disclosed herein may further include one or more mutations that enhance the double-stranded nuclease activity of the reference CRISPR nuclease into which the mutation is introduced. Examples are provided in Table 19 below.

[0019] Alternatively or additionally, the engineered CRISPR nuclease polypeptides disclosed herein may comprise, or may further comprise, one or more mutations to reduce PAM recognition stringency. In some instances, the one or more mutations to reduce PAM recognition stringency can be at positions D61, A68, H494, L1117, D1144, S1145, G1227, E1228, S1327, A1332, R1343, R1345, and / or T1347 of SEQ ID NO: 1. In some examples, such mutations can be (i) one or more arginine and / or lysine substitutions at positions D61, A68, H494, L1117, G1227, S1327, A1332, and / or T1347 of SEQ ID NO: 1, optionally arginine substitutions, (ii) one or more amino acid substitutions at positions D1144, S1145, E1228, R1343, and / or R1345 of SEQ ID NO: 1, or (iii) a combination of (i) and (ii). In a specific example, the one or more amino acid substitutions of (ii) can optionally comprise D1144L, S1145W, E1228Q, R1343P, R1345V, and / or R1345Q relative to SEQ ID NO: 1.

[0020] In specific examples, a constructed CRISPR nuclease polypeptide with less PAM recognition stringency may include the following combinations of mutations: L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and A68R relative to SEQ ID NO: 1. In other specific examples, a constructed CRISPR nuclease polypeptide with less PAM recognition stringency may include the following combinations of mutations: L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and D61R relative to SEQ ID NO: 1. In further specific examples, a constructed CRISPR nuclease polypeptide with less PAM recognition stringency may include the following combinations of mutations: L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and H494R relative to SEQ ID NO: 1. Other exemplary constructed CRISPR nuclease polypeptides can be seen in Table 20, each of which is within the scope of this disclosure.

[0021] The CRISPR nuclease polypeptides prepared herein recognize the 5'-NDR-3' PAM sequence, where N represents A, C, G, or U, D represents A, G, or T, and R represents G or A. In some cases, the CRISPR nuclease polypeptides prepared with reduced PAM recognition stringency, as disclosed herein, may recognize the 5'-NGN-3' PAM sequence. See Example 10 below. In some examples, the PAM is 5'-NRG-3' or 5'-NRR-3', where N and R are as defined herein. In some specific examples, the PAM is 5'-NGG-3', where N represents any nucleotide. In other specific examples, the PAM may be 5'-TGC-3' or 5'-GGA-3'.

[0022] Any of the CRISPR nuclease polypeptides may comprise any of the arginine and / or lysine substitutions disclosed herein, as well as any of the nickase mutations in any of the HNH or RuvC nuclease domains disclosed herein, any of the mutations that result in reduced PAM recognition stringency, or combinations thereof. For example, a CRISPR nuclease polypeptide can comprise (a) one or more nickase mutations in the HNH nuclease domain at positions D844, H845, and / or N868 relative to SEQ ID NO: 1 (e.g., the mutation is at position H845), and (b) one or more arginine and / or lysine substitutions (e.g., positions I857, L784, and K736) relative to SEQ ID NO: 1. For example, in some instances, a CRISPR nuclease polypeptide can comprise (e.g., consist of) a nickase mutation at position H845 (e.g., an H845A mutation) and an arginine and / or lysine substitution at position I857 (e.g., an I857R substitution) relative to SEQ ID NO: 1. In other examples, the CRISPR nuclease polypeptide can comprise or further comprise one or more mutations that result in reduced PAM recognition stringency (e.g., one or more positions of L1117, D1144, G1227, E1228, A1332, R1345, and / or T1347 of SEQ ID NO: 1, and optionally one or more positions of D61, A68, and H494 of SEQ ID NO: 1).

[0023] Any of the CRISPR nuclease polypeptides provided herein can comprise an amino acid sequence that is at least 90% identical to SEQ ID NO: 1. In some examples, the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 95% identical to SEQ ID NO: 1. In still other examples, the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 98% identical to SEQ ID NO: 1.

[0024] The RT polypeptide is Moloney murine leukemia virus (MMLV)-RT, or a variant thereof. In one example, the MMLV-RT comprises the amino acid sequence of SEQ ID NO: 53.

[0025] In some embodiments, the gene editing systems disclosed herein include a fusion polypeptide. Alternatively, the system includes a first nucleic acid encoding the fusion polypeptide. In some examples, the first nucleic acid is located on a vector, which is optionally a viral vector. In other examples, the first nucleic acid is a first messenger RNA (mRNA).

[0026] In any of the gene editing systems disclosed herein, the spacer sequence in the gRNA of (b) can be 15 to 30 nucleotides in length. In one example, the spacer sequence may be 15 to 20 nucleotides in length. In one specific example, the spacer sequence may be about 17 nucleotides in length.

[0027] In some embodiments, the scaffold sequence includes a nucleotide sequence that is at least 85% identical to SEQ ID NO: 2. In one example, the scaffold sequence includes the nucleotide sequence of SEQ ID NO: 2.

[0028] In some embodiments, the PBS in the RT donor RNA of the RNA molecule is 5 to 50 nucleotides long. In some examples, the PBS is 5 to 20 nucleotides long. In specific examples, the PBS is 7 to 17 nucleotides long. In some embodiments, the PBS binds to a PBS target site that is adjacent to or overlaps with the target sequence. For example, the PBS target site is adjacent to or overlaps with the target sequence. In other examples, the PBS target site is adjacent to the 5' side of the PAM.

[0029] In some embodiments, the template sequence in the RT donor RNA of the RNA molecule can be 5 to 100 nucleotides long. In some examples, the template sequence can be 15 to 25 nucleotides long. In some embodiments, the template sequence in the RT donor RNA is homologous to the genomic site of interest and contains one or more nucleotide variations relative to the genomic site of interest. In some examples, at least one nucleotide variation may be located within the target sequence. Alternatively, at least one nucleotide variation may be located within the PAM.

[0030] In some embodiments, any of the RNA molecules of (b) provided herein may further include a 3' end extension. In some examples, the RNA molecule may further include a 5' end protection fragment, a 3' end protection fragment, or both, each of which forms a secondary structure that is optionally a hairpin, a cyclic, a pseudoknot, or a triple-stranded structure.

[0031] In some examples, the RNA molecule in (b) includes a spacer sequence, a scaffold sequence, a template sequence, and PBS from the 5' end to the 3' end. In other examples, the RNA molecule may include a spacer sequence, a scaffold sequence, a template sequence, PBS, and a 3' extension from the 5' end to the 3' end.

[0032] In some embodiments, the gene editing systems disclosed herein may include a gRNA, an RT donor RNA, and an RNA molecule comprising, optionally, one or more additional elements disclosed herein. Alternatively, the gene editing system may include a nucleic acid encoding the RNA molecule. In some examples, the nucleic acid is located on a vector, which is optionally a viral vector.

[0033] In some embodiments, the gene editing systems disclosed herein may comprise one or more lipid nanoparticles (LNPs) associated with one or more of elements (a) to (b). Alternatively, the gene editing system may comprise one or more viral vectors, for example, one or more adeno-associated virus (AAV) vectors that optionally encode one or more of elements (a) to (b).

[0034] Furthermore, pharmaceutical compositions comprising any of the gene editing systems provided herein, as well as kits comprising elements (a) and (b) of the gene editing systems disclosed herein, are also provided herein.

[0035] In other embodiments, the disclosure features a gene editing method comprising delivering a gene editing system disclosed herein to a host cell to edit a genomic site targeted by the gRNA of the gene editing system. In some embodiments, the host cell is cultured in vitro. In other embodiments, the host cell is located at a target requiring gene editing.

[0036] Furthermore, this disclosure provides a fusion polypeptide comprising one of the CRISPR nuclease polypeptides described herein and one of the reverse transcriptase polypeptides described herein. Such a fusion polypeptide may comprise the amino acid sequence of SEQ ID NO: 55 or 57.

[0037] In addition, this disclosure provides nucleic acids encoding the fusion polypeptides disclosed herein. Such nucleic acids may comprise the nucleotide sequence of SEQ ID NO: 54 or 56. In some examples, the nucleic acid is a vector, such as an expression vector.

[0038] Details of one or more embodiments of the present invention are described below. Other features or advantages of the present invention will be apparent from the following drawings and detailed descriptions of some embodiments, as well as from the appended claims.

[0039] The following drawings form part of this specification and are included to further demonstrate certain aspects of the present disclosure, which can be better understood by referring to the drawings in conjunction with the detailed description of the specific embodiments presented herein. [Brief explanation of the drawing]

[0040] [Figure 1] This figure shows the gene-editing efficacy of reference CRISPR nuclease SEQ ID NO: 1 against exemplary target genes AAVS1, EMX1, and VEGFA. [Figure 2A] Includes gel images showing quantification of nuclease activity. These are gel images captured using a 700 nm channel, showing in vitro cleavage of the target strand of the target DNA substrate (labeled with IR700 dye at the 5' end) by a reference CRISPR nuclease, putative HNH knockout nickasase, or putative RuvC knockout nickasase. [Figure 2B] Includes gel images showing quantification of nuclease activity. These are gel images captured using an 800 nm channel, showing in vitro cleavage of the target strand of the target DNA substrate (labeled with IR800 dye at the 5' end) by a reference CRISPR nuclease, putative HNH knockout nickasase, or putative RuvC knockout nickasase. [Figure 2C] Includes gel images showing quantification of nuclease activity. These are overlay images captured using the 700 nm and 800 nm channels in Figures 2A and 2B. [Figure 2D] Includes gel images showing quantification of nuclease activity. Quantification of percentages of cleaved target and non-target DNA produced by the reference CRISPR nuclease, putative HNH knockout nickasase, and putative RuvC knockout nickasase tested in Example 4. [Figure 3A]Includes a figure showing the efficiency of reverse transcription-mediated gene editing using a reference CRISPR nuclease-reverse transcriptase fusion polypeptide or a CRISPR nickerse variant-reverse transcriptase fusion polypeptide. Percentage of NGS reads containing indels (black bars) and edits encoded by the editing template RNA (gray bars) using a reference CRISPR nuclease-reverse transcriptase fusion polypeptide. [Figure 3B] Includes a figure showing the efficiency of reverse transcription-mediated gene editing using a reference CRISPR nuclease-reverse transcriptase fusion polypeptide or a CRISPR nickasase variant-reverse transcriptase fusion polypeptide. Percentage of NGS reads containing indels (black bars) and edits encoded by the editing template RNA (gray bars) using a CRISPR nickasase variant-reverse transcriptase fusion polypeptide. [Figure 4A] Tables 16 and 17 include figures illustrating the gene editing efficiency of CRISPR nuclease-reverse transcriptase fusion polypeptides with various configurations. The gene editing efficiency constructs of fusion polypeptides containing the CRISPR nuclease of SEQ ID NO: 1 correspond to those listed in Table 17, except that the nickase sequence within them is replaced with the reference nuclease of SEQ ID NO: 1. [Figure 4B] Tables 16 and 17 include figures showing the gene editing efficiencies of CRISPR nuclease-reverse transcriptase fusion polypeptides with various configurations. Table 17 lists the gene editing efficiencies of the fusion polypeptides, each containing CRISPR niccasse SEQ ID NO: 32. [Figure 5] This figure shows the gene editing efficiency of exemplary CRISPR nicksase-reverse transcriptase fusion polypeptides listed in Table 17, indicated by the percentage of eGFP-positive cells relative to mCherry-positive cells. [Figure 6] Table 17 shows the gene editing efficiency of the exemplary CRISPR-nickase-reverse transcriptase fusion polypeptides listed for each exemplary CRISPR-nickase-reverse transcriptase fusion polypeptide, as a percentage of eGFP-positive cells measured for each exemplary CRISPR-nickase-reverse transcriptase fusion polypeptide, in relation to the percentage of eGFP-positive cells measured for each exemplary CRISPR-nickase-reverse transcriptase fusion polypeptide. [Figure 7] This figure shows the gene editing efficiency at the EMX1_T2 site by the exemplary CRISPR-nickase-reverse transcriptase fusion polypeptides listed in Table 17, relative to the percentage of eGFP-positive cells measured for each exemplary CRISPR-nickase-reverse transcriptase fusion polypeptide. [Modes for carrying out the invention]

[0041] This disclosure provides a gene editing system comprising both a CRISPR nuclease polypeptide and a reverse transcriptase (RT) polypeptide, as well as a guide RNA that directs gene editing at a desired genomic site, and an RT donor RNA that serves as an RNA template for the RT polypeptide that synthesizes a DNA strand carrying the desired base substitution. In some embodiments, the gene editing system provided herein may comprise a fusion polypeptide comprising a CRISPR nuclease fragment and an RT fragment, or a nucleic acid encoding the fusion polypeptide. Alternatively or in addition, the gene editing system may comprise a single RNA molecule comprising the guide RNA and the RT donor RNA, or a nucleic acid encoding a single RNA molecule.

[0042] The gene editing systems provided herein have demonstrated successful nucleotide substitution at target sites. See the following examples. Such gene editing systems are effective in introducing desired nucleotide substitutions at target gene sites, thereby achieving desired therapeutic effects (e.g., correction of gene defects). The gene editing systems provided herein can also be used in other fields, such as animal and plant reproduction and genome function research.

[0043] I. RT-CRISPR-mediated gene editing system In some embodiments, RT-CRISPR-mediated gene editing systems are provided herein, comprising at least two protein components, namely a CRISPR nuclease polypeptide and an RT polypeptide, and at least two RNA components, namely a guide RNA and an RT donor RNA. In specific embodiments, the two protein components may be located on a fusion polypeptide. Alternatively, or in addition, the two RNA components may be located on a single RNA molecule. In some cases, the gene editing system may include protein components and / or RNA components. In other cases, the gene editing system may include nucleic acids encoding protein components and / or nucleic acids encoding RNA components.

[0044] In some cases, RT-CRISPR-mediated gene editing systems include CRISPR nucleases having nickase activity as disclosed herein. Such gene editing systems are expected to achieve precise gene editing at desired genomic target sites.

[0045] A. Protein components The gene editing systems provided herein involve at least two enzymes, a CRISPR nuclease and an RT. In some embodiments, the gene editing system comprises two enzymes. In specific examples, the gene editing system may comprise a fusion polypeptide comprising two enzyme components. Alternatively, the gene editing system may comprise one or more nucleic acids encoding components of the two enzymes. For example, the gene editing system may comprise one or more expression vectors (e.g., viral vectors, such as retroviral vectors, adenovirus vectors, or adeno-associated virus vectors) capable of expressing a CRISPR nuclease, an RT, or a fusion polypeptide comprising them. In other examples, the gene editing system may comprise one or more mRNA molecules encoding a CRISPR nuclease, an RT, or a fusion polypeptide comprising them.

[0046] In some embodiments, CRISPR nuclease polypeptides and RT polypeptides as disclosed herein may form complexes that are heterodimers of two protein components via a dimerizing domain (e.g., a leucine zipper), an antibody, a nanobody, or an aptamer.

[0047] (i) CRISPR nuclease polypeptide The CRISPR nuclease polypeptides for use in the gene editing systems disclosed herein may be a CRISPR nuclease (reference CRISPR nuclease) containing the amino acid sequence of SEQ ID NO: 1, or a variant thereof.

[0048] As used herein, the term “CRISPR nuclease” refers to an RNA-guided effector capable of binding to nucleic acids and introducing single-strand or double-strand breaks. CRISPR nucleases typically comprise multiple functional domains, such as nuclease domains (e.g., RuvC and HNH), bridge helix (BH) domains, nucleic acid recognition (REC) domains, phosphate-locked loop (PLL) domains, wedge domains (WED), PAM interaction domains (PID), or combinations thereof. As used herein, the term “domain” refers to a characteristic functional and / or structural unit of a polypeptide. In some cases, functional domains may be linear. In other cases, functional domains may be discontinuous and stereostructural. In some embodiments, a domain may comprise a conserved amino acid sequence across different CRISPR nucleases.

[0049] The reference CRISPR nuclease of SEQ ID NO: 1 (see Table 1 below) is a CRISPR nuclease containing both a RuvC nuclease domain (located at residues 1-59, 722-771, and 927-1101 of SEQ ID NO: 1) and an HNH domain (located at residues 772-926 of SEQ ID NO: 1). The RuvC and HNH nuclease domains coordinate DNA strand cleavage adjacent to a 5'-NDR-3' PAM motif, where N represents any nucleotide, D represents A, G, or T, and R represents G or A. In some examples, PAM is 5'-NRG-3' or 5'-NRR'3', where N and R are as defined herein. In one specific example, PAM is 5'-NGG-3', where N represents any nucleotide. Positions D10, E763, and D991 are considered to be active sites in the RuvC domain, and positions D844, H845, and N868 are considered to be active sites in the HNH domain. In addition to the nuclease domain, the reference CRISPR nuclease of SEQ ID NO: 1 also contains a BH domain (residues 60-93 of SEQ ID NO: 1), a REC domain (residues 94-721 of SEQ ID NO: 1), a PLL domain (residues 1102-1148 of SEQ ID NO: 1), a WED domain (residues 1149-1208 of SEQ ID NO: 1), and a PID domain (residues 1209-1378 of SEQ ID NO: 1).

[0050] In some embodiments, the gene editing systems disclosed herein include variant CRISPR nuclease polypeptides derived from SEQ ID NO: 1, for example, variant CRISPR nuclease polypeptides comprising one or more arginine or lysine substitutions, one or more mutations in one of the nuclease domains, for example, in the HNH nuclease domain or the RuvC nuclease domain, or a combination thereof, relative to the reference CRISPR nuclease polypeptide of SEQ ID NO: 1. Such protein components may form complexes with gRNAs in the same gene editing system.

[0051] Variants of CRISPR nuclease polypeptides In some embodiments, the CRISPR nuclease polypeptides in the gene editing systems provided herein are variants of the reference CRISPR nuclease of SEQ ID NO: 1, for example, by introducing one or more mutations into the reference CRISPR nuclease to modulate (e.g., enhance or reduce) one or more activities of the nuclease. As used herein, the term “variant CRISPR nuclease polypeptide” refers to a CRISPR nuclease polypeptide that, compared to the reference CRISPR nuclease (SEQ ID NO: 1), contains changes, e.g., substitutions, insertions, deletions, and / or fusions at one or more residue positions.

[0052] A variant CRISPR nuclease polypeptide may contain one or more mutations (e.g., arginine substitutions) compared to the reference CRISPR nuclease. Alternatively or in addition, a variant CRISPR nuclease polypeptide may contain one or more mutations in either the RuvC nuclease domain or the HNH nuclease domain. Such mutations may reduce or eliminate the nuclease activity of either the RuvC or HNH nuclease domain, resulting in a variant exhibiting nickase activity. A variant CRISPR nuclease polypeptide may share high sequence homology (e.g., at least 85% sequence identity) with respect to the reference CRISPR nuclease.

[0053] As used herein, the term “nickase” refers to an enzyme that cleaves one strand of a double-stranded DNA at a specific recognition nucleotide sequence (e.g., a target sequence disclosed herein). A nickase may interact with one strand of a DNA double helix to produce a DNA molecule that is cleaved (also known as nicked) on one strand. In some embodiments, the nickase is a variant of a CRISPR nuclease containing an inactivated HNH domain. In some embodiments, the nickase is a variant of a CRISPR nuclease containing an inactivated RuvC domain. In other embodiments, the variant CRISPR nuclease polypeptide may contain one or more mutations that result in reduced PAM recognition stringency compared to SEQ ID NO: 1, compared to a counterpart CRISPR nuclease polypeptide that does not have such mutations. The variant CRISPR nuclease polypeptide may share high sequence homology (e.g., at least 85% sequence identity) with respect to a reference CRISPR nuclease.

[0054] The variant CRISPR nuclease polypeptides provided herein are expected to have advantageous characteristics compared to reference CRISPR nucleases, such as exhibiting nickas activity and / or higher nuclease activity. Therefore, the variant CRISPR nuclease polypeptides disclosed herein are expected to exhibit improved gene editing, such as higher efficacy and precision in gene editing involving strand replacement, compared to reference CRISPR nucleases.

[0055] In some embodiments, the variant CRISPR nuclease polypeptides provided herein comprise one or more mutations in either the RuvC nuclease domain or the HNH nuclease domain (e.g., in the HNH nuclease domain) to reduce or eliminate nuclease activity compared to reference CRISPR nuclease SEQ ID NO: 1, and / or comprise one or more arginine and / or lysine substitutions to improve nuclease characteristics suitable for use in gene editing.

[0056] The variant CRISPR nuclease polypeptides provided herein are expected to exhibit one or more regulated activities (e.g., enhanced or reduced) relative to a reference CRISPR nuclease. As used herein, the term “activity” refers to biological activity. In some embodiments, activity includes enzymatic activity, e.g., catalytic capacity of an effector. For example, activity may include nuclease activity. In some embodiments, activity may include nickasing activity. For example, a variant CRISPR nuclease polypeptide may substantially cleave only one strand of a target DNA double helix.

[0057] In some embodiments, the activity includes binding activity, e.g., binding of an effector (e.g., a CRISPR nuclease) to an RNA guide and / or target nucleic acid. In some examples, the variant CRISPR nuclease polypeptides disclosed herein have enhanced binding to the corresponding guide RNA (gRNA) compared to a reference CRISPR nuclease, for example, having binding activity that is at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 2x, 2x, 5x, 10x, or more greater than that of the reference CRISPR nuclease. The corresponding gRNA refers to a gRNA that has a scaffold recognizable by a CRISPR nuclease.

[0058] In some cases, the variant CRISPR nuclease polypeptides disclosed herein have enhanced enzymatic activity relative to a reference CRISPR nuclease, for example, having at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 2x, 2x, 5x, 10x, or more greater enzymatic activity than that of the reference CRISPR nuclease. In other cases, the variant CRISPR nuclease polypeptides disclosed herein have reduced enzymatic activity relative to a reference CRISPR nuclease (e.g., enzymatic activity for cleaving both strands of a target DNA double helix), for example, having at least 20%, 30%, 40%, 50%, 60%, or 70% lower enzymatic activity than that of the reference CRISPR nuclease. In some cases, the reduced enzymatic activity is achieved by reducing or decreasing the nuclease activity of the RuvC domain. In other cases, reduced enzyme activity is achieved by decreasing or lowering the nuclease activity of the HNH domain.

[0059] In some cases, the variant CRISPR nuclease polypeptides disclosed herein have enhanced indel activity compared to a reference CRISPR nuclease. As used herein, the term “indel activity” refers to the ability of a CRISPR nuclease to introduce indels (insertions / deletions) into a sequence (e.g., a genomic target). For example, in some embodiments, a CRISPR nuclease introduces a double-strand break into a sequence (e.g., a genomic target within a cell), and an indel is created via a DNA repair mechanism.

[0060] In some embodiments, the variant CRISPR nuclease polypeptides provided herein share high sequence homology with a reference CRISPR nuclease. For example, a variant CRISPR nuclease polypeptide may contain at least 70% (e.g., at least 80%, 85%, 90%, 95%, or more) the same amino acid sequence as SEQ ID NO: 1. In some cases, a variant CRISPR nuclease polypeptide may contain at least 90% the same amino acid sequence as SEQ ID NO: 1. In some cases, a variant CRISPR nuclease polypeptide may contain at least 95% the same amino acid sequence as SEQ ID NO: 1. In other cases, a variant CRISPR nuclease polypeptide may contain at least 97% (e.g., 98%, 99%, 99.5%, or more) the same amino acid sequence as SEQ ID NO: 1.

[0061] The "percent identity" (also known as sequence identity) of two nucleic acids or two amino acid sequences is determined using the algorithm of Karlin and Altschul Proc.Natl.Acad.Sci.USA 87:2264-68, 1990, as modified in Karlin and Altschul Proc.Natl.Acad.Sci.USA 90:5873-77, 1993. Such an algorithm is incorporated into the NBLAST and XBLAST programs (version 2.0) of Altschul, et al. J.Mol.Biol.215:403-10, 1990. BLAST nucleotide search can be performed using the NBLAST program, score=100, word length=12 to obtain nucleotide sequences homologous to the nucleic acid molecule of the present invention. BLAST protein search can be performed using the XBLAST program, score=50, word length=3 to obtain amino acid sequences homologous to the protein molecule of the present invention. If a gap exists between two sequences, gap BLAST can be used as described in Altschul et al., Nucleic Acids Res. 25(17):3389-3402, 1997. When using the BLAST and Gapped BLAST programs, the initial settings parameters for each program (e.g., XBLAST and NBLAST) can be used.

[0062] The variant CRISPR nuclease polypeptides provided herein may contain one or more modifications to the reference CRISPR nuclease of SEQ ID NO: 1, such as one or more amino acid residue substitutions, one or more deletions, one or more insertions, fusions, or combinations thereof. In some cases, the modifications may be introduced within the BH domain, PLL domain, WED domain, PID domain, or combinations thereof.

[0063] In some embodiments, the variant CRISPR nuclease polypeptides provided herein may contain one or more arginine substitutions relative to SEQ ID NO: 1. “Arginine substitution” or “lysine substitution” refers to replacing a non-arginine or non-lysine residue in SEQ ID NO: 1 with an arginine or lysine residue. In some examples, the variant CRISPR nuclease polypeptide may contain up to 20 arginine and / or lysine substitutions, e.g., up to 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or two arginine and / or lysine substitutions. In specific examples, the variant CRISPR nuclease polypeptide may contain 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or two arginine and / or lysine substitutions.

[0064] In some cases, one or more of the substituted arginine residues may be replaced by conserved amino acid residues, such as lysine or histidine. In some embodiments, the variant CRISPR nuclease polypeptides provided herein may comprise one or more arginine substitutions, one or more lysine substitutions, or a combination thereof.

[0065] In some cases, arginine substitutions can be located in the BH domain, PLL domain, WED domain, PID domain, or any combination thereof. In some cases, variant CRISPR nuclease polypeptides have one or more arginine and / or lysine substitutions at the positions I79, E331, Y348, S473, F501, I581, D720, A730, G731, Q741, V752, M753, Q809, Q840, Q849, S872, S898, E982, K918, D985, Y986, Y1015, E1037, K1091, S1094, It may contain one or more of the following: P1096, N1099, T1104, E1105, I1106, T1108, L1117, K1131, I1147, E1179, M1205, P1208E1214, A1226, Q1230, A1236, P1238, F1241, L1281, D1284, F1285, A1292, N1295, K1298, G1329, A1333, K1344, S1348, Q1360, and I1370. In specific examples, arginine and / or lysine substitutions can be located at one or more of the following positions in SEQ ID NO: I857R, K736, L784, N813, Q812, I857, and A919.

[0066] In some specific examples, the variant CRISPR nuclease polypeptides for SEQ ID NO: I79R, E331R, Y348R, S473R, F501R, I581R, D720R, A730R, G731R, Q741R, V752R, M753R, Q809R, Q840R, Q849R, S872R, S898R, E982R, K918R, D985R, Y986R, Y1015R, E1037R, K1091R, S1094R, P1096R, N1099R, It may contain one or more arginine substitutions of T1104R, E1105R, I1106R, T1108R, L1117R, K1131R, I1147R, E1179R, M1205R, P1208, E1214R, A1226R, Q1230R, A1236R, P1238R, F1241R, L1281R, D1284R, F1285R, A1292R, N1295R, K1298R, G1329R, A1333R, K1344R, S1348R, Q1360R, and I1370R.

[0067] In some cases, variant CRISPR nuclease polypeptides may contain one or more arginine substitutions at one or more of the positions described above. Examples include I857R, N813R, L784R, K736R, A919R, Q812R, or combinations thereof. In other cases, variant CRISPR nuclease polypeptides may contain one or more lysine substitutions at one or more of the positions described above. Examples include I857K, N813K, L784K, A919K, Q812K, or combinations thereof.

[0068] In other examples, variant CRISPR nuclease polypeptides use arginine and / or lysine substitution combinations as shown in SEQ ID NO: I79, E331, Y348, S473, F501, I581, D720, A730, G731, Q741, V752, M753, Q809, Q840, Q849, S872, S898, E982, K918, D985, Y986, Y1015, E1037, K1091, S1094, May be contained in P1096, N1099, T1104, E1105, I1106, T1108, L1117, K1131, I1147, E1179, M1205, P1208E1214, A1226, Q1230, A1236, P1238, F1241, L1281, D1284, F1285, A1292, N1295, K1298, G1329, A1333, K1344, S1348, Q1360, and / or I1370. In specific examples, variant CRISPR nuclease polypeptides may contain combinations of arginine and / or lysine substitutions (e.g., combinations of arginine substitutions) in I857R, K736, L784, N813, Q812, I857, and / or A919 of SEQ ID NO: 1.

[0069] In specific examples, the CRISPR nuclease polypeptide produced may contain arginine and / or lysine substitutions at the following positions relative to SEQ ID NO: 1: (a) I857, L784, and K736; (b) I857, A919, and K736; (c) I857, N813, and L784; (d) I857, L784, and A919; (e) I857, N813, and K736; (f) I857 and N813; (g) L784, A919, and K736; (h) I857 and L784; or (i) I857 and A919. In some cases, the CRISPR nuclease polypeptide produced may contain arginine substitutions at any of the positional combinations in SEQ ID NO: 1. In one specific example, the resulting CRISPR nuclease polypeptide may contain the arginine substitutions I857R, L784R, and K736R relative to SEQ ID NO: 1. Other examples of arginine and / or lysine substitutions can be seen in Table 4 below.

[0070] Alternatively, or in addition, the variant CRISPR nuclease polypeptides provided herein may contain one or more mutations within either the RuvC or HNH nuclease domain to reduce or eliminate the nuclease activity of a target domain, thereby producing a variant with nickase activity. Such mutations (nickase mutations) may be deletions, insertions, amino acid substitutions, or combinations thereof. In some embodiments, the mutation within either the RuvC or HNH nuclease domain is an amino acid substitution, and the substituted amino acid residue of the amino acid substitution is not a conserved substitution of the native amino acid residue at the site of the mutation. For example, if the native amino acid residue is R, the substituted residue may be any amino acid residue except K. Similarly, if the native amino acid residue is K, the substituted residue may be any amino acid residue except R. Bases for conserved amino acid residue substitutions are provided herein.

[0071] In some cases, one or more nickase mutations may occur within the HNH nuclease domain, for example, at D844, H845, and / or N868 in SEQ ID NO: 1. In some cases, the mutation may be an amino acid residue substitution, and the native amino acid residue in SEQ ID NO: 1 may be replaced by an amino acid residue of a different type than the native residue. For example, a positively charged residue may be replaced by an uncharged amino acid residue, or vice versa. In some cases, the amino acid residue substitution at D844 may be D844G, D844A, D844L, or D844S. In one specific example, the mutation may be D844A. In another example, the amino acid residue substitution at H845 may be H845G, H845A, H845I, H845L, H845M, H845V, or H845S. In one specific example, the mutation at position H845 may be H845A. Alternatively, or in addition, amino acid residue substitutions may occur at position N868, for example, N868G, N868A, N868L, or N868S. In one example, the mutation at position N868 is N868A.

[0072] In some cases, one or more mutations may be associated with the RuvC nuclease domain in SEQ ID NO: 1, for example, at positions D10, E763, D991, or a combination thereof (e.g., positions E763 and / or D991). In some cases, the mutations may be amino acid residue substitutions, and native amino acid residues in SEQ ID NO: 1 may be replaced by amino acid residues of a different type than the native residues. For example, a positively charged residue may be replaced by an uncharged amino acid residue, or vice versa. In some cases, the amino acid residue substitution at D10 may be D10G, D10A, D10L, or D10S. In some cases, the amino acid residue substitution at E763 may be E763G, E763A, E765L, or E763S. Alternatively or in addition, the amino acid residue substitution at D991 may be D991G, D991A, D991L, or D991S.

[0073] In some examples, the variant CRISPR nuclease polypeptides provided herein may be nickase variants, which include one or more mutations in a single nuclease domain (e.g., an HNH nuclease domain at position H845, e.g., H845A). Exemplary nickase variants are provided in Table 5 below. Such nickase variants may further include one or more arginine or lysine substitutions (e.g., arginine substitutions) to enhance certain features, such as indel activity. For example, exemplary arginine and / or lysine substitutions at one or more of the positions I857R, K736, L784, N813, Q812, I857, and A919 of SEQ ID NO: 1 (e.g., arginine substitutions at positions I857, L784, and K736) are provided herein. In some examples, a CRISPR nuclease polypeptide may contain (or consist of) a nickase mutation at position H845 (e.g., an H845A mutation) and an arginine and / or lysine substitution at position I857 (e.g., an I857R substitution) relative to SEQ ID NO: 1.

[0074] In some cases, the variant CRISPR nuclease polypeptides disclosed herein exhibit enhanced double-stranded nuclease activity. Examples of such CRISPR nuclease polypeptides are provided in Table 19 below, each of which is within the scope of this disclosure.

[0075] In some embodiments, the variant CRISPR nuclease polypeptides disclosed herein may contain one or more mutations that reduce the PAM recognition stringency compared to a reference CRISPR nuclease. In some examples, one or more mutations that reduce the PAM recognition stringency may be located at positions L1117, D1144, S1145, G1227, E1228, S1327, A1332, R1343, R1345, and / or T1347 of SEQ ID NO: 1. In some examples, one or more mutations may include (i) one or more arginine and / or lysine substitutions, optionally arginine substitutions, at positions L1117, G1227, S1327, A1332, and / or T1347 of SEQ ID NO: 1; (ii) one or more amino acid substitutions at positions D1144, S1145, E1228, R1343, and / or R1345 of SEQ ID NO: 1; or (iii) a combination of (i) and (ii).

[0076] In some cases, such variant CRISPR nuclease polypeptides exhibit mutations (e.g., amino acid residue substitutions) at positions D61 (e.g., D61R or D61K), L1117 (e.g., L1117R or L1117K), D1144 (e.g., D1144V, D1144A, D1144G, or D1144S), S1145 (e.g., S1145W, S1145Y, or S1145F), G1227 (e.g., G1227K or G1227R), E 1228 (e.g., E1228F, E1228Y, or E1228W), A1327 (e.g., A1327R, or A1332K), A1332 (e.g., A1332R, or A1332K), R1345 (e.g., R1345Q, or R1345N), R1345 (e.g., R1345V, R1345A, R1345G, or R1345S, or R1345Q, R1345N), T1347 (e.g., T1347R, or T1347K), or combinations thereof may be contained.

[0077] In a specific example, one or more amino acid substitutions in (ii) may include D1144L, S1145W, E1228Q, R1343P, R1345V, and / or R1345Q for SEQ ID NO: 1. Alternatively, the substituted amino acid residues at one or more of positions D1114, S1145, E1228, R1343, and R1345 may be conserved substitutions of L, W, Q, P, V, and Q, respectively. For example, a substitution at position D1114 may be D1114M, D1114I, or D1114V. A substitution at position S1145 may be S1145F or S1145Y. A substitution at position E1228 may be E1228N. The substitution at position R1345 may be R1345M, R1345I, R1345L, or R1345N.

[0078] Specific examples of such variant CRISPR nuclease polypeptides are provided in Table 20, each of which is within the scope of this disclosure.

[0079] In some specific examples, the CRISPR nuclease polypeptide produced contains (or consists of) L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, and T1347R relative to SEQ ID NO: 1. Such a CRISPR nuclease polypeptide has the amino acid sequence described in SEQ ID NO: 236.

[0080] In some examples, the CRISPR nuclease polypeptides produced herein (e.g., those having mutations at positions L1117, D1144, G1227, E1228, A1332, R1343, R1345, and / or T1347) have the same properties as SEQ ID NO: 236, with the same properties as D61, A68, H494, L64, S410, T67, Q849, G1110, F501, T659, L784, Y516, G55, E1037, N57, D720, A919, A1294, Q812, N700, H657, T73, Q899, I857, K751, D3 27, I581, D462, E331, A589, D471, I699, N1295, T470, I1147, E130, S473, A353, K40, K334, A60, S1348, K3 67, A1118, K31, Q349, K341, Q83, K585, Q840, G660, K527, G727, Y42, L1281, L122, Q123, T1108, E41, K113 1, K30, S872, I1206, D1132, K460, L80, E459, K1182, M696, K918, K126, N721, Q809, K1091, K736, K783, N4 This may further include one or more single arginine substitutions (e.g., a single arginine substitution) in 98, K723, H1119, F463, L594, D472, K744, E365, G595, K45, Y348, K964, S1181, N813, D407, S839, Y658, E586, G754, A730, Y1015, D903, A1333, S461, or H1359.

[0081] In some specific examples, the CRISPR nuclease polypeptide produced contains L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and A68R relative to SEQ ID NO: 1. In other specific examples, the CRISPR nuclease polypeptide produced contains L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and D61R relative to SEQ ID NO: 1. In yet another specific example, the CRISPR nuclease polypeptide produced contains L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and H494R relative to SEQ ID NO: 1.

[0082] In some cases, variant CRISPR nuclease polypeptides exhibiting less stringent recognition of PAM sequences may recognize 5'-NGN-3', 5'-NRN-3', or 5'-NYN-3' PAM sequences, where N represents any nucleotide, R represents A or G, and Y represents C or T.

[0083] In some cases, a variant CRISPR nuclease polypeptide may contain one or more of the mutations disclosed herein (e.g., one or more arginine and / or lysine substitutions, one or more nickase mutations, and / or one or more mutations resulting in less PAM recognition stringency). Any of the variant CRISPR nuclease polypeptides disclosed herein may share at least 90% (e.g., 95%, 97%, 98%, 99%, 99.5%, or more) sequence identity with SEQ ID NO: 1.

[0084] In some cases, variant CRISPR nuclease polypeptides may contain one or more conserved amino acid residue substitutions, in addition to mutations in the HNH or RuvC nuclease domain, arginine / lysine substitutions, and / or mutations that result in reduced PAM recognition stringency.

[0085] As used herein, “conservative amino acid substitution” refers to an amino acid substitution that does not alter the relative charge or size properties of the protein in which the substitution is made. Variants can be prepared by methods for altering polypeptide sequences known to those skilled in the art, for example, as found in references summarizing such methods, e.g., Molecular Cloning: A Laboratory Manual, J. Sambrook, et al., eds., Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, 1989, or Current Protocols in Molecular Biology, FMAusubel, et al., eds., John Wiley & Sons, Inc., New York. Conservative amino acid substitutions include substitutions made with amino acids in the following groups: (a) M, I, L, V; (b) F, Y, W; (c) K, R, H; (d) A, G; (e) S, T; (f) Q, N; and (g) E, D.

[0086] Exemplary CRISPR nuclease polypeptides for use in the gene editing systems provided herein are disclosed in Tables 1, 5, 19, and 20 below, each of which is within the scope of this disclosure.

[0087] In some embodiments, the CRISPR nuclease polypeptide in the gene editing systems disclosed herein (e.g., the reference nuclease of SEQ ID NO: 1 or any of its variants as disclosed herein) may be a fusion polypeptide comprising a CRISPR nuclease and one or more additional functional moieties. As used herein, the terms “fusion” and “fused” refer to the linkage of at least two nucleotides or protein molecules. For example, “fusion” and “fused” may refer to the linkage of at least two polypeptide domains that are naturally encoded by distinct genes. The fusion may be an N-terminal fusion, a C-terminal fusion, or an intramolecular fusion. In some embodiments, the domains are transcribed and translated to produce a single polypeptide.

[0088] Exemplary additional functional segments that may be included within the fusion polypeptide include peptide tags, fluorescent proteins, base editing domains, DNA methylation domains, histone residue modification domains, localization factors, transcription modifiers, photo-gate regulators, chemoinducible factors, chromatin visualization factors, or combinations thereof.

[0089] In some embodiments, additional functional components may include a nuclear localization signal (NLS), a nuclear export signal (NES), or a combination thereof. In some examples, the fusion polypeptide may contain an NLS, which may be located at either the N-terminus or the C-terminus. In specific examples, the fusion polypeptide may contain a first NLS located at the N-terminus and a second NLS located at the C-terminus. The first and second NLS fragments may be identical. Alternatively, the two NLS fragments may be different. In some embodiments, the fusion polypeptide may contain the NLS near the N-terminus and / or C-terminus (e.g., within about one, two, three, four, or five amino acids from the first or last amino acid of the CRISPR nuclease). In some embodiments, the fusion polypeptide may contain the NLS within the mobile loop of the CRISPR nuclease.

[0090] In some embodiments, the additional functional component may be a flexible peptide linker, such as an XTEN peptide linker or a G / S rich peptide linker. An example of such a peptide linker is provided in Example 1 below.

[0091] In some embodiments, the gene editing systems provided herein may comprise a CRISPR nuclease polypeptide, which may form a ribonucleoprotein (RNP) complex with a corresponding guide RNA. As used herein, the term “complex” refers to a grouping of two or more molecules. In some embodiments, the complex comprises polypeptides and nucleic acid molecules that interact with each other (e.g., bound, in contact, or adhered).

[0092] In other embodiments, the gene editing systems provided herein may include nucleic acids encoding CRISPR nuclease polypeptides. In some examples, the nucleotide sequences encoding CRISPR nuclease polypeptides described herein can be codon-optimized for use in specific host cells or organisms. For example, nucleic acids can be codon-optimized for any non-human eukaryote, including mice, rats, rabbits, dogs, livestock, or non-human primates. Codon usage tables are readily available, for example, in the "Codon Usage Database" available on the worldwide website kazusa.orjp / codon / , and these tables can be adapted in several formats. See Nakamura et al. Nucl. Acids Res. 28:292 (2000), which is incorporated herein by reference in its entirety. Computer algorithms for codon-optimizing specific sequences for expression in specific host cells, such as Gene Forge (Aptagen, Jacobus, PA), are also available. In some examples, the nucleic acids encoding the CRISPR nuclease polypeptides disclosed herein may be mRNA molecules that can be codon-optimized. Exemplary codon-optimized nucleotide sequences encoding exemplary CRISPR nuclease polypeptides can be found in Tables 1 and 5 below, any of which are within the scope of this disclosure.

[0093] In some cases, gene editing systems may include vectors encoding CRISPR nuclease polypeptides (e.g., viral vectors, such as AAV vectors, AdV vectors, or retroviral vectors).

[0094] (ii) Reverse transcriptase polypeptide The gene editing systems disclosed herein also include reverse transcriptase (RT) polypeptides, which may be wild-type RT or variants thereof. In some cases, the RT polypeptides and CRISPR nuclease polypeptides disclosed herein may form fusion proteins. As used herein, the terms “reverse transcriptase” and “RT” typically refer to a multifunctional enzyme having three enzymatic activities, including RNA and DNA-dependent DNA polymerization activity, as well as RNase H activity that catalyzes the cleavage of RNA in RNA-DNA hybrids. Reverse transcriptase can generate DNA from template RNA.

[0095] In some embodiments, the reverse transcriptase polypeptide is any wild-type reverse transcriptase, which is obtained from any naturally occurring organism or virus, or from a commercially available or uncommercial source. The reverse transcriptase polypeptide may also be a variant reverse transcriptase polypeptide.

[0096] Reverse transcriptase polypeptides can be obtained from several different sources. For example, the gene may be obtained from eukaryotic cells infected with a retrovirus, or from a plasmid containing either a portion or the whole of a retroviral genome. In addition, RNA containing the reverse transcriptase gene can be obtained from a retrovirus. In some embodiments, the reverse transcriptase is provided either when expressed or otherwise as individual components, i.e., not as a fusion protein with the CRISPR nuclease polypeptide provided herein.

[0097] Reverse transcriptases are known in the art and include, but are not limited to, Moloney's mouse leukemia virus (MMLV) reverse transcriptase, human immunodeficiency virus (HIV) reverse transcriptase, and avian sarcoma leukemia virus (ASLV) reverse transcriptase. ASLV reverse transcriptase includes, but is not limited to, Roussarcoma virus (RSV) reverse transcriptase, avian myeloblastosis virus (AMV) reverse transcriptase, avian erythroblastosis virus (AEV) helper virus MCAV reverse transcriptase, avian myelocytosis virus MC29 helper virus MCAV reverse transcriptase, avian reticuloendotheliopathy virus (REV-T) helper virus REV-A reverse transcriptase, avian sarcoma virus UR2 helper virus UR2AV reverse transcriptase, avian sarcoma virus Y73 helper virus YAV reverse transcriptase, Rouss-associated virus (RAV) reverse transcriptase, and myeloblastosis-associated virus (MAV) reverse transcriptase. Those skilled in the art will recognize that these reverse transcriptases can be suitably used in the compositions described herein.

[0098] In some embodiments, the reverse transcriptase is MMLV-RT, marathon RT from Eubacterium rectale, or RTX reverse transcriptase, or a variant of MMLV-RT, marathon RT, or RTX reverse transcriptase. In some embodiments, the reverse transcriptase is the sequence shown in Table 10, its variant, or its orthologue. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4]

[0099] In some embodiments, the reverse transcriptase polypeptide is an “error-prone” reverse transcriptase variant. Error-prone reverse transcriptases known and / or available in the art may be used. It will be understood that none of the reverse transcriptases naturally possess proofreading activity, and therefore the error rate of the reverse transcriptase is generally higher than that of DNA polymerases that contain proofreading activity. In some embodiments, a reverse transcriptase is considered “error-prone” if it has an error rate of less than one error in the approximately 15,000 nucleotides synthesized.

[0100] In some embodiments, the reverse transcriptase polypeptide has one or more mutations in the RNase H domain. In some embodiments, the reverse transcriptase polypeptide does not contain the RNase H domain (e.g., the RNase H domain is removed from the reverse transcriptase polypeptide). In some embodiments, the RNase H domain is cleaved in the reverse transcriptase polypeptide. In some embodiments, the reverse transcriptase polypeptide has one or more mutations in the RNA-dependent DNA polymerase domain. In some embodiments, the reverse transcriptase polypeptide is a variant with modified thermal stability characteristics. The ability of reverse transcriptase to withstand high temperatures is an important aspect of cDNA synthesis. Elevated reaction temperatures promote the denaturation of RNA with strong secondary structures and / or high GC content, allowing the reverse transcriptase to read the entire sequence. As a result, reverse transcription at higher temperatures enables full-length cDNA synthesis and higher yields. Wild-type M-MLV reverse transcriptase typically has an optimal temperature range of 37–48°C, but mutations can be introduced to enable reverse transcription activity at higher temperatures, including above 48°C, 49°C, 50°C, 51°C, 52°C, 53°C, 54°C, 55°C, 56°C, 57°C, 58°C, 59°C, 60°C, 61°C, 62°C, 63°C, 64°C, 65°C, and above 66°C.

[0101] The variant reverse transcriptase polypeptides used herein may be at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to any reference reverse transcriptase polypeptide, and the reference reverse transcriptase polypeptide includes any wild-type reverse transcriptase, mutant reverse transcriptase, or reverse transcriptase fragment or other reverse transcriptase variant disclosed or intended herein or known in the art. In some embodiments, the reverse transcriptase variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or up to 100, or up to 200, or up to 300, or up to 400, or up to 500 or more amino acid changes compared to the reference reverse transcriptase.In some embodiments, the reverse transcriptase variant comprises a fragment of a reference reverse transcriptase, wherein the fragment is identical to the corresponding fragment of the reference reverse transcriptase by at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9%.

[0102] Variant reverse transcriptases, including error-prone reverse transcriptases, thermostable reverse transcriptases, and reverse transcriptases with increased processing capacity, can be produced by a variety of conventional strategies, including mutagenesis or evolutionary processes. In some cases, variants can be produced by introducing a single mutation. In other cases, variants may require two or more mutations. For these variants containing two or more mutations, the effect of a given mutation can be evaluated by introducing the identified variant into the wild-type gene by site-directed mutagenesis, isolated from other mutations derived from the particular variant. Screening assays of the single variant thus produced then allow for the determination of the effect of this mutation alone.

[0103] In some embodiments, the reverse transcriptase polypeptide contains or is fused to a domain for improving the elongation rate and / or efficiency of the reverse transcriptase. In some embodiments, the reverse transcriptase polypeptide is fused to an Sso7d polypeptide, e.g., an Sso7d polypeptide from Sulfolobus solfataricus. See, for example, Wang et al., Nucleic Acids Res. 32(3):1197-207 (2004).

[0104] In some embodiments, the reverse transcriptase in any one of the embodiments described herein interacts with a ligase, integrase, and / or recombinase. In some embodiments, the reverse transcriptase in any one of the embodiments described herein is fused to a ligase, integrase, and / or recombinase. In some embodiments, the ligase, integrase, and / or recombinase is fused to the N-terminus or C-terminus of the reverse transcriptase. In some embodiments, the ligase, integrase, and / or recombinase is fused internally to the reverse transcriptase. In some embodiments, the integrase is serine integrase. In some embodiments, the integrase is Bxb1, TP901, or PhiBT1 integrase. In some embodiments, the recombinase is serine recombinase or tyrosine recombinase. In some embodiments, the recombinase is CRE recombinase. In some embodiments, reverse transcriptases that interact with or are fused to ligases, integrases, and / or recombinases further interact with or are fused to CRISPR nuclease polypeptides disclosed herein.

[0105] In other embodiments, the gene editing systems provided herein may include nucleic acids encoding RT polypeptides. In some examples, the nucleotide sequences encoding RT polypeptides described herein can be codon-optimized for use in specific host cells or organisms. In some examples, the nucleic acids encoding RT polypeptides may be mRNA molecules, which can be codon-optimized. Exemplary codon-optimized nucleotide sequences encoding exemplary RT polypeptides can be found in Table 7 below, which are within the scope of this disclosure. In some examples, the gene editing system may include vectors encoding RT polypeptides (e.g., viral vectors, e.g., AAV vectors, AdV vectors, or retroviral vectors).

[0106] (iii) Fusion polypeptide In some embodiments, the gene editing systems provided herein include a fusion polypeptide comprising both CRISPR nuclease polypeptides and RT polypeptides, as similarly disclosed herein. Alternatively, the gene editing system may include a nucleic acid encoding the fusion polypeptide (e.g., a vector such as an expression vector).

[0107] As used herein, the terms “fused” and “fused” refer to the linkage of at least two nucleotides or protein molecules. For example, “fused” and “fused” refer to a linkage of separate genes in nature (e.g., provided herein). This can refer to the fusion of at least two polypeptide domains encoded by CRISPR nuclease polypeptides and reverse transcriptase polypeptides. The fusion can be an N-terminal fusion, a C-terminal fusion, or an intramolecular fusion.

[0108] In some embodiments, the fusion polypeptide may comprise a reverse transcriptase polypeptide at its N-terminus and a CRISPR nuclease polypeptide downstream of the RT polypeptide. In other embodiments, the fusion polypeptide may comprise a CRISPR nuclease polypeptide at its N-terminus and an RT polypeptide downstream of the CRISPR nuclease polypeptide. In some embodiments, the RT polypeptide may fuse with the CRISPR nuclease polypeptide at an intramolecular position within the RT polypeptide, for example, the CRISPR nuclease polypeptide may be within the loop of the reverse transcriptase polypeptide.

[0109] Any of the CRISPR nuclease polypeptides and any of the RT polypeptides disclosed herein may be used to construct a fusion polypeptide. In some cases, the CRISPR nuclease polypeptide may be a CRISPR nuclease such as SEQ ID NO: 1. Alternatively, the CRISPR nuclease polypeptide may be a variant of the reference CRISPR nuclease of SEQ ID NO: 1. For example, the variant may be a nickase having one of the above-described mutations in the HNH domain relative to the reference CRISPR nuclease, such as the nickase of SEQ ID NO: 32. In other cases, the variant may include a combination of mutation(s) in the HNH domain and one or more arginine and / or lysine substitutions, such as those disclosed herein.

[0110] In some embodiments, any of the CRISPR nuclease-RT fusion polypeptides disclosed herein may include one or more additional functional elements, e.g., those provided herein. In some cases, the additional functional elements may be one or more NLS elements. In some examples, the fusion polypeptide may contain NLS at its N-terminus, its C-terminus, or both. Alternatively or in addition, the additional functional element may be a flexible peptide linker, which can be located between the CRISPR nuclease polypeptide and the RT polypeptide. Preferred peptide linkers include, but are not limited to, G / S-rich peptide linkers and XTEN peptide linkers. Examples of NLS and peptide linkers are provided in Tables 1, 5, and 15 below. See also Examples 1 and 7.

[0111] In some examples, the CRISPR nuclease-RT fusion polypeptides provided herein include a peptide linker (first peptide linker) located between the CRISPR nuclease polypeptide and the RT polypeptide. In some examples, the CRISPR nuclease polypeptide is N-terminus relative to the RT polypeptide. In other examples, the CRISPR nuclease polypeptide is C-terminus relative to the RT polypeptide. In some examples, an additional peptide linker and / or one or more NLS signals may be located between the CRISPR nuclease polypeptide and the RT polypeptide. For example, the additional peptide linker and NLS may be located between the CRISPR nuclease polypeptide and the RT polypeptide in addition to the first peptide linker. In some specific examples, the peptide linker between the CRISPR nuclease polypeptide and the RT polypeptide is, for example, at least 20-aa in length, ranging from about 20 to 100 amino acids.

[0112] Alternatively or in addition, the CRISPR nuclease-RT fusion polypeptides provided herein may comprise at least two NLSs (a first NLS and a second NLS), at least one of which is located at the N-terminus or C-terminus of the fusion polypeptide. In some examples, one of the two NLSs is located at the N-terminus and the other at the C-terminus. In other examples, both NLSs are located at the N-terminus. In yet another example, both NLSs are located at the C-terminus.

[0113] In some cases, the CRISPR nuclease-RT fusion polypeptides provided herein may include one or more additional NLSs (e.g., a third NLS and optionally a fourth NLS). Such additional NLS(s) may be located between the CRISPR nuclease polypeptide and the RT nuclease. In other cases, additional NLS(s) may be located between the CRISPR nuclease / RT polypeptide and the terminal NLS, optionally via a peptide linker.

[0114] In some examples, the CRISPR nuclease-RT fusion polypeptides disclosed herein may comprise, from N-terminus to C-terminus, a first NLS, a CRISPR nuclease, a peptide linker, an RT polypeptide, and a second NLS. In other examples, the CRISPR nuclease-RT polypeptides disclosed herein may comprise, from N-terminus to C-terminus, a first NLS, an RT nuclease, a peptide linker, a CRISPR nuclease polypeptide, and a second NLS (which may be identical to the first NLS), a CRISPR nuclease polypeptide, a first peptide linker, a second NLS (which may be identical to the first NLS), a second peptide linker, an RT nuclease, a third peptide linker (which may be identical to the first peptide linker), a third NLS (which may be identical to the first NLS), a fourth peptide linker, and a fourth NLS. Examples of such CRISPR nuclease-RT polypeptides are provided in Table 8 below, and all of them are within the scope of this disclosure.

[0115] In other instances, the CRISPR nuclease-RT fusion polypeptides disclosed herein may have the configurations listed in Table 16 below (from N-terminus to C-terminus). In some cases, the CRISPR nuclease-RT fusion polypeptide does not contain the FLAG motif in any of the configurations listed in Table 16. In one instance, the CRISPR nuclease-RT fusion polypeptide has the configuration corresponding to configuration 4, except that the FLAG motif is removed. In another instance, the CRISPR nuclease-RT fusion polypeptide has the configuration corresponding to configuration 5, except that the FLAG motif is removed. In yet another instance, the CRISPR nuclease-RT fusion polypeptide has the configuration corresponding to configuration 7, except that the FLAG motif is removed. In yet another instance, the CRISPR nuclease-RT fusion polypeptide has the configuration corresponding to configuration 9, except that the FLAG motif is removed. Alternatively, the CRISPR nuclease-RT fusion polypeptide has the configuration corresponding to configuration 10, except that the FLAG motif is removed. Additional exemplary CRISPR nuclease-RT fusion polypeptides are provided in Table 17. Variants of these fusion polypeptides with the FLAG motif removed are also within the scope of this disclosure.

[0116] In some cases, any of the CRISPR nuclease-RT fusion polypeptides may contain a CRISPR nuclease variant. In some specific cases, the CRISPR variant is a nickase, e.g., one provided herein (including, e.g., mutations as listed in Table 5). In some cases, the nickase contains, for example, one or more mutations in the HNH domain of the reference nuclease of SEQ ID NO: 1, located at position H845, e.g., H845A.

[0117] In other embodiments, the gene editing systems provided herein may include nucleic acids encoding a CRISPR nuclease-RT fusion polypeptide. In some examples, the nucleotide sequences encoding the fusion polypeptides described herein can be codon-optimized for use in specific host cells or organisms. In some examples, the nucleic acids encoding the fusion polypeptides disclosed herein may be mRNA molecules, which can be codon-optimized. Exemplary codon-optimized nucleotide sequences encoding exemplary CRISPR nuclease-RT fusion polypeptides can be found in Tables 8 and 17 below, any of which are within the scope of this disclosure. See also Examples 5 and 7. Coded sequences of variants of exemplary CRISPR nuclease-RT fusion polypeptides with the FLAG motif removed are also within the scope of this disclosure.

[0118] In some cases, gene editing systems may include vectors encoding CRISPR nuclease-RT fusion polypeptides (e.g., viral vectors such as AAV vectors, AdV vectors, or retroviral vectors).

[0119] (iv) Preparation of protein components The CRISPR nuclease polypeptides, RT polypeptides, or CRISPR nuclease-RT fusion polypeptides disclosed herein may be prepared by conventional methods or by methods disclosed herein. For example, CRISPR nuclease polypeptides, RT polypeptides, or CRISPR nuclease-RT fusion polypeptides may be prepared by culturing host cells capable of producing nuclease polypeptides, such as bacterial or mammalian cells, isolating the nuclease polypeptides thus produced, and optionally purifying the nuclease polypeptides. CRISPR nuclease polypeptides, RT polypeptides, or CRISPR nuclease-RT fusion polypeptides may also be prepared by in vitro coupling transcription-translation systems.

[0120] The host cells that can be used for the preparation of CRISPR nuclease polypeptides, RT polypeptides, or CRISPR nuclease-RT fusion polypeptides are not particularly limited, as long as they are capable of producing CRISPR nuclease polypeptides, RT polypeptides, or CRISPR nuclease-RT fusion polypeptides. Some non-limiting examples of host cells include bacterial cells (e.g., E. coli cells), yeast cells, insect cells, or mammalian cells.

[0121] vector This disclosure provides vectors for expressing CRISPR nuclease polypeptides, RT polypeptides, or CRISPR nuclease-RT fusion polypeptides. In some embodiments, the vectors disclosed herein include a nucleotide sequence encoding a CRISPR nuclease polypeptide, an RT polypeptide, or a CRISPR nuclease-RT fusion polypeptide. In some embodiments, the vector includes a Pol II promoter or a Pol III promoter.

[0122] The expression of native or synthetic polynucleotides is typically achieved by operably ligating a polynucleotide encoding a CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide to a promoter and incorporating the construct into an expression vector. The expression vector is not particularly limited as long as it contains a polynucleotide encoding a CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide, and may be suitable for replication and incorporation in eukaryotic cells.

[0123] A typical expression vector contains transcription and translation terminators, start sequences, and promoters useful for the expression of a desired polynucleotide. For example, plasmid vectors containing recognition sequences for RNA polymerase (e.g., pSP64, pBluescript) can be used. Vectors containing retroviruses, such as those derived from lentiviruses, are suitable tools for achieving long-term gene transfer, as they allow for the long-term stable integration of the transgene and its proliferation in daughter cells. Examples of vectors include expression vectors, replication vectors, probe-generating vectors, and sequencing vectors. Expression vectors can be supplied to cells in the form of viral vectors.

[0124] Viral vector technology is well-known in the art and is described in various virology and molecular biology manuals. Useful viruses as vectors include, but are not limited to, phage viruses, retroviruses, adenoviruses, adeno-associated viruses, herpesviruses, and lentiviruses. Generally, a suitable vector contains a functional replication origin, promoter sequence, convenient restriction endonuclease site, and one or more selectable markers in at least one organism.

[0125] The type of vector is not particularly limited, and any vector that can be expressed in host cells can be appropriately selected. More specifically, depending on the type of host cell, a promoter sequence is appropriately selected to ensure the expression of a polypeptide(s) from a polynucleotide, and this promoter sequence and polynucleotide are inserted into one of various plasmids or other materials for preparing the expression vector.

[0126] Additional promoter elements, such as enhancing sequences, regulate the frequency of transcription initiation. Typically, these are located 30–110 bp upstream of the start site, although some promoters have recently been shown to also contain functional elements downstream of the start site. Depending on the promoter, individual elements appear to be able to activate transcription either cooperatively or independently.

[0127] Furthermore, this disclosure should not be limited to the use of constitutive promoters. Inducible promoters are also intended as part of this disclosure. The use of inducible promoters provides a molecular switch that can turn on the expression of a polynucleotide sequence that is operatively ligated when such expression is desired, or turn off the expression when expression is not desired. Examples of inducible promoters include, but are not limited to, metallothione promoters, glucocorticoid promoters, progesterone promoters, and tetracycline promoters.

[0128] The introduced expression vector may also contain either a selectable marker gene or a reporter gene, or both, to facilitate the identification and selection of expressing cells from a population of cells intended for transfection or infection via the viral vector. In other embodiments, the selectable marker may be supported on a separate DNA fragment and used in a simultaneous transfection procedure. Both the selectable marker and reporter gene may be flanked by appropriate transcriptional regulatory sequences to enable expression in host cells. Examples of such markers include dihydrofolate reductase genes and neomycin resistance genes for eukaryotic cell culture, and tetracycline resistance genes and ampicillin resistance genes for E. coli and other bacterial cultures. The use of such selectable markers allows for confirmation that the polynucleotide encoding the polypeptide(s) of the present invention has been transferred into host cells and subsequently expressed without failure.

[0129] The preparation methods using recombinant expression vectors are not particularly limited and include methods using plasmids, phages, or cosmids.

[0130] Method of expression This disclosure includes methods for protein expression, comprising translating a CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide as described herein.

[0131] In some embodiments, the host cells described herein are used to express CRISPR nuclease polypeptides, RT polypeptides, or CRISPR nuclease-RT fusion polypeptides. The host cells are not particularly limited, and a variety of known cells can be preferably used. Specific examples of host cells include bacteria, e.g., E. coli, yeast (budding yeast Saccharomyces cerevisiae and fission yeast Schizosaccharomyces pombe), nematodes (Caenorhabditis elegans), Xenopus laevis oocytes, and animal cells (e.g., CHO cells, COS cells, and HEK293 cells). The method for transferring the expression vector into the host cells, i.e., the transformation method, is not particularly limited, and known methods, e.g., electroporation, calcium phosphate methods, liposome methods, and DEAE dextran methods can be used.

[0132] After the host has been transformed with an expression vector, host cells can be cultured, cultivated, or propagated for the production of CRISPR nuclease polypeptides, RT polypeptides, or CRISPR nuclease-RT fusion polypeptides. After expression, host cells can be collected, and the CRISPR nuclease polypeptides, RT polypeptides, or CRISPR nuclease-RT fusion polypeptides can be purified from the culture or other source according to conventional methods (e.g., filtration, centrifugation, cell disruption, gel filtration chromatography, ion exchange chromatography, etc.).

[0133] Various methods can be used to determine the level of production of mature CRISPR nuclease polypeptides, mature RT polypeptides, or mature CRISPR nuclease-RT fusion polypeptides in host cells. Such methods include, but are not limited to, the use of protein-specific polyclonal or monoclonal antibodies, or labeling tags as described elsewhere herein. Exemplary methods include, but are not limited to, enzyme-linked immunosorbent assays (ELISA), radioimmunoassays (RIA), fluorescence immunoassays (FIA), and fluorescence-activated cell sorting (FACS). These and other assays are well known in the art (see, for example, Maddox et al., J.Exp.Med.158:1211

[1983] ).

[0134] This disclosure provides a method for in vivo expression of a CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide (and optionally, gRNA and / or RT donor RNA in a gene editing system disclosed herein). Such a method may include providing a polyribonucleotide encoding a CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide to a host cell in a subject (e.g., a human subject), the polyribonucleotide encoding the CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide, and causing the CRISPR nuclease polypeptide, RT polypeptide, or CRISPR nuclease-RT fusion polypeptide to be expressed from the cell.

[0135] B.RNA components The gene editing systems provided herein also comprise at least two RNA components, a guide RNA (gRNA) that directs gene editing at a desired gene site, and an RT donor RNA that functions as an RNA template for an RT polypeptide in reverse transcription. The RT donor RNA contains a desired nucleotide substitution to be inserted into the gene site of interest. In some embodiments, the gene editing system comprises two RNA molecules. In a specific example, the gene editing system comprises a single RNA molecule containing the gRNA and the RT donor RNA. Alternatively, the gene editing system may comprise one or more nucleic acids encoding the two RNA components. For example, the gene editing system may comprise one or more expression vectors (e.g., viral vectors such as retroviral vectors, adenovirus vectors, or adeno-associated virus vectors) capable of producing the gRNA, the RT donor RNA, or a single RNA molecule containing them.

[0136] In some embodiments, the gRNA and RT donor RNA disclosed herein may form a complex.

[0137] (i) Guide RNA The gene editing systems disclosed herein further comprise one or more gRNAs or encoding nucleic acids. As used herein, the terms “RNA guide,” “RNA guide sequence,” or “guide RNA (gRNA)” refer to an RNA molecule or modified RNA molecule that facilitates the CRISPR nuclease described herein to target a genomic site of interest. For example, an RNA guide may be a molecule comprising a spacer sequence and a scaffold sequence. The spacer sequence recognizes (e.g., binds to) a site on a non-PAM strand that is complementary to the target sequence on the PAM strand, for example, a site on a non-PAM strand designed to be complementary to a specific nucleic acid sequence. The scaffold sequence contains a nuclease-binding sequence for binding to the CRISPR nuclease. In some embodiments, the scaffold is an RNA sequence.

[0138] In some cases, the gRNAs disclosed herein may further include a linker sequence, a 5' end and / or a 3' end protection fragment, or a combination thereof.

[0139] Spacer array As used herein, the terms “spacer” and “spacer sequence” (also known as DNA-binding sequence) refer to a portion (DNA sequence) of an RNA guide that is the RNA equivalent of the target sequence. A spacer contains a sequence that can bind to the non-PAM strand via base pairing at a site complementary to the target sequence (which is located on the PAM strand). Such spacers are also known to be specific to the target sequence. In some cases, a spacer may be at least 75% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99%) identical to the target sequence, except for the RNA-DNA sequence difference. In some cases, a spacer may be 100% identical to the target sequence, except for the RNA-DNA sequence difference.

[0140] The gene editing systems disclosed herein comprise one or more gRNAs, each comprising a spacer sequence for targeting a genomic site of interest and a scaffold recognizable by a variant CRISPR nuclease polypeptide contained in the gene editing system. The target sequence can be adjacent to a 5'-NDR-3' protospacer facilitation motif (PAM), where N represents any nucleotide, D represents A, G, or T, and R represents G or A. In some examples, the PAM is 5'-NRR-3', where N represents any nucleotide and R represents A or G. For example, the PAM could be 5'-NRG-3', where N represents any nucleotide and R represents A or G. In a specific example, the PAM motif is 5'-NGG-3', where N represents any nucleotide. In some cases, PAM sequences recognizable by CRISPR nuclease polypeptides are 5'-NGN-3', 5'-NRN-3', or 5'-NYN-3', where N represents any nucleotide, R represents A or G, and Y represents C or T.

[0141] The PAM motif is located on the 3' end of the target sequence. As used herein, the terms “protospacer-adjacent motif” or “PAM sequence” refer to a DNA sequence adjacent to the target sequence. In some embodiments, the PAM sequence is required for CRISPR nuclease binding and / or indel activity. In a double-stranded DNA molecule, the strand containing the PAM motif is referred to as the “PAM strand,” and the complementary strand is referred to as the “non-PAM strand.” The gRNA binds to a site on the non-PAM strand that is complementary to the target sequence disclosed herein, and the PAM sequence described herein is present on the PAM strand. The PAM motif may be located upstream of the target sequence.

[0142] As used herein, the term “adjacent to” means that a nucleotide or amino acid sequence is in close proximity to another nucleotide or amino acid sequence. In some embodiments, if there are no nucleotides separating the two sequences, the nucleotide sequences are adjacent to (i.e., directly adjacent to) another nucleotide sequence. In some embodiments, a nucleotide sequence is adjacent to another nucleotide sequence, but it is when a small number of nucleotides (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides) separate the two sequences.

[0143] Spacers disclosed herein may have a length of about 15 to about 30 nucleotides. For example, a spacer may have a length of about 15 to about 20 nucleotides, about 15 to about 25 nucleotides, about 20 to about 25 nucleotides, or about 20 to about 30 nucleotides. In some embodiments, spacers in gRNA may be designed to generally have a length of 15 to 25 nucleotides (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, and 25) and to be complementary to a specific target sequence. In some embodiments, the spacer sequence may be designed to have a length of 18 to 22 nucleotides (e.g., 20 nucleotides).

[0144] In some embodiments, the spacer sequence may have at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.5% sequence identity with the target sequence described herein, and may bind to a complementary region of the target sequence via base pairing.

[0145] In some embodiments, the spacer sequence contains only RNA bases. In some embodiments, the spacer sequence contains DNA bases (e.g., the spacer contains at least one thymine). In some embodiments, the spacer sequence contains both RNA bases and DNA bases (e.g., the DNA-binding sequence contains at least one thymine and at least one uracil).

[0146] Scaffolding arrangement Scaffold sequences in gRNAs are similarly recognizable by variant CRISPR nuclease polypeptides in gene editing systems. In some cases, the scaffold sequence includes SEQ ID NO: 2, which is the corresponding scaffold for the reference CRISPR nuclease SEQ ID NO: 1. GUUUUAGAGCUGUGCUGAAAAGCACAGCACGUUAAAAUAAGGCAGUGAUUGAAAAAUCCAGUCCGUAUUCAGCUUGAAAAAGUGAGCACCGAAUCGGUGCUU(Sequence ID 2)

[0147] In other instances, the scaffold sequence may be a variant derived from SEQ ID NO: 2. Such a variant scaffold sequence may contain nucleotide sequences identical to SEQ ID NO: 2 by at least 80% (e.g., at least 85%, 90%, 95%, 98%, or more). Alternatively or in addition, the variant scaffold sequence may include deletions, nucleotide substitutions, or combinations thereof. The variant CRISPR nuclease polypeptide may have increased binding to the variant scaffold sequence compared to the scaffold of SEQ ID NO: 2. In some examples, the variant scaffold may be a fragment of SEQ ID NO: 2 disclosed herein or a variant thereof. For example, the variant scaffold for use in gRNA provided herein may have a length in the range of 100 to 150 nucleotides.

[0148] In gRNA, the scaffold may be located at the 3' end of the spacer. In some cases, the scaffold and spacer are directly linked. In other cases, the scaffold and spacer may be linked via a nucleotide linker.

[0149] (ii) RT donor RNA As used herein, the terms “reverse transcription donor RNA” and “RT donor RNA” refer to an RNA molecule containing a reverse transcription template sequence (RTT sequence) and a primer binding site (PBS). RT donor RNA may be fused to an RNA guide at either its 5' or 3' end.

[0150] Any of the RT donor RNAs disclosed herein comprises (i) a primer binding site (PBS) and (ii) an RTT sequence. In some cases, the RT donor RNA may further comprise (iii) a nucleotide linker sequence, (iv) a 5' end and / or 3' end protection fragment (see disclosure herein), or a combination thereof. In some cases, the 5' end or 3' end protection fragment (e.g., 3' extension) may comprise a pseudoknot motif for protection from 3' exonuclease activity.

[0151] In some embodiments, the RT donor RNA includes an aptamer. In some embodiments, the aptamer recruits a reverse transcriptase polypeptide.

[0152] Primer binding site (PBS) In some embodiments, PBS in the RT donor RNA disclosed herein is an RNA sequence capable of binding to a DNA strand via base pairing. The DNA strand is nicked or cleaved by a CRISPR nuclease polypeptide of the gene editing system disclosed herein, or can be nicked or cleaved. In some embodiments, PBS comprises an RNA sequence capable of binding to a DNA strand (PBS target site) via base pairing. The DNA strand may have a free 3' end, or the 3' free end may be generated by cleavage by a CRISPR nuclease contained in the same gene editing system. In some examples, the PBS target site may be located on the same DNA strand as the PAM sequence (PAM strand).

[0153] In some embodiments, PBS can be 5 to 50 nucleotides long. For example, PBS can be about 5 to 40, 5 to 30, or 5 to 20 nucleotides long. In specific examples, PBS can be 5 to 20 (e.g., 7 to 17) nucleotides long. In some examples, PBS may contain 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides.

[0154] As used herein, the term “PBS target site” refers to the region to which PBS binds. The PBS target site may be adjacent to the PAM (e.g., it may be upstream of it). In gene editing systems containing a CRISPR nuclease polypeptide that is a nickase mutation (e.g., containing a disrupted HNH nuclease domain as disclosed herein), PBS in the RT donor RNA may bind to a region on the PAM strand (PBS target site). In other embodiments, the PBS target site may partially or completely overlap with the target sequence. In some cases, the PBS target site may be located upstream of the PAM sequence. For example, the PBS target site may be up to 100 nucleotides upstream of the PAM sequence, e.g., up to 50 nucleotides, up to 30 nucleotides, up to 25 nucleotides, up to 20 nucleotides, up to 15 nucleotides, up to 10 nucleotides, or up to 5 nucleotides upstream of the PAM sequence. In specific examples, the PBS target site may begin approximately 3 to 10 nucleotides upstream of the PAM sequence (i.e., the outermost 5' nucleotide of PBS may bind to approximately 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides upstream of the PAM). In specific examples, the PBS target site may begin 1, 1-2, 1-3, 1-4, or 1-5 nucleotides upstream of the PAM sequence. When a free 3' end is generated in or near the target sequence by a CRISPR nuclease polypeptide in the gene editing system, PBS binding to the PAM chain at a site upstream of the PAM sequence can efficiently facilitate DNA synthesis by the RT polypeptide in the gene editing system, and begins from the free 3' end generated in the PAM chain.

[0155] Reverse transcription template (RTT) sequence A reverse transcription template sequence (RTT sequence) serves as a template for reverse transcription mediated by an RT polypeptide in the gene editing systems disclosed herein. In some embodiments, the RTT sequence comprises a sequence having at least one encoded edit. In some embodiments, the RTT sequence comprises sequence homology with a target sequence or its complementary region having at least one encoded edit. In some embodiments, the RTT sequence is the length of at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 60 nucleotides, at least 70 nucleotides, at least 80 nucleotides, at least 90 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 300 nucleotides, at least 400 nucleotides, or at least 500 nucleotides. In some embodiments, the RTT sequence is approximately 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, or 120 nucleotides in length, or any length in between.

[0156] In some embodiments, the RTT sequence consists of approximately 10 nucleotides. In some embodiments, the RTT sequence consists of approximately 11 nucleotides. In some embodiments, the RTT sequence consists of approximately 12 nucleotides. In some embodiments, the RTT sequence consists of approximately 13 nucleotides. In some embodiments, the RTT sequence consists of approximately 14 nucleotides. In some embodiments, the RTT sequence consists of approximately 15 nucleotides. In some embodiments, the RTT sequence consists of approximately 16 nucleotides. In some embodiments, the RTT sequence consists of approximately 17 nucleotides. In some embodiments, the RTT sequence consists of approximately 18 nucleotides. In some embodiments, the RTT sequence consists of approximately 19 nucleotides. In some embodiments, the RTT sequence consists of approximately 20 nucleotides. In some embodiments, the RTT sequence consists of approximately 21 nucleotides. In some embodiments, the RTT sequence consists of approximately 22 nucleotides. In some embodiments, the RTT sequence consists of approximately 23 nucleotides. In some embodiments, the RTT sequence consists of approximately 24 nucleotides. In some embodiments, the RTT sequence consists of approximately 25 nucleotides. In some embodiments, the RTT sequence consists of approximately 26 nucleotides. In some embodiments, the RTT sequence consists of approximately 27 nucleotides. In some embodiments, the RTT sequence consists of approximately 28 nucleotides. In some embodiments, the RTT sequence consists of approximately 29 nucleotides. In some embodiments, the RTT sequence consists of approximately 30 nucleotides.

[0157] In some embodiments, the reverse transcription template sequence includes at least one coded edit (e.g., at least two) relative to the target sequence. In some embodiments, the at least one coded edit includes at least one substitution, insertion, and / or deletion. In some embodiments, the edit in the target sequence includes substitution, insertion, and / or deletion relative to the sequence of the target sequence. In some embodiments, the reverse transcription template sequence includes at least one LoxP site.

[0158] In some embodiments, the editing can be a single or multiple nucleotide substitution, for example, a G-to-T substitution, a G-to-A substitution, a G-to-C substitution, a T-to-G substitution, a T-to-A substitution, a T-to-C substitution, a C-to-G substitution, a C-to-T substitution, a C-to-A substitution, an A-to-T substitution, an A-to-G substitution, or an A-to-C substitution. In some embodiments, the change in the sequence can be a conversion of a G:C base pair to a T:A base pair, a G:C base pair to an A:T base pair, a G:C base pair to a C:G base pair, a T:A base pair to a G:C base pair, a T:A base pair to an A:T base pair, a C:G base pair to a C:G base pair, a C:G base pair to a G:C base pair, a C:G base pair to a T:A base pair, a C:G base pair to an A:T base pair, a T:A base pair to a G:C base pair, or an A:T base pair to a C:G base pair.

[0159] In some embodiments, the template sequences described herein may further introduce one or more silent mutations. As used herein, a silent mutation refers to a mutation that does not alter the amino acid residue encoded by the codon containing the mutation. The RTT sequence can be transcribed into DNA by the reverse transcriptase of the gene editing system described herein. In some embodiments, the RTT sequence is transcribed into the DNA of the PAM strand from the 5' end to the 3' end.

[0160] In some embodiments, the RTT sequence is on the 5' end of the PBS. In some embodiments, the RTT sequence is on the 3' end of the PBS. In some cases, the PBS and RTT sequences in the RT donor RNA provided herein may be linked via a linker sequence. In some embodiments, the RTT and a terminal protection fragment (e.g., a 3' terminal protection fragment) may be linked via a linker sequence to avoid steric hindrance between the two RNA components.

[0161] (iii) Single RNA molecule In some embodiments, the gene editing systems provided herein include a single RNA molecule containing both gRNA and RT donor RNA, or a nucleic acid encoding a single RNA molecule. Such a single RNA molecule can mediate cleavage at a target sequence within a desired genomic site by a CRISPR nuclease polypeptide, and synthesis of DNA fragments from the free 3' end of a free DNA strand generated by CRISPR nuclease polypeptide cleavage based on the RTT sequence in the single RNA molecule.

[0162] In some embodiments, the single RNA molecule may optionally include an RNA guide bound to the RT donor RNA via a linker. In some examples, the single RNA molecule includes, from 5' to 3', a spacer sequence, a scaffold sequence recognizable by a CRISPR nuclease polypeptide, an RTT sequence, and a PBS. In specific examples, the single RNA molecule may include, from 5' to 3', a spacer sequence, a scaffold sequence, an RTT sequence, a PBS, and a protective fragment.

[0163] Any of the single RNA molecules provided herein may further include a linker that may be located between the scaffold sequence and the RTT or after the PBS. In some examples, the linker may include a hairpin structure. In some examples, the linker may include an aptamer domain.

[0164] In some cases, the 5' and / or 3' ends of a single RNA molecule, or gRNA and / or RT donor RNA, may contain protective fragments that can enhance the RNA molecule's resistance to exonuclease activity. In some cases, the terminal protective fragments may include nucleotide sequences capable of forming secondary structures such as hairpins, circularizations, pseudoknots, or triple-stranded structures. In other cases, the terminal protective fragments may include sequences of exoribonuclease-resistant RNA (xrRNA), transfer RNA (tRNA), or cleaved tRNA. In some embodiments, modifications include Zika-like pseudoknots, mouse leukemia virus pseudoknot (MLV-PK) sequences, red clover necrotic mosaic virus (RCNMV) sequences, sweet clover necrotic mosaic virus (SCNMV) sequences, carnation ring-spot virus (CRSV) sequences, pre-Q1 aptamer sequences, truncated pre-Q1 aptamer sequences, boxB RNA sequences, or RNA bacteriophage MS2 sequences.

[0165] In some examples, the 5' end of a single RNA molecule (or, if a separate RNA molecule is used, the 5' end of gRNA and / or RT donor RNA) may contain a 5' elongation motif that can be one of the protective fragments disclosed herein. In some examples, the 3' end of a single RNA molecule (or, if a separate RNA molecule is used, the 3' end of gRNA and / or RT donor RNA) may contain a 3' elongation motif that can be one of the protective fragments disclosed herein. One specific example of a 3' elongation motif is provided in Example 6 below.

[0166] (iv) Modification of nucleic acids Any RNA component in the gene editing systems disclosed herein, such as a single RNA molecule, gRNA, and / or RT donor RNA, may include one or more modifications. Exemplary modifications may include any modifications to sugars, nucleic acid bases, nucleoside bonds (e.g., binding phosphate / phosphodiester bond / phosphodiester skeleton), and any combination thereof. Some of the exemplary modifications provided herein are described in detail below.

[0167] Any of the RNA components disclosed herein may include any useful modifications, e.g., modifications to sugars, nucleic acid bases, or nucleoside bonds (e.g., binding phosphate / phosphodiester bond / phosphodiester skeleton). One or more atoms of pyrimidine nucleic acid bases may be replaced, or may be replaced, with optionally substituted aminos, optionally substituted thiols, optionally substituted alkyls (e.g., methyl or ethyl), or halos (e.g., chloro or fluoro). One or more atoms of purine nucleic acid bases may be replaced, or may be replaced, with optionally substituted aminos, optionally substituted thiols, optionally substituted alkyls (e.g., methyl or ethyl), or halos (e.g., chloro or fluoro). In certain embodiments, modifications (e.g., one or more modifications) are present in each of the sugar and nucleoside bonds. Modifications may include modifications from ribonucleic acid (RNA) to deoxyribonucleic acid (DNA), threose nucleic acid (TNA), glycol nucleic acid (GNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), or hybrids thereof. Additional modifications are described herein.

[0168] In some embodiments, any of the RNA components in the gene editing systems disclosed herein include a debasic site (i.e., a position without purines or pyrimidines). In some embodiments, the debasic site (also referred to as a deaplynated / deapyrimidine site) is located within the editing template RNA. For example, the debasic site may be located within the RTT of the editing template RNA. In some embodiments, the activity of the reverse transcriptase is terminated at or near the debasic site.

[0169] In some embodiments, modifications may include chemical or cell-inducible modifications. For example, some non-limiting examples of intracellular RNA modifications are described by Lewis and Pan in “RNA modifications and structures cooperate to guide RNA-protein interactions” from Nat Reviews Mol Cell Biol, 2017, 18:202-210.

[0170] Different sugar modifications, nucleotide modifications, and / or nucleoside bonds (e.g., skeletal structures) can be present at various positions in the sequence. Those skilled in the art will understand that nucleotide analogs or other modifications may be located at any position in the sequence such that the function of the sequence is not substantially diminished. The sequence may consist of approximately 1% to 100% modified nucleotides (relative to the total nucleotide content, or to one or more types of nucleotides, i.e., A, G, U, or C), or any intervening percentage (e.g., 1%-20%, 1%-25%, 1%-50%, 1%-60%, 1%-70%, 1%-80%, 1%-90%, 1%-95%, 10%-20%, 10%-25%, 10%-50%, 10%-60%, 10%-70%, 10%-80%, 10%-90%, 10%) It may contain modified nucleotides of ~95%, 10%~100%, 20%~25%, 20%~50%, 20%~60%, 20%~70%, 20%~80%, 20%~90%, 20%~95%, 20%~100%, 50%~60%, 50%~70%, 50%~80%, 50%~90%, 50%~95%, 50%~100%, 70%~80%, 70%~90%, 70%~95%, 70%~100%, 80%~90%, 80%~95%, 80%~100%, 90%~95%, 90%~100%, and 95%~100%.

[0171] In some embodiments, sugar modifications (e.g., at the 2' or 4' position) or sugar substitutions in one or more ribonucleotides of a sequence may include modifications or substitutions of phosphodiester bonds, as well as skeletal modifications. Specific examples of sequences include, but are not limited to, sequences containing a modified skeleton or nucleoside modifications, such as modifications or substitutions of phosphodiester bonds, which are unnatural nucleoside-to-nucleoside bonds. Sequences having a modified skeleton include, among other things, those that do not have a phosphorus atom in the skeleton. For the purposes of this application and as is often referred to in the art, modified RNAs that do not have a phosphorus atom in the nucleoside-to-nucleoside skeleton can also be considered oligonucleosides. In certain embodiments, the sequence comprises ribonucleotides, each having a phosphorus atom in its nucleoside-to-nucleoside skeleton.

[0172] The modified sequence skeleton may include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkyl phosphotriesters, methyl and other alkylphosphonates, such as 3'-alkylene phosphonates and chiral phosphonates, phosphinates, phosphoramides, such as 3'-aminophosphoramides and aminoalkylphosphoramides, thionophosphoramides, thionoalkyl phosphonates, thionoalkyl phosphotriesters, as well as boranophosphates having the usual 3'-5' bond, their 2'-5' bond analogues, and those having reverse polarity with adjacent pairs of nucleoside units bonded from 3'-5' to 5'-3' or 2'-5' to 5'-2'. Also included are various salts, mixed salts, and free acid forms. In some embodiments, the sequences may be negatively or positively charged.

[0173] Modified nucleotides that can be incorporated into a sequence can be modified on nucleoside bonds (e.g., phosphate backbone). In this specification, the terms “phosphate” and “phosphodiester” are used interchangeably in the context of polynucleotide backbones. A phosphate group in a backbone can be modified by substituting one or more oxygen atoms with different substituents. Furthermore, modified nucleosides and nucleotides can include extensive substitutions of the unmodified phosphate moiety at other nucleoside bonds, as described herein. Examples of modified phosphate groups include, but are not limited to, phosphorothioates, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramides, phosphorodiamidates, alkyl or arylphosphonates, and phosphotryesters. Phosphorodithioates have both sulfur-substituted and unbound oxygen atoms. Phosphate linkers can also be modified by substituting bound oxygen with nitrogen (bridged phosphoramide), sulfur (bridged phosphorothioate), and carbon (bridged methylene phosphonate).

[0174] To confer stability to RNA and DNA polymers via non-natural phosphorothioate backbone binding, α-thio-substituted phosphate moieties are provided. Phosphothioate DNA and RNA exhibit increased nuclease resistance and subsequently have a longer half-life in the cellular environment.

[0175] In specific embodiments, the modified nucleoside includes alpha-thio-nucleosides (e.g., 5'-O-(1-thiophosphate)-adenosine, 5'-O-(1-thiophosphate)-cytidine (α-thiocytidine), 5'-O-(1-thiophosphate)-guanosine, 5'-O-(1-thiophosphate)-uridine, or 5'-O-(1-thiophosphate)-psoidouridine).

[0176] Other nucleoside bonds that may be used in the present invention, including nucleoside bonds that do not contain phosphorus atoms, are described herein.

[0177] In some embodiments, the sequence may contain one or more cytotoxic nucleosides. For example, cytotoxic nucleosides may be incorporated into the sequence, such as through bifunctional modification. Cytotoxic nucleosides may include, but are not limited to, adenosine arabinoside, 5-azacitidine, 4'-thio-aracitidine, cyclopentenylcytosine, cladribine, clofarabine, cytarabine, cytosine arabinoside, 1-(2-C-cyano-2-deoxy-beta-D-arabino-pentofuranosyl)-cytosine, decitabine, 5-fluorouracil, fludarabine, phloxuridine, gemcitabine, combinations of tegafur and uracil, tegafur((RS)-5-fluoro-1-(tetrahydrofuran-2-yl)pyrimidine-2,4(1H,3H)-dione), troxacitabine, tezacitabine, 2'-deoxy-2'-methylidencytidine (DMDC), and 6-mercaptopurine. Additional examples include fludarabine phosphate, N4-behenoyl-1-beta-D-arabinofuranosylcytosine, N4-octadecyl-1-beta-D-arabinofuranosylcytosine, N4-palmitoyl-1-(2-C-cyano-2-deoxy-beta-D-arabino-pentofuranosyl)cytosine, and P-4055 (cytarabine 5'-elaidic acid ester).

[0178] In some embodiments, the sequence includes one or more post-transcriptional modifications (e.g., capping, cleavage, polyadenylation, splicing, poly-A sequence, methylation, acylation, phosphorylation, methylation and acetylation of lysine and arginine residues, and nitrosylation of thiol and tyrosine residues). One or more post-transcriptional modifications can be any of the more than 100 different nucleoside modifications identified in RNA (Rozenski, J, Crain, P, and McCloskey, J. (1999). The RNA Modification Database: 1999 update. Nucleoside Acids Res 27: 196-197). In some embodiments, the first isolated nucleic acid includes messenger RNA (mRNA). In some embodiments, mRNA is pyridine-4-onribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-psoiduridine, 2-thio-psoiduridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyluridine, 1-carboxymethyl-psoiduridine, 5-propynyluridine, 1-propynyl-psoiduridine, 5-taurinomethyluridine, 1-taurinomethyl-psoiduridine, 5-taurinomethyl-2-thiouridine, 1-taurinomethyl-4-thiouridine, 5 It comprises at least one nucleoside selected from the group consisting of -methyluridine, 1-methylpsoiduridine, 4-thio-1-methylpsoiduridine, 2-thio-1-methylpsoiduridine, 1-methyl-1-deaz-psoiduridine, 2-thio-1-methyl-1-deaz-psoiduridine, dihydrouridine, dihydropsoiduridine, 2-thio-dihydrouridine, 2-thio-dihydropsoiduridine, 2-methoxyuridine, 2-methoxy-4-thiouridine, 4-methoxypsoiduridine, and 4-methoxy-2-thiopsoiduridine.In some embodiments, mRNA is 5-aza-cytidine, pseudoisocytidine, 3-methylcytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-psoidisocytidine, pyrrolo-cytidine, pyrrolo-psoidisocytidine, 2-thiocytidine, 2-thio-5-methylcytidine, 4-thio-psoidisocytidine, 4-thio-1-methyl-psoidisocytidine, 4- It comprises at least one nucleoside selected from the group consisting of thio-1-methyl-1-deazal-psoidisocytidine, 1-methyl-1-deazal-psoidisocytidine, zebralin, 5-aza-zebralin, 5-methyl-zebralin, 5-aza-2-thio-zebralin, 2-thio-zebralin, 2-methoxy-cytidine, 2-methoxy-5-methylcytidine, 4-methoxy-psoidisocytidine, and 4-methoxy-1-methyl-psoidisocytidine. In some embodiments, mRNA is 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyiso It comprises at least one nucleoside selected from the group consisting of pentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, and 2-methoxy-adenine.In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of inosine, 1-methylinosine, waiosin, waibutosin, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine.

[0179] The sequence may or may not be uniformly modified along the entire length of the molecule. For example, one or more or all types of nucleotides (e.g., spontaneous nucleotides, purines, or pyrimidines, or one or more or all of A, G, U, C, I, pU) may or may not be uniformly modified in the sequence or in a given given sequence region. In some embodiments, the sequence contains pseudouridine. In some embodiments, the sequence contains inosine, which may assist the immune system in characterizing the sequence as endogenous antiviral RNA. Inosine incorporation may also mediate improved RNA stability / reduced degradation. See, for example, Yu, Z. et al. (2015) RNA editing by ADAR1 marks dsRNA as “self”. Cell Res. 25, 1283-1284, which is incorporated in its entirety by reference.

[0180] In some embodiments, any RNA sequence described herein, such as an edited template RNA, may include terminal modifications (e.g., 5'-terminal modifications or 3'-terminal modifications). In some embodiments, the terminal modifications are chemical modifications. In some embodiments, the terminal modifications are structural modifications. See the disclosures herein for further details.

[0181] When the gene editing systems disclosed herein include nucleic acids encoding CRISPR nucleases and / or RT polypeptides, such as mRNA molecules, such nucleic acid molecules may contain any of the modifications disclosed herein, where applicable.

[0182] C. Exemplary gene editing systems The exemplary gene editing systems described herein are intended to be illustrative only. (a) A fusion polypeptide or nucleic acid encoding the same, wherein the fusion polypeptide comprises any of the CRISPR nuclease polypeptides, any of the RT polypeptides, and optionally one or more NLSs, one or more peptide linkers, or a combination thereof, which may be located at the N-terminus and / or C-terminus, (b) A single RNA molecule may include a guide RNA, an RT donor RNA, and optionally one or more nucleotide linkers, one or more 5' or 3' terminal protective elements, or a combination thereof.

[0183] In some embodiments, the CRISPR nuclease polypeptide in the fusion polypeptide may be a nickase variant (e.g., an H845 mutation, such as a nickase variant containing H845A, as provided in Table 5 below). Alternatively or in addition, the CRISPR nuclease polypeptide in the fusion polypeptide may include one or more arginine and / or lysine substitutions at the positions disclosed herein (e.g., K736, L784, Q812, N813, I857, and / or A919 in SEQ ID NO: 1). In some specific examples, the CRISPR nuclease polypeptide in the fusion polypeptide may include combinations of arginine and / or lysine substitutions at the positions provided herein, e.g., I857, L784, and K736. Alternatively, or in addition, the CRISPR nuclease polypeptide in the fusion polypeptide may contain one or more mutations to reduce PAM recognition stringency, for example, at positions D61, A68, H494, L1117, D1144, S1145, G1227, E1228, S1327, A1332, R1343, R1345, and / or T1347 of SEQ ID NO: 1. In some specific examples, such mutations may include (i) one or more arginine and / or lysine substitutions, optionally arginine substitutions, at positions D61, A68, H494, L1117, G1227, S1327, A1332, and / or T1347 of SEQ ID NO: 1; (ii) one or more amino acid substitutions at positions D1144, S1145, E1228, R1343, and / or R1345 of SEQ ID NO: 1; or (iii) a combination of (i) and (ii).

[0184] In some embodiments, the RT polypeptide in the fusion polypeptide may be an MMLV variant, for example, SEQ ID NO: 53 provided in Example 5 below.

[0185] In some examples, the fusion polypeptides provided herein may comprise an N-terminal CRISPR nuclease polypeptide at the N-terminus and a C-terminal RT polypeptide. Optionally, the fusion polypeptide may include a peptide linker (e.g., a G / S rich linker or an XTEN peptide linker) between the CRISPR nuclease polypeptide and the RT polypeptide. Alternatively or in addition, the fusion polypeptide may include NLS at the N-terminus and / or C-terminus. In some examples, the fusion polypeptide may comprise two different NLS, one at the N-terminus and the other at the C-terminus.

[0186] In some examples, the fusion polypeptides provided herein may have any of the configurations disclosed herein, for example, those listed in Table 16 below (e.g., configuration 4, configuration 5, configuration 7, configuration 9, or configuration 10), except that the FLAG motif is removed.

[0187] A single RNA molecule contained in an exemplary gene editing system may contain guide RNA and RT donor RNA in any direction. Optionally, the single RNA molecule may contain one or more nucleotide linkers between the gRNA and the RT donor RNA, and / or between functional domains in the gRNA (e.g., between the spacer and the scaffold sequence) and / or between functional domains in the RT donor RNA (e.g., between the PBS and the RTT sequence). In some examples, the single RNA molecule may further contain protective fragments (e.g., those disclosed herein) at its 5' and / or 3' ends.

[0188] In a specific example, a single RNA molecule contained in an exemplary gene editing system may include a spacer sequence, a scaffold sequence, RTT, PBS, and a 3' extension from the 5' end to the 3' end, which may have pseudoknot motifs.

[0189] In some examples, the exemplary gene editing systems provided herein may include one of the CRISPR-RT fusion polypeptides provided in Table 8 or Table 17 below, or a variant thereof from which the FLAG motif has been removed, and a single RNA molecule provided herein (e.g., comprising a spacer sequence, scaffold sequence, RTT, PBS, and 3' extension from 5' to 3', which may have a pseudoknot motif).

[0190] In other examples, the exemplary gene editing systems provided herein may include any of the CRISPR-RT fusion polypeptides provided in Table 8 or Table 17 below, or variants thereof from which the FLAG motif has been removed, and nucleic acids (vectors such as viral vectors) encoding single RNA molecules provided herein (e.g., including a spacer sequence, scaffold sequence, RTT, PBS, and 3' extension from 5' to 3', which may have a pseudoknot motif).

[0191] In further examples, the exemplary gene editing systems provided herein may include a nucleic acid encoding one of the CRISPR-RT fusion polypeptides provided in Table 8 below, and a single RNA molecule provided herein (e.g., including a spacer sequence, scaffold sequence, RTT, PBS, and 3' extension from 5' to 3', which may have a pseudoknot motif). The nucleic acid may include a codon-optimized coding nucleotide sequence, e.g., those provided in Table 8 or Table 17, or variants thereof with the FLAG motif removed.

[0192] In further examples, the exemplary gene editing systems provided herein may include a nucleic acid encoding one of the CRISPR-RT fusion polypeptides provided in Table 8 or Table 17 below, or a variant thereof with the flag motif removed, and a nucleic acid (e.g., a vector such as a viral vector) encoding a single RNA molecule provided herein (e.g., comprising a spacer sequence, scaffold sequence, RTT, PBS, and 3' extension from 5' to 3', which may have a pseudoknot motif). The nucleic acid encoding the fusion polypeptide may be codon-optimized, for example, those provided in Table 8 or Table 17, or a variant thereof with the FLAG motif removed. In some cases, the gene editing system may include two vectors, one encoding the fusion polypeptide and the other encoding a single RNA molecule. Alternatively, the gene editing system may include a single vector encoding both the fusion polypeptide and the single RNA molecule.

[0193] III. Gene Editing Methods Any of the gene editing systems can be used to genetically modify (edit) a target nucleic acid, which can be a gene site of interest, for example, a gene site where gene editing is needed, for example, to repair a gene mutation, to introduce a protective mutation, or to introduce a modification to regulate gene expression.

[0194] A. Delivery of gene editing systems to target cells Any component of the gene editing systems disclosed herein may be formulated and delivered to cells (e.g., mammalian cells) by known methods, including carriers and / or polymer carriers, e.g., liposomes. Such methods include, but are not limited to, transfection (e.g., lipid-mediated, cationic polymers, calcium phosphate, dendrimers), electroporation or other methods of membrane disruption (e.g., nucleofection), viral delivery (e.g., lentivirus, retrovirus, adenovirus, adeno-associated virus (AAV)), microinjection, microprojectile bombardment ("gene gun"), fugene, direct sonic loading, cell squeezing, phototransfection, protoplast fusion, imparefection, magnetofection, exome-mediated delivery, lipid nanoparticle-mediated delivery, and any combination thereof. In some examples, the delivery method involves the use of lipid nanoparticles to mediate the delivery of one or more components of the gene editing systems disclosed herein.

[0195] In some embodiments, the method involves delivering one or more nucleic acids (e.g., a fusion polypeptide comprising a CRISPR nuclease polypeptide, an RT polypeptide, or both, an RNA guide, an RT donor RNA, or a single RNA molecule comprising both), one or more transcripts thereof, and / or a pre-formed RNA guide / CRISPR nuclease polypeptide / RT polypeptide complex to a cell, wherein a ternary complex is formed in the cell. In some embodiments, the RNA guide and / or RT donor RNA, or their fusion, and the fusion polypeptide comprising RNA encoding a CRISPR nuclease polypeptide or an RT polypeptide, or both, are delivered together in a single composition. In some embodiments, the RNA guide and the RNA encoding the CRISPR nuclease polypeptide are delivered in separate compositions. In some embodiments, the RNA guide / RT donor RNA and the RNA encoding the CRISPR nuclease polypeptide / RT polypeptide delivered in separate compositions are delivered using the same delivery technology. In some embodiments, the RNA guide / RT donor RNA and the RNA encoding the CRISPR nuclease polypeptide / RT polypeptide delivered in separate compositions are delivered using different delivery technologies.

[0196] In some embodiments, one or more protein components and one or more RNA components are delivered together. For example, a CRISPR nuclease and / or RT polypeptide, as well as an RNA guide and / or RT donor RNA, are packaged together within a single AAV particle. In another example, a CRISPR nuclease and / or RT polypeptide, as well as an RNA guide and / or RT donor RNA, are delivered together via lipid nanoparticles (LNPs). In some embodiments, a CRISPR nuclease and / or RT polypeptide, as well as an RNA guide and / or RT donor RNA, are delivered separately. For example, a CRISPR nuclease and / or RT polypeptide, as well as an RNA guide and / or RT donor RNA, are packaged in separate AAV particles. In another example, a CRISPR nuclease and / or RT polypeptide is delivered by a first delivery mechanism, and an RNA guide and / or RT donor RNA is delivered by a second delivery mechanism.

[0197] Exemplary intracellular delivery methods include, but are not limited to, viruses such as AAV or viral-like drugs; chemical-based transfection methods such as those using calcium phosphate, dendrimers, liposomes, or cationic polymers (e.g., DEAE-dextran or polyethyleneimine); non-chemical methods such as microinjection, electroporation, cell squeezing, sonoporation, phototransfection, imparefection, protoplast fusion, bacterial conjugation, plasmid or transposon delivery; particle-based methods such as gene guns, magnetofection or magnetic-assisted transfection, particle collisions; and hybrid methods such as nucleofection. In some embodiments, lipid nanoparticles include mRNA encoding a CRISPR nuclease-RT fusion polypeptide, an edited template RNA, or mRNA encoding it. In some embodiments, the application further provides cells produced by such methods, and organisms (e.g., animals, plants, or fungi) containing or produced from such cells.

[0198] B. Genetically modified cells Any of the gene editing systems disclosed herein can be delivered to various cells (e.g., mammalian cells such as mouse cells, non-human primate cells, or human cells). In some embodiments, the cells are in cell culture or co-culture of two or more cell types. In some embodiments, the cells are ex vivo. In some embodiments, the cells are obtained from a viable organism and maintained in cell culture.

[0199] In some embodiments, the cells are derived from cell lines. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, 293T, MF7, K562, HeLa, CHO, and their transgenic variants. Cell lines are available from various sources known to those skilled in the art (see, for example, the American Type Culture Collection (ATCC) (Manassas, Va.)). In some embodiments, the cells are immortal or immortalized cells. In some embodiments, the cells are primary cells. In some embodiments, the cells are stem cells, e.g., totipotent stem cells (e.g., universally pluripotent), pluripotent stem cells, duplexitic stem cells, oligopotent stem cells, or unipotent stem cells. In some embodiments, the cells are induced pluripotent stem cells (iPSCs) or derived from iPSCs. In some embodiments, the cells are differentiated cells. In some embodiments, the cells are mammalian cells, e.g., human cells or mouse cells. In some embodiments, mouse cells are derived from wild-type mice, immunosuppressed mice, or disease-specific mouse models. In some embodiments, cells are living tissue, organs, or cells within an organism.

[0200] Furthermore, any genetically modified cells produced using any of the gene editing systems disclosed herein are within the scope of this disclosure. Such modified cells may contain disrupted target genes.

[0201] Any of the gene editing systems, compositions containing the same, vectors, nucleic acids, RNA guides, and cells disclosed herein may be used in therapy. The gene editing systems, compositions, vectors, nucleic acids, RNA guides, and cells disclosed herein may be used in methods for treating a disease or condition in a subject. Any preferred delivery or administration method known in the art may be used to deliver the compositions, vectors, nucleic acids, RNA guides, and cells disclosed herein. Such methods may involve contacting a target sequence with the compositions, vectors, nucleic acids, or RNA guides disclosed herein. Such methods may involve methods for editing the target sequence, such as those disclosed herein. In some embodiments, cells prepared using the RNA guides disclosed herein are used for ex vivo gene therapy.

[0202] IV. Therapeutic uses Any gene editing system, such as those disclosed herein, or any modified cells produced using such a gene editing system, may be used to treat diseases associated with a target gene, for example, a genetic defect in the target gene.

[0203] In some embodiments, methods for treating a target disease as disclosed herein, comprising administering one of the gene editing systems disclosed herein to a subject in need of treatment (e.g., a human patient). The gene editing system may be delivered to a specific tissue or a specific cell type in which gene editing is required. The gene editing system may include an LNP comprising one or more of its components, one or more vectors encoding one or more of its components (e.g., viral vectors), or a combination thereof. The components of the gene editing system may be formulated to form a pharmaceutical composition, which may further include one or more pharmaceutically acceptable carriers.

[0204] In some embodiments, modified cells produced using any of the gene editing systems disclosed herein may be administered to subjects in need of treatment (e.g., human patients). Modified cells may include substitutions, insertions, and / or deletions described herein. In some examples, modified cells may include cell lines modified with CRISPR nuclease polypeptides, RT polypeptides, or CRISPR nuclease-RT fusion polypeptides, as well as single RNA molecules including RNA guide and RT donor RNA, or both. In some cases, modified cells may be a heterogeneous population, which includes cells with different types of gene editing. Alternatively, modified cells may include a substantially homogeneous cell population (e.g., at least 80% of the cells in the whole population) containing one specific gene editing in a target gene. In some examples, cells may be suspended in a suitable culture medium.

[0205] In some embodiments, compositions comprising a gene editing system or its components are provided herein. Such compositions may be pharmaceutical compositions. Useful pharmaceutical compositions may be prepared, packaged, or marketed in formulations suitable for oral, rectal, vaginal, parenteral, topical, pulmonary, intranasal, lesional, buccal, ocular, intravenous, intra-organal, or other routes of administration. Pharmaceutical compositions of this disclosure may be prepared, packaged, and / or marketed in bulk as single unit doses or as multiple single unit doses. As used herein, “unit dose” refers to a separate amount of the pharmaceutical composition (e.g., a gene editing system or its components) administered to a subject, or a convenient fraction of such dose, e.g., half or one-third of such dose.

[0206] Formulations of pharmaceutical compositions suitable for parenteral administration may include an activator (e.g., a gene editing system or its components, or modified cells) combined with a pharmaceutically acceptable carrier, such as sterile water or sterile isotonic saline. Such formulations may be prepared, packaged, or sold in forms suitable for bolus or serial administration. Some injectable formulations may be prepared, packaged, or sold in unit dosage forms, for example, in ampoules or multi-dose containers containing preservatives. Some formulations for parenteral administration include, but are not limited to, suspensions, solutions, emulsions in oily or aqueous vehicles, pastes, and embedded sustained-release formulations or biodegradable formulations. Some formulations may further include, but are not limited to, one or more additional components, including suspending agents, stabilizers, or dispersants.

[0207] Pharmaceutical compositions may be in the form of sterile, injectable aqueous or oily suspensions or solutions. These suspensions or solutions may be formulated by known techniques and may contain, in addition to cells, additional components, such as dispersants, wetting agents, or suspending agents as described herein. Such sterile, injectable formulations may be prepared using non-toxic, parenterally acceptable diluents or solvents, such as water or saline. Other acceptable diluents and solvents include, but are not limited to, Ringer's solution, isotonic sodium chloride solution, and fixative oils such as synthetic monoglycerides or diglycerides. Other useful parenterally administered formulations may include those that contain cells in packaged form, within liposome preparations, or as components of biodegradable polymer systems. Some compositions for sustained release or embedding may contain pharmaceutically acceptable polymers or hydrophobic materials, such as emulsions, ion exchange resins, sparingly soluble polymers, or sparingly soluble salts.

[0208] V. Kits and their use This disclosure also provides kits that can be used, for example, to carry out the methods described herein for genetic modification of target genes. In some embodiments, the kit comprises a single RNA molecule containing an RNA guide and / or RT donor RNA, a CRISPR nuclease polypeptide, and an RT polypeptide or a fusion polypeptide thereof. In some embodiments, the kit comprises a single RNA molecule and a CRISPR nuclease-RT fusion polypeptide. In some embodiments, the kit comprises a polynucleotide encoding a CRISPR nuclease polypeptide, an RT polypeptide, or a CRISPR nuclease-RT fusion polypeptide, and optionally, the polynucleotide is contained in a vector, for example, as described herein. In some embodiments, the kit comprises a polynucleotide encoding an RNA component disclosed herein. The CRISPR nuclease polypeptide, RT polypeptide, or a fusion polypeptide thereof (or the polynucleotide encoding them) and the RNA component (e.g., as a ribonucleoprotein) may be packaged in the same or other containers within the kit, or in separate vials or other containers, and their contents may be mixed before use.

[0209] The CRISPR nuclease polypeptides, RT polypeptides, and RNA components can be packaged in the same or other containers within the kit, or in separate vials or other containers, and their contents can be mixed before use. The kit may also optionally include buffers and / or instructions on how to use the RNA components, CRISPR nuclease polypeptides, and RT polypeptides, or their fusion polypeptides.

[0210] General technology The practices described herein will employ conventional techniques within the scope of the art, including molecular biology (including recombinant techniques), microbiology, cell biology, biochemistry, and immunology, unless otherwise specified. Such techniques are described in Molecular Cloning: A Laboratory Manual, second edition (Sambrook, et al., 1989) Cold Spring Harbor Press, Oligonucleotide Synthesis (MJ Gait, ed. 1984), Methods in Molecular Biology, Humana Press, Cell Biology: A Laboratory Notebook (JECellis, ed., 1989) Academic Press, Animal Cell Culture (RIFreshney, ed. 1987), Introduction to Cell and Tissue Culture (JP Mather and PE Roberts, 1998) Plenum Press, Cell and Tissue Culture: Laboratory Procedures (A. Doyle, JBGriffiths, and DG Newell, eds. 1993-8) J. Wiley and Sons, Methods in Enzymology (Academic Press, Inc.), Handbook of Experimental Immunology (DMWeir and CCBlackwell, eds.): Gene Transfer Vectors for Mammalian Cells (JMMiller and MPCalos, eds., 1987), Current Protocols in Molecular Biology (FMAusubel, et al. eds. 1987); PCR: The Polymerase Chain Reaction, (Mullis, et al., eds. 1994), Current Protocols in Immunology (JEColigan et al., eds., 1991), Short Protocols in Molecular Biology (Wiley and Sons, 1999), Immunobiology (C. A. Janeway and P. Travers, 1997), Antibodies (P. Finch, 1997), Antibodies: a practice approach (D. Catty., ed., IRL Press, 1988-1989), Monoclonal antibodies: a practical approach (P. Shepherd and C. Dean, eds., Oxford University Press, 2000), Using antibodies: a laboratory manual (E. Harlow and D. Lane (Cold Spring Harbor Laboratory Press, 1999), The Antibodies (M. Zanetti and J. D. Capra, eds. Harwood Academic Publishers, 1995), DNA Cloning: A practical Approach, Volumes I and II (D. N. Glover ed. 1985), Nucleic Acid Hybridization (B. D. Hames & S. J. Higgins eds. (1985)); Transcription and Translation (B. D. Hames & S. J. Higgins, eds. (1984)); Animal Cell Culture (R. I. Freshney, ed. (1986)); Immobilized Cells and Enzymes (IRL Press, (1986)), and B. Perbal, A practical Guide To Molecular Cloning (1984); F. M. Ausubel et al. (eds.), are fully described in the literature.

[0211] Without further detail, those skilled in the art will be able to make the most of this disclosure based on the above description. Therefore, the following specific embodiments should be construed as illustrative only and will not limit the remainder of this disclosure. All publications referenced herein are incorporated by reference for the purposes or subjects referenced herein. [Examples]

[0212] The following embodiments are provided to further illustrate some embodiments of the present disclosure, but are not intended to limit the scope of the present disclosure. It will be understood that, by their exemplary nature, other procedures, methods, or techniques known to those skilled in the art may be used as alternatives.

[0213] Example 1: CRISPR nuclease-mediated editing of human target genes in HEK293T cells This example describes genome editing of exemplary target genes, including the AAVS1, EMX1, and VEGFA genes, by introducing the CRISPR nuclease of SEQ ID NO: 1 into HEK293T cell lines via lipid-based transient transfection.

[0214] CRISPR nucleases were tagged with the N-terminal SV40 nuclear localization signal (NLS) and the C-terminal XTEN linker immediately upstream of the nucleoplasmin NLS. Their coding sequences were converted to human codon-optimized DNA sequences, synthesized, and cloned into a pcDNA3.1 vector (Invitrogen) containing a CMV promoter for expression. The reference and NLS-tagged sequences used are shown in Table 1. Plasmids were purified using a midiprep kit. [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4]

[0215] RNA guides were designed and cloned into the pUC19 plasmid, followed by the U6 PolIII promoter, and terminated with a 6x polyT sequence. The RNA guides were designed to be specific to target sequences within the coding exons of AAVS1, EMX1, and VEGFA, which have a 5'-NGG-3'PAM sequence (the PAM sequence is located at the 3' end of the target sequence). For more efficient transcription, the U6 PolIII promoter used +1 G at the start of the transcript (i.e., the 5' end of the RNA), which is excluded from the sequences described here. See Table 2 for all RNA guide sequences. Plasmids were purified using a midiprep kit. [Table 3]

[0216] Approximately 16 hours prior to transfection, 25,000 HEK293T cells were seeded in DMEM / 10% FBS+Pen / Strep (D10 medium) into each well of a 96-well plate. On the day of transfection, the cells were 70–90% confluent. For each well to be transfected, a mixture of Lipofectamine 2000® (ThermoFisher Scientific) and Opti-MEM® (ThermoFisher Scientific) was prepared and incubated at room temperature for 5 minutes (Solution 1). After incubation, the Lipofectamine 2000®:Opti-MEM® mixture was added to a separate mixture containing a CRISPR nuclease plasmid (NLS tagged), an RNA guide plasmid, and Opti-MEM® (Solution 2). In the case of a negative control, the CRISPR nuclease plasmid was excluded. Solutions 1 and 2 were mixed by pipetting up and down, and then incubated at room temperature for 25 minutes. After incubation, the mixture of solutions 1 and 2 was added dropwise to each well of a 96-well plate containing cells. Approximately 72 hours after transfection, the cells were trypsinized by adding TrypLE® (Thermo Fisher Scientific) to the center of each well and incubating at 37°C for approximately 5 minutes. D10 medium was then added to each well and mixed to resuspend the cells. The resuspended cells were centrifuged for 10 minutes to obtain a pellet, and the supernatant was discarded. The cell pellet was then resuspended in QuickExtract® buffer (Lucigen®), and the cells were incubated at 65°C for 15 minutes, 68°C for 15 minutes, and 98°C for 10 minutes.

[0217] Next-generation sequencing (NGS) samples were prepared by two rounds of PCR. Three technical replicates were analyzed for each target, reference, and variant. The first round (PCR1) was used to amplify specific genomic regions in a target-dependent manner. The second round of PCR (PCR2) was performed to add the Illumina adapter and index. The reaction products were then pooled and purified by column purification. Sequencing was performed using a 150-cycle NextSeq 500 / 550 intermediate or high-power v2.5 kit or a 200-cycle NovaSeq 6000 SP or S1 reagent kit v1.5.

[0218] For NGS analysis, the indel mapping function used the sample's fastq file, amplicon reference sequence, and forward primer sequence. For each read, the KMER scanning algorithm was used to calculate the editing behavior (match, mismatch, insertion, deletion) between the read and the reference sequence. To remove small amounts of primer dimers present in some samples, the first 30 nt of each read was required to match the reference, and reads with more than half of the mapping nucleotides being mismatched were filtered out. Up to 50,000 reads that passed these filters were used for analysis, and if a read contained an insertion or deletion, it was counted as an indel read. The QC criterion for the minimum number of filtered reads was 10,000.

[0219] For each target, the indel ratio, which refers to the percentage of NGS reads containing indels, was calculated for each sample and its corresponding protein-free control. The higher percentage of indels in targets when CRISPR nucleases were included in transfection demonstrated the DNA editing results in cells.

[0220] As shown in Figure 1, each of the six targets tested demonstrated a greater level of indels than was observed when the CRISPR nuclease plasmid was present. Therefore, this example demonstrates that the CRISPR nuclease of Sequence ID No. 1 edited a human gene.

[0221] Example 2: Effect of variant CRISPR nuclease for targeting mammalian genes This example describes an indel evaluation for an exemplary mammalian target using a CRISPR nuclease variant transfected into HEK293T cells.

[0222] Arginine scanning mutagenesis was performed to individually substitute selected non-arginine residues with arginine in the reference CRISPR nuclease (SEQ ID NO: 1). SEQ ID NO: 1 is referred to herein as the reference sequence. This resulted in 372 single-arginine substitution variants. The reference CRISPR nuclease variants and the nucleic acids encoding each CRISPR nuclease variant were then individually cloned into a pcDNA3.1 backbone (Invitrogen™) to prepare plasmids, which were then diluted. The plasmids contained a CMV promoter, a first NLS upstream of the coding sequence (MKRTADGSEFESPKKKRKV, SEQ ID NO: 3), an XTEN linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGSS, SEQ ID NO: 8), and a second NLS downstream of the coding sequence (KRPAATKKAGQAKKKK, SEQ ID NO: 4). See also Example 1 above.

[0223] Exemplary RNA guides for VEGFA-T6 and EMX1-T7 were used in this study. Details of these gRNAs are provided in Table 2 above. The RNA guides were cloned into a pUC19 backbone (New England Biolabs®). Plasmids were purified and diluted using a maxi-prep kit. Cells were transfected, and samples were prepared for NGS as described in Example 1. Indel ratios, which refer to the percentage of NGS reads containing indels, were calculated for the reference and for each variant. The indel ratio used for multiplier change calculations was the average of two technical replicates. Then, to calculate the multiplier change in the indel ratio, the indel ratio for each variant was divided by the indel ratio for the reference. Table 3 shows the multiplier change in the indel ratio for each target tested. The numbering is relative to the reference nuclease (i.e., without NLS) of Sequence ID No. 1.

[0224] As shown in Table 3, six of the 372 variants with a single arginine substitution (left column) were characterized as resulting in an increase of at least 1.5 times the indel ratio compared to the reference indel ratio when averaged across two targets (right column). [Table 4]

[0225] Fifty-five variants with a single arginine substitution were analyzed to have indel ratios 1 to 1.5 times that of the reference indel ratio: G1329R, Q741R, P1238R, A1236R, Q1230R, E1214R, A730R, Q849R, S473R, D985R, M753R, K918R, I1106R, M1205R, Y1015R, Q1360R, K1344R, F501R, K1091R, G731R, N1295R, F1241R, V752R, N1099 R, E1179R, D720R, F1285R, I1370R, E1037R, E982R, S1094R, S872R, P1096R, A1226R, Q840R, I1147R, L1117R, E331R, T1108R, K1298R, Y986R, S898R, A1333R, P1208R, E1105R, Q809R, L1281R, A1292R, K1131R, I581R, I79R, D1284R, T1104R, Y348R, and S1348R. The remaining variants with a single arginine substitution (311 variants) resulted in a reduced indel ratio relative to the reference indel ratio (a multiplier change in indel ratios less than 1.0).

[0226] We selected the following variants that exhibited at least a 1.5-fold increase in the indel ratio relative to the reference, and created combined variants: I857R, N813R, L784R, K736R, A919R, and Q812R.

[0227] Example 3: Efficacy of combined CRISPR nuclease variants for targeting mammalian genes This example describes indel evaluation for mammalian targets using CRISPR nuclease variants containing two or more substitutions identified in Example 2 as increasing indel activity. Thirty-five combinations of CRISPR nuclease variants were tested.

[0228] Each CRISPR nuclease variant and RNA guide was cloned as described in Example 2. Exemplary RNA guides for VEGFA-T6 and EMX1-T7 were used in this study. Details of these gRNAs are provided in Table 2 above. HEK293T cells were further transfected as described in Example 2, followed by NGS analysis. For each target, the indel ratio, which refers to the percentage of NGS reads containing indels, was calculated for the reference CRISPR nuclease (SEQ ID NO: 1) and for each variant CRISPR nuclease. The indel ratios shown in Table 4 were calculated as the average of two biological replicas, each containing two technical replicas. [Table 5]

[0229] As shown in Table 4, each CRISPR nuclease variant with amino acid substitution combinations exhibited higher indel activity than the reference CRISPR nuclease (SEQ ID NO: 1). Nine CRISPR nuclease variants yielded indel ratios greater than 0.25 when averaged across both targets, indicating that more than 25% of NGS reads contained indels. These nine CRISPR nuclease variants included the following substitution combinations: a) I857R, L784R, K736R; b) I857R, A919R, K736R; c) I857R, N813R, L784R; d) I857R, L784R, A919R; e) I857R, N813R, K736R; f) I857R, N813R; g) L784R, A919R, K736R; h) I857R, L784R; and i) I857R, A919R. Eight of the CRISPR nuclease variants yielded indel ratios of 0.2–0.24 when averaged across both targets, indicating that 20%–24% of NGS reads contained indels. The 18 CRISPR nuclease variants yielded indel ratios of 0.1–0.19 when averaged across both targets, indicating that 10–19% of NGS reads contained indels. The average indel ratio across both targets exceeded the average indel ratio of the reference for all variants tested.

[0230] Based on this experiment, the highest-performing CRISPR nuclease variants, including substituted I857R, L784R, and K736R, were selected for further testing. These CRISPR nuclease variants exhibited a 2.5-fold increase in indel activity compared to the reference CRISPR nuclease.

[0231] Example 4: Production and effects of CRISPR nickerse variants for targeting mammalian genes This example describes introducing a mutation into the CRISPR nuclease of SEQ ID NO: 1 that disrupts either the HNH or RuvC domain to produce a functional nickase. D844, H845, and N868 were identified as putative catalytic residues in the HNH domain. D10, E763, and D991 were identified as putative catalytic residues in the RuvC domain. These positions were identified by analyzing models generated with AlphaFold 2 (Jumper et al., Nature 596:583-9 (2021)) for structural regions similar to known HNH and RuvC active sites, and / or by sequence alignment against other nucleases for which candidate sites had been previously identified. Examples of reference structures used to identify HNH and RuvC active sites are represented by the following Protein Databank (PDB) identifiers: 5h0m, 7eu9, 6ltu, 7odf, 7lys, 8dc2, 4cmp, 4oo8, 7z4j, 5axw, 5b2o, 6kc8, 7utn, 8csz, 8ctl, and 8dmb.

[0232] The coding sequences of CRISPR nucleases were converted to E. coli codon-optimized DNA sequences, synthesized, and cloned into a pET-28a(+) vector (Novagen) containing lac and T7 RNA polymerase promoters for gene expression. To test nickase activity, individual alanine variants were cloned for each of the positions identified as putative active site residues in the HNH and RuvC domains. A leucine variant was also cloned for position H845. Research-grade plasmids were obtained from GenScript. The constructed nickase sequences are shown in Table 5. Codons encoding substituted residues are shown in underlined bold capital letters in the nucleotide sequence, and substituted residues are shown in underlined bold letters in the amino acid sequence. The putative HNH knockout nickase was expected to cleave non-target strands but not target strands. The putative RuvC knockout nickase was expected to cleave target strands but not non-target strands. [Table 6-1] Table 6-2 Table 6-3 Table 6-4 Table 6-5 Table 6-6 Table 6-7 Table 6-8 Table 6-9 Table 6-10 Table 6-11 Table 6-12 Table 6-13

[0233] A linear DNA template encoding an RNA guide was designed, having a T7 promoter upstream and a T7Te terminator sequence downstream. The RNA guide was designed to be specific to a previously tested target sequence as described in the foregoing Example 1 and Table 2, which is located within the coding exon of EMX1 having a 5'-NGG-3' PAM sequence (PAM is located on the 3' side of the target sequence). For more efficient transcription, the T7 promoter uses a +1 G at the start of the transcript (i.e., at the 5' end of the RNA), which is shown for SEQ ID NO: 44. The sequences of the encoded RNA guide and its individual components are shown in Table 6. Table 7

[0234] DNA targets were designed and ordered as synthetic linear DNA fragments. The target sequence from EMX1 and 10 bases upstream and downstream within the exon was flanked by 200 bases of irrelevant sequence upstream and 100 bases of irrelevant sequence downstream. Extra sequences were added, whereby cleaved and uncleaved products were well separated on a gel. The target and non-target strands were labeled with 5' IR700 and 5' IR800 labels respectively via PCR amplification using labeled primers. The sequences of the DNA targets, individual components of the DNA targets, and labeled PCR primers are set forth in Table 7. Table 8

[0235] The cleavage activity of the reference CRISPR nuclease (SEQ ID NO: 1) and putative nickase was evaluated using an in vitro cleavage assay. Each polypeptide was individually co-expressed in vitro with an RNA guide by incubating a plasmid encoding the target protein from Table 5 and a linear DNA template for the T7 transcription EMX1-T2 sgRNA from Table 6 in PURExpress® solution (NEB) containing the SUPERase·In® RNase inhibitor (Invitrogen) for 2 hours at 37°C. The unpurified polypeptide / RNA solution was then diluted in a 1x solution of NEB Buffer 2 (NEB) containing approximately 1 ng / μl of labeled DNA target amplicon. The solution was then incubated at 37°C for 1 hour. The reaction was stopped by incubation with RNase Cocktail (Invitrogen; final concentration approximately 1 U / μl) at 37°C for 15 minutes, followed by incubation with proteinase K (NEB; final concentration approximately 0.04 U / μl) at 55°C for 30 minutes. The DNA was then purified using CleanNGS DNA & RNA Clean-Up Magnetic Beads (Bulldog Bio).

[0236] The cleaved and uncleaved products of the target and non-target strands were separated by running the sample on a 10% TBE-urea PAGE gel. The gel was imaged using a LI-COR Odyssey M imaging system with 700 nm and 800 nm channels to visualize 5'IR700 and 5'IR800 labeling in the target and non-target strands of the target DNA substrate. Band intensities were quantified using ImageJ software.

[0237] Gel images are shown in Figures 2A–2C, and the quantification of the percentage of cleaved target and non-target chains is shown in Figure 2D. Uncleaved chain, HNH-cleaved chain, and RuvC-cleaved chain are shown. Figure 2A is a gel image captured using a 700 nm channel, showing cleavage of the target chain. Figure 2B is a gel image captured using an 800 nm channel, showing cleavage of the non-target chain. Figure 2C is an overlay of gel images from Figures 2A and 2B. As shown in Figures 2A–2D, the reference CRISPR nuclease (SEQ ID NO: 1) cleaved both the target and non-target chains, as expected. Three of the four HNH knockout nickase constructs (H845A, H845L, and N868A) showed significantly reduced activity in the target chain while retaining activity in the non-target chain. Each of the three RuvC knockout nickase constructs (D10A, E763A, and D991A) exhibited significantly reduced activity in the non-target chain while retaining activity in the target chain (Figures 2A-2D).

[0238] Therefore, this example demonstrates the successful production of HNH knockout nickase and RuvC knockout nickase. The H845A variant was selected to introduce editing into a human gene target, as described in Example 5 below.

[0239] Example 5: Fusion of CRISPR nuclease and CRISPR niccas to reverse transcriptase In this example, the reverse transcriptase polypeptide was fused to the C-terminus of the CRISPR nuclease of SEQ ID NO: 1, or to the H845A nicasse variant of the CRISPR nuclease (SEQ ID NO: 32). See Tables 1 and 5 below.

[0240] The sequence encoding a CRISPR nuclease-reverse transcriptase fusion polypeptide was cloned into a pcDNA3.1 vector (Invitrogen) containing a CMV promoter. The fusion contained the following components arranged from N-terminus to C-terminus: 1) SV40 NLS, 2) CRISPR nuclease of SEQ ID NO: 1, 3) XTEN linker, 4) variant Moloney's mouse leukemia virus (MMLV) reverse transcriptase, and 5) nucleoplasmin NLS. The nucleotide and amino acid sequences of SV40 NLS, CRISPR nuclease of SEQ ID NO: 1, XTEN linker, and nucleoplasmin NLS are shown in Table 1 of Example 1, and the nucleotide and amino acid sequence of variant MMLV is shown in Table 9 below. The variant MMLV reverse transcriptase was a human codon-optimized DNA sequence. Research-grade plasmids were obtained from GenScript. [Table 9]

[0241] Next, a CRISPR nuclease-reverse transcriptase fusion polypeptide plasmid DNA was used to introduce the H845A nickase mutation (see Example 4) using a site-directed mutagenesis kit (New England Biolabs®). The nucleotide and amino acid sequences of the CRISPR nuclease-reverse transcriptase fusion polypeptide and the CRISPR nickase-reverse transcriptase fusion polypeptide are shown in Table 8. Codons encoding the substituted H845A residue are shown in underlined bold capital letters in the nucleotide sequence, and the substituted residue is shown in underlined bold letters in the amino acid sequence. The sequence-validated plasmids were then purified using the Qiagen Maxiprep kit. [Table 10-1] [Table 10-2] [Table 10-3] [Table 10-4] [Table 10-5] [Table 10-6]

[0242] This example describes how CRISPR nuclease-reverse transcriptase fusion polypeptide and CRISPR niccasse (H845A)-reverse transcriptase fusion polypeptide constructs were cloned. These constructs were used in Example 6 to introduce editing into a human target gene.

[0243] Example 6: Editing of human gene RNA in HEK293T cells using CRISPR nuclease-reverse transcriptase fusion polypeptide and CRISPR niccas-reverse transcriptase fusion polypeptide. This example demonstrates the genetic modification of a human gene using the CRISPR nuclease-reverse transcriptase fusion polypeptide and CRISPR niccas-reverse transcriptase fusion polypeptide constructed in Example 5. Specifically, the fusion polypeptide was used to introduce a six-nucleotide sequence substitution into a human target gene.

[0244] The editing template RNA was designed to be specific for the target sequences shown in Table 11, and cloned into the pUC19 plasmid containing the U6 PolIII promoter and a 6×polyT terminator sequence. The editing template was synthesized by GenScript, and from the 5' end to the 3' end, it comprised the following five components: 1) a spacer sequence, 2) a scaffold motif, 3) a reverse transcription template (RTT) encoding six nucleotide substitutions, 4) a primer binding site (PBS), and 5) a 3' extension motif. The RTT was 23 nucleotides in length. Two different PBS sequences of different lengths were tested: 7 nucleotides (editing templates 1 to 5) and 9 nucleotides (editing templates 6 to 10). The 3' extension motif contains a short linker sequence and a pseudoknot. The linker sequence was added to prevent steric collision between the PBS and the pseudoknot motif. The pseudoknot was added to protect the editing template RNA from 3' exonuclease activity.

[0245] Guides AAVS1-T3, EMX1-T2, EMX1-T7, VEGFA-T3, and VEGFA-T6 were used in this example. PAM, target, spacer, and gRNA sequences are provided in Table 2 above. The sequences of each component and the full-length editing template RNA sequences are shown in Table 11 and Table 12 below. The U6 PolIII promoter uses a +1 G at the start of the transcript (i.e., the 5' end of the RNA) for more efficient transcription, and this is excluded from the sequences set forth in Table 12.

Table 11-1

Table 11-2

Table 12-1

Table 12-2

[0246] Approximately 16 hours prior to transfection, 25,000 HEK293T cells were seeded in DMEM / 10% FBS+Pen / Strep (D10 medium) into each well of a 96-well plate. On the day of transfection, the cells were 50–70% confluent. For each well to be transfected, a mixture of Lipofectamine 2000® (ThermoFisher Scientific) and Opti-MEM® (ThermoFisher Scientific) was prepared and incubated at room temperature for 5 minutes (Solution 1). After incubation, the Lipofectamine 2000®:Opti-MEM® mixture was added to a separate mixture containing CRISPR nuclease-reverse transcriptase fusion polypeptide, edited template RNA, and Opti-MEM® (Solution 2), or CRISPR nickase-reverse transcriptase fusion polypeptide, edited template RNA, and Opti-MEM® (Solution 2). Solutions 1 and 2 were mixed by pipetting up and down, and then incubated at room temperature for 25 minutes. After incubation, the mixture of solutions 1 and 2 was added dropwise to each well of a 96-well plate containing cells. Approximately 72 hours after transfection, the cells were trypsinized by adding TrypLE® (Thermo Fisher Scientific) to the center of each well and incubating at 37°C for approximately 5 minutes. D10 medium was then added to each well and mixed to resuspend the cells. The resuspended cells were centrifuged for 10 minutes to obtain a pellet, and the supernatant was discarded. The cell pellet was then resuspended in Quick Extract® buffer (Lucigen®), and the cells were incubated at 65°C for 15 minutes, 68°C for 15 minutes, and 98°C for 10 minutes.

[0247] Samples were prepared for NGS and analyzed as described in Example 1. For each target, fractions of NGS reads containing indels were calculated for each sample and its corresponding protein-free control. To determine the percentage of edits introduced into the target gene, sequencing reads containing six nucleotide substitutions encoded by the editing template RNA were analyzed and quantified. The percentages of NGS reads containing indels and six nucleotide edits are shown in Tables 13 and 14, and further in Figures 3A and 3B. In Figures 3A and 3B, the percentage of NGS reads is shown on the y-axis, total edits are shown as black bars, and six nucleotide edits are shown as gray bars. The data in Tables 13, 14, Figure 3A, and Figure 3B are the mean of three technical replicates. [Table 13] [Table 14]

[0248] As shown in Figures 3A and 3B, the CRISPR nuclease-reverse transcriptase fusion polypeptide and the CRISPR nickase-reverse transcriptase fusion polypeptide introduced substitutions encoded by the tested editing template RNA at the AAVS1, EMX1, and VEGFA target loci, respectively. For the CRISPR nuclease-reverse transcriptase fusion polypeptide, the mean percentage of NGS reads containing indels ranged from 11.67% to 41.32%, and the mean percentage of NGS reads containing coded edits ranged from 0.29% to 12.80% (Table 13 and Figure 3A). Editing template RNA 3 showed the lowest incorporation of coded edits, while editing template RNA 2 showed the highest indel and coded edit introductions. For CRISPR nickase-reverse transcriptase fusion polypeptides, the average percentage of NGS reads containing indels ranged from 0.29% to 0.74%, and the average percentage of NGS reads containing coded edits ranged from 0.05% to 10.66% (Table 14 and Figure 3B). For CRISPR nickase-reverse transcriptase fusion polypeptides, low indel incorporation (less than 1%) confirmed that the H845A substitution converted CRISPR nuclease to CRISPR nickase. For both CRISPR nuclease-reverse transcriptase fusion polypeptides and CRISPR nickase-reverse transcriptase fusion polypeptides, the edit template RNA2 showed the introduction of the largest coded edits.

[0249] Conversely, controls consisting of the CRISPR nuclease of Sequence ID No. 1 along with the RNA guide in Table 2 or the editing templates in Tables 13 and 14 induced indel formation but did not incorporate the six nucleotide substitutions encoded by the editing template RNA. Furthermore, controls consisting of the CRISPR nuclease-reverse transcriptase fusion polypeptide or the CRISPR nickase-reverse transcriptase fusion polypeptide along with the RNA guide in Table 2 did not result in the incorporation of the six nucleotide substitutions encoded by the editing template RNA.

[0250] The editing efficiency mediated by CRISPR nuclease-reverse transcriptase fusion polypeptides and CRISPR nickase-reverse transcriptase fusion polypeptides was further tested with edit template RNAs having PBS lengths of 11, 13, 15, and 17 nucleotides, and RTT lengths of 18 and 23 nucleotides. These edit template RNAs were found to behave similarly to edit template RNAs containing PBS with lengths of 7 or 9 nucleotides and RTT with lengths of 23 nucleotides.

[0251] Overall, this example demonstrates that CRISPR nuclease-reverse transcriptase fusion polypeptides and CRISPR nickase-reverse transcriptase fusion polypeptides incorporated substitutions encoded by editing template RNA into human genes.

[0252] Example 7: Optimization of CRISPR nuclease and CRISPR niccas fusion into reverse transcriptase This example describes the preparation of a fusion of a reverse transcriptase polypeptide with the CRISPR nuclease of SEQ ID NO: 1 or the H845A nickasase variant of the CRISPR nuclease (SEQ ID NO: 32).

[0253] Plasmid libraries containing various combinations and orientations of NLS tags, mobile linkers, CRISPR nuclease of SEQ ID NO: 1, or CRISPR nickase of SEQ ID NO: 32, variant reverse transcriptase of SEQ ID NO: 53, and FLAG tags were designed and synthesized using GenScript. The sequences of the individual NLS, linker, and FLAG tag components are shown in Table 15. The resulting configurations are shown in Table 16. [Table 15-1] [Table 15-2] [Table 15-3] [Table 15-4] [Table 16]

[0254] The plasmid library was screened in HEK293T cells using the lipid-based transient transfection method described in Example 5. Each well of a 96-well plate was transfected with plasmids encoding a proprietary CRISPR nuclease-reverse transcriptase fusion polypeptide or a CRISPR nickase-reverse transcriptase fusion polypeptide. The edited template RNA sequence of SEQ ID NO: 70, designed to introduce six nucleotide substitutions to the EMX1_T2 target, was also transfected into each well. Quantification of the edits was performed as described in Example 6.

[0255] Figures 4A and 4B show the CRISPR nuclease-reverse transcriptase fusion polypeptides and CRISPR nickase-reverse transcriptase fusion polypeptides that introduced the highest percentage of edits encoded by the editing template RNA of all tested constructs, respectively. The results are the average of the two technical replicates. The dotted line shows the percentage of reads containing six nucleotide substitutions introduced by the control CRISPR nuclease-reverse transcriptase fusion polypeptide (Figure 4A) or the CRISPR nickase-reverse transcriptase fusion polypeptide (Figure 4B). As shown in Figures 4A and 4B, several fusion constructs resulted in an increased introduction of six nucleotide substitutions compared to the control CRISPR nuclease-reverse transcriptase fusion polypeptide and the CRISPR nickase-reverse transcriptase fusion polypeptide. The sequences of the highest-performing CRISPR nickase-reverse transcriptase fusion polypeptides are shown in Table 17. [Table 17-1] [Table 17-2] [Table 17-3] [Table 17-4] [Table 17-5] [Table 17-6] [Table 17-7] [Table 17-8] [Table 17-9] [Table 17-10] [Table 17-11] [Table 17-12] [Table 17-13]

[0256] Therefore, this example demonstrates that the incorporation of editing into the target can be improved by optimizing the composition of the CRISPR nuclease-reverse transcriptase fusion polypeptide and the CRISPR niccas-reverse transcriptase fusion polypeptide.

[0257] Example 8: Further screening of CRISPR niccas-reverse transcriptase fusion polypeptides This example describes the design and implementation of a reporter-based HEK293T stable cell line that can be used to measure the activity of a CRISPR niccas-reverse transcriptase fusion polypeptide. This system is an orthogonal readout to the NGS-based assay used in the previous example.

[0258] A stable cell line with an integrated Blue Fluorescent Protein (BFP) reporter target was created using a modified version of the traffic light reporter (TLR) described in Glaser et al., Mol Ther Nucleic Acids 5(7):e334 (2016). Editing by converting BFP to eGFP via two amino acid changes (SH>TY) results in the ability to measure eGFP intensity as a percentage of the total cell population. In this assay, cells containing the integrated reporter are mCherry positive.

[0259] Table 18 shows the sequences of the BFP target and the editing template RNA designed to convert BFP to eGFP. The editing template RNA was cloned and transfected into stable cell lines as described in Example 6. The CRISPR niccas-reverse transcriptase fusion polypeptide library described in the previous example was also transfected. [Table 18]

[0260] TLR screening analysis was performed 72 hours after transfection by imaging live cells using Operetta CLS (Perkin Elmer) and its Harmony software. eGFP was quantified and compared to the overall mCherry-positive cell population. The mCherry population represents the total number of cells containing the embedded reporter. Imaging data was collected and quantified as the percentage of eGFP-positive cells relative to the mCherry-positive cell population.

[0261] Figure 5 shows the highest-performing CRISPR-nickase-reverse transcriptase fusion polypeptides from TLR screening, expressed as the percentage of eGFP-positive cells in the mCherry-positive population. Hits are ranked by activity compared to control CRISPR-nickase-reverse transcriptase polypeptides and correlate with NGS data from Figures 4A and 4B. One additional trend observed was that increased mobile linker length between the CRISPR-nickase and reverse transcriptase components (linker_16xGGGGS > linker_8xGGGGS > linker_4xGGGGS > linker_1xGGGGS) resulted in increased BFP→eGFP editing, which translated into an increase in eGFP-positive cells in the reporter assay. Therefore, increasing the linker length between the CRISPR-nickase and reverse transcriptase components of the CRISPR-nickase-reverse transcriptase fusion polypeptides in Table 17 may be beneficial in further increasing editing efficiency.

[0262] Next, to compare the robustness of reporter-based eGFP quantification from TLR screening, the highest-performing hits were graphed and compared with NGS quantification editing for BFP to eGFP conversion (Figure 6). The percentage of editing versus the percentage of eGFP-positive cells showed a positive correlation, with higher editing introduction resulting in a higher percentage of eGFP-positive cells. Finally, editing at the EMX1_T2 target from previous examples was compared with BFP-eGFP editing for the highest-performing CRISPR-nickase-reverse transcriptase fusion polypeptide. As shown in Figure 7, for each CRISPR-nickase-reverse transcriptase fusion polypeptide, there was a positive correlation between editing at both targets. Furthermore, editing with each CRISPR-nickase-reverse transcriptase fusion polypeptide outperformed editing with the control CRISPR-nickase-reverse transcriptase fusion polypeptide.

[0263] Therefore, this example demonstrates that the optimized CRISPR niccas-reverse transcriptase fusion polypeptide introduces editing at multiple target gene loci.

[0264] Example 9: Design of additional CRISPR nucleases This example describes the design of an additional CRISPR nuclease having desired bioactivity.

[0265] Further variants of the CRISPR nickase described herein are prepared. Each individual amino acid residue of the CRISPR nickase having the H845A substitution from Example 3 is replaced with the remaining 19 available amino acids (excluding the H845A residue). The nucleic acid encoding the CRISPR nickase variant is cloned into a pcDNA3.1 vector (Invitrogen) containing the CMV promoter. The expression vector is introduced into host cells for the expression of the CRISPR nickase variant. The CRISPR nickase variants expressed in host cells are purified, and their nickase activity is evaluated according to the procedure provided in Example 3 above.

[0266] Additional CRISPR nuclease variants are prepared to increase double-stranded nuclease activity. The variants in Table 19 are cloned and evaluated as described in Example 2. [Table 19-1] [Table 19-2] [Table 19-3] [Table 19-4] [Table 19-5] [Table 19-6] [Table 19-7] [Table 19-8] Table 19-9 Table 19-10 Table 19-11 Table 19-12 Table 19-13 Table 19-14 Table 19-15 Table 19-16 Table 19-17 Table 19-18 Table 19-19 Table 19-20 Table 19-21 Table 19-22 Table 19-23 Table 19-24 Table 19-25 Table 19-26 Table 19-27 Table 19-28 Table 19-29 Table 19-30 Table 19-31 Table 19-32 Table 19-33 Table 19-34 Table 19-35 Table 19-36 Table 19-37 Table 19-38 Table 19-39 Table 19-40 Table 19-41 [Table 19-42] [Table 19-43] [Table 19-44] [Table 19-45]

[0267] We will construct additional CRISPR nuclease variants and evaluate their ability to recognize less stringent PAM sequences. We will clone the variants in Table 20 and evaluate them as described in Example 2 using target sequences adjacent to 5'-NGN-3', 5'-NRN-3', or 5'-NYN-3' PAM sequences (where N represents any nucleotide, R represents G or A, and Y represents C or T). [Table 20-1] [Table 20-2] [Table 20-3]

[0268] Example 10: Effect of a variant of a CRISPR nuclease having relaxed PAM stringency for targeting exemplary mammalian genes. This example describes an indel evaluation for an exemplary mammalian target using a variant of a CRISPR nuclease with a relaxed PAM, transfected into HEK293T cells.

[0269] Arginine scanning mutagenesis was performed to individually substitute selected non-arginine residues with arginine in the CRISPR nuclease variant of SEQ ID NO: 236. This resulted in 372 single-arginine substitution variants. The variants were cloned and evaluated using target sequences adjacent to the 5'-NGN-3'PAM sequence, summarized in Table 21, as described in Example 2. [Table 21]

[0270] HEK293T cells were further transfected as described in Example 2, followed by NGS analysis. The indel activity of the CRISPR nuclease variant of SEQ ID NO: 236 is shown in Table 22. The data in Table 22 are the average of 10 control samples, each of which had two biological replicas and two technical replicas. [Table 22]

[0271] Next, for each target, the indel ratio, which refers to the percentage of NGS reads containing indels, was calculated for the variant CRISPR nuclease (SEQ ID NO: 236) and for each variant CRISPR nuclease. Then, to calculate the multiplier change in the indel ratio, the indel ratio for each variant was divided by the indel ratio for the variant CRISPR nuclease of SEQ ID NO: 236. The indel ratio used for the multiplier change calculation was the average of the two technical replicas. As shown in Table 23, three of the 372 variants with a single arginine substitution (left column) were characterized as resulting in at least a twofold increase in the indel ratio compared to the indel ratio for the variant CRISPR nuclease of SEQ ID NO: 236 when averaged across the two targets (right column). [Table 23]

[0272] Eleven variants with a single arginine substitution were analyzed, assuming they had indel ratios 1.5 to 2 times the reference indel ratio: L64R, S410R, T67R, Q849R, G1110R, F501R, T659R, L784R, Y516R, G55R, and E1037R. 92 variants exhibited indel ratios 1 to 1.4 times the reference indel ratio: N57R, D720R, A919R, A1294R, Q812R, N700R, H657R, T73R, Q899R, T1347R, I857R, K751R, D327R, I581R, D462R, E331R, A589R, D471R, I699R. N1295R, T470R, I1147R, E130R, S473R, A353R, K40R, K334R, A60R, S1348R, K367R, A1118R, K 31R, Q349R, K341R, Q83R, K585R, Q840R, G660R, K527R, G727R, Y42R, L1281R, L122R, Q123R, T 1108R, E41R, K1131R, K30R, S872R, I1206R, D1132R, K460R, L80R, E459R, K1182R, L1117R, M 696R, K918R, K126R, N721R, G1227R, Q809R, K1091R, K736R, A1332R, K783R, N498R, K723R, E 1228R, H1119R, F463R, L594R, D472R, K744R, E365R, G595R, K45R, Y348R, K964R, S1181R, N813R, D407R, S839R, Y658R, E586R, G754R, A730R, Y1015R, D903R, A1333R, S461R, and H1359R. The remaining variants with a single arginine substitution (266 variants) resulted in a decrease in the indel ratio compared to the variant CRISPR nuclease of SEQ ID NO: 236 (indel ratio multiplier change less than 1.0).

[0273] Therefore, this embodiment demonstrates that the CRISPR nuclease variant of SEQ ID NO: 236 is an active nuclease capable of editing a target sequence adjacent to 5'-NGN-3'PAM (where N represents A, C, G, or U), and that certain further arginine substitutions (e.g., D61R, A68R, and / or H494R) increase nuclease activity.

[0274] Example 11: mRNA-mediated editing of target sequences in primary human hepatocytes This example describes genome editing of the EMX1 and VEGFA genes using mRNA encoding a CRISPR nuclease-reverse transcriptase fusion polypeptide or a CRISPR niccas-reverse transcriptase fusion polypeptide.

[0275] The nucleic acids encoding the CRISPR nuclease-reverse transcriptase fusion polypeptide of SEQ ID NO: 55 and the CRISPR nickase-reverse transcriptase fusion polypeptide of SEQ ID NO: 57 were individually cloned into an in vitro transcription (IVT) backbone containing the T7 promoter. Research-grade and sequence-validated plasmids were obtained using the maxi prep kit (Qiagen). mRNA was generated by in vitro transcription of the IVT backbone with the addition of a 5' cap and a 3' poly-A tail. Full-length mRNA sequences are shown in Table 24. Work-grade solutions of each mRNA were prepared in water. [Table 24-1] [Table 24-2] [Table 24-3] [Table 24-4] [Table 24-5] [Table 24-6]

[0276] Editing template RNAs 1 and 2 (Table 25) were designed to introduce three nucleotide insertions in the EMX1 and VEGFA target genes. The editing template RNAs were ordered from GenScript as desalting synthesis guides with the following chemical modifications: 2'-O-methyl groups for the first three bases and the last three bases, as well as phosphorothioate bonds between the first three and last three bases, as shown in bold in the Guide (RNA) column of Table 25. [Table 25]

[0277] PHH cells from human donors (Thermo Fisher Scientific or Lonza) were rapidly thawed from liquid nitrogen in a 37°C water bath. The cells were added to pre-warmed hepatocyte recovery medium (Thermo Fisher Scientific, CM7000) and centrifuged. The cell pellet was resuspended in an appropriate volume of Williams E medium (Thermo Fisher Scientific) supplemented with hepatocyte plating supplement pack (serum-containing) (Thermo Fisher Scientific). The cells were counted using a trypan blue viability counter and a Vi-CELL BLU cell counter. The desired number of viable cells were then washed in PBS and resuspended in P3 buffer + supplement (Lonza, V4SP-3096) and transfection enhancer oligo. The resuspended cells were distributed into a Lonza 96-well electroporation plate. mRNA effector (1 mg / mL in water) was mixed with synthetic RNA guide (1 mM in water) in a 1:1 volume ratio. mRNA / guide RNA mixtures were added to each reactant at a final mRNA concentration of 25 nM. The plates were electroporated using an electroporation device (Program DS-150, Lonza 4D-nucleofector). Immediately after electroporation, pre-warmed hepatocyte plating medium was added to each well and mixed very gently. For each technical replica plate, 125,000 diluted nucleofected cells were plated in a pre-warmed, collagen-coated 96-well plate (Thermo Fisher Scientific) containing hepatocyte plating medium. The cells were then incubated at 37°C. After 4 hours, the medium was changed to hepatocyte maintenance medium (Williams Medium E, Thermo Fisher Scientific) supplemented with Williams E medium cell maintenance cocktail (Thermo Fisher Scientific).

[0278] Three days after electroporation, cells were collected from the wells using Accutase (Thermo Fisher Scientific), transferred to a 96-well twin.tec® PCR plate (Eppendorf), and centrifuged. The culture medium was gently flicked off, and the cells were resuspended in DNA extraction buffer (QuickExtract®). The samples were circulated in a PCR machine at 65°C for 15 minutes, 68°C for 15 minutes, and 98°C for 10 minutes. The samples were then frozen at -20°C and analyzed by NGS as described in Example 1.

[0279] Edit incorporation into PHH using mRNA encoding the CRISPR nuclease-reverse transcription fusion polypeptide of SEQ ID NO: 55 (SEQ ID NO: 243) or mRNA encoding the CRISPR nickase-reverse transcription fusion polypeptide of SEQ ID NO: 57 (SEQ ID NO: 244) is shown in Table 26 or Table 27, respectively. For the CRISPR nuclease-reverse transcriptase fusion polypeptide, the mean percentage of NGS reads containing indels ranged from 12.6% to 13.2%, and the mean percentage of NGS reads containing three nucleotide insertions ranged from 0.57% to 0.95% (Table 26). For the CRISPR nickase-reverse transcription fusion polypeptide, the mean percentage of NGS reads containing indels ranged from 0.38% to 2.71%, and the mean percentage of NGS reads containing three nucleotide insertions ranged from 0.17% to 1.05% (Table 27). Overall, the use of mRNA encoding CRISPR niccas-reverse transcriptase fusion polypeptides and mRNA encoding CRISPR nuclease-reverse transcriptase fusion polypeptides resulted in similar levels of three-nucleotide insertions at the target locus. [Table 26] [Table 27]

[0280] Therefore, this example demonstrates that the above mRNA, combined with a synthetic guide, edits non-dividing human cells.

[0281] Example 12: Editing of mouse genes using RNA as a template in mice using CRISPR nuclease-reverse transcriptase fusion polypeptide and CRISPR niccas-reverse transcriptase fusion polypeptide. This example demonstrates genetic modification of the DNMT1 gene in mice using mRNA encoding a CRISPR nuclease-reverse transcriptase fusion polypeptide or a CRISPR niccas-reverse transcriptase fusion polypeptide.

[0282] The mRNA molecules provided in Table 24 of Example 11 were used. The edit template RNAs provided in Table 28 below were designed to introduce G>C substitutions or CCC insertions into the DNMT1 locus. G>C substitutions (edit template RNA 1) were introduced using 12 nucleotide RTTs, and CCC insertions (edit template RNA 2) were introduced using 15 nucleotide RTTs. The edit template RNAs provided in Table 28 were tested in the presence of the RNA guides provided in Table 29. These edit template guide RNAs and RNA guides were ordered from GenScript as HPLC-purified synthesis guides with the following chemical modifications: 2'-O-methyl in the first three bases and the last three bases, as well as a phosphorothioate bond between the first three bases and the last three bases, as shown in bold in the Guide (RNA) column of Table 28. [Table 28] [Table 29]

[0283] The following components were loaded into lipid nanoparticles (LNPs): a) mRNA encoding CRISPR nuclease-reverse transcriptase (SEQ ID NO: 55) (SEQ ID NO: 243), or mRNA encoding CRISPR nickase-reverse transcriptase fusion polypeptide (SEQ ID NO: 57) (SEQ ID NO: 244), b) Editing template RNA provided in Table 28, and c) RNA guide provided in Table 29.

[0284] LNP contained 46.3% cationic lipid 6-((2-hexyldecanoyl)oxy)-N-(6-((2-hexyldecanoyl)oxy)hexyl)-N-(4-hydroxybutyl)hexane-1-aminium, 9.4% phospholipid 1,2-distearoyl-sn-glycerol-3-phosphocholine (DSPC), 42.7% cholesterol, and 1.6% PEG lipid 2-[(polyethylene glycol)-2000]-N,N-ditetradecylacetamide, formulated in a molar N / P ratio of approximately 6. LNP was prepared according to the general procedure described in Schoenmaker, IJPharm, 601:120586, 2021, and its relevant disclosures are incorporated herein by reference for the subject and purposes referenced herein.

[0285] Male C57BL6 mice (6 weeks old, Jackson Laboratories, Bar Harbor, ME) were used in these studies. The animals were acclimated to the housing facility for at least 3 days prior to the start of the study. The animals' body weight was measured before administration. Study Group 1 was treated with mRNA encoding a CRISPR nickase-reverse transcriptase fusion polypeptide. Study Group 2 was treated with mRNA encoding a CRISPR nuclease-reverse transcriptase fusion polypeptide. The ratio columns in Tables 30 and 31 refer to the ratio of mRNA encoding the fusion polypeptide to the RNA guide to the editing template RNA. On day 0 of the study, the editing template, formulated in LNP, was administered intravenously via retroorbital injection in a volume of 200 μl. [Table 30] [Table 31]

[0286] Seven days after LNP administration, the animals were euthanized and perfused with PBS. Left liver lobe fragments containing 30–50 mg of tissue were collected and frozen on dry ice for tissue processing. The tissue was processed using TissueLyser II in Quick Extract® buffer (Lucigen®) with 5 mm beads, one bead per sample, and then resuspended in Quick Extract® buffer (Lucigen®). Cells were incubated at 65°C for 15 minutes and at 98°C for 2 minutes.

[0287] Samples were prepared for NGS and analyzed as described in Example 1. The percentage of NGS reads containing indels or edits encoded by the edited template RNA and incorporated into the DNMT1 locus was quantified. The results are shown in Tables 32 and 33. [Table 32] [Table 33]

[0288] For Study 1 using CRISPR nickase-reverse transcriptase fusion polypeptides, the mean percentage of NGS reads containing indels ranged from 0.42% to 0.98%, and the mean percentage of NGS reads containing edit introduction ranged from 1.13% to 5.41% (Table 32). For both edit template RNAs, there was a dose-dependent increase in the incorporation of encoded edits (Table 32). For Study 2 using CRISPR nuclease-reverse transcriptase fusion polypeptides, the mean percentage of NGS reads containing indels ranged from 5.92% to 13.02%, and the mean percentage of NGS reads containing edit introduction ranged from 0.186% to 1.12% (Table 33). Overall, the use of CRISPR nickase-reverse transcriptase fusion polypeptides resulted in higher edit introduction at the DNMT1 locus compared to CRISPR nuclease-reverse transcriptase fusion polypeptides.

[0289] Therefore, this example demonstrates that editing can be introduced at the DNMT1 locus in mice by CRISPR nuclease-reverse transcriptase fusion polypeptides and CRISPR niccas-reverse transcriptase fusion polypeptides.

[0290] Other Embodiments All features disclosed herein can be combined in any combination. Each feature disclosed herein may be replaced by an alternative feature that serves the same, equivalent, or similar purpose. Thus, unless otherwise expressly stated, each feature disclosed is merely an example of a general set of equivalent or similar features.

[0291] From the above description, those skilled in the art will readily grasp the essential features of the present invention and can adapt it to various uses and conditions by making various changes and modifications without departing from its spirit and scope. Therefore, other embodiments are also within the scope of the claims.

[0292] equivalent While several embodiments of the present invention are described and illustrated herein, various other means and / or structures for carrying out the functions described herein and / or obtaining one or more of the results and / or benefits will be readily conceivable to those skilled in the art, and each of such variations and / or modifications will be considered to fall within the scope of the embodiments of the present invention described herein. More generally, those skilled in the art will readily understand that all parameters, dimensions, materials and arrangements described herein are intended to be illustrative, and that actual parameters, dimensions, materials and / or arrangements will depend on the specific application or multiple application in which the teachings of the present invention are used. Those skilled in the art will be able to recognize or grasp many equivalents to the specific embodiments of the present invention described herein by means of conventional experimentation alone. Therefore, it should be understood that the embodiments described herein are presented merely as examples, and within the scope of the appended claims and equivalents thereto, embodiments of the present invention may be practiced in ways other than those specifically described and claimed. Embodiments of the present invention in this disclosure cover each individual feature, system, article, material, kit and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present invention as long as such features, systems, articles, materials, kits, and / or methods are not inconsistent with each other.

[0293] All definitions defined and used herein should be understood to take precedence over dictionary definitions, definitions incorporated by reference in documents, and / or the ordinary meanings of the defined terms.

[0294] All references, patents, and patent applications disclosed herein are incorporated by reference with respect to the subject matter they cite, which may in some cases encompass the entire document.

[0295] The indefinite articles “a” or “an” as used herein and in the claims should be understood to mean “at least one” unless explicitly stated otherwise.

[0296] The phrase “and / or” as used herein and in the claims should be understood to mean “either or both” of the elements thus combined, that is, elements that may coexist in some cases and exist separately in other cases. Multiple elements listed using “and / or” should also be interpreted in the same manner, that is, “one or more” of the elements thus combined. Other elements other than those specifically identified by the “and / or” clause may exist at their discretion, whether related to or unrelated to those specifically identified elements. Thus, as a non-restrictive example, a reference to “A and / or B,” when used in conjunction with an unrestrictive expression such as “including,” may refer in one embodiment to A only (optionally including elements other than B), in another embodiment to B only (optionally including elements other than A), in yet another embodiment to both A and B (optionally including other elements), and so on.

[0297] Where used herein and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as inclusive, that is, inclusion of several elements or a list of elements and at least one additional item not listed by choice, but also including two or more. Only terms that clearly indicate the opposite intention, such as “one of” or “exactly one of,” or, when used in the claims, “consisting of,” would refer to the inclusion of just one element from several elements or a list of elements. In general, where used herein, the term “or” shall be interpreted only as indicating exclusive substitution (i.e., “one or the other, but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “one of,” or “exactly one of.” Where used in the claims, “substantially consists of” has its usual meaning as used in the field of patent law.

[0298] As used herein in this specification and in the claims, the phrase “at least one” should be understood to mean at least one element selected from any one or more elements in a list of elements, but not necessarily including at least one of each element specifically enumerated in the list of elements or all elements, nor excluding any combination of elements in the list of elements. This definition also allows for the existence of elements other than those specifically identified in the list of elements to which the phrase “at least one” refers, whether related to or unrelated to those specifically identified elements, at the discretion of the definition. Therefore, as a non-restrictive example, “at least one of A and B” (or equivalently, “at least one of A or B,” or equivalently, “at least one of A and / or B”) could mean, in one embodiment, at least one A (and optionally including elements other than B) in which B is absent and optionally two or more are included; in another embodiment, at least one B (and optionally including elements other than A) in which A is absent and optionally two or more are included; in yet another embodiment, at least one A and optionally two or more are included, and at least one B (and optionally including other elements), and so on.

[0299] Furthermore, unless otherwise explicitly stated, it should be understood that in any method claimed herein that includes two or more steps or actions, the order of the steps or actions of the method is not necessarily limited to the order in which the steps or actions of the method are enumerated.

Claims

1. It is a gene editing system, (a) A fusion polypeptide comprising a CRISPR nuclease polypeptide and a reverse transcriptase (RT) polypeptide, or a first nucleic acid encoding the fusion polypeptide, wherein the CRISPR nuclease polypeptide comprises the amino acid sequence of SEQ ID NO: 1 or is a variant of SEQ ID NO: 1, and the variant is (i) One or more mutations in the HNH nuclease domain or the RuvC nuclease domain of Sequence ID No. 1 that reduce or eliminate the nuclease activity thereof. (ii) One or more arginine and / or lysine substitutions, optionally one or more arginine substitutions, (iii) One or more mutations to reduce PAM recognition stringency, or A fusion polypeptide or first nucleic acid comprising any combination of (iv), (i), (ii), and (iii), (b) An RNA molecule comprising guide RNA (gRNA) and reverse transcription donor RNA (RT donor RNA), or a second nucleic acid encoding the RNA molecule, The gRNA comprises a scaffold sequence recognizable by the CRISPR nuclease and a spacer sequence specific to a target sequence within the desired genomic region, wherein the target sequence is located upstream of a protospacer adjacent motif (PAM). A gene editing system comprising an RNA molecule or a second nucleic acid, wherein the RT donor RNA includes a primer binding site (PBS) and a template sequence.

2. The gene editing system according to claim 1, wherein the fusion polypeptide further comprises one or more nuclear localization signals (NLS) located upstream or downstream of the CRISPR nuclease polypeptide, the RT polypeptide, or both.

3. The gene editing system according to claim 2, wherein the fusion polypeptide comprises a first NLS, the CRISPR nuclease polypeptide, the RT polypeptide, and a second NLS from the N-terminus to the C-terminus, and optionally further comprises a peptide linker between the CRISPR nuclease polypeptide and the RT polypeptide.

4. The gene editing system according to claim 2, wherein the fusion polypeptide comprises a first peptide linker located between the CRISPR nuclease polypeptide and the RT polypeptide, and the fusion polypeptide comprises a first NLS and a second NLS located at the N-terminus and / or C-terminus of the fusion polypeptide.

5. The gene editing system according to claim 4, wherein the fusion polypeptide further comprises a third NLS and optionally a fourth NLS.

6. The gene editing system according to claim 4, wherein the fusion polypeptide further comprises a second peptide linker that connects the CRISPR nuclease polypeptide or the RT polypeptide to the first or second NLS(s), and optionally further comprises a third peptide linker that connects two NLSs.

7. The aforementioned fusion polypeptide, from the N-terminus to the C-terminus, (i) the first NLS, the second NLS, the CRISPR nuclease polypeptide, the first peptide linker, and the RT polypeptide, (ii) the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the first NLS, and the second NLS, (iii) The first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the third NLS, the second peptide linker, and the second NLS, (iv) the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the third NLS, the second NLS, the second peptide linker, and the first NLS, (v) The first NLS, the second peptide linker, the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the third peptide linker, and the second NLS, (vi) the first NLS, the second peptide linker, the CRISPR nuclease, the first peptide linker, the RT polypeptide, the third peptide linker, the third NLS, the fourth NLS, and the second NLS, (vii) the first NLS, the second peptide linker, the RT polypeptide, the first peptide linker, the CRISPR nuclease polypeptide, the third peptide linker, the third NLS, the fourth NLS, and the second NLS, (viiii) The RT polypeptide, the first peptide linker, the CRISPR nuclease polypeptide, the third NLS, the second NLS, the second peptide linker, and the first NLS, (ix) the first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the second peptide linker, the second NLS, and the RT polypeptide, (x) the first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the third NLS, the second peptide linker, the RT polypeptide, the third peptide linker, and the second NLS, (xi) the first NLS, the RT polypeptide, the first peptide linker, the second NLS, the second peptide linker, and the CRISPR nuclease polypeptide, (xi) The first NLS, the RT polypeptide, the first peptide linker, the third NLS, the second peptide linker, the CRISPR nuclease polypeptide, the third peptide linker, and the second NLS, (xiii) The first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, and the second NLS, or (xiv) A gene editing system according to any one of claims 4 to 6, comprising the first NLS, the CRISPR nuclease polypeptide, the first peptide linker, the RT polypeptide, the second peptide linker, the third peptide linker, and the second NLS.

8. The gene editing system according to claim 7, wherein the fusion polypeptide comprises (iv), (v), (vii), (ix), or (x) from the N-terminus to the C-terminus.

9. The gene editing system according to any one of claims 4 to 8, wherein the peptide linker(s) between the CRISPR nuclease polypeptide and the RT polypeptide is approximately 20 to 80 amino acids in length.

10. The gene editing system according to any one of claims 1 to 9, wherein the CRISPR nuclease polypeptide is the variant of SEQ ID NO:

1.

11. The gene editing system according to claim 10, wherein the CRISPR nuclease polypeptide is a variant of SEQ ID NO: 1, comprising one or more mutations in the HNH nuclease domain at positions D844, H845, and / or N868 relative to SEQ ID NO: 1, and optionally the mutation is at position H845, which is optionally H845A.

12. (a) The mutation in D844 is an amino acid substitution of D844A, D844G, D844L, or D844S, (b) The mutation in H845 is an amino acid substitution of H845A, H845G, H845L, or H845S, (c) The gene editing system according to claim 11, wherein the mutation in N868 is an amino acid substitution of N868A, N868G, N868L, or N868S.

13. The gene editing system according to any one of claims 10 to 12, wherein the CRISPR nuclease polypeptide comprises a bridge helix (BH) domain, a nucleic acid recognition (REC) domain, a phosphate-locked loop (PLL) domain, a wedge (WED) domain, and a PAM interaction (PID) domain, wherein one or more arginine and / or lysine substitutions, optionally an arginine substitution, is located in the BH domain, the REC domain, the PLL domain, the WED domain, the PID domain, or a combination thereof.

14. The gene editing system according to claim 10, wherein the CRISPR nuclease variant is one of those listed in Table 19 or Table 20.

15. The gene editing system according to any one of claims 10 to 14, wherein the CRISPR nuclease polypeptide contains up to 20 arginine and / or lysine substitutions relative to SEQ ID NO: 1, optionally, the CRISPR nuclease polypeptide contains up to 15 arginine and / or lysine substitutions relative to SEQ ID NO: 1, preferably, one or more arginine and / or lysine substitutions are located at positions K736, L784, Q812, N813, I857, and / or A919 of SEQ ID NO: 1, optionally located at position I857, and optionally I857R.

16. The gene editing system according to claim 15, wherein the CRISPR nuclease polypeptide contains at least two arginine and / or lysine substitutions relative to SEQ ID NO: 1, and the at least two arginine and / or lysine substitutions are located at positions K736, L784, Q812, N813, I857, and / or A919 of SEQ ID NO:

1.

17. The CRISPR nuclease polypeptide is subjected to arginine and / or lysine substitutions at the following positions relative to SEQ ID NO: 1: (a) I857, L784, and K736, (b) I857, A919, and K736, (c) I857, N813, and L784, (d) I857, L784, and A919, (e) I857, N813, and K736, (f) I857 and N813, (g) L784, A919, and K736, (h) I857 and L784, and (i) The gene editing system according to claim 16, comprising I857 and A919.

18. The CRISPR nuclease polypeptide is subjected to the following arginine substitutions in SEQ ID NO: 1: (a) I857R, L784R, and K736R, (b) I857R, A919R, and K736R, (c) I857R, N813R, and L784R, (d) I857R, L784R, and A919R, (e) I857R, N813R, and K736R, (f) I857R and N813R, (g) L784R, A919R, and K736R, (h) I857R and L784R, and (i) Including I857R and A919R, The gene editing system according to claim 17, wherein the CRISPR nuclease polypeptide optionally comprises the arginine substitution of (a).

19. The gene editing system according to claim 1, wherein the CRISPR nuclease polypeptide comprises, optionally, a nickase mutation at position H845, H845A, and optionally, an arginine and / or lysine substitution at position I857, I857R, relative to SEQ ID NO:

1.

20. The gene editing system according to any one of claims 1 to 19, wherein the CRISPR nuclease polypeptide comprises one or more mutations for reducing the PAM recognition stringency, and optionally, the one or more mutations are located at positions D61, A68, H494, L1117, D1144, S1145, G1227, E1228, S1327, A1332, R1343, R1345, and / or T1347 of SEQ ID NO:

1.

21. The one or more mutations described above (i) One or more arginine and / or lysine substitutions at positions D61, A68, H494, L1117, G1227, S1327, A1332, and / or T1347 of SEQ ID NO: 1, optionally arginine substitutions. (ii) One or more amino acid substitutions at positions D1144, S1145, E1228, R1343, and / or R1345 of SEQ ID NO: 1, optionally D1144L, S1145W, E1228Q, R1343P, R1345V, and / or R1345Q, or The gene editing system according to claim 20, comprising a combination of (iii)(i) and (ii).

22. The following mutation combinations for sequence number 1: (i) L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and A68R, (ii) L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and D61R, or (iii) The gene editing system according to claim 21, comprising L1117R, D1144V, G1227R, E1228F, A1332R, R1345V, T1347R, and H494R.

23. A gene editing system according to any one of claims 10 to 22, wherein the CRISPR nuclease polypeptide comprises (a) one or more mutations in the HNH nuclease domain at positions D844, H845, and / or N868 relative to SEQ ID NO: 1, and optionally one or more mutations where the mutation is at position H845; and (b) one or more arginine and / or lysine substitutions relative to SEQ ID NO: 1, and optionally one or more arginine and / or lysine substitutions where the arginine and / or lysine substitution is at positions I857, L784, and K736.

24. The gene editing system according to any one of claims 10 to 23, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 90% identical to SEQ ID NO:

1.

25. The gene editing system according to claim 24, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 95% identical to that of SEQ ID NO:

1.

26. The gene editing system according to claim 25, wherein the CRISPR nuclease polypeptide comprises an amino acid sequence that is at least 98% identical to that of SEQ ID NO:

1.

27. The gene editing system according to any one of claims 1 to 26, wherein the RT polypeptide is Moloney mouse leukemia virus (MMLV)-RT, and optionally the MMLV-RT comprises the amino acid sequence of SEQ ID NO:

53.

28. The gene editing system according to claim 1, wherein the fusion polypeptide is listed in Table 8 or Table 17.

29. The gene editing system according to any one of claims 1 to 28, wherein the system comprises the fusion polypeptide.

30. The gene editing system according to any one of claims 1 to 29, wherein the system comprises the first nucleic acid encoding the fusion polypeptide.

31. The gene editing system according to claim 30, wherein the first nucleic acid is optionally located on a viral vector.

32. The gene editing system according to claim 31, wherein the first nucleic acid is a first messenger RNA (mRNA).

33. The gene editing system according to any one of claims 1 to 32, wherein the spacer sequence in the gRNA of (b) is 15 to 30 nucleotides long, or optionally 15 to 20 nucleotides long.

34. A gene editing system according to any one of claims 1 to 33, wherein the PAM is 5'-NDR-3' or 5'-NGN-3', where N represents any nucleotide, D represents A, G, or T, and R represents G or A, and optionally the PAM is 5'-NRG-3' or 5'-NRR-3', where N and R are as defined herein, preferably the PAM is 5'-NGG-3', and N represents any nucleotide.

35. The gene editing system according to any one of claims 1 to 34, wherein the scaffold sequence includes a nucleotide sequence that is at least 85% identical to sequence number 2.

36. The gene editing system according to claim 35, wherein the scaffold sequence includes the nucleotide sequence of Sequence ID No.

2.

37. The gene editing system according to any one of claims 1 to 36, wherein the PBS in the RT donor RNA is 5 to 50 nucleotides long, and optionally 5 to 20 nucleotides long.

38. The gene editing system according to any one of claims 1 to 37, wherein the PBS binds to a PBS target site adjacent to or overlapping with the target sequence.

39. The gene editing system according to any one of claims 1 to 38, wherein the PBS target site is adjacent to or overlaps with the target sequence.

40. The gene editing system according to claim 39, wherein the PBS target site is adjacent to the 5' side of the PAM, and optionally, the 3' terminal nucleotide of the PBS target site is located about 2 to 15 nucleotides upstream of the PAM.

41. The gene editing system according to any one of claims 1 to 40, wherein the template sequence in the RT donor RNA is 5 to 100 nucleotides long, or optionally 15 to 25 nucleotides long.

42. The gene editing system according to any one of claims 1 to 41, wherein the template sequence in the RT donor RNA is homologous to the target genomic region, and includes one or more nucleotide variations relative to the target genomic region.

43. The gene editing system according to claim 42, wherein at least one nucleotide variation is located within the target sequence and / or at least one nucleotide variation is located within the PAM.

44. The gene editing system according to any one of claims 1 to 43, wherein the RNA molecule in (b) further comprises 3' end extension.

45. The gene editing system according to any one of claims 1 to 44, wherein the RNA molecule in (b) further comprises a 5' end protection fragment, a 3' end protection fragment, or both, and each of the 5' end protection fragment and the 3' end protection fragment optionally forms a secondary structure which is a hairpin, a pseudoknot, a circularized or triple-stranded structure.

46. (b) The RNA molecule moves from the 5' end to the 3' end, (i) the spacer arrangement, the scaffolding arrangement, the mold arrangement, and the PBS, or (ii) The gene editing system according to any one of claims 1 to 45, comprising the spacer sequence, the scaffold sequence, the template sequence, the PBS, and the 3' extension.

47. The gene editing system according to any one of claims 1 to 46, wherein the system comprises the RNA molecule of (b).

48. The gene editing system according to any one of claims 1 to 46, wherein the system comprises the second nucleic acid encoding the RNA molecule.

49. The gene editing system according to claim 48, wherein the nucleic acid is optionally located on a vector which is a viral vector.

50. The gene editing system according to any one of claims 1 to 49, wherein the system comprises one or more lipid nanoparticles (LNPs) associated with one or more of elements (a) to (b).

51. The gene editing system according to any one of claims 1 to 50, wherein the system comprises one or more viral vectors and one or more adeno-associated virus (AAV) vectors that optionally encode one or more of elements (a) to (b).

52. A pharmaceutical composition comprising a gene editing system according to any one of claims 1 to 51.

53. A kit comprising the elements (a) to (b) of the gene editing system according to any one of claims 1 to 51.

54. A gene editing method comprising delivering the gene editing system according to any one of claims 1 to 51 to a host cell and editing a genomic site targeted by the gRNA of the gene editing system.

55. The gene editing method according to claim 54, wherein the host cells are cultured in vitro.

56. The gene editing method according to claim 55, wherein the host cell is located within the target for which gene editing is required.

57. A fusion polypeptide comprising a CRISPR nuclease polypeptide according to any one of claims 1 to 26 and a reverse transcriptase polypeptide according to claim 1 or claim 27.

58. The fusion polypeptide according to claim 57, comprising the amino acid sequence of SEQ ID NO: 55 or 57.

59. A nucleic acid encoding the fusion polypeptide according to claim 57 or claim 58.

60. The nucleic acid according to claim 59, comprising the nucleotide sequence of sequence number 54, 243, 56, or 244.

61. The nucleic acid according to claim 60, which is a vector and optionally an expression vector.