Dual nuclease systems for genetic editing, methods of producing and using same

By using a chimeric nuclease system, utilizing the I-TevI ​​domain and RNA-guided nuclease domain, combined with viral vector delivery, the efficiency and accuracy issues of gene editing in existing technologies have been solved, achieving highly efficient and precise gene editing results.

CN121712894APending Publication Date: 2026-03-20SPECIFIC BIOLOGICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480054471.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-08
Filing Date
2024-08-22
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing gene editing technologies are limited in their ability to introduce gene deletions of a specific length or to precisely insert DNA sequences into a sufficient number of cells, and suffer from inefficiency and inaccuracy in delivering gene editors and repair templates.

Method used

Employing a chimeric nuclease system, containing polynucleotides encoding the I-TevI ​​domain and the RNA-guided nuclease domain, as well as polynucleotides encoding the first guide RNA and tRNA, efficient and precise gene editing is achieved through sequential arrangement and delivery via viral vectors.

Benefits of technology

It enables efficient and precise gene editing in cells, reduces unwanted editing effects, and improves the predictability and controllability of target site repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121712894A_ABST
    Figure CN121712894A_ABST
Patent Text Reader

Abstract

Provided herein are dual cleavage chimeric nuclease comprising an I-TevI domain and an RNA guided domain, such as Cas9, for use in a gene editing method, and chimeric nuclease systems comprising a single nucleic acid encoding the chimeric nuclease, one or more guide RNAs (gRNAs), one or more tRNAs, and optionally one or more donor polynucleotides. Also provided herein are methods of making and using such chimeric nucleases, chimeric nucleases systems, and nucleic acids encoding such chimeric nucleases and chimeric nucleases systems.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-referencing

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 631,166, filed April 8, 2024, and U.S. Provisional Application No. 63 / 590,771, filed October 16, 2023, and U.S. Provisional Application No. 63 / 578,032, filed August 22, 2023, all of which are incorporated herein by reference in their entirety. background

[0002] Existing gene editing technologies, such as RNA-programmable gene editors (CRISPR-Cas9 and Cas9 fusions), meganucleases, zinc finger proteins, IIS-type restriction endonucleases (FokI and FokI fusions), and TALENs, are limited in their ability to introduce gene deletions of specific lengths or to precisely insert DNA sequences into a sufficient number of cells. For therapeutic gene editing, it is crucial to deliver the gene editor to the target tissue or cells in an organism and to effectively and precisely modify the target site. To disrupt genes using an RNA-programmable gene editor, it is important that the gene editor and one or more guide RNAs are co-delivered into the cell, and that the gene editor causes predictable deletions or insertions of small and / or large sequences to achieve the desired editing outcome. To insert new sequences using an RNA-programmable gene editor, it is important that the gene editor, one or more guide RNAs, and one or more repair templates are co-delivered into the cell, and that the gene editor replaces or inserts sequences at the target site while producing minimal collateral effects at the target site, such as additional insertions and deletions, to achieve the desired repair outcome.

[0003] Current gene editing technologies typically rely on inefficient repair pathways in cells, such as error-prone non-homologous end joining (NHEJ), homologous directed repair (HDR), or base excision repair (BER), to modify target sites in the genome, resulting in low-target-site repair, cell cycle-specific repair, or unpredictable editing outcomes.

[0004] Several technologies have been developed for in vivo delivery of gene editors and repair templates, including viral and nonviral delivery. However, gene editors are often too large to be co-packaged with all the regulatory elements required for expression and stabilization of the gene editor, guide RNA, and donor DNA sequence or repair template into a single adeno-associated virus vector (AAV). Similarly, for nonviral delivery with messenger RNA, co-expression of all essential elements for effective and predictable target site disruption and repair is limited. For ribonucleoprotein (RNP) delivery, current versions of gene editors are limited to co-delivering the donor nucleic acid sequence along with the protein version of the gene editor into the same cell to ensure high-precision repair. In the case of gene editors expressed in cells after viral delivery, controlling the expression of nucleases over time to control treatment duration and mitigate any unwanted editing due to constitutive expression is necessary. Overview

[0005] Each of the aspects and implementations described herein can be used together unless explicitly or clearly excluded from the context of the implementation or aspect.

[0006] In one aspect, a nucleic acid is provided comprising (i) a polynucleotide encoding a chimeric nuclease comprising an I-TevI ​​domain and an RNA-directed nuclease domain; (ii) a polynucleotide encoding a first guide RNA (gRNA); and (iii) a polynucleotide encoding tRNA; wherein the polynucleotides in (ii)-(iii) are in sequence.

[0007] In one aspect, a nucleic acid is provided comprising (i) a polynucleotide encoding a chimeric nuclease comprising a GIY-YIG nuclease domain and an RNA-directed nuclease domain; (ii) a polynucleotide encoding a first guide RNA (gRNA); and (iii) a polynucleotide encoding tRNA; wherein the polynucleotides in (ii)-(iii) are in sequence.

[0008] In some implementations, the nucleic acid also includes (iv) a stable RNA polynucleotide located downstream of (i).

[0009] In one aspect, a nucleic acid is provided comprising (i) a polynucleotide encoding a first guide RNA (gRNA); and (ii) a polynucleotide encoding tRNA; wherein the polynucleotides in (i)-(ii) are in sequence.

[0010] In some implementations, the nucleic acid further comprises two or more donor polynucleotides tandemly positioned; and a ribozyme polynucleotide; wherein the donor polynucleotides are sequential, and the upstream donor polynucleotide contains a ribozyme cleavage site sequence at its 3' end.

[0011] In some embodiments, the nucleic acid is DNA or RNA. In some embodiments, the DNA is circular plasmid DNA, linear double-stranded DNA, single-stranded DNA, or a chimeric RNA and DNA. In some embodiments, the RNA is mRNA. In some embodiments, the mRNA comprises a nucleic acid mimic selected from the group consisting of peptide nucleic acids (PNA), morpholinonucleotides, cyclohexenylnucleotides (CeNA), and locked nucleic acids (LNA). In some embodiments, the mRNA comprises a modified sugar moiety, optionally wherein the modified sugar moiety is selected from the group consisting of N1-methylpseuuridine, 9-methyladenine, 2'-O-(2-methoxyethyl), 2'-dimethylaminooxyethoxy, 2'-dimethylaminoethoxyethoxy, 2'-O-methyl, and 2'-fluoro. In some embodiments, the mRNA comprises modified nucleotides, optionally wherein the modified nucleotides are selected from the group consisting of: 5-methylcytosine; 5-hydroxymethylcytosine; xanthine; hypoxanthine; 2-aminoadenine; a 6-methyl derivative of adenine; a 6-methyl derivative of guanine; a 2-propyl derivative of adenine; a 2-propyl derivative of guanine; 2-thiouracil; 2-thiothymine; 2-thiocytosine; 5-halogenuridine; 5-halogenuridine; 5-propynyluracil; 5-propynylcytosine; 6-azouracil; 6-azocytosine; 6-azothymine; pseudouracil; 4-thiouracil; 8-halogen; 8-amino; 8-thiol; 8-thioalkyl; 8-hydroxy ; 5-halogen; 5-bromine; 5-trifluoromethyl; 5-substituted uracil; 5-substituted cytosine; 7-methylguanine; 7-methyladenine; 2-F-adenine; 2-aminoadenine; 8-azaguanine; 8-azaadenine; 7-deadenine; 7-deadenine; 3-deadenine; 3-deadenine; tricyclic pyrimidine; phenyloxazinecytidine; phenothiazinecytidine; substituted phenyloxazinecytidine; carbazolecytidine; pyridineindolecytidine; 7-deadenine; 7-deadenine; 2-aminopyridine; 2-pyridone; 5-substituted pyrimidine; 6-azapyrimidine; N-2, N-6 or O-6 substituted purine; 2-aminopropyladenine; 5-propynyluracil; and 5-propynylcytosine. In some embodiments, the mRNA comprises a non-naturally occurring or non-natural internucleotide link selected from the group consisting of: thiophosphates, phosphoramides, non-phosphodiesters, heteroatoms, chiral thiophosphates, dithiophosphates, triphosphates, aminoalkyl phosphates, 3'-alkylphosphonates, 5'-alkylphosphonates, chiral phosphonates, hypophosphonates, 3'-aminophosphonamides, aminoalkylphosphonamides, phosphodiesteramides, thiocarbonylphosphonamides, thiocarbonylalkylphosphonates, thiocarbonylalkyl phosphates, selenophosphates, and boron phosphates.

[0012] In some implementations, the RNA-directed nuclease is selected from the group consisting of Staphylococcus aureus (Staphylococcus aureus). Staphylococcus aureus Cas9 (“saCas9”), Streptococcus pyogenes ( Streptococcus pyogenes Cas9, Acidaminococcus, Cas12, Deltaproteobacteria, CasX, and Eubacterium rectum ( Eubacterium rectale Cas12a. In some embodiments, Cas is inactivated Cas (dCas). In some embodiments, Cas is a nickase or dCas.

[0013] In some embodiments, I-TevI ​​is a nicking enzyme. In some embodiments, the I-TevI ​​nicking enzyme domain contains mutations at amino acid residues R27A, V117F, K135R, and N140S. In some embodiments, I-TevI ​​is inactivated. In some embodiments, the inactivating mutation of I-TevI ​​is the R27A mutation.

[0014] In some implementations, the ribozyme polynucleotide is located within the tRNA polynucleotide.

[0015] In some implementations, the nucleic acid also includes a polynucleotide encoding a second gRNA, optionally wherein the second gRNA polynucleotide is located at the 3' of the tRNA polynucleotide.

[0016] In some embodiments, the sequence of polynucleotides is chimeric nuclease, RNA-stable polynucleotide, guide RNA, tRNA, and guide RNA. In other embodiments, the sequence of polynucleotides is chimeric nuclease, RNA-stable polynucleotide, guide RNA, tRNA / ribozyme, guide RNA, donor polynucleotide 1, and donor polynucleotide 2.

[0017] In some embodiments, the stable RNA polynucleotide comprises the 3' sequence of metastasis-associated lung adenocarcinoma transcript 1 (MALAT) or the 3' end of multiple endocrine vegetation β transcript (MEN β). In some embodiments, the stable RNA sequence comprises the 3' end of a triple-helical RNA structure or an RNA transcript lacking typical polyadenylation signals.

[0018] In some embodiments, the donor polynucleotide is single-stranded or double-stranded. In some embodiments, the donor polynucleotide is DNA or RNA. In some embodiments, one strand of the double-stranded donor polynucleotide is DNA, and the other strand is RNA. In some embodiments, the donor polynucleotide comprises a cis-acting single-stranded RNA polynucleotide annealed to a complementary single-stranded DNA polynucleotide. In some embodiments, the single-stranded or double-stranded donor polynucleotide includes a 2 to 18 nucleotide overhang at its 3' end. In some embodiments, the single-stranded or double-stranded donor polynucleotide includes a 14 nucleotide overhang at its 3' end. In some embodiments, the 3' overhang is a single-stranded RNA polynucleotide.

[0019] In some embodiments, the nucleic acid additionally includes a second guide RNA capable of targeting the 5' region of a donor polynucleotide target site. In some embodiments, the first guide RNA enables a first chimeric nuclease to target and cleave at a first I-TevI ​​or Cas9 target site in the cellular genome, and the second guide RNA enables a second chimeric nuclease to target and cleave at a second I-TevI ​​target site in the cellular genome, wherein the cleavage produces a nucleotide overhang at the second I-TevI ​​target site.

[0020] In some implementations, the 3' end of the donor polynucleotide is complementary to the overhang generated by cleaving the second I-TevI ​​at the second I-TevI ​​target site.

[0021] In some implementations, the guide RNA and donor polynucleotides target mutations in the CFTR gene. In some implementations, the guide RNA and donor polynucleotides target and replace CFTR mutations such as c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), c.1652G>A (p.Gly551Asp), c.1657C>T (p.Arg553Ter), or c.3846G>A (p.Trp1282Ter).

[0022] In some implementations, the guide RNA and donor polynucleotide target mutations in the SERPINA1 gene. In some implementations, the guide RNA and donor polynucleotide target and replace the SERPINA1 c.1096G>A (p.Glu342Lys) mutation.

[0023] In some embodiments, the nucleic acid also includes a promoter. In some embodiments, the promoter is selected from the group consisting of: CMV promoter, SV40 promoter, minimal cytomegalovirus (CMV) promoter, and human extension factor-1α (EF1a) promoter. In some embodiments, the promoter is selected from the group consisting of: muscle-specific synthesis promoters, SPc5-12 neuron-specific promoters hSYN1, aldh1L1, cTNT, α-MHC, SPc5-12, MUC2, Ksp-cadherin, albumin, HAS, insulin, rhodopsin, rNSE, and cone-opsin promoters.

[0024] In some implementations, tRNA includes glycine tRNA, arginine tRNA, asparagine tRNA, aspartic acid tRNA, cysteine ​​tRNA, glutamine tRNA, glutamate tRNA, histidine tRNA, isoleucine tRNA, leucine tRNA, lysine tRNA, methionine tRNA, phenylalanine tRNA, proline tRNA, serine tRNA, threonine tRNA, tryptophan tRNA, tyrosine tRNA, or valine tRNA.

[0025] In some implementations, the ribozyme includes hammerhead ribozyme or hepatitis D virus (HDV) ribozyme.

[0026] In some embodiments, the donor polynucleotide includes a trans-acting double-stranded RNA polynucleotide having a 3'- or 5'-protruding nucleotide; a cis-acting single-stranded RNA polynucleotide having sequence similarity to the target strand of a nuclease; a cis-acting single-stranded RNA polynucleotide having sequence similarity to the non-target strand of a nuclease; one or more binding sites for genomic modifying factors, optionally for binding sites for site-specific recombinases (such as serine recombinases) or LoxP target sites; a repair template for a protein-coding sequence; one or more exons having a splice acceptor and a donor sequence; one or more optional sequences selected from the group consisting of NeoR, BsdR, HygR, PuroR, and BleoR genes; one or more drug-inducible regulatory sequences for controlled gene expression; and / or a two-nucleotide protrusion at the 3' end.

[0027] In some implementations, the nucleic acid also includes a polyadenylation signal. In some implementations, the polyadenylation signal includes simian virus 40 (SV40), α-globin, β-globin, human growth hormone (hGH), bovine growth hormone (BGH), herpes simplex virus type 1 thymidine kinase (HSV TK), or synthetic polyadenylation (Synt poly A) polyadenylation signal.

[0028] In some implementations, the nucleic acid also includes a self-inactivating sequence.

[0029] In some implementations, the nucleic acid is approximately 5 kb in length. In other implementations, the nucleic acid is less than 5 kb in length.

[0030] In some implementations, the nucleic acid is packaged within a virus. In some implementations, the virus is a lentivirus, adeno-associated virus (AAV), adenovirus, retrovirus, or modified herpes simplex virus (HSV). In some implementations, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-DJ, AAV2.5T, or AAVmyo.

[0031] In one aspect, a carrier is provided that comprises the nucleic acid of the present disclosure.

[0032] In one aspect, a viral vector is provided that comprises the nucleic acid of this disclosure.

[0033] In one aspect, an AAV virus is provided, which contains the nucleic acid of this disclosure.

[0034] In one aspect, a cell is provided that comprises the nucleic acid of this disclosure, the vector of this disclosure, the viral vector of this disclosure, or the AAV of this disclosure.

[0035] In one aspect, a composition is provided comprising a chimeric nuclease polypeptide containing an I-TevI ​​domain and an RNA-directed nuclease domain, and the nucleic acid of this disclosure.

[0036] In one aspect, a composition is provided comprising a chimeric nuclease nucleic acid encoding a chimeric nuclease comprising an I-TevI ​​domain and an RNA-directed nuclease domain, and the nucleic acid of this disclosure.

[0037] In some implementations, the chimeric nuclease nucleic acid is mRNA.

[0038] In one aspect, an LNP composition is provided, the LNP composition comprising the nucleic acid of the present disclosure or the composition of the present disclosure.

[0039] In one aspect, a pharmaceutical composition is provided comprising a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure or an AAV of the present disclosure, a composition of the present disclosure or an LNP composition of the present disclosure, and an excipient.

[0040] In one aspect, a method is provided for delivering messenger RNA encoding a chimeric nuclease comprising an I-TevI ​​domain and an RNA-guided nuclease domain to a cell, the method comprising contacting the cell with a polynucleotide and a polynucleotide donor encoding one or more guide RNAs.

[0041] In one aspect, a method for genetically modifying a cell genome is provided, the method comprising contacting the cell with a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure.

[0042] In some embodiments, the modification includes insertion, deletion, substitution, or mutation of the genome. In some embodiments, inserting a donor polynucleotide into the cell's genome results in the removal of the sequence between the I-TevI ​​target site and the Cas9 target site. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell.

[0043] In one aspect, a method is provided for inserting or replacing sequences at chimeric nuclease target sites in the genome of a cell, the method comprising: contacting a cell with a nucleic acid comprising one or more nucleic acid sequences encoding a chimeric nuclease comprising a Cas9 domain and an I-TevI ​​domain, and a nucleic acid comprising a director polynucleotide and a donor polynucleotide; wherein the director polynucleotide and the chimeric nuclease form a complex, and the complex binds to and cleaves genomic DNA at the Cas9 target site and the I-TevI ​​target site; wherein, after I-TevI ​​cleavage, the 3' end of the donor polynucleotide contains at least two bases of complementarity at the 5' end of the I-TevI ​​target site; and wherein the donor polynucleotide is incorporated into the chimeric nuclease target site at the 5' position of the Cas9 target site.

[0044] In some implementations, the 3' end of the guide polynucleotide is linked to the 5' end of the donor polynucleotide.

[0045] In some embodiments, the cellular polymerase is targeted at a chimeric nuclease target site. In some embodiments, the cellular polymerase is polymerase θ.

[0046] In one aspect, a method is provided for replacing at least a portion of the CFTR gene in the genome of a cell, the method comprising contacting a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure with the cell. In some embodiments, a guide RNA and a donor polynucleotide target mutations in the CFTR gene.

[0047] In some implementations, the guide RNA and donor polynucleotides target and replace CFTR c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), c.1652G>A (p.Gly551Asp), c.1657C>T (p.Arg553Ter), or c.3846G>A (p.Trp1282Ter) mutations.

[0048] In one aspect, a method for treating cystic fibrosis in a patient with a corresponding need is provided, the method comprising administering to the patient a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure.

[0049] In some implementations, the guide RNA and donor polynucleotides target mutations in the CFTR gene. In some implementations, the guide RNA and donor polynucleotides target and replace CFTR mutations such as c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), c.1652G>A (p.Gly551Asp), c.1657C>T (p.Arg553Ter), or c.3846G>A (p.Trp1282Ter).

[0050] A method is provided for replacing at least a portion of the SERPINA1 gene in the genome of a cell, the method comprising contacting a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure with the cell.

[0051] In some implementations, the guide RNA and donor polynucleotide target mutations in the SERPINA1 gene. In some implementations, the guide RNA and donor polynucleotide target and replace the SERPINA1 c.1096G>A (p.Glu342Lys) mutation.

[0052] In one aspect, a method is provided for treating α-1-antitrypsin deficiency in patients with corresponding needs, the method comprising administering to the patient a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure. In some embodiments, a guide RNA and a donor polynucleotide target a mutation in the SERPINA1 gene. In some embodiments, the guide RNA and the donor polynucleotide target and replace the SERPINA1 c.1096G>A (p.Glu342Lys) mutation.

[0053] In one aspect, a method is provided for replacing at least a portion of a DMPK gene in the genome of a cell, the method comprising contacting a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure with the cell. In some embodiments, one or more guide RNAs target mutations in the DMPK gene. In some embodiments, one or more guide RNAs target CAG triplet polynucleotide sequences in the 3' untranslated region of the DMPK gene.

[0054] In one aspect, a method is provided for treating type 1 myotonic dystrophy in patients with corresponding needs, the method comprising administering to the patient a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure.

[0055] In one aspect, a method is provided for replacing at least a portion of the C9ORF72 gene in the genome of a cell, the method comprising contacting a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure with the cell. In some embodiments, one or more guide RNAs target mutations in the C9ORF72 gene. In some embodiments, one or more guide RNAs target the GGGGCC hexanucleotide repeat sequence between exon 1a and exon 1b of the C9ORF72 gene.

[0056] In one aspect, a method is provided for treating patients with amyotrophic lateral sclerosis or frontotemporal dementia who have a corresponding need, the method comprising administering to the patient a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure. Brief description of the attached diagram

[0057] The features of this disclosure are specifically set forth in the appended claims. A better understanding of the features and advantages of this disclosure will be obtained by referring to the following detailed description of illustrative embodiments utilizing the principles of this disclosure, along with the accompanying drawings, in which: Figure 1 A schematic diagram of an exemplary AAV box coding element comprising a CMV promoter, Dualase, a MALAT sequence, gRNA1, tRNA', gRNA2, a repair template (RT), and a multi-A sequence (synt[A]) for target site disruption or repair is shown. The minimal CMV promoter drives transcription of this exemplary 3-in-1 construct, which comprises Dualase, long non-coding RNA MALAT-1, a synthetic transfer RNA (tRNA') containing a sequence-specific ribozyme, a gRNA fused to the repair template (RT), and a synthetic multi-A signal sequence (synt[A]). Post-transcriptional, RNA maturation at the 3' end of MALAT and both ends of tRNA' leads to the dissociation of Dualase-MALAT RNA, tRNA, and gRNA. U-rich repeats on MALAT help protect the Dualase RNA from degradation, while tRNA' recognizes and cleaves in regions between RTs and between RT and synt[A] sequences. After translation, Dualase can form a complex (RNP) with gRNA and cleave DNA at the intended target site. The presence of an appropriate in-place repair template (cis-sense and antisense) fused to the 3' end of the gRNA, or abundant free-floating repair template (trans-antsense and trans-sense), serves as a bridge between the two cleavage sites and a local reference template for cellular repair mechanisms.

[0058] Figure 2A schematic diagram of an exemplary mRNA cassette encoding elements for efficient and accurate target site disruption or repair is shown. The T7 promoter drives transcription of an exemplary 3-in-1 construct containing Dualase, a long non-coding RNA MALAT-1, a synthetic transfer RNA (tRNA') containing a sequence-specific ribozyme, a gRNA fused to a repair template (RT), and a synthetic multi-A signal sequence (synt[A]). Upon cellular delivery, RNA maturation at the 3' end of MALAT and both ends of tRNA' leads to the dissociation of mRNAs encoding Dualase-MALAT RNA, tRNA, and gRNA. U-rich repeats on the MALAT sequence protect the Dualase RNA from degradation, while tRNA' recognizes and cleaves regions between the RT and between the RT and synt[A] sequences. Post-translational, Dualase can form a complex (RNP) with the gRNA and cleave at the intended target site. Optionally, a GFP-coding sequence is included to track cellular uptake and is translated and separated from the cellular Dualase protein via an N-terminal T2A (thoseaasigna virus 2A) peptide sequence skipped by ribosomes. The presence of an appropriate repair template (cis-sense and antisense) fused to the 3' end of the gRNA or an abundant free-floating repair template (trans-antisense and trans-sense) serves as a bridge between the two cleavage sites and a local reference template for cellular repair mechanisms.

[0059] Figure 3 A schematic diagram of an exemplary dual-cleavage nuclease ribonucleoprotein (RNP) complex (Dualase) with an integrated guide RNA repair cassette is shown. Dualase and 2-in-1 gRNA-RT are incubated to form the ribonucleoprotein complex (RNP) and delivered to target cells, where they can interact and cleave the intended target site after translocation to the cell nucleus.

[0060] Figure 4A and Figure 4B The use of AAV-Dualase- was demonstrated. AAVS1 An exemplary representation of Dualase's genome editing results is achieved through decomposition tracking of insertions / deletions (TIDE) data. Deletions and insertions of defined lengths, along with unmodified reads, are represented as bar graphs. Figure 4A (The dashed line indicates that P < 0.01.) Figure 4B The results show the difference between the AAV-Dualase- and reference sequences. AAVS1 Alternatively, use the chromatogram of the sequencing sample treated with reagents only, where the difference in peaks is used to determine the restored sequence.

[0061] Figures 5A-5CBar graphs and sequence plots of next-generation sequencing (NGS) results from cell-only, cis-antisense, cis-sense, and trans-RT are shown. NGS data were generated using two bioinformatics tools ( Figure 5A CRIS.Py and Geneious alignment platforms and ( Figure 5B Analysis was performed using CRISPRESSO2. "Insertion / Deletion" indicates the repair event that led to the insertion or deletion. "Exact Repair" indicates a sequencing read that perfectly aligned with the repaired sequence. "Repair + SNP" indicates a sequence alignment with both the repair and another nucleotide change. "Other" indicates other sequence modifications in the sample. Editing results were identified using total reads and aligned reads. The alignment of NGS sequencing reads of cis-antisense repaired sequences is shown in... Figure 5C The text indicates the number and percentage of reads. Dashed lines indicate reads below the generally accepted sequencing detection limit, and I-TevI ​​and Cas9 target sites are indicated above the sequence.

[0062] Figures 6A-6C Images of the gels are shown, depicting the blocking effect in the presence of escalating doses ( Figure 6A Non-homologous end connections (NHEJ), ( Figure 6B ) Homologous directional repair (HDR) or ( Figure 6C In the case of inhibitors of the Rad52-dependent pathway, using lipid transfection to express Dualase and target... AAVS1 AAV (integrated) of gRNA-RT and expression targeting co-delivered with indicated repair templates AAVS1 Editing efficiency of Dualase AAV control (co-delivery) in HEK293 cells by restriction enzyme (RE) digestion. Editing efficiency of samples was analyzed by restriction enzyme (RE) digestion efficiency using unique REs at the insertion repair template site. NHEJ (DNA) = non-homologous double-stranded DNA, HDR (DNA) = double-stranded DNA with homologous arms & NHEJ (RNA) = non-homologous single-stranded RNA.

[0063] Figures 7A-7B Gel images are shown, depicting the repair template and the targeted RNA repair template combined with the repair template via directed polymerase chain reaction (PCR). AAVS1 The directed insertion of the Dualase ribonucleoprotein (RNP) complex, which guides RNA formation. Figure 7AThe results of restriction enzyme digestion of HEK293 cells transfected with Dualase along with separately delivered guide RNA and repair template (co-delivery) or Dualase along with fused guide RNA and RNA repair template (“integrated”), lipid transfection reagent only (“reagent only”), or cell only (“mimetic”) are shown. For each reaction, the results are shown in the diagram. AAVS1 Undigested PCR amplicon (“Sub”) at the site and AAVS1 Digested PCR amplicon at the site (“digested”). Figure 7B The diagram shows that if the repair template is inserted correctly, the PCR product is present in the correct orientation ("correctly oriented RT"), and if the repair template is not inserted correctly, the PCR product is present in the incorrect orientation ("incorrectly oriented RT").

[0064] Figure 8 The images shown depict the targeted insertion of repair templates into the target area via directed polymerase chain reaction (PCR). AAVS1 Dualase integrated mRNA with cis-sense or cis-antisense repair templates. Results of restriction enzyme digestion (BglII digestion) or PCR reactions of HEK293 cells lipid-transfected with Dualase (Tev[VKN]-SaCas9[WT] or Tev[VKN]-SaCas9[D10E]), SaCas9[WT], or lipid-only transfection reagents are shown. For each reaction, the results are shown. AAVS1 Undigested PCR amplicon (“full-length”) at the locus AAVS1 The digestion of the PCR amplicon at the target site (“BglII digest”), the PCR product if the repair template is inserted only in the correct orientation (“correctly oriented RT”), and the PCR product if the repair template is inserted only in the incorrect orientation (“misoriented RT”). Targeting is also shown. AAVS1 Control cells treated with Dualase mRNA at the site and co-delivered repair template (co-delivery).

[0065] Figure 9 Gel images depicting cell repair following Dualase cleavage and treatment with NHEJ and Rad52 inhibitors, as well as trans-dsRNA repair templates, are shown. In the presence of escalating doses of inhibitors blocking non-homologous end joining (NHEJ inhibitors) or Rad52 pathway inhibitors (“Rad52 inhibitors”), cell repair was achieved via lipid transfection with AAV expressing Dualase and targeting… AAVS1 gRNA-RT (integrated) and expression targeting co-delivered with the indicated repair template AAVS1HEK293 cells were treated with a control of Dualase AAV (co-delivered). NHEJ (DNA) = non-homologous double-stranded DNA, HDR (DNA) = double-stranded DNA with homologous arms & NHEJ (RNA) = non-homologous single-stranded RNA.

[0066] Figure 10A A schematic diagram illustrates the precise removal of large repetitive sequences using dual-guided TevCas9 nucleases, where the guide RNA targets the opposing strands of the double-stranded DNA to align the two I-TevI ​​domains head-to-head. A first TevCas9 nuclease (1), containing an inactivated D10A+H557A mutation but an active I-TevI ​​domain (2), is targeted at the 5' end of the repetitive sequence (3) using the guide RNA. A second TevCas9 nuclease (4), containing an inactivated D10A+H557A mutation but an active I-TevI ​​domain, is targeted at the 3' end of the repetitive sequence, but on the opposing strands, such that the I-TevI ​​domains of both nucleases point towards the repetitive sequence. After binding and cleavage (5) via the I-TevI ​​nuclease domains, two complementary 3' 2-nucleotide overhangs (6 and 7) are retained. These complementary overhangs (8) can be repaired by the cell via a non-homologous end joining pathway, and the large repetitive sequence (9) is removed, leaving a defined number of repetitive sequences (10) in the genomic DNA.

[0067] Figure 10B A schematic diagram illustrates the precise removal of large repetitive sequences using dual-guided TevCas9 nucleases, where the guide RNA targets the same strand of the double-stranded DNA to tandemly align the two I-TevI ​​domains. A first TevCas9 nuclease (1), containing an inactivated D10A+H557A mutation but an active I-TevI ​​domain (2), is targeted at the 5' end of the repetitive sequence (3) using the first guide RNA. A second TevCas9 nuclease (4), containing an inactivated D10A+H557A mutation but an active I-TevI ​​domain, is targeted downstream of the 3' end of the repetitive sequence on the same strand, such that the I-TevI ​​domains of the two chimeric nucleases are aligned in the same direction. After binding and cleaving (5) via the two I-TevI ​​nuclease domains, two complementary 3' 2-nucleotide overhangs (6 and 7) are retained. These complementary overhangs (8) can be repaired by the cell via a non-homologous end joining pathway, and the large repetitive sequence (9) is removed, leaving a defined number of repetitive sequences (10) in the genomic DNA.

[0068] Figures 11A-11G This demonstrates the use of an integrated AAV encoding TevCas9 along with dual-guided removal of the 5' and 3' ends of the targeted repeat amplification. DMPKThe result of repeated amplification of the CAG triplet in the 3'-UTR of (myotonic dystrophy protein kinase, also labeled DmpkI). Figure 11A A schematic diagram of the active I-TevI ​​domain and the inactivated Cas9 domain is shown, where the D10A+H557A mutation is targeted by two guide RNAs to orient the I-TevI ​​domain into and downstream of the CAG repeat sequence. CAG repeats greater than 75 result in skeletal muscle diseases such as myotonic dystrophy. Figure 11B A schematic diagram of an expression cassette encoding dual-directed Tev[KTQ]-SaCas9[D10A+H557A] and conditionally cleavable guide RNA is shown, wherein the expressed MiniCMV is the promoter sequence, HHribo is the hammerhead ribozyme sequence, sgRNA1 is the 5' guide RNA targeting sequence, Gly tRNA is the glycine tRNA sequence, sgRNA2 is the 3' guide RNA targeting sequence, HDVribo is the hepatitis D virus ribozyme sequence, and multiple A is the polyadenylated sequence (SEQ ID NO:118). Figure 11C This demonstrates the use of equimolar ratios with the target DMPK A schematic diagram of the expected product from the in vitro lysis reaction of purified Tev[KTQ]-saCas9[D10A+H557A] protein with a dual-guider complex of CAG repeat sequences. The right side shows the expected product from the in vitro lysis reaction of the Tev[KTQ]-saCas9[D10A+H557A] protein with a dual-guider complex. DMPK Agarose gel of the products from the in vitro lysis reaction of CAG repeat DNA substrate. The expected reaction products by size are indicated next to the gel image. Figure 11D The image shows the Sanger sequencing reads of clonal amplicon 1–10, along with expected lysis, generated from cells transduced with an integrated AAV encoding TevCas9 along with dual instructions. [CAG] 51+ The lowercase letter indicates the presence of repeating sequences, and indicates the exact number of CAG repeats after TevCas9 cleavage. Figure 11E This demonstrates the transduction using an integrated AAV TevCas9 and SaCas9 dual guide. DMPK MUT The workflow for quantifying repeat collapse using repeat-induced PCR in CAG repeat cells is as follows: Transduced cells are harvested and genomic DNA is extracted for PCR amplification using primers for both external and internal repeat amplification. The resulting PCR products are analyzed on an Agilent Bioanalyzer, and peaks corresponding to repeats are compared with those for repeat-only cells. DMPK MUTQuantification of CAG repeat cells was performed. In TevCas9-treated cells, more than 65% of repeat sequences collapsed, while in SaCas9-treated cells, ~30% of repeat sequences collapsed. Figure 11F The transduced cells underwent quantitative RT-PCR to quantify the transcription after each treatment using primers specific to each transcript. DMPK MUT and DMPK WT Transcription levels DMPK A graph showing the fold change. Statistical significance of the change is indicated by horizontal bars: ns = no significant change, * = p < 0.05, ** = p < 0.01, or *** = p < 0.001. Figure 11G This diagram shows the fold change in missplicing variants of the CLCN1 gene in exon 6 (E6) in cells transduced with TevCas9 dual-guide (TevCas9dg), TevCas9 single-guide (TevCas9sg), and Cas9 dual-guide (Cas9dg) relative to untreated cells (“cells only”). Missplicing of the human ClC-1 chloride channel causes myotonic dystrophy. Both TevCas9 dual-guide and TevCas9 single-guide significantly reduced missplicing of CLCN1 exon 6, while Cas9 did not significantly alter missplicing. Statistical significance of the changes is indicated by horizontal bars: ns = no significant change, ** = p < 0.01, or *** = p < 0.001. Figures 11H-11T This demonstrates the use of an integrated AAV TevCas9 along with dual guides targeting the 5' and 3' ends for repetitive amplification. C9ORF72 The result of the amplification of the GGGGCC hexanucleotide repeat between exon 1a and exon 1b of the gene. Figure 11H A schematic diagram of the active I-TevI ​​domain and the inactivated Cas9 domain is shown, where the D10A+H557A mutation is targeted by two guide RNAs to orient the I-TevI ​​domain into the GGGGCC repeat sequence. GGGGCC repeats greater than 24 result in motor neuron diseases such as ALS. Figure 11I This demonstrates the use of equimolar ratios with the target C9ORF72 A schematic diagram of the expected product from the in vitro cleavage reaction of purified Tev[VKN]-saCas9[D10A+H557A] protein with a dual-guider complex of the GGGGCC repeat sequence. The right side shows the expected product from the in vitro cleavage reaction of the Tev[VKN]-saCas9[D10A+H557A] protein with the dual-guider complex. C9ORF72Agarose gel of the products from the in vitro lysis reaction of GGGGCC repetitive DNA substrate. The expected reaction products by size are indicated next to the gel image. Figure 11J Micrographs show C9ORF72 repeating amplified motor neuron progenitor cells maturing into motor neurons. Western blot analysis of the motor neuron-specific marker ISL1, comparing the maturation process with the housekeeping gene β-tubulin, also confirms motor neuron maturation. Figure 11K Summary results show that repeat-induced PCR was used to quantify the percentage of large repeat products in motor neurons of motor diseases transduced with integrated AAV encoding TevCas9 or Cas9 along with dual guides. Figure 11L The results of monoclonal sequencing of collapsed repeat sequences from motor neurons transduced by TevCas9 AAV are shown, indicating the exact collapsed repeat products in 4 of the 9 sequences indicated by asterisks (*). Figure 11M This paper presents a summary of the editing percentages determined by deep sequencing analysis of potential off-target sites in AAV-transduced motor neurons expressing the TevSaCas9 dual guide. Hash lines indicate the detection limits of deep sequencing for potential off-target effects; no off-target effects were detected in the transduced cells. Figure 11N This paper presents a summary of the editing percentages determined by deep sequencing analysis of potential off-target sites in AAV-transduced motor neurons expressing the dual guide SaCas9. Off-target events were indicated in chromosome 11 with active SaCas9 (OT24). Statistical significance of changes was marked as nd = no significant difference, or * = p < 0.001. Figure 11O The left side shows the results of Western blot analysis in motor neurons transduced with AAVs expressing TevCas9 dual guides (TevCas9dg), TevCas9 single guides (TevCas9sg), and Cas9 dual guides (Cas9dg), as well as in cell-only controls using anti-C9ORF72 antibody and anti-GAPDH housekeeping gene control antibody. The right side shows a summary of the fold change in C9ORF72 expression in three replicates of Western blot analysis normalized to GAPDH housekeeping gene expression. Treatment of motor neurons with AAVs expressing both TevCas9 and C9ORF72 dual guides resulted in a significant >2-fold increase in expression. Statistical significance of the change was marked as * = p < 0.01. Figure 11PThe left side shows representative anti-multi-GR dipeptide dot blots of cell lysates from AAV-transduced motor neurons expressing TevCas9 dual guides (TevCas9dg), TevCas9 single guides (TevCas9sg), and Cas9 dual guides (Cas9dg), as well as from cells treated with the simulated TevCas9. The right side shows a summary quantification of the relative intensity of the dot blots from cells treated with the simulated TevCas9 c9orf72 dual guide, showing a reduction in the amount of detected multi-GR. Figure 11Q This demonstrates the intracranial (ICV) injection of two doses of the dual-guided expression agent TevCas9 AAV (low = 1 × 10⁻⁶) into mice with GGGGCC repeat amplification of human C9ORF72 copies. 13 One viral genome and height = 2 × 10 13 A schematic diagram (n=3). The right figure shows the results of RT-qPCR detection of TevCas9 expression in cerebellar samples from mice injected 28 days after injection. Control mice were injected with phosphate-buffered saline (PBS). Quantitative reverse transcriptase PCR (RT-qPCR) showed dose-dependent expression of TevCas9 in most cerebellar tissues relative to PBS-injected control mice (n=3). Figure 11R The results of PCR detection of the size of GGGGCC repeat sequences in the cerebellum of humanized mice expressing dual-guided TevCas9 AAV, injected with phosphate-buffered saline (PBS) or two doses, are shown. The left side shows representative agarose gel amplifications of PCR products amplified from genomic DNA and normal-sized repeat sequences. The right side shows a summary of the quantification of normal-sized C9ORF72 PCR products relative to the housekeeping gene (n=3). Figure 11S The results of reverse transcriptase quantitative PCR (RT-qPCR) of C9ORF72 transcripts in the cerebellum of humanized mice expressing dual-guided TevCas9 AAV, injected with PBS or two doses, are shown. Figure 11T The results of Western blotting of C9ORF72 protein in the cerebellum of humanized mice with AAV expressing dual-guided TevCas9, injected with PBS or two doses, are shown.

[0069] Figure 12A A diagram depicting the structure of the self-inactivating Dualase vector is shown. The construct contains nucleotide sequences encoding a promoter, a human codon-optimized TevSaCas9, a polyadenylation signal (“polyA”), and a guide RNA sequence (“gRNA”). The self-inactivation target site can be located in the region between the promoter and the TevSaCas9 site (denoted as “promoter”) or at the end of the TevSaCas9 coding sequence and the beginning of the polyA sequence (denoted as “polyA”). Figure 12BThe images show samples harvested at 24, 48, and 72 hours post-transfection with a product targeting β-2-microglobulin (…). B2M Gel assay results of T7E1 editing in HEK293 cells transfected with plasmid DNA of the TevSaCas9 gene. Lanes marked "None" do not contain self-inactivation sequences in the vector, and lanes marked "Promoter" contain sequences between the promoter and TevSaCas9 sequences. B2M1 The TevSaCas9 target site, marked "multi-A", contains the space between the TevSaCas9 terminus and the multi-A signal sequence. B2M1 TevSaCas9 target site. The level of editing, as determined by the amount of digestion product relative to the substrate, is comparable across constructs over time. Figure 12C Showing from Figure 12B The same treated cells were blotted with the hemagglutinin (α-HA) encoded at the 3' end of TevSaCas9 (“Dualase”). Lanes marked “None” do not contain self-inactivating sequences in the vector, while lanes marked “Promoter” contain sequences between the promoter and TevSaCas9 sequences. B2M1 The TevSaCas9 target site, marked "multi-A", contains the space between the TevSaCas9 terminus and the multi-A signal sequence. B2M1 TevSaCas9 target site. β-actin (α-Act) was imprinted on the same membrane as a loading control. Figure 12D The target was shown B2M The production of multi-A self-inactivated AAV2 virus was achieved, with a titer of 2.4 × 10⁻⁶ measured by quantitative reverse transcription PCR (RT-qPCR). 12 One viral genome (vg) / mL, and a Western blot of the hemagglutinin (α-HA) encoded by the 3' end of TevSaCas9 (“Dualase”) in virus-transduced cells. β-actin (α-Act) was blotted on the same membrane as a loading control. Figure 12E It shows the time from the use of the product within 14 days to generate Figure 12D The plasmid DNA form of the vector that self-inactivates AAV (“self-inactivating”) and the vector without the self-inactivating sequence (“non-self-inactivating”) and the plasmid expressing GFP (“pAAV-GFP”) were used as controls for the Western blot results of saCas9 and GAPDH transfected cells.

[0070] Figures 13A-13C This demonstrates the use of Dualase and targeted AAVS1 Before the guidance RNA cleavage ( Figure 13A ),after( Figure 13B )of AAVS1A schematic diagram of the target site, and its association with precise repair. AAVS1 The relationship between gRNA repair template at the site ( Figure 13C ).

[0071] Figure 14 Western blots of HA-tagged Dualase or SaCas9 protein precipitated from cells treated with Dualase and gRNA-RT or SaCas9 and gRNA-RT are shown. The presence of co-immunoprecipitated Rad52 or PolQ (polymerase θ; Polθ) was determined using antibodies specific to Rad52 (αRad52) or PolQ (αPolQ). Rad52 was co-precipitated with both Dualase-treated and Cas9-treated cells, but PolQ was co-precipitated only with Dualase-treated cells.

[0072] Figure 15A The diagram depicts the purified TevCas9 mutant protein and its targeting... AAVS1 Gel images showing the editing efficiency of restriction enzyme (RE) in HEK293 cells treated with gRNA-RT. The “Tev[WT]-dCas9” lane contains purified TevCas9 protein with an inactivated D10A+H557A mutation but containing an active I-TevI ​​domain. The “Tev[R27A]-dCas9” lane contains purified TevCas9 protein with an inactivated D10A+H557A mutation and an I-TevI ​​domain inactivated with an R27A mutation. Control lanes are also shown: saCas9[WT], reagent only, cells only, and control lanes of cells transfected with Tev-Cas9 3-in-1 mRNA control lipids. Figure 15B Gel images depicting the editing efficiency of restriction enzymes (REs) in HEK293 cells treated with mRNA types of SaCas9 (“Tev-free”) and those containing I-TevI ​​domains with mutations in R27A, V117F, K135R, and N140S (WT = wild-type; D10A+H557A = inactive; D10A = nickase; H557A = nickase).

[0073] Figure 16A This demonstrates the use of rep-gRNA to insert eGFP into the frame. CFTR A schematic diagram of the coding sequence. eGFP is expressed by the promoter of the endogenous gene and is separated from the truncated endogenous protein by the T2A cleavable peptide sequence. Figure 16B Photomicrographs of cells transfected with saCas9 or Dualase mRNA and rep-gRNA lipids encoding eGFP are shown. Phase differences and GFP are shown, including cell-only and rep-gRNA-only controls. Figure 16C An image of a gel with insertion sites amplified using PCR is shown, depicting a larger inserted sequence as if edited and no insertion as if unedited. Figure 16D Demonstrates the use of targeting CFTR F508 site guide RNA and coding target CFTR A schematic diagram of the insertion of GFP's rep-gRNA "bridging rep-gRNA" at the G542 site into eGFP. The bridging rep-gRNA and upstream... CFTR The F508 site has 14 base homologs. The ~28 kilobase (kb) region between the F508 and G542 sites was removed and replaced with a GFP sequence encoded by the bridging rep-gRNA. Figure 16E Photomicrographs of cells transfected with an equimolar mixture of lipids containing Dualase (“TevSaCas9”) or SaCas9 and upstream guide RNA and bridging rep-gRNA are shown. Figure 16F Images of the gel are shown, with unedited target sites indicated by “~28 kb region” and inserted sequences indicated by “removed and replaced sequences”. A summary of the fractions of deep sequencing reads at the 5' and 3' join points of the inserted sequences is also shown, where the repaired join points are correct (“exact repair”) or there is a detected insertion or deletion (“insertion / deletion”).

[0074] Figures 17A-17E The illustration shows the design and prediction of TevSaCas9 protein binding to its target site in a schematic RNA-templated DNA repair (“repair editing”) scheme. Figure 17A This diagram illustrates the repair editing at the TevSaCas9 target site, which has separate DNA targeting components, domain structures, and rep-gRNA (also known as "gRNA-RT"). The I-TevI ​​linker zinc finger is indicated by yellow dots. Cleavage of the I-TevI ​​domain leaves a 2-nt 3' overhang complementary to the 3' end of the rep-gRNA, while cleavage of SaCas9 produces a flat DNA end. Figure 17B It shows AAVS1 The target site, the expected repair product, and the structure of rep-gRNA and its interaction with the target site. Figure 17C The identification and distribution of the Tev CNNNG nuclease motif upstream of the SaCas9 binding site in the human genome are shown. Figure 17D Schematic and predicted processing of the integrated construct, with individual components indicated (not to scale). MALAT, RNA stability element from MALAT1 noncoding RNA; HH, hammerhead ribozyme; T2A, self-splicing peptide linker; IRES, internal ribosome entry site. Figure 17EThe left side shows the in vitro processing of the monolithic mRNA transcript from HEK293 cell extracts, analyzed by incubation time indicated and resolved on a 1% agarose gel. The right side shows eGFP activity in HEK293 cells transfected with the monolithic construct.

[0075] Figures 18A-18H .exist AAVS1 Repair and editing of the safe harbor site. Figure 18A This shows PCR amplification from treated HEK293 cells. AAVS1 Representative agarose gel of BglII digests at the target site. AAV; adeno-associated virus 2; pDNA, plasmid DNA; ribonucleoprotein particles of RNP, TevSaCas9 and rep-gRNA; mRNA, monolithic construct. Figure 18B It shows AAVS1 The editing results at the site were summarized through analysis methods. Figure 18C This shows a summary of the editing results by delivery method. The bar chart is the mean of all replicates, with the whiskers representing the standard deviation from the mean. The dots represent individual replicates. Figure 18D The images show cells from edited data (blue triangles, 6 replicates) or cells simulated by transfection (orange circles, 4 replicates). AAVS1 A graph showing the obvious nucleotide substitutions from deep sequencing of PCR amplicon data, illustrating the fidelity of edits on the edit window. Points represent the mean, with whisker lines indicating the standard deviation from the mean. Targeted editing and repair areas are demarcated by vertical dashed lines. Figure 18E This shows cells edited with TevSaCas9 / rep-gRNA. AAVS1 An exemplary read from deep sequencing of the site, where the reference (wild-type) sequence is shown at the top and the repair product is shown below, with the number of reads for each sequence indicated on the right. Differences relative to the wild-type sequence are indicated by nucleotide coloring. Figure 18F It shows that it is designed to be in AAVS1 A schematic diagram showing the deletion of 13-bp and the generation of rep-gRNA at the HindIII site. Figure 18G The image shows cells from edited or simulated treatments relative to unedited cells. AAVS1 A graph showing the proportion of deep sequencing reads with varying target site lengths. Figure 18H The alignment of sequencing reads is shown, with gaps indicated by dashed lines (-) and the number of reads indicated on the right, where >48% of the reads show the expected 13-bp deletion.

[0076] Figure 19A This diagram illustrates a rep-gRNA used to test base pairing between the 3' end of the rep-gRNA and the I-TevI ​​cleavage overhang. All rep-gRNAs target... AAVS1 Safe harbor site. Figure 19B The effect of rep-gRNAs with different lengths of 3' and 5' overlap on editing efficiency is shown. This is illustrated in cells treated with [the following]. AAVS1 A bar graph showing the ratio of repaired to unrepaired sites as determined by quantitative PCR at the target site. The bars represent the average of two biological replicates, with the error bars indicating the standard deviation from the mean. Dark green bars represent cells treated with TevSaCas9 / rep-gRNA, while light green bars represent cells treated with protected rep-gRNA, in which single-stranded DNA oligonucleotides complementary to the overhangs are co-delivered with the rep-gRNA. Figure 19C The deep sequencing results of TevCas9 / rep-gRNA-treated cells with various overhang lengths are shown, where nucleotide differences from unmodified (WT) sequences are colored. Figure 19D The effect of the last two nucleotides of rep-gRNA on editing is shown. This is illustrated from treated HEK293 cells. AAVS1 A representative gel of BglII digest of the target site amplicon. Figure 19E This illustrates mispairing between the crRNA portions of rep-gRNA. AAVS1 The effect of repair at the site, where rep-gRNA precisely matches the I-TevI ​​overhang (CC) or wobble base pair (UC). Nucleotide mismatches in crRNA relative to the target site are indicated by their position above the gel image.

[0077] Figure 20A A schematic diagram of a method for identifying off-target sites of TevSaCas9 is shown. Figure 20B This demonstrates an indicator mismatch with the gRNA moiety. AAVS1 Distribution map of CNNNG motifs upstream of off-target sites. (This can be used...) AAVS1 The CNNNG supporting repair at the 3'-GG end of the rep-gRNA is highlighted. Figure 20C The off-target and on-target sites are shown. AAVS1 Site alignment. Identical nucleotides are colored, and CNNNG motifs in the upstream region are underlined. Figure 20D The percentage of edits is shown as determined by CRISPRaltRations analysis. Points represent individual experiments. ns = not significant, as calculated by paired t-test of treatment. Figure 20E The results show the effects of AAV transduction (yellow) or mRNA lipid transfection (blue) using TevSaCas9 / AAVS1The percentage of edits at target and off-target sites in HEK293 treated with rep-gRNA was determined by CRISPECTOR analysis. The graph was divided into insertions / deletions mapped near the predicted SaCas9 cleavage site and the Tev cleavage site at each target site. The bars are the average of three replicates, and the whisker lines represent the 95% confidence interval of the edit rate calculated from CRISPECTOR.

[0078] Figure 21 A demonstrates the identification of repair proteins involved in repair editing using small molecule inhibitors. The results are shown in treated HEK293 cells. AAVS1 Agarose gel edited at the target site. On the right, HEK293 cells were treated with inhibitors of B02, SCR7, D-103, or ART558, and transduced with AAV2-capsidated TevSaCas9 and rep-gRNA. Increased inhibitor concentrations are indicated by blue triangles. AAVS1 Target site PCR amplification followed by BglII digestion; BglII digestion indicates successful repair events. Left side: Editing as amplified in HEK293 cells transduced with AAV2-TevSaCas9-rep-gRNA and different co-delivered repair sequences as indicated, along with indicated small molecule inhibitors. AAVS1 The target site was determined by BglII digestion. Figure 21 B. Extracts from HEK293 cells transduced with AAV2-TevSaCas9-rep-gRNA or AAV2-SaCas9-rep-gRNA were repeatedly co-immunoprecipitated with anti-HA antibody, followed by Western blotting with anti-Polθ or anti-RAD52 antibody. Dimensions (in kd) are indicated on the left side of the gel image. Figure 21 C is used for the repair editing model. Cleavage of TevSaCas9-rep-gRNA recruits Rad52, which can bind to RNA or dsRNA or mediate RNA:DNA strand exchange at the I-TevI ​​cleavage site. Formation of a hybrid RNA:DNA structure with 3'-OH recruits Polθ. Repair intermediates are elucidated through a gap-filling repair process involving DNA ligases and unidentified DNA polymerases.

[0079] Figure 22A It shows the target CFTR A schematic diagram of rep-gRNA design for three mutations in the gene: 3nt deletion DF508, G>A conversion mutation G542X, and G>C transversion mutation W1282X. The expected edit products are shown below, where silent substitutions are indicated by hash symbols (#) and corrective edits are indicated by circles marked with x. Figure 22BRepresentative gels are shown for SspI (DF508 or G542X) or HindIII (W1282X) restriction digestion analysis of PCR amplicones in 16HBEge cells. Cells treated with TevSaCas9 / rep-gRNA are indicated by (+), and cells treated with only the transfection reagent are indicated by (-).

[0080] Figure 23A This shows the total 16 HBEs transfected with TevSaCas9 / rep-gRNA. CFTR Alignment of universal reads from deep sequencing of PCR products amplified from G542X cells (cells were not enriched or selected prior to analysis). Figure 23B This demonstrates the formulation of an integrated messenger RNA (mRNA) expressing TevCas9 and a conditionally cleavable repair template guide RNA (rep-gRNA) using lipid nanoparticles for intratracheal delivery to humanized organisms. CFTR G542X A schematic diagram of a mouse. The TevCas9 mRNA also includes a cleavable GFP tag to allow for quantitative uptake into lung cells. Figure 23C The humanization process was demonstrated by treatment with TevCas9 integrated mRNA and GFP-expressing mRNA as controls. CFTR G542X Results of fluorescence-activated cell sorting (FACS) of harvested lung tissue from mice. In TevCas9-treated mice, approximately 31% of the cells indicated in the boxed area were counted as GFP-positive, compared to approximately 25% of the cells in the GFP-only control mice. Figure 23D The results of PCR detection of repair sequences in lung tissue harvested from TevCas9-treated mice and GFP control mice are shown. PCR was calibrated using lung epithelial cell samples with known repair (“calibrated cell controls”). Repair products were detected only in lung tissue from mice treated with TevCas9 integrated mRNA, but not in GFP control mice.

[0081] Figure 24A This demonstrates a method for targeting humans at the amino acid position E342K. SERPINA1 A schematic diagram of the rep-gRNA design of the gene, and an alignment of universal reads from deep sequencing of PCR products amplified from liver fibroblasts of GM11423 patients transfected with TevSaCas9 / rep-gRNA (cells were not enriched or selected prior to analysis). Figure 24B This demonstrates the humanization of mutations by treating an integrated TevSaCas9 / rep-gRNA encapsulated with lipid nanoparticles. SERPINA1 Administration and sampling protocol for E342K mice. Figure 24C The results of enzyme-linked immunosorbent assay (ELISA) of human α-1-antitrypsin (A1AT) in serum samples from treated mice and mice used for the first time in the experiment are shown at indicated times, before and after administration. Figure 24D The results of serum biochemical analysis of aspartate aminotransferase (AST) and alanine aminotransferase (ALT) in serum samples from treated mice and mice used for the first time in the experiment are shown at indicated times, before and after administration. Figure 24E The results of reverse transcriptase quantitative PCR (RT-qPCR) targeting the Cas9 domain of TevCas9 to confirm the expression of TevCas9 in liver samples from mice used for the first time in experiments and from treated mice are shown.

[0082] Figure 25A A schematic diagram is shown showing the removal and replacement of 74 nucleotides between amino acids R553 and G542 in human CFTR using a guide RNA and bridging rep-gRNA strategy. Figure 25B The image shows 16HBEge cells treated with MfeI and SspI, as well as SaCas9+rep-gRNA and TevSaCas9+rep-gRNA digested with restriction enzymes. CFTR Gel photograph of PCR amplicon at the target site. The presence of digestion products indicates repair.

[0083] Figures 26A-26C Exemplary combinations and orientations of conditionally cleavable nucleotide sequences of this disclosure are shown. Figure 26A A schematic diagram is shown of exemplary constructs using human transfer RNA (tRNA) or transfer RNA (tRNA') with a trans-acting ribozyme encoded in an anticodon loop as both DNA and RNA constructs compatible. Cleavage sites for the RNA form of the construct, after translation from the DNA form into mRNA or introduction into cells from the in vitro transcribed RNA form of the construct, are indicated by solid triangles (▲). Cleavage sites for trans-acting elements are indicated by lines ending with arrows. The messenger RNA form of the construct has a 5'-UTR, and the DNA form of the construct has a promoter sequence. Table 10 lists exemplary human transfer tRNAs that undergo tRNA maturation when expressed or transfected into cells. Some exemplary forms include a MALAT stabilizing sequence to stabilize the I-TevI ​​and Cas9 coding sequences after mRNA cleavage. Figure 26B and Figure 26CA schematic diagram is shown of an exemplary construct, such as one compatible with an adeno-associated virus vector, intended solely as a DNA construct. Ribozymes can only encode DNA because they self-cleave during transcription into RNA. Ribozymes cleave only on one side, and their orientation is indicated by the direction of the arrows. Transfer RNA or trans-acting transfer RNA can be combined with ribozymes used to cleave or cleave guide RNA. Figure 26B An exemplary construct without a MALAT stable sequence is depicted, and Figure 26C An exemplary schematic diagram of a construct with a stable MALAT sequence is depicted. In addition to the hammerhead (HH) ribozyme and the hepatitis D virus (HDV) ribozyme, Table 9 lists other exemplary ribozymes active in human cells.

[0084] Figure 27 illustrates the development of a new AAV construct with conditionally spliable elements. Figure 27A A schematic diagram illustrates the original DNA construct and the conditionally cleavable features redesigned in the second DNA construct. The mRNA transcribed from the construct to release targeted mRNA is also shown. DMPK Two guide RNAs for the CAG repeat sequence. Figure 27B It shows in DMPK A schematic diagram of the experimental workflow for testing AAV2 capsidated primitive (“AAV6”) and redesigned constructs (“AAV67”) in CAG repeat fibroblasts is shown. The duration of transduction with constructs before harvesting cells for genomic DNA analysis is illustrated, as well as the polymerase chain reaction (PCR) used to analyze genomic DNA in the presence or absence of CAG repeat sequences. Figure 27C Agarose gels show PCR products of genomic DNA extracted from AAV2-transduced cells infected with 5,000, 10,000, or 25,000 multiplicity of infection (MOI) at indicated times. Lanes labeled "L" contain ladder-like DNA bands, lanes labeled "E" are empty, and lanes labeled "C" contain PCR products of genomic DNA from untreated cells. Figure 27D Results are shown for quantifying the ratio of PCR products with more than 1,000 base pairs (“high molecular weight / MW”) to PCR products with less than 1,000 base pairs (“low molecular weight / LW”) by density determination of agarose gel.

[0085] Figure 28A and Figure 28B This shows a comparison of general readings of deep sequencing of PCR products from HEK293 cells amplified with Tev[V117F+K135R+N140S]-saCas9[D10E] variant and AAVS1-target repair template-guided RNA transfection (cells were not enriched or selected prior to analysis). Figure 28AA summary of the proportions of unmodified, insertion / missing (“NHEJ”), accurately repaired (“HDR”), imperfectly repaired (“imperfect HDR”), and indeterminate reads is shown. Figure 28B The alignment of the most common sequencing reads from transfected cells is shown, indicating the number of reads (read #) and the percentage of reads. Expected repair products are indicated by an asterisk (*).

[0086] Figures 29A-29C It shows the use in B2M Intracellular editing of genes containing chimeric nucleases of I-TevI ​​and erCas12a. Figure 29A A schematic diagram is shown highlighting the chimeric nuclease features, including the I-TevI ​​nuclease domain, the linker domain, and the erCas12a nuclease domain (“Tev-erCas12a”). Figure 29B It shows B2M DNA target sites in the gene are highlighted, showing the erCas12 guide RNA binding site (“B2M6 crRNA”), the adjacent motif of the erCas12 protospacer region (“PAM”), the DNA spacer sequence between the erCas12a site and the I-TevI ​​site (“B2M6 Tev spacer”), and the I-TevI ​​target site (“B2M6 I-TevI ​​site”). I-TevI ​​and erCas12a sense and antisense cleavage sites are also shown. Figure 29C The illustration depicts the use of coded targeting B2M Gel images showing the editing efficiency of plasmid DNA from the Tev-erCas12a, erCas12a, TevSaCas9, and saCas9 genes in HEK293 by T7 endonuclease I (T7E1). PCR amplicon from the B2M target site (“Sub”) and the product of T7E1 digestion (“Edit”) are shown. The percentage of editing is indicated below the gel images. Detailed Explanation

[0087] This article provides, in particular, compositions and methods for chimeric nucleases and chimeric nuclease systems, as well as nucleic acids encoding chimeric nucleases and chimeric nuclease systems.

[0088] This disclosure is based in part on the discovery that the chimeric nucleases and nuclease systems of this disclosure can be expressed from a single nucleic acid using a single promoter in a cell. The chimeric nuclease system of this disclosure and the nucleic acid encoding the chimeric nuclease system can be used to edit the cellular genome. The nucleic acid encoding the chimeric nuclease system of this disclosure can be transcribed in a cell into a single mRNA containing components of the chimeric nuclease system. The mRNA can be further processed in a cell into its individual components, such as a Cas9 nuclease, a guide RNA, and a donor polynucleotide. The chimeric nuclease system of this disclosure may contain one or more chimeric nucleases and one or more guide polynucleotides. Additionally, the chimeric nuclease system may contain a donor polynucleotide. The nucleic acid encoding the chimeric nuclease system of this disclosure may contain a single promoter sequence, a sequence encoding the chimeric nuclease, and a sequence encoding the guide polynucleotide.

[0089] Additionally, nucleic acids may contain nucleic acid sequences that facilitate the processing of single mRNAs into smaller fragments and / or RNA stabilizing sequences that avoid the need for stable polynucleotides. The transcribed mRNA sequence may contain a nucleic acid sequence encoding a chimeric nuclease, one or more guide RNAs and one or more donor polynucleotides, additional sequences such as RNA stabilizing sequences, tRNA, ribozymes, and ribozyme cleavage sites.

[0090] This disclosure is also based in part on the finding that, in the absence of exogenous reverse transcriptase, RNA-templated DNA repair in mammalian cells, chimeric nucleases and chimeric nuclease systems can efficiently and precisely replace sequences in the cellular genome.

[0091] The nucleic acids disclosed herein have the advantage of being small compared to multi-promoter nucleic acid constructs and can be packaged into viral genomes, such as AAV genomes. Another advantage of the nucleic acid design provided herein is that the nucleic acids can be designed to ensure precise 5' and 3' ends of the nucleic acid components after transcription and further processing. Another advantage of the chimeric nucleases and chimeric nuclease systems and nucleic acids disclosed herein is that the system can be used for RNA-mediated repair without the need for delivery of exogenous reverse transcriptase to cells. Due to the small size of the nucleic acids encoding the chimeric nucleases and chimeric nuclease systems of this disclosure, the nucleic acids can be packaged into viral vectors for efficient cellular delivery of these integrated editing systems.

[0092] Before describing embodiments of this disclosure, it should be understood that such embodiments are provided by way of example only, and various alternatives to the embodiments of this disclosure described herein can be used to practice the invention. Many variations, modifications, and substitutions will now occur to those skilled in the art without departing from the invention.

[0093] Unless otherwise defined herein, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Various scientific dictionaries including those containing the terms herein are well known and available to those skilled in the art. While any methods and materials similar to or equivalent to those described herein may be used to practice or test the contents of this disclosure, some preferred methods and materials are described. Therefore, the terms defined below are described more fully with reference to the specification as a whole.

[0094] definition All terms are intended to be understood in the manner that would be understood by one of ordinary skill in the art to which this disclosure pertains. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0095] The following definitions supplement those in the art and are specific to this application, and should not be attributed to any related or unrelated circumstances, such as any jointly owned patents or applications. While any methods and materials similar to or equivalent to those described herein may be used in the practice of testing the contents of this disclosure, preferred materials and methods are described herein. Therefore, the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be restrictive.

[0096] In this application, unless otherwise expressly stated, the use of the singular includes the plural. It must be noted that, unless the context clearly indicates otherwise, as used in the specification, the singular forms “a,” “an,” and “the” include plural references.

[0097] In this application, unless otherwise stated, the use of “or” means “and / or”. The terms “and / or” and “any combination thereof” and their grammatical equivalents are used interchangeably as used herein. The meaning of any combination is specifically contemplated. For illustrative purposes only, the phrases “A, B and / or C” or “A, B, C or any combination thereof” can mean “A alone; B alone; C alone; A and B; B and C; A and C; and A, B and C.” The term “or” can be used conjunctively or disjunctively unless the context explicitly indicates disjunctively.

[0098] Furthermore, the use of the term "including" and other forms such as "include", "includes", and "included" is not restrictive.

[0099] The use of terms such as "some implementation schemes," "implementation schemes," "an implementation scheme," or "other implementation schemes" in the specification means that a particular feature, structure, or characteristic described in connection with the implementation scheme is included in at least some of the implementation schemes of this disclosure, but not necessarily in all implementation schemes.

[0100] As used in this specification and claims, the terms "comprising" (and any form of "comprising," such as "comprise" and "comprises"), "having" (and any form of "having," such as "have" and "has"), "including" (and any form of "including," such as "includes" and "include"), or "containing" (and any form of "containing," such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unlisted elements or method steps. It is contemplated that any embodiments discussed in this specification can be implemented with respect to any method or composition of this disclosure, and vice versa. Furthermore, the compositions of this disclosure can be used to implement the methods of this disclosure.

[0101] The terms "about" or "approximately" mean within an acceptable range of error for a particular value, as determined by a person skilled in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" may, according to practice in the art, mean within one or more standard deviations. Alternatively, "about" may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. In another instance, the quantity "about 10" includes 10 and any quantity from 9 to 11. In yet another instance, the term "about" with respect to a reference value may also include a range of values ​​plus or minus 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%. Alternatively, particularly for biological systems or processes, the term "about" may mean within an order of magnitude of the value, preferably within five times the value, and more preferably within two times the value. When a particular value is described in this application and claims, unless otherwise stated, the term “about” shall be assumed to mean an acceptable range of error for the particular value.

[0102] The term “at least” followed by a number is used in this document to indicate the starting point of a range that begins with that number (which can be a range with or without an upper limit, depending on the definition of the variable). For example, “at least 1” means 1 or more.

[0103] The term "at most" followed by a number is used herein to indicate the endpoint of a range that ends with that number (which can be a range with a lower limit of 1 or 0, or a range without a lower limit, depending on the defined variable). For example, "at most 4" means 4 or less, and "at most 40%" means 40% or less. In this specification, when a range is given as "(first number) to (second number)" or "(first number) - (second number)", this means a range whose lower limit is the first number and whose upper limit is the second number. For example, 25 mm to 100 mm means a range whose lower limit is 25 mm and whose upper limit is 100 mm.

[0104] As may be used herein, the terms “nucleic acid,” “nucleic acid molecule,” “nucleic acid oligomer,” “oligonucleotide,” “nucleic acid sequence,” “nucleic acid fragment,” and “polynucleotide” are used interchangeably and are intended to include, but are not limited to, nucleotides, deoxyribonucleotides, or ribonucleotides, or analogs, derivatives, or modifications thereof, that are covalently linked together in polymeric forms of various lengths. Different polynucleotides may have different three-dimensional structures and may perform a variety of known or unknown functions. Non-limiting examples of polynucleotides include genes, gene fragments, exons, introns, intergenic DNA (including, but not limited to, heterochromatin DNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched-chain polynucleotides, plasmids, vectors, sequence-isolated DNA, sequence-isolated RNA, nucleic acid probes, and primers. Polynucleotides useful in the methods of this disclosure may include natural nucleic acid sequences and their variants, artificial nucleic acid sequences, or combinations of such sequences.

[0105] In the context of two or more nucleic acid or peptide sequences, the term "identical" or "percentage of identity" refers to the fact that two or more sequences or subsequences are identical or have a specified percentage of identical amino acid residues or nucleotides when compared and aligned for maximum correspondence. Sequence alignment methods used for comparison are well known in the art. During alignment, the number of matches is determined by counting the number of positions in both sequences where identical nucleotide or amino acid residues are present. The percentage of sequence identity is determined by dividing the number of matches in the alignment by the length of the reference sequence and then multiplying the result by 100. For example, a peptide sequence with 1166 matches when aligned with a test sequence having 1554 amino acids is 75.0% identical to the test sequence (1166 ÷ 1554 * 100 = 75.0). As used herein, vacancies in the alignment do not reduce the percentage of sequence identity. Unless otherwise stated, the optimal alignment of sequences used for comparison was performed using the global alignment algorithm of Needleman and Wunsch, Mol. Biol. 48:443 (1970), as implemented by EMBOSS Needle (on the World Wide Web, at ebi.ac.uk / Tools / psa / emboss_needle / ) (Madeira et al.) Nucleic Acids Res. 50(W1):W276-W279 (2022)). In the implementation, other alignment methods may be used, including but not limited to those described below: Devereux, et al., Nucleic Acids Res. 12:387-95 (1984); Altschul et al., J.Mol. Biol. 215:403-10 (1990) (BLAST); Carrillo and Lipman Siam J. Appl. Math.48(5) (1988); Computational Molecular Biology (Lesk, AM, ed., 1989); Biocomputing Informatics and Genome Projects, (Smith, DW, ed., 1993); Computer Analysis of Sequence Data, Part I, (Griffin and Griffin, ed., 1994); Sequence Analysis in Molecular Biology (von Heijne, 2012); Sequence Analysis Primer (Gribskov and Devereux, J., ed., 1993). Sequence identity was calculated using an implementation of the Needleman-Wunsch algorithm provided by the National Library of Medicine (on the World Wide Web, at blast.ncbi.nlm.nih.gov / Blast.cgi?PAGE_TYPE=BlastSearch&BLAST_SPEC=GlobalAln).

[0106] For example, sequence identity can be determined using standard methods commonly used to compare the similarity of two polypeptide or two polynucleotide sequences. Using computer programs such as the EMBOSS Needle or BLAST, two polypeptide or two polynucleotide sequences are aligned to best match their respective residues (along the full length of one or both sequences, or along a predetermined portion of one or both sequences). The programs provide default open and default empty penalties, as well as scoring matrices such as PAM 250 (the standard scoring matrix; see Dayhoff et al., in Atlas of Protein Sequence and Structure, Vol. 5, Supplement 3 (1978)) that can be used in conjunction with the computer program.

[0107] "Binding" means attachment via covalent or non-covalent bonds. Non-covalent bonds include those formed by van der Waals forces, hydrogen bonds, ionic bonds, encapsulation or physical encapsulation, absorption, adsorption, and / or other intermolecular forces. Binding can be achieved by any useful means, such as enzymatic binding (e.g., enzymatic linkage) or chemical binding (e.g., chemical linkage).

[0108] As used herein, the term "expression" generally refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (such as into mRNA or other RNA transcripts), or the process by which transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene products." If the polynucleotides originate from genomic DNA, expression can include the splicing of mRNA in eukaryotic cells.

[0109] As used herein, “operably linked,” “operable linkage,” “operatively linked,” or their syntactic equivalents generally refer to the juxtaposition of genetic elements (e.g., promoters, enhancers, polyadenylated sequences, etc.) in a relationship that allows them to operate in the intended manner. For example, if a regulatory element (which may include a promoter or enhancer sequence) helps initiate transcription of a coding sequence, then the regulatory element is operatively linked to the coding region. Intercalation residues may exist between the regulatory element and the coding region as long as this functional relationship is maintained.

[0110] As used herein, "vector" generally refers to a macromolecule or macromolecular conjugate that contains or is associated with a polynucleotide and can be used to mediate the delivery of polynucleotides into cells. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery mediators. Vectors typically contain genetic elements, such as regulatory elements, operatively linked to a gene to promote gene expression at a target.

[0111] As used herein, "expression cassette" and "nucleic acid cassette" are generally used interchangeably and refer to a combination of nucleic acid sequences or elements that are expressed together or operatively linked together for expression. In some cases, an expression cassette refers to a combination of a regulatory element and one or more genes operatively linked thereto for expression.

[0112] As used herein, an "engineered" object generally indicates an object that has been modified through human intervention. By way of non-limiting examples: nucleic acids can be modified by altering their sequence to a sequence not found in nature; nucleic acids can be modified by linking a nucleic acid to a nucleic acid to which it does not associate in nature, so that the ligation product has a function not present in the original nucleic acid; engineered nucleic acids can be synthesized in vitro using sequences not found in nature; proteins can be modified by altering the amino acid sequence of a protein to a sequence not found in nature; engineered proteins can acquire new functions or properties. An "engineered" system contains at least one engineered component.

[0113] This disclosure includes variants of any endonuclease described herein that have one or more conserved amino acid substitutions. Such conserved substitutions can be made in the amino acid sequence of a polypeptide without disrupting its three-dimensional structure or function. Conservative substitutions can be achieved by substituting amino acids with each other having similar hydrophobicity, polarity, and R-chain length. Alternatively or additionally, by comparing aligned sequences of homologous proteins from different species, conserved substitutions can be identified by locating amino acid residues that have been mutated between species (e.g., non-conserved residues) without altering the fundamental function of the encoded protein. Such conserved substitution variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any of the endonuclease protein sequences described herein. In some embodiments, such conserved substitution variants are functional variants. Such functional variants may encompass sequences having substituted sequences that do not impair the activity of one or more key active site residues or guide RNA binding residues of the endonuclease.

[0114] Conserved substitutions of functionally similar amino acids are provided from various references (see, for example, Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd edition (December 1993)). The following eight groups each contain amino acids that are conserved substitutions for each other: 1) Alanine (A), glycine (G); 2) Aspartic acid (D), glutamic acid (E); 3) Asparagine (N), glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); 6) Phenylalanine (F), tyrosine (Y), tryptophan (W); 7) Serine (S), threonine (T); and 8) Cysteine ​​(C), Methionine (M).

[0115] The "position" of an amino acid or nucleotide base is represented by a number that sequentially identifies each amino acid (or nucleotide base) in the reference sequence based on its position relative to the N-terminus (or 5' end). Because deletions, insertions, truncations, fusions, etc., must be considered when determining optimal alignment, the number of amino acid residues in the test sequence determined simply by counting from the N-terminus is not necessarily the same as the number of their corresponding positions in the reference sequence. For example, if a variant has a deletion relative to the aligned reference sequence, the amino acid in the variant will not be present at the position corresponding to the deletion site in the reference sequence. If an insertion is present in the aligned reference sequence, the insertion will not correspond to the numbered amino acid position in the reference sequence. In the case of truncation or fusion, there may be amino acid segments in the reference or aligned sequences that do not correspond to any amino acid in the corresponding sequence.

[0116] When used in the context of numbering a given amino acid or polynucleotide sequence, the term "reference number" or "corresponds to" refers to the numbering of residues in a specified reference sequence when the given amino acid or polynucleotide sequence is compared to a reference sequence.

[0117] As used herein, the term "one or more target sites" specifically refers, for the purposes of this invention, to a location in a gene that will be bound and / or cleaved by a nuclease when targeted by a nuclease. Target sites may comprise multiple nucleotides, such as cleavage sites of I-TevI ​​or CRISPR / Cas nucleases, or DNA-binding sites of I-TevI ​​nucleases or CRISPR / Cas guide RNA. Target sites can be located in gene or intergenic regions in any of various cell types and organisms.

[0118] As used herein, the terms “target” or “targets” or “to target” or “targeting” refer to the invention relating to this application, which uses, for example, selected or engineered DNA-binding domains or guide RNA to target or direct nucleases to specific, selected DNA sequences.

[0119] As used herein, the term "viral vector" refers to a tool commonly used to deliver genetic material into cells. This process can take place in a living organism (in vivo) or in cell culture (in vitro). Examples of viral vectors include AAV vectors, lentiviral vectors, and adenovirus vectors.

[0120] As used herein, the term "simultaneously" means the administration of the CRISPR / Cas9 complex with multiple gRNAs at the same or substantially the same time. It should be understood that some procedures can be performed in sequential but closely spaced steps. For example, electroporating cells and then administering donor DNA within approximately one hour.

[0121] It should be understood that, for clarity, certain features of this disclosure described in the context of individual embodiments may also be provided in combination in a single embodiment. Conversely, for brevity, multiple features of this disclosure described in the context of a single embodiment may also be provided individually or in any suitable sub-combination. All combinations of embodiments relating to this disclosure are specifically covered by this disclosure and disclosed herein, as if each and every combination were individually and explicitly disclosed herein. Furthermore, all sub-combinations of various embodiments and their elements are also specifically covered by this disclosure and disclosed herein, as if each and every such sub-combination were individually and explicitly disclosed herein.

[0122] Chimeric nucleases and chimeric nuclease systems This article specifically provides chimeric nucleases and chimeric nuclease systems that contain two or more nucleases and guide RNA. Chimeric nucleases and chimeric nuclease systems can be used to introduce modifications into the cellular genome, such as insertions, deletions, or mutations.

[0123] In some embodiments, the chimeric nuclease system comprises a CRISPR Cas nuclease domain, a guide RNA, and a GIY-YIG nuclease domain. In some embodiments, the chimeric nuclease comprises a CRISPR Cas nuclease domain and a GIY-YIG nuclease domain. In some embodiments, the chimeric nuclease comprises a fusion protein of the CRISPR Cas nuclease domain and the GIY-YIG nuclease domain. In some embodiments, the CRISPR Cas nuclease domain is located at the N-terminus or C-terminus (e.g., the N-terminus) of the GIY-YIG nuclease domain. In some embodiments, the chimeric nuclease comprises a linker between the CRISPR Cas nuclease domain and the GIY-YIG nuclease domain.

[0124] CRISPR-Cas nuclease CRISPR-Cas nucleases are programmable RNA-directed nucleases that have been described as playing a role in the adaptive immune system in microorganisms. Nuclease targeting of a specific target nucleic acid sequence typically requires both of the following: (i) complementary hybridization between the first 6–8 nucleic acids of the target (target seed) and the crRNA; and (ii) the presence of a protospacer adjacent motif (PAM) sequence in the defined vicinity of the target seed. CRISPR-Cas systems are generally organized into 2 classes, 5 types, and 16 subtypes based on shared functional characteristics and evolutionary similarity. A review of the structure, nomenclature, and classification of CRISPR systems is available in Makarova et al., Evolution and classification of the CRISPR-Cas systems. Nature Reviews Microbiology. 2011 June; 9(6): 467–477.

[0125] Type I CRISPR-Cas systems have large multi-subunit effector complexes and are classified into types I, III, and IV. Type I CRISPR systems contain a multi-protein complex called Cascade (a CRISPR-associated complex for antiviral defense), consisting of subunits CasA, B, C, D, and E, and crRNA. The Cascade-crRNA complex recognizes the target nucleic acid by hybridizing it with crRNA. The bound nucleoprotein complex recruits Cas3 helicase / nuclease to facilitate the cleavage of the target nucleic acid. Type III CRISPR systems include the RAMP superfamily of endonucleases (e.g., Cas6), which cleaves the pre-crRNA array with the aid of one or more CRISPR polymerase-like proteins. Type IV CRISPR-Cas systems have effector complexes composed of two genes from a highly reduced group of large subunit nucleases (csfl), Cas5 (csf3), and Cas7 (csf2) RAMP proteins, and, in some cases, genes from smaller subunits; such systems are typically found on endogenous plasmids.

[0126] Type II CRISPR Cas systems typically possess single-peptide, multi-domain nuclease effectors and are categorized into types II, V, and VI. Type II CRISPR Cas systems typically contain a Cas9 nuclease, crRNA, and trans-activating CRISPR RNA (tracrRNA). tracrRNA hybridizes to a repetitive crRNA sequence. The tracrRNA / crRNA complex can associate with a nuclease (e.g., Cas9). The crRNA-tracrRNA-Cas9 complex recognizes the target nucleic acid through hybridization with crRNA. The hybridization of crRNA with the target nucleic acid activates the Cas9 nuclease for cleavage of the target nucleic acid. Type V CRISPR systems contain different groups of Cas-like genes, including Cas12 nucleases such as Cas12a (Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, and Cas14. Type VI CRISPR Cas systems are RNA-directed RNA endonucleases.

[0127] In some embodiments, the chimeric nuclease comprises a Cas9 CRISPR Cas nuclease. Cas9 is a programmable class II CRISPR Cas nuclease that forms a complex with homologous crRNA and tracrRNA. In some embodiments, the chimeric nuclease system comprises a Cas9 nuclease, crRNA, and tracrRNA. In some embodiments, the chimeric nuclease comprises a nuclear localization signal (NLS). In some embodiments, the chimeric nuclease comprises two C-terminal nucleoplasmic protein NLS separated by a human influenza hemagglutinin (HA) sequence, or two SV40 NLS separated by an HA sequence, or an HA sequence followed by two SV40 NLS.

[0128] In some embodiments, the Cas9 nuclease contains conserved amino acid substitutions. In some embodiments, the Cas9 nuclease contains a mutation in the catalytic domain. In some embodiments, the Cas9 nuclease contains a mutation that results in nicking enzyme activity (nCas9). In some embodiments, the Cas9 nuclease contains a mutation that produces catalytically inactive Cas9 (dCas9, inactivated Cas9). In some embodiments, the Cas nuclease is the Cas9 domain.

[0129] In some implementations, the Cas9 nuclease can be derived from or originate from: Staphylococcus aureus ( Staphylococcus aureus Streptococcus pyogenes () Streptococcus pyogenes Streptococcus thermophilus () Streptococcus thermophilus ), Species of Streptococcus ( Streptococcus sp. ), Nocardia dassonvillei ( Nocardiopsis dassonvillei), Streptomyces coccidioides ( Streptomyces pristinespiralis ), green color-producing Streptomyces ( Streptomyces viridochromogenes ), Rosacea ( Streptosporangium pink ), Acid-heat cyclophosphamide ( Alicyclobacillus acidocaldarius ), Pseudomycium-like Bacillus ( Bacillus pseudomycoides ), selenium-reducing Bacillus ( Bacillus selenite reducing ), Thalictrum microbacterium ( Exiguobacterium sibiricum Lactobacillus delbrueckii (), Lactobacillus delbrueckii ), Lactobacillus salivarius ( Lactobacillus salivarius Marine microoscillator bacteria ( Marine microscilla Burkholderia ( ) Burkholderiales bacteria ), Naphthylaminoprotozoa ( Polaromonas naphthalenivorans ), species of the genus *Pteromonas* ( Polaromonas sp. ), *Cyclocarya var. vara* ( Crocosphaera watsonii ), species of the genus *Cymbidium* ( Cyanothece sp. Microcystis aeruginosa ( Microcystis aeruginosa ), species of the genus Synechococcus ( Synechococcus sp. ), Arabinose ( Acetohalobium arabicum ), Ammonite of the Degens Babella pyrolyticus ( Caldicellosiruptor becscii ), Desulfurized candidates Clostridium botulinum ( Clostridium botulinum Clostridium difficile ( Clostridium difficile ), Griffon's bacterium ( Finally big goldfish ), Thermophilic Neisseria ( Natranaerobius thermophilus ), thermophilic propionic acid anaerobic enterobacteria ( Pelotomaculum thermopropionicum ), Thiobacillus thermophila ( Acidithiobacillus caldus ), Acidophilic ferrothiobacillus ( Acidithiobacillus ferrooxidans ), Allochromatium vinous Species of the genus *Hymenobacter* ( Marinobacter sp. ), halophilic nitrosococci ( Nitrosococcus halophilus ), Nitrostrophus warwickii ( Nitrosococcus watsoni ), Pseudomonas aeruginosa ( Pseudoalteromonas haloplanktis ), Ktedonobacter racemifer , Methanohalobium evestigatum Variable fish algae ( Anabaena variable ), Foamy Glomerula ( Foamy nodule ), species of the genus Nostoc ( Nostoc sp. Spirulina macrophylla () Arthrospira maxima Spirulina platensis () Arthrospira platensis Spirulina species Arthrospira sp. ), species of the genus *Cyrtomia* ( Lyngbya sp .), Prototype Microsheatha ( Microcoleus chthonoplastes ), species of the genus Oscillatoria ( Oscillatoria sp. ), Petrotoga mobilis African thermocline bacteria ( Thermosipho africanus )or Acaryochloris marina .

[0130] In some implementations, the Cas9 nuclease is selected from the following Cas9 nucleases: *F. novicida*, *T. denticola*, *Campylobacter jejuni*, *Alicyclobacillus acidoterrestris*, *Prevotella* and *Francisella*, *Acidaminococcus sp.* BV3L6, *Eubacterium rectale*, *SpCas12f1* (497 aa), *AsCas12f1* (422 aa), *S. lugdunensis* (Slu), *S. hyicus* (Shy), *S. microti* (Smi), and *S. pasteuri*. (Spa), marine microbial communities, freshwater microbial communities from Lake Mendota, Planctomycetes, agricultural soil microbial communities from Utah for nitrogen management research (Steercompost 2015), thermophilic microbial communities enriched in rice / straw / compost from the Joint Bioenergy Institute, California, USA (eDNA_2), Eggerthella sp. YY7918, Finegoldia magna ATCC_29328, Lactobacillus rhamnosus LOCK900, Nitratifractor salsuginis, Streptococcus gordonii str. Challis substr. CH1, Tissierellia bacterium KA00581, Turicibacter species or engineered SpyCas9.

[0131] In some implementations, the Cas9 nuclease is derived from Staphylococcus aureus, Streptococcus pyogenes, and Neisseria meningitidis. Neisseria meningitidis Campylobacter jejuni, Streptococcus pasteurellis Streptococcus pasteurianus Clostridium cellulose () Clostridium cellulolyticum ) or thermophilic denitrifying Bacillus ( Geobacillus thermodenitrificans ) T1 .

[0132] In some embodiments, the Cas9 domain is derived from Staphylococcus aureus (SaCas9). In some embodiments, the Cas9 domain is derived from Streptococcus pyogenes (spCas9). In some embodiments, the Cas9 domain is derived from Neisseria meningitidis (NmCas9). In some embodiments, the Cas9 domain is derived from Campylobacter jejuni (CjCas9). In some embodiments, the Cas9 domain is derived from Streptococcus pasteurellii (SpCas9). In some embodiments, the Cas9 domain is derived from Clostridium cellulosum (CcCas9). In some embodiments, the Cas9 domain is derived from Bacillus thermophilus Tl (GtCas9).

[0133] Exemplary Cas9 nucleases and homologous PAM sequences are shown in Table A.

[0134] Table A

[0135] In some embodiments, the Staphylococcus aureus Cas9 domain contains a mutation or amino acid substitution or combination thereof corresponding to a position selected from any of the following: D10, H557, N580, H840, D1135, R1335, T1337, T267, L325, V327, D333, A336, I341, E345, D348, K352, S360, T368, N369, N371, S372, E373, K386, N393, H408, N410, I414, A415, T438, Y467, N471. D485, M489, E506, R409, T510, N515, Y518, A539, ​​F550, N551, S596, T602, A611, I617, T620, R650, G654, N667, R685, K695, I706, K722, A723, K724, M731, F732, K735, S739, P741, E742, E746, Q747, I754, T755, H757, K760, H761, P778, E781, I783, N784, D785, T 786L, L787, Y788, K792, D794, T798, L799, V801, N803, L804, N805, G806, D813, K814, L818, I819, S822, E824, L841, G847, D848, Y857, V875, I876, N884, A888, L890, D894, D895, P897, V903, G920, F924, N929, E936, N937, V941, N942, S943, C945, E947, K951, L 952, S956, N957, Q958, A959, N974, G975, V983, N984, N985, D986, I991, V993, M995, I996, T999, Y1000, R1001, E1002, L1004, E1 005, N1006, M1007, D1009, K1010, R1011, P1012, P1013, I1015, I1016, A1020, S1021, Q1024, K1027, E1039, H1045, I0148, K1050.

[0136] In some implementations, the Staphylococcus aureus Cas9 domain contains mutations or substitutions or combinations thereof corresponding to any one or more of the following: D10A, D10E, H557A, N580A, H840A, D1135E, R1335Q, T1337R, T267A, L325F, V327I, D333G, A336S, I341L, E345D, D348N, K352E, S360A, T368A, N369E, N371E, S372P, E373K, K386T, N393R, H408N, N410S, I414M, A415T, T438S, Y467F, N471K, D485E, M 489F, E506K, R409K, T510E, N515K, Y518F, A539P, F550Y, N551H, S596A, T60 2I, A611S, I617V, T620K, R650K, G654E, N667D, R685K, K695Q, I706V, K722T , A723T, K724N, M731T, F732V, K735Q, S739N, P741L, E742G, E746D, Q747D, I 754D, T755I, H757R, K760Q, H761S, P778I, E781K, I783V, N784D, D785E, T786 L, L787V, Y788H, K792E, D794T, T798R, L799I, V801I, N803S, L804I, N805K, G806N, D813G, K814E, L818I, I819F, S822P, E824G, L841T, G847S, D848N, Y8 57H, V875I, I876V, N884K, A888V, L890R, D894G, D895H, P897L, V903I, G920 D. F924L, N929Y, E936D, N937G, V941I, N942D, S943L, C945A, E947K, K951R, L 952Q, S956N, N957E, Q958K, A959S, N974D, G975K, V983A, N984S, N985D, D98 6G, I991V, V993L, M995F, I996V, T999N, Y1000K, R1001E, E1002D, L1004I, E 1005K, N1006M, M1007N, D1009L, K1010S, R1011T, P1012S, P1013F, I1015L, I1016R, A1020G, S1021K, Q1024K, K1027S, E1039K, H1045K, I0148M, K1050M.

[0137] In some embodiments, Cas9 contains a mutation at residue 10 corresponding to amino acid position SEQ ID NO:36 or SEQ ID NO:86. In some embodiments, Cas9 contains a mutation at residue 557 corresponding to amino acid position SEQ ID NO:36. In some embodiments, Cas9 contains a mutation at residue 580 corresponding to amino acid position SEQ ID NO:36. In some embodiments, Cas9 contains a mutation at residue 650 corresponding to amino acid position SEQ ID NO:36. In some embodiments, Cas9 contains a mutation at residue 840 corresponding to amino acid position SEQ ID NO:86. In some embodiments, Cas9 contains a mutation at residue 1135 corresponding to amino acid position SEQ ID NO:86. In some embodiments, Cas9 contains a mutation at residue 1335 corresponding to amino acid position SEQ ID NO:86. In some embodiments, Cas9 contains a mutation at residue 1337 corresponding to amino acid position SEQ ID NO:86.

[0138] In some embodiments, Cas9 contains a mutation at the residue corresponding to D10 in SEQ ID NO:36 or SEQ ID NO:86. In some embodiments, Cas9 contains a D10E mutation in SEQ ID NO:36 or SEQ ID NO:86. In some embodiments, Cas9 contains a D10A mutation in SEQ ID NO:36 or SEQ ID NO:86. In some embodiments, Cas9 contains an H557A mutation in SEQ ID NO:36. In some embodiments, Cas9 contains an N580A mutation in SEQ ID NO:36. In some embodiments, Cas9 contains an R650K mutation in SEQ ID NO:36. In some embodiments, Cas9 contains an H840A mutation in SEQ ID NO:86. In some embodiments, Cas9 contains a D1135E mutation in SEQ ID NO:86. In some embodiments, Cas9 contains an R1335Q mutation in SEQ ID NO:86. In some embodiments, Cas9 contains a T1337R mutation in SEQ ID NO:86.

[0139] In some embodiments, Cas9 contains the D10E and H557A mutations in SEQ ID NO:36. In some embodiments, Cas9 contains the D10A and H557A mutations in SEQ ID NO:36. In some embodiments, Cas9 contains the D10E and N580A mutations in SEQ ID NO:36. In some embodiments, Cas9 contains the D10A and N580A mutations in SEQ ID NO:36. In some embodiments, Cas9 contains the D10E and H840A mutations in SEQ ID NO:86. In some embodiments, Cas9 contains the D10E and D1135E mutations in SEQ ID NO:86. In some embodiments, Cas9 contains the D10A and D1135E mutations in SEQ ID NO:86. In some embodiments, Cas9 contains the D10E and R1335Q mutations in SEQ ID NO:86. In some embodiments, Cas9 contains the D10A and R1335Q mutations in SEQ ID NO:86. In some embodiments, Cas9 contains the D10E and T1337R mutations in SEQ ID NO:86. In some embodiments, Cas9 contains the D10A and T1337R mutations in SEQ ID NO:86. In some embodiments, Cas9 contains the D10E, D1135E, R1335Q, and T1337R mutations in SEQ ID NO:86. In some embodiments, Cas9 contains the D10E, H840A, D1135E, R1335Q, and T1337R mutations in SEQ ID NO:86.

[0140] In some embodiments, the Cas9 nuclease is a Staphylococcus aureus (SaCas9) nuclease. In some embodiments, SaCas9 contains 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence of SEQ ID NO:36. In some embodiments, SaCas9 contains 90% of the amino acid sequence of SEQ ID NO:36. In some embodiments, SaCas9 contains 91% of the amino acid sequence of SEQ ID NO:36. In some embodiments, SaCas9 contains 92% of the amino acid sequence of SEQ ID NO:36. In some embodiments, SaCas9 contains 93% of the amino acid sequence of SEQ ID NO:36. In some embodiments, SaCas9 contains 94% of the amino acid sequence of SEQ ID NO:36. In some embodiments, SaCas9 contains 95% of the amino acid sequence of SEQ ID NO:36. In some embodiments, SaCas9 contains 96% of the amino acid sequence of SEQ ID NO:36. In some embodiments, SaCas9 contains 97% of the same amino acid sequence as SEQ ID NO:36. In some embodiments, SaCas9 contains 98% of the same amino acid sequence as SEQ ID NO:36. In some embodiments, SaCas9 contains 99% of the same amino acid sequence as SEQ ID NO:36. In some embodiments, SaCas9 contains 100% of the same amino acid sequence as SEQ ID NO:36.

[0141] In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence identical to SEQ ID NO:37. In some embodiments, SaCas9 comprises 90% of the amino acid sequence identical to SEQ ID NO:37. In some embodiments, SaCas9 comprises 91% of the amino acid sequence identical to SEQ ID NO:37. In some embodiments, SaCas9 comprises 92% of the amino acid sequence identical to SEQ ID NO:37. In some embodiments, SaCas9 comprises 93% of the amino acid sequence identical to SEQ ID NO:37. In some embodiments, SaCas9 comprises 94% of the amino acid sequence identical to SEQ ID NO:37. In some embodiments, SaCas9 comprises 95% of the amino acid sequence identical to SEQ ID NO:37. In some embodiments, SaCas9 comprises 96% of the amino acid sequence identical to SEQ ID NO:37. In some embodiments, SaCas9 comprises 97% of the amino acid sequence identical to SEQ ID NO:37. In some embodiments, SaCas9 contains 98% of the same amino acid sequence as SEQ ID NO:37. In some embodiments, SaCas9 contains 99% of the same amino acid sequence as SEQ ID NO:37. In some embodiments, SaCas9 contains 100% of the same amino acid sequence as SEQ ID NO:37.

[0142] In some implementations, SaCas9 contains the same amino acid sequence as SEQ ID NO:38 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.

[0143] In some embodiments, SaCas9 contains 90% of the same amino acid sequence as SEQ ID NO:38. In some embodiments, SaCas9 contains 91% of the same amino acid sequence as SEQ ID NO:38. In some embodiments, SaCas9 contains 92% of the same amino acid sequence as SEQ ID NO:38. In some embodiments, SaCas9 contains 93% of the same amino acid sequence as SEQ ID NO:38. In some embodiments, SaCas9 contains 94% of the same amino acid sequence as SEQ ID NO:38. In some embodiments, SaCas9 contains 95% of the same amino acid sequence as SEQ ID NO:38. In some embodiments, SaCas9 contains 96% of the same amino acid sequence as SEQ ID NO:38. In some embodiments, SaCas9 contains 97% of the same amino acid sequence as SEQ ID NO:38. In some embodiments, SaCas9 contains 98% of the same amino acid sequence as SEQ ID NO:38. In some embodiments, SaCas9 contains 99% of the same amino acid sequence as SEQ ID NO:38. In some implementations, SaCas9 contains the same amino acid sequence as SEQ ID NO:38.

[0144] In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence identical to SEQ ID NO:39. In some embodiments, SaCas9 comprises 90% of the amino acid sequence identical to SEQ ID NO:39. In some embodiments, SaCas9 comprises 91% of the amino acid sequence identical to SEQ ID NO:39. In some embodiments, SaCas9 comprises 92% of the amino acid sequence identical to SEQ ID NO:39. In some embodiments, SaCas9 comprises 93% of the amino acid sequence identical to SEQ ID NO:39. In some embodiments, SaCas9 comprises 94% of the amino acid sequence identical to SEQ ID NO:39. In some embodiments, SaCas9 comprises 95% of the amino acid sequence identical to SEQ ID NO:39. In some embodiments, SaCas9 comprises 96% of the amino acid sequence identical to SEQ ID NO:39. In some embodiments, SaCas9 comprises 97% of the amino acid sequence identical to SEQ ID NO:39. In some embodiments, SaCas9 contains 98% of the same amino acid sequence as SEQ ID NO:39. In some embodiments, SaCas9 contains 99% of the same amino acid sequence as SEQ ID NO:39. In some embodiments, SaCas9 contains 100% of the same amino acid sequence as SEQ ID NO:39.

[0145] In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence of SEQ ID NO:40. In some embodiments, SaCas9 comprises 90% of the amino acid sequence of SEQ ID NO:40. In some embodiments, SaCas9 comprises 91% of the amino acid sequence of SEQ ID NO:40. In some embodiments, SaCas9 comprises 92% of the amino acid sequence of SEQ ID NO:40. In some embodiments, SaCas9 comprises 93% of the amino acid sequence of SEQ ID NO:40. In some embodiments, SaCas9 comprises 94% of the amino acid sequence of SEQ ID NO:40. In some embodiments, SaCas9 comprises 95% of the amino acid sequence of SEQ ID NO:40. In some embodiments, SaCas9 comprises 96% of the amino acid sequence of SEQ ID NO:40. In some embodiments, SaCas9 comprises 97% of the amino acid sequence of SEQ ID NO:40. In some embodiments, SaCas9 contains 98% of the same amino acid sequence as SEQ ID NO:40. In some embodiments, SaCas9 contains 99% of the same amino acid sequence as SEQ ID NO:40. In some embodiments, SaCas9 contains 100% of the same amino acid sequence as SEQ ID NO:40.

[0146] In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence identical to SEQ ID NO:46. In some embodiments, SaCas9 comprises 90% of the amino acid sequence identical to SEQ ID NO:46. In some embodiments, SaCas9 comprises 91% of the amino acid sequence identical to SEQ ID NO:46. In some embodiments, SaCas9 comprises 92% of the amino acid sequence identical to SEQ ID NO:46. In some embodiments, SaCas9 comprises 93% of the amino acid sequence identical to SEQ ID NO:46. In some embodiments, SaCas9 comprises 94% of the amino acid sequence identical to SEQ ID NO:46. In some embodiments, SaCas9 comprises 95% of the amino acid sequence identical to SEQ ID NO:46. In some embodiments, SaCas9 comprises 96% of the amino acid sequence identical to SEQ ID NO:46. In some embodiments, SaCas9 comprises 97% of the amino acid sequence identical to SEQ ID NO:46. In some embodiments, SaCas9 contains 98% of the same amino acid sequence as SEQ ID NO:46. In some embodiments, SaCas9 contains 99% of the same amino acid sequence as SEQ ID NO:46. In some embodiments, SaCas9 contains 100% of the same amino acid sequence as SEQ ID NO:46.

[0147] In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence identical to SEQ ID NO:47. In some embodiments, SaCas9 comprises 90% of the amino acid sequence identical to SEQ ID NO:47. In some embodiments, SaCas9 comprises 91% of the amino acid sequence identical to SEQ ID NO:47. In some embodiments, SaCas9 comprises 92% of the amino acid sequence identical to SEQ ID NO:47. In some embodiments, SaCas9 comprises 93% of the amino acid sequence identical to SEQ ID NO:47. In some embodiments, SaCas9 comprises 94% of the amino acid sequence identical to SEQ ID NO:47. In some embodiments, SaCas9 comprises 95% of the amino acid sequence identical to SEQ ID NO:47. In some embodiments, SaCas9 comprises 96% of the amino acid sequence identical to SEQ ID NO:47. In some embodiments, SaCas9 comprises 97% of the amino acid sequence identical to SEQ ID NO:47. In some embodiments, SaCas9 contains 98% of the same amino acid sequence as SEQ ID NO:47. In some embodiments, SaCas9 contains 99% of the same amino acid sequence as SEQ ID NO:47. In some embodiments, SaCas9 contains 100% of the same amino acid sequence as SEQ ID NO:47.

[0148] In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence identical to SEQ ID NO:48. In some embodiments, SaCas9 comprises 90% of the amino acid sequence identical to SEQ ID NO:48. In some embodiments, SaCas9 comprises 91% of the amino acid sequence identical to SEQ ID NO:48. In some embodiments, SaCas9 comprises 92% of the amino acid sequence identical to SEQ ID NO:48. In some embodiments, SaCas9 comprises 93% of the amino acid sequence identical to SEQ ID NO:48. In some embodiments, SaCas9 comprises 94% of the amino acid sequence identical to SEQ ID NO:48. In some embodiments, SaCas9 comprises 95% of the amino acid sequence identical to SEQ ID NO:48. In some embodiments, SaCas9 comprises 96% of the amino acid sequence identical to SEQ ID NO:48. In some embodiments, SaCas9 comprises 97% of the amino acid sequence identical to SEQ ID NO:48. In some embodiments, SaCas9 contains 98% of the same amino acid sequence as SEQ ID NO:48. In some embodiments, SaCas9 contains 99% of the same amino acid sequence as SEQ ID NO:48. In some embodiments, SaCas9 contains 100% of the same amino acid sequence as SEQ ID NO:48.

[0149] In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence identical to SEQ ID NO:49. In some embodiments, SaCas9 comprises 90% of the amino acid sequence identical to SEQ ID NO:49. In some embodiments, SaCas9 comprises 91% of the amino acid sequence identical to SEQ ID NO:49. In some embodiments, SaCas9 comprises 92% of the amino acid sequence identical to SEQ ID NO:49. In some embodiments, SaCas9 comprises 93% of the amino acid sequence identical to SEQ ID NO:49. In some embodiments, SaCas9 comprises 94% of the amino acid sequence identical to SEQ ID NO:49. In some embodiments, SaCas9 comprises 95% of the amino acid sequence identical to SEQ ID NO:49. In some embodiments, SaCas9 comprises 96% of the amino acid sequence identical to SEQ ID NO:49. In some embodiments, SaCas9 comprises 97% of the amino acid sequence identical to SEQ ID NO:49. In some embodiments, SaCas9 contains 98% of the same amino acid sequence as SEQ ID NO:49. In some embodiments, SaCas9 contains 99% of the same amino acid sequence as SEQ ID NO:49. In some embodiments, SaCas9 contains 100% of the same amino acid sequence as SEQ ID NO:49.

[0150] In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence identical to SEQ ID NO:50. In some embodiments, SaCas9 comprises 90% of the amino acid sequence identical to SEQ ID NO:50. In some embodiments, SaCas9 comprises 91% of the amino acid sequence identical to SEQ ID NO:50. In some embodiments, SaCas9 comprises 92% of the amino acid sequence identical to SEQ ID NO:50. In some embodiments, SaCas9 comprises 93% of the amino acid sequence identical to SEQ ID NO:50. In some embodiments, SaCas9 comprises 94% of the amino acid sequence identical to SEQ ID NO:50. In some embodiments, SaCas9 comprises 95% of the amino acid sequence identical to SEQ ID NO:50. In some embodiments, SaCas9 comprises 96% of the amino acid sequence identical to SEQ ID NO:50. In some embodiments, SaCas9 comprises 97% of the amino acid sequence identical to SEQ ID NO:50. In some embodiments, SaCas9 contains 98% of the same amino acid sequence as SEQ ID NO:50. In some embodiments, SaCas9 contains 99% of the same amino acid sequence as SEQ ID NO:50. In some embodiments, SaCas9 contains 100% of the same amino acid sequence as SEQ ID NO:50.

[0151] In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence encoded by SEQ ID NO:15. In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence encoded by SEQ ID NO:16. In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence encoded by SEQ ID NO:17. In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence encoded by SEQ ID NO:18. In some embodiments, SaCas9 comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence encoded by SEQ ID NO:19.

[0152]

[0153] In some embodiments, the Cas9 nuclease is a Streptococcus pyogenes (SpCas9) nuclease. In some embodiments, SpCas9 contains 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence of SEQ ID NO:86. In some embodiments, SpCas9 contains 90% of the amino acid sequence of SEQ ID NO:86. In some embodiments, SpCas9 contains 91% of the amino acid sequence of SEQ ID NO:86. In some embodiments, SpCas9 contains 92% of the amino acid sequence of SEQ ID NO:86. In some embodiments, SpCas9 contains 93% of the amino acid sequence of SEQ ID NO:86. In some embodiments, SpCas9 contains 94% of the amino acid sequence of SEQ ID NO:86. In some embodiments, SpCas9 contains 95% of the amino acid sequence of SEQ ID NO:86. In some embodiments, SpCas9 contains 96% of the amino acid sequence of SEQ ID NO:86. In some embodiments, SpCas9 contains 97% of the same amino acid sequence as SEQ ID NO:86. In some embodiments, SpCas9 contains 98% of the same amino acid sequence as SEQ ID NO:86. In some embodiments, SpCas9 contains 99% of the same amino acid sequence as SEQ ID NO:86. In some embodiments, SpCas9 contains 100% of the same amino acid sequence as SEQ ID NO:86.

[0154] In some implementations, the Cas9 domain of Streptococcus pyogenes contains the domain corresponding to SEQ ID.Mutations at any of the following positions of NO:86, or combinations thereof: D10, S29, F32, D39, R40, H41, S42, I48, C80, S87, K112, H113, K132, K141, D147, L158, E171, P176, I186, V189, Q190, Q194, N199, I201, N202, A203, S204, R205, A210, Q228, L229, G231, S245, T249, S254, D261, T270, N295, T300, D304, V308, N309, I312, T333, A337, E345, F352, Q3 54. S355, K356, G366, A367, E396, L398, I414, D428, F429, D435, K468, S469 , E470, T472, E480, A486, S490, F498, K500, N501, N504, K528, V530, E532, G5 33. A538, T555, K570, F575, D605, E611, R629, E634, T638, R655, R664, R671 , K705, E706, Q709, K710, S714, G7115, G717, H721, H723, A725, N726, V743, L 747, V748, K772, K775, N776, I788, G792, K797, Y799, T804, N808, L811, R82 0. N831, R832, ​​V842, L847, N869, E874, N881, Q885, N888, T893, L911, Y945, D946, L949, E952, A1023, Y1036, G1067, G1077, R1078, N1093, R1114, N1115 , D1117, A1121, D1125, P1128, K1129, V1146, S1154, S1159, L1164, S1172, N1 177, P1178, I1179, D1180, K1211, M1213, G1218, N1234, E1243, K1244, E125 3. E1260, K1263, H1264, E1271, Q1272, E1275, V1290, L1291, S1292, A1293, N 1295, H1297, R1298, D1299, K1300, R1303, E1307, N1308, I1309, I1310, H13 11. L1312, L1315, T1316, N1317, Y1326, D1328, V1342, A1345, I1360, S1363.

[0155] In some implementations, the Cas9 domain of Streptococcus pyogenes contains the domain corresponding to SEQ ID. Mutations of any one of the following or combinations thereof in NO:86: D10E, D10A, S29T, F32M, D39N, R40K, H41Q, S42T, I48L, C80R, S87A, K112D, H113N, K132N, K141E, D147E, L158V, E171Q, P176S, I186K, V189L, Q190H, Q194E, N199R, I201L, N202E, A203E, S204I, R205K, A210G, Q228A, L229F, G231N, S245A, T249M, S254A, D261N, T270S, N295K. T300I, D304G, V308A, N309D, I312V, T333A, A337V, E345K, F352S, Q354K, S355T, K356T, G366K, A367T, E396D, L398F, I414V, D428A, F429Y, D435E, K 468Q, S469R, E470N, T472A, E480D, A486T, S490L, F498V, K500E, N501H, N 504T, K528R, V530I, E532D, G533E, A538E, T555A, K570Q, F575C, D605E, E6 11D, R629K, E634K, T638K, R655H, R664K, R671K, K705V, E706D, Q709K, K7 10A, S714F, G7115E, G717K, H721K, H723Q, A725S, N726A, V743I, L747I, V7 48I, K772Q, K775R, N776R, I788M, G792R, K797E, Y799H, T804A, N808D, L8 11R, R820K, N831D, R832H, V842I, L847I, N869D, E874A, N881S, Q885R, N88 8K, T893S, L911A, Y945H, D946G, L949P, E952A, A1023G, Y1036R, G1067E, G1077E, R1078K, N1093T, R1114G, N1115E, D1117A, A1121P, D1125G, P1128 T, K1129T, V1146I, S1154T, S1159P, L1164V, S1172N, N1177D, P1178S, I1 179V, D1180S, K1211R, M1213L, G1218T, N1234H, E1243D, K1244T, E1253K,E1260D、K1263Q、H1264Y、E1271D、Q1272W、E1275H、V1290L、L1291R、S1292A、A1293T、N1295E、H1297N、R1298T、D1299H、K1300L、R1303S、E1307D、N1308S、I1309M、I1310L、H1311N、L1312A、L1315F、T1316S、N1317R、Y1326F、D1328N、V1342I、A1345S、I1360L、S1363N。、

[0156]

[0157] In some embodiments, the Cas9 domain is the Neisseria meningitidis Cas9 nuclease. In some embodiments, the Neisseria meningitidis Cas9 domain comprises the amino acid sequence listed in SEQ ID NO:148. Other Neisseria meningitidis Cas9s can be found at www.uniprot.org / uniprot / , with accession numbers C9X1G5, A1IQ68, E0NB23, A9M1K5, or C6S593.

[0158] In some embodiments, the Neisseria meningitidis Cas9 domain contains a mutation or combination thereof corresponding to any of the following positions in SEQ ID NO:148: I9, D16, D30, E31, A94, I103, P124, N164, I213, G229, T241, S376, E393, G454, K471, G490, D660, C665, K764, T770, P803, A841, H842, K843, D844, L846, R847, K854, H855, N856, K858, K862, W865, E868, I869. , A872, D873, N876, Y880, G883, I886, E887, E890, R895, A898, Y899, G900, G901, N902, A903, K904, Q905, D908, N912, K917, G919, L921, V927, K929, T930, E932, S933, L936, L937, N938, K939, K940, Y943, T944, G949, D950, C958, K965, N 966, Q967, F969, A975, E980, N981, I986, D987, C988, K989, G990, Y991, R992, I993, D994, Y997, T998, C1000, S1002, H1004, K1005, Y1006, A1010, F1011, Q1012, K1013, D1014, E1015, K1018, V1019, E1020, F1021, A1022, Y1024, I1025, N1026, C1027, D1028, S1029, S1030, N1031, R1033, F1034, Y1035, L1036, A1037, W1038, K1041, G1042, K1044, E1045, Q1046, Q1047, F1048, R1049, I1050, S1051, T1052, Q1053, N1054, L1055, V1056, L1057, I1058, Y1061, V1063, N1064.

[0159] In some embodiments, the Neisseria meningitidis Cas9 domain contains a mutation or combination thereof corresponding to any one of the following in SEQ ID NO:148: I9M, D16E, D30E, E31K, A94D, I103V, P124C, N164D, I213N, G229D, T241A, S376T, E393K, G454C, K471E, G490C, D660E, C665R, K764E, T770A, P803S, A841Q, H842G, K843H, D844E, L846V, R847K, K854R, H855L, N856D, K858G, K862L, W865P, E868Q, I869L, A 872K, D873G, N876K, Y880R, G883E, I886P, E887K, E890E, R895Q, A898T, Y899H, G900K, G901D, N902D, A903P, K904T, Q905K, D908A, N912E, K917Y, G919T, L921Q, V927I, K929Q, T930V, E932K, S933T, L936W, L937V, N938R, K939N, K940H, Y943N, T944G, G949A, D950T, C958E, K965G , N966G, Q967K, F969Y, A975S, E980K, N981G, I986R, D987A, C988V, K989V, G990A, Y991F, R992K, I993D, D994E, Y997F, T998E, C1000R, S10 02I, H1004Y, K1005A, Y1006N, A1010K, F1011L, Q1012T, K1013A, D1014K, E1015K, K1018N, V1019E, E1020F, F1021L, A1022G, Y1024F, I102 5V, N1026S, C1027L, D1028N, S1029R, S1030A, N1031T, R1033A, F1034I, Y1035D, L1036I, A1037R, W1038T, K1041T, G1042D, K1044T, E1045 K, Q1046G, Q1047E, F1048Q, R1049S, I1050V, S1051G, T1052V, Q1053K, N1054T, L1055A, V1056L, L1057S, I1058F, Y1061N, V1063I, N1064D.

[0160] In some embodiments, the RNA-directed nuclease Neisseria meningitidis Cas9 domain comprises an amino acid sequence having at least 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO:148. In some embodiments, the RNA-directed nuclease Neisseria meningitidis Cas9 domain comprises an amino acid sequence having sequence identity between 85%-90%, 90%-95%, 95%-97%, 97%-98%, or 98%-99% with SEQ ID NO:148.

[0161]

[0162] In some embodiments, the Cas9 domain is a Campylobacter jejuni Cas9 nuclease. In some embodiments, the Campylobacter jejuni Cas9 domain comprises the amino acid sequence listed in SEQ ID NO:149. Other Campylobacter jejuni Cas9s can be found at www.uniprot.org / uniprot / , accession numbers Q0P897, A7H5P1, A0A2U0QR81, A0A5Y4VLH1, or A0A381CRM8.In some implementations, the RNA-directed nuclease Campylobacter jejuni Cas9 domain contains a mutation or combination thereof corresponding to any of the following positions in SEQ ID NO:149: L5, A6, D8, I9, S12, S13, F18, S19, L24, K25, I31, T40, E42, L50, L58, A59, R61, L58, L65, H67A N74, K77, L98, I99, P101, N110, L113, A119, A126, R128, I134, K140, A144, K147, Q151, L156, V184, S190, F199, D202, G203, R212, F214, K221, E223, Y232, A235, V243, S247, D251, P256, L261, T269, N276, N277, L285, T287, L2 91. K300, T305, Q308, L312, G314, Y335, K336, I339, H345, D351, N353, E354, I362, K370, D383E, S384, K391, I3 96. L403, T405, K413, N419, L421, D430, K432, A437, L453, K457, V462, A465, K472, N477, A492, E495, L525, K526 , L527, K531, E532, E542, Q550, E556, H559, Y561, S564, M572, V577, Q581, N587, N596, K600, Q602, K603, Q616, K617, N623, Y624, K633, D634, Y642, N649, D656, L660, D662, K667, V677, E680, K682, L686, H692, T693, V712, I7 14. V722, K723, S736, L739, K742, L747, N751, F756, R763, Q764, E772, K777, A786, E790, F792, Q800, S801, G80 4. L812, E813, V833, I835, T841, Y845, A855, L856, A863, V864, D879, E883, D900, Q902, K927, F928, V971, T972.

[0163] In some implementations, the RNA-directed nuclease Campylobacter jejuni Cas9 domain contains the domain corresponding to SEQ ID NO.Mutations of any one of the following or combinations thereof in NO:149: L5I, A6G, D8N, D8E, I9L, S12A, S13N, F18L, S19R, L24I, K25I, I31V, T40N, E42N, L50E, L58V, A59K, R61K, L58V, L65M, H67A, N74K, K77N, L98T, I99Q, P101I, N110S, L113I, A119S, A126V, R128H, I134S, K140N, A144T, K147E, Q151K, L156M, V184I, S190D, F199L, D202Q, G203E, R212K F214L, K221K, E223K, Y232F, A235P, V243I, S247I, D251N, P256A, L261S, T2 69G, N276K, N277S, L285V, T287E, L291I, K300D, T305S, Q308K, L312I, G314N , Y335L, K336N, I339K, H345T, D351I, N353D, E354S, I362T, K370E, D383E, S 384K, K391N, I396L, L403Q, T405I, K413R, N419E, L421C, D430E, K432S, A437 L, L453I, K457C, V462L, A465D, K472S, N477H, A492K, E495I, L525Q, K526I, L527V, K531E, E532D, E542L, Q550D, E556V, H559Y, Y561R, S564N, M572S, V57 7T, Q581L, N587G, N596E, K600L, Q602A, K603E, Q616R, K617F, N623F, Y624F , K633T, D634E, Y642W, N649S, D656S, L660I, D662E, K667A, V677Q, E680V, ​​K6 82S, L686I, H692N, T693F, V712I, I714V, V722I, K723F, S736K, L739F, K742 N, L747S, N751L, F756L, R763K, Q764E, E772N, K777H, A786T, E790L, F792P, Q 800N, S801T, G804D, L812V, E813K, V833S, I835L, T841K, Y845H, A855S, L85 6T, A863T, V864P, D879N, E883N, D900G, Q902K, K927N, F928Y, V971L, T972S.

[0164] In some embodiments, the Campylobacter jejuni Cas9 domain comprises an amino acid sequence having at least 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO:149. In some embodiments, the Campylobacter jejuni Cas9 domain comprises an amino acid sequence having sequence identity with SEQ ID NO:149 between 85%-90%, 90%-95%, 95%-97%, 97%-98%, or 98%-99%.

[0165]

[0166] In some embodiments, the Cas9 domain is a Pasteurella multocida Cas9 nuclease. In some embodiments, the Pasteurella multocida Cas9 domain comprises the amino acid sequence listed in SEQ ID NO:150. Other Pasteurella multocida Cas9s can be found at www.uniprot.org / uniprot / , accession number F5X275.

[0167] In some embodiments, the Pasteurella Cas9 domain contains a mutation or combination thereof corresponding to any of the following positions in SEQ ID NO:150: D11, E85, A88, T92, E96, Y100, T109, D110, D113, E115, R116, D125, I127, K128, E132, S147, I185, A187, K228, Y229, T232, M255, S271, N273, A294, A327. , E355, K357, N379, T380, S382, A385, D439, R440, S464, H469, Y519, I528, N569, I581, A60 7. K632, D633, H635, E636, A647, D648, T703, P705, K712, S713, A724, V750, D882, S951, D9 77. E979, S1014, H1027, I1030, E1081, D1082, D1086, K1088, S1089, N1090, R1092, T1093, I1094, C1095, A1138, Y1139, D1141, T1142, F1158, A1168, E1190, E1198, H1202, I1204, R1 205, I1210, K1224, S1232, M1240, V1241, I1242, P1243, G1424, K1248, Q1254, N1257, S125 8. T1262, K1263, Y1264, D1266, A1270, K1277, D1284, L1288, V1302, N1316, T1346, I1374.

[0168] In some embodiments, the Pasteurella Cas9 domain contains a mutation or combination thereof corresponding to any one of the following in SEQ ID NO:150: D11E, D11A, E85D, A88T, T92A, E96D, Y100Q, T109D, D110N, D113N, E115D, R116S, D125E, I127D, K128A, E132K, S147T, I185L, A187T, K228N, Y229N, T232K, M255T, S271T, N273E, A294S, A32 7V, E355K, K357Q, N379G, T380I, S382T, A385N, D439E, R440E, S464A, H469R, Y519F, I528V, N569D, I581V, A607S, K632R, D633E, H635Q, E636Q, A647K, D648Q, T703A, P705S, K712E, S713A, A724T, V750I, D882G, S951 R, D977E, E979K, S1014P, H1027R, I1030V, E1081G, D1082E, D1086N, K1088R, S1089T, N1090D, R1092E, T10 93K, I1094V, C1095R, A1138V, Y1139L, D1141E, T1142P, F1158L, A1168T, E1190K, E1198K, H1202Q, I1204V, R1205Q, I1210M, K1224R, S1232T, M1240I, V1241M, I1242L, P1243S, G1424A, K1248A, Q1254H, N1257G, S12 58N, T1262A, K1263E, Y1264H, D1266K, A1270E, K1277E, D1284N, L1288V, V1302A, N1316D, T1346N, I1374L.

[0169] In some embodiments, the Pasteurella Cas9 domain comprises an amino acid sequence having at least 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO:150. In some embodiments, the Pasteurella Cas9 domain comprises an amino acid sequence having sequence identity with SEQ ID NO:150 between 85% and 90%-95%, 95%-97%, 97%-98%, or 98%-99%.

[0170]

[0171] In some embodiments, the Cas9 domain is the Clostridium cellulose-degrading Cas9 domain. In some embodiments, the Clostridium cellulose-degrading Cas9 domain comprises the amino acid sequence listed in SEQ ID NO:151. Other Clostridium cellulose-degrading Cas9 domains can be found at www.uniprot.org / uniprot / , accession number B8I085.

[0172] In some embodiments, the Clostridium cellulose Cas9 domain contains a mutation or combination thereof corresponding to any of the following positions in SEQ ID NO:151: T4, D10, V9, D20, K21, I27, C33, K36, A47, A49, S64, Q65, E102, L103, T122, I124, K131, D137, R163, G166, I169, F170, V183, D184, I187, E 193. K200, K208, L209, D221, N224, E227, F228, S234, V242, K244, L252, T256, C25 8. S261, V413, M415, K416, R417, K424, Y426, K427, S429, D430, A468, T470, A472, A 478, Q481, K482, L485, A497, L535, W540, R541, E544, G554, P556, I570, Y574, M58 0. Y584, M585, T592, D593, V606, W607, I647, N650, S693, L697, E702, S704, A713, V 714, I715, D776, L847, G850, G853, A854, R860, I900, H904, M905, I906, E921, Q92 3. S929, T930, H931, Q939, N994, I997, N1000, K1001, S1002, I1003, K1005, P1008.

[0173] In some embodiments, the Clostridium cellulose Cas9 domain comprises a mutation or combination thereof corresponding to any one of the following in SEQ ID NO:151: T4S, D10E, V9I, D20N, K21E, I27E, C33I, K36V, A47S, A49P, S64R, Q65H, E102L, L103V, T122V, I124F, K131Q, D137E, R163Q, G166S, I169L, F170L, V183G, D184G, I187T, E19 3S, K200Q, K208A, L209Y, D221K, N224Q, E227S, F228S, S234T, V242I, K244N, L252K, T256K, C258T , S261F, V413K, M415L, K416R, R417N, K424Q, Y426I, K427P, S429H, D430Q, A468S, T470S, A472V, A4 78G, Q481K, K482R, L485S, A497M, L535H, W540Y, R541K, E544Q, G554F, P556S, I570V, Y574I, M580 F, Y584N, M585N, T592A, D593A, V606W, W607F, I647R, N650H, S693K, L697F, E702Q, S704N, A713V, V 714I, I715V, D776E, L847A, G850P, G853A, A854P, R860K, I900V, H904D, M905V, I906L, E921Y, Q92 3E, S929D, T930E, H931Y, Q939P, N994Q, I997P, N1000R, K1001M, S1002N, I1003K, K1005H, P1008K.

[0174] In some embodiments, the Clostridium cellulose-degrading Cas9 domain comprises an amino acid sequence having at least 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO:151. In some embodiments, the Clostridium cellulose-degrading Cas9 domain comprises an amino acid sequence having sequence identity with SEQ ID NO:151 between 85%-90%, 90%-95%, 95%-97%, 97%-98%, or 98%-99%.

[0175]

[0176] In some embodiments, the thermophilic denitrifying Bacillus T1 Cas9 domain comprises the amino acid sequence listed in SEQ ID NO:152. Other thermophilic denitrifying Bacillus T1 Cas9s can be found at www.uniprot.org / uniprot / , accession number A0A1W6VMQ3.

[0177] In some embodiments, the thermophilic denitrifying Bacillus T1 Cas9 domain contains a mutation or combination thereof corresponding to any of the following positions of SEQ ID NO:152: K2, D8, I14, D35, K41, F74, V75, K91, I117, R128, T136, Q151, S152, S156, A161, V164, S171, E178, D179, V185, R192, K195, A199, Y204, I207, V208, A212, H215, S219, F227. T260, V261, V271, G274, I276, A278, L279, D282, I287, K289, H293, F299, V302, N307, R313, L317, L3 18. V331, G337, K341, S348, A354, A355, K356, R359, M372, T377, R380, E395, D399, E404, S416, T441 , R445, N464, E504, S508, M515, Q516, E520, G521, V534, L545, K559, T578, K603, T612, L619, S621, N656, N660, L673, D685, I699, N708, N717, R737, V738, S752, D756, Q771, N777, N792, E793, I811, I8 24. K839, Q845, K848, T849, L895, I902, T908, V929, I943, I946, M948, F990, T995, V1000, Q1014, D1 017, S1019, N1020, G1021, S1024, N1030, N1031, R1035, S1036, I1037, V1067, S1071, A1075, I1079.

[0178] In some embodiments, the *Bacillus thermophilus* T1 Cas9 domain contains a mutation corresponding to any one of positions D8, D179, D282, D399, D685, D756, and D1071. In some embodiments, the RNA-directed nuclease *Bacillus thermophilus* T1 Cas9 domain contains a mutation corresponding to SEQ ID [SEQ ID]. The following mutations or combinations thereof in NO:152: K2R, D8E, D8A, I14V, D35E, K41Q, F74V, V75I, K91E, I117V, R128K, T136S, Q151R, S152A, S156G, A161G, V164I, S171A, E178G, D179E, V185I, R192H, K195R, A199S, Y204F, I207M, V208S, A212K, H215N, S219T, F227V. T260I, V261A, V271I, G274S, I276A, A278G, L279P, D282E, I287L, K289E, H293Q, F299Y, V302I, N307R, R313Y, L317I, L 318V, V331I, G337D, K341Q, S348K, A354K, A355S, K356S, R359L, M372L, T377A, R380H, E395P, D399N, E404N, S416T, T44 1S, R445K, N464T, E504D, S508T, M515T, Q516K, E520D, G521E, V534M, L545H, K559R, T578V, K603R, T612I, L619V, S621 T, N656M, N660S, L673F, D685E, I699V, N708E, N717D, ​​R737K, V738I, S752A, D756E, Q771R, N777H, N792D, E793Q, I811V, I824V, K839T, Q845K, K848A, T849S, L895P, I902V, T908K, V929V, I943V, I946M, M948I, F990L, T995I, V1000G, Q1014K, D1017H, S1019G, N1020T, G1021A, S1024E, N1030C, N1031S, R1035S, S1036G, I1037V, V1067L, S1071A, A1075T, I1079V.

[0179] In some embodiments, the RNA-directed nuclease *Bacillus thermodextrin* T1 Cas9 domain comprises a mutation corresponding to any one of positions D8E, D179E, D282E, D399N, D685E, D756E, and D1071H in SEQ ID NO:152. In some embodiments, the RNA-directed nuclease *Bacillus thermodextrin* T1 Cas9 domain comprises an amino acid sequence having at least 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO:152. In some embodiments, the RNA-directed nuclease *Bacillus thermodextrin* T1 Cas9 domain comprises an amino acid sequence having sequence identity between 85%-90%, 90%-95%, 95%-97%, 97%-98%, and 98%-99% with SEQ ID NO:152.

[0180] In some implementations, the CRISPR Cas nuclease is δ-Proteobacterium CasX, aminococcus Cas12, or Eubacterium rectum Cas12a.

[0181] In some implementations, the CRISPR Cas nuclease contains 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence of C. Cas12, C. δ-Proteobacteria, or C. Cas12a.

[0182]

[0183] In some embodiments, the CasX domain is derived from Planctomycetes bacterium. In some embodiments, the CasX domain is derived from the class Deltaproteobacteria. In some embodiments, the CasX domain contains at least about 85%, 90%, 95%, 97%, 98%, or 99% of the same amino acid sequence as or identical to SEQ ID NO:153. In some embodiments, the CasX domain contains amino acid sequences corresponding to SEQ ID NO:153. Mutations at any of the following positions of NO:153 or combinations thereof: R11, R12, V14, K15, S17, N18, A22, G23, T25, P38, K41, E42, N46, L47, N53, I54, P57, T61, S62, R63, A64, E75, H82, Q89, P104, N106, I113, N199, S124, S125, C133, Y137, N145, D146, H151, S161, R1 65. N177, L180, R202, N205, G215, C219, V236, T241, L248, I254, S269, I290, E291, V297, Q299, I314, E318, Q32 3. L333, E359, D360, K362, Q366, N367, L368, A369, G370, Y371, H404, H409, G410, E411, Y417, V428, E429, S432, K433, L437, S443, A451, I464, A470, I502, L503, I531, G537, L540, N553, I559, S563, V571, N579, H589, S607, L 608, L620, R623, R624, L644, S646, M652, I657, R679, L684, N686, H689, S696, T702, T737, L742, Y744, Q748, M7 51. I753, A771, R777, P792, S818, R823, V824, E826, K827, A832, T833, M836, I839, G841, V846, N860, V862, D86 4. V867, V877, S883, S889, G890, S894, K908, N913, F916, T918, R936, Q938, Y940, K942, S963, R966, K967, K968.

[0184] In some implementations, the CasX domain contains any one or more mutations or substitutions or combinations thereof corresponding to SEQ ID NO:153, including: R11K, R12K, V14S, K15A, S17N, N18A, A22V, G23S, T25S, P38D, K41K, E42K, N46K, L47R, N53V, I54M, P57V, T61N, S62A, R63A, A64N, E75K, H82Q, Q89K, P104S, N106K, I113K, N199K, S124T, S125A, C133G, Y137F, N145S, D146E, H151Y, and S161A. R165K, N177S, L180A, R202K, N205T, G215A, C219Y, V236I, T241S, L248I, I254V, S269G, I290V, E291D, V297I, Q299R, I314L, E318D, Q3 23L, L333V, E359D, D360M, K362R, Q366S, N367G, L368V, A369T, G370A, Y371E, H404Y, H409Y, G410A, E411G, Y417F, V428I, E429A, S432 T, K433S, L437R, S443A, A451V, I464L, A470M, I502V, L503V, I531L, G537K, L540I, N553S, I559L, S563G, V571L, N579Q, H589T, S607L, L608I, L620I, R623K, R624K, L644V, S646P, M652V, I657V, R679E, L684S, N686G, H689D, S696G, T702A, T737S, L742F, Y744H, Q748H, M7 51V, I753V, A771T, R777K, P792T, S818T, R823G, V824M, E826V, K827R, A832S, T833D, M836A, I839L, G841N, V846A, N860T, V862E, D864 E, V867A, V877G, S883K, S889R, G890D, S894F, K908Q, N913D, F916H, T918V, R936N, Q938N, Y940F, K942S, S963A, R966K, K967R, K968R.

[0185]

[0186] In some embodiments, the Cas12 domain comprises an amino acid sequence that is at least about 85%, 90%, 95%, 97%, 98%, or 99% identical to or the same as SEQ ID NO:154. In some embodiments, the Cas12 domain comprises an amino acid sequence corresponding to SEQ ID NO:154.Mutations at any of the following positions in NO:154: T1, Q2, E4, G5, N8, L9, K28, H29, I30, Q31, E32, Q33, F35, I36, E37, E38, A41, N43, D44, H45, E48, I52, R55, T59, Y60, A61, D62, Q63, C64, Q66, L67, Q69, L70, N74, S76, A77, D80, S81, Y82, E85, E88, T90, R91, N92, A93, I95, E97, A99, T100, Y101, N103, A104, H106, D107, I110, R112, T 113, D114, R159, S169, S185, A187, I192, D195, K201, T212, R218, N223, I2 28. S233, I236, E237, V239, F242, Q249, Y257, V279, I284, F305, N313, S324 , I329, S331, T337, L338, L345, E349, S357, I358, N386, I393, L396, I400, S 403, V408, Q409, G427, K428, Q436, L442, S468, Q469, S472, L473, L479, E48 7. S488, A497, L510, A516, K522, Q535, M536, S541, V545, K549, N550, G552 , V557, N559, S586, Y596, A601, I604, A613, S628, E637, A657, K660, G663, Q 665, C673, L683, L697, A711, L717, Q723, A733, E735, Y740, K751, K756, G7 66. I778, R793, L844, I858, S865, I874, H898, I903, I916, L931, K941, N945 , V951, S958, V959, D965, I938, H984, A1009, C1024, G1037, T1049, G1055, T1056, Y1068, L1075, V1083, K1085, L1097, H1104, D1106, D1111, L1122, A1 134, V1138, D1147, V1160, P1161, R1171, R1173, Y1176, N1205, D1207, S1220, V1221, A1230, N1237, L1243, M1259, Q1274, G1291, Q1295, A1299 or L1304.

[0187] In some implementations, the Cas12 domain contains the corresponding SEQ ID The following mutations or substitutions of NO:154: T1S, Q2N, E4S, G5E, N8H, L9K, K28E, H29N, I30L, Q31T, E32A, Q33Y, F35M, I36V, E37N, E38D, A41L, N43S, D44E, H45N, E48K, I52V, R55K, T59Y, Y60F, A61I, D62E, Q63E, C64T, Q66K, L67H, Q69A, L70I, N74P, S76Y, A77K, D80T, S81A, Y82F, E85D, E88L, T90N, R91N, N92T, A93N , I95R, E97I, A99D, T100N, Y101C, N103K, A104S, H106A, D107G, I110E, R1 12K, T113V, D114P, R159K, S169V, S185A, A187S, I192L, D195E, K201I, T21 2K, R218N, N223T, I228T, S233G, I236L, E237D, V239I, F242V, Q249C, Y257 F, V279T, I284V, F305Y, N313S, S324N, I329L, S331A, T337E, L338K, L345I , E349Q, S357L, I358A, N386D, I393V, L396A, I400L, S403N, V408I, Q409E , G427D, K428D, Q436A, L442I, S468V, Q469L, S472A, L473V, L479T, E487D, S488D, A497V, L510I, A516V, K522Q, Q535S, M536N, S541D, V545E, K549Q, N 550Q, G552C, V557E, N559E, S586N, Y596Q, A601S, I604L, A613D, S628N, E6 37T, A657D, K660R, G663N, Q665K, C673H, L683V, L697V, A711G, L717F, Q7 23E, A733L, E735D, Y740F, K751E, K756A, G766A, I778V, R793P, L844F, I85 8V, S865T, I874L, H898N, I903V, I916A, L931F, K941N, N945Q, V951I, S958 T, V959A, D965E, I938V, H984Q, A1009S, C1024Y, G1037S, T1049E, G1055R,T1056N, Y1068F, L1075A, V1083R, K1085G, L1097I, H1104K, D1106N, D1111N, L1122K, A1134D, V1138I, D1147A, V1160E, P1161F, R1171Q, R1173E, Y1176L, N1205T, D1207N, S1220L, V1221T, A1230E, N1237S, L1243I, M1259K, Q1274L, G1291A, Q1295N, A1299N, or L1304K.

[0188] The amino acid sequences of the other Cas9 nucleases are listed in Table 1.

[0189] Table 1. Exemplary Cas9 nuclease amino acid sequences.

[0190]

[0191] In some embodiments, the chimeric nuclease comprises a Cas nuclease or nuclease domain selected from Table 1. In some embodiments, the chimeric nuclease comprises a Cas nuclease or nuclease domain selected from SEQ ID NO:161-191. In some embodiments, the Cas nuclease or nuclease domain comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence identical to that of SEQ ID NO:161-191.

[0192] In some embodiments, the chimeric nuclease also comprises additional protein fusion compounds. In some embodiments, the chimeric nuclease comprises a self-cleaving peptide, such as the T2A peptide.

[0193] GIY-YIG nuclease In some embodiments, the chimeric nuclease comprises a GIY-YIG nuclease. In some embodiments, the chimeric nuclease comprises a GIY-YIG nuclease catalytic domain.

[0194] GIY-YIG nucleases are a family of homing endonucleases that cleave DNA. GIY-YIG nucleases fold into two structural and functional domains: an N-terminal catalytic domain and a C-terminal DNA-binding domain separated by a flexible linker. In some embodiments, the chimeric nuclease includes the N-terminal catalytic domain of the GIY-YIG nuclease. In some embodiments, the chimeric nuclease includes both the N-terminal catalytic domain and the flexible linker of the GIY-YIG nuclease, but lacks the C-terminal DNA-binding domain of the GIY-YIG nuclease.

[0195] In some embodiments, the GIY-YIG nuclease is an I-TevI ​​nuclease. I-TevI ​​is a site-specific, sequence-tolerant homing endonuclease encoded by the td intron of phage T4 (Uniprot ID A0A7S9XH31, SEQ ID NO: 155). In some embodiments, the chimeric nuclease contains the catalytic domain of the I-TevI ​​nuclease.

[0196] In some embodiments, I-TevI ​​is modified I-TevI. In some embodiments, I-TevI ​​is the catalytic domain of I-TevI. In some embodiments, I-TevI ​​contains conserved amino acid substitutions. In some embodiments, I-TevI ​​contains mutations in the catalytic domain. In some embodiments, I-TevI ​​is a nickase.

[0197] In some embodiments, the I-TevI ​​nuclease comprises the following amino acid sequence: MKSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIADATFGDTCSTHPLKEEIIKKRSETVKAKMLKLGPDGRKALYSKPGSKNGRWNPETHKFCKCGVRIQTSAYTCSKCRNRSGENNSFFNHKHSDITKSKISEKMKGKKPSNIKKISCDGVIFDCAADAARHFKISSGLVTYRVKSDKWNWFYINA (SEQ ID NO:155).

[0198] In some implementations, the I-TevI ​​nuclease contains the same amino acid sequence as SEQ ID NO:155 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.

[0199] In some embodiments, the I-TevI ​​nuclease comprises the following amino acid sequence: MKSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIADATFGDTCSTHPLKEEIIKKRSETVKAKMLKLGPDGRKALYSKPGSKNGRWNPETHKFCKCGVRIQTSAYTCSKCRN (SEQ ID NO:741).

[0200] In some implementations, the I-TevI ​​nuclease contains the same amino acid sequence as SEQ ID NO:741 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.

[0201] In some embodiments, the I-TevI ​​nuclease comprises the following amino acid sequence: MGKSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIADATFGDTCSTHPLKEEIIKKRSETVKAKMLKLGPDGRKALYSKPGSKNGRWNPETHKFCKCGVRIQTSAYTCSKCRNRSGENNSFFNHKHSDITKSKISEKMKGKKPSNIKKISCDGVIFDCAADAARHFKISSGLVTYRVKSDKWNWFYINA (SEQ ID NO:156).

[0202] In some implementations, the I-TevI ​​nuclease contains the same amino acid sequence as SEQ ID NO:156 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.

[0203] In some embodiments, the I-TevI ​​nuclease comprises the following amino acid sequence: KSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIADATFGDTCSTHPLKEEIIKKRSETVKAKMLKLGPDGRKALYSKPGSKNGRWNPETHKFCKCGVRIQTSAYTCSKCRNRSGENNSFFNHKHSDITKSKISEKMKGKKPSNIKKISCDGVIFDCAADAARHFKISSGLVTYRVKSDKWNWFYINA (SEQ ID NO:157).

[0204] In some implementations, the I-TevI ​​nuclease contains the same amino acid sequence as SEQ ID NO:157 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.

[0205] In some embodiments, the I-TevI ​​nuclease catalytic domain comprises the following amino acid sequence: MKSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIA (SEQ ID NO:24).

[0206] In some embodiments, the I-TevI ​​nuclease catalytic domain comprises the following amino acid sequence: MGKSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIA (SEQ ID NO:158).

[0207] In some embodiments, the I-TevI ​​nuclease catalytic domain contains the following amino acid sequence: KSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIA (SEQ ID NO:159).

[0208] In some embodiments, the I-TevI ​​nuclease catalytic domain comprises the following amino acid sequence: MGKSGIYQIKNTLNNKVYVGSAKDFERRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIADASFGDTCSTHPLKEEIIKKRSETVKAKMLKLGPDGRKALYSKPGSKNGRWNPETHKFCKCGVRIRTSAYTCSKCRN (SEQ ID NO:272).

[0209] In some implementations, the I-TevI ​​nuclease contains the same amino acid sequence as SEQ ID NO:272 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.

[0210] In some embodiments, the I-TevI ​​nuclease catalytic domain comprises the following amino acid sequence: MGKSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIADATFGDTCSTHPLKEEIIKKRSETFKAKMLKLGPDGRKALYSRPGSKSGRWNPETHKFCKCGVRIQTSAYTCSKCRN (SEQ ID NO:740).

[0211] In some implementations, the I-TevI ​​nuclease contains the same amino acid sequence as SEQ ID NO:740 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.

[0212] In some embodiments, the I-TevI ​​sequence contains a mutation at the amino acid position corresponding to K26 of SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains a mutation at the amino acid position corresponding to T95 of SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains a mutation at the amino acid position corresponding to Q158 of SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains a mutation at the amino acid position corresponding to V117 of SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains a mutation at the amino acid position corresponding to K135 of SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains a mutation at the amino acid position corresponding to N140 of SEQ ID NO:155.

[0213] In some embodiments, the I-TevI ​​sequence contains mutations at the amino acid positions corresponding to K26 and T95 of SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains mutations at the amino acid positions corresponding to K26 and Q158 of SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains mutations at the amino acid positions corresponding to T95 and Q158 of SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains mutations at the amino acid positions corresponding to K26, T95, and Q158 of SEQ ID NO:155.

[0214] In some embodiments, the I-TevI ​​sequence contains mutations at the amino acid positions corresponding to V117, K135, and N140 of SEQ ID NO:155.

[0215] In some embodiments, the I-TevI ​​sequence contains a K26 mutation at the amino acid position corresponding to SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains a T95S mutation at the amino acid position corresponding to SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains a Q158R mutation at the amino acid position corresponding to SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains a V117F mutation at the amino acid position corresponding to SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains a K135R mutation at the amino acid position corresponding to SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains an N140S mutation at the amino acid position corresponding to SEQ ID NO:155.

[0216] In some embodiments, the I-TevI ​​sequence contains K26R and T95S mutations at the amino acid position corresponding to SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains K26R and Q158R mutations at the amino acid position corresponding to SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains T95S and Q158R mutations at the amino acid position corresponding to SEQ ID NO:155. In some embodiments, the I-TevI ​​sequence contains K26R, T95S, and Q158R mutations at the amino acid position corresponding to SEQ ID NO:155.

[0217] In some embodiments, the I-TevI ​​sequence contains V117F, K135R, and N140S mutations at the amino acid positions corresponding to SEQ ID NO:155.

[0218] In some embodiments, the I-TevI ​​sequence contains mutations at the amino acid positions corresponding to R27, V117, K135, and N140 of SEQ ID NO:155.

[0219] In some embodiments, the I-TevI ​​nickase domain contains mutations in R27A, V117F, K135R, and N140S. In some embodiments, I-TevI ​​contains a mutation that produces catalytically inactive I-TevI. In some embodiments, I-TevI ​​contains a mutation that produces catalytically inactive I-TevI ​​at the residue corresponding to amino acid position R27A in SEQ ID NO:24. In some embodiments, I-TevI ​​contains a mutation that produces catalytically inactive I-TevI ​​at the residue corresponding to amino acid position R28A in SEQ ID NO:25.

[0220] In some embodiments, I-TevI ​​contains a mutation at residue 26 corresponding to amino acid position 26 of SEQ ID NO:24. In some embodiments, I-TevI ​​contains the K26R mutation of SEQ ID NO:24. In some embodiments, I-TevI ​​contains a mutation at residue 27 corresponding to amino acid position 27 of SEQ ID NO:25. In some embodiments, I-TevI ​​contains the K27R mutation of SEQ ID NO:25.

[0221] In some embodiments, the modified I-TevI ​​nuclease domain comprises one or more substitutions selected from T11V, V16I, N14G, E25D, K26R, E36S, K37N, G38N, C39V, S41H, L45F, F49Y, I60V, and E81I at the amino acid positions corresponding to those in SEQ ID NO:155.

[0222] In some embodiments, the I-TevI ​​nuclease comprises the same amino acid sequence as SEQ ID NO:156 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the I-TevI ​​nuclease comprises the same amino acid sequence as SEQ ID NO:157 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the I-TevI ​​nuclease comprises the same amino acid sequence as SEQ ID NO:158 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the I-TevI ​​nuclease comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence of SEQ ID NO:159. In some embodiments, the I-TevI ​​nuclease comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence of SEQ ID NO:24. In some embodiments, I-TevI ​​comprises 90% of the amino acid sequence of SEQ ID NO:24. In some embodiments, I-TevI ​​comprises 91% of the amino acid sequence of SEQ ID NO:24. In some embodiments, I-TevI ​​comprises 92% of the amino acid sequence of SEQ ID NO:24. In some embodiments, I-TevI ​​comprises 93% of the amino acid sequence of SEQ ID NO:24. In some embodiments, I-TevI ​​contains 94% of the same amino acid sequence as SEQ ID NO:24. In some embodiments, I-TevI ​​contains 95% of the same amino acid sequence as SEQ ID NO:24. In some embodiments, I-TevI ​​contains 96% of the same amino acid sequence as SEQ ID NO:24. In some embodiments, I-TevI ​​contains 97% of the same amino acid sequence as SEQ ID NO:24. In some embodiments, I-TevI ​​contains 98% of the same amino acid sequence as SEQ ID NO:24. In some embodiments, I-TevI ​​contains 99% of the same amino acid sequence as SEQ ID NO:24. In some embodiments, I-TevI ​​contains 100% of the same amino acid sequence as SEQ ID NO:24.

[0223] In some embodiments, the I-TevI ​​nuclease comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence of SEQ ID NO:25. In some embodiments, I-TevI ​​comprises 90% of the amino acid sequence of SEQ ID NO:25. In some embodiments, I-TevI ​​comprises 91% of the amino acid sequence of SEQ ID NO:25. In some embodiments, I-TevI ​​comprises 92% of the amino acid sequence of SEQ ID NO:25. In some embodiments, I-TevI ​​comprises 93% of the amino acid sequence of SEQ ID NO:25. In some embodiments, I-TevI ​​comprises 94% of the amino acid sequence of SEQ ID NO:25. In some embodiments, I-TevI ​​comprises 95% of the amino acid sequence of SEQ ID NO:25. In some embodiments, I-TevI ​​comprises 96% of the amino acid sequence of SEQ ID NO:25. In some embodiments, I-TevI ​​contains 97% of the same amino acid sequence as SEQ ID NO:25. In some embodiments, I-TevI ​​contains 98% of the same amino acid sequence as SEQ ID NO:25. In some embodiments, I-TevI ​​contains 99% of the same amino acid sequence as SEQ ID NO:25. In some embodiments, I-TevI ​​contains 100% of the same amino acid sequence as SEQ ID NO:25.

[0224] In some implementations, the GIY-YIG nuclease is the I-BmoI nuclease. I-BmoI is encoded by group I introns of the thymidine synthase (TS) gene (thyA) of Bacillus mojavensis s87-18. In some embodiments, the I-BmoI nuclease comprises the following amino acid sequence: MKSGVYKITNKNTGKFYIGSSEDCESRLKVHFRNLKNNRHINRYLNNSFNKHGEQVFIGEVIHILPIEEAIAKEQWYIDNFYEEMYNISKSAYHGGDLTSYHPDKRNIILKRADSLKKVYLKMTSEEKAKRWQCVQGENNPMFGRKHTETTKLKISNHNKLYYSTHKNPFKGKKHSEESKTKLSEYASQRVGEKNPFYGKTHSDEFKTYMSKKFKGRKPKNSRPVIIDGTEYESATEASRQLNVVPATILHRIKSKNEKYSGYFYK (SEQ ID NO:737).

[0225] In some implementations, the I-BmoI nuclease contains the same amino acid sequence as SEQ ID NO:737 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.

[0226] In some implementations, the GIY-YIG nuclease is an Eco29kI nuclease. An exemplary Eco29kI nuclease is derived from the genus *Exothiospira* (…). Ectothiorhodospira magna Eco29kI restriction endonuclease.

[0227] In some embodiments, the Eco29kI nuclease comprises the following amino acid sequence: MTDDKVIPFNPLDKRHLGESVGQAMLRQPVVPMAKLSRFRGAGIYAIYYTGNFEAYQGIAACNRDDRFAAPIYVGKAVPKGARKGSGSLDTSPGAVLFSRLAQHGKSIQEVKNLDINDFYCRYLIVDDIWIPLGESLLIAKFNPLWNSVLDGFGNHDPGKGRHAGLRPRWDVVHPGRAWAGRCQAREETAEKILREAVNFLASNPPPGDW (SEQ ID NO:738).

[0228] In some implementations, the Eco29kI nuclease contains the same amino acid sequence as SEQ ID NO:738 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.

[0229] In some implementations, the Eco29kI nuclease is derived from Escherichia coli (E. coli). Escherichia coli Eco29kI restriction endonuclease.

[0230] In some embodiments, the Eco29kI nuclease comprises the following amino acid sequence: MTGKVIPFNPLDKQNLGASVAEALLSKDAHPLEELTSFQGAGIYAIYYTGDHPAYRQLAELNRDGQFRLPIYVGKAVPAGARMGLTNPDKVGNVLFRRLKEHAESIRAAENLSIEDFYCRFLVVDDIWIPLGESLVISRFKPIWNSSIDGFGNHDPGKHRYTGLRPRWDFMHPGRGWAQNLRERDETVDELIRDSIQYLQNLPPCLAQKFIEAEGD (SEQ ID NO:739).

[0231] In some implementations, the Eco29kI nuclease contains the same amino acid sequence as SEQ ID NO:739 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.

[0232] connector In some embodiments, the chimeric nuclease includes a linker domain. In some embodiments, the linker domain is located between the CRISPR-Cas nuclease and the GIY-YIG nuclease. In some embodiments, the linker domain is located between the CRISPR-Cas nuclease domain and the GIY-YIG nuclease domain.

[0233] In some embodiments, the linker comprises an I-TevI ​​linker (amino acids 93-150 of I-TevI ​​(SEQ ID NO: 155)). The linker may optionally or further comprise a flexible amino acid linker containing 10 to 100 amino acids. Such a linker may be unstructured or comprise a Gly-Ser linker.

[0234] In some embodiments, the linker comprises an amino acid substitution selected from the positions of residues T95S, S101Y, A119D, K120N, K135N, K135R, P126S, D127K, N140S, T147I, Q158R, A161V, or S165G corresponding to SEQ ID NO:155. In some embodiments, the linker comprises a substitution selected from the positions of residues T95S, V117F, K135R, N140S, or Q158R corresponding to SEQ ID NO:155.

[0235] In some embodiments, the adapter contains a mutation selected from any one of T95S, S101Y, A119D, K120N, K135N, K135R, P126S, D127K, N140S, T147I, Q158R, A161V, V117F, S165G corresponding to SEQ ID NO:155, or a combination thereof.

[0236] In some embodiments, the linker domain comprises the same amino acid sequence as 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of SEQ ID NO:26. In some embodiments, the linker domain comprises the same amino acid sequence as 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of SEQ ID NO:27. In some embodiments, the linker domain comprises the same amino acid sequence as 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of SEQ ID NO:28. In some embodiments, the linker domain comprises the same amino acid sequence as 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of SEQ ID NO:29. In some embodiments, the linker domain comprises the same amino acid sequence as 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of SEQ ID NO:30. In some embodiments, the linker domain comprises the same amino acid sequence as 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of SEQ ID NO:31. In some embodiments, the linker domain comprises the same amino acid sequence as 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of SEQ ID NO:32. In some embodiments, the linker domain comprises the same amino acid sequence as 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of SEQ ID NO:33. In some embodiments, the adapter domain comprises the same amino acid sequence as SEQ ID NO:34 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the adapter domain comprises the same amino acid sequence as SEQ ID NO:35 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.

[0237] In some embodiments, the connector additionally includes a hinge sequence. In some embodiments, the hinge sequence includes the amino acid sequence GGSGGTGGSG.

[0238] In some embodiments, the I-TevI ​​domain and adapter sequence comprise SEQ ID NO:225. In some embodiments, the I-TevI ​​domain and adapter sequence comprise 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the same amino acid sequence as SEQ ID NO:225.

[0239] In some implementations, I-TevI ​​is I-TevI ​​[R27A] and Cas9 is SaCas9 [D10A+H557A].

[0240]

[0241] In some implementations, the chimeric nuclease comprises an amino acid sequence selected from Table 2.

[0242] Table 2. Exemplary chimeric nuclease amino acid sequences

[0243] In some embodiments, the chimeric nuclease comprises a chimeric nuclease sequence selected from Table 2. In some embodiments, the chimeric nuclease comprises a chimeric nuclease sequence selected from SEQ ID NO:192-223. In some embodiments, the chimeric nuclease comprises 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the amino acid sequence identical to SEQ ID NO:192-223. In some embodiments, the chimeric nuclease has cleaving activity against the I-TevI ​​recognition site in double-stranded DNA. In some embodiments, the chimeric nuclease has cleaving activity against both the I-TevI ​​recognition site and the Cas9 recognition site in double-stranded DNA. In some embodiments, the chimeric nuclease has binding activity against the I-TevI ​​recognition site in double-stranded DNA. In some embodiments, the chimeric nuclease has cleaving activity against both the I-TevI ​​recognition site and the Cas9 recognition site in double-stranded DNA. In some implementations, the chimeric nuclease has nicking enzyme activity for both the I-TevI ​​recognition site and the Cas9 recognition site in double-stranded DNA.

[0244] Guided polynucleotides In some embodiments, the chimeric nuclease or chimeric nuclease system further comprises a guide polynucleotide. In some embodiments, the guide polynucleotide is guide RNA. In some embodiments, the guide RNA targets genomic target sites in cells. In some embodiments, the guide RNA targets pathogenic mutations in mammalian cells. In some embodiments, the mammalian cell is a human cell. In some embodiments, the guide RNA targets bacterial or viral sequences.

[0245] In some implementations, the guide RNA is a single guide RNA (sgRNA) (e.g., a fusion of crRNA and tracrRNA) or a dual guide RNA (crRNA and tracrRNA).

[0246] A single guide RNA (sgRNA) may include, in the 5' to 3' direction, an optional spacer extension sequence, a spacer sequence, a minimal CRISPR repeat sequence, a single guide adapter, a minimal tracrRNA sequence, a 3' tracrRNA sequence, and / or an optional tracrRNA extension sequence. The scaffold sequence may include all elements without a spacer sequence. The optional tracrRNA extension sequence may include elements that contribute additional function (e.g., stability) to the guide RNA. The single guide adapter may connect the minimal CRISPR repeat sequence and the minimal tracrRNA sequence to form a hairpin structure. The optional tracrRNA extension sequence may include one or more hairpins. In a particular embodiment, this disclosure provides an sgRNA comprising a spacer sequence and a tracrRNA sequence. In some implementations, crRNA and / or tracrRNA are derived from Staphylococcus aureus, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus species, Nocardia dassonvillei, Streptococcus spp., Streptococcus viridans, Streptococcus rosea, Bacillus cyclophosphamide, Bacillus pseudofungiformis, Bacillus selenium-reducing, Microbacterium thalianae, Lactobacillus delbrueckii, Lactobacillus salivarius, Microospora marineis, Burkholderia, Polarmonas naphthalene-eating, Polarmonas species, Alligator cephalopoda, Blue-stem algae species, Microcystis aeruginosa, Synechococcus species, and Arabinacillus arabinose. Ammonifex degensii Bab's pyrolytic cellulose bacteria Candidates desulforudis Clostridium botulinum, Clostridium difficile, Clostridium falciparum, Neisseria thermophile, Propionibacterium thermophilum, Thiobacillus thermophilus, Thiobacillus acidophilus, Allochromatium vinosum Species of the genus *Hymenobacter*, *Hydroxynitrosococcus*, *Nitrosococcus wari*, *Pseudomonas aeruginosa*, Ktedonobacter racemifer , Methanohalobium evestigatum Species of *Anabaena*, *Nostoc*, *Spirulina maxima*, *Spirulina platensis*, *Spirulina*, *Microcoleina*, *Oscillatoria* Petrotoga mobilis African thermocline bacteria or Acaryochloris marina Cas9 system.

[0247] In some implementations, variants of the guide RNA scaffold sequence may be used. In some implementations, the scaffold sequence followed by the guide sequence at its 3' end is SEQ ID NO:75.

[0248] In some embodiments, the guide RNA comprises the nucleic acid sequence listed in SEQ ID NO:53. In some embodiments, the guide RNA comprises 95%, 96%, 97%, 98%, 99%, or 100% identical nucleic acid sequence to SEQ ID NO:53. In some embodiments, the guide RNA comprises 95% identical nucleic acid sequence to SEQ ID NO:53. In some embodiments, the guide RNA comprises 96% identical nucleic acid sequence to SEQ ID NO:53. In some embodiments, the guide RNA comprises 97% identical nucleic acid sequence to SEQ ID NO:53. In some embodiments, the guide RNA comprises 98% identical nucleic acid sequence to SEQ ID NO:53. In some embodiments, the guide RNA comprises 99% identical nucleic acid sequence to SEQ ID NO:53. In some embodiments, the guide RNA comprises 100% identical nucleic acid sequence to SEQ ID NO:53.

[0249] In some embodiments, the guide RNA comprises the nucleic acid sequence listed in SEQ ID NO:76. In some embodiments, the guide RNA comprises 95%, 96%, 97%, 98%, 99%, or 100% identical nucleic acid sequence to SEQ ID NO:76. In some embodiments, the guide RNA comprises 95% identical nucleic acid sequence to SEQ ID NO:76. In some embodiments, the guide RNA comprises 96% identical nucleic acid sequence to SEQ ID NO:76. In some embodiments, the guide RNA comprises 97% identical nucleic acid sequence to SEQ ID NO:76. In some embodiments, the guide RNA comprises 98% identical nucleic acid sequence to SEQ ID NO:76. In some embodiments, the guide RNA comprises 99% identical nucleic acid sequence to SEQ ID NO:76. In some embodiments, the guide RNA comprises 100% identical nucleic acid sequence to SEQ ID NO:76.

[0250] In some embodiments, the guide RNA comprises the nucleic acid sequence listed in SEQ ID NO:77. In some embodiments, the guide RNA comprises 95%, 96%, 97%, 98%, 99%, or 100% identical nucleic acid sequence to SEQ ID NO:77. In some embodiments, the guide RNA comprises 95% identical nucleic acid sequence to SEQ ID NO:77. In some embodiments, the guide RNA comprises 96% identical nucleic acid sequence to SEQ ID NO:77. In some embodiments, the guide RNA comprises 97% identical nucleic acid sequence to SEQ ID NO:77. In some embodiments, the guide RNA comprises 98% identical nucleic acid sequence to SEQ ID NO:77. In some embodiments, the guide RNA comprises 99% identical nucleic acid sequence to SEQ ID NO:77. In some embodiments, the guide RNA comprises 100% identical nucleic acid sequence to SEQ ID NO:77.

[0251] In some embodiments, the guide RNA comprises the nucleic acid sequence listed in SEQ ID NO:78. In some embodiments, the guide RNA comprises 95%, 96%, 97%, 98%, 99%, or 100% identical nucleic acid sequence to SEQ ID NO:78. In some embodiments, the guide RNA comprises 95% identical nucleic acid sequence to SEQ ID NO:78. In some embodiments, the guide RNA comprises 96% identical nucleic acid sequence to SEQ ID NO:78. In some embodiments, the guide RNA comprises 97% identical nucleic acid sequence to SEQ ID NO:78. In some embodiments, the guide RNA comprises 98% identical nucleic acid sequence to SEQ ID NO:78. In some embodiments, the guide RNA comprises 99% identical nucleic acid sequence to SEQ ID NO:78. In some embodiments, the guide RNA comprises 100% identical nucleic acid sequence to SEQ ID NO:78.

[0252] In some embodiments, the guide RNA comprises the nucleic acid sequence listed in SEQ ID NO:79. In some embodiments, the guide RNA comprises 95%, 96%, 97%, 98%, 99%, or 100% identical nucleic acid sequence to SEQ ID NO:79. In some embodiments, the guide RNA comprises 95% identical nucleic acid sequence to SEQ ID NO:79. In some embodiments, the guide RNA comprises 96% identical nucleic acid sequence to SEQ ID NO:79. In some embodiments, the guide RNA comprises 97% identical nucleic acid sequence to SEQ ID NO:79. In some embodiments, the guide RNA comprises 98% identical nucleic acid sequence to SEQ ID NO:79. In some embodiments, the guide RNA comprises 99% identical nucleic acid sequence to SEQ ID NO:79. In some embodiments, the guide RNA comprises 100% identical nucleic acid sequence to SEQ ID NO:79.

[0253] In some embodiments, the guide RNA comprises the nucleic acid sequences listed in SEQ ID NO:76-85. In some embodiments, the guide RNA comprises the nucleic acid sequences listed in SEQ ID NO:234-235.

[0254] In some embodiments, the guide RNA is modified. In some embodiments, the guide RNA comprises one or more of the following: non-natural nucleoside internucleotides, nucleic acid mimics, modified sugar moieties, and modified nucleobases. In some embodiments, the guide RNA comprises modified nucleotides. In some embodiments, the modified nucleotide includes one or more of the following: 5-methylcytosine; 5-hydroxymethylcytosine; xanthine; hypoxanthine; 2-aminoadenine; a 6-methyl derivative of adenine; a 6-methyl derivative of guanine; a 2-propyl derivative of adenine; a 2-propyl derivative of guanine; 2-thiouracil; 2-thiothymine; 2-thiocytosine; 5-halogenuridine; 5-halogenuridine; 5-propynyluracil; 5-propynylcytosine; 6-azouracil; 6-azocytosine; 6-azothymine; pseudouracil; 4-thiouracil; 8-halogen; 8-amino; 8-thiol; 8-thioalkyl; 8-hydroxy; 5-halogen; 5-Bromo; 5-Trifluoromethyl; 5-Substituted Uracil; 5-Substituted Cytosine; 7-Methylguanine; 7-Methyladenine; 2-F-Adenine; 2-Aminoadenine; 8-Zazaguanine; 8-Zazaadenine; 7-Denitroguanine; 7-Denitroadenine; 3-Denitroguanine; 3-Denitroadenine; Tricyclic Pyrimidine; Phenyrazinecytidine; Phenothiazinecytidine; Substituted Phenyrazinecytidine; Carbazolecytidine; Pyridineindolecytidine; 7-Denitroadenine; 7-Denitroguanine; 2-Aminopyridine; 2-Pyridone; 5-Substituted Pyrimidine; 6-Zazapyrimidine; N-2, N-6 or O-6 substituted purine; 2-Aminopropyladenine; 5-Protyynyluracil; or 5-Protyynylcytosine.

[0255] In some embodiments, the non-natural nucleoside internucleotide bond comprises one or more of the following: thiophosphate, phosphoramide, non-phosphodiester, heteroatom, chiral thiophosphate, dithiophosphate, triphosphate, aminoalkyl phosphate triester, 3'-alkylphosphonate, 5'-alkylphosphonate, chiral phosphonate, hypophosphonate, 3'-aminophosphamide, aminoalkylphosphamide, phosphodiamid, thiocarbonylphosphamide, thiocarbonylalkylphosphamide, thiocarbonylalkyl phosphate triester, selenophosphate, and boron phosphate. In some embodiments, the nucleic acid mimic comprises one or more of peptide nucleic acid (PNA), morpholino nucleic acid, cyclohexenyl nucleic acid (CeNA), or locked nucleic acid (LNA). In some embodiments, the modified sugar moiety comprises one or more of 2'-O-(2-methoxyethyl), 2'-dimethylaminooxyethoxy, 2'-dimethylaminoethoxyethoxy, 2'-O-methyl, or 2'-fluorine.

[0256] In some embodiments, the chimeric nuclease targets a site within the gene. In some embodiments, the chimeric nuclease targets a site within an intron or exon of the gene. In some embodiments, the gene is... B2M or AAVS1 In some implementations, genes are... CFTR In some implementations, genes are... SERPINA1 (Encoding α-1-antitrypsin). In some implementations, the gene is... DMPK In some implementations, genes are... C9ORF72 In some implementations, the guide RNA contains a target B2M or AAVS1 The spacer region containing the target site. In some implementations, the guide RNA contains the target... CFTR The spacer region containing the target site. In some implementations, the guide RNA contains the target... SERPINA 1 The spacer region containing the target site. In some implementations, the guide RNA contains the target... DMPK The spacer region containing the target site. In some implementations, the guide RNA contains the target... C9ORF72 The spacer region of the target site.

[0257] In some implementation schemes, genes are C9ORF72 or DMPK In some implementations, the spacer region contains nucleic acid sequences listed in SEQ ID NO:20 to 23.

[0258] In some embodiments, the spacer region contains the nucleic acid sequence listed in SEQ ID NO:41 or SEQ ID NO:51. In some embodiments, the spacer region contains the nucleic acid sequences listed in SEQ ID NO:76 to SEQ ID NO:85.

[0259] It should be understood that in the sequence, if the nucleic acid sequence is DNA, then T is thymine, and if the nucleic acid sequence is RNA, then T is uracil.

[0260] CFTR Cystic fibrosis transmembrane conduction regulator ( CFTR The gene encodes a member of the ATP-binding cassette (ABC) transporter superfamily. The encoded protein functions as a chloride ion channel, making it unique within this family, and controls the secretion and uptake of ions and water in epithelial tissues. Channel activation is mediated by a cycle of phosphorylation of the regulatory domain, ATP binding of the nucleotide-binding domain, and ATP hydrolysis. Mutations in this gene cause cystic fibrosis (CF), the most common and fatal genetic disease in Nordic populations. CFTRMutations in the gene (GRCh38.p14GCF_000001405.40, NM_000492.4, NC_000007.14 refer to GRCh38.p14) can lead to suboptimal ion transport and fluid retention, resulting in prominent clinical manifestations of abnormal thickening of mucus in the lungs and pancreatic insufficiency. The most common mutation in cystic fibrosis, δF508, results in impaired folding and transport of the encoded protein. CFTR proteins are present in a wide range of organs, including the pancreas, kidneys, liver, lungs, gastrointestinal tract, and reproductive tract, making CF a multi-organ disease. In the lungs, dysfunctional CFTR can impair mucociliary clearance, making the organ susceptible to bacterial infection and inflammation, ultimately leading to airway obstruction, respiratory failure, and premature death. CF remains the most common and deadliest genetic disease in Caucasian populations, affecting an estimated 70,000–100,000 people worldwide, highlighting the real need to develop better treatments.

[0261] CFTR Mutations in the process include, but are not limited to, c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter) or c.3846G>A (p.Trp1282Ter) .

[0262] In some implementation schemes, guidance is provided for polynucleotide targeting. CFTR Sites in the intron or exon regions of a gene.

[0263] In some implementations, the target region of the chimeric nuclease is located at Chr7:117,559,464-117,587,836 relative to the hg38 genome. CFTR The F508–R553X region. In some embodiments, the chimeric nuclease directly cleaves the 5', 3' and / or interior regions adjacent to Chr7:117,559,464-117,587,836.

[0264] SERPINA1 Depend on SERPINA1 The gene (GRCh38.p14 (GCF_000001405.40), NG_008290.1 ​​RefSeqGene) encodes a member of the serine protease inhibitor family A, a serine protease inhibitor belonging to the serine protease inhibitor superfamily. Its targets include elastase, plasmin, thrombin, trypsin, chymotrypsin, and plasminogen activator. SERPINA1 protein is produced in the liver and bone marrow by lymphocytes and monocytes in lymphoid tissues, and by Panette cells in the intestine. Defects in this gene are associated with chronic obstructive pulmonary disease, emphysema, and chronic liver disease, and play a role in α1-antitrypsin deficiency. Several transcript variants of this gene encoding the same protein have been identified. SERPINA1The E342K mutation is involved in α1-antitrypsin deficiency.

[0265] In some implementation schemes, guidance is provided for polynucleotide targeting. SERPINA1 Sites in the intron or exon regions of a gene.

[0266] DMPK DM1 protein kinase (DMPK) is a serine-threonine kinase that is closely associated with other kinases that interact with Rho family members of small GTPases. Its substrates include myopoietin, the β subunit of L-type calcium channels, and phospholeman. DMPK The gene (NG_009784.1 RefSeqGene, GRCh38.p14 (GCF_000001405.40)) is located at 19q13.32. The 3' untranslated region of this gene contains 5–38 copies of the (CTG)n*(CAG)n trinucleotide repeat sequence. Expansion of this unstable motif to 40–5,000 copies causes type I myotonic dystrophy, the severity of which increases with the copy number of the repeat element. Repeat expansion is associated with condensation of local chromatin structure that disrupts gene expression in this region. Several alternative splicing transcript variants of this gene have been described. Removal of excess (CTG)n*(CAG)n trinucleotide repeat sequences can stabilize cells. DMPK The 3' untranslated region. In some implementations, one or more guide polynucleotides target... DMPK The (CAG)n site in the 3' untranslated region of the gene is excised.

[0267] In some embodiments, the spacer region of the guiding polynucleotide comprises a nucleic acid sequence listed in SEQ ID NO:22 to 23. In some embodiments, the spacer region of the guiding polynucleotide comprises a nucleic acid sequence listed in any one of SEQ ID NO:136-147.

[0268] In some implementations, the chimeric nuclease system targets the enzymes listed in SEQ ID NO:4. DMPK The intron region.

[0269] In some embodiments, the chimeric nuclease system is a dual-directed chimeric nuclease. In some embodiments, the chimeric nuclease system is a dual-directed Tev-dCas9. The dual-directed nuclease of this disclosure can be used to cleave... DMPKExemplary excision sequences are listed in Table 3A. Exemplary excision regions are listed in Table 3B. In some embodiments, the excision sequence includes SEQ ID NO:236. In some embodiments, the excision sequence includes any one of SEQ ID NO:237 to 247.

[0270] Table 3A lists sequences in the 3' untranslated region of Dmpk that can be targeted for excision using dual-guided Tev-dCas9.

[0271] *[CAG]40+ represents 40 or more CAG repetitions. Table 3B shows the range of chromosome locations that can be cut using dual Tev-dCas9 in Dmpk.

[0272] Table 4A shows the target DPMK The nucleic acid sequence of SEQ ID NO:236 in the gene is an exemplary dual-guided chimeric nuclease used for excision.

[0273] Table 4A Target DMPK An exemplary nucleotide sequence of a combination of nuclease + guide RNA + conditionally cleavable element.

[0274]

[0275] Table 4B shows the targets DPMK The nucleic acid sequence of SEQ ID NO:240 in the gene is an exemplary dual-guided chimeric nuclease for excision.

[0276] Table 4B Targets DMPK An exemplary nucleotide sequence of a combination of nuclease + guide RNA + conditionally cleavable element.

[0277]

[0278] Table 5. Exemplary amino acid sequences of chimeric nucleases and nucleic acid sequences of two types of guide RNA after transcription, processing, and translation.

[0279]

[0280] C9ORF72 The C9orf72-SMCR8 complex subunit (C9ORF72) plays a crucial role in regulating endosomal transport and has been shown to interact with Rab proteins involved in autophagy and endocytic transport. In transcripts from this gene, the GGGGCC repeat, amplified from 2–22 copies to 700–1600 copies in intronic sequences alternating between 5' exons, is associated with 9p-linked ALS (amyotrophic lateral sclerosis) and FTD (frontotemporal dementia). Studies have shown that hexanucleotide amplification can lead to selective stabilization of pre-mRNAs containing repeat sequences and the accumulation of insoluble dipeptide repeat aggregates that may be pathogenic in FTD-ALS patients (PMID: 23393093). Alternative splicing produces multiple transcript variants encoding different isotypes.

[0281] In some implementation schemes, genes are C9ORF72 In some implementations, the spacer region contains the nucleic acid sequence listed in SEQ ID NO:20 or 21.

[0282] In some implementations, the spacer region contains a nucleic acid sequence listed in any of SEQ ID NO:132-135.

[0283] In some implementations, the chimeric nuclease targets the enzyme listed in SEQ ID NO:2 or 3. C9ORF72 The intron region. In some implementations, the removed C9orf72 G4C2 repeat sequence includes the [G4C2]24+ repeat sequence.

[0284] Table 6. Range of chromosome locations that can be cleaved in C9orf72 using dual Tev-dCas9.

[0285] Table 7. Exemplary nucleotide sequences of combinations of nuclease + guide RNA + conditionally cleavable element targeting C9orf72.

[0286] In some implementation schemes, genes are C9ORF72 In some implementations, the spacer region contains nucleic acid sequences listed in SEQ ID NO:20 to 23.

[0287] In some embodiments, the chimeric nuclease system is a dual-directed chimeric nuclease system. A dual-directed nuclease system comprises two chimeric nucleases that bind and cleave at two sites in the cellular genome, preferably on the same chromosome. In some embodiments, the two chimeric nucleases are identical chimeric nucleases having two different guide polynucleotides that can bind upstream or downstream of each other on the DNA sequence. In some embodiments, the two chimeric nucleases are different chimeric nucleases having two different guide polynucleotides.

[0288] In some embodiments, the chimeric nuclease system comprises a first chimeric nuclease and a first guide RNA, as well as a second chimeric nuclease and a second guide RNA. In some embodiments, the first chimeric nuclease and the first guide RNA target a first target site, and the second chimeric nuclease and the second guide RNA target a second target site. In some embodiments, the first chimeric nuclease comprises an active I-TevI ​​nuclease and an inactive dCas nuclease. In some embodiments, the second chimeric nuclease comprises an active I-TevI ​​nuclease and an inactive dCas nuclease. In some embodiments, the first chimeric nuclease comprises an active I-TevI ​​nuclease and an inactive dCas nuclease, and the second chimeric nuclease comprises an active I-TevI ​​nuclease and an inactive dCas nuclease.

[0289] In some embodiments, the first guide RNA and the second guide RNA target different target sites in the cell's genome. In some embodiments, the distance between the first target site and the second target site is 100 bases. In some embodiments, the distance between the first target site and the second target site is 200 bases. In some embodiments, the distance between the first target site and the second target site is 300 bases. In some embodiments, the distance between the first target site and the second target site is 400 bases. In some embodiments, the distance between the first target site and the second target site is 500 bases. In some embodiments, the distance between the first target site and the second target site is 600 bases. In some embodiments, the distance between the first target site and the second target site is 700 bases. In some embodiments, the distance between the first target site and the second target site is 800 bases. In some embodiments, the distance between the first target site and the second target site is 900 bases. In some embodiments, the distance between the first target site and the second target site is 1000 bases. In some embodiments, the distance between the first target site and the second target site is 2000 bases. In some embodiments, the distance between the first target site and the second target site is 3000 bases. In some embodiments, the distance between the first target site and the second target site is 4000 bases. In some embodiments, the distance between the first target site and the second target site is 5000 bases. In some embodiments, the distance between the first target site and the second target site is 6000 bases. In some embodiments, the distance between the first target site and the second target site is 7000 bases. In some embodiments, the distance between the first target site and the second target site is 8000 bases. In some embodiments, the distance between the first target site and the second target site is 9000 bases. In some embodiments, the distance between the first target site and the second target site is 10000 bases. In some embodiments, the distance between the first target site and the second target site is 11000 bases. In some embodiments, the distance between the first target site and the second target site is 12000 bases. In some embodiments, the distance between the first target site and the second target site is 13,000 bases. In some embodiments, the distance between the first target site and the second target site is 14,000 bases. In some embodiments, the distance between the first target site and the second target site is 15,000 bases. In some embodiments, the distance between the first target site and the second target site is 20,000 bases. In some embodiments, the distance between the first target site and the second target site is 25,000 bases. In some embodiments, the distance between the first target site and the second target site is 28,000 bases. In some embodiments, the distance between the first target site and the second target site is 30,000 bases.In some embodiments, the distance between the first target site and the second target site is more than 30,000 bases. In some embodiments, the distance between the first target site and the second target site is between 100 and 1,000 bases. In some embodiments, the distance between the first target site and the second target site is between 100 and 10,000 bases. In some embodiments, the distance between the first target site and the second target site is between 100 and 20,000 bases. In some embodiments, the distance between the first target site and the second target site is between 100 and 30,000 bases. In some embodiments, the distance between the first target site and the second target site is between 1,000 and 10,000 bases. In some embodiments, the distance between the first target site and the second target site is between 1,000 and 20,000 bases. In some embodiments, the distance between the first target site and the second target site is between 1,000 and 30,000 bases. In some implementations, the distance between the first target site and the second target site is between 5,000 and 30,000 bases.

[0290] In some embodiments, the first and second guide RNAs target the same strand of genomic DNA. In some embodiments, the first and second guide RNAs target opposite strands of genomic DNA. In some embodiments, the first guide RNA targets a first I-TevI ​​target site in the cell's genome with a first chimeric nuclease and cleaves it at the first I-TevI ​​target site, and the second guide RNA targets a second I-TevI ​​target site in the cell's genome with a second chimeric nuclease and cleaves it at the second I-TevI ​​target site, wherein the cleavage produces nucleotide overhangs at the second I-TevI ​​target site.

[0291] In some implementations, a first guide RNA targets a first chimeric nuclease to a Cas9 target site in the cell's genome and cleaves it at the Cas9 target site, and a second guide RNA targets a second chimeric nuclease to an I-TevI ​​target site in the cell's genome and cleaves it at the I-TevI ​​target site, wherein the cleavage produces nucleotide overhangs at the I-TevI ​​target site.

[0292] In some embodiments, the second guide RNA targets the 5' region of the donor polynucleotide target site. In some embodiments, the 3' end of the donor polynucleotide is complementary to a protrusion generated by cleaving the second I-TevI ​​at the second I-TevI ​​target site.

[0293] In some embodiments, the guide RNA further comprises a donor polynucleotide. In some embodiments, the donor polynucleotide is attached to the 3' end of the guide RNA. In some embodiments, the donor polynucleotide comprises a repair template. In some embodiments, the repair template comprises a 2-nucleotide overhang compared to the target site. In some embodiments, the repair template is 36 bases long.

[0294] In some embodiments, the guide RNA further comprises one or more synthetic tRNA sequences encoding a trans-acting ribozyme sequence. In some embodiments, the tRNA is located at the 5' end of the guide RNA. In some embodiments, the tRNA is located at the 5' end of the guide RNA. In some embodiments, the ribozyme is capable of cleaving the tRNA from the guide RNA.

[0295] Nucleic acid This article provides, in particular, the nucleic acids encoding the chimeric nucleases and chimeric nuclease systems described herein.

[0296] In one aspect, the nucleic acid comprises a polynucleotide encoding a chimeric nuclease (the chimeric nuclease comprising a GIY-YIG nuclease domain and an RNA-directed nuclease domain); a polynucleotide encoding guide RNA (gRNA); and a polynucleotide encoding tRNA. In some embodiments, the nucleic acid comprises a polynucleotide encoding a chimeric nuclease (the chimeric nuclease comprising an I-TevI ​​nuclease domain and an RNA-directed nuclease domain); a polynucleotide encoding guide RNA (gRNA); and a polynucleotide encoding tRNA.

[0297] Exemplary designs of nucleic acids encoding the chimeric nucleases and chimeric nuclease systems described herein are shown in Figure 1 , Figure 2 , Figure 26A , Figure 26B , Figure 26C and Figure 27A middle.

[0298] In another respect, nucleic acids contain polynucleotides that encode guide RNA and polynucleotides that encode tRNA.

[0299] In some implementations, the nucleic acid also contains one or more additional guide RNAs.

[0300] In some implementations, the nucleic acid is DNA or RNA. In some implementations, the DNA is circular plasmid DNA, linear double-stranded DNA, single-stranded DNA, or a combination of chimeric RNA and DNA.

[0301] In some embodiments, the RNA is mRNA. In some embodiments, the mRNA comprises a nucleic acid mimic selected from the group consisting of peptide nucleic acids (PNA), morpholinonucleotides, cyclohexenylnucleotides (CeNA), and locked nucleic acids (LNA). In some embodiments, the mRNA comprises a modified sugar moiety, optionally wherein the modified sugar moiety is selected from the group consisting of N1-methylpseudouridine, 9-methyladenine, 2'-O-(2-methoxyethyl), 2'-dimethylaminooxyethoxy, 2'-dimethylaminoethoxyethoxy, 2'-O-methyl, and 2'-fluoro. In some embodiments, the mRNA comprises modified nucleotides, optionally wherein the modified nucleotides are selected from the group consisting of: 5-methylcytosine; 5-hydroxymethylcytosine; xanthine; hypoxanthine; 2-aminoadenine; a 6-methyl derivative of adenine; a 6-methyl derivative of guanine; a 2-propyl derivative of adenine; a 2-propyl derivative of guanine; 2-thiouracil; 2-thiothymine; 2-thiocytosine; 5-halogenuridine; 5-halogenuridine; 5-propynyluracil; 5-propynylcytosine; 6-azouracil; 6-azocytosine; 6-azothymine; pseudouracil; 4-thiouracil; 8-halogen; 8-amino; 8-thiol; 8-thioalkyl; 8-hydroxy ; 5-halogen; 5-bromine; 5-trifluoromethyl; 5-substituted uracil; 5-substituted cytosine; 7-methylguanine; 7-methyladenine; 2-F-adenine; 2-aminoadenine; 8-azaguanine; 8-azaadenine; 7-deadenine; 7-deadenine; 3-deadenine; 3-deadenine; tricyclic pyrimidine; phenyloxazine cytosine; phenothiazine cytosine; substituted phenyloxazine cytosine; carbazole cytosine; pyridine indole cytosine; 7-deadenine; 7-deadenine; 2-aminopyridine; 2-pyridone; 5-substituted pyrimidine; 6-azapyrimidine; N-2, N-6 or O-6 substituted purine; 2-aminopropyladenine; 5-propynyluracil; or 5-propynylcytosine. In some implementations, the mRNA comprises a non-natural or non-natural internucleotide link selected from the group consisting of: thiophosphates, phosphoramides, non-phosphodiesters, heteroatoms, chiral thiophosphates, dithiophosphates, triphosphates, aminoalkyl phosphates, 3'-alkylphosphonates, 5'-alkylphosphonates, chiral phosphonates, hypophosphonates, 3'-aminophosphamides, aminoalkylphosphamides, phosphodiesteramides, thiocarbonylphosphamides, thiocarbonylalkylphosphonates, thiocarbonylalkyl phosphates, selenophosphates, or boron phosphates.

[0302] In some implementations, the nucleic acid is approximately 5 kb in length. In some implementations, the nucleic acid is less than 5 kb in length. In some implementations, the nucleic acid can be packaged into an AAV.

[0303] In some embodiments, nucleic acids are transcribed into a single mRNA encoding a chimeric nuclease, one or more guide RNAs, and any additional donor polynucleotides. In some embodiments, the single mRNA is processed into individual components at one or more cleavage sites or trans-acting cleavage sites via ribozymes or cellular mRNA processing mechanisms.

[0304] In some embodiments, the cleavage site is a ribozyme cleavage site. In some embodiments, the cleavage site is a tRNA cleavage site. In some embodiments, the guide RNA sequence is directly adjacent to the 5' cleavage site. In some embodiments, the guide RNA sequence is directly adjacent to the 3' cleavage site.

[0305] In some implementations, the guide RNA sequence is directly adjacent to the 5' and 3' cleavage sites. Cleavage sites in the mRNA can increase the efficiency and quantity of guide RNA sequences available in the cell after transcription.

[0306] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter (when the nucleic acid is DNA) or untranslated region (UTR) (when the nucleic acid is RNA) sequence, a chimeric nuclease sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA3 sequence, a cleavage site, and a polyA sequence.

[0307] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter / untranslated region (UTR) sequence, a chimeric nuclease sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, a gRNA2 sequence, a trans-acting target / cleavage site, and a multi-A sequence.

[0308] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter / untranslated region (UTR) sequence, a chimeric nuclease sequence, a trans-acting target / cleavage site, a gRNA1 sequence, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, and a multi-A sequence.

[0309] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter or UTR sequence, a chimeric nuclease sequence, a MALAT sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA3 sequence, a cleavage site, and a multi-A sequence.

[0310] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter / UTR sequence, a chimeric nuclease sequence, a MALAT sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, a gRNA2 sequence, a trans-acting target / cleavage site, and a multi-A sequence.

[0311] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter / untranslated region (UTR) sequence, a chimeric nuclease sequence, a MALAT sequence, a trans-acting target / cleavage site, a gRNA1 sequence, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, and a multi-A sequence.

[0312] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, and a multi-A sequence.

[0313] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a Ribo2 sequence, and a multiple A sequence.

[0314] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a trans-acting target / cleavage site, and a multi-A sequence.

[0315] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a trans-acting target / cleavage site, a gRNA1 sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a Ribo2 sequence, and a multi-A sequence.

[0316] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a trans-acting target / cleavage site, a Ribo1 sequence, a cleavage site, a gRNA2 sequence, a Ribo2 sequence, and a multi-A sequence.

[0317] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a Ribo1 sequence, a trans-acting target / cleavage site, a gRNA2 sequence, a Ribo2 sequence, and a multiple A sequence.

[0318] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a trans-acting target / cleavage site, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, and a multi-A sequence.

[0319] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, and a multi-A sequence.

[0320] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a Ribo2 sequence, and a multi-A sequence.

[0321] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a trans-acting target / cleavage site, and a multi-A sequence.

[0322] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a trans-acting target / cleavage site, a gRNA1 sequence, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a Ribo2 sequence, and a multi-A sequence.

[0323] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a trans-acting target / cleavage site, a Ribo1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a Ribo2 sequence, and a multi-A sequence.

[0324] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a Ribo1 sequence, a trans-acting target / cleavage site, a gRNA2 sequence, a Ribo2 sequence, and a multi-A sequence.

[0325] In some implementations, from the 5' end to the 3' end, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a trans-acting target / cleavage site, a Ribo1 sequence, a gRNA1 sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, and a multi-A sequence.

[0326] In some implementations, the trans-acting target / cleavage site is a ribozyme or tRNA cleavage site.

[0327] In some implementations, the nucleic acid additionally contains one or more donor nucleotide sequences.

[0328] In some embodiments, the nucleic acid comprises nucleic acid sequences as listed in SEQ ID NO:15-19. In some embodiments, the nucleic acid comprises nucleic acid sequences as listed in SEQ ID NO:113, 114, 115, 116, 117, or 118.

[0329] In some embodiments, the nucleic acid also includes a self-inactivating nucleic acid sequence. In some embodiments, the self-inactivating nucleic acid sequence is included at a nuclease binding and cleavage site between the promoter and start codon of the nuclease, or between the nuclease and a polyadenylated sequence or 5' untranslated region.

[0330] Donor polynucleotides In some embodiments, the nucleic acid further comprises a nucleic acid sequence encoding a donor polynucleotide. In some embodiments, the nucleic acid comprises two or more donor polynucleotides. In some embodiments, the donor polynucleotide is a single-stranded nucleic acid. In some embodiments, the donor polynucleotide is a double-stranded nucleic acid. In some embodiments, the donor polynucleotide is a partially double-stranded nucleic acid. In some embodiments, the donor polynucleotide is RNA. In some embodiments, the first strand of the double-stranded donor polynucleotide is DNA, and the second strand is RNA. In some embodiments, the donor polynucleotide is a cis-acting repair template encoding the sense strand of a repair sequence. In some embodiments, the donor polynucleotide is a cis-acting repair template encoding a complementary or inverse complementary sequence of the repair sequence. In some embodiments, the donor polynucleotide is a trans-acting repair template of the repair sequence. In some embodiments, the donor polynucleotide comprises a cis-acting single-stranded RNA polynucleotide annealed to a complementary single-stranded DNA polynucleotide.

[0331] In some embodiments, the repair template is 20 bases long. In some embodiments, the repair template is 25 bases long. In some embodiments, the repair template is 30 bases long. In some embodiments, the repair template is 35 bases long. In some embodiments, the repair template is 40 bases long. In some embodiments, the repair template is 45 bases long. In some embodiments, the repair template is 50 bases long. In some embodiments, the repair template is 55 bases long. In some embodiments, the repair template is 60 bases long. In some embodiments, the repair template is 65 bases long. In some embodiments, the repair template is 70 bases long. In some embodiments, the repair template is 75 bases long. In some embodiments, the repair template is 80 bases long. In some embodiments, the repair template is 85 bases long. In some embodiments, the repair template is 90 bases long. In some embodiments, the repair template is 95 bases long. In some embodiments, the repair template is 100 bases long. In some embodiments, the repair template is 150 bases long. In some embodiments, the repair template is 200 bases long. In some embodiments, the repair template is 250 bases long. In some embodiments, the repair template is 300 bases long. In some embodiments, the repair template is 350 bases long. In some embodiments, the repair template is 400 bases long. In some embodiments, the repair template is 450 bases long. In some embodiments, the repair template is 500 bases long. In some embodiments, the repair template is 600 bases long. In some embodiments, the repair template is 700 bases long. In some embodiments, the repair template is 800 bases long. In some embodiments, the repair template is 900 bases long. In some embodiments, the repair template is 1000 bases long. In some embodiments, the repair template is at least 900 bases long. In some embodiments, the repair template is between 20 and 100 bases long. In some embodiments, the repair template is between 20 and 200 bases long. In some embodiments, the repair template is between 20 and 300 bases long. In some embodiments, the repair template is between 20 and 400 bases long. In some embodiments, the repair template is between 20 and 500 bases long. In some embodiments, the repair template is between 20 and 600 bases long. In some embodiments, the repair template is between 20 and 700 bases long. In some embodiments, the repair template is between 20 and 800 bases long. In some embodiments, the repair template is between 20 and 900 bases long. In some embodiments, the repair template is between 20 and 1000 bases long.

[0332] In some embodiments, the donor polynucleotide is attached to the guide RNA. In some embodiments, the donor polynucleotide is attached to the guide RNA via its 5' end and the 3' end of the guide RNA. In some embodiments, the donor polynucleotide is attached to the guide RNA via a linker. In some embodiments, the linker is attached to both the 5' end of the donor polynucleotide and the 3' end of the guide RNA.

[0333] In some embodiments, the linker is one nucleotide long. In some embodiments, the linker is two nucleotides long. In some embodiments, the linker is three nucleotides long. In some embodiments, the linker is four nucleotides long. In some embodiments, the linker is five nucleotides long. In some embodiments, the linker is six nucleotides long. In some embodiments, the linker is seven nucleotides long. In some embodiments, the linker is eight nucleotides long. In some embodiments, the linker is nine nucleotides long. In some embodiments, the linker is ten nucleotides long. In some embodiments, the linker is eleven nucleotides long. In some embodiments, the linker is twelve nucleotides long. In some embodiments, the linker is thirteen nucleotides long. In some embodiments, the linker is fourteen nucleotides long. In some embodiments, the linker is fifteen nucleotides long. In some embodiments, the linker is sixteen nucleotides long. In some embodiments, the linker is seventeen nucleotides long. In some embodiments, the linker is eighteen nucleotides long. In some embodiments, the linker is nineteen nucleotides long. In some embodiments, the linker is twenty nucleotides long. In some embodiments, the linker is twenty-five nucleotides long. In some embodiments, the linker is thirty nucleotides long.

[0334] In some implementations, one or more donor polynucleotides are separated by one or more ribozyme cleavage sites.

[0335] In some embodiments, the target site of the donor polynucleotide is the I-TevI ​​cleavage site. In some embodiments, the target site of the donor polynucleotide is the Cas cleavage site.

[0336] In some embodiments, one or more donor polynucleotides contain a 2-nucleotide overhang compared to the corresponding base adjacent to the 3' of the target site. In some embodiments, one or more donor polynucleotides contain a 3-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides contain a 4-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides contain a 5-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides contain a 6-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides contain a 7-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides contain an 8-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides contain a 9-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides contain a 10-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides contain 11 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 12 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 13 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 14 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 15 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 16 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 17 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 18 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 19 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 20 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 2 to 16 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 2 to 18 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 2 to 20 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 4 to 16 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides contain 4 to 18 nucleotide overhangs compared to the target site.In some implementations, one or more donor polynucleotides contain 4 to 20 nucleotide overhangs compared to the target site.

[0337] In some embodiments, the second guide RNA targets the 5' region of the donor polynucleotide target site. In some embodiments, the 3' end of the donor polynucleotide is complementary to a protrusion generated by cleaving the second I-TevI ​​at the second I-TevI ​​target site.

[0338] In some embodiments, one or more donor polynucleotides are at least 30 bases long. In some embodiments, one or more donor polynucleotides are between at least 30 and 1500 bases long. In some embodiments, one or more donor polynucleotides are between at least 30 and 100 bases long. In some embodiments, one or more donor polynucleotides are between at least 40 and 90 bases long. In some embodiments, one or more donor polynucleotides are between at least 50 and 80 bases long. In some embodiments, one or more donor polynucleotides are between at least 60 and 70 bases long. In some embodiments, one or more donor polynucleotides are between at least 100 and 1000 bases long. In some embodiments, one or more donor polynucleotides are between at least 100 and 1500 bases long. In some embodiments, one or more donor polynucleotides are between at least 200 and 800 bases long. In some embodiments, one or more donor polynucleotides are between 300 and 700 bases in length. In some embodiments, one or more donor polynucleotides are between 400 and 600 bases in length. In some embodiments, one or more donor polynucleotides are at least 500 bases in length.

[0339] In some implementations, donor polynucleotides can be inserted. AAVS1 The safe harbor site is located. AAVS1 The nucleic acid sequence of the integration site is listed in SEQ ID NO:1.

[0340] In some implementations, the donor polynucleotide comprises a donor polynucleotide sequence selected from Table 8.

[0341] In some implementations, the donor polynucleotide comprises one of the nucleic acid sequences listed in SEQ ID NO:12-14.

[0342] In some implementations, the donor polynucleotide comprises one of the nucleic acid sequences listed in SEQ ID NO:42-45.

[0343] In some implementations, the donor polynucleotide contains B2M Inactivated sequence SEQ ID NO:52.

[0344] In some implementations, donor polynucleotides can be inserted. CFTR In genes. In some implementations, donor polynucleotides can be used for repair. CFTR Gene. In some embodiments, the donor polynucleotide comprises nucleic acid sequences SEQ ID NO:67 to 69. In some embodiments, the donor polynucleotide comprises 95%, 96%, 97%, 98%, 99%, or 100% identical nucleic acid sequences to SEQ ID NO:67 to 69 or 71-74.

[0345] In some implementations, donor polynucleotides can be inserted. SERPINA1 In genes. In some implementations, donor polynucleotides can be used for repair. SERPINA 1 gene. In some implementations, the donor polynucleotide contains SERPINA1 The sequence is SEQ ID NO:70. In some embodiments, the donor polynucleotide comprises 95%, 96%, 97%, 98%, 99%, or 100% identical nucleic acid sequences to any one of SEQ ID NO:70 or 75. In some embodiments, the donor polynucleotide comprises 95%, 96%, 97%, 98%, 99%, or 100% identical nucleic acid sequences to any one of SEQ ID NO:226 to 229. In some embodiments, the donor polynucleotide comprises sequences selected from SEQ ID NO:119-122. SERPINA1 Nucleic acid sequence. In some embodiments, the donor polynucleotide comprises 95%, 96%, 97%, 98%, 99%, or 100% identical nucleic acid sequences to any one of SEQ ID NO:119-122.

[0346] In some embodiments, the guide sequence having the donor polynucleotide comprises a sequence selected from SEQ ID NO:123-126. SERPINA1 Nucleic acid sequence. In some embodiments, the donor polynucleotide comprises 95%, 96%, 97%, 98%, 99%, or 100% identical nucleic acid sequences to any one of SEQ ID NO:123-126.

[0347] Table 8 shows exemplary sequences of target sites and donor polynucleotides. All sequences are paired with the exemplary scaffold sequence GTTTTAGTACTCTGGAAACAGAATCTACTAAAACAAGGCAAAATGCCGTGTTTATCTCGTCAACTTGTTGGCGAGAT (SEQ ID NO:87).

[0348] In some embodiments, the guiding polynucleotide and the donor polynucleotide comprise 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequence of SEQ ID NO:88-100. In some embodiments, the guiding polynucleotide and the donor polynucleotide comprise 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequence of SEQ ID NO:101-103 or 105-108. In some embodiments, the guiding polynucleotide and the donor polynucleotide comprise 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequence of SEQ ID NO:104 or 109. In some embodiments, the guiding polynucleotide and the donor polynucleotide comprise 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequence of SEQ ID NO:230-233.

[0349] Table 8: Exemplary donor polynucleotide sequences and target sites

[0350] In some embodiments, the donor polynucleotide encodes a protein or a fragment thereof. In some embodiments, the protein is a fluorescent protein. In some embodiments, the fluorescent protein is GFP, eGFP, RFP, YFP, BFP, or CFP.

[0351] In some embodiments, the donor polynucleotide further comprises a coding sequence for a self-cleaving peptide. Exemplary self-cleaving peptides include T2A, P2A, E2A, and F2A. In some embodiments, the T2A self-cleaving peptide comprises the sequence of EGRGSLLTCGDVEENPGP. In some embodiments, the P2A self-cleaving peptide comprises the sequence of ATNFSLLKQAGDVEENPGP. In some embodiments, the E2A self-cleaving peptide comprises the sequence of QCTNYALLKLAGDVESNPGP. In some embodiments, the F2A self-cleaving peptide comprises the sequence of VKQTLNFDLLKLAGDVESNPGP. In some embodiments, the T2A self-cleaving peptide comprises the sequence of EGRGSLLTCGDVEENPGP. Any of the foregoing may also comprise an N-terminal GSG linker. For example, the T2A self-cleaving peptide may also comprise the sequence of GSGEGRGGSLLTCGDVEENPGP.

[0352] In some embodiments, the donor polynucleotide further comprises a trans-acting double-stranded RNA polynucleotide having a 3'- or 5'-overhanging nucleotide. In some embodiments, the donor polynucleotide further comprises a cis-acting single-stranded RNA polynucleotide with sequence similarity to the target strand of the nuclease. In some embodiments, the donor polynucleotide further comprises a cis-acting single-stranded RNA polynucleotide with sequence similarity to the non-target strand of the nuclease. In some embodiments, the donor polynucleotide further comprises one or more binding sites for genomic modifying factors, optionally binding sites for site-specific recombinases (such as serine recombinases or LoxP target sites). In some embodiments, the donor polynucleotide further comprises a repair template for protein-coding sequences. In some embodiments, the donor polynucleotide further comprises one or more exons having a splice acceptor and a donor sequence. In some embodiments, the donor polynucleotide further comprises one or more optional sequences selected from the group consisting of NeoR, BsdR, HygR, PuroR, and BleoR genes. In some embodiments, the donor polynucleotide further comprises one or more drug-inducible regulatory sequences for controlled gene expression. In some implementations, the donor polynucleotide comprises a combination of the above.

[0353] In some implementations, the donor polynucleotide contains a self-complementary single-stranded RNA repair template separated by a ribozyme cleavage site, which acts as a trans-acting double-stranded nucleic acid during transcription to repair the target sequence via the NHEJ pathway.

[0354] promoter In another aspect, this document provides certain expression control sequences that can be used to express the recombinant nucleic acids provided herein. Non-limiting exemplary embodiments of the recombinant nucleic acids of this disclosure may include one or more of the following features.

[0355] In some implementations, any of the recombinant nucleic acids provided herein can be operatively linked, for example, to other structural elements (e.g., promoter sequences) required for the expression of such recombinant nucleic acids, such as when placed in a host cell, in a subject, or in an ex vivo cell-free expression system.

[0356] As used herein, the terms “promoter” and “promoter sequence” are used interchangeably and refer to a DNA sequence that promotes the expression of a protein-coding open reading frame or a nucleotide sequence encoding a functional RNA (e.g., a polynucleotide or donor polynucleotide). Those skilled in the art will understand that different promoters direct gene expression in different tissues or cell types, at different developmental stages, or in response to different environmental or physiological conditions.

[0357] In some implementations, the nucleic acid also contains one or more promoters.

[0358] In some embodiments, the nucleic acid further comprises a promoter that drives the expression of the chimeric nuclease and the guide RNA together. In some embodiments, the promoter is operatively linked to the nucleic acid encoding the chimeric nuclease. In some embodiments, the promoter is operatively linked to the nucleic acid encoding the guide RNA. In some embodiments, the promoter is operatively linked to the nucleic acid sequence encoding the entire operon of this disclosure.

[0359] In some embodiments, the promoter is a CMV promoter, an SV40 promoter, a minimal cytomegalovirus (CMV) promoter, or a human extension factor-1α (EF1a) promoter. In some embodiments, the promoter is a miniature CMV promoter. In some embodiments, the promoter is an MND promoter. In some embodiments, the promoter is a U6 promoter. In some embodiments, the promoter is a T7 promoter.

[0360] In some implementations, the promoter is a tissue- or cell-specific promoter. In some implementations, the promoter is a muscle-specific synthesis promoter, the SPc5-12 neuron-specific promoter hSYN1, the aldh1L1 promoter, the cTNT promoter, the α-MHC promoter, the SPc5-12 promoter, the MUC2 promoter, the Ksp-cadherin promoter, the albumin promoter, the HAS promoter, the insulin promoter, the rhodopsin promoter, the rNSE promoter, or the cone protein promoter. In some implementations, the promoter is the T7 promoter.

[0361] In some implementations, the promoter is the CMV promoter listed in SEQ ID NO:6.

[0362] In some embodiments, a donor polynucleotide is inserted into the genome within a frame of an endogenous coding sequence. In some embodiments, the donor polynucleotide is inserted into the genome within a frame for expression from an endogenous promoter.

[0363] Ribozymes and transfer RNA In some embodiments, the nucleic acid also includes a nucleic acid sequence encoding one or more self-cleaving ribozymes or endogenous RNA processing sequences (e.g., tRNA). A self-cleaving ribozyme is a catalytic RNA molecule that cleaves its own phosphodiester backbone. Introducing a self-cleaving ribozyme ensures that the 5' or 3' end of the guide RNA is clean, depending on the placement of the ribozyme within the mRNA. In some embodiments, the ribozyme is placed at the 5' end of the guide RNA. In some embodiments, the ribozyme is placed directly adjacent to the 5' end of the guide RNA.

[0364] In some embodiments, the ribozyme is a hammerhead ribozyme or a hepatitis D virus (HDV) ribozyme. In some embodiments, the ribozyme is a hammerhead ribozyme encoded by the nucleic acid sequence according to SEQ ID NO:8, SEQ ID NO:9, or SEQ ID NO:111. In some embodiments, the ribozyme is an HDV ribozyme encoded by the nucleic acid sequence according to SEQ ID NO:10 or SEQ ID NO:112.

[0365] Exemplary ribozyme sequences are listed in Table 9. In some embodiments, the ribozyme is encoded by a nucleic acid sequence according to any one of SEQ ID NO: 289 to 301.

[0366] Table 9. Exemplary ribozyme sequences for conditionally cleavable elements

[0367] In some embodiments, the ribozyme sequence is located within the tRNA sequence and is trans-acting. In some embodiments, the tRNA sequence containing the ribozyme sequence is listed in SEQ ID NO:5. It should be understood that the DNA sequence is transcribed into an RNA sequence, and the ribozyme is active when present as RNA. In some embodiments, the tRNA is a glycine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine tRNA. In some embodiments, one or more tRNAs separate one or more guide RNAs. When expressed as RNA through an endogenous tRNA maturation process, the tRNA is conditionally cleaved within the cell, resulting in RNA fragmentation into two or more parts. In some embodiments, one or more tRNAs separate an RNA stability sequence from one or more tRNAs. In some embodiments, the tRNA encodes a nucleic acid sequence according to SEQ ID NO:127 or SEQ ID NO:128. In some embodiments, the tRNA encodes a nucleic acid sequence according to the nucleic acid sequences in Table 10. In some embodiments, the tRNA encodes a nucleic acid sequence according to any one of SEQ ID NO:302 to SEQ ID NO:733.

[0368] Table 10. Exemplary tRNA sequences for conditionally cleavable elements.

[0369]

[0370] stable RNA sequence In some implementations, nucleic acids also include RNA-stable polynucleotides.

[0371] In some implementations, the nucleic acid contains a polyadenylation signal. In some implementations, the polyadenylation signal includes simian virus 40 (SV40), α-globin, β-globin, human growth hormone (hGH), bovine growth hormone (BGH), herpes simplex virus type 1 thymidine kinase (HSV TK), or synthetic polyadenylation (Synt poly A) polyadenylation signal.

[0372] In some embodiments, the RNA-stable polynucleotide is a multi-A signal. The multi-A tail acts as a binding site for multi-A binding proteins. Multi-A binding proteins promote export from the nucleus and translation, and inhibit degradation. In some embodiments, the multi-A signal is listed in SEQ ID NO:11. Typically, the multi-A signal is located at the 3' end of the mRNA transcript. The multi-A signal cannot typically be located at or inside the 5' end of the mRNA transcript. Ribozyme cleavage exposes the free 3' end of the mRNA transcribed from the nucleic acid described herein. To protect the 3' end from degradation, a second RNA-stable polynucleotide in addition to the multi-A signal can be introduced.

[0373] In some embodiments, the stable RNA polynucleotide is the 3' sequence of metastasis-associated lung adenocarcinoma transcript 1 (MALAT) or the 3' end of multiple endocrine vegetation β transcript (MEN β). In some embodiments, the MALAT sequence is a nucleic acid sequence as listed in SEQ ID NO:7.

[0374] In some implementations, the RNA stable polynucleotide is the 3' end of a triple-helical RNA structure or an RNA transcript lacking typical polyadenylation signals.

[0375] In some embodiments, the stable RNA polynucleotide is located at the 3' end of the polynucleotide encoding the chimeric nuclease. In some embodiments, the stable RNA polynucleotide is located at the 5' end of the ribozyme cleavage site. In some embodiments, the stable RNA polynucleotide protects the mRNA encoding the polypeptide portion of the chimeric nuclease from degradation.

[0376] In some implementations, the nucleic acid also includes a 5' UTR sequence and / or a 3' UTR sequence.

[0377] In some embodiments, a composition is provided comprising a chimeric nuclease polypeptide containing an I-TevI ​​domain and an RNA-directed nuclease domain, and the nucleic acid of this disclosure.

[0378] In some embodiments, a composition is provided comprising a chimeric nuclease nucleic acid encoding a chimeric nuclease comprising an I-TevI ​​domain and an RNA-directed nuclease domain, and the nucleic acid of this disclosure. In some embodiments, the chimeric nuclease nucleic acid is mRNA.

[0379] Other components In another aspect, the nucleic acids provided herein may also contain additional regulatory elements. In some embodiments, the additional regulatory elements may be WPRE sequences, multiple A sequences, and / or viral replication and packaging control sequences (long terminal repeat (LTR) sequences).

[0380] In some implementations, the recombinant nucleic acid includes a polyadenylation (poly A) signal. In some implementations, the polyadenylation signal includes simian virus 40 (SV40), α-globin, β-globin, human growth hormone (hGH), bovine growth hormone (BGH), herpes simplex virus type 1 thymidine kinase (HSV TK), or synthetic polyadenylation (Synt poly A) signal.

[0381] In some implementations, the viral replication and packaging control sequences are long terminal repeat (LTR) sequences. In some implementations, the viral replication and packaging control sequences are inverted terminal repeat (ITR) sequences.

[0382] carrier In some embodiments, the nucleic acids of this disclosure may be incorporated into expression vectors or vectors for viral production. The vector may be a plasmid, bacteriophage, or granule, into which another DNA segment may be inserted to induce replication of the inserted segment. In some embodiments, the expression vector may be an integration vector. Therefore, vectors, plasmids, or viruses encoding one or more nucleic acids of any of the nucleic acids disclosed herein are also provided herein. The nucleic acids described above may be contained within vectors capable of directing their expression in, for example, cells already transduced with the vector. Suitable vectors for use in eukaryotic and prokaryotic cells are known in the art and are commercially available or readily prepared by those skilled in the art. Additional vectors may also be found, for example, in Ausubel, FM, et al., Current Protocols in Molecular Biology, (Current Protocol, 1994) and Sambrook et al., “Molecular Cloning: A Laboratory Manual,” 2nd ed. (1989).

[0383] Viral vector In some embodiments, the nucleic acids of this disclosure may be packaged in viral particles or viral vectors. Methods for generating viral vectors from various virus types are known in the art. Exemplary types of viral particles that can be recombinantly engineered as delivery media include retroviruses, lentiviruses (e.g., HIV and its derivatives and SIV), adeno-associated viruses, adenoviruses, MMLV retroviruses, MSCV retroviruses, baculoviruses, vesicular stomatitis viruses, herpes simplex viruses, and poxviruses. Examples include adeno-associated virus (AAV) particles for gene therapy. In some embodiments, the virus may be AAV. In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-DJ, AAV2.5T, or AAVmyo. In some embodiments, the virus may be lentivirus.

[0384] Pharmaceutical Composition In another aspect, this document provides compositions and pharmaceutical compositions suitable for administration to human subjects, said compositions and pharmaceutical compositions comprising any of the chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors, or viral vectors disclosed and described herein. The chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors, or viral vectors of this disclosure can be formulated into compositions, including pharmaceutical compositions. Such compositions typically comprise one or more chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors, or viral vectors as disclosed and described herein, and pharmaceutically acceptable excipients, such as transport vectors. In some embodiments, the compositions of this disclosure are formulated for the treatment or management of a health condition. In some embodiments, the health condition is a congenital disease. For example, the compositions of this disclosure can be formulated into therapeutic compositions or pharmaceutical compositions or mixtures thereof comprising pharmaceutically acceptable excipients. In some embodiments, the compositions of this disclosure are formulated as a therapy for cystic fibrosis. In some embodiments, the compositions of this disclosure are formulated as a therapy for α-1-antitrypsin deficiency. In some embodiments, the compositions of this disclosure are formulated as a therapy for myotonic dystrophy type 1. In some embodiments, the compositions of this disclosure are formulated for use as a treatment for amyotrophic lateral sclerosis or frontotemporal dementia.

[0385] Therefore, in one aspect, this document provides pharmaceutical compositions comprising pharmaceutically acceptable excipients and nucleic acids, vectors, or viral vectors of the present disclosure.

[0386] In some embodiments, a pharmaceutical composition is provided comprising the chimeric nuclease, chimeric nuclease system, nucleic acid, vector, viral vector, AAV of the present disclosure, composition of the present disclosure, or LNP composition of the present disclosure; and excipients.

[0387] In some embodiments, the compositions described herein, such as chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors or viral vectors, and / or pharmaceutical compositions, are incorporated into a therapeutic composition for use in methods of preventing or treating subjects with cystic fibrosis, suspected of having cystic fibrosis, or who may be at high risk of developing cystic fibrosis. In some embodiments, the compositions described herein, such as chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors or viral vectors, and / or pharmaceutical compositions, are incorporated into a therapeutic composition for use in methods of preventing or treating subjects with alpha-1-antitrypsin deficiency, suspected of having alpha-1-antitrypsin deficiency, or who may be at high risk of developing alpha-1-antitrypsin deficiency. In some embodiments, the compositions described herein, such as chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors or viral vectors, and / or pharmaceutical compositions, are incorporated into a therapeutic composition for use in methods of preventing or treating subjects with myotonic dystrophy type 1, suspected of having myotonic dystrophy type 1, or who may be at high risk of developing myotonic dystrophy type 1. In some embodiments, the compositions described herein, such as chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors or viral vectors and / or pharmaceutical compositions, are incorporated into a therapeutic composition for use in methods of preventing or treating subjects who have amyotrophic lateral sclerosis or frontotemporal dementia, are suspected of having amyotrophic lateral sclerosis or frontotemporal dementia, or may be at high risk of developing amyotrophic lateral sclerosis or frontotemporal dementia.

[0388] In these cases, the composition should be sterile, formulated for easy administration to human subjects, and stable under manufacturing and storage conditions to resist contamination by microorganisms such as bacteria and fungi.

[0389] In some embodiments, the composition is formulated for one or more of intramuscular, intratumoral, intravenous, intratracheal, intraperitoneal, or intracranial administration.

[0390] Cell delivery Chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors, or viral vectors disclosed and described herein can be delivered to cells for genome editing.

[0391] In some embodiments, delivery is either non-viral or viral. In some embodiments, the nuclease can be delivered as a ribonucleoprotein complex, DNA encoding the nuclease system, or messenger RNA encoding the nuclease system, or a combination thereof. For example, the polypeptide portion of the nuclease can be delivered to the cell as a protein, and the guide RNA / donor can be delivered as RNA.

[0392] In some embodiments, this disclosure provides a vector comprising the nucleic acids described herein. In some embodiments, a chimeric nuclease system and the nucleic acid encoding the chimeric nuclease system can be delivered to cells using lipid transfection or polymer-based transfection. In some embodiments, the nucleic acid is contained in lipid nanoparticles (LNPs). In some embodiments, the chimeric nuclease system is contained in lipid nanoparticles.

[0393] In some embodiments, viral delivery can be used to deliver nucleic acids encoding a chimeric nuclease system into cells. In some embodiments, the nucleic acid is packaged into a viral vector. Mammalian cells can be transduced using the nucleic acids of this disclosure using a viral vector. In some embodiments, the virus is a lentivirus, adeno-associated virus (AAV), adenovirus, retrovirus, or modified herpes simplex virus (HSV). In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-DJ, AAV2.5T, or AAVmyo.

[0394] In some embodiments, the cells are mammalian cells. In some embodiments, the cells are human cells. In some embodiments, the cells are immune cells (such as T cells), hematopoietic stem cells, mesenchymal stem cells, or induced pluripotent stem cells (iPSCs). Cell types that can be modified in vitro by the methods and nucleases described herein include immune cells, such as T cells or NK cells; or pluripotent cells, such as mesenchymal stem cells, hematopoietic stem cells, or cells additionally induced to be pluripotent using techniques known in the art. In some embodiments, the cells are muscle cells. In some embodiments, the cells are brain cells.

[0395] Methods using chimeric nuclease systems In one aspect of this disclosure, methods are provided for editing cellular genomes using the chimeric nuclease system described herein and the nucleic acid encoding the chimeric nuclease system.

[0396] In some embodiments, the method includes contacting cells with the nucleic acid of this disclosure. In some embodiments, the nucleic acid is DNA. In some embodiments, the nucleic acid is RNA. In some embodiments, the nucleic acid is mRNA. In some embodiments, the method includes contacting cells with a virus containing the nucleic acid of this disclosure. In some embodiments, the method includes contacting cells with a vector containing the nucleic acid of this disclosure. In some embodiments, the method includes contacting cells with an LNP containing the nucleic acid of this disclosure.

[0397] In some embodiments, a method is provided for removing a precise length of DNA from the genome of a cell, the method comprising providing the cell with a chimeric nuclease that cleaves at two sites to remove the precise length of DNA and leave two distinct DNA ends. The two DNA ends provide the necessary target sites for accurate repair via error-free NHEJ or RNA-dependent Polθ-mediated repair, thereby enabling repair at all cell cycle stages, including G1 / G0.

[0398] In some embodiments, a method is provided for editing the genome of a cell at a chimeric nuclease target site, the method comprising providing the cell with a chimeric nuclease having a repair polynucleotide having a two-base overhang corresponding to a 3' base adjacent to an I-TevI ​​target site in the cell's genome. In some embodiments, a method is provided for editing the genome of a cell at a target site, the method comprising providing the cell with a chimeric nuclease having a repair polynucleotide having a two- to eighteen-base overhang corresponding to a 3' base adjacent to an I-TevI ​​target site in the cell's genome. In some embodiments, a method is provided for editing the genome of a cell at a target site, the method comprising providing the cell with a chimeric nuclease having a repair polynucleotide having a fourteen-base overhang corresponding to a 3' base adjacent to an I-TevI ​​target site in the cell's genome. In some embodiments, the repair polynucleotide comprises a donor polynucleotide. In some embodiments, the repair polynucleotide comprises a guide polynucleotide.

[0399] In some implementations, a method is provided for delivering messenger RNA encoding a chimeric nuclease comprising an I-TevI ​​domain and an RNA-guided nuclease domain to a cell, the method comprising contacting the cell with a polynucleotide encoding one or more guide RNAs and a polynucleotide donor.

[0400] In some embodiments, a method for genetically modifying a cell genome is provided, the method comprising contacting the cell with a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure.

[0401] In some implementations, modifications include the insertion, deletion, substitution, or mutation of the cell's genome.

[0402] In some implementations, inserting the donor polynucleotide into the cell's genome results in the removal of the sequence between the I-TevI ​​target site and the Cas9 target site.

[0403] In some implementations, the cell is a mammalian cell. In some implementations, the cell is a human cell.

[0404] In some embodiments, a method is provided for inserting or replacing sequences at chimeric nuclease target sites in the genome of a cell, the method comprising: contacting a cell with a nucleic acid comprising a chimeric nuclease containing a Cas9 domain and an I-TevI ​​domain, and a nucleic acid comprising a director polynucleotide and a donor polynucleotide; wherein the director polynucleotide and the chimeric nuclease form a complex, and the complex binds to and cleaves genomic DNA at the Cas9 target site and the I-TevI ​​target site; wherein the 3' end of the donor polynucleotide contains at least two bases of complementarity at the 5' end of the I-TevI ​​target site; and wherein the donor polynucleotide is incorporated into the chimeric nuclease target site at the 5' position of the Cas9 target site.

[0405] In some embodiments, the cellular polymerase targets the chimeric nuclease target site. In some embodiments, the cellular polymerase is polymerase θ.

[0406] In some implementations, the 3' end of the guide polynucleotide is linked to the 5' end of the donor polynucleotide.

[0407] In some embodiments, a method is provided for editing the genome of a cell at a target site, the method comprising providing the cell with a chimeric nuclease having a repair polynucleotide, said repair polynucleotide being a cis-acting single-stranded nucleic acid encoding a sense strand or an antisense strand of a repair sequence, and repair being performed via an RNA-dependent Rad52-mediated repair pathway.

[0408] In some embodiments, a method is provided for editing the genome of a cell at a target site, the method comprising providing the cell with a purified protein-chimeric nuclease complexed with one or more synthetic guide RNA sequences, the one or more synthetic guide RNA sequences being separated by one or more synthetic tRNA sequences encoding trans-acting ribozyme sequences, and further comprising a donor DNA sequence separated by one or more ribozyme cleavage sites.

[0409] In some implementations, a method is provided for editing the genome of a cell at a target site to produce a large, predictable deletion at the target site via error-free NHEJ, the method comprising providing the cell with a dual-directed chimeric nuclease system comprising a non-catalytically active Cas9 domain, a catalytically active I-TevI ​​domain, and two guide RNAs targeting two different target sites.

[0410] In some implementations, a method is provided to replace the genome of a cell. CFTRA method for at least a portion of a gene, said method comprising contacting a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure with a cell. In some embodiments, the guide RNA and donor polynucleotide are targeted. CFTR Mutations in genes. In some implementations, chimeric nucleases with guide RNA and donor polynucleotides target and replace... CFTR c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), c.1652G>A (p.Gly551Asp), c.1657C>T (p.Arg553Ter) or c.3846G>A (p.Trp1282Ter) mutation.

[0411] In some implementations, a method is provided to replace the genome of a cell. SERPINA1 A method for at least a portion of a gene, said method comprising contacting a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure with a cell. In some embodiments, the guide RNA and donor polynucleotide are targeted. SERPINA1 Mutations in genes. In some implementations, guide RNA and donor polynucleotides target and replace... SERPINA1 c.1096G>A (p.Glu342Lys) mutation.

[0412] In some implementations, a method is provided to excise the genome of a cell. DMPK A method for at least a portion of a gene, said method comprising contacting a nucleic acid of this disclosure, a vector of this disclosure, a viral vector of this disclosure, or an AAV of this disclosure, a composition of this disclosure, or an LNP composition of this disclosure with a cell. In some embodiments, a dual-guided chimeric nuclease excises and deletes... DMPK The (CTG)n * (CAG)n trinucleotide repeat sequence in the 3' untranslated region of the gene.

[0413] In some implementations, a method is provided to excise the genome of a cell. C9ORF72 A method for at least a portion of a gene, said method comprising contacting a nucleic acid of this disclosure, a vector of this disclosure, a viral vector of this disclosure, or an AAV of this disclosure, a composition of this disclosure, or an LNP composition of this disclosure with a cell. In some embodiments, a dual-guided chimeric nuclease excises and deletes... C9ORF72 The GGGGCC repeat sequence in the intron sequence between alternating 5' exons of a gene. In some implementations, chimeric nucleases, guide RNA, and donor polynucleotides target and remove... C9ORF72 A large hexanucleotide GGGGCC repeat sequence between exon 1a and exon 1b of the gene.

[0414] In some implementations, a method is provided to remove genome from cells DMPK A method for at least a portion of a gene, said method comprising contacting a nucleic acid of this disclosure, a vector of this disclosure, a viral vector of this disclosure, or an AAV of this disclosure, a composition of this disclosure, or an LNP composition of this disclosure with a cell. In some embodiments, two guide RNAs of this disclosure target... DMPK Two distinct mutations in the gene. In some implementations, chimeric nucleases, guide RNA, and donor polynucleotides target and remove... DMPK The large triplet CAG repeat sequence in the 3' untranslated region of the gene.

[0415] In some embodiments, a method is provided for tunably editing the genome of a cell at a target site, the method comprising providing the cell with nucleic acid expressing a chimeric nuclease system described herein, the nucleic acid further comprising a self-inactivating sequence cleaved by the chimeric nuclease. In some embodiments, the genomic target site is B2M, and the target sequence is SEQ ID NO:51 or SEQ ID NO:130. In some embodiments, the self-inactivating sequence is SEQ ID NO:52 or SEQ ID NO:131.

[0416] In some embodiments, a method is provided for inserting a sequence into a target site of a cell’s genome at a target site, the method comprising providing the cell with nucleic acids comprising the nuclease system described herein and donor polynucleotides.

[0417] In some embodiments, a method is provided for inserting a polypeptide sequence into a target site in the genome of a cell at a target site, the method comprising providing the cell with nucleic acids comprising the nuclease system described herein and donor polynucleotides.

[0418] In some embodiments, a method is provided for removing disease-causing large repetitive DNA sequences from the genome of a cell, the method comprising providing a chimeric nuclease that cleaves at two sites to remove the large repetitive DNA sequences.

[0419] In some embodiments, a method is provided for inserting a donor polynucleotide into an endogenous promoter frame, the method comprising providing a cell with a nucleic acid comprising a nuclease system described herein and a donor polynucleotide encoding a polypeptide and a T2A cleavable peptide sequence.

[0420] In some implementations, a method is provided to replace the genome in the cell. DMPKA method for at least a portion of a gene, said method comprising contacting a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure with a cell. In some embodiments, one or more guide RNAs target... DMPK Mutations in genes. In some implementations, one or more guide RNAs target... DMPK The CAG triplet polynucleotide sequence in the 3' untranslated region of the gene.

[0421] In some embodiments, the length of the cut or deletion is 30 base pairs. In some embodiments, the length of the deletion is 35 base pairs. In some embodiments, the length of the deletion is 40 base pairs. In some embodiments, the length of the deletion is 45 base pairs. In some embodiments, the length of the deletion is 50 base pairs. In some embodiments, the length of the deletion is 60 base pairs. In some embodiments, the length of the deletion is 70 base pairs. In some embodiments, the length of the deletion is 80 base pairs. In some embodiments, the length of the deletion is 90 base pairs. In some embodiments, the length of the deletion is 150 base pairs. In some embodiments, the length of the deletion is 30 base pairs. In some embodiments, the length of the deletion is 200 base pairs. In some embodiments, the length of the deletion is 300 base pairs. In some embodiments, the length of the deletion is 400 base pairs. In some embodiments, the length of the deletion is 500 base pairs. In some embodiments, the length of the deletion is 600 base pairs. In some embodiments, the deletion length is 700 base pairs. In some embodiments, the deletion length is 800 base pairs. In some embodiments, the deletion length is 900 base pairs. In some embodiments, the deletion length is 1000 base pairs. In some embodiments, the deletion length is 1500 base pairs. In some embodiments, the deletion length is 2000 base pairs. In some embodiments, the deletion length is 2500 base pairs. In some embodiments, the deletion length is 3000 base pairs. In some embodiments, the deletion length is 4000 base pairs. In some embodiments, the deletion length is 5000 base pairs. In some embodiments, the deletion length is 6000 base pairs. In some embodiments, the deletion length is 7000 base pairs. In some embodiments, the deletion length is 8000 base pairs. In some embodiments, the deletion length is 9000 base pairs. In some embodiments, the deletion length is 10000 base pairs. In some implementations, the missing length is more than 10,000 base pairs.

[0422] Treatment In another respect, this article provides methods for treating human subjects in need using chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors, or viral vectors disclosed and described herein.

[0423] The chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors, or viral vectors disclosed and described herein are suitable for use as gene therapies for treating subjects with corresponding needs. In some embodiments, the chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors, or viral vectors disclosed and described herein include therapeutic agents for use in methods of treating subjects who have one or more health conditions or diseases treatable with such gene therapies, are suspected of having one or more health conditions or diseases treatable with such gene therapies, or are at high risk of developing one or more health conditions or diseases treatable with such gene therapies. Exemplary health conditions or diseases may include, but are not limited to, congenital diseases such as cystic fibrosis, α-1-antitrypsin deficiency, myotonic dystrophy type 1, amyotrophic lateral sclerosis, or frontotemporal dementia.

[0424] Cystic fibrosis Cystic fibrosis (CF) is caused by the encoding of epithelial anion channels. CFTR Cystic fibrosis (CF) is an autosomal recessive disease caused by mutations in the gene. The CFTR protein, a transmembrane transport regulator of cystic fibrosis, is found in a wide range of organs, including the pancreas, kidneys, liver, lungs, gastrointestinal tract, and reproductive tract, making CF a multi-organ disease. CFTR Mutations in the gene (GRCh38.p14GCF_000001405.40, NM_000492.4) can lead to suboptimal ion transport and fluid retention, resulting in prominent clinical manifestations of abnormal thickening of mucus in the lungs and pancreatic insufficiency. In the lungs, dysfunctional CFTR can impede mucociliary clearance, making the organ susceptible to bacterial infection and inflammation, ultimately leading to airway obstruction, respiratory failure, and premature death. CF remains the most common and deadliest genetic disease in Caucasian populations, affecting an estimated 70,000–100,000 people worldwide, highlighting the real need to develop better treatments.

[0425] CFTR Mutations in the process include, but are not limited to, c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter) or c.3846G>A (p.Trp1282Ter) .

[0426] In some embodiments, a method for treating cystic fibrosis in patients with a corresponding need is provided, the method comprising administering a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure.

[0427] In some implementations, the RNA and donor polynucleotides are directed to target CFTR Mutations in genes.

[0428] In some implementations, the guide RNA and donor polynucleotides are targeted and replaced. CFTR c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), c.1652G>A (p.Gly551Asp), c.1657C>T (p.Arg553Ter) or c.3846G>A (p.Trp1282Ter) mutation.

[0429] α1-Antipase Deficiency Alpha-1 antitrypsin deficiency is a hereditary disorder that affects the lungs and sometimes the liver. The gene affected in alpha-1 antitrypsin deficiency is located on chromosome 14. SERPINA1 (GRCh38.p14 (GCF_000001405.40)) (encoding α-1-antitrypsin). The protein encoded by this gene is a serine protease inhibitor belonging to the serine protease inhibitor superfamily. Its targets include elastase, plasmin, thrombin, trypsin, chymotrypsin, and plasminogen activator. This protein is produced in the liver and bone marrow by lymphocytes and monocytes in lymphoid tissues, and by Panette cells in the intestine. Defects in this gene are associated with chronic obstructive pulmonary disease, emphysema, and chronic liver disease.

[0430] In some embodiments, a method is provided for treating α-1-antitrypsin deficiency in patients with corresponding needs, the method comprising administering to the patient a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure.

[0431] In some implementations, the RNA and donor polynucleotides are directed to target SERPINA1 Mutations in genes. In some implementations, guide RNA and donor polynucleotides target and replace... SERPINA1 c.1096G>A (p.Glu342Lys) mutation.

[0432] Type 1 myotonic dystrophy Myotonic dystrophy type 1 (DM1) is a multisystem disorder affecting skeletal and smooth muscle, as well as the eyes, heart, endocrine system, and central nervous system. Clinical findings spanning a continuous range from mild to severe have been classified into three somewhat overlapping phenotypes: mild, classic, and congenital. DM1 is caused by... DMPK The amplification of CTG trinucleotide repeats in the non-coding region causes this. Diagnosis of DM1 is suspected in individuals with characteristic muscle weakness, and is diagnosed through... DMPK Molecular genetic testing confirmed this. A CTG repeat length exceeding 34 repeats is abnormal. Molecular genetic testing can detect the pathogen in almost 100% of affected individuals.

[0433] In some embodiments, a method is provided for treating type 1 myotonic dystrophy in patients with corresponding needs, the method comprising administering to the patient a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure.

[0434] Amyotrophic Lateral Sclerosis Amyotrophic lateral sclerosis (ALS) is a neurological disorder affecting motor neurons (nerve cells) that control voluntary muscle movement and respiration in the brain and spinal cord. C9orf72 frontotemporal dementia and / or amyotrophic lateral sclerosis (C9orf72-FTD / ALS) is most commonly characterized by frontotemporal dementia (FTD) as well as upper motor neuron disease and lower motor neuron disease (MND). C9orf72-FTD / ALS shows a heterozygous aberration in C9orf72, specifically an amplified hexanucleotide repeat of the G4C2 (GGGGCC, C4C2) hexanucleotide repeat, which can be identified by molecular genetic testing.

[0435] In some embodiments, a method is provided for treating amyotrophic lateral sclerosis or frontotemporal dementia in patients with corresponding needs, the method comprising administering to the patient a nucleic acid of the present disclosure, a vector of the present disclosure, a viral vector of the present disclosure, or an AAV of the present disclosure, a composition of the present disclosure, or an LNP composition of the present disclosure.

[0436] Reagent test kit This document also provides various kits for carrying out the methods described herein, along with written instructions for their preparation and use. In particular, some embodiments relate to kits for methods of treating diseases in subjects with corresponding needs. For example, in some embodiments, kits are provided herein comprising one or more nucleic acid and / or pharmaceutical compositions as provided and described herein, along with written instructions for their use. In some embodiments, the kits of this disclosure also include one or more means for administering engineered immune cells and / or pharmaceutical compositions to a subject. For example, in some embodiments, the kits of this disclosure also include one or more bags, syringes (including pre-filled syringes) for administering any of the provided engineered immune cells and / or pharmaceutical compositions to a subject.

[0437] In some embodiments, the kit may also include instructions for use in carrying out the methods disclosed herein using the components of the kit. For example, the kit may include a packaging insert containing information about the pharmaceutical compositions and dosage forms in the kit. Typically, such information helps patients and physicians use the included pharmaceutical compositions and dosage forms effectively and safely. For example, the insert may provide information about combinations of the present disclosure, including: pharmacokinetics, pharmacodynamics, clinical studies, efficacy parameters, indications and usage, contraindications, warnings, precautions, adverse reactions, overdose, appropriate dosage and administration, method of supply, appropriate storage conditions, references, manufacturer / distributor information, and intellectual property information.

[0438] Instructions for use in implementing the method are typically documented on a suitable recording medium. For example, instructions for use may be printed on a substrate such as paper or plastic. Instructions for use may be present as a packaging insert in the kit, on a label of a container of the kit or its components (e.g., associated with the package or sub-package), etc. Instructions for use may exist as an electronic storage data file on a suitable computer-readable storage medium (e.g., CD-ROM, floppy disk, flash drive, etc.). In some cases, no actual instructions for use are included in the kit, but means for obtaining the instructions for use from a remote resource (e.g., via the Internet) are provided. An example of this implementation is a kit that includes a URL at which the instructions for use can be viewed and / or downloaded. Like the instructions for use, such means for obtaining the instructions for use may be documented on a suitable substrate.

[0439] Further embodiments are disclosed in detail in the following examples, which are provided by way of illustration and are not intended in any way to limit the scope of this disclosure or the claims. Example

[0440] These embodiments are provided for illustrative purposes only and are not intended to limit the scope of the claims provided herein.

[0441] Example 1. Design of an integrated chimeric nuclease kit This example describes the design and synthesis of an integrated chimeric nuclease kit.

[0442] In short, a chimeric nuclease cassette is designed, comprising a promoter, a chimeric nuclease, a stable mRNA sequence, gRNA, a modified tRNA containing a ribozyme sequence, a repair template, and a multi-A tail. An exemplary nucleic acid expression cassette is shown in... Figure 1 , Figure 2 , Figure 26A , Figure 26B and Figure 26CThe transcription of a 3-in-1 construct, driven by a minimal CMV promoter, comprises Dualase (a chimeric nuclease), long non-coding RNA MALAT-1, a synthetic transfer RNA (tRNA') containing a sequence-specific ribozyme, a gRNA fused to a repair template (RT), and a synthetic multi-A signal sequence (synt[A]). Posttranscriptionally, RNA maturation at the 3' end of MALAT and both ends of the tRNA leads to the dissociation of Dualase-MALAT RNA, tRNA, and gRNA. U-rich repeats on MALAT help protect the Dualase RNA from degradation, while tRNA' recognizes and cleaves in regions between RTs and between RTs and the synt[A] sequence. Posttranslational, Dualase can form a complex (RNP) with the gRNA and cleave the target site. The presence of a suitable repair template (cis-sense and antisense) fused to the 3' end of the gRNA, or abundant free-floating repair templates (trans-antisense and trans-sense), serves as a bridge between the two cleavage sites and a local reference template for cellular repair mechanisms.

[0443] Another exemplary chimeric nuclease cassette was designed, comprising a promoter, a chimeric nuclease, a self-cleaving peptide, a protein tag, a stable mRNA sequence, gRNA, a modified tRNA containing a ribozyme sequence, a repair template, and a multi-A tail. An exemplary nucleic acid expression cassette is shown in... Figure 2 The integrated mRNA cassette encodes essential elements for efficient and accurate target site disruption or repair. T7 promoter drives transcription of a 3-in-1 construct containing Dualase, the long non-coding RNA MALAT-1, a synthetic transfer RNA (tRNA') containing a sequence-specific ribozyme, a gRNA fused to a repair template (RT), and a synthetic multi-A signal sequence (synt[A]). Conditions are established where tRNA, MALAT maturation, and ribozyme cleavage activity are inactive in storage buffer. Upon delivery, RNA maturation at the 3' end of MALAT and both ends of tRNA' leads to the separation of Dualase-MALAT RNA, tRNA, and gRNA. U-rich repeats on MALAT help protect the Dualase RNA from degradation, while tRNA' recognizes and cleaves in regions between RTs and between RTs and the synt[A] sequence. After translation, Dualase forms a complex (RNP) with gRNA and cleaves the target site. The presence of appropriate repair templates (cis-sense and antisense) fused to the 3' end of gRNA or abundant free-floating repair templates (trans-antisense and trans-sense) serves as a bridge between two cleavage sites and a local reference template for cellular repair mechanisms.

[0444] Example 2. Delivery of chimeric nuclease This example describes the delivery of a chimeric nuclease system into cells.

[0445] In short, the chimeric nuclease is expressed and purified in *E. coli*. This produces RNA containing gRNA fused to the repair template. The purified chimeric nuclease and guide RNA / repair template are assembled into a ribonucleoprotein (RNP) complex and co-delivered to cells, for example, using LNP or via electroporation.

[0446] Figure 3 A diagram of a dual-cleavage nuclease ribonucleoprotein (RNP) complex with an integrated guide RNA repair cassette is shown. Dualase and 2-in-1 gRNA-RT were incubated to form the RNP complex and deliver it to target cells. Following nuclear translocation, they interacted within the target cells and cleaved the target site.

[0447] Example 3. Gene editing using chimeric nucleases This example describes the use of target cells. AAVS1 Chimeric nucleases at gene loci perform genome editing in cells.

[0448] In short, using AAV-Dualase- AAVS1 HEK293 cells were processed. Genomic DNA was collected, purified, and sequenced. Data were analyzed by trace insertion / deletion (TIDE) degradation. Results showed that AAV-Dualase- AAVS1 A 35 bp deletion. For those derived from AAV-Dualase- AAVS1 Amplicon sequencing of the processed samples was performed using Sanger sequencing and analyzed using the TIDE method. Deletions, insertions, and unmodified reads of defined lengths were characterized as bar graphs. Dashed line P < 0.01 (B) Compared to the reference sequence, AAV-Dualase- AAVS1 Chromatograms of sequencing samples treated with reagents only (or) Figure 4A and Figure 4B ).

[0449] Example 4. Gene editing using an integrated chimeric nuclease kit This example describes the use of target cells. AAVS1 Integrated chimeric nucleases at gene loci perform genome editing in cells.

[0450] In short, Dualase, along with guide RNA and donor nucleic acid, is co-expressed in HEK293 cells to co-deliver exogenous donor nucleic acid from a nucleotide sequence operably ligated to the promoter sequence, one or more self-cleaving guide RNA sequences, and one or more self-cleaving repair template sequences. Genomic DNA was collected, purified, and sequenced. Next-generation sequencing (NGS) showed highly accurate target site repair. NGS data were analyzed using two bioinformatics tools ( Figure 5A CRIS.Py and Geneious alignment platforms and ( Figure 5B Analysis was performed using CRISPRESSO2. "Insertion / Deletion" indicates a repair event leading to an insertion or deletion. "Exact Repair" indicates a sequencing read that perfectly aligned with the repair sequence. "Repair + SNP" indicates a sequence alignment with both repair and another nucleotide change. "Other" indicates other sequence modifications in the sample. Analysis of individual sequences using CRIS.Py and Geneious identified ~10% of reads with repair + SNP or other sequences possibly caused by sequencing errors, as only cell samples contained the same number of reads. Editing results were identified using total reads and aligned reads (p < 0.001). Alignment of sequencing reads of cis-antisense RNA repair template-guided RNA showed repair accuracy at I-TevI ​​and Cas9 sites, with no changes other than the expected repair results exceeding the sequencing detection limit indicated by the dashed line. Figure 5C The results showed that the gene editor inserted RNA sequences into target sites with high efficiency and accuracy. Figure 5A , Figure 5B and Figure 5C ).

[0451] Example 5. Gene editing using an integrated chimeric nuclease kit and DNA repair inhibitor This example describes the use of target cells. AAVS1 The integrated AAV chimeric nuclease at the locus, together with the NHEJ inhibitor, performs genome editing in cells.

[0452] In short, in the presence of escalating doses of blockade ( Figure 6A Non-homologous end connections (NHEJ), ( Figure 6B ) Homologous directional repair (HDR) or ( Figure 6C In the case of inhibitors of the Rad52-dependent pathway, lipid transfection was used to express Dualase and target... AAVS1 The gRNA-RT AAV (integrated) and the targeted expression of Dualase co-delivered with the indicated repair template. AAVS1HEK293 cells were treated with a control (co-delivered) of AAV. Editing efficiency of NHEJ, Rad51, and Rad52 inhibitors versus cis-sense and cis-antisense repair template samples was analyzed by RE digestion efficiency using a unique restriction enzyme (RE) at the insertion site of the repair template. Results showed that repair using the cis-sense and cis-antisense integrated construct was inhibited only by the Rad52 inhibitor. Figures 6A-6C The control inhibition of co-delivered repair templates indicated that the inhibitors functioned as expected. NHEJ(DNA) = non-homologous double-stranded DNA, HDR(DNA) = double-stranded DNA with homologous arms, and NHEJ(RNA) = non-homologous single-stranded RNA.

[0453] Example 6. Gene editing using an integrated chimeric nuclease kit and repair template This embodiment describes the targeted insertion of a repair template via directed polymerase chain reaction (PCR) into a target assembly combined with an RNA repair template. AAVS1 The Dualase ribonucleoprotein (RNP) complex that guides RNA formation.

[0454] In short, HEK293 cells were transfected with Dualase along with separately delivered guide RNA and repair template (co-delivery) or Dualase along with fused guide RNA and RNA repair template (“integrated”), lipid transfection reagent only (“reagent only”), or cell only (“mimic”). For each reaction, [details omitted]. AAVS1 Undigested PCR amplicon (“Sub”) at the site and AAVS1 Digested PCR amplicon at the insertion site (“digested”). Genomic DNA was extracted and purified. Editing efficiency of the samples was analyzed by RE digestion efficiency using a unique restriction enzyme (RE) at the insertion repair template site. Results showed that the PCR product was present in the correct orientation only if the repair template was correctly inserted (“correctly oriented RT”), and the PCR product was present in the incorrect orientation if the repair template was incorrectly inserted (“incorrectly oriented RT”). The accurate insertion of the Dualase RNP complex with the integrated repair template was demonstrated by the correctly oriented PCR product. Figure 7A Using a co-delivered repair template, the repair template is inserted in both correct and incorrect orientations. Figure 7B ).

[0455] Example 7. Gene editing using an integrated chimeric nuclease kit and cis-sense or cis-antisense repair template This example describes the analysis of targeted polymerase chain reaction (PCR) for... AAVS1 The Dualase integrated mRNA repair template is inserted into the directional insertion of cis-sense or cis-antisense repair template.

[0456] In short, HEK293 cells were transfected with Dualase (Tev[VKN]-SaCas9[WT] or Tev[VKN]-SaCas9[D10E]), SaCas9[WT], or lipid transfection reagents alone. Genomic DNA was extracted and purified. Editing efficiency of the samples was analyzed by RE digestion (BglII digestion) efficiency using a unique restriction enzyme (RE) at the insertion repair template site. For each reaction, the results are shown. AAVS1 Undigested PCR amplicon (“full-length”) at the locus AAVS1 PCR amplicon digested at the site (“BglII digest”), PCR product if the repair template is inserted only in the correct orientation (“correctly oriented RT”), and PCR product if the repair template is inserted only in the incorrect orientation (“incorrectly oriented RT”). Figure 8 It also demonstrates the use of targeted... AAVS1 Control cells were treated with Dualase mRNA at the site and co-delivered repair template (co-delivery). Accurate insertion of the cis-sense or cis-antisense repair template into the integrated Dualase mRNA was demonstrated by correctly oriented PCR products. No insertion was detected with saCas9[WT]-only cis-sense or antisense repair templates, as demonstrated by the absence of BglII digestion or oriented PCR products. Co-delivered Dualase mRNA and repair template were inserted in both correct and incorrect orientations.

[0457] Example 8. Gene editing using an integrated chimeric nuclease kit and NHEJ and Rad52 inhibitors. This example describes an analysis of the editing efficiency in cells using messenger RNA (mRNA) Dualase, NHEJ and Rad52 inhibitors, and trans-dsRNA to repair templates.

[0458] In short, in the presence of escalating doses of inhibitors that block non-homologous terminal linkages (NHEJ inhibitors) or inhibitors of the Rad52 pathway (“Rad52 inhibitors”), lipid transfection is used to express Dualase and target… AAVS1 The AAV (integrated) of gRNA-RT (also known as repair template guide RNA or "rep-gRNA") and the targeting of Dualase expression co-delivered with the indicated repair template. AAVS1 HEK293 cells were treated with a control (co-delivered) of AAV. Editing efficiency of the samples was analyzed by examining the RE digestion efficiency using a unique restriction enzyme (RE) at the insertion site of the inserted repair template. Repair using the trans-RT integrated construct was inhibited by both NHEJ and Rad52 inhibitors, indicating that both pathways can be used with self-complementary trans-repair templates. Figure 9The control inhibition of the co-delivered repair template indicated that the inhibitor functioned as expected. NHEJ (DNA) = non-homologous double-stranded DNA, HDR (DNA) = double-stranded DNA with homologous arms, &NHEJ (RNA) = non-homologous single-stranded RNA.

[0459] Example 9. Deletion gene editing using dual-guided chimeric nucleases This example describes the precise removal of large repetitive sequences using a dual-guided TevCas9 nuclease.

[0460] Figure 10A A schematic diagram illustrates the precise removal of large repetitive sequences using dual-guided TevCas9 nucleases, where the guide RNA targets the opposing strands of the double-stranded DNA to align the two I-TevI ​​domains head-to-head. A TevCas9 nuclease (1), containing an inactivated D10A+H557A mutation but an active I-TevI ​​domain (2), is targeted at the 5' end of the repetitive sequence (3) using synthetic guide RNA. A second TevCas9 nuclease (4), containing an inactivated D10A+H557A mutation but an active I-TevI ​​domain, is targeted at the 3' end of the repetitive sequence, but on the opposing strands, so that the I-TevI ​​domains of both nucleases point towards the repetitive sequence. After binding and cleavage (5) via the I-TevI ​​nuclease domains, two complementary 3' 2-nucleotide overhangs (6 and 7) are retained. These complementary overhangs (8) can be repaired by the cell via a non-homologous end joining pathway, and the large repetitive sequence (9) is removed, leaving a defined number of repetitive sequences (10) in the genomic DNA.

[0461] Figure 10B A schematic diagram illustrates the precise removal of large repetitive sequences using dual-guided TevCas9 nucleases, where the guide RNA targets the same strand of the double-stranded DNA to tandemly align the two I-TevI ​​domains. A TevCas9 nuclease (1), containing an inactivated D10A+H557A mutation but an active I-TevI ​​domain (2), is targeted at the 5' end of the repetitive sequence (3) using synthetic guide RNA. A second TevCas9 nuclease (4), containing an inactivated D10A+H557A mutation but an active I-TevI ​​domain, is targeted downstream of the 3' end of the repetitive sequence on the same strand, such that the I-TevI ​​domains of both nucleases are aligned in the same direction. After binding and cleavage (5) via the I-TevI ​​nuclease domains, two complementary 3' 2-nucleotide overhangs (6 and 7) are retained. These complementary overhangs (8) can be repaired by the cell via a non-homologous end joining pathway, and the large repetitive sequence (9) is removed, leaving a defined number of repetitive sequences (10) in the genomic DNA.

[0462] Example 10. Deletion of CAG triplet or... using dual-guided chimeric nucleases in cells and humanized mice Gene editing of the GGGGCC hexanucleotide repeat sequence This embodiment describes the use of an integrated AAV encoding TevCas9, along with dual guides targeting the 5' and 3' ends of the target for in vivo repetitive amplification, to precisely remove... C9ORF72 Between exon 1a and 1b DMPK Amplification of CAG triplet repeats in 3'-UTR or GGGGCC hexanucleotide repeats.

[0463] In short, transduction was achieved using an integrated AAV TevCas9 and SaCas9 dual guide (SEQ ID NO:117). DMPK MUT CAG repeating cells. Figure 11A The dual-guided Tev[KTQ]-Cas9[D10A+H557A] is shown in... DMPK A diagram showing the orientation of repeating target sites in CAG. Figure 11B A diagram showing the sequences in the integrated TevCas9 AAV is presented. It shows the hammerhead ribozyme (HHribo), single guide RNA (sgRNA), glycine tRNA (Gly tRNA), hepatitis D virus ribozyme (HDVribo), and polyadenylated sequences (polyA). Figure 11C This demonstrates the dual-guided TevCas9 pair. DMPK MUT In vitro activity of the CAG repeat DNA sequence demonstrated I-TevI ​​domain activity at this target. Figure 11D This image shows the Sanger sequencing reads of amplicon generated from cells transduced with an integrated AAV encoding TevCas9 along with dual guides, and the expected cleavage showing the removal of duplicated CAG sequences. Figure 11E This demonstrates the transduction of TevCas9 and SaCas9 using an integrated AAV dual guide. DMPK MUT The workflow for quantifying repeat collapse using repeat-induced PCR in CAG repeat cells is as follows: Transduced cells are harvested and genomic DNA is extracted for PCR amplification using primers for both external and internal repeat amplification. The resulting PCR products are analyzed on an Agilent Bioanalyzer, and peaks corresponding to repeats are compared with those for repeat-only cells. DMPK MUT Quantification of CAG repeats was performed. In TevCas9-treated cells, over 65% of repeat sequences collapsed, while in SaCas9-treated cells, ~30% of repeat sequences collapsed. Figure 11F This demonstrates the use of primers specific to each transcript for quantitative RT-PCR of transduced cells to quantify the transduced cells after each treatment. DMPKMUT and DMPK WT Transcript levels. Compared to untreated cells, TevCas9 treatment resulted in... DMPK WT Statistically significant increase in transcripts and DMPK MUT Decreased transcript levels (p < 0.01). Dual guide treatment with TevCas9 and SaCas9, each containing a single target guide RNA, did not result in statistically significant changes in transcript levels relative to cellular-only treatment. Figure 11G The results show that RT-PCR was performed on transduced cells using primers specific to the missplicing sequence to quantify the levels of misspliced ​​CLCN1 exon 6 transcripts after each treatment. TevCas9 single-guide or TevCas9 dual-guide resulted in a statistically significant reduction in misspliced ​​CLCN1 exon 6.

[0464] for C9ORF72 GGGGCC repeat sequence removal ( Figure 11H ), Figure 11I The dual-guided TevCas9 pair is shown. C9ORF72 MUT In vitro activity of the GGGGCC repeat DNA sequence was demonstrated, showing I-TevI ​​domain activity at this target site. This resulted in the maturation of patient motor neuron progenitor cells containing amplified large GGGGCC repeats into motor neurons. Figure 11J Then, using TevCas9 or saCas9 encoding and targeting... C9ORF72 Integrated AAV transduction of dual-guided receptors in the repetitive region. Maturation is confirmed by protein blotting through the expression of the motor neuron-specific marker (ISL1). Figure 11K The results of repeat-triggered PCR for the quantitative removal of large repeat sequences in TevCas9 or Cas9-transduced cells are shown, with approximately 50% removal in TevCas9-treated cells and less than 5% removal in Cas9-treated cells. Figure 11L The sequenced collapsed repeat sequences in TevCas9-treated cells are shown, with 4 out of 9 sequences exhibiting precise repeat collapse, as observed in the sequences marked with an asterisk (*). Figure 11M This demonstrates the C9ORF72 guide RNA pair targeting Tev[VKN]-saCas9[D10A+H557A]. C9ORF72 The target site is specific because no off-target effects in the genome were detected at levels exceeding the detection limits of reported depth sequencing assays. Figure 11N This demonstrates the need for TevSaCas9, because it is composed of the same C9ORF72SaCas9, which targets RNA, exhibits detectable off-target effects on chromosome 11. Figure 11O The display shows that when using the expression Tev[VKN]-saCas9[D10A+H557A] and dual C9ORF72 In cases of AAV transduction guided by the drug, C9ORF72 protein expression can be restored using mature motor neurons derived from the patient. Figure 11P It was demonstrated that transducing patient-derived motor neurons with AAVs (SEQ ID NO:115) expressing Tev[VKN]-saCas9[D10A+H557A] and dual C9ORF72 guides can reduce the incidence of motor neuron disease caused by cerebral infarction. C9ORF72 Accumulation of multi-GR dipeptides expressed by repetitive sequences. The toxic accumulation of these dipeptides contributes to the degeneration of motor neurons in the disease. Figure 11Q Humanization of two doses of Tev[VKN]-saCas9[D10A+H557A] and dual C9ORF72-guided formulations via intracranial injection is demonstrated. C9ORF72 The expression of TevCas9 in the cerebellum of repeat-amplified mice via RT-qPCR demonstrated brain tissue transduction. Figure 11R This demonstrates intracranial injection of two doses of Tev[VKN]-saCas9[D10A+H557A] and dual... C9ORF72 Normal-sized humanized mice of the guide C9orf72 Dose-dependent recovery of repetitive sequences. Figure 11S This demonstrates the humanization of mice by intracranial injection of two doses of Tev[VKN]-saCas9[D10A+H557A] and dual C9ORF72 guides. C9ORF72 Recovery of C9ORF72 transcripts in cerebellar tissue of repeat amplified mice via RT-qPCR. Figure 11T The restoration of C9ORF72 protein expression in cerebellar tissue by Western blot was demonstrated in mice with humanized C9ORF72 repeat sequence amplification via intracranial injection of two doses of Tev[VKN]-saCas9[D10A+H557A] and dual C9ORF72 guides.

[0465] Example 11. Gene editing using self-inactivating chimeric nucleases This embodiment describes the use of self-deactivating TevSaCas9 in... B2M Gene editing is performed within the gene.

[0466] In short, the target β-2-microglobulin was harvested at 24, 48, and 72 hours post-transfection. B2M HEK293 cells were transfected with plasmid DNA of the TevSaCas9 gene. Figure 12AThe structure of the self-inactivation vector was depicted. The construct contains nucleotide sequences encoding a promoter, a human codon-optimized TevSaCas9, a polyadenylation signal (“polyA”), and a guide RNA sequence (“gRNA”). The self-inactivation target site can be located in the region between the promoter and the TevSaCas9 site (denoted as “promoter”) or at the end of the TevSaCas9 coding sequence and the beginning of the polyA sequence (denoted as “polyA”). Vectors targeting β-2-microglobulin were harvested at 24, 48, and 72 hours post-transfection. B2M The results of the T7E1 editing assay in HEK293 cells transfected with plasmid DNA of the TevSaCas9 gene are shown in the figure. Figure 12B Lanes marked "None" do not contain self-inactivating sequences in the vector; lanes marked "Promoter" contain the B2M1TevSaCas9 target site between the promoter sequence and the TevSaCas9 sequence; and lanes marked "Multi-A" contain the B2M1TevSaCas9 target site between the TevSaCas9 terminus and the multi-A signal sequence. The editing level, as determined by the amount of digestion product relative to the substrate, is comparable across constructs over time. Figure 12C The results of Western blotting of hemagglutinin (α-HA) encoded at the 3' end of TevSaCas9 (“Dualase”) from cells treated in the same manner as in B are shown. Lanes labeled “None” do not contain the self-inactivating sequence in the vector, lanes labeled “Promoter” contain the B2M1 TevSaCas9 target site between the promoter sequence and the TevSaCas9 sequence, and lanes labeled “Multi-A” contain the B2M1 TevSaCas9 target site between the TevSaCas9 terminus and the multi-A signal sequence. The presence of bands in the α-HA blot indicates that the TevSaCas9 protein is expressed. TevSaCas9 expression levels were stable over time in vectors without the inactivating sequence, decreased within 72 hours in vectors containing the promoter inactivating sequence, and decreased to undetectable levels within 72 hours in vectors containing the multi-A inactivating sequence. Figure 12D The generation and titer of the AAV2 capsid encapsulating the self-inactivating TevCas9 targeting B2M are shown, along with an image of an SDS-polyacrylamide gel demonstrating successful capsid assembly with VP1, VP2, and VP3 capsid structural proteins. As observed in the Western blot of the HA tag on the TevCas9 protein, the level of TevCas9 protein was significantly reduced after 72 hours compared to the β-actin housekeeping gene, indicating inactivation. Figure 12E The image shows the source used to generate Figure 12DThe plasmid DNA form of the self-inactivating AAV vector (“self-inactivating”), the vector without the self-inactivating sequence (“non-self-inactivating”), and the plasmid expressing GFP (“pAAV-GFP”) were used as controls for Western blot results of saCas9 and GAPDH in lipid-transfected cells. Transfected cells were harvested at 48 h (48 h), 72 h (72 h), 7 d (7 d), and 14 d (14 d) and lysed to recover total protein. The presence of expressed Cas9 was detected using a Cas9 antibody, and the housekeeping gene GAPDH was used to assess the consistency of the loaded samples. The SaCas9-sized band was detected only in the non-self-inactivating control samples at 48 h and 72 h, but not in the self-inactivating samples, indicating that the presence of the self-inactivating target site in the expression vector limited TevSaCas9 expression.

[0467] Example 12. Gene editing using TevCas9 and rep-gRNA to recruit endogenous polymerase θ (PolQ or Polθ) This example describes the use of the TevCas9 gene editor and rep-gRNA to recruit endogenous polymerase in cells.

[0468] In short, using TevCas9 or Cas9 mRNA and targeting AAVS1 rep-gRNA lipid transfection of HEK293 cells ( Figures 13A-13C Cells were lysed, and extracts were precipitated using magnetic beads conjugated with anti-HA antibodies that specifically recognize the C-terminal HA tag on TevCas9 or Cas9. After washing the magnetic beads, the precipitated proteins were resolved on a polyacrylamide gel and Western blotted using various antibodies. Figure 14 Western blots using Rad52-specific antibodies and polymerase θ-specific (PolQ or Polθ) antibodies show that polymerase θ co-precipitates only with TevCas9.

[0469] Example 13. Gene editing using inactivated or nicking enzyme-based chimeric nucleases This example describes the use of a nuclease inactivation or nicking enzyme gene editor, with TevCas9, for... AAVS1 Gene editing is performed at specific sites.

[0470] In short, using purified TevCas9 and Cas9 proteins and targeting... AAVS1The two-in-one guide RNA lipid transfection site was used to transfect HEK293 cells. The editing efficiency of the samples was analyzed by looking for RE digestion efficiency using the unique restriction enzyme (RE) at the insertion site of the inserted repair template. Purified TevCas9 (“Tev[WT]-dCas9”) containing the inactivated D10A+H557A mutation in the Cas9 domain and TevCas9 (“Tev[R27A]-dCas9”) containing the inactivated D10A+H557A mutation in the Cas9 domain and the inactivated R27A mutation in the I-TevI ​​domain were used. Figure 15A Digestion was observed. No digestion was observed in cells treated with Cas9 alone, reagent alone, or untreated cells (“cells only”). Digestion was observed in controls of cells transfected with Tev-Cas9 3-in-1 mRNA lipids. Digestion was also observed in the mRNA form of TevCas9 containing the mutant R27A+V117F+K135R+N140S and the Cas9 domain with a D10A or H557A nickase mutation. Figure 15B ).

[0471] Example 14. In-frame insertion of GFP into the cellular genome This example describes the in-frame insertion of GFP into the genome of a mammalian cell using TevCas9.

[0472] In short, a rep-gRNA containing the T2A sequence and the eGFP coding sequence was designed for insertion into... CFTR G542 site of the gene ( Figure 16A To cover the G542X, R553X, and G551D mutations, cells were transfected with TevCas9 and rep-RNA encoding eGFP, or saCas9 and rep-RNA encoding eGFP lipids. Results showed that GFP was generated endogenously under TevCas9-only conditions. CFTR Promoter expression, but not under SaCas9, rep-gRNA, or reagent-only conditions ( Figure 16B eGFP separates from truncated endogenous proteins by self-cleaving the T2A peptide. Figure 16C It shows that when amplified by PCR CFTR When the target region was reached, the inserted sequence was detected as a large band. PCR results for the target region showed an edited to unedited ratio of ~50%. Another approach uses a combination of upstream (5') guide RNA of the rep-gRNA to remove a large ~28 kilobase sequence and replace it with a 921-base-pair (bp) GFP-coding sequence. The rep-gRNA encoding GFP spans... CFTR The region between F508 and G542 was cut to extract a ~28 kb region between exon 11 and exon 12 and replaced with GFP. Figure 16DThe target was shown CFTR The F508 region of the guide RNA (sgRNA) and target RNA CFTR A schematic diagram of the rep-gRNA in the G542 region. Lung epithelial cells were transfected with TevCas9 or Cas9 mRNA lipids combined with an equimolar mixture of guide RNA and rep-gRNA. Results showed that GFP was generated endogenously only under TevCas9 conditions and not under Cas9 conditions. CFTR Promoter expression ( Figure 16E ). Figure 16F A gel image of the PCR amplicon at the target site is shown, demonstrating the removal and replacement of the ~28 kb region between the F508 and G542 sites, as evidenced by the ~1500 base pair product in the TevCas9 lane.

[0473] Example 15. Modified gRNA enables targeted DSB repair in human cells. This example describes the identification of targeted DSB repair in human cells via modified gRNA and the TevSaCas9 nuclease.

[0474] To test directed RNA templated repair using the TevSaCas9 nuclease, modified gRNA is generated by fusing an RNA repair sequence to the 3' end of a standard SaCas9 single guide RNA (sgRNA) (which contains a crRNA portion and a tracrRNA portion that are complementary to the target site sequence) to produce rep-gRNA (rep-templated repair gRNA). Figure 17A and Figure 17B In the proposed repair protocol, the 3' end of the rep-gRNA is a complementary sequence to the 2-nt 3' overhang generated by the I-TevI ​​nuclease domain (hereinafter referred to as Tev), which will form the initiation site for repair (possibly through reverse transcription of the RNA repair template). Figure 17A The targeted editing site can be anywhere between the Tev and SaCas9 cleavage sites, extending into the scaffold region. First, rep-gRNA is designed, with the crRNA portion targeting... AAVS1 Loci in the safe harbor locus ( Figure 17B ). AAVS1 rep-gRNA contains a 37-nucleotide-long string of molecules. AAVS1 The sequences are dissimilar and contain a repair sequence at the diagnostic BglII site. Crucially, the 3' end of the rep-gRNA has two nucleotides (5'-CC-3') complementary to the 2-nt 3' overhang left by Tev cleavage at the appropriately positioned 5'-CA↑GG↓G-3' site (the 2-nt 3' overhang product is 5'-GG-3'). Figure 17BTo confirm that TevSaCas9 and SaCas9 are in AAVS1 The activity at the site was assessed using sgRNA targeting that site, and robust activity was observed in vitro and in HEK293 cells, with the main editing product being a 35-bp deletion corresponding to the distance between the Tev and SaCas9 cleavage sites.

[0475] To determine the number of putative genomic targets for CNNNG motif cleavage via Tev in TevSaCas9 constructs, target sites were predicted, with CNNNG motifs spaced 5–31 bp from gRNA binding sites, consistent with the spacing preferences of native I-TevI ​​nucleases and other Tev-based chimeric nucleases. It is estimated that ~75% of SaCas9 sites in the human genome (defined by the presence of NNGRRTPAM sites and 21-nt gRNA binding sites) have CNNNG motifs with appropriate spacing to support Tev cleavage. Figure 17C To deliver TevSaCas9 and rep-gRNA into cells, an integrated construct was developed, in which a single ~4.5 kb RNA polymerase II transcript contains the TevSaCas9 coding sequence and rep-gRNA, and it will be adapted to a single AAV viral vector ( Figure 17D This setup avoids the need for a separate RNA polymerase III promoter, which is typically used to express gRNA. Because this setup removes the multi(A) tail from the TevSaCas9 coding sequence, the MALAT1 sequence is included at the 3' end of TevSaCas9 for stability. To stimulate post-transcriptional processing of the integrated mRNA, an alanine tRNA sequence (ala tRNA) was inserted between the TevSaCas9-eGFP coding region and the rep-gRNA, and a trans-acting hammerhead (HH) ribozyme sequence was inserted into the hairpin of the alanine tRNA to act on the HH cleavage site located between the 3' end of the rep-gRNA and the multi(A) signal. To confirm correct processing, the in vitro transcribed TevSaCas9-eGFP / rep-gRNA integrated mRNA was incubated with HEK293 cell extract, and processing into two predicted RNA products (TevSaCas9-eGFP and rep-gRNA) was observed on a 1% agarose gel. Figure 17E HEK293 cells were transfected with TevSaCas9-eGFP / rep-gRNA integrated mRNA, and robust eGFP activity was observed, indicating that MALAT sequencing can provide RNA stability (Fig. 17F).

[0476] Example 16. Rep editing in human cell lines This example describes the identification of targeted DSB repair in human cells via modified gRNA and the TevSaCas9 nuclease.

[0477] The target was delivered using four different methods (AAV2 transduction, lipid nanoparticles containing mRNA, plasmid DNA (pDNA) transfection, or purified RNP). AAVS1 TevSaCas9 / rep-gRNA or SaCas9 / rep-gRNA integrated constructs at safe harbor sites were introduced into HEK293 cells. Figure 18A ). Target sites amplified by PCR from genomic DNA (cells not enriched prior to DNA isolation) were observed to be effective with TevSaCas9 / rep-gRNA in all four delivery modalities, including BglII digestion (56% + / - 12%) and deep sequencing (52% + / - 14%). AAVS1 Robust editing at the site, but not observed with SaCas9 / rep-gRNA ( Figure 18A and Figure 18B AAV2 delivery resulted in the highest edit rate (67% + / - 10%), while RNP transfection resulted in the lowest edit rate (43% + / - 11% across all analysis methods). Figure 18C This may reflect lower RNP transfection and AAV2 transduction efficiency. Using PCR primers that correctly distinguished incorrectly oriented editing events, TevSaCas9 / rep-gRNA editing showed that it produced correctly oriented repair products. Analysis of deep sequencing reads from six independent TevSaCas9-rep-gRNA editing events confirmed the fidelity of oriented insertion and editing, with 42% + / - 6% of reads corresponding to the expected repair products without other templated nucleotide changes. Figure 18D and Figure 18E Less than 0.1% of reads corresponded to incorrect repair products, while 1.6% + / - 1.1% of reads had insertions / deletions. In cases where repair extended beyond the RNA repair template region to include sequences in stem-loop structures within the scaffold region of the rep-gRNA, no reads were identified, indicating that editing and repair were specified only by the RNA repair template portion of the rep-gRNA. Furthermore, the obvious nucleotide substitutions in the sequencing reads of edited cells were indistinguishable from those in simulated-treatment cells. Figure 18D ).

[0478] Design rep-gRNA to... AAVS1 Introducing a 13-bp deletion at the site will produce the HindIII site ( Figure 18F44% of the deep sequencing reads were consistent with the exact 13-bp deletion, and 4.1% of the reads were consistent with other deletion lengths, compared to 1.4% of the sequencing reads from simulated treatment cells that showed length differences. Figure 18G and Figure 18H ).

[0479] Example 17. Determining the length of overhangs used for repair with Rep-gRNA This example describes the impact of determining the length of the overhang at the 3' end of the rep-gRNA and its sequence complementarity with the target sequence on editing efficiency.

[0480] In short, TevCas9 messenger RNA (mRNA) is complexed with a repair template guide RNA having or not having a complementary single-stranded DNA (“protected rep-gRNA”) with overhangs of the indicated length targeting the I-TevI ​​cleavage site and the Cas9 cleavage site. Figure 19A The cells were transfected with lipids. The treated cells were harvested and total RNA was extracted from the cells. The repaired RNA in each sample was compared using quantitative reverse transcriptase PCR (RT-qPCR). AAVS1 The levels of transcripts and non-repair transcripts showed the maximum repair length of 14 nucleotides at the 3'-overhead length. Figure 19C ).

[0481] To further investigate the impact of nucleotide identity between the 3' end of rep-gRNA and the target on repair efficiency, a series of [tests / tests] were designed. AAVS1 Targeting rep-gRNA, the length of which is complementary to the sequence upstream of the Tev cleavage site increases in increments of 2 nt to a maximum of 18 nt. Figure 19A (This is referred to as 3' overlap). The complementarity variation from the SaCas9 cleavage site to the PAM site reaches a maximum of 6 nt ( Figure 19A (referred to as 5' overlap). To enable analysis of multiple designs, a lipid transfection strategy was used, employing mRNA encoding TevSaCas9 and in vitro transcribed rep-gRNA, and quantitative reverse transcription PCR was used to measure the concentration of TevSaCas9 in each sample. AAVS1 The ratio of correctly repaired to unrepaired sequences in the transcript. Results showed that 2-nt complementarity with the Tev 2-nt overhang at the 3' end of the rep-gRNA was essential for repair, and complementarity exceeding 14 nt reduced repair efficiency. Figure 19B (Dark green bar). A similar increase in efficiency was observed in deep sequencing, but the editing fidelity did not change with increasing 3' overlap. Figure 19C In contrast, repair was not increased with 5' overlapping rep-gRNAs that had increased complementarity PAMs near the Cas9 cleavage site. Figure 19BIt was assumed that the 3' end of the rep-gRNA would be readily degraded by 3' exonucleases, or that a self-inhibiting structure within the rep-gRNA might affect repair efficiency, as suggested by initiating pegRNA editing. To test this, protective oligonucleotides were included that were complementary in length to the 3' repair template portion of the rep-gRNA. Repeated lipid transfection with the protected rep-gRNA resulted in a modest increase in repair efficiency for rep-gRNAs with 12, 14, or 16 nt 3' overlap. Figure 19B (Light green bar). Detection of mismatch between the 3' end and the Tev overhang of rep-gRNA and... AAVS1 Tolerance to target site mismatch ( Figure 19A , Figure 19D and Figure 19E The results showed that, apart from the exact 2-nt match between the 3' rep-gRNA (5'-CC-3') and the Tev overhang (5'-GG-3'), Figure 19D Besides lane 2, the only other rep-gRNA that promotes repair has a 5'-CU-3' sequence that generates an AU wobble at the 3' end. Figure 19D Lane 5). Interestingly, rep-gRNAs with a 5'-UC-3' sequence that produces a GU wobble at a base inside the 3' end do not support repair ( Figure 19D (lane 17). rep-gRNA and AAVS1 Single nucleotide mismatch at the site but using rep-gRNA with only the 5'-CC-3' sequence ( Figure 19E Lanes 1-8 also support repair, while rep-gRNA with a 5'-UC-3' sequence does not support repair. Figure 19E (Lane 13-19). AAVS1 Dinucleotide or trinucleotide mismatches at a site do not support repair using any rep-gRNA. Figure 19E (Lane 9-19). In summary, this data suggests that the stability between the 3' end of rep-gRNA and the Tev cleavage overhang guides repair, and that hybridization of protective oligonucleotides with rep-gRNA can moderately increase repair.

[0482] ...

Claims

1. A nucleic acid, comprising (i) A polynucleotide encoding a chimeric nuclease containing an I-TevI ​​domain and an RNA-directed nuclease domain; (ii) Polynucleotides encoding first guide RNA (gRNA); and (iii) Polynucleotides encoding tRNA; The polynucleotides in (ii)-(iii) are in sequence.

2. A nucleic acid, comprising (i) A polynucleotide encoding a chimeric nuclease containing a GIY-YIG nuclease domain and an RNA-directed nuclease domain; (ii) Polynucleotides encoding first guide RNA (gRNA); and (iii) Polynucleotides encoding tRNA; The polynucleotides in (ii)-(iii) are in sequence.

3. The nucleic acid according to claim 1 or claim 2, further comprising... (iv) RNA stable polynucleotides located downstream of (i).

4. A nucleic acid, comprising (i) Polynucleotides encoding the first guide RNA (gRNA); and (ii) Polynucleotides encoding tRNA; The polynucleotides in (i)-(ii) are in sequence.

5. The nucleic acid according to any one of claims 1 to 4, further comprising: Two or more donor polynucleotides tandemly positioned; and Ribozyme polynucleotide; The donor polynucleotides are arranged in sequence, and the upstream donor polynucleotide contains a ribozyme cleavage site sequence at its 3' end.

6. The nucleic acid according to any one of claims 1 to 5, wherein the nucleic acid is DNA or RNA.

7. The nucleic acid according to claim 6, wherein the DNA is circular plasmid DNA, linear double-stranded DNA, single-stranded DNA, or chimeric RNA and DNA.

8. The nucleic acid according to claim 6, wherein the RNA is mRNA.

9. The nucleic acid according to claim 8, wherein the mRNA comprises a nucleic acid mimic selected from the group consisting of peptide nucleic acid (PNA), morpholino nucleic acid, cyclohexenyl nucleic acid (CeNA), and locked nucleic acid (LNA).

10. The nucleic acid according to claim 8 or claim 9, wherein the mRNA comprises a modified sugar moiety, optionally wherein the modified sugar moiety is selected from the group consisting of: N1-methylpseuuridine, 9-methyladenine, 2'-O-(2-methoxyethyl), 2'-dimethylaminooxyethoxy, 2'-dimethylaminoethoxyethoxy, 2'-O-methyl, and 2'-fluoro.

11. The nucleic acid according to any one of claims 8 to 10, wherein the mRNA comprises a modified nucleotide, optionally wherein the modified nucleotide is selected from the group consisting of: 5-methylcytosine; 5-hydroxymethylcytosine; xanthine; hypoxanthine; 2-aminoadenine; a 6-methyl derivative of adenine; a 6-methyl derivative of guanine; a 2-propyl derivative of adenine; a 2-propyl derivative of guanine; 2-thiouracil; 2-thiothymine; 2-thiocytosine; 5-halogenuridine; 5-halogenuridine; 5-propynyluracil; 5-propynylcytosine; 6-azouracil; 6-azocytosine; 6-azothymine; pseudouracil; 4-thiouracil; 8-halogen; 8-amino; 8-thiol; 8 -Thioalkyl; 8-hydroxy; 5-halogen; 5-bromine; 5-trifluoromethyl; 5-substituted uracil; 5-substituted cytosine; 7-methylguanine; 7-methyladenine; 2-F-adenine; 2-aminoadenine; 8-azaguanine; 8-azaadenine; 7-deadenine; 7-deadenine; 3-deadenine; 3-deadenine; tricyclic pyrimidine; phenyloxazinecytidine; phenothiazinecytidine; substituted phenyloxazinecytidine; carbazolecytidine; pyridineindolecytidine; 7-deadenine; 7-deadenine; 2-aminopyridine; 2-pyridone; 5-substituted pyrimidine; 6-azapyrimidine; N-2, N-6 or O-6 substituted purine; 2-aminopropyladenine; 5-propynyluracil; and 5-propynylcytosine.

12. The nucleic acid according to any one of claims 8 to 11, wherein the mRNA comprises a non-natural or non-natural internucleotide bond selected from the group consisting of: thiophosphate, phosphoramide, non-phosphodiester, heteroatom, chiral thiophosphate, dithiophosphate, triphosphate, aminoalkyl phosphate triester, 3'-alkylphosphonate, 5'-alkylphosphonate, chiral phosphonate, hypophosphonate, 3'-aminophosphamide, aminoalkylphosphamide, phosphodiamid, thiocarbonylphosphamide, thiocarbonylalkylphosphamide, thiocarbonylalkyl phosphate triester, selenophosphate, and boron phosphate.

13. The nucleic acid according to any one of claims 1 to 12, wherein the RNA-directed nuclease is selected from the group consisting of Staphylococcus aureus (…). Staphylococcus aureus Cas9 ("saCas9"), Streptococcus pyogenes ( Streptococcus pyogenes Cas9, Acidaminococcus; Cas12, Deltaproteobacteria; CasX and Eubacterium rectum ( Eubacterium rectale ) Cas12a.

14. The nucleic acid according to claim 13, wherein the Cas is inactivated Cas (dCas).

15. The nucleic acid according to claim 13, wherein the Cas is a nickase (nCas) or dCas.

16. The nucleic acid according to any one of claims 1 to 15, wherein the I-TevI ​​is a nicking enzyme, and the I-TevI ​​is an I-TevI ​​nicking enzyme domain.

17. The nucleic acid according to claim 16, wherein the I-TevI ​​nickase domain comprises mutations at amino acid residues R27A, V117F, K135R and N140S corresponding to the amino acid residues in SEQ ID NO:

155.

18. The nucleic acid according to any one of claims 1 to 15, wherein the I-TevI ​​is inactivated.

19. The nucleic acid according to claim 18, wherein the I-TevI ​​inactivation mutation is an R27A mutation at an amino acid residue corresponding to the amino acid residue in SEQ ID NO:

155.

20. The nucleic acid according to any one of claims 5 to 19, wherein the ribozyme polynucleotide is located within the tRNA polynucleotide.

21. The nucleic acid according to any one of claims 1 to 20, further comprising a polynucleotide encoding a second gRNA, optionally wherein the second gRNA polynucleotide is located at the 3' of the tRNA polynucleotide.

22. The nucleic acid according to any one of claims 3 or 5 to 21, wherein the polynucleotide is in the order of chimeric nuclease, RNA-stable polynucleotide, guide RNA, tRNA, and guide RNA.

23. The nucleic acid according to any one of claims 5 to 21, wherein the polynucleotide is in the following order: chimeric nuclease, RNA stable polynucleotide, guide RNA, tRNA / ribozyme, guide RNA, donor polynucleotide 1, and donor polynucleotide 2.

24. The nucleic acid according to any one of claims 3 or 5 to 23, wherein the RNA-stabilized polynucleotide comprises the 3' sequence of metastasis-associated lung adenocarcinoma transcript 1 (MALAT) or the 3' end of multiple endocrine neoplasm β transcript (MEN β).

25. The nucleic acid according to any one of claims 3 or 5 to 23, wherein the stable RNA sequence comprises a triple-helical RNA structure or the 3' end of an RNA transcript lacking typical polyadenylation signals.

26. The nucleic acid according to any one of claims 5 to 25, wherein the donor polynucleotide is single-stranded or double-stranded.

27. The nucleic acid according to any one of claims 5 to 25, wherein the donor polynucleotide is DNA or RNA.

28. The nucleic acid according to any one of claims 5 to 25, wherein one strand of the double-stranded donor polynucleotide is DNA and the other strand is RNA.

29. The nucleic acid according to any one of claims 1 to 25, wherein the donor polynucleotide comprises a cis-acting single-stranded RNA polynucleotide annealed to a complementary single-stranded DNA polynucleotide.

30. The nucleic acid according to any one of claims 1 to 28, wherein the single-stranded donor polynucleotide or the double-stranded donor polynucleotide comprises a 3' end with 2 to 18 nucleotide overhangs.

31. The nucleic acid according to any one of claims 5 to 29, wherein the single-stranded donor polynucleotide or the double-stranded donor polynucleotide comprises a 14-nucleotide overhang at the 3' end.

32. The nucleic acid according to any one of claims 26 to 31, wherein the overhang at the 3' end is a single-stranded RNA polynucleotide.

33. The nucleic acid according to any one of claims 1 to 32, further comprising a second guide RNA capable of targeting the region of the donor polynucleotide target site 5'.

34. The nucleic acid of claim 33, wherein the first guide RNA enables the first chimeric nuclease to target and cleave at a first I-TevI ​​or Cas9 target site in the cell genome, and the second guide RNA enables the second chimeric nuclease to target and cleave at a second I-TevI ​​target site in the cell genome, wherein the cleavage produces a nucleotide overhang at the second I-TevI ​​target site.

35. The nucleic acid of claim 34, wherein the 3' end of the donor polynucleotide is complementary to the overhang generated by cleaving the second I-TevI ​​at the second I-TevI ​​target site.

36. The nucleic acid according to any one of claims 1 to 35, wherein the guide RNA and the donor polynucleotide target CFTR Mutations in genes.

37. The nucleic acid according to any one of claims 1 to 35, wherein the guide RNA and donor polynucleotide target and replace CFTR c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), c.1652G>A (p.Gly551Asp), c.1657C>T (p.Arg553Ter) or c.3846G>A (p.Trp1282Ter) mutation.

38. The nucleic acid according to any one of claims 1 to 35, wherein the guide RNA and the donor polynucleotide target SERPINA1 Mutations in genes.

39. The nucleic acid according to any one of claims 1 to 35, wherein the guide RNA and donor polynucleotide target and replace SERPINA1 c.1096G>A (p.Glu342Lys) mutation.

40. The nucleic acid according to any one of claims 1 to 39, further comprising a promoter.

41. The nucleic acid according to claim 40, wherein the promoter is selected from the group consisting of: CMV promoter, SV40 promoter, minimal cytomegalovirus (CMV) promoter and human extension factor-1α (EF1a) promoter.

42. The nucleic acid according to claim 41, wherein the promoter is selected from the group consisting of: muscle-specific synthesis promoters, SPc5-12 neuron-specific promoters hSYN1, aldh1L1, cTNT, α-MHC, SPc5-12, MUC2, Ksp-cadherin, albumin, HAS, insulin, rhodopsin, rNSE, and cone protein promoters.

43. The nucleic acid according to any one of claims 1 to 42, wherein the tRNA comprises glycine tRNA, arginine tRNA, asparagine tRNA, aspartic acid tRNA, cysteine ​​tRNA, glutamine tRNA, glutamate tRNA, histidine tRNA, isoleucine tRNA, leucine tRNA, lysine tRNA, methionine tRNA, phenylalanine tRNA, proline tRNA, serine tRNA, threonine tRNA, tryptophan tRNA, tyrosine tRNA, or valine tRNA.

44. The nucleic acid according to any one of claims 1 to 42, wherein the ribozyme comprises a hammerhead ribozyme or a hepatitis D virus (HDV) ribozyme.

45. The nucleic acid according to any one of claims 1 to 42, wherein the donor polynucleotide comprises Trans-acting double-stranded RNA polynucleotides with 3'- or 5'-protruding nucleotides; A cis-acting single-stranded RNA polynucleotide with sequence similarity to the target strand of the nuclease; Cis-acting single-stranded RNA polynucleotides that have sequence similarity to the non-target strand of the nuclease; For one or more binding sites of genomic modifying factors, optionally for binding sites of site-specific recombinases such as serine recombinases or LoxP target sites. Templates used for repairing protein-coding sequences; One or more exons having splice acceptor and donor sequences; One or more optional sequences selected from the group consisting of: NeoR gene, BsdR gene, HygR gene, PuroR gene, and BleoR gene; One or more drug-inducible regulatory sequences for controlled gene expression; and / or The two nucleotide overhangs at the 3' end.

46. ​​The nucleic acid according to any one of claims 1 to 45 further comprises a polyadenylation signal.

47. The nucleic acid according to claim 46, wherein the polyadenylation signal comprises simian virus 40 (SV40), α-globin, β-globin, human growth hormone (hGH), bovine growth hormone (BGH), herpes simplex virus type 1 thymidine kinase (HSVTK), or synthetic polyadenylation (Synt poly A) polyadenylation signal.

48. The nucleic acid according to any one of claims 1 to 47, further comprising a self-inactivating sequence.

49. The nucleic acid according to any one of claims 1 to 48, wherein the length of said nucleic acid is about 5 kb.

50. The nucleic acid according to any one of claims 1 to 48, wherein the length of the nucleic acid is less than 5 kb.

51. The nucleic acid according to any one of claims 1 to 50, wherein the nucleic acid is packaged in a virus.

52. The nucleic acid according to claim 51, wherein the virus is a lentivirus, adeno-associated virus (AAV), adenovirus, retrovirus, or modified herpes simplex virus (HSV).

53. The nucleic acid according to claim 52, wherein the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-DJ, AAV2.5T, or AAVmyo.

54. A vector comprising the nucleic acid according to any one of claims 1 to 53.

55. A viral vector comprising the nucleic acid according to any one of claims 1 to 53.

56. An AAV virus comprising the nucleic acid according to any one of claims 1 to 53.

57. A cell comprising a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56.

58. A composition comprising a chimeric nuclease polypeptide containing an I-TevI ​​domain and an RNA-directed nuclease domain and a nucleic acid according to any one of claims 4 to 53.

59. A composition comprising a chimeric nuclease nucleic acid encoding a chimeric nuclease comprising an I-TevI ​​domain and an RNA-directed nuclease domain, and a nucleic acid according to any one of claims 4 to 53.

60. The composition according to claim 59, wherein the chimeric nuclease nucleic acid is mRNA.

61. An LNP composition comprising a nucleic acid according to any one of claims 1 to 50 or a composition according to any one of claims 58 to 60.

62. A pharmaceutical composition comprising a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55 or an AAV according to claim 56, a composition according to claim 58 or an LNP composition according to claim 61; and an excipient.

63. A method for delivering messenger RNA encoding a chimeric nuclease comprising an I-TevI ​​domain and an RNA-directed nuclease domain to a cell, the method comprising: The cell is brought into contact with a polynucleotide encoding one or more guide RNAs and a polynucleotide donor.

64. A method for genetically modifying the genome of a cell, the method comprising contacting the cell with a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.

65. The method of claim 64, wherein the modification includes insertion, deletion, substitution or mutation of the genome.

66. The method according to any one of claims 63 to 65, wherein inserting the donor polynucleotide into the genome of the cell results in the removal of the sequence between the I-TevI ​​target site and the Cas9 target site.

67. The method according to claim 64 or 65, wherein the cell is a mammalian cell.

68. The method according to any one of claims 64 to 65, wherein the cell is a human cell.

69. A method for inserting or replacing a sequence at a chimeric nuclease target site in the genome of a cell, the method comprising: Contact the cell with a nucleic acid containing one or more sequences encoding the following nucleic acid sequences: Chimeric nucleases containing Cas9 and I-TevI ​​domains; and Nucleic acids containing both guide polynucleotides and donor polynucleotides; The guide polynucleotide and the chimeric nuclease form a complex, and the complex binds to and cleaves the genomic DNA at the Cas9 target site and the I-TevI ​​target site; Following I-TevI ​​cleavage, the 3' end of the donor polynucleotide contains at least two bases of complementarity at the 5' end of the I-TevI ​​target site; and The donor polynucleotide is incorporated into the chimeric nuclease target site at the 5' position of the Cas9 target site.

70. The method of claim 69, wherein the 3' end of the guiding polynucleotide and the 5' end of the donor polynucleotide are linked.

71. The method of claim 69, wherein the cell polymerase is targeted at the chimeric nuclease target site.

72. The method according to claim 71, wherein the cell polymerase is polymerase θ.

73. A replacement for the genome in cells CFTR A method for at least a portion of a gene, the method comprising contacting the cell with a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.

74. The method of claim 73, wherein the guide RNA and donor polynucleotide target the CFTR Mutations in genes.

75. The method according to any one of claims 73 to 74, wherein the guide RNA and donor polynucleotide target and replace CFTR c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), c.1652G>A (p.Gly551Asp), c.1657C>T (p.Arg553Ter) or c.3846G>A (p.Trp1282Ter) mutation.

76. A method for treating cystic fibrosis in a patient with a corresponding need, the method comprising administering to the patient a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.

77. The method of claim 76, wherein the guide RNA and donor polynucleotide targeting CFTR Mutations in genes.

78. The method of claim 76, wherein the guide RNA and donor polynucleotide target and replace CFTR c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), c.1652G>A (p.Gly551Asp), c.1657C>T (p.Arg553Ter) or c.3846G>A (p.Trp1282Ter) mutation.

79. A replacement for the genome in cells SERPINA1 A method for at least a portion of a gene, the method comprising contacting the cell with a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.

80. The method of claim 79, wherein the guide RNA and donor polynucleotide target the SERPINA1 Mutations in genes.

81. The method of claim 79, wherein the guide RNA and donor polynucleotide target and replace SERPINA1 c.1096G>A (p.Glu342Lys) mutation.

82. A method for treating α-1-antitrypsin deficiency in a patient with corresponding need, the method comprising administering to the patient a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.

83. The method of claim 82, wherein the guide RNA and donor polynucleotide targeting SERPINA1 Mutations in genes.

84. The method of claim 82, wherein the guide RNA and donor polynucleotide target and replace SERPINA1 c.1096G>A (p.Glu342Lys) mutation.

85. A replacement for the genome in cells DMPK A method for at least a portion of a gene, the method comprising contacting the cell with a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.

86. The method of claim 85, wherein the one or more guide RNAs target the DMPK Mutations in genes.

87. The method of claim 85, wherein the one or more guide RNAs target the DMPK The CAG triplet polynucleotide sequence in the 3' untranslated region of the gene.

88. A method for treating type 1 myotonic dystrophy in patients with corresponding needs, the method comprising administering to the patient a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.

89. A replacement for the genome in cells C9ORF72 A method for at least a portion of a gene, the method comprising contacting the cell with a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.

90. The method of claim 89, wherein the one or more guide RNAs target the C9ORF72 Mutations in genes.

91. The method of claim 89, wherein the one or more guide RNAs target the C9ORF72 The GGGGCC hexanucleotide repeat sequence between exon 1a and exon 1b of the gene.

92. A method for treating amyotrophic lateral sclerosis or frontotemporal dementia in patients with corresponding needs, the method comprising administering to the patient a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.