Dual nuclease system for gene editing, method for its construction, and method for its use.
Patent Information
- Application Number
- JP2026512617
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-08
- Filing Date
- 2024-08-22
- Publication Date
- 2026-09-01
Smart Images

Figure 2026529711000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the interests of U.S. Provisional Application No. 63 / 631,166 filed on 8 April 2024, U.S. Provisional Application No. 63 / 590,771 filed on 16 October 2023, and U.S. Provisional Application No. 63 / 578,032 filed on 22 August 2023, which are incorporated herein by reference in their entirety. [Background technology]
[0002] Existing gene editing technologies such as RNA-programmable gene editors (CRISPR-Cas9 and Cas9 fusions), meganucleases, zinc finger proteins, type IIS restriction endonucleases (FokI and FokI fusions), and TALENS have limited ability to introduce gene deletions of specific lengths or to accurately insert DNA sequences into a sufficient number of cells. For therapeutic gene editing, it is crucial to deliver the gene editor to the target tissue or cells of the organism and to efficiently and accurately modify the target site. To disrupt a gene using an RNA-programmable gene editor, it is important to co-deliver the gene editor and one or more guide RNAs to the cells, and for the gene editor to produce predictable deletions or insertions of small sequences and / or large sequences to achieve the desired editing outcome. To insert a new sequence using an RNA-programmable gene editor, it is important to co-deliver the gene editor, one or more guide RNAs, and one or more repair templates to the cell, and that the gene editor replaces or inserts the sequence at the target site with minimal collateral effects, such as additional insertions and deletions, to achieve the desired repair outcome.
[0003] Current gene editing technologies often rely on inefficient repair pathways in cells, such as error-prone non-homologous end joining (NHEJ), homology-directed repair (HDR), or base excision repair (BER), to modify target sites in the genome, resulting in poor target site repair, cell cycle-specific repair, or unpredictable editing outcomes.
[0004] Several techniques have been developed for delivering gene editors and repair templates in vivo, such as viral and nonviral delivery. However, gene editors are often too large to co-package into a single adeno-associated viral vector (AAV) along with the gene editor, guide RNA, and all the regulatory elements required to express and stabilize the donor DNA sequence or repair template. Similarly, in the case of nonviral delivery using messenger RNA, there are limitations to the co-expression of all the elements necessary for efficient and predictable target site disruption and repair. For ribonucleoprotein (RNP) delivery, current versions of gene editors are limited to co-delivering the donor nucleic acid sequence along with the protein version of the gene editor into the same cell to ensure high-precision repair. In the case of gene editors expressed in cells after viral delivery, it is essential to control nuclease expression over time to control the duration of treatment and mitigate any undesirable editing caused by constitutive expression. [Overview of the project]
[0005] Each of the embodiments and aspects described herein may be used together unless explicitly or explicitly excluded from the context of the embodiment or aspect.
[0006] In one embodiment, a nucleic acid is provided comprising (i) a polynucleotide encoding a chimeric nuclease containing an I-TevI domain and an RNA-inducible nuclease domain, (ii) a polynucleotide encoding a first guide RNA (gRNA), and (iii) a polynucleotide encoding tRNA, wherein the polynucleotides in (ii) to (iii) are in sequential order.
[0007] In one embodiment, a nucleic acid is provided comprising (i) a polynucleotide encoding a chimeric nuclease containing a GIY-YIG nuclease domain and an RNA-inducible nuclease domain, (ii) a polynucleotide encoding a first guide RNA (gRNA), and (iii) a polynucleotide encoding tRNA, where the polynucleotides in (ii) to (iii) are in a contiguous order.
[0008] In some embodiments, the nucleic acid further comprises an RNA-stabilized polynucleotide located downstream of (iv)(i).
[0009] In one embodiment, a nucleic acid is provided comprising (i) a polynucleotide encoding a first guide RNA (gRNA) and (ii) a polynucleotide encoding tRNA, where the polynucleotides in (i) and (ii) are in a consecutive order.
[0010] In some embodiments, the nucleic acid further comprises two or more donor polynucleotides arranged in tandem, and a ribozyme polynucleotide, where the donor polynucleotides are in a contiguous order, and the upstream donor polynucleotide has a ribozyme cleavage site sequence at its 3' end.
[0011] In some embodiments, nucleic acids are DNA or RNA. In some embodiments, DNA is circular plasmid DNA, linear double-stranded DNA, single-stranded DNA, or chimeric RNA and DNA. In some embodiments, RNA is mRNA. In some embodiments, mRNA comprises nucleic acid mimetics selected from the group consisting of peptide nucleic acid (PNA), morpholino nucleic acid, cyclohexenyl nucleic acid (CeNA), and locked nucleic acid (LNA). In some embodiments, mRNA comprises modified sugar moieties, which are optionally selected from the group consisting of N1-methylpseudridine, 9-methyladenine, 2'-O-(2-methoxyethyl), 2'-dimethylaminooxyethoxy, 2'-dimethylaminoethoxyethoxy, 2'-O-methyl, and 2'-fluoro.In some embodiments, mRNA contains modified nucleobases, which are optionally 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl derivative of adenine, 6-methyl derivative of guanine, 2-propyl derivative of adenine, 2-propyl derivative of guanine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 6-azouracil, 6-azocytosine, 6-azocytosine, pseudouracil, 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl, 5-halo, 5-bromo, 5-trifluoromethyl, 5-substitution The following are selected from the group: uracil, 5-substituted cytosine, 7-methylguanine, 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine, 8-azaadenine, 7-deazaguanine, 7-deazaadenine, 3-deazaguanine, 3-deazaadenine, tricyclic pyrimidine, phenoxazinecytidine, phenothiazinecytidine, substituted phenoxazinecytidine, carbazolecytidine, pyridoindolecytidine, 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine, 2-pyridone, 5-substituted pyrimidine, 6-azapyrimidine, N-2, N-6, or O-6 substituted purines, 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. In some embodiments, the mRNA contains internucleoside bonds that are not naturally occurring or are unnatural, selected from the group consisting of phosphorothioates, phosphoramidates, nonphosphodiesters, heteroatoms, chiral phosphorothioates, phosphorodithioates, phosphotryesters, aminoalkylphosphotryesters, 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphinates, 3'-aminophosphoramidates, aminoalkylphosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotryesters, selenophosphates, and boranophosphates.
[0012] In some embodiments, the RNA-inducible nuclease is selected from the group consisting of Staphylococcus aureus Cas9 ("saCas9"), Streptococcus pyogenes Cas9, Acidaminococcus Cas12, and Deltaproteobacteria CasX, as well as Eubacterium rectale Cas12a. In some embodiments, Cas is an inactivated Cas (dCas). In some embodiments, Cas is a nickase or a dCas.
[0013] In some embodiments, I-TevI is a nickase. In some embodiments, the I-TevI nickase domain contains mutations at amino acid residues R27A, V117F, K135R, and N140S. In some embodiments, I-TevI is inactivated. In some embodiments, the I-TevI inactivating mutation is the R27A mutation.
[0014] In some embodiments, the ribozyme polynucleotide is located within the tRNA polynucleotide.
[0015] In some embodiments, the nucleic acid further comprises a polynucleotide encoding a second gRNA, which is optionally located at the 3' position of the tRNA polynucleotide.
[0016] In some embodiments, the polynucleotide sequence is chimeric nuclease, RNA-stabilized polynucleotide, guide RNA, tRNA, and guide RNA. In some embodiments, the polynucleotide sequence is chimeric nuclease, RNA-stabilized polynucleotide, guide RNA, tRNA / ribozyme, guide RNA, donor polynucleotide 1, and donor polynucleotide 2.
[0017] In some embodiments, the RNA-stabilized polynucleotide comprises the 3' sequence of metastasis-associated lung adenocarcinoma transcript 1 (MALAT) or the 3' end of multiple endocrine neoplasia beta transcript (MENβ). In some embodiments, the RNA-stabilized sequence comprises the 3' end of a triple-helix RNA structure or an RNA transcript lacking a standard polyadenylation signal.
[0018] In some embodiments, the donor polynucleotide is single-stranded or double-stranded. In some embodiments, the donor polynucleotide is DNA or RNA. In some embodiments, one strand of the double-stranded donor polynucleotide is DNA and the other strand is RNA. In some embodiments, the donor polynucleotide comprises a cis-acting single-strand RNA polynucleotide annealed to a complementary single-strand DNA polynucleotide. In some embodiments, the single-stranded or double-stranded donor polynucleotide has a 2-18 nucleotide overhang at its 3' end. In some embodiments, the single-stranded or double-stranded donor polynucleotide has a 14 nucleotide overhang at its 3' end. In some embodiments, the overhang at the 3' end is a single-stranded RNA polynucleotide.
[0019] In some embodiments, the nucleic acid further comprises a second guide RNA capable of targeting the 5' region to a donor polynucleotide target site. In some embodiments, the first guide RNA targets a first chimeric nuclease to a first I-TevI or Cas9 target site and is capable of cleaving at the first I-TevI or Cas9 target site in the genome of a cell, and the second guide RNA targets a second chimeric nuclease to a second I-TevI target site in the genome of a cell and is capable of cleaving at the second I-TevI target site, wherein the cleavage produces a nucleotide overhang at the second I-TevI target site.
[0020] In some embodiments, the 3' end of the donor polynucleotide is complementary to the overhang produced by cleavage of the second I-TevI at the second I-TevI target site.
[0021] In some embodiments, the guide RNA and the donor polynucleotide target a mutation in the CFTR gene. In some embodiments, the guide RNA and the donor polynucleotide target and replace the CFTR mutation selected from the group consisting of c.1521_1523del(p.Phe508del), c.1624G>T(p.Gly542Ter), c.1652G>A(p.Gly551Asp), c.1657C>T(p.Arg553Ter), and c.3846G>A(p.Trp1282Ter).
[0022] In some embodiments, the guide RNA and the donor polynucleotide target a mutation in the SERPINA1 gene. In some embodiments, the guide RNA and the donor polynucleotide target and replace the SERPINA1 c.1096G>A(p.Glu342Lys) mutation.
[0023] In some embodiments, the nucleic acid further comprises a promoter. In some embodiments, the promoter is selected from the group consisting of a CMV promoter, an SV40 promoter, a minimal cytomegalovirus (CMV) promoter, and a human elongation factor-1α (EF1a) promoter. In some embodiments, the promoter is selected from the group consisting of muscle-specific synthetic promoter SPc5-12, neuron-specific promoter hSYN1, aldh1L1, cTNT, alpha-MHC, SPc5-12, MUC2, Ksp-cadherin, albumin, HAS, insulin, rhodopsin, rNSE, and cone-opsin promoter.
[0024] In some embodiments, the tRNA comprises glycine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine tRNA.
[0025] In some embodiments, the ribozyme includes a hammerhead ribozyme or a hepatitis delta virus (HDV) ribozyme. In some embodiments, the donor polynucleotide includes a trans-acting double-stranded RNA polynucleotide having a 3' or 5' overhanging nucleotide, a cis-acting single-stranded RNA polynucleotide having sequence similarity to the target strand of a nuclease, a cis-acting single-stranded RNA polynucleotide having sequence similarity to the non-target strand of a nuclease, one or more binding sites for genome modifying factors, optionally, a binding site for site-specific recombinases such as serine recombinase or LoxP target sites, one or more exons having a protein coding sequence repair template, a splice acceptor sequence and a donor sequence, one or more selectable sequences selected from the group of NeoR, BsdR, HygR, PuroR, and BleoR genes, one or more drug-inducible regulatory sequences for controlled gene expression, and / or a 2-nucleotide overhang at the 3' end.
[0026] In some embodiments, the nucleic acid further comprises a polyadenylation signal. In some embodiments, the polyadenylation signal comprises Simian virus 40 (SV40), α-globin, β-globin, human growth hormone (hGH), bovine growth hormone (BGH), herpes simplex virus type 1 thymidine kinase (HSV TK), or a synthetic polyadenylation (Synt poly A) polyadenylation signal.
[0027] In some embodiments, the nucleic acid further comprises a self-inactivating sequence.
[0028] In some embodiments, the nucleic acid is approximately 5 kb in length. In some embodiments, the nucleic acid is less than 5 kb in length.
[0029] In some embodiments, the nucleic acid is packaged in a virus. In some embodiments, the virus is a lentivirus, adeno-associated virus (AAV), adenovirus, retrovirus, or modified herpes simplex virus (HSV). In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV7, AAV8, AAV9, AAV10, AAV-DJ, AAV2.5T, or AAVmyo.
[0030] In one embodiment, a vector comprising the nucleic acid of the present disclosure is provided.
[0031] In one embodiment, a viral vector comprising the nucleic acid of the present disclosure is provided.
[0032] In one embodiment, an AAV virus comprising the nucleic acid of the present disclosure is provided.
[0033] In one embodiment, cells comprising the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure are provided.
[0034] In one embodiment, a composition is provided comprising a chimeric nuclease polypeptide containing an I-TevI domain and an RNA-inducible nuclease domain, and the nucleic acid of the present disclosure.
[0035] In one embodiment, a composition is provided comprising a chimeric nuclease nucleic acid encoding a chimeric nuclease including an I-TevI domain and an RNA-inducible nuclease domain, and the nucleic acid of the present disclosure.
[0036] In some embodiments, the chimeric nuclease nucleic acid is mRNA.
[0037] In one embodiment, an LNP composition comprising the nucleic acid or composition of the Disclosure is provided.
[0038] In one embodiment, a pharmaceutical composition is provided comprising a nucleic acid of the Disclosure, a vector of the Disclosure, a viral vector of the Disclosure, or an AAV of the Disclosure, a composition of the Disclosure, or an LNP composition of the Disclosure, and an excipient.
[0039] One embodiment provides a method for delivering messenger RNA encoding a chimeric nuclease comprising an I-TevI domain and an RNA-inducible nuclease domain to a cell, the method comprising contacting the cell with a polynucleotide encoding one or more guide RNAs and a polynucleotide donor.
[0040] In one embodiment, a method is provided for genetically modifying the genome of a cell, the method comprising contacting the cell with the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure.
[0041] In some embodiments, the modification includes genomic insertions, deletions, substitutions, or mutations. In some embodiments, the insertion of a donor polynucleotide into the cell's genome results in the removal of a sequence between an I-TevI target site and a Cas9 target site. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell.
[0042] In one embodiment, a method is provided for inserting or substituting a sequence into a chimeric nuclease target site in the genome of a cell, the method comprising contacting a nucleic acid comprising a chimeric nuclease comprising a Cas9 domain and an I-TevI domain, and one or more nucleic acid sequences encoding a nucleic acid comprising a guide polynucleotide and a donor polynucleotide, wherein the guide polynucleotide and the chimeric nuclease form a complex, the complex binding to genomic DNA at the Cas9 target site and the I-TevI target site, and cleaving the genomic DNA, where the 3' end of the donor polynucleotide contains at least two bases complementary to the I-TevI target site after I-TevI cleavage, and the donor polynucleotide is incorporated into the chimeric nuclease target site at position 5' relative to the Cas9 target site.
[0043] In some embodiments, the 3' end of the guide polynucleotide and the 5' end of the donor polynucleotide are ligated together.
[0044] In some embodiments, the cell polymerase is targeted to a chimeric nuclease target site. In some embodiments, the cell polymerase is a polymerase theta.
[0045] In one embodiment, a method is provided for replacing at least a portion of the CFTR gene in the genome of a cell, the method comprising contacting the cell with the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure. In some embodiments, the guide RNA and donor polynucleotides target mutations in the CFTR gene.
[0046] In some embodiments, the guide RNA and donor polynucleotides target and replace the following mutations in CFTR: c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), c.1652G>A (p.Gly551Asp), c.1657C>T (p.Arg553Ter), or c.3846G>A (p.Trp1282Ter).
[0047] In one embodiment, a method is provided for treating cystic fibrosis in a patient requiring treatment for cystic fibrosis, the method comprising administering to the patient the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure.
[0048] In some embodiments, the guide RNA and donor polynucleotides target mutations in the CFTR gene. In some embodiments, the guide RNA and donor polynucleotides target and replace the following mutations in CFTR: c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), c.1652G>A (p.Gly551Asp), c.1657C>T (p.Arg553Ter), or c.3846G>A (p.Trp1282Ter).
[0049] A method is provided for replacing at least a portion of the SERPINA1 gene in the genome of a cell, the method comprising contacting the cell with the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure.
[0050] In some embodiments, the guide RNA and donor polynucleotides target mutations in the SERPINA1 gene. In some embodiments, the guide RNA and donor polynucleotides target and replace the SERPINA1 c.1096G>A(p.Glu342Lys) mutation.
[0051] In one embodiment, a method is provided for treating α-1-antitrypsin deficiency in a patient requiring treatment for α-1-antitrypsin deficiency, the method comprising administering to the patient the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure. In some embodiments, the guide RNA and donor polynucleotide target mutations in the SERPINA1 gene. In some embodiments, the guide RNA and donor polynucleotide target and substitute the SERPINA1 c.1096G>A(p.Glu342Lys) mutation.
[0052] In one embodiment, a method is provided for replacing at least a portion of the DMPK gene in the genome of a cell, comprising the step of contacting the cell with the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure. In some embodiments, one or more guide RNAs target a mutation in the DMPK gene. In some embodiments, one or more guide RNAs target a CAG triplet polynucleotide sequence in the 3' untranslated region of the DMPK gene.
[0053] In one embodiment, a method is provided for treating myotonic dystrophy type 1 in a patient requiring treatment for myotonic dystrophy type 1, the method comprising administering to the patient the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure.
[0054] In one embodiment, a method is provided for replacing at least a portion of the C9ORF72 gene in the genome of a cell, comprising the step of contacting the cell with the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure. In some embodiments, one or more guide RNAs target mutations in the C9ORF72 gene. In some embodiments, one or more guide RNAs target the GGGGCC hexanucleotide repeat sequence between exon 1a and exon 1b of the C9ORF72 gene.
[0055] In one embodiment, a method is provided for treating amyotrophic lateral sclerosis or frontotemporal dementia in a patient requiring treatment for the condition, the method comprising administering the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the patient's LNP composition of the Disclosure. [Brief explanation of the drawing]
[0056] The features of this disclosure are described in detail in the attached claims. A better understanding of the features and advantages of this disclosure will be obtained by referring to the following detailed description, which describes exemplary embodiments in which the principles of this disclosure are utilized, and to the attached drawings. [Figure 1]A schematic diagram of an exemplary AAV cassette encoding elements for disruption or repair of a target site is shown, including a CMV promoter, Dualase, MALAT sequence, gRNA1, tRNA', gRNA2, repair templates (RT), and a poly(A) sequence (synt[A]). The minimal CMV promoter drives the transcription of this exemplary 3-in-1 construct containing Dualase, long non-coding RNA MALAT-1, synthetic transfer RNA (tRNA') containing a sequence-specific ribozyme, gRNA (RT) fused with the repair template, and synthetic poly(A) signal sequence (synt[A]). During transcription, RNA maturation at the 3' end of MALAT and both ends of tRNA' results in the separation of Dualase-MALAT RNA, tRNA, and gRNA. U-rich repeats on MALAT may play a role in protecting Dualase RNA from degradation, while tRNA' is recognized and cleaved in the regions between RTs and between RTs and the synt[A] sequence. After translation, Dualase forms a complex with gRNA (RNP) and can cleave DNA at the intended target site. The presence of an in-place repair template (cis-sense and antisense) fused to the 3' end of the gRNA or an abundant free-floating repair template (trans-antisense and trans-sense) acts as a bridge between the two cleavage sites and a local reference template for cellular repair machinery. [Figure 2]A schematic diagram of an exemplary mRNA cassette encoding elements for efficient and precise target site disruption or repair is shown. The T7 promoter drives the transcription of an exemplary 3-in-1 construct containing Dualase, a long non-coding RNA MALAT-1, a synthetic transfer RNA (tRNA') containing a sequence-specific ribozyme, a gRNA (RT) fused with a repair template, and a synthetic poly(A) signaling sequence (synt[A]). Upon cell delivery, RNA maturation at the 3' end of MALAT and both ends of tRNA' results in the separation of the mRNA encoding Dualase-MALAT RNA, tRNA, and gRNA. U-rich repeats on the MALAT sequence can protect Dualase RNA from degradation, while tRNA' can recognize and cleave regions between RTs, as well as between RTs and the synt[A] sequence. Post-translation, Dualase can complex with gRNA (RNP) and cleave at the intended target site. Optionally, the GFP coding sequence is isolated from the Dualase protein in cells by an N-terminal T2A (thosea asigna virus 2A) peptide sequence that is included to track cellular uptake, translated, and skipped by ribosomes. The presence of in-place repair templates (cis-sense and antisense) fused to the 3' end of the gRNA or abundant free-floating repair templates (trans-antisense and trans-sense) serves as a bridge between the two cleavage sites and a local reference template for the cellular repair mechanism. [Figure 3] A schematic diagram of an exemplary dual-cleavage nuclease ribonucleoprotein (RNP) complex (Dualase) with an all-in-one guide RNA repair cassette is shown. Dualase and 2-in-1 gRNA-RT are incubated to form the ribonucleoprotein complex (RNP), which is then delivered to target cells where they can interact and cleave the intended target site after translocation to the nucleus. [Figure 4A]Figure 4A shows exemplary Tracking of Indels by DEcomposition (TIDE) data for genome editing results using AAV-Dualase-AAVS1 Dualase. Deletions and insertions of defined length, as well as unmodified reads, are characterized as bar graphs. (Figure 4A). Dotted line P<0.01. (Figure 4B) Chromatograms of sequenced samples treated with AAV-Dualase-AAVS1 or reagent alone are shown compared to a reference sequence for which peak differences are used to determine the repaired sequence. [Figure 4B] Figure 4A shows exemplary tracer (TIDE) data of indels due to DE degradation resulting from genome editing with AAV-Dualase-AAVS1 Dualase. Defined length deletions and insertions, as well as unmodified reads, are characterized as bar graphs. (Figure 4A). Dotted line P<0.01. (Figure 4B) Chromatograms of sequenced samples treated with AAV-Dualase-AAVS1 or reagent alone are shown compared to a reference sequence for which peak differences are used to determine the repaired sequence. [Figure 5A]Bar graphs and sequence plots of results from next-generation sequencing (NGS) of cells only, cis-antisense, cis-sense, and trans-RT are shown. NGS data were analyzed using two bioinformatics tools (Figure 5A) CRIS.Py and the Geneious alignment platform, as well as (Figure 5B) CRISPRESSO2. "Indels" represent repair events resulting in insertions or deletions. "Precise repair" represents sequencing reads with complete alignment to the repaired sequence. "Repair + SNP" represents sequence alignment with repair and another nucleotide change. "Other" represents other sequence modifications in the sample. Editing results were identified using all reads and aligned reads. Alignment of NGS sequencing reads of cis-antisense repair sequences is shown in Figure 5C, displaying the number and percentage of reads. The dotted lines indicate reads below the generally acceptable limit for sequencing detection, and I-TevI and Cas9 target sites are shown above the sequences. [Figure 5B]Bar graphs and sequence plots of results from next-generation sequencing (NGS) of cells only, cis-antisense, cis-sense, and trans-RT are shown. NGS data were analyzed using two bioinformatics tools (Figure 5A) CRIS.Py and the Geneious alignment platform, as well as (Figure 5B) CRISPRESSO2. "Indel" represents a repair event resulting in an insertion or deletion. "Precise Repair" represents a sequencing read with perfect alignment to the repaired sequence. "Repair + SNP" represents a sequence alignment with repair and another nucleotide change. "Other" represents other sequence modifications in the sample. Edit results were identified using all reads and aligned reads. Alignment of NGS sequencing reads of cis-antisense repair sequences is shown in Figure 5C, with the number and percentage of reads displayed. Dotted lines indicate reads below the generally acceptable limit for sequencing detection, and I-TevI and Cas9 target sites are indicated above the sequence. [Figure 5C] Bar graphs and sequence plots of results from next-generation sequencing (NGS) of cells only, cis-antisense, cis-sense, and trans-RT are shown. NGS data were analyzed using two bioinformatics tools (Figure 5A) CRIS.Py and the Geneious alignment platform, as well as (Figure 5B) CRISPRESSO2. "Indel" represents a repair event resulting in an insertion or deletion. "Precise Repair" represents a sequencing read with perfect alignment to the repaired sequence. "Repair + SNP" represents a sequence alignment with repair and another nucleotide change. "Other" represents other sequence modifications in the sample. Edit results were identified using all reads and aligned reads. Alignment of NGS sequencing reads of cis-antisense repair sequences is shown in Figure 5C, with the number and percentage of reads displayed. Dotted lines indicate reads below the generally acceptable limit for sequencing detection, and I-TevI and Cas9 target sites are indicated above the sequence. [Figure 6A]Figure 6A shows a gel photograph illustrating the editing efficiency of HEK293 cells treated with an all-in-one AAV expressing Dualase and gRNA-RT that targets AAVS1, and a control of an AAV expressing Dualase that targets AAVS1, co-delivered by lipofection with a specified repair template. The treatment was performed under conditions of increased doses of inhibitors blocking non-homologous end-joining (NHEJ), homologous-directed repair (HDR), or Rad52-dependent pathways. (Figure 6C) NHEJ (DNA) = non-homologous double-stranded DNA, HDR (DNA) = double-stranded DNA with homology arms, and NHEJ (RNA) = non-homologous single-stranded RNA. [Figure 6B] Figure 6A shows a gel photograph illustrating the restriction enzyme (RE) editing efficiency in HEK293 cells treated with AAV (all-in-one) expressing Dualase and gRNA-RT targeting AAVS1, and a control of AAV expressing Dualase targeting AAVS1, co-delivered by lipofection with a specified repair template (co-delivered), under conditions of non-homologous end joining (NHEJ), (Figure 6B) homology-dependent repair (HDR), or (Figure 6C) increased doses of inhibitors blocking the Rad52-dependent pathway. Samples were analyzed for editing efficiency by restriction enzyme (RE) digestion efficiency using RE specific to the inserted repair template site. NHEJ(DNA) = non-homologous double-stranded DNA, HDR(DNA) = double-stranded DNA with homologous arms, and NHEJ(RNA) = non-homologous single-stranded RNA. [Figure 6C]Figure 6A shows a gel photograph illustrating the restriction enzyme (RE) editing efficiency in HEK293 cells treated with AAV (all-in-one) expressing Dualase and gRNA-RT targeting AAVS1, and a control of AAV expressing Dualase targeting AAVS1, co-delivered by lipofection with a specified repair template (co-delivered), under conditions of non-homologous end joining (NHEJ), (Figure 6B) homology-dependent repair (HDR), or (Figure 6C) increased doses of inhibitors blocking the Rad52-dependent pathway. Samples were analyzed for editing efficiency by restriction enzyme (RE) digestion efficiency using RE specific to the inserted repair template site. NHEJ(DNA) = non-homologous double-stranded DNA, HDR(DNA) = double-stranded DNA with homologous arms, and NHEJ(RNA) = non-homologous single-stranded RNA. [Figure 7A]Figure 7A shows a gel photograph illustrating directional insertion of an RNA repair template by a Dualase ribonucleoprotein (RNP) complex formed with AAVS1-targeted guide RNA combined with an RNA repair template, via directional polymerase chain reaction (PCR). Restriction enzyme digestion of HEK293 cells lipofected with Dualase, either with separately delivered guide RNA and repair template (co-delivered), fused guide RNA and RNA repair template ("all-in-one"), lipofection reagent only ("reagent only"), or cells only ("mock"). For each reaction, the undigested PCR amplicon of the AAVS1 site ("Sub") and the digested PCR amplicon of the AAVS1 site ("Digested") are shown. (Figure 7B) This shows that the PCR product is in the correct orientation ("Right-oriented RT") only when the repair template is correctly inserted, and in the wrong orientation ("Wrong-oriented RT") when the repair template is incorrectly inserted. [Figure 7B]Figure 7A shows a gel photograph illustrating the directional insertion of an RNA repair template by a Dualase ribonucleoprotein (RNP) complex formed with AAVS1-targeted guide RNA combined with an RNA repair template via directional polymerase chain reaction (PCR). Restriction enzyme digestion of HEK293 cells lipofected with Dualase with separately delivered guide RNA and repair template (co-delivery), fused guide RNA and RNA repair template ("all-in-one"), lipofection reagent only ("reagent only"), or cells only ("mock") is shown. For each reaction, the undigested PCR amplicon of the AAVS1 site ("Sub") and the digested PCR amplicon of the AAVS1 site ("Digested") are shown. Figure 7B shows that the PCR product is correctly oriented ("Right-oriented RT") only when the repair template is correctly inserted, and incorrectly oriented ("Wrong-oriented RT") when the repair template is incorrectly inserted. [Figure 8]The images show gels demonstrating directional insertion of repair templates using AAVS1-targeted Dualase all-in-one mRNA with cis-sense or cis-antisense repair templates via directional polymerase chain reaction (PCR). The results of restriction enzyme digestion (BglII digestion) or PCR reactions of HEK293 cells lipofected with Dualase (Tev[VKN]-SaCas9[WT] or Tev[VKN]-SaCas9[D10E]), SaCas9[WT], or lipofection reagents alone are shown. For each reaction, the undigested PCR amplicon of the AAVS1 site ("Full length"), the digested PCR amplicon of the AAVS1 site ("BglII digest"), the PCR product when the repair template was inserted only in the correct orientation ("correctly oriented RT"), and the PCR product when the repair template was inserted only in the wrong orientation ("mis-oriented RT") are shown. The image also shows control cells treated with Dualase mRNA targeting the AAVS1 site via a co-delivered repair template (co-delivery). [Figure 9] The image shows a gel representing cell repair after dualase cleavage and treatment with NHEJ and Rad52 inhibitors and a trans-dsRNA repair template. HEK293 cells were treated by lipofection with AAV (all-in-one) expressing AAVS1-targeting Dualase and gRNA-RT, and a control (co-delivered) of AAV expressing AAVS1-targeting Dualase, co-delivered with the shown repair template, under increasing doses of an inhibitor that blocks non-homologous end joining (NHEJ inhibitor) or an inhibitor that blocks the Rad52 pathway ("Rad52 inhibitor"). NHEJ(DNA) = non-homologous double-stranded DNA, HDR(DNA) = double-stranded DNA with homologous arms, and NHEJ(RNA) = non-homologous single-stranded RNA. [Figure 10A]A schematic diagram of the precise removal of large repetitive sequences using dual-guide TevCas9 nucleases is shown, where the guide RNA targets the opposite strands of double-stranded DNA and orients the two I-TevI domains head-to-head. The first TevCas9 nuclease (1), containing an inactivated D10A+H557A mutation but an active I-TevI domain (2), is targeted to the 5' end of the repetitive sequence (3) using the guide RNA. The second TevCas9 nuclease (4), containing an inactivated D10A+H557A mutation but an active I-TevI domain, is targeted to the 3' end of the repetitive sequence, but both nucleases are on the opposite strand so that the I-TevI domains are oriented towards the repetitive sequence. Two complementary 3'2-nucleotide overhangs remain (6 and 7) upon binding and cleavage by the I-TevI nuclease domains (5). Through non-homologous end joining pathways, cells can repair these complementary overhangs (8), removing large repetitive sequences (9) and leaving a predetermined number of repeats in the genomic DNA (10). [Figure 10B] A schematic diagram of the precise removal of large repetitive sequences using dual-guide TevCas9 nucleases is shown, where the guide RNA targets the same strand of double-stranded DNA and orients the two I-TevI domains in tandem. The first TevCas9 nuclease (1), containing an inactivated D10A+H557A mutation but an active I-TevI domain (2), is targeted at the 5' end of the repetitive sequence (3) using the first guide RNA. The second TevCas9 nuclease (4), containing an inactivated D10A+H557A mutation but an active I-TevI domain, targets downstream of the 3' end of the repetitive sequence on the same strand, so that the I-TevI domains of both chimeric nucleases are oriented in the same direction. Two complementary 3' 2-nucleotide overhangs remain (6 and 7) upon binding and cleavage by the two I-TevI nuclease domains (5). Through non-homologous end joining pathways, cells can repair these complementary overhangs (8), removing large repetitive sequences (9) and leaving a predetermined number of repeats in the genomic DNA (10). [Figure 11A] Figure 11A shows the results of removing CAG triplet repeat extensions in the 3'-UTR of DMPK (myotonic dystrophy protein kinase), (also labeled DmpkI), using an all-in-one AAV encoding TevCas9, along with dual guides targeting the 5' and 3' ends of the repeat extension. The active I-TevI domain and inactivated Cas9 domain with the D10A+H557A mutation are targeted by two guide RNAs to orient the I-TevI domain downstream of the CAG repeat sequence and CAG repeats. More than 75 CAG repeats result in skeletal muscle diseases such as myotonic dystrophy. [Figure 11B] Figure 11B shows the results of removing the CAG triplet repeat extension in the 3'-UTR of DMPK (myotonic dystrophy protein kinase, also labeled as DmpkI) using an all-in-one AAV encoding TevCas9, along with dual guides targeting the 5' and 3' ends of the repeat extension. Figure 11B shows a schematic diagram of the expression cassette encoding dual-guide Tev[KTQ]-SaCas9[D10A+H557A] and conditionally cleavable guide RNA, where expressed miniCMV (miniCMV) is the promoter sequence, HHribo is the hammerhead ribozyme sequence, sgRNA1 is the 5' guide RNA target sequence, Gly tRNA is the glycine tRNA sequence, sgRNA2 is the 3' guide RNA target sequence, HDVribo is the hepatitis delta virus ribozyme sequence, and polyA is the polyadenylated sequence (SEQ ID NO: 118). [Figure 11C]Figure 11C shows the results of removing CAG triplet repeat extensions in the 3'-UTR of DMPK (myotonic dystrophy protein kinase, also labeled DmpkI) using an all-in-one AAV encoding TevCas9, along with a dual guide that targets the 5' and 3' ends of the repeat extension. Figure 11C shows a schematic diagram of the expected products of the in vitro cleavage reaction using purified Tev[KTQ]-saCas9[D10A+H557A] protein complexed with a dual guide that targets DMPK CAG repeats in equimolar ratios. To the right is an agarose gel containing the products of the in vitro cleavage reaction using a DMPK CAG repeat DNA substrate mixed with the Tev[KTQ]-saCas9[D10A+H557A] dual guide ribonucleoprotein complex. Expected products of the reaction by size are shown next to the gel image. [Figure 11D] Figure 11D shows the results of removing CAG triplet repeat extensions in the 3'-UTR of DMPK (myotonic dystrophy protein kinase, also labeled DmpkI) using an all-in-one AAV encoding TevCas9 with dual guides targeting the 5' and 3' ends of the repeat extension. Alignment of Sanger sequencing reads of clonal amplicons 1-10 generated from cells transduced with the all-in-one AAV encoding TevCas9 with dual guides, and the expected cleavage. [CAG]51+ indicates the presence of the repeat sequence, and the lowercase letter indicates the exact number of CAG repeats after TevCas9 cleavage. [Figure 11E]Figure 11E shows the results of removing CAG triplet repeat extensions in the 3'-UTR of DMPK (myotonic dystrophy protein kinase, also labeled DmpkI) using an all-in-one AAV encoding TevCas9, along with a dual guide that targets the 5' and 3' ends of the repeat extension. The experimental workflow for quantifying repeat collapse using repeat-prime PCR in DMPKMUT CAG repeat cells transduced with the all-in-one AAV TevCas9 and SaCas9 dual guide is shown. Transduced cells are collected, and genomic DNA is extracted for PCR amplification using outer and inner primers for repeat amplification. The resulting PCR products are analyzed with an Agilent Bioanalyzer, and peaks corresponding to repeats are quantified only for DMPKMUT CAG repeat cells. Over 65% of repeats collapsed in TevCas9-treated cells, compared to approximately 30% in SaCas9-treated cells. [Figure 11F] Figure 11F shows the results of removing the CAG triplet repeat extension in the 3'-UTR of DMPK (myotonic dystrophy protein kinase, also labeled DmpkI) using an all-in-one AAV encoding TevCas9, along with dual guides targeting the 5' and 3' ends of the repeat extension. Graphs show the multiplicative changes in DMPK in transduced cells subjected to quantitative RT-PCR to quantify the levels of DMPKMUT and DMPKWT transcripts after each treatment, using primers specific to each transcript. Statistical significance of the changes is indicated by horizontal bar labels as ns = no significant change, *p = <0.05, ** = <0.01, or *** = p < 0.001. [Figure 11G]Figure 11G shows the results of removing CAG triplet repeat elongation in the 3'-UTR of DMPK (myotonic dystrophy protein kinase, also labeled as DmpkI) using an all-in-one AAV encoding TevCas9, along with a dual guide that targets the 5' and 3' ends of the repeat elongation. Figure 11G shows the plicatilization (fold) of TevCas9 dual guide (TevCas9dg), TevCas9 single guide (TevCas9sg), and Cas9 dual guide (Cas9dg) transduced cells compared to untreated cells ("Cell Only") with a missplicing mutant of the CLCN1 gene at exon 6 (E6). The graph shows the change. Missplicing of human ClC-1 chloride channels causes myotonic dystrophy. Both TevCas9 dual-guide and TevCas9 single-guide significantly reduced missplicing of CLCN1 exon 6, while Cas9 did not significantly alter missplicing. The statistical significance of the change is indicated by horizontal bar labels as ns = no significant change, ** = <0.01, or *** = p < 0.001. [Figure 11H] Figures 11H-11T show the results of removing a CAG triplet repeat extension in the 3'-UTR of DMPK (myotonic dystrophy protein kinase), (also labeled DmpkI), using an all-in-one AAV encoding TevCas9 with dual guides targeting the 5' and 3' ends of the repeat extension. Figures 11H-11T show the results of removing a GGGGCC hexanucleotide repeat extension between exons 1a and 1b of the C9ORF72 gene using the all-in-one AAV TevCas9 with dual guides targeting the 5' and 3' ends of the repeat extension. Figure 11H shows a schematic diagram of the active I-TevI domain and inactivated Cas9 domain with the D10A+H557A mutation, targeted by two guide RNAs to orient the I-TevI domain to the GGGGCC repeat sequence. More than 24 GGGGCC repeats result in motor neuron diseases such as ALS. [Figure 11I]Figures 11H-T show the results of removing the CAG triplet repeat extension in the 3'-UTR of the C9ORF72 gene using an all-in-one AAV encoding TevCas9 with a dual guide that targets the 5' and 3' ends of the repeat extension. Figure 11I shows the results of removing the GGGGCC hexanucleotide repeat extension between exons 1a and 1b of the C9ORF72 gene using an all-in-one AAV TevCas9 with a dual guide that targets the 5' and 3' ends of the repeat extension. A schematic diagram of the expected products of the in vitro cleavage reaction is shown. To the right is an agarose gel containing the products of the in vitro cleavage reaction using a C9ORF72 GGGGCC repeat DNA substrate mixed with the Tev[VKN]-saCas9[D10A+H557A] dual-guide ribonucleoprotein complex. The expected products of the reaction based on size are shown next to the gel image. [Figure 11J] Figures 11H-T show the results of removing the CAG triplet repeat extension in the 3'-UTR of the C9ORF72 gene using an all-in-one AAV encoding TevCas9, along with a dual guide that targets the 5' and 3' ends of the repeat extension. Figure 11J shows the results of removing the GGGGCC hexanucleotide repeat extension between exons 1a and 1b of the C9ORF72 gene using the all-in-one AAV TevCas9, along with a dual guide that targets the 5' and 3' ends of the repeat extension. The results of removing the GGGGCC hexanucleotide repeat extension between exons 1a and 1b of the C9ORF72 gene using an all-in-one AAV encoding TevCas9, along with a dual guide that targets the 5' and 3' ends of the repeat extension. The results of removing the GGGGCC hexanucleotide repeat extension between exons 1a and 1b of the C9ORF72 gene are shown. [Figure 11K]Figures 11H-T show the results of removing the CAG triplet repeat extension in the 3'-UTR of DMPK (myotonic dystrophy protein kinase, also labeled DmpkI) using an all-in-one AAV encoding TevCas9 with dual guides targeting the 5' and 3' ends of the repeat extension. Figure 11K shows the results of removing the GGGGCC hexanucleotide repeat extension between exons 1a and 1b of the C9ORF72 gene using the all-in-one AAV TevCas9 with dual guides targeting the 5' and 3' ends of the repeat extension. [Figure 11L] Figures 11H-T show the results of removing a CAG triplet repeat extension in the 3'-UTR of DMPK (myotonic dystrophy protein kinase, also labeled DmpkI) using an all-in-one AAV encoding TevCas9 with dual guides targeting the 5' and 3' ends of the repeat extension. Figure 11L shows the results of removing a GGGGCC hexanucleotide repeat extension between exons 1a and 1b of the C9ORF72 gene using the all-in-one AAV TevCas9 with dual guides targeting the 5' and 3' ends of the repeat extension. Figure 11L shows the results of sequencing a single clone of folded repeat sequences from a TevCas9 AAV-transduced motor neuron, showing the exact folded repeat products in four of the nine sequences indicated by Asterix(*). [Figure 11M]Figures 11H-T show the results of removing the CAG triplet repeat extension in the 3'-UTR of the C9ORF72 gene using an all-in-one AAV encoding TevCas9 with a dual guide that targets the 5' and 3' ends of the repeat extension. Figure 11M shows the results of removing the GGGGCC hexanucleotide repeat extension between exons 1a and 1b of the C9ORF72 gene using the all-in-one AAV TevCas9 with a dual guide that targets the 5' and 3' ends of the repeat extension. Figure 11M shows an overview of the edit percentages determined by deep sequencing analysis of potential off-target sites in motor neurons transduced with an AAV expressing the TevCas9 dual guide. The hashed line indicates the detection limit of deep sequencing for potential off-target effects for which no off-targets were detected in the transduced cells. [Figure 11N] Figures 11H-T show the results of removing a CAG triplet repeat elongation in the 3'-UTR of the C9ORF72 gene using an all-in-one AAV encoding TevCas9 with dual guides targeting the 5' and 3' ends of the repeat elongation. Figure 11N shows the results of removing a GGGGCC hexanucleotide repeat elongation between exons 1a and 1b of the C9ORF72 gene using the all-in-one AAV TevCas9 with dual guides targeting the 5' and 3' ends of the repeat elongation. Figure 11N shows a summary of edit percentages determined by deep sequencing analysis of potential off-target sites in motor neurons transduced with an AAV expressing a SaCas9 dual guide. An off-target site (OT24) with active SaCas9 is shown on chromosome 11. Statistical significance of changes is labeled as nd = no significant difference or * = p < 0.001. [Figure 11O]Figure 11H-T shows the results of removing CAG triplet repeat extensions in the 3'-UTR of DMPK (myotonic dystrophy protein kinase, also labeled DmpkI) using an all-in-one AAV encoding TevCas9, along with dual guides targeting the 5' and 3' ends of the repeat extension. Figure 110 shows the results of removing the GGGGCC hexanucleotide repeat extension between exons 1a and 1b of the C9ORF72 gene using TevCas9. The left Western blot shows the results using anti-C9ORF72 antibody and anti-GAPDH housekeeping gene control antibody in AAV-transduced motor neurons expressing TevCas9 dual guide (TevCas9dg), TevCas9 single guide (TevCas9sg), and Cas9 dual guide (Cas9dg), as well as in cell-only controls. The right Western blot shows a quantified summary of the C9ORF72 expression change multiplier in three replicas of Western blots normalized to GAPDH housekeeping gene expression. Treatment of motor neurons with AAV-expressing TevCas9 and C9ORF72 dual guide significantly increased expression by a >2-fold factor. Statistical significance of the change is labeled a*=p<0.01. [Figure 11P]Figure 11H-T shows the results of removing a CAG triplet repeat elongation in the 3'-UTR of the C9ORF72 gene using an all-in-one AAV encoding TevCas9 with a dual guide that targets the 5' and 3' ends of the repeat elongation. Figure 11P shows the results of removing a GGGGCC hexanucleotide repeat elongation between exons 1a and 1b of the C9ORF72 gene using an all-in-one AAV TevCas9 with a dual guide that targets the 5' and 3' ends of the repeat elongation. Figure 11P shows representative anti-polyGR dipeptide dot blots on the left side of cell lysates from motor neurons transduced with AAVs expressing TevCas9 dual guide (TevCas9dg), TevCas9 single guide (TevCas9sg), and Cas9 dual guide (Cas9dg), as well as mock-treated cells. TevCas9 shows a reduction in the amount of polyGR detected. The summary quantification of the relative intensity of dot blots for cells mocked with the c9orf72 dual guide is shown on the right. [Figure 11Q]Figures 11H-T show the results of removing the CAG triplet repeat elongation in the 3'-UTR of the C9ORF72 gene using an all-in-one AAV encoding TevCas9 with dual guides that target the 5' and 3' ends of the repeat elongation. Figure 11Q shows the results of removing the GGGGCC hexanucleotide repeat elongation between exons 1a and 1b of the C9ORF72 gene using an all-in-one AAV TevCas9 with dual guides that target the 5' and 3' ends of the repeat elongation. A schematic diagram of intracranial (ICV) injection of 1 × 10¹³ viral genomes (High = 2 × 10¹³) is shown. On the right are the results of RT-qPCR to detect TevCas9 expression in cerebellar samples from injected mice 28 days after injection. Control mice were injected with phosphate-buffered saline (PBS). Reverse transcriptase quantitative PCR (RT-qPCR) shows dose-dependent expression of TevCas9 in bulk cerebellar tissue compared to PBS-injected control mice (n=3). [Figure 11R]Figures 11H-T show the results of removing a CAG triplet repeat extension in the 3'-UTR of the C9ORF72 gene using an all-in-one AAV encoding TevCas9 with a dual guide that targets the 5' and 3' ends of the repeat extension. Figure 11R shows the results of removing a GGGGCC hexanucleotide repeat extension between exons 1a and 1b of the C9ORF72 gene using an all-in-one AAV TevCas9 with a dual guide that targets the 5' and 3' ends of the repeat extension. Figure 11R shows the results of PCR to detect the size of GGGGCC repeats in the cerebellum of humanized mice injected with phosphate-buffered saline (PBS) or two doses of AAV expressing dual-guide TevCas9. Representative extensions and normal-sized repeats from agarose gel of PCR products amplified from genomic DNA are shown on the left. A summary of the quantification of normal-sized C9ORF72 PCR products for housekeeping genes (n=3) is shown on the right. [Figure 11S] Figures 11H-T show the results of removing the CAG triplet repeat extension in the 3'-UTR of the C9ORF72 gene using an all-in-one AAV encoding TevCas9 with dual guides that target the 5' and 3' ends of the repeat extension. Figure 11S shows the results of removing the GGGGCC hexanucleotide repeat extension between exons 1a and 1b of the C9ORF72 gene using an all-in-one AAV TevCas9 with dual guides that target the 5' and 3' ends of the repeat extension. [Figure 11T]Figures 11H-11T show the results of removing the CAG triplet repeat extension in the 3'-UTR of the C9ORF72 gene using an all-in-one AAV encoding TevCas9 with dual guides that target the 5' and 3' ends of the repeat extension. Figure 11H-11T shows the results of removing the GGGGCC hexanucleotide repeat extension between exons 1a and 1b of the C9ORF72 gene using an all-in-one AAV TevCas9 with dual guides that target the 5' and 3' ends of the repeat extension. Figure 11T shows the results of Western blotting of the C9ORF72 protein in the cerebellum of humanized mice injected with PBS or two doses of AAV expressing dual-guide TevCas9. [Figure 12A] This figure shows the structure of a self-inactivating Dualase vector. The construct includes a promoter, human codon-optimized TevSaCas9, a polyadenylation signal ("PolyA"), and a nucleotide sequence encoding a guide RNA sequence ("gRNA"). The self-inactivation target site can be located in the region between the promoter and the TevSaCas9 site (indicated as "promoter"), or at the end of the TevSaCas9 coding sequence and the beginning of the PolyA sequence (indicated as "PolyA"). [Figure 12B] The gel shows the results of the T7E1 editing assay in HEK293 cells transfected with plasmid DNA containing TevSaCas9 targeting the beta-2-microglobulin (B2M) gene, collected 24, 48, and 72 hours after transfection. Lanes marked "None" contain no self-inactivating sequences in the vector, lanes marked "Promoter" contain the B2M1 TevSaCas9 target site between the promoter sequence and the TevSaCas9 sequence, and lanes marked "Poly-A" contain the B2M1 TevSaCas9 target site between the end of TevSaCas9 and the poly-A signal sequence. The level of editing, determined by the amount of digested product relative to the substrate, is comparable over time between constructs. [Figure 12C]Figure 12B shows Western blots of hemagglutinin (α-HA) encoded at the 3' end of TevSaCas9 ("Dualase") from the same treated cells. Lanes marked "None" do not contain a self-inactivating sequence in the vector; lanes marked "Promoter" contain the B2M1 TevSaCas9 target site between the promoter sequence and the TevSaCas9 sequence; and lanes marked "Poly-A" contain the B2M1 TevSaCas9 target site between the end of TevSaCas9 and the poly-A signal sequence. Beta-actin (α-Act) was blotted on the same membrane as the loading control. [Figure 12D] This document demonstrates the production of poly(A) self-inactivated AAV2 virus targeting the B2M gene, with titers measured by Western blotting for hemagglutinin (α-HA) encoded at the 3' end of TevSaCas9 ("Dualase") in virally transduced cells, and beta-actin (α-Act) was blotted on the same membrane as the loading control. [Figure 12E] Figure 12D shows the plasmid DNA version of the vector used to produce self-inactivating AAV ("Self-inactivating"), as well as the results of 14 days of Western blots for saCas9 and GAPDH from cells lipofected with a GFP-expressing plasmid ("pAAV-GFP") as a control, along with a vector that does not contain the self-inactivating sequence ("non-self-inactivating") and a control. [Figure 13A] Schematic diagrams of the AAVS1 target site are shown, before (Figure 13A), after (Figure 13B), and associated with gRNA repair templates for the AAVS1 site for accurate repair using Dualase and guide RNA targeting AAVS1 (Figure 13C). [Figure 13B]Schematic diagrams of the AAVS1 target site are shown, before (Figure 13A), after (Figure 13B), and associated with gRNA repair templates for the AAVS1 site for accurate repair using Dualase and guide RNA targeting AAVS1 (Figure 13C). [Figure 13C] Schematic diagrams of the AAVS1 target site are shown, before (Figure 13A), after (Figure 13B), and associated with gRNA repair templates for the AAVS1 site for accurate repair using Dualase and guide RNA targeting AAVS1 (Figure 13C). [Figure 14] Western blots of immunoprecipitated HA-tagged Dualase or SaCas9 from cells treated with Dualase and gRNA-RT or SaCas9 and gRNA-RT are shown. The presence of co-immunoprecipitated Rad52 or PolQ (polymerase theta; Polθ) was determined using antibodies specific to Rad52 (αRad52) or PolQ (αPolQ). Rad52 co-precipitates with Dualase and Cas9-treated cells, while PolQ co-precipitates only with Dualase-treated cells. [Figure 15A] The image shows a gel illustrating the restriction enzyme (RE) editing efficiency in HEK293 cells treated with purified TevCas9 mutant protein and gRNA-RT targeting AAVS1. The "Tev[WT]-dCas9" lane contains purified TevCas9 protein containing an inactivated D10A+H557A mutation but an active I-TevI domain. The "Tev[R27A]-dCas9" lane contains purified TevCas9 protein containing an inactivated D10A+H557A mutation and an I-TevI domain inactivated by the R27A mutation. Control lanes for saCas9[WT], reagent only, cells only, and cells lipofected with Tev-Cas9 3-in-1 mRNA control are also shown. [Figure 15B]The image shows a gel photograph illustrating the restriction enzyme (RE) editing efficiency in HEK293 cells treated with an mRNA version of SaCas9 (WT=wild-type, D10A+H557A=inactive, D10A=nickase, H557A=nickase) that lacks the I-TevI domain ("No Tev") but contains an I-TevI domain with R27A, V117F, K135R, and N140S mutations. [Figure 16A] A schematic diagram shows the insertion of eGFP in-frame into the CFTR coding sequence using rep-gRNA. eGFP is expressed from an endogenous gene promoter and isolated from the endogenous protein cleaved by a T2A-cleavable peptide sequence. [Figure 16B] Micrographs of cells lipofected with saCas9 or dualase mRNA and rep-gRNA encoding eGFP are shown. Phase contrast and GFP images are shown, including controls of cells only and rep-gRNA only. [Figure 16C] The images show gels where the insertion site was amplified using PCR, with larger insertion sequences shown as edited and sequences without insertions shown as unedited. [Figure 16D] A schematic diagram shows the insertion of eGFP using a guide RNA targeting the CFTR F508 site and a rep-gRNA encoding GFP that targets the CFTR G542 site, known as a "crosslinked rep-gRNA." The crosslinked rep-gRNA has 14 nucleotide homology to the upstream CFTR F508 site. The ~28 kilobase (kb) region between the F508 and G542 sites is removed and replaced with the GFP sequence encoded by the crosslinked rep-gRNA. [Figure 16E] Micrographs of cells lipofected with Dualase ("TevSaCas9") or an equimolar mixture of SaCas9, upstream guide RNA, and cross-linked rep-gRNA are shown. [Figure 16F]The images show gels with unedited target sites indicated by "~28kb region" and inserted sequences indicated by "Sequence removed and replaced". Also shown is an overview of the percentage of deep sequencing reads across the 5'- and 3'-junctions of inserted sequences where the repaired junction is correct ("Exact repair") or where a detected insertion or deletion exists ("Indel"). [Figure 17A] This shows a schematic RNA template DNA repair ("Rep editing") design and its predicted engagement with the TevSaCas9 protein at its target site. Figure 17A shows a schematic diagram of rep editing outlining the TevSaCas9 target site using individual DNA targeting components, domain structures, and rep-gRNA (also called "gRNA-RT"). The I-TevI linker zinc finger is indicated by a yellow dot. Cleavage by the I-TevI domain leaves a 2-nt 3' overhang complementary to the 3' end of the rep-gRNA, while cleavage by SaCas9 produces a blunt DNA end. [Figure 17B] This shows a schematic RNA template DNA repair ("Rep editing") design and its predicted TevSaCas9 protein engagement with the target site. Figure 17B shows the AAVS1 target site, the expected repair product, and the structure and interaction between rep-gRNA and the target site. [Figure 17C] This shows a schematic RNA template DNA repair ("Rep editing") design and its predicted engagement with the TevSaCas9 protein at its target site. Figure 17C shows the identification and distribution of the Tev CNNNG nuclease motif upstream of the SaCas9 binding site in the human genome. [Figure 17D]This shows a schematic RNA template DNA repair ("Rep-editing") design and its predicted engagement with the TevSaCas9 protein at its target site. Figure 17D shows a schematic diagram of the all-in-one construct with its individual components shown (not to scale) and predicted processing: MALAT, RNA stability element derived from MALAT1 non-coding RNA; HH, hammerhead ribozyme; T2A, self-splicing peptide linker; IRES, internal ribosome entry site. [Figure 17E] This shows a schematic RNA template DNA repair ("Rep editing") design and its predicted engagement with the TevSaCas9 protein at its target site. Figure 17E shows, on the left, the in vitro processing of the all-in-one mRNA transcript by HEK293 cell extract after incubation for the indicated time and separation on a 1% agarose gel. On the right, it shows eGFP activity in HEK293 cells transfected with the all-in-one construct. [Figure 18A] This shows rep-editing at the AAVS safe harbor site. Figure 18A shows a representative agarose gel of BglII digestion of the AAVS1 target site PCR-amplified from treated HEK293 cells; AAV, adeno-associated virus 2; pDNA, plasmid DNA; RNP, TevSaCas9 and rep-gRNA ribonucleoprotein particles; mRNA, all-in-one construct. [Figure 18B] This shows Rep-editing at the AAVS safe harbor site. Figure 18B shows an overview of the editing results at the AAVS1 site using the analysis method. [Figure 18C] This shows replica editing in the AAVS safe harbor area. Figure 18C shows an overview of the editing results by delivery method. The bar plots represent the mean of all replicas, and the whiskers represent the standard deviation from the mean. The dots represent individual replicas. [Figure 18D]Figure 18D shows Rep-editing at the AAVS safe harbor site. It plots apparent nucleotide substitutions from deep sequencing of AAVS1 PCR amplicons from edited cells (blue triangles, 6 replications) or simulated transfected cells (orange circles, 4 replications), indicating the fidelity of editing across the editing window. The dots represent the mean, and the whiskers represent the standard deviation from the mean. Regions targeted for editing and repair are enclosed by dashed vertical lines. [Figure 18E] This shows rep-editing at the AAVS safe harbor site. Figure 18E shows exemplary reads from deep sequencing of the AAVS1 site in TevSaCas9 / rep-gRNA edited cells, with the reference (wild-type) sequence shown on the top line, the repair product on the bottom, and the number of reads for each sequence on the right. Differences from the wild-type sequence are colored by nucleotides. [Figure 18F] This shows rep-editing at the AAVS safe harbor site. Figure 18F shows a schematic diagram of a reg-gRNA designed to create the HindIII site in AAVS1 by deleting 13 bp. [Figure 18G] This shows rep-editing at the AAVS safe harbor site. Figure 18G shows a plot of the percentage of deep sequencing reads with a length difference relative to the length of the unedited AAVS1 target site from edited cells or mock-treated cells. [Figure 18H] Rep-editing at the AAVS safe harbor site is shown. Figure 18H shows the alignment of sequencing reads with gaps indicated by dashed lines (-) and read counts shown on the right, where >48% of reads show the intended 13bp deletion. [Figure 19A] A schematic diagram of the rep-gRNA used to test base pairing between the 3' end of the rep-gRNA and the I-TevI cleavage overhang is shown. All rep-gRNAs targeted the AAVS1 safe harbor site. [Figure 19B] This shows the effect of rep-gRNAs with 3' and 5' duplications of different lengths on editing efficiency. A bar graph shows the repair site to unrepair site ratio, determined by quantitative PCR of the AAVS1 target site in treated cells. The bar graph represents the mean of two biological copies, and the error bars represent the standard deviation from the mean. Dark green bars represent cells treated with TevSaCas9 / rep-gRNA, and light green bars represent cells treated with protective rep-gRNA, in which a single-stranded DNA oligonucleotide complementary to the overhang was co-delivered with the rep-gRNA. [Figure 19C] The results of deep sequencing of TevCas9 / rep-gRNA-treated cells with various overhang lengths are shown, with nucleotide differences compared to the unmodified (WT) sequence highlighted in color. [Figure 19D] This shows the effect of the last two nucleotides of rep-gRNA on editing. A representative gel of BglII digest of AAVS1 target site amplicons from treated HEK293 cells is shown. [Figure 19E] This shows the effect of mismatches between rep-gRNA and the crRNA portion of rep-gRNA on repair at the AAVS1 site, with rep-gRNA precisely matching I-TevI overhangs (CC) or wobble base pairs (UC). Nucleotide mismatches in crRNA relative to the target site are indicated by their position on the gel image. [Figure 20A] A schematic diagram of the method used to identify off-target sites in TevSaCas9 is shown. [Figure 20B] This plot shows the distribution of CNNNG motifs upstream of AAVS1 off-target sites exhibiting mismatches to the gRNA portion. It highlights CNNNGs that can support repair using the AAVS1 rep-gRNA 3'-GG terminus. [Figure 20C] This shows the alignment of the off-target site to the on-target AAVS1 site. The same nucleotides are colored, and the CNNNG motif in the upstream region is underlined. [Figure 20D] The edit percentages are shown as determined by CRISPRaltRations analysis. Each point represents an individual experiment. ns = calculated by a paired t-test of treatments and is not statistically significant. [Figure 20E] The plot shows the edit percentages at on- and off-target sites, determined by CRISPECTOR analysis of HEK293 treated with TevSaCas9 / AAVS1 rep-gRNA by AAV transduction (yellow) or mRNA lipofection (blue). The plot is separated into indels mapped near the predicted SaCas9 and Tev cleavage sites for each target site. The bars represent the mean of three replicates, and the whiskers represent the 95% confidence interval from the CRISPECTOR calculated edit percentage. [Figure 21A] This paper demonstrates the identification of repair proteins involved in rep editing using small molecule inhibitors. Agarose gels showing editing at the AAVS1 target site in treated HEK293 cells are shown. In the right figure, HEK293 cells were treated with B02, SCR7, D-103, or ART558 inhibitors simultaneously with transduction with AAV2, which capsids TevSaCas9 and rep-gRNA. Increasing inhibitor concentrations are indicated by blue triangles. The AAVS1 target site was PCR amplified and digested with BglII; BglII digestion indicates a successful repair event. The left figure shows editing determined by BglII digestion of the amplified AAVS1 target site in HEK293 cells transduced with AAV2-TevSaCas9-rep-gRNA and different co-delivered repair sequences as shown, along with small molecule inhibitors. [Figure 21B] This image shows replication co-immunoprecipitation with anti-HA antibody, followed by Western blotting with anti-Polθ or anti-Rad52 antibody, of extracts from HEK293 cells transduced with AAV2-TevSaCas9-rep-gRNA or AAV2-SaCas9-rep-gRNA. The cell size (kd) is shown on the left side of the gel image. [Figure 21C]A model for rep editing is shown. Cleavage by TevSaCas9-rep-gRNA recruits Rad52, which can bind to RNA or dsRNA, or mediate RNA:DNA strand exchange at the I-TevI cleavage site. Formation of a hybrid RNA:DNA structure with a 3'-OH recruits Polθ. Degradation of repair intermediates by a fill-in gap repair process involving DNA ligases and an unidentified DNA polymerase. [Figure 22A] This diagram shows schematic diagrams of rep-gRNA designs to target three mutations in the CFTR gene: the 3nt deletion DF508, the G>A transition mutation G542X, and the G>C transversion mutation W1282X. The expected edit products are shown below, with silent substitutions indicated by a hash symbol (#) and corrective edits by an x-marked circle. [Figure 22B] Representative gels for SspI (DF508 or G542X) or HindIII (W1282X) restriction digestion analysis of PCR amplicons from 16HBEge cells are shown. Cells treated with TevSaCas9 / rep-gRNA are indicated with (+), and cells treated with transfection reagents only are indicated with (-). [Figure 23A] This shows the alignment of dominant reads from deep sequencing of PCR products amplified from whole 16HBE CFTR G542X cells transfected with TevSaCas9 / rep-gRNA (cells were not enriched or selected before analysis). [Figure 23B] A schematic diagram shows the formulation of an all-in-one messenger RNA (mRNA) expressing TevCas9 in lipid nanoparticles and a conditionally cleavable repair template guide RNA (rep-gRNA) for intratracheal delivery to humanized CFTRG542X mice. The TevCas9 mRNA also contains a cleavable GFP tag that allows for quantification of uptake into lung cells. [Figure 23C] The results of fluorescence-activated cell sorting (FACS) of harvested lung tissue from humanized CFTRG542X mice treated with TevCas9 all-in-one mRNA and GFP-expressing mRNA control are shown. In the TevCas9-treated mice, approximately 31% of the cells shown in the boxed area were counted as GFP-positive, while in the GFP-only control mice, approximately 25% of the cells shown in the boxed area were counted as GFP-positive. [Figure 23D] The results of PCR detecting repair sequences in lung tissue collected from TevCas9-treated mice and GFP-controlled mice are shown. PCR is calibrated with lung epithelial cell samples that have known repairs ("corrected cell control"). Repair products were detected only in lung tissue from TevCas9 all-in-one mRNA-treated mice and not in GFP-controlled mice. [Figure 24A] A schematic diagram shows the alignment of dominant reads from deep sequencing of PCR products amplified in GM11423 patient liver fibroblasts transfected with rep-gRNA design for targeting the human SERPINA1 gene at amino acid position E342K, and TevSaCas9 / rep-gRNA (cells were not enriched or selected before analysis). [Figure 24B] The administration and sampling schedule for humanized SERPINA1 E342K mice treated with lipid nanoparticle-encapsulated all-in-one TevSaCas9 / re-gRNA to correct mutations is shown. [Figure 24C] The results of enzyme-linked immunosorbent assay (ELISA) for human alpha-1-antitrypsin (A1AT) from pre- and post-administration serum samples from treated and naive mice at the indicated time points are shown. [Figure 24D] The results of serum biochemical analyses for aspartate transaminase (AST) and alanine transaminase (ALT) from pre- and post-administration serum samples from treated and naive mice at the indicated time points are shown. [Figure 24E] The results of reverse transcriptase quantitative PCR (RT-qPCR) against the Cas9 domain of TevCas9, used to confirm TevSaCas9 expression in liver samples from naive and treated mice, are shown. [Figure 25A] A schematic diagram is shown illustrating the removal and substitution of 74 nucleotides between amino acids R553 and G542 in human CFTR using guide RNA and cross-linking rep-gRNA strategies. [Figure 25B] The image shows gel photographs of PCR amplicons of CFTR target sites from 16HBEge cells treated with MfeI and SspI, as well as SaCas9+rep-gRNA and TevSaCas9+rep-gRNA digested with restriction enzymes. The presence of digestion products indicates repair. [Figure 26A]This disclosure shows exemplary combinations and orientations of conditionally cleavable nucleotide sequences. Figure 26A shows a schematic diagram of exemplary constructs that fit as both DNA and RNA constructs, using human transfer RNA (tRNA) or transfer RNA (tRNA') having a trans-acting ribozyme encoded in an anticodon loop. The DNA version of the encoded construct shows the cleavage sites after translation to mRNA, or the RNA version of the construct transcribed in vitro is introduced into cells, and the RNA cleavage sites are indicated by filled triangles (▲). Cleavage sites of trans-acting elements are indicated by lines ending in arrows. The messenger RNA version of the construct has a 5'-UTR, and the DNA version of the construct has a promoter sequence. Table 10 lists exemplary human transfer tRNAs that undergo tRNA maturation when expressed or transfected into cells. Some exemplary versions contain MALAT stabilization sequences to stabilize the I-TevI and Cas9 coding sequences after mRNA cleavage. [Figure 26B] This disclosure shows exemplary combinations and orientations of conditionally cleavable nucleotide sequences. Figures 26B and 26C show schematic diagrams of exemplary constructs that are suitable only as DNA constructs, such as using an adeno-associated virus vector. Ribozymes can only encode as DNA because they self-cleave upon transcription to RNA. Ribozymes cleave only on one side, and the orientation of the ribozyme is indicated by the direction of the arrow. Transfer RNA or trans-acting transfer RNA can be combined with the ribozyme to cleave the cleavage guide RNA. Figure 26B shows an exemplary construct without a MALAT-stabilizing sequence, and Figure 26C shows an exemplary schematic diagram of a construct with a MALAT-stabilizing sequence. In addition to hammerhead (HH) ribozymes and hepatitis delta virus (HDV) ribozymes, Table 9 lists other exemplary ribozymes that have activity in human cells. [Figure 26C]This disclosure shows exemplary combinations and orientations of conditionally cleavable nucleotide sequences. Figures 26B and 26C show schematic diagrams of exemplary constructs that are suitable only as DNA constructs, such as using an adeno-associated virus vector. Ribozymes can only encode as DNA because they self-cleave upon transcription to RNA. Ribozymes cleave only on one side, and the orientation of the ribozyme is indicated by the direction of the arrow. Transfer RNA or trans-acting transfer RNA can be combined with the ribozyme to cleave the cleavage guide RNA. Figure 26B shows an exemplary construct without a MALAT-stabilizing sequence, and Figure 26C shows an exemplary schematic diagram of a construct with a MALAT-stabilizing sequence. In addition to hammerhead (HH) ribozymes and hepatitis delta virus (HDV) ribozymes, Table 9 lists other exemplary ribozymes that have activity in human cells. [Figure 27A] This demonstrates the development of a novel AAV construct with conditionally cleavable elements. Figure 27A shows a schematic diagram of the original DNA construct and the conditionally cleavable features redesigned in the second DNA construct. mRNA resulting from the transcription of the construct to release two guide RNAs that target the DMPK CAG repeat sequence is also shown. [Figure 27B] This demonstrates the development of a novel AAV construct with conditionally cleavable elements. Figure 27B shows a schematic diagram of the experimental workflow for testing AAV2, which capsidizes the original construct ("AAV6") and the redesigned construct ("AAV67") in DMPK CAG repeat fibroblasts. The duration of transduction by the construct before cell harvesting for analysis of genomic DNA, and polymerase chain reaction (PCR) of the genomic DNA used to analyze the presence or absence of CAG repeat sequences are shown. [Figure 27C]This demonstrates the development of novel AAV constructs with conditionally cleavable elements. Figure 27C shows agarose gels of PCR products of genomic DNA extracted from cells transduced at 5,000, 10,000, or 25,000 multiplicities of infection (MOI) of AAV2 at the indicated time. Lanes marked "L" contain DNA ladders, lanes marked "E" are empty, and lanes marked "C" contain PCR products from genomic DNA of untreated cells. [Figure 27D] This demonstrates the development of a novel AAV construct with conditionally cleavable elements. Figure 27D shows the results of quantifying the ratio of PCR products with more than 1,000 base pairs ("high molecular weight / MW") to PCR products with less than 1,000 base pairs ("low molecular weight / LW") by agarose gel densitometry. [Figure 28A] Figure 28A shows the alignment of dominant reads from deep sequencing of PCR products amplified from HEK293 cells transfected with the Tev[V117F+K135R+N140S]-saCas9[D10E] variant and AAVS1 targeted repair template guide RNA (cells were not enriched or selected before analysis). Figure 28A provides an overview of the percentages of unmodified, insertion / deletion ("NHEJ"), accurately repaired ("HDR"), incompletely repaired ("incomplete HDR"), and ambiguous reads. [Figure 28B] Figure 28B shows the alignment of dominant reads from deep sequencing of PCR products amplified from HEK293 cells transfected with the Tev[V117F+K135R+N140S]-saCas9[D10E] variant and AAVS1 targeted repair template guide RNA (cells were not enriched or selected before analysis). Figure 28B shows the alignment of the most dominant sequencing reads from the transfected cells, along with the number of reads (#reads) and the percentage of reads. Expected repair products are indicated by Asterix(*). [Figure 29A]This shows cell editing using a chimeric nuclease containing I-TevI and erCas12a in the B2M gene. Figure 29A shows a schematic diagram of the chimeric nuclease characteristics highlighting the I-TevI nuclease domain, linker domain, and erCas12a nuclease domain ("Tev-erCas12a"). [Figure 29B] This shows cell editing using chimeric nucleases containing I-TevI and erCas12a in the B2M gene. Figure 29B highlights the DNA target sites in the B2M gene, including the erCas12 guide RNA binding site ("B2M6 crRNA"), the erCas12 protospacer flanking motif ("PAM"), the DNA spacer sequence between the erCas12a site and the I-TevI site ("B2M6 Tev Spacer"), and the I-TevI target site ("B2M6 I-TevI site"). The I-TevI and erCas12a sense and antisense cleavage sites are also shown. [Figure 29C] This shows cell editing using chimeric nucleases including I-TevI and erCas12a in the B2M gene. Figure 29C shows a gel photograph showing the editing efficiency of HEK293 with T7 endonuclease I (T7E1) using plasmid DNA encoding Tev-erCas12a, erCas12a, TevSaCas9, and saCas9 targeting the B2M gene. PCR amplicons from the B2M target site ("Sub") and T7E1 digested product ("Edited") are shown. The editing percentage is shown below the gel photograph. Detailed description of the invention
[0057] This specification provides, in particular (inter alia), compositions and methods for chimeric nucleases and chimeric nuclease systems, as well as nucleic acids encoding chimeric nucleases and chimeric nuclease systems.
[0058] This disclosure is partly based on the discovery that the chimeric nucleases and nuclease systems of this disclosure can be expressed in a cell from a single nucleic acid using a single promoter. The chimeric nuclease systems and the nucleic acids encoding the chimeric nuclease systems of this disclosure can be used to edit the genome of a cell. The nucleic acids encoding the chimeric nuclease systems of this disclosure can be transcribed in a cell into a single mRNA containing the components of the chimeric nuclease system. The mRNA can be further processed in the cell into its individual components, e.g., Cas9 nuclease, guide RNA, and donor polynucleotide. The chimeric nuclease systems of this disclosure may comprise one or more chimeric nucleases and one or more guide polynucleotides. In addition, the chimeric nuclease systems may comprise donor polynucleotides. The nucleic acids encoding the chimeric nuclease systems of this disclosure may comprise a single promoter sequence, a sequence encoding the chimeric nuclease, and a sequence encoding the guide polynucleotide.
[0059] In addition, nucleic acids may contain nucleic acid sequences that help process a single mRNA into smaller fragments and / or obviate RNA stabilizing sequences that eliminate the need for stabilizing poly A. The transcribed mRNA sequence may contain additional sequences such as chimeric nucleases, one or more guide RNAs, and nucleic acid sequences encoding one or more donor polynucleotides, RNA stabilizing sequences, tRNAs, ribozymes, and ribozyme cleavage sites.
[0060] This disclosure is also in part based on the discovery that chimeric nucleases and chimeric nuclease systems efficiently and accurately replace sequences in the cellular genome using RNA template DNA repair in mammalian cells in the absence of exogenous reverse transcriptase.
[0061] An advantage of the nucleic acids of this disclosure is that they are smaller than multi-promoter nucleic acid constructs and can be packaged into viral genomes, such as the AAV genome. A further advantage of the nucleic acid designs provided herein is that the nucleic acids can be designed to ensure the precise 5' and 3' ends of the nucleic acid components after transcription and further processing. Another advantage of the chimeric nucleases and chimeric nuclease systems and nucleic acids of this disclosure is that the system can be used for RNA-mediated repair without the delivery of exogenous reverse transcriptase to cells. Due to the small size of the nucleic acids encoding the chimeric nucleases and chimeric nuclease systems of this disclosure, the nucleic acids can be packaged into viral vectors for efficient cell delivery of these all-in-one editing systems.
[0062] Before describing embodiments of this disclosure, it should be understood that such embodiments are provided only as examples, and that various alternatives to the embodiments of this disclosure described herein may be adopted in practice of the invention. Numerous modifications, alterations, and substitutions will come to mind for those skilled in the art without departing from the invention.
[0063] Unless otherwise defined herein, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art in which this disclosure pertains. Various scientific dictionaries containing the terms included herein are well known and available to those skilled in the art. Any methods and materials similar to or equivalent to those described herein may be used in the practice or testing of this disclosure, but several preferred methods and materials are described. Thus, the terms defined below are more fully explained by referring to this specification as a whole.
[0064] definition All terms are intended to be understood as they are understood by those skilled in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as they are commonly understood by those skilled in the art in which this disclosure relates.
[0065] The following definitions are supplementary to those skilled in the art, apply to this application, and do not belong to any related or unrelated cases, such as any jointly owned patents or applications. Any methods and materials similar or equivalent to those described herein may be used in carrying out the tests of this disclosure, but preferred materials and methods are described herein. Accordingly, the terms used herein are for the sole purpose of describing specific embodiments and are not intended to limit them.
[0066] In this application, the use of the singular form includes the plural form unless otherwise specified. Note that, as used herein, the singular forms "a," "an," and "the" refer to multiple objects unless the context clearly indicates otherwise.
[0067] In this application, the use of “or” means “and / or” unless otherwise specified. As used herein, the terms “and / or” and “any combination thereof,” as well as their grammatical equivalents, can be used interchangeably. These terms can convey that any combination is specifically intended. For illustrative purposes only, the following phrases “A, B, and / or C” or “A, B, C, or any combination thereof” may mean “A individually; B individually; C individually; A and B; B and C; A and C; and A, B, and C.” The term “or” can be used disjunctively or disjunctively unless the context specifically refers to a disjunctive use.
[0068] Furthermore, the use of the term "including," as well as other forms such as "include," "includes," and "included," is not limited.
[0069] References to “some embodiments,” “an embodiment,” “one embodiment,” or “other embodiments” in this specification mean that certain features, structures, or characteristics described in relation to an embodiment are included in at least some embodiments of this disclosure, but not necessarily in all embodiments.
[0070] As used herein and in the claims, the terms “comprising” (including derivatives of “comprising,” such as “comprise” and “comprises”), “having” (including derivatives of “having,” such as “have” and “has”), “including” (including derivatives of “including,” such as “includes” and “include”), or “containing” (including derivatives of “containing,” such as “contains” and “contain”) are inclusive or non-exclusive and do not exclude additional elements or steps of methods not enumerated. Any embodiment discussed herein may be implemented with respect to any method or composition of the Disclosure, and vice versa. Furthermore, compositions of the Disclosure may be used to achieve the methods of the Disclosure.
[0071] The terms “about” or “approximately” mean within an acceptable margin of error of a particular value as determined by those skilled in the art, which in part depends on how that value is measured or determined, i.e., the limitations of the measuring system. For example, “about” can mean within or above one standard deviation, according to convention in the art. Alternatively, “about” can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. In another example, the quantity “about 10” includes 10 and any quantities from 9 to 11. In yet another example, the term “about” in relation to a reference numerical value can also include a range of values plus or minus 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of that value. Alternatively, particularly in relation to biological systems or processes, the term “about” can mean within one order of magnitude of the value, preferably within five times, and more preferably within two times. Where a specific value is stated in this application and claims, unless otherwise stated, the term “about” should be assumed to mean within the allowable margin of error of that specific value.
[0072] The term “at least” followed by a number is used herein to indicate the beginning of a range that starts with that number (which may be an upper limit range or an upper limit range, depending on the variable being defined). For example, “at least 1” means 1 or greater than 1.
[0073] The term “at most” followed by a number is used herein to indicate the end of a range ending at that number (which may be a range with a lower limit of 1 or 0, or a range with no lower limit, depending on the variable being defined). For example, “at most 4” means 4 or less, and “at most 40%” means 40% or less. Where a range is given herein as “(a first number) ~ (a second number)” or “(a first number) - (a second number)”, this means a range where the lower limit is the first number and the upper limit is the second number. For example, 25 ~ 100 mm means a range with a lower limit of 25 mm and an upper limit of 100 mm.
[0074] Where used herein, the terms “nucleic acid,” “nucleic acid molecule,” “nucleic acid oligomer,” “oligonucleotide,” “nucleic acid sequence,” “nucleic acid sequence,” and “polynucleotide” are interchangeable and are intended to include, but are not limited to, polymeric forms of covalently bonded nucleotides, which may have varying lengths of deoxyribonucleotides or ribonucleotides, or their analogs, derivatives, or modifications. Different polynucleotides may have different three-dimensional structures and may perform a variety of known or unknown functions. Non-exclusive examples of polynucleotides include genes, gene fragments, exons, introns, intergenetic DNA (including, but not limited to, heterochromatin DNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA sequences, isolated RNA sequences, nucleic acid probes, and primers. Polynucleotides useful in the methods of this disclosure may include natural nucleic acid sequences and their variants, artificial nucleic acid sequences, or combinations of such sequences.
[0075] In the context of two or more nucleic acid or polypeptide sequences, the terms “identical” or “percent identity” refer to two or more sequences or subsequences that, when compared and aligned for maximum correspondence, are identical or have a specific proportion of the same amino acid residues or nucleotides. Methods for aligning sequences for comparison are well known in the art. Once aligned, the number of matches is determined by counting the number of positions where identical nucleotides or amino acid residues exist in both sequences. The sequence identity percentage is determined by dividing the number of matches in the alignment by the length of the reference sequence and then multiplying the resulting value by 100 (multiplying). For example, when aligned with a test sequence having 1554 amino acids, a peptide sequence with 1166 matches is 75.0 percent identical to the test sequence (1166 ÷ 1554 * 100 = 75.0). When these terms are used herein, gaps in the alignment do not reduce the sequence identity percentage. Unless otherwise specified, the optimal alignment of sequences for comparison is performed using the global alignment algorithm (Needleman and Wunsch, Mol. Biol. 48:443 (1970)) implemented by EMBOSS Needle (ebi.ac.uk / Tools / psa / emboss_needle / on the World Wide Web) (Madeira et al. Nucleic Acids Res. 50(W1):W276-W279 (2022)).In embodiments, Devereux, et al., Nucleic Acids Res.12:387-95 (1984), Altschul et al., J. Mol. Biol. 215:403-10 (1990) (BLAST), Carrillo and Lipman Siam J. Appl. Biology (Lesk, AM, ed., 1989), Biocomputing Informatics and Genome Projects, (Smith, DW, ed., 1993), Computer Analysis of Sequence Data, Part I, (Griffin and Griffin, eds., 1994), Sequence Analysis in Molecular Biology (von Heijne, 2012), Sequence Analysis Primer (Gribskov and Other alignment methods may be used, including but not limited to those described in Devereux, J., eds. 1993. Sequence identity is calculated using an implementation of the Needleman-Wunsch algorithm provided by the National Library of Medicine (blast.ncbi.nlm.nih.gov / Blast.cgi?PAGE_TYPE=BlastSearch&BLAST_SPEC=GlobalAln on the World Wide Web).
[0076] For example, sequence identity can be determined by standard methods commonly used to compare the similarity of two polypeptide or polynucleotide sequences. Using computer programs such as EMBOSS Needle or BLAST, two polypeptide or polynucleotide sequences are aligned for optimal matching of their respective residues (along the entire length of one or both sequences, or along a given portion of one or both sequences). The program provides a default opening penalty and a default gap penalty, as well as a scoring matrix such as PAM250 (a standard scoring matrix; see Dayhoff et al., in Atlas of Protein Sequence and Structure, vol. 5, supp. 3 (1978)) that can be used in conjunction with the computer program.
[0077] "Binding" refers to joining by covalent or non-covalent bonds. Non-covalent bonds include those formed by van der Waals forces, hydrogen bonds, ionic bonds, encapsulation or physical encapsulation, absorption, adsorption, and / or other intermolecular forces. Bonding can be achieved by any useful means, such as enzymatic bonding (e.g., enzymatic ligation) or chemical bonding (e.g., chemical ligation).
[0078] As used herein, the term "expression" generally refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (to mRNA or other RNA transcripts, etc.), or the process by which the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and the polypeptides they encode may collectively be referred to as "gene products." When polynucleotides originate from genomic DNA, expression may include the splicing of mRNA in eukaryotic cells.
[0079] As used herein, “operably linked,” “operable linkage,” “operatively linked,” or their grammatical equivalents generally refer to the juxtaposition of genetic elements, such as promoters, enhancers, and polyadenylation sequences (juxtapositions), which are related in a way that enables them to function as expected. For example, a regulatory element, which may include a promoter or enhancer sequence, is operably linked to a coding region if the regulatory element helps initiate transcription of the coding sequence. Intervening residues may exist between the regulatory element and the coding region as long as this functional relationship is maintained.
[0080] As used herein, “vector” generally refers to a macromolecule or association of macromolecules that contains or associates with polynucleotides and can be used to mediate the delivery of polynucleotides to cells. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. Vectors generally contain gene elements, such as regulatory elements, that are operablely linked to a gene in order to promote gene expression at a target.
[0081] As used herein, “an expression cassette” and “a nucleic acid cassette” are used interchangeably to generally refer to a combination of nucleic acid sequences or elements that are expressed together or operably linked for expression. In some cases, an expression cassette refers to a combination of regulatory elements and the gene(s) to which they are operably linked for expression.
[0082] As used herein, the term “engineered” generally indicates that an object has been modified by human intervention. To give a non-limiting example, nucleic acids may be modified by altering their sequence to one that does not exist in nature; nucleic acids may be modified by ligating them to nucleic acids that do not naturally associate such that the ligated product has a function not present in the original nucleic acid; engineered nucleic acids may be synthesized in vitro using sequences that do not exist in nature; proteins may be modified by altering their amino acid sequence to one that does not exist in nature; engineered proteins may acquire new functions or properties. An “engineered” system includes at least one engineered component.
[0083] This disclosure includes variants of any of the endonucleases described herein having one or more conserved amino acid substitutions. Such conserved substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids having similar hydrophobicity, polarity, and R-chain length with each other. In addition, or alternatively, by comparing aligned sequences of homologous proteins from different species, conserved substitutions can be identified by locating the interspecies mutated amino acid residues (e.g., non-conserved residues) without altering the fundamental function of the encoded protein. Such conservatively substituted mutants may include mutants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% identity to any one of the endonuclease protein sequences described herein. In some embodiments, such conservatively substituted mutants are functional mutants. Such functional mutants may include sequences having substitutions such that the activity of one or more important active site residues or guide RNA binding residues of the endonuclease is not disrupted.
[0084] Conservative substitution tables, which provide functionally similar amino acids, are available from various references (e.g., Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd edition (December 1993)). The following eight groups each contain amino acids that are conserved substitutions for each other: 1) Alanine (A), Glycine (G); 2) Aspartic acid (D), glutamic acid (E); 3) Asparagine (N), glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), leucine (L), methionine (M), valine (V); 6) Phenylalanine (F), tyrosine (Y), tryptophan (W); 7) Serine (S), threonine (T); and 8) Cysteine (C), Methionine (M).
[0085] The "position" of an amino acid or nucleotide base is indicated by a number that sequentially identifies each amino acid (or nucleotide base) in the reference sequence based on its position relative to the N-terminus (or 5'-terminus). Due to deletions, insertions, cleavages, fusions, etc., which must be considered when determining the optimal alignment, the number of amino acid residues in a test sequence, determined simply by counting from the N-terminus, is not necessarily the same as the number of corresponding positions in the reference sequence. For example, if a mutant has a deletion relative to the aligned reference sequence, the amino acid corresponding to the position in the reference sequence at the deletion site is not present in the mutant. If there is an insertion in the aligned reference sequence, that insertion does not correspond to a numbered amino acid position in the reference sequence. In the case of deletions or fusions, there may be stretches of amino acids in either the reference sequence or the aligned sequence that do not correspond to any amino acid in the corresponding sequence.
[0086] When used in the context of numbering a given amino acid or polynucleotide sequence, the terms "numbered with reference to" or "corresponding to" refer to the numbering of residues in a particular reference sequence when a given amino acid or polynucleotide sequence is compared to that reference sequence.
[0087] The terms “target site” or “target sites,” as used herein, refer, in relation to the present invention, to a location within a gene that, when targeted by a nuclease, is bound to and / or cleaved by the nuclease. Target sites can encompass a number of nucleotides, such as cleavage sites of I-TevI or CRISPR / Cas nucleases, or DNA binding sites of I-TevI nucleases or CRISPR / Cas guide RNAs. Target sites can be located within genes or intergenetic regions in any of the various cell types and organisms.
[0088] When used herein, the terms “target,” “targets,” “to target,” or “targeting” refer, for example, to direct or orient a nuclease to a specific selected DNA sequence using a selected or manipulated DNA-binding domain or guide RNA, with respect to the invention of this application.
[0089] As used herein, the term "viral vector" refers to a tool typically used to deliver genetic material to cells. This process can be carried out in vivo or in vitro. Examples of viral vectors include AAV vectors, lentiviral vectors, and adenovirus vectors.
[0090] The term "simultaneously," as used herein, refers to the administration of a CRISPR / Cas9 complex containing multiple gRNAs simultaneously or substantially simultaneously. It will be understood that certain procedures may be carried out in steps that are sequential but separated by small intervals. For example, administering donor DNA within approximately one hour following electroporation of cells.
[0091] For clarity, it should be understood that certain features of the Disclosure described in the context of separate embodiments may also be provided in combination in a single embodiment. Conversely, for brevity, various features of the Disclosure described in the context of a single embodiment may also be provided separately or in any suitable partial combination. All combinations of embodiments relating to the Disclosure are specifically encompassed by the Disclosure and are disclosed herein as if each and every combination were individually and expressly disclosed. Furthermore, all partial combinations of various embodiments and their elements are also specifically encompassed by the Disclosure and are disclosed herein as if each and every such partial combination were individually and expressly disclosed herein.
[0092] Chimeranucleases and chimeranucleases This specification provides, in particular, chimeric nucleases and chimeric nuclease systems comprising two or more nucleases and guide RNAs. Chimeric nucleases and chimeric nuclease systems can be used to introduce modifications, such as insertions, deletions, or mutations, into the genome of a cell.
[0093] In some embodiments, the chimeric nuclease system comprises a CRISPR Cas nuclease domain, a guide RNA, and a GIY-YIG nuclease domain. In some embodiments, the chimeric nuclease comprises a CRISPR Cas nuclease domain and a GIY-YIG nuclease domain. In some embodiments, the chimeric nuclease comprises a fusion protein of a CRISPR Cas nuclease domain and a GIY-YIG nuclease domain. In some embodiments, the CRISPR Cas nuclease domain is located at the N-terminus or C-terminus (e.g., N-terminus) of the GIY-YIG nuclease domain. In some embodiments, the chimeric nuclease comprises a linker between the CRISPR Cas nuclease domain and the GIY-YIG nuclease domain.
[0094] CRISPR Cas nuclease CRISPR-Cas nucleases are programmable RNA-directed nucleases described as functioning as an adaptive immune system in microorganisms. Nuclease targeting of specific target nucleic acid sequences generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (target seed) and a crRNA guide, and (ii) the presence of a protospacer-adjacent motif (PAM) sequence within a defined neighborhood of the target seed. CRISPR-Cas systems are typically organized into two classes, five types, and 16 subtypes based on common functional characteristics and evolutionary similarities. The structure, nomenclature, and classification of CRISPR systems are outlined in Makarova st al, Evolution and classification of the CRISPR-Cas systems. Nature Reviews Microbiology. 2011 June;9(6):467-477.
[0095] Class I CRISPR Cas systems have large multi-subunit effector complexes and include types I, III, and IV. Type I CRISPR systems include a multi-protein complex called Cascade (a CRISPR-associated complex for antiviral defense) composed of subunits CasA, B, C, D, and E, as well as crRNA. The Cascade-crRNA complex recognizes target nucleic acids by hybridization with the crRNA. The bound nucleoprotein complex recruits Cas3 helicase / nuclease to facilitate cleavage of the target nucleic acid. Type III CRISPR systems include the RAMP superfamily of endoribonucleases (e.g., Cas6) that cleave pre-crRNA arrays with the help of one or more CRISPR polymerase-like proteins. The type IV CRISPR-Cas system has an effector complex consisting of two genes for RAMP proteins from the highly reduced large subunit nuclease (csf1), Cas5 (csf3), and Cas7 (csf2) groups, as well as, in some cases, a gene for a predicted small subunit. Such systems are typically found on endogenous plasmids.
[0096] Class 2 CRISPR-Cas systems generally possess a single polypeptide multidomain nuclease effector and include types II, V, and VI. Type II CRISPR-Cas systems typically include Cas9 nuclease, crRNA, and transactivating CRISPR RNA (tracrRNA). The tracrRNA hybridizes to crRNA repeats. The tracrRNA / crRNA complex can associate with a nuclease, such as Cas9. The crRNA-tracrRNA-Cas9 complex recognizes the target nucleic acid by hybridization with the crRNA. Hybridization of crRNA to the target nucleic acid activates the Cas9 nuclease for target nucleic acid cleavage. The V-type CRISPR system includes a different set of Cas-like genes, including Cas12 nucleases such as Cas12a (Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, and Cas14. The VI-type CRISPR Cas system is an RNA-induced RNA endonuclease.
[0097] In some embodiments, the chimeric nuclease comprises a Cas9 CRISPR Cas nuclease. Cas9 is a programmable class II CRISPR Cas nuclease that forms complexes with congeneral crRNAs and tracrRNAs. In some embodiments, the chimeric nuclease system comprises a Cas9 nuclease, a crRNA, and a tracrRNA. In some embodiments, the chimeric nuclease comprises a nuclear localization signal (NLS). In some embodiments, the chimeric nuclease comprises two C-terminal nucleoplasmin NLSs isolated by a human influenza hemagglutinin (HA) sequence, or two SV40 NLSs isolated by an HA sequence, followed by two SV40 NLSs.
[0098] In some embodiments, the Cas9 nuclease contains a conserved amino acid substitution. In some embodiments, the Cas9 nuclease contains a mutation in the catalytic domain. In some embodiments, the Cas9 nuclease contains a mutation (nCas9) that results in nickase activity. In some embodiments, the Cas9 nuclease contains a mutation that results in a Cas9 that has lost catalytic activity (catalytically dead Cas9) (dCas9, inactivated Cas9). In some embodiments, the Cas nuclease is the Cas9 domain.
[0099] In some embodiments, the Cas9 nuclease is Staphylococcus aureus, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus desulforudis, Clostridium botulinum, Clostridium difficile, Fine goldia magna, Natranaerobius thermophilus, Pelotomaculum the rmopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp.It may be derived from or originate from Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, or Acaryochloris marina.
[0100] In some embodiments, Cas9 nuclease is used in studies of F. novicida, T. denticola, Campylobacter jejuni, Alicyclobacillus acidoterrestris, Prevotella and Francisella Acidaminococcus sp.BV3L6, Eubacterium rectale, SpCas12f1 (497 aa), AsCas12f1 (422 aa), S. lugdunensis (Slu), S. hyicus (Shy), S. microti (Smi), S. pasteuri (Spa), Seawater microbial communities, Freshwater microbial communities from Lake Mendota, Plantomycetes, Agricultural soil microbial communities from Utah to study Nitrogen management - Steer compost 2015, Thermophilic microbial communities from the Joint Bioenergy Institute, California, USA of rice / straw / compost enrichment - Selected from eDNA_2, Eggerthella_sp._YY7918, Finegoldia_magna_ATCC_29328, Lactobacillus_rhamnosus_LOCK900, Nitratifractor salsuginis, Streptococcus_gordonii_str._Challis_substr._CH1, Tissierellia bacterium KA00581, Turicibacter sp., or Cas9 nucleases derived from manipulated SpyCas9.
[0101] In some embodiments, the Cas9 nuclease is derived from Staphylococcus aureus, Streptococcus pyogenes, Neisseria meningitidis, Campylobacter jejuni, Streptococcus pasteurianus, Clostridium cellulolyticum, or Geobacillus thermodenitrificans T1.
[0102] In some embodiments, the Cas9 domain is derived from Staphylococcus aureus (SaCas9). In some embodiments, the Cas9 domain is derived from Streptococcus pyogenes (spCas9). In some embodiments, the Cas9 domain is derived from Neisseria meningitidis (NmCas9). In some embodiments, the Cas9 domain is derived from Campylobacter jejuni (CjCas9). In some embodiments, the Cas9 domain is derived from Streptococcus pasteurianus (SpCas9), and in some embodiments, the Cas9 domain is derived from Clostridium cellulolyticum (CcCas9). In some embodiments, the Cas9 domain is derived from Geobacillus thermodenitrificans T1 (GtCas9).
[0103] Exemplary Cas9 nucleases and related PAM sequences are shown in Table A.
[0104] [Table 1]
[0105] In some embodiments, the Staphylococcus aureus Cas9 domains are D10, H557, N580, H840, D1135, R1335, T1337, T267, L325, V327, D333, A336, I341, E345, D348, K352, S360, T368, N369, N371, S372, E373, K386, N393, H408, N410, I414, A415, T438, Y467, N471, D485, M489, E506, R409, T510, N515, Y518, A539, F550, N551, S596, T602, A611, I617, T620, R650, G654, N667, R685, K695, I706, K722, A723, K724, M731, F732, K735, S73 9, P741, E742, E746, Q747, I754, T755, H757, K760, H761, P778, E781, I783, N784, D785, T786L, L787, Y788, K792, D794, T798, L799, V 801, N803, L804, N805, G806, D813, K814, L818, I819, S822, E824, L841, G847, D848, Y857, V875, I876, N884, A888, L890, D894, D895 , P897, V903, G920, F924, N929, E936, N937, V941, N942, S943, C945, E947, K951, L952, S956, N957, Q958, A959, N974, G975, V983, N9 Includes mutations or amino acid substitutions corresponding to any one of the following positions: 84, N985, D986, I991, V993, M995, I996, T999, Y1000, R1001, E1002, L1004, E1005, N1006, M1007, D1009, K1010, R1011, P1012, P1013, I1015, I1016, A1020, S1021, Q1024, K1027, E1039, H1045, I0148, K1050, or any combination thereof.
[0106] In some embodiments, the Staphylococcus aureus Cas9 domains are D10A, D10E, H557A, N580A, H840A, D1135E, R1335Q, T1337R, T267A, L325F, V327I, D333G, A336S, I341L, E345D, D348N, K352E, S360A, T368A, N369E, N371E, S37 2P, E373K, K386T, N393R, H408N, N410S, I414M, A415T, T438S, Y467F, N471K, D485E, M489F, E506 K, R409K, T510E, N515K, Y518F, A539P, F550Y, N551H, S596A, T602I, A611S, I617V, T620K, R650K,G654E, N667D, R685K, K695Q, I706V, K722T, A723T, K724N, M731T, F732V, K735Q, S739N, P741L, E742G, E7 46D, Q747D, I754D, T755I, H757R, K760Q, H761S, P778I, E781K, I783V, N784D, D785E, T786L, L787V, Y788H , K792E, D794T, T798R, L799I, V801I, N803S, L804I, N805K, G806N, D813G, K814E, L818I, I819F, S822P, E 824G, L841T, G847S, D848N, Y857H, V875I, I876V, N884K, A888V, L890R, D894G, D895H, P897L, V903I, G920 D, F924L, N929Y, E936D, N937G, V941I, N942D, S943L, C945A, E947K, K951R, L952Q, S956N, N957E, Q958K, A959S, N974D, G975K, V983A, N984S, N985D, D986G, I991V, V993L, M995F, I996V, T999N, Y1000K, R1001E, E Includes mutations or substitutions corresponding to one or more of the following: 1002D, L1004I, E1005K, N1006M, M1007N, D1009L, K1010S, R1011T, P1012S, P1013F, I1015L, I1016R, A1020G, S1021K, Q1024K, K1027S, E1039K, H1045K, I0148M, K1050M, or any combination thereof.
[0107] In some embodiments, Cas9 includes a mutation in the residue corresponding to amino acid position 10 of SEQ ID NO: 36 or SEQ ID NO: 86. In some embodiments, Cas9 includes a mutation in the residue corresponding to amino acid position 557 of SEQ ID NO: 36. In some embodiments, Cas9 includes a mutation in the residue corresponding to amino acid position 580 of SEQ ID NO: 36. In some embodiments, Cas9 includes a mutation in the residue corresponding to amino acid position 650 of SEQ ID NO: 36. In some embodiments, Cas9 includes a mutation in the residue corresponding to amino acid position 840 of SEQ ID NO: 86. In some embodiments, Cas9 includes a mutation in the residue corresponding to amino acid position 1135 of SEQ ID NO: 86. In some embodiments, Cas9 includes a mutation in the residue corresponding to amino acid position 1335 of SEQ ID NO: 86. In some embodiments, Cas9 includes a mutation in the residue corresponding to amino acid position 1337 of SEQ ID NO: 86.
[0108] In some embodiments, Cas9 contains a mutation in the residue corresponding to D10 in SEQ ID NO: 36 or SEQ ID NO: 86. In some embodiments, Cas9 contains the D10E mutation in SEQ ID NO: 36 or SEQ ID NO: 86. In some embodiments, Cas9 contains the D10A mutation in SEQ ID NO: 36 or SEQ ID NO: 86. In some embodiments, Cas9 contains the H557A mutation in SEQ ID NO: 36. In some embodiments, Cas9 contains the N580A mutation in SEQ ID NO: 36. In some embodiments, Cas9 contains the R650K mutation in SEQ ID NO: 36. In some embodiments, Cas9 contains the H840A mutation in SEQ ID NO: 86. In some embodiments, Cas9 contains the D1135E mutation in SEQ ID NO: 86. In some embodiments, Cas9 contains the R1335Q mutation in SEQ ID NO: 86. In some embodiments, Cas9 contains the T1337R mutation in SEQ ID NO: 86.
[0109] In some embodiments, Cas9 includes the D10E mutation and the H557A mutation in SEQ ID NO: 36. In some embodiments, Cas9 includes the D10A mutation and the H557A mutation in SEQ ID NO: 36. In some embodiments, Cas9 includes the D10E mutation and the N580A mutation in SEQ ID NO: 36. In some embodiments, Cas9 includes the D10A mutation and the N580A mutation in SEQ ID NO: 36. In some embodiments, Cas9 includes the D10E mutation and the H840A mutation in SEQ ID NO: 86. In some embodiments, Cas9 includes the D10A mutation and the H840A mutation in SEQ ID NO: 86. In some embodiments, Cas9 includes the D10E mutation and the D1135E mutation in SEQ ID NO: 86. In some embodiments, Cas9 includes the D10E mutation and the R1335Q mutation in SEQ ID NO: 86. In some embodiments, Cas9 includes the D10A and R1335Q mutations in SEQ ID NO: 86. In some embodiments, Cas9 includes the D10E and T1337R mutations in SEQ ID NO: 86. In some embodiments, Cas9 includes the D10A and T1337R mutations in SEQ ID NO: 86. In some embodiments, Cas9 includes the D10E, D1135E, R1335Q, and T1337R mutations in SEQ ID NO: 86. In some embodiments, Cas9 includes the D10E, H840A, D1135E, R1335Q, and T1337R mutations in SEQ ID NO: 86.
[0110] In some embodiments, the Cas9 nuclease is the Staphylococcus aureus (SaCas9) nuclease. In some embodiments, SaCas9 contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 36. In some embodiments, SaCas9 contains an amino acid sequence that is 90% identical to SEQ ID NO: 36. In some embodiments, SaCas9 contains an amino acid sequence that is 91% identical to SEQ ID NO: 36. In some embodiments, SaCas9 contains an amino acid sequence that is 92% identical to SEQ ID NO: 36. In some embodiments, SaCas9 contains an amino acid sequence that is 93% identical to SEQ ID NO: 36. In some embodiments, SaCas9 contains an amino acid sequence that is 94% identical to SEQ ID NO: 36. In some embodiments, SaCas9 contains an amino acid sequence that is 95% identical to SEQ ID NO: 36. In some embodiments, SaCas9 contains an amino acid sequence that is 96% identical to SEQ ID NO: 36. In some embodiments, SaCas9 contains an amino acid sequence that is 97% identical to SEQ ID NO: 36. In some embodiments, SaCas9 contains an amino acid sequence that is 98% identical to SEQ ID NO: 36. In some embodiments, SaCas9 contains an amino acid sequence that is 99% identical to SEQ ID NO: 36. In some embodiments, SaCas9 contains an amino acid sequence that is 100% identical to SEQ ID NO: 36.
[0111] In some embodiments, SaCas9 includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 37. In some embodiments, SaCas9 includes an amino acid sequence that is 90% identical to SEQ ID NO: 37. In some embodiments, SaCas9 includes an amino acid sequence that is 91% identical to SEQ ID NO: 37. In some embodiments, SaCas9 includes an amino acid sequence that is 92% identical to SEQ ID NO: 37. In some embodiments, SaCas9 includes an amino acid sequence that is 93% identical to SEQ ID NO: 37. In some embodiments, SaCas9 includes an amino acid sequence that is 94% identical to SEQ ID NO: 37. In some embodiments, SaCas9 includes an amino acid sequence that is 95% identical to SEQ ID NO: 37. In some embodiments, SaCas9 includes an amino acid sequence that is 96% identical to SEQ ID NO: 37. In some embodiments, SaCas9 contains an amino acid sequence that is 97% identical to SEQ ID NO: 37. In some embodiments, SaCas9 contains an amino acid sequence that is 98% identical to SEQ ID NO: 37. In some embodiments, SaCas9 contains an amino acid sequence that is 99% identical to SEQ ID NO: 37. In some embodiments, SaCas9 contains an amino acid sequence that is 100% identical to SEQ ID NO: 37.
[0112] In some embodiments, SaCas9 contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 38.
[0113] In some embodiments, SaCas9 includes an amino acid sequence that is 90% identical to SEQ ID NO: 38. In some embodiments, SaCas9 includes an amino acid sequence that is 91% identical to SEQ ID NO: 38. In some embodiments, SaCas9 includes an amino acid sequence that is 92% identical to SEQ ID NO: 38. In some embodiments, SaCas9 includes an amino acid sequence that is 93% identical to SEQ ID NO: 38. In some embodiments, SaCas9 includes an amino acid sequence that is 94% identical to SEQ ID NO: 38. In some embodiments, SaCas9 includes an amino acid sequence that is 95% identical to SEQ ID NO: 38. In some embodiments, SaCas9 includes an amino acid sequence that is 96% identical to SEQ ID NO: 38. In some embodiments, SaCas9 includes an amino acid sequence that is 97% identical to SEQ ID NO: 38. In some embodiments, SaCas9 includes an amino acid sequence that is 98% identical to SEQ ID NO: 38. In some embodiments, SaCas9 includes an amino acid sequence that is 99% identical to SEQ ID NO: 38. In some embodiments, SaCas9 includes an amino acid sequence that is 100% identical to SEQ ID NO: 38.
[0114] In some embodiments, SaCas9 includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 39. In some embodiments, SaCas9 includes an amino acid sequence that is 90% identical to SEQ ID NO: 39. In some embodiments, SaCas9 includes an amino acid sequence that is 91% identical to SEQ ID NO: 39. In some embodiments, SaCas9 includes an amino acid sequence that is 92% identical to SEQ ID NO: 39. In some embodiments, SaCas9 includes an amino acid sequence that is 93% identical to SEQ ID NO: 39. In some embodiments, SaCas9 includes an amino acid sequence that is 94% identical to SEQ ID NO: 39. In some embodiments, SaCas9 includes an amino acid sequence that is 95% identical to SEQ ID NO: 39. In some embodiments, SaCas9 includes an amino acid sequence that is 96% identical to SEQ ID NO: 39. In some embodiments, SaCas9 contains an amino acid sequence that is 97% identical to SEQ ID NO: 39. In some embodiments, SaCas9 contains an amino acid sequence that is 98% identical to SEQ ID NO: 39. In some embodiments, SaCas9 contains an amino acid sequence that is 99% identical to SEQ ID NO: 39. In some embodiments, SaCas9 contains an amino acid sequence that is 100% identical to SEQ ID NO: 39.
[0115] In some embodiments, SaCas9 contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 40. In some embodiments, SaCas9 contains an amino acid sequence that is 90% identical to SEQ ID NO: 40. In some embodiments, SaCas9 contains an amino acid sequence that is 91% identical to SEQ ID NO: 40. In some embodiments, SaCas9 contains an amino acid sequence that is 92% identical to SEQ ID NO: 40. In some embodiments, SaCas9 contains an amino acid sequence that is 93% identical to SEQ ID NO: 40. In some embodiments, SaCas9 contains an amino acid sequence that is 94% identical to SEQ ID NO: 40. In some embodiments, SaCas9 contains an amino acid sequence that is 95% identical to SEQ ID NO: 40. In some embodiments, SaCas9 contains an amino acid sequence that is 96% identical to SEQ ID NO: 40. In some embodiments, SaCas9 contains an amino acid sequence that is 97% identical to SEQ ID NO: 40. In some embodiments, SaCas9 contains an amino acid sequence that is 98% identical to SEQ ID NO: 40. In some embodiments, SaCas9 contains an amino acid sequence that is 99% identical to SEQ ID NO: 40. In some embodiments, SaCas9 contains an amino acid sequence that is 100% identical to SEQ ID NO: 40.
[0116] In some embodiments, SaCas9 contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 46. In some embodiments, SaCas9 contains an amino acid sequence that is 90% identical to SEQ ID NO: 46. In some embodiments, SaCas9 contains an amino acid sequence that is 91% identical to SEQ ID NO: 46. In some embodiments, SaCas9 contains an amino acid sequence that is 92% identical to SEQ ID NO: 46. In some embodiments, SaCas9 contains an amino acid sequence that is 93% identical to SEQ ID NO: 46. In some embodiments, SaCas9 contains an amino acid sequence that is 94% identical to SEQ ID NO: 46. In some embodiments, SaCas9 contains an amino acid sequence that is 95% identical to SEQ ID NO: 46. In some embodiments, SaCas9 contains an amino acid sequence that is 96% identical to SEQ ID NO: 46. In some embodiments, SaCas9 contains an amino acid sequence that is 97% identical to SEQ ID NO: 46. In some embodiments, SaCas9 contains an amino acid sequence that is 98% identical to SEQ ID NO: 46. In some embodiments, SaCas9 contains an amino acid sequence that is 99% identical to SEQ ID NO: 46. In some embodiments, SaCas9 contains an amino acid sequence that is 100% identical to SEQ ID NO: 46.
[0117] In some embodiments, SaCas9 includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 47. In some embodiments, SaCas9 includes an amino acid sequence that is 90% identical to SEQ ID NO: 47. In some embodiments, SaCas9 includes an amino acid sequence that is 91% identical to SEQ ID NO: 47. In some embodiments, SaCas9 includes an amino acid sequence that is 92% identical to SEQ ID NO: 47. In some embodiments, SaCas9 includes an amino acid sequence that is 93% identical to SEQ ID NO: 47. In some embodiments, SaCas9 includes an amino acid sequence that is 94% identical to SEQ ID NO: 47. In some embodiments, SaCas9 includes an amino acid sequence that is 95% identical to SEQ ID NO: 47. In some embodiments, SaCas9 includes an amino acid sequence that is 96% identical to SEQ ID NO: 47. In some embodiments, SaCas9 contains an amino acid sequence that is 97% identical to SEQ ID NO: 47. In some embodiments, SaCas9 contains an amino acid sequence that is 98% identical to SEQ ID NO: 47. In some embodiments, SaCas9 contains an amino acid sequence that is 99% identical to SEQ ID NO: 47. In some embodiments, SaCas9 contains an amino acid sequence that is 100% identical to SEQ ID NO: 47.
[0118] In some embodiments, SaCas9 contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 48. In some embodiments, SaCas9 contains an amino acid sequence that is 90% identical to SEQ ID NO: 48. In some embodiments, SaCas9 contains an amino acid sequence that is 91% identical to SEQ ID NO: 48. In some embodiments, SaCas9 contains an amino acid sequence that is 92% identical to SEQ ID NO: 48. In some embodiments, SaCas9 contains an amino acid sequence that is 93% identical to SEQ ID NO: 48. In some embodiments, SaCas9 contains an amino acid sequence that is 94% identical to SEQ ID NO: 48. In some embodiments, SaCas9 contains an amino acid sequence that is 95% identical to SEQ ID NO: 48. In some embodiments, SaCas9 contains an amino acid sequence that is 96% identical to SEQ ID NO: 48. In some embodiments, SaCas9 contains an amino acid sequence that is 97% identical to SEQ ID NO: 48. In some embodiments, SaCas9 contains an amino acid sequence that is 98% identical to SEQ ID NO: 48. In some embodiments, SaCas9 contains an amino acid sequence that is 99% identical to SEQ ID NO: 48. In some embodiments, SaCas9 contains an amino acid sequence that is 100% identical to SEQ ID NO: 48.
[0119] In some embodiments, SaCas9 contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 49. In some embodiments, SaCas9 contains an amino acid sequence that is 90% identical to SEQ ID NO: 49. In some embodiments, SaCas9 contains an amino acid sequence that is 91% identical to SEQ ID NO: 49. In some embodiments, SaCas9 contains an amino acid sequence that is 92% identical to SEQ ID NO: 49. In some embodiments, SaCas9 contains an amino acid sequence that is 93% identical to SEQ ID NO: 49. In some embodiments, SaCas9 contains an amino acid sequence that is 94% identical to SEQ ID NO: 49. In some embodiments, SaCas9 contains an amino acid sequence that is 95% identical to SEQ ID NO: 49. In some embodiments, SaCas9 contains an amino acid sequence that is 96% identical to SEQ ID NO: 49. In some embodiments, SaCas9 contains an amino acid sequence that is 97% identical to SEQ ID NO: 49. In some embodiments, SaCas9 contains an amino acid sequence that is 98% identical to SEQ ID NO: 49. In some embodiments, SaCas9 contains an amino acid sequence that is 99% identical to SEQ ID NO: 49. In some embodiments, SaCas9 contains an amino acid sequence that is 100% identical to SEQ ID NO: 49.
[0120] In some embodiments, SaCas9 includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 50. In some embodiments, SaCas9 includes an amino acid sequence that is 90% identical to SEQ ID NO: 50. In some embodiments, SaCas9 includes an amino acid sequence that is 91% identical to SEQ ID NO: 50. In some embodiments, SaCas9 includes an amino acid sequence that is 92% identical to SEQ ID NO: 50. In some embodiments, SaCas9 includes an amino acid sequence that is 93% identical to SEQ ID NO: 50. In some embodiments, SaCas9 includes an amino acid sequence that is 94% identical to SEQ ID NO: 50. In some embodiments, SaCas9 includes an amino acid sequence that is 95% identical to SEQ ID NO: 50. In some embodiments, SaCas9 includes an amino acid sequence that is 96% identical to SEQ ID NO: 50. In some embodiments, SaCas9 contains an amino acid sequence that is 97% identical to SEQ ID NO: 50. In some embodiments, SaCas9 contains an amino acid sequence that is 98% identical to SEQ ID NO: 50. In some embodiments, SaCas9 contains an amino acid sequence that is 99% identical to SEQ ID NO: 50. In some embodiments, SaCas9 contains an amino acid sequence that is 100% identical to SEQ ID NO: 50.
[0121] In some embodiments, SaCas9 includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the SaCas9 amino acid sequence encoded by SEQ ID NO: 15. In some embodiments, SaCas9 includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the SaCas9 amino acid sequence encoded by SEQ ID NO: 16. In some embodiments, SaCas9 includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the SaCas9 amino acid sequence encoded by SEQ ID NO: 17. In some embodiments, SaCas9 includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the SaCas9 amino acid sequence encoded by SEQ ID NO: 18. In some embodiments, SaCas9 includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the SaCas9 amino acid sequence encoded by SEQ ID NO: 19.
[0122]
[0123] In some embodiments, the Cas9 nuclease is a Streptococcus pyogenes (SpCas9) nuclease. In some embodiments, SpCas9 contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 86. In some embodiments, SpCas9 contains an amino acid sequence that is 90% identical to SEQ ID NO: 86. In some embodiments, SpCas9 contains an amino acid sequence that is 91% identical to SEQ ID NO: 86. In some embodiments, SpCas9 contains an amino acid sequence that is 92% identical to SEQ ID NO: 86. In some embodiments, SpCas9 contains an amino acid sequence that is 93% identical to SEQ ID NO: 86. In some embodiments, SpCas9 contains an amino acid sequence that is 94% identical to SEQ ID NO: 86. In some embodiments, SpCas9 includes an amino acid sequence that is 95% identical to SEQ ID NO: 86. In some embodiments, SpCas9 includes an amino acid sequence that is 96% identical to SEQ ID NO: 86. In some embodiments, SpCas9 includes an amino acid sequence that is 97% identical to SEQ ID NO: 86. In some embodiments, SpCas9 includes an amino acid sequence that is 98% identical to SEQ ID NO: 86. In some embodiments, SpCas9 includes an amino acid sequence that is 99% identical to SEQ ID NO: 86. In some embodiments, SpCas9 includes an amino acid sequence that is 100% identical to SEQ ID NO: 86.
[0124] In some embodiments, the *Staphylococcus pyogenes* Cas9 domain is located at positions D10, S29, F32, D39, R40, H41, S42, I48, C80, S87, K112, H113, K132, K141, D147, L158, E171, P176, I186, V189, Q190, Q194, N199, I201, N202, A203, S204, R205, A210, Q228, L229, G231, S245, T249, S254, D261, T270, N295, T300, D304, V308, N309, I312, T333, A3 37, E345, F352, Q354, S355, K356, G366, A367, E396, L398, I414, D428, F4 29, D435, K468, S469, E470, T472, E480, A486, S490, F498, K500, N501, N5 04, K528, V530, E532, G533, A538, T555, K570, F575, D605, E611, R629, E6 34, T638, R655, R664, R671, K705, E706, Q709, K710, S714, G7115, G717, H7 21, H723, A725, N726, V743, L747, V748, K772, K775, N776, I788, G792, K7 97, Y799, T804, N808, L811, R820, N831, R832, V842, L847, N869, E874, N8 81, Q885, N888, T893, L911, Y945, D946, L949, E952, A1023, Y1036, G1067 , G1077, R1078, N1093, R1114, N1115, D1117, A1121, D1125, P1128, K1129 , V1146, S1154, S1159, L1164, S1172, N1177, P1178, I1179, D1180, K1211 , M1213, G1218, N1234, E1243, K1244, E1253, E1260, K1263, H1264, E1271 , Q1272, E1275, V1290, L1291, S1292, A1293, N1295, H1297, R1298, D1299 , K1300, R1303, E1307, N1308, I1309, I1310, H1311, L1312, L1315, T1316,Includes mutations corresponding to any one of the following: N1317, Y1326, D1328, V1342, A1345, I1360, S1363, or any combination thereof.
[0125] In some embodiments, the Streptococcus pyogenes Cas9 domain is D10E, D10A, S29T, F32M, D39N, R40K, H41Q, S42T, I48L, C80R, S87A, K112D, H113N, K132N, K141E, D147E, L158V, E171Q, P176S, I186K, V189L, Q190H, Q194E, N199R, I201L, N202E, A203E, S204I, R205K, A210G, Q228A, L229F, G231N, S245A, T249M, S254A, D261N, T27 0S, N295K, T300I, D304G, V308A, N309D, I312V, T333A, A337V, E345K, F35 2S, Q354K, S355T, K356T, G366K, A367T, E396D, L398F, I414V, D428A, F42 9Y, D435E, K468Q, S469R, E470N, T472A, E480D, A486T, S490L, F498V, K50 0E, N501H, N504T, K528R, V530I, E532D, G533E, A538E, T555A, K570Q, F57 5C, D605E, E611D, R629K, E634K, T638K, R655H, R664K, R671K, K705V, E70 6D, Q709K, K710A, S714F, G7115E, G717K, H721K, H723Q, A725S, N726A, V7 43I, L747I, V748I, K772Q, K775R, N776R, I788M, G792R, K797E, Y799H, T8 04A, N808D, L811R, R820K, N831D, R832H, V842I, L847I, N869D, E874A, N8 81S, Q885R, N888K, T893S, L911A, Y945H, D946G, L949P, E952A, A1023G, Y 1036R, G1067E, G1077E, R1078K, N1093T, R1114G, N1115E, D1117A, A1121 P, D1125G, P1128T, K1129T, V1146I, S1154T, S1159P, L1164V, S1172N, N1 177D, P1178S, I1179V, D1180S, K1211R, M1213L, G1218T, N1234H, E1243D,Includes mutations corresponding to any one of the following: K1244T, E1253K, E1260D, K1263Q, H1264Y, E1271D, Q1272W, E1275H, V1290L, L1291R, S1292A, A1293T, N1295E, H1297N, R1298T, D1299H, K1300L, R1303S, E1307D, N1308S, I1309M, I1310L, H1311N, L1312A, L1315F, T1316S, N1317R, Y1326F, D1328N, V1342I, A1345S, I1360L, S1363N, or any combination thereof.
[0126]
[0127] In some embodiments, the Cas9 domain is the Neisseria meningitidis Cas9 nuclease. In some embodiments, the Neisseria meningitidis Cas9 domain contains the amino acid sequence described in Sequence ID No. 148. Other Neisseria meningitidis Cas9s can be found at www.uniprot.org / uniprot / under accession numbers C9X1G5, A1IQ68, E0NB23, A9M1K5, or C6S593.
[0128] In some embodiments, the meningococcal Cas9 domain is located at position I9, D16, D30, E31, A94, I103, P124, N164, I213, G229, T241, S376, E393, G454, K471, G490, D660, C665, K764, T770, P803, A841, H842, K843, D844, L846, R847, K854, H855, N856, K858, K862, W865, E868, I869, A872, D873, N876, Y880, G883, I886, E887, E890, R895, A898, Y899, G900, G901, N902, A903, K904, Q905, D908, N912, K917, G919, L9 21, V927, K929, T930, E932, S933, L936, L937, N938, K939, K940, Y943, T944, G949, D950, C958, K965, N966, Q967, F969, A975, E980, N981, I986, D987, C988, K989, G990, Y991, R992, I993, D994, Y997, T998, C1000, S1002, H1004 , K1005, Y1006, A1010, F1011, Q1012, K1013, D1014, E1015, K1018, V1019, E1020, F1021, A1022, Y1024, I1025, N Includes mutations corresponding to any one of the following: 1026, C1027, D1028, S1029, S1030, N1031, R1033, F1034, Y1035, L1036, A1037, W1038, K1041, G1042, K1044, E1045, Q1046, Q1047, F1048, R1049, I1050, S1051, T1052, Q1053, N1054, L1055, V1056, L1057, I1058, Y1061, V1063, N1064, or any combination thereof.
[0129] In some embodiments, the meningococcal Cas9 domain is I9M, D16E, D30E, E31K, A94D, I103V, P124C, N164D, I213N, G229D, T241A, S376T, E393K, G454C, K471E, G490C, D660E, C665R, K764E, T770A, P803S, A841Q, H842G, K843H, D844E, L846V, R847K, K854R, H855L, N856D, K858G, K862L, W865P, E868Q, I869L , A872K, D873G, N876K, Y880R, G883E, I886P, E887K, E890E, R895Q, A898 T, Y899H, G900K, G901D, N902D, A903P, K904T, Q905K, D908A, N912E, K91 7Y, G919T, L921Q, V927I, K929Q, T930V, E932K, S933T, L936W, L937V, N9 38R, K939N, K940H, Y943N, T944G, G949A, D950T, C958E, K965G, N966G, Q9 67K, F969Y, A975S, E980K, N981G, I986R, D987A, C988V, K989V, G990A, Y 991F, R992K, I993D, D994E, Y997F, T998E, C1000R, S1002I, H1004Y, K10 05A, Y1006N, A1010K, F1011L, Q1012T, K1013A, D1014K, E1015K, K1018N , V1019E, E1020F, F1021L, A1022G, Y1024F, I1025V, N1026S, C1027L, D10 Includes mutations corresponding to any one of the following: 28N, S1029R, S1030A, N1031T, R1033A, F1034I, Y1035D, L1036I, A1037R, W1038T, K1041T, G1042D, K1044T, E1045K, Q1046G, Q1047E, F1048Q, R1049S, I1050V, S1051G, T1052V, Q1053K, N1054T, L1055A, V1056L, L1057S, I1058F, Y1061N, V1063I, N1064D, or any combination thereof.
[0130] In some embodiments, the RNA-induced nuclease Neisseria meningitidis Cas9 domain includes an amino acid sequence having at least 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 148.
[0131]
[0132] In some embodiments, the Cas9 domain is the Campylobacter jejuniCas9 nuclease. In some embodiments, the Campylobacter jejuniCas9 domain contains the amino acid sequence shown in SEQ ID NO: 149. Other Campylobacter jejuniCas9s can be found at www.uniprot.org / uniprot / under accession numbers Q0P897, A7H5P1, A0A2U0QR81, A0A5Y4VLH1, or A0A381CRM8.In some embodiments, the RNA-inducible nuclease Campylobacter jejuni Cas9 domain is located at positions L5, A6, D8, I9, S12, S13, F18, S19, L24, K25, I31, T40, E42, L50, L58, A59, R61, L58, L65, H67A in SEQ ID NO: 149 N74, K77, L98, I99, P101, N110, L113, A119, A126, R128, I134, K140, A144, K147, Q151, L156, V184, S190, F199, D202 , G203, R212, F214, K221, E223, Y232, A235, V243, S247, D251, P256, L261, T269, N276, N277, L285, T287, L291, K300 , T305, Q308, L312, G314, Y335, K336, I339, H345, D351, N353, E354, I362, K370, D383E, S384, K391, I396, L403, T40 5, K413, N419, L421, D430, K432, A437, L453, K457, V462, A465, K472, N477, A492, E495, L525, K526, L527, K531, E532 , E542, Q550, E556, H559, Y561, S564, M572, V577, Q581, N587, N596, K600, Q602, K603, Q616, K617, N623, Y624, K633 , D634, Y642, N649, D656, L660, D662, K667, V677, E680, K682, L686, H692, T693, V712, I714, V722, K723, S736, L739 This includes mutations corresponding to any one of the following: K742, L747, N751, F756, R763, Q764, E772, K777, A786, E790, F792, Q800, S801, G804, L812, E813, V833, I835, T841, Y845, A855, L856, A863, V864, D879, E883, D900, Q902, K927, F928, V971, T972, or any combination thereof.
[0133] In some embodiments, the RNA-inducible nuclease Campylobacter jejuni Cas9 domain is L5I, A6G, D8N, D8E, I9L, S12A, S13N, F18L, S19R, L24I, K25I, I31V, T40N, E42N, L50E, L58V, A59K, R61K, L58V, L65M, H67A, N74K, K77N, L98T, I99Q, P101I, N110S, L113I, A119S, A126V, R128H, I134S, K140N, A144T, K147E, Q151K, L156M, V 184I, S190D, F199L, D202Q, G203E, R212K, F214L, K221K, E223K, Y232F, A23 5P, V243I, S247I, D251N, P256A, L261S, T269G, N276K, N277S, L285V, T287E, L291I, K300D, T305S, Q308K, L312I, G314N, Y335L, K336N, I339K, H345T, D3 51I, N353D, E354S, I362T, K370E, D383E, S384K, K391N, I396L, L403Q, T405I , K413R, N419E, L421C, D430E, K432S, A437L, L453I, K457C, V462L, A465D, K 472S, N477H, A492K, E495I, L525Q, K526I, L527V, K531E, E532D, E542L, Q55 0D, E556V, H559Y, Y561R, S564N, M572S, V577T, Q581L, N587G, N596E, K600L , Q602A, K603E, Q616R, K617F, N623F, Y624F, K633T, D634E, Y642W, N649S, D6 56S, L660I, D662E, K667A, V677Q, E680V, K682S, L686I, H692N, T693F, V712 I, I714V, V722I, K723F, S736K, L739F, K742N, L747S, N751L, F756L, R763K, Q 764E, E772N, K777H, A786T, E790L, F792P, Q800N, S801T, G804D, L812V, E81 3K, V833S, I835L, T841K, Y845H, A855S, L856T, A863T, V864P, D879N, E883N,Includes mutations corresponding to any one of the following: D900G, Q902K, K927N, F928Y, V971L, T972S, or any combination thereof.
[0134] In some embodiments, the Campylobacter jejuniCas9 domain includes an amino acid sequence having at least 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 149. In some embodiments, the Campylobacter jejuniCas9 domain includes an amino acid sequence having 85-90%, 90-95%, 95-97%, 97-98%, or 98-99% sequence identity with SEQ ID NO: 149.
[0135]
[0136] In some embodiments, the Cas9 domain is the Streptococcus pasteurianusCas9 nuclease. In some embodiments, the Streptococcus pasteurianusCas9 domain comprises the amino acid sequence described in SEQ ID NO: 150. Other Streptococcus pasteurianusCas9 can be found at www.uniprot.org / uniprot / under accession number F5X275.
[0137] In some embodiments, the Streptococcus pasteurianusCas9 domain is located at positions D11, E85, A88, T92, E96, Y100, T109, D110, D113, E115, R116, D125, I127, K128, E132, S147, I185, A187, K228, Y229, T232, M255, S271, N273, A294, A327, E355, K357, N379, T380, S382, A385, D439, R440, S464, H469, Y519, I528, N569, I581, A607, K632 , D633, H635, E636, A647, D648, T703, P705, K712, S713, A724, V750, D882, S951, D977, E979, S10 14, H1027, I1030, E1081, D1082, D1086, K1088, S1089, N1090, R1092, T1093, I1094, C1095, A113 8, Y1139, D1141, T1142, F1158, A1168, E1190, E1198, H1202, I1204, R1205, I1210, K1224, S1232 This includes mutations corresponding to M1240, V1241, I1242, P1243, G1424, K1248, Q1254, N1257, S1258, T1262, K1263, Y1264, D1266, A1270, K1277, D1284, L1288, V1302, N1316, T1346, I1374, or any one combination thereof.
[0138] In some embodiments, the Streptococcus pasturianusCas9 domain is D11E, D11A, E85D, A88T, T92A, E96D, Y100Q, T109D, D110N, D113N, E115D, R116S, D125E, I127D, K128A, E132K, S147T, I185L, A187T, K228N, Y229N, T232K, M255T, S271T, N273E, A294S, A327V of Sequence ID No. 150 , E355K, K357Q, N379G, T380I, S382T, A385N, D439E, R440E, S464A, H469R, Y519F, I528V, N569D, I581V, A607S, K 632R, D633E, H635Q, E636Q, A647K, D648Q, T703A, P705S, K712E, S713A, A724T, V750I, D882G, S951R, D977E, E97 9K, S1014P, H1027R, I1030V, E1081G, D1082E, D1086N, K1088R, S1089T, N1090D, R1092E, T1093K, I1094V, C1095 R, A1138V, Y1139L, D1141E, T1142P, F1158L, A1168T, E1190K, E1198K, H1202Q, I1204V, R1205Q, I1210M, K1224R This includes mutations corresponding to any one of the following: S1232T, M1240I, V1241M, I1242L, P1243S, G1424A, K1248A, Q1254H, N1257G, S1258N, T1262A, K1263E, Y1264H, D1266K, A1270E, K1277E, D1284N, L1288V, V1302A, N1316D, T1346N, I1374L, or any combination thereof.
[0139] In some embodiments, the Streptococcus pasteurianusCas9 domain includes an amino acid sequence having at least 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 150.
[0140]
[0141] In some embodiments, the Cas9 domain is the Clostridium cellulolyticum Cas9 domain. In some embodiments, the Clostridium cellulolyticum Cas9 domain includes the amino acid sequence described in Sequence ID No. 151. Other Clostridium cellulolyticum Cas9 domains can be found at www.uniprot.org / uniprot / under accession number B8I085.
[0142] In some embodiments, the Clostridium cellulolyticum Cas9 domain is located at positions T4, D10, V9, D20, K21, I27, C33, K36, A47, A49, S64, Q65, E102, L103, T122, I124, K131, D137, R163, G166, I169, F170, V183, D 184, I187, E193, K200, K208, L209, D221, N224, E227, F228, S234, V242, K244, L252 , T256, C258, S261, V413, M415, K416, R417, K424, Y426, K427, S429, D430, A468, T4 70, A472, A478, Q481, K482, L485, A497, L535, W540, R541, E544, G554, P556, I570, Y574, M580, Y584, M585, T592, D593, V606, W607, I647, N650, S693, L697, E702, S70 Includes mutations corresponding to any one of the following: 4, A713, V714, I715, D776, L847, G850, G853, A854, R860, I900, H904, M905, I906, E921, Q923, S929, T930, H931, Q939, N994, I997, N1000, K1001, S1002, I1003, K1005, P1008, or any combination thereof.
[0143] In some embodiments, the Clostridium cellulolyticum Cas9 domain is T4S, D10E, V9I, D20N, K21E, I27E, C33I, K36V, A47S, A49P, S64R, Q65H, E102L, L103V, T122V, I124F, K131Q, D137E, R163Q, G166S, I169L, F170L, V183G, D184G, I187T, E19 3S,K200Q,K208A,L209Y,D221K,N224Q,E227S,F228S,S234T,V242I,K244N,L252K,T256K,C258T,S261 F,V413K,M415L,K416R,R417N,K424Q,Y426I,K427P,S429H,D430Q,A468S,T470S,A472V,A478G,Q481K, K482R,L485S,A497M,L535H,W540Y,R541K,E544Q,G554F,P556S,I570V,Y574I,M580F,Y584N,M585N,T 592A,D593A,V606W,W607F,I647R,N650H,S693K,L697F,E702Q,S704N,A713V,V714I,I715V,D776E,L84 Contains mutations corresponding to any one of the following: 7A, G850P, G853A, A854P, R860K, I900V, H904D, M905V, I906L, E921Y, Q923E, S929D, T930E, H931Y, Q939P, N994Q, I997P, N1000R, K1001M, S1002N, I1003K, K1005H, P1008K, or any combination thereof.
[0144] In some embodiments, the Clostridium cellulolyticumCas9 domain includes an amino acid sequence having at least 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 151. In some embodiments, the Clostridium cellulolyticumCas9 domain includes an amino acid sequence having 85-90%, 90-95%, 95-97%, 97-98%, or 98-99% sequence identity with SEQ ID NO: 151.
[0145]
[0146] In some embodiments, the Geobacillus thermodenitrificans T1 Cas9 domain includes the amino acid sequence described in Sequence ID No. 152. Other Geobacillus thermodenitrificans T1 Cas9 can be found at www.uniprot.org / uniprot / under accession number A0A1W6VMQ3.
[0147] In some embodiments, the Geobacillus photodenitrificans T1 Cas9 domain is located at positions K2, D8, I14, D35, K41, F74, V75, K91, I117, R128, T136, Q151, S152, S156, A161, V164, S171, E178, D179, V185, R192, K195, A199, Y204, I207, V208, A212, H215, S219, F227, T260, V261, V271, G274, I276, A278, L279, D282, I287, K289, H293, F299, V302, N307, R313, L317, L318, V331, G337, K341, S348, A354, A355, K356, R359, M372, T377, R380, E395, D399, E404, S416, T441, R445, N464, E504, S508, M515, Q516, E520, G521, V534, L545, K559, T578, K603, T612, L619, S621, N656, N660, L673, D685, I699, N708, N717, R737, V738, S752, D756, Q771, N777, N792, E793, I811, I824, K839, Q845, K848, T849, Includes mutations corresponding to any one of the following: L895, I902, T908, V929, I943, I946, M948, F990, T995, V1000, Q1014, D1017, S1019, N1020, G1021, S1024, N1030, N1031, R1035, S1036, I1037, V1067, S1071, A1075, I1079, or any combination thereof.
[0148] In some embodiments, the Geobacillus thermodenitrificans T1 Cas9 domain contains a mutant corresponding to one of the following positions: D8, D179, D282, D399, D685, D756, or D1071. In some embodiments, the RNA-inducible nuclease Geobacillus thermodenitrificans T1 The Cas9 domain includes K2R, D8E, D8A, I14V, D35E, K41Q, F74V, V75I, K91E, I117V, R128K, T136S, Q151R, S152A, S156G, A161G, V164I, S171A, E178G, D179E, V185I, R192H, K195R, A199S, Y204F, I207M, V208S, A212K, H215N, S219T, F227V, T260I, and V261. A, V271I, G274S, I276A, A278G, L279P, D282E, I287L, K289E, H293Q, F299Y, V302I, N307R, R313Y, L317I, L318V, V331I, G33 7D, K341Q, S348K, A354K, A355S, K356S, R359L, M372L, T377A, R380H, E395P, D399N, E404N, S416T, T441S, R445K, N464T, E5 04D, S508T, M515T, Q516K, E520D, G521E, V534M, L545H, K559R, T578V, K603R, T612I, L619V, S621T, N656M, N660S, L673F, D 685E, I699V, N708E, N717D, R737K, V738I, S752A, D756E, Q771R, N777H, N792D, E793Q, I811V, I824V, K839T, Q845K, K848A, Includes mutations corresponding to any one of the following: T849S, L895P, I902V, T908K, V929V, I943V, I946M, M948I, F990L, T995I, V1000G, Q1014K, D1017H, S1019G, N1020T, G1021A, S1024E, N1030C, N1031S, R1035S, S1036G, I1037V, V1067L, S1071A, A1075T, I1079V, or any combination thereof.
[0149] In some embodiments, the RNA-inducible nuclease Geobacillus thermodenitrificans T1 Cas9 domain includes a mutation corresponding to one of the positions D8E, D179E, D282E, D399N, D685E, D756E, or D1071H of SEQ ID NO: 152. In some embodiments, the RNA-inducible nuclease Geobacillus thermodenitrificans T1 Cas9 domain includes an amino acid sequence having at least 85%, 90%, 95%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 152. In some embodiments, the RNA-inducible nuclease Geobacillus thermodenitrificans T1 Cas9 domain includes an amino acid sequence having 85-90%, 90-95%, 95-97%, 97-98%, or 98-99% sequence identity to SEQ ID NO: 152.
[0150] In some embodiments, the CRISPR Cas nuclease is Deltaproteobacteria CasX, Acidaminococcus Cas12, or Eubacterium rectale Cas12a. In some embodiments, the CRISPR Cas nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to Acidaminococcus Cas12, Deltaproteobacteria CasX, or Eubacterium rectale Cas12a.
[0151] In some embodiments, the Cas domain is the CasX domain. In some embodiments, the CasX domain isAmino acid sequence: (SEQ ID NO: 153) is included.
[0152] In some embodiments, the CasX domain is derived from Plantomycetes Bacterinum. In some embodiments, the CasX domain is derived from Deltaproteobacteria. In some embodiments, the CasX domain contains an amino acid sequence that is at least about 85%, 90%, 95%, 97%, 98%, or 99% identical to SEQ ID NO: 153. In some embodiments, the CasX domain is located at positions R11, R12, V14, K15, S17, N18, A22, G23, T25, P38, K41, E42, N46, L47, N53, I54, P57, T61, S62, R63, A64, E75, H82, Q89, P104, N106, I113, N199, S124, S125, C133, Y137, N145, D146, H151, S161, R165 , N177, L180, R202, N205, G215, C219, V236, T241, L248, I254, S269, I290, E291, V297, Q299, I314, E318, Q323, L333, E359, D360, K362, Q366, N367, L368, A369, G370, Y371, H404, H409, G410, E411, Y417, V428, E429, S432, K433, L437, S4 43, A451, I464, A470, I502, L503, I531, G537, L540, N553, I559, S563, V571, N579, H589, S607, L608, L620, R623, R62 4, L644, S646, M652, I657, R679, L684, N686, H689, S696, T702, T737, L742, Y744, Q748, M751, I753, A771, R777, P792, Includes mutations corresponding to S818, R823, V824, E826, K827, A832, T833, M836, I839, G841, V846, N860, V862, D864, V867, V877, S883, S889, G890, S894, K908, N913, F916, T918, R936, Q938, Y940, K942, S963, R966, K967, K968, or any one of these combinations.
[0153] In some embodiments, the CasX domain is R11K, R12K, V14S, K15A, S17N, N18A, A22V, G23S, T25S, P38D, K41K, E42K, N46K, L47R, N53V, I54M, P57V, T61N, S62A, R63A, A64N, E75K, H82Q, Q89K, P104S, N106K, I113K, N199K, S124T, S125A, C133G, Y137F, N145S, D146E, H151Y, S161A, R165K, N 177S, L180A, R202K, N205T, G215A, C219Y, V236I, T241S, L248I, I254V, S269G, I290V, E291D, V297I, Q299R, I314L, E318D, Q323L, L333V, E3 59D, D360M, K362R, Q366S, N367G, L368V, A369T, G370A, Y371E, H404Y, H409Y, G410A, E411G, Y417F, V428I, E429A, S432T, K433S, L437R, S44 3A, A451V, I464L, A470M, I502V, L503V, I531L, G537K, L540I, N553S, I559L, S563G, V571L, N579Q, H589T, S607L, L608I, L620I, R623K, R624 K, L644V, S646P, M652V, I657V, R679E, L684S, N686G, H689D, S696G, T702A, T737S, L742F, Y744H, Q748H, M751V, I753V, A771T, R777K, P792T This includes mutations or substitutions corresponding to one or more of the following: S818T, R823G, V824M, E826V, K827R, A832S, T833D, M836A, I839L, G841N, V846A, N860T, V862E, D864E, V867A, V877G, S883K, S889R, G890D, S894F, K908Q, N913D, F916H, T918V, R936N, Q938N, Y940F, K942S, S963A, R966K, K967R, K968R, or any combination thereof.
[0154]
[0155] In some embodiments, the Cas12 domain contains an amino acid sequence that is at least approximately 85%, 90%, 95%, 97%, 98%, or 99% identical to or identical to SEQ ID NO: 154. In some embodiments, the Cas12 domain is located at the position of SEQ ID NO: 154.T1, Q2, E4, G5, N8, L9, K28, H29, I30, Q31, E32, Q33, F35, I36, E37, E38, A41 , N43, D44, H45, E48, I52, R55, T59, Y60, A61, D62, Q63, C64, Q66, L67, Q69, L 70, N74, S76, A77, D80, S81, Y82, E85, E88, T90, R91, N92, A93, I95, E97, A99 , T100, Y101, N103, A104, H106, D107, I110, R112, T113, D114, R159, S169, S 185, A187, I192, D195, K201, T212, R218, N223, I228, S233, I236, E237, V2 39, F242, Q249, Y257, V279, I284, F305, N313, S324, I329, S331, T337, L338 , L345, E349, S357, I358, N386, I393, L396, I400, S403, V408, Q409, G427, K 428, Q436, L442, S468, Q469, S472, L473, L479, E487, S488, A497, L510, A51 6, K522, Q535, M536, S541, V545, K549, N550, G552, V557, N559, S586, Y596 , A601, I604, A613, S628, E637, A657, K660, G663, Q665, C673, L683, L697, A 711, L717, Q723, A733, E735, Y740, K751, K756, G766, I778, R793, L844, I85 8, S865, I874, H898, I903, I916, L931, K941, N945, V951, S958, V959, D965, Contains mutations corresponding to one of the following: I938, H984, A1009, C1024, G1037, T1049, G1055, T1056, Y1068, L1075, V1083, K1085, L1097, H1104, D1106, D1111, L1122, A1134, V1138, D1147, V1160, P1161, R1171, R1173, Y1176, N1205, D1207, S1220, V1221, A1230, N1237, L1243, M1259, Q1274, G1291, Q1295, A1299, or L1304.
[0156] In some embodiments, the Cas12 domain is T1S, Q2N, E4S, G5E, N8H, L9K, K28E, H29N, I30L, Q31T, E32A, Q33Y, F35M, I36V, E37N, E38D, A41L, N43S, D44E, H45N, E48K, I52V, R55K, T59Y, Y60F, A61I, D62E, Q63E, C64T, Q66K, L67H, Q69A, L70I, N74P, S76Y, A77K, D80T, S81A, Y82F, E85D, E88L, T90N, R91N, N92T, A93N of sequence number 154. , I95R, E97I, A99D, T100N, Y101C, N103K, A104S, H106A, D107G, I110E, R112 K, T113V, D114P, R159K, S169V, S185A, A187S, I192L, D195E, K201I, T212K, R 218N, N223T, I228T, S233G, I236L, E237D, V239I, F242V, Q249C, Y257F, V27 9T, I284V, F305Y, N313S, S324N, I329L, S331A, T337E, L338K, L345I, E349Q, S357L, I358A, N386D, I393V, L396A, I400L, S403N, V408I, Q409E, G427D, K4 28D, Q436A, L442I, S468V, Q469L, S472A, L473V, L479T, E487D, S488D, A497V , L510I, A516V, K522Q, Q535S, M536N, S541D, V545E, K549Q, N550Q, G552C, V 557E, N559E, S586N, Y596Q, A601S, I604L, A613D, S628N, E637T, A657D, K660 R, G663N, Q665K, C673H, L683V, L697V, A711G, L717F, Q723E, A733L, E735D, Y740F, K751E, K756A, G766A, I778V, R793P, L844F, I858V, S865T, I874L, H89 8N, I903V, I916A, L931F, K941N, N945Q, V951I, S958T, V959A, D965E, I938V , H984Q, A1009S, C1024Y, G1037S, T1049E, G1055R, T1056N, Y1068F, L1075A,Includes mutations or substitutions corresponding to one or more of the following: V1083R, K1085G, L1097I, H1104K, D1106N, D1111N, L1122K, A1134D, V1138I, D1147A, V1160E, P1161F, R1171Q, R1173E, Y1176L, N1205T, D1207N, S1220L, V1221T, A1230E, N1237S, L1243I, M1259K, Q1274L, G1291A, Q1295N, A1299N, or L1304K.
[0157] Additional Cas9 nuclease amino acid sequences are listed in Table 1.
[0158] [Table 2-1]
[0159] [Table 2-2]
[0160] [Table 2-3]
[0161] [Table 2-4]
[0162] [Table 2-5]
[0163] [Table 2-6]
[0164] [Table 2-7]
[0165] [Table 2-8]
[0166] [Table 2-9]
[0167] [Table 2-10]
[0168] [Table 2-11]
[0169] [Table 2-12]
[0170] [Table 2-13]
[0171] [Table 2-14]
[0172] In some embodiments, the chimeric nuclease comprises a Cas nuclease or nuclease domain selected from Table 1. In some embodiments, the chimeric nuclease comprises a Cas nuclease or nuclease domain selected from Sequence IDs 161-191. In some embodiments, the Cas nuclease or nuclease domain comprises an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to Sequence IDs 161-191.
[0173] In some embodiments, the chimeric nuclease further comprises a further protein fusion. In some embodiments, the chimeric nuclease comprises a self-cleaving peptide, such as the T2A peptide.
[0174] GIY-YIG nuclease In some embodiments, the chimeric nuclease comprises a GIY-YIG nuclease. In some embodiments, the chimeric nuclease comprises a GIY-YIG nuclease catalytic domain.
[0175] GIY-YIG nucleases are a family of homing endonucleases that cleave DNA. GIY-YIG nucleases fold into two structural and functional domains, an N-terminal catalytic domain and a C-terminal DNA-binding domain, separated by a flexible linker. In some embodiments, chimeric nucleases include the N-terminal catalytic domain of a GIY-YIG nuclease. In some embodiments, chimeric nucleases include the N-terminal catalytic domain and the mobile linker of a GIY-YIG nuclease. In some embodiments, chimeric nucleases include the N-terminal catalytic domain and the mobile linker of a GIY-YIG nuclease but lack the C-terminal DNA-binding domain of a GIY-YIG nuclease.
[0176] In some embodiments, the GIY-YIG nuclease is the I-TevI nuclease. I-TevI is a site-specific, sequence-tolerant homing endonuclease encoded by the td intron of bacteriophage T4 (Uniprot ID A0A7S9XH31, SEQ ID NO: 155). In some embodiments, the chimeric nuclease contains the catalytic domain of the I-TevI nuclease.
[0177] In some embodiments, I-TevI is modified I-TevI. In some embodiments, I-TevI is the catalytic domain of I-TevI. In some embodiments, I-TevI contains a conserved amino acid substitution. In some embodiments, I-TevI contains a mutation in the catalytic domain. In some embodiments, I-TevI is a nickase.
[0178] In some embodiments, the I-TevI nuclease contains the amino acid sequence (SEQ ID NO: 155) MKSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIADATFGDTCSTHPLKEEIIKKRSETVKAKMLKLGPDGRKALYSKPGSKNGRWNPETHKFCKCGVRIQTSAYTCSKCRNRSGENNSFFNHKHSDITKSKISEKMKGKKPSNIKKISCDGVIFDCAADAARHFKISSGLVTYRVKSDKWNWFYINA.
[0179] In some embodiments, the I-TevI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 155.
[0180] In some embodiments, the I-TevI nuclease contains the amino acid sequence (SEQ ID NO: 741) MKSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIADATFGDTCSTHPLKEEIIKKRSETVKAKMLKLGPDGRKALYSKPGSKNGRWNPETHKFCKCGVRIQTSAYTCSKCRN.
[0181] In some embodiments, the I-TevI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 741.
[0182] In some embodiments, the I-TevI nuclease contains the amino acid sequence (SEQ ID NO: 156) MGKSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIADATFGDTCSTHPLKEEIIKKRSETVKAKMLKLGPDGRKALYSKPGSKNGRWNPETHKFCKCGVRIQTSAYTCSKCRNRSGENNSFFNHKHSDITKSKISEKMKGKKPSNIKKISCDGVIFDCAADAARHFKISSGLVTYRVKSDKWNWFYINA.
[0183] In some embodiments, the I-TevI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 156.
[0184] In some embodiments, the I-TevI nuclease contains the amino acid sequence KSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIADATFGDTCSTHPLKEEIIKKRSETVKAKMLKLGPDGRKALYSKPGSKNGRWNPETHKFCKCGVRIQTSAYTCSKCRNRSGENNSFFNHKHSDITKSKISEKMKGKKPSNIKKISCDGVIFDCAADAARHFKISSGLVTYRVKSDKWNWFYINA (SEQ ID NO: 157).
[0185] In some embodiments, the I-TevI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 157.
[0186] In some embodiments, the I-TevI nuclease catalytic domain contains the amino acid sequence MKSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIA (SEQ ID NO: 24).
[0187] In some embodiments, the I-TevI nuclease catalytic domain contains the amino acid sequence MGKSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIA (SEQ ID NO: 158).
[0188] In some embodiments, the I-TevI nuclease catalytic domain comprises the amino acid sequence KSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIA (SEQ ID NO: 159).
[0189] In some embodiments, the I-TevI nuclease catalytic domain comprises the amino acid sequence (SEQ ID NO: 272) MGKSGIYQIKNTLNNKVYVGSAKDFERRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIADASFGDTCSTHPLKEEIIKKRSETVKAKMLKLGPDGRKALYSKPGSKNGRWNPETHKFCKCGVRIRTSAYTCSKCRN.
[0190] In some embodiments, the I-TevI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 272.
[0191] In some embodiments, the I-TevI nuclease catalytic domain comprises the amino acid sequence (SEQ ID NO: 740) MGKSGIYQIKNTLNNKVYVGSAKDFEKRWKRHFKDLEKGCHSSIKLQRSFNKHGNVFECSILEEIPYEKDLIIERENFWIKELNSKINGYNIADATFGDTCSTHPLKEEIIKKRSETFKAKMLKLGPDGRKALYSRPGSKSGRWNPETHKFCKCGVRIQTSAYTCSKCRN.
[0192] In some embodiments, the I-TevI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 740.
[0193] In some embodiments, the I-TevI sequence includes a mutation at the amino acid position corresponding to K26 in SEQ ID NO: 155. In some embodiments, the I-TevI sequence includes a mutation at the amino acid position corresponding to T95 in SEQ ID NO: 155. In some embodiments, the I-TevI sequence includes a mutation at the amino acid position corresponding to Q158 in SEQ ID NO: 155. In some embodiments, the I-TevI sequence includes a mutation at the amino acid position corresponding to V117 in SEQ ID NO: 155. In some embodiments, the I-TevI sequence includes a mutation at the amino acid position corresponding to K135 in SEQ ID NO: 155. In some embodiments, the I-TevI sequence includes a mutation at the amino acid position corresponding to N140 in SEQ ID NO: 155.
[0194] In some embodiments, the I-TevI sequence comprises a mutation at amino acid positions corresponding to K26 and T95 of SEQ ID NO: 155. In some embodiments, the I-TevI sequence comprises a mutation at amino acid positions corresponding to K26 and Q158 of SEQ ID NO: 155. In some embodiments, the I-TevI sequence comprises a mutation at amino acid positions corresponding to T95 and Q158 of SEQ ID NO: 155. In some embodiments, the I-TevI sequence comprises a mutation at amino acid positions corresponding to K26, T95, and Q158 of SEQ ID NO: 155.
[0195] In some embodiments, the I-TevI sequence comprises a mutation at amino acid positions corresponding to V117, K135, and N140 of SEQ ID NO: 155.
[0196] In some embodiments, the I-TevI sequence comprises the K26R mutation at the amino acid position corresponding to SEQ ID NO: 155. In some embodiments, the I-TevI sequence comprises the T95S mutation at the amino acid position corresponding to SEQ ID NO: 155. In some embodiments, the I-TevI sequence comprises the Q158R mutation at the amino acid position corresponding to SEQ ID NO: 155. In some embodiments, the I-TevI sequence comprises the V117F mutation at the amino acid position corresponding to SEQ ID NO: 155. In some embodiments, the I-TevI sequence comprises the K135R mutation at the amino acid position corresponding to SEQ ID NO: 155. In some embodiments, the I-TevI sequence comprises the N140S mutation at the amino acid position corresponding to SEQ ID NO: 155.
[0197] In some embodiments, the I-TevI sequence comprises the K26R and T95S mutations at the amino acid positions corresponding to SEQ ID NO: 155. In some embodiments, the I-TevI sequence comprises the K26R and Q158R mutations at the amino acid positions corresponding to SEQ ID NO: 155. In some embodiments, the I-TevI sequence comprises the T95S and Q158R mutations at the amino acid positions corresponding to SEQ ID NO: 155. In some embodiments, the I-TevI sequence comprises the K26R, T95S, and Q158R mutations at the amino acid positions corresponding to SEQ ID NO: 155.
[0198] In some embodiments, the I-TevI sequence comprises the V117F, K135R, and N140S mutations at the amino acid positions corresponding to SEQ ID NO: 155.
[0199] In some embodiments, the I-TevI sequence comprises mutations at the amino acid positions corresponding to R27, V117, K135, and N140 of SEQ ID NO: 155.
[0200] In some embodiments, the I-TevI nickase domain comprises the R27A, V117F, K135R, and N140S mutations. In some embodiments, the I-TevI comprises a mutation that results in I-TevI that has lost catalytic activity. In some embodiments, the I-TevI comprises a mutation that results in I-TevI that has lost catalytic activity at the residue corresponding to amino acid position R27A of SEQ ID NO: 24. In some embodiments, the I-TevI comprises a mutation that results in I-TevI that has lost catalytic activity at the residue corresponding to amino acid position R28A of SEQ ID NO: 25.
[0201] In some embodiments, I-TevI contains a mutation at the residue corresponding to amino acid position 26 in SEQ ID NO: 24. In some embodiments, I-TevI contains the K26R mutation in SEQ ID NO: 24. In some embodiments, I-TevI contains a mutation at the residue corresponding to amino acid position 27 in SEQ ID NO: 25. In some embodiments, I-TevI contains the K27R mutation in SEQ ID NO: 25.
[0202] In certain embodiments, the modified I-TevI nuclease domain includes substitutions selected from one or more of the following: T11V, V16I, N14G, E25D, K26R, E36S, K37N, G38N, C39V, S41H, L45F, F49Y, I60V, and E81I, corresponding to the amino acid positions of SEQ ID NO: 155.
[0203] In some embodiments, the I-TevI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 156. In some embodiments, the I-TevI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 157. In some embodiments, the I-TevI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 158. In some embodiments, the I-TevI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 159. In some embodiments, the I-TevI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 24. In some embodiments, I-TevI contains an amino acid sequence that is 90% identical to SEQ ID NO: 24. In some embodiments, I-TevI contains an amino acid sequence that is 91% identical to SEQ ID NO: 24. In some embodiments, I-TevI contains an amino acid sequence that is 92% identical to SEQ ID NO: 24. In some embodiments, I-TevI contains an amino acid sequence that is 93% identical to SEQ ID NO: 24. In some embodiments, I-TevI contains an amino acid sequence that is 94% identical to SEQ ID NO: 24. In some embodiments, I-TevI contains an amino acid sequence that is 95% identical to SEQ ID NO: 24. In some embodiments, I-TevI contains an amino acid sequence that is 96% identical to SEQ ID NO: 24. In some embodiments, I-TevI contains an amino acid sequence that is 97% identical to SEQ ID NO: 24. In some embodiments, I-TevI contains an amino acid sequence that is 98% identical to SEQ ID NO: 24.In some embodiments, I-TevI contains an amino acid sequence that is 99% identical to SEQ ID NO: 24. In some embodiments, I-TevI contains an amino acid sequence that is 100% identical to SEQ ID NO: 24.
[0204] In some embodiments, the I-TevI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 25. In some embodiments, I-TevI contains an amino acid sequence that is 90% identical to SEQ ID NO: 25. In some embodiments, I-TevI contains an amino acid sequence that is 91% identical to SEQ ID NO: 25. In some embodiments, I-TevI contains an amino acid sequence that is 92% identical to SEQ ID NO: 25. In some embodiments, I-TevI contains an amino acid sequence that is 93% identical to SEQ ID NO: 25. In some embodiments, I-TevI contains an amino acid sequence that is 94% identical to SEQ ID NO: 25. In some embodiments, I-TevI contains an amino acid sequence that is 95% identical to SEQ ID NO: 25. In some embodiments, I-TevI contains an amino acid sequence that is 96% identical to SEQ ID NO: 25. In some embodiments, I-TevI contains an amino acid sequence that is 97% identical to SEQ ID NO: 25. In some embodiments, I-TevI contains an amino acid sequence that is 98% identical to SEQ ID NO: 25. In some embodiments, I-TevI contains an amino acid sequence that is 99% identical to SEQ ID NO: 25. In some embodiments, I-TevI contains an amino acid sequence that is 100% identical to SEQ ID NO: 25.
[0205] In some embodiments, the GIY-YIG nuclease is the I-BmoI nuclease. I-BmoI is encoded by a group I intron that interrupts the thymidylate synthase (TS) gene (thyA) in Bacillus mojavensis s87-18. In some embodiments, the I-BmoI nuclease contains the amino acid sequence MKSGVYKITNKNTGKFYIGSSEDCESRLKVHFRNLKNNRHINRYLNNSFNKHGEQVFIGEVIHILPIEEAIAKEQWYIDNFYEEMYNISKSAYHGGDLTSYHPDKRNIILKRADSLKKVYLKMTSEEKAKRWQCVQGENNPMFGRKHTETTKLKISNHNKLYYSTHKNPFKGKKHSEESKTKLSEYASQRVGEKNPFYGKTHSDEFKTYMSKKFKGRKPKNSRPVIIDGTEYESATEASRQLNVVPATILHRIKSKNEKYSGYFYK (SEQ ID NO: 737).
[0206] In some embodiments, the I-BmoI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 737.
[0207] In some embodiments, the GIY-YIG nuclease is an Eco29kI nuclease. An exemplary Eco29kI nuclease is an Eco29kI restriction endonuclease derived from Ectothiorhodospira magna.
[0208] In some embodiments, the Eco29kI nuclease contains the amino acid sequence MTDDKVIPFNPLDKRHLGESVGQAMLRQPVVPMAKLSRFRGAGIYAIYYTGNFEAYQGIAACNRDDRFAAPIYVGKAVPKGARKGSGSLDTSPGAVLFSRLAQHGKSIQEVKNLDINDFYCRYLIVDDIWIPLGESLLIAKFNPLWNSVLDGFGNHDPGKGRHAGLRPRWDVVHPGRAWAGRCQAREETAEKILREAVNFLASNPPPGDW (SEQ ID NO: 738).
[0209] In some embodiments, the Eco29kI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 738.
[0210] In some embodiments, the Eco29kI nuclease is an Eco29kI restriction endonuclease derived from Escherichia coli.
[0211] In some embodiments, Eco29kI nuclease is
[0212] Contains the amino acid sequence MTGKVIPFNPLDKQNLGASVAEALLSKDAHPLEELTSFQGAGIYAIYYTGDHPAYRQLAELNRDGQFRLPIYVGKAVPAGARMGLTNPDKVGNVLFRRLKEHAESIRAAENLSIEDFYCRFLVVDDIWIPLGESLVISRFKPIWNSSIDGFGNHDPGKHRYTGLRPRWDFMHPGRGWAQNLRERDETVDELIRDSIQYLQNLPPCLAQKFIEAEGD (Sequence ID 739).
[0213] In some embodiments, the Eco29kI nuclease contains an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 739.
[0214] Linker In some embodiments, the chimeric nuclease includes a linker domain. In some embodiments, the linker domain is located between the CRISPR Cas nuclease and the GIY-YIG nuclease. In some embodiments, the linker domain is located between the CRISPR Cas nuclease domain and the GIY-YIG nuclease domain.
[0215] In some embodiments, the linker comprises an I-TevI linker (amino acids 93-150 of I-TevI (SEQ ID NO: 155)). Alternatively, the linker may further comprise a flexible amino acid linker comprising 10-100 amino acids. Such linkers may be unstructured or may comprise a Gly-Ser linker.
[0216] In certain embodiments, the linker includes an amino acid substitution selected from one or more positions corresponding to residues T95S, S101Y, A119D, K120N, K135N, K135R, P126S, D127K, N140S, T147I, Q158R, A161V, or S165G of SEQ ID NO: 155. In certain embodiments, the linker includes a substitution selected from one or more positions corresponding to residues T95S, V117F, K135R, N140S, or Q158R of SEQ ID NO: 155.
[0217] In some embodiments, the linker includes a mutation selected from the mutations corresponding to any one of the following mutations of SEQ ID NO: T95S, S101Y, A119D, K120N, K135N, K135R, P126S, D127K, N140S, T147I, Q158R, A161V, V117F, S165G, or any combination thereof.
[0218] In some embodiments, the linker domain includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 26. In some embodiments, the linker domain includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 27. In some embodiments, the linker domain includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 28. In some embodiments, the linker domain includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 29. In some embodiments, the linker domain includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 30. In some embodiments, the linker domain includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 31. In some embodiments, the linker domain includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 32. In some embodiments, the linker domain includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 33. In some embodiments, the linker domain includes an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 34.In some embodiments, the linker domain comprises an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 35.
[0219] In some embodiments, the linker further comprises a hinge sequence. In some embodiments, the hinge sequence comprises the amino acid sequence GGSGGTGGSG.
[0220] In some embodiments, the I-TevI domain and the linker sequence comprise SEQ ID NO: 225. In some embodiments, the I-TevI domain and the linker sequence comprise an amino acid sequence that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 225.
[0221] In some embodiments, the I-TevI is I-TevI[R27A], and the Cas9 is SaCas9[D10A+H557A].
[0222] In some embodiments, the I-TevI is I-TevI WT, and the Cas9 is CasX. In some embodiments, the chimeric nuclease is
[0223] In some embodiments, the chimeric nuclease comprises an amino acid sequence selected from Table 2.
[0224] [Table 3-1]
[0225] [Table 3-2]
[0226] [Table 3-3]
[0227] [Table 3-4]
[0228] [Table 3-5]
[0229] [Table 3-6]
[0230] [Table 3-7]
[0231] [Table 3-8]
[0232] [Table 3-9]
[0233] Table 3-10
[0234] Table 3-11
[0235] Table 3-12
[0236] Table 3-13
[0237] Table 3-14
[0238] Table 3-15
[0239] Table 3-16
[0240] Table 3-17
[0241] In some embodiments, the chimeranucleases include chimeranucleases selected from Table 2. In some embodiments, the chimeranucleases include those selected from SEQ ID NOs: 192-223. In some embodiments, the chimeranucleases include amino acid sequences that are 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 192-223. In some embodiments, the chimeranucleases have cleavage activity against I-TevI recognition sites in double-stranded DNA. In some embodiments, the chimeranucleases have cleavage activity against I-TevI recognition sites in double-stranded DNA and cleavage activity against Cas9 recognition sites in double-stranded DNA. In some embodiments, the chimeranucleases have binding activity against I-TevI recognition sites in double-stranded DNA. In some embodiments, the chimeric nuclease has cleavage activity against the I-TevI recognition site in double-stranded DNA and binding activity against the Cas9 recognition site in double-stranded DNA. In some embodiments, the chimeric nuclease has nickase activity against the I-TevI recognition site in double-stranded DNA and nickase activity against the Cas9 recognition site in double-stranded DNA.
[0242] Guide polynucleotides In some embodiments, the chimeric nuclease or chimeric nuclease system further comprises a guide polynucleotide. In some embodiments, the guide polynucleotide is a guide RNA. In certain embodiments, the guide RNA targets a genomic target site within a cell. In certain embodiments, the guide RNA targets disease-causing mutations in mammalian cells. In certain embodiments, the mammalian cells are human cells. In certain embodiments, the guide RNA targets bacteria or biological sequences.
[0243] In some embodiments, the guide RNA is either a single guide RNA (sgRNA) (e.g., a fusion of crRNA and tracrRNA) or a dual guide RNA (crRNA and tracrRNA).
[0244] A single guide RNA (sgRNA) may include, in the 5' to 3' direction, an optional spacer extension sequence, a spacer sequence, a minimal CRISPR repeat sequence, a single guide linker, a minimal tracrRNA sequence, a 3' tracrRNA sequence, and / or an optional tracrRNA extension sequence. The scaffold sequence may include all elements except the spacer sequence. The optional tracrRNA extension may include elements that contribute to further functionality (e.g., stability) of the guide RNA. The single guide linker can ligate the minimal CRISPR repeat and the minimal tracrRNA sequence to form a hairpin structure. Any tracrRNA extension may include one or more hairpins. In certain embodiments, this disclosure provides an sgRNA including a spacer sequence and a tracrRNA sequence. In some embodiments, the crRNA and / or tracrRNA are Staphylococcus aureus, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus desulforudis, Clostridium botulinum, Clostridium difficile, Fine goldia magna, Natranaerobius thermophilus, Pelotomaculum the rmopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira It originates from sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, or Acaryochloris marina Cas9 system.
[0245] In some embodiments, variants of the guide RNA scaffold sequence may be used. In some embodiments, the scaffold sequence following the guide sequence at its 3' end is SEQ ID NO: 75.
[0246] In some embodiments, the guide RNA includes the nucleic acid sequence described in SEQ ID NO: 53. In some embodiments, the guide RNA includes a nucleic acid sequence that is 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 53. In some embodiments, the guide RNA includes a nucleic acid sequence that is 95% identical to SEQ ID NO: 53. In some embodiments, the guide RNA includes a nucleic acid sequence that is 96% identical to SEQ ID NO: 53. In some embodiments, the guide RNA includes a nucleic acid sequence that is 97% identical to SEQ ID NO: 53. In some embodiments, the guide RNA includes a nucleic acid sequence that is 98% identical to SEQ ID NO: 53. In some embodiments, the guide RNA includes a nucleic acid sequence that is 99% identical to SEQ ID NO: 53. In some embodiments, the guide RNA includes a nucleic acid sequence that is 100% identical to SEQ ID NO: 53.
[0247] In some embodiments, the guide RNA includes the nucleic acid sequence described in SEQ ID NO: 76. In some embodiments, the guide RNA includes a nucleic acid sequence that is 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 76. In some embodiments, the guide RNA includes a nucleic acid sequence that is 95% identical to SEQ ID NO: 76. In some embodiments, the guide RNA includes a nucleic acid sequence that is 96% identical to SEQ ID NO: 76. In some embodiments, the guide RNA includes a nucleic acid sequence that is 97% identical to SEQ ID NO: 76. In some embodiments, the guide RNA includes a nucleic acid sequence that is 98% identical to SEQ ID NO: 76. In some embodiments, the guide RNA includes a nucleic acid sequence that is 99% identical to SEQ ID NO: 76. In some embodiments, the guide RNA includes a nucleic acid sequence that is 100% identical to SEQ ID NO: 76.
[0248] In some embodiments, the guide RNA includes the nucleic acid sequence described in SEQ ID NO: 77. In some embodiments, the guide RNA includes a nucleic acid sequence that is 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 77. In some embodiments, the guide RNA includes a nucleic acid sequence that is 95% identical to SEQ ID NO: 77. In some embodiments, the guide RNA includes a nucleic acid sequence that is 96% identical to SEQ ID NO: 77. In some embodiments, the guide RNA includes a nucleic acid sequence that is 97% identical to SEQ ID NO: 77. In some embodiments, the guide RNA includes a nucleic acid sequence that is 98% identical to SEQ ID NO: 77. In some embodiments, the guide RNA includes a nucleic acid sequence that is 99% identical to SEQ ID NO: 77. In some embodiments, the guide RNA includes a nucleic acid sequence that is 100% identical to SEQ ID NO: 77.
[0249] In some embodiments, the guide RNA includes the nucleic acid sequence described in SEQ ID NO: 78. In some embodiments, the guide RNA includes a nucleic acid sequence that is 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 78. In some embodiments, the guide RNA includes a nucleic acid sequence that is 95% identical to SEQ ID NO: 78. In some embodiments, the guide RNA includes a nucleic acid sequence that is 96% identical to SEQ ID NO: 78. In some embodiments, the guide RNA includes a nucleic acid sequence that is 97% identical to SEQ ID NO: 78. In some embodiments, the guide RNA includes a nucleic acid sequence that is 98% identical to SEQ ID NO: 78. In some embodiments, the guide RNA includes a nucleic acid sequence that is 99% identical to SEQ ID NO: 78. In some embodiments, the guide RNA includes a nucleic acid sequence that is 100% identical to SEQ ID NO: 78.
[0250] In some embodiments, the guide RNA includes the nucleic acid sequence described in SEQ ID NO: 79. In some embodiments, the guide RNA includes a nucleic acid sequence that is 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 79. In some embodiments, the guide RNA includes a nucleic acid sequence that is 95% identical to SEQ ID NO: 79. In some embodiments, the guide RNA includes a nucleic acid sequence that is 96% identical to SEQ ID NO: 79. In some embodiments, the guide RNA includes a nucleic acid sequence that is 97% identical to SEQ ID NO: 79. In some embodiments, the guide RNA includes a nucleic acid sequence that is 98% identical to SEQ ID NO: 79. In some embodiments, the guide RNA includes a nucleic acid sequence that is 99% identical to SEQ ID NO: 79. In some embodiments, the guide RNA includes a nucleic acid sequence that is 100% identical to SEQ ID NO: 79.
[0251] In some embodiments, the guide RNA includes the nucleic acid sequences described in SEQ ID NOs. 76-85. In some embodiments, the guide RNA includes the nucleic acid sequences described in SEQ ID NOs. 234-235.
[0252] In some embodiments, the guide RNA is modified. In certain embodiments, the guide RNA includes one or more of the following: non-natural nucleoside bonds, nucleic acid mimetic, modified sugar moiety, and modified nucleobase. In some embodiments, the guide RNA includes modified nucleotides. In certain embodiments, the modified nucleotides include 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl derivative of adenine, 6-methyl derivative of guanine, 2-propyl derivative of adenine, 2-propyl derivative of guanine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 6-azouracil, 6-azocytosine, 6-azocymine, pseudouracil, 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl, 5-halo, 5-bromo, 5-trifluoromethyl, 5-substituted uracil, 5-substituted cytosine, 7-methylguanine, 7-methyladenine, 2-F-adenine, and 2-amino-adenine. , comprising one or more of the following: 8-azaguanine, 8-azaadenine, 7-deazaguanine, 7-deazaadenine, 3-deazaguanine, 3-deazaadenine, tricyclic pyrimidine, phenoxazinecytidine, phenothiazinecytidine, substituted phenoxazinecytidine, carbazolecytidine, pyridroinchlorcytidine, 7-creazaadenine, 7-dreazaguanosine, 2-aminopyridine, 2-pyridone, 5-substituted pyrimidine, 6-azapyrimidine, N-2, N-6, or 0-6 substituted purines, 2-aminopropyladenine, 5-propynyluracil, or 5-propynylcytosine.
[0253] In certain embodiments, the non-natural nucleoside bond comprises one or more of the following: phosphorothioates, phosphoramidates, nonphosphodiesters, heteroatoms, chiral phosphorothioates, phosphorodithioates, phosphotryesters, aminoalkylphosphotryesters, 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphinates, 3'-aminophosphoramidates, aminoalkylphosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotryesters, selenophosphates, and boranophosphates. In certain embodiments, the nucleic acid mimetic comprises one or more of the following: peptide nucleic acids (PNAs), morpholino nucleic acids, cyclohexenyl nucleic acids (CeNAs), or locked nucleic acids (LNAs). In certain embodiments, the modified sugar moiety comprises one or more of 2'-O-(2-methoxyethyl), 2'-dimethylaminooxyethoxy, 2'-dimethylaminoethoxyethoxy, 2'-O-methyl, or 2'-fluoro.
[0254] In some embodiments, the chimeric nuclease targets a site within a gene. In some embodiments, the chimeric nuclease targets a site in an intron or exon of a gene. In some embodiments, the gene is B2M or AAVS1. In some embodiments, the gene is CFTR. In some embodiments, the gene is SERPINA1 (encoding alpha-1-antitrypsin). In some embodiments, the gene is DMPK. In some embodiments, the gene is C9ORF72. In some embodiments, the guide RNA includes a spacer that targets the target site in B2M or AAVS1. In some embodiments, the guide RNA includes a spacer that targets the target site in CFTR. In some embodiments, the guide RNA includes a spacer that targets the target site in SERPINA1. In some embodiments, the guide RNA includes a spacer that targets the target site in DMPK. In some embodiments, the guide RNA includes a spacer that targets the target site in C9ORF72.
[0255] In some embodiments, the gene is C9ORF72 or DMPK. In some embodiments, the spacer includes the nucleic acid sequences described in SEQ ID NOs. 20-23.
[0256] In some embodiments, the spacer includes the nucleic acid sequence described in SEQ ID NO: 41 or SEQ ID NO: 51. In some embodiments, the spacer includes the nucleic acid sequences described in SEQ ID NOs: 76 to 85.
[0257] In this sequence, please understand that if the nucleic acid sequence is DNA, T represents thymine, and if the nucleic acid sequence is RNA, T represents uracil.
[0258] CFTR The cystic fibrosis transmembrane conductance regulator (CFTR) gene encodes a member of the ATP-binding cassette (ABC) transporter superfamily. The encoded protein functions as a chloride channel, unique among members of this protein family, and regulates the secretion and absorption of ions and water in epithelial tissues. Channel activation is mediated by a cycle of regulatory domain phosphorylation, ATP binding by the nucleotide-binding domain, and ATP hydrolysis. Mutations in this gene cause cystic fibrosis (CF), the most common fatal genetic disorder in people of Nordic descent. Mutations in the CFTR gene (GRCh38.p14 GCF_000001405.40,NM_000492.4,NC_000007.14 Reference GRCh38.p14) can result in suboptimal ion transport and fluid retention, leading to marked clinical manifestations of abnormal mucus thickening in the lungs and pancreatic dysfunction. DeltaF508, the most frequently occurring mutation in cystic fibrosis, leads to impaired folding and transport of the encoded protein. CFTR proteins are found across a wide range of organs, including the pancreas, kidneys, liver, lungs, gastrointestinal tract, and genitals, making CF a multi-organ disease. In the lungs, dysfunctional CFTRs interfere with mucociliary clearance, making the organ more susceptible to bacterial infection and inflammation, ultimately leading to airway obstruction, respiratory failure, and premature death. CF remains the most common and fatal genetic disorder in the Caucasian population, with an estimated 70,000 to 100,000 patients worldwide, highlighting the urgent need for the development of better treatments.
[0259] Mutations in CFTR include, but are not limited to, c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), or c.3846G>A (p.Trp1282Ter). In some embodiments, the guide polynucleotide targets a site within the intronic or exonic region of the CFTR gene.
[0260] In some embodiments, the target region of the chimeric nuclease is the CFTR F508-R553X region on Chr7:117,559,464-117,587,836 relative to the hg38 genome. In some embodiments, the chimeric nuclease cleaves directly adjacent to the 5', 3', and / or inner regions of the Chr7:117,559,464-117,587,836 region.
[0261] SERPINA1 The Serpin Family A Member 1 protein, encoded by the SERPINA1 gene (GRCh38.p14(GCF_000001405.40), NG_008290.1RefSeqGene), is a serine protease inhibitor belonging to the serpin superfamily. Its targets include elastase, plasmin, thrombin, trypsin, chymotrypsin, and plasminogen activator. The SERPINA1 protein is produced in the liver and bone marrow by lymphocytes and monocytes in lymphoid tissues, as well as by Paneth cells in the intestine. Deficiency of this gene is associated with chronic obstructive pulmonary disease, emphysema, and chronic liver disease, and plays a role in alpha-1 antitrypsin deficiency. Several transcriptional variants encoding the same protein have been found for this gene. The SERPINA1 E342K mutation is associated with alpha-1 antitrypsin deficiency.
[0262] In some embodiments, the guide polynucleotide targets a site within an intronic or exonic region of the SERPINA1 gene.
[0263] DMPK DM1 protein kinase (DMPK) is a serine-threonine kinase closely related to other kinases that interact with members of the Rho family of small GTPases. Substrates for this enzyme include myogenin, the β-subunit of L-type calcium channels, and phosphoremane. The DMPK gene (NG_009784.1 RefSeqGene, GRCh38.p14(GCF_000001405.40)) is located at 19q13.32. The 3' untranslated region of this gene contains 5–38 copies of (CTG)n*(CAG)n trinucleotide repeats. Extension of this unstable motif to 40–5,000 copies leads to myotonic dystrophy type I, with increasing severity with increasing repeat element copy number. Repeat extension is associated with the condensation of local chromatin structures that disrupt gene expression in this region. Several alternatively spliced transcript variants of this gene have been described. Removal of excess (CTG)n*(CAG)n trinucleotide repeat sequences can stabilize the 3' untranslated region of DMPK in cells. In some embodiments, one or more guide polynucleotide target (CAG)n sites in the 3' untranslated region of the DMPK gene are used for excision.
[0264] In some embodiments, the spacer of the guide polynucleotide includes the nucleic acid sequence described in SEQ ID NOs: 22-23. In some embodiments, the spacer of the guide polynucleotide includes the nucleic acid sequence described in any of SEQ ID NOs: 136-147.
[0265] In some embodiments, the chimeric nuclease system targets the intron region of DMPK as described in SEQ ID NO: 4.
[0266] In some embodiments, the chimeric nuclease system is a dual-guide chimeric nuclease. In some embodiments, the chimeric nuclease system is a dual-guide Tev-dCas9. Exemplary excision sequences in DMPK that can be excised using the dual-guide nucleases of this disclosure are listed in Table 3A. Exemplary regions for excision are listed in Table 3B. In some embodiments, the excised sequence includes sequence number 236. In some embodiments, the excised sequence includes any one of sequence numbers 237-247.
[0267] [Table 4-1]
[0268] [Table 4-2]
[0269] [Table 5]
[0270] Table 4A shows an exemplary dual-guide chimeric nuclease that targets the nucleic acid sequence of sequence number 236 in the DPMK gene for excision.
[0271] [Table 6-1]
[0272] [Table 6-2]
[0273] [Table 6-3]
[0274] Table 4B shows an exemplary dual-guide chimeric nuclease that targets the nucleic acid sequence of sequence number 240 in the DPMK gene for excision.
[0275] [Table 7-1]
[0276] [Table 7-2]
[0277] [Table 7-3]
[0278] [Table 8]
[0279] C9ORF72 The C9ORF72-SMCR8 complex subunit (C9ORF72) plays a crucial role in regulating endosomal trafficking and has been shown to interact with Rab proteins involved in autophagy and endocytotic transport. Expansion of GGGGCC repeats from 2–22 copies to 700–1600 copies in the intron sequence between alternative 5' exons in transcripts derived from this gene is associated with 9p-linked ALS (amyotrophic lateral sclerosis) and FTD (frontotemporal dementia). Studies suggest that hexanucleotide elongation may lead to selective stabilization of repeat-containing premRNA and accumulation of potentially pathogenic insoluble dipeptide repeat protein aggregates in FTD-ALS patients (PMID:23393093). Alternative splicing results in multiple transcript variants encoding different isoforms.
[0280] In some embodiments, the gene is C9ORF72. In some embodiments, the spacer includes the nucleic acid sequence described in SEQ ID NO: 20 or 21.
[0281] In some embodiments, the spacer includes a nucleic acid sequence described in any one of sequence numbers 132 to 135.
[0282] In some embodiments, the chimeric nuclease targets the intronic region of C9ORF72 as shown in SEQ ID NO: 2 or 3. In some embodiments, the removed C9orf72 G4C2 repeats include [G4C2]24+ repeats.
[0283] [Table 9]
[0284] [Table 10-1]
[0285] [Table 10-2]
[0286] [Table 10-3]
[0287] In some embodiments, the gene is C9ORF72. In some embodiments, the spacer includes the nucleic acid sequences described in SEQ ID NOs: 20-23.
[0288] In some embodiments, the chimeric nuclease system is a dual-guided chimeric nuclease system. The dual-guided nuclease system comprises two chimeric nucleases that bind and cleave at two locations in the cell's genome, preferably on the same chromosome. In some embodiments, the two chimeric nucleases are the same chimeric nuclease, having two different guide polynucleotides that can bind to each other upstream or downstream on the DNA sequence. In some embodiments, the two chimeric nucleases are different chimeric nucleases, having two different guide polynucleotides.
[0289] In some embodiments, the chimeric nuclease system comprises a first chimeric nuclease and a first guide RNA, and a second chimeric nuclease and a second guide RNA. In some embodiments, the first chimeric nuclease and the first guide RNA target a first target site, and the second chimeric nuclease and the second guide RNA target a second target site. In some embodiments, the first chimeric nuclease comprises an active I-TevI nuclease and an inactive dCas nuclease. In some embodiments, the second chimeric nuclease comprises an active I-TevI nuclease and an inactive dCas nuclease. In some embodiments, the first chimeric nuclease comprises an active I-TevI nuclease and an inactive dCas nuclease, and the second chimeric nuclease comprises an active I-TevI nuclease and an inactive dCas nuclease.
[0290] In some embodiments, the first and second guide RNAs target different target sites in the cell's genome. In some embodiments, the distance between the first and second target sites is 100 bases. In some embodiments, the distance between the first and second target sites is 200 bases. In some embodiments, the distance between the first and second target sites is 300 bases. In some embodiments, the distance between the first and second target sites is 400 bases. In some embodiments, the distance between the first and second target sites is 500 bases. In some embodiments, the distance between the first and second target sites is 600 bases. In some embodiments, the distance between the first and second target sites is 700 bases. In some embodiments, the distance between the first and second target sites is 800 bases. In some embodiments, the distance between the first and second target sites is 900 bases. In some embodiments, the distance between the first target site and the second target site is 1000 bases. In some embodiments, the distance between the first target site and the second target site is 2000 bases. In some embodiments, the distance between the first target site and the second target site is 3000 bases. In some embodiments, the distance between the first target site and the second target site is 4000 bases. In some embodiments, the distance between the first target site and the second target site is 5000 bases. In some embodiments, the distance between the first target site and the second target site is 6000 bases. In some embodiments, the distance between the first target site and the second target site is 7000 bases. In some embodiments, the distance between the first target site and the second target site is 8000 bases. In some embodiments, the distance between the first target site and the second target site is 9000 bases. In some embodiments, the distance between the first target site and the second target site is 10,000 bases. In some embodiments, the distance between the first target site and the second target site is 11,000 bases. In some embodiments, the distance between the first target site and the second target site is 12,000 bases.In some embodiments, the distance between the first target site and the second target site is 13,000 bases. In some embodiments, the distance between the first target site and the second target site is 14,000 bases. In some embodiments, the distance between the first target site and the second target site is 15,000 bases. In some embodiments, the distance between the first target site and the second target site is 20,000 bases. In some embodiments, the distance between the first target site and the second target site is 25,000 bases. In some embodiments, the distance between the first target site and the second target site is 28,000 bases. In some embodiments, the distance between the first target site and the second target site is 30,000 bases. In some embodiments, the distance between the first target site and the second target site is greater than 30,000 bases. In some embodiments, the distance between the first target site and the second target site is between 100 and 1,000 bases. In some embodiments, the distance between the first target site and the second target site is 100 to 10,000 bases. In some embodiments, the distance between the first target site and the second target site is 100 to 20,000 bases. In some embodiments, the distance between the first target site and the second target site is 100 to 30,000 bases. In some embodiments, the distance between the first target site and the second target site is 1,000 to 10,000 bases. In some embodiments, the distance between the first target site and the second target site is 1,000 to 20,000 bases. In some embodiments, the distance between the first target site and the second target site is 1,000 to 30,000 bases. In some embodiments, the distance between the first target site and the second target site is 5,000 to 30,000 bases.
[0291] In some embodiments, the first guide RNA and the second guide RNA target the same strand of genomic DNA. In some embodiments, the first guide RNA and the second guide RNA target the opposite strand of genomic DNA. In some embodiments, the first guide RNA targets the first chimeric nuclease to the first I-TevI target site, causing it to cleave at the first I-TevI target site in the cell's genome, while the second guide RNA targets the second chimeric nuclease to the second I-TevI target site in the cell's genome, causing it to cleave at the second I-TevI target site, and the cleavage generates a nucleotide overhang at the second I-TevI target site.
[0292] In some embodiments, a first guide RNA targets a first chimeric nuclease to a Cas9 target site, causing it to cleave at the Cas9 target site in the cell's genome, while a second guide RNA targets a second chimeric nuclease to an I-TevI target site in the cell's genome, causing it to cleave at the I-TevI target site, which generates a nucleotide overhang at the I-TevI target site.
[0293] In some embodiments, the second guide RNA targets the donor polynucleotide target site with its 5' region. In some embodiments, the 3' end of the donor polynucleotide is complementary to the overhang resulting from the cleavage of the second I-TevI at the second I-TevI target site.
[0294] In some embodiments, the guide RNA further comprises a donor polynucleotide. In some embodiments, the donor polynucleotide is bound to the 3' end of the guide RNA. In some embodiments, the donor polynucleotide comprises a repair template. In some embodiments, the repair template comprises a 2-nucleotide overhang compared to the target site. In some embodiments, the repair template is 36 nucleotides long.
[0295] In some embodiments, the guide RNA further comprises one or more synthetic tRNA sequences encoding a trans-acting ribozyme sequence. In some embodiments, the tRNA is located at the 5' end of the guide RNA. In some embodiments, the tRNA is located at the 5' end of the guide RNA. In some embodiments, the ribozyme can cleave the tRNA from the guide RNA.
[0296] nucleic acid This specification provides, in particular, nucleic acids encoding chimeric nucleases and chimeric nuclease systems described herein.
[0297] In one embodiment, the nucleic acid comprises a polynucleotide encoding a chimeric nuclease containing a GIY-YIG nuclease domain and an RNA-inducible nuclease domain, a polynucleotide encoding a guide RNA (gRNA), and a polynucleotide encoding tRNA. In some embodiments, the nucleic acid comprises a polynucleotide encoding a chimeric nuclease containing an I-TevI nuclease domain and an RNA-inducible nuclease domain, a polynucleotide encoding a guide RNA (gRNA), and a polynucleotide encoding tRNA.
[0298] Exemplary designs of nucleic acids encoding chimeric nucleases and chimeric nuclease systems described herein are shown in Figures 1, 2, 26A, 26B, 26C, and 27A.
[0299] In another embodiment, the nucleic acid includes a polynucleotide encoding a guide RNA and a polynucleotide encoding a tRNA.
[0300] In some embodiments, the nucleic acid further comprises one or more additional guide RNAs.
[0301] In some embodiments, the nucleic acid is DNA or RNA. In some embodiments, the DNA is circular plasmid DNA, linear double-stranded DNA, single-stranded DNA, or chimeric RNA and DNA.
[0302] In some embodiments, RNA is mRNA. In some embodiments, mRNA comprises a nucleic acid mimetic selected from the group consisting of peptide nucleic acids (PNA), morpholino nucleic acids, cyclohexenyl nucleic acids (CeNA), and locked nucleic acids (LNA). In some embodiments, mRNA comprises a modified sugar moiety, which is optionally selected from the group consisting of N1-methylpseudridine, 9-methyladenine, 2'-O-(2-methoxyethyl), 2'-dimethylaminooxyethoxy, 2'-dimethylaminoethoxyethoxy, 2'-O-methyl, and 2'-fluoro. In some embodiments, the mRNA contains modified nucleic acid bases, which are optionally 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl derivative of adenine, 6-methyl derivative of guanine, 2-propyl derivative of adenine, 2-propyl derivative of guanine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 5-halouracil, 5-halosyl Tosine, 5-propynyluracil, 5-propynylcytosine, 6-azouracil, 6-azocytosine, 6-azothimine, pseudouracil, 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl, 5-halo, 5-bromo, 5-trifluoromethyl, 5-substituted uracil, 5-substituted cytosine, 7-methylguanine, 7-methyladenine, 2-F-adenine, 2-amino-adenine Selected from the group consisting of 8-azaguanine, 8-azaadenine, 7-deazaguanine, 7-deazaadenine, 3-deazaguanine, 3-deazaadenine, tricyclic pyrimidine, phenoxazinecytidine, phenothiazinecytidine, substituted phenoxazinecytidine, carbazolecytidine, pyridoindolecytidine, 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine, 2-pyridone, 5-substituted pyrimidine, 6-azapyrimidine, N-2, N-6, or O-6 substituted purines, 2-aminopropyladenine, 5-propynyluracil, or 5-propynylcytosine.In some embodiments, the mRNA contains internucleoside bonds that are not naturally occurring or are unnatural, selected from the group consisting of phosphorothioates, phosphoramidates, nonphosphodiesters, heteroatoms, chiral phosphorothioates, phosphorodithioates, phosphotryesters, aminoalkylphosphotryesters, 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphinates, 3'-aminophosphoramidates, aminoalkylphosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotryesters, selenophosphates, or boranophosphates.
[0303] In some embodiments, the nucleic acid is approximately 5 kb in length. In some embodiments, the nucleic acid is less than 5 kb in length. In some embodiments, the nucleic acid may be packaged in AAV.
[0304] In some embodiments, the nucleic acid is transcribed into a single mRNA encoding a chimeric nuclease, one or more guide RNAs, and any additional donor polynucleotides. In some embodiments, the single mRNA is processed into individual components at one or more cleavage or trans-cleavage sites by a ribozyme or cellular mRNA processing mechanism.
[0305] In some embodiments, the cleavage site is a ribozyme cleavage site. In some embodiments, the cleavage site is a tRNA cleavage site. In some embodiments, the guide RNA sequence is directly adjacent to the 5' cleavage site. In some embodiments, the guide RNA sequence is directly adjacent to the 3' cleavage site.
[0306] In some embodiments, the guide RNA sequence is directly adjacent to the 5' and 3' cleavage sites. Cleavage sites in mRNA can increase the efficiency and quantity of the guide RNA sequence available in the cell after transcription.
[0307] In some embodiments, the nucleic acid includes a promoter sequence (if the nucleic acid is DNA) or an untranslated region (UTR) sequence (if the nucleic acid is RNA) at the 5' to 3' end, a chimeric nuclease sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA3 sequence, a cleavage site, and a poly(A) sequence.
[0308] In some embodiments, the nucleic acid includes a promoter / untranslated region (UTR) sequence, a chimeric nuclease sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, a gRNA2 sequence, a trans-acting target site / cleavage site, and a polyA sequence at its 5' to 3' end.
[0309] In some embodiments, the nucleic acid includes a promoter / untranslated region (UTR) sequence, a chimeric nuclease sequence, a transactive target site / cleavage site, a gRNA1 sequence, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, and a polyA sequence at its 5' to 3' end.
[0310] In some embodiments, the nucleic acid includes a promoter or UTR sequence, a chimeric nuclease sequence, a MALAT sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA3 sequence, a cleavage site, and a polyA sequence at its 5' to 3' end.
[0311] In some embodiments, the nucleic acid includes a promoter / UTR sequence, a chimeric nuclease sequence, a MALAT sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, a gRNA2 sequence, a trans-acting target site / cleavage site, and a polyA sequence at its 5' to 3' end.
[0312] In some embodiments, the nucleic acid includes, at its 5' to 3' end, a promoter / untranslated region (UTR) sequence, a chimeric nuclease sequence, a MALAT sequence, a trans-acting target site / cleavage site, a gRNA1 sequence, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, and a polyA sequence.
[0313] In some embodiments, the nucleic acid includes, at its 5' to 3' end, a promoter sequence, a chimeric nuclease sequence, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, and a polyA sequence.
[0314] In some embodiments, the nucleic acid includes, at its 5' to 3' end, a promoter sequence, a chimeric nuclease sequence, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a Ribo2 sequence, and a polyA sequence.
[0315] In some embodiments, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a trans-acting target site / cleavage site, and a polyA sequence at its 5' to 3' end.
[0316] In some embodiments, the nucleic acid includes, at its 5' to 3' end, a promoter sequence, a chimeric nuclease sequence, a transactive target site / cleavage site, a gRNA1 sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a Ribo2 sequence, and a polyA sequence.
[0317] In some embodiments, the nucleic acid includes, at its 5' to 3' end, a promoter sequence, a chimeric nuclease sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a transactive target site / cleavage site, a Ribo1 sequence, a cleavage site, a gRNA2 sequence, a Ribo2 sequence, and a poly(A) sequence.
[0318] In some embodiments, the nucleic acid includes, at its 5' to 3' end, a promoter sequence, a chimeric nuclease sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a Ribo1 sequence, a trans-acting target site / cleavage site, a gRNA2 sequence, a Ribo2 sequence, and a poly(A) sequence.
[0319] In some embodiments, the nucleic acid includes, at its 5' to 3' end, a promoter sequence, a chimeric nuclease sequence, a transactive target site / cleavage site, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, and a polyA sequence.
[0320] In some embodiments, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, and a polyA sequence at its 5' to 3' end.
[0321] In some embodiments, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a Ribo2 sequence, and a polyA sequence at its 5' to 3' end.
[0322] In some embodiments, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a Ribo1 sequence, a cleavage site, a gRNA1 sequence, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a trans-acting target site / cleavage site, and a polyA sequence at its 5' to 3' end.
[0323] In some embodiments, the nucleic acid includes a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a trans-acting target site / cleavage site, a gRNA1 sequence, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a Ribo2 sequence, and a polyA sequence at its 5' to 3' end.
[0324] In some embodiments, the nucleic acid includes, at its 5' to 3' end, a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a trans-acting target site / cleavage site, a Ribo1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a Ribo2 sequence, and a polyA sequence.
[0325] In some embodiments, the nucleic acid includes, at its 5' to 3' end, a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a tRNA1 sequence, a cleavage site, a gRNA1 sequence, a cleavage site, a Ribo1 sequence, a trans-acting target site / cleavage site, a gRNA2 sequence, a Ribo2 sequence, and a polyA sequence.
[0326] In some embodiments, the nucleic acid includes, at its 5' to 3' end, a promoter sequence, a chimeric nuclease sequence, a MALAT sequence, a trans-acting target site / cleavage site, a Ribo1 sequence, a gRNA1 sequence, a cleavage site, a tRNA1 sequence, a cleavage site, a gRNA2 sequence, a cleavage site, a tRNA2 sequence, a cleavage site, and a polyA sequence.
[0327] In some embodiments, the transactive target site / cleavage site is a cleavage site of a ribozyme or tRNA.
[0328] In some embodiments, the nucleic acid further comprises one or more donor nucleotide sequences.
[0329] In some embodiments, the nucleic acid comprises the nucleic acid sequences described in SEQ ID NOs: 15-19. In some embodiments, the nucleic acid comprises the nucleic acid sequences described in SEQ ID NOs: 113, 114, 115, 116, 117, or 118.
[0330] In some embodiments, the nucleic acid further comprises a self-inactivating nucleic acid sequence. In some embodiments, the self-inactivating nucleic acid sequence includes nuclease-binding sites and cleavage sites between the nuclease promoter and the start codon, between the nuclease and the polyadenylated sequence, or between the 5' untranslated region.
[0331] Donor polynucleotides In some embodiments, the nucleic acid further comprises a nucleic acid sequence encoding a donor polynucleotide. In some embodiments, the nucleic acid comprises two or more donor polynucleotides. In some embodiments, the donor polynucleotide is a single-stranded nucleic acid. In some embodiments, the donor polynucleotide is a double-stranded nucleic acid. In some embodiments, the donor polynucleotide is a partially double-stranded nucleic acid. In some embodiments, the donor polynucleotide is RNA. In some embodiments, the first strand of the double-stranded donor polynucleotide is DNA and the second strand is RNA. In some embodiments, the donor polynucleotide is a cis-acting repair template encoding the sense strand of the repair sequence. In some embodiments, the donor polynucleotide is a cis-acting repair template encoding the complement or reverse complement of the repair sequence. In some embodiments, the donor polynucleotide is a trans-acting repair template of the repair sequence. In some embodiments, the donor polynucleotide comprises a cis-acting single-stranded RNA polynucleotide annealed to a complementary single-stranded DNA polynucleotide.
[0332] In some embodiments, the repair template is 20 bases long. In some embodiments, the repair template is 25 bases long. In some embodiments, the repair template is 30 bases long. In some embodiments, the repair template is 35 bases long. In some embodiments, the repair template is 40 bases long. In some embodiments, the repair template is 45 bases long. In some embodiments, the repair template is 50 bases long. In some embodiments, the repair template is 55 bases long. In some embodiments, the repair template is 60 bases long. In some embodiments, the repair template is 65 bases long. In some embodiments, the repair template is 70 bases long. In some embodiments, the repair template is 75 bases long. In some embodiments, the repair template is 80 bases long. In some embodiments, the repair template is 85 bases long. In some embodiments, the repair template is 90 bases long. In some embodiments, the repair template is 95 bases long. In some embodiments, the repair template is 100 bases long. In some embodiments, the repair template is 150 bases long. In some embodiments, the repair template is 200 base pairs long. In some embodiments, the repair template is 250 base pairs long. In some embodiments, the repair template is 300 base pairs long. In some embodiments, the repair template is 350 base pairs long. In some embodiments, the repair template is 400 base pairs long. In some embodiments, the repair template is 450 base pairs long. In some embodiments, the repair template is 500 base pairs long. In some embodiments, the repair template is 600 base pairs long. In some embodiments, the repair template is 700 base pairs long. In some embodiments, the repair template is 800 base pairs long. In some embodiments, the repair template is 900 base pairs long. In some embodiments, the repair template is 1000 base pairs long. In some embodiments, the repair template is at least 900 base pairs long.In some embodiments, the repair template is 20 to 100 base pairs long. In some embodiments, the repair template is 20 to 200 base pairs long. In some embodiments, the repair template is 20 to 300 base pairs long. In some embodiments, the repair template is 20 to 400 base pairs long. In some embodiments, the repair template is 20 to 500 base pairs long. In some embodiments, the repair template is 20 to 600 base pairs long. In some embodiments, the repair template is 20 to 700 base pairs long. In some embodiments, the repair template is 20 to 800 base pairs long. In some embodiments, the repair template is 20 to 900 base pairs long. In some embodiments, the repair template is 20 to 1000 base pairs long.
[0333] In some embodiments, the donor polynucleotide is bound to the guide RNA. In some embodiments, the donor polynucleotide is bound to the guide RNA via the 5' end of the donor polynucleotide and the 3' end of the guide RNA. In some embodiments, the donor polynucleotide is bound to the guide RNA via a linker. In some embodiments, the linker is bound to the 5' end of the donor polynucleotide and the 3' end of the guide RNA.
[0334] In some embodiments, the linker is 1 nucleotide long. In some embodiments, the linker is 2 nucleotide long. In some embodiments, the linker is 3 nucleotide long. In some embodiments, the linker is 4 nucleotide long. In some embodiments, the linker is 5 nucleotide long. In some embodiments, the linker is 6 nucleotide long. In some embodiments, the linker is 7 nucleotide long. In some embodiments, the linker is 8 nucleotide long. In some embodiments, the linker is 9 nucleotide long. In some embodiments, it is 10 nucleotide long. In some embodiments, the linker is 11 nucleotide long. In some embodiments, the linker is 12 nucleotide long. In some embodiments, the linker is 13 nucleotide long. In some embodiments, the linker is 14 nucleotide long. In some embodiments, the linker is 15 nucleotide long. In some embodiments, the linker is 16 nucleotide long. In some embodiments, the linker is 17 nucleotide long. In some embodiments, the linker is 18 nucleotide long. In some embodiments, the linker is 19 nucleotide long. In some embodiments, the linker is 20 nucleotides long. In some embodiments, the linker is 25 nucleotides long. In some embodiments, the linker is 30 nucleotides long.
[0335] In some embodiments, one or more donor polynucleotides are separated by one or more ribozyme cleavage sites.
[0336] In some embodiments, the target site of the donor polynucleotide is an I-TevI cleavage site. In some embodiments, the target site of the donor polynucleotide is a Cas cleavage site.
[0337] In some embodiments, one or more donor polynucleotides include a 2-nucleotide overhang compared to the corresponding 3' base adjacent to the target site. In some embodiments, one or more donor polynucleotides include a 3' nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 4-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 5-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 6-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 7-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include an 8-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 9-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 10-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include an 11-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 12-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 13-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 14-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 15-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 16-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 17-nucleotide overhang compared to the target site.In some embodiments, one or more donor polynucleotides include an 18-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 19-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include a 20-nucleotide overhang compared to the target site. In some embodiments, one or more donor polynucleotides include 2 to 16 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides include 2 to 18 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides include 2 to 20 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides include 4 to 16 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides include 4 to 18 nucleotide overhangs compared to the target site. In some embodiments, one or more donor polynucleotides include 4 to 20 nucleotide overhangs compared to the target site.
[0338] In some embodiments, the second guide RNA targets the donor polynucleotide target site with its 5' region. In some embodiments, the 3' end of the donor polynucleotide is complementary to the overhang resulting from the cleavage of the second I-TevI at the second I-TevI target site.
[0339] In some embodiments, one or more donor polynucleotides are at least 30 bases long. In some embodiments, one or more donor polynucleotides are at least 30 to 1500 bases long. In some embodiments, one or more donor polynucleotides are at least 30 to 100 bases long. In some embodiments, one or more donor polynucleotides are at least 40 to 90 bases long. In some embodiments, one or more donor polynucleotides are at least 50 to 80 bases long. In some embodiments, one or more donor polynucleotides are at least 60 to 70 bases long. In some embodiments, one or more donor polynucleotides are at least 100 to 1000 bases long. In some embodiments, one or more donor polynucleotides are at least 100 to 1500 bases long. In some embodiments, one or more donor polynucleotides are at least 200 to 800 bases long. In some embodiments, one or more donor polynucleotides are at least 300 to 700 bases long. In some embodiments, one or more donor polynucleotides are at least 400 to 600 bases long. In some embodiments, one or more donor polynucleotides are at least 500 base pairs long.
[0340] In some embodiments, a donor polynucleotide can be inserted into the AAVS1 safe harbor site. The nucleic acid sequence of the AAVS1 integration site is shown in SEQ ID NO: 1.
[0341] In some embodiments, the donor polynucleotide comprises a donor polynucleotide sequence selected from Table 8.
[0342] In some embodiments, the donor polynucleotide comprises one of the nucleic acid sequences described in SEQ ID NOs: 12-14.
[0343] In some embodiments, the donor polynucleotide comprises one of the nucleic acid sequences described in SEQ ID NOs. 42-45.
[0344] In some embodiments, the donor polynucleotide includes B2M, which inactivates the sequence of SEQ ID NO: 52.
[0345] In some embodiments, the donor polynucleotide may be inserted into the CFTR gene. In some embodiments, the donor polynucleotide may be used to repair the CFTR gene. In some embodiments, the donor polynucleotide comprises the nucleic acid sequence of SEQ ID NOs. 67–69. In some embodiments, the donor polynucleotide comprises a nucleic acid sequence that is 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs. 67–69 or 71–74.
[0346] In some embodiments, a donor polynucleotide can be inserted into the SERPINA1 gene. In some embodiments, a donor polynucleotide can be used to repair the SERPINA1 gene. In some embodiments, the donor polynucleotide includes SERPINA1 sequence SEQ ID NO: 70. In some embodiments, the donor polynucleotide includes a nucleic acid sequence that is 95%, 96%, 97%, 98%, 99%, or 100% identical to either SEQ ID NO: 70 or 75. In some embodiments, the donor polynucleotide includes a nucleic acid sequence that is 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 226-229. In some embodiments, the donor polynucleotide includes a SERPINA1 nucleic acid sequence selected from SEQ ID NOs: 119-122. In some embodiments, the donor polynucleotide includes a nucleic acid sequence that is 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 119-122.
[0347] In some embodiments, the guide sequence having a donor polynucleotide includes a SERPINA1 nucleic acid sequence selected from SEQ ID NOs. 123-126. In some embodiments, the donor polynucleotide includes a nucleic acid sequence that is 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs. 123-126.
[0348] Table 8 shows exemplary sequences of target sites and donor polynucleotides. All sequences are paired with the exemplary scaffold sequence GTTTTAGTACTCTGGAAACAGAATCTACTAAAACAAGGCAAAATGCCGTGTTTATCTCGTCAACTTGTTGGCGAGAT (SEQ ID NO: 87).
[0349] In some embodiments, the guide polynucleotide and donor polynucleotide include nucleic acid sequences that are 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs. 88-100. 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs. 101-103 or 105-108. 106%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs. 107. 108. 109. 230-233.
[0350] [Table 11-1]
[0351] [Table 11-2]
[0352] [Table 11-3]
[0353] [Table 11-4]
[0354] In some embodiments, the donor polynucleotide encodes a protein or a fragment thereof. In some embodiments, the protein is a fluorescent protein. In some embodiments, the fluorescent protein is GFP, eGFP, RFP, YFP, BFP, or CFP.
[0355] In some embodiments, the donor polynucleotide further comprises the coding sequence of the self-cleaving peptide. Exemplary self-cleaving peptides include T2A, P2A, E2A, and F2A self-cleaving peptides. In some embodiments, the T2A self-cleaving peptide comprises the sequence EGRGSLLTCGDVEENPGP. In some embodiments, the P2A self-cleaving peptide comprises the sequence ATNFSLLKQAGDVEENPGP. In some embodiments, the E2A self-cleaving peptide comprises the sequence QCTNYALLKLAGDVESNPGP. In some embodiments, the F2A self-cleaving peptide comprises the sequence VKQTLNFDLLKLAGDVESNPGP. In some embodiments, the T2A self-cleaving peptide comprises the sequence EGRGSLLTCGDVEENPGP. Any of the above may also comprise an N-terminal GSG linker. For example, the T2A self-cleaving peptide may also comprise the sequence GSGEGRGSLLTCGDVEENPGP.
[0356] In some embodiments, the donor polynucleotide further comprises a trans-acting double-stranded RNA polynucleotide having a 3' or 5' overhanging nucleotide. In some embodiments, the donor polynucleotide further comprises a cis-acting single-stranded RNA polynucleotide having sequence similarity to the target strand of the nuclease. In some embodiments, the donor polynucleotide further comprises a cis-acting single-stranded RNA polynucleotide having sequence similarity to the non-target strand of the nuclease. In some embodiments, the donor polynucleotide further comprises one or more binding sites for genome modifiers, optionally including binding sites for site-specific recombinases such as serine recombinase or LoxP target sites. In some embodiments, the donor polynucleotide further comprises a repair template for a protein-coding sequence. In some embodiments, the donor polynucleotide further comprises one or more exons having a splice acceptor and a donor sequence. In some embodiments, the donor polynucleotide further comprises one or more selectable sequences selected from the group of NeoR, BsdR, HygR, PuroR, and BleoR genes. In some embodiments, the donor polynucleotide further comprises one or more drug-inducible regulatory sequences for controlled gene expression. In some embodiments, the donor polynucleotide comprises the above combination.
[0357] In some embodiments, the donor polynucleotide comprises a self-complementary single-stranded RNA repair template separated by a ribozyme cleavage site, which, when transcribed, acts as a transactive double-stranded nucleic acid to repair the target sequence via the NHEJ pathway.
[0358] promoter In another embodiment, specific expression regulatory sequences useful for the expression of recombinant nucleic acids provided herein are provided herein. Non-limiting exemplary embodiments of recombinant nucleic acids of this disclosure may include one or more of the following features:
[0359] In some embodiments, any of the recombinant nucleic acids provided herein may be operably ligated to other structural elements (e.g., promoter sequences) required for the expression of such recombinant nucleic acids in a host cell, in a subject, or in an ex-vivo cell-free expression system, and may be, for example, brought under their control.
[0360] As used herein, the terms “promoter” and “promoter sequence” are used interchangeably to refer to DNA sequences that promote the expression of a protein encoding an open reading frame or a nucleotide sequence encoding an ana functional RNA (e.g., a guide polynucleotide or donor polynucleotide). Those skilled in the art will understand that different promoters induce gene expression in different tissues or cell types, at different developmental stages, or in response to different environmental or physiological conditions.
[0361] In some embodiments, the nucleic acid further comprises one or more promoters.
[0362] In some embodiments, the nucleic acid further comprises a promoter that jointly drives the expression of the chimeric nuclease and the guide RNA. In some embodiments, the promoter is operably ligated to the nucleic acid encoding the chimeric nuclease. In some embodiments, the promoter is operably ligated to the nucleic acid encoding the guide RNA. In some embodiments, the promoter is operably ligated to the nucleic acid sequence encoding the entire operon of the guide RNA of this disclosure.
[0363] In some embodiments, the promoter is the CMV promoter, the SV40 promoter, the minimal cytomegalovirus (CMV) promoter, or the human elongation factor-1 alpha (EF1a) promoter. In some embodiments, the promoter is the miniCMV promoter. In some embodiments, the promoter is the MND promoter. In some embodiments, the promoter is the U6 promoter. In some embodiments, the promoter is the T7 promoter.
[0364] In some embodiments, the promoter is a tissue or cell-specific promoter. In some embodiments, the promoter is a muscle-specific synthetic promoter, SPc5-12, a neuron-specific promoter, hSYN1, aldh1L1 promoter, cTNT promoter, alpha-MHC promoter, SPc5-12 promoter, MUC2 promoter, Ksp-cadherin promoter, albumin promoter, HAS promoter, insulin promoter, rhodopsin promoter, rNSE promoter, or cone-opsin promoter. In some embodiments, the promoter is a T7 promoter.
[0365] In some embodiments, the promoter is the CMV promoter described in Sequence ID No. 6.
[0366] In some embodiments, the donor polynucleotide is inserted into the genome in frame with the endogenous coding sequence. In some embodiments, the donor polynucleotide is inserted into the genome in frame for expression from the endogenous promoter.
[0367] Ribozymes and Transcription RNA In some embodiments, the nucleic acid further comprises one or more self-cleaving ribozymes or nucleic acid sequences encoding endogenous RNA processing sequences, such as tRNA. Self-cleaving ribozymes are catalytic RNA molecules that cleave their own phosphodiester backbone. The introduction of self-cleaving ribozymes can ensure a clean 5' or 3' end of the guide RNA, depending on the placement of the ribozyme in the mRNA. In some embodiments, the ribozyme is positioned at the 5' end of the guide RNA. In some embodiments, the ribozyme is positioned directly adjacent to the 5' end of the guide RNA.
[0368] In some embodiments, the ribozyme is a hammerhead ribozyme or a hepatitis delta virus (HDV) ribozyme. In some embodiments, the ribozyme is a hammerhead ribozyme encoded by the nucleic acid sequence described in SEQ ID NO: 8, SEQ ID NO: 9, or SEQ ID NO: 111. In some embodiments, the ribozyme is an HDV ribozyme encoded by the nucleic acid sequence described in SEQ ID NO: 10 or SEQ ID NO: 112.
[0369] Exemplary ribozyme sequences are listed in Table 9. In some embodiments, the ribozyme is encoded by a nucleic acid sequence described in any of sequence numbers 289-301.
[0370] [Table 12]
[0371] In some embodiments, the ribozyme sequence is located inside the tRNA sequence and is transactive. In some embodiments, the tRNA sequence containing the ribozyme sequence is described in Sequence ID No. 5. It should be understood that the DNA sequence is transcribed into the RNA sequence, and the ribozyme is active when it exists as RNA. In some embodiments, the tRNA is glycinearginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine tRNA. In some embodiments, one or more tRNAs separate one or more guide RNAs. If the tRNA is expressed as RNA, it is conditionally cleaved in the cell through the endogenous tRNA maturation process, resulting in cleavage into two or more parts of the RNA. In some embodiments, one or more tRNAs separate one or more tRNAs from an RNA stable sequence. In some embodiments, the tRNA is encoded by the nucleic acid sequence described in SEQ ID NO: 127 or SEQ ID NO: 128. In some embodiments, the tRNA is encoded by the nucleic acid sequence described in Table 10. In some embodiments, the tRNA is encoded by the nucleic acid sequence described in any one of SEQ ID NOs: 302 to 733.
[0372] [Table 13-1]
[0373] [Table 13-2]
[0374] [Table 13-3]
[0375] [Table 13-4]
[0376] Table 13-5
[0377] Table 13-6
[0378] Table 13-7
[0379] Table 13-8
[0380] Table 13-9
[0381] Table 13-10
[0382] Table 13-11
[0383] Table 13-12
[0384] Table 13-13
[0385] Table 13-14
[0386] Table 13-15
[0387] Table 13-16
[0388] Table 13-17
[0389] Table 13-18
[0390] Table 13-19
[0391] Table 13-20
[0392] Table 13-21
[0393] Table 13-22
[0394] Table 13-23
[0395] Table 13-24
[0396] [Table 13-25]
[0397] RNA stabilization sequence In some embodiments, the nucleic acid further comprises RNA-stabilized polynucleotides.
[0398] In some embodiments, the nucleic acid includes a polyadenylation signal. In some embodiments, the polyadenylation signal includes Simian virus 40 (SV40), α-globin, β-globin, human growth hormone (hGH), bovine growth hormone (BGH), herpes simplex virus type 1, thymidine kinase (HSV TK), or a synthetic polyadenylation (Synt polyA) polyadenylation signal.
[0399] In some embodiments, the RNA-stabilizing polynucleotide is a poly(A) signal. The poly(A) tail acts as a binding site for poly(A)-binding proteins. Poly(A)-binding proteins promote export from the nucleus and translation, and inhibit degradation. In some embodiments, a polyadenylation signal is described in SEQ ID NO: 11. Generally, the poly(A) signal is located at the 3' end of the mRNA transcript. The poly(A) signal cannot typically be located at the 5' end or inside the mRNA transcript. Ribozyme cleavage of a single mRNA transcribed from the nucleic acids described herein exposes the free 3' end on the mRNA. A second RNA-stabilizing polynucleotide other than poly(A) may be introduced to protect the 3' end from degradation.
[0400] In some embodiments, the RNA-stabilized polynucleotide is the metastasis-associated lung adenocarcinoma transcript 1 (MALAT) 3' sequence or the 3' end of the multiple endocrine neoplasm β transcript (MENβ). In some embodiments, the MALAT sequence is the nucleic acid sequence shown in Sequence ID No. 7.
[0401] In some embodiments, the RNA-stabilized polynucleotide is the 3' end of an RNA transcript lacking a triple-helix RNA structure or a standard polyadenylation signal.
[0402] In some embodiments, the RNA-stabilizing polynucleotide is located at the 3' end of the polynucleotide encoding the chimeric nuclease. In some embodiments, the RNA-stabilizing polynucleotide is located at the 5' end of the ribozyme cleavage site. In some embodiments, the RNA-stabilizing polynucleotide protects the mRNA encoding the polypeptide portion of the chimeric nuclease from degradation.
[0403] In some embodiments, the nucleic acid further comprises 5'UTR and / or 3'UTR sequences.
[0404] In some embodiments, compositions are provided comprising a chimeric nuclease polypeptide containing an I-TevI domain and an RNA-inducible nuclease domain, and the nucleic acid of the present disclosure.
[0405] In some embodiments, compositions are provided comprising a chimeric nuclease nucleic acid encoding a chimeric nuclease containing an I-TevI domain and an RNA-inducible nuclease domain, and the nucleic acid of the present disclosure. In some embodiments, the chimeric nuclease nucleic acid is mRNA.
[0406] Additional elements In another embodiment, the nucleic acids provided herein may further include additional regulatory elements. In some embodiments, the additional regulatory elements may be WPRE sequences, poly(A) sequences, and / or viral replication sequences, as well as packaging control sequences and long terminal repeat (LTR) sequences.
[0407] In some embodiments, recombinant nucleic acids include a polyadenylation (polyA) signal. In some embodiments, the polyadenylation signal includes Simian virus 40 (SV40), α-globin, β-globin, human growth hormone (hGH), bovine growth hormone (BGH), herpes simplex virus type 1, thymidine kinase (HSV TK), or a synthetic polyadenylation (Synt polyA) signal.
[0408] In some embodiments, the viral replication and packaging control sequence is a long terminal repeat (LTR) sequence. In some embodiments, the viral replication and packaging control sequence is an inverted terminal repeat (ITR) sequence.
[0409] vector In some embodiments, the nucleic acids of this disclosure can be incorporated into an expression vector or a vector used for virus production. The vector may be a plasmid, phage, or cosmid into which another DNA segment may be inserted to result in replication of the inserted segment. In some embodiments, the expression vector may be an incorporation vector. Thus, the Specified also provides vectors, plasmids, or viruses containing one or more nucleic acids encoding any of the nucleic acids disclosed herein. The nucleic acids may be contained in a vector that can direct their expression in cells transduced with the vector, for example. Vectors suitable for use in eukaryotic and prokaryotic cells are known in the Art, commercially available, or readily prepared by those skilled in the art. Additional vectors can also be found, for example, in Ausubel, FM, et al., Current Protocols in Molecular Biology, (Current Protocol, 1994) and Sambrook et al., “Molecular Cloning: A Laboratory Manual,” 2nd ED. (1989).
[0410] Viral vector In some embodiments, the nucleic acids of this disclosure can be packaged in viral particles or viral vectors. Methods for constructing viral vectors from various viral types are known in the art. Exemplary types of viral particles that can be recombinantly operated as delivery vehicles include retroviruses, lentiviruses (e.g., HIV and its derivatives, as well as SIV), adeno-associated viruses, adenoviruses, MMLV retroviruses, MSCV retroviruses, baculoviruses, vesicular stomatitis viruses, herpes simplex viruses, and vaccinia viruses. An example is adeno-associated virus (AAV) particles used in gene therapy. In some embodiments, the virus may be AAV. In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV7, AAV8, AAV9, AAV10, AAV-DJ, AAV2.5T, or AAVmyo. In some embodiments, the virus may be a lentivirus.
[0411] Pharmaceutical composition In other embodiments, compositions and pharmaceutical compositions suitable for administration to human subjects are provided herein, comprising any of the chimeranucleases and chimeranucleases, nucleic acids, vectors, or viral vectors disclosed and described herein. The chimeranucleases and chimeranucleases, nucleic acids, vectors, or viral vectors of this disclosure may be formulated as compositions comprising pharmaceutical compositions. Such compositions generally comprise one or more chimeranucleases and chimeranucleases, nucleic acids, vectors, or viral vectors disclosed and described herein, as well as pharmaceutically acceptable excipients, such as carriers. In some embodiments, the compositions of this disclosure are formulated for the treatment or management of a health condition. In some embodiments, the health condition is a congenital disorder. For example, the compositions of this disclosure may be formulated as therapeutic compositions, or pharmaceutical compositions comprising pharmaceutically acceptable excipients, or mixtures thereof. In some embodiments, the compositions of this disclosure are formulated for use as a therapy for cystic fibrosis. In some embodiments, the compositions of the present disclosure are formulated for use as a therapy for alpha-1-antitrypsin deficiency. In some embodiments, the compositions of the present disclosure are formulated for use as a therapy for myotonic dystrophy type 1. In some embodiments, the compositions of the present disclosure are formulated for use as a therapy for amyotrophic lateral sclerosis or frontotemporal dementia.
[0412] Accordingly, in one embodiment, a pharmaceutical composition comprising a pharmaceutically acceptable excipient and a nucleic acid, vector, or viral vector of the present disclosure is provided herein.
[0413] In some embodiments, pharmaceutical compositions are provided comprising a chimeric nuclease, a chimeric nuclease system, a nucleic acid, a vector, a viral vector, an AAV of the Disclosure, a composition of the Disclosure, or an LNP composition of the Disclosure, and an excipient.
[0414] In some embodiments, the compositions described herein, for example, chimeranucleases and chimeranucleases, nucleic acids, vectors, or viral vectors, and / or pharmaceutical compositions, are incorporated into therapeutic compositions for use in methods of preventing or treating subjects who have, are suspected of having, or are at high risk of developing cystic fibrosis. In some embodiments, the compositions described herein, for example, chimeranucleases and chimeranucleases, nucleic acids, vectors, or viral vectors, and / or pharmaceutical compositions, are incorporated into therapeutic compositions for use in methods of preventing or treating subjects who have, are suspected of having, or are at high risk of developing alpha-1-antitrypsin deficiency. In some embodiments, the compositions described herein, for example, chimeranucleases and chimeranucleases, nucleic acids, vectors, or viral vectors, and / or pharmaceutical compositions, are incorporated into therapeutic compositions for use in methods of preventing or treating subjects who have, are suspected of having, or are at high risk of developing myotonic dystrophy type 1. In some embodiments, the compositions described herein, such as chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors, or viral vectors, and / or pharmaceutical compositions, are incorporated into therapeutic compositions for use in methods of preventing or treating subjects who have, are suspected of having, or are at high risk of developing amyotrophic lateral sclerosis or frontotemporal dementia.
[0415] In these cases, the composition must be sterile, formulated to facilitate administration to human subjects, and stable under manufacturing and storage conditions that allow it to withstand contamination by microorganisms such as bacteria and fungi.
[0416] In some embodiments, the composition is formulated for one or more of the following: intramuscular, intranodal, intravenous, intratracheal, intraperitoneal, or intracranial administration.
[0417] cell delivery The chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors, or viral vectors disclosed and described herein can be delivered to cells for genome editing.
[0418] In some embodiments, delivery is nonviral or viral. In some embodiments, the nuclease can be delivered as a ribonucleoprotein complex, as DNA encoding the nuclease system, as messenger RNA encoding the nuclease system, or as a combination thereof. For example, the polypeptide portion of the nuclease can be delivered to the cell as a protein, and the guide RNA / donor can be delivered as RNA.
[0419] In some embodiments, this disclosure provides vectors comprising nucleic acids described herein. In some embodiments, the chimeric nuclease system and the nucleic acid encoding the chimeric nuclease system can be delivered to cells using lipofection or polymer-based transfection. In some embodiments, the nucleic acid is contained in lipid nanoparticles (LNPs). In some embodiments, the chimeric nuclease system is contained in lipid nanoparticles.
[0420] In some embodiments, nucleic acids encoding a chimeric nuclease system can be delivered to cells using viral delivery. In some embodiments, the nucleic acids are packaged in a viral vector. The viral vector can be used to transduce the nucleic acids of this disclosure into mammalian cells. In some embodiments, the virus is a lentivirus, adeno-associated virus (AAV), adenovirus, retrovirus, or modified herpes simplex virus (HSV). In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-DJ, AAV2.5T, or AAVmyo.
[0421] In some embodiments, the cells are mammalian cells. In some embodiments, the cells are human cells. In some embodiments, the cells are immune cells (such as T cells), hematopoietic stem cells, mesenchymal stem cells, or induced pluripotent stem cells (iPSCs). Cell types that can be modified ex vivo by the methods and nucleases described herein include immune cells such as T cells or NK cells, or pluripotent cells such as mesenchymal stem cells or hematopoietic stem cells, or cells otherwise induced to be pluripotent using techniques known in the art. In some embodiments, the cells are muscle cells. In some embodiments, the cells are brain cells.
[0422] How to use chimeric nucleases One aspect of this disclosure provides a method for editing the genome of a cell using a chimeric nuclease system and a nucleic acid encoding a chimeric nuclease system as described herein.
[0423] In some embodiments, the method includes the step of contacting cells with the nucleic acid of the Disclosure. In some embodiments, the nucleic acid is DNA. In some embodiments, the nucleic acid is RNA. In some embodiments, the nucleic acid is mRNA. In some embodiments, the method includes the step of contacting cells with a virus comprising the nucleic acid of the Disclosure. In some embodiments, the method includes the step of contacting cells with a vector comprising the nucleic acid of the Disclosure. In some embodiments, the method includes the step of contacting cells with an LNP comprising the nucleic acid of the Disclosure.
[0424] In some embodiments, a method is provided for removing DNA of a precise length from the genome of a cell, wherein the method provides the cell with a chimeric nuclease that cuts at two sites and leaves two distinct DNA ends in order to remove DNA of a precise length. The two DNA ends provide target sites necessary for accurate repair by error-free NHEJ or RNA-dependent Polθ-mediated repair, enabling repair at all cell cycle phases, including G1 / G0.
[0425] In some embodiments, a method is provided for editing the genome of a cell at a chimeric nuclease target site, the method comprising the step of providing the cell with a chimeric nuclease having a repair polynucleotide having a 2-base overhang corresponding to the base 3' adjacent to the I-TevI target site in the cell's genome. In some embodiments, a method is provided for editing the genome of a cell at a target site, the method comprising the step of providing the cell with a chimeric nuclease having a repair polynucleotide having a 2- to 18-base overhang corresponding to the base 3' adjacent to the I-TevI target site in the cell's genome. In some embodiments, a method is provided for editing the genome of a cell at a target site, the method comprising the step of providing the cell with a chimeric nuclease having a repair polynucleotide having a 14-base overhang corresponding to the base 3' adjacent to the I-TevI target site in the cell's genome. In some embodiments, the repair polynucleotide comprises a donor polynucleotide. In some embodiments, the repair polynucleotide comprises a guide polynucleotide.
[0426] In some embodiments, a method is provided for delivering messenger RNA encoding a chimeric nuclease comprising an I-TevI domain and an RNA-inducible nuclease domain to cells, the method comprising the step of contacting the cells with a polynucleotide encoding one or more guide RNAs and a polynucleotide donor.
[0427] In some embodiments, methods are provided for genetically modifying the genome of a cell, the method comprising contacting the cell with the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure.
[0428] In some embodiments, the modification includes insertions, deletions, substitutions, or mutations in the cell's genome.
[0429] In some embodiments, the insertion of a donor polynucleotide into the cell's genome results in the removal of a sequence between the I-TevI target site and the Cas9 target site.
[0430] In some embodiments, the cells are mammalian cells. In some embodiments, the cells are human cells.
[0431] In some embodiments, methods are provided for inserting or substituting a sequence into a chimeric nuclease target site in the genome of a cell, the method comprising contacting the cell with a nucleic acid comprising a chimeric nuclease containing a Cas9 domain and an I-TevI domain, a guide polynucleotide and a donor polynucleotide, the guide polynucleotide and the chimeric nuclease forming a complex, the complex binding to and cleaving genomic DNA at the Cas9 target site and the I-TevI target site, the 3' end of the donor polynucleotide having at least two base complements to the 5' end of the I-TevI target site, and the donor polynucleotide being incorporated into the chimeric nuclease target site at the 5' position relative to the Cas9 target site.
[0432] In some embodiments, the cell polymerase targets the chimeric nuclease target site. In some embodiments, the cell polymerase is a polymerase theta.
[0433] In some embodiments, the 3' end of the guide polynucleotide and the 5' end of the donor polynucleotide are ligated together.
[0434] In some embodiments, a method is provided for editing the genome of a cell at a target site, the method comprising providing the cell with a chimeric nuclease having a repair polynucleotide, which is a cis-acting single-stranded nucleic acid encoding a sense strand or an antisense strand of a repair sequence, and which is repaired via an RNA-dependent Rad52-mediated repair pathway.
[0435] In some embodiments, a method is provided for editing the genome of a cell at a target site, the method comprising the step of providing a cell with a purified protein chimeric nuclease that forms a complex with one or more synthetic guide RNA sequences isolated by one or more synthetic tRNA sequences encoding a trans-acting ribozyme sequence, and further comprising a donor DNA sequence isolated by one or more ribozyme cleavage sites.
[0436] In some embodiments, a method is provided for editing the genome of a cell at a target site to generate a large, predictable deletion at the target site via an error-free NHEJ, the method comprising the step of providing the cell with a dual-guide chimeric nuclease system comprising a catalytically inactive Cas9 domain, a catalytically active I-TevI domain, and two guide RNAs targeting two different target sites.
[0437] In some embodiments, methods are provided for replacing at least a portion of the CFTR gene in the genome of a cell, comprising the step of contacting the cell with the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure. In some embodiments, the guide RNA and donor polynucleotide target mutations in the CFTR gene. In some embodiments, a chimeric nuclease having the guide RNA and donor polynucleotide targets and replaces the CFTR c.1521_1523del(p.Phe508del), c.1624G>T(p.Gly542Ter), c.1652G>A(p.Gly551Asp), c.1657C>T(p.Arg553Ter), or c.3846G>A(p.Trp1282Ter) mutations.
[0438] In some embodiments, methods are provided for replacing at least a portion of the SERPINA1 gene in the genome of a cell, the methods comprising contacting the cell with the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure. In some embodiments, the guide RNA and donor polynucleotides target mutations in the SERPINA1 gene. In some embodiments, the guide RNA and donor polynucleotides target and replace the SERPINA1 c.1096G>A(p.Glu342Lys) mutation.
[0439] In some embodiments, a method is provided for excising at least a portion of the DMPK gene in the genome of a cell, the method comprising contacting the cell with the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the Composition of the Disclosure, or the LNP Composition of the Disclosure. In some embodiments, a dual-guide chimeric nuclease excises and deletes the (CTG)n*(CAG)n trinucleotide repeat in the 3' untranslated region of the DMPK gene.
[0440] In some embodiments, a method is provided for excising at least a portion of the C9ORF72 gene in the genome of a cell, the method comprising contacting the cell with the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure. In some embodiments, a dual-guide chimeric nuclease excises and deletes GGGGCC repeats in the intron sequence between alternative 5' exons of the C9ORF72 gene. In some embodiments, a chimeric nuclease, guide RNA, and donor polynucleotides target and remove the large hexanucleotide GGGGCC repeat between exon 1a and exon 1b of the C9ORF72 gene.
[0441] In some embodiments, methods are provided for removing at least a portion of the DMPK gene from the genome in a cell, the method comprising contacting the cell with the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure. In some embodiments, two guide RNAs of the Disclosure target two different mutations in the DMPK gene. In some embodiments, a chimeric nuclease, guide RNA, and donor polynucleotide target and remove a large triplet CAG repeat in the 3' untranslated region of the DMPK gene.
[0442] In some embodiments, a method is provided for tunably editing the genome of a cell at a target site, the method comprising the step of providing a cell with a nucleic acid expressing a chimeric nuclease system described herein, the nucleic acid further comprising a self-inactivating sequence that is cleaved by the chimeric nuclease. In some embodiments, the genome target site is B2M, and the targeting sequence is SEQ ID NO: 51 or SEQ ID NO: 130. In some embodiments, the self-inactivating sequence is SEQ ID NO: 52 or SEQ ID NO: 131.
[0443] In some embodiments, a method is provided for inserting a sequence into a target site in the genome of a cell at a target site, the method comprising the step of providing the cell with nucleic acids comprising the nuclease system and donor polynucleotides described herein.
[0444] In some embodiments, a method is provided for inserting a polypeptide sequence into a target site in the genome of a cell at a target site, the method comprising providing the cell with nucleic acids comprising a nuclease system and donor polynucleotides as described herein.
[0445] In some embodiments, a method is provided for removing disease-causing large repeat sequences of DNA from the genome of a cell, the method comprising the step of providing a chimeric nuclease that cuts the large repeat sequences of DNA at two sites in order to remove them.
[0446] In some embodiments, a method is provided for inserting a donor polynucleotide in frame with an endogenous promoter, the method comprising the step of providing cells with a nucleic acid comprising a donor polynucleotide encoding a polypeptide and a T2A cleavable peptide sequence as described herein.
[0447] In some embodiments, a method for replacing at least a portion of the DMPK gene in the genome of a cell comprises the step of contacting the cell with the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure. In some embodiments, one or more guide RNAs target mutations in the DMPK gene. In some embodiments, one or more guide RNAs target a CAG triplet polynucleotide sequence in the 3' untranslated region of the DMPK gene.
[0448] In some embodiments, the deletion is 30 base pairs long. In some embodiments, the deletion is 35 base pairs long. In some embodiments, the deletion is 40 base pairs long. In some embodiments, the deletion is 45 base pairs long. In some embodiments, the deletion is 50 base pairs long. In some embodiments, the deletion is 60 base pairs long. In some embodiments, the deletion is 70 base pairs long. In some embodiments, the deletion is 80 base pairs long. In some embodiments, the deletion is 90 base pairs long. In some embodiments, the deletion is 150 base pairs long. In some embodiments, the deletion is 30 base pairs long. In some embodiments, the deletion is 200 base pairs long. In some embodiments, the deletion is 300 base pairs long. In some embodiments, the deletion is 400 base pairs long. In some embodiments, the deletion is 500 base pairs long. In some embodiments, the deletion is 600 base pairs long. In some embodiments, the deletion is 700 base pairs long. In some embodiments, the deletion is 800 base pairs long. In some embodiments, the deletion is 900 base pairs long. In some embodiments, the deletion is 1000 base pairs long. In some embodiments, the deletion is 1500 base pairs long. In some embodiments, the deletion is 2000 base pairs long. In some embodiments, the deletion is 2500 base pairs long. In some embodiments, the deletion is 3000 base pairs long. In some embodiments, the deletion is 4000 base pairs long. In some embodiments, the deletion is 5000 base pairs long. In some embodiments, the deletion is 6000 base pairs long. In some embodiments, the deletion is 7000 base pairs long. In some embodiments, the deletion is 8000 base pairs long. In some embodiments, the deletion is 9000 base pairs long. In some embodiments, the deletion is 10000 base pairs long. In some embodiments, the deletion is longer than 10,000 base pairs.
[0449] Treatment method In another embodiment, methods for using chimeric nucleases and chimeric nuclease systems, nucleic acids, vectors, or viral vectors disclosed and described herein for the treatment of human subjects requiring them are provided herein.
[0450] Such chimeranucleases and chimeranucleases, nucleic acids, vectors, or viral vectors disclosed and described herein are suitable for use as gene therapies for the treatment of subjects requiring them. In some embodiments, the chimeranucleases and chimeranucleases, nucleic acids, vectors, or viral vectors disclosed and described herein include therapeutic agents for use in methods of treating subjects who have, are suspected of having, or are at high risk of developing one or more health conditions or diseases that can be treated with such gene therapy. Exemplary health conditions or diseases may include, but are not limited to, congenital disorders such as cystic fibrosis, alpha-1 antitrypsin deficiency, myotonic dystrophy type 1, amyotrophic lateral sclerosis, or frontotemporal dementia.
[0451] Cystic fibrosis Cystic fibrosis (CF) is an autosomal recessive disease caused by mutations in the CFTR gene, which encodes epithelial anion channels. The CFTR protein, a cystic fibrosis membrane conductance regulator, is found in a wide range of organs, including the pancreas, kidneys, liver, lungs, gastrointestinal tract, and reproductive organs, making CF a multi-organ disease. Mutations in the CFTR gene (GRCh38.p14 GCF_000001405.40, NM_000492.4) result in suboptimal ion transport and fluid retention, which can lead to abnormal mucus thickening in the lungs and marked clinical manifestations of pancreatic dysfunction. In the lungs, dysfunctional CFTR interferes with mucociliary clearance, making the organ more susceptible to bacterial infection and inflammation, which can ultimately lead to airway obstruction, respiratory failure, and premature death. CF remains the most common and deadly genetic disorder in the Caucasian population, with an estimated 70,000 to 100,000 patients worldwide, highlighting the real need for the development of better treatments.
[0452] Mutations in CFTR include, but are not limited to, c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), or c.3846G>A (p.Trp1282Ter).
[0453] In some embodiments, a method is provided for treating cystic fibrosis in a patient requiring treatment for cystic fibrosis, the method comprising the step of administering the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure.
[0454] In some embodiments, guide RNA and donor polynucleotides target mutations in the CFTR gene.
[0455] In some embodiments, the guide RNA and donor polynucleotides target and replace the following mutations in CFTR: c.1521_1523del (p.Phe508del), c.1624G>T (p.Gly542Ter), c.1652G>A (p.Gly551Asp), c.1657C>T (p.Arg553Ter), or c.3846G>A (p.Trp1282Ter).
[0456] Alpha-1 antitrypsin deficiency Alpha-1 antitrypsin deficiency is a genetic disorder affecting the lungs and sometimes the liver. The affected gene in alpha-1 antitrypsin deficiency is SERPINA1 (encoding alpha-1 antitrypsin) on chromosome 14 (GRCh38.p14(GCF_000001405.40)). The protein encoded by this gene is a serine protease inhibitor belonging to the serpine superfamily, and its targets include elastase, plasmin, thrombin, trypsin, chymotrypsin, and plasminogen activator. This protein is produced in the liver, bone marrow, lymphocytes and monocytes in lymphoid tissues, and Paneth cells in the intestines. Deficiency of this gene is associated with chronic obstructive pulmonary disease, emphysema, and chronic liver disease.
[0457] In some embodiments, a method is provided for treating alpha-1-antitrypsin deficiency in a patient requiring treatment for alpha-1-antitrypsin deficiency, the method comprising administering to the patient the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure.
[0458] In some embodiments, the guide RNA and donor polynucleotides target mutations in the SERPINA1 gene. In some embodiments, the guide RNA and donor polynucleotides target and replace the SERPINA1 c.1096G>A(p.Glu342Lys) mutation.
[0459] Myotonic dystrophy type 1 Myotonic dystrophy type 1 (DM1) is a multisystem disorder affecting skeletal and smooth muscles, as well as the eyes, heart, endocrine system, and central nervous system. Clinical findings, ranging from mild to severe, are classified into three somewhat overlapping phenotypes: mild, classic, and congenital. DM1 is caused by elongation of CTG trinucleotide repeats in the non-coding region of DMPK. A diagnosis of DM1 is suspected in individuals with characteristic muscle weakness and confirmed by molecular genetic testing of DMPK. A CTG repeat length exceeding 34 repeats is abnormal. Molecular genetic testing can detect pathogenic variants in nearly 100% of affected individuals.
[0460] In some embodiments, a method for treating myotonic dystrophy type 1 in a patient requiring treatment for myotonic dystrophy type 1, the method comprising the step of administering to the patient the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure.
[0461] Amyotrophic lateral sclerosis (ALS) Amyotrophic lateral sclerosis (ALS) is a neurological disorder that affects motor neurons, which are nerve cells in the brain and spinal cord that control voluntary muscle movement and respiration. C9orf72 frontotemporal dementia and / or amyotrophic lateral sclerosis (C9orf72-FTD / ALS) is most often characterized by frontotemporal dementia (FTD) and superior-subordinate motor neuron disease (MND). C9orf72-FTD / ALS can be identified by molecular genetic testing, which shows heterozygous abnormal G4C2 (GGGGCC,C4C2) hexanucleotide repeat elongation in C9orf72.
[0462] In some embodiments, a method for treating amyotrophic lateral sclerosis or frontotemporal dementia in a patient requiring treatment for the condition, the method comprising administering to the patient the nucleic acid of the Disclosure, the vector of the Disclosure, the viral vector of the Disclosure, or the AAV of the Disclosure, the composition of the Disclosure, or the LNP composition of the Disclosure.
[0463] kit This specification also provides various kits for the practice of the methods described herein, as well as written instructions for their preparation and use. In particular, several embodiments relate to kits for methods of treating diseases in subjects requiring treatment of the disease. For example, in some embodiments, a kit is provided herein that includes one or more nucleic acids and / or pharmaceutical compositions provided and described herein, as well as written instructions for their use. In some embodiments, the kits of the disclosure further include one or more means useful for administering the engineered immune cells and / or pharmaceutical compositions to a subject. For example, in some embodiments, the kits of the disclosure further include one or more bags, syringes (including pre-filled syringes) used for administering any one of the engineered immune cells and / or pharmaceutical compositions provided to a subject.
[0464] In some embodiments, the kit may further include instructions for carrying out the methods disclosed herein using the components of the kit. For example, the kit may include a package insert containing information about the pharmaceutical compositions and dosage forms contained in the kit. Generally, such information helps patients and physicians to use the enclosed pharmaceutical compositions and dosage forms effectively and safely. For example, the following information relating to the combination of the disclosures may be supplied in the insert: pharmacokinetics, pharmacodynamics, clinical trials, efficacy parameters, indications and uses, contraindications, warnings, precautions, adverse reactions, overdose, appropriate dosage and administration, supply methods, appropriate storage conditions, references, manufacturer / distributor information, and intellectual property information.
[0465] Instructions for carrying out the method are typically recorded on a suitable recording medium. For example, the instructions may be printed on a substrate such as paper or plastic. The instructions may be present in the kit as an accompanying document, on the label of the kit's container or its components (e.g., in relation to the packaging or sub-packaging). The instructions may exist as an electronic storage data file on a suitable computer-readable storage medium, such as a CD-ROM, diskette, or flash drive. In some examples, the actual instructions are not present in the kit, but means for obtaining the instructions from a remote source (e.g., via the Internet) may be provided. An example of this embodiment is a kit that includes a web address from which the instructions can be viewed and / or downloaded. Similar to the instructions, this means for obtaining the instructions may be recorded on a suitable substrate. Further embodiments are disclosed in more detail in the following embodiments, which are provided as examples and are not intended in any way to limit the scope of this disclosure or the claims. [Examples]
[0466] These examples are provided for illustrative purposes only and do not limit the scope of the claims provided herein.
[0467] Example 1. Design of an all-in-one chimeric nuclease cassette. This example describes the design and synthesis of an all-in-one chimeric nuclease cassette.
[0468] In short, we design a chimeric nuclease cassette containing a promoter, chimeric nuclease, mRNA stabilization sequence, gRNA, modified tRNA including a ribozyme sequence, repair template, and poly(A) tail. Exemplary nucleic acid expression cassettes are shown in Figures 1, 2, 26A, 26B, and 26C. The minimal CMV promoter drives the transcription of a 3-in-1 construct containing Dualase (chimeric nuclease), long non-coding RNA MALAT-1, synthetic transfer RNA (tRNA') containing a sequence-specific ribozyme, gRNA (RT) fused with a repair template, and synthetic poly(A) signal sequence (synt[A]). During transcription, RNA maturation at the 3' end of MALAT and both ends of tRNA' results in the separation of Dualase-MALAT RNA, tRNA, and gRNA. U-rich repeats on MALAT help protect Dualase RNA from degradation, while tRNA' recognizes and cleaves in regions between RTs, as well as between RTs and synt[A] sequences. Post-translation, Dualase can form a complex with gRNA(RNP) and cleave the target site. The presence of in-place repair templates (cis-sense and antisense) fused to the 3' end of the gRNA or abundant free-floating repair templates (trans-antisense and trans-sense) acts as a bridge between the two cleavage sites and local reference templates for cellular repair mechanisms.
[0469] Another exemplary chimeric nuclease cassette is designed to include a promoter, chimeric nuclease, self-cleaving peptide, protein tag, mRNA stabilization sequence, gRNA, modified tRNA containing a ribozyme sequence, repair template, and poly(A) tail. An exemplary nucleic acid expression cassette is shown in Figure 2. An all-in-one mRNA cassette encoding elements necessary for efficient and precise target site disruption or repair. The T7 promoter drives the transcription of a Dualase-containing 3-in-1 construct, long non-coding RNA MALAT-1, synthetic transfer RNA (tRNA') containing a sequence-specific ribozyme, gRNA fused with a repair template (RT), and synthetic poly(A) signal sequence (synt[A]). Conditions are established for tRNA, MALAT maturation, and ribozyme cleavage activity to be inactive in storage buffer. Upon delivery, RNA maturation at the 3' end of MALAT and both ends of tRNA' results in the separation of Dualase-MALAT RNA, tRNA, and gRNA. U-rich repeats on MALAT help protect Dualase RNA from degradation, while tRNA' recognizes and cleaves in regions between RTs, as well as between RTs and synt[A] sequences. Post-translation, Dualase forms a complex with gRNA (RNP) and cleaves the target site. The presence of in-place repair templates (cis-sense and antisense) fused to the 3' end of the gRNA or abundant free-floating repair templates (trans-antisense and trans-sense) acts as a bridge between the two cleavage sites and local reference templates for cellular repair mechanisms.
[0470] Example 2. Delivery of chimeric nuclease This example describes the delivery of a chimeric nuclease system to cells.
[0471] In short, a chimeric nuclease is expressed in E. coli and purified. RNA containing gRNA fused to a repair template is produced. The purified chimeric nuclease and guide RNA / repair template are assembled into a ribonucleoprotein (RNP) complex and co-delivered to cells, for example, together with LNPs or by electroporation.
[0472] Figure 3 shows a diagram of the dual-cleavage nuclease ribonucleoprotein (RNP) complex with an all-in-one guide RNA repair cassette. Dualase and 2-in-1 gRNA-RT are incubated to form the ribonucleoprotein (RNP) complex, which is then delivered to target cells where they interact and cleave the target site after nuclear translocation.
[0473] Example 3. Gene editing using chimeric nuclease The example describes in-cellular genome editing using a chimeric nuclease that targets the AAVS1 gene locus in cells.
[0474] In short, HEK293 cells were treated with AAV-Dualase-AAVS1. Genomic DNA was recovered, purified, and sequenced. Data were analyzed by indel tracking by DE composition (TIDE). The results showed a 35 bp deletion by AAV-Dualase-AAVS1. Amplicons from AAV-Dualase-AAVS1 treated samples were Sanger sequenced and analyzed using the TIDE approach. Deletion portions, insertions of specified length, and unmodified reads were characterized as bar graphs. Chromatograms of sequenced samples treated with AAV-Dualase-AAVS1 or reagent alone compared to the dotted line P<0.01(B) reference sequence (Figure 4A and B).
[0475] Example 4. Gene editing using an all-in-one chimeric nuclease cassette. This example describes in-cellular genome editing using an all-in-one chimeric nuclease that targets the AAVS1 gene locus in cells.
[0476] In short, exogenous donor nucleic acids were co-delivered in HEK293 cells by co-expressing Dualase along with guide RNA and donor nucleic acids, using one nucleotide sequence operably bound to a promoter sequence, one or more self-cleavage guide RNA sequences, and one or more self-cleavage repair template sequences. Genomic DNA was recovered, purified, and sequenced. Next-generation sequencing (NGS) demonstrated highly accurate target site repair. NGS data were analyzed using two bioinformatics tools (Figure 5A) CRIS.Py and Geneious alignment platforms, as well as (Figure 5B) CRISPRESSO2. "Indels" represent repair events resulting in insertions or deletions. "Precise repair" represents sequencing reads with complete alignment to the repaired sequence. "Repair+SNP" represents sequence alignment with repair and another nucleotide change. "Other" represents other sequence modifications in the sample. Analysis of individual sequences using CRIS.Py and Geneious identified approximately 10% of reads containing repair + SNPs or other sequences, which is likely due to sequencing errors, as cell-only samples contained the same number of reads. Editing results were identified using all reads and aligned reads (p<0.001). Alignment of sequencing reads of the cis-antisense RNA repair template guide RNA shows the accuracy of repair at I-TevI and Cas9 sites with no changes other than the expected repair result, exceeding the detection limit of sequencing, as indicated by the dotted line (Figure 5C). The results demonstrate that gene editors insert RNA sequences at target sites with high efficiency and accuracy (Figures 5A, 5B, and C).
[0477] Example 5. Gene editing using an all-in-one chimeric nuclease cassette and a DNA repair inhibitor. This example describes genome editing in cells using an all-in-one AAV chimeric nuclease that targets the AAVS1 gene locus in cells in conjunction with an NHEJ inhibitor.
[0478] In short, HEK293 cells were lipofection with (Figure 6A) non-homologous end-joining (NHEJ), (Figure 6B) homologous-directed repair (HDR), or (Figure 6C) inhibitors blocking the Rad52-dependent pathway, using AAV expressing Dualase, a gRNA-RT (all-in-one) targeting AAVS1, and a control (co-delivered) of AAV expressing Dualase targeting AAVS1 co-delivered with the indicated repair template. NHEJ, Rad51, and Rad52 inhibitors with cis-sense and cis-antisense repair template samples were analyzed for editing efficiency at the inserted repair template site by restriction enzyme (RE) digestion efficiency using specific REs. The results indicate that repair using cis-sense and cis-antisense all-in-one constructs was inhibited only by the Rad52 inhibitor. (Figure 6A-C) Control inhibition of co-delivered repair templates indicates that the inhibitors functioned as expected. NHEJ(DNA) = non-homologous double-stranded DNA, HDR(DNA) = double-stranded DNA with homologous arms, and NHEJ(RNA) = non-homologous single-stranded RNA.
[0479] Example 6. Gene editing using an all-in-one chimeric nuclease cassette and repair template. This example describes directional insertion of an RNA repair template by a Dualase ribonucleoprotein (RNP) complex formed from an AAVS1-targeted guide RNA combined with the RNA repair template via directional polymerase chain reaction (PCR).
[0480] In short, HEK293 cells were lipofected with Dualase with separately delivered guide RNA and repair template (co-delivery), with Dualase with fused guide RNA and RNA repair template ("all-in-one"), with lipofection reagent only ("reagent only"), or with cells only ("mock"). For each reaction, the undigested PCR amplicon of the AAVS1 site ("Sub") and the digested PCR amplicon of the AAVS1 site ("Digested") are shown. Genomic DNA was extracted and purified. Samples were analyzed for editing efficiency by restriction enzyme (RE) digestion efficiency using RE specific to the inserted repair template site. The results show that the PCR product is correctly oriented ("correctly oriented RT") only when the repair template is correctly inserted, and incorrectly oriented ("mis-oriented RT") when the repair template is incorrectly inserted. The Dualase RNP complex with the all-in-one repair template is accurately inserted, as demonstrated by the correctly oriented PCR product (Figure 7A). This involves using a co-delivered repair template, which is inserted in both correct and incorrect orientations (Figure 7B).
[0481] Example 7. Gene editing using an all-in-one chimeric nuclease cassette and a cis-sense or cis-antisense repair template. This example describes the analysis of directional insertion of repair templates by AAVS1-targeted Dualase all-in-one mRNA having cis-sense or cis-antisense repair templates using directional polymerase chain reaction (PCR).
[0482] In short, HEK293 cells were lipofected with Dualase (Tev[VKN]-SaCas9[WT] or Tev[VKN]-SaCas9[D10E]), SaCas9[WT], or lipofection reagent alone. Genomic DNA was extracted and purified. The editing efficiency of the samples was analyzed by restriction enzyme (RE) digestion (BglII digestion) efficiency using a unique RE at the inserted repair template site. The results shown for each reaction are the undigested PCR amplicon of the AAVS1 site ("full length"), the digested PCR amplicon of the AAVS1 site ("BglII digestion"), the PCR product when the repair template was inserted only in the correct orientation ("correctly oriented RT"), and the PCR product when the repair template was inserted only in the wrong orientation ("mis-oriented RT") (Figure 8). Control cells treated with Dualase mRNA targeting the AAVS1 site with co-delivered repair template (co-delivery) are also shown. As is evident from the precisely oriented PCR products, cis-sense or cis-antisense repair templates were accurately inserted with the all-in-one Dualase mRNA. As is evident from the absence of BglII digestion and directional PCR products, no insertions were detected with cis-sense or antisense repair templates containing only saCas9[WT]. Co-delivered Dualase mRNA and repair templates were inserted in both correct and incorrect orientations.
[0483] Example 8. Gene editing using an all-in-one chimeric nuclease cassette and NHEJ and Rad52 inhibitors. This example describes the analysis of editing efficiency in cells using messenger RNA (mRNA) dualase, NHEJ and Rad52 inhibitors, and a trans-dsRNA repair template.
[0484] In short, HEK293 cells were treated by lipofection with an all-in-one AAV expressing Dualase and gRNA-RT (also known as repair template guide RNA or "rep-gRNA") that targets AAVS1, and a control AAV expressing Dualase that targets AAVS1 and is co-delivered with the indicated repair template (co-delivered). Sample editing efficiency was analyzed by examining restriction enzyme (RE) digestion efficiency using specific restriction enzymes (RE) at the insertion site of the inserted repair template. Repair using the trans-RT all-in-one construct was inhibited by both NHEJ and Rad52 inhibitors, demonstrating that both of these pathways can be used with the self-complementary trans-repair template (Figure 9). Control inhibition of the co-delivered repair template indicates that the inhibitors functioned as expected. NHEJ (DNA) = non-homologous double-stranded DNA, HDR (DNA) = double-stranded DNA with homology arms, and NHEJ (RNA) = non-homologous single-stranded RNA.
[0485] Example 9. Gene editing using dual-guide chimeric nuclease for deletion This embodiment describes the precise removal of large repetitive sequences using a dual-guide TevCas9 nuclease.
[0486] Figure 10A shows a schematic diagram of the precise removal of a large repetitive sequence using dual-guide TevCas9 nucleases, where the guide RNA targets the opposite strand of the double-stranded DNA, orienting the two I-TevI domains opposite each other. TevCas9 nuclease (1), containing an inactivated D10A+H557A mutation but an active I-TevI domain (2), is targeted to the 5' end of the repetitive sequence (3) using synthetic guide RNA. A second TevCas9 nuclease (4), containing an inactivated D10A+H557A mutation but an active I-TevI domain, is targeted to the 3' end of the repetitive sequence but is targeted to the opposite strand, so that the I-TevI domains of both nucleases are directed to the repetitive sequence. Upon binding and cleavage by the I-TevI nuclease domains (5), two complementary 3' 2-nucleotide overhangs remain (6 and 7). Through non-homologous end joining pathways, cells can repair these complementary overhangs (8), large repetitive sequences are removed (9), and a predetermined number of repeats remain in the genomic DNA (10).
[0487] Figure 10B shows a schematic diagram of the precise removal of a large repetitive sequence using dual-guide TevCas9 nucleases, where the guide RNA targets the same strand of double-stranded DNA and orients the two I-TevI domains in series. A TevCas9 nuclease (1) containing an inactivated D10A+H557A mutation but an active I-TevI domain (2) is targeted to the 5' end of the repetitive sequence (3) using synthetic guide RNA. A second TevCas9 nuclease (4) containing an inactivated D10A+H557A mutation but an active I-TevI domain is targeted downstream of the 3' end of the repetitive sequence on the same strand so that the I-TevI domains of both nucleases are oriented in the same direction. Two complementary 3' 2-nucleotide overhangs remain upon binding and cleavage by the I-TevI nuclease domains (5) (6 and 7). Through non-homologous end joining pathways, cells can repair these complementary overhangs (8), removing large repetitive sequences (9) and leaving a predetermined number of repeats in the genomic DNA (10).
[0488] Example 10. Gene editing using a dual-guided chimeric nuclease for deletion of CAG triplets or GGGGCC hexanucleotide repeats in cells and humanized mice. This embodiment describes the precise removal of CAG triplet repeat extensions in DMPK 3'-UTR or GGGGCC hexanucleotide repeats between exons 1a and 1b in C9ORF72 using an all-in-one AAV encoding TevCas9, along with dual guides that target the 5' and 3' ends of the repeat extension in vivo.
[0489] In short, DMPK MUT CAG repeat cells were transduced with the all-in-one AAV TevCas9 and SaCas9 dual guide (SEQ ID NO: 117). (Figure 11A) A diagram showing the orientation of the dual guide Tev[KTQ]-Cas9[D10A+H557A] on the DMPK CAG repeat target site. (Figure 11B) A diagram showing the sequence in the all-in-one TevCas9 AAV. Hammerhead ribozyme (HHribo), single guide RNA (sgRNA), glycine tRNA (Gly tRNA), hepatitis delta virus ribozyme (HDVribo), and polyadenylated sequence (PolyA) are shown. (Figure 11C) DMPK showing I-TevI domain activity at this target. MUT The in vitro activity of dual-guide TevCas9 against CAG repeat DNA sequences is shown. (Figure 11D) Alignment of Sanger sequencing reads of amplicons generated from cells transduced with an all-in-one AAV encoding TevCas9 together with the dual guide is shown, along with predicted cleavage indicating removal at the repeat CAG sequence. (Figure 11E) DMPK transduced with all-in-one AAV TevCas9 and SaCas9 dual guide MUTShows the workflow of an experiment for quantifying repeat collapse using repeat-primed PCR in CAG repeat cells. Transduced cells are collected, and genomic DNA is extracted for PCR amplification using primers outside and inside the repeat amplification. The obtained PCR products are analyzed with an Agilent Bioanalyzer, and peaks corresponding to repeats are quantified for DMPK MUT only in CAG repeat cells. More than 65% of repeats were collapsed in TevCas9-treated cells, while approximately 30% of repeats were collapsed in SaCas9-treated cells. (Figure 11F) Transduced cells were subjected to quantitative RT-PCR to quantify transcript levels of DMPK MUT and DMPK WT after each treatment using primers specific to each transcript. Compared with untreated cells, TevCas9 treatment resulted in a statistically significant increase in DMPK WT transcript and a decrease in DMPK MUT transcript (p<0.01). Neither TevCas9 with a single targeting guide RNA nor SaCas9 dual guide treatment resulted in a statistically significant change in transcription levels compared with cells alone. (Figure 11G) shows that transduced cells were subjected to RT-PCR to quantify the level of mis-spliced CLCN1 exon 6 transcript after each treatment using primers specific to the mis-spliced sequence. TevCas9 single guide or TevCas9 dual guide resulted in a statistically significant reduction in mis-spliced CLCN1 exon 6.
[0490] For C9ORF72 GGGGCC repeat removal (Figure 11H), Figure 11I demonstrates I-TevI domain activity at this target for C9ORF72 MUTThis document demonstrates the in vitro activity of dual-guide TevCas9 against GGGGCC repeat DNA sequences. Patient motor neuron progenitor cells containing large GGGGCC repeat extensions were matured into motor neurons (Figure 11J) and then transduced with an all-in-one AAV encoding TevCas9 or saCas9 and a dual guide targeting the C9ORF72 repeat region. Maturation was confirmed by Western blotting with expression of a motor neuron-specific marker (ISL1). Figure 11K quantifies the removal of larger repeat sequences in TevCas9 or Cas9 transduced cells, showing repeat prime PCR results with approximately 50% removal in TevCas9-treated cells and less than 5% removal in Cas9-treated cells. Figure 11L shows that sequencing of folded repeats in TevCas9-treated cells reveals four out of nine sequences with precise repeat folding, as seen in sequences marked with an asterisk (*). Figure 11M shows the C9ORF72 guide RNA that targets Tev[VKN]-saCas9[D10A+H557A], which is specific to the C9ORF72 target site, because off-targets in the genome are not detected at levels exceeding the reported detection limits of the deep sequencing assay. Figure 11N shows the need for TevSaCas9, as active SaCas9 targeted by the same C9ORF72 targeting guide RNA has detectable off-targets on chromosome 11. Figure 11O shows that C9ORF72 protein expression can be recovered in patient-derived mature motor neurons by transduction of AAVs expressing Tev[VKN]-saCas9[D10A+H557A] and dual C9ORF72 guides. Figure 11P shows that transduction of patient-derived motor neurons into AAVs expressing Tev[VKN]-saCas9[D10A+H557A] and dual C9ORF72 guides (SEQ ID NO: 115) can reduce the accumulation of polyGR dipeptides expressed from C9ORF72 repeat sequences. The accumulation of dipeptide toxins (toxic build-up) contributes to motor neuron degeneration in disease.Figure 11Q shows RT-qPCR-mediated TevCas9 expression in the cerebellum of humanized C9ORF72 repeat elongation mice intracranially injected with two doses of Tev[VKN]-saCas9[D10A+H557A] and a dual C9ORF72 guide, demonstrating transduction of brain tissue. Figure 11R shows dose-dependent recovery of normal-sized C9ORF72 repeat sequences in humanized mice intracranially injected with two doses of Tev[VKN]-saCas9[D10A+H557A] and a dual C9ORF72 guide. Figure 11S shows the recovery of C9ORF72 transcripts by RT-qPCR in cerebellar tissue of humanized C9ORF72 repeat elongation mice injected intracranially with two doses of Tev[VKN]-saCas9[D10A+H557A] and dual C9ORF72 guides. Figure 11T shows the recovery of C9ORF72 protein expression by Western blotting in cerebellar tissue of humanized C9ORF72 repeat elongation mice injected intracranially with two doses of Tev[VKN]-saCas9[D10A+H557A] and dual C9ORF72 guides.
[0491] Example 11. Gene editing using self-inactivating chimeric nuclease This example illustrates gene editing of the B2M gene using self-inactivating TevSaCas9.
[0492] In short, HEK293 cells transfected with plasmid DNA containing TevSaCas9 targeting the beta-2-microglobulin (B2M) gene were harvested 24, 48, and 72 hours after transfection. Figure 12A shows the structure of the self-inactivating vector. The construct includes a promoter, human codon-optimized TevSaCas9, a polyadenylation signal ("PolyA"), and nucleotide sequences encoding a guide RNA sequence ("gRNA"). The self-inactivating target site can be located in the region between the promoter and the TevSaCas9 site (indicated as "promoter") or at the end of the TevSaCas9 coding sequence and the beginning of the PolyA sequence (indicated as "PolyA"). Figure 12B shows the results of the T7E1 editing assay in HEK293 cells transfected with plasmid DNA containing TevSaCas9 targeting the β-2-microglobulin (B2M) gene, harvested 24, 48, and 72 hours after transfection. Lanes marked "None" do not contain any self-inactivating sequences in the vector; lanes marked "Promoter" contain the B2M1 TevSaCas9 target site between the promoter sequence and the TevSaCas9 sequence; and lanes marked "PolyA" contain the B2M1 TevSaCas9 target site between the end of TevSaCas9 and the polyA signal sequence. The level of editing, determined by the amount of digested product relative to the substrate, is equivalent over time between constructs. Figure 12C shows the results of a Western blot for hemagglutinin (α-HA) encoded at the 3' end of TevSaCas9 ("dualase") from the same treated cells in B. Lanes marked "None" do not contain a self-inactivating sequence in the vector; lanes marked "Promoter" contain the B2M1 TevSaCas9 target site between the promoter sequence and the TevSaCas9 sequence; and lanes marked "Poly-A" contain the B2M1 TevSaCas9 target site between the end of TevSaCas9 and the poly-A signal sequence. The presence of a band in the α-HA blot indicates that the TevSaCas9 protein is expressed.TevSaCas9 expression levels remained stable over time in vectors without inactivating sequences, decreased over 72 hours in vectors containing promoter-inactivating sequences, and became undetectable after 72 hours in vectors containing poly(A)-inactivating sequences. Figure 12D shows a photograph of an SDS-polyacrylamide gel illustrating the production and titer of the AAV2 capsid encapsulating B2M-targeting self-inactivated TevCas9, as well as the successful assembly of the capsid with VP1, VP2, and VP3 capsid structural proteins. As seen in the Western blot for the HA tag on the TevCas9 protein, after 72 hours, the level of TevCas9 protein was significantly reduced compared to the beta-actin housekeeping genes, which exhibited the resulting inactivation. Figure 12E shows Western blot results for saCas9 and GAPDH from lipofected cells transfected with a plasmid DNA version of the vector used to produce self-inactivating AAV in Figure 12D ("self-inactivating"), a vector without the self-inactivating sequence ("non-self-inactivating"), and a plasmid expressing GFP as a control ("pAAV-GFP"). Transfected cells were harvested at 48 hours (48h), 72 hours (72h), 7 days (7d), and 14 days (14d), lysed, and total protein was recovered. The presence of expressed Cas9 was detected with a Cas9 antibody, and the consistency of the loaded sample was assessed using the housekeeping gene GAPDH. SaCas9-sized bands were detected only in the non-self-inactivating control samples at 48 and 72 hours, and not in the self-inactivating samples, demonstrating that the presence of a self-inactivating target site in the expression vector limits the expression of TevSaCas9.
[0493] Example 12. Gene editing using TevCas9 and rep-gRNA to recruit endogenous polymerase theta (PolQ or Polθ) This example illustrates the use of a TevCas9 gene editor and rep-gRNA in cells to recruit endogenous polymerase.
[0494] In short, HEK293 cells were lipofected with TevCas9 or Cas9 mRNA and AAVS1-targeted rep-gRNA (Figure 13A-C). The cells were lysed, and the extracts were precipitated using magnetic beads conjugated with anti-HA antibodies that specifically recognize TevCas9 or the C-terminal HA tag on Cas9. After washing the magnetic beads, the precipitated proteins were separated on a polyacrylamide gel and Western blotted with various antibodies. Figure 14 shows Western blots using Rad52-specific antibodies and polymerase theta-specific (Pol Q or Pol θ) antibodies, demonstrating that polymerase θ co-precipitates only with TevCas9.
[0495] Example 13. Gene editing using inactivated or nickas chimelanucleases This example describes gene editing of the AAVS1 site by TevCas9 using nuclease inactivation or a nickase gene editor.
[0496] In short, HEK293 cells were lipofected with purified TevCas9 and Cas9 proteins, as well as a 2-in-1 guide RNA targeting the AAVS1 site. The editing efficiency of the samples was analyzed by examining restriction enzyme (RE) digestion efficiency using restriction enzymes specific to the inserted repair template insertion site. Digestion was observed using purified TevCas9 containing an inactivated D10A+H557A mutation in the Cas9 domain ("Tev[WT]-dCas9"), and TevCas9 containing an inactivated D10A+H557A mutation in Cas9, along with an inactivated R27A mutation in the I-TevI domain ("Tev[R27A]-dCas9", Figure 15A). No digestion was observed in cells treated with Cas9 alone, reagent alone, or untreated cells ("Cell only"). Digestion was observed in controls of cells lipofected with Tev-Cas9 3-in-1 mRNA. Digestion was also observed in the mRNA version of TevCas9 containing the mutation R27A+V117F+K135R+N140S, and in Cas9 domains with either the D10A or H557A nickase mutation (Figure 15B).
[0497] Example 14. In-frame insertion of GFP into the cell genome. This example describes the in-frame insertion of GFP into the genome of a mammalian cell containing TevCas9.
[0498] In short, to cover the G542X, R553X, and G551D mutations, rep-gRNAs containing the T2A sequence and the eGFP coding sequence were designed for insertion into the G542 site of the CFTR gene (Figure 16A). Cells were lipofected with TevCas9 and the eGFP-coding rep-RNA, or with saCas9 containing the eGFP-coding rep-RNA. The results show that GFP is expressed from the endogenous CFTR promoter only under TevCas9 conditions and not under SaCas9, rep-gRNA, or reagent-only conditions (Figure 16B). eGFP is isolated from the endogenous protein cleaved by the self-cleaving T2A peptide. Figure 16C shows that the inserted sequence is detected as a larger band when the CFTR target site region is amplified by PCR. PCR results for the target region show that the ratio of edited to unedited is approximately 50%. Another approach uses a combination of upstream (5') guide RNAs of rep-gRNA to remove a large sequence of approximately 28 kilobases and replace it with a 921 base pair (bp) GFP coding sequence. The GFP-coding rep-gRNA extends to the region between CFTR F508 ...
Claims
1. (i) A polynucleotide encoding a chimeric nuclease containing an I-TevI domain and an RNA-inducible nuclease domain, (ii) A polynucleotide encoding the first guide RNA (gRNA), (iii) Polynucleotides that represent tRNA and A nucleic acid comprising the polynucleotides in (ii) to (iii) in a consecutive order.
2. (i) A polynucleotide encoding a chimeric nuclease containing a GIY-YIG nuclease domain and an RNA-inducible nuclease domain, (ii) A polynucleotide encoding the first guide RNA (gRNA), (iii) Polynucleotides that represent tRNA and A nucleic acid comprising the polynucleotides in (ii) to (iii) in a consecutive order.
3. (iv) RNA-stabilized polynucleotide located downstream of (i) The nucleic acid according to claim 1 or claim 2, further comprising:
4. (i) A polynucleotide encoding the first guide RNA (gRNA), (ii) Polynucleotides that encode tRNA and A nucleic acid comprising (i) to (ii), wherein the polynucleotides in (i) to (ii) are in a consecutive order.
5. The nucleic acid is Two or more donor polynucleotides arranged in tandem, Ribozyme polynucleotides and The nucleic acid according to any one of claims 1 to 4, further comprising, wherein the donor polynucleotides are in a continuous order, and the upstream donor polynucleotide includes a ribozyme cleavage site sequence at its 3' end.
6. The nucleic acid according to any one of claims 1 to 5, wherein the nucleic acid is DNA or RNA.
7. The nucleic acid according to claim 6, wherein the DNA is circular plasmid DNA, linear double-stranded DNA, single-stranded DNA, or chimeric RNA and DNA.
8. The nucleic acid according to claim 6, wherein the RNA is mRNA.
9. The nucleic acid according to claim 8, wherein the mRNA comprises a nucleic acid mimetic selected from the group consisting of peptide nucleic acids (PNA), morpholino nucleic acids, cyclohexenyl nucleic acids (CeNA), and locked nucleic acids (LNA).
10. The nucleic acid according to claim 8 or 9, wherein the mRNA comprises a modified sugar moiety, and optionally the modified sugar moiety is selected from the group consisting of N1-methylpseudridine, 9-methyladenine, 2'-O-(2-methoxyethyl), 2'-dimethylaminooxyethoxy, 2'-dimethylaminoethoxyethoxy, 2'-O-methyl, and 2'-fluoro.
11. The mRNA contains modified nucleic acid bases, and optionally the modified nucleic acid bases are 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, adenine 6-methyl derivative, guanine 6-methyl derivative, adenine 2-propyl derivative, guanine 2-propyl derivative, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 6-azouracil, 6-azocytosine, 6-azocytosine, pseudouracil, 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl, 5-halo, 5-bromo, 5-trifluoromethyl, 5-substituted uracil, 5-substituted uracil A nucleic acid according to any one of claims 8 to 10, selected from the group consisting of substituted cytosine, 7-methylguanine, 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine, 8-azaadenine, 7-deazaguanine, 7-deazaadenine, 3-deazaguanine, 3-deazaadenine, tricyclic pyrimidine, phenoxazinecytidine, phenothiazinecytidine, substituted phenoxazinecytidine, carbazolecytidine, pyridoindolecytidine, 7-deazaadenine, 7-deazaguanosine, 2-aminopyridine, 2-pyridone, 5-substituted pyrimidine, 6-azapyrimidine, N-2, N-6, or O-6 substituted purines, 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine.
12. The nucleic acid according to any one of claims 8 to 11, wherein the mRNA comprises a nucleoside bond that is not naturally occurring or is unnatural, selected from the group consisting of phosphorothioates, phosphoramidates, nonphosphodiesters, heteroatoms, chiral phosphorothioates, phosphorodithioates, phosphotryesters, aminoalkylphosphotryesters, 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphines, 3'-aminophosphoramidates, aminoalkylphosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotryesters, selenophosphates, and boranophosphates.
13. The nucleic acid according to any one of claims 1 to 12, wherein the RNA-inducible nuclease is selected from the group consisting of Staphylococcus aureus Cas9 ("saCas9"), Streptococcus pyogenes Cas9, Acidaminococcus Cas12, Deltaproteobacteria CasX, and Eubacterium rectare Cas12a.
14. The nucleic acid according to claim 13, wherein Cas is inactivated Cas (dCas).
15. The nucleic acid according to claim 13, wherein Cas is nicase (nCas) or dCas.
16. The nucleic acid according to any one of claims 1 to 15, wherein I-TevI is a nickase, and the I-TevI is an I-TevI nickase domain.
17. The nucleic acid according to claim 16, wherein the I-TevI nicckerse domain includes mutations in amino acid residues R27A, V117F, K135R, and N140S, which correspond to the amino acid residues in SEQ ID NO:
155.
18. The nucleic acid according to any one of claims 1 to 15, wherein I-TevI is inactivated.
19. The nucleic acid according to claim 18, wherein the I-TevI inactivating mutation is an R27A mutation in an amino acid residue corresponding to the amino acid residue in SEQ ID NO:
155.
20. The nucleic acid according to any one of claims 5 to 19, wherein the ribozyme polynucleotide is located within the tRNA polynucleotide.
21. The nucleic acid according to any one of claims 1 to 20, further comprising a polynucleotide encoding a second gRNA, wherein optionally the second gRNA polynucleotide is located at 3' of the tRNA polynucleotide.
22. The nucleic acid according to any one of claims 3 or 5 to 21, wherein the order of the polynucleotides is chimeric nuclease, RNA-stabilizing polynucleotide, guide RNA, tRNA, and guide RNA.
23. The nucleic acid according to any one of claims 5 to 21, wherein the order of the polynucleotides is chimeric nuclease, RNA-stabilizing polynucleotide, guide RNA, tRNA / ribozyme, guide RNA, donor polynucleotide 1, and donor polynucleotide 2.
24. The nucleic acid according to any one of claims 3 or 5 to 23, wherein the RNA-stabilized polynucleotide comprises the 3' sequence of metastasis-associated lung adenocarcinoma transcript 1 (MALAT) or the 3' end of multiple endocrine neoplasm β transcript (MENβ).
25. The nucleic acid according to any one of claims 3 or 5 to 23, wherein the RNA stabilization sequence comprises the 3' end of an RNA transcript lacking a triple-helix RNA structure or a standard polyadenylation signal.
26. The nucleic acid according to any one of claims 5 to 25, wherein the donor polynucleotide is single-stranded or double-stranded.
27. The nucleic acid according to any one of claims 5 to 25, wherein the donor polynucleotide is DNA or RNA.
28. The nucleic acid according to any one of claims 5 to 25, wherein one strand of the double-stranded donor polynucleotide is DNA and the other strand is RNA.
29. The nucleic acid according to any one of claims 1 to 25, wherein the donor polynucleotide comprises a cis-acting single-stranded RNA polynucleotide annealed to a complementary single-stranded DNA polynucleotide.
30. The nucleic acid according to any one of claims 1 to 28, wherein a single-stranded or double-stranded donor polynucleotide includes a 2 to 18-nucleotide overhang at its 3' end.
31. The nucleic acid according to any one of claims 5 to 29, wherein a single-stranded or double-stranded donor polynucleotide includes a 14-nucleotide overhang at its 3' end.
32. The nucleic acid according to any one of claims 26 to 31, wherein the overhang at the 3' end is a single-stranded RNA polynucleotide.
33. The nucleic acid according to any one of claims 1 to 32, further comprising a second guide RNA capable of targeting region 5' to a donor polynucleotide target site.
34. The nucleic acid according to claim 33, wherein the first guide RNA can target a first chimeric nuclease to a first I-TevI or Cas9 target site and cleave the first I-TevI or Cas9 target site in the cell's genome, and the second guide RNA can target a second chimeric nuclease to a second I-TevI target site in the cell's genome and cleave the second I-TevI target site, the cleavage generating a nucleotide overhang at the second I-TevI target site.
35. The nucleic acid according to claim 34, wherein the 3' end of the donor polynucleotide is complementary to the overhang resulting from the cleavage of the second I-TevI at the second I-TevI target site.
36. The nucleic acid according to any one of claims 1 to 35, wherein the guide RNA and donor polynucleotide target mutations in the CFTR gene.
37. The nucleic acid according to any one of claims 1 to 35, wherein the guide RNA and donor polynucleotide target and substitute for mutations in CFTR c. 1521_1523del (p. Phe508del), c. 1624G>T (p. Gly542Ter), c. 1652G>A (p. Gly551Asp), c. 1657C>T (p. Arg553Ter), or c. 3846G>A (p. Trp1282Ter).
38. The nucleic acid according to any one of claims 1 to 35, wherein the guide RNA and donor polynucleotide target mutations in the SERPINA1 gene.
39. The nucleic acid according to any one of claims 1 to 35, wherein the guide RNA and donor polynucleotide target and substitute for the mutation in SERPINA1 c. 1096G>A (p.Glu342Lys).
40. The nucleic acid according to any one of claims 1 to 39, further comprising a promoter.
41. The nucleic acid according to claim 40, wherein the promoter is selected from the group consisting of the CMV promoter, the SV40 promoter, the minimal cytomegalovirus (CMV) promoter, and the human elongation factor-1α (EF1a) promoter.
42. The nucleic acid according to claim 41, wherein the promoter is selected from the group consisting of the muscle-specific synthetic promoter SPc5-12, the neuron-specific promoter hSYN1, aldh1L1, cTNT, alpha-MHC, SPc5-12, MUC2, Ksp-cadherin, albumin, HAS, insulin, rhodopsin, rNSE, and the cone-opsin promoter.
43. The nucleic acid according to any one of claims 1 to 42, wherein the tRNA comprises glycinearginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine tRNA.
44. The nucleic acid according to any one of claims 1 to 42, wherein the ribozyme comprises a hammerhead ribozyme or a hepatitis delta virus (HDV) ribozyme.
45. Donor polynucleotides Trans-acting double-stranded RNA polynucleotide having a 3' or 5' overhanging nucleotide, A cis-acting single-stranded RNA polynucleotide having sequence similarity to the target strand of a nuclease, A cis-acting single-stranded RNA polynucleotide having sequence similarity to the non-target strand of the nuclease, One or more binding sites to genome modifying factors, optionally, binding sites to site-specific recombinases such as serine recombinase or LoxP target sites. Protein coding sequence repair template, One or more exons having a splice acceptor and a donor sequence, One or more selectable sequences selected from the group of NeoR, BsdR, HygR, PuroR, and BleoR genes, One or more drug-inducible regulatory sequences for controlled gene expression, and / or Overhang of two nucleotides at the 3' end A nucleic acid according to any one of claims 1 to 42, comprising:
46. The nucleic acid according to any one of claims 1 to 45, further comprising a polyadenylation signal.
47. The nucleic acid according to claim 46, wherein the polyadenylation signal comprises Simian virus 40 (SV40), α-globin, β-globin, human growth hormone (hGH), bovine growth hormone (BGH), herpes simplex virus type 1 thymidine kinase (HSV TK), or a synthetic polyadenylation (Synth poly A) polyadenylation signal.
48. The nucleic acid according to any one of claims 1 to 47, further comprising a self-inactivating sequence.
49. A nucleic acid according to any one of claims 1 to 48, having a length of approximately 5 kb.
50. A nucleic acid according to any one of claims 1 to 48, having a length of less than 5 kb.
51. A nucleic acid according to any one of claims 1 to 50, which is packaged in a virus.
52. The nucleic acid according to claim 51, wherein the virus is a lentivirus, adeno-associated virus (AAV), adenovirus, retrovirus, or modified herpes simplex virus (HSV).
53. The nucleic acid according to claim 52, wherein the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV7, AAV8, AAV9, AAV10, AAV-DJ, AAV2.5T, or AAVmyo.
54. A vector comprising the nucleic acid described in any one of claims 1 to 53.
55. A viral vector comprising the nucleic acid described in any one of claims 1 to 53.
56. AAV virus comprising the nucleic acid described in any one of claims 1 to 53.
57. A cell comprising the nucleic acid according to any one of claims 1 to 53, the vector according to claim 54, the viral vector according to claim 55, or the AAV according to claim 56.
58. A composition comprising a chimeric nuclease polypeptide containing an I-TevI domain and an RNA-inducible nuclease domain, and a nucleic acid according to any one of claims 4 to 53.
59. A composition comprising a chimeric nuclease nucleic acid encoding a chimeric nuclease containing an I-TevI domain and an RNA-inducible nuclease domain, and the nucleic acid according to any one of claims 4 to 53.
60. The composition according to claim 59, wherein the chimeric nuclease nucleic acid is mRNA.
61. An LNP composition comprising a nucleic acid according to any one of claims 1 to 50, or a composition according to any one of claims 58 to 60.
62. A pharmaceutical composition comprising a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61, and an excipient.
63. A method for delivering messenger RNA encoding a chimeric nuclease comprising an I-TevI domain and an RNA-inducible nuclease domain to a cell, the method comprising the step of bringing the cell into contact with a polynucleotide encoding one or more guide RNAs and a polynucleotide donor.
64. A method for genetically modifying the genome of a cell, the method comprising the step of contacting the cell with a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.
65. The method according to claim 64, wherein the modification includes an insertion, deletion, substitution, or mutation of the genome.
66. The method according to any one of claims 63 to 65, wherein the insertion of a donor polynucleotide into the genome of the cell results in the removal of a sequence between the I-TevI target site and the Cas9 target site.
67. The method according to claim 64 or 65, wherein the cells are mammalian cells.
68. The method according to any one of claims 64 to 65, wherein the cells are human cells.
69. A method for inserting or substituting a sequence into a chimeric nuclease target site in the genome of a cell, wherein the method is: The aforementioned cells twice Chimeric nucleases containing Cas9 domain and I-TevI domain, and Nucleic acids including guide polynucleotides and donor polynucleotides, The process includes contacting a nucleic acid containing one or more nucleic acid sequences encoding, The guide polynucleotide and the chimeric nuclease form a complex, which binds to the genomic DNA at the Cas9 target site and the I-TevI target site, and cleaves the genomic DNA. The 3' end of the donor polynucleotide, after I-TevI cleavage, contains at least two bases complementary to the 5' end of the I-TevI target site. A method comprising incorporating the donor polynucleotide into the chimeric nuclease target site at the 5' position relative to the Cas9 target site.
70. The method according to claim 69, wherein the 3' end of the guide polynucleotide and the 5' end of the donor polynucleotide are linked.
71. The method according to claim 69, wherein the cell polymerase is targeted to the chimeric nuclease target site.
72. The method according to claim 71, wherein the cell polymerase is polymerase theta.
73. A method for replacing at least a portion of a CFTR gene in the genome of a cell, the method comprising the step of contacting the cell with a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.
74. The method according to claim 73, wherein the guide RNA and donor polynucleotide target mutations in the CFTR gene.
75. The method according to any one of claims 73 to 74, wherein the guide RNA and donor polynucleotides target and substitute mutations in CFTR c. 1521_1523del (p. Phe508del), c. 1624G>T (p. Gly542Ter), c. 1652G>A (p. Gly551Asp), c. 1657C>T (p. Arg553Ter), or c. 3846G>A (p. Trp1282Ter).
76. A method for treating cystic fibrosis in a patient requiring treatment for cystic fibrosis, the method comprising the step of administering to the patient a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.
77. The method according to claim 76, wherein the guide RNA and donor polynucleotides target mutations in the CFTR gene.
78. The method according to claim 76, wherein the guide RNA and the donor polynucleotide target and substitute for the following mutations: CFTR c. 1521_1523del (p. Phe508del), c. 1624G>T (p. Gly542Ter), c. 1652G>A (p. Gly551Asp), c. 1657C>T (p. Arg553Ter), or c. 3846G>A (p. Trp1282Ter).
79. A method for replacing at least a portion of the SERPINA1 gene in the genome of a cell, the method comprising the step of contacting the cell with a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.
80. The method according to claim 79, wherein the guide RNA and donor polynucleotide target mutations in the SERPINA1 gene.
81. The method according to claim 79, wherein the guide RNA and donor polynucleotide target and substitute for the mutation in SERPINA1 c. 1096G>A (p.Glu342Lys).
82. A method for treating alpha-1-antitrypsin deficiency in a patient requiring treatment for alpha-1-antitrypsin deficiency, the method comprising the step of administering to the patient a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.
83. The method according to claim 82, wherein the guide RNA and donor polynucleotide target mutations in the SERPINA1 gene.
84. The method according to claim 82, wherein the guide RNA and donor polynucleotide target and substitute for the mutation in SERPINA1 c. 1096G>A (p.Glu342Lys).
85. A method for replacing at least a portion of the DMPK gene in the genome of a cell, the method comprising the step of contacting the cell with a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.
86. The method according to claim 85, wherein one or more guide RNAs target mutations in the DMPK gene.
87. The method according to claim 85, wherein one or more guide RNAs target a CAG triplet polynucleotide sequence in the 3' untranslated region of the DMPK gene.
88. A method for treating myotonic dystrophy type 1 in a patient requiring treatment for myotonic dystrophy type 1, the method comprising the step of administering to the patient a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.
89. A method for replacing at least a portion of the C9ORF72 gene in the genome of a cell, the method comprising the step of contacting the cell with a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.
90. The method according to claim 89, wherein one or more guide RNAs target mutations in the C9ORF72 gene.
91. The method according to claim 89, wherein one or more guide RNAs target the GGGGCC hexanucleotide repeat sequence between exon 1a and exon 1b of the C9ORF72 gene.
92. A method for treating amyotrophic lateral sclerosis or frontotemporal dementia in a patient requiring treatment for the condition, the method comprising the step of administering to the patient a nucleic acid according to any one of claims 1 to 53, a vector according to claim 54, a viral vector according to claim 55, or an AAV according to claim 56, a composition according to claim 58, or an LNP composition according to claim 61.