Supplementation of liver enzyme expression

JP2025522292A5Pending Publication Date: 2026-06-01METAGENOMI INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
METAGENOMI INC
Filing Date
2023-05-25
Publication Date
2026-06-01
Patent Text Reader

Abstract

Methods, compositions, and systems derived from useful uncultured microorganisms for supplementing liver enzyme deficiencies are described herein. Disclosed herein is an engineered nuclease system comprising: a) an endonuclease comprising an amino acid sequence; b) an engineered guide polynucleotide that forms a complex with the endonuclease and is configured to hybridize to a target nucleic acid sequence within an albumin gene or within an intron of an albumin gene; and c) a donor template comprising a nucleic acid sequence encoding a Factor VIII (FVIII) gene or a functional fragment thereof.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 345,526, filed May 25, 2022, U.S. Provisional Patent Application No. 63 / 359,288, filed Jul. 8, 2022, and U.S. Provisional Patent Application No. 63 / 396,421, filed Aug. 9, 2022, each of which is hereby incorporated by reference in its entirety.

[0002] Sequence Listing The contents of the electronic sequence listing (MTG-017WO_SL.xml, size: 384,067 bytes, and creation date: May 17, 2023) are hereby incorporated by reference in its entirety.

Summary of the Invention

[0003] Various disorders are caused by deficiencies in liver-produced factors (e.g., liver enzymes), and the deficiencies themselves result from genetically inherited mutations. For example, hemophilia A and hemophilia B are genetic disorders caused by mutations in genes encoding clotting factors such as factor VIII produced by the liver. Individuals with hemophilia have low levels of these clotting factors and, as a result, cannot properly clot their blood. Treating these disorders with a gene therapy approach that incorporates into the genome a liver gene encoding a functional liver enzyme (e.g., factor VIII) in a patient enables continuous expression of the functional liver enzyme. Accordingly, methods, compositions, and systems for the replenishment of liver enzymes via gene therapy methods are described herein.

[0004] In certain embodiments, an engineered nuclease system is disclosed herein, comprising: a) an endonuclease comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide that forms a complex with the endonuclease and is configured to hybridize to a target nucleic acid sequence within the albumin gene or within an intron of the albumin gene; and c) a donor template comprising a nucleic acid sequence encoding a Factor VIII (FVIII) gene or a functional fragment thereof. In some embodiments, the target nucleic acid sequence within the albumin gene is within intron 1 of the albumin gene. In some embodiments, the sequence encoding the Factor VIII (FVIII) gene or a functional fragment thereof is linked to a splice acceptor sequence that targets exon 1 of the albumin gene. In some embodiments, the endonuclease comprises a sequence having at least 90% sequence identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease comprises a sequence having 100% sequence identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having 100% sequence identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, 98. In some embodiments, the engineered guide polynucleotide comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, 98.In some embodiments, the target nucleic acid sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 2, and 8. In some embodiments, the target nucleic acid sequence comprises a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1, 2, and 8. In some embodiments, the target nucleic acid sequence comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 1, 2, and 8. In some embodiments, the target nucleic acid sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3-6. In some embodiments, the target nucleic acid sequence comprises a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 3-6. In some embodiments, the target nucleic acid sequence comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 3-6. In some embodiments, the donor template further comprises a polyadenylation signal. In some embodiments, the donor template further comprises a nuclear targeting sequence. In some embodiments, the nuclear targeting sequence comprises a plurality of transcription factor binding sites. In some embodiments, the transcription factor is TCF1, HNF1, NFY, CEBP, OCT1, AP1, HNF1-α, HNF1-β, CEBPA, LEF-1, FOX D1, IRF1, HNF3, HNF4, HNF5, Tal1β / E47, or MyoD. In some embodiments, the nuclear targeting sequence is on the 5' and 3' ends of the donor template. In some embodiments, the donor template further comprises a recognition site sequence for an endonuclease on the 5' or 3' end. In some embodiments, when the donor template is adjacent at the 5' end, the nuclear targeting sequence is on the 5' side of the recognition site sequence. In some embodiments, when the donor template is adjacent at the 3' end, the nuclear targeting sequence is on the 3' side of the recognition site sequence.In some embodiments, the donor template comprises NTS(1)-NRS(1)-SA-FVIII-NRS(2)-NTS(2) in the 5' to 3' direction, where NTS(1) represents a first nuclear targeting sequence, NTS(2) represents a second nuclear targeting sequence, NRS(1) represents a first nuclease recognition site sequence, NRS(2) represents a second nuclease recognition site sequence, SA represents a splice acceptor sequence targeting exon 1 of the albumin gene, and FVIII represents the factor VIII gene or a fragment thereof. In some embodiments, the 5' to 3' orientation of NRS(1) and NRS(2) follows (a) forward, forward, (b) reverse, reverse, (c) forward, reverse, (d) reverse, forward, where forward indicates the same 5' to 3' orientation as the target nucleic acid sequence and reverse indicates the opposite 5' to 3' orientation to the target nucleic acid sequence. In some embodiments, the donor template comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 12, 13, 16-23, 32, 33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 12, 13, 16-23, 32, 33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least 100% sequence identity to any one of SEQ ID NOs: 12, 13, 16-23, 32, 33, 56-59, 81-88, and 90-94. In some embodiments, the FVIII gene or a functional fragment thereof is codon-optimized to remove at least one cytosine-guanine (CG or CpG) motif. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least 90% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having 100% identity to any one of SEQ ID NOs: 10, 71-79, and 89.In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least 80% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least 90% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having 100% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof is modified to comprise a B domain comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 71-79 and 89. In some embodiments, the FVIII gene or a functional fragment thereof is modified to comprise a B domain comprising a sequence having 100% identity to any one of SEQ ID NOs: 71-79 and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprising a modified B domain comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 86, 87, and 90. In some embodiments, the FVIII gene or a functional fragment thereof comprising a modified B domain comprises a sequence having 100% identity to any one of SEQ ID NOs: 86, 87, and 90.

[0005] In certain embodiments, a method for doing so in a subject in need of replenishing hepatic enzyme expression, the method comprising administering to the subject: a) an endonuclease comprising an amino acid sequence having at least 80% sequence identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence within the albumin gene or within an intron of the albumin gene; and c) a donor template comprising a nucleic acid sequence encoding Factor VIII (FVIII) gene or a functional fragment thereof, thereby replenishing hepatic enzyme expression in the subject. In some embodiments, the target nucleic acid sequence within the albumin gene is within intron 1 of the albumin gene. In some embodiments, the sequence encoding the Factor VIII (FVIII) gene or a functional fragment thereof is operably linked to a splice acceptor sequence that targets exon 1 of the albumin gene. In some embodiments, the endonuclease comprises a sequence having at least 90% sequence identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease comprises a sequence having 100% sequence identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having 100% sequence identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, 98.In some embodiments, the engineered guide polynucleotide comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, 98. In some embodiments, the target nucleic acid sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 2, and 8. In some embodiments, the target nucleic acid sequence comprises a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1, 2, and 8. In some embodiments, the target nucleic acid sequence comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 1, 2, and 8. In some embodiments, the target nucleic acid sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3-6. In some embodiments, the target nucleic acid sequence comprises a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 3-6. In some embodiments, the target nucleic acid sequence comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 3-6. In some embodiments, the donor template further comprises a polyadenylation signal. In some embodiments, the donor template further comprises a nuclear targeting sequence. In some embodiments, the nuclear targeting sequence comprises a plurality of transcription factor binding sites. In some embodiments, the transcription factor is TCF1, HNF1, NFY, CEBP, OCT1, AP1, HNF1-α, HNF1-β, CEBPA, LEF-1, FOX D1, IRF1, HNF3, HNF4, HNF5, Tal1β / E47, or MyoD. In some embodiments, the nuclear targeting sequence is on the 5' and 3' ends of the donor template. In some embodiments, the donor template further comprises a recognition site sequence for an endonuclease on the 5' or 3' end. In some embodiments, when the donor template is adjacent at the 5' end, the nuclear targeting sequence is on the 5' side of the recognition site sequence. In some embodiments, when the donor template is adjacent at the 3' end, the nuclear targeting sequence is on the 3' side of the recognition site sequence.In some embodiments, the donor template comprises, in the 5′ to 3′ direction, NTS(1)-NRS(1)-SA-FVIII-NRS(2)-NTS(2), where NTS(1) represents a first nuclear targeting sequence, NTS(2) represents a second nuclear targeting sequence, NRS(1) represents a first nuclease recognition site sequence, NRS(2) represents a second nuclease recognition site sequence, SA represents a splice acceptor sequence that targets exon 1 of the albumin gene, and FVIII represents the factor VIII gene or a fragment thereof. In some embodiments, the 5′ to 3′ orientation of NRS(1) and NRS(2) follows (a) forward, forward, (b) reverse, reverse, (c) forward, reverse, (d) reverse, forward, where forward indicates the same 5′ to 3′ orientation as the target nucleic acid sequence and reverse indicates the opposite 5′ to 3′ orientation as the target nucleic acid sequence. In some embodiments, the donor template comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 12, 13, 16-23, 32, 33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 12, 13, 16-23, 32, 33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least 100% sequence identity to any one of SEQ ID NOs: 12, 13, 16-23, 32, 33, 56-59, 81-88, and 90-94. In some embodiments, the FVIII gene or a functional fragment thereof is codon-optimized to remove at least one cytosine-guanine (CG or CpG) motif. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least 90% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having 100% identity to any one of SEQ ID NOs: 10, 71-79, and 89.In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least 80% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least 90% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having 100% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof is modified to comprise a B domain comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 71-79 and 89. In some embodiments, the FVIII gene or a functional fragment thereof is modified to comprise a B domain comprising a sequence having 100% identity to any one of SEQ ID NOs: 71-79 and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprising a modified B domain comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 86, 87, and 90. In some embodiments, the FVIII gene or a functional fragment thereof comprising a modified B domain comprises a sequence having 100% identity to any one of SEQ ID NOs: 86, 87, and 90.

[0006] In certain embodiments, cells comprising an engineered nuclease system disclosed herein are disclosed. In some embodiments, the cells are liver cells. In some embodiments, the cells are eukaryotic cells. In some embodiments, the cells are mammalian cells. In some embodiments, the cells are immortalized cells. In some embodiments, the cells are insect cells. In some embodiments, the cells are yeast cells. In some embodiments, the cells are plant cells. In some embodiments, the cells are fungal cells. In some embodiments, the cells are prokaryotic cells. In some embodiments, the cells are A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC1, BSC40, BMT10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof. In some embodiments, the cells are engineered cells. In some embodiments, the cells are stable cells.

[0007] In certain embodiments, lipid nanoparticles (LNPs) comprising components (a) and (b), or components (a), (b), and (c) of the engineered nuclease system disclosed herein are disclosed. In some embodiments, the LNP comprises a cationic lipid, a neutral lipid, cholesterol or a cholesterol analog, and a PEG-conjugated lipid. In some embodiments, the cationic lipid comprises C12-200 (1,1‘-((2-(4-(2-((2-(bis(2-hydroxydodecyl)amino)ethyl)(2-hydroxydodecyl)amino)ethyl)piperazin-1-yl)ethyl)azanediyl)bis(dodecane-2-ol)), the neutral lipid comprises 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), or the PEG-conjugated lipid comprises 1,2-dimyristoyl-rac-glycero-3-methoxypolyethylene glycol-2000 (DMG-PEG-2000).

[0008] In certain embodiments, viral vectors comprising the engineered nuclease systems disclosed herein are disclosed herein. In some embodiments, the viral vector is an adeno-associated virus (AAV) vector. In some embodiments, the AAV comprises AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh10, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, AAV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, AAV-HSC16, or a derivative thereof. In some embodiments, the AAV is AAV6. In some embodiments, the AAV is AAV8.

[0009] In certain embodiments, a method for doing so in an individual in need of replenishing hepatic enzyme expression, the method comprising administering to the individual: (a) an endonuclease comprising a nucleic acid encoding a RuvC domain or an endonuclease, the endonuclease having at least 80% sequence identity to the nucleic acid sequence of SEQ ID NO: 30; (b) an engineered guide RNA comprising: (i) a spacer sequence configured to complex with the endonuclease and (ii) configured to hybridize to at least a portion of a target nucleic acid sequence within or in an intron of the albumin gene; and (c) a donor template comprising a nucleic acid sequence encoding a therapeutic gene encoding a functional hepatic enzyme (e.g., Factor VIII), thereby replenishing hepatic enzyme (e.g., Factor VIII) expression in the individual. In some embodiments, the spacer sequence is configured to hybridize to intron 1 of the albumin gene. In some embodiments, the nucleic acid sequence encoding the therapeutic gene is linked to a splice acceptor sequence that targets exon 1 of the albumin gene.

[0010] In certain embodiments, a method for doing so in an individual in need of replenishing hepatic enzyme expression, the method comprising administering to the individual: (a) an endonuclease comprising a nucleic acid encoding a RuvC domain or an endonuclease, the endonuclease having at least 80% sequence identity to the nucleic acid sequence of SEQ ID NO: 31; (b) an engineered guide RNA comprising: (i) a spacer sequence configured to complex with the endonuclease and (ii) configured to hybridize to a target nucleic acid sequence within the albumin gene; and (c) a donor template comprising a nucleic acid sequence encoding a functional hepatic enzyme (e.g., Factor VIII) linked to a splice acceptor sequence that targets exon 1 of the albumin gene, thereby replenishing hepatic enzyme expression in the subject.

[0011] In some embodiments, the endonuclease induces a single-strand break or a double-strand break at or proximal to the target nucleic acid sequence. In some embodiments, the endonuclease induces a double-strand break at or proximal to the target nucleic acid sequence. In some embodiments, the donor template is integrated into the target nucleic acid sequence at the double-strand break. In some embodiments, the donor template is integrated into the target nucleic acid sequence at the double-strand break via non-homologous end joining (NHEJ). In some embodiments, the donor template is integrated into the target nucleic acid sequence at the double-strand break via homology-directed repair (HDR).

[0012] In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 80% identity to any one of SEQ ID NOs: 3-6. In some embodiments, the engineered guide RNA comprises a sequence by any one of SEQ ID NOs: 24-27.

[0013] In some embodiments, the donor template further comprises a polyadenylation signal. In some embodiments, the donor template is a closed-ended linear double strand. In some embodiments, the donor template further comprises a recognition site sequence for an endonuclease on the 5' or 3' end of the donor template. In some embodiments, the donor template further comprises a nuclear targeting sequence that is (a) on the 5' side of the recognition site sequence when the coding sequence is adjacent to the 5' end of the donor template, or (b) on the 3' side of the recognition site sequence when the coding sequence is adjacent to the 3' end of the donor template. In some embodiments, the donor template comprises nuclear targeting sequences on both the 5' and 3' ends of the donor template. In some embodiments, the donor template comprises NTS(1)-NRS(1)-SA-TG-NRS(2)-NTS(2) in the 5' to 3' direction, where NTS(1) represents a first nuclear targeting sequence, NTS(2) represents a second nuclear targeting sequence, NRS(1) represents a first nuclease recognition site sequence, NRS(2) represents a second nuclease recognition site sequence, SA represents a splice acceptor sequence targeting exon 1 of the albumin gene, and TG represents a therapeutic gene. In some embodiments, the 5' to 3' orientation of NRS(1) and NRS(2) follows (a) forward, forward, (b) reverse, reverse, (c) forward, reverse, (d) reverse, forward, where forward indicates the same 5' to 3' orientation as the target nucleic acid sequence and reverse indicates the opposite 5' to 3' orientation to the target nucleic acid sequence. In some embodiments, the first nuclear targeting sequence or the second nuclear targeting sequence comprises a plurality of binding sites for LEF / TCF1, HNF1, NFY, CEBP, OCT1, AP1, HNF1-A, HNF1-B, CEBPA, LEF-1, FOX D1, or IRF1. In some embodiments, the first nuclear targeting sequence or the second nuclear targeting sequence comprises a sequence having at least 80% identity to SEQ ID NO: 1 or SEQ ID NO: 1 having SEQ ID NO: 2 added to the 5' or 3' end of the nuclear targeting sequence.

[0014] In some embodiments, the donor template comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 12, 13, 16-23, and 32, 33. In some embodiments, the donor template comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 16-23. In some embodiments, the donor template comprises a sequence having at least 75% identity to any one of SEQ ID NOs: 16-19. In some embodiments, the donor template comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 20-23.

[0015] In some embodiments, the therapeutic gene is the Factor VIII (FVIII) gene or a functional fragment thereof. In some embodiments, the FVIII gene or a functional fragment thereof is codon-optimized to remove at least one cytosine-guanine (CG or CpG) motif. In some embodiments, the codon-optimized FVIII gene or a functional fragment thereof comprises a sequence having at least 80% identity to SEQ ID NO: 10.

[0016] In certain embodiments, vectors comprising the endonuclease systems disclosed herein are disclosed herein. In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is an adeno-associated virus (AAV) vector. In some embodiments, the vector is a lipid nanoparticle (LNP). In some embodiments, the LNP comprises a cationic lipid, a neutral lipid, cholesterol or a cholesterol analog, or a PEG-conjugated lipid. In some embodiments, the cationic lipid comprises C12-200 (1,1‘-((2-(4-(2-((2-(bis(2-hydroxydodecyl)amino)ethyl)(2-hydroxydodecyl)amino)ethyl)piperazin-1-yl)ethyl)azanediyl)bis(dodecan-2-ol)), the neutral lipid comprises 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), or the PEG-conjugated lipid comprises 1,2-dimyristoyl-rac-glycero-3-methoxypolyethylene glycol-2000 (DMG-PEG-2000).

[0017] In certain embodiments, a system is disclosed herein that comprises (a) an endonuclease capable of cleaving at least one strand of a target nucleic acid within a first nuclease recognition site sequence or a second nuclease recognition site sequence, and (b) in the 5' to 3' direction, (i) a first nuclear targeting sequence, and (ii) a coding sequence for a therapeutic gene encoding a functional hepatic enzyme (e.g., factor VIII), wherein the coding sequence is adjacent at the 5' end by the first nuclease recognition sequence and / or at the 3' end by the second nuclease recognition sequence. In some embodiments, the nucleic acid is double-stranded DNA. In some embodiments, the double-stranded DNA is a closed-ended linear double-strand. In some embodiments, the coding sequence for the coding sequence for the therapeutic gene is adjacent at the 5' end by the first nuclease recognition sequence, and the first recognition site sequence is on the 3' side of the first nuclear targeting sequence. In some embodiments, the coding sequence for the therapeutic gene is adjacent at the 3' end by the second nuclease recognition site sequence. In some embodiments, the nucleic acid further comprises a second nuclear targeting sequence. In some embodiments, the nucleic acid comprises a first nuclear targeting sequence on the 5' end of the nucleic acid and a second nuclear targeting sequence on the 3' end. In some embodiments, the nucleic acid comprises NTS(1)-NRS(1)-SA-TG-NRS(2)-NTS(2) in the 5' to 3' direction, where NTS(1) represents the first nuclear targeting sequence, NTS(2) represents the second nuclear targeting sequence, NRS(1) represents the first nuclease recognition site sequence, NRS(2) represents the second nuclease recognition site sequence, SA represents a splice acceptor sequence targeting exon 1 of the albumin gene, and TG represents the therapeutic gene. In some embodiments, the 5' to 3' orientation of NRS(1) and NRS(2) follows (a) forward, forward, (b) reverse, reverse, (c) forward, reverse, (d) reverse, forward, where forward indicates the same 5' to 3' orientation as the target nucleic acid sequence and reverse indicates the opposite 5' to 3' orientation as the target nucleic acid sequence. In some embodiments, the nucleic acid comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 12, 13, 16-23, or 32, 33.In some embodiments, the first nuclear targeting sequence or the second nuclear targeting sequence comprises a binding site for LEF / TCF1, HNF1, NFY, CEBP, OCT1, AP1, HNF1-A, HNF1-B, CEBPA, LEF-1, FOX D1, or IRF1. In some embodiments, the first nuclear targeting sequence or the second nuclear targeting sequence comprises a sequence having at least 80% identity to SEQ ID NO: 1, or SEQ ID NO: 1 having SEQ ID NO: 2 added to the 5' or 3' end at the 5' or 3' end of the nuclear targeting sequence. In some embodiments, the therapeutic gene is the Factor VIII (FVIII) gene or a functional fragment thereof. In some embodiments, the FVIII gene or a functional fragment thereof is codon-optimized to remove at least one cytosine-guanine (CG or CpG) motif. In some embodiments, the codon-optimized FVIII gene or a functional fragment thereof comprises a sequence having at least 80% identity to SEQ ID NO: 10.

[0018] In certain embodiments, in the 5' to 3' direction, it includes: (a) a nucleic acid sequence having at least 80% identity to the first intron 1 sequence of the albumin gene and containing KTTN (K = G or T, N = any base) or AAANNN sequence (N = any base); (b) a splice acceptor sequence targeting exon 1 of the albumin gene; (c) a coding sequence for a therapeutic gene encoding a functional liver enzyme (such as factor VIII) linked to a polyadenylation signal and a splice acceptor sequence; (d) a nucleic acid sequence having at least 80% identity to the second intron 1 sequence of the albumin gene and containing KTTN (K = G or T, N = any base) or AAANNN (N = any base), wherein the second intron 1 sequence is on the 3' side of the first intron 1 sequence of the albumin gene. A viral vector as disclosed herein is provided. In some embodiments, when the coding sequence is adjacent to the 5' end of the nucleic acid, the nucleic acid further includes a first nuclear targeting sequence (NTS) on the 5' side of the recognition site sequence, or when the coding sequence is adjacent to the 3' end of the nucleic acid, the nucleic acid further includes a first nuclear targeting sequence (NTS) on the 3' side of the recognition site sequence. In some embodiments, the nucleic acid further includes a second nuclear targeting sequence (NTS). In some embodiments, the nucleic acid includes a first nuclear targeting sequence on the 5' end of the nucleic acid and a second nuclear targeting sequence on the 3' end of the nucleic acid. In some embodiments, the nucleic acid includes NTS(1)-NRS(1)-SA-TG-NRS(2)-NTS(2) in the 5' to 3' direction, where NTS(1) represents the first nuclear targeting sequence, NTS(2) represents the second nuclear targeting sequence, NRS(1) represents the first nuclease recognition site sequence, NRS(2) represents the second nuclease recognition site sequence, SA represents a splice acceptor sequence targeting exon 1 of the albumin gene, and TG represents a coding sequence for a therapeutic gene. In some embodiments, the 5' to 3' orientation of NRS(1) and NRS(2) follows (a) forward, forward, (b) reverse, reverse, (c) forward, reverse, (d) reverse, forward, where forward indicates the same 5' to 3' orientation as the target nucleic acid sequence and reverse indicates the opposite 5' to 3' orientation to the target nucleic acid sequence.In some embodiments, the nucleic acid comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 12, 13, 16-23, or 32, 33. In some embodiments, the first nuclear targeting sequence or the second nuclear targeting sequence comprises a binding site for LEF / TCF1, HNF1, NFY, CEBP, OCT1, AP1, HNF1-A, HNF1-B, CEBPA, LEF-1, FOX D1, or IRF1. In some embodiments, the first nuclear targeting sequence or the second nuclear targeting sequence comprises a sequence having at least 80% identity to SEQ ID NO: 1 having SEQ ID NO: 2 added to the 5' or 3' end. In some embodiments, the therapeutic gene is the Factor VIII (FVIII) gene or a functional fragment thereof. In some embodiments, the FVIII gene or a functional fragment thereof is codon-optimized to remove at least one cytosine-guanine (CG or CpG) motif. In some embodiments, the codon-optimized FVIII gene or a functional fragment thereof comprises a sequence having at least 80% identity to SEQ ID NO: 10.

[0019] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, which illustrates only exemplary embodiments of the present disclosure. As will be understood, the present disclosure is capable of other and different embodiments, and some of the details thereof are capable of modification in various obvious respects without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.

[0020] The novel features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained from the following detailed description, which illustrates exemplary embodiments in which the principles of the present disclosure are utilized, and from the appended drawings.

Brief Description of the Drawings

[0021]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

[0022] Brief Description of the Sequence Listing The sequence listing submitted with this specification provides exemplary polynucleotide sequences and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. The following is an exemplary description of the sequences therein.

[0023] SEQ ID NO: 1 shows the DNA sequence of the human albumin target site.

[0024] SEQ ID NO: 2 shows the DNA sequence of the transcriptional enhancer.

[0025] SEQ ID NOs: 3-6 show the DNA sequences of the mouse albumin target sites.

[0026] SEQ ID NO: 7 shows the nucleotide sequence of the 20bp spacer.

[0027] SEQ ID NO: 8 shows the DNA sequences of the human albumin target site and the transcriptional enhancer.

[0028] SEQ ID NO: 9 shows the DNA sequence of the splice acceptor.

[0029] SEQ ID NO: 10 shows the DNA sequence of the human Factor VIII coding sequence.

[0030] SEQ ID NO: 11 shows the DNA sequences of the stop codon, spacer, and polyadenylation signal.

[0031] SEQ ID NOs: 12-13, 16-23, and 32-33 show the DNA sequences of the Factor VIII donor template.

[0032] SEQ ID NOs: 14, 15, and 24, 25 show the nucleotide sequences of the MG29-1 sgRNA targeting mouse albumin.

[0033] SEQ ID NOs: 26-27 show the nucleotide sequences of the MG3-6 / 3-4 sgRNA targeting mouse albumin.

[0034] SEQ ID NOs: 28-29 show the nucleotide sequences of the PCR primers.

[0035] SEQ ID NO: 30 shows the DNA sequence encoding MG3-6 / 3-4 mRNA.

[0036] SEQ ID NO: 31 shows the DNA sequence encoding MG29-1 mRNA.

[0037] SEQ ID NOs: 37, 39, and 41 show the nucleotide sequences of the transcription factor binding sequences.

[0038] SEQ ID NOs: 43-45 show the nucleotide sequences of the MG29-1 sgRNA targeting mouse albumin.

[0039] SEQ ID NOs: 46-49 show the nucleotide sequences of the PCR primers.

[0040] SEQ ID NO: 50 shows the nucleotide sequence of the MG29-1 sgRNA targeting mouse albumin.

[0041] SEQ ID NO: 51 shows the DNA sequence encoding spCas9 mRNA.

[0042] SEQ ID NO: 52 shows the spCas9 amino acid sequence.

[0043] SEQ ID NO: 53 shows the DNA sequence encoding MG29-1 mRNA.

[0044] SEQ ID NO: 54 shows the MG29-1 amino acid sequence.

[0045] SEQ ID NO: 55 shows the nucleotide sequence of the MG29-1 sgRNA targeting mouse albumin.

[0046] SEQ ID NOs: 56-59 show the DNA sequence of the Factor VIII donor template.

[0047] SEQ ID NOs: 60-63 show the nucleotide sequence of the MG29-1 sgRNA targeting mouse albumin.

[0048] SEQ ID NOs: 64-68 show the nucleotide sequence of the MG29-1 sgRNA targeting human albumin intron 1.

[0049] SEQ ID NO: 69 shows the protein sequence of the protease recognition site.

[0050] SEQ ID NOs: 71-79 show the protein sequence intended to replace the B domain of FVIII.

[0051] SEQ ID NO: 80 shows the protein sequence of the SQ linker.

[0052] SEQ ID NOs: 81-85 show the nucleotide sequence of the human FVIII donor cassette.

[0053] SEQ ID NOs: 86-87 show the protein sequences of cynomolgus macaque FVIII sequences with a substituted B domain.

[0054] SEQ ID NO: 88 shows the nucleotide sequence of the cynomolgus macaque FVIII donor cassette.

[0055] SEQ ID NO: 89 shows the protein sequence intended to replace the B domain of FVIII.

[0056] SEQ ID NO: 90 shows the protein sequence of the human FVIII sequence with a substituted B domain.

[0057] SEQ ID NOs: 91-94 show the nucleotide sequences of the human FVIII donor cassette.

[0058] SEQ ID NO: 95 shows the nucleotide sequence of the messenger RNA encoding MG3-6 / 3-4 nuclease with nuclear localization signals added at both the N and C termini.

[0059] SEQ ID NO: 96 shows the protein sequence of MG3-6 / 3-4 nuclease with nuclear localization signals added at both the N and C termini.

[0060] SEQ ID NOs: 97, 98 show the nucleotide sequences of MG3-6 / 3-4 sgRNA targeting mouse albumin intron 1.

DETAILED DESCRIPTION OF THE INVENTION

[0061] Although various embodiments of the present disclosure are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, modifications, and substitutions may occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be used.

[0062] The practice of several of the methods disclosed herein uses, unless otherwise indicated, the techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012), the series Current Protocols in Molecular Biology (F.M. Ausubel, et al. eds.), the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J. MacPherson, B.D. Hames and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)).

[0063] As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Further, the terms "including", "includes", "having", "has", "with", or variants thereof, as used in any of the detailed description and / or claims, are intended to be inclusive in the same manner as the term "comprising".

[0064] The terms "about" or "approximately" mean within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., on the limitations of the measurement system. For example, "about" can mean within one or more standard deviations in accordance with the practice in the art. Alternatively, "about" can mean within a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.

[0065] As used herein, the term "nucleotide" refers to a base-sugar-phosphate combination. Contemplated nucleotides include naturally occurring nucleotides and synthetic nucleotides. Nucleotides are the monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term "nucleotide" includes ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytidine triphosphate (CTP), guanosine triphosphate (GTP) and deoxyribonucleoside triphosphates, e.g., dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives include, for example, [αS]dATP, 7-deaza-dGTP and 7-deaza-dATP, and nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term "nucleotide" encompasses dideoxyribonucleoside triphosphates (ddNTPs) and derivatives thereof. Illustrative examples of ddNTPs include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides can be unlabeled or detectably labeled, such as by using a portion that is optically detectable (e.g., a fluorophore) or a portion that includes a quantum dot. Detectable labels include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels of nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2’7’-dimethoxy-4’5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N’,N’-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4’dimethylaminophenylazo)benzoic acid (DABCYL), cascade blue, Oregon green, Texas red, cyanine and 5-(2’-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, available from Perkin Elmer, Foster City, Calif; FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP, available from Amersham, Arlington Heights, IL; fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP, available from Boehringer Mannheim, Indianapolis, Ind; and chromosome-labeling nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, cascade blue-7-UTP, cascade blue-7-dUTP, fluorescein-12-UTP, fluorescein-12-dUTP, Oregon Green 488-5-dUTP, rhodamine green-5-UTP, rhodamine green-5-dUTP, tetramethylrhodamine-6-UTP, tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP, available from Molecular Probes, Eugene, Oreg. The term nucleotide includes chemically modified nucleotides. An exemplary chemically modified nucleotide is biotin-dNTP.Non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., biotin-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0066] The terms "polynucleotide", "oligonucleotide", and "nucleic acid" are used interchangeably to refer to a polymeric form of nucleotides of any length, in single-stranded, double-stranded, or multi-stranded form, either deoxyribonucleotides or ribonucleotides, or analogs thereof. The polynucleotides contemplated include genes or fragments thereof. Exemplary polynucleotides include, but are not limited to, DNA, RNA, coding or non-coding regions of genes or gene fragments, loci (gene loci) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, cell-free DNA (cfDNA) and cell-free RNA (cfRNA), cell-free polynucleotides including cells, nucleic acid probes, and primers. When referring to T, in polynucleotides, T means U (uracil) in RNA and T (thymine) in DNA. Polynucleotides may be exogenous or endogenous to a cell and / or may be present in a cell-free environment. The term polynucleotide includes modified polynucleotides (e.g., modified backbones, sugars, or nucleobases). When present, modifications to the nucleotide structure are imparted either before or after assembly of the polymer. Non-limiting examples of modifications include 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotide, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein conjugated to a sugar), thiol-containing nucleotides, biotin-conjugated nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. The sequence of nucleotides may be interrupted by non-nucleotide components.

[0067] The terms "transfection" or "transfected" refer to the introduction of a polynucleotide into a cell by a non-viral or virus-based method. The polynucleotide can be a gene sequence encoding a full-length protein or a functional portion thereof. See, for example, Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88.

[0068] The terms "peptide", "polypeptide", and "protein" are used interchangeably herein to refer to a polymer of at least two amino acid residues joined by peptide bonds. The term is not meant to denote a specific length of the polymer and is not intended to imply or distinguish whether the peptide is produced using recombinant techniques, chemical or enzymatic synthesis, or is naturally occurring. The term applies to both naturally occurring amino acid polymers and amino acid polymers containing at least one modified amino acid. In some cases, the polymer is interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins and proteins (e.g., domains) that do or do not have secondary or tertiary structure. The term also encompasses amino acid polymers modified by any other operation, such as disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and conjugation with a labeling component. As used herein, the terms "amino acid" and "amino acids" refer to natural and non-natural amino acids, including but not limited to modified amino acids. Modified amino acids include amino acids that have been chemically modified to contain a group or chemical moiety not naturally present on the amino acid. The term "amino acid" includes both D-amino acids and L-amino acids.

[0069] As used herein, "non-natural" refers to a nucleic acid or polypeptide sequence that does not occur in nature. Non-natural refers to a nucleic acid or polypeptide sequence that does not occur in nature and includes modifications such as mutations, insertions, or deletions. The term non-natural encompasses fusion nucleic acids or polypeptides that encode or exhibit the activity of a nucleic acid or polypeptide sequence (e.g., enzyme activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.) to which a non-natural sequence is fused. Non-natural nucleic acid or polypeptide sequences include those that are linked by genetic manipulation to a nucleic acid or polypeptide sequence (or variant thereof) that occurs in nature to produce a chimeric nucleic acid or polypeptide sequence that encodes a chimeric nucleic acid or polypeptide.

[0070] As used herein, the term "promoter" refers to a regulatory DNA region that controls the transcription or expression of a polynucleotide (e.g., a gene) and can be located adjacent to or overlapping a nucleotide or region of nucleotides at which RNA transcription is initiated. A promoter may often contain a specific DNA sequence that binds to a protein factor, often called a transcription factor, which facilitates the binding of RNA polymerase to the DNA, thereby resulting in gene transcription. Eukaryotic basal promoters typically, but not always, contain a TATA-box and / or a CAAT box.

[0071] As used herein, the term "expression" refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (such as into mRNA or other RNA transcripts), and / or the process by which the transcribed mRNA is later translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product". When a polynucleotide is derived from genomic DNA, the term expression includes mRNA splicing in eukaryotic cells.

[0072] As used herein, "operably linked," "operable linkage," "operatively linked," or grammatical equivalents thereof, refers to the arrangement of genetic elements, such as a promoter, enhancer, polyadenylation sequence, etc., where the operation (e.g., movement or activation) of a first genetic element has some effect on a second genetic element. The effect on the second genetic element may be of the same type as the operation of the first genetic element, but need not be. For example, if the movement of a first element causes the activation of a second element, the two genetic elements are operably linked. For example, a regulatory element, which may include a promoter sequence and / or enhancer sequence, is operably linked to a coding region if it aids in initiating transcription of the coding sequence. There may be intervening residues between the regulatory element and the coding region as long as this functional relationship is maintained.

[0073] As used herein, "vector" refers to a macromolecule or associated macromolecules that contain or are related to a polynucleotide and mediate the delivery of the polynucleotide into a cell. Examples of vectors include nucleus-based vectors (e.g., plasmids and viral vectors) and liposomes. Exemplary nucleic acid-based vectors generally contain genetic elements, such as regulatory elements, operably linked to a gene to facilitate expression of the gene in a target.

[0074] As used herein, "expression cassette" and "nucleic acid cassette" are used interchangeably to refer to components of a vector that include a combination of nucleic acid sequences or elements (e.g., a therapeutic gene, a promoter, and a terminator) that are either co-expressed or operably linked for expression. The terms encompass expression cassettes that include a combination of regulatory elements operably linked for expression and a gene or genes.

[0075] A "functional fragment" of a DNA or protein sequence refers to a fragment that retains a biological activity (either functional or structural) that is substantially similar to the biological activity of the full-length DNA or protein sequence. The biological activity of a DNA sequence includes the ability to affect expression in a manner resulting from the full-length sequence.

[0076] The terms "engineered", "synthetic", and "artificial" are used interchangeably herein to refer to an object modified by human intervention. For example, the term refers to a polynucleotide or polypeptide that does not occur in nature. An engineered peptide need not have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) to a naturally occurring human protein. For example, the VPR domain and the VP64 domain are synthetic transactivation domains. Non-limiting examples include nucleic acids modified by changing their sequence to a sequence that does not occur in nature, nucleic acids modified by ligating to nucleic acids not naturally associated such that the ligation product possesses a function not present in the original nucleic acid, engineered nucleic acids synthesized in vitro using sequences that do not occur in nature, proteins modified by changing their amino acid sequence to a sequence that does not occur in nature, and engineered proteins that acquire a new function or property. An "engineered" system includes at least one engineered component.

[0077] The term "tracrRNA" or "tracr sequence" means to trans-activate CRISPR RNA. The tracrRNA interacts with CRISPR (cr)RNA to form the guide (g)RNA of type II and subtype V-B CRISPR-Cas systems. When the tracrRNA is engineered, it can have about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity and / or sequence similarity to an exemplary wild-type tracrRNA sequence (e.g., tracrRNA derived from S. pyogenes, S. aureus). The tracrRNA can refer to a modified form of tracrRNA that can include nucleotide changes such as deletions, insertions, or substitutions, variants, mutations, or chimeras. The term tracrRNA encompasses nucleic acids that can be at least about 60% identical to an exemplary wild-type tracrRNA sequence (e.g., tracrRNA derived from S. pyogenes, S. aureus, etc.) over a stretch of at least six contiguous nucleotides. For example, the tracrRNA sequence can have at least about 60% identity, at least about 65% identity, at least about 70% identity, at least about 75% identity, at least about 80% identity, at least about 85% identity, at least about 90% identity, at least about 95% identity, at least about 98% identity, at least about 99% identity, or 100% identity to an exemplary wild-type tracrRNA sequence (e.g., tracrRNA derived from S. pyogenes, S. aureus, etc.) over a stretch of at least six contiguous nucleotides. The type II tracrRNA sequence can be predicted on the genomic sequence by identifying a region having complementarity to a part of the repeat sequence in the adjacent CRISPR array.

[0078] As used herein, the term “guide nucleic acid” or “guide polynucleotide” refers to a nucleic acid that hybridizes to a target nucleic acid, thereby directing an associated nuclease to the target nucleic acid. The guide nucleic acid is, without limitation, RNA (guide RNA or gRNA), DNA, or a mixture of RNA and DNA. The guide nucleic acid may include a crRNA or a tracrRNA, or a combination of both. The term guide nucleic acid encompasses engineered guide nucleic acids and programmable guide nucleic acids that specifically bind to a target nucleic acid. A portion of the target nucleic acid may be complementary to a portion of the guide nucleic acid. The strand of the double-stranded target polynucleotide that is complementary to and hybridizes with the guide nucleic acid is the complementary strand. The strand of the double-stranded target polynucleotide that is complementary to the complementary strand and thus not complementary to the guide nucleic acid is called the non-complementary strand. A guide nucleic acid having a polynucleotide strand is a “single guide nucleic acid”. A guide nucleic acid having two polynucleotide strands is a “double guide nucleic acid”. Otherwise, the term “guide nucleic acid” is inclusive and refers to both single guide nucleic acids and double guide nucleic acids. The guide nucleic acid may include a segment referred to as a “nucleic acid targeting segment” or “nucleic acid targeting sequence” or “spacer”. The nucleic acid targeting segment may include a sub-segment referred to as a “protein binding segment” or “protein binding sequence” or “Cas protein binding segment”.

[0079] As used herein, the term “Cas12a” refers to a class 2, type V-A Cas endonuclease family that (a) uses a relatively small guide RNA (about 42-44 nucleotides) that is processed by the nuclease itself after transcription from a CRISPR array and (b) cleaves DNA to leave staggered cleavage sites.

[0080] As used herein, the term "RuvC_III domain" refers to the third discontinuous segment of the RuvC endonuclease domain (the RuvC nuclease domain is composed of three discontinuous segments, RuvC_I, RuvC_II, and RuvC_III). The RuvC domain or segment thereof can generally be identified by alignment to a documented domain sequence, structural alignment to a protein with an annotated domain, or comparison to a hidden Markov model (HMM) constructed based on a documented domain sequence (e.g., Pfam HMM PF18541 for RuvC_III).

[0081] As used herein, the term "wedge" (WED) domain refers to a domain that mainly interacts with the sgRNA and PAM duplex: anti-duplex (e.g., present in Cas proteins). The WED domain can generally be identified by alignment to a documented domain sequence, structural alignment to a protein with an annotated domain, or comparison to a hidden Markov model (HMM) constructed based on a documented domain sequence.

[0082] As used herein, the term "PAM interaction domain" or "PI domain" refers to a domain that interacts with the protospacer adjacent motif (PAM) outside the seed sequence within the region targeted by the Cas protein. Examples of PAM interaction domains include, but are not limited to, the topoisomerase homology (TOPO) domain and the C-terminal domain (CTD) present in Cas proteins. The PAM interaction domain or segment thereof can generally be identified by alignment to a documented domain sequence, structural alignment to a protein with an annotated domain, or comparison to a hidden Markov model (HMM) constructed based on a documented domain sequence.

[0083] As used herein, the term "REC domain" refers to a domain (e.g., present in a Cas protein) that includes at least one of two segments (REC1 or REC2), which are alpha helix domains thought to contact guide RNA. The REC domain or segment thereof can generally be identified by alignment to a documented domain sequence, structural alignment to a protein with an annotated domain, or comparison to a hidden Markov model (HMM) constructed based on a documented domain sequence (e.g., Pfam PF19501 for domain REC1).

[0084] As used herein, the term "BH domain" refers to a domain (e.g., present in a Cas protein) that is a bridging helix between the NUC lobe and the REC lobe of a type II Cas enzyme. The BH domain or segment thereof can generally be identified by alignment to a documented domain sequence, structural alignment to a protein with an annotated domain, or comparison to a hidden Markov model (HMM) constructed based on a documented domain sequence (e.g., Pfam PF16593 for domain BH).

[0085] As used herein, the term "HNH domain" refers to an endonuclease domain having characteristic histidine and asparagine residues. The HNH domain can generally be identified by alignment to a documented domain sequence, structural alignment to a protein with an annotated domain, or comparison to a hidden Markov model (HMM) constructed based on a documented domain sequence (e.g., Pfam HMM PF01844 for domain HNH).

[0086] "Factor VIII" or "FVIII" refers to the antihemophilic factor (i.e., a blood coagulation or clotting protein) encoded by the FVIII gene. In its inactive form, Factor VIII is bound to von Willebrand factor. In response to injury, the two factors separate, and FVIII activates and interacts with FIX to initiate a chain of chemical reactions leading to a blood clot. A genetic deficiency of Factor VIII results in hemophilia A.

[0087] The term "donor template" refers to a polynucleotide comprising an exogenous polynucleotide sequence (e.g., the nucleic acid sequence of a therapeutic gene) and one or more polynucleotide sequences for mediating recombination by non-homologous end joining (NHEJ) or homology-directed repair (HDR), etc.

[0088] As used herein, the term "complex" refers to the joining of at least two components. Each of the two components may retain the properties / activities it had prior to forming the complex or acquire properties as a result of forming the complex. The joining includes, but is not limited to, covalent bonding, non-covalent bonding (i.e., hydrogen bonding, ionic interactions, van der Waals interactions, and hydrophobic bonding), the use of linkers, fusion, or any other suitable method. The intended components of the complex include polynucleotides, polypeptides, or combinations thereof. For example, a complex includes an endonuclease and a guide polynucleotide.

[0089] The terms "sequence identity" or "identity rate" in the context of two or more nucleic acid or polypeptide sequences, when measured using a sequence comparison algorithm, refer to sequences that are identical or have a specific percentage of identical amino acid residues or nucleotides when compared and aligned for maximum correspondence over a local or global comparison window (e.g., in pairwise alignment) of two or (e.g., in multiple sequence alignment) more sequences. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using the BLOSUM62 scoring matrix with parameters of 3 word length (W), 10 expectation value (E), and gap costs set at 11 existence, 1 extension, and using conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using the PAM30 scoring setting gap costs of 9 for open gaps and 1 for extension gaps for sequences shorter than 30 residues with parameters of 2 word length (W), 1000000 expectation value (E) (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW using Smith-Waterman homology search algorithm parameters of 2 match, -1 mismatch, and -1 gap; MUSCLE using default parameters; MAFFT using parameters of 2 trees and 1000 maximum iterations; Novafold using default parameters; HMMER hmmalign using default parameters.

[0090] In the context of two or more nucleic acid or polypeptide sequences, the term "optimally aligned" refers to two (e.g., in pairwise alignment) or more (e.g., in multiple sequence alignment) sequences aligned with a maximum match of amino acid residues or nucleotides, determined, for example, by an alignment that generates the highest or "optimized" percent identity score.

[0091] Variants of any of the enzymes described herein having one or more conservative amino acid substitutions are included in the present disclosure. Such conservative substitutions can be made in the amino acid sequence of the polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids with similar hydrophobicity, polarity, and R-chain length to each other. Additionally, or alternatively, by comparing the aligned sequences of homologous proteins from different species, conservative substitutions can be identified by finding amino acid residues that vary between species (e.g., non-conserved residues) without changing the basic function of the encoded protein. Such conservatively substituted variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of the endonuclease protein sequences described herein (e.g., the MG3 or MG29 family endonucleases described herein, or any other family nuclease described herein). In some embodiments, such conservatively substituted variants are functional variants. Such functional variants can include sequences with substitutions such that the activity of one or more important active site residues or guide RNA binding residues of the endonuclease is not disrupted. In some embodiments, any functional variant of the proteins described herein lacks substitutions of at least one conserved residue or functional residue.

[0092] Also included in the present disclosure are variants of any of the enzymes described herein (e.g., activity-reduced variants) having substitutions of one or more catalytic residues to reduce or remove the activity of the enzyme. In some embodiments, an activity-reduced variant as a protein described herein includes disruptive substitutions of at least one, at least two, or all three catalytic residues.

[0093] Tables of conservative substitutions that result in functionally similar amino acids are available from a variety of references (see, e.g., Creighton, Proteins: Structures and Molecular Properties (W H Freeman & Co.; 2nd edition (December 1993))). The following eight groups each contain amino acids that are conservative substitutions for one another: 1) Alanine (A), Glycine (G); 2) Aspartic acid (D), Glutamic acid (E); 3) Asparagine (N), Glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W); 7) Serine (S), Threonine (T); and 8) Cysteine (C), Methionine (M)

[0094] Summary Treatment of diseases and disorders caused by genetic deficiencies can involve using a gene editing system to correct the genetic deficiency. This can be done by integrating a copy of the gene into the genome at an appropriate site and in appropriate cells or tissues such that the gene is expressed and produces a functional protein.

[0095] Hemophilia A is caused by mutations in the FVIII gene that reduce the expression of factor VIII (FVIII) or inactivate the function of the FVIII protein. Because many different mutations in FVIII cause hemophilia A, gene therapy approaches that are not specific to individual mutations are preferred. A promising approach is the complementation of a defective genomic copy of FVIII with a functional transgenic copy of the gene. In one embodiment of this approach, the functional copy of the FVIII gene is delivered to hepatocytes in the liver by systemic (e.g., intravenous) administration of a vector comprising a nuclease and a guide polynucleotide that targets a target nucleic acid at a safe harbor locus such as the albumin locus. The FVIII gene is then integrated into the target nucleic acid upon cleavage by the nuclease via non-homologous end joining (NHEJ) or homology-directed repair (HDR), a combination of two DNA repair mechanisms, or by other DNA repair mechanisms.

[0096] CRISPR / Cas enzymes The discovery of new Cas enzymes with unique functionality and structure has the potential to improve further gene editing techniques, speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of clustered regularly interspaced short palindromic repeat (CRISPR) systems in microorganisms and the full diversity of microbial species, there are relatively few CRISPR / Cas enzymes that have been functionally characterized in the literature. This is in part because under laboratory conditions, a vast number of microbial species are not easily cultured. Metagenomic sequencing from natural environmental niches containing a large number of microbial species has the potential to dramatically increase the number of characterized new CRISPR / Cas systems and accelerate the discovery of new oligonucleotide editing functions. A recent and fruitful example of such an approach is demonstrated by the discovery in 2016 of the CasX / CasY CRISPR system from metagenomic analysis of natural microbial communities.

[0097] The CRISPR / Cas system is an RNA-guided nuclease complex that functions as an adaptive immune system in microorganisms. In their natural context, CRISPR / Cas systems occur at CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) operons or loci, which generally consist of two parts: (i) an array of short repeat sequences (30 - 40 bp) separated by short spacer sequences that encode RNA-based targeting elements, and (ii) an ORF that encodes a Cas nuclease. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6 - 8 nucleic acids of the target nucleic acid and the crRNA guide, and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a specific vicinity of the target nucleic acid sequence that depends on a specific Cas nuclease (PAM is typically a sequence not commonly represented within the host genome). Depending on the exact function and composition of the system, CRISPR-Cas systems are generally classified into two classes, five types, and sixteen subtypes based on shared functional features and evolutionary similarities.

[0098] Class 1 CRISPR-Cas systems have large multi-subunit effector complexes and include type I, III, and IV Cas nucleases. Class 2 CRISPR-Cas systems generally have single polypeptide multi-domain nuclease effectors and include type II, V, and VI Cas nucleases.

[0099] Type I CRISPR-Cas systems are considered to have moderate complexity in terms of components. In Type I CRISPR-Cas systems, an array of RNA targeting elements is transcribed as a long precursor crRNA (pre-crRNA) that is processed by repetitive elements, and short mature crRNAs that orient the nuclease complex to nucleic acid targets are subsequently released following a suitable short consensus sequence called the protospacer adjacent motif (PAM). This processing occurs via the endoribonuclease subunit (Cas6) of a large endonuclease complex called Cascade, which also includes the nuclease (Cas3) protein component of the crRNA-guided nuclease complex. Cas I nucleases function primarily as DNA nucleases.

[0100] Type III CRISPR systems are characterized by the presence of a central nuclease known as Cas10, along with repeated associated mysterious proteins (RAMP) that include Csm or Cmr protein subunits. Similar to Type I systems, mature crRNAs are processed from pre-crRNAs using a Cas6-like enzyme. Unlike Type I and II systems, Type III systems are thought to target and cleave DNA-RNA duplexes (such as the DNA strand used as a template for RNA polymerase).

[0101] Type IV CRISPR-Cas systems possess effector complexes composed of a highly reduced large subunit nuclease (csf1), two genes of the RAMP protein group of Cas5 (csf3) and Cas7 (csf2), and in some cases, genes for predicted small subunits, and such systems are generally found on endogenous plasmids.

[0102] Class 2 CRISPR-Cas systems generally have a single polypeptide multi-domain nuclease effector and include Types II, V, and VI.

[0103] Type II CRISPR-Cas systems are considered the simplest in terms of components. In Type II CRISPR-Cas systems, processing of the CRISPR array into mature crRNAs does not require the presence of a special endonuclease subunit. Instead, it is a small trans-encoded crRNA (tracrRNA) with regions complementary to the array repeat sequences. The tracrRNA interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequences to form a precursor dsRNA structure and is cleaved by endogenous RNase III to generate a mature effector enzyme loaded with both tracrRNA and crRNA. Cas II nuclease is identified as a DNA nuclease. Type II effectors generally exhibit a structure that includes a RuvC-like endonuclease domain that fits into an RNase H fold with an unrelated HNH nuclease domain inserted into the fold of the RuvC-like nuclease domain. The RuvC-like domain is involved in cleavage of the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is involved in cleavage of the displaced DNA strand.

[0104] Type V CRISPR-Cas systems are characterized by a nuclease effector (e.g., Cas12) structure similar to that of Type II effectors, including a RuvC-like domain. Similar to Type II, most (but not all) Type V CRISPR systems use tracrRNA to process pre-crRNA into mature crRNAs. However, unlike Type II systems that require RNase III to cleave pre-crRNA into multiple crRNAs, Type V systems can use the effector nuclease itself to cleave pre-crRNA. Similar to Type II CRISPR-Cas systems, Type V CRISPR-Cas systems are also identified as DNA nucleases. Different from Type II CRISPR-Cas systems, some Type V enzymes (e.g., Cas12a) appear to have robust single-stranded non-specific deoxyribonuclease activity that is activated by the first crRNA-directed cleavage of double-stranded target sequences.

[0105] Type VI CRISPR-Cas systems have RNA-guided RNA endonucleases. Instead of an RuvC-like domain, the single polypeptide effector of the Type VI system (e.g., Cas13) contains two HEPN ribonuclease domains. Unlike both Type II and V systems, the Type VI system does not appear to require a tracrRNA to process pre-crRNA into crRNA. However, similar to the Type V system, some Type VI systems (e.g., C2C2) are thought to have robust single-stranded non-specific nuclease (ribonuclease) activity that is activated by the first crRNA-directed cleavage of the target RNA.

[0106] Gene editing system In certain embodiments, engineered nuclease systems are disclosed herein that include: a) an endonuclease; b) an engineered guide polynucleotide that forms a complex with the endonuclease and is configured to hybridize to a target nucleic acid sequence within or in an intron of an albumin gene; and c) a donor template comprising a nucleic acid sequence encoding a Factor VIII (FVIII) gene or a functional fragment thereof.

[0107] MG endonuclease Systems and methods for the replenishment of liver enzymes are disclosed herein. In some embodiments, the systems and methods include an endonuclease. In some embodiments, the endonuclease is functional in prokaryotic or eukaryotic cells for in vitro, in vivo, or ex vivo applications. In some embodiments, the endonuclease is a nucleic acid-guided nuclease, a chimeric nuclease, or a fusion nuclease.

[0108] In some embodiments, the endonuclease is MG29-1 (i.e., SEQ ID NO: 54). In some embodiments, the endonuclease is MG3-6 / 3-4 (i.e., SEQ ID NO: 96). MG29-1 is a type V CRISPR nuclease, and MG3-6 / 3-4 is a type II CRISPR nuclease created by exchanging the PAM interaction domain (PID) of MG3-6 with that of MG3-4 to alter the PAM recognition specificity. In some embodiments, the PAM of MG29-1 is functionally defined as KTTN (K = G or T, N = any base) in mammalian cells. In some embodiments, the PAM of MG3-6 or MG3-4 is functionally defined as AAANN (N = any base) in mammalian cells.

[0109] In some embodiments, the endonuclease comprises a sequence having a sequence identity of at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease comprises a sequence having at least about 70% identity to either one of SEQ ID NO: 54 and 96. In some embodiments, the endonuclease comprises a sequence having at least about 75% identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease comprises a sequence having at least about 80% identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease comprises a sequence having at least about 85% identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease comprises a sequence having at least about 90% identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease comprises a sequence having at least about 95% identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease comprises a sequence having at least about 96% identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease comprises a sequence having at least about 97% identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease comprises a sequence having at least about 98% identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease comprises a sequence having at least about 99% identity to SEQ ID NO: 54 or SEQ ID NO: 96. In some embodiments, the endonuclease comprises a sequence having 100% identity to SEQ ID NO: 54 or SEQ ID NO: 96.

[0110] In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least about 70% identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least about 75% identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least about 80% identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least about 85% identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least about 90% identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least about 95% identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least about 96% identity to any one of SEQ ID NOs: 30, 31, 53, and 95.In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least about 97% identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least about 98% identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having at least about 99% identity to any one of SEQ ID NOs: 30, 31, 53, and 95. In some embodiments, the endonuclease is encoded by a nucleic acid sequence having 100% identity to any one of SEQ ID NOs: 30, 31, 53, and 95.

[0111] In some embodiments, the endonuclease comprises one or more fragments or domains of a nuclease, such as a nucleic acid-guided nuclease. In some embodiments, the endonuclease comprises one or more fragments or domains of a nuclease from an ortholog of a organism, genus, species, or other taxonomic group described herein. In some embodiments, the endonuclease comprises one or more fragments or domains from nuclease orthologs of different species.

[0112] In some embodiments, the endonuclease comprises one or more fragments or domains of a nuclease, such as a nucleic acid-guided nuclease. In some embodiments, the endonuclease comprises one or more fragments or domains of a nuclease from an ortholog of a organism, genus, species, or other taxonomic group described herein. In some embodiments, the endonuclease comprises one or more fragments or domains from nuclease orthologs of different species.

[0113] In some embodiments, the endonuclease comprises one or more fragments or domains from at least two different nucleases. In some embodiments, the endonuclease comprises one or more fragments or domains from at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more different nucleases. In some embodiments, the endonuclease comprises one or more fragments or domains from at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleases from different species. In some embodiments, the endonuclease comprises two fragments or domains, each from a different nuclease. In some embodiments, the endonuclease comprises three fragments or domains, each from a different nuclease. In some embodiments, the endonuclease comprises four fragments or domains, each from a different nuclease. In some embodiments, the endonuclease comprises five fragments or domains, each from a different nuclease. In some embodiments, the endonuclease comprises three fragments or domains, with at least one fragment or domain being from a different nuclease. In some embodiments, the endonuclease comprises four fragments or domains, with at least one fragment or domain being from a different nuclease. In some embodiments, the endonuclease comprises five fragments or domains, with at least one fragment or domain being from a different nuclease.

[0114] In some embodiments, the linkage between fragments or domains from different nucleases or species occurs in a stretch of unstructured region. An unstructured region in a polynucleotide includes regions that do not have predicted secondary structure elements such as, for example, an alpha helix or a beta strand. The unstructured region can include, for example, regions that are exposed within the protein structure, loop regions, or regions that are not conserved within various protein orthologs, as predicted by sequence or structural alignment.

[0115] In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLSs) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises any one of the sequences of SEQ ID NOs: 144-159. In some embodiments, the NLS comprises any one of the sequences of SEQ ID NOs: 144-159, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of the sequences of SEQ ID NOs: 144-159. In some embodiments, the NLS comprises a sequence having at least about 80% identity to the sequences of SEQ ID NOs: 144-159. In some embodiments, the NLS comprises a sequence having at least about 85% identity to the sequences of SEQ ID NOs: 144-159. In some embodiments, the NLS comprises a sequence having at least about 90% identity to the sequences of SEQ ID NOs: 144-159. In some embodiments, the NLS comprises a sequence having at least about 91% identity to the sequences of SEQ ID NOs: 144-159. In some embodiments, the NLS comprises a sequence having at least about 92% identity to the sequences of SEQ ID NOs: 144-159. In some embodiments, the NLS comprises a sequence having at least about 93% identity to the sequences of SEQ ID NOs: 144-159. In some embodiments, the NLS comprises a sequence having at least about 94% identity to the sequences of SEQ ID NOs: 144-159. In some embodiments, the NLS comprises a sequence having at least about 95% identity to the sequences of SEQ ID NOs: 144-159. In some embodiments, the NLS comprises a sequence having at least about 96% identity to the sequences of SEQ ID NOs: 144-159. In some embodiments, the NLS comprises a sequence having at least about 97% identity to the sequences of SEQ ID NOs: 144-159.In some embodiments, the NLS comprises a sequence having at least about 98% identity to SEQ ID NOs: 144-159. In some embodiments, the NLS comprises a sequence having at least about 99% identity to SEQ ID NOs: 144-159. In some embodiments, the NLS comprises a sequence having 100% identity to SEQ ID NOs: 144-159.

[0116]

Table 1A

[0117] Guide polynucleotide The systems and methods for supplementing liver enzymes described herein may include a guide polynucleotide, e.g., a guide ribonucleic acid (gRNA) for supplementing liver enzymes, a single gRNA, or a dual guide RNA. When referring to T, in a polynucleotide, T means U (uracil) in RNA and T (thymine) in DNA.

[0118] In some embodiments, the target gene or locus is albumin. In some embodiments, the guide polynucleotide targets or hybridizes to a target nucleic acid sequence in albumin. In some embodiments, the guide polynucleotide targets or hybridizes to a target nucleic acid sequence in intron 1 of albumin. In some embodiments, the guide polynucleotide targets or hybridizes to a target nucleic acid sequence in exon 1 of albumin.

[0119] In some embodiments, the target gene or locus is albumin. In some embodiments, the guide polynucleotide targeting albumin is any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98, or is encoded by a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 80% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 85% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 90% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98.In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 95% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 96% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 97% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 98% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 99% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide is encoded by a sequence having 100% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98.

[0120] In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to a target nucleic acid sequence (e.g., SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98) within the albumin gene or within an intron of the albumin gene. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to a sequence having at least about 80% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to a sequence having at least about 85% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to a sequence having at least about 90% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to a sequence having at least about 95% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to a sequence having at least about 96% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98.In some embodiments, the guide polynucleotide hybridizes to, or targets, a sequence complementary to a sequence having at least about 97% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide hybridizes to, or targets, a sequence complementary to a sequence having at least about 98% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide hybridizes to, or targets, a sequence complementary to a sequence having at least about 99% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98. In some embodiments, the guide polynucleotide hybridizes to, or targets, a sequence complementary to a sequence having 100% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98.

[0121] In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to a target nucleic acid sequence (e.g., SEQ ID NOs: 3-6) within the albumin gene or within an intron of the albumin gene. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence by any one of SEQ ID NOs: 3-6, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 3-6. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 80% identity to any one of SEQ ID NOs: 3-6. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 85% identity to any one of SEQ ID NOs: 3-6. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 90% identity to any one of SEQ ID NOs: 3-6. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 95% identity to any one of SEQ ID NOs: 3-6. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 96% identity to any one of SEQ ID NOs: 3-6. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 97% identity to any one of SEQ ID NOs: 3-6. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 98% identity to any one of SEQ ID NOs: 3-6. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 99% identity to any one of SEQ ID NOs: 3-6.In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having 100% identity to any one of SEQ ID NOs: 3-6.

[0122] In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to a target nucleic acid sequence (e.g., SEQ ID NO: 1, 2, and 8) within the albumin gene or within an intron of the albumin gene. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence by any one of SEQ ID NO: 1, 2, and 8, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NO: 1, 2, and 8. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 80% identity to any one of SEQ ID NO: 1, 2, and 8. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 85% identity to any one of SEQ ID NO: 1, 2, and 8. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 90% identity to any one of SEQ ID NO: 1, 2, and 8. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 95% identity to any one of SEQ ID NO: 1, 2, and 8. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 96% identity to any one of SEQ ID NO: 1, 2, and 8. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 97% identity to any one of SEQ ID NO: 1, 2, and 8. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 98% identity to any one of SEQ ID NO: 1, 2, and 8. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 99% identity to any one of SEQ ID NO: 1, 2, and 8.In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having 100% identity to any one of SEQ ID NOs: 1, 2, and 8.

[0123] In some embodiments, the guide polynucleotide is configured to form a complex with an endonuclease. In some embodiments, the guide polynucleotide binds to the endonuclease to form a complex. In some embodiments, the guide polynucleotide binds to the endonuclease (e.g., non-covalently through electrostatic interactions or hydrogen bonds) to form a complex. In some embodiments, the guide polynucleotide fuses to the endonuclease to form a complex.

[0124] In some embodiments, the guide polynucleotide comprises a spacer sequence. In some embodiments, the spacer sequence is configured to hybridize to a target nucleic acid sequence. In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence.

[0125] In some embodiments, the guide polynucleotide (e.g., gRNA) targets a gene or locus in a cell. In some embodiments, the guide polynucleotide targets a gene or locus in a mammalian cell. In some embodiments, the mammalian cell is a porcine, bovine, caprine, ovine, rodent, rat, mouse, non-human primate, or human cell.

[0126] In some embodiments, a guide polynucleotide (e.g., guide RNA) includes various structural elements including, but not limited to, a spacer sequence that binds to a protospacer sequence (target sequence), crRNA, and optionally a tracrRNA. In some embodiments, the genome editing system includes a CRISPR guide RNA. In some embodiments, the guide RNA includes a crRNA that includes a spacer sequence. In some embodiments, the guide RNA additionally includes a tracrRNA or a modified tracrRNA.

[0127] In some embodiments, the systems provided herein include one or more guide polynucleotides. In some embodiments, the guide polynucleotide includes a sense sequence. In some embodiments, the guide polynucleotide includes an antisense sequence. In some embodiments, the guide polynucleotide includes a nucleotide sequence other than a region that is complementary or substantially complementary to a region of the target sequence. For example, crRNA is part of or considered part of the guide polynucleotide or is included in a guide polynucleotide, such as a crRNA:tracrRNA chimera.

[0128] In some embodiments, the guide polynucleotide includes synthetic or modified nucleotides. In some embodiments, the guide polynucleotide includes one or more internucleoside linkers modified from natural phosphodiesters. In some embodiments, all of the internucleoside linkers of the guide polynucleotide, or consecutive nucleotide sequences thereof, are modified. For example, in some embodiments, the internucleoside linkage includes sulfur (S), such as a phosphorothioate internucleoside linkage. In some embodiments, the guide polynucleotide includes more than about 10%, 25%, 50%, 75%, or 90% modified internucleoside linkers. In some embodiments, the guide polynucleotide includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 modified internucleoside linkers (e.g., phosphorothioate internucleoside linkages).

[0129] In some embodiments, the guide polynucleotide comprises a modification to the ribose sugar or nucleobase. In some embodiments, the guide polynucleotide comprises one or more nucleosides comprising a modified sugar moiety, where the modified sugar moiety is a modification of the sugar moiety as compared to the ribose sugar moiety found in deoxyribonucleic acid (DNA) and RNA. In some embodiments, the modification is within the ribose ring structure. Exemplary modifications include replacement with a hexose ring (HNA), a bicyclic ring having a biradical bridge between the C2 and C4 carbons on the ribose ring (e.g., locked nucleic acid (LNA)), or an unlinked ribose ring typically lacking a bond between the C2 and C3 carbons (e.g., UNA), but are not limited thereto. In some embodiments, the sugar-modified nucleoside comprises a bicyclohexose nucleic acid or a tricyclic nucleic acid. In some embodiments, the modified nucleoside comprises a nucleoside where the sugar moiety is replaced with a non-sugar moiety, such as a peptide nucleic acid (PNA) or a morpholino nucleic acid.

[0130] In some embodiments, the guide polynucleotide comprises one or more modified sugars. In some embodiments, the sugar modification comprises a modification made by altering a substituent on the ribose ring to a group other than hydrogen or to the 2'-OH group naturally found in DNA and RNA nucleosides. In some embodiments, the substituent is introduced at the 2', 3', 4', 5' positions, or combinations thereof. In some embodiments, the nucleoside having a modified sugar moiety comprises a 2'-modified nucleoside, such as a 2'-substituted nucleoside. The 2'-sugar modified nucleoside is, in some embodiments, a nucleoside having a substituent other than H or -OH at the 2'-position (2'-substituted nucleoside), or comprises a 2'-linked biradical and includes 2'-substituted nucleosides and LNA (2'-4'-biradical bridged) nucleosides. Examples of 2'-substituted modified nucleosides include, but are not limited to, 2'-O-alkyl-RNA, 2'-O-methyl-RNA, 2'-alkoxy-RNA, 2'-O-methoxyethyl-RNA (MOE), 2'-amino-DNA, 2'-fluoro-RNA, and 2'-F-ANA nucleosides. In some embodiments, the modification in the ribose group comprises a modification at the 2'-position of the ribose group. In some embodiments, the modification at the 2'-position of the ribose group is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-deoxy, and 2'-O-(2-methoxyethyl).

[0131] In some embodiments, the guide polynucleotide comprises one or more modified sugars. In some embodiments, the guide polynucleotide comprises only modified sugars. In some embodiments, the guide polynucleotide comprises greater than about 10%, 25%, 50%, 75%, or 90% modified sugars. In some embodiments, the modified sugar is a bicyclic sugar. In some embodiments, the modified sugar comprises 2'-O-methyl. In some embodiments, the modified sugar comprises 2'-fluoro. In some embodiments, the modified sugar comprises a 2'-O-methoxyethyl group. In some embodiments, the guide polynucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 modified sugars (e.g., comprising 2'-O-methyl or 2'-fluoro).

[0132] In some embodiments, the guide polynucleotide comprises both internucleoside linker modifications and nucleoside modifications. In some embodiments, the guide polynucleotide comprises greater than about 10%, 25%, 50%, 75%, or 90% modified internucleoside linkers, and greater than about 10%, 25%, 50%, 75%, or 90% modified sugars. In some embodiments, the guide polynucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 modified internucleoside linkers (e.g., phosphorothioate internucleoside linkages), and 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 modified sugars (e.g., including 2'-O-methyl or 2'-fluoro).

[0133] In some embodiments, the guide polynucleotide comprises a sequence that is complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some embodiments, the guide polynucleotide comprises a sequence that is complementary to a eukaryotic genomic polynucleotide sequence. In some embodiments, the guide polynucleotide comprises a sequence that is complementary to a fungal genomic polynucleotide sequence. In some embodiments, the guide polynucleotide comprises a sequence that is complementary to a plant genomic polynucleotide sequence. In some embodiments, the guide polynucleotide comprises a sequence that is complementary to a mammalian genomic polynucleotide sequence. In some embodiments, the guide polynucleotide comprises a sequence that is complementary to a human genomic polynucleotide sequence.

[0134] In some embodiments, the guide polynucleotide has a length of 30 to 250 nucleotides. In some embodiments, the guide polynucleotide has a length of more than 90 nucleotides. In some embodiments, the guide polynucleotide has a length of less than 245 nucleotides. In some embodiments, the guide polynucleotide has a length of 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, or more than 240 nucleotides. In some embodiments, the guide polynucleotide has a length of about 30 to about 40, about 30 to about 50, about 30 to about 60, about 30 to about 70, about 30 to about 80, about 30 to about 90, about 30 to about 100, about 30 to about 120, about 30 to about 140, about 30 to about 160, about 30 to about 180, about 30 to about 200, about 30 to about 220, about 30 to about 240, about 50 to about 60, about 50 to about 70, about 50 to about 80, about 50 to about 90, about 50 to about 100, about 50 to about 120, about 50 to about 140, about 50 to about 160, about 50 to about 180, about 50 to about 200, about 50 to about 220, about 50 to about 240, about 100 to about 120, about 100 to about 140, about 100 to about 160, about 100 to about 180, about 100 to about 200, about 100 to about 220, about 100 to about 240, about 160 to about 180, about 160 to about 200, about 160 to about 220, or about 160 to about 240 nucleotides.

[0135] MG gene editing system In certain embodiments, an engineered nuclease system is disclosed herein, the engineered nuclease system comprising: a) an endonuclease; b) an engineered guide polynucleotide that forms a complex with the endonuclease and is configured to hybridize to a target nucleic acid sequence within or in an intron of the albumin gene; and c) a donor template comprising a nucleic acid sequence encoding a Factor VIII (FVIII) gene or a functional fragment thereof.

[0136] In some embodiments, the endonuclease induces a single-strand break at or proximal to the target nucleic acid sequence. In some embodiments, the endonuclease induces a double-strand break at or proximal to the target nucleic acid sequence. In some embodiments, the donor template is integrated into the target nucleic acid sequence at the double-strand break. In some embodiments, the donor template is integrated into the target nucleic acid sequence at the double-strand break via non-homologous end joining (NHEJ). In some embodiments, the donor template is integrated into the target nucleic acid sequence at the double-strand break via homology-directed repair (HDR).

[0137] In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 70% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 75% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 80% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 85% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 90% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 95% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide; and c) a donor template comprising a sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 96% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 97% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 98% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 99% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having 100% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.

[0138] In some embodiments, the engineered guide polynucleotide is a single guide nucleic acid. In some embodiments, the engineered guide polynucleotide is a dual guide nucleic acid. In some embodiments, the engineered guide polynucleotide is RNA. In some embodiments, the endonuclease is complexed with the engineered guide polynucleotide. In some embodiments, the endonuclease is bound to the engineered guide polynucleotide.

[0139] In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 70% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within the albumin gene or an intron of the albumin gene, the engineered guide polynucleotide comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 75% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within the albumin gene or an intron of the albumin gene, the engineered guide polynucleotide comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 80% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or in an intron of the albumin gene, the engineered guide polynucleotide comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 85% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or in an intron of the albumin gene, the engineered guide polynucleotide comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 90% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within the albumin gene or an intron of the albumin gene, the engineered guide polynucleotide comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 95% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within the albumin gene or an intron of the albumin gene, the engineered guide polynucleotide comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 96% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or in an intron of the albumin gene, the engineered guide polynucleotide comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 97% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or in an intron of the albumin gene, the engineered guide polynucleotide comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 98% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within the albumin gene or an intron of the albumin gene, the engineered guide polynucleotide comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising a sequence having at least about 99% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within the albumin gene or an intron of the albumin gene, the engineered guide polynucleotide comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.In some embodiments, the engineered nuclease system comprises: a) an endonuclease comprising 100% identity to SEQ ID NO: 54 or SEQ ID NO: 96; b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to at least a portion of a target nucleic acid sequence within or in an intron of the albumin gene, the engineered guide polynucleotide comprising 100% identity to any one of SEQ ID NO: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to any one of SEQ ID NO: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NO: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98, and c) the donor template comprises a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.

[0140] In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 70% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 75% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 80% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 85% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 90% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 95% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 96% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 97% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 98% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 99% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease having 100% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.

[0141] In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 70% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or intronic to the albumin gene, the engineered guide polynucleotide being encoded by a nucleic acid sequence having at least about 70% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 75% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or intronic to the albumin gene, the engineered guide polynucleotide being encoded by a nucleic acid sequence having at least about 75% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 80% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or in an intron of the albumin gene, the engineered guide polynucleotide being encoded by a nucleic acid sequence having at least about 80% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 85% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or in an intron of the albumin gene, the engineered guide polynucleotide being encoded by a nucleic acid sequence having at least about 85% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 90% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or in an intron of the albumin gene, the engineered guide polynucleotide being encoded by a nucleic acid sequence having at least about 90% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 95% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or in an intron of the albumin gene, the engineered guide polynucleotide being encoded by a nucleic acid sequence having at least about 95% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 96% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within the albumin gene or an intron of the albumin gene, the engineered guide polynucleotide being encoded by a nucleic acid sequence having at least about 96% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 97% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within the albumin gene or an intron of the albumin gene, the engineered guide polynucleotide being encoded by a nucleic acid sequence having at least about 97% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 98% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or in an intron of the albumin gene, the engineered guide polynucleotide being encoded by a nucleic acid sequence having at least about 98% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the engineered nuclease system comprises: a) an endonuclease encoded by a nucleic acid sequence having at least about 99% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or in an intron of the albumin gene, the engineered guide polynucleotide being encoded by a nucleic acid sequence having at least about 99% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.In some embodiments, the engineered nuclease system comprises: a) an endonuclease having 100% identity to any one of SEQ ID NOs: 30, 31, 53, and 95; b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to at least a portion of a target nucleic acid sequence within or within an intron of the albumin gene, the engineered guide polynucleotide having 100% identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98; and c) a donor template comprising a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 14, 15, 24-27, 43-45, 50, 55, 60-68, 97, and 98, and c) the donor template comprises a nucleic acid sequence encoding the FVIII gene or a functional fragment thereof.

[0142] In some embodiments, the donor template comprises a polyadenylation signal. In some embodiments, the polyadenylation signal is at the C-terminus of the Factor VIII gene or a fragment thereof. In some embodiments, the polyadenylation signal is linked to the Factor VIII gene or a fragment thereof. In some embodiments, the polyadenylation signal is fused to the Factor VIII gene or a fragment thereof.

[0143] In some embodiments, the donor template includes a nuclear targeting sequence (NTS). In some embodiments, the nuclear targeting sequence includes a plurality of transcription factor binding sites. In some embodiments, the transcription factor (TF) binding site is the SV40 enhancer region. In some embodiments, the transcription factor is TCF1, HNF1, NFY, CEBP, OCT1, AP1, HNF1-α, HNF1-β, CEBPA, LEF-1, FOX D1, IRF1, HNF3, HNF4, HNF5, Tal1β / E47, or MyoD. Table 1B includes examples of TFs found in the promoters and enhancers of genes highly expressed in the liver. Common liver-specific transcription factors include HNF3, HNF4, HNF5, C / EBP, HNF1α, LEF1, FOX, IRF, and TCF.

[0144]

Table 1B

[0145] In some embodiments, a functional fragment of the albumin promoter is used in the donor template. In some embodiments, the functional fragment is located 5' of the transcription start site and includes a sequence of about 100 bp containing binding sites for several TFs including HNF1, CEBP, LEF-1, FOX, IRF1, and LEF1.

[0146] In some embodiments, the transcription factor binding sequence comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 37, 39, and 41, or any of the sequences listed in Table 1B. In some embodiments, the transcription factor binding sequence comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 37, 39, and 41, or any of the sequences listed in Table 1B. In some embodiments, the transcription factor binding sequence comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 37, 39, and 41, or any of the sequences listed in Table 1B. In some embodiments, the transcription factor binding sequence comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 37, 39, and 41, or any of the sequences listed in Table 1B. In some embodiments, the transcription factor binding sequence comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 37, 39, and 41, or any of the sequences listed in Table 1B. In some embodiments, the transcription factor binding sequence comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 37, 39, and 41, or any of the sequences listed in Table 1B. In some embodiments, the transcription factor binding sequence comprises a sequence having 100% identity to any one of SEQ ID NOs: 37, 39, and 41, or any of the sequences listed in Table 1B.

[0147] In some embodiments, the nuclear targeting sequence is at the 5' end of the donor template. In some embodiments, the nuclear targeting sequence is at the 3' end of the donor template. In some embodiments, the nuclear targeting sequence is at both the 5' and 3' ends of the donor template.

[0148] In some embodiments, the donor template further comprises a recognition site sequence for an endonuclease at the 5' or 3' end of the donor template. In some embodiments, the nuclear targeting sequence is on the 5' side of the recognition site sequence when the donor template is flanked at the 5' end by the nuclear targeting sequence. In some embodiments, the nuclear targeting sequence is on the 3' side of the recognition site sequence when the donor template is flanked at the 3' end by the nuclear targeting sequence.

[0149] In some embodiments, the donor template comprises a splice acceptor sequence. In some embodiments, the splice acceptor sequence targets the albumin gene. In some embodiments, the splice acceptor sequence targets exon 1 of the albumin gene. In some embodiments, the splice acceptor sequence targets intron 1 of the albumin gene. In some embodiments, the splice acceptor sequence is linked to the factor VIII gene or a fragment thereof. In some embodiments, the splice acceptor sequence is an intron sequence linked to the exon sequence of the factor VIII gene or a fragment thereof. In some embodiments, the splice acceptor sequence is linked to the factor VIII gene or a fragment thereof at the 5' end of the factor VIII gene or fragment. In some embodiments, the splice acceptor sequence is linked to the factor VIII gene or a fragment thereof at the 3' end of the factor VIII gene or fragment. In some embodiments, the splice acceptor sequence is linked to the factor VIII gene or a fragment thereof using a linker. In some embodiments, the linker comprises a sequence having at least 80% sequence identity to SEQ ID NO: 80. In some embodiments, the linker comprises a sequence having at least 90% sequence identity to SEQ ID NO: 80. In some embodiments, the linker comprises a sequence having at least 95% sequence identity to SEQ ID NO: 80. In some embodiments, the linker comprises a sequence having at least 96% sequence identity to SEQ ID NO: 80. In some embodiments, the linker comprises a sequence having at least 97% sequence identity to SEQ ID NO: 80. In some embodiments, the linker comprises a sequence having at least 98% sequence identity to SEQ ID NO: 80. In some embodiments, the linker comprises a sequence having at least 99% sequence identity to SEQ ID NO: 80. In some embodiments, the linker comprises a sequence having 100% sequence identity to SEQ ID NO: 80. In some embodiments, the splice acceptor sequence is fused to the factor VIII gene or a fragment thereof.

[0150] In some embodiments, the donor template comprises NTS(1)-NRS(1)-SA-FVIII-NRS(2)-NTS(2) in the 5' to 3' direction, where NTS(1) represents a first nuclear targeting sequence, NTS(2) represents a second nuclear targeting sequence, NRS(1) represents a first nuclease recognition site sequence, NRS(2) represents a second nuclease recognition site sequence, SA represents a splice acceptor sequence that targets exon 1 of the albumin gene, and FVIII represents the Factor VIII gene or a fragment thereof. In some embodiments, the 5' to 3' orientation of NRS(1) and NRS(2) follows (a) forward, forward, (b) reverse, reverse, (c) forward, reverse, (d) reverse, forward, where forward indicates the same 5' to 3' orientation as the target nucleic acid sequence and reverse indicates the opposite 5' to 3' orientation of the target nucleic acid sequence. In some embodiments, the donor template comprises a KTTN (K = G or T, N = any base) or AAANNN (N = any base) sequence.

[0151] In some embodiments, the donor template comprises a sequence having a sequence identity of at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 12-13, 16-23, 32-33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 12-13, 16-23, 32-33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 12-13, 16-23, 32-33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 12-13, 16-23, 32-33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 12-13, 16-23, 32-33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 12-13, 16-23, 32-33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 12-13, 16-23, 32-33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 12-13, 16-23, 32-33, 56-59, 81-88, and 90-94.In some embodiments, the donor template comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 12-13, 16-23, 32-33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 12-13, 16-23, 32-33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 12-13, 16-23, 32-33, 56-59, 81-88, and 90-94. In some embodiments, the donor template comprises a sequence having 100% identity to any one of SEQ ID NOs: 12-13, 16-23, 32-33, 56-59, 81-88, and 90-94.

[0152] In some embodiments, the FVIII gene or a functional fragment thereof is codon-optimized to remove at least one cytosine-guanine (CG or CpG) motif.

[0153] In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having a sequence identity of at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 10, 71-79, and 89.In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 10, 71-79, and 89. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having 100% identity to any one of SEQ ID NOs: 10, 71-79, and 89.

[0154] In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having a sequence identity of at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 70% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 75% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 80% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 85% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 90% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 95% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 96% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 97% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 98% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having at least about 99% identity to SEQ ID NO: 10. In some embodiments, the FVIII gene or a functional fragment thereof comprises a sequence having 100% identity to SEQ ID NO: 10.

[0155] In some embodiments, the FVIII gene or a functional fragment thereof is modified to include a B domain that includes a sequence having at least about 90% identity to any one of SEQ ID NOs: 71-79 and 89. In some embodiments, the FVIII gene or a functional fragment thereof is modified to include a B domain that includes a sequence having at least about 95% identity to any one of SEQ ID NOs: 71-79 and 89. In some embodiments, the FVIII gene or a functional fragment thereof is modified to include a B domain that includes a sequence having at least about 96% identity to any one of SEQ ID NOs: 71-79 and 89. In some embodiments, the FVIII gene or a functional fragment thereof is modified to include a B domain that includes a sequence having at least about 97% identity to any one of SEQ ID NOs: 71-79 and 89. In some embodiments, the FVIII gene or a functional fragment thereof is modified to include a B domain that includes a sequence having at least about 98% identity to any one of SEQ ID NOs: 71-79 and 89. In some embodiments, the FVIII gene or a functional fragment thereof is modified to include a B domain that includes a sequence having at least about 99% identity to any one of SEQ ID NOs: 71-79 and 89. In some embodiments, the FVIII gene or a functional fragment thereof is modified to include a B domain that includes a sequence having 100% identity to any one of SEQ ID NOs: 71-79 and 89.

[0156] In some embodiments, the FVIII gene comprising the modified B domain or a functional fragment thereof comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 86, 87, and 90. In some embodiments, the FVIII gene comprising the modified B domain or a functional fragment thereof comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 86, 87, and 90. In some embodiments, the FVIII gene comprising the modified B domain or a functional fragment thereof comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 86, 87, and 90. In some embodiments, the FVIII gene comprising the modified B domain or a functional fragment thereof comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 86, 87, and 90. In some embodiments, the FVIII gene comprising the modified B domain or a functional fragment thereof comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 86, 87, and 90. In some embodiments, the FVIII gene comprising the modified B domain or a functional fragment thereof comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 86, 87, and 90. In some embodiments, the FVIII gene comprising the modified B domain or a functional fragment thereof comprises a sequence having 100% identity to any one of SEQ ID NOs: 86, 87, and 90.

[0157] Method of Use Methods for replenishing hepatic enzymes using engineered nuclease systems described herein are provided herein. The method for replenishing a hepatic enzyme involves integrating a liver gene into the genome at an appropriate site and in an appropriate cell or tissue such that the gene is expressed and produces a functional protein (e.g., functional FVIII) and thus replenishes the deficiency.

[0158] In some embodiments, the engineered nuclease system described herein is used to integrate the Factor VIII gene into the genome of an individual in need thereof. In some embodiments, the engineered nuclease system described herein is used to integrate the Factor VIII gene into the genome of an individual in need thereof, thereby treating hemophilia A. In some embodiments, the engineered nuclease system described herein is used to effectuate this in an individual in need of treating hemophilia A.

[0159] In some embodiments, the FVIII gene or a functional fragment thereof is delivered to hepatocytes of the liver by systemic (e.g., intravenous) administration of a vector comprising an endonuclease and a guide polynucleotide that targets a target nucleic acid at a safe harbor locus such as the albumin locus.

[0160] In some embodiments, the site of integration into the FVIII gene is within the albumin gene. In some embodiments, the site of integration into the FVIII gene is intron 1 of the albumin gene. Since the albumin gene is expressed at high levels in hepatocytes, the albumin gene promoter is expected to drive efficient expression of the integrated FVIII gene. This method of integration into an intron has the advantage that double-strand breaks that are subsequently repaired by error-prone NHEJ do not adversely affect the function of the albumin gene. In some embodiments, a splice acceptor site is included at the 5' end of the donor template immediately prior to the N-terminus of the FVIII protein coding sequence to capture transcription initiated from the albumin promoter. This splice acceptor captures a portion of the splicing event from albumin exon 1 and can result in an mRNA that includes the 5'UTR and exon 1 of albumin fused in-frame to the coding sequence of FVIII.

[0161] Delivery and Vectors In some embodiments, nucleic acid sequences encoding the engineered nuclease systems or components thereof described herein (e.g., endonucleases, engineered guide polynucleotides, or donor templates) are disclosed herein.

[0162] In some embodiments, the nucleic acid encoding the engineered nuclease system described herein is DNA, e.g., linear DNA, plasmid DNA, or minicircle DNA. In some embodiments, the nucleic acid encoding the engineered nuclease system or components thereof described herein is RNA, e.g., mRNA.

[0163] In some embodiments, the nucleic acid encoding the engineered nuclease system or components thereof described herein is delivered by a nucleic acid-based vector. In some embodiments, the nucleic acid-based vector is a plasmid (e.g., a circular DNA molecule that can replicate autonomously inside a cell), a cosmid (e.g., pWE or sCos vector), an artificial chromosome, a human artificial chromosome (HAC), a yeast artificial chromosome (YAC), a bacterial artificial chromosome (BAC), a P1-derived artificial chromosome (PAC), a phagemid, a phage derivative, a bacmid, or a virus. In some embodiments, the nucleic acid-based vector is selected from the list consisting of pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His, pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 vector, pEF1a-tdTomato vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), pSF-CMV-PURO-NH2-CMYC, pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, pSF-Tac, pRI 101-AN DNA, pCambia2301, pTYB21, pKLAC2, pAc5.1 / V5-His A, and pDEST8.

[0164] In some embodiments, the nucleic acid-based vector comprises a promoter. In some embodiments, the promoter is selected from the group consisting of a mini-promoter, an inducible promoter, a constitutive promoter, and derivatives thereof. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, synapsin, CaMKII, GRK1, and derivatives thereof. In some embodiments, the promoter is the U6 promoter. In some embodiments, the promoter is the CAG promoter.

[0165] In some embodiments, the nucleic acid-based vector is a virus. In some embodiments, the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpes virus, poxvirus, anellovirus, bocavirus, vaccinia virus, or retrovirus. In some embodiments, the virus is an alphavirus. In some embodiments, the virus is a parvovirus. In some embodiments, the virus is an adenovirus. In some embodiments, the virus is AAV. In some embodiments, the virus is a baculovirus. In some embodiments, the virus is a dengue virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is a herpes virus. In some embodiments, the virus is a poxvirus. In some embodiments, the virus is an anellovirus. In some embodiments, the virus is a bocavirus. In some embodiments, the virus is a vaccinia virus. In some embodiments, the virus is or a retrovirus.

[0166] In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh10, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, AAV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, AAV-HSC16, or a derivative thereof. In some embodiments, the herpes virus is HSV type 1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.

[0167] In some embodiments, the virus is AAV1 or a derivative thereof. In some embodiments, the virus is AAV2 or a derivative thereof. In some embodiments, the virus is AAV3 or a derivative thereof. In some embodiments, the virus is AAV4 or a derivative thereof. In some embodiments, the virus is AAV5 or a derivative thereof. In some embodiments, the virus is AAV6 or a derivative thereof. In some embodiments, the virus is AAV7 or a derivative thereof. In some embodiments, the virus is AAV8 or a derivative thereof. In some embodiments, the virus is AAV9 or a derivative thereof. In some embodiments, the virus is AAV10 or a derivative thereof. In some embodiments, the virus is AAV11 or a derivative thereof. In some embodiments, the virus is AAV12 or a derivative thereof. In some embodiments, the virus is AAV13 or a derivative thereof. In some embodiments, the virus is AAV14 or a derivative thereof. In some embodiments, the virus is AAV15 or a derivative thereof. In some embodiments, the virus is AAV16 or a derivative thereof. In some embodiments, the virus is AAV-rh8 or a derivative thereof. In some embodiments, the virus is AAV-rh10 or a derivative thereof. In some embodiments, the virus is AAV-rh20 or a derivative thereof. In some embodiments, the virus is AAV-rh39 or a derivative thereof. In some embodiments, the virus is AAV-rh74 or a derivative thereof. In some embodiments, the virus is AAV-rhM4-1 or a derivative thereof. In some embodiments, the virus is AAV-hu37 or a derivative thereof. In some embodiments, the virus is AAV-Anc80 or a derivative thereof. In some embodiments, the virus is AAV-Anc80L65 or a derivative thereof. In some embodiments, the virus is AAV-7m8 or a derivative thereof. In some embodiments, the virus is AAV-PHP-B or a derivative thereof. In some embodiments, the virus is AAV-PHP-EB or a derivative thereof.In some embodiments, the virus is AAV-2.5 or a derivative thereof. In some embodiments, the virus is AAV-2tYF or a derivative thereof. In some embodiments, the virus is AAV-3B or a derivative thereof. In some embodiments, the virus is AAV-LK03 or a derivative thereof. In some embodiments, the virus is AAV-HSC1 or a derivative thereof. In some embodiments, the virus is AAV-HSC2 or a derivative thereof. In some embodiments, the virus is AAV-HSC3 or a derivative thereof. In some embodiments, the virus is AAV-HSC4 or a derivative thereof. In some embodiments, the virus is AAV-HSC5 or a derivative thereof. In some embodiments, the virus is AAV-HSC6 or a derivative thereof. In some embodiments, the virus is AAV-HSC7 or a derivative thereof. In some embodiments, the virus is AAV-HSC8 or a derivative thereof. In some embodiments, the virus is AAV-HSC9 or a derivative thereof. In some embodiments, the virus is AAV-HSC10 or a derivative thereof. In some embodiments, the virus is AAV-HSC11 or a derivative thereof. In some embodiments, the virus is AAV-HSC12 or a derivative thereof. In some embodiments, the virus is AAV-HSC13 or a derivative thereof. In some embodiments, the virus is AAV-HSC14 or a derivative thereof. In some embodiments, the virus is AAV-HSC15 or a derivative thereof. In some embodiments, the virus is AAV-TT or a derivative thereof. In some embodiments, the virus is AAV-DJ / 8 or a derivative thereof. In some embodiments, the virus is AAV-Myo or a derivative thereof. In some embodiments, the virus is AAV-NP40 or a derivative thereof. In some embodiments, the virus is AAV-NP59 or a derivative thereof. In some embodiments, the virus is AAV-NP22 or a derivative thereof. In some embodiments, the virus is AAV-NP66 or a derivative thereof. In some embodiments, the virus is AAV-HSC16 or a derivative thereof.

[0168] In some embodiments, the virus is HSV-1 or a derivative thereof. In some embodiments, the virus is HSV-2 or a derivative thereof. In some embodiments, the virus is VZV or a derivative thereof. In some embodiments, the virus is EBV or a derivative thereof. In some embodiments, the virus is CMV or a derivative thereof. In some embodiments, the virus is HHV-6 or a derivative thereof. In some embodiments, the virus is HHV-7 or a derivative thereof. In some embodiments, the virus is HHV-8 or a derivative thereof.

[0169] In some embodiments, the nucleic acid encoding the engineered nuclease system or a component thereof described herein is delivered by a non-nucleic acid-based delivery system (e.g., a non-viral delivery system). In some embodiments, the non-viral delivery system is a liposome. In some embodiments, the nucleic acid is lipid-associated. Lipid-associated nucleic acids are, in some embodiments, encapsulated within the aqueous interior of a liposome, dispersed within the lipid bilayer of a liposome, attached to a liposome via a binding molecule associated with both the liposome and the nucleic acid, entrapped within a liposome, complexed with a liposome, dispersed in a lipid-containing solution, mixed with a lipid, combined with a lipid, included as a suspension in a lipid, included as a micelle, or complexed with a micelle, or otherwise lipid-associated. In some embodiments, the nucleic acid is included in lipid nanoparticles (LNPs).

[0170] In some embodiments, the engineered nuclease system or a component thereof described herein is introduced into cells in any suitable manner, either stably or transiently. In some embodiments, the engineered nuclease system or a component thereof described herein is transfected into cells. In some embodiments, cells are transduced or transfected with a nucleic acid construct encoding the engineered endonuclease or a component thereof. For example, cells are transduced (e.g., using a virus encoding the engineered nuclease system or a component thereof described herein) or transfected (e.g., using a plasmid encoding the engineered nuclease system or a component thereof described herein) with the engineered nuclease system or a component thereof described herein, or with a nucleic acid encoding the translated engineered nuclease system or a component thereof described herein. In some embodiments, the transduction is stable or transient transduction. In some embodiments, cells expressing the engineered nuclease system or a component thereof described herein, or cells containing the engineered nuclease system or a component thereof described herein, are transduced or transfected with, for example, one or more gRNA molecules when the engineered nuclease system or a component thereof described herein comprises a CRISPR nuclease. In some embodiments, plasmids expressing the engineered nuclease system or a component thereof described herein are introduced into cells through electroporation, transient (e.g., lipofection) and stable genomic integration (e.g., piggybac), as well as viral transduction (e.g., lentivirus or AAV), or other methods known to those skilled in the art. In some embodiments, the gene editing system is introduced into cells as one or more polypeptides. In some embodiments, delivery is achieved through the use of RNP complexes. Methods for delivering polypeptides and / or RNPs to cells are known in the art, for example, by electroporation or by cell squeezing.

[0171] Exemplary methods of nucleic acid delivery include lipofection, nucleofection, electroporation, stable genomic integration (e.g., piggybac), microinjection, biolistic, virosome, liposome, immunoliposome, polycation or lipid nucleic acid conjugate, naked DNA, artificial virion, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Patent Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™, Lipofectin™, and SF Cell Line 4D-Nucleofector X Kit™ (Lonza)). Cationic and neutral lipids suitable for efficient receptor recognition lipofection of polynucleotides include the lipids of WO91 / 17424 and WO91 / 16024. In some embodiments, the delivery is to a cell (e.g., in vitro or ex vivo administration) or a target tissue (e.g., in vivo administration). In some embodiments, the nucleic acid is contained in liposomes or nanoparticles that specifically target host cells.

[0172] Additional methods for delivery of nucleic acids to cells are known to those of skill in the art. See, for example, US2003 / 0087817.

[0173] In some embodiments, the disclosure provides a cell comprising a vector or nucleic acid described herein. In some embodiments, the cell expresses a gene editing system or a part thereof. In some embodiments, the cell is a human cell. In some embodiments, the cell is genome edited ex vivo. In some embodiments, the cell is genome edited in vivo.

[0174] Lipid nanoparticles In certain embodiments, lipid nanoparticles comprising the engineered nuclease system of the disclosure for delivery of the engineered nuclease system to cells are disclosed herein.

[0175] In some embodiments, the lipid nanoparticles comprise an engineered nuclease system or a nucleic acid encoding an engineered nuclease system. In some embodiments, the lipid nanoparticles comprise one or more components of an engineered nuclease system. In some embodiments, the lipid nanoparticles comprise an endonuclease or a nucleic acid encoding an endonuclease. In some embodiments, the lipid nanoparticles comprise an engineered guide polynucleotide. In some embodiments, the lipid nanoparticles comprise a donor template.

[0176] In some embodiments, the lipid nanoparticles are linked to an engineered nuclease system.

[0177] The lipid nanoparticles described herein can be four-component lipid nanoparticles. Such nanoparticles can be configured for delivery of RNA or other nucleic acids (e.g., synthetic RNA, mRNA, or mRNA synthesized in vitro) and can generally be formulated as described in WO2012 / 135805 (A2). Such nanoparticles can generally comprise (a) a cationic lipid, (b) a neutral lipid (e.g., DSPC or DOPE), (c) a sterol (e.g., cholesterol or a cholesterol analog), or (d) a PEGylated lipid (e.g., PEG-DMG).

[0178] The cationic lipid referred to as "C12-200" in this specification is disclosed by Love et al., Proc Natl Acad Sci USA. 2010 107:1864-1869 and Liu and Huang, Molecular Therapy. 2010 669-670. Cationic lipid formulations can include particles containing any of three or four or more components in addition to polynucleotides, primary constructs, or RNA (e.g., mRNA). As an example, a formulation containing a particular cationic lipid includes, but is not limited to, 98N12-5, and can contain 42% lipidoid, 48% cholesterol, and 10% PEG (alkyl chain length C14 or greater). As another example, a formulation having a particular lipidoid includes, but is not limited to, C12-200, and can contain 50% cationic lipid, 10% distearoyl phosphatidylcholine, 38.5% cholesterol, and 1.5% PEG-DMG.

[0179] In some embodiments, the cationic lipid nanoparticles comprise a cationic lipid, a PEGylated lipid, a sterol, and a non-cationic lipid. In some embodiments, the cationic lipid nanoparticles have a molar ratio of about 20-60% cationic lipid: about 5-25% non-cationic lipid: about 25-55% sterol, and about 0.5-15% PEGylated lipid. In some embodiments, the cationic lipid nanoparticles comprise a molar ratio of about 50% cationic lipid, about 1.5% PEGylated lipid, about 38.5% cholesterol, and about 10% non-cationic lipid. In some embodiments, the cationic lipid nanoparticles comprise a molar ratio of about 55% cationic lipid, about 2.5% PEGylated lipid, about 32.5% cholesterol, and about 10% non-cationic lipid. In some embodiments, the cationic lipid is an ionic cationic lipid, the non-cationic lipid is a neutral lipid, and the sterol is cholesterol. In some embodiments, the cationic lipid nanoparticles have a molar ratio of 50:38.5:10:1.5 of cationic lipid:cholesterol:PEG2000-DMG:DSPC or DMG:DOPE. In some embodiments, the lipid nanoparticles described herein can comprise cholesterol, 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,1‘-((2-(4-(2-((2-(bis(2-hydroxydodecyl)amino)ethyl)(2-hydroxydodecyl)amino)ethyl)piperazin-1-yl)ethyl)azanediyl)bis(dodecan-2-ol) (C12-200), and DMG-PEG-2000 in a molar ratio of 47.5:16:35:1.5.

[0180] cell In certain embodiments, cells comprising the engineered nuclease systems described herein are described herein.

[0181] In some embodiments, the cell is a eukaryotic cell (e.g., a plant cell, an animal cell, a protist cell, or a fungal cell), a mammalian cell (Chinese hamster ovary (CHO) cell, baby hamster kidney (BHK), human embryonic kidney (HEK), mouse myeloma (NS0), or human retinal cells), an immortalized cell (e.g., HeLa cells, COS cells, HEK-293T cells, MDCK cells, 3T3 cells, PC12 cells, Huh7 cells, HepG2 cells, K562 cells, N2a cells, or SY5Y cells), an insect cell (e.g., Spodoptera frugiperda cells, Trichoplusia ni cells, Drosophila melanogaster cells, S2 cells, or Heliothis virescens cells), a yeast cell (e.g., Saccharomyces cerevisiae cells, Cryptococcus cells, or Candida cells), a plant cell (e.g., parenchyma cells, collenchyma cells, or sclerenchyma cells), a fungal cell (e.g., Saccharomyces cerevisiae cells, Cryptococcus cells, or Candida cells), or a prokaryotic cell (e.g., E. coli cells, streptococcus bacterial cells, streptomyces soil bacterial cells, or archaebacterial cells). In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell.

[0182] In some embodiments, the cell is A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC1, BSC40, BMT10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof.

[0183] In some embodiments, the cell is a liver cell.

[0184] Kit In some embodiments, the disclosure provides a kit comprising one or more nucleic acid constructs encoding various components of the engineered nuclease system described herein. In some embodiments, the nucleotide sequence comprises a heterologous promoter that drives the expression of the engineered nuclease system component.

[0185] In some embodiments, the engineered nuclease system or its components disclosed herein are incorporated into a pharmaceutical, diagnostic, or research kit to facilitate their use in therapeutic, diagnostic, or research applications. The kit may include one or more containers containing any of the vectors disclosed herein, and instructions for use.

[0186] The kit can be designed to facilitate the use of the methods described herein by researchers and can take many forms. Each of the compositions of the kit can be provided in liquid form (e.g., in solution) or in solid form (e.g., as a dry powder) where applicable. In certain cases, some of the compositions can be constituted or otherwise treatable (e.g., into an active form) by the addition of a suitable solvent or other species (e.g., water or cell culture medium), for example, when provided with the kit and when not. As used herein, "instructions" defines the components of the instructions and / or promotional materials and can typically be accompanied by written instructions on or associated with the packaging of the disclosure. The instructions can also include any oral or electronic instructions provided in any manner such that the user clearly recognizes that the instructions are related to the kit, for example, visual (e.g., videotape, DVD, etc.), Internet, and / or web-based communications. The written instructions are, in some embodiments, in a form prescribed by a government agency that regulates the manufacture, use, or sale of pharmaceuticals or biological products, and the instructions can also reflect approval by an agency for manufacture, use, or sale for animal administration.

Examples

[0187] Example 1 - Design of a DNA donor template optimized for non - viral delivery for the purpose of integrating FVIII into the genome and producing functional FVIII protein An 88bp sequence from the human albumin promoter containing several transcription factor (TF) binding sites was selected as part of the donor template. This 88bp sequence extends from 111bp 5’ of the transcription start site to just 8bp 3’ of the transcription start site and contains binding sites for the transcription factors LEF / TCF1, HNF1, NFY, and CEBP. This sequence was designated as the human albumin nuclear targeting sequence (hANTS) and contains the following sequence: 5′-TGAATTTTGTAATCGGTTGGCAGCCAATGAAATACAAAGATGAGTCTAGTTAATAATCTACAATTATTGGTTAAAGAAGTATATTAGT-3′ (SEQ ID NO: 1). Additionally, a 72bp SV40 enhancer: ATGCTTTGCATACTTCTGCCTGCTGGGGAGCCTGGGGACTTTCCACACCCTAACTGACACACATTCCAC (SEQ ID NO: 2) was optionally included in the template.

[0188] The donor DNA template used for non - viral delivery was designed using the following general structure: 5’ closed end (CE) - nuclear - targeting DNA sequence (NTDS) - spacer - target site for CRISPR nuclease or other sequence - specific nuclease - therapeutic gene (TG) - polyA signal spacer target site for CRISPR nuclease or other sequence - specific nuclease - spacer - nuclear - targeting DNA sequence (NTDS) - closed end (CE). The target site for the CRISPR nuclease or other sequence - specific nuclease in the donor DNA is the same as the target site in the genomic locus selected as the integration site within the genome. The closed end (CE) indicates that the DNA sequence is synthesized such that the 5’ and 3’ ends of the double - strand are covalently joined, which increases stability against nuclease degradation.

[0189] Specific sequences of the NTDS can potentially affect the integration efficiency of the donor template, so donor DNAs with different NTDSs were designed. In one embodiment, the NTDS included a single copy of hANTS (SEQ ID NO: 1) present at one or both ends of the DNA donor. In another embodiment, the NTDS included a single copy of SV40e (SEQ ID NO: 2) present at one or both ends of the DNA donor. In yet another embodiment, the NTDS included a single copy of hANTS (SEQ ID NO: 1) and a single copy of SV40e (SEQ ID NO: 2) present at one or both ends of the DNA donor.

[0190] In the case of the Factor VIII gene where the target locus in the genome is the albumin locus (particularly intron 1 of the albumin gene), the donor DNA template was designed with the following components in the following order: 5' closed end (CE) - nuclear targeting sequence (NTS) - 20 bp spacer - target sites for guides 8 and 12 for nuclease MG29-1 of mouse albumin intron 1 - spacer (37 bp) - splice acceptor - FVIII CDS - polyA - spacer (36 nt) - target sites for guides 8 and 12 for CRISPR nuclease MG29-1 of mouse albumin intron 1 - 20 bp spacer - nuclear targeting sequence (NTS) - closed end (CE) - 3'. Twelve possible combinations of guide polynucleotide target site orientations and NTDS sequences were designed, which included each of the four guide orientations combined with either hANTS alone, SV40e alone, or a combination of nANTS and SV40e at both ends of the donor. The individual sequence components are listed in SEQ ID NOs: 3 - 11.

[0191] Since the relative orientation of the cleavage sites within the donor with respect to the genomic site was considered a factor affecting integration efficiency, examples with all combinations of cleavage site orientations within the donor DNA were designed. Four possible orientations were envisioned: (1) forward-forward, (2) reverse-reverse, (3) forward-reverse, and (4) reverse-forward. Forward means that the target site is in the same orientation as in the genome, and reverse means that the target site is the reverse complement of the sequence in the genome. Two examples of the full sequence of the donor template for FVIII are the F-F orientation (pMG4010, SEQ ID NO: 12) of the guide cleavage site with an NTDS containing both hANTS and CMVe, and the R-R orientation (pMG4011, SEQ ID NO: 13) of the guide cleavage site with an NTDS containing both hANTS and CMVe. The two other possible orientations of the guide cleavage site are brought about by inverting the orientation of the cleavage site on the appropriate end to create the F-R (pMG4022, SEQ ID NO: 32) and R-F (pMG4023, SEQ ID NO: 33) variants where all other sequence elements remain unchanged. pMG4010 and pMG4011 are 4931 bp in length. In one embodiment, the FVIII coding sequence was codon-optimized to improve the expression of the FVIII gene after integration into the target locus within the genome. The innate immune system can recognize and eliminate DNA through the recognition of unmethylated CG dinucleotides (CpG motifs). Codon optimization typically involves the selection of the codons most frequently used for each amino acid. In the case of the FVIII coding sequence of SEQ ID NO: 10, all of the CpG motifs were removed by careful selection of alternative codons after codon optimization. Additionally, the spacer sequences were designed without CpG residues. There are a total of six CpG residues in each of pMG4010 and pMG4011.

[0192] Example 2 - In Vivo (Predictive) Testing of Non-Viral Delivery of a DNA Donor Template Containing an FVIII Gene Cassette Adjacent to NTDS and Guide RNA Target Sites To evaluate whether the DNA donor templates designed in Example 1 (e.g., SEQ ID NOs: 12, 13, 32, and 33) can function in vivo after non-viral delivery, the donor DNA templates are synthesized with closed ends. This DNA is encapsulated within lipid nanoparticles. The lipids are dissolved in ethanol. The donor DNA is prepared in water and then diluted in 100 mM sodium acetate (pH 4.0) to create a DNA working stock. Four lipid components are combined in ethanol at the desired ratio to create a lipid working stock. A typical lipid mixture contains cholesterol, 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,1‘-((2-(4-(2-((2-(bis(2-hydroxydecyl)amino)ethyl)(2-hydroxydecyl)amino)ethyl)piperazin-1-yl)ethyl)azanediyl)bis(dodecan-2-ol) (C12-200), and 1,2-dimyristoyl-rac-glycero-3-methoxypolyethylene glycol-2000 (DMG-PEG-2000) in a molar ratio of 47.5:16:35:1.5. The lipid working stock and the DNA working stock are combined in a microfluidic mixing device (Precision Nanosystems) at a flow rate of 12 mL / min and a ratio of 1 volume of lipid working stock to 3 volumes of DNA working stock. The mass ratio of C12-200 to DNA in the formulation is 5:1 to 20:1. The formulated LNP is diluted 1:1 with 1×PBS and then dialyzed twice in 1×PBS for 1 hour each, followed by concentration in an Amicon spin concentrator. The final LNP is formulated in 1×PBS buffer, filter sterilized through a 0.2 μM filter, and stored at 4°C. The concentrations of DNA inside and outside the LNP are measured. The average diameter and polydispersity of the LNP are measured by dynamic light scattering on the final concentrated LNP. The expected size range of the LNP is 80 - 100 nanometers with a PDI < 0.15 and a DNA encapsulation ratio exceeding 90%.

[0193] Separate LNPs are formulated with mRNA encoding MG29-1 nuclease and guide RNA8 or guide RNA12, both of which target sites within mouse albumin intron 1. The guide RNAs contained a specific set of chemical modifications optimized for stability and in vivo efficacy (SEQ ID NOs: 14 and 15). The guide RNAs and mRNAs are packaged separately. The lipids are dissolved in ethanol. The mRNA or guide RNA is prepared in water and then diluted in 100 mM sodium acetate (pH 4.0) to create an RNA working stock. Four lipid components are combined in ethanol at the desired ratio to create a lipid working stock. A typical lipid mixture contains cholesterol, DOPE, C12-200, and DMG-PEG-2000 in a molar ratio of 47.5:16:35:1.5. The lipid working stock and the RNA working stock are combined in a microfluidic mixing device at a flow rate of 12 mL / min and a ratio of 1 volume of lipid working stock to 3 volumes of RNA working stock. The mass ratio of C12-200 to RNA in the formulation is 10 to 1. The formulated LNP is diluted 1:1 with 1×PBS and then dialyzed twice in 1×PBS for 1 hour each, followed by concentration in a spin concentrator. The final LNP is formulated in 1×PBS buffer, filter sterilized through a 0.2 μM filter, and stored at 4°C. The concentration of the inner and outer RNAs of the LNP is measured. The average diameter and polydispersity of the LNP are measured by dynamic light scattering in the final concentrated LNP. Typically, the LNP is in the size range of 80-100 nanometers with a PDI < 0.15 and an RNA encapsulation ratio of greater than 90%. The LNP encapsulating guide RNA mAlb29-8-50 or mAlb29-12-50 is mixed with the LNP encapsulating MG29-1 mRNA at an RNA mass ratio of 1:1 (guide RNA:mRNA).

[0194] To initiate gene therapy, wild-type C57Bl6 mice are intravenously injected (0.1 mL via the tail vein) with LNP-encapsulated donor DNA pMG4010 or pMG4011 (or other variants of the donor described in Example 1) at a dose of 0.1 mg / kg to 2 mg / kg. The timing of administration of the LNP encapsulating MG29-1 mRNA and guide RNA relative to the DNA-donor LNP is evaluated. In one dosing regimen, the DNA donor-LNP and the MG29-1 mRNA / guide RNA LNP are pre-mixed and administered in a single injection. In another dosing regimen, the DNA donor-LNP is administered to the mice 1 to 48 hours prior to the administration of the MG29-1 mRNA / guide RNA LNP. In yet another dosing regimen, the MG29-1 mRNA / guide RNA LNP is administered to the mice 1 to 48 hours prior to the administration of the DNA donor-LNP. The dose of the DNA-donor LNP ranges from 0.1 mg / kg to 2 mg / kg of DNA in a total volume of 0.1 mL per mouse, and the dose of the MG29-1 mRNA / guide RNA LNP is initially set at 0.25 mg / kg for the purpose of achieving editing (cleavage) of the target site in the range of 20% to 40%. Plasma is collected from the mice on days 7 and 14 after administration, and the human FVIII levels are assayed using a capture CoA assay to detect the activity of human FVIII against the background of mouse FVIII. On day 14, the mice are sacrificed, the entire liver is flash-frozen and stored at -80°C. The entire left lateral lobe of the liver is homogenized in a bead mill using 0.4 mL of buffer per 100 mg of tissue weight. Genomic DNA is purified from an aliquot of the homogenate. The genomic DNA is analyzed for integration at the predicted target site in albumin intron 1 by in-out PCR using one primer complementary to the sequence in the next genome of the target site and one primer in the DNA donor template. The frequency of integration in the correct or reverse orientation is measured using either the quantitative real-time PCR or droplet digital PCR version of the in-out PCR assay.The percentage of cells expressing the integrated human FVIII mRNA was measured on liver sections by in situ hybridization using a fluorescent probe designed to detect hybrid mRNA transcripts at the junction between albumin mRNA and FVIII mRNA.

[0195] Example 3 - Design (predictive) of a DNA donor template optimized for viral delivery by AAV for the purpose of integrating a functional FVIII gene into the genome that produces a functional FVIII protein The donor DNA template cassette is designed using the following sequence elements in the 5' to 3' direction: target site for CRISPR nuclease - spacer - splice acceptor - therapeutic gene - polyA signal - spacer - target site for CRISPR nuclease. This donor template cassette is adjacent to the AAV inverted terminal repeat (ITR) and enables packaging into any AAV viral serotype of interest. For studies in mice, AAV serotype 8 or AAV serotype 6 is selected. In this example, the target site for CRISPR nuclease or other sequence - specific nuclease is selected from the target sites of nucleases MG29 - 1 or MG3 - 6 / 3 - 4. The specific guide RNA target sites for MG29 - 1 or MG3 - 6 / 3 - 4 are selected from the guide target sites of genomic target loci identified by screening for active guides that promote efficient DSB formation. In this specific example, the target locus in the genome is albumin and the specific region of albumin to be targeted is intron 1. It is envisioned that other genomic target sites can be selected to integrate the FVIII gene into different genomic loci. For both MG29 - 1 and MG3 - 6 / 3 - 4, screening for active guides targeting the mouse and human albumin intron 1 identified several highly active guide RNAs.

[0196] The orientation of the guide RNA target site in the donor template can be designed to be the same as the target site in albumin intron 1 or another target site (forward orientation, F), or the reverse complement of the target site in albumin intron 1 or another target site (reverse orientation, R). Thus, there are four possible combinations of guide target site orientations designated as FF, RR, FR, and RF. The FVIII donor cassette can be integrated into albumin intron 1 in either orientation, and only the "forward" orientation where the splice acceptor in the donor is positioned proximal to the albumin promoter is expected to result in the production of functional FVIII protein. The orientation of the guide RNA target site in the donor template can affect the efficiency with which the donor template cassette is integrated into albumin intron 1 in the forward orientation that can result in the production of functional FVIII protein. The MG29-1 and MG3-6 / 3-4 nucleases have not been tested in this context for the integration of DNA donor templates into hepatocytes in vivo. MG29-1 is a type V CRISPR nuclease that cleaves in an alternating pattern at the target site, which is in contrast to the blunt-end cleavage generated by Cas9 nuclease and MG3-6 / 3-4. The ultimate outcome of DSB repair generated by MG29-1 and MG3-6 / 3-4 in mammalian cells in the absence of donor DNA is mainly deletions, with very few alleles having inserted bases. In the case of MG29-1, deletions in the size range of 1 to 15 bases are most frequent, while MG3-6 / 3-4 tends to generate smaller deletions in the range of 1 to 6 bases on average. The profile of insertions and deletions (INDELS) arising from DSBs (INDEL profile) reflects the DNA repair process used by the cell to repair the DSB. Larger deletions generally indicate the alternative NHEJ repair pathway, while shorter deletions indicate the canonical NHEJ repair pathway. MG29-1 generates staggered cuts within the genome, leaving a 5' overhang that covers approximately 18 to 22 bases 3' from the PAM end of the target site, and the DNA donor template shown in Figure 1. The size of the single-stranded 5' overhang is predicted to be 3 to 6 nucleotides based on the in vitro assay shown in Figure 1.A single-stranded overhang can be efficiently joined by complementary single-stranded overhangs due to base-pairing interactions, in the same way that "sticky-end" ligation is more efficient than blunt-end ligation in conventional cloning methodologies. Thus, including guide target sites on both sides of the donor cassette can not only generate a linear double-stranded template in vivo within the nucleus of hepatocytes, but the resulting product can have single-stranded overhangs that are compatible with the overhangs at the double-strand break in the genome. In some cases, one of the four possible orientations of the guide RNA target site in the donor will result in an improvement in the frequency of integration in the forward orientation, which can, therefore, be determined empirically.

[0197] In the situation where both guide target sites in the donor DNA template are in the same orientation as the same target site in the genome and the alternating cuts generated by MG29-1 are joined to the donor via complementarity of the single-stranded 5' overhangs, it should be noted that the guide target sites can be completely reformed at both junctions and the DNA can be re-cut, thereby releasing the donor DNA (Tables 2 and 3).

[0198] [Table 2]

[0199] [Table 3]

[0200] When the donor DNA and the 5' overhang in target annealing promote integration, the impact on integration can be hypothesized as follows. For a donor with the FF orientation of the guide target site, integration in the forward orientation only (not the reverse orientation) is expected to occur via annealing of complementary ends, which is expected to result in the reformation of the guide target site, which is then re-cut, thereby excising the donor template. The alignment of the cut donor template in the reverse orientation to the cut genome does not result in complementary 5' overhangs at either junction.

[0201] For a donor with the RR orientation of the guide target site, integration in the reverse orientation can occur via annealing of complementary ends, while integration in the forward orientation is disadvantaged due to the lack of annealing of complementary ends. For a donor with the FR orientation of the guide target site, integration in both orientations can occur via annealing of complementary ends only at the 5' junction. For a donor with the RF orientation of the guide target site, integration in both orientations can occur via annealing of complementary ends only at the 3' junction. For each junction formed by the complete annealing of complementary overhangs, the guide target site can be reformed, which is then re-cut, thereby excising the donor template, which can impair integration efficiency. However, the NHEJ repair process is error-prone and introduces insertions and deletions at the site of DSB during the repair process, so these integration events may not occur without the DNA insertions or deletions simply described above. This is clearly seen in the cases of MG29-1 and MG3-6 / 3-4 by the observed INDEL profiles.

[0202] One possible mechanism for donor DNA integration is an iterative process of end joining and re-cleavage by a nuclease until an insertion / deletion (indel) occurs at the junction that prevents cleavage by the nuclease. When complementary 5’ overhangs at the ends of the donor template and the DSB in the genome facilitate annealing as an essential part of the NHEJ-driven repair process, the junction formed by annealing of the complementary ends can be fixed and stabilized when sufficient indels are introduced into the junction to block re-cleavage. Given the complexity of the DNA repair process, it may be necessary to empirically determine the optimal design of the donor template, particularly in the case of a novel nuclease.

[0203] For MG29-1, two guides called mAlb29-8-50b (SEQ ID NO: 24) and mAlb29-12-50b (SEQ ID NO: 25) were selected based on their editing activities in the mouse liver cell line Hepa1-6 and in vivo in the mouse liver. These guides contain a 20 nucleotide spacer and various chemical modifications, as well as an additional stem-loop at the 5’ end, which together significantly improve the stability and efficacy of the guide RNA in vivo. The target sites for both mAlb29-8-50b and mAlb29-12-50b are present in mouse albumin intron 1 in the forward orientation (with the PAM sequence defined as being 5’ to the guide target site, or closest to exon 1).

[0204] For MG3-6 / 3-4, two guides called mAlb3634-34 (SEQ ID NO: 26) and mAlb3634-59 (SEQ ID NO: 27) were selected based on their editing activities in the mouse hepatocyte cell line Hepa1-6 and in vivo in the mouse liver. These guides contain chemical modifications that improve guide stability and efficacy in vivo. The target site of guide mAlb3634-34 is present in mouse albumin intron 1 in the reverse orientation (the PAM sequence is on the 3' side of the guide target site, or is defined as being closest to exon 2). The target site of guide mAlb3634-59 is present in mouse albumin intron 1 in the forward orientation (the PAM sequence is on the 5' side of the guide target site, or is defined as being closest to exon 1). Four FVIII donor DNA template cassettes containing four guide RNA target site orientations were designed for each of the nucleases MG29-1 and MG3-6 / 3-4. For MG29-1, donor DNA templates having the orientations FF, RR, FR, and RF were designated as pMG4006 (SEQ ID NO: 16), pMG4007 (SEQ ID NO: 17), pMG4008 (SEQ ID NO: 18), and pMG4009 (SEQ ID NO: 19), respectively. The four MG29-1 donors contain the target sites of both guides mAlb29-8 and mAlb29-12 on both sides of the donor cassette.

[0205] For use with the MG3-6 / 3-4 nuclease, donor DNA templates having the guide target site orientations FF, RR, FR, and RF were designated as pMG4012 (SEQ ID NO: 20), pMG4013 (SEQ ID NO: 21), pMG4014 (SEQ ID NO: 22), and pMG4015 (SEQ ID NO: 23), respectively. The four MG3-6 / 3-4 donors contain the target sites of both guides mAlb3634-34 and mAlb3634-59 on both sides of the donor cassette. The notation "F" means the same orientation as the target site in the genome, so the F orientation of mAlb3634-34 means the PAM site on the 3' side of the target site, and the F orientation of mAlb3634-59 means the PAM site on the 5' side of the target site.

[0206] An overview of the relative orientation of different guide target sites in the donor and genomic target with respect to the location of PAM is shown in Tables 4 and 5 below.

[0207]

Table 4

[0208]

Table 5

[0209] Example 4 - Testing (predictive) of the FVIII donor template delivered by AAV for integration into albumin intron 1 in mice Human FVIII donor templates pMG4006 - pMG4009 (MG29 - 1 targeted) and pMG4012 - pMG4015 (MG3 - 6 / 3 - 4 targeted) are packaged into AAV8 virus capsids using molecular biology techniques. The AAV virus is purified using a CsCl gradient and its purity is analyzed by protein gel electrophoresis to visualize the virus capsid protein. The titer of each virus is determined by quantitative PCR using primers directed against the inverted terminal repeat (ITR) and expressed as vector genome copies per milliliter (vg / mL).

[0210] Messenger RNAs encoding MG29-1 nuclease and MG3-6 / 3-4 nuclease are generated by in vitro transcription of a linearized plasmid template using T7 RNA polymerase, a mixture of ribonucleotides rATP, rCTP, and rGTP, N1-methylpseudouridine, and a capping reagent. An SV40-derived nuclear localization sequence (PKKKRKVGGGGS (SEQ ID NO: 103)) followed by a short linker is included at the N-terminus of the coding sequences of both MG3-6 / 3-4 and MG29-1. A nuclear localization signal from nucleoplasmin preceded by a short linker (SGGKRPAATKKAGQAKKKK (SEQ ID NO: 104)) is added to the C-terminus of the coding sequences of both MG3-6 / 3-4 and MG29-1. Thus, the same nuclear localization signal is used for both MG29-1 and MG3-6 / 3-4. The plasmid also encodes a polyA tail of approximately 100 nt at the 3’ end of both the MG3-6 / 3-4 and MG29-1 coding sequences, which generates a polyA tail in the mRNA. The coding sequences of both MG3-6 / 3-4 and MG29-1 are codon-optimized. The DNA sequence encoding MG3-6 / 3-4 mRNA is shown in SEQ ID NO: 30, and the DNA sequence encoding MG29-1 mRNA is shown in SEQ ID NO: 31. The mRNA is purified on a spin column, its concentration is determined by absorbance at 260 nM, its purity is determined, and is equivalent for both MG3-6 / 3-4 mRNA and MG29-1 mRNA. For in vivo delivery to mice, MG3-6 / 3-4 mRNA or MG29-1 mRNA and their corresponding guide RNAs are separately packaged within lipid nanoparticles (LNPs). The guide RNA and mRNA are packaged separately for both MG3-6 / 3-4 and MG29-1. The lipids are dissolved in ethanol. The mRNA or guide RNA is prepared in water and then diluted in 100 mM sodium acetate (pH 4.0) to create an RNA working stock. Four lipid components are combined in ethanol at the desired ratio to create a lipid working stock. A typical lipid mixture contains cholesterol, DOPE, C12-200, and DMG-PEG-2000 in a molar ratio of 47.5:16:35:1.5.Combine the lipid working stock and the RNA working stock in a microfluidic mixing device at a flow rate of 12 mL / min and a ratio of 1 volume of lipid working stock to 3 volumes of RNA working stock. The mass ratio of C12-200 to RNA in the formulation is 10 to 1. Dilute the formulated LNP 1:1 with 1×PBS, then dialyze twice in 1×PBS for 1 hour each, and subsequently concentrate in a spin concentrator. Formulate the final LNP in 1×PBS buffer, filter sterilize through a 0.2 μM filter, and store at 4°C. Measure the concentration of RNA inside and outside the LNP. Measure the average diameter and polydispersity of the LNP by dynamic light scattering on the final concentrated LNP. Typically, the LNP is in the size range of 80 - 100 nanometers with a PDI < 0.15 and an RNA encapsulation ratio exceeding 90%.

[0211] Mix LNPs encapsulating guide RNA mAlb29-8-50b (SEQ ID NO: 24) or mAlb29-12-50b (SEQ ID NO: 25) and MG29-1 mRNA at a 1:1 RNA mass ratio (guide RNA:mRNA). Mix LNPs encapsulating guide RNA mAlb3634-34 (SEQ ID NO: 26) or mAlb3634-59 (SEQ ID NO: 27) and MG3-6 / 3-4 mRNA at a 1:1 RNA mass ratio (guide RNA:mRNA). Intravenously inject a mixture of the guide RNA LNP and the matching mRNA LNP into wild-type C57Bl / 6 mice via the tail vein at a total RNA dose of 1 mg / kg - 0.25 mg / kg of RNA in a total volume of 0.1 mL per mouse (N = 5 mice per LNP dose).

[0212] The optimal order and timing of administration of AAV-encapsulated donor templates and LNP-encapsulated nuclease mRNA and guide RNA are determined empirically. After AAV transduction of mammalian cells, the single-stranded DNA packaged within AAV (referred to as the AAV genome) is converted to double-stranded DNA in the nucleus of mammalian cells by annealing of complementary strands (positive and negative single strands) that are packaged at equal rates in bulk AAV preparations, or by de novo synthesis of new complementary strands, or by a combination of these two mechanisms. Once converted to double-stranded DNA, the AAV genome undergoes concatemerization, whereby multiple copies of the AAV genome are joined to form circular concatemers that can persist for years as episomes in non-dividing cells in vivo. Double-stranded DNA is a substrate for NHEJ-mediated integration at DSBs, while single-stranded DNA is not a substrate for this mechanism of integration. Given the time it takes for the single-stranded AAV genome to be converted to double-stranded DNA and the fact that the double-stranded AAV genome persists for months to years in vivo, it is logical to administer the AAV virus before editing the nuclease and guide RNA. This is particularly important when the editing nuclease is encoded in the mRNA and delivered together with the guide RNA in the LNP, as this results in transient expression of the nuclease by translation of a limited amount of mRNA into the nuclease protein. The nuclease protein has a limited lifespan within the cell. Therefore, it may be beneficial to first administer the AAV-encapsulated donor template and then the LNP encapsulating the nuclease mRNA and guide RNA. The optimal time between AAV administration and LNP administration is determined empirically and can vary between mammalian species. In the case of mice, the LNP is administered 24 hours to 21 days after the AAV has been evaluated. In one potential study design, wild-type C57Bl6 mice or hemophilia A mice (deficient in murine FVIII) are injected intravenously with AAV encapsulating different FVIII donor cassettes at 5×10 11 vg / kg to 1×10 13Administer at a dose of vg / kg. One day, seven days, fourteen days, or twenty-one days after AAV administration, inject the same mice with an appropriate LNP encapsulating nuclease mRNA and a guide RNA that induces an RNA that matches the guide target site in the previously administered donor template. Plasma is collected from the mice on days 7 and 14 after LNP administration and assayed for human FVIII protein or human FVIII activity using an appropriate assay. For the detection of human FVIII protein, a human FVIII-specific ELISA assay may be used. For the detection of human FVIII activity, a capture Coatest assay may be used in which human FVIII in the mouse plasma sample is first captured on the surface of a 96-well plate by a human FVIII-specific antibody that does not bind to mouse FVIII. After washing away unbound mouse FVIII in the plasma, bound human FVIII can be quantified using a human FVIII activity assay. A human FVIII standard curve is generated in each assay run using purified human FVIII protein diluted in naive mouse plasma such that the levels of mouse plasma are the same as those in the sample during the capture step. Alternatively, if a strain of hemophilia A mice lacking mouse FVIII is used, FVIII activity can be measured directly in the plasma. On day 14 after LNP administration, the mice are sacrificed, the entire liver is flash-frozen, and stored at -80°C. The entire liver lobe is homogenized in genomic digestion using 0.4 mL of buffer per 100 mg of tissue weight in a bead mill. Genomic DNA is purified from an aliquot of the homogenate. To measure the total editing efficiency at the target site of albumin intron 1, the albumin intron 1 region is PCR amplified from 50 ng of genomic DNA in a reaction containing 0.5 micromolar each of primers mAlb90F (SEQ ID NO: 28, CTCCTCTTCGTCTCCGGC) and mAlb1073R (SEQ ID NO: 29, CTGCCACATTGCTCAGCAC) and 1X high-fidelity PCR master mix. The resulting 984 bp PCR product spanning the entire intron 1 of mouse albumin is purified using a column-based purification kit. The PCR product is sequenced by next-generation sequencing (NGS).When a nuclease causes a double-strand break (DSB) in DNA within a living cell, the DSB can be repaired by the cellular DNA repair machinery. When cells such as transformed mammalian cells in culture are actively dividing and there is no repair template, this repair can occur via the NHEJ pathway. The NHEJ pathway can be a process prone to errors that introduce base insertions or deletions at the site of the double-strand break. Thus, these insertions and deletions (INDELS) are characteristics of double-strand breaks that occur and are subsequently repaired, and are widely used as a readout of nuclease editing or cleavage efficiency. Sequence reads are analyzed that align each sequence read to the wild-type target sequence (in this case, albumin intron 1) and calculate the number of reads containing at least one INDEL regardless of the INDEL size within a 10-base pair window on either side of the predicted on-target cleavage site of the nuclease. The same liver genomic DNA is analyzed for integration at the predicted target site in albumin intron 1 by in-out PCR using one primer complementary to a sequence in the genome next to the target site and one primer in the DNA donor template. The frequency of integration in the correct or reverse orientation is measured using either the quantitative real-time PCR or droplet digital PCR version of the in-out PCR assay. The percentage of cells expressing the integrated human FVIII mRNA is measured on liver sections by in situ hybridization using a fluorescent probe designed to detect hybrid mRNA transcripts at the junction between albumin mRNA and FVIII mRNA.

[0213] Example 5 - Comparison of in vivo editing efficiency between MG29-1 and spCas9 To compare the in vivo editing efficiency of MG29-1 nuclease with that of spCas9, dose responses were performed in wild-type C57Bl6 mice. Albumin intron 1 was selected as the genomic target locus for both spCas9 and MG29-1. In silico searches of spCas9 guide target sites in mouse intron 1 using the Chop-Chop algorithm identified a total of 39 potential guides, which were ranked according to their efficiency scores and off-target predictions. Furthermore, guide target sites located within 50 bp of exon 1 or exon 2 were excluded. The top three guides from this ranking were designated mAlbR1 (SEQ ID NO: 43), mAlbR2 (SEQ ID NO: 44), and mAlbR3 (SEQ ID NO: 45) and chemically synthesized using chemical modifications at both the 5' and 3' ends, including methylated bases (represented by the nomenclature mA, mC, mG, and mU) and phosphorothioate backbone linkages (represented by the nomenclature A*, C*, G*, and U*). The editing efficiency of these three guides was evaluated in the mouse liver cell line Hepa1-6 by nucleofection of ribonucleoprotein complexes formed by mixing the guide RNA and spCas9 protein at a 1:2.5 molar ratio (protein to guide RNA). 20 μmol of spCas9 protein was mixed with 50 μmol of guide RNA, and then an electroporation device using program setting EH100 was used to 2×10 5They were nucleofected into Hepa1-6 cells. The nucleofected cells were each transferred to a well of a 48-well plate in fresh growth medium and cultured for 48 hours in a humidified incubator at 5% CO2 / 37 °C. Genomic DNA was purified from the cells and analyzed for editing at the target site of albumin intron 1 by PCR amplification of the target locus using primers mAlb90F and mAlb1073R (SEQ ID NOs: 46 and 47), and a high-fidelity PCR enzyme mix. The PCR products were subjected to Sanger sequencing using primers mAlb282F or mAlb460F (SEQ ID NOs: 48 and 49). The Sanger sequencing chromatograms were analyzed for insertions and deletions (“indels”). The presence of indels at the target site is the result of the generation of double-strand breaks in the DNA, which are then repaired by the cellular repair machinery that is prone to introduce errors that introduce insertions and deletions. The results of the TIDE analysis are shown in Table 6. All three guides generated indel frequencies exceeding 90%, indicating that all three guides are highly active.

[0214]

Table 6

[0215] Guide mALbR2 with extensive chemical modifications was synthesized. The chemical modifications include modifications of 3 bases at the 5'-end and 3 bases at the 3'-end, and have phosphorothioate linkages between the 2'-O-methyl bases and the 3 bases at the 5'-end and the 3 bases at the 3'-end. Furthermore, 33 of the internal bases are modified with 2'-O-methyl (SEQ ID NO: 50). These chemical modifications of the guide RNA of spCas9 have been reported to enable efficient in vivo editing in the mouse liver after delivery of the mRNA of spCas9 and the guide RNA in lipid nanoparticles.

[0216] Guide screening was performed for a guide that targets MG29-1 nuclease to mouse albumin intron 1 and promotes cleavage and indel formation. When the nuclease was delivered as mRNA, the two guides with the highest editing activity in Hepa1-6 cells were mALb29-8 and mAlb29-12. Guide mALb29-8 was selected for in vivo comparison with spCas9 guide mAlbR2 in mice. Chemical and structural modifications to the guide RNA of MG29-1 were optimized by evaluating the effects of different chemical modifications, including 2’O-methyl and 2’-fluoro modified bases, phosphorothioate linkages, and additional stem-loops on guide stability and editing activity.

[0217] Experiments on guide chemistry optimization showed that guide chemical #50 was the most active guide chemical among those tested. When delivered in vivo to mice using MG29-1 mRNA and an LNP encapsulating the same guide RNA sequence that targets mouse albumin intron 1 but has two different guide chemistries (#37 and #50), chemical #50 was approximately 4-fold more potent than chemical #37 at a dose of 0.5 mg / kg. Thus, MG29-1 guide chemical #50 was selected for in vivo testing compared to spCas9 with its cognate guide mAlbR2 (SEQ ID NO: 50).

[0218] Messenger RNAs encoding MG29-1 nuclease or spCas9 nuclease were generated by in vitro transcription of linearized plasmid templates using T7 RNA polymerase, along with a mixture of ribonucleotides rATP, rCTP, and rGTP, N1-methylpseudouridine, and a CleanCAP capping reagent. A nuclear localization sequence derived from SV40 (PKKKRKVGGGGS (SEQ ID NO: 103)) was followed by a short linker and included at the N-terminus of the coding sequences for both spCas9 and MG29-1. A nuclear localization signal from nucleoplasmin preceded by a short linker (SGGKRPAATKKAGQAKKKK (SEQ ID NO: 104)) was added to the C-terminus of the coding sequences for both spCas9 and MG29-1. Thus, the same nuclear localization signal was used for both MG29-1 and spCas9. The plasmid also encoded a polyA tail of approximately 100 nt at the 3’ end of both the spCas9 and MG29-1 coding sequences, which generated a polyA tail in the mRNA. The coding sequences for both spCas9 and MG29-1 were codon-optimized using the same algorithm. The DNA sequence encoding spCas9 mRNA is in SEQ ID NO: 51, and the amino acid sequence encoded by spCas9 mRNA is in SEQ ID NO: 52. The mRNA was purified on a commercial spin column, and the concentration was determined by absorbance at 260 nM, and the purity was determined and found to be equivalent for both spCas9 mRNA and MG29-1 mRNA. For in vivo delivery to mice, spCas9 mRNA / mAlbR2 guide or MG29-1 mRNA / mAlb29-8-50 guide was packaged within lipid nanoparticles (LNPs). The guide RNA and mRNA were packaged separately for both spCas9 and MG29-1. The lipids were dissolved in ethanol. The mRNA or guide RNA was prepared in water and then diluted in 100 mM sodium acetate (pH 4.0) to make an RNA working stock. Four lipid components were combined in specific ratios in ethanol to make a lipid working stock.The exemplary lipid mixture contained cholesterol, neutral lipids such as 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), cationic lipids such as 1,1‘-((2-(4-(2-((2-(bis(2-hydroxydodecyl)amino)ethyl)(2-hydroxydodecyl)amino)ethyl)piperazin-1-yl)ethyl)azanediyl)bis(dodecan-2-ol) (C12-200), and PEG-conjugated lipids such as 1,2-dimyristoyl-rac-glycero-3-methoxypolyethylene glycol-2000 (DMG-PEG-2000) in a molar ratio of 47.5:16:35:1.5. The lipid working stock and the RNA working stock were combined in a microfluidic mixing device at a flow rate of 12 mL / min and a ratio of 1 volume of lipid working stock to 3 volumes of RNA working stock. The mass ratio of C12-200 to RNA in the formulation was 10 to 1. The formulated LNP was diluted 1:1 with 1×PBS, then dialyzed twice in 1×PBS for 1 hour each, and subsequently concentrated in a spin concentrator. The resulting LNP was formulated in 1×PBS buffer, filter-sterilized through a 0.2 μM filter, and stored at 4°C. The concentrations of the inner and outer RNAs of the LNP were measured. The average diameter and polydispersity of the LNP were measured by dynamic light scattering with the resulting concentrated LNP. Representative LNP had a size range of 80 - 100 nanometers with a PDI < 0.15 and an RNA encapsulation ratio of over 90%. The average diameter, polydispersity, and RNA encapsulation efficiency are shown in Table 7 below.

[0219]

Table 7

[0220] LNPs encapsulating guide RNA mAlb29-8-50 and MG29-1 mRNA were mixed at a 1:1 RNA mass ratio. LNPs encapsulating guide RNA mAlbR1 and spCas9 mRNA were mixed at a 1:1 RNA mass ratio. Both LNP mixtures were intravenously injected into wild-type C57Bl / 6 mice via the tail vein at a total RNA dose of 1 mg / kg, 0.5 mg / kg, or 0.25 mg / kg of RNA at a total volume of 0.1 mL per mouse (N = 5 mice per LNP dose). On day 5 after administration, the mice were sacrificed, and the whole liver was flash-frozen and stored at -80 °C. The entire left lateral lobe of the liver was homogenized in a bead mill using 0.4 mL of buffer per 100 mg of tissue weight. Genomic DNA was purified from an aliquot of the homogenate. The albumin intron 1 region was PCR amplified from 50 ng of genomic DNA in a reaction containing 0.5 μmol each of primers mAlb90F (SEQ ID NO: 46, CTCCTCTTCGTCTCCGGC) and mAlb1073R (SEQ ID NO: 47, CTGCCACATTGCTCAGCAC) and 1× high-fidelity PCR master mix. The resulting 984-bp PCR product spanning the entire intron 1 of mouse albumin was purified using a column-based purification kit. The PCR product was sequenced by next-generation sequencing (NGS), and the creation of indels in the target sequence was analyzed and used as an indicator of the creation of double-strand breaks by the Cas enzyme and the involvement of the NHEJ pathway.

[0221] The sequence determination reads were analyzed to align each sequence read to the wild-type target sequence (in this case, albumin intron 1) and calculate the number of reads containing at least one indel regardless of the indel size within a 10-base pair window on either side of the predicted on-target cleavage site of the nuclease. The editing efficiency (indel frequency) in each of five mice per group, as well as the group mean and standard deviation, are summarized in Figure 2. No editing was detected in the control mice injected with PBS buffer. Both spCas9 mRNA / mAlbR2 LNP and MG29-1 mRNA / mAlb29-8-50 LNP resulted in dose-dependent editing. At all three doses, the editing efficiency was higher with MG29-1 mRNA / mAlb29-8-50 LNP than with spCas9 mRNA / mAlbR2 LNP. The average editing efficiencies at the three doses are summarized in Table 8. At a dose of 1 mg / kg (0.5 mg / kg mRNA and 0.5 mg / kg guide RNA), MG29-1 was slightly more potent than spCas9 and resulted in approximately 15% more indels. At a dose of 0.5 mg / kg (0.25 mg / kg mRNA and 0.25 mg / kg guide RNA), MG29-1 resulted in approximately 50% more indels. At a dose of 0.25 mg / kg (0.125 mg / kg mRNA and 0.125 mg / kg guide RNA), MG29-1 resulted in 100% more indels. These data indicate that MG29-1 nuclease, in combination with a properly optimized guide RNA, using the same LNP for delivery and mRNA produced using the same process, is more potent than spCas9 nuclease and a properly modified guide RNA. The superior in vivo editing efficiency of MG29-1 was particularly evident at the lowest dose tested, and MG29-1 was twice as potent as spCas9 at the same dose. These results suggest that the MG29-1 nuclease and a properly modified guide RNA exemplified by chemical #50 may have advantages for in vivo gene editing using LNP delivery.

[0222]

Table 8

[0223] Integration of the FVIII gene cassette in albumin intron 1 in the liver of mice mediated by sequence-specific double-stranded DNA cleavage by the Example 6-MG29-1 RNA-guided nuclease To evaluate whether the human FVIII gene can be integrated into albumin intron 1 in the liver of mice and produce human FVIII protein in the blood of mice, a dual-vector approach was utilized in which the FVIII cassette was delivered in AAV8 virus and the mRNAs encoding the MG29-1 nuclease and the albumin intron 1-targeting sgRNA were delivered using lipid nanoparticles. The human FVIII gene cassette was designed as described in Example 3.

[0224] The FVIII gene cassettes pMG4006 (SEQ ID NO: 56), pMG4007 (SEQ ID NO: 57), pMG4008 (SEQ ID NO: 58), and pMG4009 (SEQ ID NO: 59) were packaged into adeno-associated virus serotype 8 (AAV8) using molecular biology techniques. These viruses were titrated by quantitative PCR measurement of the encapsulated DNA and expressed as genome copies per mL. Synthetic mRNA encoding the MG29-1 nuclease adjacent to the nuclear localization signal (NLS) was produced as described above in Example 1. This MG29-1 mRNA (SEQ ID NO: 53) encoding the amino acid sequence of SEQ ID NO: 54, and sgRNA mA29-8b-50 (SEQ ID NO: 60) or mA29-12b-50 (SEQ ID NO: 61) were co-formulated into LNPs at an RNA mass ratio of 1:1 (mRNA:sgRNA) using the methodology described above in Example 5. After systemic administration of the AAV8 virus by intravenous injection, it is mainly taken up by the liver but also by other tissues. After the AAV transduces hepatocytes, the single-stranded AAV genome is converted into a double-stranded form, which then concatemerizes to form circular head-to-tail and head-to-head multimers, which are thought to represent the major episomal form of the AAV genome that persists in the nucleus for a long time. The process of cell transduction and conversion to a stable concatemerized AAV genome occurs over several days to weeks and may at least partially explain why it can take several weeks for the expression of the transgene encoded by the AAV genome to reach maximal levels. In contrast, LNP delivery of mRNA is a rapid process with maximal gene expression occurring within 24 hours of systemic IV administration in mice, as demonstrated with mRNA encoding the reporter protein luciferase. Therefore, the AAV8-FVIII virus was injected 21 days prior to the administration of the LNP-encapsulated MG29-1 mRNA / sgRNA to allow time for efficient AAV transduction of hepatocytes in the liver and conversion of the single-stranded AAV genome to a stable double-stranded form. Each of four AAV8 viruses encapsulating the FVIII gene cassettes pMG4006, pMG4007, pMG4008, and pMG4009 was intravenously (IV) injected via the tail vein into a group of 5 wild-type C57Bl / 6 mice at 1 × 10 per kg body weight13 Performed with the dose of vector genome (vg). After 21 days, the same mice were intravenously (IV) injected with LNP (formulated at a mass ratio of 1:1 mRNA:sgRNA) encapsulating either MG29-1 mRNA and sgRNA mA29-12b-50, or MG29-1 mRNA and sgRNA mA29-8b-50, at a total RNA dose of 0.5 mg per kg body weight. On day 7 after LNP administration, plasma was collected from each mouse for analysis of human FVIII levels. On day 12 after LNP administration, the mice were sacrificed, plasma was collected from all mice by cardiac puncture, and liver samples were collected for genomic DNA extraction.

[0225] Liver samples were homogenized using 0.4 mL of buffer per 100 mg of tissue weight in a bead mill. Genomic DNA was purified from an aliquot of the homogenate. The albumin intron 1 region was PCR amplified from 50 ng of genomic DNA in a reaction containing 0.5 micromolar each of primers mAlb90F (SEQ ID NO: 46, CTCCTCTTCGTCTCCGGC) and mAlb1073R (SEQ ID NO: 47, CTGCCACATTGCTCAGCAC) and 1X high-fidelity PCR master mix. The resulting 984 bp PCR product spanning the entire intron 1 of mouse albumin was purified using a column-based purification kit. The PCR product was sequenced by next-generation sequencing (NGS).

[0226] Array determination reads were analyzed with a custom Python script that aligned each array read to the wild-type target sequence (in this case, albumin intron 1) and calculated the number of reads containing at least one INDEL regardless of the INDEL size within a 10-base pair window on either side of any of the predicted on-target cleavage sites of the nuclease. The data are summarized in Figure 3, which plots the editing efficiency (INDEL frequency) in each of five mice per group, as well as the group mean and standard deviation. The editing data are also summarized in Table 9 below. The mean INDEL frequencies of eight groups treated with LNPs encapsulating MG29-1 mRNA and albumin-targeting sgRNA ranged from 43% to 53%. There was no significant difference in INDEL frequency between individual groups or between groups 1-4 edited with guide mA29-12b-50 and groups 5-8 edited with guide mA29-8b-50.

[0227]

Table 9

[0228] Human FVIII in plasma from mice in groups 1 - 8 was measured using a human FVIII - specific ELISA kit (documented to have no cross - reactivity to mouse FVIII). Plasma from control untreated mice (group 9) was included in each assay, and values were subtracted from the experimental samples. The background signal from untreated mouse plasma was low. A standard curve consisting of plasma - derived full - length human FVIII pharmaceutical diluted in a matrix of 25% naive C57BL / 6 mouse plasma was run twice on each assay plate over the range of 10 mIU / mL to 200 mIU / mL. Plasma samples from treated mice were also diluted in 25% plasma before being added to the wells of the ELISA plate twice. After binding to the capture antibody and washing with PBS / 0.05% Tween, the bound human FVIII was detected using the biotin - labeled detection antibody in the kit, followed by addition of streptavidin - HRP and TMB substrate, and the absorbance at 450 nM was measured with a plate reader. The concentration of human FVIII (mIU FVIII per mL of plasma) in plasma samples from treated mice was interpolated from the standard curve. On day 7 after LNP administration, human FVIII levels in treated mice were in the range of 13 mIU / mL to 46 mIU / mL, which represents 1.3% - 4.6% of normal FVIII levels in humans (Tables 10 and 11). On day 12 after LNP administration, human FVIII levels in treated mice were in the range of 39 mIU / mL to 120 mIU / mL, which represents 3.9% - 12% of normal FVIII levels (Tables 10 and 11).

[0229]

Table 10

[0230]

Table 11

[0231] Plasma samples on day 7 from mice in groups 5 - 8 (treated with LNP containing guide mA29 - 8b - 50) were reassayed with the same human FVIII - specific ELISA kit, using recombinant human B - domain - deleted FVIII pharmaceutical (Xyntha). The mean human FVIII levels determined using Xyntha as the standard were 67 ± 36 mIU / mL, 33 ± 18 mIU / mL, 104 ± 62 mIU / mL, and 78 ± 53 mIU / mL in groups 5, 6, 7, and 8, respectively. The FVIII levels determined using Xyntha as the standard were on average 2.8 - fold higher than when octanoate was used as the standard. The FVIII gene delivered to the mice encodes a B - domain - deleted FVIII protein, which has the same protein sequence in all four viruses (pMG4006, pMG4007, pMG4008, pMG4009). The B - domain of FVIII contains 908 amino acids (38% of the full - length FVIII protein) and thus constitutes a significant proportion of the FVIII protein in octanoate. In summary, these data demonstrate that delivery of a human FVIII donor cassette to mice, followed by an LNP encapsulating MG29 - 1 mRNA and an sgRNA targeting albumin intron 1, results in measurable human FVIII protein expression that is detectable in the blood of mice on days 7 and 12 after LNP administration.

[0232] The FVIII gene cassette packaged in AAV lacks a promoter to drive the expression of the FVIII gene from the episomal AAV genome. However, since the AAV ITR has weak promoter activity, there is a possibility that RNA encoding the FVIII coding sequence can be transcribed from the episomal genome. However, it is unlikely that episome-derived RNA can be translated into FVIII protein secreted into the blood for two reasons. First, the FVIII gene cassette does not contain an in-frame translation start codon (ATG) and also lacks a translation start consensus sequence (KOZAK sequence) at the 5' end of the cassette. Therefore, any FVIII protein is unlikely to be translated from any putative RNA produced from the episomal AAV genome. Second, the FVIII protein produced from the episomal AAV genome does not contain a signal peptide at the N-terminus. Since the signal peptide is required to induce the secretion of the protein from the cell, any FVIII protein that can be expressed from the episomal AAV genome will not be secreted into the blood. The FVIII gene cassette was designed to express the FVIII protein after integration into albumin intron 1. After integration into albumin intron 1 in the forward orientation, transcription from the endogenous albumin promoter will produce a pre-mRNA encoding albumin exon 1, followed by intron 1 and a portion of the human FVIII coding sequence. Splicing of the pre-mRNA from the albumin exon 1 splice donor to the splice acceptor contained in the FVIII gene cassette will result in a hybrid mRNA transcript in which the albumin exon is fused in-frame to the mature B domain-deleted human FVIII coding sequence. Since albumin exon 1 encodes the signal peptide of albumin, this will provide the signal peptide necessary for the secretion of FVIII into the blood.

[0233] The frequency of FVIII gene integration in albumin intron 1 in the liver of mice was measured using a quantitative assay designed to measure both forward and reverse integrations. These assays used digital droplet PCR (dd-PCR) technology, in which sample genomic DNA was encapsulated into individual droplets such that the droplets either contained a single molecule of DNA or no DNA. Droplets containing DNA that was positive for the integration junction (detected by PCR-based amplification incorporating a fluorescent probe) were scored as positive, and droplets lacking the integration junction were scored as negative. The dd-PCR measurements counted the number of positive droplets and used an algorithm to determine the absolute number of copies of the target amplicon (in this case, the integration junction) in the original genomic DNA sample. The PCR primers and probes were optimized to confirm that the assay was specific and quantitative. An internal control assay for the cytochrome C1 gene was used to normalize to the number of copies of genomic DNA in each sample. The results are presented as the integration rate calculated as the number of copies of the FVIII gene integration junction per 100 copies of cytochrome C1. Two assays for FVIII integration were used, one detecting the 5' junction of the forward integration product and the other detecting the 5' junction of the reverse integration product. The sequences of the primers and probe for quantifying the forward integration product are KAS_401_F1-Fwd: 5’GCACAGATATAAACACTTAACGGGT3’ (SEQ ID NO: 105), KAS_401_F1-Rev: 5’GGAGGAAATCTAGCATCCACAG3’ (SEQ ID NO: 106), KAS_401_F1-Probe: 5’+C+CACCAGAAGA+TAT+T+ACCTG3’ 6-FAM / 3′IBFQ (SEQ ID NO: 107).The primer and probe sequences for quantifying the reverse incorporated product are KAS_501_R2-Fwd: 5’GCACAGATATAAACACTTAACGGG3’ (SEQ ID NO: 108), KAS_501_R2-Rev: 5’TGCTCTGAGAATGGAAGTGC3’ (SEQ ID NO: 109), KAS_501_R2-Probe: 5’+C+GATCAGT+AGAGGTCCTGAGC3’ 6-FAM / 3’IBFQ (SEQ ID NO: 110) (+ indicates a locked nucleic acid base). These results for individual mice are summarized in Table 12.

[0234]

Table 12-1

[0235]

Table 12-2

[0236] Figure 7 shows the forward integration frequencies in individual mice from each group. Animals m1, m4, and m5 were control mice injected with PBS buffer only, and no integration was detected, demonstrating that the assay had no background signal. All mice that received AAV8-pMG4006, AAV8-pMG4007, AAV8-pMG4008, or AAV8-pMG4009, followed by LNPs encapsulating MG29-1 mRNA and guide RNA 12 (mA29-12b-50), had measurable integration in the forward orientation in the range of 0.25% - 2% (0.25 - 2 copies per 100 copies of cytochrome C1). The data for mouse m21 were not reported due to technical problems with the dd-PCR assay for that sample. The average forward integration frequency per group is shown in Figure 8. There was no significant difference among the four groups administered four AAV donors (AAV8-pMG4006, AAV8-pMG4007, AAV8-pMG4008, AAV8-pMG4009) with different orientations of the guide RNA target sites adjacent to the donor, but pMG4008 exhibited the most consistent frequency among mice and the highest average forward integration of 2%. The frequency of integration of the FVIII cassette into albumin intron 1 in the reverse orientation was measured and plotted in Figure 9 together with the forward integration frequency. The reverse integration frequency was in the range of 1% - 2%, and there was no statistically significant difference among the groups that received the four AAV donors (AAV8-pMG4006, AAV8-pMG4007, AAV8-pMG4008, AAV8-pMG4009). Overall, the reverse integration frequency was similar to the forward integration frequency in each mouse, demonstrating that there was no preferential integration of the FVIII cassette in either the forward or reverse direction.

[0237] Figure 10 shows the forward integration frequencies in individual mice from each group administered with AAV8-pMG4006, AAV8-pMG4007, AAV8-pMG4008, or AAV8-pMG4009, followed by an LNP encapsulating MG29-1 mRNA and guide RNA8 (mA29-8b-50). All mice had measurable integration in the forward orientation in the range of 0.25% - 2% (0.25 - 2 copies per 100 copies of cytochrome C1). The average forward integration frequency per group is shown in Figure 11. There was no significant difference among the four groups administered with four AAV donors (AAV8-pMG4006, AAV8-pMG4007, AAV8-pMG4008, AAV8-pMG4009) with different orientations of the guide RNA target sites adjacent to the donor. The frequency of integration of the FVIII cassette into albumin intron 1 in the reverse orientation was measured and plotted in Figure 12 together with the forward integration frequency. The reverse integration frequency was in the range of 0.5% - 3%, and there was no statistically significant difference among the groups receiving the four AAV donors (AAV8-pMG4006, AAV8-pMG4007, AAV8-pMG4008, AAV8-pMG4009). Overall, the reverse integration frequency was similar to the forward integration frequency in each mouse, demonstrating no preferential integration of the FVIII cassette in either the forward or reverse direction.

[0238] The expression of the predicted FVIII-encoding mRNA from the integrated FVIII cassette in the liver of the same mouse was quantified using a dd-PCR assay. Integration of the FVIII cassette in the forward orientation (defined as the 5′ end of the FVIII cassette adjacent to albumin exon 1) at the double-strand break created by the MG29-1 nuclease and guide RNA was predicted to produce a hybrid mRNA resulting from RNA splicing between the albumin exon 1 splice donor and splice acceptor at the 5′ end of the FVIII cassette. Thus, this hybrid mRNA would contain a novel sequence junction between albumin exon 1 and the 5′ end of the coding sequence of mature FVIII, as shown in FIG. 13. A dd-PCR assay was designed where the forward primer was complementary to a sequence within albumin exon 1, the reverse primer was complementary to a sequence within the 5′ end of human FVIII, and the probe spanned the predicted junction between albumin exon 1 and the 5′ end of human FVIII after correct splicing. The sequences of the primers and probe were MG101-set1-FWD: 5′TAACCTTTCTCCTCCTCCTCTT3′ (SEQ ID NO: 111), MG101-set1-REV: 5′TCCACAGCTCCCAGGTAATA3′ (SEQ ID NO: 112), MG101-set1-Probe: 5′TCTTCTGGTGGCCAGTGCTTCTC3′ FAM, ZEN / 3′IBFQ (SEQ ID NO: 113).

[0239] Total RNA was purified from the left lateral lobe of the liver from each mouse. After DNase digestion to remove residual genomic DNA, cDNA was prepared from 500 ng of total RNA. The cDNA was assayed for albumin-FVIII fusion mRNA and cytochrome C1 mRNA. The absolute copy of albumin-FVIII hybrid mRNA was divided by the absolute copy of cytochrome C1 mRNA using the level of cytochrome C1 mRNA, and the quality and quantity of the assayed mRNA were normalized by expressing this as a percentage. The results of mice that received any of four AAV8-FVIII donors (AAV8-pMG4006, AAV8-pMG4007, AAV8-pMG4008, AAV8-pMG4009), as well as LNPs encapsulating MG29-1 mRNA and guide RNA mA29-8b-50 (guide 8), are shown in Figure 14. No signal was detected in PBS-injected control mice, demonstrating that the assay was specific. The levels of albumin-FVIII fusion mRNA were in the range of 20% to 30% of cytochrome C1 in the four groups. The difference between groups was not statistically significant, but there was a tendency for higher levels of albumin-FVIII fusion mRNA in mice that received AAV8-pMG4008 or AAV8-pMG4009.

[0240] The results of mice that received any of four AAV8-FVIII donors (AAV8-pMG4006, AAV8-pMG4007, AAV8-pMG4008, AAV8-pMG4009), as well as LNPs encapsulating MG29-1 mRNA and guide RNA mA29-12b-50 (guide 12), are shown in Figure 15. No signal was detected in PBS-injected control mice, demonstrating that the assay was specific. The levels of albumin-FVIII fusion mRNA were in the range of 5% to 10% of cytochrome C1 in the four groups. The differences between the groups were not statistically significant, but there was a tendency for higher levels of albumin-FVIII fusion mRNA in mice that received AAV8-pMG4007, AAV8-pMG4008, or AAV8-pMG4009. Overall, the levels of albumin-FVIII fusion mRNA were lower in the set of mice that received guide 12 (mA29-12b-50) compared to the set of mice that received guide 8 (mA29-8b-50). The higher levels of albumin-FVIII fusion mRNA observed in mice in which the FVIII gene was integrated into the guide 8 target site did not correlate with a higher frequency of FVIII gene integration in the forward orientation. The highest frequency of FVIII integration in the forward orientation was observed in mice that received pMG4008 and guide 12, and the albumin-FVIII fusion mRNA frequency was 10%, lower than 30% measured in mice that received pMG4008 and pMG4008 and guide 8.

[0241] Example 7 - Demonstration of Integration of the Human FVIII Gene Cassette into Albumin Intron 1 in Mice A cohort of 15 adult C57BL / 6 mice (8 weeks old) was administered 5 × 10 12IV injection of pMG4006 at vg / kg was performed (Groups 2, 3, 4). Five mice in a group were not injected with AAV (Group 1). Twenty-one days later, five mice in Group 3 were given an IV injection of LNP encapsulating MG29-1 mRNA and sgRNA mA29-8-37 (SEQ ID NO: 62), and five mice in Group 4 were given an IV injection of LNP encapsulating MG29-1 mRNA and sgRNA mA29-12-37 (SEQ ID NO: 63). sgRNA mA29-12-37 targets the genomic sequence within albumin intron 1 that is the same as mA29-12b-50 (SEQ ID NO: 61), and sgRNA mA29-8-37 targets the same genomic sequence within albumin intron 1 that is the same as mA29-8b-50 (SEQ ID NO: 60). Mice in Groups 1 and 2 were given an IV injection of PBS buffer as a control. On the 12th day after LNP administration, plasma was collected from all mice by cardiac puncture, and liver tissue was collected for genomic DNA analysis.

[0242] The editing frequency at the target site in the albumin intron in genomic DNA purified from mouse liver was measured by NGS as described in the above examples (Table 13).

[0243]

Table 13

[0244] Groups 1 and 2, which did not receive the LNP, had background-level editing as expected. Mice in Group 3 that were injected with the LNP encapsulating MG29-1 mRNA and sgRNA mA29-8-37 had INDEL frequencies in the range of 43% to 51%. Two mice (Mouse #16 and #19) in Group 4 that were injected with the LNP encapsulating MG29-1 mRNA and sgRNA mA29-12-37 had high INDEL frequencies of 45% and 49%, similar to those observed in Group 3. One mouse (Mouse #18) in Group 4 had a moderate INDEL frequency of 19%. Two mice (Mouse #17 and #20) in Group 4 had low or undetectable INDEL frequencies similar to those of the non-injected mice, indicating that LNP administration likely did not succeed in these two mice.

[0245] The human FVIII levels in plasma samples were measured using a capture chromogenic activity assay in which human FVIII in plasma was captured on the surface of plates and then FVIII activity was measured using a human FVIII-specific antibody that does not bind to mouse FVIII. A 96-well plate was first coated with a mixture of two human-specific anti-FVIII antibodies. The plate surface was blocked with milk powder in PBS buffer and washed with PBS + 0.05% tween. After that, appropriately diluted mouse plasma samples and standards were added to the wells and incubated at 37 °C for 2 hours. The standards were prepared by diluting the European Pharmacopoeia (EP) Reference Standard derived from human plasma in naive C57BL / 6 mouse plasma. The wells were washed three times with PBS containing 0.05% Tween, and the FVIII activity bound to each well was assayed using a commercially available FVIII chromogenic assay kit according to the manufacturer's protocol. The results were back-calculated to the percentage of FVIII levels in normal human plasma present in the mouse plasma. The results (Table 13) demonstrate that no human FVIII activity was measured in the plasma of mice from Group 1 that received neither AAV nor LNP. Mice in Group 2 that received AAV encoding the human FVIII gene cassette but did not receive LNP (and thus were not edited at the albumin locus) also had no detectable human FVIII activity in their blood. This demonstrates that AAV virus alone, which is expected to deliver the FVIII gene cassette and result in an episomal AAV genome, was unable to produce active human FVIII protein. Three out of five mice in Group 3 that received both AAV and LNP (including guide mA29-8-37) had detectable human FVIII activity in their blood at approximately 4% of normal human levels. Two out of five mice in Group 4 that received both AAV and LNP (including guide mA29-12-37) had detectable human FVIII activity in their blood at 2.6% and 13% of normal human levels. Two mice in Group 4 that had no INDELS had no detectable FVIII.In summary, the data from groups 3 and 4 demonstrate that 5 out of 8 mice edited at the albumin locus expressed detectable human FVIII activity in their blood on day 12 after LNP administration.

[0246] To confirm that the expression of human FVIII in the blood of these mice is associated with the integration of the FVIII gene cassette into the sgRNA target site in albumin intron 1, an in-out PCR assay was performed. A pair of PCR primers was designed that was complementary to the DNA sequence adjacent to the predicted junction between albumin intron 1 and the 5' end of the human FVIII gene cassette packaged in the AAV virus. The primer that binds to albumin intron 1 was located 5' to the target site of the sgRNA. When the FVIII gene cassette was integrated at the sgRNA target site within albumin intron 1, genomic DNA purified from the livers of the mice was used as a template to generate a 286 bp PCR product for group 3 (integration into the mA29-8-37 target site) and a 221 bp product for group 4 (integration into the mA29-12-37 target site). The products of the PCR reaction were fractionated on an agarose gel and imaged by staining. As shown in Figure 4, in-out PCR analysis of genomic DNA from the mice in group 1 (lanes 1-3) did not result in bands, indicating the absence of background PCR amplification. Four out of five mice (mice 6, 7, 9, 10) from group 2, which was injected with AAV only (without LNP), also could not generate PCR products. However, mouse 8 from group 2 produced a faint band that was not the correct size for the integrated FVIII gene. This demonstrates that the FVIII gene cassette was not integrated into the guide RNA target site of the mice injected with AAV carrying the FVIII gene cassette and was not edited by the LNP delivery of MG29-1 and the cognate sgRNA. Liver genomic DNA from all five mice (mice #11-#15) in group 3 produced a PCR product that matched the predicted size of 286 bp for the FVIII cassette integrated in the forward (expression competent) orientation at the mA29-8-37 target site. Since this PCR assay is not quantitative, the PCR band intensity does not represent the relative integration frequency among different mice.Liver genomic DNA from 3 out of 5 mice in Group 4 (Mice #16, #18, #19) produced PCR products that matched the expected size of 221 bp for the FVIII cassette integrated in the forward (expression-competent) orientation at the mA29-12-37 target site. Mice #17 and #20 in Group 4 did not produce PCR products, indicating that in these two mice, integration of the FVIII cassette in the forward (expression-competent) orientation was not detected using this assay. This is consistent with the fact that Mice #17 and #20 did not have INDELS at the guide target site of albumin intron 1 (presumably due to failed LNP administration), providing further evidence that integration is dependent on editing by MG29-1.

[0247] In summary, these data demonstrate that MG29-1 nuclease combined with appropriately designed guide RNA can induce the integration of an appropriately designed human FVIII gene cassette into the genome in the liver of mice at the albumin intron 1. Furthermore, this integration results in the expression of functional human FVIII protein in the blood.

[0248] Example 8 - Efficient Editing by MG29-1 Nuclease in the Liver of Non-Human Primates after Systemic Administration of MRG29-1 mRNA and a Suitable Guide RNA Packaged in Lipid Nanoparticles Five guide RNAs for MG29-1 targeting human albumin intron 1 were selected based on testing of 23 guides in the hepatic cell line Hep3B. These five guide RNAs, designated chA29-74B-50 (SEQ ID NO: 64), cA29-78B-50 (SEQ ID NO: 65), chA29-83B-50 (SEQ ID NO: 66), cA29-84B-50 (SEQ ID NO: 67), and cA29-87B-50 (SEQ ID NO: 68), were synthesized with specific chemical modifications to the RNA to improve the in vitro stability and in vivo efficacy of gene editing against MG29-1 (Compound 50). The nomenclature for the modifications to the RNA is as follows, where m indicates that the base has a 2'-O-methyl group (e.g., mG), f indicates that the base has a 2'-fluoro group (e.g., fG), and * indicates that the backbone contains a phosphorothioate bond (e.g., G*G).

[0249] The messenger RNA encoding the MG29-1 nuclease was generated by in vitro transcription of a linearized plasmid template using a mixture of T7 RNA polymerase, and ribonucleotides rATP, rCTP, and rGTP, N1-methylpseudouridine, and a CleanCAP capping reagent. An SV40-derived nuclear localization sequence (PKKKRKVGGGGS (SEQ ID NO: 103)), followed by a short linker, was included at the N-terminus of the coding sequence. A nuclear localization signal from nucleoplasmin preceded by a short linker (SGGKRPAATKKAGQAKKKK (SEQ ID NO: 104)) was added to the C-terminus of the coding sequence. The plasmid also encoded a polyA tail of approximately 100 nt at the 3' end, which generated a polyA tail in the mRNA. The coding sequence of MG29-1 was codon-optimized. The DNA sequence encoding the MG29-1 mRNA was the same as SEQ ID NO: 53, and the amino acid sequence encoded by the MG29-1 mRNA was the same as SEQ ID NO: 54. The mRNA was column-purified, and the concentration was determined by absorbance at 260 nm and the purity was determined.

[0250] The relative efficacy of five guide RNAs was evaluated in primary hepatocytes derived from cynomolgus monkeys (PCH). Cryopreserved PCH cells were seeded into 24-well plates according to the supplier's instructions. Each of MG29-1 mRNA and the guide was mixed at a 1:20 mRNA:guide molar ratio, mixed with transfection reagent according to the manufacturer's protocol, and then applied to PCH. The medium on the cells was changed every 24 hours after transfection, and after 48 hours, genomic DNA was purified from the cells. The target region of albumin intron 1 was PCR amplified using a primer pair and a high-fidelity PCR master mix. The purified PCR products of appropriate size were sequenced by next-generation sequencing. The sequence reads were aligned to the reference sequence of the albumin gene from cynomolgus macaque, and a custom script was used to count the number of reads containing insertions or deletions (indels) at the target site of a specific guide RNA. The editing frequency was defined as the percentage of the total sequence reads containing indels. The results (Figure 5) demonstrate that all five guides resulted in editing at the predicted target sites in a dose-dependent manner. The ranking of guide efficacy from highest to lowest was cA29-87B-50 > chA29-74B-50 > cA29-78B-50 > cA29-84B-50 > chA29-83B-50. The editing efficiencies of these guides and MG29-1 mRNA in single-dose primary human hepatocytes were 38%, 43%, 55%, 40%, and 40.5% for guides chA29-74B-50, cA29-78B-50, chA29-83B-50, cA29-84B-50, and cA29-87B-50, respectively. Thus, in primary human hepatocytes, guide chA29-83B-50 appeared to have the highest efficacy. Based on these data for PCH and PHH, two guides, cA29-87B-50 and chA29-83B-50, were selected to be tested in non-human primates.

[0251] MG29-1 mRNA and either guide chA29-83B-50 (SEQ ID NO: 66) or guide cA29-87B-50 (SEQ ID NO: 68) were co-formulated at a mass ratio of 1:1 (mRNA:guide RNA) in two lipid nanoparticle formulations called L1 and L2. The lipid nanoparticles (LNPs) had an average diameter of less than 100 nm, a polydispersity index of less than 0.12 as measured by dynamic light scattering, and an encapsulation of greater than 85% as measured by the Ribogreen assay. After formulation, the LNPs were buffer-exchanged into phosphate-buffered saline containing sucrose, frozen, and stored at -80 °C for several weeks prior to administration.

[0252] Purpose-bred naive cynomolgus monkeys (Macaca fascicularis) with an average body weight of 1.8 kg ± 0.16 kg were acclimated for 2 weeks prior to dosing. The LNP formulation was thawed and diluted to 0.3 mg total RNA per mL by dilution with sterile 0.9% sodium chloride. Each of four LNP preparations was administered to groups of 3 monkeys by intravenous injection into the tail vein (5 mL / kg based on body weight) over 60 minutes. All animals were monitored by veterinary staff and no serious adverse events were observed. The hold of the dosing solution was assayed for RNA concentration using the Ribogreen assay with a standard containing the guide and mRNA used to produce the LNP. Based on this analysis, the actual doses of RNA administered to each group were 1.4, 1.3, 1.25, and 1.25 mg / kg for groups 1, 2, 3, and 4, respectively, as shown in Table 14.

[0253]

Table 14

[0254] On the 8th day after administration, all groups were sacrificed and samples of different tissues were collected for analysis. Samples from each of the five lobes of the liver were collected separately from each animal and cryopreserved before extraction and purification of genomic DNA. The target region of albumin intron 1 was PCR amplified from the genomic DNA purified from each of the five liver lobes of each animal using a pair of primers and a high-fidelity PCR master mix. The purified PCR products of appropriate sizes were sequenced by next-generation sequencing on an Illumina MiSeq instrument. The sequence reads were aligned to the reference sequence of the albumin gene from Macaca fascicularis, and a custom script was used to count the number of reads containing insertions or deletions (indels) at the target sites of specific guide RNAs. The editing frequency was defined as the percentage of the total sequence reads containing indels. Editing at the predicted target sites for each guide was detected in the livers of all 12 animals (Figure 6). The average editing percentage from the five liver lobes of each animal is presented in Table 15 along with the average of the editing for each group. The average editing in the groups ranged from 29% to 50%. Group 4 (treated with LNP L2 encapsulating MG29-1 mRNA and guide cA29-87B-50) exhibited the highest average editing (50%) and the most consistent editing among three animals per group. Animal 2502 presented squamous skin at the injection site and showed an increase in cytokines on days 1 to 3 after administration, indicating that this animal experienced an acute but self-resolving inflammatory response to the test substance. This suggests that the low level of editing (3%) in this animal was caused by the inflammatory response that prevented efficient delivery to the liver. No animals received anti-inflammatory drugs before or during the study.

[0255] Overall, these data demonstrated that the MG29-1 nuclease, when combined with an appropriate guide RNA and delivered systemically as RNA packaged in LNP, can mediate efficient editing in the livers of non-human primates.

[0256]

Table 15

[0257] Integration of the FVIII gene cassette in albumin intron 1 in the liver of mice mediated by sequence-specific double-stranded DNA cleavage by the Example 9 - MG3-6 / 3-4 RNA-guided nuclease A dual vector approach was utilized to evaluate whether the type II CRISPR system MG3-6 / 3-4 could mediate the integration of the human FVIII gene into albumin intron 1 in the liver of mice and generate human FVIII protein in the blood of mice. The FVIII cassette was delivered in the AAV8 virus, and the mRNA encoding the MG3-6 / 3-4 nuclease and the albumin intron 1-targeting sgRNA was delivered using lipid nanoparticles. The human FVIII gene cassette contained the same human FVIII coding sequence as pMG4006 (SEQ ID NO: 16), pMG4007 (SEQ ID NO: 17), pMG4008 (SEQ ID NO: 18), and pMG4009 (SEQ ID NO: 19), but the flanking sequences were modified to include the target sites of two MG3-6 / 3-4 guide RNAs called mA364-34 and mA364-59 that target intron 1 of mouse albumin. These guide RNAs were selected based on the editing (INDEL) efficiency from a screen of the guides of mG3-6 / 3-4 in the mouse liver cell line Hepa1-6. The FVIII cassettes in pMG4012 (SEQ ID NO: 20), pMG4013 (SEQ ID NO: 21), pMG4014 (SEQ ID NO: 22), and pMG4015 (SEQ ID NO: 23) had different orientations of the flanking guide RNA target sites: pMG4012 (forward-forward orientation), pMG4013 (reverse-reverse orientation), pMG4014 (forward-reverse orientation), and pMG4015 (reverse-forward orientation).

[0258] The FVIII cassettes in pMG4012 (SEQ ID NO: 20), pMG4013 (SEQ ID NO: 21), pMG4014 (SEQ ID NO: 22), and pMG4015 (SEQ ID NO: 23) were packaged into adeno-associated virus serotype 8 (AAV8) using standard methodologies. These viruses were titrated by quantitative PCR measurement of the encapsulated DNA and expressed as genome copies per milliliter. Synthetic mRNA encoding the MG3-6 / 3-4 nuclease adjacent to the nuclear localization signal (NLS) was produced as described above in Example 1. This MG3-6 / 3-4 mRNA (SEQ ID NO: 95) encoding the amino acid sequence of SEQ ID NO: 96, and sgRNA mA364-34-1 (SEQ ID NO: 97) or mA364-59-1 (SEQ ID NO: 98) were co-formulated into LNPs at a 1:1 (mRNA:sgRNA) RNA mass ratio using the methodology described above in Example 5.

[0259] Each of four AAV8 viruses encapsulating the FVIII gene cassettes pMG4012, pMG4013, pMG4014, and pMG4015 was intravenously (iv) injected via the tail vein into a group of five immunodeficient mice at a dose of 1×10 13 vector genomes (vg) per kilogram body weight. Twenty-one days later, the same mice were given an iv injection of LNP encapsulating either MG3-6 / 3-4 mRNA and sgRNA mA364-34-1, or MG3-6 / 3-4 mRNA and sgRNA mA364-59-1, at a dose of 0.5 mg or 0.7 mg total RNA per kilogram body weight (formulated at a 1:1 mRNA:sgRNA mass ratio). Additional mice were given the four AAV viruses but no LNP.

[0260] At the end of the study when the mice were sacrificed (about 5 months after LNP administration), genomic DNA was purified from the liver tissue of each mouse, and analyzed for editing at the target site of the guide RNA by NGS. The results are shown in Fig. 16. Mice administered with LNP encapsulating MG3-6 / 3-4 mRNA and guide RNA mA364-34-1 at either 0.5 or 0.7 mg per kg of body weight had 63% - 70% INDELS throughout the liver. Hepatocytes account for about 60 - 70% of the cells in the liver, and since the LNP used was mainly delivered to hepatocytes, 70% of the indels throughout the liver represent saturation editing of hepatocytes. In contrast, mice administered with LNP encapsulating MG3-6 / 3-4 mRNA and guide RNA mA364-59-1 at either 0.5 or 0.7 mg per kg of body weight had 0% - 35% INDELS throughout the liver. Two groups that received AAV8-pMG4014 and AAV8-pMG4015, as well as LNP encapsulating MG3-6 / 3-4 mRNA and guide RNA mA364-59-1 at 0.5 mpk showed little editing, which was due to failed administration. Overall, the editing efficiency of guide mA364-34-1 was at least twice higher than that of guide mA364-59-1.

[0261] Using dd-PCR assays, integration of the FVIII gene cassette into the target site of albumin intron 1 was measured in either the forward or reverse orientation. Genomic DNA purified from mouse livers was digested with EcoRI. For quantification of the forward orientation, the primers used to amplify the 5′ forward junction were KAS_401_F1-Fwd: 5′GCACAGATATAAACACTTAACGGGT3′ (SEQ ID NO: 105), KAS_401_F1-Rev: 5′GGAGGAAATCTAGCATCCACAG3′ (SEQ ID NO: 106), KAS_401_F1-Probe: 5′+C+CACCAGAAGA+TAT+T+ACCTG3′ 6-FAM / 3′IBFQ (SEQ ID NO: 107). For quantification of the reverse orientation, the primers used to amplify the 5′ reverse junction were KAS_501_R2-Fwd: 5′GCACAGATATAAACACTTAACGGG3′ (SEQ ID NO: 108), KAS_501_R2-Rev: 5′TGCTCTGAGAATGGAAGTGC3′ (SEQ ID NO: 109), KAS_501_R2-Probe: 5′+C+GATCAGT+AGAGGTCCTGAGC3′ 6-FAM / 3′IBFQ (SEQ ID NO: 110) (+ indicates a locked nucleic acid base).

[0262] Using the dd-PCR assay for cytochrome C1, the copy of genomic DNA in each sample was corrected. The results were calculated as the percentage of forward integration (copy of forward integration junction per 100 copies of cytochrome C1). Mice administered with any of the four AAV8 FVIII donors (pMG4012, pMG4013, pMG4014, or pMG4015) and subsequently administered with LNP encapsulating MG3-6 / 3-4 mRNA and sgRNA mA364-34-1 were analyzed for the integration of the FVIII gene cassette at the end of the study. Mice from the group treated with AAV and subsequently with LNP encapsulating MG3-6 / 3-4 mRNA and sgRNA mA364-59-1 were not assayed for integration due to the low level of editing achieved. The results demonstrate that forward integration occurred at a frequency of 0.2% - 1.3% among the 24 mice analyzed (Figure 17). The average forward integration frequency in the group was approximately 0.5% (Figure 18), and no significant difference was observed among the groups. The reverse integration frequency ranged from 0.05% to 0.8% among the 24 mice analyzed (Figure 19). In each of the individual mice that received pMG4013 (a total of 10 mice, 5 receiving 0.7 mpk of LNP and 5 receiving 0.5 mpk of LNP) and pMG4014 (a total of 5 mice), the reverse integration frequency was lower than the forward integration frequency (Figure 19). In contrast, the forward and reverse integration frequencies were similar in each individual mouse of the groups that received AAV8-pMG4012 and AAV8-pMG4015. For AAV8-pMG4012 and AAV8-pMG4015, the guide 34 (mA364-34-1) target sites adjacent to the FVIII cassette are in the forward-forward and reverse-reverse orientations, respectively. For AAV8-pMG4013 and AAV8-pMG4014, which resulted in preferential integration into the desired forward orientation, the guide 34 (mA364-34-1) target sites adjacent to the FVIII cassette are in the reverse-reverse and forward-reverse orientations, respectively.These data demonstrate that including guide target sites of MG3-6 / 3-4 adjacent to the FVIII cassette in an inverse-inverse or forward-inverse orientation results in preferred preferential integration of the FVIII cassette in the forward orientation of albumin intron 1. The forward orientation of integration is preferred because it can be expressed from the albumin promoter to produce mRNA encoding FVIII.

[0263] To evaluate the expression of the FVIII transgene driven by the endogenous albumin promoter, a dd-PCR assay was used. Integration of the FVIII cassette in the forward orientation (defined as the 5' end of the FVIII cassette adjacent to albumin exon 1) at the double-strand break created by the MG3-6 / 3-4 nuclease and guide RNA is predicted to produce a hybrid mRNA resulting from RNA splicing between the albumin exon 1 splice donor and splice acceptor at the 5' end of the FVIII cassette. Thus, this hybrid mRNA will contain a novel sequence junction between albumin exon 1 and the 5' end of the coding sequence of mature FVIII, as shown in Figure 13. A dd-PCR assay was designed where the forward primer is complementary to a sequence within albumin exon 1, the reverse primer is complementary to a sequence within the 5' end of human FVIII, and the probe spans the predicted junction between albumin exon 1 and the 5' end of human FVIII after correct splicing. The sequences of the primers and probe are MG101-set1-FWD: 5’TAACCTTTCTCCTCCTCCTCTT3’ (SEQ ID NO: 111), MG101-set1-REV: 5’TCCACAGCTCCCAGGTAATA3’ (SEQ ID NO: 112), MG101-set1-Probe: 5’TCTTCTGGTGGCCAGTGCTTCTC3’ FAM, ZEN / 3’IBFQ (SEQ ID NO: 113).

[0264] Total RNA was purified from the left lateral lobe of the liver from each mouse using the QIAGEN RNeasy Plus Mini Kit with a genomic DNA elimination column. After ezDNase digestion to remove residual genomic DNA, cDNA was prepared from 500 ng of total RNA. The cDNA was assayed for albumin-FVIII fusion mRNA and cytochrome C1 mRNA. The absolute copy of albumin-FVIII hybrid mRNA was divided by the absolute copy of cytochrome C1 mRNA and expressed as a percentage to normalize for the quality and quantity of the assayed mRNA. In mice that received one of the AAV8-FVIII donor viruses (AAV8-pMG4012, AAV8-pMG4013, AAV8-pMG4014, or AV8-pMG4015) and an LNP encapsulating MG3-6 / 3-4 mRNA and guide RNA mA364-34-1, the albumin-FVIII fusion mRNA levels ranged from 1% to 20% of the endogenous cytochrome C1 mRNA level (Figure 20). Two mice had very low levels of albumin-FVIII fusion mRNA, and the remaining mice exhibited levels between 3% and 20% of the endogenous cytochrome C1 mRNA level. The average albumin-FVIII fusion mRNA level was 7% - 11%, and there was no significant difference between groups (Figure 21).

[0265] Example 10 - Evaluation of FVIII B Domain Substitution Sequences Incorporating Different Numbers of N-Linked Glycosylation Sites and Furin Cleavage Sites The untreated full-length wild-type FVIII protein contains six domains in the order A1-A2-B-A3-C1-C2 (from N-terminus to C-terminus). During post-translational processing, the full-length FVIII protein is cleaved by furin protease, which results in the removal of most of the B domain, generating the mature two-chain form of FVIII in which the heavy and light chains are held together by metal bridges. The B domain of FVIII contains most of the N-linked glycosylation sites in FVIII but is not required for the biological activity of the protein and is generally absent in recombinantly produced FVIII (commonly referred to as B-domain deleted FVIII) that is marketed as a drug for treating patients with hemophilia A. In these B-domain deleted FVIII proteins, the B domain is generally replaced by a short linker sequence called the "SQ linker" that is derived from the N and C termini of the B domain and retains the native furin cleavage site (RHQR, SEQ ID NO: 69), thereby ensuring that the protein is processed into the native two-chain FVIII protein during expression.

[0266] The consensus sequence of the N-linked glycosylation site is the triplet amino acid sequence N-X-S / T, where X represents any amino acid, N represents asparagine, and S / T represents either a serine or threonine residue at the third position. The glycan chain is attached to asparagine. Not all N-X-S / T sequences in a protein are glycosylated; certain sequences are glycosylated, but it cannot be accurately predicted from the surrounding sequences. The N6 amino acid sequence (SFSQNATNVSNNSNTSNDSNVSPPVLKRHQR, SEQ ID NO: 70) has been shown to function in mice.

[0267] Nine possible designs of B domain replacement arrays that meet these criteria were created as shown in Figure 22. These designs, called VAR1-VAR9 (SEQ ID NOs: 71-79), contain 1-3 N-linked glycosylation sites and 0-1 amino acid changes relative to wild-type human FVIII. Four of these designs, VAR2 (ENRSFSQNPPVLKRHQR, SEQ ID NO: 72), VAR3 (EPRSFSQNCSQNPPVLKRHQR, SEQ ID NO: 73), VAR4 (EPRNFSQNCSQNPPVLKRHQR, SEQ ID NO: 74), and VAR8 (ENRSNFSQNCSQNPPVLKRHQR, SEQ ID NO: 78), were selected for experimental evaluation. VAR2 and VAR3 contain one N-linked glycan site and one amino acid difference relative to wild-type FVIII. VAR4 contains two N-linked glycan sites and one amino acid difference relative to wild-type FVIII. VAR8 contains three N-linked glycan sites and three amino acid differences relative to wild-type FVIII. All four of these B domain replacement designs, as well as the SQ linker (EPRSFSQNPPVLKRHQR, SEQ ID NO: 80) and the N6 glycan B domain replacement containing the sequence ENRSFSQNATNVSNNSNTSNASNVSPPVLKRHQR (SEQ ID NO: 99), were inserted in place of the B domain of the human FVIII coding sequence to create constructs called pMG4017 (SQ linker, SEQ ID NO: 81), pMG4018 (VAR2, SEQ ID NO: 82), pMG4019 (VAR3, SEQ ID NO: 83), pMG4020 (VAR4, SEQ ID NO: 84), and pMG4021 (VAR8, SEQ ID NO: 85). It should be noted that for each of these B domain replacements (SEQ ID NOs: 71-79), only the sequence between SFSQN (SEQ ID NO: 114) and PPVLKRHQR (SEQ ID NO: 115) is part of the sequence defined above as the linker that replaces the B domain in B domain-deleted FVIII. The sequence listing includes 3 or 4 residues before "SFSQN (SEQ ID NO: 114)" because some of the novel variants contain amino acid changes or additions within this sequence before "SFSQN (SEQ ID NO: 114)".Since each of these B-domain substitutions contains the native furin cleavage site (RHQR (SEQ ID NO: 69)), it is predicted that the FVIII proteins expressed from all of these constructs will be cleaved by furin to generate two-chain FVIII. The human FVIII coding sequences are identical in these constructs except for the differences in the B-domain substitution sequences and consist of co1 codon-optimized DNA sequences as described below.

[0268] The human FVIII coding sequence used herein lacks the signal peptide and is codon-optimized, which is designed to improve expression in the selected species by selecting more frequently used codons and by removing potential splice sites and other undesirable sequence features. When applied to the FVIII coding sequence for use in human cells, this sequence optimization increased the number of CpG dinucleotides in the FVIII coding sequence from 53 in native human FVIII to 210. Since CpG dinucleotides are recognized by the innate immune response, all CpG dinucleotides were removed by changing any of the codons adjacent to the next most frequent codons with CpG dinucleotides removed. This codon optimization was designated copt1. The overall G / C content of the copt1 codon-optimized B-domain deleted FVIII sequence with the SQ linker is 51%, which is the same as that of native human FVIII. The DNA sequence identity between FVIII-BDD after copt1 codon optimization and native human FVIII was approximately 80%. Since most of the DNA base changes are at the last position of the codon, the 80% identity represents a change to approximately 3 × 20% = 60% of the codons in native FVIII.

[0269] Each of the constructs pMG4017, pMG4018, pMG4019, pMG4020, and pMG4021 contained the same sequence adjacent to the FVIII gene cassette. The target sites of the forward-oriented MG29-1 guide RNAs 8 (mAlb29-8-50) and 12 (mAlb29-12b-50), followed by a spacer sequence, a splice acceptor sequence, and the dinucleotide TG (necessary to maintain the correct reading frame after RNA splicing between albumin exon 1 and the splice acceptor site), were added to the 5’ end of the FVIII coding sequence. A polyadenylation signal, a short spacer sequence, the target site of the reverse-oriented MG29-1 guide RNA 12 (mAlb29-12b-50), followed by the target site of MG29-1 guide RNA 8 (mAlb29-8-50), were added after the stop codon at the 3’ end of the FVIII coding sequence.

[0270] AAV8 virus was produced from constructs pMG4017, pMG4018, pMG4019, pMG4020, and pMG4021 using standard methodologies, and the viral genome copy number was measured by PCR. Adult wild-type C57BL / 6 mice were given an IV injection of each of the AAV8 viruses (AAV8-pMG4017, AAV8-pMG4018, AAV8-pMG4019, AAV8-pMG4020, AAV8-pMG4021) at a dose of 1×10 13 vg / kg. Three weeks later, all mice were given an IV injection of hepatotropic LNPs encapsulating MG29-1 mRNA and guide RNA mA29-8b-50 at a dose of 0.5 mg / kg (total RNA dose per kg body weight). The LNPs were prepared as described in the above examples. All mice were sacrificed 16 days after LNP administration, plasma was collected by cardiac puncture, and liver tissue was collected for purification of genomic DNA or total RNA. Editing across the whole liver was in the range of 40% - 50% in the five groups, without editing detected in the PBS-injected control mice (Figure 23).

[0271] Total RNA was purified from the left lateral lobe of the liver from each mouse using the QIAGEN RNeasy Plus Mini Kit with a genomic DNA elimination column. After ezDNase digestion to remove residual genomic DNA, cDNA was prepared from 500 ng of total RNA. The cDNA was assayed for albumin-FVIII fusion mRNA using the primers MG101-set1-FWD (5’TAACCTTTCTCCTCCTCCTCTT3’ (SEQ ID NO: 111)) and MG101-set1-REV (5’TCCACAGCTCCCAGGTAATA3’ (SEQ ID NO: 112)), and MG101-set1-Probe: (5’TCTTCTGGTGGCCAGTGCTTCTC3’ FAM, ZEN / 3’IBFQ (SEQ ID NO: 113)) and cytochrome C1 mRNA. The level of cytochrome C1 mRNA was used to normalize the absolute copy of albumin-FVIII hybrid mRNA by dividing it by the absolute copy of cytochrome C1 mRNA and expressing this as a percentage, for the quality and quantity of the assayed mRNA. The results (Figure 24) demonstrated that mice administered AAV8-pMG4020 had the highest levels of albumin-FVIII fusion mRNA expressed from the integrated FVIII gene cassette. The levels of albumin-FVIII fusion mRNA were higher in all 4 constructs containing a B domain substitution that included one or more N-linked glycosylation sites, compared to AAV8-pMG4017 in which the B domain was replaced by an SQ linker lacking an N-linked glycosylation site. These data demonstrate that including 1, 2, or 3 N-linked glycosylation sites instead of the B domain of FVIII improved the level of expressed mRNA. AAV8-pMG4020 (containing 2 N-linked glycan sites) had higher levels of mRNA expression than AAV8-pMG4018 and AAV8-pMG4019 (both containing 1 N-linked glycan), but AAV8-pMG4021 (containing 3 N-linked glycans) had levels of mRNA expression similar to constructs with a single N-linked glycan. Thus, more N-linked glycans are not necessarily associated with higher mRNA expression.Based on these results, the B domain var4 substitution design present in pMG4020 (which contains only two N-linked glycans and one amino acid change compared to wild-type FVIII) produces the highest level of albumin-FVIII fusion mRNA encoding the incorporated FVIII protein.

[0272] Plasma from the same mice collected 16 days after LNP administration was assayed for human FVIII protein using a human FVIII-specific ELISA assay. Recombinant human FVIII (Xyntha) spiked into the same percentage of naive mouse plasma was used for the standard curve. FVIII above the background of PBS-injected control mice was detectable in mice administered AAV8-pMG4018, AAV8-pMG4019, AAV8-pMG4020, and AAV8-pMG4021, but not in plasma from mice administered AAV8-pMG4017 (which contains an SQ linker instead of the B domain). Thus, inclusion of the B domain substitution sequences var2, var3, var4, and var8 (which contain one, two, or three N-linked glycosylation sites) enabled detectable levels of human FVIII expression after integration into albumin intron 1 (Figure 25). The highest level of FVIII was seen in mice that received AAV8-pMG4020 (var4), which correlates with a higher level of albumin-FVIII fusion mRNA (Figure 24). AAV8-pMG4020 (var4) contains a B domain substitution by two N-glycan sites.

[0273] Example 11 - Design and Evaluation of the Cynomolgus FVIII Gene Donor Cassette in Mice The cynomolgus macaque (cyno) is an accepted preclinical model that more accurately predicts the behavior of drugs in humans. Human FVIII protein is immunogenic when expressed or administered to non-human primates such as cynomolgus macaques due to differences in the amino acid sequence between the human FVIII protein and the cynomolgus FVIII protein, resulting in the production of neutralizing antibodies within the first 1-2 months after administration. The genome editing approach contemplated herein for the treatment of hemophilia A integrates the FVIII gene into the genome for the purpose of providing a sustained curative therapy from a single administration. To evaluate the durability of this approach in NHPs, the FVIII gene encoding cynomolgus FVIII protein is used. The B-domain deleted form of the Macca Fasicularis FVIII protein sequence was generated by alignment to human FVIII BDD-SQ, in which the B-domain is replaced with a so-called "SQ" linker containing the sequence "SFSQNPPVLKRHQR (SEQ ID NO: 116)". Cynomolgus FVIII has the same sequence as human FVIII around the B-domain junction such that deletion of the B-domain from cynomolgus FVIII results in a junction identical to the SQ linker. Alignment of the B-domain deleted versions (both with the SQ linker) of human and cynomolgus FVIII revealed that the cynomolgus FVIII protein sequence has 28 amino acid differences from human FVIII. To enable detection of cynomolgus FVIII protein derived from the integrated transgene, a single amino acid change F2196K (lysine is shown underlined and in bold in the sequence: ASSY KTNM (SEQ ID NO: 117) was introduced into the B domain-deleted (SQ linker) cynomolgus FVIII protein sequence (SEQ ID NO: 86). Incorporating lysine instead of phenylalanine at residue 2196 (numbered according to full-length human FVIII) blocks the binding of the neutralizing anti-FVIII monoclonal antibody BO2C11. This enables the activity of cynomolgus FVIII-F2196K protein to be measured after neutralizing the endogenous cynomolgus FVIII protein by the BO2C11 antibody. To generate the DNA sequence encoding cynomolgus FVIII-F2196K, the DNA sequence encoding human B domain-deleted FVIII having the var 4 B domain substitution (present in pMG4020) identified in Example 10 was modified to select the most frequently occurring codons that encode cynomolgus FVIII amino acids but do not create CG dinucleotides in the DNA sequence, thereby changing the codons of 28 amino acids different from cynomolgus FVIII. In three cases, the previous codon was changed to avoid creating CG dinucleotides (CpG). Additionally, residue F2196 was changed to lysine, and the most frequently occurring codon of lysine was selected, resulting in protein, SEQ ID NO: 87. The sequence changes made to create protein, SEQ ID NO: 87, are shown in Table 16.

[0274]

Table 16-1

[0275]

Table 16-2

[0276] To enable the integration and expression of FVIII from albumin intron 1 in cynomolgus monkeys, guide cA29-87b (guide 87) was selected for nuclease MG29-1, and guide ch364-58 (guide 58) was selected for nuclease MG3-6 / 3-4. The target sites of both guides were included in the donor DNA sequence adjacent to the FVIII donor cassette. When explaining the orientation of the guide target sites in the FVIII donor cassette, the orientation is relative to the target site in the genome. Therefore, the forward orientation means the same orientation as the target site of that guide in the genome, in this case, the target site within albumin intron 1. For MG29-1, based on the studies described in Example 6, the orientation of the adjacent guide had no significant effect on the integration efficiency, and the forward-FVIII gene-reverse orientation was selected. For the MG29-1 guide cA29-87b, the orientation of the two guide target sites within the donor was forward (5' side from the PAM) at the 5' end of the FVIII cassette and reverse (3' side from the PAM) at the 3' end of the FVIII cassette.

[0277] For the MG3-6 / 3-4 nuclease, including the guide cleavage site adjacent to the FVIII donor in the reverse-FVII gene-reverse or forward-FVIII gene-reverse orientation resulted in a higher frequency of integration of the FVIII gene cassette into albumin 1 in the preferred forward orientation (see Example 9, Figure 19). Based on this data, the reverse-FVIII gene-reverse orientation of the MG3-6 / 3-4 guide was selected for use in the cynomolgus monkey FVIII construct. For the MG3-6 / 3-4 guide ch364-58, the orientation of the two guide target sites within the donor was reverse (3' side from the PAM) at the 5' end of the FVIII cassette and reverse (3' side from the PAM) at the 3' end of the FVIII cassette.

[0278] In addition to the guide RNA cleavage site, the same splice acceptor sequence used in pMG4008 was inserted at the 5’ end of the FVIII coding sequence, followed by insertion of the dinucleotide TG, which maintains the correct reading frame in the mRNA after splicing from albumin exon 1. The same polyadenylation signal was inserted 3’ of the stop codon of the FVIII coding sequence. In addition, short spacer sequences were inserted between the guide target site (TS) and the splice acceptor, and between the polyA signal and the guide target site. These spacer sequences are designed to separate the FVIII gene cassette from small deletions at the guide cleavage site that can occur as part of the NHEJ-driven repair process prior to integration.

[0279] The DNA sequence elements present in the full-length cynomolgus FVIII donor cassette designated as pMG4016 (SEQ ID NO: 88) are as follows. 5’ cA29-87b TS - chA364-58 TS - spacer (16bp) - splice acceptor - TG - mature cynomolgus FVIII - F2196K coding sequence including var4 B domain substitution - stop codon - polyadenylation signal - spacer (10bp) - chA364-58 TS - cA29-87b TS - 3’ (Figure 26).

[0280] To evaluate the cynomolgus FVIII donor cassette in mice, pMG4016 was packaged into AAV6 or AAV8 virus using either a HEK293 (HK)-based packaging system or an sf9 insect cell packaging system (sf) using standard methodologies. Wild-type C57Bl / 6 mice received 1 × 10 13Mice were given an IV injection of 1×10¹¹ vg / kg of AAV6(sf)-pMG4016, AAV8(sf)-pMG4016, or AAV8(HK)-pMG4016. Twenty-one days later, the same mice were given an IV injection of 0.5 mg / kg of LNP A and 0.5 mg / kg of LNP B. LNP A contained MG29-1 mRNA and mA29-8b-50 guide RNA co-formulated at a 1:1 mass ratio. LNP B contained MG29-1 mRNA and cAlb29-87b-50 guide RNA co-formulated at a 1:1 mass ratio. The guide RNA in LNP A targeted mouse albumin intron 1 in the mouse genome. The guide RNA in LNP B targeted the MG29-1 guide RNA target site adjacent to the FVIII gene cassette in pMG4016. Two weeks later, the mice were sacrificed and liver tissue was collected for genomic DNA purification and total RNA purification.

[0281] Editing of the albumin intron 1 target site (targeted by guide mA29-8b-50) in the mouse liver ranged from 45% to 50% across the whole liver, demonstrating efficient delivery of the MG29-1 editing system and the expected editing of the genomic target (Figure 27).

[0282] Integration of the cyno_FVIII gene cassette into the mouse albumin intron 1 target site was quantified using a dd-PCR assay. Genomic DNA purified from mouse liver was digested with EcoRI. A primer and probe set called MG401_set3 that detects the 5' integration junction was used for quantification of integration in the forward orientation. The sequences of the primers and probe were MG401-set3-FWD: 5’TCTTGAGTTTGAATGCACAGAT3’ (SEQ ID NO: 118), MG401-set3-REV: 5’TAGTCCCAGCTCAGTTCCA3’ (SEQ ID NO: 119), MG401-set3-Probe: 5’TGGCCACCAGAAGATATTACCTGGGA3’ FAM, ZEN / 3’IBFQ (SEQ ID NO: 120).

[0283] The dd-PCR assay of cytochrome C1 was used to correct the copy of genomic DNA in each sample. The results were calculated as the percentage of forward integration (copy of forward integration junction per 100 copies of cytochrome C1). The integration of the cynomolgus FVIII gene in the forward orientation was detected in all mice at different levels ranging from about 0.1% to about 2% (Figure 28). Comparing the average forward integration in a group of 5 mice that received 3 different AAV viruses, the forward integration was highest in the group that received AAV8(HK)-pMG4016 (mice #22 - #25), with an average integration of 1% (Figure 28). The average integration frequencies of AAV8(sf)-pMG4016 and AAV6(sf)-pMG4016 were about 0.5%.

[0284] The expression of cynomolgus FVIII-encoded mRNA predicted from the integrated cynomolgus FVIII cassette in the liver of the same mouse was quantified using a dd-PCR assay. Integration of the cynomolgus FVIII cassette in the forward orientation (defined as the 5'-end of the cynomolgus FVIII cassette adjacent to albumin exon 1) at the double-strand break created by MG29-1 nuclease and guide RNA is predicted to produce a hybrid mRNA resulting from RNA splicing between the albumin exon 1 splice donor and splice acceptor at the 5'-end of the cynomolgus FVIII cassette. Thus, this hybrid mRNA will contain a novel sequence junction between albumin exon 1 and the 5'-end of the coding sequence of mature FVIII, as shown in Figure 13. A dd-PCR assay was designed where the forward primer is complementary to a sequence within albumin exon 1, the reverse primer is complementary to a sequence within the 5'-end of cynomolgus FVIII, and the probe spans the predicted junction between albumin exon 1 and the 5'-end of cynomolgus FVIII after correct splicing. The sequences of the primers and probe are MG113-set1: MG113-set1-FWD: 5’CTCTTCGTCTCCGGCTCT3’ (SEQ ID NO: 121), MG113-set1-REV: 5’TCCACAGCTCCCAGGTAATA3’ (SEQ ID NO: 112), MG113-set1-Probe: 5’TCTTCTGGTGGCCAGTGCTTCTC3’ FAM, ZEN / 3’IBFQ (SEQ ID NO: 113).

[0285] Total RNA was purified from the left lateral lobe of the liver from each mouse. After DNase digestion to remove residual genomic DNA, cDNA was prepared from 500 ng of total RNA. The cDNA was assayed for albumin-FVIII fusion mRNA and cytochrome C1 mRNA. The absolute copy of albumin-FVIII hybrid mRNA was divided by the absolute copy of cytochrome C1 mRNA and expressed as a percentage to normalize for the quality and quantity of the assayed mRNA. Albumin-cynomolgus FVIII fusion mRNA was detected in all mice and levels ranged from 2% to 40% of cytochrome C1 (Figure 29). The mean albumin-cynomolgus FVIII fusion mRNA levels for each group of 5 mice administered 3 different AAV viruses were 5% for AAV6(sf)-pMG4016, 10% for AAV8(sf)-pMG4016, and 18% for AAV8(HK)-pMG4016. These mRNA levels correlated with the forward integration frequency measured in the livers of the same mice. The observation that AAV8-pMG4016 packaged in the HEK293 production system resulted in a higher frequency of forward integration and higher levels of albumin-cynomolgus FVIII fusion mRNA suggests that differences in AAV virus quality are an important factor in optimizing integration using this gene editing approach. Overall, these results validated the pMG4016 cynomolgus FVIII cassette design.

[0286] Example 12 - Design of an additional FVIII donor sequence consisting of a human FVIII coding sequence encoding a single-chain FVIII protein The native FVIII protein in circulation is composed of two separate protein chains called the heavy and light chains, which are generated from a single FVIII protein by post-translational cleavage by a protease called furin. To generate a single-chain version of B-domain-deleted human FVIII containing a var4 B-domain substitution, the furin cleavage site at the C-terminus of the B-domain substitution sequence was inactivated. The sequence of the var4 B-domain substitution sequence containing two N-linked glycosylation sites is NFSQNCSQNPPVLK RHQR (SEQ ID NO: 100), and the furin cleavage site is underlined. Three of the four amino acids that make up the furin cleavage site were deleted to obtain a sequence called var4sc (NFSQNCSQNPPVLKR, SEQ ID NO: 89). When incorporated into human B-domain-deleted FVIII, this creates a protein with SEQ ID NO: 90, which has a 1aa substitution (S~N resulting in one of the two N-glycan sites), an insertion of four amino acids (SQNC (SEQ ID NO: 122), which is the normal sequence within the B-domain of FVIII), and a deletion of three residues (RHQ) compared to native FVIII. This sequence was partially selected to minimize the sequence differences with the native human FVIII protein or recombinant B-domain-deleted FVIII proteins.

[0287] Example 13 - Design of a DNA sequence encoding a single-chain FVIII with two N-linked glycans instead of a B-domain consisting of alternative codon optimization To evaluate the functionality of single-chain B-domain-deleted human FVIII containing the var4sc B-domain substitution (SEQ ID NO: 90), two DNA sequences encoding this protein were generated for the purpose of optimizing expression. A given protein amino acid sequence can be encoded by a number of unique DNA sequences due to the redundancy of the genetic code in which several codons encode each amino acid. On average, there are approximately four codons for each amino acid, some amino acids are encoded by two codons, some amino acids are encoded by five or six codons, while one amino acid (methionine) is encoded by a single codon. The number of DNA sequences that can encode a given protein sequence is calculated by summing the number of codons for each amino acid. For example, a protein having only 10 amino acids where each amino acid can be encoded by four possible codons is 4 10 encoded by 1,048,576 DNA molecules. Even the number of possible DNA molecules that can encode a short protein composed of 100 amino acids is very large, approximately 2×10 60 (4 100 ). Large proteins such as B-domain-deleted FVIII (1438 amino acids) can be encoded by approximately 4 1438 DNA molecules. Thus, there are a number of possible DNA sequences that can be used to encode the FVIII protein.

[0288] In living organisms, the frequency of use of each codon in naturally produced proteins tends to correlate with the cellular abundance of the transfer RNA (tRNA) for that codon, and it is generally accepted that codons that occur at lower frequencies can limit translation efficiency in mammalian cells due to limiting the concentration of the corresponding tRNA. Codon usage frequencies are significantly different between prokaryotes and eukaryotes, and adjusting codon usage to match the codon usage of the organism in which a heterologous gene is expressed can improve the level of the protein produced, particularly when a gene identified in prokaryotes is expressed in a eukaryotic system, and vice versa. It has also been reported that codon usage differs in genes highly expressed in specific tissues such as the liver compared to the average codon usage of a large set of genes in the same organism, again suggesting that codon usage has been selected during evolution to allow for high levels of expression. Codons that are utilized at particularly low frequencies in expressed human genes are defined herein as having a codon usage frequency of less than 10 according to the published human codon usage table. In this case, the codon frequency is defined as the number of times the codon is used per 1,000 amino acids.

[0289] This list of rare codons consists of the following 12 codons, with the amino acid and frequency shown in parentheses after each codon. GCG (A, 7.6), TGT (C, 9.6), CAT (H, 9.9), ATA (I, 5.5), CTA (L, 5.9), TTA (L, 5.7), CCG (P, 6.2), CGA (R, 6.1), CGT (R, 4.5), TCG (S, 4.9), ACG (T, 6.4), and GTA (V, 6.9). Alternatively, the rarest codon for each amino acid can be considered, adding an additional 8 codons. GAA (E, 27.5), TTT (F, 17.1), GGT (G, 11 / .), AAA (K, 24.8), CAA (Q, 10.4), TAT (Y, 13.1). GAT (D, 22), and AAT (N, 17.5).

[0290] The native human FVIII DNA sequence in the human genome (with the B domain artificially removed) contains 18 of these rare codons. Comparing the codon usage frequency within the native human FVIII DNA sequence with that of liver-expressed human genes (Table 17) reveals codons that are used at higher or lower frequencies in FVIII compared to genes expressed in the liver.

[0291]

Table 17-1

[0292]

Table 17-2

[0293]

Table 17-3

[0294] In particular, GCT (A), GCG (A), CCG (P), CGG (R), TCG (S), and ACG (T) are used at less than 50% of the frequency that is averaged in liver-expressed genes. In contrast, GAT (D), TTT (F), CAT (H), ATT (I), ACT (T), and GTA (V) are used in FVIII at a frequency more than 150% of the frequency that is averaged in liver-expressed genes. This indicates that native FVIII contains a significant number of codons that occur at frequencies different from those in the average gene expressed in the liver. This difference in codon usage can negatively affect the ability of liver cells to express human FVIII. This provides a basis for modifying the codon usage of the artificially generated FVIII DNA sequence to improve expression without changing the encoded amino acid sequence.

[0295] The mature B-domain deleted FVIII amino acid sequence containing a var4 B-domain substitution was codon-optimized and then all CG dinucleotides were removed using a custom design algorithm that selected the next most frequent codon for any given CG sequence and for any of the resulting codons. The rationale for removing CG dinucleotides is that CpG can be recognized by the innate immune system as part of the cellular response to foreign viral and bacterial DNA and can thus have an adverse effect on DNA delivery or potential expression in vivo. This algorithm removed a total of 210 CG sequences and in the process changed approximately 210 codons while retaining the same encoded amino acid sequence. The resulting DNA sequence was designed as codon optimization 1 (copt1) and is contained in construct pMG4026 (SEQ ID NO: 91). Analysis of codon usage in the copt1 DNA sequence revealed that the copt1 sequence changed the frequency of certain codons compared to the native human FVIII DNA sequence. Of the 12 codons defined as rare in liver-expressed human genes (<10 codons per 1000 amino acids), 6 had reduced frequencies in copt1 compared to native FVIII, while the CAT codon frequency was reduced by 50% and the TGT codon frequency did not change, and the frequencies of 4 codons were already present at low frequencies (1 or 2 occurrences in the entire sequence) in native FVIII. Codons that are utilized at the lowest frequency for their amino acid in liver-expressed genes but are not rare in liver-expressed genes (i.e., have a frequency >10 per 1000 amino acids) all had reduced frequencies in copt1 compared to native FVIII, but by only approximately 50% on average. These differences are summarized in Table 18.

[0296]

Table 18-1

[0297]

Table 18-2

[0298] As shown in Table 19, the copt1 sequence still contains several rare codons that make up 27% - 48% of the codons encoding its specific amino acids. This is likely not optimal for protein expression from this FVIII DNA sequence. Therefore, further reducing or removing these rare codons was included in a second codon optimization design called copt4.

[0299]

Table 19

[0300] Removal of all CpG dinucleotides from a codon-optimized array by using the most frequent codons particularly affects the distribution of arginine codons because 4 out of 6 Arg codons contain a CG sequence and 2 of them are in the most frequently used Arg codons. The mature FVIII-BDD sequence has 70 arginine codons including var 4 B domain substitutions. CpG can be recognized by the innate immune system as part of the cellular response to foreign viral and bacterial DNA. However, CpG also naturally exists in mammalian genomic DNA but has a different distribution and different methylation patterns than in foreign DNA (for example, bacterial CpG is not methylated while eukaryotes tend to methylate CpG). Thus, removal of all CpG may not be necessary to minimize the innate immune response. Sequence elements in the mammalian genome called CpG islands (defined as 200 bp regions of DNA with a G+C content of more than 50%) or high densities of CpG are likely to be problematic.

[0301] The use of arginine codons in cop1 is summarized in Table 20.

[0302]

Table 20

[0303] Analysis of the arginine (R) codons in copt1 revealed an almost exclusive use of the codon AGA (66 occurrences), with the other codons being used only 4 times (AGG). This is in contrast to native FVIII, where all 6 codons are used with a bias towards the AGA and AGG codons. Considering that the codon usage frequency of arginine in liver-expressed genes varies over only a 3-fold range (4.55 - 11.11) and 4 of the 6 codons have similar frequencies, a more uniform distribution of the R codons among the 4 most frequently used codons (AGA, AGG, CGC, CGG) is likely to result in improved expression. Therefore, adjusting the use of R codons was included in a second codon optimization design called copt4, which necessarily increases the CG dinucleotide content. CpGs can be immunogenic, but on the other hand, they are naturally present in mammalian genomic DNA but have a different distribution and different methylation patterns compared to foreign DNA (e.g., bacterial CpGs are not methylated, while in eukaryotes, CpGs tend to be methylated). Therefore, complete removal of all CpGs may not be necessary to minimize the innate immune response. Sequence elements in mammalian genomes called CpG islands (defined as 200 bp regions of DNA with a G + C content of over 50%) or high densities of CpGs are likely to be problematic and were avoided. Another consideration when optimizing the coding sequence for codons is that some codon pairs are used more frequently in the coding region of the gene, while other codon pairs are avoided. Additionally, if possible, out-of-frame TAA or TAG stop codons should be avoided.

[0304] Using a probabilistic design approach, the copt1 sequence was modified and ultimately, through several iterative steps, codon-optimized copt4 encoding B domain-deleted FVIII with var4 B domain substitution was created. After each step, the impact on CpG content and codon usage was evaluated. The resulting DNA sequence encoding B domain-deleted FVIII with var4 B domain substitution is designated copt4 and is included in construct pMG4029 (SEQ ID NO: 92). Table 21 shows the codon usage in copt4 compared to native B domain-deleted FVIII.

[0305]

Table 21-1

[0306]

Table 21-2

[0307] Compared to native FVIII, while only TGT and CAT were retained, the frequency of all 12 rare codons was reduced, while the other 10 were completely removed. Of the 6 codons that are the lowest frequency for the corresponding amino acids but are not rare (i.e., present >10 per 1000 codons in liver-expressed genes), all were significantly reduced (2 - 3 fold) in occurrence compared to native FVIII.

[0308] Codon usage for arginine is similar to that of native FVIII except that the rare codon CGT present in 7 of the 70 R codons in native FVIII was removed. The total CG dinucleotide content of copt4 is 22 compared to 53 for native B domain-deleted FVIII. Thus, a more representative mixture of arginine codons in terms of their use in liver-expressed genes was included, removing one rare arginine codon and maintaining a CpG content lower than that of native FVIII.

[0309] Another notable change in the copt4 sequence compared to the copt1 sequence was within the valine codons. In copt1, 78 out of 88 valine amino acids were encoded by the GTG codon, which is not the most frequently used codon for valine. The most frequently used codons for valine in liver-expressed genes (codons per 1,000 amino acids listed in parentheses) are GTC (29.9), followed by GTG (15), GTT (11.3), and GTA (6.8). In copt4, 44 of the valine residues are encoded by GTC and 44 are encoded by GTG. There are six codons for serine, and three codons (AGC, TCC, TCT) are most frequently used in liver-expressed genes. copt1 utilized only these three most frequent codons (in contrast to native FVIII, which contains 23 AGT codons and 26 TCA codons), and the distribution of these codons was highly skewed towards two codons: AGC occurred 61 times and TCT occurred 43 times. This was adjusted in copt4 such that the codons AGC, TCC, and TCT were distributed at 61, 28, and 29 occurrences, respectively.

[0310] These two DNA sequences encoding the same B-domain deleted human FVIII protein containing a B-domain replacement were synthesized with appropriate flanking sequences that allow expression after integration into albumin intron 1 via RNA splicing from albumin exon 1. Specifically, a splice acceptor site, followed by the dinucleotide TG, was included on the 5'-side of the N-terminus of the FVIII coding sequence. The TG dinucleotide is necessary to maintain the correct reading frame after splicing from albumin exon 1 to the splice acceptor. In addition, the guide target sites of MG29-1 guide 8 (mAb29-8) and MG3-6 / 3-4 guide 12 (mA364-12), as well as appropriate spacer sequences, were placed on the 5'-side of the splice acceptor. A stop codon and polyadenylation signal were added to the C-terminus of the FVIII coding sequence, followed by a spacer, and the target sites of MG29-1 guide 8 (mAb29-8) and MG3-6 / 3-4 guide 12 (mA364-12). The resulting constructs were designated pMG4026 (SEQ ID NO: 91) and pMG4029 (SEQ ID NO: 92), and the codon optimization of the FVIII coding sequences used was copt1 or copt4, respectively (Table 22).

[0311]

Table 22

[0312] AAV8 viruses encapsulating the DNA sequences of pMG4026 and pMG4029m were produced in HEK293 cells using standard protocols. These AAV viruses can be tested in mice for their integration efficiency and levels of circulating albumin-FVIII fusion protein and FVIII protein using the methods described herein.

[0313] Example 14 - Design of an additional FVIII donor sequence consisting of a human FVIII coding sequence encoding a single-chain FVIII protein containing a B-domain replacement with two N-linked glycosylation sites with the amino acid change F309S Passage of FVIII protein through the ER-Golgi apparatus is known to be a limiting factor for FVIII protein secretion. Two DNA sequences encoding a single-chain FVIII protein containing the var4sc B-domain substitution were modified to contain serine instead of phenylalanine at position 309 (amino acid numbering according to full-length FVIII). The resulting DNA sequences were designated pMG4027_co1_F309S (SEQ ID NO: 93), which utilizes alternative codon optimization designated copt2, and pMG4029_co4_F309S (SEQ ID NO: 94), which utilizes copt4 codon optimization. Both sequences encode the same amino acid sequence. Both pMG4027_co2_F309S and pMG4029_co4_F309S contain the same sequences flanking the FVIII coding sequences present in pMG4026 and pMG4029, which are required for splicing from albumin exon 1 after integration at the target site of the guide.

[0314]

Table 23-1

[0315]

Table 23-2

[0316]

Table 23-3

[0317]

Table 23-4

[0318]

Table 23-5

[0319]

Table 23-6

[0320]

Table 23-7

[0321]

Table 23-8

[0322]

Table 23-9

[0323]

Table 23-10

[0324]

Table 23-11

[0325]

Table 23-12

[0326]

Table 23-13

[0327]

Table 23-14

[0328]

Table 23-15

[0329]

Table 23-16

[0330]

Table 23-17

[0331]

Table 23-18

[0332]

Table 23-19

[0333]

Table 23-20

[0334]

Table 23-21

[0335]

Table 23-22

[0336]

Table 23-23

[0337]

Table 23-24

[0338]

Table 23-25

[0339]

Table 23-26

[0340]

Table 23-27

[0341]

Table 23-28

[0342]

Table 23-29

[0343]

Table 23-30

[0344]

Table 23-31

[0345]

Table 23-32

[0346]

Table 23-33

[0347]

Table 23-34

[0348]

Table 23-35

[0349]

Table 23-36

[0350]

Table 23 - 37

[0351]

Table 23 - 38

[0352]

Table 23 - 39

[0353]

Table 23 - 40

[0354]

Table 23 - 41

[0355]

Table 23 - 42

[0356]

Table 23 - 43

[0357]

Table 23 - 44

[0358]

Table 23 - 45

[0359]

Table 23 - 46

[0360]

Table 23-47

[0361]

Table 23-48

[0362]

Table 23-49

[0363]

Table 23-50

[0364]

Table 23-51

[0365]

Table 23-52

[0366]

Table 23-53

[0367]

Table 23-54

[0368]

Table 23-55

[0369]

Table 23-56

[0370]

Table 23-57

[0371]

Table 23-58

[0372]

Table 23-59

[0373]

Table 23-60

[0374]

Table 23-61

[0375]

Table 23-62

[0376]

Table 23-63

[0377]

Table 23-64

[0378]

Table 23-65

[0379]

Table 23-66

[0380]

Table 23-67

[0381]

Table 23-68

[0382]

Table 23-69

[0383]

Table 23-70

[0384]

Table 23-71

[0385]

Table 23-72

[0386]

Table 23-73

[0387]

Table 23-74

[0388]

Table 23-75

[0389]

Table 23-76

[0390]

Table 23-77

[0391]

Table 23 - 78

[0392]

Table 23 - 79

[0393]

Table 23 - 80

[0394]

Table 23 - 81

[0395]

Table 23 - 82

[0396]

Table 23 - 83

[0397]

Table 23 - 84

[0398]

Table 23 - 85

[0399]

Table 23 - 86

[0400]

Table 23 - 87

[0401]

Table 23 - 88

[0402]

Table 23 - 89

[0403]

Table 23 - 90

[0404]

Table 23 - 91

[0405]

Table 23 - 92

[0406]

Table 23 - 93

[0407]

Table 23 - 94

[0408]

Table 23 - 95

[0409]

Table 23 - 96

[0410]

Table 23 - 97

[0411]

Table 23-98

[0412]

Table 23-99

[0413]

Table 23-100

[0414]

Table 23-101

[0415]

Table 23-102

[0416]

Table 23-103

[0417]

Table 23-104

[0418]

Table 23-105

[0419]

Table 23-106

[0420] Preferred embodiments of the present disclosure are shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present disclosure be limited by the specific examples provided within this specification. The present disclosure is described with reference to the foregoing specification, but the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, modifications, and substitutions will occur to those skilled in the art without departing from the present disclosure. Further, it should be understood that all aspects of the present disclosure are not limited to the spe...

Claims

1. A manipulated nuclease system, a) an endonuclease or nucleic acid encoding the endonuclease, comprising an amino acid sequence having at least 80% sequence identity with SEQ ID NO: 54 or SEQ ID NO: 96, b) an engineered guide polynucleotide which can form a complex with the endonuclease and ii) a target nucleic acid sequence in or within an intron of the albumin gene, c) A donor template polynucleotide comprising a nucleic acid encoding the factor VIII (FVIII) gene or a functional fragment thereof, having at least 80% sequence identity with any one of sequence numbers 12, 13, 16-23, 32, 33, 56-59, 81-88, and 91-94, and A modified nuclease system, including [specific component].

2. a) The target nucleic acid sequence in the albumin gene is located within intron 1 of the albumin gene, and / or b) The nucleic acid encoding the factor VIII (FVIII) gene or a functional fragment thereof is linked to a splice acceptor sequence that targets exon 1 of the albumin gene, The manipulated nuclease system according to claim 1.

3. The manipulated nuclease system according to claim 1, wherein the nucleic acid encoding the endonuclease has at least 80%, at least 90%, or 100% sequence identity with any one of SEQ ID NOs: 31, 53, 30, and 95.

4. The manipulated nuclease system according to claim 1, wherein the manipulated guide polynucleotide comprises a sequence having at least 90% or 100% sequence identity with any one of SEQ ID NOs: 14-15, 24-27, 43-45, 50, 55, 60-68, and 97-98.

5. The manipulated nuclease system according to claim 1, wherein the target nucleic acid sequence includes a sequence having at least 80%, at least 90%, or 100% sequence identity with any one of sequence numbers 1, 2-6, and 8.

6. a) The FVIII gene or its functional fragment is codon-optimized to remove at least one cytosine-guanine (CG or CpG) motif, b) The FVIII gene or its functional fragment has at least 90% identity with any one of SEQ ID NOs: 10, 71-79, and 89, and / or c) The FVIII gene or its functional fragment containing the modified B domain contains a sequence that has at least about 90% or 100% identity with any one of SEQ ID NOs: 86-87 and 90. The manipulated nuclease system according to claim 1.

7. The modified nuclease system according to claim 1, wherein the donor template polynucleotide further comprises a polyadenylation signal and / or a nuclear targeting sequence.

8. A method for supplementing liver enzyme expression in a subject that requires such supplementation, wherein the subject a) an endonuclease or nucleic acid encoding an endonuclease, comprising an amino acid sequence having at least 80% sequence identity with SEQ ID NO: 54 or SEQ ID NO: 96, b) an engineered guide polynucleotide which can form a complex with the endonuclease and ii) a target nucleic acid sequence in or within an intron of the albumin gene, c) A donor template polynucleotide comprising a nucleic acid encoding the factor VIII (FVIII) gene or a functional fragment thereof, wherein the nucleic acid has at least 80% sequence identity with any one of sequence numbers 12, 13, 16-23, 32, 33, 56-59, 81-88, and 91-94, and A method comprising administering a substance to thereby supplement liver enzyme expression in the subject.

9. a) The target nucleic acid sequence in the albumin gene is located within intron 1 of the albumin gene, and / or b) The nucleic acid encoding the factor VIII (FVIII) gene or a functional fragment thereof is operably linked to a splice acceptor sequence that targets exon 1 of the albumin gene, The method according to claim 8.

10. The method according to claim 8, wherein the nucleic acid encoding the endonuclease has at least 80%, at least 90%, or 100% sequence identity with any one of sequence numbers 31, 53, 30, and 95.

11. The method according to claim 8, wherein the manipulated guide polynucleotide comprises a sequence having at least 90% or 100% sequence identity with any one of SEQ ID NOs: 14-15, 24-27, 43-45, 50, 55, 60-68, and 97-98.

12. The method according to claim 8, wherein the target nucleic acid sequence includes a sequence having at least 80%, at least 90%, or 100% sequence identity with any one of sequence numbers 1, 2-6, and 8.

13. a) The FVIII gene or its functional fragment is codon-optimized to remove at least one cytosine-guanine (CG or CpG) motif, b) The FVIII gene or its functional fragment has at least 90% identity with any one of SEQ ID NOs: 10, 71-79, and 89, and / or c) The FVIII gene or its functional fragment containing the modified B domain contains a sequence that has at least 90% or 100% identity with any one of sequence numbers 86-87 and 90, The method according to claim 8.

14. The method according to claim 8, wherein the donor template polynucleotide further comprises a polyadenylation signal and / or a nuclear targeting sequence.

15. A cell comprising the manipulated nuclease system according to any one of claims 1 to 7.

16. The cell is a) Liver cells, b) eukaryotic cells; c) mammalian cells; d) immortalized cells; e) insect cells, f) yeast cells; g) plant cells; h) fungal cells; i) prokaryotic cells; j) A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC1, BSC40, BMT10, WI38, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or their derivatives. k) Manipulated cells, or l) Stable cells The cell according to claim 15.

17. Lipid nanoparticles (LNPs) comprising components (a) and (b) of the manipulated nuclease system according to any one of claims 1 to 7, or components (a), (b), and (c).

18. The lipid nanoparticle according to claim 17, wherein the LNP comprises a cationic lipid, a neutral lipid, cholesterol or a cholesterol analog, and a PEG-bound lipid.

19. The lipid nanoparticle according to claim 18, wherein the cationic lipid comprises C12-200(1,1'-((2-(4-(2-((2-((bis(2-hydroxydodecyl)amino)ethyl)(2-hydroxydodecyl)amino)ethyl)piperazine-1-yl)ethyl)azandiyl)bis(dodecane-2-ol)), the neutral lipid comprises 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), or the PEG-bound lipid comprises 1,2-dimiristoyl-rac-glycero-3-methoxypolyethylene glycol-2000 (DMG-PEG-2000).

20. A viral vector comprising the manipulated nuclease system according to any one of claims 1 to 7.

21. The viral vector according to claim 20, wherein the viral vector is an adeno-associated virus (AAV) vector.

22. The AAV is AAV8, AAV6, AAV1, AAV2, AAV3, AAV4, AAV5, AAV7, AAV9, AAV10, AA V11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh10, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65 , AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, AAV-3B, AAV-LK0 3. The viral vector according to claim 21, which is AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, AAV-HSC16, or a derivative thereof.

23. A method for supplementing liver enzyme expression in a subject that requires such supplementation, wherein the subject a) Lipid nanoparticles (LNPs) comprising (i) an endonuclease or a nucleic acid encoding the endonuclease, wherein the endonuclease has the amino acid sequence of SEQ ID NO: 54, and (ii) an engineered guide polynucleotide that can form a complex with the endonuclease and can hybridize to a target nucleic acid sequence in intron 1 of the albumin gene, b) An AAV8 vector comprising a donor template polynucleotide comprising a nucleic acid or functional fragment thereof encoding the factor VIII (FVIII) gene, wherein the nucleic acid has at least about 90% identity with SEQ ID NOs. 12, 13, 16-23, 32, 33, or 56-59, and A method including administering [a substance].