Homology-independent targeted integration for gene editing

JP2025515030A5Pending Publication Date: 2026-03-19FOND AZIONE TELETHON
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
FOND AZIONE TELETHON
Filing Date
2023-05-02
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Current gene therapy methods, such as those using adeno-associated viruses (AAVs), face challenges in achieving stable and efficient liver transgene expression, particularly in pediatric patients and in tissues with limited regenerative potential.

Method used

The use of Homology-Independent Target Integration (HITI) systems that target the albumin gene locus in hepatocytes, allowing for the stable integration and high-level expression of therapeutic transgenes without interfering with the expression of the albumin gene.

Benefits of technology

HITI systems achieve stable, high-level expression of therapeutic transgenes in the liver, maintaining expression even in differentiated cells and during liver regeneration, thus addressing the limitations of existing gene therapy approaches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000124_0000
    Figure 00000124_0000
  • Figure 00000124_0001
    Figure 00000124_0001
  • Figure 00000124_0002
    Figure 00000124_0002
Patent Text Reader

Abstract

The present invention relates to a method for integrating an exogenous DNA sequence into the genome of a cell, the method comprising contacting a cell with a donor nucleic acid comprising said exogenous DNA sequence and optionally one or more albumin exons, said donor nucleic acid flanked at 5' and 3' by inverted target sequences, a complementary oligonucleotide homologous to the target sequence, and a nuclease that recognizes said target sequence, wherein said target sequence is located at the 3' end of the albumin gene in a region selected from intron 12, intron 13 and intron 14 of said albumin gene. The present invention also relates to systems, vectors and pharmaceutical compositions comprising said donor nucleic acid and / or complementary oligonucleotides homologous to said target sequence and / or nucleases, and to their medical uses.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for integrating a foreign DNA sequence into the genome of a cell by contacting the cell with a donor nucleic acid together with a nuclease that recognizes the target sequence and a complementary oligonucleotide that is homologous to the target sequence. The present invention also relates to constructs, vectors, systems and pharmaceutical compositions comprising the donor nucleic acid and / or the complementary oligonucleotide that is homologous to the target sequence and / or the nuclease, as well as medical uses thereof. [Background technology]

[0002] In recent years, genome editing has emerged as a viable option for treating genetic diseases. Genome editing uses endonucleases, typically CRISPR / Cas9 [1, 2]. CRISPR-Cas9 is a ribonucleoprotein that binds to a sequence called a guide RNA (gRNA) and uses it to recognize a target DNA sequence by Watson-Crick base complementarity. This target DNA sequence must be adjacent to a protospacer adjacent motif (PAM) sequence, which allows Cas9 to bind to DNA and cleave the target sequence [3]. Cas9's RNA-based targeting facilitates its design to target different genetic loci and even allows it to target two different sequences by delivering Cas9 and two different gRNAs into the same cell. After Cas9 targets a specific location in the genome, it generates a double-strand break (DSB), which is repaired by one of two repair mechanisms:

[0003] Non-homologous end joining (NHEJ) is the most prevalent mechanism in most cell types. It is active in all phases of the cell cycle and consists of random insertion or deletion of bases at the DSB site for repair. This random insertion or deletion (INDEL) often causes a change in the reading frame, knocking out the expression of the target gene [3].

[0004] Homology-directed repair (HDR) is a process that primarily occurs during the G and S2 phases of the cell cycle and precisely corrects DSBs using a homologous template, which can be provided by external donor DNA or another allele [3]. HDR-based gene repair has been successfully used in vitro [4] and in vivo [5-8], even in the absence of Cas9 [6]. However, its efficiency is limited in vivo by the low activity of the homologous recombination pathway in differentiated cells [9]. Therefore, the field is seeking alternative therapeutic gene replacement strategies that enable gene repair in tissues and differentiated cells that are not undergoing active regeneration.

[0005] Homology-independent targeted integration (HITI) was recently developed to overcome the limitations of both HDR-based gene repair and allele-specific knockout [10, 11]. HITI uses donor DNA flanked by identical gRNA target sequences within the gene of interest. After Cas9 cleaves both the gene and the donor DNA, the cellular NHEJ mechanism can incorporate the donor DNA into the repair of the break, with surprisingly high integration rates (60–80%) without INDELs. The possibility of the donor DNA integrating in the reverse orientation is avoided by inverting its gRNA target sequence, and if reverse integration occurs, Cas9 can re-recognize and cleave the target sequence. Because HITI uses NHEJ, it is effective in terminally differentiated cells such as neurons or tissues such as the liver (e.g., both adult and pediatric tissues), regardless of their regenerative potential

[11] . Furthermore, HITI-mediated insertion of a wild-type copy of a therapeutic gene has the potential to confer therapeutic efficacy independent of the specific disease-causing mutation and the underlying proliferative state of the target cell

[11] .

[0006] The inventors previously discovered that HITI can be used to convert the liver into a factory for delivering high levels of therapeutic proteins throughout the body, which is desirable for treating many inherited and common conditions caused by loss of function or condition, such as hemophilia, LSD, and diabetes, where the factor to be replaced must be secreted from the liver and / or reach other target organs through the blood to perform its function, overcoming the limitations of currently available treatments such as inefficient enzyme replacement therapy, conventional gene therapy, and gene editing.

[0007] Adeno-associated virus (AAV)-based vectors are most frequently used for in vivo gene therapy applications due to their safety profile, broad tropism, and ability to provide long-term transgene expression.

[12] However, given the episomal state of the AAV genome, hepatic transgene expression from AAV can be lost over time in the developing liver or if there is liver damage.

[13] This has resulted in limited success, for example, in pediatric patients.

[0008] Therefore, there is a need for more stable and efficient transgene expression in the liver. The HITI developed by the present inventors overcomes these limitations by inserting the coding sequence of a secreted protein of interest into the highly transcribed albumin locus [5-8], providing long-term expression of high levels of systemically secreted protein while maintaining endogenous expression of albumin protein.

[0009] Mucopolysaccharidosis type VI (MPS VI) is a rare lysosomal storage disease (LSD) caused by arylsulfatase B (ARSB) deficiency, resulting in widespread accumulation and urinary excretion of toxic glycosaminoglycans (GAGs). Clinically, the MPS VI phenotype is characterized by growth retardation, coarse facial features, skeletal deformities, joint stiffness, corneal opacities, cardiac valve thickening, and organomegaly, without primary cognitive impairment.

[14] Treatment of MPS relies on normal lysosomal hydrolases being secreted and then internalized by most cells via the mannose-6-phosphate receptor pathway.

[0010] We previously demonstrated that a single systemic administration of a recombinant AAV vector, serotype 8 (AAV2 / 8), encoding ARSB under the transcriptional control of the liver-specific thyroxine-binding globulin (TBG) promoter (AAV2 / 8.TBG.hARSB) resulted in sustained liver transduction and phenotypic improvement in an MPS VI animal model [15-21]. We also showed that this was at least as effective in MPS VI mice as weekly administration of enzyme replacement therapy (ERT), the current standard of care for this condition [22-24]. We recently initiated a phase I / II clinical trial (ClinicalTrials.gov identifier: NCT03173521) to test both the safety and efficacy of this approach in MPS VI patients.

[0011] Hemophilia A (HemA) is a severe bleeding disorder caused by deficient or complete absence of activity of clotting factors VIII or 8 (FVIII or F8). It is the most common inherited X-linked recessive coagulation disorder, with an incidence of approximately 1 in 5,000 male births worldwide.

[0012] Approximately 50% of HemA cases are severe, i.e., circulating FVIII levels are below 1%

[25] . Severe forms of HemA are clinically characterized by spontaneous musculoskeletal and soft tissue bleeding and the inability to achieve hemostasis after trauma without infusion of clotting factor concentrates. Current treatment involves prophylactic administration of recombinant or plasma-derived coagulation FVIII. Frequent infusions (two or three times per week) are required and can be burdensome. Furthermore, they do not prevent spontaneous bleeding and always carry a high risk of developing neutralizing anti-FVIII antibodies (inhibitors) (approximately 25–30% of patients). Therefore, gene therapy as an alternative holds great promise for providing a lifelong cure with a single administration.

[0013] The use of HemA gene therapy has been extensively studied over the past two decades since it was observed that even small improvements in FVIII levels (1-2%) can significantly reduce the risk of spontaneous bleeding and decrease the need for FVIII replacement infusions. Furthermore, gene therapy has a wide therapeutic window, eliminating the need for strict control of gene expression and providing a readily quantifiable therapeutic endpoint (FVIII plasma levels). Several gene transfer strategies for FVIII replacement have been evaluated, and adeno-associated virus (AAV) vectors have emerged as the most promising due to their excellent safety profile and ability to direct long-term transgene expression from postmitotic tissues such as the liver. However, HemA poses significant challenges for AAV gene therapy because the size of the F8 gene coding sequence (7 kb) delivered exceeds the standard carrying capacity of AAV (4.7 kb). Previously, a 5 kb expression cassette containing the B-domain deleted (BDD) F8 transgene and both a short liver-specific promoter and poly(A) signal was packaged into AAV5 and shown to deliver therapeutic levels of FVIII in mice and cynomolgus monkeys, as well as in HemA patients

[26]

[27] . However, the genome of this vector is slightly oversized and packaged in the AAV capsid as a library of heterogeneously truncated genomes, resulting in ineffective transduction upon reassembly in target cells. The efficiency of oversized AAV vectors is lower than that of normal-sized vectors, and the quality of such products with heterogeneously truncated genomes may hinder their further development. All AAV-based products currently in clinical development consist of a B-domain deleted (BDD) version of the F8 transgene, approximately 4.4 kb in size

[28] . Nevertheless, such a large transgene limits the space available for necessary regulatory elements in the vector, thereby limiting the choice of promoter and poly(A) signal. Furthermore, all of these vector genomes are at the limit of the normal carrying capacity of AAV and are at risk of being improperly packaged as a library of heterogeneous truncated genomes.Despite the ability of such oversized vectors to successfully express large proteins, their long-term efficiency and safety have yet to be confirmed [28-32].

[0014] HITI in hepatocytes of the high-transcription albumin locus has the potential to overcome some of the limitations of other safe and effective liver gene therapies using AAV, including: (i) the level of transgene expression (especially high expression from the albumin locus); (ii) stability of transgene expression, ensured by insertion of the therapeutic coding sequence into a genomic locus that replicates in the event of hepatocyte loss and in the developing liver (enabling treatment in pediatric patients);

[0015] A conventional approach for integrating foreign DNA sequences into the genome of a cell based on HITI is disclosed in WO2020079033, which specifically targets the second exon of the albumin gene, in which the targeted albumin locus is disrupted by insertion of donor DNA carrying the transgene of interest. There remains a need for gene therapy strategies for diseases that require stable systemic expression of therapeutic proteins. Summary of the Invention

[0016] The inventors have discovered a new approach for integrating foreign DNA sequences into the genome of cells by utilizing the HITI system to target the albumin gene without interfering with its expression. Evidence of efficacy was provided in MPS VI mice using ARSB, which encodes arylsulfatase B, a lysosomal enzyme deficient in mucopolysaccharidosis VI (MPS VI), and in hemophilic mice using the F8 CodopV3 transgene, which encodes factor VIII, deficient in hemophilia A.

[0017] The inventors have found that integration of a transgene, such as ARSB or F8, into the 3' end of mouse albumin (mAlb) via the novel HITI system results in increased levels and / or activity of the defective enzyme. In particular, integration of ARSB resulted in supraphysiological levels of circulating enzyme at two doses tested, while a single dose induced lower levels. Furthermore, MPS VI mice showed phenotypic improvement up to 36 weeks after neonatal administration, while F8 activity levels in hemophilic mice treated with the inventive system increased by 20% compared to unaffected controls. The inventors have also demonstrated in vitro integration of a reporter gene into the 3' end of human albumin (hALB).

[0018] Overall, the novel HITI results in stable, high expression levels of therapeutic transgenes from the liver accompanied by hepatocyte proliferation.

[0019] The present invention relies on the insertion of a sequence of interest into the albumin gene locus, a gene that is highly expressed in the liver. The gene of interest is expressed under the albumin promoter, resulting in high-level expression of the gene of interest, albeit from a relatively small number of cells in the liver parenchyma. Therefore, the expression of the gene of interest is high enough to achieve a therapeutic effect. Furthermore, since the gene of interest is stably integrated into the liver genome, the expression of the gene of interest is not lost during tissue regeneration (in children or upon liver injury). The albumin gene locus is targeted using a strategy that allows the maintenance of albumin gene expression.

[0020] An object of the present invention is a method for integrating a foreign DNA sequence into the genome of a cell, comprising contacting the cell with: a) a donor nucleic acid comprising: - the foreign DNA sequence - optionally one or more albumin exons; wherein the donor nucleic acid is flanked at 5' and 3' by inverted target sequences, b) a complementary oligonucleotide homologous to the target sequence; c) a nuclease that recognizes the target sequence wherein said target sequence is located at the 3' end of the albumin gene in a region selected from intron 9, intron 11, intron 12, intron 13 and intron 14 of said albumin gene. In the present invention, the donor nucleic acid preferably comprises one or more albumin exons, said exons being exon 10 and / or exon 11 and / or exon 12 and / or exon 13 and / or exon 14 or fragments thereof.

[0021] An object of the present invention is a method for integrating a foreign DNA sequence into the genome of a cell, comprising contacting the cell with: a) a donor nucleic acid comprising: - the foreign DNA sequence - optionally one or more albumin exons; wherein the donor nucleic acid is flanked at 5' and 3' by inverted target sequences, b) a complementary oligonucleotide homologous to the target sequence; c) a nuclease that recognizes the target sequence wherein said target sequence is located at the 3' end of the albumin gene in a region selected from intron 12, intron 13 and intron 14 of said albumin gene. In the present invention, the donor nucleic acid preferably comprises one or more albumin exons, said exons being exon 13 and / or exon 14 or fragments thereof.

[0022] In a preferred embodiment, said albumin exon is present and is exon 10 and / or exon 11 and / or exon 12 and / or exon 13 and / or exon 14 or fragments thereof. Preferably, it is located at the 5' end of the foreign DNA sequence and can be separated therefrom by a ribosomal skipping sequence.

[0023] In a preferred embodiment, the albumin exon is present and is exon 13 and / or exon 14 or a fragment thereof. Preferably, it is located at the 5' end of the foreign DNA sequence and can be separated therefrom by a ribosomal skipping sequence. In the present invention, the donor nucleic acid preferably comprises one or more albumin exons, wherein the exons are exon 10 and / or exon 11 and / or exon 12 and / or exon 13 and / or exon 14 or a fragment thereof.

[0024] In the present invention, the donor nucleic acid preferably comprises one or more albumin exons, said exons being exon 13 and / or exon 14 or fragments thereof.

[0025] The albumin exons can be from an albumin gene of any origin, but preferably they are from a human or mouse albumin gene.

[0026] Preferably, the complementary strand oligonucleotide homologous to the target sequence is a guide RNA that hybridizes to a target sequence located within intron 9, intron 11, intron 12, intron 13 or intron 14 of the albumin gene, or its complementary strand, and preferably, the guide RNA is adjacent to a protospacer adjacent motif (PAM) sequence.

[0027] Preferably, the complementary oligonucleotide homologous to the target sequence is a guide RNA that hybridizes to a target sequence located within intron 12, intron 13 or intron 14 of the albumin gene, or its complementary strand, and preferably the guide RNA is adjacent to a protospacer adjacent motif (PAM) sequence.

[0028] Preferably, the target sequence is a guide RNA (gRNA) target site, and the complementary strand oligonucleotide homologous to the target sequence is a guide RNA hybridizing to a target sequence located within intron 12, intron 13, or intron 14, or its complementary strand. Thus, the oligonucleotide directs the nuclease to cleave within intron 12, 13, or 14 of the albumin gene. Preferably, the guide RNA is adjacent to a protospacer adjacent motif (PAM) sequence.

[0029] Preferably, the target sequence is a guide RNA (gRNA) target site, and the complementary strand oligonucleotide homologous to the target sequence is a guide RNA hybridizing to a target sequence located within intron 9, intron 11, intron 12, intron 13, or intron 14, or its complementary strand. Thus, the oligonucleotide directs the nuclease to cleave within intron 9, 11, 12, 13, or 14 of the albumin gene. Preferably, the guide RNA is adjacent to a protospacer adjacent motif (PAM) sequence.

[0030] In the context of the present invention, the albumin gene is preferably a human or mouse gene.

[0031] Preferably, the target sequence comprises or essentially has a sequence having at least 95% identity to any one of SEQ ID NOs: 1 to 2, 9 to 18, 54, 92 to 98, or a fragment thereof, and / or the complementary strand oligonucleotide homologous to the target sequence comprises or essentially has a sequence having at least 95% identity to any one of SEQ ID NOs: 1 to 2, 9 to 18, 54, 92 to 98, or a fragment thereof.

[0032] In this embodiment, a guide RNA comprising or having at least 95% identity to any one of SEQ ID NO:2, SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:14, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:54, 92-98, or a fragment thereof, is capable of binding to a target sequence comprising or having at least 95% identity to SEQ ID NO:1, SEQ ID NO:9, SEQ ID NO:11, SEQ ID NO:13, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17 or SEQ ID NO:18, or a fragment thereof, or its complementary strand.

[0033] In this embodiment, the guide RNA comprising or having at least 95% identity to SEQ ID NO:2, SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:14, SEQ ID NO:16, SEQ ID NO:17 or 18 or a fragment thereof is capable of binding to a target sequence comprising or having at least 95% identity to SEQ ID NO:1, SEQ ID NO:9, SEQ ID NO:11, SEQ ID NO:13 or SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17 or 18 or a fragment thereof, or its complementary strand.

[0034] Preferably, the fragment is at least 15 nucleotides in length.

[0035] In one embodiment, the target sequence and each gRNA are as described in Tables 1 and 3.

[0036] The present invention also includes embodiments in which the above-mentioned sequences, i.e., SEQ ID NOs: 1, 2, 9-18, 54, 92-98, have reversed sequences, i.e., 3' to 5' orientation.

[0037] The present invention also includes embodiments in which the above-referenced sequences, i.e., SEQ ID NOs: 1, 2, 9-18, 54, 92-98, are RNA.

[0038] Said gRNA is also the subject of the present invention.

[0039] A further object of the present invention is an isolated guide ribonucleic acid (gRNA) comprising or consisting of a sequence selected from SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17 or SEQ ID NO:18, 54, 92-98, and a sequence substantially complementary to or fully annealing to at least a 15 nucleotide long portion thereof. Preferably, the albumin gene is of human or mouse origin.

[0040] The albumin gene introns may be present in the albumin gene of any origin, but preferably they are present in the human or mouse albumin gene.

[0041] In one embodiment, the complementary strand oligonucleotide homologous to the target sequence is under the control of a promoter, preferably under the control of a U6 promoter.

[0042] In the present invention, guide RNA or gRNA can be used as a synonym for a complementary strand oligonucleotide that is homologous to a target sequence.

[0043] In a preferred embodiment, the foreign DNA sequence is a coding sequence of the arylsulfatase B (ARSB) gene. Preferably, the ARSB coding sequence is of human origin. Preferably, it comprises or essentially has a sequence having at least 95% identity to SEQ ID NO: 33. The coding sequence can encode variants of arylsulfatase B (ARSB), for example, it can contain additions, deletions, or substitutions to the coding sequence of the wild-type arylsulfatase B (ARSB) gene, so long as these protein variants retain substantially the same relevant functional activity as the original ARSB. The coding sequence can also encode fragments of arylsulfatase B (ARSB), so long as the fragments retain substantially the same relevant functional activity as the original ARSB.

[0044] Preferably, the coding sequence may be codon-optimized for expression in humans.

[0045] In a further preferred embodiment, the exogenous DNA sequence is a coding sequence for a factor VIII (F8) gene or a B-domain deleted (BDD) F8 gene. Preferably, the BDD F8 coding sequence is of human origin, and more preferably, it comprises or essentially has a sequence having at least 95% identity to SEQ ID NO: 36 or 55. Preferably, the F8 coding sequence comprises or essentially has a sequence having at least 95% identity to SEQ ID NO: 36 or 55. The coding sequence can encode a variant of BDD F8 or F8. For example, it can include additions, deletions, or substitutions to the coding sequence of the wild-type BDD F8 gene or the wild-type F8 gene, so long as these protein variants retain substantially the same relevant functional activity as the original BDD F8 or F8, respectively. The coding sequence can also encode a fragment of BDD F8 or F8, so long as the fragment retains substantially the same relevant functional activity as the original BDD F8 or F8, respectively.

[0046] In a further preferred embodiment, the exogenous DNA sequence is the coding sequence of a gene that is subject to loss-of-function mutations in patients with recessive genetic disorders, thus allowing the liver to be used as a therapeutic target.

[0047] In a further preferred embodiment, the foreign DNA sequence is the coding sequence of the F9 gene or genes mutated in alpha 1-antitrypsin (AAT) deficiency, Wilson's disease, OAT deficiency, or MPSVII.

[0048] In the context of the present invention, the inverted target sequences are located one upstream and one downstream of the donor nucleic acid, which is the DNA construct that is cleaved and integrated into the target locus. The inverted target sequence is completely identical to the sequence recognized by the guide RNA in the target genomic locus, i.e., the target sequence, but is inverted or reversed relative to the genomic sequence. Inverted or reversed means that if the target sequence has a specific 5'-3' sequence, the inverted target sequence has the same sequence but in a 3'-5' direction. Therefore, the inverted target sequence is complementary to the guide RNA but inverted. This allows for unidirectional integration, as is known in the HITI method. In other words, if the donor DNA is integrated in the opposite direction, a nuclease such as Cas9 can again recognize and cleave the target site. If integration occurs in the correct direction, the nuclease can no longer cleave the target site. The inverted target sequence is preferably an inverted sequence relative to a target sequence located at the 3' end of the albumin gene in a region selected from intron 9, intron 11, intron 12, intron 13, and intron 14. Preferably, each of the inverted target sequences is linked at its 3' end to a protospacer adjacent motif (PAM) sequence.

[0049] The foreign DNA sequence may also comprise a reporter gene, preferably selected from at least one of Discosoma Red, green fluorescent protein (GFP), red fluorescent protein (RFP), luciferase, αβ-galactosidase, and αβ-glucuronidase.

[0050] In one embodiment, the donor nucleic acid further comprises one or more of the following: - post-transcriptional regulatory elements (preferably located at the 3' end of the foreign DNA sequence); - a transcription termination sequence (preferably located at the 3' end of the post-transcriptional regulatory element or at the 3' end of the foreign DNA sequence); - a splice acceptor sequence (preferably located at the 3' end of the donor nucleic acid, e.g. linked to the albumin exon, if present); - a ribosome skipping sequence (preferably located between the foreign DNA sequence and the albumin exon).

[0051] In one embodiment, the donor nucleic acid further comprises one or more of the following: - a splice acceptor sequence (preferably located at the 5' end of the donor nucleic acid and, for example, linked to the albumin exon, if present); - a ribosome skipping sequence (preferably located between the albumin exon and the foreign DNA sequence; - post-transcriptional regulatory elements (preferably located at the 3' end of the foreign DNA sequence); - a transcription termination sequence (preferably located at the 3' end of the post-transcriptional regulatory element or at the 3' end of the foreign DNA sequence).

[0052] Preferably, the ribosome skipping sequence is T2A, P2A, E2A, F2A, preferably the T2A sequence, which allows the protein of interest to be separated from albumin when expressed in cells.

[0053] Preferably, the post-transcriptional regulatory element is the woodchuck hepatitis virus post-transcriptional regulatory element (WPRE).

[0054] Preferably, the transcription termination sequence is a polyadenylation signal sequence, preferably bovine growth hormone polyA (BGH polyA), most preferably a short synthetic polyA. As described above, the donor DNA sequence is flanked on the 5' and 3' by the same gRNA target sites that the gRNA recognizes but in an inverted position (e.g., inverted target sites).

[0055] Preferably, the target sequence comprises or essentially has a sequence having at least 95% identity with one of the sequences mentioned herein or a functional fragment thereof, and / or the complementary strand oligonucleotide homologous to the target sequence comprises or essentially has a sequence having at least 95% identity with one of the sequences mentioned herein or a functional fragment thereof.

[0056] Preferably, the inverted target sequence comprises or essentially has a sequence having at least 95% identity to SEQ ID NO: 20, 54 or 77, or SEQ ID NO: 2 or 1, or any one of SEQ ID NOs: 9-18, 92-98 or 54.

[0057] Preferably, the donor nucleic acid comprises: - an inverted target sequence with its protospacer adjacent motif (PAM) sequence; - splice acceptor sequence; - one or more albumin exons (preferably one or more exons selected from exons 13 and 14); - a ribosome skipping sequence (preferably T2A); - a foreign DNA sequence (preferably the coding sequence of the marker dsRed, the human ARSB gene or the BDD F8 gene); - a transcription termination sequence; and - a further inverted target sequence with its protospacer adjacent motif (PAM) sequence.

[0058] Preferably, the donor nucleic acid comprises: - an inverted target sequence with its protospacer adjacent motif (PAM) sequence; - splice acceptor sequence; - one or more albumin exons (preferably one or more exons selected from exon 10, exon 11, exon 12, exon 13 and exon 14); - a ribosome skipping sequence (preferably T2A); - a foreign DNA sequence (preferably the coding sequence of the marker dsRed, the human ARSB gene or the BDD F8 gene); - a transcription termination sequence; and - a further inverted target sequence with its protospacer adjacent motif (PAM) sequence.

[0059] Preferably, the elements are in 5'-3' order as listed, although other orders may be suitable as well.

[0060] Preferably, the transcription termination sequence comprises or essentially comprises a sequence having at least 95% identity to SEQ ID NO:26, SEQ ID NO:37, SEQ ID NO:48 or SEQ ID NO:65.

[0061] Preferably, the ribosome skipping sequence comprises or essentially comprises a sequence having at least 95% identity to SEQ ID NO: 23 or 63.

[0062] Preferably, said albumin exon comprises or essentially comprises a sequence having at least 95% identity to SEQ ID NO: 22 and / or 78 and / or 79.

[0063] Preferably, the splice acceptor sequence comprises or essentially consists of a sequence having at least 95% identity to SEQ ID NO:21.

[0064] Preferably, said inverted target sequence comprises or essentially comprises a sequence having at least 95% identity to SEQ ID NO: 1 or 2 or 20 or 54 or 77, and preferably does not comprise a PAM sequence.

[0065] The nuclease can be provided as a protein or as a nucleic acid encoding the nuclease. The nucleic acid can be DNA or RNA, for example it can be an mRNA for the nuclease, or it can be a cDNA or DNA coding sequence for the nuclease, or a DNA construct encoding the nuclease.

[0066] Preferably, the nucleic acid encoding the nuclease is a DNA construct comprising a nucleic acid encoding Cas9 or spCas9, preferably under the control of a tissue-specific promoter, e.g., a liver-specific promoter such as the hybrid liver promoter (HLP). The construct may further comprise a polyA, preferably a short synthetic polyA (synt. polyA). All such elements are well known in the art and may have conventional nucleotide sequences. Exemplary DNA constructs encoding nucleases comprise or essentially have a sequence having at least 95% identity to SEQ ID NO: 43, 47, or 52.

[0067] The nuclease is preferably selected from the group consisting of CRISPR nucleases, TALENs, DNA-guided nucleases, meganucleases and zinc finger nucleases, and preferably the nuclease is a CRISPR nuclease selected from the group consisting of Cas9, Cpf1, Cas12b (C2c1), Cas13a (C2c2), Cas3, Csf1, Cas13b (C2c6) and C2c3, or variants thereof such as SaCas9, VQR-Cas9-HF1 or dcas9.

[0068] Preferably, the cell is contacted with a nucleic acid encoding said nuclease, and preferably said nuclease encoding nucleic acid is under the control of a tissue-specific promoter, such as the liver-specific hybrid liver promoter (HLP).

[0069] Preferably, the nucleic acid encoding said nuclease is under the control of a tissue-specific promoter, most preferably a liver-specific promoter, such as a hepatocyte-specific promoter, such as the liver-specific hybrid liver promoter (HLP).

[0070] Preferably, the donor nucleic acid, the complementary oligonucleotide homologous to the target sequence, and the nucleic acid encoding the nuclease are contained in a DNA construct. Preferably, a first DNA construct contains the donor nucleic acid and the complementary oligonucleotide homologous to the target sequence, and a second DNA construct contains the nucleic acid encoding the nuclease that recognizes the target sequence. Alternatively, the first DNA construct contains the donor nucleic acid, and the second DNA construct contains the complementary oligonucleotide homologous to the target sequence and the nucleic acid encoding the nuclease that recognizes the target sequence. As a further option, three constructs are provided: a first construct contains the donor nucleic acid, a second construct contains the complementary oligonucleotide homologous to the target sequence, and a third construct contains the nucleic acid encoding the nuclease that recognizes the target sequence.

[0071] Said constructs are also the subject of the present invention.

[0072] Preferably, one or more of the DNA constructs are contained in a vector, preferably a viral vector, more preferably a lentiviral vector or an adeno-associated viral vector. Alternatively, all or part of the DNA constructs may be inserted into a non-viral vector, which is selected from polymer-based, particle-based, lipid-based, peptide-based delivery vehicles, such as cationic polymers, micelles, liposomes, exosomes, microparticles, and nanoparticles, including lipid nanoparticles (LNPs), or combinations thereof.

[0073] Said vectors are also the subject of the present invention.

[0074] The complementary oligonucleotide, the donor nucleic acid, and the nucleic acid encoding the nuclease can be contained in one or more viral or non-viral vectors, and preferably the viral vector is selected from adeno-associated virus, lentivirus, retrovirus, and adenovirus. This means that they can be in the same vector or different vectors. Preferably, the cells are selected from the group consisting of hepatocytes, one or more lymphocytes, monocytes, neutrophils, eosinophils, basophils, endothelial cells, epithelial cells, liver cells, bone cells, platelets, adipocytes, cardiac myocytes, neurons, retinal cells, smooth muscle cells, skeletal muscle cells, spermatocytes, oocytes, and pancreatic cells, induced pluripotent stem cells (iPS cells), stem cells, hematopoietic stem cells, and hematopoietic progenitor cells, and preferably the cells are liver cells of a subject. Another subject of the present invention is a cell obtainable by the method defined above.

[0075] The cells can be used as a medicine or for the treatment of genetic diseases or recessive genetic diseases and common diseases (preferably, the diseases include diabetes, lysosomal storage diseases including mucopolysaccharidoses (MPS I, MPS II, MPS IIIA, MPS IIIB, MPS IIIC, MPS I V A, MPS I V B, MPS V I, MPS V II), sphingolipidoses (Fabry disease, Gaucher disease, Niemann-Pick disease, GM1 gangliosidosis), lipofuscinoses (Batten disease and others) and mucolipidoses; gyriform chorioretinal atrophy, adenylosuccinic acid deficiency, hemophilia A and B, ALA dehydrogenase deficiency, adrenoleukodystrophy).

[0076] The cells of the invention can be used to treat diseases in which both the wild-type and mutant alleles are replaced with the correct copy of the gene provided by the donor DNA, or recessive genetic diseases and common diseases due to loss of function (preferably, said diseases include hemophilia, diabetes, lysosomal storage diseases including mucopolysaccharidoses, e.g., MPS I, MPS II, MPS IIIA, MPS IIIB, MPS IIIC, MPS I V A, MPS I V B, MPS V I and MPS V II, sphingolipidoses, e.g., Fabry disease, Gaucher disease, Niemann-Pick disease and GM1 gangliosidosis, lipofuscinoses, e.g., Batten disease, and mucolipidoses; gyriform chorioretinal atrophy, adenylosuccinic acid deficiency, hemophilia A and B, ALA dehydrogenase deficiency, adrenoleukodystrophy).

[0077] Thus, the cells of the invention can be used to treat hemophilia, diabetes, lysosomal storage diseases such as mucopolysaccharidoses, e.g., MPS1, MPSII, MPSIIIA, MPSIIIB, MPSIIIC, MPSIVA, MPSIVB, MPSVI and MPSVII, sphingolipidoses, e.g., Fabry disease, Gaucher disease, Niemann-Pick disease and GM1 gangliosidosis, lipofuscinoses, e.g., Batten disease, and mucolipidoses; gyriform chorioretinal atrophy, adenylosuccinic acid deficiency, hemophilia A and B, ALA dehydrogenase deficiency or adrenoleukodystrophy.

[0078] The cells obtainable by the present invention express the foreign DNA sequence, and preferably also express the full-length albumin coding sequence.

[0079] Once the donor nucleic acid contacts the cell, it is inserted into the target gene, usually via non-homologous end joining.

[0080] A further object of the present invention is a system comprising: a) a donor nucleic acid comprising: - foreign DNA sequences; - optionally, one or more albumin exons; wherein the donor nucleic acid is flanked at 5' and 3' by inverted target sequences, b) a complementary oligonucleotide homologous to the target sequence, and c) a nuclease that recognizes the target sequence wherein the target sequence is located at the 3' end of the albumin gene in a region selected from intron 9, intron 11, intron 12, intron 13 and intron 14.

[0081] A further object of the present invention is a system comprising: a) a donor nucleic acid comprising: - foreign DNA sequences; - optionally, one or more albumin exons; wherein the donor nucleic acid is flanked at 5' and 3' by inverted target sequences, b) a complementary oligonucleotide homologous to the target sequence, and c) a nuclease that recognizes the target sequence (wherein the target sequence is located at the 3' end of the albumin gene in a region selected from intron 12, intron 13 and intron 14).

[0082] The nuclease can be provided as a protein or as a nucleic acid encoding the nuclease. The nucleic acid can be DNA or RNA, for example it can be an mRNA for the nuclease, or it can be a cDNA or DNA coding sequence for the nuclease, or a DNA construct encoding the nuclease.

[0083] In one embodiment, the nuclease is encoded by a nucleic acid, and the donor nucleic acid, the complementary strand oligonucleotide, and the nucleic acid encoding the nuclease are located on a DNA construct, preferably the donor nucleic acid and the complementary strand oligonucleotide are located on the same DNA construct, while the nucleic acid encoding the nuclease is located on another DNA construct. Preferably, the construct comprising the donor nucleic acid and the complementary strand oligonucleotide comprises or essentially comprises a sequence having at least 95% identity to a sequence comprising the following sequences: SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:33, SEQ ID NO:26, SEQ ID NO:20, SEQ ID NO:27, SEQ ID NO:1, and SEQ ID NO:28.

[0084] In another embodiment, the construct comprising the donor nucleic acid comprises or essentially comprises a sequence having at least 95% identity to a sequence comprising the following sequences: SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:20.

[0085] A construct comprising the donor nucleic acid and the complementary strand oligonucleotide may comprise or essentially comprise a sequence having at least 95% identity to SEQ ID NO:34.

[0086] A construct comprising the donor nucleic acid can comprise or essentially comprise a sequence having at least 95% identity to SEQ ID NO:38.

[0087] In a preferred embodiment, the donor nucleic acid and / or foreign DNA sequence and / or albumin exon and / or inverted target sequence and / or target sequence and / or complementary strand oligonucleotide and / or nuclease and / or intron are as defined above or as defined herein.

[0088] Another object of the present invention is a process for preparing viral vector particles, which comprises introducing such a DNA construct into a host cell.Furthermore, the obtaining of viral vector particles is also an object of the present invention.

[0089] In a preferred embodiment, the donor nucleic acid and / or foreign DNA sequence and / or target sequence and / or complementary strand oligonucleotide and / or nuclease are as defined above.

[0090] In a preferred embodiment, the donor nucleic acid and / or foreign DNA sequence and / or albumin exon and / or inverted target sequence and / or target sequence and / or complementary strand oligonucleotide and / or nuclease and / or intron are as defined above or herein.

[0091] In the context of the present invention, the donor nucleic acid and / or the nucleic acid encoding the complementary strand oligonucleotide and / or nuclease are comprised in one or more viral or non-viral vectors, preferably said viral vectors are selected from adeno-associated viruses, retroviruses, adenoviruses and lentiviruses; said non-viral vectors are preferably selected from polymer-based, particle-based, lipid-based, peptide-based delivery vehicles or combinations thereof, such as cationic polymers, micelles, liposomes, exosomes, microparticles and nanoparticles, including lipid nanoparticles (LNPs).

[0092] Preferably, the first vector comprises a donor nucleic acid and a complementary oligonucleotide homologous to the target sequence, and the second vector comprises a nucleic acid encoding a nuclease that recognizes the target sequence. Alternatively, the first vector comprises a donor nucleic acid, and the second vector comprises a complementary oligonucleotide homologous to the target sequence and a nucleic acid encoding a nuclease that recognizes the target sequence. As a further option, three vectors are provided: the first vector comprises a donor nucleic acid, the second vector comprises a complementary oligonucleotide homologous to the target sequence, and the third vector comprises a nucleic acid encoding a nuclease that recognizes the target sequence.

[0093] Preferably, the system comprises a first vector comprising a nucleic acid that expresses a nuclease, and a second vector comprising a donor nucleic acid and a complementary oligonucleotide homologous to a target sequence, these elements being as defined above or herein.

[0094] Preferably, the system comprises a first vector comprising a donor nucleic acid and a second vector comprising a complementary oligonucleotide homologous to the target sequence and a nucleic acid encoding a nuclease, these elements being as defined above or herein. The system according to the invention is preferably for medical use, preferably for use in the treatment of genetic diseases or loss-of-function genetic and common diseases (preferably, said diseases include diabetes, lysosomal storage diseases including mucopolysaccharidoses (MPS I, MPS II, MPS IIIA, MPS IIIB, MPS IIIC, MPS I V A, MPS I V B, MPS V I, MPS V II), sphingolipidoses (Fabry disease, Gaucher disease, Niemann-Pick disease, GM1 gangliosidosis), lipofuscinoses (Batten disease and others) and mucolipidoses; gyrus chorioretinal atrophy, adenylosuccinic acid deficiency, hemophilia A and B, ALA dehydrogenase deficiency, adrenoleukodystrophy).

[0095] The system of the present invention can be used to treat diseases in which both the wild-type and mutant alleles are replaced with the correct copy of the gene provided by the donor DNA, or recessive genetic diseases and common diseases due to loss of function (preferably, said diseases include hemophilia, diabetes mellitus, lysosomal storage diseases including mucopolysaccharidoses, e.g., MPS I, MPS II, MPS IIIA, MPS IIIB, MPS IIIC, MPS I V A, MPS I V B, MPS V I and MPS V II, sphingolipidoses, e.g., Fabry disease, Gaucher disease, Niemann-Pick disease and GM1 gangliosidosis, lipofuscinoses, e.g., Batten disease, and mucolipidoses; gyriform chorioretinal atrophy, adenylosuccinic acid deficiency, hemophilia A and B, ALA dehydrogenase deficiency, adrenoleukodystrophy).

[0096] Thus, the system of the present invention can be used to treat hemophilia, diabetes, lysosomal storage diseases such as mucopolysaccharidoses, e.g., MPS1, MPSII, MPSIIIA, MPSIIIB, MPSIIIC, MPSIVA, MPSIVB, MPSVI and MPSVII, sphingolipidoses, e.g., Fabry disease, Gaucher disease, Niemann-Pick disease and GM1 gangliosidosis, lipofuscinoses, e.g., Batten disease, and mucolipidoses; gyriform chorioretinal atrophy, adenylosuccinic acid deficiency, hemophilia A and B, ALA dehydrogenase deficiency or adrenoleukodystrophy.

[0097] A further subject of the present invention is a vector comprising a donor nucleic acid as defined above or herein and / or a complementary strand oligonucleotide homologous to the target sequence and / or a nucleic acid encoding a nuclease that recognizes the target sequence.

[0098] In the context of the present invention, preferably the donor nucleic acid and / or foreign DNA sequence and / or albumin exon and / or inverted target sequence and / or target sequence and / or complementary strand oligonucleotide and / or nuclease and / or intron are as defined above or as defined herein.

[0099] Preferably, the vector is a viral vector, preferably a lentiviral vector or an adeno-associated viral vector, or a non-viral vector, preferably selected from polymer-based, particle-based, lipid-based, peptide-based delivery vehicles, or combinations thereof, such as cationic polymers, micelles, liposomes, exosomes, nanoparticles, including microparticles and lipid nanoparticles (LNPs). Preferably, an AAV2 / 8 vector is used.

[0100] In a preferred embodiment of the invention, one vector containing a nucleic acid expressing Cas9, as defined above, is used in conjunction with a second vector containing donor DNA and a complementary oligonucleotide homologous to the target sequence, i.e., gRNA.

[0101] In another preferred embodiment of the invention, one vector containing a nucleic acid expressing Cas9 and a complementary oligonucleotide homologous to the target sequence, i.e., gRNA, as defined above, is used along with a second vector containing donor DNA.

[0102] Preferably, the first vector contains a nucleic acid encoding Cas9 or spCas9, preferably under the control of a tissue-specific promoter, e.g., a liver-specific promoter such as the hybrid liver promoter (HLP). The vector may further contain polyA, preferably a short synthetic polyA (synt. polyA). Preferably, the second vector contains the gRNA expression cassette defined above and donor DNA. Preferably, the gRNA expression cassette contains the gRNA defined above under the control of a U6 promoter. Preferably, the donor DNA is flanked at its 3' and 5' by inverted targeting sequences, preferably linked to each PAM.

[0103] Preferably, the vector is a viral vector, and preferably, the viral vector is a lentiviral vector, an adeno-associated viral vector, an adenoviral vector, a retroviral vector, a polioviral vector, a murine Moloney-based viral vector, an alphaviral vector, a poxviral vector, a herpesviral vector, a vaccinia viral vector, a baculoviral vector, or a parvoviral vector, and preferably, the adeno-associated virus is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAVSH19, AAVPHP.B, AAV2, AAV9, AAV1, AAVSH19, AAVPHP.B, AAV8, AAV6.

[0104] Preferably, the viral vector or vector further comprises a 5' terminal repeat (5'-TR) nucleotide sequence and a 3' terminal repeat (3'-TR) nucleotide sequence, and preferably, the 5'-TR is a 5' inverted terminal repeat (5'-ITR) nucleotide sequence, and the 3'-TR is a 3' inverted terminal repeat (3'-ITR) nucleotide sequence, and preferably, the ITRs are derived from the same viral serotype or different viral serotypes, and preferably, the virus is AAV, and preferably, the serotype is 2.

[0105] In one embodiment, the viral vector comprising the gRNA expression cassette and the donor DNA further comprises a 5' inverted terminal repeat (ITR) sequence (preferably located at the 5' end of the construct comprising the gRNA expression cassette and the donor DNA, preferably of AAV) and a 3' inverted terminal repeat (ITR) sequence (preferably located at the 3' end of the construct comprising the gRNA expression cassette and the donor DNA, preferably of AAV).

[0106] Preferably, the ITR comprises or has a sequence having at least 95% identity with SEQ ID NO: 110, SEQ ID NO: 29 or 66.

[0107] Preferably, the viral vector comprising the gRNA expression cassette and the donor DNA comprises the following: - AAV 5' inverted terminal repeat (5'-ITR) sequence; - An inverted target sequence linked to its protospacer adjacent motif (PAM) sequence; - Splice acceptor sequence; - One or more albumin exons (preferably one or more exons selected from exon 10, exon 11, exon 12, exon 13 and exon 14); - Ribosome skipping sequence (preferably T2A); - Foreign DNA sequence (preferably the coding sequence of the human ARSB gene); - Transcription termination sequence; - A further inverted target sequence linked to its protospacer adjacent motif (PAM) sequence; - A complementary strand oligonucleotide homologous to the target sequence, under the control of a promoter (preferably the U6 promoter); - Chimeric gRNA scaffold; and - AAV 3' inverted terminal repeat (3'-ITR) sequence.

[0108] Preferably, the viral vector containing the gRNA expression cassette and donor DNA comprises the following: - AAV 5' inverted terminal repeat (5'-ITR) sequence; - An inverted target sequence linked to its protospacer adjacent motif (PAM) sequence; - Splice acceptor sequence; - One or more albumin exons (preferably one or more exons selected from exon 13 and 14); - Ribosome skipping sequence (preferably T2A); - Foreign DNA sequence (preferably the coding sequence of the human ARSB gene); - Transcription termination sequence; - A further inverted target sequence linked to its protospacer adjacent motif (PAM) sequence; - A complementary strand oligonucleotide homologous to the target sequence, under the control of a promoter (preferably the U6 promoter); - a chimeric gRNA scaffold; and - AAV 3' inverted terminal repeat (3'-ITR) sequence.

[0109] Alternatively, the viral vector comprising the donor DNA comprises: - AAV 5' inverted terminal repeat (5'-ITR) sequence; - an inverted target sequence linked to its protospacer adjacent motif (PAM) sequence; - splice acceptor sequence; - one or more albumin exons (preferably one or more exons selected from exon 10, exon 11, exon 12, exon 13 and exon 14); - a ribosome skipping sequence (preferably T2A); - a foreign DNA sequence (preferably the coding sequence of the human BDD F8 gene); - transcription termination sequence; - an additional inversion target sequence linked to its protospacer adjacent motif (PAM) sequence; - AAV 3' inverted terminal repeat (3'-ITR) sequence.

[0110] Alternatively, the viral vector comprising the donor DNA comprises: - AAV 5' inverted terminal repeat (5'-ITR) sequence; - an inverted target sequence linked to its protospacer adjacent motif (PAM) sequence; - splice acceptor sequence; - one or more albumin exons (preferably one or more exons selected from exons 13 and 14); - a ribosome skipping sequence (preferably T2A); - a foreign DNA sequence (preferably the coding sequence of the human BDD F8 gene); - transcription termination sequence; - an additional inversion target sequence linked to its protospacer adjacent motif (PAM) sequence; - AAV 3' inverted terminal repeat (3'-ITR) sequence.

[0111] Preferably, when the vector does not contain a gRNA expression cassette, the cassette is contained in a vector containing a nucleic acid that expresses a nuclease.

[0112] The above elements may be in the 5' to 3' order defined above, but as will be understood by those skilled in the art, other orders are equally suitable.

[0113] The vector may further contain additional viral sequences such as additional AAV sequences.

[0114] Preferably, the vector containing the donor nucleic acid and the complementary strand oligonucleotide contains, or consists essentially of, a sequence having at least 95% identity with a sequence containing the following sequences: SEQ ID NO: 29, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 33, SEQ ID NO: 26, SEQ ID NO: 20, SEQ ID NO: 27, SEQ ID NO: 2, SEQ ID NO: 28, and SEQ ID NO: 29.

[0115] In another embodiment, the vector containing the donor nucleic acid contains, or consists essentially of, a sequence having at least 95% identity with a sequence containing the following sequences: SEQ ID NO: 110, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 20, and SEQ ID NO: 29.

[0116] Another object of the present invention is a host cell containing the construct, vector, vector system, or system defined as above.

[0117] Another object of the present invention is a viral particle containing the construct, vector, vector system, or system defined as above.

[0118] Preferably, the viral vector defined herein includes viral vector particles.

[0119] The term "virus particle" or "viral particle" is intended to mean the extracellular form of a non-pathogenic virus, particularly a viral vector, in which genetic material made of either DNA or RNA is enclosed in a protein coat called a capsid, and in some cases an envelope derived from a portion of the host cell membrane that contains viral glycoproteins.

[0120] As used herein, viral vector also refers to a viral vector particle.

[0121] The viral vectors encompassed by the present invention are suitable for gene therapy.

[0122] Preferably, the viral particle comprises the capsid protein of AAV.

[0123] Preferably, the viral particle comprises capsid proteins of one or more AAV serotypes selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAVSH19, and AAVPHP.B, and preferably comprises capsid proteins of AAV of the AAV2 or AAV8 serotype.

[0124] Another subject of the present invention is a pharmaceutical composition comprising one of the following: a system as defined above, one or more vectors, host cells or viral particles, and a pharmaceutically acceptable carrier.

[0125] Another subject of the present invention is a kit comprising: a DNA construct, a system or one or more vectors or host cells or viral particles as defined above, or a pharmaceutical composition as defined above, in one or more containers, and optionally further comprising instructions or packaging material describing how to administer the nucleic acid construct, vector, host cell, viral particle or pharmaceutical composition to a patient.

[0126] The system or one or more vectors or host cells or viral particles as defined above, or the pharmaceutical composition as defined above, is preferably for use as a medicament, preferably for use in the treatment of the diseases mentioned herein, preferably liver diseases, lysosomal storage diseases including mucopolysaccharidoses (e.g. MPS I, MPS II, MPS IIIA, MPS IIIB, MPS IIIC, MPS I V A, MPS I V B, MPS V I, MPS V I), sphingolipidoses (Fabry disease, Gaucher disease, Niemann-Pick disease, GM1 gangliosidosis), lipofuscinoses (Batten disease and others) and mucolipidoses; other diseases in which the liver can be used as a factory for the production and secretion of therapeutic proteins, e.g. diabetes, gyriform chorioretinal atrophy, adenylosuccinic acid deficiency, hemophilia A and B, ALA dehydrogenase deficiency, adrenoleukodystrophy.

[0127] The system or one or more vectors or host cells or viral particles as defined above, or the pharmaceutical composition as defined above, can be used for the treatment of diseases in which both the wild type and mutant alleles are replaced with the correct gene copy provided by donor DNA, or recessive loss-of-function diseases and common diseases (preferably, said diseases include hemophilia, diabetes, lysosomal storage diseases including mucopolysaccharidoses, e.g. MPS I, MPS II, MPS IIIA, MPS IIIB, MPS IIIC, MPS I V A, MPS I V B, MPS I V and MPS V II, sphingolipidoses, e.g. Fabry disease, Gaucher disease, Niemann-Pick disease and GM1 gangliosidosis, lipofuscinoses, e.g. Batten disease, and mucolipidoses; gyriform chorioretinal atrophy, adenylosuccinic acid deficiency, hemophilia A and B, ALA dehydrogenase deficiency, adrenoleukodystrophy).

[0128] Thus, the system or one or more vectors or host cells or viral particles as defined above, or the pharmaceutical composition as defined above, can be used for the treatment of hemophilia, diabetes, lysosomal storage diseases such as mucopolysaccharidoses, e.g. MPS I, MPS II, MPS IIIA, MPS IIIB, MPS IIIC, MPS I V A, MPS I V B, MPS V I and MPS V II, sphingolipidoses, e.g. Fabry disease, Gaucher disease, Niemann-Pick disease and GM1 gangliosidosis, lipofuscinoses, e.g. Batten disease, and mucolipidoses; gyriform chorioretinal atrophy, adenylosuccinic acid deficiency, hemophilia A and B, ALA dehydrogenase deficiency or adrenoleukodystrophy.

[0129] A further subject of the present invention is a construct as defined above for the production of viral particles.

[0130] An object of the present invention is also a method for treating a subject suffering from a disease mentioned herein, preferably a genetic disease caused by loss of gene function, comprising administering to the subject an effective amount of a vector system or vector or host cell or viral particle or pharmaceutical composition as defined above. Preferably, said disease is a lysosomal storage disease such as mucopolysaccharidoses (MPS I, MPS II, MPS IIIA, MPS IIIB, MPS IIIC, MPS IVA, MPS I VB, MPS V I, MPS V II) or hemophilia A or B.

[0131] Preferably, the sequences mentioned herein are the subject of the present invention.

[0132] Preferably, the donor DNA cassette element and / or gRNA expression cassette element and / or promoter sequence and / or U6 promoter for gRNA expression and / or gRNA and / or gRNA target site and / or inverted target sequence and / or Cas9 and / or foreign DNA sequence and / or post-transcriptional regulatory element and / or transcription termination sequence and / or splice acceptor sequence and / or ribosomal skipping sequence are the sequences shown in SEQ ID NOs: 1 to 109 below.

[0133] Another subject of the present invention is a DNA construct comprising a donor nucleic acid as defined above or as defined herein and / or a complementary strand oligonucleotide homologous to a target sequence and / or a nucleic acid encoding a nuclease recognizing said target sequence.

[0134] In one embodiment, the method of the invention is ex vivo or in vitro.

[0135] In one embodiment, in the methods of the invention, the cells are isolated cells from a subject or patient.

[0136] The sequence of albumin is preferably described under the following accession number AC140220.4 (GeneBank, NCBI, database; latest version) or under the following accession number NC_000004.12.

[0137] The present invention also provides a pharmaceutical composition comprising a nucleic acid as defined above or a nucleotide sequence as defined above or a vector as defined above, and a pharmaceutically acceptable diluent and / or excipient and / or carrier.

[0138] Preferably, the composition further comprises a therapeutic agent, preferably said therapeutic agent is selected from the group consisting of enzyme replacement therapy and small molecule therapy.

[0139] Preferably, the pharmaceutical composition is administered via a route selected from the group consisting of parenteral, intravenous (e.g., via the temporal vein), intraperitoneal, intratumoral, intrahepatic, or any combination thereof. Preferably, the vector of the present invention is administered via the intravenous or parenteral route.

[0140] The present invention also provides a medical use of a vector as defined above, wherein said vector is administered via a route selected from the group consisting of parenteral, intravenous (e.g., via the temporal vein), intraperitoneal, intratumoral, intrahepatic, or any combination thereof. Preferably, the vector of the present invention is administered via the intravenous or parenteral route.

[0141] In a preferred embodiment of the present invention, - if the target sequence is located within intron 9 of the albumin gene, albumin exons 10, 11, 12, 13 and 14 or fragments thereof are present; - if the target sequence is located within intron 11 of the albumin gene, albumin exons 12, 13 and 14 or fragments thereof are present; - if the target sequence is located within intron 12 of the albumin gene, albumin exons 13 and 14 or fragments thereof are present; - if the target sequence is located within intron 13 of the albumin gene, albumin exon 14 or a fragment thereof is present; If the target sequence is located within intron 14 of the albumin gene, then no albumin exon or fragment thereof is present.

[0142] In the present invention, a 3XFLAG sequence may be present, and preferably it comprises or essentially comprises a sequence having at least 95% identity to SEQ ID NO: 56 or 62.

[0143] In the present invention, "at least 80% identity" means that there may be at least 80%, 85%, 90%, 95%, or 100% sequence identity to the referenced sequence. This applies to all stated % identities. In the present invention, "at least 95% identity" means that there may be at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the referenced sequence. This applies to all stated % identities. In the present invention, "at least 98% identity" means that there may be at least 98%, 99%, or 100% sequence identity to the referenced sequence. This applies to all stated % identities. Preferably, the % identity is relative to the entire length of the referenced sequence.

[0144] The present invention also includes nucleic acid sequences derived from the nucleotide sequences referred to herein, such as functional fragments, mutants, variants, derivatives, analogs, and sequences having at least 80% identity to the sequences referred to herein. DETAILED DESCRIPTION OF THE INVENTION

[0145] <Definition> As used herein, the terms "comprising," "comprises," and "comprised of" are synonymous with "including," "includes," "containing," or "contains," and are inclusive or open-ended and do not exclude additional, unrecited members, elements, or steps. The terms "comprising," "comprises," and "comprised of" also include the term "consisting of."

[0146] In the present invention, "at least 80% identity" means that there may be at least 80%, 85%, 90%, 95%, or 100% sequence identity to the referenced sequence. This applies to all stated % identities. In the present invention, "at least 95% identity" means that there may be at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to the referenced sequence. This applies to all stated % identities. In the present invention, "at least 98% identity" means that there may be at least 98%, 99%, or 100% sequence identity to the referenced sequence. This applies to all stated % identities. Preferably, the % identity is relative to the entire length of the referenced sequence.

[0147] The present invention also includes nucleic acid sequences derived from the nucleotide sequences referred to herein, such as functional fragments, mutants, variants, derivatives, analogs, and sequences having at least 80% identity to the sequences referred to herein, so long as such fragments, mutants, variants, derivatives and analogs maintain the function of the sequences from which they are derived. [Brief explanation of the drawings]

[0148] [Figure 1]In vivo integration and expression of the DsRed transgene at the 3' mAlb locus. Wild-type mice received a mixture of AAV8-SpCas9 and AAV8-donor gRNA (or scRNA as a negative control) via the temporal vein at postnatal day 1 (p1). (A) Schematic of the HITI construct and integration at the mouse 3' albumin locus. (B) Representative indel analysis by T7 endonuclease (T7) cleavage assay. Expected band sizes and average indel frequencies are shown (n = 5). (C) PCR of DNA extracted from liver samples showing precise 5' (left panel) and 3' (right panel) ligation products at the 3' mAlb locus using specific primer combinations. Black arrows indicate expected band sizes. (D) Representative fluorescent microscopy images at 20x magnification of liver cryosections from n = 5 gRNA-treated and n = 5 scRNA-treated animals. gRNA = HITI donor-gRNA + SpCas9; scRNA = HITI donor-scRNA + SpCas9. In the right panel, the percentage of Ds-Red positive hepatocytes is reported. [Figure 2] Integration into the 3' end of the albumin gene after neonatal AAV-HITI administration ameliorates the phenotype of the MPSVI mouse model. (A) Schematic diagram of the AAV-gRNA-HITI donor and AAV-Cas9 constructs. SAS: synthetic splicing acceptor signal; Exon14: exon 14 of mouse albumin; T2A: Thosea asigna virus 2A skipping peptide; spA: synthetic bovine growth hormone polyA (B). Serum arylsulfatase B (ARSB) activity measured in normal (NR), untreated MPS VI mice (AF NT), and gRNA-treated MPS VI mice (AF gRNA) is reported. Values ​​are reported on a logarithmic scale. (C) Urinary glycosaminoglycans (GAGs) were measured in normal (NR), untreated MPS VI mice (AF NT), and gRNA-treated MPS VI mice (AF gRNA). Values ​​are reported as a percentage of age-matched scrambled controls. [Figure 3]HITI-mediated integration of F8 codopV3 into the 3' end of the mouse Alb (mAlb) locus in neonatal hemophilic mice. (A) Schematic diagram of the construct. U6 = U6 promoter; gRNA = gRNA expression cassette; HLP = hybrid liver promoter; Cas9 = SpCas9 coding sequence; pA = polyadenylation signal; SAS = synthetic splicing acceptor; Ex14 = mouse albumin exon 14; T2A = Thosea asigna virus 2A skipping peptide; F8 = CodopV3 coding sequence; pA = polyadenylation signal. (B) AF = untreated hemophilia-affected mouse; NR = unaffected control; HITI gRNA = affected animal treated with Cas9 + U6 gRNA expression cassette and HITI donor; HITI scRNA = chromogenic assay performed on plasma samples from affected animal treated with Cas9 + U6 scRNA expression cassette and HITI donor. Each dot corresponds to a different animal. F8 activity is reported in international units per deciliter (IU / dl). [Figure 4] Serum albumin levels. Serum albumin levels were measured by p360 after treatment in all animals treated or untreated with AAV-HITI. AFgRNA = diseased animals treated with AAV-HITI guide (gRNA) vector; AFscRNA = diseased animals treated with AAV-HITI scrambled RNA (scRNA) vector; NR = non-diseased, untreated animals. Each bar corresponds to the albumin level of a single animal. Serum albumin is expressed as (mg albumin) / (ml serum). [Figure 5] Mouse alpha-fetoprotein (AFP) levels. AFP levels were measured in serum samples collected 360 days after treatment. AFgRNA = diseased animals treated with AAV-HITI guide (gRNA) vector; AFscRNA = diseased animals treated with AAV-HITI scrambled RNA (scRNA) vector; AF = untreated diseased animals; NR = non-diseased untreated animals. Each bar corresponds to the AFP level of a single animal. AFP levels are expressed as (ng of AFP) / (ml of serum). [Figure 6]CAST-seq analysis of AAV-HITI samples. Visual representation of CAST-seq analysis performed on genomic DNA extracted from liver samples from three different AAV-HITI gRNA-treated MPSVI mice. [Figure 7] Dose response of AAV-HITI to treat MPSVI mice. Serum active ARSB levels are shown. Treatment, genotype, and time point are indicated below the graph. Each dot represents one mouse, and the mean level is indicated within each bar (above each bar for LD treatment). Normal levels in unaffected mice expressing ARSB are indicated by the dotted line and are reported as the mean ± standard error of the mean (from Alliegro & Ferla et al., 2016). NT = untreated; HD = high dose; MD = medium dose; LD = low dose; AF = affected or Arsb- / - mice; p30 = postnatal day 30. [Figure 8] Assessment of INDELs at the 3'ALB locus. Representative indel analysis by T7 endonuclease (T7) cleavage assay. Expected band sizes are indicated by black arrows, and the % of indels is reported below each gRNA lane and is shown as the mean ± standard error of the mean (n = 3 independent experiments). Molecular weight markers are 100 bp markers. scRNA = scrambled RNA; + = T7 enzyme-treated sample; - = T7 enzyme-untreated sample; SEM = standard error of the mean. [Figure 9]HITI-mediated integration into the 3'Alb or 3'ALB locus in vitro. Quantification of DsRed-positive (DsRed+) cells resulting from integration of donor DNA containing a promoterless DsRed coding sequence into the 3'Alb or 3'ALB locus. The number of DsRed+ cells resulting from gRNA-induced integration was normalized to samples that received scrambled RNA (scRNA) and is reported as the percentage of EGFP-positive cells ligated to Cas9. The cell line, gRNA ID, and target intron of Alb or ALB are reported below the graph. Each dot represents a biological replicate of transfected cells. scRNA = scrambled RNA; HEPA 1-6 = mouse hepatoma cell line 1-6; HUH7 = human hepatoma cell line 7; Alb = mouse albumin; intr = intron; ALB = human albumin. [Figure 10A] Molecular characterization of AAV-HITI at on-target and off-target sites in mice. A) Indel rate (%) at on-target sites determined by Illumina-seq NGS analysis. AAV-HITI gRNA-treated mice showed 29% indels at on-target sites, whereas AAV-HITI scRNA-treated mice had a near-zero indel rate. [Figure 10B] B) Illumina-seq-generated reads were aligned to the on-target site to detect potential AAV genome integrations. Reads containing ITR sequences were enriched when either the AAV-HITI or AAV-Cas9 genome was used as a reference. [Figure 10C] C) NGS off-target analysis performed on the top 10 predicted off-target sites in DNA samples derived from AAV-HITI gRNA and AAV-HITI scRNA mice showed that our selected gRNA was specific for the on-target site (mouse albumin intron 13). [Figure 10D]

[0149] <Albumin gene> The albumin gene is the target genomic locus recognized by the gRNA of the present invention for inserting a foreign DNA sequence to be expressed under the control of the albumin promoter.

[0150] The sequence of albumin is preferably set forth under the following accession number AC140220.4 or under the following accession number NC_000004.12.

[0151] The albumin gene (ENSMUSG00000029368) is located on chromosome 5 and has three alternative transcript variants, only one of which (ENSMUST00000031314.10, containing 15 exons) encodes the albumin protein (P07724, 608 aa).

[0152] Albumin protein is abundant in plasma and functions as a carrier protein for various molecules, such as steroids and fatty acids, in the blood, and is essential for maintaining oncotic pressure. This gene is primarily expressed in the liver, and the encoded protein undergoes proteolytic processing before being secreted into plasma. [Provided by RefSeq; Oct 2015] Therapeutic genes and proteins Therapeutic genes of the present invention are genes responsible for one or more genetic disorders, such as lysosomal storage diseases including mucopolysaccharidoses (MPS I, MPS II, MPS IIIA, MPS IIIB, MPS IIIC, MPS I V A, MPS I V B, MPS V II), sphingolipidoses (Fabry disease, Gaucher disease, Niemann-Pick disease, GM1 gangliosidosis), lipofuscinoses (Batten disease and others) and mucolipidoses; gyriform chorioretinal atrophy, diabetes mellitus, adenylosuccinic acid deficiency, hemophilia A and B, ALA dehydrogenase deficiency, adrenoleukodystrophy.

[0153] Particularly preferred therapeutic genes of the present invention are genes that can be expressed by hepatocytes to correct defects in the same or other tissues.

[0154] Advantageously, according to the present invention, the liver can be used as a factory for the production and secretion of therapeutic proteins to correct genetic defects in the liver or in different tissues.

[0155] The therapeutic genes of the present invention are also genes that exhibit loss of function in recessive diseases (autosomal or sex-linked).

[0156] <Factor VIII> The factor VIII gene (ENSG00000185010, gene alias: FVIII, F8, DXS1253E, F8C, or HEMA) is located on the X chromosome (Xq28) and encodes coagulation factor VIII, which is involved in the intrinsic pathway of blood coagulation. Factor VIII is a cofactor for factor IXa and is a Ca +2 In the presence of phospholipids and phospholipids, this gene converts factor X to the active form Xa. This gene generates two alternatively spliced ​​transcripts. Transcript variant 1 (ENST00000360256.9, 26 exons) encodes isoform a, a large glycoprotein that circulates in plasma and forms a noncovalent complex with von Willebrand factor. This protein undergoes multiple cleavages. Transcript variant 2 (ENST00000330287.10, 5 exons) encodes isoform b, a putative small protein consisting primarily of the phospholipid-binding domain of factor VIIIc. This binding domain is essential for coagulation activity. At least seven alternative transcripts have been annotated (Ensembl.org). Defects in this gene result in hemophilia A, a common recessive X-linked coagulation disorder. [Ref Seq. Provided; July 2008] The sequence of Factor VIII is preferably set forth below under accession number NM_000132.4. Several modifications of factor VIII have been engineered to improve its stability and activity, as described, for example, in Miao, HZ et al., "Bioengineering of coagulation factor VIII for improved secretion." Blood (2004). In addition to the B domain deletion, which deletes amino acids 740-1649 (B domain) of the WT F8 protein, linkers (e.g., N6 linkers) have been engineered to further improve F VIII secretion by mimicking some of the normally occurring post-translational modifications, as described in Miao, HZ et al., "Bioengineering of coagulation factor VIII for improved secretion." Blood (2004) and Ward et al. (Ward, NJ et al., "Codon optimization of human factor VIII cDNAs leads to high-level expression." Blood (2011)).

[0157] Suitably, fragments of the Factor VIII coding sequence are within the scope of the present invention.

[0158] Modified Factor VIII is also within the scope of the present invention.

[0159] Suitably, codon-optimised versions of the Factor VIII coding sequence or fragments thereof, such as the BDD Factor VIII coding sequence, are within the scope of the present invention.

[0160] <Arylsulfatase B (ARSB)> The gene (ENSG00000113273) encoding arylsulfatase B (ARSB) is located on chromosome 5, and at least 7 alternative transcripts are annotated (ensembl.org). Isoform 1 (ENST00000264914.10, 8 exons; corresponding to RefSeq NM_000046.5) encodes a 533aa protein (P15848-1). Arylsulfatase B encoded by this gene belongs to the sulfatase family. The arylsulfatase B homodimer hydrolyzes the sulfate groups of N-acetyl-D-galactosamine, chondroitin sulfate, and dermatan sulfate. This protein targets lysozyme. Mucopolysaccharidosis type VI is an autosomal recessive lysosomal storage disorder caused by a deficiency of arylsulfatase B. (Provided by RefSeq; December 2016).

[0161] The sequence of arylsulfatase B (ARSB) is preferably described by the following accession number NM_000046.5.

[0162] <DNA construct> <Foreign DNA sequence> The foreign DNA sequence referred to above includes a DNA fragment to be integrated into the genomic DNA of the target genome. In some embodiments, the foreign DNA includes at least a portion of a gene. The foreign DNA may include a coding sequence, e.g., cDNA, related to the wild-type gene or a "codon-optimized" sequence of the factor to be expressed. In some embodiments, the foreign DNA includes at least one exon of the gene. In some embodiments, the foreign DNA includes an enhancer element of the gene. In some embodiments, the foreign DNA includes a discontinuous sequence of the gene, including the 5' portion of the gene fused to the 3' portion of the gene. In some embodiments, the foreign DNA includes a wild-type gene sequence. In some embodiments, the foreign DNA includes a mutant gene sequence. In some embodiments, the foreign DNA includes a wild-type gene sequence. In some embodiments, the foreign DNA sequence includes a reporter gene. In some embodiments, the reporter gene is selected from at least one of Discosoma Red (Dsred), green fluorescent protein (GFP), red fluorescent protein (RFP), luciferase, β-galactosidase, and β-glucuronidase. In some embodiments, the foreign DNA sequence comprises a transcriptional regulatory element of a gene, which may include, for example, an enhancer sequence. In some embodiments, the foreign DNA sequence comprises one or more exons or fragments thereof. In some embodiments, the foreign DNA sequence comprises one or more introns or fragments thereof. In some embodiments, the foreign DNA sequence comprises at least a portion of a 3' untranslated region or a 5' untranslated region. In some embodiments, the foreign DNA sequence comprises an artificial DNA sequence. In some embodiments, the foreign DNA sequence comprises a nuclear localization sequence and / or a nuclear export sequence. In some embodiments, the foreign DNA sequence comprises a signal peptide sequence. In some embodiments, the foreign DNA sequence comprises a segment of nucleic acid to be integrated into the target genomic locus. In some embodiments, the foreign DNA sequence comprises one or more polynucleotides of interest. In some embodiments, the foreign DNA sequence comprises one or more expression cassettes.In some embodiments, such expression cassettes include a foreign DNA sequence of interest, a polynucleotide encoding a selectable marker and / or reporter gene, and regulatory components that affect expression. In some embodiments, the foreign DNA sequence includes genomic nucleic acid. The genomic nucleic acid may be derived from an animal, mouse, human, non-human, rodent, non-human, rat, hamster, rabbit, pig, cow, deer, sheep, goat, chicken, cat, dog, ferret, primate (e.g., marmoset, rhesus monkey), domesticated or agricultural mammal, bird, bacterium, archaea, virus, or any other organism of interest, or a combination thereof. Foreign DNA sequences of any suitable size may be integrated into the target genome. In some embodiments, the foreign DNA sequence integrated into the genome is less than 0.5, about 0.5, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 kilobases (kb) in length, hi some embodiments, the foreign DNA sequence integrated into the genome is at least about 0.5, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5 kb in length.

[0163] <Genome insertion site> The double-strand break (DSB) site can be specifically introduced using any suitable technique, for example, the CRISPR / Cas9 system and guide RNA disclosed herein. In the present invention, the DSB is introduced into intron 9, intron 11, intron 12, intron 13, or intron 14 of the albumin gene. Exemplary genomic insertion sites are position 733 of intron 9, position 152 of intron 11, position 538 of intron 12, position 927 of intron 12, position 173 of intron 13, position 456 of intron 13, or position 123 of intron 14 of the human albumin gene, where the positions indicate the first nucleotide of each intron.

[0164] The nuclease is preferably guided to the insertion site by a gRNA comprising or consisting of a sequence selected from SEQ ID NOs: 1 to 2 and SEQ ID NOs: 9 to 18.

[0165] [Ribosome skipping sequence: 2A self-cleaving peptide] Ribosomal skipping sequence is used herein as a synonym for 2A self-cleaving peptide or 2A peptide.

[0166] These are 18-22 aa long peptides that can induce cleavage of recombinant proteins within cells. The 2A peptides are derived from the 2A region of the viral genome.

[0167] Four members of the 2A peptide family are frequently used in life science research: P2A, E2A, F2A, and T2A. F2A is derived from foot-and-mouth disease virus 18; E2A is derived from equine rhinitis A virus; P2A is derived from porcine teschovirus-1 2A; and T2A is derived from Thosea asigna virus 2.

[0168] [Table 1]

[0169] Any ribosome skipping sequence can be used within the meaning of the present invention. A preferred one is T2A. The ribosome skipping peptide, such as the 2A peptide, is preferably located between the albumin exon and the foreign DNA sequence.

[0170] <Splicing acceptor sequence> RNA splicing is a form of RNA processing in which newly made precursor messenger RNA (pre-mRNA) transcripts are converted into mature messenger RNA (mRNA). During splicing, introns (non-coding regions) are removed and exons (coding regions) are joined together.

[0171] Within an intron, a donor site (at the 5' end of the intron), a branch point site (near the 3' end of the intron), and an acceptor site (at the 3' end of the intron) are required for splicing. The splice donor site contains a nearly invariant GU sequence at the 5' end of the intron within a larger, less highly conserved region. The splice acceptor site at the 3' end of the intron terminates the intron with a nearly invariant AG sequence. Upstream (5') of the AG is a pyrimidine (C and U)-rich region, or polypyrimidine tract. Further upstream of the polypyrimidine tract is the branch point.

[0172] A "splice acceptor sequence" is a nucleotide sequence that can function as an acceptor site at the 3' end of an intron. The consensus sequences and frequencies of human splice site regions are described in Ma, SL, et al., 2015. PLoS One, 10(6), p.e0130729.

[0173] Preferably, the splicing acceptor sequence has the nucleotide sequence (Y) n NYAG (n is 10 to 20), or a variant having at least 90% or at least 95% sequence identity. Preferably, the splicing acceptor sequence is the sequence (Y) n It may comprise NCAG (n is 10 to 20), or a variant having at least 90% or at least 95% sequence identity.

[0174] <Regulatory elements> The constructs of the invention may contain one or more regulatory elements, which may act pre- or post-transcriptionally. The one or more regulatory elements may facilitate expression in the cells of the invention.

[0175] A "regulatory element" is any nucleotide sequence that facilitates expression of a polypeptide, e.g., increases expression of a transcript or enhances mRNA stability. Suitable regulatory elements include, for example, promoters, enhancer elements, post-transcriptional regulatory elements, and polyadenylation sites.

[0176] The present invention also relates to a construct that may contain regulatory elements that are functional in the host cell in which the vector containing the construct is intended to be expressed.Those skilled in the art can select regulatory elements for use in suitable host cells, such as mammalian or human host cells.Regulatory elements include, for example, promoters, transcription termination sequences, translation termination sequences, enhancers, signal peptides, degradation signals, and polyadenylation elements.

[0177] The constructs of the present invention may optionally include post-transcriptional regulatory elements such as transcription termination sequences, translation termination sequences, signal peptide sequences, internal ribosome entry sites (IRES), enhancer elements, and / or woodchuck hepatitis virus (WHV) post-transcriptional regulatory elements (WPRE). Transcriptional termination regions can typically be obtained from the 3' untranslated region of eukaryotic or viral gene sequences. Transcriptional termination sequences can be located downstream of the coding sequence to provide efficient termination. In the systems of the present invention, a transcriptional termination site is usually included.

[0178] <Promoter> The nucleic acid construct of the present invention may comprise a promoter sequence operably linked to a nucleic acid sequence encoding a desired polypeptide. The term "operably linked" means that the parts (e.g., transgene and promoter) are linked in a manner that allows both to perform their functions substantially unimpeded.

[0179] A promoter within the meaning of the present invention may be a ubiquitous promoter, i.e., a promoter that drives the expression of a gene in a wide range of cells and tissues. Further promoters of the present invention are tissue-specific promoters that exhibit selective activity in one or a group of tissues, but exhibit low or no activity in other tissues. The promoter may exhibit inducible expression in response to the presence of another factor, such as a factor present in the host cell.

[0180] When a vector containing the construct is administered therapeutically, it is preferred that the promoter be functional in the target cells (eg, hepatocytes).

[0181] In some embodiments, the promoter is a ubiquitous promoter or a liver-specific promoter, preferably a hepatocyte-specific promoter. Promoters contemplated for use in the present invention include, but are not limited to, native gene promoters or fragments thereof, such as the cytomegalovirus (CMV) promoter (KF853603.1, bp 149-735), the U6 promoter [37, 38], the thyroxine-binding globulin (TBG) promoter, and the hybrid liver-specific promoter (HLP). However, any suitable promoter known in the art can be used. In preferred embodiments, the promoter is a CMV, HLP, or U6 promoter.

[0182] In a preferred embodiment, the promoter is a U6 promoter, such as the promoter of SEQ ID NO: 27 or a fragment thereof.

[0183] Preferably the promoter nucleotide sequence has at least 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% nucleotide identity to SEQ ID NO: 27, 46, 59 or 61, or a fragment thereof, and preferably the promoter substantially retains the native function of the promoter of SEQ ID NO: 27, 46, 59 or 61.

[0184] The promoter can be incorporated into the construct using standard techniques known in the art. Multiple copies of a promoter or multiple promoters can be used in the constructs of the present invention. In one embodiment, the promoter can be positioned approximately the same distance from the transcription start site as it is from the transcription start site in its natural genetic environment. Some variation in this distance can be tolerated without substantial loss of promoter activity.

[0185] <Polyadenylation sequence> The nucleic acid construct of the present invention may contain a polyadenylation sequence. Preferably, the transgene is operably linked to the polyadenylation sequence. The polyadenylation sequence may be inserted downstream of the transgene to improve expression of the transgene.

[0186] A polyadenylation sequence typically includes a polyadenylation signal, a polyadenylation site, and downstream elements; the polyadenylation signal contains a sequence motif recognized by an RNA cleavage complex; the polyadenylation site is the cleavage site where a polyA tail is added to the mRNA; and the downstream element is typically a GT-rich region located immediately downstream of the polyadenylation site and is important for efficient processing.

[0187] In some embodiments, the polyadenylation sequence is the bovine growth hormone (bGH) polyadenylation sequence or the SV40 polyadenylation sequence; or a fragment thereof that retains the native function of the polyadenylation sequence.

[0188] In a preferred embodiment, the polyadenylation sequence is the bovine growth hormone (bGH) polyadenylation sequence, most preferably a short synthetic polyA.

[0189] Preferred polyadenylation sequences of the present invention are SEQ ID NO:26 or SEQ ID NO:37 or SEQ ID NO:48 or SEQ ID NO:65.

[0190] In some embodiments, the polyadenylation sequence comprises, or consists of, a nucleic acid sequence having at least 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% nucleic acid identity with SEQ ID NO: 26 or SEQ ID NO: 37 or SEQ ID NO: 48 or SEQ ID NO: 65, and preferably, the polyadenylation sequence substantially retains the natural function of the polyadenylation sequence of SEQ ID NO: 26 or SEQ ID NO: 37 or SEQ ID NO: 48 or SEQ ID NO: 65.

[0191] <Post-transcriptional regulatory element> The nucleic acid construct of the present invention may include a post-transcriptional regulatory element. Preferably, the protein coding sequence is operably linked to one or more additional post-transcriptional regulatory elements that may improve gene expression.

[0192] The construct of the present invention may include a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE).

[0193] Suitable WPRE sequences are well known to those skilled in the art (see, for example, Zufferey et al. (1999) Journal of Virology 73:2886-2892; Zanta-Boussif et al. (2009) Gene Therapy 16:605-619). Preferably, the WPRE is a wild-type WPRE or a mutant WPRE. For example, the WPRE can be mutated to inhibit the translation of the woodchuck hepatitis virus X protein (WHX), for example, by mutating the WHX ORF translation initiation codon.

[0194] Preferably, the WPRE comprises, or consists essentially of, a sequence having at least 95% identity with SEQ ID NO: 25.

[0195] <Kozak sequence> The nucleic acid construct of the present invention may include a Kozak sequence. The Kozak sequence is operably linked. The Kozak sequence can be inserted before the start codon to improve the initiation of translation.

[0196] Suitable Kozak sequences are well known to those skilled in the art (see, eg, Kozak (1987) Nucleic Acids Research 15:8125-8148).

[0197] <Guide RNA> "Guide RNA" (gRNA) confers target sequence specificity to RNA-guided nucleases. Guide RNAs are short, non-coding RNA sequences that bind to complementary target DNA sequences. For example, in the CRISPR / Cas9 system, the guide RNA first binds to the Cas9 enzyme, and the gRNA sequence guides the resulting complex through base pairing to a specific location on the DNA, where Cas9 exerts its nuclease activity by cleaving the target DNA strand.

[0198] The term "guide RNA" encompasses not only gRNAs that are compatible with particular nucleases, such as Cas9, but also any suitable gRNA that can be used with any RNA-guided nuclease.

[0199] The guide RNA may comprise a transactivating CRISPR RNA (tracrRNA) that provides a stem-loop structure and a target-specific CRISPR RNA (crRNA) designed to cleave the desired gene target site. The tracrRNA and crRNA can be annealed, for example, by heating them to 95°C for 5 minutes and slowly cooling them to room temperature over 10 minutes. Alternatively, the guide RNA may be a single guide RNA (sgRNA) that contains both the crRNA and the tracrRNA as a single construct.

[0200] The guide RNA may include a 3' end that forms a scaffold for nuclease binding and a programmable 5' end that targets different DNA sites. For example, the target specificity of CRISPR-Cas9 can be determined by the sequence of 15-25 bp at the 5' end of the guide RNA. The desired target sequence usually precedes the protospacer adjacent motif (PAM), which is a short DNA sequence with a length of 2-6 bp following the DNA region targeted for cleavage by a CRISPR system such as CRISPR-Cas9. The PAM is required for the Cas nuclease to cleave and is usually found 3-4 bp downstream of the cleavage site. After base pairing with the target of the guide RNA, Cas9 mediates a double-strand break approximately 3-nt upstream of the PAM.

[0201] There are numerous tools for designing guide RNAs (e.g., Cui, Y., et al., 2018. Interdisciplinary Sciences: Computational Life Sciences, 10(2), pp. 455-465). For example, COSMID is a web-based tool for identifying and validating guide RNAs (Cradick TJ, et al. Mol Ther-Nucleic Acids. 2014;3(12):e214).

[0202] <Chimeric RNA scaffold> The chimeric gRNA scaffold is a double-stranded RNA structure that induces the Cas9 endonuclease to introduce site-specific double-strand breaks in the target DNA and is said to enhance the efficiency of the Cas nuclease (Martin Jinek# et al. 2012 A programmable dual RNA-guided DNA endonuclease in adaptive bacterial immunity). Within the scope of the present invention, the preferred chimeric RNA scaffolds of SEQ ID NO: 28 or 60 are used.

[0203] <RNA-induced gene editing> The vector system of the present invention can be used to deliver a foreign DNA sequence into a cell. The foreign DNA sequence can then be introduced into the genome of the cell at the site of the double-strand break (DSB) by non-homologous end joining (NHEJ). The site of the double-strand break (DSB) can be specifically introduced by any suitable technique, for example, using an RNA-guided gene editing system.

[0204] An "RNA-guided gene editing system" can be used to introduce DSBs and typically includes a guide RNA and an RNA-guided nuclease. The CRISPR / Cas9 system is one example of a commonly used RNA-guided gene editing system, although other RNA-guided gene editing systems can also be used.

[0205] <Nuclease> Nucleases that recognize target sequences are known by those skilled in the art, and include but are not limited to zinc finger nucleases (ZFN), transcription activation-like effector nucleases (TALEN), clustered regularly interspaced short palindromic repeats (CRISPR) nucleases and meganucleases.The nucleases that can be found in compositions and used in the methods disclosed herein are described in more detail below.

[0206] <Zinc finger nuclease (ZFN)> A "zinc finger nuclease" or "ZFN" is a fusion of a Fokl cleavage domain and a DNA recognition domain containing three or more zinc finger motifs. Heterodimerization of two individual ZFNs at specific locations in DNA with precise orientation and spacing results in double-stranded breaks in the DNA. In some cases, ZFNs fuse a cleavage domain to the C-terminus of each zinc finger domain. To enable the two cleavage domains to dimerize and cleave the DNA, the two individual ZFNs bind to opposite strands of DNA with their C-termini spaced a certain distance apart. In some cases, the linker sequence between the zinc finger domain and the cleavage domain requires that the 5' ends of each binding site be approximately 5-7 bp apart. Exemplary ZFNs useful in the present invention include those described in Urnov et al., Nature Reviews Genetics, 2010, 11:636-646; Gaj et al., Nat Methods,2012,9(8):805-7;US Patent No. 6,534,261;6,607,882;6,746,838;6,794,13 No. 6; No. 6,824,978; No. 6,866,997; No. 6,933,113; No. 6,979,539; No. 7,013,219; No. 7,030,215; 7, The ZFNs include, but are not limited to, those described in US Patent Application Publication Nos. 2003 / 0232410 and 2009 / 0203140. In some embodiments, the ZFN is a zinc finger nickase, and in some embodiments, it is an engineered ZFN that induces site-specific single-strand DNA cleavage or nick. The description of zinc finger nickases can be found, for example, in Ramirez et al., Nucl Acids Res, 2012, 40(12):5560-8; Kim et al., Genome Res, 2012, 22(7):1327-33. <talen> "TALEN" or "TAL effector nuclease" is an engineered transcriptional activator-like effector nuclease that includes a central domain of DNA-binding tandem repeats, a nuclear localization signal, and a C-terminal transcriptional activation domain. In some embodiments, the DNA-binding tandem repeats include 33-35 amino acids in length and contain two hypervariable amino acid residues at positions 12 and 13 that recognize one or more specific DNA base pairs. TALENs are produced by fusing the TAL effector DNA-binding domain to a DNA cleavage domain. For example, the TALE protein can be fused to a nuclease such as a wild-type or mutant Fokl endonuclease, or the catalytic domain of Fokl. For its use in TALENs, several mutations have been added to Fokl, which, for example, improve cleavage specificity or activity. Such TALENs are designed to bind to any desired DNA sequence. TALENs are often used to create double-strand breaks in the target DNA sequence and then cause gene modification by undergoing NHEJ or HDR. In some cases, a single-stranded donor DNA repair template is provided to facilitate HDR. Detailed descriptions of TALENs and their use for gene editing can be found, for example, in U.S. Patent Nos. 8,440,431; 8,440,432; 8,450,471; 8,586,363; and 8,697,853; Scharenberg et al., Curr Gene Ther, 2013, 13(4):291-303; Gaj et al., Nat Methods, 2012, 9(8):805-7; Beurdeley et al., Nat Commun, 2013, 4:1762; and, Joung and Sander, Nat Rev Mol Cell Biol, 2013, 14(l):49-55. <DNA-induced nuclease> A "DNA-guided nuclease" is a nuclease that uses complementary nucleotides of single-stranded DNA to guide the nuclease to the correct location in the genome by hybridizing to another nucleic acid (e.g., a target nucleic acid in the genome of a cell). In some embodiments, the DNA-guided nuclease comprises an Argonaute nuclease. In some embodiments, the DNA-guided nuclease is selected from TtAgo, PfAgo, and NgAgo. In some embodiments, the DNA-guided nuclease is NgAgo.

[0207] <Meganuclease> A "meganuclease" is, in certain embodiments, a highly specific, rare-cutting or homing endonuclease that recognizes a DNA target site in the range of at least 12 base pairs in length, e.g., 12-40 base pairs or 12-60 base pairs.

[0208] Any meganuclease that may be used in this specification includes, but is not limited to, I-Scel, I-Scell, I-SceIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-Ceul, I-CeuAIIP, I-Crel, I-CrepsblP, I-CrepsbllP, I-CrepsbIIIP, I-CrepsbIVP, I-Tlil, I-Ppol, PI-PspI, F-Scel, F-Scell, F-Suvl, F-Tevl, F-TevII, I-Amal, I-Anil, I-Chul, I-Cmoel, I-Cpal, I-CpaII, I-Csml, I-Cvul, I-CvuAIP, I-Ddil, I-DdiII, I-Dirl, I-Dmol, I-Hmul, I-HmuII, I-HsNIP, I-Llal, I-Msol, I-Naal, I-Nanl, I-NcIIP, I-NgrIP, I-Nitl, I-Njal, I-Nsp236IP, I-Pakl, I-PboIP, I-PcuIP, I-PcuAI, I-PcuVI, I-PgrlP, 1-PobIP, I-Porl, I-PorIIP, I-PbpIP, I-SpBetaIP, I-Scal, I-SexIP, 1-SneIP, I-Spoml, I-SpomCP, I-SpomIP, I-SpomIIP, I-SquIP, I-Ssp6803I, I-SthPhiJP, I-SthPhiST3P, I-SthPhiSTe3bP, I-TdeIP, I-Tevl, I-TevII, I-TevIII, I-UarAP, I-UarHGPAIP, I-UarHGPA13P, I-VinlP, 1-ZbiIP, PI-MtuI, PI-MtuHIP, PI-MtuHIIP, PI-PfuI, PI-PfuII, PI-PkoI, Pl-PkoII, PI-Rma43812IP, PI-SpBetaIP, PI-SceI, PI-Tful, PI-TfuII, PI-Thyl, PI-Tlil, PI-THII, I-Crel meganuclease, I-Ceul meganuclease, I-Msol meganuclease, I-Scel meganuclease, or any active variant, fragment, mutant or derivative thereof.

[0209] <crispr> The CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) / Cas (CRISPR-associated protein) nuclease system is an artificial nuclease system based on a bacterial system used for genome engineering. It is based, in part, on the adaptive immune response of many bacteria and archaea. When a virus or plasmid invades a bacterium, a segment of the invader's DNA is converted into CRISPR RNA (crRNA) by the "immune" response. The crRNA then binds to another type of RNA, called tracrRNA, through a region of partial complementarity and guides a Cas (e.g., Cas9) nuclease to a "protospacer," a region in the target DNA that is homologous to the crRNA. The Cas (e.g., Cas9) nuclease cleaves DNA at a site specified by a 20-nucleotide complementary sequence contained in the crRNA transcript, generating a blunt end at the double-strand break. Cas (e.g., Cas9) nucleases, in some embodiments, require both a crRNA and a tracrRNA for site-specific DNA recognition and cleavage. This system is currently designed in certain embodiments where the crRNA and tracrRNA are combined into a single molecule (a "single guide RNA" or "sgRNA"), and the crRNA-equivalent portion of the single guide RNA guides the Cas (e.g., Cas9) nuclease to target any desired sequence (see, e.g., Jinek et al. (2012) Science 337:816-821; Jinek et al. (2013) eLife 2:e00471; Segal (2013) eLife 2:e00563).

[0210] As used herein, tracrRNA is also defined as a scaffold gRNA. Thus, CRISPR / Cas systems can be designed to generate double-strand breaks at desired targets in a cell's genome and utilize the cell's endogenous mechanisms to repair the induced breaks by homology-directed repair (HDR) or non-homologous end joining (NHEJ). In some embodiments, the Cas nuclease has DNA cleavage activity. In some embodiments, the Cas nuclease directs cleavage of one or both strands at a position in the target DNA sequence. For example, in some embodiments, the Cas nuclease is a nickase with one or more inactivating catalytic domains that cleave a single strand of the target DNA sequence. Non-limiting examples of Cas nucleases include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cpf1, C2c3, C2c2 and C2c1, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, C Cas nucleases include Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Cpf1, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, their homologs, their variants, their mutants, and their derivatives. There are three major types of Cas nucleases (type I, type II, and type III) and ten subtypes, including five type I, three type II, and two type III proteins (see, e.g., Hochstrasser and Doudna, Trends Biochem Sci, 2015:40(1):58-66). Type II Cas nucleases include, but are not limited to, Cas1, Cas2, Csn2, and Cas9, and are well known to those skilled in the art.For example, the amino acid sequence of a Streptococcus pyogenes wild-type Cas9 polypeptide is described, for example, in NBCI Ref.Seq.No.NP269215, and the amino acid sequence of a Streptococcus thermophilus wild-type Cas9 polypeptide is described, for example, in NBCI Ref.Seq.No.WP_011681470. Cas nucleases, e.g., Cas9 polypeptides, are derived from various bacterial species in some embodiments. "Cas9" refers to an RNA-guided double-stranded DNA-binding nuclease protein or nickase protein. Wild-type Cas9 nucleases have two functional domains, e.g., RuvC and HNH, which cleave different DNA strands. When both functional domains are active, Cas9 can induce double-strand breaks in genomic DNA (target DNA). In some embodiments, the Cas9 enzyme comprises one or more catalytic domains of a Cas9 protein derived from bacteria belonging to the group consisting of Corynebacterium, Sutterella, Legionella, Treponema, Filifactor, Eubacterium, Streptococcus, Lactobacillus, Mycoplasma, Bacteroides, Flavivora, Flavobacterium, Sphaerochaeta, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, Staphylococcus, Nitratifactor, and Campylobacter. In some embodiments, the Cas9 is a fusion protein, e.g., the two catalytic domains are derived from different bacterial species. Useful variants of the Cas9 nuclease include RuvC. - or HNH - The Cas9 nickase includes a single inactive catalytic domain, such as an enzyme or nickase. Cas9 nickases have only one active functional domain and, in some embodiments, cleave only one strand of the target DNA, thereby generating a single-strand break or nick. In some embodiments, the Cas9 nickase is a mutant Cas9 nuclease with at least a D10A mutation. In other embodiments, the Cas9 nickase is a mutant Cas9 nuclease with at least an H840A mutation. Examples of other mutations present in Cas9 nickases include, but are not limited to, N854A and N863A. When at least two DNA-targeting RNAs targeting opposite DNA strands are used, a double-strand break is introduced using Cas9 nickases. The double-nick-induced double-strand break is repaired by NHEJ or HDR. This gene editing strategy favors HDR and reduces the frequency of indel mutations at off-target DNA sites. In some embodiments, the Cas9 nuclease or nickase is codon-optimized for the target cell or organism. In some embodiments, the Cas nuclease is a Cas9 polypeptide containing two silencing mutations in the RuvC1 and HNH nuclease domains (D10A and H840A), referred to as dCas9. In one embodiment, the dCas9 polypeptide from Streptococcus pyogenes contains at least one mutation at positions D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, A987, or any combination thereof. A description of such dCas9 polypeptides and variants thereof is provided, for example, in International Patent Publication No. WO2013 / 176772. In some embodiments, the dCas9 enzyme contains a mutation at D10, E762, H983, or D986, as well as a mutation at H840 or N863. In some embodiments, the dCas9 enzyme comprises a D10A or D10N mutation, and alternatively, the dCas9 enzyme comprises a H840A, H840Y, or H840N mutation.In some embodiments, the dCas9 enzyme of the present invention comprises the following substitutions: D10A and H840A; D10A and H840Y; D10A and H840N; D10N and H840A; D10N and H840Y; or D10N and H840N. The substitutions are alternatively conservative or non-conservative substitutions to render the Cas9 polypeptide catalytically inactive and capable of binding to target DNA. For genome editing methods, the Cas nuclease, in some embodiments, comprises a Cas9 fusion protein, such as a polypeptide comprising the catalytic domain of the type IIS restriction enzyme Fok1 linked to dCas9. The Fok1-dCas9 fusion protein (fCas9) can use two guide RNAs to bind to a single strand of target DNA to generate a double-strand break.

[0211] <Target sequence> A target sequence herein is a nucleic acid sequence that is recognized and cleaved by a nuclease. In some embodiments, the target sequence is about 9 to about 12 nucleotides, about 12 to about 18 nucleotides, about 18 to about 21 nucleotides, about 21 to about 40 nucleotides, about 40 to about 80 nucleotides, or any combination of subranges thereof (e.g., 9 to 18, 9 to 21, 9 to 40, and 9 to 80 nucleotides) in length. In some embodiments, the target sequence comprises a nuclease binding site. In some embodiments, the target sequence comprises a nick / cleavage site. In some embodiments, the target sequence comprises a protospacer adjacent motif (PAM) sequence. In some embodiments, the target nucleic acid sequence (e.g., protospacer) is 20 nucleotides. In some embodiments, the target nucleic acid is less than 20 nucleotides. In some embodiments, the target nucleic acid is at least 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or more nucleotides in length. In some embodiments, the target nucleic acid is at most 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or more nucleotides. In some embodiments, the target nucleic acid sequence is 16, 17, 18, 19, 20, 21, 22, or 23 bases immediately 5' to the first nucleotide of the PAM. In some embodiments, the target nucleic acid sequence is 16, 17, 18, 19, 20, 21, 22, or 23 bases immediately 3' to the last nucleotide of the PAM. In some embodiments, the target nucleic acid sequence is 20 bases immediately 5' to the first nucleotide of the PAM. In some embodiments, the target nucleic acid sequence is 20 bases immediately 3' to the last nucleotide of the PAM. In some embodiments, the target nucleic acid sequence is 5' or 3' to the PAM. In some embodiments, the target sequence comprises a nucleic acid sequence present in the target nucleic acid to which the nucleic acid targeting segment of the complementary strand nucleic acid binds. For example, in some embodiments, the target sequence comprises a sequence to which a complementary strand of nucleic acid is designed to base pair. The target sequence comprises a nuclease cleavage site. In some embodiments, the target sequence is adjacent to the nuclease cleavage site.In some embodiments, a nuclease cleaves a nucleic acid at a site inside or outside the nucleic acid sequence present in the target nucleic acid to which the complementary nucleic acid target sequence binds. The cleavage site, in some embodiments, includes the position in the nucleic acid where the nuclease makes a single-strand or double-strand break. For example, the formation of a nuclease complex containing a complementary nucleic acid hybridized to a protease recognition sequence and complexed with a protease results in cleavage of one or both strands within or near (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 19, 20, 23, 50, or more base pairs) the nucleic acid sequence present in the target nucleic acid to which the spacer region of the complementary nucleic acid binds. In some embodiments, the cleavage site is located in only one strand of the nucleic acid or in both strands. In some embodiments, the cleavage site is located at the same position on both strands of the nucleic acid (creating a blunt end) or at a different site on each strand (creating a staggered end). In some embodiments, site-specific cleavage of a target nucleic acid by a nuclease protein occurs at a position determined by base-pairing complementarity between the complementary strand nucleic acid and the target nucleic acid. In some embodiments, site-specific cleavage of a target nucleic acid by a nuclease protein occurs at a position determined by a short motif in the target nucleic acid called a protospacer adjacent motif (PAM). For example, the PAM flanks the recognition sequence at its 3' end. In some cases, cleavage generates a blunt end. In some cases, cleavage generates a staggered end or a sticky end with a 5' overhang. In some cases, cleavage generates a staggered end or a sticky end with a 3' overhang. Orthologs of various nuclease proteins utilize different PAM sequences. For example, in some embodiments, different Cas proteins recognize different PAM sequences. For example, in S. pyogenes, the PAM is a sequence in the target nucleic acid comprising the sequence 5'-XRR-3', where R is either A or G, X is any nucleotide, and X is immediately 3' to the target nucleic acid sequence targeted by the spacer sequence. The PAM sequence for S. pyogenes Cas9 (SpyCas9) is 5'-XGG-3', where X is any DNA nucleotide and is immediately 3' to the nuclease recognition sequence on the non-complementary strand of the target DNA.The PAM of Cpf1 is 5'-TTX-3', where X is any DNA nucleotide and is immediately 5' to the nuclease recognition sequence. Preferably, the Cas9 / sgRNA complex introduces a DSB three base pairs upstream of the PAM sequence in the genomic target sequence, resulting in two blunt ends. The exact same Cas9 / sgRNA target sequence is loaded onto the donor DNA in the reverse orientation. The target genomic locus and donor DNA are cleaved by Cas9 / gRNA, and the linearized donor DNA is integrated into the target site via the NHEJ DSB repair pathway. If the donor DNA is integrated in the correct orientation, the binding sequence is protected from further cleavage by Cas9 / gRNA. If the donor DNA is integrated in the reverse orientation, an intact Cas9 / gRNA target site is present, allowing Cas9 / gRNA to excise the integrated donor DNA.

[0212] In an embodiment of the invention, the PAM has a sequence selected from TGG, AGG, GGG, CGG.

[0213] <Vector> The present invention also relates to vectors comprising the nucleic acid constructs described herein.

[0214] Such a vector may therefore contain any of the elements described above in connection with the construct, in particular it may include one or more regulatory elements, in particular as defined above, such as a promoter, a transcription termination sequence, a translation termination sequence, an enhancer, a signal peptide, a degradation signal and a polyadenylation element.

[0215] Vectors suitable for delivery and expression of nucleic acids into cells for gene therapy are encompassed by the invention.

[0216] The vectors of the present invention include viral and non-viral vectors.

[0217] Non-viral vectors include non-viral agents commonly used to introduce or maintain nucleic acids in cells, including polymer-, particle-, lipid-, peptide-based delivery vehicles, or combinations thereof, such as cationic polymers, micelles, liposomes, exosomes, microparticles, and nanoparticles, including lipid nanoparticles (LNPs), among others.

[0218] Among viral delivery methods, genetically engineered viruses such as adeno-associated viruses are currently one of the most common tools for gene delivery. The concept of viral-based gene delivery is to engineer viruses so that they can express a gene of interest or regulatory sequences such as promoters and introns. Depending on the specific application and type of virus, most viral vectors contain mutations that prevent them from replicating freely in the host as wild-type viruses. Several different families of viruses have been modified to generate viral vectors for gene delivery. These viruses include retroviruses, lentiviruses, adenoviruses, adeno-associated viruses, herpesviruses, baculoviruses, picornaviruses, and alphaviruses.

[0219] Viral vectors of the present invention can be derived from non-pathogenic parvoviruses such as adeno-associated viruses (AAV), retroviruses such as gammaretroviruses, spumaviruses and lentiviruses, adenoviruses, poxviruses and herpesviruses.

[0220] Particularly preferred viruses according to the present invention are lentiviruses and adeno-associated viruses.

[0221] Viral vectors are essentially capable of penetrating cells and delivering a nucleic acid of interest into the cell through a process known as viral transduction.

[0222] As used herein, the term "viral vector" refers to a non-replicating, non-pathogenic virus engineered to deliver genetic material into cells. Viral genes essential for replication and pathogenicity are replaced with an expression cassette for the desired transgene. Thus, the viral vector genome contains an expression cassette for the transgene flanked by viral sequences required for viral vector production.

[0223] The term "virus particle" or "viral particle" is intended to mean the extracellular form of a non-pathogenic virus, particularly a viral vector, in which genetic material made of either DNA or RNA is enclosed in a protein coat called a capsid, and in some cases an envelope derived from a portion of the host cell membrane that contains viral glycoproteins.

[0224] As used herein, viral vector also refers to a viral vector particle.

[0225] The viral vectors encompassed by the present invention are suitable for gene therapy.

[0226] Viral particles can be obtained, for example, using a vector capable of harboring a gene of interest and a helper cell capable of providing viral structural proteins and enzymes that allow the production of vector-containing infectious viral particles.

[0227] <Adeno-associated virus (AAV)> Adeno-associated viruses are a family of viruses that differ in nucleotide sequence, amino acid sequence, genome structure, pathogenicity, and host range. This diversity provides an opportunity to develop various therapeutic applications using viruses with different biological properties.

[0228] An ideal adeno-associated virus-based vector for gene delivery must be efficient, cell-specific, controlled, and safe. The efficiency of delivery can determine the effectiveness of therapy. Current efforts aim to achieve cell-type-specific infection and gene expression with adeno-associated virus vectors. Furthermore, because long-term or controlled expression may be required for therapy, adeno-associated virus vectors have been developed to control the expression of the gene of interest.

[0229] Adeno-associated virus (AAV) is a small virus that infects humans and some other primates. AAV is not currently known to cause disease, and as a result, the virus induces a very mild immune response. Gene therapy vectors using AAV can infect both dividing and quiescent cells and persist extrachromosomally without integrating into the host cell genome. These characteristics make AAV a very attractive candidate for generating viral vectors for gene therapy and for generating syngeneic human disease models.

[0230] Wild-type AAV has attracted considerable interest from gene therapy researchers due to a number of features. Chief among these is the virus's apparent lack of pathogenicity. It can also infect non-dividing cells and has the ability to stably integrate into the host cell genome at a specific site on human chromosome 19 (designated AAVS1). However, the development of AAV as a gene therapy vector eliminated this integration ability by removing the rep and cap sequences from the vector DNA. The desired gene, along with the promoter driving gene transcription, is inserted between the ITRs, which support concatemer formation in the nucleus after single-stranded vector DNA is converted to double-stranded DNA by the host cell DNA polymerase complex. AAV-based gene therapy vectors form episomal concatemers in the host cell nucleus. In non-dividing cells, these concatemers remain intact throughout the life of the host cell. In dividing cells, episomal DNA is not replicated along with the host cell DNA, resulting in the loss of AAV DNA through cell division. Random integration of AAV DNA into the host genome is detectable but occurs at a very low frequency. Furthermore, AAVs exhibit very low immunogenicity, seemingly limited to the generation of neutralizing antibodies, while they do not induce clearly defined cytotoxic responses, characteristics that, together with their ability to infect quiescent cells, make AAVs particularly suitable for human gene therapy.

[0231] The AAV genome is constructed of approximately 4.7 kilobases of single-stranded deoxyribonucleic acid (ssDNA), which can be either positive or negative strand. The genome contains inverted terminal repeats (ITRs) at both ends of the DNA strand and two open reading frames (ORFs): rep and cap. The former consists of four overlapping genes encoding the Rep protein required for the AAV life cycle, while the latter contains overlapping nucleotide sequences for the capsid proteins: VP1, VP2, and VP3, which interact to form a capsid with icosahedral symmetry.

[0232] Inverted terminal repeat (ITR) sequences are so named because of their symmetry, which has been shown to be required for efficient amplification of the AAV genome. Another property of these sequences is their ability to form hairpins that contribute to so-called self-priming, which allows primase-independent synthesis of the second DNA strand. ITRs have also been shown to be required for efficient encapsidation of AAV DNA in combination with the generation of fully assembled, deoxyribonuclease-resistant AAV particles.

[0233] For gene therapy, the ITRs appear to be the only cis sequences required next to the therapeutic gene: the structural gene (cap) and packaging gene (rep) can be delivered in trans. Based on this assumption, many methods have been established to efficiently produce recombinant AAV (rAAV) vectors containing reporter or therapeutic genes.

[0234] An AAV vector comprises an AAV capsid capable of transducing a target cell of interest. The AAV capsid can be derived from one or more natural or artificial AAV serotypes.

[0235] AAVs are sometimes referred to by their serotypes. Serotypes correspond to AAV variants with specific reactivity that can be used to distinguish them from other variants based on the expression profile of capsid surface antigens. Usually, AAV vector particles with a specific AAV serotype do not effectively cross-react with neutralizing antibodies specific to any other AAV serotype.

[0236] All known serotypes are capable of infecting cells from multiple diverse tissue types. Tissue specificity is determined by the capsid serotype, and pseudotyping AAV vectors to alter their tropism range impacts their use in therapy.

[0237] The inverted terminal repeat (ITR) sequence used in the AAV vector system of the present invention can be any AAV ITR. The ITRs used in the AAV vector can be the same or different. For example, the vector can include an ITR of AAV serotype 2 and an ITR of AAV serotype 5. In one embodiment of the vector of the present invention, the ITRs are derived from AAV serotypes 2, 4, 5, or 8. The ITRs of AAV serotype 2 are preferred in the present invention. AAV ITR sequences are well known in the art (e.g., for ITR2, see GenBank accession numbers AF043303.1; NC_001401.2; J01901.1; JN898962.1; for ITR5, see GenBank accession number NC_006152.1).

[0238] Serotype 2 (AAV2) has been the most extensively studied to date and displays natural tropism for skeletal muscle, neurons, vascular smooth muscle cells, and hepatocytes.

[0239] Three cellular receptors have been described for AAV2: heparan sulfate proteoglycans (HSPGs), αVβ5 integrin, and fibroblast growth factor receptor 1 (FGFR-1). The first functions as a primary receptor, whereas the latter two have co-receptor activity, allowing AAV to enter cells by receptor-mediated endocytosis. Although HSPGs function as primary receptors, their abundance in the extracellular matrix can remove AAV particles, potentially impairing infection efficiency.

[0240] Although AAV2 is the most common serotype in various AAV-based studies, other serotypes have been shown to be effective as gene delivery vectors. For example, AAV6 appears to be much better at infecting airway epithelial cells, and AAV7 exhibits very high transduction rates of mouse skeletal muscle cells (similar to AAV1 and AAV5). AAV8 is highly effective at transducing hepatocytes and photoreceptor cells, and AAV1 and AAV5 have been shown to be highly efficient at gene delivery to vascular endothelial cells. In the brain, most AAV serotypes exhibit neuronal tropism, but AAV5 also transduces astrocytes. AAV6, a hybrid of AAV1 and AAV2, also exhibits lower immunogenicity than AAV2.

[0241] Serotypes can differ with respect to the receptors they bind: for example, AAV4 and AAV5 transduction can be inhibited by soluble sialic acid (in different forms for each of these serotypes), and AAV5 has been shown to enter cells via the platelet-derived growth factor receptor.

[0242] Methods for preparing virions containing virus and heterologous polynucleotide or construct are known in the art.In the case of AAV, cells can be co-infected or transfected with adenovirus or polynucleotide constructs containing suitable adenovirus genes for AAV helper function.Examples of materials and methods are described in, for example, U.S. Patent Nos. 8,137,962 and 6,967,018.The AAV virus or AAV vector of the present invention can be any AAV serotype, including but not limited to AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, and AAV11, AAV-PhP.B and AAV-PhP.eB serotypes.

[0243] In certain embodiments, AAV2 or AAV5 or AAV7 or AAV8 or AAV9 serotypes are utilized. Preferably, AAV2-8 is used.

[0244] Preferably, the AAV genome is derivatized for the purpose of administration to a patient. Such derivatization is standard in the art, and the present invention encompasses the use of any known AAV genome derivative, and derivatives that can be produced by applying techniques known in the art. The AAV genome can be a derivative of any naturally occurring AAV. Preferably, the AAV genome is a derivative of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, or AAV11.

[0245] Derivatives of the AAV genome include any truncated or modified form of the AAV genome that allows for in vivo expression of a transgene from an AAV vector of the invention. In one embodiment, the AAV serotype provides one or more tyrosine to phenylalanine (YF) mutations on the capsid surface.

[0246] The DNA constructs described above can be used to generate AAV vectors of the present invention. AAV vectors can be produced, for example, by triple transfection of producer cells, such as HEK293 cells, a method known in the art in which a plasmid containing a gene of interest is transfected into producer cells, where viral particles are produced, along with two additional plasmids.

[0247] <Plasmid> Plasmids for the production of viral vectors as defined herein are also within the scope of the present invention.

[0248] The plasmid may comprise the DNA construct described above. Plasmids typically further comprise backbone elements, such as a bacterial replication origin, a bacterial promoter, and an antibiotic resistance gene, that are typically required for large-scale production of the plasmid in bacteria.

[0249] The use of said plasmids for the generation of the vectors of the present invention is within the scope of the present invention.

[0250] Vectors, such as AAV vectors, can be produced, for example, by triple transfection of producer cells, such as HEK293 cells, a method known in the art where a plasmid containing the DNA construct of interest is transfected into producer cells where viral particles are produced together with two additional plasmids.

[0251] <HITI Genome Editing System> As used herein, a "genome editing system" is preferably a system that includes all the components necessary to edit the genome, using the constructs or vectors of the present invention.

[0252] Within the scope of the present invention, a genome editing system is a system that includes a donor nucleic acid optionally containing one or more exons of a foreign DNA sequence and an albumin gene, a complementary strand oligonucleotide homologous to a target sequence (e.g., a gRNA homologous to a target sequence within the albumin gene as defined herein, preferably intron 12, 13 or 14 of the albumin gene), and a nuclease that recognizes the target sequence.

[0253] Preferably, the genome editing system of the present invention includes the nucleotide sequences, DNA constructs, vectors of the present invention, such as non-viral or viral vectors, and / or viral particles.

[0254] <Host Cell> The present invention also relates to host cells containing the viral vectors of the present invention. Host cells can be cultured or primary cells, i.e., cells isolated directly from an organism (e.g., a human). Host cells can be adherent or suspension cells, i.e., cells growing in suspension. Suitable host cells known in the art include, for example, DH5α, E. coli cells, Chinese hamster ovary cells, monkey VERO cells, COS cells, HEK293 cells, and the like. The cells can be human cells or derived from other animals. In one embodiment, the cells are retinal cells, particularly photoreceptors, RPE cells, or cone cells. The cells can also be hepatocytes, particularly hepatocytes. Selection of an appropriate host is considered within the capabilities of one skilled in the art given the teachings herein. Preferably, the host cells are animal cells, most preferably human cells. The cells are capable of expressing the nucleotide sequence provided in the viral vectors of the present invention.

[0255] Those skilled in the art are familiar with standard methods for incorporating polynucleotides or vectors into host cells, such as, for example, transfection, lipofection, electroporation, microinjection, viral infection, heat shock, transformation after chemical permeabilization of the membrane or cell fusion.

[0256] The term "host cell or genetically engineered host cell" as used herein relates to a host cell that has been transduced, transformed or transfected by a viral vector of the present invention.

[0257] <Composition> Pharmaceutical compositions within the meaning of the present invention include those in which the system of the present invention, one or more vectors, host cells, or viral particles are combined with a pharmaceutically acceptable carrier, diluent, excipient, or adjuvant. The choice of pharmaceutical carrier, excipient, or diluent can be selected based on the intended route of administration and standard pharmaceutical practice. The pharmaceutical composition can include, as or in addition to a carrier, excipient, or diluent, any suitable binder, lubricant, suspending agent, coating agent, solubilizer, and other carrier agent (e.g., lipid delivery system, etc.) that can assist or enhance the entry of the virus into the target site. The vector can be administered in vivo or ex vivo.

[0258] Pharmaceutical compositions suitable for parenteral administration, containing an amount of the compound of the present invention, constitute a preferred embodiment. For parenteral administration, the composition is best used in the form of a sterile aqueous solution, which may contain other substances, such as sufficient salts or simple sugars, to make the solution isotonic with blood. In a preferred embodiment, the vector or pharmaceutical composition is delivered systemically, for example, by intravenous injection.

[0259] The methods of the present invention can be used with humans and other animals. As used herein, the terms "patient" and "subject" are used interchangeably and are intended to include such human and non-human species. Similarly, the in vitro methods of the present invention can be performed on cells of such human and non-human species.

[0260] <Kit> The present invention also relates to kits comprising a DNA construct, system, one or more vectors, host cells, or viral particles of the present invention in one or more containers. The kits of the present invention may optionally include a pharmaceutically acceptable carrier and / or diluent. In one embodiment, the kits of the present invention include one or more other components, additives, or adjuvants described herein. In one embodiment, the kits of the present invention include instructions or packaging materials describing how to administer the vector system of the kit. The containers of the kit can be made of any suitable material, such as glass, plastic, or metal, and of any suitable size, shape, or configuration. In one embodiment, the viral vectors or host cells of the present invention are provided in the kit as solids. In another embodiment, the viral vectors or host cells of the present invention are provided in the kit as liquids or solutions. In one embodiment, the kit includes an ampoule or syringe containing the viral vectors or host cells of the present invention in liquid or solution form.

[0261] <Delivery> The vectors of the present invention can be administered to a patient. The administration can be "in vivo" or "ex vivo." Those skilled in the art will be able to determine the appropriate dosage. The term "administered" includes delivery by viral or non-viral techniques. Viral delivery mechanisms include, but are not limited to, the adenoviral vectors, adeno-associated viral (AAV) vectors, herpesvirus vectors, retroviral vectors, lentiviral vectors, and baculoviral vectors described above. Non-viral delivery systems include polymer-based, particle-based, lipid-based, peptide-based delivery vehicles, or combinations thereof, such as cationic polymers, micelles, liposomes, exosomes, nanoparticles, including microparticles and lipid nanoparticles (LNPs), and DNA transfection, such as electroporation. Delivery of one or more therapeutic genes by the vector systems of the present invention can be used alone or in combination with other therapies or components of the therapy.

[0262] Any suitable delivery method for delivering the compositions of the present invention is contemplated. In some embodiments, the individual components of the HITI genome editing system (e.g., gRNA, nuclease, and / or foreign DNA sequence) are delivered simultaneously or temporally separated. The choice of genetic modification method depends on the type of cell to be transformed and / or the context in which the transformation is performed (e.g., in vitro, ex vivo, or in vivo). A general description of these methods can be found in Ausubel, et al., Short Protocols in Molecular Biology, 3rd ed., Wiley & Sons, 1995.

[0263] The term "contacting a cell" includes all delivery methods disclosed herein. In some embodiments, the methods disclosed herein include contacting a target DNA or introducing into a cell (or population of cells) one or more nucleic acids including a nucleotide sequence encoding a complementary strand nucleic acid (e.g., gRNA), a site-specific modifying polypeptide (e.g., Cas protein) or a nucleic acid encoding the same, and / or an exogenous DNA sequence.

[0264] <Method of genome DNA editing> The term "genome editing" refers to a type of genetic engineering that uses one or more nucleases and / or nickases to insert, replace, or remove DNA into target DNA, e.g., the genome of a cell.

[0265] Provided herein are methods and compositions for homology-independent targeted integration (HITI) to modify nucleic acids, such as genomic DNA, including the genomic DNA of dividing or non-dividing cells or terminally differentiated cells. The methods herein, at least in some embodiments, use non-homologous end joining to insert foreign DNA into target DNA (e.g., the genomic DNA of a cell, such as a dividing or non-dividing or terminally differentiated cell) without relying on homology. In some embodiments, the methods herein include a method for integrating a foreign DNA sequence into the genome of a dividing or non-dividing cell, comprising contacting the non-dividing cell with a composition comprising one or more targeting constructs comprising the foreign DNA sequence and a target sequence, a complementary oligonucleotide homologous to the target sequence, and a nuclease. The foreign DNA sequence comprises at least one nucleotide difference compared to the genome, and the target sequence is recognized by the nuclease. In some embodiments of the HITI methods disclosed herein, the foreign DNA sequence is a DNA fragment comprising a desired sequence to be inserted into the genome of the target or host cell. At least a portion of the foreign DNA sequence has a sequence homologous to a portion of the genome of the target cell or host cell, and at least a portion of the foreign DNA sequence has a sequence that is not homologous to a portion of the genome of the target cell or host cell. For example, in some embodiments, the foreign DNA sequence can include a portion of the host cell genomic DNA sequence having a mutation therein. Thus, when the foreign DNA sequence is integrated into the genome of the host cell or target cell, the mutation found in the foreign DNA sequence is carried to the host cell or target cell genome. In some embodiments of the HITI methods disclosed herein, the foreign DNA sequence is flanked by at least one target sequence. In some embodiments, the foreign DNA sequence is flanked by two target sequences. The target sequence comprises a specific DNA sequence recognized by at least one nuclease. In some embodiments, the target sequence is recognized by the nuclease in the presence of a complementary strand oligonucleotide having a sequence homologous to the target sequence. In some embodiments, in the HITI methods disclosed herein, the target sequence comprises a nucleotide sequence recognized and cleaved by the nuclease.Nucleases that recognize target sequences are known by those skilled in the art, and include, but are not limited to, zinc finger nucleases (ZFNs), transcription activation-like effector nucleases (TALENs), and clustered regularly interspaced short palindromic repeats (CRISPR) nucleases. In some embodiments, ZFNs comprise zinc finger DNA binding domains and DNA cleavage domains, which are fused to create sequence-specific nucleases. In some embodiments, TALENs comprise TAL effector DNA binding domains and DNA cleavage domains, which are fused to create sequence-specific nucleases. In some embodiments, CRISPR nucleases are naturally occurring nucleases that recognize DNA sequences homologous to clustered regularly interspaced short palindromic repeats, which are commonly found in prokaryotic DNA. CRISPR nucleases include, but are not limited to, Cas9, Cpf1, C2c3, C2c2, and C2c1. Advantageously, the Cas9 of the present invention is a mutant with reduced off-target activity, SpCas9D10A (Ran, FA, et al., Genome engineering using the CRISPR-Cas9 system. Nat Protoc, 2013. 8(11): pp. 2281-2308.) (with inactivation of RuvC domain cleavage activity); SpCas9N863A (Ran, FA, et al., Genome engineering using the CRISPR-Cas9 system. Nat Protoc, 2013. 8(11): pp. 2281-2308). (inactivation of HNH domain cleavage activity); SpCas9-HF1 (Kleinstiver, B.P., et al., High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects. Nature, 2016. 529(7587): pp. 490-5) (reducing Cas9 binding energy through protein engineering); eSpCas9 (Laymaker, I.M., et al., Rationally engineered Cas9 nucleases with improved specificity. Science, 2016. 351(6268): pp. 84-8) (reducing the positive charge of Cas9); EvoCas9 (Asini, A., et al., A highly specific SpCas9 variant is identified by in vivo screening in yeast. Nat Biotechnol, 2018. 36(3): pp. 265-271) (mutagenesis of the REC3 domain); KamiCas9 (Merienne, N., et al. al., The Self-Inactivating KamiCas9 System for the Editing of CNS Disease Genes. Cell Rep, 2017. 20(12): p. 2980-2991) (knockout of Cas9 after expression). The HITI DNA genome editing methods disclosed herein, in some embodiments, can introduce foreign DNA sequences into a host genome or a target genome. In some embodiments, the insertion comprises a specific number of nucleotides ranging from 1 to 4700 base pairs, e.g., 1 to 10, 5 to 20, 15 to 30, 20 to 50, 40 to 80, 50 to 100, 100 to 1000, 500 to 2000, or 1000 to 4700 base pairs. In some embodiments, the methods comprise removing at least one gene, or a fragment thereof, e.g., one or more exons or fragments thereof, from the host genome or the target genome. In some embodiments, the methods comprise introducing a foreign gene (also defined herein as a foreign DNA sequence or gene of interest) or a fragment thereof into the host genome or the target genome. The HITI genome editing methods disclosed herein have increased capabilities for modifying genomic DNA in dividing and non-dividing cells. Non-dividing cells include, but are not limited to, cells of the central nervous system, including neurons, oligodendrocytes, microglia, and ependymal cells; sensory transduction cells; autonomic nerve cells; sensory and peripheral nerve support cells; retinal cells, including photoreceptors, rods, and cones; kidney cells, including mural cells, glomerular podocytes, proximal tubule brush border cells, Henle's loop tubule cells, distal tubule cells, and collecting duct cells; hematopoietic cells, including lymphocytes, monocytes, neutrophils, eosinophils, basophils, and platelets. Preferred non-dividing cells of the present invention are hepatocytes, including hepatocytes, stellate cells, Kupffer cells, and hepatic endothelial cells, and preferably hepatocytes. In some embodiments, the HITI genome editing method disclosed herein provides a method for modifying genomic DNA in dividing cells, and the method has higher efficiency than previous methods disclosed in the art. In some embodiments, the donor nucleic acid, complementary oligonucleotide, and / or polynucleotide encoding the nuclease for the HITI methods described herein are introduced into target or host cells by a virus. In some embodiments, the virus infects the target cell and expresses the targeting construct, complementary oligonucleotide, and nuclease, allowing the foreign DNA of the targeting construct to be integrated into the host genome.In some embodiments, the virus comprises a Sendai virus, retrovirus, lentivirus, baculovirus, adenovirus, or adeno-associated virus. In some embodiments, the virus is a pseudotyped virus. In some embodiments, the donor nucleic acid, complementary oligonucleotide, and / or polynucleotide encoding a nuclease for the HITI method described herein are introduced into target or host cells by non-viral gene delivery methods. In some embodiments, non-viral gene delivery methods deliver genetic material (including DNA, RNA, and proteins) into target cells, express the donor nucleic acid, complementary oligonucleotide, and nuclease, and allow the foreign DNA of the donor nucleic acid to be integrated into the host genome. In some embodiments, non-viral methods include transfection reagents (including nanoparticles) for DNA, mRNA, or protein, or electroporation.

[0266] <Methods for treating diseases> Also provided herein are methods and compositions for treating diseases such as genetic diseases. Genetic diseases are diseases caused by genetic DNA mutations. In some embodiments, genetic diseases are caused by genomic DNA mutations. Genetic mutations are known by those skilled in the art and include single base pair changes or point mutations, insertions, and deletions. In some embodiments, the methods provided herein include methods for treating a genetic disease in a subject in need thereof, the genetic disease being caused by a mutant gene having at least one altered nucleotide compared to a wild-type gene, the method comprising contacting at least one cell of the subject with a composition comprising a DNA construct, a vector (e.g., a non-viral or viral vector), or a system according to the present invention, in which a donor nucleic acid comprising a foreign DNA sequence and, optionally, one or more exons of an albumin gene, a complementary oligonucleotide homologous to a target sequence (e.g., a gRNA homologous to the target sequence), and a nuclease that recognizes the target sequence are introduced into the cell, the target sequence being located at the 3' end of the albumin gene in a region selected from intron 9, intron 11, intron 12, intron 13, and intron 14 of the albumin gene. The donor DNA is then inserted into the target locus by NHEJ, the albumin gene is rearranged, and a therapeutic gene is expressed in the target cell under the control of the albumin promoter.

[0267] Genetic diseases treated by the methods disclosed herein include, but are not limited to, lysosomal storage diseases, including mucopolysaccharidoses (MPS I, MPS II, MPS IIIA, MPS IIIB, MPS IIIC, MPS I V A, MPS I V B, MPS V II), sphingolipidoses (Fabry disease, Gaucher disease, Niemann-Pick disease, GM1 gangliosidosis), lipofuscinoses (Batten disease and others), and mucolipidoses; other diseases in which the liver can be used as a factory for the production and / or secretion of therapeutic proteins, such as diabetes, gyrus chorioretinal atrophy, adenylosuccinic acid deficiency, hemophilia A and B, ALA dehydrogenase deficiency, adrenoleukodystrophy.

[0268] The term "genome editing" refers to a type of genetic engineering that uses one or more nucleases and / or nickases to insert, replace, or remove DNA into target DNA, e.g., the genome of a cell.

[0269] The term "non-homologous end joining" or "NHEJ" refers to a pathway for repairing double-stranded DNA breaks in which the broken ends are directly joined, without the need for a homologous template.

[0270] The terms "polynucleotide," "oligonucleotide," "nucleic acid," "nucleotide," and "nucleic acid molecule" can be used interchangeably and refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and polymers thereof in single-, double-, or multi-stranded form. The term includes, but is not limited to, single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and / or pyrimidine bases, or other naturally occurring, chemically modified, biochemically modified, non-natural, synthetic, or derivatized bases. It also includes modifications, such as methylation and / or capping, as well as unmodified forms of polynucleotides. More specifically, the terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule" refer to polydeoxyribonucleotides (containing 2-deoxy-D-ribose), polyribonucleotides (containing D-ribose), any other type of polynucleotide (containing N- or C-glycosides of purine or pyrimidine bases and non-nucleotide backbones, such as polyamide (e.g., peptide nucleic acid (PNA)) and polymorpholino (commercially available as Niugene from Anti-Virals, Inc., Corvallis, Oregon) polymers), as well as other synthetic sequence-specific nucleic acid polymers (polymers containing nucleobases in a configuration that allows for base pairing and base stacking, such as that found in DNA and RNA). In some embodiments, nucleic acids can include mixtures of DNA and RNA and their analogs. Unless otherwise specified, the term encompasses nucleic acids containing known natural nucleotide analogs that have similar binding properties to the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise specified, a particular nucleotide sequence also implicitly encompasses conservatively modified variants (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), and complementary sequences, as well as the explicitly stated sequence. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues.The term nucleic acid is used interchangeably with gene, cDNA, and mRNA encoded by a gene. The term "gene" or "nucleic acid sequence encoding a polypeptide" refers to a segment of DNA involved in producing a polypeptide chain. The DNA segment can include regions before and after the coding region (leader and trailer) involved in transcription / translation of the gene product and regulation of transcription / translation, as well as intervening sequences (introns) between individual coding segments (exons).

[0271] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. These terms apply to amino acid polymers in which one or more amino acid residues are artificial chemical mimics of corresponding naturally occurring amino acids, as well as naturally occurring amino acid polymers and non-natural amino acid polymers. As used herein, these terms encompass any length of amino acid chain, including full-length proteins in which amino acid residues are linked by covalent peptide bonds.

[0272] A "recombinant expression vector" is a recombinantly or synthetically produced nucleic acid construct that has a set of specified nucleic acid elements that allow for transcription of a specific polynucleotide sequence in a host cell. An expression vector can be part of a plasmid, a viral genome, or a nucleic acid fragment. Typically, an expression vector contains a polynucleotide to be transcribed operably linked to a promoter.

[0273] As used herein, the term "administering" includes oral administration, topical contact, administration as a suppository, intravenous administration, intraperitoneal administration, intramuscular administration, intralesional administration, intrathecal administration, nasal administration, or subcutaneous administration to a subject. Administration can be by any route, including parenteral and transmucosal (e.g., buccal, sublingual, palatal, gingival, nasal, vaginal, rectal, or transdermal). Parenteral administration includes, for example, intravenous, intramuscular, intraarterial, intradermal, subcutaneous, intraperitoneal, intraventricular, and intracranial. Other delivery modes include, but are not limited to, the use of liposomal formulations, intravenous infusion, transdermal patches, and the like.

[0274] The term "treating" refers to an approach to obtaining beneficial or desired results, including, but not limited to, therapeutic benefit and / or prophylactic benefit. Therapeutic benefit refers to any therapeutically relevant improvement or effect in one or more diseases, conditions, or symptoms under treatment. Slowing the progression of a disease is considered a therapeutic improvement within the meaning of this invention. For prophylactic benefit, the composition can be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject who reports one or more physiological symptoms of a disease, even though the disease, condition, or symptom may not yet be manifest. The term "effective amount" or "sufficient amount" refers to an amount of an agent (e.g., a DNA nuclease) sufficient to produce a beneficial or desired result. A therapeutically effective amount may vary depending on one or more of the subject and disease state being treated, the subject's weight and age, the severity of the disease state, the mode of administration, etc., and can be readily determined by one of ordinary skill in the art. The particular amount may vary depending on one or more of the particular agent selected, the type of target cell, the location of the target cell within the subject, the administration regimen to be followed, whether it is administered in combination with other compounds, the timing of administration, and the physical delivery system in which it is delivered.

[0275] The term "pharmaceutically acceptable carrier" refers to a substance that aids in the administration of a drug (e.g., a DNA nuclease) to a cell, organism, or subject. A "pharmaceutically acceptable carrier" refers to a carrier or excipient that can be included in a composition or formulation and does not cause significant adverse toxic effects to the patient. Non-limiting examples of pharmaceutically acceptable carriers include water, NaCl, normal saline, lactated Ringer's solution, normal sucrose, normal glucose, binders, fillers, disintegrants, lubricants, coating agents, sweeteners, flavors, and colorants. Those skilled in the art will recognize that other pharmaceutical carriers are useful in the present invention.

[0276] <Variants, derivatives, analogs, and fragments> In addition to the specific proteins and nucleic acids described herein, the present invention also encompasses variants, derivatives, and fragments thereof.

[0277] In the context of the present invention, a "variant" of any given sequence is a sequence in which a specific sequence of residues (either amino acid or nucleic acid residues) has been modified so that the polypeptide or polynucleotide in question retains at least one of its endogenous functions. Variant sequences can be obtained by adding, deleting, replacing, modifying, substituting and / or changing at least one residue present in the naturally occurring polypeptide or polynucleotide.

[0278] The term "derivative" as used herein in relation to a protein or polypeptide of the present invention includes any substitution, alteration, modification, replacement, deletion and / or addition of one (or more than one) amino acid residue from the sequence, provided that the resulting protein or polypeptide retains at least one of its endogenous functions.

[0279] Typically, amino acid substitutions can be made, for example from 1, 2 or 3 to 10 or 20 substitutions, as long as the altered sequence retains the required activity or ability. Amino acid substitutions can include the use of non-naturally occurring analogues.

[0280] The proteins used in the present invention may also have deletions, insertions, or substitutions of amino acid residues, resulting in silent changes and functionally equivalent proteins. Deliberate amino acid substitutions can be made based on similarities in polarity, charge, solubility, hydrophobicity, hydrophilicity, and / or amphipathicity of the residues, as long as the intrinsic function is maintained. For example, negatively charged amino acids include aspartic acid and glutamic acid; positively charged amino acids include lysine and arginine; and amino acids with uncharged polar head groups with similar hydrophilicity values ​​include asparagine, glutamine, serine, threonine, and tyrosine.

[0281] Conservative substitutions may be made, for example, according to the following table: Amino acids in the same block in the second column and in the same line in the third column may be substituted for each other.

[0282] [Table 2]

[0283] Typically, the variant may have a certain identity to the wild-type amino acid sequence or the wild-type nucleotide sequence.

[0284] As used herein, a variant sequence is taken to include an amino acid sequence that may have at least 50%, 55%, 65%, 75%, 85% or 90% identity to the subject sequence, preferably at least 95%, 96%, 97%, 98% or 99% identity. Although variants can also be considered in terms of similarity (i.e., amino acid residues that have similar chemical properties / functions), in the context of the present invention they are preferably expressed in terms of sequence identity.

[0285] As used herein, a variant sequence is taken to include a nucleotide sequence which may have at least 50%, 55%, 65%, 75%, 85% or 90% identity to the subject sequence, preferably at least 95%, 96%, 97%, 98% or 99% identity. Although variants can also be considered in terms of similarity, in the context of the present invention they are preferably expressed in terms of sequence identity.

[0286] Suitably, reference to a sequence having a given percent identity with any one of the SEQ ID NOs detailed herein refers to a sequence having the stated percent identity over the entire length of the referenced SEQ ID NO.

[0287] Sequence identity comparisons can be performed by eye, or more usually, with the aid of readily available sequence comparison programs. These commercially available computer programs can calculate the percent identity between two or more sequences.

[0288] Percent identity can be calculated over contiguous sequences, i.e., one sequence is aligned with the other, and each amino acid or nucleotide in one sequence is directly compared to the corresponding amino acid or nucleotide in the other sequence, one residue at a time. This is called an "ungapped" alignment. Typically, such ungapped alignments are performed over only a relatively small number of residues.

[0289] While this is a very simple and consistent method, it does not take into account the possibility that, in a pair of identical sequences, a single insertion or deletion of an amino acid or nucleotide sequence may cause the subsequent residue or codon to misalign, resulting in a significant decrease in percent identity when performing a global alignment. As a result, most sequence comparison methods are designed to produce optimal alignments that take into account possible insertions and deletions without excessively penalizing the overall identity score. This is achieved by inserting "gaps" into the sequence alignment to attempt to maximize local identity.

[0290] However, these more complex methods assign a "gap penalty" to each gap that occurs in the alignment, such that, for the same number of identical amino acids or nucleotides, a sequence alignment with as few gaps as possible will achieve a higher score than one with many gaps, reflecting a higher relatedness between the two compared sequences. "Affine gap costs" are typically used, which impose a relatively high cost for the presence of a gap and a relatively low penalty for each subsequent residue after the gap. This is the most commonly used gap scoring system. Higher gap penalties will naturally result in an optimized alignment with fewer gaps. Most alignment programs allow for the modification of gap penalties. However, it is preferable to use the default values ​​when comparing sequences using such software. For example, when using the Bestfit package from GCG Wisconsin, the default gap penalty for amino acid sequences is -12 for a single gap and -4 for each extension.

[0291] Therefore, calculating maximum percent identity requires first generating optimal alignment, taking into account gap penalties.A suitable computer program for carrying out such alignment is the Bestfit package of GCG Wisconsin (University of Wisconsin, USA; Devereux et al. (1984) Nucleic Acids Research 12:387).Other software examples that can carry out sequence comparison include but are not limited to the BLAST package (see Ausubel et al. (1999) ibid-Ch.18), FASTA (Atschul et al. (1990) J.Mol.Biol.403-410), EMBOSS Needle (Madeira, F., et al., 2019.Nucleic acids research, 47(W1), pp.W636-W641) and GENEWORKS comparison tool suite. Both BLAST and FASTA are available for offline and online searching (see Ausubel et al. (1999) ibid; pp. 7-58 to 7-60). However, for some applications it is preferable to use the GCG Bestfit program. Another tool, BLAST 2 Sequences, is also available for comparing protein and nucleic acid sequences (FEMS Microbiol. Lett. (1999) 174(2):247-50; FEMS Microbiol. Lett. (1999) 177(1):187-8).

[0292] Although a final percent identity can be measured, the alignment process itself is typically not based on an all-or-nothing pairwise comparison. Instead, a scaled similarity score matrix is ​​commonly used that assigns a score to each pairwise comparison based on chemical similarity or evolutionary distance. One example of such a matrix commonly used is the BLOSUM62 matrix (the default matrix for the BLAST suite of programs). GCG Wisconsin programs typically use either the public default values ​​or a custom symbol comparison table if provided (see user manual for details). For some applications, it is preferable to use the public default values ​​for the GCG package or a default matrix such as BLOSUM62 for other software.

[0293] Once the software has produced an optimal alignment, it is possible to calculate the percent sequence identity. The software typically does this as part of the sequence comparison and generates a numerical result. The percent sequence identity can be calculated as the number of identical residues as a percentage of the total number of residues in the referenced SEQ ID NO.

[0294] "Fragments" are also variants, and the term generally refers to a selected region of a polypeptide or polynucleotide of interest functionally or, for example, in an assay. A "fragment" therefore refers to an amino acid or nucleic acid sequence that is a portion of a full-length polypeptide or polynucleotide.

[0295] Such variants, derivatives, and fragments can be prepared using standard recombinant DNA techniques such as site-directed mutagenesis. When an insertion is to be made, synthetic DNA encoding the insertion can be made along with the 5' and 3' flanking regions corresponding to the naturally occurring sequences on both sides of the insertion site. The flanking regions contain convenient restriction sites corresponding to sites in the naturally occurring sequences so that the sequences can be cut with appropriate enzymes and the synthetic DNA can be ligated to the cut sites. Then, the DNA is expressed according to the present invention to produce the encoded protein. These methods are merely exemplary of a number of standard techniques for the manipulation of DNA sequences, and other known techniques can also be used.

[0296] The present invention is illustrated by the following examples.

Examples

[0297] <Materials and Methods> <Plasmids Used as Cas9 Templates> The plasmid used for AAV vector production is derived from the pAAV2.1 plasmid containing the inverted terminal repeats of AAV serotype 2.

[0298] The AAV vector plasmid required to generate AAV-SpCas9 contains a hybrid liver promoter (HLP) and a synthetic pA sequence.

[0299] The AAV vector plasmid required to generate AAV-gRNA-donorDsRed contains a U6 promoter, a specific gRNA and PAM sequence, and a chimeric gRNA scaffold; a splicing acceptor signal, exon 14 of mAlb, a T2A linker, a DsRed coding sequence [CDS (NCBI ref.MK301207.1)], a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), a bovine growth hormone polyA (BGH polyA), and a stop codon surrounded by an inverted gRNA and PAM sequence.

[0300] The AAV vector plasmid required to generate AAV-gRNA-donorARSB contains a U6 promoter, a specific gRNA and PAM sequence, and a chimeric gRNA scaffold; a splicing acceptor signal, exon 14 of mAlb, a T2A linker, the human ARSB CDS (NCBI ref.NM_000046.5), BGH polyA, and a stop codon surrounded by an inverted gRNA and PAM sequence.

[0301] The AAV vector plasmid required to generate AAV-gRNA-Cas9 contains a U6 promoter, a specific gRNA, and a chimeric gRNA scaffold; a hybrid liver promoter (HLP), spCas9 and a synthetic pA sequence.

[0302] The AAV vector plasmid required to generate AAV-donorFVIII contains a splicing acceptor signal, exon 14 of mAlb, a T2A linker, a human FVIII B domain deletion codon-optimized sequence (published in

[33] ), BGH polyA, and a stop codon surrounded by an inverted gRNA and PAM sequence.

[0303] Mouse albumin (mAlb) gRNAs (Table 1, 3) were designed using the Benchling gRNA design tool (www.benchling.com) by selecting the gRNAs with the best predicted on-target and off-target scores that target intron 13 of mAlb or intron 12, 13 or 14 of human albumin (hALB). Scrambled gRNAs were designed to not align with any sequences in the mouse genome.

[0304] <Production and Characterization of AAV Vectors> AAV vectors were produced by the TIGEM AAV Vector Core by two CsCl2 purifications following triple transfection of HEK293 cells

[34] . For each viral preparation, the physical titer (GC / mL) was determined by averaging the titers achieved by dot blot analysis

[39] and PCR quantification using TaqMan (Applied Biosystems, Carlsbad, CA, USA)

[34] . The probes used for dot blot and PCR analysis were designed to anneal to the IRBP promoter for the pAAV2.1-IRBP-SpCas9-spA vector, the HLP promoter for the pAAV2.1-HLP-SpCas9-spA vector, and the bGHpA region for the donor DNA vector. The length of the probes varied between 200 and 700 bp.

[0305] <Culture and Transfection of HEK293 Cells> HEK293 cells were maintained in DMEM containing 10% fetal bovine serum (FBS) and 2 mM L-glutamine (Gibco, Thermo Fisher Scientific, Waltham, MA, USA). Cells were seeded in 6-well plates (1×10 6 cells / well) and transfected 16 h later with plasmids encoding Cas9 and various gRNAs and donor DNA by the calcium phosphate method (1–2 μg / 1×10 6 cells); the medium was changed 4 h later. The maximum amount of material transfected was 3 μg. In all cases, the amount of plasmid DNA was equalized between wells using empty vectors as needed.

[0306] <Flow Cytometry Analysis> HEK293 cells seeded in 6-well plates were washed once with PBS, detached with trypsin 0.05% EDTA (Thermo Fisher Scientific, Waltham, MA, USA), washed twice with PBS, and resuspended in sorting solution containing PBS, 5% FBS, and 2.5 mM EDTA. Cells were analyzed on a BD FACS ARIA III (BD Biosciences, San Jose, CA, USA) equipped with BD FACSDiva software (BD Biosciences) using appropriate excitation and detection settings for EGFP and DsRed. The threshold for fluorescence detection was set on untransfected cells, and a minimum of 10,000 cells were analyzed per sample. A minimum of 50,000 GFP+ or GFP+ / DsRed+ cells per sample were sorted and used for DNA extraction.

[0307] <Mouse liver frozen section and fluorescent image> To evaluate DsRed expression in the liver after HITI, C57BL / 6 mice were injected at p1 and sacrificed by cardiac perfusion 1 month after injection. Small pieces of each liver lobe were dissected and fixed overnight in 4% paraformaldehyde. After fixation, the pieces were infiltrated in 15% sucrose for 1 day and 30% sucrose overnight before being embedded in OCT matrix (Kaltek, Padua, Italy) and cryosectioned. Liver cryosections were cut at 6 μm thickness, distributed onto slides, and mounted with Vectashield supplemented with DAPI (Vector Lab, Peterborough, UK, #H-1200). Cryosections were then analyzed at 20x magnification under a confocal microscope LSM 700 (Leica Microsystems, Wetzlar, Germany) using appropriate excitation and detection settings.

[0308] Characterization of integrated binding DNA was extracted from liver tissue using the DNeasy Blood and Tissue Kit (Qiagen, Hilden, Germany) according to the manufacturer's protocol.

[0309] For the T7 cleavage assay, 100 ng of DNA was used to PCR amplify the region containing the Cas9 target site in mouse Alb intron 13 using specific primers (Table 2). These primers generate a 652-bp PCR product. The PCR product was examined by T7 endonuclease I assay according to the manufacturer's recommendations. Briefly, DNA was deannealed and reannealed using a gentle temperature gradient in a thermocycler using NEBuffer2 (New England BioLabs, Ipswich, MA, USA). Samples were then incubated with 1 μL of T7 endonuclease (NEB, #M0302L) at 37°C for 30 minutes and analyzed on a 2% agarose gel. PCR products were also used for Sanger sequencing (Eurofins Genomics, Ebersberg, Germany) and then processed using SYNTHEGO software (https: / / ice.synthego.com / # / ) to analyze indel frequencies.

[0310] To detect donor integration into the 3' mouse albumin locus (end of intron 13), 100 ng of extracted DNA was used for PCR amplification of the HITI junction using specific primers (Table 2). PCR products were analyzed on a 2% agarose gel and further cloned in PCR-Blunt II-TOPO (Invitrogen, Carlsbad, CA, USA). Single clones were then used for Sanger sequencing (Eurofins Genomics, Ebersberg, Germany) to characterize donor integration.

[0311] <Plasma collection and F8 assay> Nine parts of blood were collected retroorbitally in one part of buffered trisodium citrate 0.109 M (5T31.363048; BD, Franklin Lakes, NJ, USA). Plasma samples were collected after centrifugation at 3000 rpm for 15 min at 4°C.

[0312] To assess F8 activity, a chromogenic assay was performed on plasma samples using the Coatest™ SP4 FVIII kit (K824094; Chromogenix, Werfen, Milan, Italy) according to the manufacturer's instructions. A standard curve was generated by serial dilutions of commercially available human F8 (Refacto, Pfizer). Results are expressed as international units (IU) per deciliter (dl).

[0313] <Result> [Example 1] HITI-mediated DsRed integration into the mouse 3'Alb (mAlb) locus in neonatal wild-type mice As a proof-of-concept, we performed in vivo experiments to knock in a reporter DsRed transgene 3' to the mAlb locus in wild-type neonatal mice (Figure 1A). To this end, we generated three different AAV8 vectors: one vector encoding SpCas9 under the expression of the hybrid liver promoter (HLP); one vector containing the HITI donor DsRed coding sequence (CDS); and one vector containing either a U6-gRNA or U6-scRNA expression cassette. Specifically, the donor DNA cassette contains a synthetic splicing acceptor signal (SAS), the last albumin exon (ex14), and the coding sequence for DsRed followed by a T2A sequence. The sequences are shown below.

[0314] Wild-type (WT) C57BL / 6 mice were divided into two treatment groups (gRNA or scRNA) and administered a mixture of vectors at a 1:1 ratio via the temporal vein on postnatal day 1 (p1). To integrate the DsRed CDS into the 3' mAlb locus, the gRNA group was injected with a vector encoding SpCas9 along with a U6 promoter and gRNA sequence, and a vector carrying a HITI donor. As a negative control, the scRNA group was treated according to the same experimental scheme, except that the vector carrying the HITI donor contained a U6-scRNA expression cassette. To evaluate the cleavage efficiency of SpCas9 and PCR-amplify the integration junction using specific primers (Table 2), animals were sacrificed 4 weeks after injection, and DNA was extracted from liver samples. As expected, SpCas9 cleavage occurred only in gRNA-treated animals, but not in the scRNA group (Figure 1B). Furthermore, PCR analysis demonstrated 5' and 3' ligation products at the target site in gRNA-treated animals but not in scRNA-treated animals (Figure 1C). Consistent with these results, microscopic images of liver cryosections revealed that DsRed was highly expressed in gRNA-treated animals but completely absent in scRNA-treated controls (Figure 1D). Taken together, all these data demonstrate that HITI is a suitable platform for the integration and expression of DsRed at the 3' mAlb locus.

[0315] [Example 2] HITI-mediated ARSB delivery in neonatal MPSVI mice Next, the inventors verified whether the level of lysosomal hydrolase arylsulfatase B (ARSB), which is deficient in lysosomal storage disease mucopolysaccharidosis VI (MPS VI), could be stabilized and therapeutically appropriate by HITI at the 3’mAlb locus of neonatal mice. Since ARSB is secreted into the bloodstream and can be measured non-invasively, it can be used as an indicator of liver transduction

[35] . ARSB deficiency results in the accumulation and urinary secretion of abnormal glycosaminoglycans (GAGs), which are useful biomarkers for MPS VI

[36] . The inventors generated an AAV vector carrying the above donor DNA cassette: a synthetic splicing acceptor signal (SAS), the last albumin exon (ex14), the coding sequence of human ARSB (hARSB) followed by a T2A sequence, and a gRNA expression cassette containing either a gRNA or a scrambled sequence as a control (Figure 2A). The gRNA donor vector or the scRNA vector was co-administered systemically to neonatal MPS VI mice (p1-2) in combination with the HLP-SpCas9 vector (Figure 2A). Serum ARSB activity was measured in gRNA-treated MPS VI mice at levels higher than those of normal littermates (Figure 2B) and remained stable over time until 1 year of age. Serum ARSB activity was not detected in scrambled- or untreated MPS VI mice. Importantly, although no significant difference in urinary GAG was observed between the scrambled and gRNA treatment groups at p60, AAV-HITI-mediated ARSB expression was able to normalize urinary GAG from p90 to p360 (Figure 2C).

[0316] <AAV-HITI Dose Response> The inventors investigated arylsulfatase B (Arsb) in mucopolysaccharidosis VI (MPSVI) - / - We tested three doses of AAV-HITI to treat a mouse model by integrating donor DNA containing the promoterless coding sequence of ARSB into the 3' albumin locus. Animals were administered three doses of AAV-HITI on days 1-2 of age: 1.2E+14 total genome copies (GC) / kg (high dose or HD), 3.9E+13 total GC / kg (medium dose or MD), and 1.2E+13 total GC / kg (low dose or LD). Preliminary results from our ARSB immunoassay on serum samples indicate that HD and MD achieve supraphysiological levels of secreted active ARSB by 30 days of age (Figure 7). LD also induced secretion of active ARSB at various levels (Figure 7).

[0317] [Example 3] HITI-mediated F8 codopV3 integration into the mouse 3'Alb (mAlb) locus in neonatal hemophilic mice.

[0318] The inventors conducted in vivo experiments to knock-in the F8 CodopV3 transgene into the 3’mAlb locus of neonatal hemophilic mice. For this purpose, three different AAV8 vectors were generated: two vectors encoding SpCas9 with U6-gRNA or U6-scRNA expression cassettes under the expression of a hybrid liver promoter (HLP); and one vector containing the HITI donor F8CodopV3 coding sequence (CDS). Specifically, the donor DNA cassette included a synthetic splicing acceptor signal (SAS), the last albumin exon (ex14), followed by a T2A sequence and the coding sequence of F8 CodopV3 (Figure 3A). Hemophilic mice were divided into two different treatment groups (gRNA or scRNA), and a mixture of vectors was administered at a ratio of 1:1 via the jugular vein on the first day after birth (p1). To integrate the F8 CDS into the 3’mAlb locus, the gRNA group was injected with a vector encoding SpCas9 together with the U6 promoter and gRNA sequence, and a vector carrying the HITI donor. As a negative control, the scRNA group was treated according to the same experimental scheme, and the vector carrying SpCas9 contained a U6-scRNA expression cassette. Plasma samples were collected 4 weeks after vector administration.

[0319] When F8 activity was monitored using a functional chromogenic assay, the F8 activity level was shown to be 20% compared to non-affected controls (Figure 3B).

[0320] <gRNA sequence> [Table 3]

[0321] [Table 4]

[0322] gRNAs with the best predicted on-target and off-target scores targeting the 9th, 11th, 12th, 13th, or 14th intron of human albumin (hALB) were designed and reported in Table 3.

[0323] [Table 5]

[0324] The gRNAs were designed on either strand of the genomic DNA and are shown in the 5'-3' orientation. The sequences are shown in the 5'-3' orientation.

[0325] <Serum albumin level> Serum albumin levels were measured from blood samples collected at p360 from treated and control mice using an ELISA kit (Abcam, 108791, Cambridge, UK) according to the manufacturer's instructions (Figure 4). Serum albumin levels were found to be similar regardless of treatment group, indicating that our AAV-HITI does not affect endogenous protein expression.

[0326] <Alpha-fetoprotein level> Elevated alpha-fetoprotein (AFP) levels have been reported to be associated with hepatocellular carcinoma (HCC) in mice (Ferla et al., Molecular Therapy: Methods & Clinical Development 2021). We measured AFP levels in serum samples from p360-treated and control mice using a Mouse Alpha-fetoprotein / AFP Quantikine Elisa Kit (R&D Systems, Minneapolis, MN, USA) according to the manufacturer's instructions. Mouse AFP levels were increased in AAV-HITI gRNA-treated mice but not in scRNA-treated or control mice (Figure 5).

[0327] <Off-target analysis> To investigate potential chromosomal aberrations (as translocation events) through on-target sites (OMT), we performed CAST-seq analysis, a previously described technique (Turchiano et al., 2021), on our AAV-HITI gRNA liver DNA samples, while AAV-HITI scRNA and untreated liver DNA samples served as controls. The CAST-seq analysis data showed that our AAV-HITI gRNA samples exhibited deletion events at the on-target sites, but no OMT was observed (Figure 6).

[0328] [Example 4] Selection of gRNAs targeting the 3' human albumin (ALB) locus Using Benchling and / or CHOPCHOP software, we selected one gRNA targeting intron 13 of ALB and eight gRNAs targeting introns 9 and 11–13 of ALB (Table 4). In-silico selection was based on i) a low number of predicted off-targets and ii) high efficiency of targeting the desired locus. A plasmid encoding Cas9-EGFP under the CBh promoter and one of the selected gRNAs or scRNAs under the human U6 promoter were transfected into HEPA 1-6 cells or HEK293 cells to target the Alb or ALB locus, respectively. HEPA 1-6 cells were transfected with 1 μg of plasmid DNA using Lipofectamine LTX (Thermo Fisher Scientific, Waltham, MA, USA), while HEK293 cells were transfected with 1 μg of plasmid DNA using calcium phosphate. DNA was extracted from sorted cells expressing Cas9-EGFP, and the genomic region recognized by the gRNA was PCR-amplified. The PCR product was digested with T7 enzyme (Neb, Ipswich, MA, USA) to detect Cas9-mediated INDELs. The same PCR products were Sanger sequenced and INDELs were quantified using Synthego's ICE software. gRNA0 (ALB intron 13) and gRNAs 3 and 5 (ALB intron 12) induced high Cas9-mediated INDELs, while lower levels were detected using gRNA2 (ALB intron 13). No INDELs were detected using either gRNA1 (ALB intron 13) or gRNA4 (ALB intron 14) (Table 4 and Figure 8). The allelic variant frequencies of the sequences recognized by gRNAs targeting the ALB locus were analyzed for selected gRNAs using the Aggregated Human Genome Database (gnomAD) version 3.1.2. The highest detected allelic variant frequency was 10. 3 There is one SNP per allele (gRNAs 3 and 6) and, importantly, no homozygous mutations.

[0329] <Integration efficiency of the 3'Alb locus and 3'ALB locus> We evaluated the efficiency of HITI-mediated integration into the 3'Alb and 3'ALB loci by generating HITI donors flanked by the inverted sequences of gRNA0 for integration into the 3'Alb locus or gRNA3 or 5 for integration into the 3'ALB locus. The donors encoded a synthetic splice acceptor signal, exon 14 of Alb or exons 13-14 of ALB, linked via a T2A skipping peptide to the fluorescent reporter DsRed coding sequence. We transfected HEPA1-6 cells with 1 μg of plasmid DNA encoding Cas9-EGFP and gRNA0 and 1 μg of plasmid DNA encoding donor DNA flanking gRNA0 using Lipofectamine LTX (Thermo Fisher Scientific, Waltham, MA, USA). Similarly, we transfected human hepatoma cell line 7 (HUH7) with 1 μg of plasmid encoding Cas9-EGFP and gRNA3 plus 1 μg of plasmid encoding a HITI donor flanking gRNA3, or 1 μg of plasmid encoding Cas9-EGFP and gRNA5 plus 1 μg of plasmid encoding a HITI donor flanking gRNA5. Cells transfected with the HITI donor and Cas9-EGFP-encoding plasmid DNA and scRNA were used to normalize DsRed fluorescence and quantify only productive HITI donor integration. Fluorescence-activated cell sorting analysis showed that gRNA0 and gRNA3 induced productive integration of the HITI donor at the 3′Alb and 3′ALB loci, respectively (Figure 9 ).

[0330] [Table 6]

[0331] [Example 5] Accuracy of the AAV-HITI platform at the 3' mAlb locus To evaluate the accuracy of our AAV-HITI strategy, we performed several molecular analyses. First, we examined the cleavage efficiency (indel %) of our selected gRNAs. We performed Illumina-seq NGS analysis on genomic DNA extracted from the livers of AAV-HITI gRNA- or AAV-HITI scRNA-treated MPSVI mice. We found 29% indels only in AAV-HITI gRNA-treated mice (Figure 10A). Furthermore, we also evaluated whether we could find the portion of the AAV vector genome integrated into the on-target site upon Cas9-induced double-strand break. Using the entire AAV vector genome of either the HITI donor DNA or Cas9 as a reference sequence, we aligned the reads generated from the Illumina-seq NGS experiment. We were able to align reads covering different portions of a given AAV vector genome, with the majority of reads covering the ITR region (Figure 10B). Next, we evaluated HITI-mediated integration compared with ITR-mediated integration. To this end, we generated a donor DNA (ITR donor) with the same structure as the HITI gRNA donor DNA but without inverted gRNA sites flanking its 5' and 3' ends. This construct was then produced as AAV8 and used in vivo. Wild-type mice were injected via the temporal vein at p1-2 with a mixture of AAV-Cas9 and the AAV-ITR donor containing the Ds-Red coding sequence (as previously described for the HITI donor DNA). In parallel, a second group of mice was injected with a combination of AAV-Cas9 and the AAV-HITI gRNA donor. Both groups were sacrificed one month after treatment, and DNA was extracted from the livers of all treated animals and used for further molecular analysis.

[0332] Specific primers (Table 2) were used to PCR amplify both the 5' and 3' junction sites between the inserted donor DNA and the endogenous locus.

[0333] As seen on the agarose gel (Figure 10C), PCR amplification of junction bands (5' and 3') of the expected size was achieved according to the donor (HITI or ITR donor). Interestingly, in HITI-treated mice, a fainter upper band was also observed, similar in size to that observed in the AAV-ITR donor gRNA-treated samples. Sanger sequencing analysis of both the upper band and the band of the expected size revealed that donor DNA integration also occurred via the ITR in HITI donor DNA-treated animals.

[0334] Next, we evaluated potential gRNA off-target activity. To this end, we selected the top 10 predicted off-targets using CRISPOR (Table 5). NGS analysis of PCR bands obtained from liver genomic DNA revealed very low or undetectable off-target editing events for each of the selected off-target loci (both in gRNA- and scRNA-treated samples) (Figure 10D).

[0335] Table 5 [Table 7]

[0336] <array> <Sequence of Example 1 above> 5'-ITR

[0337] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT [SEQ ID NO: 110]

[0338] Mouse albumin intron 13 gRNA sequence

[0339] CACTGCTGCCTATTAAATAC [SEQ ID NO: 1]

[0340] scRNA sequences

[0341] gactcgcgcgagtcgaggag [SEQ ID NO: 111]

[0342] Inverted gRNA sequence of mouse albumin intron 13 without PAM

[0343] CACTGCTGCCTATTAAATAC [SEQ ID NO: 1]

[0344] Inverted gRNA sequence of mouse albumin intron 13 + PAM sequence (underlined)

[0345] CCACACTGCTGCCTATTAAATAC [SEQ ID NO: 20]

[0346] splice acceptor sequence

[0347] GATAGGCACCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAG [SEQ ID NO: 21]

[0348] Exon 14 mouse albumin

[0349] GGTCCAAACCTTGTCACTAGATGCAAAGACGCCTTAGCC [SEQ ID NO: 22]

[0350] Thosea asigna virus 2A (T2A) skipping peptide

[0351] GGAAGCGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTCGAGGAGAATCCTGGACCT [SEQ ID NO: 23]

[0352] Discosoma Red (DsRed) coding sequence

[0353] ATGGATAGCACTGAGAACGTCATCAAGCCCTTCATGCGCTTCAAGGTGCACATGGAGGGCTCCGTGAACGGCCACGAGTTCGAGATCGAGGGCGAGGGCGAGGGCAAGCCCTACGAGGGCACCCAGACCGCCAAGCTGCAGGTGACCAAGGGCGGCCCCCTGCCCTTCGCCTGGGACATCCTGTCCCCCCAGTTCCAGTACGGCTCCAAGGTGTACGTGAAGCACCCCGCCGACATCCCCGACTACAAGAAGCTGTCCTTCCCCGAGGGCTTCAAGTGGGAGCGCGTGATGAACTTCGAGGACGGCGGCGTGGTGACCGTGACCCAGGACTCCTCCCTGCAGGACGGCACCTTCATCTACCACGTGAAGTTCATCGGCGTGAACTTCCCCTCCGACGGCCCCGTAATGCAGAAGAAGACTCTGGGCTGGGAGCCCTCCACCGAGCGCCTGTACCCCCGCGACGGCGTGCTGAAGGGCGAGATCCACAAGGCGCTGAAGCTGAAGGGCGGCGGCCACTACCTGGTGGAGTTCAAGTCAATCTACATGGCCAAGAAGCCCGTGAAGCTGCCCGGCTACTACTACGTGGACTCCAAGCTGGACATCACCTCCCACAACGAGGACTACACCGTGGTGGAGCAGTACGAGCGCGCCGAGGCCCGCCACCACCTGTTCCAGTAG [SEQ ID NO: 24]

[0354] Woodchuck hepatitis virus post - transcriptional regulatory element (WPRE)

[0355] aatcaacctctggattacaaaatttgtgaaagattgactggtattcttaactatgttgctccttttacgctatgtggatacgctgctttaatgcctttgtatcatgctattgcttcccgtatggctttcattttctcctccttgtataaatcctggttgctgtctctttatgaggagttgtggcccgttgtcaggcaacgtggcgtggtgtgcactgtgtttgctgacgcaacccccactggttggggcattgccaccacctgtcagctcctttccgggactttcgctttccccctccctattgccacggcggaactcatcgccgcctgccttgcccgctgctggacaggggctcggctgttgggcactgacaattccgtggtgttgtcggggaaatcatcgtcctttccttggctgctcgcctgtgttgccacctggattctgcgcgggacgtccttctgctacgtcccttcggccctcaatccagcggaccttccttcccgcggcctgctgccggctctgcggcctcttccgcgtcttcg [SEQ ID NO: 25]

[0356] Bovine Growth Hormone PolyA (BGH pA)

[0357] GCCTCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGA [SEQ ID NO: 26]

[0358] Human U6 Promoter

[0359] tttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaaagatattagtacaaaatacgtgacgtagaaagta ataatttcttgggtagtttgcagttttaaaattatgttttaaaatggactatcatatgcttaccgtaacttgaaagtatttcgatttcttggctttatatatcttgtggaaaggacg [Sequence number 27]

[0360] Chimeric RNA scaffolds

[0361] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC [SEQ ID NO: 28]

[0362] 3'-ITR

[0363] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAG [SEQ ID NO: 29]

[0364] Construct p1492_pTIGEM_mAlb3'HITI donor (SAS_albex14_T2A_dsRED_bGHpA) + U6gRNA mAlb3'

[0365]

[0366] Construct p1496_pTIGEM_mAlb3'HITI donor (SAS_albex14_T2A_dsRED_bGHpA) + U6 scrambled RNA mAlb3'

[0367]

[0368] <Sequence of Example 2> 5'-ITR

[0369] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT [SEQ ID NO: 110]

[0370] Mouse albumin intron 13 gRNA sequence (positive strand 5'-3' direction)

[0371] CACTGCTGCCTATTAAATAC [SEQ ID NO: 1]

[0372] scRNA sequences

[0373] gactcgcgcgagtcgaggag [SEQ ID NO: 111]

[0374] Inverted gRNA sequence of mouse albumin intron 13 without PAM

[0375] CACTGCTGCCTATTAAATAC [SEQ ID NO: 1]

[0376] Inverted gRNA sequence of mouse albumin intron 13 + PAM sequence

[0377] CCACACTGCTGCCTATTAAATAC [SEQ ID NO: 20]

[0378] splicing acceptor sequence

[0379] GATAGGCACCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAG [SEQ ID NO: 21]

[0380] Exon 14 mouse albumin

[0381] GGTCCAAACCTTGTCACTAGATGCAAAGACGCCTTAGCC [SEQ ID NO: 22]

[0382] Thosea asigna virus 2A (T2A) skipping peptide

[0383] GGAAGCGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTCGAGGAGAATCCTGGACCT [SEQ ID NO: 23]

[0384] ARSB coding sequence

[0385]

[0386] Bovine growth hormone polyA (BGH pA)

[0387] GCCTCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGA [SEQ ID NO: 26]

[0388] Human U6 promoter

[0389] tttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaaagatattagtacaaaatacgtgacgtagaaagta ataatttcttgggtagtttgcagttttaaaattatgttttaaaatggactatcatatgcttaccgtaacttgaaagtatttcgatttcttggctttatatatcttgtggaaaggacg [Sequence number 27]

[0390] Chimeric RNA scaffolds

[0391] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC [SEQ ID NO: 28]

[0392] 3'-ITR

[0393] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAG [SEQ ID NO: 29]

[0394] Construct p1479_pTIGEM_mAlb3'HITI donor (SAS_albex14_T2A_ARSB_bGHpA) + U6gRNA mAlb3'

[0395]

[0396] Construct p1480_pTIGEM_mAlb3'HITI donor (SAS_albex14_T2A_ARSB_bGHpA) + U6scrRNA mAlb3'

[0397]

[0398] <Sequence of Example 3> 5'-ITR

[0399] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT [SEQ ID NO: 110]

[0400] Inverted gRNA sequence of mouse albumin intron 13 without PAM

[0401] CACTGCTGCCTATTAAATAC [SEQ ID NO: 1]

[0402] Inverted gRNA sequence of mouse albumin intron 13 + PAM sequence (underlined)

[0403] CCACACTGCTGCCTATTAAATAC [SEQ ID NO: 20]

[0404] splicing acceptor sequence

[0405] GATAGGCACCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAG [SEQ ID NO: 21]

[0406] Exon 14 mouse albumin

[0407] GGTCCAAACCTTGTCACTAGATGCAAAGACGCCTTAGCC [SEQ ID NO: 22]

[0408] Thosea asigna virus 2A (T2A) skipping peptide

[0409] GGAAGCGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTCGAGGAGAATCCTGGACCT [SEQ ID NO: 23]

[0410] F8_CodopV3 code sequence

[0411]

[0412] Synthetic polyadenylation signal

[0413] tcgcgaataaaagatctttattttcattagatctgtgtgttggttttttgtgtgatgcagc [SEQ ID NO: 37]

[0414] 3'ITR

[0415] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAG [SEQ ID NO: 29]

[0416] Construct p1493_pTIGEM_mAlb3'HITI donor (SAS_albex14_T2A_CodopV3_pA)

[0417]

[0418] Construct_p1139_pAAV2.1._HLP_SpCas9(HA)_spA The underlined sequence is the promoter sequence. Cas9 / Cas9-2a-GFP

[0419]

[0420] HITI3'mALb f8 (hemophilia A) p1498_pAAV_HLP_SpCas9+U6 3'malb_gRNA(5,1kb) 5'ITR

[0421] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT (SEQ ID NO: 110)

[0422] Additional AAV sequences

[0423] Tgtagttaatgattaacccgccatgctacttatctacgtagagctcttgtcgaggtcgacatttaaatgaagggcgaattcaattt (SEQ ID NO: 44)

[0424] U6 expression cassette gRNA

[0425] Gagggcctatttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaaagatattagtacaaaatacgtgacgtagaaagtaataatttcttgggtagtttgcagtttaaaattatgttttaaaatggactatcata tgcttaccgtaacttgaaagtatttcgatttcttggctttatatatcttgtggaaaggacgaaacaccgtatttaataggcagcagtggttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgcttttttgttttagagcta (Sequence number 45)

[0426] HLP promoter

[0427] TGTTTGCTGCTTGCAATGTTTGCCCATTTTAGGGTGGACACAGGACGCTGTGGTTTCTGAGCCAGGGGGCGACTCAGATCCCAGCCAGTGGACTTAGCCCCTGTTTGCTCCTCCGATAACTGGGGTGACCTTGGTTAATATTCACCAGCAGCCTCCCCCGTTGCCCCTCTGGATCCACTGCTTAAATACGGACGAGGACAGGGCCCTGTCTCCTCAGCTTCAGGCACCACCACTGACCTGGGACAGTGAAT (SEQ ID NO: 46)

[0428] spCas9

[0429]

[0430] Synthetic PolyA

[0431] Aattcaataaaagatctttattttcattagatctgtgtgttggttttttgtgtgcggcc (SEQ ID NO: 48)

[0432] 3'ITR

[0433] Aggaaccccctagtgatggagttggccactccctctctgcgcgctcgctcgctcactgaggccgggcgaccaaaggtcgcccgacgcccgggctttgcccgggcggcctcagtgagcgagcgagcgcgcag (SEQ ID NO: 29)

[0434] Full sequence p1498_pAAV_HLP_SpCas9+U6 3'malb_gRNA(5,1kb)

[0435]

[0436] p1500 / / pAAV_HLP_SpCas9+U6 3'malb_scRNA 5'ITR

[0437] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT (SEQ ID NO: 110)

[0438] Additional AAV sequences

[0439] Tgtagttaatgattaacccgccatgctacttatctacgtagagctcttgtcgaggtcgac (SEQ ID NO: 50)

[0440] U6 expression cassette scRNA

[0441] Ctgacctcgagtttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaaagatattagtacaaaatacgtgacgtagaaagtaataatttcttgggtagtttgcagttttaaaattatgttttaaaatggactatcatatgctt accgtaacttgaaagtatttcgatttcttggctttatatatcttgtggaaaggacgaaacaccggactcgcgcgagtcgaggaggttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgctttttgttttagagctagaaatagcaag (Sequence number 51)

[0442] HLP promoter

[0443] TGTTTGCTGCTTGCAATGTTTGCCCATTTTAGGGTGGACACAGGACGCTGTGGTTTCTGAGCCAGGGGCGACTCAGATCCCAGCCAGTGGACTTAGCCCCTGTTTGCTCCTCCGATAACTGGGGTGACCTTGGTTAATATTCACCAGCAGCCCTCCCCGTTGCCCCTTGGATCCCACTGCTTAAATACGGACGGACAGGCCCTGTCTCCTCAGCTTCAGCTTCAGCCACCACCACTGACCTGGGACAGTGGAATGAAT (sequence number 46)

[0444] spCas9

[0445]

[0446] Synthetic polyA

[0447] Aattcaataaaagatctttattttcattagatctgtgtgttggttttttgtgtgcggcc (SEQ ID NO: 48)

[0448] 3’ ITR

[0449] Gcaggaacccctagtgatggagttggccactccctctctgcgcgctcgctcgctcactgaggccgggcgaccaaaggtcgcccgacgcccgggctttgcccgggcggcctcagtgagcgagcgagcgcgcag (SEQ ID NO: 29)

[0450] Full sequence p1500

[0451] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT

[0452] P1617_pTIGEM_HITI 3'malb CodopV3 HITI donor 5'ITR

[0453] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT (SEQ ID NO: 110)

[0454] Inverted gRNA+PAM site

[0455] GTATTTAATAGGCAGCAGTGTGG (SEQ ID NO: 54)

[0456] Synthetic splicing acceptors

[0457] GATAGGCACCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAG (SEQ ID NO: 21)

[0458] mAlbumin exon 14

[0459] ggtccaaaccttgtcactagatgcaaagacgccttagcc (SEQ ID NO: 22)

[0460] T2A sequence

[0461] Ggaagcggagagggcagaggaagtctgctaacatgcggtgacgtcgaggagaatcctggacct (SEQ ID NO: 23)

[0462] Codopv3

[0463]

[0464] 3XFLAG

[0465] GACTACAAAGACCATGACGGTGATTATAAAGATCATGACATCGACTACAAGGATGACGATGACAAGTGA (SEQ ID NO: 56)

[0466] Synthetic PolyA

[0467] Tcgcgaataaaagatctttattttcattagatctgtgtgttggttttttgtgtgatgcagc (SEQ ID NO: 37)

[0468] Inverted gRNA and PAM

[0469] gtatttaataggcagcagtgtgg (SEQ ID NO: 54)

[0470] Additional AAV sequences

[0471] GAGCTCTTGTCGAGGTCGACATTTAAATGAATTCCAATTG (SEQ ID NO: 57)

[0472] 3'itr

[0473] Aggaaccccctagtgatggagttggccactccctctctgcgcgctcgctcgctcactgaggccgggcgaccaaaggtcgcccgacgcccgggctttgcccgggcggcctcagtgagcgagcgagcgcgcag (SEQ ID NO: 29)

[0474]

[0475] <Sequence of Example 4>

[0476] HITI3' human ALb Construct p939_pCbh-SpCas9(BB)-2A-GFP + scrambled gRNA

[0477] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACC (Sequence number 59)

[0478] Human U6 promoter

[0479] Gactcgcgcgagtcgaggag (SEQ ID NO: 111)

[0480] Scrambled RNA

[0481] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTT (SEQ ID NO: 60)

[0482] Chimeric RNA scaffolds

[0483] cgttacataacttacggtaaatggcccgcctggctgaccgcccaacgacccccgcccattgacgtcaatagtaacgccaatagggactttccattgacgtcaatgggtggagtatttacggtaaactgcccacttggcagtacatcaagtgtatcatatgccaagtacgccccctattgacgtcaatgacggtaaatggcccgcctggcattgtgcccagtacatgaccttatgggactttcctacttggcagtacatctacgtattagtcatcgctattaccatggtcgaggtgagccccacgttctgcttcactctccccatctcccccccctccccacccccaattttgtatttatttattttttaattattttgtgcagcgatgggggcggggggggggggggggcgcgcgccaggcggggcggggcggggcgaggggcggggcggggcgaggcggagaggtgcggcggcagccaatcagagcggcgcgctccgaaagtttccttttatggcgaggcggcggcggcggcggccctataaaaagcgaagcgcgcggcgggcgggagtcgctgcgacgctgccttcgccccgtgccccgctccgccgccgcctcgcgccgcccgccccggctctgactgaccgcgttactcccacaggtgagcgggcgggacggcccttctcctccgggctgtaattagctgagcaagaggtaagggtttaagggatggttggttggtggggtattaatgtttaattacctggagcacctgcctgaaatcactttttttcaggttgg (SEQ ID NO: 61)

[0484] CBH promoter

[0485] ATGGACTATAAGGACCACGACGGAGACTACAAGGATCATGATATTGATTACAAAGACGATGACGATAAG (SEQ ID NO: 62)

[0486] 3x Flag Tags

[0487]

[0488] 5' nuclear localization signal + SpCas9 + 3' nuclear localization signal

[0489] GGCAGTGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTCGAGGAGAATCCTGGCCCA (SEQ ID NO: 63)

[0490] Thosea Asigna virus T2A skipping peptide

[0491] GTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTACAAGGAATTCTAA (SEQ ID NO: 64)

[0492] EGFP fusion protein

[0493] Ctagagctcgctgatcagcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattgggaagagaatagcaggcatgctgggga (SEQ ID NO: 65)

[0494] BGH polyA

[0495] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAGCTGCCTGCAGG (SEQ ID NO: 66)

[0496] 3’ ITR

[0497]

[0498] Construct p1526_pCbh-SpCas9(BB)-2A-GFP+3'human albumin gRNA1

[0499] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACC (Sequence number 59)

[0500] Human U6 promoter

[0501] Aatctctggacggaagctca (SEQ ID NO: 10)

[0502] gRNA1 human albumin

[0503] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTT (SEQ ID NO: 60)

[0504] Chimeric RNA scaffolds

[0505] Cgttacataacttacggtaaatggcccgcctggctgaccgcccaacgacccccgcccattgacgtcaatagtaacgccaatagggactttccattgacgtcaatgggtggagtatttacggtaaactgcccacttggcagtacatcaagtgtatcatatgccaagtacgccccctattgacgtcaatgacggtaaatggcccgcctggcattgtgcccagtacatgaccttatgggactttcctacttggcagtacatctacgtattagtcatcgctattaccatggtcgaggtgagccccacgttctgcttcactctccccatctcccccccctccccacccccaattttgtatttatttattttttaattattttgtgcagcgatgggggcggggggggggggggggcgcgcgccaggcggggcggggcggggcgaggggcggggcggggcgaggcggagaggtgcggcggcagccaatcagagcggcgcgctccgaaagtttccttttatggcgaggcggcggcggcggcggccctataaaaagcgaagcgcgcggcgggcgggagtcgctgcgacgctgccttcgccccgtgccccgctccgccgccgcctcgcgccgcccgccccggctctgactgaccgcgttactcccacaggtgagcgggcgggacggcccttctcctccgggctgtaattagctgagcaagaggtaagggtttaagggatggttggttggtggggtattaatgtttaattacctggagcacctgcctgaaatcactttttttcaggttgg (SEQ ID NO: 61)

[0506] CBH promoter

[0507] ATGGACTATAAGGACCACGACGGAGACTACAAGGATCATGATATTGATTACAAAGACGATGACGATAAG (SEQ ID NO: 62)

[0508] 3x Flag Tags

[0509]

[0510] 5' nuclear localization signal + SpCas9 + 3' nuclear localization signal

[0511] GGCAGTGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTCGAGGAGAATCCTGGCCCA (SEQ ID NO: 63)

[0512] Thosea Asigna virus T2A skipping peptide

[0513] GTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTACAAGGAATTCTAA

[0514] EGFP fusion protein

[0515] Ctagagctcgctgatcagcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattgggaagagaatagcaggcatgctgggga (SEQ ID NO: 65)

[0516] BGH PolyA

[0517] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAGCTGCCTGCAGG (SEQ ID NO: 66)

[0518] 3’ ITR

[0519]

[0520] Construct p1530_pCbh-SpCas9(BB)-2A-GFP+3'human albumin gRNA2

[0521] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACC (Sequence number 59)

[0522] Human U6 promoter

[0523] Acagtatggcacaatagagc (SEQ ID NO: 12)

[0524] gRNA2 human albumin

[0525] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTT (SEQ ID NO: 60)

[0526] Chimeric RNA scaffolds

[0527] cgttacataacttacggtaaatggcccgcctggctgaccgcccaacgacccccgcccattgacgtcaatagtaacgccaatagggactttccattgacgtcaatgggtggagtatttacggtaaactgcccacttggcagtacatcaagtgtatcatatgccaagtacgccccctattgacgtcaatgacggtaaatggcccgcctggcattgtgcccagtacatgaccttatgggactttcctacttggcagtacatctacgtattagtcatcgctattaccatggtcgaggtgagccccacgttctgcttcactctccccatctcccccccctccccacccccaattttgtatttatttattttttaattattttgtgcagcgatgggggcggggggggggggggggcgcgcgccaggcggggcggggcggggcgaggggcggggcggggcgaggcggagaggtgcggcggcagccaatcagagcggcgcgctccgaaagtttccttttatggcgaggcggcggcggcggcggccctataaaaagcgaagcgcgcggcgggcgggagtcgctgcgacgctgccttcgccccgtgccccgctccgccgccgcctcgcgccgcccgccccggctctgactgaccgcgttactcccacaggtgagcgggcgggacggcccttctcctccgggctgtaattagctgagcaagaggtaagggtttaagggatggttggttggtggggtattaatgtttaattacctggagcacctgcctgaaatcactttttttcaggttgg

[0528] CBH promoter

[0529] ATGGACTATAAGGACCACGACGGAGACTACAAGGATCATGATATTGATTACAAAGACGATGACGATAAG (SEQ ID NO: 62)

[0530] 3x Flag Tags

[0531]

[0532] 5' nuclear localization signal + SpCas9 + 3' nuclear localization signal

[0533] GGCAGTGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTCGAGGAGAATCCTGGCCCA (SEQ ID NO: 63)

[0534] Thosea Asigna virus T2A skipping peptide

[0535] GTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTACAAGGAATTCTAA

[0536] EGFP fusion protein

[0537] Ctagagctcgctgatcagcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattgggaagagaatagcaggcatgctgggga (SEQ ID NO: 65)

[0538] BGH polyA

[0539] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAGCTGCCTGCAGG (SEQ ID NO: 66)

[0540] 3’ ITR

[0541]

[0542] Construct p1531_pCbh-SpCas9(BB)-2A-GFP+3'human albumin gRNA3

[0543] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACC (Sequence number 59)

[0544] Human U6 promoter

[0545] Acactacataacgtgatgag (SEQ ID NO: 13)

[0546] gRNA3 human albumin

[0547] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTT (SEQ ID NO: 60)

[0548] Chimeric RNA scaffolds

[0549] cgttacataacttacggtaaatggcccgcctggctgaccgcccaacgacccccgcccattgacgtcaatagtaacgccaatagggactttccattgacgtcaatgggtggagtatttacggtaaactgcccacttggcagtacatcaagtgtatcatatgccaagtacgccccctattgacgtcaatgacggtaaatggcccgcctggcattgtgcccagtacatgaccttatgggactttcctacttggcagtacatctacgtattagtcatcgctattaccatggtcgaggtgagccccacgttctgcttcactctccccatctcccccccctccccacccccaattttgtatttatttattttttaattattttgtgcagcgatgggggcggggggggggggggggcgcgcgccaggcggggcggggcggggcgaggggcggggcggggcgaggcggagaggtgcggcggcagccaatcagagcggcgcgctccgaaagtttccttttatggcgaggcggcggcggcggcggccctataaaaagcgaagcgcgcggcgggcgggagtcgctgcgacgctgccttcgccccgtgccccgctccgccgccgcctcgcgccgcccgccccggctctgactgaccgcgttactcccacaggtgagcgggcgggacggcccttctcctccgggctgtaattagctgagcaagaggtaagggtttaagggatggttggttggtggggtattaatgtttaattacctggagcacctgcctgaaatcactttttttcaggttgg

[0550] CBH promoter

[0551] ATGGACTATAAGGACCACGACGGAGACTACAAGGATCATGATATTGATTACAAAGACGATGACGATAAG (SEQ ID NO: 62)

[0552] 3x Flag Tags

[0553]

[0554] 5' nuclear localization signal + SpCas9 + 3' nuclear localization signal

[0555] GGCAGTGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTCGAGGAGAATCCTGGCCCA (SEQ ID NO: 63)

[0556] Thosea Asigna virus T2A skipping peptide

[0557] GTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTACAAGGAATTCTAA

[0558] EGFP fusion protein

[0559] Ctagagctcgctgatcagcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattgggaagagaatagcaggcatgctgggga (SEQ ID NO: 65)

[0560] BGH polyA

[0561] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAGCTGCCTGCAGG (SEQ ID NO: 66)

[0562] 3’ ITR

[0563]

[0564] Construct p1532_pCbh-SpCas9(BB)-2A-GFP+3'human albumin gRNA4

[0565] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACC (Sequence number 59)

[0566] Human U6 promoter

[0567] Aaatagtttagaatagtggt (SEQ ID NO: 14)

[0568] gRNA4 human albumin

[0569] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTT (SEQID NO:60)

[0570] Chimeric RNA scaffolds

[0571] cgttacataacttacggtaaatggcccgcctggctgaccgcccaacgacccccgcccattgacgtcaatagtaacgccaatagggactttccattgacgtcaatgggtggagtatttacggtaaactgcccacttggcagtacatcaagtgtatcatatgccaagtacgccccctattgacgtcaatgacggtaaatggcccgcctggcattgtgcccagtacatgaccttatgggactttcctacttggcagtacatctacgtattagtcatcgctattaccatggtcgaggtgagccccacgttctgcttcactctccccatctcccccccctccccacccccaattttgtatttatttattttttaattattttgtgcagcgatgggggcggggggggggggggggcgcgcgccaggcggggcggggcggggcgaggggcggggcggggcgaggcggagaggtgcggcggcagccaatcagagcggcgcgctccgaaagtttccttttatggcgaggcggcggcggcggcggccctataaaaagcgaagcgcgcggcgggcgggagtcgctgcgacgctgccttcgccccgtgccccgctccgccgccgcctcgcgccgcccgccccggctctgactgaccgcgttactcccacaggtgagcgggcgggacggcccttctcctccgggctgtaattagctgagcaagaggtaagggtttaagggatggttggttggtggggtattaatgtttaattacctggagcacctgcctgaaatcactttttttcaggttgg

[0572] <000,1975>CBH promoter

[0573] ATGGACTATAAGGACCACGACGGAGACTACAAGGATCATGATATTGATTACAAAGACGATGACGATAAG (SEQ ID NO: 62)

[0574] 3x Flag Tags

[0575]

[0576] 5' nuclear localization signal + SpCas9 + 3' nuclear localization signal

[0577] GGCAGTGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTCGAGGAGAATCCTGGCCCA (SEQ ID NO: 63)

[0578] Thosea Asigna virus T2A skipping peptide

[0579] GTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTACAAGGAATTCTAA

[0580] EGFP fusion protein

[0581] Ctagagctcgctgatcagcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattgggaagagaatagcaggcatgctgggga (SEQ ID NO: 65)

[0582] BGH polyA

[0583] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAGCTGCCTGCAGG (SEQ ID NO: 66)

[0584] 3’ ITR

[0585]

[0586] Construct p1556_pCbh-SpCas9(BB)-2A-GFP+3'human albumin gRNA5

[0587] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACC (Sequence number 59)

[0588] Human U6 promoter

[0589] Gtgggctgtaatcatcgtct (SEQ ID NO: 16)

[0590] gRNA5 human albumin

[0591] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTT (SEQ ID NO: 60)

[0592] Chimeric RNA scaffolds

[0593] cgttacataacttacggtaaatggcccgcctggctgaccgcccaacgacccccgcccattgacgtcaatagtaacgccaatagggactttccattgacgtcaatgggtggagtatttacggtaaactgcccacttggcagtacatcaagtgtatcatatgccaagtacgccccctattgacgtcaatgacggtaaatggcccgcctggcattgtgcccagtacatgaccttatgggactttcctacttggcagtacatctacgtattagtcatcgctattaccatggtcgaggtgagccccacgttctgcttcactctccccatctcccccccctccccacccccaattttgtatttatttattttttaattattttgtgcagcgatgggggcggggggggggggggggcgcgcgccaggcggggcggggcggggcgaggggcggggcggggcgaggcggagaggtgcggcggcagccaatcagagcggcgcgctccgaaagtttccttttatggcgaggcggcggcggcggcggccctataaaaagcgaagcgcgcggcgggcgggagtcgctgcgacgctgccttcgccccgtgccccgctccgccgccgcctcgcgccgcccgccccggctctgactgaccgcgttactcccacaggtgagcgggcgggacggcccttctcctccgggctgtaattagctgagcaagaggtaagggtttaagggatggttggttggtggggtattaatgtttaattacctggagcacctgcctgaaatcactttttttcaggttgg

[0594] CBH promoter

[0595] ATGGACTATAAGGACCACGACGGAGACTACAAGGATCATGATATTGATTACAAAGACGATGACGATAAG (SEQ ID NO: 62)

[0596] 3x Flag Tags

[0597]

[0598] 5' nuclear localization signal + SpCas9 + 3' nuclear localization signal

[0599] GGCAGTGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTCGAGGAGAATCCTGGCCCA (SEQ ID NO: 63)

[0600] Thosea Asigna virus T2A skipping peptide

[0601] GTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTACAAGGAATTCTAA [[ID=​​​​​Ctagagctcgctgatcagcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattgggaagagaatagcaggcatgctgggga (SEQ ID NO: 65)

[0604] BGH PolyA

[0605] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAGCTGCCTGCAGG (SEQ ID NO: 66)

[0606] 3’ ITR

[0607]

[0608] Construct p1545_pTIGEM_hALB3'HITI donor (SAS_albex13_ex14_T2A_dsRED_WPRE_bGHpA)

[0609] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT (SEQ ID NO: 110)

[0610] 5'-ITR

[0611] CCACTCATCACGTTATGTAGTGT (SEQ ID NO: 77)

[0612] Inverted gRNA sequence of human albumin intron 12 + PAM sequence

[0613] Gataggcacctattggtcttactgacatccactttgcctttctctccacag (SEQ ID NO: 21)

[0614] splicing acceptor sequence

[0615] TGCACTTGTTGAGCTCGTGAAACACAAGCCCAAGGCAACAAAAGAGCAACTGAAAGCTGTTATGGATGATTTCGCAGCTTTTGTAGAGAAGTGCTGCAAGGCTGACGATAAGGAGACCTGCTTTGCCGAGGAG (SEQ ID NO: 78)

[0616] Exon 13 human albumin

[0617] Ggtaaaaaacttgttgctgcaagtcaagctgccttaggctta (SEQ ID NO: 79)

[0618] Exon 14 human albumin

[0619] GGAAGCGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTCGAGGAGAATCCTGGACCT (SEQ ID NO: 23)

[0620] Thosea asigna virus 2A (T2A) skipping peptide

[0621] ATGGATAGCACTGAGAACGTCATCAAGCCCTTCATGCGCTTCAAGGTGCACATGGAGGGCTCCGTGAACGGCCACGAGTTCGAGATCGAGGGCGAGGGCGAGGGCAAGCCCTACGAGGGCACCCAGACCGCCAAGCTGCAGGTGACCAAGGGCGGCCCCCTGCCCTTCGCCTGGGACATCCTGTCCCCCCAGTTCCAGTACGGCTCCAAGGTGTACGTGAAGCACCCCGCCGACATCCCCGACTACAAGAAGCTGTCCTTCCCCGAGGGCTTCAAGTGGGAGCGCGTGATGAACTTCGAGGACGGCGGCGTGGTGACCGTGACCCAGGACTCCTCCCTGCAGGACGGCACCTTCATCTACCACGTGAAGTTCATCGGCGTGAACTTCCCCTCCGACGGCCCCGTAATGCAGAAGAAGACTCTGGGCTGGGAGCCCTCCACCGAGCGCCTGTACCCCCGCGACGGCGTGCTGAAGGGCGAGATCCACAAGGCGCTGAAGCTGAAGGGCGGCGGCCACTACCTGGTGGAGTTCAAGTCAATCTACATGGCCAAGAAGCCCGTGAAGCTGCCCGGCTACTACTACGTGGACTCCAAGCTGGACATCACCTCCCACAACGAGGACTACACCGTGGTGGAGCAGTACGAGCGCGCCGAGGCCCGCCACCACCTGTTCCAGTAG

[0622] Discosoma Red (DsRed) coding sequence

[0623] aatcaacctctggattacaaaatttgtgaaagattgactggtattcttaactatgttgctccttttacgctatgtggatacgctgctttaatgcctttgtatcatgctattgcttcccgtatggctttcattttctcctccttgtataaatcctggttgctgtctctttatgaggagttgtggcccgttgtcaggcaacgtggcgtggtgtgcactgtgtttgctgacgcaacccccactggttggggcattgccaccacctgtcagctcctttccgggactttcgctttccccctccctattgccacggcggaactcatcgccgcctgccttgcccgctgctggacaggggctcggctgttgggcactgacaattccgtggtgttgtcggggaaatcatcgtcctttccttggctgctcgcctgtgttgccacctggattctgcgcgggacgtccttctgctacgtcccttcggccctcaatccagcggaccttccttcccgcggcctgctgccggctctgcggcctcttccgcgtcttcg

[0624] Woodchuck hepatitis virus post - transcriptional regulatory element (WPRE)

[0625] gcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattgggaagacaatagcaggcatgctgggga (SEQ ID NO: 26)

[0626] Bovine growth hormone polyA (BGH pA)

[0627] CCACTCATCACGTTATGTAGTGT (SEQ ID NO: 77)

[0628] Inverted gRNA sequence of human albumin intron 12 + PAM sequence

[0629] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAG (SEQ ID NO: 29)

[0630] 3'-ITR

[0631]

[0632] Construct p1615_pCbh-SpCas9(BB)-2A-GFP+3'human albumin gRNA6

[0633] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACC (Sequence number 59)

[0634] Human U6 promoter

[0635] Tattggcagtcaaggccccg (SEQ ID NO: 17)

[0636] gRNA6 human albumin

[0637] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTT (SEQ ID NO: 60)

[0638] Chimeric RNA scaffolds

[0639] cgttacataacttacggtaaatggcccgcctggctgaccgcccaacgacccccgcccattgacgtcaatagtaacgccaatagggactttccattgacgtcaatgggtggagtatttacggtaaactgcccacttggcagtacatcaagtgtatcatatgccaagtacgccccctattgacgtcaatgacggtaaatggcccgcctggcattgtgcccagtacatgaccttatgggactttcctacttggcagtacatctacgtattagtcatcgctattaccatggtcgaggtgagccccacgttctgcttcactctccccatctcccccccctccccacccccaattttgtatttatttattttttaattattttgtgcagcgatgggggcggggggggggggggggcgcgcgccaggcggggcggggcggggcgaggggcggggcggggcgaggcggagaggtgcggcggcagccaatcagagcggcgcgctccgaaagtttccttttatggcgaggcggcggcggcggcggccctataaaaagcgaagcgcgcggcgggcgggagtcgctgcgacgctgccttcgccccgtgccccgctccgccgccgcctcgcgccgcccgccccggctctgactgaccgcgttactcccacaggtgagcgggcgggacggcccttctcctccgggctgtaattagctgagcaagaggtaagggtttaagggatggttggttggtggggtattaatgtttaattacctggagcacctgcctgaaatcactttttttcaggttgg

[0640] CBH promoter

[0641] ATGGACTATAAGGACCACGACGGAGACTACAAGGATCATGATATTGATTACAAAGACGATGACGATAAG (SEQ ID NO: 62)

[0642] 3x Flag Tags

[0643]

[0644] 5' nuclear localization signal + SpCas9 + 3' nuclear localization signal

[0645] GGCAGTGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTCGAGGAGAATCCTGGCCCA (SEQ ID NO: 63)

[0646] Thosea Asigna virus T2A skipping peptide

[0647] GTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTACAAGGAATTCTAA

[0648] EGFP fusion protein

[0649] Ctagagctcgctgatcagcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattgggaagagaatagcaggcatgctgggga (SEQ ID NO: 65)

[0650] BGH polyA

[0651] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAGCTGCCTGCAGG (SEQ ID NO: 66)

[0652] 3’ ITR

[0653]

[0654] Construct p1616_pCbh-SpCas9(BB)-2A-GFP+3'human albumin gRNA7

[0655] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACC (Sequence number 59)

[0656] Human U6 promoter

[0657] Tcgaatgtattgtgacagag (SEQ ID NO: 18)

[0658] gRNA7 human albumin

[0659] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTT (SEQ ID NO: 60)

[0660] Chimeric RNA scaffolds

[0661] cgttacataacttacggtaaatggcccgcctggctgaccgcccaacgacccccgcccattgacgtcaatagtaacgccaatagggactttccattgacgtcaatgggtggagtatttacggtaaactgcccacttggcagtacatcaagtgtatcatatgccaagtacgccccctattgacgtcaatgacggtaaatggcccgcctggcattgtgcccagtacatgaccttatgggactttcctacttggcagtacatctacgtattagtcatcgctattaccatggtcgaggtgagccccacgttctgcttcactctccccatctcccccccctccccacccccaattttgtatttatttattttttaattattttgtgcagcgatgggggcggggggggggggggggcgcgcgccaggcggggcggggcggggcgaggggcggggcggggcgaggcggagaggtgcggcggcagccaatcagagcggcgcgctccgaaagtttccttttatggcgaggcggcggcggcggcggccctataaaaagcgaagcgcgcggcgggcgggagtcgctgcgacgctgccttcgccccgtgccccgctccgccgccgcctcgcgccgcccgccccggctctgactgaccgcgttactcccacaggtgagcgggcgggacggcccttctcctccgggctgtaattagctgagcaagaggtaagggtttaagggatggttggttggtggggtattaatgtttaattacctggagcacctgcctgaaatcactttttttcaggttgg

[0662] CBH promoter

[0663] ATGGACTATAAGGACCACGACGGAGACTACAAGGATCATGATATTGATTACAAAGACGATGACGATAAG (SEQ ID NO: 62)

[0664] 3x Flag Tags

[0665]

[0666] 5' nuclear localization signal + SpCas9 + 3' nuclear localization signal

[0667] GGCAGTGGAGAGGGCAGAGGAAGTCTGCTAACATGCGGTGACGTCGAGGAGAATCCTGGCCCA (SEQ ID NO: 63)

[0668] Thosea Asigna virus T2A skipping peptide

[0669] GTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTACAAGGAATTCTAA

[0670] EGFP fusion protein

[0671] Ctagagctcgctgatcagcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattgggaagagaatagcaggcatgctgggga (SEQ ID NO: 65)

[0672] BGH polyA

[0673] AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAGCTGCCTGCAGG (SEQ ID NO: 66)

[0674] 3’ ITR

[0675]

[0676] <Sequence of Example 5 above> Sequence of the ITR donor DNA construct (donor DNA_p1547 without the 5' and 3' inverted gRNA sites)

[0677] 5'ITR

[0678] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT (SEQ ID NO: 110)

[0679] Additional aav sequences

[0680] Gctagtgctagc (SEQ ID NO: 84)

[0681] SAS (Splice Acceptor Signal)

[0682] Gataggcacctattggtcttactgacatccactttgcctttctctccacag (SEQ ID NO: 21)

[0683] Mouse albumin exon 14

[0684] ggtccaaaccttgtcactagatgcaaagacgccttagcc (SEQ ID NO: 22)

[0685] T2A sequence

[0686] Ggaagcggagagggcagaggaagtctgctaacatgcggtgacgtcgaggagaatcctggacct (SEQ ID NO: 23)

[0687] Ds-Red coding sequence

[0688] Atggatagcactgagaacgtcatcaagcccttcatgcgcttcaaggtgcacatggagggctccgtgaacggccacgagttcgagatcgagggcgagggcgagggcaagccctacgagggcacccagaccgccaagctgcaggtgaccaagggcggccccctgcccttcg cctgggacatcctgtccccccagttccagtacggctccaaggtgtacgtgaagcaccccgccgacatccccgactacaagaagctgtccttccccgagggcttcaagtgggagcgcgtgatgaacttcgaggacggcggcgtggtgaccgtgacccaggactcctccctg caggacggcaccttcatctaccacgtgaagttcatcggcgtgaacttcccctccgacggccccgtaatgcagaagaagactctgggctgggagccctccaccgagcgcctgtaccccgcgacggcgtgctgaagggcgagatccacaaggcgctgaagctgaagggcg gcggccactacctggtggagttcaagtcaatctacatggccaagaagcccgtgaagctgccccggctactactacgtggactccaagctggacatcacctcccaacgaggactacaccgtggtggagcagtacgagcgcgccgaggcccgccaccacctgttccagtag

[0689] WPRE sequence

[0690] Aatcaacctctggattacaaaatttgtgaaagattgactggtattcttaactatgttgctccttttacgctatgtggatacgctgctttaatgcctttgtatcatgctattgcttcccgtatggctttcattttctcctccttgtataaatcctggttgctgtctctttatgaggagttgtggcccgttgtcaggcaacgtggcgtggtgtgcactgtgtttgctgacgcaacccccactggttggggcattgccaccacctgtcagctcctttccgggactttcgctttccccctccctattgccacggcggaactcatcgccgcctgccttgcccgctgctggacaggggctcggctgttgggcactgacaattccgtggtgttgtcggggaaatcatcgtcctttccttggctgctcgcctgtgttgccacctggattctgcgcgggacgtccttctgctacgtcccttcggccctcaatccagcggaccttccttcccgcggcctgctgccggctctgcggcctcttccgcgtcttcg

[0691] BGH polyA

[0692] GCCTCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGA (SEQ ID NO: 26)

[0693] Additional AAV sequence

[0694] Agagctcttgtcgaggtcgacatttaaatgaagggcgaattcaattt (SEQ ID NO: 85)

[0695] U6 expression cassette

[0696] Gagggcctatttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaaagatattagtacaaaatacgtgacgtagaaagtaataatttcttgggtagtttgcagtttaaaattatgttttaaaatggacta tcatatgcttaccgtaacttgaaagtatttcgatttcttggctttatatatcttgtggaaaggacgaaacaccgtatttaataggcagcagtggttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgcttttttgt (Sequence number 86)

[0697] Additional AAV sequences

[0698] Tttagagctagaaatagcaagttaaaataaggctagtccgtttttagcgcgtgcgccaattctgcagacaaatggctctagaggtaccaattg (SEQ ID NO: 87)

[0699] 3'ITR sequence

[0700] Aggaaccccctagtgatggagttggccactccctctctgcgcgctcgctcgctcactgaggccgggcgaccaaaggtcgcccgacgcccgggctttgcccgggcggcctcagtgagcgagcgagcgcgcagt (SEQ ID NO: 29)

[0701] Full sequence_p1547

[0702] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT

[0703] References [1] Cong, L., et al., Multiplex genome engineering using CRISPR / Cas systems. Science, 2013. 339(6121): p. 819-23. [2] Jiang, W., et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nat Biotechnol, 2013. 31(3): p. 233-9. [3] Tu, Z., et al., CRISPR / Cas9: a powerful genetic engineering tool for establishing large animal models of neurodegenerative diseases. Mol Neurodegener, 2015. 10: p. 35. [4] Nishiyama, J., T. Mikuni, and R. Yasuda, Virus-Mediated Genome Editing via Homology-Directed Repair in Mitotic and Postmitotic Cells in Mammalian Brain. Neuron, 2017. 96(4): p. 755-768 e5. [5] Anguela, X.M., et al., Robust ZFN-mediated genome editing in adult hemophilic mice. Blood, 2013. 122(19): p. 3283-7. [6] Barzel, A., et al., Promoterless gene targeting without nucleases ameliorates haemophilia B in mice. Nature, 2015. 517(7534): p. 360-4. [7]Li, H., et al., In vivo genome editing restores haemostasis in a mouse model of haemophilia. Nature, 2011. 475(7355): p. 217-21. [8]Sharma, R., et al., In vivo genome editing of the albumin locus as a platform for protein replacement therapy. Blood, 2015. 126(15): p. 1777-84. [9]Bakondi, B., In vivo versus ex vivo CRISPR therapies for retinal dystrophy. Expert Rev Ophthalmol, 2016. 11(6): p. 397-400.

[10] Lackner, D.H., et al., A generic strategy for CRISPR-Cas9-mediated gene tagging. Nat Commun, 2015. 6: p. 10237.

[11] Suzuki, K., et al., In vivo genome editing via CRISPR / Cas9 mediated homology-independent targeted integration. Nature, 2016. 540(7631): p. 144-149.

[12] Brunetti-Pierri N, A.A., Gene Therapy of Human Inherited Diseases, in The Metabolic and Molecular Bases of Inherited Diseases, S. R, Editor. 2010, McGraw Hill: New York.

[13] Ehrhardt, A., H. Xu, and M.A. Kay, Episomal persistence of recombinant adenoviral vector genomes during the cell cycle in vivo. J Virol, 2003. 77(13): p. 7689-95.

[14] E Neufeld, J.M., The mucopolysaccharidoses, in The mucopolysaccharidoses, A.B. CR Scriver, WS Sly, DM Valle, Editor. 2001, McGraw-Hill: New York (2001). p. 3421-3452.

[15] Cotugno, G., et al., Impact of age at administration, lysosomal storage, and transgene regulatory elements on AAV2 / 8-mediated rat liver transduction. PLoS One, 2012. 7(3): p. e33286.

[16] Ferla, R., et al., Similar therapeutic efficacy between a single administration of gene therapy and multiple administrations of recombinant enzyme in a mouse model of lysosomal storage disease. Hum Gene Ther, 2014. 25(7): p. 609-18.

[17] Ferla, R., et al., Gene therapy for mucopolysaccharidosis type VI is effective in cats without pre-existing immunity to AAV8. Hum Gene Ther, 2013. 24(2): p. 163-9.

[18] Tessitore, A., et al., Biochemical, pathological, and skeletal improvement of mucopolysaccharidosis VI after gene transfer to liver but not to muscle. Mol Ther, 2008. 16(1): p. 30-7.

[19] Alliegro, M., et al., Low-dose Gene Therapy Reduces the Frequency of Enzyme Replacement Therapy in a Mouse Model of Lysosomal Storage Disease. Mol Ther, 2016. 24(12): p. 2054-2063.

[20] Ferla, R., et al., Non-clinical Safety and Efficacy of an AAV2 / 8 Vector Administered Intravenously for Treatment of Mucopolysaccharidosis Type VI. Mol Ther Methods Clin Dev, 2017. 6: p. 143-158.

[21] Cotugno, G., et al., Long-term amelioration of feline Mucopolysaccharidosis VI after AAV-mediated liver gene transfer. Mol Ther, 2011. 19(3): p. 461-9.

[22] Giugliani, R., et al., Natural history and galsulfase treatment in mucopolysaccharidosis VI (MPS VI, Maroteaux-Lamy syndrome)-10-year follow-up of patients who previously participated in an MPS VI Survey Study. Am J Med Genet A, 2014. 164A(8): p. 1953-64.

[23] Desnick, R.J. and E.H. Schuchman, Enzyme replacement therapy for lysosomal diseases: lessons from 20 years of experience and remaining challenges. Annu Rev Genomics Hum Genet, 2012. 13: p. 307-35.

[24] Neufeld, E.F., Lysosomal storage diseases. Annu Rev Biochem, 1991. 60: p. 257-80.

[25] Bowen, D.J., Haemophilia A and haemophilia B: molecular insights. Mol Pathol, 2002. 55(2): p.127-44. Antonarakis, S.E., et al., Molecular etiology of factor VIII deficiency in hemophilia A. Adv Exp Med Biol, 1995. 386: p. 19-34

[26] Bunting, S., et al.,Gene Therapy with BMN 270 Results in Therapeutic Levels of FVIII in Mice and Primates and Normalization of Bleeding in Hemophilic Mice. Mol Ther, 2018. 26(2): p. 496-509

[27] Rangarajan, S., et al., AAV5-Factor VIII Gene Transfer in Severe Hemophilia A. N Engl J Med, 2017. 377(26): p. 2519-2530

[28] Makris, M. Gene therapy 1 0 in haemophilia: effective and safe, but with many uncertainties. The Lancet Haematology (2020) doi:10.1016 / S2352-3026(20)30035-1

[29] Grieger, J. C. et al., Packaging Capacity of Adeno-Associated Virus Serotypes: Impact of Larger Genomes on Infectivity and Postentry Steps. J. Virol. (2005)

[30] Dong, B. et al., Characterization of genome integrity for oversized recombinant AAV vector. Mol. Ther. (2010)

[31] Hirsch, M. et al., Little vector, big gene transduction: Fragmented genome reassembly of adeno-associated virus. Molecular Therapy (2010)

[32] Wu, Z.et al., Effect of genome size on AAV vector packaging. Mol. Ther. (2010)

[33] McIntosh J, Lenting PJ, Rosales C, Lee D, Rabbanian S, Raj D, Patel N, Tuddenham EG, Christophe OD, McVey JH, Waddington S, Nienhuis AW, Gray JT, Fagone P, Mingozzi F, Zhou SZ, High KA, Cancio M, Ng CY, Zhou J, Morton CL, Davidoff AM, Nathwani AC. Therapeutic levels of FVIII following a single peripheral vein administration of rAAV vector encoding a novel human factor VIII variant. Blood. 2013 Apr 25;121(17):3335-44. doi: 10.1182 / blood-2012-10-462200. Epub 2013 Feb 20. PMID: 23426947; PMCID: PMC3637010.

[34] Doria, M., A. Ferrara, and A. Auricchio, AAV2 / 8 vectors purified from culture medium with a simple and rapid protocol transduce murine liver, muscle, and retina efficiently. Hum Gene Ther Methods, 2013. 24(6): p. 392-8.

[35] Ferla R, Claudiani P, Cotugno G, Saccone P, De Leonibus E, Auricchio A. Similar therapeutic efficacy between a single administration of gene therapy and multiple administrations of recombinant enzyme in a mouse model of lysosomal storage disease. Hum Gene Ther. 2014;25(7):609-618. doi:10.1089 / hum.2013.213

[36] Harmatz PR, Shediac R. Mucopolysaccharidosis VI: Pathophysiology, diagnosis and treatment. Front Biosci - Landmark. 2017;22(3):385-406. doi:10.2741 / 4490

[37] Paddison PJ, Caudy AA, Bernstein E, Hannon GJ, Conklin DS. Short hairpin RNAs (shRNAs) induce sequence-specific silencing in mammalian cells. Genes Dev. 2002 Apr 15;16(8):948-58. doi: 10.1101 / gad.981002. PMID: 11959843; PMCID: PMC152352.

[38] Paul CP, Good PD, Winer I, Engelke DR. Effective expression of small interfering RNA in human cells. Nat Biotechnol. 2002 May;20(5):505-8. doi: 10.1038 / nbt0502-505. PMID: 11981566.

[39] Drittanti, L., et al., High throughput production, screening and analysis of adeno-associated viral vectors. Gene Ther, 2000. 7(11): p. 924-9.< / crispr> < / talen>

Claims

1. A method for incorporating foreign DNA sequences into the genome of a cell, A method comprising bringing the aforementioned cells into contact with the following: a) Donor nucleic acids including the following: - The aforementioned foreign DNA sequence; - Any one or more albumin exons; (Here, the donor nucleic acid is sandwiched between the inverted target sequences at 5' and 3'.) b) Complementary oligonucleotides homologous to the target sequence; and c) Nuclease that recognizes the target sequence (Here, the target sequence is located at the 3' end of the albumin gene in a region selected from introns 9, 11, 12, 13, and 14 of the albumin gene.)

2. The method according to claim 1, wherein the donor nucleic acid comprises one or more albumin exons, and the exons are exon 13 and / or exon 14 or fragments thereof.

3. The method according to claim 1, wherein the donor nucleic acid comprises one or more albumin exons, the exons being exon 10 and / or exon 11 and / or exon 12 and / or exon 13 and / or exon 14 or fragments thereof.

4. The method according to any one of claims 1 to 3, wherein the complementary oligonucleotide homologous to the target sequence is a guide RNA that hybridizes to a target sequence located in intron 9, intron 11, intron 12, intron 13, or intron 14 of the albumin gene, or to its complementary strand, and preferably the guide RNA is adjacent to a protospacer fringe motif (PAM) sequence.

5. The method according to any one of claims 1 to 3, wherein the albumin gene is a human or mouse gene.

6. The method according to any one of claims 1 to 3, wherein a complementary oligonucleotide chain homologous to the target sequence is under the control of a promoter, preferably under the control of a U6 promoter.

7. The method according to any one of claims 1 to 3, wherein the foreign DNA sequence is a coding sequence for the allyl sulfatase B (ARSB) gene, and preferably the ARSB coding sequence includes or essentially includes a sequence that is at least 95% identical to sequence number 33.

8. The method according to any one of claims 1 to 3, wherein the foreign DNA sequence is the coding sequence of the factor VIII (F8) gene, and preferably the F8 coding sequence includes or essentially includes a sequence that is at least 95% identical to sequence number 36 or 55.

9. The method according to any one of claims 1 to 3, wherein the inverted target sequence is an inverted sequence with respect to a target sequence located at the 3' end of the albumin gene in a region selected from intron 9, intron 11, intron 12, intron 13, and intron 14, and preferably the inverted target sequence is ligated to a protospacer adjacent motif (PAM) sequence at its 3' end.

10. The method according to any one of claims 1 to 3, wherein the donor nucleic acid further comprises one or more of the following: - Post-transcriptional regulatory elements (preferably localized to the 3' end of the exogenous DNA sequence); - Transcription termination sequence (preferably localized to the 3' end of a post-transcriptional regulatory element or the 3' end of an exogenous DNA sequence); - Splice acceptor sequence (preferably localized at the 3' end of the donor nucleic acid, and, for example, ligated to an albumin exon if present); - Ribosome skipping sequence (preferably localized between the foreign DNA sequence and the albumin exon).

11. The method according to claim 10, wherein the ribosome skipping sequence is T2A, P2A, E2A, or F2A, preferably a T2A sequence, and / or the post-transcriptional regulatory element is a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), and / or the transcription termination sequence is a polyadenylation signal sequence, preferably bovine growth hormone polyA (BGH polyA).

12. The method according to any one of claims 1 to 3, wherein the target sequence includes or essentially includes a sequence having at least 95% identity with any one of SEQ ID NOs: 1 to 2, SEQ ID NOs: 9 to 18 or a functional fragment thereof, and / or a complementary oligonucleotide chain homologous to the target sequence includes or essentially includes a sequence having at least 95% identity with any one of SEQ ID NOs: 1 to 2, SEQ ID NOs: 9 to 18 or a functional fragment thereof.

13. The method according to any one of claims 1 to 3, wherein the donor nucleic acid comprises the following: - The first inverted target sequence having the protospacer adjacent motif (PAM) sequence; - Splice acceptor array; - Preferably one or more albumin exons selected from exons 13 and 14; - Preferably a ribosome skipping sequence which is T2A; - Preferably an exogenous DNA sequence which is the coding sequence of the human ARSB gene; - Transcription termination sequences; and - A second inverted target sequence having the protospacer adjacent motif (PAM) sequence.

14. The method according to any one of claims 1 to 3, wherein the nuclease is preferably selected from the group consisting of CRISPR nuclease, TALEN, DNA-inducing nuclease, meganuclease, and zinc finger nuclease, and preferably the nuclease is a CRISPR nuclease selected from the group consisting of Cas9, Cpf1, Cas12b (C2c1), Cas13a (C2c2), Cas3, Csf1, Cas13b (C2c6), and C2c3, or SaCas9 or variants thereof such as VQR-Cas9-HF1.

15. The method according to any one of claims 1 to 3, wherein a cell is brought into contact with a nucleic acid encoding the nuclease, preferably the nucleic acid encoding the nuclease is under the control of a tissue-specific promoter, such as a liver-specific hybrid liver promoter (HLP).

16. The method according to any one of claims 1 to 3, wherein the complementary oligonucleotide, the donor nucleic acid, and the nucleic acid encoding the nuclease are contained in a viral vector or a non-viral vector, and preferably the viral vector is selected from adeno-associated viruses, lentiviruses, retroviruses, and adenoviruses.

17. The method according to any one of claims 1 to 3, wherein the cells are selected from the group consisting of hepatocytes, one or more lymphocytes, monocytes, neutrophils, eosinophils, basophils, endothelial cells, epithelial cells, hepatocytes, osteocytes, platelets, adipocytes, cardiomyocytes, nerve cells, retinal cells, smooth muscle cells, skeletal muscle cells, spermatocytes, oocytes, and pancreatic cells, induced pluripotent stem cells (iPS cells), stem cells, hematopoietic stem cells, and hematopoietic progenitor cells, and preferably the cells are hepatocytes of the subject.

18. Cells that can be obtained by the method of any one of claims 1 to 3, Cells for use in the treatment of diseases selected from diseases in which both the mutant allele and the wild-type allele are replaced with the correct gene copy provided by donor DNA, or recessive genetic diseases and common diseases due to loss of function, preferably hemophilia, diabetes, lysosomal storage disorders such as MPSI, MPSII, MPSIIIIA, MPSIIIIB, MPSIIIIC, MPSIVA, MPSIVB, MPSVI and MPSVII, sphingolipidosis such as Fabry disease, Gaucher disease, Niemann-Pick disease and GM1 gangliosidosis, lipofuscinosis such as Batten disease and mucolipidosis; gyroretinal atrophy, adenylosuccinate deficiency, hemophilia A and B, ALA dehydrogenase deficiency, and adrenoleukodystrophy.

19. Systems including the following: a) Donor nucleic acids including the following: - Foreign DNA sequence; - Optionally, one or more albumin exons; (Here, the donor nucleic acid is sandwiched between the inverted target sequences at 5' and 3'); b) Complementary oligonucleotides homologous to the target sequence; and c) A nuclease that recognizes the target sequence; (Here, the target sequence is located at the 3' end of the albumin gene in a region selected from introns 9, 11, 12, 13, and 14.)

20. The system according to claim 19, wherein the nuclease is encoded by a nucleic acid, the donor nucleic acid, the complementary oligonucleotide, and the nucleic acid encoding the nuclease are located on a DNA construct, preferably the donor nucleic acid and the complementary oligonucleotide are located on the same DNA construct, while the nucleic acid encoding the nuclease is located on a separate DNA construct.

21. The system according to any one of claims 19 to 20, wherein the donor nucleic acid and / or foreign DNA sequence and / or target sequence and / or complementary oligonucleotide and / or nuclease are as defined in any one of claims 1 to 3.

22. The system according to any one of claims 19 to 20, wherein a nucleic acid encoding a complementary oligonucleotide and / or donor nucleic acid and / or nuclease is contained in one or more viral vectors or nonviral vectors, preferably the viral vector being selected from adeno-associated viruses, retroviruses, adenoviruses and lentiviruses.

23. The system according to claim 22, comprising a first vector containing a nucleic acid expressing a nuclease, and a second vector containing a complementary oligonucleotide chain homologous to a donor nucleic acid and a target sequence, wherein these elements are as defined in any one of claims 1 to 3.

24. The system according to claim 22, comprising a first vector containing a donor nucleic acid and a second vector containing a nucleic acid encoding a complementary oligonucleotide and a nuclease homologous to a target sequence defined in any one of claims 1 to 3, wherein these elements are as defined in any one of claims 1 to 3.

25. A system according to any one of claims 19 to 20 for use as a pharmaceutical.

26. A vector comprising a donor nucleic acid and / or a complementary oligonucleotide homologous to a target sequence as defined in any one of claims 1 to 3, and / or a nucleic acid encoding a nuclease that recognizes the target sequence.

27. The vector according to claim 26, wherein the vector is a viral vector, preferably a lentiviral vector or an adeno-associated virus vector, or a nonviral vector, preferably selected from polymer-based, particle-based, lipid-based, peptide-based delivery vehicles or combinations thereof, such as cationic polymers, micelles, liposomes, exosomes, microparticles and lipid nanoparticles (LNPs).

28. The vector according to claim 26, further comprising a 5'-terminal repeat (5'-TR) nucleotide sequence and a 3'-terminal repeat (3'-TR) nucleotide sequence, preferably where the 5'-TR is a 5'-inverted terminal repeat (5'-ITR) nucleotide sequence and the 3'-TR is a 3'-inverted terminal repeat (3'-ITR) nucleotide sequence, preferably where the ITRs are derived from the same or different viral serotypes, preferably where the virus is AAV, and preferably where serotype 2.

29. A host cell comprising the system according to any one of claims 19 to 20.

30. A virus particle comprising the system described in any one of claims 19 to 20.

31. The virus particle according to claim 30, wherein the virus particle contains the capsid protein of AAV.

32. The virus particle according to claim 31, wherein the virus particle contains a capsid protein of an AAV serotype selected from one or more of the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, and AAV10, and preferably contains a capsid protein of an AAV of serotype AAV2 or AAV8.

33. A pharmaceutical composition comprising one of the following: The system and pharmaceutically acceptable carrier according to any one of claims 19 to 20.

34. A kit comprising one or more containers containing the system according to any one of claims 19 to 20, and optionally further comprising instructions or packaging materials describing a method for administering a nucleic acid construct to a patient.

35. The vector according to claim 26 for use as a pharmaceutical.

36. The system according to any one of claims 19 to 20 for use in the treatment of liver diseases, mucopolysaccharidosis, lysosomal storage disorders such as MPSI, MPSII, MPSIIIIA, MPSIIIIB, MPSIIIIC, MPSIVA, MPSIVB, MPSVI and MPSVII, sphingolipidosis, such as Fabry disease, Gaucher disease, Niemann-Pick disease and GM1 gangliosidosis, lipofuscinosis, such as Batten disease, and mucolipidosis; and other diseases in which the liver can be used as a factory for the production and / or secretion of therapeutic proteins, such as diabetes mellitus, gyroretinal atrophy, adenylosuccinate deficiency, hemophilia A and B, ALA dehydrogenase deficiency, and adrenoleukodystrophy.

37. A system for producing viral particles according to any one of claims 19 to 20.

38. A DNA construct comprising a donor nucleic acid and / or a complementary oligonucleotide homologous to a target sequence and / or a nucleic acid encoding a nuclease that recognizes the target sequence, as defined in any one of claims 1 to 3.